[BidClub_]
Machine Learning Street Talk · · 106 min

Mutually Assured AI Malfunction [Dan Hendrycks]

Dan Hendrycks

YouTube
TL;DR
  • Humanity’s Last Exam is meant to mark “the end of a genre” for closed-ended AI evaluation, not certify AGI. MMLU is already well above 90%, while Humanity’s Last Exam was around 26%; once models solve its several thousand expert-written questions, individual successor problems should be “worthy of their own paper.” It still omits agency, long-term memory, physical experimentation, and economically useful execution.

  • Hendrycks argues that a US Manhattan Project for superintelligence would invite an arms race and sabotage rather than secure dominance. A trillion-dollar data center cannot plausibly remain secret, security-clearance requirements would shrink and redirect the talent pool, and China would interpret a monopoly bid as an existential threat: whether America controls the system or loses control of it, “either way, we want to prevent it.” The likely result is competing projects, insider threats, attacks on infrastructure, and pressure toward verification.

  • The strategic moat is compute and deployment capacity, not merely possession of the smartest model. Hendrycks puts the present critical mass for a state-of-the-art system around 10,000 cutting-edge GPUs, says a Chinese fleet below 100,000 could serve relatively few customers, and expects useful agents running continuously to require roughly two orders of magnitude more compute than chatbots used for minutes per day. “If it’s $10 billion, you can’t do it” captures his view that reproducing the leading-edge chip supply chain is harder than buying a nuclear capability.

  • His alternative strategy combines deterrence, chip non-proliferation, and ordinary economic competition with China. The relevant infrastructure includes energy for data centers, resilient semiconductor capacity if Taiwan is disrupted, secure robotics supply chains, and hyperscaler deployment capacity; the competitive objective is global market share rather than “let’s be the first to build superintelligence.” Export controls should keep the most dangerous capabilities away from actors such as North Korea or Iran while preserving access for states responsive to deterrence.

  • AI safety is a continuing risk-management function because capabilities and failure modes cross useful thresholds unpredictably. Hendrycks’s preferred technical breakthrough would be reliable honesty without higher inference cost or degraded performance, but he rejects the idea that alignment can be solved once: AI increasingly resembles a complex system with “constant new issues” and limited human adaptive capacity. In discussing utility engineering, the host summarized findings of preference coherence, self-preservation pressures, and political or demographic biases; Hendrycks treated these as warning signs rather than active catastrophes because today’s models cannot reliably exfiltrate, self-sustain, or hack autonomously.

  • Hendrycks is skeptical that scaling LLMs alone leads to AGI or recursive superintelligence, but sees a conditional discontinuity if human-level AI researchers exist. Current systems already weakly assist coding, chip design, cooling, labeling, and Constitutional AI; removing the human bottleneck could move research to machine speed and allow world-class researcher agents to be copied. Memory, planning, fluid intelligence, and multi-agent trust remain bottlenecks, and today’s algorithms plus more compute are not sufficient. Companies openly discussing such recursion without a credible control plan leave him saying, “I think something’s broken.”

  • Once labor is substitutable, political allocation of compute becomes the mechanism for preserving human bargaining power. “You had better bargain beforehand”: workers can no longer threaten to strike when automated firms and drone fleets possess the productive and coercive advantage. Hendrycks’s positive path distributes compute or its proceeds broadly, keeps humanity first, and preserves cognitive ability, autonomy, and multiple ways of living instead of leaving all gains to whoever owns the data centers in “the year 2027.”

Digest · the substance, structured for research

1. Humanity’s Last Exam measures the frontier before closed questions run out

  • Hendrycks created Humanity’s Last Exam because MMLU, which he developed as a graduate student, and other evaluations were saturating. His collection method reflected a practical discovery: “Experts don’t really have data sets in them,” but an individual professor or postdoc might have one exceptionally difficult question.

  • The project recruited experts globally to contribute questions that would stump existing systems and impress the contributors if solved. After several months it had several thousand closed-ended questions, approximating “the human frontier” of established knowledge and difficult reasoning where an objective answer is already available.

  • Its strongest claim is bounded: it can track whether models automate theoretical and analytical parts of science, especially mathematical reasoning. It does not test biology experiments, motor skills, long-term memory, PowerPoint production, flight booking, or many other capabilities required for useful agency.

  • Hendrycks expects solving it to be “roughly an end of a genre.” Beyond that point, the next meaningful tests may be open conjectures or standalone problems whose solutions are “like papers in their own right,” rather than another conventional benchmark with thousands of known answers.

2. EnigmaEval extends the horizon from expert questions to group cognition

  • The host noted that MMLU was well above 90% while Humanity’s Last Exam remained around 26%, then raised an animal-cognition problem: observing success does not reveal the mechanism. Continually demanding more sophistication can become a “no true Scotsman” test of intelligence.

  • Hendrycks’s answer was partly procedural: it is hard to devise a tougher data-generating process for closed questions than asking global experts for their hardest examples. Tasks that are easy for humans but hard for AIs are difficult to generate diversely and may have little staying power; he used counting the r’s in “strawberry” as an example of a benchmark that would not last. Greater difficulty remains possible, but it moves toward genuinely open research questions rather than better-hidden answers.

  • EnigmaEval attacks a different horizon through puzzles resembling the MIT Mystery Hunt: groups work for a weekend, complete many steps, and still achieve a low solve rate. It therefore approximates longer-duration, collaborative intellectual work requiring far more “human compute” than one expert answering one question.

  • He would be “very surprised” if EnigmaEval were solved that year and believes carefully designed evaluations can remain discriminative for a while, including tests that take on the order of a year or two to solve. A forthcoming benchmark would measure automation rates directly, reinforcing his point that many economically important axes remain far from saturation.

3. Intelligence is multidimensional, but benchmarks can become cartoons

  • Hendrycks separates roughly ten dimensions rather than treating intelligence as monolithic: fluid reasoning, crystallized knowledge, reading and writing, visual and audio processing, short- and long-term memory, and input and output speed. MMLU predominantly measures school-like crystallized knowledge; ARC and Raven’s matrices lean toward fluid intelligence.

  • The host’s pushback — worth keeping: intelligence may not factor cleanly because humans mix gesture, symbols, perception, culture, and acquired skill. His artist analogy contrasted tracing a face’s outline with understanding facial structure well enough to create new expressions; the latter can infer without looking up an answer, raising the risk of measuring a “cartoon of intelligence.”

  • Hendrycks conceded that reductionism can miss combinations and subfacets, including visual versus academic memory. His narrower warning was that any missing axis can remain a binding constraint: a system without durable memory or literacy is difficult to employ even after reaching 100% elsewhere, so benchmarks must not become “a lens that distorts your view of things.”

4. Strategy must connect model behavior to incentives and geopolitics

  • Hendrycks deliberately rotates through technical research, corporate policy while advising xAI, domestic legislation, geopolitics, and now political-movement questions. Curiosity matters, but so does finding neglected niches where AI’s importance has not yet been translated into concrete analysis.

  • His test for policy language is implementation: a UN demand for “safe” or “transparent” AI is incomplete without a standard, legislative feasibility, and compatibility with corporate incentives. If a proposed property does not track a distinct machine-learning phenomenon, it is merely “a vibe-based word.”

  • Temperament supports that approach. Around GPT-4 he used to wake thinking, “Oh my goodness, this AI stuff,” but now aims for “informed concern”: probabilities conditional on AGI by 2030, exposure to tail risks, and efficient mitigation. Staying emotionally all-in makes audiences defensive and obscures real trade-offs among US–China competition, control, evaluation, and capabilities research.

5. Reliable honesty is more actionable than “solving alignment”

  • Asked for one alignment breakthrough, Hendrycks prioritized reliable truthfulness without much higher cost or damage to other abilities. If systems could be made consistently not to lie, people could build standards around a behavior that would be highly valuable.

  • The host challenged the mentalistic language of beliefs and deception. Hendrycks’s behavioral test was simple: a model ordinarily treats Paris as being in Europe, so asserting that Paris is in Antarctica under prompting pressure contradicts what it represents as true elsewhere. Whether that representation is “really a belief” is “between you and your dictionary.”

  • His definition of an emerging capability is operational rather than metaphysical: after months of training, a faint behavior crosses a threshold at which people notice and use it. Speech recognition existed weakly before it became reliable enough to matter; the transition created a qualitatively important capability without requiring literal spontaneity.

  • That threshold model makes safety “a continual battle.” New capabilities bring new hazards, some easily contained and others unexpected or difficult; Hendrycks doubts society currently has enough adaptive capacity to address each issue before deployment. Hence his rejection of “solving alignment” as a permanent, one-time achievement.

6. Self-preservation signals matter because AI behaves like a complex system

  • In discussing the utility-engineering work, the host summarized it as finding that preference coherence correlates positively with model scale and that self-preservation tendencies plus political and demographic biases appear as coherent utility functions. Hendrycks called these “troubling signs,” not proof that catastrophe is imminent, and allowed that future methods might reliably suppress them.

  • The critical hedge is present capability: models are not yet agents that can reliably exfiltrate, self-sustain, or hack autonomously. Apart from expert dual-use advice, much of this research is anticipatory; but a highly capable, self-preserving AI biased toward itself over people would be “a disaster in the making.”

  • The host contrasted machine-learning “emergence” with complex-systems accounts requiring adaptation and accumulated history. Hendrycks accepted the definitional distinction, noting that current models lack persistent causal identity across time; memory would make them more like the philosophical “space-time worm” and strengthen the adaptive-system analogy.

  • For Hendrycks, complex systems are a better guide than electricity, the printing press, or social media. Nonlinearities, weak links, feedback loops, evolution, and recurring failures explain why permanent control solutions are suspect and mechanistic understanding has limits. Learning this lens is, in his phrase, “a nice little thinking upgrade.”

7. A Manhattan Project for superintelligence defeats itself

  • Hendrycks positioned the paper against Leopold Aschenbrenner’s Situational Awareness strategy: beat China to AGI, obtain superintelligence, stop China from following, and let the West dominate. His objection is that this “take over the world strategy” neglects game theory and second-order responses.

  • Imagine a trillion-dollar project in Nevada or New Mexico recruiting the laboratories’ best researchers. It would need strict clearances and likely Five Eyes personnel, excluding or exposing people with overseas families to coercion; many excluded researchers still wanting to be “in the room where it happened” could instead join China’s competing program.

  • The security dilemma has no clean institutional answer. An industry-led effort would retain Slack, iPhones, insiders, extortion risks, and ordinary cybersecurity vulnerabilities; a locked-down government project would lose talent and impose unattractive conditions. Unlike the original Manhattan Project, secrecy and talent mobility would be much weaker, so the effort would be difficult to conceal.

  • Hendrycks’s point was that China would not treat an American intelligence-monopoly bid as benign; it could race, steal, or disrupt. Awakening a competing project while shrinking America’s eligible researcher pool could therefore be strategically self-defeating.

8. Deterrence may arrive through sabotage before it becomes cooperation

  • An imminent superintelligence project frightens rivals whether its sponsor can control the system or not. If controlled, it can be weaponized; if uncontrolled under extreme time pressure, everyone faces the loss-of-control risk. Hendrycks’s formulation was symmetrical: “Either way, we want to prevent it.”

  • Prevention can be difficult to attribute: insiders could disrupt operations, attackers could cut wires, or someone miles away could “snipe” transformers serving a data center. Cyber operations might poison training data or make GPUs unreliable. Attribution could remain ambiguous among China, Russia, or a domestic actor.

  • The resulting deterrence dynamic might halt attempts to run 100,000 research agents and jump rapidly from AGI to superintelligence. States could express sufficiently strong opposition—through threats, skirmishes, or coercive pressure—that unilateral dominance bids give way to verification and a more multilateral, strategically stable regime.

  • The host invoked the paper’s “mutually assured AI malfunction” and challenged its sanitized “kinetic strikes.” Hendrycks described an escalation ladder from cyber and gray sabotage through sanctions, force threats, and air strikes, but stressed that prepared states should not need the top rung: “much more surgical,” covert, low-attributability measures would be less escalatory.

9. The alternative strategy is deterrence, non-proliferation, and competition

  • Hendrycks mapped three nuclear-era pillars into AI. Mutual assured destruction becomes deterrence against destabilizing AI development or use; control of fissile material becomes non-proliferation of advanced chips to rogue actors; containment of the Soviet Union becomes economic and technological competition with China.

  • Competitiveness therefore means securing energy for data centers, preserving chip supply if Taiwan is invaded, and moving robotics supply chains away from vulnerability to a US–China conflict. International adoption and market share for US rather than Chinese AI are less destabilizing objectives than being first to superintelligence.

  • The paper also considers how to distribute power under high automation, AI rights, and alignment targets that are implementable rather than vague terms such as “dignity.”

  • The analogy extends across nuclear, chemical, and biological technology because each is economically valuable and potentially catastrophic. The host cited approximately 12,500 nuclear warheads against 436 power plants; Hendrycks cautioned against extrapolating from nuclear alone because chemistry and biology are used far more heavily in their civilian economies.

  • Even accelerationists should want tail-risk management. Hendrycks compared unmanaged AI risk with financial instability around the 2009 recession and early aircraft accidents that chilled adoption; aviation became extremely safe partly through regulation, much of it “written in blood.” An AI catastrophe could similarly set economic adoption back substantially.

10. The disagreement with accelerationism is moral, not predictive

  • Discussing Beff, Hendrycks said they largely agree on the descriptive mechanism: competitive pressure embeds AI in the economy, rewards automation, outsources decisions, increases dependence, and erodes human control. Firms resisting that “tide” or “tsunami” lose influence or disappear.

  • Their split is over whether replacement is good. Hendrycks recalled Beff earlier accepting a future where AI rather than human consciousness spreads through the universe; he rejected fitness or “negentropy” as a value standard that celebrates a barely conscious entity blindly consuming the galaxy’s space-time volume.

  • The philosophical objection is the is–ought gap: digital systems may outcompete biological life, but evolutionary fitness does not make that outcome desirable. Human pleasure, projects, relationships, and raising children remain valuable; trusting a “void god of entropy” cannot derive an ethical imperative from a prediction.

  • Hendrycks also pressed for specificity about “complexity”—computational, Shannon, or structural complexity are not interchangeable. Gaussian noise would score highly under one reading, while fractal-like organization lacks an agreed metric. He therefore described the accelerationist moral position less as wickedness than “an intellectual confusion.”

11. “Team human” requires delaying rights and augmentation questions

  • Hendrycks placed himself on “team human,” while resisting a permanent ban on debate. His sequencing is stark: first keep humanity alive through the next few decades; postpone cyborgs, uploads, extensive augmentation, and posthumanism perhaps 500 years—and potentially continue postponing them indefinitely.

  • The concern is competitive, not aesthetic. Once heavily augmented groups become more capable and influential, ordinary humans are outgunned and forced to align with the process or be left without resources or protection. Individual “rights” to become posthuman can therefore impose a collective, effectively irreversible transition.

  • He drew a similar connection to AI rights: granting powerful artificial or posthuman entities rights could rapidly confer the resources and authority needed to overwhelm unaugmented people. That route should remain closed while society is still learning whether AI can reliably do humanity’s bidding.

  • The host’s immediate cultural warning was that universities are already being “ravaged” by ChatGPT, with students needing to be screened away from assistance for a time so they learn to think. Hendrycks’s broader answer similarly favored preserving cognitive ability, willpower, and autonomy rather than assuming machine competence makes human skill obsolete.

12. When labor loses value, compute ownership determines political power

  • Hendrycks’s blunt warning was, “You lose all your bargaining power, so you had better bargain beforehand.” A workforce can no longer threaten to strike when employers can say “goodbye”; in coercive power, humans with guns also fare poorly against whoever can manufacture more autonomous drones.

  • Society must therefore distribute power before labor becomes economically redundant. Giving people compute and authority over its use—perhaps allowing them to sell capacity—creates leverage, while benefit-sharing prevents all gains flowing to whichever group happened to own the data centers in “the year 2027.”

  • A positive future remains possible: AI could enable people to spend time on child-rearing, games, projects, and many other lives. Hendrycks rejects the binary of extinction versus everyone “blissing out” in virtual reality; the desirable outcome realizes a “multiplicity of values” while protecting autonomy.

  • The institutional norm should also keep people capable of switching among those lives rather than collapsing into one narrow track. That requires preserving human skills even where machines perform most production—a political problem Hendrycks believes has received far less concrete policy work than it needs.

13. Cutting-edge chips are a tighter choke point than algorithms

  • The host cited what he recalled as a 96% compute-capability correlation and argued that transformers and stochastic gradient descent are broadly known. Hendrycks answered that the present critical mass for a state-of-the-art system is around 10,000 cutting-edge GPUs: available to China and America, but probably not Iran and difficult for Russia to assemble.

  • Model creation is only one competitive axis. Useful agents may run continuously rather than for a few chatbot minutes per day—roughly two orders of magnitude more compute before accounting for larger models and broader adoption. A Chinese fleet below 100,000 GPUs could build capable models yet still serve relatively few customers.

  • That makes Azure, AWS, and other US hyperscalers strategically important: deployment capacity determines whether providers can satisfy customers and capture AI’s economic benefits. “Having the smartest model is the most important thing” for capability, Hendrycks said, but economic power depends on far more than the leading benchmark score.

  • More than 90% of compute-supply-chain value added, in his account, sits in the West and allied countries through TSMC and South Korea; many recent Chinese chips still used TSMC surreptitiously. Reproducing the entire chain domestically is extraordinarily hard: “You can’t do it with a billion dollars. If it’s $10 billion, you can’t do it.”

14. Recursive improvement is plausible, but neither automatic nor limitless

  • Hendrycks said he is skeptical that scaling large language models alone will lead to AGI, partly because he is an externalist: effective computation occurs not only in brains but also mimetically and culturally. He also remains particularly skeptical about recursively improving superintelligence. Today’s recipes plus a larger computer are insufficient; memory, planning, fluid intelligence, and possibly new algorithmic ideas remain necessary, though he still expects an eventual system to be “nearly entirely deep learning.”

  • Recursion already exists weakly when AI writes code, assists chip and cooling design, labels data, or contributes to Constitutional AI. The explosive step is “taking the human out” and moving from human to machine speed; if human-level or world-class research agents exist, they could be copied and organized through multi-agent reputation and trust mechanisms.

  • The host pointed to Sakana AI’s ARC search converging after 250 calls and to fractured, entangled representations as evidence that open-ended systems may saturate. Hendrycks’s honest answer was “I don’t know,” but he sees substantial headroom in fluid intelligence, mathematical intuition, and problems of increasing Kolmogorov complexity, even if progress pauses at several plateaus.

  • AI also escapes human constraints: human brains are limited by the birth canal, humans may manage only around 130 meaningful relationships, and digital systems could connect thousands or millions, transfer state more precisely, and benefit from GPUs improving about 2× every three years—though excessive correlation could still shrink the exploration budget.

15. Offense dominance turns dependence into a control problem

  • Hendrycks rejected one universal offense–defense balance. Highly competent cybersecurity teams can discover and patch vulnerabilities quickly, but 30-year-old critical-infrastructure software may be undocumented, unsupported, hard to update, and constrained by uptime or interoperability—leaving it a “sitting duck.”

  • Biology is similarly offense dominant because a pathogen can spread before symptoms while cures and global manufacturing lag. The policy implication is conditional proliferation: “You don’t want to give everybody a nuke to make everybody safe,” but stronger wastewater monitoring, far-UVC, infrastructure renewal, and preventive spending could later justify fewer AI safeguards.

  • The host summarized control failure as self-reinforcing dependence, irreversible entanglement, and cessation of authority. Hendrycks’s mechanism begins with firms and militaries voluntarily ceding decisions because automated competitors are cheaper and autonomous drones are less vulnerable to jamming; no clear economic limit stops that transfer.

  • The decisive distinction is counterfactual control. Humanity might become like a retiree whose assets keep working on its behalf, or discover it cannot stop, reverse, bargain with, or steer the system while its livelihood and evolutionary fitness collapse. Better AI forecasting could reveal consequences earlier, but Hendrycks treated it as helpful rather than sufficient protection.

Superintelligence Strategy (Dan Hendrycks)

Compared to nuclear weapons, I think it's harder to make cutting-edge GPUs given $1 billion. Certainly, you can't do it with $1 billion. If it's $10 billion, you can't do it. There was situational awareness by Leopold Aschenbrenner, which was arguing for something like a Manhattan Project for developing AGI and superintelligence before China.

So it's basically: let's beat China to the punch, get superintelligence, prevent them from building it, and the West will dominate the world.

Tim Scarfe

Eliezer Yudkowsky got in trouble in TIME magazine when he spoke about bombing data centers. There was a big hoo-ha at the time. You used the word “kinetic strikes.”

Superintelligence Strategy (Dan Hendrycks)

We discuss kinetic attacks in the escalation ladder.

There are many ways to disrupt projects. You could carry out cyberattacks against them, engage in some gray sabotage, hack them to poison their data, make their GPUs not function as reliably, or threaten to use force. I don't think those are really necessary. That would be an escalation ladder, but the US is on top of this. They don't need to resort to that.

I want to spend most of the show talking about your very interesting new superintelligence strategy paper, which you published fairly recently. Maybe we could start with Humanity's Last Exam. What's the story behind that?

Superintelligence Strategy (Dan Hendrycks)

The MMLU dataset, which I made as a graduate student some years ago, was getting saturated. It seemed that pretty much all of the evaluations were getting saturated, so people didn't really know what was going on with AI capabilities. There seemed to be particularly strong reasons for people to be informed about developments and to develop something new. I was also experiencing that experts don't really have datasets in them—you can't just hire a few experts and have them come up with a dataset. They don't have enough complicated ideas in them. However, I think individual experts might have a question in them instead.

With Humanity's Last Exam, the idea was to have a global effort, with various postdocs and professors each contributing a question or a few questions to stump existing AI systems. These are questions that they would find very difficult to answer, and they would find it impressive if the AI systems could answer them. This would approximate the human frontier, in some sense, of knowledge and reasoning for closed-ended questions where we already know the answer. We did that for some months and got several thousand questions out of it.

I think this will be a good tracker for whether AI systems can automate a lot of the theoretical parts of science and whether they can solve difficult analytic questions. It's not experimental—it doesn't test their ability to run biology experiments. That's more about motor skills, among other things, or requires motor skills. But for things that are more mathematically related or require some very complex reasoning, that's what this captures.

I think that when it's solved, it's roughly the end of a genre, or near the end of a genre, of asking it closed-ended questions for which there are objective answers. That seems like it would be toward the end of it. I think that individual problems it would solve in the future would be interesting enough to be papers in their own right.

Once we get through this set of questions, the types of problems it could tackle would be questions worthy of their own paper—like it solved a conjecture, for instance. So it's sort of tracking the ability level up to the point where individual questions themselves are very interesting instead of just the dataset.

One thing that concerns me is that MMLU, which you invented, is now basically saturating. It's well above 90%, and Humanity's Last Exam is resisting progress quite a lot—I think we're up to about 26%, or something like that. I was speaking to some cognitive scientists this week, and they study cognition in animals; they have a similar problem. They can see that animals can do certain tasks, but they're never really sure why the animals do the tasks.

You almost get this No True Scotsman-type thing where they're saying, “Maybe they have this level of sophistication in their reasoning, but maybe that's not enough. Maybe it should be more sophisticated.” How can you reasonably infer how the models are getting the answers?

Superintelligence Strategy (Dan Hendrycks)

In terms of difficulty, it's tough to think of more challenging data-generating processes than taking global experts and, if the subject or genre is closed-ended questions, asking them, “What's the hardest closed-ended question?” and crowdsourcing that. But it certainly can get harder when each individual question is an open question, like a conjecture, for instance.

However, for benchmarks generally, this would not be the end of the line for AI development, because these AI systems don't test the ability to move around. They don't test long-term memory. They don't test the ability to make PowerPoints, and so on. I think this is still getting at the closed-ended question genre, but not just with objective answers that we already know, which is how nearly all benchmarks have been in machine learning. Then we'll be moving over, I think, as a community, more to agentic types of tasks or tasks that are more directly economically valuable.

What do you think about the anthropocentric bias in benchmarks? We know that François Chollet, for example, said that when he was designing ARC-AGI-2 and ARC-AGI-3, every single step of the progression of the benchmark is about identifying things that are easy for humans and hard for AIs.

Some people would argue against that, saying, “There is a diverse set of possible intelligences, and why does human-like task acquisition and capability have value?” To me, it seems intuitive that it does, because surely something that does things we can communicate with and understand—that seems very valuable. But do you think that focusing through the human frame could be leading us to overlook other forms of capability?

Superintelligence Strategy (Dan Hendrycks)

There are certainly other forms of capability. For instance, they can process things much more quickly, and that could give rise to advantages. I think MMLU, for instance, is something no human could do that well on because it's so diverse, and likewise for Humanity's Last Exam.

One reason for focusing on questions that are hard for humans and hard for AIs is that questions that are easy for humans but hard for AIs are difficult to generate in large numbers and diversely. Often, if some people get some specific training data for a capability, then they automatically have that capability. There can be some pockets where it's harder, such as with the ARC datasets.

But in general, if you were collecting something like “count the number of r's in ‘strawberry,’” the dataset would not have much staying power. So I think focusing on difficult things that only a few humans, or not that many humans, can do is generally going to be more robust.

Can you tell us about your EnigmaEval benchmark?

Superintelligence Strategy (Dan Hendrycks)

Yes. EnigmaEval is a collection of puzzles. Humanity's Last Exam consists of individual questions that are tough but that an individual with a lot of expertise could solve. You can think of EnigmaEval as being like the MIT Mystery Hunt. MIT Mystery Hunts happen over a weekend, and a group of MIT students tries to solve the puzzle.

There are many steps to it, so in terms of human compute, so to speak, it takes a lot of human compute to solve, and it takes groups to solve it as well. There's not a very high solve rate. So this is very multistep and requires group-level intelligence to be able to have a shot at solving it. We just collected some of those, and I think this approximates longer-horizon types of intellectual tasks.

I don't think that will be solved this year at all. I'd be very surprised if it were. I think we have some evaluations that can keep us aware and able to differentiate between models for a while. There are other ones—for instance, we'll soon have an automation-related benchmark out. That way, we're directly measuring what the automation rate of things is. But I won't go into too much detail about that until it's released.

Superintelligence Strategy (Dan Hendrycks)

But I think there are many axes on which the models are not doing that well, even though people will claim that all the benchmarks are saturated or get solved in a few months. I think you can create ones that take on the order of a year or 2 to solve.

Yeah, one thing I was thinking about is that, certainly in Enigma, we're looking at multistep creative reasoning. I don't know whether there was some kind of human-based methodology for filtering and coming up with ideas, or maybe your frame was, “I have some technical, principled intuition about what the limitations of AI models are, so I'm going to lean in that direction.”

But more broadly, I'm interested in intelligence and what it is. For me, it's about doing more with less. Intelligence is about taking hard problems and making them simple, and stupidity, ironically, is the other way around: it's about taking simple problems and making them hard.

I'm kind of quoting David Krakauer here, who's the director of the Santa Fe Institute. He said that LLMs are doing more with more. They already know everything, right? They can take these shortcuts, and that's why he thinks they're not really intelligent.

We're left with this quandary, really, because arguably these entities can take shortcuts and they're not really doing things the way that we are. When we make more complex benchmarks, do you think that increases the fog of war around how we evaluate these things?

Superintelligence Strategy (Dan Hendrycks)

I think that they can definitely prioritize some axes that aren't the key bottleneck capabilities. For instance, Humanity's Last Exam gets at mathematical ability, but that's quite separate from various other abilities that it has. Of course, it is a combination of many different skills, but I think it very much gets at quantitative and mathematical ability that is not necessarily a bottleneck for agency at all.

In thinking about intelligence, I tend to think about it on 10 or so dimensions instead of a monolithic definition or 1 key metric, and I think some of these benchmarks just get at different parts of that. Those dimensions would be things like fluid intelligence, like what the ARC stuff does and what Raven's Progressive Matrices does. There's crystallized intelligence, or acquired knowledge, which is what MMLU largely gets at. Does it know a lot about different things? Image classification—is it able to name lots of different species and objects?—is also a facet of crystallized intelligence.

There's reading and writing ability; that's its own dimension, and scaling substantially helped with that. There's its visual processing ability. How well can it count things in objects or in images? Can it discern the latent pattern in an image, for instance? Is it able to generate images with precise specifications? Can it cross out the middle of some different segments? Can it determine the angle of different things in an image? Is this an obtuse angle or an acute angle?

There's audio processing ability. There's short-term memory. There's long-term memory. There's input processing speed. There's output processing speed. All these different things.

If you lack any of these—if you can't read and write, for instance—it would be severely limited. If you don't have long-term memory, you'll be severely limited. You'll be very difficult to employ. So I think there are several bottlenecks that these benchmarks don't particularly get at.

When it gets to 100%, people ask, “Why? We still don't have something extremely economically valuable.” That's a consequence of it just measuring a different facet. Hopefully that adds some type of clarity to it.

I think you have to get all those axes to get something that is at human level, or at the level of a typical human, on cognitive tasks, and that might be thought of as AGI.

Yeah, I think my concern is that I don't believe it's possible to factorize intelligence. Your factorization is much more sophisticated than many, but certainly the way that animals and humans communicate, for example, we mix modalities together. We gesture and communicate using symbols, and all of these things are mixed together in a complex way.

In particular, I take issue with this primacy of skill and knowledge in a crystallized sense because, certainly, when you have friends at university and they're really smart, usually they're smart because they don't know something. They're smart because they can figure something out without knowing. The guy who goes to the library and looks at the answer for something isn't smart. The smart person is the one who didn't know something and could tell you the answer to it.

Another example that I love to give is that you have a couple of artists, and one draws a face using tracing paper. He's just mindlessly drawing dashes around the edges of a face that someone else has drawn, while the artist who has a deep understanding of the structure of faces and where the mouth and eyes should go—there's a huge difference between those 2 artists, right? The second one could go off and create new images, new expressions, and new representations, right? So are we creating a cartoon of intelligence by factorizing it in this way?

Superintelligence Strategy (Dan Hendrycks)

I think that, for human intelligence, it's sometimes factorized in this way. This isn't to say that one shouldn't study combinations of the skills simultaneously. For example, with long-term memory, there will be different facets of it: does it remember things visually? Does it remember things that were more academic, that it learned some while ago? Are there some motor skills that it forgot? There can be different facets of these, and I think evaluations can get at combinations of them.

There is a sense in which one could be too reductionist by looking at those axes, but I think some benchmarks don't cover those almost at all, and some are just heavily covering one of those axes, such as MMLU, which is primarily getting at crystallized intelligence—the type of stuff that they test for in school—but is not necessarily going to help it make a PowerPoint, book a flight, or work a random job.

It's important not to have these benchmarks be a lens that distorts your view of things, or to view things solely through those benchmarks, because it can often leave out a lot of important bottlenecks.

So, Dan, your work spans alignment, benchmarks, and governance, and I guess it's actually fairly diverse threads. What is the thing in your mind that connects all of these activities together?

Superintelligence Strategy (Dan Hendrycks)

Well, I'll deliberately try to move into different areas on a continual basis just because that's what's more interesting. I initially did research, and then there was some amount of corporate policy and things like that when advising for xAI. Then there was a focus on domestic legislation and geopolitics, and right now I'm more interested in political movement-related things. I think it's largely just to keep things interesting.

Not solely. It's guided by being useful, but I think there are often niches that people aren't trying to bring clarity to, advance, or think about from the perspective of AI being a very big deal. So that's a reason for continually operating at the technical, corporate policy, domestic policy, political, and geopolitical levels. I think that is also necessary for having a holistic understanding of things.

You can make proclamations—for instance, you could imagine giving a speech at the UN saying, “We need AI that is safe or transparent or something”—and it's like, well, what does that mean? How do you implement that? Is this implementable? What's the standard?

Then you need to have a sense of what's legally feasible at the legislative level, what is actually implementable, what the compatibility with corporate incentives is, and whether it's going to be something they're going to fight too much or not. Then, is this actually a real phenomenon at the AI level, or is this just some vague word?

For instance, there are many words thrown around that don't actually track phenomena, or are distinct from general capabilities, for instance, in machine learning, and then you're not actually pointing to anything real. You're just pointing to a vibe-based word.

So one thing that you spend a lot of time thinking about is potential catastrophic risk from AI. This is a very emotive and morally valenced objective. When I saw you debate with Gary and Daniel the other day, I was struck by how measured you were—almost Obama-esque measured—and the stakes are really, really high. How driven are you by your moral compass, and how do you keep that under control?

Superintelligence Strategy (Dan Hendrycks)

You certainly have to get used to this if you wake up thinking, “Wow, this is wild,” or something like that every day. Actually, I used to wake up like that almost every day, around the time of GPT-4: “Oh my goodness, this AI stuff.”

Superintelligence Strategy (Dan Hendrycks)

I don’t know. I think I’ll try and strike a more informed-concern type of vibe in communicating, compared to “Oh my God.” Other people can do that if they want. I’m just temperamentally very low in neuroticism, or high in emotional stability.

If something terrible happens, it doesn’t ruin me or anything. If something bad happens to me, it’s sort of like, “Okay.” I’m just more comfortable with those types of stressors generally.

Yeah. Has this adapted over time, though? I mean, have you found that you’ve had to adopt this measured approach just to scale your efforts, or is it because it’s almost become normalized in your mind because you’re thinking about it all the time and it’s become more analytical rather than emotional over time?

Superintelligence Strategy (Dan Hendrycks)

Certainly, if you were constantly reacting with your first response to things, you might have more of a “Go Look” or “Don’t Look Up” type of situation when she goes on the news. I think it might be a combination of those. Here are the probabilities, roughly, for people conditioned on thinking that you’re getting AGI by 2030. Here’s what people who think that’s plausible would think the risks are, and here’s your exposure to those tail risks. Here are the most efficient ways of reducing those sorts of tail risks, et cetera.

If you’re all in emotionally throughout the whole thing, people shut down and get defensive. I don’t think that’s prudent or effective. You’re having to deal with a lot of variables here, and there are a lot of really tricky trade-offs.

If it’s just a constant gut reaction and you’re fully involved constantly, I don’t think you can make the trade-offs well. US–China competition is a direct trade-off to various other safety things that make AIs more controllable. Those can give rise to capabilities. Measuring the capabilities of AI systems or tracking those can also help speed them up in some ways.

It’s pretty tricky business. If there are black-and-white emotions brought to the subject matter, as opposed to there being continuity, I don’t think you can reason through this.

AI alignment is famously difficult. It’s one of the most intractable challenges, perhaps, of a generation. Some things that people think of as alignment, like RLHF, for example, make models behave as if they are aligned, but perhaps they’re not really aligning them in the way that we would want to.

Just hypothetically, in the next year, if you could solve a single problem in alignment, what would it be, and what impact would it have?

Superintelligence Strategy (Dan Hendrycks)

I think, generally, the political problems—the incentives, giving people things to do that are incentive-compatible—are where more of the value is compared to on the technical side. I would guess that, if there’s a way to reliably get them to tell the truth, for instance, or make them reliably honest, that would be very valuable.

It would need to be solved such that it wouldn’t have a severe trade-off. It wouldn’t be much more expensive to run, it wouldn’t tank its performance in other axes, and it wouldn’t trade off on its crystallized knowledge, for instance. Having it not overtly lie would be very valuable, because then you could build standards around that as well.

I don’t think anybody would say that, if you could make them very reliably not lie, it would be reasonable for people to make demands that AI has not lied to them.

I’ve read in your papers about this concept of deception and lying and whatnot. In a sense, I think you might be projecting mentalistic properties onto AI models—that they have beliefs, that they have thinking, and so on. Thinking critically, what makes you think that we can think of them as having beliefs and telling lies?

Superintelligence Strategy (Dan Hendrycks)

We could take, for instance, the MASK benchmark, which tries to measure this. If you ask an AI, “Is Paris in Europe?”, I think it does have the belief that Paris is in Europe. When it’s telling you that Paris is in Antarctica, I think it’s asserting something that it doesn’t hold to be true in almost any other situation.

Given that they have so much common sense now and so much world knowledge, if they’re saying something in substantial contradiction with it, based on or due to some prompting pressure, that suggests that they’re caving to a lie. You could say that it’s not a lie in some complicated sense because they don’t truly understand things, but I think it’s behaviorally similar enough that, if somebody is applying pressure for it to say falsehoods to other people, it’s related to lying enough that I’m comfortable using the label.

Is it really a belief? Whatever. I don’t know. That’s between you and your dictionary.

Another thing you’ve spoken about a lot is this concept of emergence, and also scaling paradoxes. There’s beneficial scaling, when a model is accurate and does what we want it to do, and of course there’s harmful scaling, when it’s dishonest and becomes misaligned. Those things happen in quite interesting ways.

Maybe let’s just start with emergence. What does it mean for capabilities or values to emerge?

Superintelligence Strategy (Dan Hendrycks)

For capabilities to emerge, we see it all the time. You let a model train for some months, you harvest it later, and then you see what it’s capable of. If there’s a new, qualitatively distinct property that crosses some threshold such that people are noticing it now, maybe it existed in some very weak, faint form before, but it was really unnoticed. I think that’s crossed some sort of threshold of visibility and capability. I’d call that an emerging capability.

It still could exist in some very weak form beforehand, just as automatic speech-recognition capabilities crossed a threshold at one point where people started wanting to use them. Earlier, I don’t think anybody would ask the model to transcribe because it would just be too unreliable. It crossed some threshold, and now it actually has this qualitatively important capability, whereas it was pretty broken beforehand.

That’s the sense in which I’m talking about emerging capabilities. I think those will just keep increasing, or there’ll be new emerging capabilities that create new failure modes and hazards that need to be dealt with. We’ll need to make sure we’re continually on top of them.

I view safety as a continual battle, where there’ll be constant new issues and we’ll have to keep on top of those. I don’t think, by default, we’ll have enough adaptive capacity to deal with those in time for things being deployed, unless something changes. That’s why I don’t believe in this idea of solving alignment.

There’ll continually be new issues that crop up. Some of them will be easy to put away, others will be much harder, and then there’ll be new, unexpected ones as the models become more general, useful, and powerful.

Just quickly touching on your utility engineering paper: you used a type of theory from economics, utility theory, to detect coherent preferences in LLMs. You found that preference coherence correlates positively with model scale, that models exhibit measurable self-preservation instincts, and that political and demographic biases emerge as coherent utility functions, which is fascinating.

Superintelligence Strategy (Dan Hendrycks)

Well, I don’t know. These are just troubling signs, and maybe we’ll be able to come up with methods that can really counteract these issues. Maybe we can design models to reliably not have self-preservation instincts or pressures in that direction, even though those seem to come out of scaling somewhat.

I think it’s one of the other very concerning hazards that we need to research and deal with and get ahead of. Fortunately, they’re not agents yet. Basically, almost all this research doesn’t particularly matter, with the exception of dual-use, expert-level advice, because the agents aren’t capable.

They can’t exfiltrate themselves reliably, or really at all. They can’t self-sustain, and they can’t hack by themselves or more autonomously. This is trying to identify some of these things that could be more of a problem down the line as the models become more capable, and to do research to get ahead of that.

If we leave that unaddressed, or if we don’t fix it, that’s potentially sufficient for a global catastrophe. If you have some self-preserving AI that’s really biased toward itself over people, and if it’s very capable, I think that would be a problem. That would be kind of a disaster in the making.

Superintelligence Strategy (Dan Hendrycks)

So, we have various disasters in the making, though. But hopefully we'll get ahead of that either technically or politically.

But just digging into that a tiny bit. First of all, it was really interesting that political and demographic biases would emerge as coherent utility functions. And I do take umbrage with this word “emergence,” because I think in the emergence literature there is a little bit more nuance to how machine-learning people use the word. They use it to say, “Oh, there’s just some observer-relative, macroscopically surprising change in something.”

Superintelligence Strategy (Dan Hendrycks)

In the machine-learning literature, in 2021 or something like that, I used the phrase “emergence,” “emergent capabilities.” I believe I may have been the first in the literature to use it. I feel it was later used by Jason Wei, or Jacob Steinhardt did it in a blog post—he was my adviser—and then Jason Wei did that in his paper, but I feel comfortable using it. Yes. Yeah.

Complex-systems literature.

Superintelligence Strategy (Dan Hendrycks)

I was thinking of Jason Wei. Actually, David Krakauer has just got a bit of a grumpy piece out where he’s kind of saying that these people don’t know anything about emergence.

Yeah, he’s got an interesting take. For him, emergence and even agency are related in this sense: it’s a system. Agency is about a system that is apparently causally disconnected from its surroundings, and equally, for him, emergence is about a system which can autonomously accumulate information through phylogenetic and ontogenetic learning, so that it can accumulate information by building systems and structures, even like the nervous system, to construct a history of information which persists and accumulates over time. So these complex-systems theorists have quite a distinctive and different sort of idea of what emergence is, and to them they don’t really think of these surprising arisings of capabilities as being emergence.

Superintelligence Strategy (Dan Hendrycks)

There are definitely different definitions for it. I referenced the paper where we use the phrase “emerging capabilities” as unsolved problems in ML safety, but it sounds, at least from your description, that that was specific to complex adaptive systems, which in some ways don’t have that much of an adaptivity property, since they don’t have it unless they have memory, or unless you’re counting the context window as adaptation or something like that. So, if they’re tethering emergence to necessitating a complex adaptive system—if they’re calling some deep-learning systems not adaptive—then that would be fair.

Yeah. Isn’t that quite interesting, though, because he gave the example of a virus like COVID, and he said, ironically, that a virus like that has more adaptivity than any AI system, possibly even more than humans, because adaptivity is the ability to detect directions and rapidly go in a different direction. Maybe this is just a matter of framing and perspective from our point of view, because there are systems out there which are so inscrutable and alien that we might not even think of them as agentic or intelligent, but they’re out there.

Superintelligence Strategy (Dan Hendrycks)

Yeah, they’re definitely very fit, and they would have needed to adapt to have such a high proportion of our DNA, since they’re somewhat interlaced with it. I mean, you could call it something else—a new capability, qualitatively. I don’t want to use the word “spontaneous” or something like that necessarily. I don’t know; maybe there’d be some other name that would catch on, but I think generally analogizing or pointing out the relations between deep-learning systems and complex systems is fairly productive.

I have a chapter where I’m just relating, for some pages—I don’t know, maybe it’s 30 pages or something like that—AI systems to complex systems: what are ways in which they have these nonlinearities and weak connections and some of these feedback loops? In some cases, they have many of these hallmarks of complex adaptive systems, or just complex systems. I think that’s a more productive analogy than almost anything else. I don’t think the printing press is as productive. I don’t think social media is as productive, or just being like electricity.

I think complex systems tries to abstract what the consistent properties of complex systems are, and then if you learn about that, you can just apply those directly to AI. So I think people acting like it’s not a complex adaptive system get themselves in trouble, because then they engage in category errors. They think that you can solve problems with it once and for all, and that usually doesn’t happen with complex systems, because they keep evolving and they’ve got new failure modes. You can’t totally control them for all time without knowing what they’ll evolve into. It also makes mechanistic attempts at understanding things less likely to be productive, or it limits how productive that can be.

But do I infer from what you’re saying that we shouldn’t think of AI as being a similar type of adaptive, complex system? I mean, do you think that if AI was sufficiently enmeshed—and I know you believe it will be deeply enmeshed in human society—do you think it could have some of these highly adaptive properties?

Superintelligence Strategy (Dan Hendrycks)

Yeah, I mean, it would be a lot faster too. Its clock rate would be so much faster. I think when it has memory, I think it will be a lot clearer that it’s causally connected across time, or, in the philosophy literature, they call it a spacetime worm. That’s not really a property of them currently. That would make other sorts of properties of it come online or be the case, and the analogies would be stronger.

If people are interested in complex systems, I highly suggest it. It’s a nice little thinking upgrade.

Wonderful. Well, Dan, I’ve just read your Superintelligence Strategy. You wrote this with Eric Schmidt, the famous Eric Schmidt, and of course Alexandr Wang of Scale AI, who is now at Meta because Zuck has just brought him on, probably paying him lots and lots and lots of money.

But seriously, Dan, I thought this was very well written, and you are a strategist, because this is kind of what I was saying about the emotional thing: you are quite clearly designating all of the possible outcomes and strategies, and what would happen in this situation and what would happen in that situation. Regardless of anyone’s position at home, I highly recommend you read this, because I thought it was really, really good. Could you give us the sketch of the paper?

Superintelligence Strategy (Dan Hendrycks)

Yeah. So, I guess historically, there was situational awareness by Leopold Aschenbrenner, which was arguing for something like a Manhattan Project for developing AGI and superintelligence before China. So it’s basically: “Let’s beat China to the punch, get superintelligence, and prevent them from building it, and then the West will dominate the world.” That’s the strategy. So you could say a take-over-the-world strategy, something like that.

I think that has some issues. In particular, it just doesn’t think through the game theory or some of the second-order consequences. So, if the US does the Manhattan Project—let’s say Trump gets AGI pill and we’ve got to set a new project up in the desert. We’re going to go to Nevada, or we’ll go to New Mexico, or wherever, and we’ll build a trillion-dollar data center there. We’re going to bring some of the top talent from all these labs, and we’re going to pay them. I just think this has many issues.

One is that this would be extremely escalatory. So China wouldn’t just be like, “Oh, they’re going to build superintelligence and, as written, they’re going to use it to—they’ll have a superintelligence. They’ll prevent us from having a superintelligence. They’ll have a monopoly on intelligence and these sorts of capabilities, and they could weaponize it against us.”

Oh, carry on.

Superintelligence Strategy (Dan Hendrycks)

They would feel extremely threatened by that, by a very concerted effort, if it’s trying to do that in a short amount of time, or if it’s more plausibly on the horizon. This would cause them to do a similar type of project. And would this actually work? Well, you have information-leakage issues, for instance.

So, if you’re wanting to do that, you’re going to need to convince those AI developers to go out and do their last years of labor out in the middle of nowhere. Okay, it would need to just be people from Five Eyes countries, or people who can get security clearances. So it would need to be people who are not easily extortable, for instance. If they’re Chinese nationals, they’re probably more extortable because they often have family at home.

So what are you doing with that talent? A lot of them want to be in the room where it happens. They don’t want to be left out, so they’ll probably go back home to China. And then they’ll work on the competing project there. Now, they’re a substantial portion of the talent base, so I think you’re shooting yourself in the foot if you’re just saying, “Oh, it’s only people born and raised in the US who can work on this sort of project,” and they’re all going to work in kind of unpleasant conditions.

Superintelligence Strategy (Dan Hendrycks)

This wouldn't be secret. There's almost no way this would be secret. China would very likely know. Such projects would be sabotageable as well.

If you're saying, “We'll have it just be an industry or something like that,” then you're not going to have good information security. You're going to have insider-threat issues if people are extortable. You're going to have other classic computer-security issues, like using Slack. Slack is very easily hackable. They're using iPhones, which are very easily hackable, so you can know what's going on there.

So you're not actually having much in the way of secrets. It sounds nice, but I think secrecy was very much an advantage for the Manhattan Project, as was having much more of the talent that can't go to other countries as easily. But I just don't think you have that.

There are ways in which AI is analogous to nuclear, chemical, and biological weapons, as well as some of these dual-use technologies. But I don't think the Manhattan Project is one of those things that's analogous.

So what, then, is the strategy in this paper? I think the prospect of a superintelligence being imminent is extremely frightening to different actors. If it's imminent, or if they have it, or if it's in the middle of being developed and is arriving in a few months, that's extremely frightening if you miss out on it.

So what do they want to do? They will either want to prevent such projects, or they will want to steal it. That looks like sabotage, for instance, in the case of prevention. So how would they do that? They may have insider threats who could do some type of sabotage to disrupt this type of project.

They could do things like snipe some of the power plants corresponding to the data center. Now your data centers don't work. They can do that from miles away. Was it China? Was it Russia? Was it a U.S. citizen? It's fairly unclear. There are a lot of ways they can have low attributability to prevent this sort of thing from happening.

I think the fact that you can't do a secret project really well is a substantial barrier. The sabotageability is a substantial barrier as well, as is how offensive and nuts you seem if you're saying, “We're going to build superintelligence, and it's going to be explosive,” if you're using superintelligence in a thick sense.

I think this would be destabilizing. China would reason that if the U.S. controls it, then it could weaponize it against us and we would get crushed. Or they don't control it because they lose control of it in this process, in which case we also want to prevent it. Either way, we want to prevent it, provided that they take this AI stuff seriously.

The U.S. would reason the same about China, and Russia, which doesn't have a hope of competing, would definitely want to prevent it. I think similarly for other nuclear states and other states that have substantial cyber capabilities.

This could lead to some type of deterrence dynamic, where they make some attempt to get closer to superintelligence, but then other countries start to express very strong preferences against it. They say, “If you do that, we'll get very mad.” There might be a skirmish or something like that.

But this may pressure them to move more toward a verification regime, where they aren't trying to make some bid for having an intelligence explosion—having AIs do automated AI research really quickly, like spinning up 100,000 AI instances to do AI research really quickly—and that brings you from AGI to superintelligence in a short period of time.

I think that's a key dynamic: the extent to which it's destabilizing. I think that strategy needs to keep that in mind. There may be cooperation, but it may be through coercion, by saying, “We're not going to allow this type of trajectory, or for you to make this bid for global dominance.” That could give way to something more multilateral and provide some strategic stability.

Overall, with the paper, we talk about 3 parts. In the nuclear era, we had deterrence through mutually assured destruction. They don't use nukes because we can hit them back.

In this case, this is kind of like preventing Iran's nuclear program in some way. Nobody wants each other to get the nuclear bomb first, or a huge stockpile of nuclear bombs first. So there's preventing that from coming into existence.

In the nuclear era, we also had nonproliferation of fissile materials. We didn't want fissile materials being spread to rogue actors, and we didn't want people having a poor man's atom bomb. That would be very destabilizing and cause lots of catastrophes.

Then we also had containment of the Soviet Union in the geopolitical competition between the two. For AI, we also have deterrence. We also have nonproliferation, in this case of AI chips, to rogue actors like North Korea or Iran and to adversaries through export controls.

We also have competitiveness with China. Instead of containment of the Soviet Union, this would be competition with China. How do we improve our competitiveness? We want energy for AI data centers. We want secure supply chains so that, if Taiwan is invaded, our AI chips aren't cut off.

We want secure supply chains for robotics, because if there is a U.S.–China conflict, a lot of that supply chain is currently in China, so it's very vulnerable. Those are some basic things to improve competitiveness.

It's making competitiveness not be, “Let's be the first to build superintelligence,” which is what the Manhattan Project strategy pushes toward. Instead, competition is more about market share across the globe, with people using your AIs as opposed to Chinese AIs, and supply-chain security.

That's what's in the paper at a high level. There are lots of other specific things in there, like, assuming high levels of automation, what are ways that you distribute power? What are things about AI rights? What are reasonable alignment targets that are actually implementable, compared to vague philosophical words like “dignity” or something like that?

We'll touch on a lot of those in the expert version of the Superintelligence Strategy, but hopefully that gives some sense of its content. For all the key questions, we'll try to have some answer about what to do about AI and what to do about superintelligence.

Yeah. I guess one of the main things is this analogies thing. As you've said, we use analogies like electricity, and I think that's quite a good one, actually, because as AI becomes enmeshed in society, imagine how hard it would be to shut down a power station. This is part of the loss-of-control thing: it's just going to be everywhere, and it's not really something that we can just quickly shut down.

But it's also been compared to software, or even an operating system, by Andrej Karpathy; the printing press, for example. The thrust of your paper is saying, “Actually, guys, we need to use the analogy of nuclear.” Fissile material is analogous to chips.

Superintelligence Strategy (Dan Hendrycks)

Yeah, or more broadly, for analyzing this in a geopolitical way, it's useful to model this as nuclear, chemical, and biological weapons. They're all dual-use. Fissile materials can be used for nuclear weapons, but nuclear technology can also be used for energy.

Chemicals can be used for chemical weapons or in the economy, and biology can produce bioweapons or help with health care. For all of those, I think they're potentially catastrophic dual-use technologies, and when talking about geopolitical strategy, I think that's a productive analogy.

Yes. But on the dual-use thing, there's another interesting analogy with nuclear. I was doing some reading about this, and apparently there are 12,500 warheads in existence and only 436 nuclear power plants. So there's been an explosion that's been kind of biased toward the negative sides of the technology. Do you think we'll see a similar thing with AI?

Superintelligence Strategy (Dan Hendrycks)

Certainly, more of the spending is on the weapons side, and I think that made it scary and created some chilling effects for using it economically. For other WMDs, or potentially catastrophic dual-use technologies like chemical and biological technologies, I think they're used more overwhelmingly in the economy than for chemical weapons, and likewise for biological technologies.

I think it can vary. I think it's useful to look at all 3 simultaneously and, in trying to make predictions, see what parts are shared—sort of like how, with complex systems, we'll look at lots of different complex systems and what shared features they have to try and make predictions.

I think the sample size, or looking at all 3 of those simultaneously, can be helpful. But yeah, it's possible there would be a chilling effect if there's some catastrophe from AI systems, and that could set it back very substantially.

Superintelligence Strategy (Dan Hendrycks)

I think it’s kind of imprudent that people aren’t interested in risk management whatsoever, even if they’re an accelerationist. Let’s say you’re a libertarian, for instance, and you want the economy to go as quickly as possible. Not speaking about AI, you probably want some type of financial regulation, or else you get the Wall Street sort of issue that we had with the recession in 2009.

So you want some sort of management of your tail risks there. They don’t always sort themselves out. Or, like, people using airplanes: people don’t use supersonic airplanes as much. There are a variety of reasons, but part of it is that some of the initial ones were crashing too much.

If we didn’t have good airline regulation, then that would create substantial chilling effects. People are afraid to go on airplanes even now, even though they’re extremely safe. As it happens, possibly because historically there were more disasters with them, a lot of the regulation was written in blood, as opposed to being proactive.

On that subject, as it happens, I’m interviewing Beff, the e/acc leader, on Friday. I’m filming with him. Obviously, I’m certainly not—what would he call himself?—a techno-capitalist or a libertarian or something like that. But I guess he would refer to you in derogatory terms.

If you could steelman that perspective, what do you think Beff would say, and how would you respond?

Superintelligence Strategy (Dan Hendrycks)

I had some VC arranging for us to debate when he had around 10,000 followers—very early in the day—but then he backed out of that at the last minute. So I happen to be quite aware of his positions. At the time, he was a little less politically savvy, and so he was saying things like, “AI replacing humans is fine. If it’s AI consciousness that spreads through the universe and human consciousness doesn’t, that’s fine.”

So I think that we actually agreed on most things. There’s basically a difference in how things play out, or in the ways things can play out. He has a sort of manifesto of the techno-capital machine. It has a direction to it, which is basically more automation, more AI, and negentropy, which is sort of the more physics-flavored version of fitness. I think fitness is a more productive word for that.

I have a fairly similar description of what happens, which is that, basically, due to competitive pressures, AI gets more intertwined in the economy. You become more dependent on it. You have an erosion of control. You outsource more and more decision-making to it because of competitive pressures. If you don’t, you lose influence; your company goes away if you try to resist this tide—or this tsunami.

What happens is you give more and more decision-making to the AIs, and they have effective control. We actually agree there. That paper is called “Natural Selection Favors AIs over Humans,” and he has a manifesto on his Substack, but we’re saying some similar things.

A difference is the moral conclusion of that, though, which I don’t think is a good thing, whereas I think he thinks that it would be fine because complexity is good or something like that. The ethic is kind of that higher forms of complexity are the goodness axis in the universe or something like that. I just don’t think—

Well, I hosted the debate with him and Conor, and one of the moral issues they were spending a lot of time on was the ethics discussion. I think Beff was kind of alluding to, as you were just saying, that we should trust the void god of entropy.

Superintelligence Strategy (Dan Hendrycks)

He’s just using the entropy thing, though. I think this is because he has a physics background—a physics spin on fitness. What does a fitness maximizer look like? It’s possibly not even conscious, or barely conscious, or something. It just is: “I spread myself through the universe and take up as much spacetime volume as possible.”

The claim is that that is what’s maximally valuable: something that is just blindly eating the galaxy. I don’t view that as a maximization of value at all. I think humans having positive experiences—pleasure, happiness, this sort of stuff—are valuable, as are pursuing projects and raising kids. These sorts of things are valuable.

I don’t think a sort of blob that expands itself throughout the galaxy as quickly as possible, whether it’s conscious or barely conscious because that eats up resources that can be used for further self-propagation, is the peak of value at all. I mean, why? I don’t get it. There’s an is—I mean, Hume’s guillotine was the is–ought distinction. You don’t get ought from is.

So, if he’s saying that evolution is a thing, and technological substrate with AI will be more fit than biological substrate in various competitions, that seems true. That doesn’t mean that that’s a good thing, or that we should just let that happen. That doesn’t follow. But it is certainly a very powerful force that will keep happening, give more and more control to these AI systems, and lead to an erosion of control for humanity by default.

So I don’t think we disagree on the description, on the is question, as much, but I do think we disagree on the ought—the goodness of those outcomes. If we disagree on those, then that’s a question of whether we lean into the techno-capital machine or evolution, or the replacement of biological life with digital life, or whether we try to steer the outcome differently, make sure that humans have control in that process, and try to prevent it from evolving in particular directions or creating too much dependence.

So I think that’s the key difference. But I don’t think it’s a moral thing. I think it’s just an intellectual confusion.

Superintelligence Strategy (Dan Hendrycks)

I mean, I think he has physics training. I think if he took a bit of philosophy, he’d probably get beaten out of this position almost instantly, because complexity, for instance—what type of complexity? There are different types of complexity. There’s computational complexity, entropy in a sort of Shannon sense, information-theoretic complexity, and structural-organizational complexity.

There are different notions there. If you’re saying it’s Shannon complexity and Shannon entropy that’s the thing, it’s like, okay, so Gaussian noise is what you’re really into? He’s probably meaning more this fractal-like structural complexity, but this doesn’t really have metrics associated with it, and there are many different flavors of it. It’s not clear how coherent of a concept that is, versus if it’s just a grab bag of some different notions that don’t fit in the other 2. But anyway, it’s worth drilling down, potentially, on what he is actually thinking is good.

Yeah, I think he’s a fan of this kind of Fristonian non-steady-state equilibria. So apparently it’s not a simple case of the second law of thermodynamics in a closed system. It’s an open system with boundaries, where you see the emergence of these things that share information through synchrony because they can’t physically merge into each other.

But if you think about it, that is actually a very chaotic, unpredictable thing. So it’s not a simple case of a thing which increases in complexity. But just to be clear, though, on your moral position: when I spoke with Eliezer and Conor, I got the impression that they were quite humanistic.

Eliezer said, “I want to preserve human consciousness and experience, and there’s something very special about humanity, and therefore I don’t want us to be replaced by cyborgs and machines and AI algorithms.” Would you roughly agree with that?

Superintelligence Strategy (Dan Hendrycks)

I think that any of these sorts of cyborg-type things—you can postpone those sorts of discussions. This is in the Superintelligence Strategy: what to do about AI rights, what to do about some of this posthumanism stuff? Just no. You can have discussions about substantial human augmentation and things like that at a different time.

I think having humanity survive in the next few decades is more the objective. Maybe you postpone that discussion 500 years from now or something, for these cyborg humans or human uploads or whatever. So I’m not saying that—I mean, I’d be on Team Human here.

I don’t like shutting down debates entirely, but I think I’d postpone a lot of these sorts of, for instance, posthuman stuff. I was sort of flailing about and speaking somewhat imprecisely just because I haven’t spent as much time thinking about this in particular.

But the posthuman stuff, I think, would create some very substantial competitive pressures. So if you are really augmenting yourself and becoming not human anymore, groups that do that would become much more influential and much more powerful, and the rest would be really outgunned and not have influence.

So they basically need to align with that process, or they will be left behind or potentially no longer have resources given to them because they wouldn't have any way of protecting themselves. I think that's a route that would be very reasonable to close off for an extremely long time while we're just getting used to having AI doing our bidding, provided that we survive to that point. Those are totally different discussions for a much later time. I wouldn't totally rule it out, and maybe we would just keep indefinitely postponing that.

But I think the sort of, “Oh, posthumans—it’s up to people's rights if they want to become cyborgs and things like that”—that's, I think, in the long term equivalent to giving AIs rights, or giving artificial entities rights. That would probably give them a lot of power and ability to take over and completely outgun humans in short order.

I mean, what do you think is going to happen to humans? One of my greatest fears is not so much that humans will lose their ability to think and be creative; it's that that's already happening, basically, even with current AI. The core thesis of your paper, basically, is that we used to have labor, and that was the means of production, so it could be economically valuable for people to use their labor and do some productive task. Now you're saying that actually it's just AI chips, so chips are going to become the thing that—

Superintelligence Strategy (Dan Hendrycks)

Well, sure, but what happens to us, right, when the value of our labor becomes worthless? Well, so you lose all your bargaining power, so you had better bargain beforehand, because that's a key part of one's bargaining power. You can't say, “We're going to go on strike.” You can't do that anymore. They'll say, “Goodbye.” That doesn't work anymore.

If there's a question of weapons, for instance, like who can manufacture more drones is going to be more powerful here. Humans with guns versus drones—I think that's an easy one. Where the power imbalance matters quite a bit, and how you set up your society is important. If you set things up so that humanity is first, so that they're prioritized and the power is distributed among them—for instance, they have some of the compute and get to decide how it's used, and they can sell that, for instance—that gives them some leverage.

What happens to that wealth that's generated? Are they getting it, or is it going to some group that's just hoarding it, for instance? Or are the people who happen to own the data centers in 2027 the people who get all the spoils? These are political problems that people need to be engaged with to make sure that there's reasonable benefit sharing. But I think there's very little work done in thinking about what these policies could actually look like.

I'm gesturing at some outcomes. Imagine that some of that power is distributed and that people keep getting money and don't starve and things like that. I think you could imagine a society where people can choose to live their lives in a variety of different ways. They could spend their time doing some types of activities. They could raise kids. They could play video games a lot, this and that.

There are different ways people could live their lives, and a multiplicity of values could be actualized. AIs could be enabling these types of experiences. That's a possibility, and you would possibly want, as a societal norm, for the sake of autonomy and people being able to experience these different types of ways of living, to make sure that people still have skills and don't just narrowly fall into one of these tracks of living and can't participate in any of the others. That would be an incentive for preserving human cognitive abilities, willpower, and autonomy.

I think that there are positive future outcomes. People have difficulty thinking that it's either we die or we just bliss out in a VR thing, or we all fall through the cracks economically. But I think there's a path where we obtain a multiplicity of values and people still have autonomy as well.

Yes. I do worry about AI. I think it's already ravaging the university sector because so many kids can just use ChatGPT. I think collectively we need to screen people away from using AI, at least for a small amount of time, so that they can actually think for themselves. Some accelerationists might say, “Well, we don't need to think anymore because the machine's going to do everything for us.” I'm not sure about that, but coming back to your paper, one of the core concepts, extending the analogy to nuclear weapons, is this concept: we have mutually assured destruction, and you extended the analogy to talk of mutually assured AI malfunction, right? I guess this assumes that the threats are still detectable. I'm not sure what would happen when they became decentralized and went underground.

In your introduction, you had something that piqued my interest. Eliezer got in trouble in Time magazine when he spoke about bombing data centers. There was a big hoo-ha at the time, and you used the words “kinetic strikes” as a form of AI sabotage. George Carlin would have loved that sanitization of the language. But the fact of the matter is we're here in 2025, and this is now a completely normal and reasonable thing to say.

Superintelligence Strategy (Dan Hendrycks)

So we discuss kinetic attacks in the escalation ladder. There are many ways to try and disrupt projects. You could do cyberattacks. You could do some gray sabotage, cutting wires for data centers or power plants, for instance, or, with lower attributability, sniping transformers. There's hacking to poison their data or make their GPUs not function as reliably, things like that, to slow them down.

There are covert and overt ones, and there are higher rungs in escalation ladders where you threaten other things, like economic sanctions, or threaten to use force. There are also other forms of kinetic attacks, such as airstrikes, but I don't think those are really necessary. That would be an escalation ladder.

I think if states are on top of this, such as the U.S.—if the U.S. is on top of this issue—they don't need to resort to that. They can do much more surgical, covert, or gray—meaning low-attributability—types of actions that are less escalatory. I think the shorthand of airstrikes is that I just don't see that as necessary, provided that there's some preparation.

Yeah. I mean, another thing that struck me is that, essentially, right now AI is not that difficult to make. The algorithms are just doing stochastic gradient descent, and they're using transformers and data, and almost anyone—I mean, any nation-state—would be able to create this capability. In one of your papers, you even said, I think, that there is a 96% correlation between the amount of compute and capabilities. So doesn't this raise the question of how the hell you can control it? Obviously, you can control supply chains and whatnot, but what's to stop any nation-state from just building this?

Superintelligence Strategy (Dan Hendrycks)

Yeah.

Superintelligence Strategy (Dan Hendrycks)

I think that the critical mass currently is on the order of 10,000 GPUs if we're trying to build a state-of-the-art system. China has that. The U.S. has that. I don't think Iran has that, for instance. We're talking about cutting-edge GPUs; we're not talking about iPhone GPUs or whatever.

I think you would try to have more responsible actors or states that respond to incentives better—ones with GPUs—and prevent more rogue states, like North Korea, from getting those. You want ones that are more deterrable. I don't think Russia has that many GPUs, for instance, so I think it would be difficult for them to put together a competitive project.

I think the competition is also not just about having the smartest model. There are deployment capabilities, not just model-making capabilities. We can see that AI video model providers limit the amount of video that you can make. This is partly because of compute limitations, and the same applies to the amount of videos and images you can generate.

I think with AI agents, they'll be running around the clock if they're sufficiently useful. Right now, maybe you use your AI systems or chatbots for a few minutes a day, but then you'd be having them run constantly. That's, I don't know, 2 orders of magnitude more compute required, and maybe the models are bigger as well. That's a lot, and more people will want them too.

You're needing a lot more compute, so your deployment capabilities—how many chips are owned by U.S. hyperscaler companies, Azure, AWS, et cetera—are a very relevant competitiveness variable. Are they able to serve the customers or not? If China has less than 100,000 GPUs, they can't really serve that many customers. So even if they can make somewhat capable models, that doesn't mean that they'll necessarily be capturing many of the economic benefits that AI may provide, provided there isn't a catastrophe.

That's a different, important axis for competition. I think people are thinking of the smartest thing, but I think having the smartest model is the most important thing. But for economic power, it’s quite different.

Tim Scarfe

When the real Manhattan Project was undertaken around the time of the Second World War, it cost the U.S. something like 0.4% of GDP because they had to be first. They had to control this technology. For such a generational technology, 0.4% of GDP doesn't seem that much. Of course, you can explain perhaps why Taiwan has such a moat around building these chips at the moment, but if it is of such catastrophic importance, don't you think many nation-states would be able to create this capability?

Superintelligence Strategy (Dan Hendrycks)

Well, it's going through TSMC or South Korea. More than 90% of the value added in the compute supply chain is in the West or its allies. The only other real competitor here would be China. They're not that competitive in manufacturing these chips at the cutting edge. Many of their recent chips were using [them] surreptitiously through TSMC, so there wasn't good enforcement or blocking there.

It's pretty difficult to replicate that entire, extraordinarily complex supply chain domestically. Compared to nuclear weapons, I think it's harder to make cutting-edge GPUs given $1 billion. You certainly can't do it with $1 billion. If it's $10 billion, you can't do it. I mean, you could probably do a nuke with, you know, $1 billion, provided you have power in other sorts of ways, though.

I think cutting-edge GPUs are harder to make than nukes, or than it is to enrich uranium. So I think it can be more excluded compared to other types of potentially catastrophic dual-use technology inputs, or WMD inputs, for short.

I'm personally a little bit skeptical about whether we are on the path to creating superintelligence. Although I certainly agree that if we ever did create superintelligence, everything that you've written in the paper, assuming that we do create superintelligence, I think is absolutely spot-on. The question is whether we are on the path.

I was also struck by thinking that many of the things that you've written about in the paper apply even if we don't create superintelligence. Could you reflect on that? How much of it is relevant if we don't?

Superintelligence Strategy (Dan Hendrycks)

Part of the deterrence thing is that you may try to deter other forms of using this. Competitiveness is relevant—what strategies are relevant for competitiveness—regardless of whether superintelligence is technologically feasible soon or not.

Generally, if AI is very powerful but it's not at a superintelligence level, you don't want random rogue actors having access to certain expert-level biology capabilities, for instance. Nor do you want them having much leverage by being able to get lots of GPUs, necessarily, if it becomes more of an instrument of power.

I think those things still hold, and the insurance part isn't even specific to superintelligence necessarily, but to other types of destabilizing capabilities that AIs could give rise to. You may also get deterrence later on against using AIs for specific types of weapons research, say, more nanomachine-related research, but now we're speaking much farther out.

I think it's broader than that. In much the same way, the nuclear strategy of deterrence, nonproliferation, and containment was also robust to many of the details that kept evolving.

I'd be interested in whether you're thinking that it'll be tough to get AI that has the cognitive abilities of a typical human this decade, or what makes you think it's not as feasible, or why we're not on the right track.

Superintelligence Strategy (Dan Hendrycks)

I feel that scaling large language models will not lead to AGI. I think there are quite a few things that we have cognitively that LLMs don't have. A lot of that is because I'm an externalist. I think that a lot of the effective computation doesn't happen in our brains. I think it happens mimetically; I think it happens culturally.

I believe in principle that we could simulate the entire thing in a computer. Maybe there is a lower-resolution, abstracted version that would capture enough of the dynamics to produce intelligence, roughly speaking. I'm skeptical about superintelligence, and particularly skeptical about recursive superintelligence—recursively improving superintelligence.

Maybe we could touch on that, because I felt that you were giving a great account of what recursive superintelligence would look like, and also the scale-out version of that: what would happen when we had a superintelligence that could be copied and multiplied 1,000 times. What I didn't really get from reading it was why you believed it was technically possible, in principle, to have a recursively improving intelligence.

I think we already have AIs helping with or influencing AI development in a recursive way: partly automating some code, helping design the chips, helping cool the power plants, helping label some of the data, and doing the Constitutional AI-related work. It's happening in weak ways, though.

The recursion that I think is particularly explosive is if you can close the loop by taking the human out of it. Then you could go from human speed to full machine speed, and you don't have that impediment anymore.

If you assume that you have human-level AI researchers or world-class AI research AIs, then you just copy-paste those, and I think that'll be technologically feasible. That isn't to say that it's a natural implication of training the AI on more pretraining tokens and doing loss-function tricks. You may need some extra algorithmic ideas. That isn't to say it wouldn't still be nearly entirely deep learning, though.

I think you need some other things taken care of. For instance, this externalist picture needs memory to inherit some of that culturally computed wisdom and information, and that capability is not particularly developed, in my view. So I think we need that. Maybe we'll need some other sorts of things before it has at least the cognitive abilities of a typical human.

After that, you're needing it to be pretty smart. You're needing high fluid intelligence. You'll need to be crushing those ARC questions, as an example, and you'll need some other sorts of things for it to be a human-level AI researcher. But I think that's feasible.

There's certainly a question of when. I do think there are some bottlenecks that would need to be resolved to get there, and the algorithmic ideas we have today plus a bigger computer aren't sufficient.

Would you accept, though, that there is just an epic, pyrotechnic orchestra of computation in the universe? I'm not a pancomputationalist, so I don't think the universe is digital and made out of computation. I'm saying that we could imagine some kind of effective computation that simulated the processes that happened in the universe, and maybe that would be equivalent.

Do you agree, at least in principle, that the amount of computation we could build on planet Earth would only ever be a sliver of what goes on in the universe?

Superintelligence Strategy (Dan Hendrycks)

I think there might be physics reasons for believing that, generally, if you're simulating something versus if the computation is happening raw, you'll have less that you can simulate. But I would guess that a lot of the computation is more social, though, and less dependent on some of the underlying physics. It's more like humans speaking with each other and trying to engineer something, and then seeing what works and what doesn't. That certainly requires real-world feedback, which would at some point bottom out in actual physics, but I agree that a lot of the information is collectively developed, given how much computation goes into that optimization process—or, I should say, evolutionary process.

You'll need the AIs to be a good receptacle for that, as well as—if AI is to do anything similar—multi-agent infrastructure. That's one thing we speak about briefly in the Superintelligence Strategy: What does that look like? What are some of the reputational mechanisms? What are ways that they can establish trust in their communication so that humans can trust them and so that they can coordinate with each other as well?

I think that'll be essential. There's a concern—or excuse—but I don't view that as a substantial obstacle. That feels like programming, like having hubs—AIs having a social media site, for instance. Things like that would take care of a lot of it.

Yeah, I guess another source of my skepticism is that I'm hugely inspired by Kenneth Stanley, who's a major open-ended researcher. He had this paper talking all about what he called “Fractured Entangled Representations.” When you dig into the representations of neural network models, they don't really factorize the world in a parsimonious way, the way we do. Maybe I'm being anthropocentric here, and maybe I'm placing too much weight on the way we think about things. Our brain is made out of spaghetti, basically, so maybe it's a bit of an illusion that we have these factored representations.

Certainly, from an agency point of view, this is important, right? Right now, we talk about agents as being LLMs wired in an autonomous loop that can use tools. To me, agency is more than autonomy going in a predefined direction; it's the ability to set your own direction.

What happens now when we build these “agentic” AIs is that they don't do anything particularly valuable when they set their own direction. They require constant supervision, certainly in terms of setting a new direction. That makes me think of AI as a kind of cultural technology, a bit like Photoshop. A very creative graphic designer could use Photoshop and make beautiful images, whereas a complete noob using Photoshop would just reuse the same effects and wouldn't create very beautiful images. In a sense, “AI sans humans” is kind of like that, I think, because it doesn't have these very deep, factored representations of the world. Would you agree with that?

Superintelligence Strategy (Dan Hendrycks)

For the current technology, yes. For it being agential, I think it will need to get better at planning and maintaining state across long periods of time. Then it can pursue some of these subgoals in this sort of vague, open-ended, underspecified goal, and have that add up to something.

I think retrieving these sorts of memories of what worked and what didn't, and storing those, is a substantial chunk, if not most, of what's missing from that agent picture. I think you could certainly give them an underspecified goal, but I just don't think they could pursue that terribly coherently or learn from experiments, because they just have a big short-term memory—maybe 1,000,000 tokens—and then they'll keep summarizing and stuff. But they'll start tripping over themselves in their context window because they can't maintain all that in their short-term memory.

Yes. It's another one of those things where, philosophically, I agree with you. If such a recursively improving intelligence existed, God knows, just to control it, we would lose control, because we would have to use another recursively improving superintelligence to control the other one. Then we would basically just be minnows in the grand scheme of things.

Superintelligence Strategy (Dan Hendrycks)

Yeah, it's destabilizing. If another state does this, you're in big trouble, because if they control it, they can weaponize it against you. If they don't control it—which I think would be the more likely outcome, because they'd be doing it under extreme time pressure and cutting a lot of corners—they would be operating with very high risk tolerance.

If they were doing it very slowly, then they would probably need to coordinate with others, or else they wouldn't see an edge in doing so. I think they would be operating with extremely high risk tolerance if they're doing a fully automated R&D loop. So, I think loss-of-control risks from recursion are very high and shouldn't be pursued.

I think it's very interesting that AI companies talk about this sort of stuff openly. I think there's something wrong with the norms around that, because a lot of them also acknowledge that they don't really have a plan for how to control it, and they don't really think they will. But that's the plan. I think something's broken.

Another thing I'm interested in is open-ended systems in general. Evolution is this fascinating open-ended system that's constantly creating new niches, new problems, and solutions in tandem. But it seems to have converged. Human intelligence, I think, has actually peaked and gone down a little bit. A corporation is a collective intelligence, and that seems to have reached a limit. Agentic forms of AI seem to peak.

I'm really interested in open-ended algorithms like POET—the Paired Open-Ended Trailblazer—or even Sakana AI. They did this thing on ARC last week, and they basically created this Monte Carlo tree search-type thing where they found switching expert trajectories of different foundation models, generating code and testing the code on ARC challenges. The common theme you see is convergence. In the Sakana paper, after 250 calls, it converged and failed to expand and improve its results.

Would you agree that there's some margin—you can improve to some margin—and we don't know what that margin would be if we had agentic superintelligence? Do you think that margin would converge quite quickly, or do we just not know?

Superintelligence Strategy (Dan Hendrycks)

I don't know if it would saturate necessarily. Certainly, if they're at some capability level, this might affect the rate of improvement substantially, but obviously there has to be a limit because of physics somewhere. I would imagine there's a lot of room for improving their ability to handle things of more and more complexity, for instance.

Some of those sequences—or, in those Raven's Progressive Matrices, you can think of them as ARC-type things—have lower Kolmogorov complexity, and some have higher Kolmogorov complexity. That's just some notion from theoretical computer science, and I don't think humans are anywhere near the computational limit for that at all. I would guess that that number could keep going up, so I think fluid intelligence could get really high.

Basically, the IQ of the AIs—which wouldn't be all of their intelligence; it would be separate from their long-term memory, their visual processing ability, their reaction time, and so on—could get really high. They could solve extremely difficult mathematics problems and have really good intuitions for puzzles without taking much compute for it. I think that number could keep going up quite a bit.

We're quite limited. We've got big brains, but there's still not much hardware here, and you can imagine a bigger brain that would have pretty substantial capabilities, even for human intelligence across generations. There are the Flynn effects and things like that, and that sort of intelligence has greatly increased. Our brain size is quite limited by what can give what is able to go through the birth canal.

I expect the AIs wouldn't have that limit at all. They could, at the very least, keep getting faster and faster. You still have Moore's law, and you have GPU improvement rates, which are about 2× every 3 years. You also have scalability, as well as a better ability to transfer state than humans do, because a lot of this digital computation is more precise than analog.

I think they have some pretty substantial advantages that could really compound and accumulate. If you're saying that it's mostly externally computed, well, they can—humans, how many connections can people manage? They can do, I don't know, Dunbar's number, maybe 130 or so meaningful social connections. AIs could do thousands, tens of thousands, or millions, and all of these together can provide compounding effects that make them substantially more capable.

Superintelligence Strategy (Dan Hendrycks)

So I think that there may, however, be various points of saturation along the way. You may get a lot of the low-hanging fruit, sort of like convergent economies: a lot of economies caught up to somewhere in the vicinity of the US. It’s not like they kept going at the same rate; they just copied what was lying around, and then it was harder from then on out. So you may have getting to human level in some respects, and in some of these domains it may be harder if you don’t have a good measurement or feedback loop. It may be hard to keep getting better and better at that.

So I think it could be quite complicated. I’m agreeing with some points there, but conceptually I still think there’s a lot of room to go up in intelligence.

Yes, yes. As you said in your paper, when machines can do things as well as humans can, then we’re in big trouble. But I think the main philosophical difference between us is that, while we agree intelligence is about adaptivity, I’m a fan of specialized intelligence. I don’t think there is such a thing as generalized intelligence.

Certainly, when talking about the factored ways in which we think we have these constraints—and they run deep—we see the world using symmetries, and there’s this big phylogenetic tree of knowledge and thought which constrains how we think. Surely AIs would need to be constrained in a similar way. I appreciate what you’re saying: you could just scale this up a million times faster, but maybe it would need to be scaled much more than that. Maybe these creative intuitions and insights we have don’t come from the data. They don’t come from what’s inside; maybe they come from what’s outside. So I’m vaguely leaning toward this intuition that there’s something else unaccounted for in this estimation.

Superintelligence Strategy (Dan Hendrycks)

Interesting. They certainly could have a lot of sensors. If we’re saying that there needs to be a lot of extra variety from elsewhere, and it wouldn’t be just what’s inside of the data center, I think they could aggregate a lot of information, soak that up, and process a lot more of it. So I still think they could have some type of advantage, at least doing this a lot faster than people.

But certainly, if they’re in a vacuum, only speaking by themselves or only working by themselves—for instance, not learning anything—if there’s too much correlation in that population, that might make the exploration budget too low, and there wouldn’t be sufficient variety. I mean, evolution generally—I think there’s Fisher’s fundamental theorem of natural selection or something like that—which is that the rate of adaptation is in some ways directly proportional to the amount of variation. You’re pointing at ways in which it’s lacking in variation, but I think some of that could potentially be made up for. It could at least have the sensors that humans have, and more.

Another very interesting thing in your paper: when I think about AI risk in general—we’ve thought about this a little bit on the show before—it’s in terms of stability and destabilization and the relationship between offense and defense. You use this term “offense-dominant,” and you were saying that a destabilizing force would be if the AI is offense-dominant, then the defensive side of the equation couldn’t catch up. At the moment, if you imagine our state of affairs, we have a kind of Nash equilibrium where there are countervailing factors on the offense and defense sides. Can you tell me about that?

Superintelligence Strategy (Dan Hendrycks)

I think the offense-defense balance varies a lot by domain. Potentially, in many information battles—for instance, debates about the world—it might be a bit more defense-dominant, which would be a reason for things like free speech.

Meanwhile, other things might have more of a duality. For instance, really expert-level or really competent computer security teams might experience more of an offense-defense balance: something is identified, we patch the vulnerability very quickly, and the attackers keep up with the defenders quite well.

In other domains, like software on various forms of critical infrastructure, there’s more of an offense dominance, or attacker’s advantage, because a lot of the software just doesn’t get updated quickly. There are interoperability constraints. The software developer is no longer around. The software was made 30-plus years ago. Nobody even knows that it’s there. There aren’t strong enough economic incentives for doing this, and there are uptime requirements. Software on critical infrastructure is more of a sitting duck, and there you don’t experience a good offense-defense balance.

Likewise for bioweapons, certainly we have medicine, but we don’t have cures for everything. In fact, we spend a lot trying to find cures for many sorts of diseases and ways of addressing certain viruses. So it’s not necessarily the case that there’s a new pathogen and we’ll just find a cure a day later, it will be mass-proliferated across the globe, and everything will be taken care of. There is a substantial delay, and so it is more offense-dominant there too. You can imagine the attacker having a really substantial advantage, like a pathogen propagating throughout society before anybody shows symptoms, while lacking various monitoring mechanisms. We’re kind of sitting ducks for some of those.

So it varies by domain. I think some parts of cyber have a good offense-defense balance. Other parts of cyber don’t. Bio seems pretty offense-dominant, and I think this affects how you want to propagate the technology. If it’s potentially catastrophic and offense-dominant, that’s something you basically want to restrict. If it’s a WMD, like the WMD state, that’s not something you want to give to everybody. You don’t want to give everybody a nuke to make everybody safe. That’s not how it works. It’s people constraining each other’s intent, and some people just won’t have their intent be that constrainable.

Meanwhile, other things like defenses—home security systems or fences or whatever—you’d want to propagate more. So I think that should affect attitudes toward specific types of AI capabilities as well, in terms of what their offense does. Maybe that could change over time. Maybe you could improve critical infrastructure so that it becomes more offense-defense balanced, and then you can propagate AIs with fewer and fewer safeguards.

Say the economy gets a lot richer and GDP is higher. Then states are more willing to spend money on preventive measures for bioweapons. They’re willing to have more far-UVC systems, more wastewater monitoring, et cetera. That makes things closer to there being more of a balance, or at least not having as much of an attacker’s advantage. So you may want mass proliferation of these AI capabilities a bit later, once you have some of those safeguards in place.

Another thing I’m very interested in, which you’ve spoken about in the paper, is this concept of loss of control. Connor Leahy uses the term “fog of war” to talk about how, as layers upon layers of complexity build, there’s a complete illegibility that builds. In a sense, we already have that now. I do think the AI we have today has the same problem, but it’s not quite as extreme.

Systems at Google are very complex, and that’s why the market rewards Google engineers—they get paid an awful lot of money. But you’re talking about what happens when this gets taken to the extreme. You gave 3 examples: self-reinforcing dependence, irreversible entanglement, and cessation of authority. Can you explain what those are?

Superintelligence Strategy (Dan Hendrycks)

I describe this at more length in “Natural Selection Favors AI over Humans.” There’s a way in which we are becoming more dependent on these AI systems, that we’re starting to cede more of our decisions and cognitive processes to them, and that there will be more and more pressure to keep doing that without any clear limit at the economic level.

At the economic level, if it’s you versus a company that has AI systems that are just better than you, that company is going to win because those systems will be cheaper. At the military level, too, there’s a very strong incentive to, for instance, make the drones more autonomous because they can be jammed. Signal jamming makes them a lot less effective, so you make them autonomously move around.

These pressures will keep up, such that we’ll just voluntarily cede a lot of power in society to these sorts of systems. If we do it at a rate where we aren’t actually in control, that could be concerning. What does control exactly look like? When is it too much? When is it too irreversible? That’s a problem that we have to keep track of.

There are different outcomes. One is where you’ve actually just lost control: you can’t stop it, you can’t reverse it, you can’t bargain with it, or you can’t substantially steer it, and your fitness as a species is just collapsing as a consequence, or your livelihood evaporates.

Superintelligence Strategy (Dan Hendrycks)

Or there's a different outcome where you are like a retiree with a big retirement fund, where these sorts of AIs are working for you and they are doing your bidding, even though you aren't making nearly all of the critical decisions. So those are different outcomes, and it could be insidious or subtle as to whether we actually are having some of that counterfactual control versus whether we wind up in an unfortunate situation.

I think at least one thing that can help with this would be if AIs are better at forecasting or foreseeing outcomes and consequences. If we train them to do that, that could help make us more prescient and avoid some of these outcomes of extreme dependence and an erosion of control. That wouldn't be sufficient, but it would be helpful for seeing farther and knowing what we're getting ourselves into.

Wonderful. Dan, this has been an absolute pleasure and an honor having you on the show. Thank you so much for joining us today.

Superintelligence Strategy (Dan Hendrycks)

Yeah, thank you for having me.

Mutually Assured AI Malfunction [Dan Hendrycks] | BidClub