[BidClub_]
No Priors · · 36 min

No Priors Ep. 105 | With Director of the Center of AI Safety Dan Hendrycks

Sarah GuoNeil Tiwari

YouTube
TL;DR
  • Hendrycks argues that AI safety is primarily a geopolitical and economic problem, not something laboratories can solve through model alignment alone. Labs are “predetermined to race” and can add basic safeguards, but they cannot design away labor disruption, concentrated power, or US–China strategic competition. Even perfectly obedient national systems could be integrated rapidly into competing militaries and force both sides toward higher risk tolerance.

  • Near-term capability is uneven: AI is not yet decisive across national security, but reasoning models have recently begun to create biological national-security implications and are rounding the corner toward expert-level literature knowledge and wet-lab assistance. Hendrycks doubts today’s systems could independently execute a devastating grid attack, while warning that biological capabilities have recently become more consequential. His proposed control is simple identity-gating: an unidentified user asking how to culture a virus gets refused; a legitimate biotech company can “just speak to sales.”

  • The military discussion extends well beyond chatbots into drones, electronic warfare, command-and-control, and situational awareness. Hendrycks points to drones, weapons research, and better detection of nuclear submarines or hardened launch sites, which could disrupt second-strike capabilities without itself being a weapon. Guo adds electronic warfare and a Wall Street analogy in which human judgment became rows of people clicking “accept, accept, accept.” Reliability therefore becomes both a product constraint and a strategic-security variable.

  • Neither unilateral restraint nor a clean race to superintelligence survives Hendrycks’s second-order test. A voluntary pause without verification or force merely weakens the participant, but a visible trillion-dollar desert cluster intended to secure dominance would invite espionage, cyberattack, or preemption. He says that in some places, more than 30% of employees at top AI companies are Chinese nationals; excluding that talent would damage the US effort and potentially strengthen China.

  • Mutually assured AI malfunction, or MIM, is proposed with Eric Schmidt and Alexandr Wang as deterrence against destabilizing superweapon projects rather than a halt to ordinary AI competition. States would monitor rival programs and retain cyber options to disable data centers attempting a decisive strategic breakout: “We will not build a superweapon, and we’re going to be watching for other people building them too.” Competition in chips, drones, and conventional systems would continue.

  • Compute controls should prioritize chip visibility and nonproliferation to rogue actors, not assume China can be denied advanced capability indefinitely. Guo raises DeepSeek and recent releases as a challenge to a simple 100,000-chip compute-security premise; Hendrycks agrees that efficient training undermines a Manhattan Project strategy and says China could ultimately steal model weights. He nevertheless argues that basic end-use checks might have revealed where the “10% of NVIDIA’s chips” discussed in relation to Singapore were going.

  • The key capability inflection is reliable agency, not simply higher scores on closed-ended academic benchmarks. Humanity’s Last Exam measures the remaining frontier of closed-ended expert knowledge; near-ceiling performance would imply something like a superhuman mathematician or STEM scientist. Guo also mentions Enigma, though Hendrycks’s answer focuses on HLE and the broader closed-ended/agent distinction. Agents remain “near the floor”: models can solve difficult physics while failing to book a flight, and Hendrycks expects the economic “vibes really shift” once they can reliably complete multi-hour digital work.

Digest · the substance, structured for research

1. Safety is constrained by geopolitics before model design

  • Hendrycks entered AI safety because the technology’s trajectory looked consequential while its unpleasant tail risks were “systematically under-addressed.” His objective is broader than preventing catastrophe: understand the trajectory, channel it productively, and manage disruptions that technical model work cannot settle.

  • His institutional diagnosis is blunt: labs can refuse requests such as “help me make a virus,” but they are “kind of predetermined to race.” A company that meaningfully opts out risks becoming irrelevant, while different refusal data cannot change the prospect of mass labor disruption or automated digital work.

  • Alignment is therefore a subset of safety, not its synonym. An AI reliably obedient to the US and another reliably obedient to China could still be embedded rapidly into competing militaries; strategic pressure would raise risk tolerance even if both systems performed their principals’ bidding perfectly.

  • Guo’s pushback — worth keeping: lab leaders would not say they can do nothing, and every participant has economic equity in the outcome. Hendrycks allows that companies can pursue controllability research and policy advocacy, but maintains that the broader problem is geopolitically determined.

2. Near-term danger is jagged, but access controls buy a lot

  • Hendrycks distinguishes trajectory from present capability: in many national-security domains AI is not yet powerful, though “this could well change within a year’s time.” Today’s systems probably cannot execute a devastating grid attack by a malicious actor, while reasoning models are rounding the corner toward expert-level biological literature knowledge and practical wet-lab assistance.

  • Guo points to defensive cybersecurity and biology-discovery companies as evidence that competition produces near-term benefits. Hendrycks rejects a large safety-versus-benefit tradeoff here: unidentified users requesting step-by-step virus-culturing help can be blocked, while verified biotech customers receive the capability through an enterprise account — “just speak to sales.”

  • His threat map separates actors and applications: bioweapons make more sense as a non-state risk; cyber operations matter for both states and non-state actors. Hendrycks points to drones, weapons research including exotic EMPs, and situational awareness, including locating submarines or hardened launchers. Guo extends the discussion to electronic warfare, radio, radar, targeting, and command-and-control, especially in Ukraine.

  • Guo offers the Wall Street anecdote, where human oversight degraded into rows of people clicking “accept.” Hendrycks says automating more decision-making would not surprise him and turns the issue into one of reliability research.

3. Pauses and monopoly races both fail the second-order test

  • A voluntary pause without “teeth” simply lets worse actors advance. Treaties require verification, credible enforcement, or a threat of force; cyberattacks and corporate espionage offer little evidence that norms alone would hold.

  • A straight race is defensible for chips or drones, but Hendrycks rejects racing to turn superintelligence into a weapon. Espionage makes durable monopoly unlikely, yet removing Chinese researchers would also be self-defeating: he says that in some places more than 30% of employees at top AI companies are Chinese nationals, and many could return to China. He favors easier immigration for very talented AI researchers while saying the issue should remain separate from Southern-border policy.

  • China would not watch passively as the US built a trillion-dollar compute cluster “totally visible from space” as a bid for dominance. Hendrycks compares the logic to early nuclear proposals for preempting the USSR: multinational talent, interdependence, and limited time may mean the supposed window for a secure monopoly never exists.

4. MIM turns shared cyber vulnerability into deterrence

  • Mutually assured AI malfunction, or MIM, proposed with Eric Schmidt and Alexandr Wang, applies nuclear-style shared vulnerability once AI becomes pivotal and can automate AI research. A rival approaching a decisive superweapon would face espionage, sabotage, or a cyberattack on its data center; that expectation is intended to deter the destabilizing project without requiring every form of AI development to stop.

  • The analogy has limits. Hendrycks expects coordination where interests overlap — keeping dangerous capabilities from terrorists, as with chemical and biological weapons, and avoiding projects that could let one country “crush” another — while conventional competition in drones and other systems continues.

  • Calibration matters because, if AI chips become the “currency of economic power,” turning China’s export-control pain dial fully upward could strengthen its incentive to invade Taiwan. Hendrycks notes that China already wants to; this would give it additional reason.

  • The practical package is narrower: a CIA cell tracking foreign AI programs, Cyber Command purchasing disablement options, chip-location licensing, allied notification exemptions, and prioritized end-use inspections.

  • Against Guo’s DeepSeek challenge, Hendrycks agrees efficient training undermines the desert-cluster strategy: controls cannot robustly eliminate a great power’s capability or make it abandon the economic value of AI. China may obtain a fraction of the chips or steal weights anyway; controls remain useful for visibility and rogue-actor nonproliferation, particularly if officials actually investigate routes such as Singapore. He says AI-chip export controls were not a priority for Bureau of Industry and Security leadership and that basic checks could have exposed diversion to China.

5. Evals show oracle-like intelligence before reliable agency

  • Humanity’s Last Exam extends Hendrycks’s earlier benchmark work, including MMLU and the MATH dataset. Guo also mentions Enigma, though Hendrycks’s response here focuses on HLE and the broader closed-ended/agent split. Professors and researchers submitted unusually difficult research-grade questions with definitive, closed-ended answers, creating what he describes as a possible “end of the road” for exam-style academic evaluation.

  • Near-ceiling performance would roughly exhaust that benchmark genre and signal something like a superhuman mathematician or STEM scientist on closed-ended problems. It would not establish competence at open-ended work, where defining, pursuing, and completing a task matters as much as knowing an answer.

  • Agent evaluations should instead assign real digital tasks, allow a few hours of work, and check completion. Current systems remain “extremely defective as agents” and near the floor, although Hendrycks hedges that this “could possibly change overnight.”

  • The frontier remains jagged: systems answer hard physics questions but cannot fold clothes, and may surpass humans at mathematics while failing to book a flight. For verifiable reasoning, humans might still run an AI proof through a proof checker; by contrast, if systems develop better “taste,” that would be harder to confirm.

  • Hendrycks thinks AI is on track for “really good oracle-like skills” before it can reliably act on people’s behalf. Once it acquires agent skills, few barriers remain to major economic impact, and AI becomes “in a category of its own.”

Sarah Guo

Hi listeners and welcome back to No Priors. Today I'm with Dan Hendrycks, AI researcher and director of the Center for AI Safety. He's published papers and widely used evals such as MLU and, most recently, Humanity's Last Exam. He's also published Superintelligence Strategy alongside authors including former Google CEO Eric Schmidt and Scale founder Alex Wang. We talk about AI safety and geopolitical implications, analogies to nuclear compute security, and the state of evals.

Dan, thanks for doing this.

Dan Hendrycks

Glad to be here.

Sarah Guo

How did you end up working on AI safety?

Dan Hendrycks

AI was pretty clearly going to be a big deal if I just thought through its conclusion. Early on, it seemed like other people were ignoring it because it was weirder or not that pleasant to think about. It's hard to wrap your head around, but it seemed like the most important thing during this century.

So I thought that would be a good place to develop my career toward, and that's why I started on it early on. Since it would be such a big deal, we need to make sure that we can think about it properly, channel it in a productive direction, and take care of some sort of tail risks, which are generally systematically under-addressed. That's why I got into it. It's a big deal, and people weren't really doing much about it at the time.

Sarah Guo

What do you think of as the Center for AI Safety's role versus safety efforts within the large labs?

Dan Hendrycks

There aren't that many safety efforts in the labs even now. I think the labs can focus on doing some very basic measures to refuse queries related to, "Help me make a virus," and things like that.

But I don't think labs have an extremely large role in making this go well overall. They're kind of predetermined to race. They can't really choose not to unless they would no longer be a relevant company in the arena. I think they can reduce terrorism risks or some accidents, but beyond that, I don't think they can dramatically change the outcomes in any substantial way.

A lot of this is geopolitically determined. If companies decide to act very differently, there's the prospect of competing with China, or maybe Russia will become relevant later. As that happens, this constrains their behavior substantially.

I've been interested in tackling AI at multiple levels. There are things companies can do to have some very basic antiterrorism safeguards, which are pretty easy to implement. There's also the economic effects that will need to be managed well, and companies can't really change how that goes either. It's going to cause mass disruptions to labor and automate a lot of digital labor.

If they tinker with the design choice or add some different refusal data, it doesn't change that fact. Safety—making AI go well—and risk management are much broader problems. They have some technical aspects, but I think that's a small part of it.

Sarah Guo

I don't know that the leaders of the labs would say, "We can do nothing about this," but maybe it's also a question of semantics. Everybody also has equity in this equation, right? Can you describe how you think of the difference between alignment and safety?

Dan Hendrycks

I'm just using safety as a catch-all for dealing with risks. There are other risks, like if you never get really intelligent AI systems, that poses some risks in itself. There are other sorts of risks that are not necessarily technical, like concentration of power.

I view the distinction between alignment and safety as alignment being a sort of subset of safety. Obviously, you want the value systems of the AIs to be in keeping with or compatible with, say, the US public for US AIs, or with you as an individual. But that doesn't necessarily make it safe.

If you have an AI that's reliably obedient or aligned to you, this doesn't make everything work totally well. China can have AIs that are totally aligned with them, and the US can have AIs that are totally aligned with them. You're still going to have a strategic competition between the two.

They're going to need to integrate AI into their militaries, and they're probably going to need to integrate it really quickly. Competition is going to force them to have a high risk tolerance in the process. So even if the AIs are reliably doing their principals' bidding, this doesn't necessarily make the overall situation perfectly fine.

I think it's not just a question of reliability or whether they do what you want. There are other structural pressures that cause this to be riskier, like geopolitics at the highest level.

Sarah Guo

With increasingly capable models and weights, why do we care about AI from a national security perspective? What's the most practical way it matters in geopolitics or gets used as a weapon?

Dan Hendrycks

I think AI isn't that powerful currently in many respects, so in many ways it's not actually that relevant for national security currently. This could well change within a year's time. Generally, I've been focused on the trajectory that it's on, as opposed to saying, "Right now, it is extremely concerning."

That said, for cyber, I don't think AIs are currently that relevant for being able to pull off a devastating cyberattack on the grid by a malicious actor. We should look at cyber, be prepared, and think about what its strategic implications are.

There are other capabilities, like biology. The AIs are getting very good at STEM PhD-level topics, and that includes biology. I think they're rounding the corner on being able to provide expert-level capabilities in terms of their knowledge of the literature or even helping in practical wet-lab situations.

So I do think that, on the biology aspect, they already have national security implications. But that's only very recent, with the reasoning models. In many other respects, they're not as relevant. It's more prospective that AI could become the way in which a nation might try to dominate another nation, and the backbone for not just war but also economic security.

The amount of ships that the US has versus China might be the determinant of which country is the most prosperous and which one falls behind. This is all prospective. I don't think it's just speculative. It's speculative in the same way that NVIDIA's valuation is speculative, or the valuations behind AI companies are speculative. It's something that I think a lot of people are expecting, and expecting fairly soon.

Sarah Guo

Yeah, it's quite hard to think about time horizons in AI. We invest in things that I think of as medium-term speculative, but they get pulled in quite quickly.

Just because you mentioned both cyber and bio, we're investors in companies like Cymulate or Cylance on the defensive cybersecurity side, or Chai and Somite on the biotech discovery side, modeling different systems in biology that will help us with treatments. How do you think about the balance of competition, benefits, and safety? Some of these things, I think, are working effectively in the near term on the positive side as well.

Dan Hendrycks

I don't get this big trade-off between safety and the benefits of AI. You're just taking care of a few tail risks. For bio, if you want to expose those capabilities, just talk to sales and get the enterprise account.

You can have the little refusal mechanism for biology, but if you just create an account and ask it how to culture this virus, show it a picture of your Petri dish, and ask what the next step should be, then, yeah, if you want access to those capabilities, you can speak to sales. That's basically the X in an X-risk management framework: we're just not exposing those expert-level capabilities to people whose identities we don't know. But if we do know who they are, then sure, give them access.

Likewise with cyber, I think you can very easily capture the benefits while taking care of some pretty avoidable tail risks. Once you have that, you've basically taken care of malicious use for the models behind your API, and that's about the best that you can do as a company.

You could try to influence policy by using your voice or something, but I don't see a substantial amount that companies could do. They could do some research to make the models more controllable, or try to make policymakers more aware of the situation more broadly in terms of where we're going, because I don't think policymakers have internalized what's happening at all.

They still think it's just hype, and they don't actually believe—or the companies and their employees don't actually believe—that we could get AGI in the next few years. So I don't see really substantial trade-offs there. I see much more substantial complications when we're dealing with the right level of stringency in export controls, for instance.

If you turn the pain dial all the way up for China in export controls, and if AI chips are the currency of economic power in the future, then this increases the probability that they want to invade Taiwan. They already want to; this gives them all the more reason.

If AI chips are the main thing, and they're not getting any of them—and they're not even getting the latest semiconductor manufacturing tools for making cutting-edge CPUs, let alone GPUs—those are some other types of complicated problems that we have to address and calibrate appropriately.

But in terms of just mitigating biology risks, speak to sales. If you're Genentech or a biotech startup, then you have access to those capabilities.

Sarah Guo

Perhaps what's a way you actually expect AI to get used as a weapon beyond virology and cybersecurity?

Dan Hendrycks

I wouldn't expect a bioweapon from a state actor. From a non-state actor, that would make a lot more sense.

Cyber makes sense from state actors and non-state actors. Then there are drone applications. These could disrupt other things, and they could help with other types of weapons research, such as exploring exotic EMPs. They could help create better types of drones and substantially help with situational awareness, so one might know where all the nuclear submarines are.

Some advancement in AI might be able to help with that, and that could disrupt our second-strike capabilities and mutually assured destruction. Those are some geopolitical implications. It could potentially bear on nuclear deterrence.

That's not even a weapon. The example of just heightened situational awareness and being able to pinpoint where hardened land-based nuclear launchers are, or where nuclear submarines are, is just informational, but could nonetheless be extremely disruptive and destabilizing.

Outside of that, the default conventional AI weapon would be drones. I don't know if that makes sense, or that countries would compete on that, and I think it would be a mistake if the US weren't trying more in manufacturing drones.

Sarah Guo

Yeah, I started working recently with an electronic warfare company. I think there's a massive lack of understanding of just the basic concept: we have autonomous systems, they all have communication systems, and our missile systems have targeting and communication systems.

From a battlefield-awareness and control perspective, a lot of that effort will be won with radio, radar, and related systems. I think there's an area where AI is going to be very relevant and is already very relevant in Ukraine.

Speaking about AI assisting with command and control, I remember hearing some story about how, on Wall Street, humans used to—you always had a human in the loop for each decision. At a later stage, before they removed that requirement on Wall Street, you just had rows of people clicking the "Accept, accept, accept" button. We're getting to a similar state in some contexts with AI.

Dan Hendrycks

It wouldn't surprise me if we ended up automating more of that decision-making. But this just turns into questions of reliability, and doing reliability research seems useful.

I think people are largely thinking that the push for risk management is to do some sort of pausing or something like that. An issue is that you need teeth behind an agreement. If you do it voluntarily, you just make yourself less powerful and let worse actors get ahead of you.

You could say, "We'll sign a treaty," but assuming that the treaty will be followed would be very imprudent. You would actually need some sort of threat of force or something to back it up, or some verification mechanism. But absent that, if it's entirely voluntary, this doesn't seem like a useful thing at all.

Sarah Guo

But absent that, as a proxy, there's clearly been very little compliance with either treaties or norms around cyberattacks and corporate espionage, right?

Dan Hendrycks

Yeah. Corporate espionage, for instance, was one strategy behind this voluntary-pause strategy, with people thinking that equals safety. Then maybe last year there was that paper, Situational Awareness: The Decade Ahead, written by Leopold Aschenbrenner. He's a sort of safety person.

His idea was, "Let's instead try and beat China to superintelligence as much as possible." But that has some weaknesses because it assumes that corporate espionage will not be a thing at all, which is very difficult to do.

In some places, more than 30% of the employees at these top AI companies are Chinese nationals. This is not feasible. If you're going to get rid of them, they're going to go to China, and they're probably going to beat you because they're extremely important for the US's success.

So you're going to want to keep them here, but that's going to expose you to some information-security issues. That's just too bad.

Sarah Guo

Do you have a point of view on how we should change immigration policy, if at all, given these risks?

Dan Hendrycks

I would, of course, claim that the policy on this would be totally separate from Southern border policy and broader policy. But if we're talking about AI researchers, if they're very talented, then I think you want to make it easier. I think it's probably too difficult for many of them to stay currently.

That discussion should be kept totally separate from Southern border policy.

Sarah Guo

Just in terms of broad strokes, what are things that you think won't work? Voluntary compliance and assuming that'll happen, or just straight race?

Dan Hendrycks

We want to be competitive, and I think racing in other spheres, say drones or AI chips, seems fine. If you're saying, "Let's race to superintelligence to try and turn that into a weapon," and they're not going to do the same, or they're not going to have access to it, or they're not going to prevent that from happening, that seems like quite a tall claim.

If we did have a substantially better AI, they could just co-opt it; they could just steal it, unless you had really strong information security. You could move the AI researchers out to the desert, but then you're reducing your probability of actually beating them because a lot of your best scientists would end up going back to China.

Even then, if there were signs that they were really pulling ahead and going to be able to get some powerful AI that would enable China—or that would enable the US—to crush China, they would then try to deter them from doing something like that.

They're not going to sit idly by and say, "You know what? Go ahead, develop your superintelligence or whatever, and then you can boss us around, and we'll just accept your dictates till the end of time."

That, I think, is a failure of some sort of second-order reasoning: how would China respond to this sort of maneuver if we're building a $1 trillion compute cluster in the desert, totally visible from space? The only plausible read on this is that it's a bid for dominance or a sort of monopoly on superintelligence. It reminds me of the nuclear era. There was a brief period when some people were saying, "We have to just preemptively or preventively destroy the USSR." Even people who were normally pacifists, like Bertrand Russell, were advocating for this. The opportunity window for that maybe never existed, but there was a prospect of it for some time. I don't think that opportunity window really exists here because of the complex interdependence and multinational talent dependence in the United States. I don't think you can have China be totally excluded from any awareness or ability to gain insight into or imitate what we're doing here. We're clearly nowhere close to that as a real environment right now, right?

Sarah Guo

Right. It would take years to do well, and I don't even think the timelines for some very powerful AI systems mean there might be enough time to do that securitization anyway.

You propose, along with Eric Schmidt and Alexandr Wang, a new deterrent regime: mutually assured AI malfunction, or MIM. It's a bit of a scary acronym and also a nod to mutually assured destruction. Can you explain MIM in plain language?

Dan Hendrycks

Let's think of what happened in nuclear strategy. Basically, a lot of states deterred each other from doing a first strike because they could then retaliate. They had a shared vulnerability. They were saying, "We're not going to take this really aggressive action of trying to make a bid to wipe you out, because that will end up causing us to be damaged."

Later on, when AI is more salient, when it's viewed as pivotal to the future of a nation, and when people are on the verge of making a superintelligence—when they can say, "Automate pretty much all AI research"—I think states would try to deter each other from trying to leverage that to develop something like a superweapon.

That could allow one country to crush the others, or allow those AIs to conduct a really rapid, automated AI research-and-development loop that could bootstrap them from their current levels to something superintelligent, vastly more capable than any other system out there.

I think later on it becomes so destabilizing that China just says, "We're going to do something preemptive, like a cyberattack on your data center," and the US might do that to China.

Russia, coming out of Ukraine, will reassess the situation and become situationally aware. It will think, "What's going on with the US and China? My goodness, they're so focused on AI."

Let's say it's later in the year, when a big chunk of software engineering is starting to be impacted by AI. Russia might say, "Oh, wow, this is looking pretty relevant. If you try to use this to crush us, we will prevent that by doing a cyberattack on you, and we will keep tabs on your projects."

It's pretty easy for them to do espionage. All they need to do is a zero-day attack on Slack, and then they can know what DeepMind is up to in very high fidelity, as well as OpenAI, xAI, and others. It's pretty easy for them to do espionage and sabotage.

Right now, they don't need to threaten that because it's not at the level of severity. It's not actually that potentially destabilizing; it's still too distant. A lot of decision-makers still aren't taking this AI stuff that seriously, relatively speaking, but I think that'll change as it gets more powerful.

Then I think this is how they would end up responding. This keeps us from winding up in a situation where we're doing something extremely destabilizing, like trying to create a weapon that enables one country to totally wipe out the other, as was proposed by people like Leopold Aschenbrenner.

Sarah Guo

What are the parallels here that you think make sense to nuclear weapons, and which ones don't?

Dan Hendrycks

More broadly, AI is a dual-use technology, in that it has civilian applications, military applications, and economic applications. Its economic applications are still limited in some ways, and likewise its military applications are still limited, but I think that will keep changing rapidly.

Chemical technology was important for the economy and had some military use, but countries coordinated not to go down the chemical-weapons route. Biology can be used as a weapon and has enormous economic applications, and likewise with nuclear technology.

For each of those technologies, countries did eventually coordinate to make sure they didn't wind up in the hands of rogue actors like terrorists. There have been a lot of efforts to make sure rogue actors don't get access to them and use them against their adversaries, because it's in neither side's interest.

Bioweapons and chemical weapons are a poor man's atom bomb. That's why we have the Chemical Weapons Convention and the Biological Weapons Convention. There's some shared interest there. They might be rivals in other senses, in the way that the US and the Soviet Union were rivals, but they're still able to coordinate on that because it's incentive-compatible.

It doesn't benefit them in any way if terrorists have access to these sorts of things. It's just inherently destabilizing. So I think that's an opportunity for coordination.

That isn't to say that they have an incentive to pause all forms of AI development. It may mean that they would be deterred from some particular forms of AI development, particularly those that have a very plausible prospect of enabling one country to get a decisive edge over another and crush it.

So, no superweapon-type stuff, but more conventional types of warfare, like drones, will continue. I expect that they'll continue to race and probably not even coordinate on anything like that. That's just how things will go. It's like bows and arrows and nuclear weapons: it made sense for them to develop those sorts of weapons and threaten each other with them.

Sarah Guo

If you could magically adopt some policy or action in the current administration, what is the first step here?

Dan Hendrycks

The first step is, "We will not build a superweapon, and we're going to be watching for other people building them too."

As I've been alluding to throughout this conversation, what would the companies do? Not that much. They would add some basic antiterrorism safeguards. This is technically pretty easy, unlike refusal for other things. If you're trying to deal with crimes and torts, that's harder because it's much messier and overlaps with typical everyday interaction.

I think the asks for states are not that challenging either. It's just a matter of doing them. One step would be for the CIA to have a cell doing more espionage of other states' AI programs, so that we have a better sense of what's going on and aren't caught by surprise.

Secondly, maybe some part of the government, such as Cyber Command, which has a lot of cyberoffensive capabilities, gets some cyberattacks ready to disable data centers in other countries if they're looking like they're running or creating a destabilizing project.

That's it for the deterrence. For nonproliferation of AI chips to rogue actors in particular, I think there would be some adjustments to export controls. In particular, we need to know where the AI chips are reliably, for the same reason we want to know where our fissile material is, and for the same reason that we want Russia to know where its fissile material is. That's just generally good information to collect.

That can be done with some very basic statecraft: having a licensing regime, and having allies notify you whenever chips are being shipped to a different location. They would get a license exemption on that basis, and then you would have enforcement officers prioritize some basic inspections for AI chips and end-use checks.

All of these are a few texts away or a basic document away. I think that 80/20 is a lot of it. Of course, this is always a changing situation.

Safety isn't, as I've been trying to reinforce, really that much of a technical problem. This is more of a complex geopolitical problem with technical aspects. Later on, maybe we'll need to do more. There might be some new risk sources that we need to take care of and adjust for.

But right now, I think that spies for the CIA, sabotage with Cyber Command, building up those capabilities, and buying those options would take care of a lot of the risk.

Sarah Guo

Let's talk about compute security. If we're talking about 100,000 networked, state-of-the-art chips, you can tell where that is. How do DeepSeek and the recent releases they've had factor into your view of compute security, given that export controls have clearly led to innovation toward highly compute-efficient pretraining that works on chips China can import at what might be considered an irrelevant scale—a much smaller scale today?

It seems directionally hard to see training becoming less efficient, even if we want to scale it up. Does that change your view at all?

Dan Hendrycks

No. It just sort of undermines other types of strategies, like the Manhattan Project strategy of moving people out to the desert and building a big cluster there.

What it shows is that you can't rely as much on restricting another superpower's ability to make models. You can restrict their intent, which is what deterrence does, but I don't think you can reliably or robustly restrict their capabilities.

You can restrict the capabilities of rogue actors, and that's what I would want compute security and export controls to facilitate. We should make sure it doesn't wind up in the hands of Iran or another rogue actor.

China will probably keep getting some fraction of these chips, but we should basically try to know where they are more reliably, and we can tighten things up. You could even coordinate with China to make sure the chips aren't winding up in rogue actors' hands.

I should also say that export controls weren't actually a priority among the leadership at the Bureau of Industry and Security, to my understanding. AI chips were a substantial priority for some people, but for the enforcement officers, did any of them go to Singapore to see where 10% of NVIDIA's chips were going? I think they would have very quickly found that they were going to China.

Some basic end-use checks would have taken care of that. It's not that export controls don't work. We've done nonproliferation of lots of other things, like chemical agents and fissile material, so it can be done if people care.

Even so, I still think that if you really tighten the export controls so that China can't get any of those chips at all, and this is one of your biggest priorities, they're just going to steal the weights anyway. I think it will be too difficult to totally restrict their capabilities, but I think you can restrict them through deterrence.

Sarah Guo

It also seems like either this stuff is powerful or it's not. It seems infeasible to me, given the economic opportunity, that China will say, "We don't need the capability."

Dan Hendrycks

Yeah, I fail to see a version of the world where the leadership of another great power that believes there is value here says, "We don't need that," from an economic-value perspective.

For a lot of these things, it would perhaps be nicer if everything went 3× slower and there were fewer mess-ups. If there were some magic button that would do that, maybe it would be useful. I don't know whether that's true, actually. I don't have a position on that.

Given the structural constraints and the competitive pressures between these companies and these states, a lot of these things become infeasible. Many of the gestures that could be useful for risk mitigation become less tractable when you consider the structural realities.

That said, there still would be some pausing or halting of the development of particular projects that you could potentially lose control of, or that, if controlled, would be very destabilizing because they would enable one country to crush the other.

I think people's conception of what risk management looks like is that it's a peacetime thing, or something like that—it's all kumbaya, and we just have to ignore the structural realities of operating in this space.

Instead, the right approach is more like nuclear strategy. It's an evolving situation. There are some basic things you can do: you're probably going to need a stockpile of nuclear weapons, you're going to need to secure a second strike, you're going to need to keep an eye on what the other side is doing, and you're going to need to make sure there isn't proliferation to rogue actors when the capabilities are extremely hazardous.

This is a continual battle, but it's not clearly going to be an extremely positive thing no matter what, and it's not going to be doomsday no matter what. Nuclear strategy was obviously risky business. The Cuban Missile Crisis came pretty close to an all-out nuclear war.

It depends on what we do. Some basic interventions and very basic statecraft can take care of a lot of these risks and make the situation manageable. I imagine we're left with more domestic problems, like what to do about automation and things like that, but I think maybe we'll be able to get a handle on some of the geopolitics here.

Sarah Guo

I want to change tack for our last couple of minutes and talk about evals. It's obviously very related to safety and understanding where we are in terms of capability.

You came out with the strikingly named Humanity's Last Exam eval, and then also Enigma. Why are these relevant, and where are we with evals?

Dan Hendrycks

I've been making evaluations to try to understand where we are in AI for as long as I've been doing AI research. Previously, I've done datasets like MMLU and the MATH dataset. Before ChatGPT, there were things like ImageNet and other sorts of benchmarks.

Humanity's Last Exam was basically an attempt to get at the end of the road for evaluations and benchmarks based on exam-like questions—ones that test some sort of academic knowledge. We asked professors and researchers around the world to submit a really challenging question, and then we added those questions to the dataset.

It's a big collection of what professors, for instance, would encounter as challenging problems in their research, with a definitive, closed-ended, objective answer. I think the genre of a closed-ended question, where the answer is multiple-choice or a simple short answer, will roughly be exhausted when performance on this dataset is near the ceiling.

When performance is near the ceiling, I think that would basically be an indication that you have something like a superhuman mathematician or a superhuman STEM scientist, at least in areas where closed-ended questions are useful, such as mathematics.

But it doesn't get at other things, such as the ability to perform open-ended tasks. That's more of an agent-type evaluation, and I think that will take more time. We can try to measure directly what its ability is to automate various digital tasks: collect various digital tasks, have it work on them for a few hours, and see whether it successfully completes them.

Coming out soon, we have a test for closed-ended questions that test knowledge in academia, including areas like mathematics, but they're still very bad at agent tasks. This could possibly change overnight, but they're still near the floor. I think they're still extremely defective as agents.

There will need to be more evaluations for that, but the overall approach is just to try to understand what's going on: what's the rate of development? That way, the public can at least understand what's happening.

If all evaluations are saturated, it's difficult to even have a conversation about the state of AI. Nobody really knows exactly where it is, where it's going, or what the rate of improvement is.

Sarah Guo

Is there anything that qualitatively changes when these models and model systems are just better than humans—when they're exceeding human capability? Does that change our ability to evaluate them?

Dan Hendrycks

I think the intelligence frontier is just so jagged. What they can and can't do is surprising. They still can't fold clothes, but they can answer a lot of tough physics problems. There are complicated reasons for why that is, so it's not uniform.

In some ways, they'll be better than humans. It seems totally plausible that they'll be better than humans at mathematics before too long, but still not able to book a flight. The implication is that they might be better in some limited ways, with limited influence, but that won't necessarily generalize to other things.

I do think it's possible that they'll be better at reasoning skills than us. We could still have humans checking their work because we can verify it. If an AI mathematician is better than a human, humans can still run the proof through a proof checker and confirm that it was correct.

In that way, humans can still understand what's going on in some ways. But in other ways, if they're getting better taste in things—if that makes any sense, although maybe it doesn't make philosophical sense—that would be pretty difficult for people to confirm.

I think we're on track overall to have AIs with really good oracle-like skills. You can ask them things, and they say something insightful, very nontrivial, or push the bounds of knowledge in some particular way. But they may not necessarily be able to carry out tasks on behalf of people for some while.

I think this is why we don't take AIs that seriously: they still can't do a lot of very trivial stuff. But when they get some of the agent skills, I don't think there are many barriers to their economic impact, or to people thinking that this is more than just an interesting thing.

That's an emergent property with agent skills. The vibes really shift, and it's pretty clear that this is much bigger than some prior technology, like the App Store or social media. It's in a category of its own.

Sarah Guo

Dan, thanks for doing this. It's a great conversation.

Dan Hendrycks

Glad to be here. Thank you for having me.

Sarah Guo

Find us on Twitter at No Priors Pod. Subscribe to our YouTube channel if you want to see our faces. Follow the show on Apple Podcasts, Spotify, or wherever you listen. That way you get a new episode every week. Sign up for emails or find transcripts for every episode at no-pri.com.

No Priors Ep. 105 | With Director of the Center of AI Safety Dan Hendrycks | BidClub