No Priors 第105期|对话 AI Safety 中心主任 Dan Hendrycks
Hendrycks 认为,AI 安全首先是地缘政治和经济问题,而不是实验室仅靠模型对齐就能解决的问题。 实验室“注定要竞速”,可以加上基础安全措施,但无法通过设计模型来消除劳动力重构、权力集中或美中战略竞争。即便两个国家的系统都能完美服从本国意志,也可能被迅速整合进彼此竞争的军队,迫使双方提高风险容忍度。
短期能力并不均衡:AI 尚未在所有国家安全领域形成决定性影响,但推理模型近期已开始产生生物国家安全含义,并正逼近专家级文献知识和湿实验室辅助能力。 Hendrycks 怀疑当前系统能否独立发动一次灾难性的电网攻击,但警告生物能力近期已变得更具现实影响。他提出的控制方式很简单:身份不明的用户询问如何培养病毒,就拒绝回答;合法生物科技公司则可以“直接找销售”(“just speak to sales”)。
军事讨论远不止聊天机器人,还涉及无人机、电子战、指挥控制和态势感知。 Hendrycks 提到无人机、武器研发,以及更准确地探测核潜艇或加固发射场;即使这种能力本身不是武器,也可能破坏二次打击能力。Guo 又补充了电子战,并讲了一个华尔街类比:人类判断最终退化成一排人不断点击“接受、接受、接受”。因此,可靠性既是产品约束,也是战略安全变量。
单边克制和直奔超级智能的竞赛,都经不起 Hendrycks 的二阶检验。 没有核查或强制力的自愿暂停只会削弱参与者;但一个清晰可见、耗资1万亿美元、旨在确保主导地位的沙漠集群,则会招致间谍活动、网络攻击或先发制人的打击。他说,在一些地方,顶尖 AI 公司超过30%的员工是中国籍;排除这部分人才会损害美国自身努力,并可能反而增强中国。
Hendrycks 与 Eric Schmidt、Alexandr Wang 提出的“相互确保 AI 故障”(MIM),旨在威慑破坏稳定性的超级武器项目,而不是叫停普通 AI 竞争。 各国将监控对手项目,并保留网络能力,禁用试图实现决定性战略突破的数据中心:“我们不会制造超级武器,也会盯着其他人是否在制造超级武器。” 芯片、无人机和常规系统的竞争仍将继续。
算力管控应优先提高芯片去向的可见性,并防止其扩散给流氓行为体,而不是假设可以无限期阻止中国获得先进能力。 Guo 提到 DeepSeek 和近期发布的模型,对“需要10万枚芯片才能确保算力安全”的简单前提构成挑战;Hendrycks 同意,高效训练削弱了曼哈顿计划式战略,中国最终也可能窃取模型权重。但他仍认为,基础的最终用途核查或许能揭示此前围绕新加坡讨论的“NVIDIA 芯片的10%”究竟流向了哪里。
关键能力拐点在于可靠的代理能力,而不只是封闭式学术基准上的更高分数。 Humanity’s Last Exam 衡量封闭式专家知识的剩余前沿;接近满分将意味着出现类似超人类数学家或 STEM 科学家的能力。Guo 也提到 Enigma,但 Hendrycks 的回答聚焦于 HLE,以及更广义的封闭式问题与代理能力之分。代理系统目前仍“接近地板”:模型能解高难度物理题,却不会订机票;Hendrycks 预计,等它们能可靠完成持续数小时的数字工作后,经济“体感真的会发生变化”。
1. 地缘政治先于模型设计约束 AI 安全
Hendrycks 之所以进入 AI 安全领域,是因为他认为这项技术的演进轨迹影响重大,而其中令人不安的尾部风险却被“系统性地忽视”。他的目标不只是防止灾难:还包括理解技术走向、把它引向更有益的方向,并管理那些单靠模型技术无法解决的冲击。
他对机构的判断很直接:实验室可以拒绝“帮我制造病毒”之类的请求,但它们“某种程度上注定要竞速”。一家真正选择退出的公司可能被淘汰,而不同的拒答数据也无法改变大规模劳动力重构或自动化数字工作的前景。
因此,对齐只是安全的一部分,不能与安全画等号。一个可靠服从美国、另一个可靠服从中国的 AI,仍可能被迅速嵌入彼此竞争的军队;即使两个系统都完美执行本国决策者的意志,战略压力也会推高风险容忍度。
Guo 的反驳值得保留:实验室负责人不会说自己什么都做不了,而且每个参与者都在经济利益上与结果绑定。Hendrycks 承认,公司可以推进可控性研究并参与政策倡议,但仍坚持认为,更大的问题由地缘政治决定。
2. 短期风险呈锯齿状,但访问控制能解决很多问题
Hendrycks 区分了发展轨迹与当前能力:在许多国家安全领域,AI 目前还不够强,但“这很可能在一年内改变”。当前系统大概还无法被恶意行为体用来发动灾难性的电网攻击;与此同时,推理模型正逼近专家级生物学文献知识和实用湿实验室辅助能力。
Guo 以防御性网络安全和生物发现公司为例,说明竞争也会带来短期收益。Hendrycks 不认为这里存在巨大的安全与收益权衡:身份不明的用户若要求逐步指导如何培养病毒,可以被拦截;经过验证的生物科技客户则能通过企业账户获得相应能力——“直接找销售”(“just speak to sales”)。
他的威胁地图区分了行为体和应用场景:生物武器更适合作为非国家行为体风险来理解;网络行动则同时影响国家和非国家行为体。Hendrycks 提到无人机、包括新型 EMP 在内的武器研发,以及态势感知,包括定位潜艇或加固发射装置。Guo 又把讨论扩展到电子战、无线电、雷达、目标识别和指挥控制,尤其是在乌克兰战场上的应用。
Guo 讲了一个华尔街轶事:人类监督最终退化成一排人不断点击“接受”。Hendrycks 说,进一步自动化决策并不会让他意外,问题最终转化为可靠性研究。
3. 暂停和垄断竞赛都过不了二阶检验
没有“牙齿”的自愿暂停只会让更糟糕的行为体取得进展。条约需要核查、可信的执行机制,或使用武力的威胁;网络攻击和企业间谍活动都很难证明,仅靠规范就能维持秩序。
芯片或无人机可以直接竞速,但 Hendrycks 反对竞相把超级智能武器化。间谍活动使持久垄断难以实现,但排除中国研究人员同样会适得其反:他说,在一些地方,顶尖 AI 公司超过30%的员工是中国籍,其中许多人可能回到中国。他支持放宽对顶尖 AI 研究人员的移民限制,同时认为这一问题应与南部边境政策分开处理。
如果美国建造一个“从太空完全可见”的1万亿美元算力集群,意图争夺主导地位,中国不会坐视不理。Hendrycks 将其类比为早期关于先发制人打击苏联的核战略设想:跨国人才、相互依赖和有限时间窗口,可能意味着所谓安全垄断的窗口从未真正存在。
4. MIM 将共同的网络脆弱性转化为威慑
Hendrycks 与 Eric Schmidt、Alexandr Wang 提出的“相互确保 AI 故障”(MIM),是在 AI 变得关键且能够自动化 AI 研发后,将核时代的共同脆弱性逻辑应用于 AI。一个对手若接近制造决定性超级武器,就可能面临间谍活动、破坏行动或针对其数据中心的网络攻击;这种预期旨在威慑破坏稳定性的项目,而不要求停止所有 AI 发展。
这个类比也有边界。Hendrycks 预计,各方会在共同利益重叠的地方协调——例如阻止危险能力落入恐怖分子手中,就像管控化学武器和生物武器;也会避免出现能够让一个国家“碾压”另一个国家的项目,但无人机等常规系统的竞争仍将继续。
校准力度很重要,因为如果 AI 芯片成为“经济力量的货币”,把对华出口管制的压力旋钮拧到最大,可能反而增强中国入侵台湾的动机。Hendrycks 指出,中国本来就想这么做;更严厉的管制只会再增加一个理由。
实际政策组合可以更窄:由 CIA 设立小组追踪外国 AI 项目,由网络司令部采购禁用能力,实施芯片位置许可制度,为盟友设置通知豁免,并优先开展最终用途核查。
面对 Guo 以 DeepSeek 提出的挑战,Hendrycks 同意,高效训练削弱了沙漠集群战略:管控无法稳健地消除大国的能力,也无法让其放弃 AI 的经济价值。中国可能最终仍会获得一部分芯片,或窃取模型权重;但管控对于提高可见性、阻止其扩散给流氓行为体仍有价值,尤其是在官员真正调查新加坡等转运路线的情况下。他说,AI 芯片出口管制并不是美国工业与安全局领导层的优先事项,基础核查本可以揭露流向中国的转运。
5. 评测会先展示类似神谕的智能,可靠代理能力则在之后
Humanity’s Last Exam 延续了 Hendrycks 早期的基准测试工作,包括 MMLU 和 MATH 数据集。Guo 也提到 Enigma,但 Hendrycks 在这里的回答聚焦于 HLE,以及更广义的封闭式问题与代理能力之分。教授和研究人员提交了异常困难、研究级且答案确定的封闭式问题,形成了他所说的考试式学术评测可能走到的“终点”。
接近满分的表现,大致意味着封闭式基准测试这一类型已被穷尽,并表明系统在封闭式问题上具备类似超人类数学家或 STEM 科学家的能力。但这并不能证明系统能胜任开放式工作,因为开放式工作不仅要求知道答案,还要求定义任务、推进任务并完成任务。
代理能力评测应改为分配真实的数字任务,允许系统工作数小时,再检查任务是否完成。当前系统作为代理“极其有缺陷”,仍接近地板,尽管 Hendrycks 留有余地,认为这一点“可能一夜之间改变”。
前沿能力仍然呈锯齿状:系统能回答高难度物理问题,却不会叠衣服;可能在数学上超过人类,却不会订机票。对于可验证的推理,人类或许仍能把 AI 生成的证明交给证明检查器;相比之下,如果系统形成更好的“品味”,就更难确认。
Hendrycks 认为,AI 会先获得“非常出色的、类似神谕的技能”,之后才可能可靠地替人类采取行动。一旦系统具备代理技能,产生重大经济影响的障碍就所剩无几,AI 也会“进入一个独立的类别”。
Hi listeners and welcome back to No Priors. Today I'm with Dan Hendrycks, AI researcher and director of the Center for AI Safety. He's published papers and widely used evals such as MLU and, most recently, Humanity's Last Exam. He's also published Superintelligence Strategy alongside authors including former Google CEO Eric Schmidt and Scale founder Alex Wang. We talk about AI safety and geopolitical implications, analogies to nuclear compute security, and the state of evals.
Dan, thanks for doing this.
Dan Hendrycks
Glad to be here.
How did you end up working on AI safety?
Dan Hendrycks
AI was pretty clearly going to be a big deal if I just thought through its conclusion. Early on, it seemed like other people were ignoring it because it was weirder or not that pleasant to think about. It's hard to wrap your head around, but it seemed like the most important thing during this century.
So I thought that would be a good place to develop my career toward, and that's why I started on it early on. Since it would be such a big deal, we need to make sure that we can think about it properly, channel it in a productive direction, and take care of some sort of tail risks, which are generally systematically under-addressed. That's why I got into it. It's a big deal, and people weren't really doing much about it at the time.
What do you think of as the Center for AI Safety's role versus safety efforts within the large labs?
Dan Hendrycks
There aren't that many safety efforts in the labs even now. I think the labs can focus on doing some very basic measures to refuse queries related to, "Help me make a virus," and things like that.
But I don't think labs have an extremely large role in making this go well overall. They're kind of predetermined to race. They can't really choose not to unless they would no longer be a relevant company in the arena. I think they can reduce terrorism risks or some accidents, but beyond that, I don't think they can dramatically change the outcomes in any substantial way.
A lot of this is geopolitically determined. If companies decide to act very differently, there's the prospect of competing with China, or maybe Russia will become relevant later. As that happens, this constrains their behavior substantially.
I've been interested in tackling AI at multiple levels. There are things companies can do to have some very basic antiterrorism safeguards, which are pretty easy to implement. There's also the economic effects that will need to be managed well, and companies can't really change how that goes either. It's going to cause mass disruptions to labor and automate a lot of digital labor.
If they tinker with the design choice or add some different refusal data, it doesn't change that fact. Safety—making AI go well—and risk management are much broader problems. They have some technical aspects, but I think that's a small part of it.
I don't know that the leaders of the labs would say, "We can do nothing about this," but maybe it's also a question of semantics. Everybody also has equity in this equation, right? Can you describe how you think of the difference between alignment and safety?
Dan Hendrycks
I'm just using safety as a catch-all for dealing with risks. There are other risks, like if you never get really intelligent AI systems, that poses some risks in itself. There are other sorts of risks that are not necessarily technical, like concentration of power.
I view the distinction between alignment and safety as alignment being a sort of subset of safety. Obviously, you want the value systems of the AIs to be in keeping with or compatible with, say, the US public for US AIs, or with you as an individual. But that doesn't necessarily make it safe.
If you have an AI that's reliably obedient or aligned to you, this doesn't make everything work totally well. China can have AIs that are totally aligned with them, and the US can have AIs that are totally aligned with them. You're still going to have a strategic competition between the two.
They're going to need to integrate AI into their militaries, and they're probably going to need to integrate it really quickly. Competition is going to force them to have a high risk tolerance in the process. So even if the AIs are reliably doing their principals' bidding, this doesn't necessarily make the overall situation perfectly fine.
I think it's not just a question of reliability or whether they do what you want. There are other structural pressures that cause this to be riskier, like geopolitics at the highest level.
With increasingly capable models and weights, why do we care about AI from a national security perspective? What's the most practical way it matters in geopolitics or gets used as a weapon?
Dan Hendrycks
I think AI isn't that powerful currently in many respects, so in many ways it's not actually that relevant for national security currently. This could well change within a year's time. Generally, I've been focused on the trajectory that it's on, as opposed to saying, "Right now, it is extremely concerning."
That said, for cyber, I don't think AIs are currently that relevant for being able to pull off a devastating cyberattack on the grid by a malicious actor. We should look at cyber, be prepared, and think about what its strategic implications are.
There are other capabilities, like biology. The AIs are getting very good at STEM PhD-level topics, and that includes biology. I think they're rounding the corner on being able to provide expert-level capabilities in terms of their knowledge of the literature or even helping in practical wet-lab situations.
So I do think that, on the biology aspect, they already have national security implications. But that's only very recent, with the reasoning models. In many other respects, they're not as relevant. It's more prospective that AI could become the way in which a nation might try to dominate another nation, and the backbone for not just war but also economic security.
The amount of ships that the US has versus China might be the determinant of which country is the most prosperous and which one falls behind. This is all prospective. I don't think it's just speculative. It's speculative in the same way that NVIDIA's valuation is speculative, or the valuations behind AI companies are speculative. It's something that I think a lot of people are expecting, and expecting fairly soon.
Yeah, it's quite hard to think about time horizons in AI. We invest in things that I think of as medium-term speculative, but they get pulled in quite quickly.
Just because you mentioned both cyber and bio, we're investors in companies like Cymulate or Cylance on the defensive cybersecurity side, or Chai and Somite on the biotech discovery side, modeling different systems in biology that will help us with treatments. How do you think about the balance of competition, benefits, and safety? Some of these things, I think, are working effectively in the near term on the positive side as well.
Dan Hendrycks
I don't get this big trade-off between safety and the benefits of AI. You're just taking care of a few tail risks. For bio, if you want to expose those capabilities, just talk to sales and get the enterprise account.
You can have the little refusal mechanism for biology, but if you just create an account and ask it how to culture this virus, show it a picture of your Petri dish, and ask what the next step should be, then, yeah, if you want access to those capabilities, you can speak to sales. That's basically the X in an X-risk management framework: we're just not exposing those expert-level capabilities to people whose identities we don't know. But if we do know who they are, then sure, give them access.
Likewise with cyber, I think you can very easily capture the benefits while taking care of some pretty avoidable tail risks. Once you have that, you've basically taken care of malicious use for the models behind your API, and that's about the best that you can do as a company.
You could try to influence policy by using your voice or something, but I don't see a substantial amount that companies could do. They could do some research to make the models more controllable, or try to make policymakers more aware of the situation more broadly in terms of where we're going, because I don't think policymakers have internalized what's happening at all.
They still think it's just hype, and they don't actually believe—or the companies and their employees don't actually believe—that we could get AGI in the next few years. So I don't see really substantial trade-offs there. I see much more substantial complications when we're dealing with the right level of stringency in export controls, for instance.
If you turn the pain dial all the way up for China in export controls, and if AI chips are the currency of economic power in the future, then this increases the probability that they want to invade Taiwan. They already want to; this gives them all the more reason.
If AI chips are the main thing, and they're not getting any of them—and they're not even getting the latest semiconductor manufacturing tools for making cutting-edge CPUs, let alone GPUs—those are some other types of complicated problems that we have to address and calibrate appropriately.
But in terms of just mitigating biology risks, speak to sales. If you're Genentech or a biotech startup, then you have access to those capabilities.
Perhaps what's a way you actually expect AI to get used as a weapon beyond virology and cybersecurity?
Dan Hendrycks
I wouldn't expect a bioweapon from a state actor. From a non-state actor, that would make a lot more sense.
Cyber makes sense from state actors and non-state actors. Then there are drone applications. These could disrupt other things, and they could help with other types of weapons research, such as exploring exotic EMPs. They could help create better types of drones and substantially help with situational awareness, so one might know where all the nuclear submarines are.
Some advancement in AI might be able to help with that, and that could disrupt our second-strike capabilities and mutually assured destruction. Those are some geopolitical implications. It could potentially bear on nuclear deterrence.
That's not even a weapon. The example of just heightened situational awareness and being able to pinpoint where hardened land-based nuclear launchers are, or where nuclear submarines are, is just informational, but could nonetheless be extremely disruptive and destabilizing.
Outside of that, the default conventional AI weapon would be drones. I don't know if that makes sense, or that countries would compete on that, and I think it would be a mistake if the US weren't trying more in manufacturing drones.
Yeah, I started working recently with an electronic warfare company. I think there's a massive lack of understanding of just the basic concept: we have autonomous systems, they all have communication systems, and our missile systems have targeting and communication systems.
From a battlefield-awareness and control perspective, a lot of that effort will be won with radio, radar, and related systems. I think there's an area where AI is going to be very relevant and is already very relevant in Ukraine.
Speaking about AI assisting with command and control, I remember hearing some story about how, on Wall Street, humans used to—you always had a human in the loop for each decision. At a later stage, before they removed that requirement on Wall Street, you just had rows of people clicking the "Accept, accept, accept" button. We're getting to a similar state in some contexts with AI.
Dan Hendrycks
It wouldn't surprise me if we ended up automating more of that decision-making. But this just turns into questions of reliability, and doing reliability research seems useful.
I think people are largely thinking that the push for risk management is to do some sort of pausing or something like that. An issue is that you need teeth behind an agreement. If you do it voluntarily, you just make yourself less powerful and let worse actors get ahead of you.
You could say, "We'll sign a treaty," but assuming that the treaty will be followed would be very imprudent. You would actually need some sort of threat of force or something to back it up, or some verification mechanism. But absent that, if it's entirely voluntary, this doesn't seem like a useful thing at all.
But absent that, as a proxy, there's clearly been very little compliance with either treaties or norms around cyberattacks and corporate espionage, right?
Dan Hendrycks
Yeah. Corporate espionage, for instance, was one strategy behind this voluntary-pause strategy, with people thinking that equals safety. Then maybe last year there was that paper, Situational Awareness: The Decade Ahead, written by Leopold Aschenbrenner. He's a sort of safety person.
His idea was, "Let's instead try and beat China to superintelligence as much as possible." But that has some weaknesses because it assumes that corporate espionage will not be a thing at all, which is very difficult to do.
In some places, more than 30% of the employees at these top AI companies are Chinese nationals. This is not feasible. If you're going to get rid of them, they're going to go to China, and they're probably going to beat you because they're extremely important for the US's success.
So you're going to want to keep them here, but that's going to expose you to some information-security issues. That's just too bad.
Do you have a point of view on how we should change immigration policy, if at all, given these risks?
Dan Hendrycks
I would, of course, claim that the policy on this would be totally separate from Southern border policy and broader policy. But if we're talking about AI researchers, if they're very talented, then I think you want to make it easier. I think it's probably too difficult for many of them to stay currently.
That discussion should be kept totally separate from Southern border policy.
Just in terms of broad strokes, what are things that you think won't work? Voluntary compliance and assuming that'll happen, or just straight race?
Dan Hendrycks
We want to be competitive, and I think racing in other spheres, say drones or AI chips, seems fine. If you're saying, "Let's race to superintelligence to try and turn that into a weapon," and they're not going to do the same, or they're not going to have access to it, or they're not going to prevent that from happening, that seems like quite a tall claim.
If we did have a substantially better AI, they could just co-opt it; they could just steal it, unless you had really strong information security. You could move the AI researchers out to the desert, but then you're reducing your probability of actually beating them because a lot of your best scientists would end up going back to China.
Even then, if there were signs that they were really pulling ahead and going to be able to get some powerful AI that would enable China—or that would enable the US—to crush China, they would then try to deter them from doing something like that.
They're not going to sit idly by and say, "You know what? Go ahead, develop your superintelligence or whatever, and then you can boss us around, and we'll just accept your dictates till the end of time."
That, I think, is a failure of some sort of second-order reasoning: how would China respond to this sort of maneuver if we're building a $1 trillion compute cluster in the desert, totally visible from space? The only plausible read on this is that it's a bid for dominance or a sort of monopoly on superintelligence. It reminds me of the nuclear era. There was a brief period when some people were saying, "We have to just preemptively or preventively destroy the USSR." Even people who were normally pacifists, like Bertrand Russell, were advocating for this. The opportunity window for that maybe never existed, but there was a prospect of it for some time. I don't think that opportunity window really exists here because of the complex interdependence and multinational talent dependence in the United States. I don't think you can have China be totally excluded from any awareness or ability to gain insight into or imitate what we're doing here. We're clearly nowhere close to that as a real environment right now, right?
Right. It would take years to do well, and I don't even think the timelines for some very powerful AI systems mean there might be enough time to do that securitization anyway.
You propose, along with Eric Schmidt and Alexandr Wang, a new deterrent regime: mutually assured AI malfunction, or MIM. It's a bit of a scary acronym and also a nod to mutually assured destruction. Can you explain MIM in plain language?
Dan Hendrycks
Let's think of what happened in nuclear strategy. Basically, a lot of states deterred each other from doing a first strike because they could then retaliate. They had a shared vulnerability. They were saying, "We're not going to take this really aggressive action of trying to make a bid to wipe you out, because that will end up causing us to be damaged."
Later on, when AI is more salient, when it's viewed as pivotal to the future of a nation, and when people are on the verge of making a superintelligence—when they can say, "Automate pretty much all AI research"—I think states would try to deter each other from trying to leverage that to develop something like a superweapon.
That could allow one country to crush the others, or allow those AIs to conduct a really rapid, automated AI research-and-development loop that could bootstrap them from their current levels to something superintelligent, vastly more capable than any other system out there.
I think later on it becomes so destabilizing that China just says, "We're going to do something preemptive, like a cyberattack on your data center," and the US might do that to China.
Russia, coming out of Ukraine, will reassess the situation and become situationally aware. It will think, "What's going on with the US and China? My goodness, they're so focused on AI."
Let's say it's later in the year, when a big chunk of software engineering is starting to be impacted by AI. Russia might say, "Oh, wow, this is looking pretty relevant. If you try to use this to crush us, we will prevent that by doing a cyberattack on you, and we will keep tabs on your projects."
It's pretty easy for them to do espionage. All they need to do is a zero-day attack on Slack, and then they can know what DeepMind is up to in very high fidelity, as well as OpenAI, xAI, and others. It's pretty easy for them to do espionage and sabotage.
Right now, they don't need to threaten that because it's not at the level of severity. It's not actually that potentially destabilizing; it's still too distant. A lot of decision-makers still aren't taking this AI stuff that seriously, relatively speaking, but I think that'll change as it gets more powerful.
Then I think this is how they would end up responding. This keeps us from winding up in a situation where we're doing something extremely destabilizing, like trying to create a weapon that enables one country to totally wipe out the other, as was proposed by people like Leopold Aschenbrenner.
What are the parallels here that you think make sense to nuclear weapons, and which ones don't?
Dan Hendrycks
More broadly, AI is a dual-use technology, in that it has civilian applications, military applications, and economic applications. Its economic applications are still limited in some ways, and likewise its military applications are still limited, but I think that will keep changing rapidly.
Chemical technology was important for the economy and had some military use, but countries coordinated not to go down the chemical-weapons route. Biology can be used as a weapon and has enormous economic applications, and likewise with nuclear technology.
For each of those technologies, countries did eventually coordinate to make sure they didn't wind up in the hands of rogue actors like terrorists. There have been a lot of efforts to make sure rogue actors don't get access to them and use them against their adversaries, because it's in neither side's interest.
Bioweapons and chemical weapons are a poor man's atom bomb. That's why we have the Chemical Weapons Convention and the Biological Weapons Convention. There's some shared interest there. They might be rivals in other senses, in the way that the US and the Soviet Union were rivals, but they're still able to coordinate on that because it's incentive-compatible.
It doesn't benefit them in any way if terrorists have access to these sorts of things. It's just inherently destabilizing. So I think that's an opportunity for coordination.
That isn't to say that they have an incentive to pause all forms of AI development. It may mean that they would be deterred from some particular forms of AI development, particularly those that have a very plausible prospect of enabling one country to get a decisive edge over another and crush it.
So, no superweapon-type stuff, but more conventional types of warfare, like drones, will continue. I expect that they'll continue to race and probably not even coordinate on anything like that. That's just how things will go. It's like bows and arrows and nuclear weapons: it made sense for them to develop those sorts of weapons and threaten each other with them.
If you could magically adopt some policy or action in the current administration, what is the first step here?
Dan Hendrycks
The first step is, "We will not build a superweapon, and we're going to be watching for other people building them too."
As I've been alluding to throughout this conversation, what would the companies do? Not that much. They would add some basic antiterrorism safeguards. This is technically pretty easy, unlike refusal for other things. If you're trying to deal with crimes and torts, that's harder because it's much messier and overlaps with typical everyday interaction.
I think the asks for states are not that challenging either. It's just a matter of doing them. One step would be for the CIA to have a cell doing more espionage of other states' AI programs, so that we have a better sense of what's going on and aren't caught by surprise.
Secondly, maybe some part of the government, such as Cyber Command, which has a lot of cyberoffensive capabilities, gets some cyberattacks ready to disable data centers in other countries if they're looking like they're running or creating a destabilizing project.
That's it for the deterrence. For nonproliferation of AI chips to rogue actors in particular, I think there would be some adjustments to export controls. In particular, we need to know where the AI chips are reliably, for the same reason we want to know where our fissile material is, and for the same reason that we want Russia to know where its fissile material is. That's just generally good information to collect.
That can be done with some very basic statecraft: having a licensing regime, and having allies notify you whenever chips are being shipped to a different location. They would get a license exemption on that basis, and then you would have enforcement officers prioritize some basic inspections for AI chips and end-use checks.
All of these are a few texts away or a basic document away. I think that 80/20 is a lot of it. Of course, this is always a changing situation.
Safety isn't, as I've been trying to reinforce, really that much of a technical problem. This is more of a complex geopolitical problem with technical aspects. Later on, maybe we'll need to do more. There might be some new risk sources that we need to take care of and adjust for.
But right now, I think that spies for the CIA, sabotage with Cyber Command, building up those capabilities, and buying those options would take care of a lot of the risk.
Let's talk about compute security. If we're talking about 100,000 networked, state-of-the-art chips, you can tell where that is. How do DeepSeek and the recent releases they've had factor into your view of compute security, given that export controls have clearly led to innovation toward highly compute-efficient pretraining that works on chips China can import at what might be considered an irrelevant scale—a much smaller scale today?
It seems directionally hard to see training becoming less efficient, even if we want to scale it up. Does that change your view at all?
Dan Hendrycks
No. It just sort of undermines other types of strategies, like the Manhattan Project strategy of moving people out to the desert and building a big cluster there.
What it shows is that you can't rely as much on restricting another superpower's ability to make models. You can restrict their intent, which is what deterrence does, but I don't think you can reliably or robustly restrict their capabilities.
You can restrict the capabilities of rogue actors, and that's what I would want compute security and export controls to facilitate. We should make sure it doesn't wind up in the hands of Iran or another rogue actor.
China will probably keep getting some fraction of these chips, but we should basically try to know where they are more reliably, and we can tighten things up. You could even coordinate with China to make sure the chips aren't winding up in rogue actors' hands.
I should also say that export controls weren't actually a priority among the leadership at the Bureau of Industry and Security, to my understanding. AI chips were a substantial priority for some people, but for the enforcement officers, did any of them go to Singapore to see where 10% of NVIDIA's chips were going? I think they would have very quickly found that they were going to China.
Some basic end-use checks would have taken care of that. It's not that export controls don't work. We've done nonproliferation of lots of other things, like chemical agents and fissile material, so it can be done if people care.
Even so, I still think that if you really tighten the export controls so that China can't get any of those chips at all, and this is one of your biggest priorities, they're just going to steal the weights anyway. I think it will be too difficult to totally restrict their capabilities, but I think you can restrict them through deterrence.
It also seems like either this stuff is powerful or it's not. It seems infeasible to me, given the economic opportunity, that China will say, "We don't need the capability."
Dan Hendrycks
Yeah, I fail to see a version of the world where the leadership of another great power that believes there is value here says, "We don't need that," from an economic-value perspective.
For a lot of these things, it would perhaps be nicer if everything went 3× slower and there were fewer mess-ups. If there were some magic button that would do that, maybe it would be useful. I don't know whether that's true, actually. I don't have a position on that.
Given the structural constraints and the competitive pressures between these companies and these states, a lot of these things become infeasible. Many of the gestures that could be useful for risk mitigation become less tractable when you consider the structural realities.
That said, there still would be some pausing or halting of the development of particular projects that you could potentially lose control of, or that, if controlled, would be very destabilizing because they would enable one country to crush the other.
I think people's conception of what risk management looks like is that it's a peacetime thing, or something like that—it's all kumbaya, and we just have to ignore the structural realities of operating in this space.
Instead, the right approach is more like nuclear strategy. It's an evolving situation. There are some basic things you can do: you're probably going to need a stockpile of nuclear weapons, you're going to need to secure a second strike, you're going to need to keep an eye on what the other side is doing, and you're going to need to make sure there isn't proliferation to rogue actors when the capabilities are extremely hazardous.
This is a continual battle, but it's not clearly going to be an extremely positive thing no matter what, and it's not going to be doomsday no matter what. Nuclear strategy was obviously risky business. The Cuban Missile Crisis came pretty close to an all-out nuclear war.
It depends on what we do. Some basic interventions and very basic statecraft can take care of a lot of these risks and make the situation manageable. I imagine we're left with more domestic problems, like what to do about automation and things like that, but I think maybe we'll be able to get a handle on some of the geopolitics here.
I want to change tack for our last couple of minutes and talk about evals. It's obviously very related to safety and understanding where we are in terms of capability.
You came out with the strikingly named Humanity's Last Exam eval, and then also Enigma. Why are these relevant, and where are we with evals?
Dan Hendrycks
I've been making evaluations to try to understand where we are in AI for as long as I've been doing AI research. Previously, I've done datasets like MMLU and the MATH dataset. Before ChatGPT, there were things like ImageNet and other sorts of benchmarks.
Humanity's Last Exam was basically an attempt to get at the end of the road for evaluations and benchmarks based on exam-like questions—ones that test some sort of academic knowledge. We asked professors and researchers around the world to submit a really challenging question, and then we added those questions to the dataset.
It's a big collection of what professors, for instance, would encounter as challenging problems in their research, with a definitive, closed-ended, objective answer. I think the genre of a closed-ended question, where the answer is multiple-choice or a simple short answer, will roughly be exhausted when performance on this dataset is near the ceiling.
When performance is near the ceiling, I think that would basically be an indication that you have something like a superhuman mathematician or a superhuman STEM scientist, at least in areas where closed-ended questions are useful, such as mathematics.
But it doesn't get at other things, such as the ability to perform open-ended tasks. That's more of an agent-type evaluation, and I think that will take more time. We can try to measure directly what its ability is to automate various digital tasks: collect various digital tasks, have it work on them for a few hours, and see whether it successfully completes them.
Coming out soon, we have a test for closed-ended questions that test knowledge in academia, including areas like mathematics, but they're still very bad at agent tasks. This could possibly change overnight, but they're still near the floor. I think they're still extremely defective as agents.
There will need to be more evaluations for that, but the overall approach is just to try to understand what's going on: what's the rate of development? That way, the public can at least understand what's happening.
If all evaluations are saturated, it's difficult to even have a conversation about the state of AI. Nobody really knows exactly where it is, where it's going, or what the rate of improvement is.
Is there anything that qualitatively changes when these models and model systems are just better than humans—when they're exceeding human capability? Does that change our ability to evaluate them?
Dan Hendrycks
I think the intelligence frontier is just so jagged. What they can and can't do is surprising. They still can't fold clothes, but they can answer a lot of tough physics problems. There are complicated reasons for why that is, so it's not uniform.
In some ways, they'll be better than humans. It seems totally plausible that they'll be better than humans at mathematics before too long, but still not able to book a flight. The implication is that they might be better in some limited ways, with limited influence, but that won't necessarily generalize to other things.
I do think it's possible that they'll be better at reasoning skills than us. We could still have humans checking their work because we can verify it. If an AI mathematician is better than a human, humans can still run the proof through a proof checker and confirm that it was correct.
In that way, humans can still understand what's going on in some ways. But in other ways, if they're getting better taste in things—if that makes any sense, although maybe it doesn't make philosophical sense—that would be pretty difficult for people to confirm.
I think we're on track overall to have AIs with really good oracle-like skills. You can ask them things, and they say something insightful, very nontrivial, or push the bounds of knowledge in some particular way. But they may not necessarily be able to carry out tasks on behalf of people for some while.
I think this is why we don't take AIs that seriously: they still can't do a lot of very trivial stuff. But when they get some of the agent skills, I don't think there are many barriers to their economic impact, or to people thinking that this is more than just an interesting thing.
That's an emergent property with agent skills. The vibes really shift, and it's pretty clear that this is much bigger than some prior technology, like the App Store or social media. It's in a category of its own.
Dan, thanks for doing this. It's a great conversation.
Dan Hendrycks
Glad to be here. Thank you for having me.
Find us on Twitter at No Priors Pod. Subscribe to our YouTube channel if you want to see our faces. Follow the show on Apple Podcasts, Spotify, or wherever you listen. That way you get a new episode every week. Sign up for emails or find transcripts for every episode at no-pri.com.