超级智能前夜的国家安全战略与 AI 评测:Dan Hendrycks
- Hendrycks 的核心判断是,AI 安全首先是国家战略问题,而不只是实验室层面的对齐问题。 实验室“注定要竞速”;它们可以低成本限制专家级生物学能力的访问、应对网络风险,但即使美国和中国的系统都服从指令,战略竞争、军事快速融合、劳动力自动化以及不断上升的风险承受度仍然存在。
- 眼下的国家安全图景参差不齐:推理模型正逼近专家级病毒学能力,但 Hendrycks 不认为 AI 目前已足以帮助恶意行为者发动毁灭性电网攻击。 这“很可能在1年内改变”。讨论还涉及网络防御、无人机、电子战、指挥控制可靠性和算力安全等近期应用与风险领域。
- Hendrycks 否定自愿暂停,也否定美国单方面冲刺超级智能,因为两者都经不起对手反制。 条约需要核查机制或武力保障,模型权重可能被窃取,而一座“从太空完全可见的1万亿美元沙漠算力集群”会被视为争夺支配地位的行动,中国将试图加以遏制。
- 他提出的“相互确保 AI 失灵”(MAIM)机制建立在共同脆弱性之上:各国不会推进破坏稳定的超级武器项目,因为对手可以监视或瘫痪其数据中心。 芯片和无人机竞争仍将继续;协调重点在于阻止流氓行为者获得能力,类似核、生化武器机制将大国竞争与共同防扩散利益分开。
- DeepSeek 式效率削弱了对华能力封锁,因此 Hendrycks 会主要用出口管制追踪芯片、约束流氓行为者,再用威慑约束大国意图。 他的80/20方案很具体:设立 CIA AI 间谍小组,由网络司令部准备破坏选项,实施许可和发货通知制度,并核查最终用途——包括调查此前提到的10% NVIDIA 芯片究竟流向何处。
- 当模型从令人印象深刻的“预言机式能力”迈向可靠的代理能力时,经济将迎来相变。 Humanity’s Last Exam 的目标是在封闭式学术基准接近上限时让其退出历史舞台——这将意味着模型在部分数学和 STEM 领域达到超人水平——但代理仍然“接近地板”;一旦它们能够完成数字化工作,Hendrycks 预计“整个氛围会彻底改变”(the vibes really shift)。
1. AI 安全远不止让模型服从指令
Hendrycks 说,他很早就选择 AI 安全,是因为沿着 AI 的发展轨迹观察,它看起来会是“本世纪最重要的事情”,而它的诡异之处和令人不适的含义又让其他人望而却步。这项使命从来不只是防止系统失效:还包括保持清醒判断、把技术导向生产性用途,并覆盖那些“系统性投入不足”的尾部风险。
他对机构分工的判断很直接:“即使现在,实验室里的安全工作也没有那么多”,而各家公司“某种程度上注定要竞速”,否则就会失去相关性。它们可以拒绝“帮我制造病毒”的请求,减少恐怖主义或某些事故,并研究可控性;但地缘政治竞争对最终结果的决定性更大。
在他的分类中,对齐只是安全的一部分。AI 可以可靠地服从美国或中国,却仍然处于一场迫使各方快速推进军事融合、提高风险承受度的竞争之中;“它们是否按你的要求行事”并不能消除权力集中、战略压力,也不能消除永远无法获得高度智能系统的风险。
Guo 反驳说,实验室领导者不会声称自己什么都做不了,而且每个人“在这个方程式里都有自己的利益”。Hendrycks 的回应是结构性的:设计调整和拒答数据无法阻止劳动力遭遇大规模冲击,也无法阻止数字化工作的自动化;与此同时,政策制定者仍认为 AI “只是在贩卖炒作”,不相信企业员工真的预计 AGI 会在未来几年出现。
2. 国家安全能力正在不均衡地到来
Hendrycks 将当前能力与发展轨迹分开看:他不认为 AI 目前已与恶意行为者发动毁灭性电网攻击有关,但这“很可能在1年内改变”。病毒学进展更快;推理模型正逼近专家级文献知识,甚至可能协助实际湿实验,因此已经具备国家安全相关性。
Guo 提到防御性网络安全领域的 Culminate 和 Sibyl,以及生物科技领域的 Chai 和 Somite,说明收益也在同步出现。Hendrycks 看不到太大取舍:把专家级病毒学能力放在企业级合作关系之后——“直接找销售”——而不是向一个新账户开放,让对方询问如何培养病毒。
除了生物安全,Hendrycks 关注的能力栈还包括国家和非国家网络攻击、无人机开发、非传统电磁脉冲研究,以及更强的态势感知。仅仅定位加固后的发射场或核潜艇,就可能威胁二次打击能力和相互确保摧毁;这是一种信息能力提升,本身不是武器,却“极具破坏性”,并会造成不稳定。
3. 自愿暂停与单方面冲刺超级智能都行不通
Hendrycks 反对自愿放慢发展的理由是执行问题:没有核查机制或“某种武力威胁”,克制只会让守规矩的一方变弱,而更糟糕的行为者继续前进。Guo 将其类比为网络攻击和企业间谍活动:条约和规范并没有带来多少可靠的合规。
封闭式的美国冲刺则有相反的缺陷。Hendrycks 指出,一些顶级 AI 公司有“30%以上”的员工是中国籍;排除他们可能把不可替代的科学家送到中国,降低美国的胜算,而留下他们又必然带来信息安全暴露。他在移民问题上的偏好仍然是,让极具天赋的 AI 研究者更容易留在美国。
更深层的失败在于没有进行二阶推理:中国不会坐视美国建造一座公开可见、旨在制造能够“对我们发号施令……直到时间尽头”的系统的1万亿美元算力集群。Hendrycks 将这一想法类比为核时代短暂出现过的先发摧毁苏联论,但同时指出,AI 的跨国人才和知识流动意味着这样的机会窗口实际上并不存在。
“相互确保 AI 失灵”(MAIM)将这种脆弱性转化为威慑。如果快速自动化研发看起来是在争取决定性支配地位,对手就可能威胁攻击数据中心,或通过“在 Slack 上利用一个零日漏洞”监控实验室;共同暴露在攻击面前,可以遏制超级武器项目,同时不终结无人机、芯片或常规系统领域的竞争。
4. 威慑约束意图,算力管制遏制扩散
Hendrycks 眼下主张的国家战略并不华丽:设立一个专门关注外国 AI 项目的 CIA 小组,然后让网络司令部准备瘫痪海外数据中心的选项,以应对破坏稳定的项目。间谍活动提供预警;建设这些能力并“购买这些选项”,才能让威慑真正具备执行力。
在芯片方面,他提出仿照裂变材料追踪的许可制度:如果盟友在硬件改变位置时通知美国,就可以获得豁免;执法人员则开展有优先级的最终用途核查。他质疑的是执行,而不是可行性——他问,有没有人去新加坡查过此前提到的10% NVIDIA 芯片究竟流向哪里;一次基本的最终用途核查就应该能查明目的地。
DeepSeek 的高效训练并没有推翻这一判断,反而削弱了“曼哈顿计划”式幻想,即某个超级大国可以稳健地阻止另一个国家构建模型。中国很可能仍会获得一部分芯片;即使把管制彻底收紧,中国也可能转而窃取模型权重。管制的主要目标应是阻止能力扩散到伊朗等流氓行为者,必要时争取中国合作;而威慑则用来约束破坏稳定的意图。
Guo 的挑战在于经济层面:如果 AI 足够强大,没有哪个大国会放弃这种能力。Hendrycks 表示同意,但他还没有确定答案,不知道一种神奇的3倍减速是否有帮助;如果把出口管制的“疼痛旋钮”完全拧向中国,反而可能提高其入侵台湾的动机,因为一旦 AI 芯片成为未来经济力量的货币,芯片就会变得更加关键。
5. 评测揭示出一个惊人却仍无法行动的前沿
Hendrycks 将 Humanity’s Last Exam 放在从 MMLU、MATH 数据集到 ChatGPT 之前 ImageNet-C 等工作的演进路径上。全球各地的教授和研究者贡献了极其困难、答案客观且封闭的问题——也就是他们在研究中会遇到的那类问题——试图为考试式学术基准走到终点。
如果模型表现接近上限,就意味着封闭式答案题这一类型“大致已经走到尽头”,并可能表明模型在相关领域已经具备超人级数学家或 STEM 科学家的能力。但这并不能证明模型具备开放式能力,因此不能把基准饱和与系统能够自主完成真实工作混为一谈。
下一代评测因此会向模型提供数字化任务,让它们工作数小时,再检查是否完成。当前系统作为代理仍然“极其有缺陷”,接近地板水平,不过 Hendrycks 也承认,这“可能一夜之间改变”;如果没有尚未饱和的测试,公众既看不到能力处于什么水平,也看不到能力提升的速度。
这个前沿极不平整:系统可以解决困难的物理问题,却不会叠衣服;可能在数学上超越人类,却仍然无法预订机票。证明检查器可以验证超人级数学,而更好的“品味”等特质可能难以核验;就目前而言,Hendrycks 预计模型会先展现强大的“预言机式能力”,然后才获得可靠的代理能力。
Hi, listeners, and welcome back to No Priors. Today, I'm with Dan Hendrycks, AI researcher and director of the Center for AI Safety. He's published papers and widely used evals such as MMLU and, most recently, Humanity's Last Exam. He's also published Superintelligence Strategy alongside authors including former Google CEO Eric Schmidt and Scale founder Alex Wang. We talk about AI safety and geopolitical implications, analogies to nuclear, compute security, and the state of evals. Dan, thanks for doing this.
Glad to be here.
How’d you end up working on AI safety?
AI was pretty clearly going to be a big deal if one would just think through its conclusion. Early on, it seemed like other people were ignoring it because it was weirder or not that pleasant to think about. It’s hard to wrap your head around, but it seemed like the most important thing during this century. I thought that would be a good place to develop my career toward, and that’s why I started on it early on.
Since it’d be such a big deal, we’d need to make sure that we can think about it properly, channel it in a productive direction, and take care of some sort of tail risks, which are generally systematically under-addressed. That’s why I got into it: it’s a big deal, and people weren’t really doing much about it at the time.
What do you think of as the center’s role versus safety efforts within the large labs?
Well, there aren’t that many safety efforts in the labs even now. I think the labs can just focus on doing some very basic measures to refuse queries like, “Help me make a virus,” and things like that. But I don’t think labs have an extremely large role in safety overall or in making this go well.
They’re kind of predetermined to race. They can’t really choose not to unless they would no longer be a relevant company in the arena. I think they can reduce terrorism risks or some accidents, but beyond that, I don’t think they can dramatically change the outcomes in too substantial of a way.
Because a lot of this is geopolitically determined, if companies decide to act very differently, there’s the prospect of competing with China, or maybe Russia will become relevant later. As that happens, this constrains their behavior substantially.
I’ve been interested in tackling AI at multiple levels. There are things companies can do to have some very basic anti-terrorism safeguards, which are pretty easy to implement. There are also the economic effects that will need to be managed well, and companies can’t really change how that goes either.
It’s going to cause mass disruptions to labor and automate a lot of digital labor. If they tinker with the design choice or add some different refusal data, it doesn’t change that fact. Making AI go well and managing the risks is much more of a broader problem. It’s got some technical aspects, but I think that’s a small part of it.
I don’t know that the leaders of the labs would say, “We can do nothing about this,” but maybe it’s also a question of everybody having equity in this equation, right? Maybe it’s also a question of semantics. Can you describe how you think about the difference between alignment and safety?
I’m just using safety as a sort of catchall for dealing with risks. There are other risks, like if you never get really intelligent AI systems, that poses some risks in itself. There are other sorts of risks that aren’t necessarily technical, like concentration of power.
So I view the distinction between alignment and safety as alignment being a sort of subset of safety. Obviously, you want the value systems of the AIs to be in keeping with or compatible with, say, the US public for US AIs, or with you as an individual, but that doesn’t necessarily make it safe.
If you have an AI that’s reliably obedient or aligned to you, this doesn’t make everything work totally well. China can have AIs that are totally aligned with them. The US can have AIs that are totally aligned with them. You still are going to have a strategic competition between the 2.
They’re going to need to integrate it into their militaries. They’re probably going to need to integrate it really quickly. This competition is going to force them to have a higher risk tolerance in the process. So even if the AIs are reliably doing their principals’ bidding, this doesn’t necessarily make the overall situation perfectly fine.
I think it’s not just a question of reliability or whether they do what you want. There are other structural pressures that cause this to be riskier, like geopolitics.
At the highest level, with a bundle of weights that’s increasingly capable, why do we care about AI from a national security perspective? What’s the most practical way it matters in geopolitics or gets used as a weapon?
I think that AI isn’t that powerful currently in many respects. So, in many ways, it’s not actually that relevant for national security currently. This could well change within a year’s time. Generally, I’ve been focused on the trajectory that it’s on, as opposed to saying that right now it is extremely concerning.
That said, for instance, in cyber, I don’t think AIs are that relevant for being able to pull off a devastating cyberattack on the grid by a malicious actor currently. That said, we should look at cyber, be prepared, and think about its strategic implications.
There are other capabilities, like virology. The AIs are getting very good at STEM PhD-level topics, and that includes virology. So I think they’re sort of rounding the corner on being able to provide expert-level capabilities in terms of their knowledge of the literature or even helping in practical wet-lab situations.
I do think that, on the virology aspect, they already have national security implications, but that’s only very recently with the reasoning models. In many other respects, they’re not as relevant.
It’s more prospective that AI could well become the way in which a nation might try to dominate another nation, and the backbone for not just war but also economic security. The number of ships that the US has versus China might be the determinant of which country is the most prosperous and which one falls behind.
This is all prospective. I don’t think it’s just speculative. It’s speculative in the same way that NVIDIA’s valuation is speculative, or the valuations behind AI companies are speculative. It’s something that I think a lot of people are expecting, and expecting fairly soon.
Yeah, it’s quite hard to think about time horizons in AI. We invest in things that I think of as medium-term speculative, but they get pulled in quite quickly.
Just because you mentioned both cyber and bio, we’re investors in companies like Culminate or Sibyl on the defensive cybersecurity side, or Chai and Somite on the biotech discovery side, modeling different systems in biology that will help us with treatments. How do you think about the balance of competition, benefits, and safety? Some of these things, I think, are working effectively in the near term on the positive side as well.
Yeah, I don’t get this big trade-off. For bio, if you want to expose those capabilities, just talk to sales and get the enterprise account. Here, you can have the little refusal thing for virology.
But if you just created an account a second ago and you’re asking it how to culture this virus, saying, “Here’s your picture of your petri dish. What’s the next step that you should do?”—if you want access to those capabilities, you can speak to sales. That’s basically in xAI’s risk management framework: we’re not exposing those expert-level capabilities to people who we don’t know.
But if we do, then sure, have them. Likewise with cyber, I think you can very easily capture the benefits while taking care of some of these pretty avoidable tail risks. Once you have that, you’ve basically taken care of malicious use for the models behind your API, and that’s about the best that you can do as a company.
You could try to influence policy by using your voice or something, but I don’t see a substantial amount that they can do. They could do some research to try to make the models more controllable, or try to make policymakers more aware of the situation more broadly in terms of where we’re going.
I don’t think policymakers have internalized what’s happening in AI at all. They still think it’s just selling hype, and they don’t actually believe that the companies’ employees actually believe that we could get AGI in, so to speak, the next few years.
I don’t see really substantial trade-offs there. I think the complications really come about when we’re dealing with what the right stringency in export controls is, for instance. That’s complicated.
If you turn the pain dial all the way up for China in export controls, and if AI chips are the currency of economic power in the future, then this increases the probability that they want to invade Taiwan. They already want to. This would give them all the more reason if AI chips are the main thing and they’re not getting any of it, and they’re not even getting the latest semiconductor manufacturing tools for making cutting-edge CPUs, let alone GPUs.
Those are some other types of complicated problems that we have to address, think about, and calibrate appropriately. But in terms of just mitigating virology risks, if you’re at Genentech or a biotech startup, just speak to sales, and then you have access to those capabilities. Problem solved.
What is a way you actually expect AI to get used as a weapon beyond virology and security?
I wouldn’t expect a bioweapon from a state actor. From a non-state actor, that would make a lot more sense. I think cyber makes sense from state actors and non-state actors. Then there are drone applications. These could disrupt other things. They could help with other types of weapons research, like exploring exotic EMPs, and could help create better types of drones. They could substantially help with situational awareness, so that one might know where all the nuclear submarines are.
Some advances in AI might be able to help with that, and that could disrupt our second-strike capabilities and mutually assured destruction. Those are some geopolitical implications. It could potentially bear on nuclear deterrence, and that’s not even a weapon. The example of heightened situational awareness and being able to pinpoint where hardened land-based nuclear launch sites are, or where nuclear submarines are, is just informational but could nonetheless be extremely disruptive and destabilizing.
Outside of that, the default conventional AI weapon would be drones. I don’t know if it makes sense that countries would compete on that, and I think that would be a mistake if the US weren’t trying to do more in manufacturing drones.
I started working recently with an electronic warfare company. I think there’s a massive lack of understanding of just the basic concept: We have autonomous systems, and they all have communication systems. Our missile systems have targeting and communication systems. From a battlefield-awareness and control perspective, a lot of that fight will be won with radio, radar, and related systems, right?
Mm-hmm.
And so I think there’s an area where AI is going to be very relevant and is already very relevant in Ukraine.
Speaking about AI assisting with command and control, I was hearing a story about how, on Wall Street, you always had a human in the loop for each decision. At a later stage, before they removed that requirement on Wall Street, you just had rows of people clicking the “Accept, accept, accept” button. We’re getting to a similar state in some contexts with AI.
It wouldn’t surprise me if we ended up automating more of that decision-making. This just turns into questions of reliability, and doing some reliability research seems useful. To return to that larger question of where the safety trade-offs are, I think people are largely thinking that the push for risk management is to do some sort of pausing or something like that.
An issue is that you need teeth behind an agreement. If you do it voluntarily, you just make yourself less powerful, and you let the worst actors get ahead of you. You could say, “Well, we’ll sign a treaty.” We will not assume that the treaty will be followed. That would be very imprudent. You would actually need some sort of threat of force or something to back it up, or some verification mechanism.
But absent that, if it’s entirely voluntary, then this doesn’t seem like a useful thing at all. I think people’s conflation of safety is: What we must do is voluntarily slow it down. It just doesn’t make as much geopolitical sense unless you have some threat of force to back it up or some very strong verification mechanism. But in the absence of that—
As a proxy, there’s clearly been very little compliance with either treaties or norms around cyberattacks and corporate espionage, right?
Yeah. I mean, corporate espionage, for instance—that was one strategy, the sort of voluntary-pause strategy. People are thinking that equals safety. Then maybe last year there was that paper, “Situational Awareness,” written by Leopold Aschenbrenner, and he’s a sort of safety person. His idea was, “Let’s instead try to beat China to superintelligence as much as possible.”
But that has some weaknesses because it assumes that corporate espionage will not be a thing at all, which is very difficult to do. I mean, at some of these top AI companies, 30% or more of the employees are Chinese nationals. This is not feasible. If you get rid of them, they’re going to go to China, and then they’re probably going to beat you because they’re extremely important for US success.
So you’re going to want to keep them here, but that’s going to expose you to some information-security issues. But that’s just too bad.
Do you have a point of view on how we should change immigration policy, if at all, given these risks?
I would, of course, claim that the policy on this should be totally separate from southern border policy and broader policy. But if we’re talking about AI researchers, if they’re very talented, then I think you’d want to make it easier, and I think it’s probably too difficult for many of them to stay currently. I think that discussion should be kept totally separate from southern border policy.
Just in terms of broad strokes, things that you think won’t work: voluntary compliance and assuming that will happen, or just a straight race?
We want to be competitive, and I think racing in other sorts of spheres, say drones or AI chips, seems fine. If you’re saying, “Let’s race to superintelligence to try and get ahead and turn that into a weapon to crush them,” and they’re not going to do the same, or they’re not going to have access to it, or they’re not going to prevent that from happening, that seems like quite a tall claim.
I mean, if we did have a substantially better AI, they could just co-opt it. They could just steal it, unless you had really, really strong information security—for example, if you moved the AI researchers out to the desert. But then you’re reducing your probability of actually beating them because a lot of your best scientists ended up going back to China.
Even then, if there were signs that they were really pulling ahead and going to be able to get some powerful AI that would enable the US to crush China, they would then try to deter them from doing something like that. They’re not going to sit idly by and say, “You know what? Yeah, go ahead. Develop your superintelligence or whatever, and then you can boss us around, and we’ll just accept your dictates till the end of time.”
I think there is a failure of some sort of second-order reasoning going on there, which is: How would China respond to this sort of maneuver if we’re building a trillion-dollar compute cluster in the desert, totally visible from space? Basically, the only plausible read on this is that this is a bid for dominance or a sort of monopoly on superintelligence.
It reminds me of the nuclear era. There was a brief period where some people were saying, “You know what? We’ve got to just preemptively destroy or preventively destroy the USSR. We’ve got to nuke ’em.” Even pacifists, or people who are normally pacifists, like Bertrand Russell, were advocating for this. The window of opportunity for that maybe never existed, but there was a prospect of it for some time.
I don’t think the opportunity window really exists here because of the complex interdependence and multinational talent dependence in the United States. I don’t think you can have China be totally severed from any awareness or any ability to gain insight or imitate what we’re doing here.
We’re clearly nowhere close to that in a real environment right now, right?
No, it would take years. It would take years to do well, and I don’t even think the timelines for some very powerful AI systems leave enough time to do that securitization anyway.
Yeah.
So, okay, in reaction, you propose, along with some other esteemed authors and friends, Eric Schmidt and Alex Wang, a new deterrence regime: Mutually Assured AI Malfunction. I think that’s the right name.
MAIM—a bit of a scary acronym, and also a nod to mutually assured destruction. Can you explain MAIM in plain language?
Let's think about what happened in nuclear strategy. Basically, a lot of states deterred each other from carrying out a first strike because they could then retaliate, so they had a shared vulnerability. They were saying, “We're not going to take this really aggressive action of trying to wipe you out, because that will end up causing us to be damaged.”
We have a somewhat similar situation later on, when AI is more salient, when it is viewed as pivotal to the future of a nation. When people are on the verge of making a superintelligence—when they can, say, automate pretty much all AI research—I think states would try to deter each other from leveraging that to develop something like a superweapon that would allow one country to crush the others, or from using those AIs to conduct a really rapid, automated AI research-and-development loop that could bootstrap from its current levels to something superintelligent, vastly more capable than any other system out there.
I think that later on, it becomes so destabilizing that China just says, “We're going to do something preemptive, like a cyberattack on your data center.” The U.S. might do that to China. Russia, coming out of Ukraine, will reassess the situation, get situationally aware, and think, “What's going on with the U.S. and China? Oh my goodness, they're so far ahead on AI. AI is looking like a big deal.”
Let's say it's later in the year, when a big chunk of software engineering is starting to be impacted by AI. Russia might think, “Oh, wow, this is looking pretty relevant. If you try to use this to crush us, we will prevent that by doing a cyberattack on you, and we will keep tabs on your projects,” because it is pretty easy for them to conduct that espionage. All they need to do is find a zero-day in Slack, and then they can know what DeepMind, OpenAI, xAI, and others are doing with very high fidelity.
It's pretty easy for them to conduct espionage and sabotage. Right now, they don't need to be threatening that because it's not at the level of severity; it's not actually that potentially destabilizing. The capabilities are still too distant, and a lot of decision-makers still aren't taking this AI stuff that seriously, relatively speaking.
But I think that will change as it gets more powerful, and then I think this is how they would end up responding. This makes sure we don't wind up in a situation where we are doing something extremely destabilizing, like trying to create a weapon that enables one country to totally wipe out the other, as was proposed by people like Leo.
What are the parallels here that you think make sense to nuclear and don't?
I think that, more broadly, as a dual-use technology, AI has civilian applications and military applications. Its economic applications are still, in some ways, limited, and likewise its military applications are still limited, but I think that will keep changing rapidly.
Chemical technology was important for the economy. It had some military use, but countries coordinated not to go down the chemical route. Biology as well can be used as a weapon and has enormous economic applications, and likewise with nuclear technology. I think AI has some of those properties.
For each of those technologies, countries did eventually coordinate to make sure that they didn't wind up in the hands of rogue actors like terrorists. There have been a lot of efforts to make sure that rogue actors don't get access to them and use them against those countries, because it's in neither country's interest.
Basically, biological weapons, for instance, and chemical weapons are a poor man's atom bomb, and this is why we have the Chemical Weapons Convention and the Biological Weapons Convention. That's where there's some shared interest. They might be rivals in other senses, in the way that the U.S. and the Soviet Union were rivals, but there's still coordination on that because it was incentive-compatible.
It doesn't benefit them in any way if terrorists have access to these sorts of things. It's just inherently destabilizing. So I think that's an opportunity for coordination. That isn't to say that they have an incentive to pause all forms of AI development, but it may mean that they would be deterred from some particular forms of AI development, in particular ones that have a very plausible prospect of enabling one country to get a decisive edge over another and crush them.
So, no superweapon-type stuff. But for more conventional types of warfare, like drones and things like that, I expect that they'll continue to race and probably not even coordinate on anything like that. That's just how things will go. That's like bows and arrows and nuclear weapons: it made sense for them to develop those sorts of weapons and threaten each other with them.
If you all could propose and magically adopt some policy or action for the current administration, what is the first step here? Is it the—
Yeah.
“We will not build a superweapon, and we're going to be watching for other people building them, too”?
As I've been alluding to throughout this whole conversation, what would the companies do? Not that much. They could add some basic antiterrorism safeguards, but I think this is pretty technically easy.
This is unlike refusal for other things. Refusal robustness for other things is harder. If you're trying to get at crimes and torts, that's harder because it's a lot messier. It overlaps with typical everyday interaction.
I think, likewise, the asks for states are not that challenging either. It's just a matter of them doing it. One step would be for the CIA to have a cell that's conducting more espionage of other states' AI programs, so that they have a better sense of what's going on and aren't caught by surprise.
Secondly, maybe some part of the government—let's say Cyber Command, which has a lot of cyberoffensive capabilities—gets some cyberattacks ready to disable other data centers in other countries if they look like they're running or creating a destabilizing AI project. That's it for deterrence.
For the nonproliferation of AI chips to rogue actors in particular, I think there would be some adjustments to export controls. In particular, we should reliably know where the AI chips are, for the same reason we want to know where our fissile material is, and for the same reason that we want Russia to know where its fissile material is. It's just generally good information to collect, and that can be done with some very basic statecraft by having a licensing regime.
For allies, they just notify you whenever the chips are being shipped to a different location, and they get a license exemption on that basis. Then you have enforcement officers prioritize doing some basic inspections of AI chips for end-use checks.
I think all of these are a few texts away or a basic document away, and I think that kind of 80/20 is a lot of it. Of course, this is always a changing situation. Safety isn't, as I've been trying to reinforce, really that much of a technical problem. This is more of a complex geopolitical problem with technical aspects.
Later on, maybe we'll need to do more. Maybe there will be some new risk sources that we need to take care of and adjust. But right now, I think espionage through the CIA, sabotage with Cyber Command, building up those capabilities, and buying those options seems like it takes care of a lot of the risk.
Let's talk about compute security.
Mm-hmm.
If we're talking about 100,000 networked, state-of-the-art chips, you can tell where that is. How do DeepSeek and the recent releases they've had factor into your view of compute security, given that export controls have clearly led to innovation toward highly compute-efficient pretraining that works on chips that China can import at what one might consider an irrelevant scale—a much smaller scale today?
It's hard for me to see, directionally, training becoming less efficient, even if people want to scale it up. Does that change your view at all?
No. I think it just undermines other types of strategies, like this Manhattan Project-type strategy of, “Let's move people out to the desert and do a big cluster there.” What it shows is that you can't rely as much on restricting another superpower's capabilities—their ability to make models.
You can restrict their intent, which is what deterrence does, but I don't think you can reliably or robustly restrict their capabilities. You can restrict the capabilities of rogue actors, and that's what I would want things like compute security and export controls to facilitate. Make sure it doesn't wind up in the hands of Iran or something.
China will probably keep getting some fraction of these chips, but we should basically just try to know more about where they're at, and we can tighten things up.
You could even coordinate with China to make sure that the chips aren't winding up in rogue actors' hands. I should also say that the export controls weren't actually a substantial priority among leadership at BIS, to my understanding. The AI chips were a substantial priority for some people, but not for the enforcement officers.
Did any of them go to Singapore to see where those 10% of Nvidia's chips were going? I think they would've very quickly found, “Oh, they were going to China.” Some basic end-use check would've taken care of that. I don't think this means that export controls don't work.
We've done nonproliferation of lots of other things, like chemical agents and fissile material, so it can be done if people care. But even so, I still think that if you really tightened the export controls and made it so that China couldn't get any of those chips at all, and this was one of your biggest priorities, they'd just steal the weights anyway. I think it'll be too difficult to totally restrict their capabilities, but I think you can restrict their intent through deterrence.
It also seems like either stuff is powerful or it's not. It seems infeasible to me, given the economic opportunity, that China will say, “We don't need the capability.”
I fail to see a version of the world where leadership in another great power that believes there's value here says, “We don't need that,” from an economic-value perspective.
Yeah, that's right. Just for a lot of these, maybe it would be nicer if everything went 3× slower, and maybe there'd be fewer mess-ups if there were some magic button that would do that. I don't know whether that's true or not, actually. I don't have a position on that.
Given the structural constraints and the competitive pressures between these companies and between these states, it just makes a lot of these things infeasible. A lot of these other gestures that could be useful for risk mitigation, when you consider them in light of the structural realities, just become a lot less tractable.
That said, there still would be, in some ways, some pausing or halting of development of particular projects that you could potentially lose control of, or that, if controlled, would be very destabilizing because it would enable one country to crush the other. I think people's conception of what risk management looks like is that it's a peacenik thing or something like that. It's all kumbaya, and we just have to ignore the structural realities of operating in this space.
I think instead the right approach toward this is sort of like nuclear strategy. It is an evolving situation; it depends. There are some basic things you can do. You're probably going to need to stockpile nuclear weapons, secure a second strike, keep an eye on what they're doing, and make sure that there isn't proliferation to rogue actors when the capabilities are extremely hazardous.
This is a continual battle, but it's not going to be clearly an extremely positive thing no matter what. It's not going to be doomsday no matter what for nuclear strategy. It was obviously risky business. The Cuban Missile Crisis became pretty close to an all-out nuclear war.
It depends on what we do. I think some basic interventions and some very basic statecraft can take care of a lot of these sorts of risks and make it manageable. I imagine then we're left with more domestic-type problems, like what to do about automation and things like that. But I think maybe we'll be able to get a handle on some of the geopolitics here.
I want to change tack for our last couple of minutes and talk about evals. It's obviously very related to safety and understanding where we are in terms of capability. You came out with a triggeringly named Humanity's Last Exam eval, and then also Enigma. Why are these relevant, and where are we in evals?
Yeah. For context, I've been making evaluations to try to understand where we're at in AI for about as long as I've been doing AI research. Previously, I've done some datasets like MMLU and the MATH dataset. Before that, before ChatGPT, there were things like ImageNet-C and other sorts of things.
Humanity's Last Exam was basically an attempt at getting at what would be the end of the road for evaluations and benchmarks based on exam-like questions—ones that test some sort of academic knowledge. For this, we asked professors and researchers around the world to submit a really challenging question, and then we would add that to the dataset.
It's a big collection of what professors, for instance, would encounter as challenging problems in their research that have a definitive, closed-ended, objective answer. With that, I think the genre of closed-ended questions, where there's just a multiple-choice or simple short answer, will roughly have expired when performance on this dataset is near the ceiling.
When performance is near the ceiling, I think that would basically be an indication that you have something like a superhuman mathematician or a superhuman STEM scientist in many ways, when closed-ended questions are very useful, such as in math. But it doesn't get at other things to measure, such as its ability to perform open-ended tasks.
That's more agent-type evaluation, and I think that will take more time. So we'll try to measure directly its ability to automate various digital tasks: collect various digital tasks, have it work on them for a few hours, and see if it successfully completed them—something like that coming out soon.
We have a test for closed-ended questions, things that test knowledge in academia, like mathematics. But they still are very bad at agent stuff. This could possibly change overnight, but it's still near the floor. I think they're still extremely defective as agents, so there'll need to be more evaluations for that.
The overall approach is just to try to understand what's going on, what's the rate of development, so that the public can at least understand what's happening. If all the evaluations are saturated, it's difficult to even have a conversation about the state of AI. Nobody really knows exactly where it's at, where it's going, or what the rate of improvement is.
Is there anything that qualitatively changes when, let's say, these models and model systems are just better than humans—exceeding human capability—in how we do evals? Does it change our ability to evaluate them?
The intelligence frontier is just so jagged. What things they can do and can't do is often surprising. They still can't fold clothes. They can answer a lot of tough physics problems, though. Why that is, there are complicated reasons.
So it's not all uniform. In some ways, they'll be better than humans. It seems totally plausible that they'll be better than humans at mathematics not too long from now, but still not able to book a flight.
The implications of that are that, when you have them being better, they might just be better in some limited ways. That might have limited influence in its domain, but not necessarily generalize to other sorts of things.
But I do think it's possible that they'll be better at reasoning skills than us. We still could have humans checking, because they can still verify. If an AI mathematician is better than a human, humans can still run the proof through a proof checker and then confirm that it was correct. So in that way, humans can still understand what's going on in some ways.
But in other ways, like if they're getting better taste in things—if that makes any sense; maybe it doesn't make any philosophical sense—that would be pretty difficult for people to confirm. I think we're on track overall to have AIs that have really good oracle-like skills. You can ask them things and just think, “Wow, it just totally said something insightful or very nontrivial, or pushed the boundaries of knowledge in some particular way.”
But they won't necessarily be able to carry out tasks on behalf of people for some while. I think this is why we don't take the AIs that seriously, because they still can't do a lot of very trivial stuff.
But when they get some of the agent skills, then I don't think there will be many barriers between their economic impacts and people thinking that this is kind of an interesting thing and thinking that it's the most important thing. I think that's an emergent property with agent skills: the vibes really shift, and it's pretty clear that this is much bigger than some prior technology, like the App Store or social media. It's in a category of its own.
Well, Dan, thanks for doing this.
It was a great conversation.
Yeah, glad. Thank you for having me.
Find us on Twitter at nopriorspod. Subscribe to our YouTube channel if you wanna see our faces. Follow the show on Apple Podcasts, Spotify, or wherever you listen. That way you get a new episode every week. And sign up for emails or find transcripts for every episode at no-priors.com.