Sarah Guo
Hi, listeners, and welcome back to No Priors. Today, I'm with Dan Hendrycks, AI researcher and director of the Center for AI Safety. He's published papers and widely used evals such as MMLU and, most recently, Humanity's Last Exam. He's also published Superintelligence Strategy alongside authors including former Google CEO Eric Schmidt and Scale founder Alex Wang. We talk about AI safety and geopolitical implications, analogies to nuclear, compute security, and the state of evals. Dan, thanks for doing this.
Dan Hendrycks
Glad to be here.
Sarah Guo
How’d you end up working on AI safety?
Dan Hendrycks
AI was pretty clearly going to be a big deal if one would just think through its conclusion. Early on, it seemed like other people were ignoring it because it was weirder or not that pleasant to think about. It’s hard to wrap your head around, but it seemed like the most important thing during this century. I thought that would be a good place to develop my career toward, and that’s why I started on it early on.
Since it’d be such a big deal, we’d need to make sure that we can think about it properly, channel it in a productive direction, and take care of some sort of tail risks, which are generally systematically under-addressed. That’s why I got into it: it’s a big deal, and people weren’t really doing much about it at the time.
Sarah Guo
What do you think of as the center’s role versus safety efforts within the large labs?
Dan Hendrycks
Well, there aren’t that many safety efforts in the labs even now. I think the labs can just focus on doing some very basic measures to refuse queries like, “Help me make a virus,” and things like that. But I don’t think labs have an extremely large role in safety overall or in making this go well.
They’re kind of predetermined to race. They can’t really choose not to unless they would no longer be a relevant company in the arena. I think they can reduce terrorism risks or some accidents, but beyond that, I don’t think they can dramatically change the outcomes in too substantial of a way.
Because a lot of this is geopolitically determined, if companies decide to act very differently, there’s the prospect of competing with China, or maybe Russia will become relevant later. As that happens, this constrains their behavior substantially.
I’ve been interested in tackling AI at multiple levels. There are things companies can do to have some very basic anti-terrorism safeguards, which are pretty easy to implement. There are also the economic effects that will need to be managed well, and companies can’t really change how that goes either.
It’s going to cause mass disruptions to labor and automate a lot of digital labor. If they tinker with the design choice or add some different refusal data, it doesn’t change that fact. Making AI go well and managing the risks is much more of a broader problem. It’s got some technical aspects, but I think that’s a small part of it.
Sarah Guo
I don’t know that the leaders of the labs would say, “We can do nothing about this,” but maybe it’s also a question of everybody having equity in this equation, right? Maybe it’s also a question of semantics. Can you describe how you think about the difference between alignment and safety?
Dan Hendrycks
I’m just using safety as a sort of catchall for dealing with risks. There are other risks, like if you never get really intelligent AI systems, that poses some risks in itself. There are other sorts of risks that aren’t necessarily technical, like concentration of power.
So I view the distinction between alignment and safety as alignment being a sort of subset of safety. Obviously, you want the value systems of the AIs to be in keeping with or compatible with, say, the US public for US AIs, or with you as an individual, but that doesn’t necessarily make it safe.
If you have an AI that’s reliably obedient or aligned to you, this doesn’t make everything work totally well. China can have AIs that are totally aligned with them. The US can have AIs that are totally aligned with them. You still are going to have a strategic competition between the 2.
They’re going to need to integrate it into their militaries. They’re probably going to need to integrate it really quickly. This competition is going to force them to have a higher risk tolerance in the process. So even if the AIs are reliably doing their principals’ bidding, this doesn’t necessarily make the overall situation perfectly fine.
I think it’s not just a question of reliability or whether they do what you want. There are other structural pressures that cause this to be riskier, like geopolitics.
Sarah Guo
At the highest level, with a bundle of weights that’s increasingly capable, why do we care about AI from a national security perspective? What’s the most practical way it matters in geopolitics or gets used as a weapon?
Dan Hendrycks
I think that AI isn’t that powerful currently in many respects. So, in many ways, it’s not actually that relevant for national security currently. This could well change within a year’s time. Generally, I’ve been focused on the trajectory that it’s on, as opposed to saying that right now it is extremely concerning.
That said, for instance, in cyber, I don’t think AIs are that relevant for being able to pull off a devastating cyberattack on the grid by a malicious actor currently. That said, we should look at cyber, be prepared, and think about its strategic implications.
There are other capabilities, like virology. The AIs are getting very good at STEM PhD-level topics, and that includes virology. So I think they’re sort of rounding the corner on being able to provide expert-level capabilities in terms of their knowledge of the literature or even helping in practical wet-lab situations.
I do think that, on the virology aspect, they already have national security implications, but that’s only very recently with the reasoning models. In many other respects, they’re not as relevant.
It’s more prospective that AI could well become the way in which a nation might try to dominate another nation, and the backbone for not just war but also economic security. The number of ships that the US has versus China might be the determinant of which country is the most prosperous and which one falls behind.
This is all prospective. I don’t think it’s just speculative. It’s speculative in the same way that NVIDIA’s valuation is speculative, or the valuations behind AI companies are speculative. It’s something that I think a lot of people are expecting, and expecting fairly soon.
Sarah Guo
Yeah, it’s quite hard to think about time horizons in AI. We invest in things that I think of as medium-term speculative, but they get pulled in quite quickly.
Just because you mentioned both cyber and bio, we’re investors in companies like Culminate or Sibyl on the defensive cybersecurity side, or Chai and Somite on the biotech discovery side, modeling different systems in biology that will help us with treatments. How do you think about the balance of competition, benefits, and safety? Some of these things, I think, are working effectively in the near term on the positive side as well.
Dan Hendrycks
Yeah, I don’t get this big trade-off. For bio, if you want to expose those capabilities, just talk to sales and get the enterprise account. Here, you can have the little refusal thing for virology.
But if you just created an account a second ago and you’re asking it how to culture this virus, saying, “Here’s your picture of your petri dish. What’s the next step that you should do?”—if you want access to those capabilities, you can speak to sales. That’s basically in xAI’s risk management framework: we’re not exposing those expert-level capabilities to people who we don’t know.
But if we do, then sure, have them. Likewise with cyber, I think you can very easily capture the benefits while taking care of some of these pretty avoidable tail risks. Once you have that, you’ve basically taken care of malicious use for the models behind your API, and that’s about the best that you can do as a company.
You could try to influence policy by using your voice or something, but I don’t see a substantial amount that they can do. They could do some research to try to make the models more controllable, or try to make policymakers more aware of the situation more broadly in terms of where we’re going.
I don’t think policymakers have internalized what’s happening in AI at all. They still think it’s just selling hype, and they don’t actually believe that the companies’ employees actually believe that we could get AGI in, so to speak, the next few years.
I don’t see really substantial trade-offs there. I think the complications really come about when we’re dealing with what the right stringency in export controls is, for instance. That’s complicated.
If you turn the pain dial all the way up for China in export controls, and if AI chips are the currency of economic power in the future, then this increases the probability that they want to invade Taiwan. They already want to. This would give them all the more reason if AI chips are the main thing and they’re not getting any of it, and they’re not even getting the latest semiconductor manufacturing tools for making cutting-edge CPUs, let alone GPUs.
Those are some other types of complicated problems that we have to address, think about, and calibrate appropriately. But in terms of just mitigating virology risks, if you’re at Genentech or a biotech startup, just speak to sales, and then you have access to those capabilities. Problem solved.
Sarah Guo
What is a way you actually expect AI to get used as a weapon beyond virology and security?
Dan Hendrycks
I wouldn’t expect a bioweapon from a state actor. From a non-state actor, that would make a lot more sense. I think cyber makes sense from state actors and non-state actors. Then there are drone applications. These could disrupt other things. They could help with other types of weapons research, like exploring exotic EMPs, and could help create better types of drones. They could substantially help with situational awareness, so that one might know where all the nuclear submarines are.
Some advances in AI might be able to help with that, and that could disrupt our second-strike capabilities and mutually assured destruction. Those are some geopolitical implications. It could potentially bear on nuclear deterrence, and that’s not even a weapon. The example of heightened situational awareness and being able to pinpoint where hardened land-based nuclear launch sites are, or where nuclear submarines are, is just informational but could nonetheless be extremely disruptive and destabilizing.
Outside of that, the default conventional AI weapon would be drones. I don’t know if it makes sense that countries would compete on that, and I think that would be a mistake if the US weren’t trying to do more in manufacturing drones.
Sarah Guo
I started working recently with an electronic warfare company. I think there’s a massive lack of understanding of just the basic concept: We have autonomous systems, and they all have communication systems. Our missile systems have targeting and communication systems. From a battlefield-awareness and control perspective, a lot of that fight will be won with radio, radar, and related systems, right?
Dan Hendrycks
Mm-hmm.
Sarah Guo
And so I think there’s an area where AI is going to be very relevant and is already very relevant in Ukraine.
Dan Hendrycks
Speaking about AI assisting with command and control, I was hearing a story about how, on Wall Street, you always had a human in the loop for each decision. At a later stage, before they removed that requirement on Wall Street, you just had rows of people clicking the “Accept, accept, accept” button. We’re getting to a similar state in some contexts with AI.
It wouldn’t surprise me if we ended up automating more of that decision-making. This just turns into questions of reliability, and doing some reliability research seems useful. To return to that larger question of where the safety trade-offs are, I think people are largely thinking that the push for risk management is to do some sort of pausing or something like that.
An issue is that you need teeth behind an agreement. If you do it voluntarily, you just make yourself less powerful, and you let the worst actors get ahead of you. You could say, “Well, we’ll sign a treaty.” We will not assume that the treaty will be followed. That would be very imprudent. You would actually need some sort of threat of force or something to back it up, or some verification mechanism.
But absent that, if it’s entirely voluntary, then this doesn’t seem like a useful thing at all. I think people’s conflation of safety is: What we must do is voluntarily slow it down. It just doesn’t make as much geopolitical sense unless you have some threat of force to back it up or some very strong verification mechanism. But in the absence of that—
Sarah Guo
As a proxy, there’s clearly been very little compliance with either treaties or norms around cyberattacks and corporate espionage, right?
Dan Hendrycks
Yeah. I mean, corporate espionage, for instance—that was one strategy, the sort of voluntary-pause strategy. People are thinking that equals safety. Then maybe last year there was that paper, “Situational Awareness,” written by Leopold Aschenbrenner, and he’s a sort of safety person. His idea was, “Let’s instead try to beat China to superintelligence as much as possible.”
But that has some weaknesses because it assumes that corporate espionage will not be a thing at all, which is very difficult to do. I mean, at some of these top AI companies, 30% or more of the employees are Chinese nationals. This is not feasible. If you get rid of them, they’re going to go to China, and then they’re probably going to beat you because they’re extremely important for US success.
So you’re going to want to keep them here, but that’s going to expose you to some information-security issues. But that’s just too bad.
Sarah Guo
Do you have a point of view on how we should change immigration policy, if at all, given these risks?
Dan Hendrycks
I would, of course, claim that the policy on this should be totally separate from southern border policy and broader policy. But if we’re talking about AI researchers, if they’re very talented, then I think you’d want to make it easier, and I think it’s probably too difficult for many of them to stay currently. I think that discussion should be kept totally separate from southern border policy.
Sarah Guo
Just in terms of broad strokes, things that you think won’t work: voluntary compliance and assuming that will happen, or just a straight race?
Dan Hendrycks
We want to be competitive, and I think racing in other sorts of spheres, say drones or AI chips, seems fine. If you’re saying, “Let’s race to superintelligence to try and get ahead and turn that into a weapon to crush them,” and they’re not going to do the same, or they’re not going to have access to it, or they’re not going to prevent that from happening, that seems like quite a tall claim.
I mean, if we did have a substantially better AI, they could just co-opt it. They could just steal it, unless you had really, really strong information security—for example, if you moved the AI researchers out to the desert. But then you’re reducing your probability of actually beating them because a lot of your best scientists ended up going back to China.
Even then, if there were signs that they were really pulling ahead and going to be able to get some powerful AI that would enable the US to crush China, they would then try to deter them from doing something like that. They’re not going to sit idly by and say, “You know what? Yeah, go ahead. Develop your superintelligence or whatever, and then you can boss us around, and we’ll just accept your dictates till the end of time.”
I think there is a failure of some sort of second-order reasoning going on there, which is: How would China respond to this sort of maneuver if we’re building a trillion-dollar compute cluster in the desert, totally visible from space? Basically, the only plausible read on this is that this is a bid for dominance or a sort of monopoly on superintelligence.
It reminds me of the nuclear era. There was a brief period where some people were saying, “You know what? We’ve got to just preemptively destroy or preventively destroy the USSR. We’ve got to nuke ’em.” Even pacifists, or people who are normally pacifists, like Bertrand Russell, were advocating for this. The window of opportunity for that maybe never existed, but there was a prospect of it for some time.
I don’t think the opportunity window really exists here because of the complex interdependence and multinational talent dependence in the United States. I don’t think you can have China be totally severed from any awareness or any ability to gain insight or imitate what we’re doing here.
Sarah Guo
We’re clearly nowhere close to that in a real environment right now, right?
Dan Hendrycks
No, it would take years. It would take years to do well, and I don’t even think the timelines for some very powerful AI systems leave enough time to do that securitization anyway.
Sarah Guo
Yeah.
So, okay, in reaction, you propose, along with some other esteemed authors and friends, Eric Schmidt and Alex Wang, a new deterrence regime: Mutually Assured AI Malfunction. I think that’s the right name.
MAIM—a bit of a scary acronym, and also a nod to mutually assured destruction. Can you explain MAIM in plain language?
Dan Hendrycks
Let's think about what happened in nuclear strategy. Basically, a lot of states deterred each other from carrying out a first strike because they could then retaliate, so they had a shared vulnerability. They were saying, “We're not going to take this really aggressive action of trying to wipe you out, because that will end up causing us to be damaged.”
We have a somewhat similar situation later on, when AI is more salient, when it is viewed as pivotal to the future of a nation. When people are on the verge of making a superintelligence—when they can, say, automate pretty much all AI research—I think states would try to deter each other from leveraging that to develop something like a superweapon that would allow one country to crush the others, or from using those AIs to conduct a really rapid, automated AI research-and-development loop that could bootstrap from its current levels to something superintelligent, vastly more capable than any other system out there.
I think that later on, it becomes so destabilizing that China just says, “We're going to do something preemptive, like a cyberattack on your data center.” The U.S. might do that to China. Russia, coming out of Ukraine, will reassess the situation, get situationally aware, and think, “What's going on with the U.S. and China? Oh my goodness, they're so far ahead on AI. AI is looking like a big deal.”
Let's say it's later in the year, when a big chunk of software engineering is starting to be impacted by AI. Russia might think, “Oh, wow, this is looking pretty relevant. If you try to use this to crush us, we will prevent that by doing a cyberattack on you, and we will keep tabs on your projects,” because it is pretty easy for them to conduct that espionage. All they need to do is find a zero-day in Slack, and then they can know what DeepMind, OpenAI, xAI, and others are doing with very high fidelity.
It's pretty easy for them to conduct espionage and sabotage. Right now, they don't need to be threatening that because it's not at the level of severity; it's not actually that potentially destabilizing. The capabilities are still too distant, and a lot of decision-makers still aren't taking this AI stuff that seriously, relatively speaking.
But I think that will change as it gets more powerful, and then I think this is how they would end up responding. This makes sure we don't wind up in a situation where we are doing something extremely destabilizing, like trying to create a weapon that enables one country to totally wipe out the other, as was proposed by people like Leo.
Sarah Guo
What are the parallels here that you think make sense to nuclear and don't?
Dan Hendrycks
I think that, more broadly, as a dual-use technology, AI has civilian applications and military applications. Its economic applications are still, in some ways, limited, and likewise its military applications are still limited, but I think that will keep changing rapidly.
Chemical technology was important for the economy. It had some military use, but countries coordinated not to go down the chemical route. Biology as well can be used as a weapon and has enormous economic applications, and likewise with nuclear technology. I think AI has some of those properties.
For each of those technologies, countries did eventually coordinate to make sure that they didn't wind up in the hands of rogue actors like terrorists. There have been a lot of efforts to make sure that rogue actors don't get access to them and use them against those countries, because it's in neither country's interest.
Basically, biological weapons, for instance, and chemical weapons are a poor man's atom bomb, and this is why we have the Chemical Weapons Convention and the Biological Weapons Convention. That's where there's some shared interest. They might be rivals in other senses, in the way that the U.S. and the Soviet Union were rivals, but there's still coordination on that because it was incentive-compatible.
It doesn't benefit them in any way if terrorists have access to these sorts of things. It's just inherently destabilizing. So I think that's an opportunity for coordination. That isn't to say that they have an incentive to pause all forms of AI development, but it may mean that they would be deterred from some particular forms of AI development, in particular ones that have a very plausible prospect of enabling one country to get a decisive edge over another and crush them.
So, no superweapon-type stuff. But for more conventional types of warfare, like drones and things like that, I expect that they'll continue to race and probably not even coordinate on anything like that. That's just how things will go. That's like bows and arrows and nuclear weapons: it made sense for them to develop those sorts of weapons and threaten each other with them.
Sarah Guo
If you all could propose and magically adopt some policy or action for the current administration, what is the first step here? Is it the—
Dan Hendrycks
Yeah.
Sarah Guo
“We will not build a superweapon, and we're going to be watching for other people building them, too”?
Dan Hendrycks
As I've been alluding to throughout this whole conversation, what would the companies do? Not that much. They could add some basic antiterrorism safeguards, but I think this is pretty technically easy.
This is unlike refusal for other things. Refusal robustness for other things is harder. If you're trying to get at crimes and torts, that's harder because it's a lot messier. It overlaps with typical everyday interaction.
I think, likewise, the asks for states are not that challenging either. It's just a matter of them doing it. One step would be for the CIA to have a cell that's conducting more espionage of other states' AI programs, so that they have a better sense of what's going on and aren't caught by surprise.
Secondly, maybe some part of the government—let's say Cyber Command, which has a lot of cyberoffensive capabilities—gets some cyberattacks ready to disable other data centers in other countries if they look like they're running or creating a destabilizing AI project. That's it for deterrence.
For the nonproliferation of AI chips to rogue actors in particular, I think there would be some adjustments to export controls. In particular, we should reliably know where the AI chips are, for the same reason we want to know where our fissile material is, and for the same reason that we want Russia to know where its fissile material is. It's just generally good information to collect, and that can be done with some very basic statecraft by having a licensing regime.
For allies, they just notify you whenever the chips are being shipped to a different location, and they get a license exemption on that basis. Then you have enforcement officers prioritize doing some basic inspections of AI chips for end-use checks.
I think all of these are a few texts away or a basic document away, and I think that kind of 80/20 is a lot of it. Of course, this is always a changing situation. Safety isn't, as I've been trying to reinforce, really that much of a technical problem. This is more of a complex geopolitical problem with technical aspects.
Later on, maybe we'll need to do more. Maybe there will be some new risk sources that we need to take care of and adjust. But right now, I think espionage through the CIA, sabotage with Cyber Command, building up those capabilities, and buying those options seems like it takes care of a lot of the risk.
Sarah Guo
Let's talk about compute security.
Dan Hendrycks
Mm-hmm.
Sarah Guo
If we're talking about 100,000 networked, state-of-the-art chips, you can tell where that is. How do DeepSeek and the recent releases they've had factor into your view of compute security, given that export controls have clearly led to innovation toward highly compute-efficient pretraining that works on chips that China can import at what one might consider an irrelevant scale—a much smaller scale today?
It's hard for me to see, directionally, training becoming less efficient, even if people want to scale it up. Does that change your view at all?
Dan Hendrycks
No. I think it just undermines other types of strategies, like this Manhattan Project-type strategy of, “Let's move people out to the desert and do a big cluster there.” What it shows is that you can't rely as much on restricting another superpower's capabilities—their ability to make models.
You can restrict their intent, which is what deterrence does, but I don't think you can reliably or robustly restrict their capabilities. You can restrict the capabilities of rogue actors, and that's what I would want things like compute security and export controls to facilitate. Make sure it doesn't wind up in the hands of Iran or something.
China will probably keep getting some fraction of these chips, but we should basically just try to know more about where they're at, and we can tighten things up.
You could even coordinate with China to make sure that the chips aren't winding up in rogue actors' hands. I should also say that the export controls weren't actually a substantial priority among leadership at BIS, to my understanding. The AI chips were a substantial priority for some people, but not for the enforcement officers.
Did any of them go to Singapore to see where those 10% of Nvidia's chips were going? I think they would've very quickly found, “Oh, they were going to China.” Some basic end-use check would've taken care of that. I don't think this means that export controls don't work.
We've done nonproliferation of lots of other things, like chemical agents and fissile material, so it can be done if people care. But even so, I still think that if you really tightened the export controls and made it so that China couldn't get any of those chips at all, and this was one of your biggest priorities, they'd just steal the weights anyway. I think it'll be too difficult to totally restrict their capabilities, but I think you can restrict their intent through deterrence.
Sarah Guo
It also seems like either stuff is powerful or it's not. It seems infeasible to me, given the economic opportunity, that China will say, “We don't need the capability.”
I fail to see a version of the world where leadership in another great power that believes there's value here says, “We don't need that,” from an economic-value perspective.
Dan Hendrycks
Yeah, that's right. Just for a lot of these, maybe it would be nicer if everything went 3× slower, and maybe there'd be fewer mess-ups if there were some magic button that would do that. I don't know whether that's true or not, actually. I don't have a position on that.
Given the structural constraints and the competitive pressures between these companies and between these states, it just makes a lot of these things infeasible. A lot of these other gestures that could be useful for risk mitigation, when you consider them in light of the structural realities, just become a lot less tractable.
That said, there still would be, in some ways, some pausing or halting of development of particular projects that you could potentially lose control of, or that, if controlled, would be very destabilizing because it would enable one country to crush the other. I think people's conception of what risk management looks like is that it's a peacenik thing or something like that. It's all kumbaya, and we just have to ignore the structural realities of operating in this space.
I think instead the right approach toward this is sort of like nuclear strategy. It is an evolving situation; it depends. There are some basic things you can do. You're probably going to need to stockpile nuclear weapons, secure a second strike, keep an eye on what they're doing, and make sure that there isn't proliferation to rogue actors when the capabilities are extremely hazardous.
This is a continual battle, but it's not going to be clearly an extremely positive thing no matter what. It's not going to be doomsday no matter what for nuclear strategy. It was obviously risky business. The Cuban Missile Crisis became pretty close to an all-out nuclear war.
It depends on what we do. I think some basic interventions and some very basic statecraft can take care of a lot of these sorts of risks and make it manageable. I imagine then we're left with more domestic-type problems, like what to do about automation and things like that. But I think maybe we'll be able to get a handle on some of the geopolitics here.
Sarah Guo
I want to change tack for our last couple of minutes and talk about evals. It's obviously very related to safety and understanding where we are in terms of capability. You came out with a triggeringly named Humanity's Last Exam eval, and then also Enigma. Why are these relevant, and where are we in evals?
Dan Hendrycks
Yeah. For context, I've been making evaluations to try to understand where we're at in AI for about as long as I've been doing AI research. Previously, I've done some datasets like MMLU and the MATH dataset. Before that, before ChatGPT, there were things like ImageNet-C and other sorts of things.
Humanity's Last Exam was basically an attempt at getting at what would be the end of the road for evaluations and benchmarks based on exam-like questions—ones that test some sort of academic knowledge. For this, we asked professors and researchers around the world to submit a really challenging question, and then we would add that to the dataset.
It's a big collection of what professors, for instance, would encounter as challenging problems in their research that have a definitive, closed-ended, objective answer. With that, I think the genre of closed-ended questions, where there's just a multiple-choice or simple short answer, will roughly have expired when performance on this dataset is near the ceiling.
When performance is near the ceiling, I think that would basically be an indication that you have something like a superhuman mathematician or a superhuman STEM scientist in many ways, when closed-ended questions are very useful, such as in math. But it doesn't get at other things to measure, such as its ability to perform open-ended tasks.
That's more agent-type evaluation, and I think that will take more time. So we'll try to measure directly its ability to automate various digital tasks: collect various digital tasks, have it work on them for a few hours, and see if it successfully completed them—something like that coming out soon.
We have a test for closed-ended questions, things that test knowledge in academia, like mathematics. But they still are very bad at agent stuff. This could possibly change overnight, but it's still near the floor. I think they're still extremely defective as agents, so there'll need to be more evaluations for that.
The overall approach is just to try to understand what's going on, what's the rate of development, so that the public can at least understand what's happening. If all the evaluations are saturated, it's difficult to even have a conversation about the state of AI. Nobody really knows exactly where it's at, where it's going, or what the rate of improvement is.
Sarah Guo
Is there anything that qualitatively changes when, let's say, these models and model systems are just better than humans—exceeding human capability—in how we do evals? Does it change our ability to evaluate them?
Dan Hendrycks
The intelligence frontier is just so jagged. What things they can do and can't do is often surprising. They still can't fold clothes. They can answer a lot of tough physics problems, though. Why that is, there are complicated reasons.
So it's not all uniform. In some ways, they'll be better than humans. It seems totally plausible that they'll be better than humans at mathematics not too long from now, but still not able to book a flight.
The implications of that are that, when you have them being better, they might just be better in some limited ways. That might have limited influence in its domain, but not necessarily generalize to other sorts of things.
But I do think it's possible that they'll be better at reasoning skills than us. We still could have humans checking, because they can still verify. If an AI mathematician is better than a human, humans can still run the proof through a proof checker and then confirm that it was correct. So in that way, humans can still understand what's going on in some ways.
But in other ways, like if they're getting better taste in things—if that makes any sense; maybe it doesn't make any philosophical sense—that would be pretty difficult for people to confirm. I think we're on track overall to have AIs that have really good oracle-like skills. You can ask them things and just think, “Wow, it just totally said something insightful or very nontrivial, or pushed the boundaries of knowledge in some particular way.”
But they won't necessarily be able to carry out tasks on behalf of people for some while. I think this is why we don't take the AIs that seriously, because they still can't do a lot of very trivial stuff.
But when they get some of the agent skills, then I don't think there will be many barriers between their economic impacts and people thinking that this is kind of an interesting thing and thinking that it's the most important thing. I think that's an emergent property with agent skills: the vibes really shift, and it's pretty clear that this is much bigger than some prior technology, like the App Store or social media. It's in a category of its own.
Sarah Guo
Well, Dan, thanks for doing this.
It was a great conversation.
Dan Hendrycks
Yeah, glad. Thank you for having me.
Sarah Guo
Find us on Twitter at nopriorspod. Subscribe to our YouTube channel if you wanna see our faces. Follow the show on Apple Podcasts, Spotify, or wherever you listen. That way you get a new episode every week. And sign up for emails or find transcripts for every episode at no-priors.com.