Adam Gleave
It’s great to be back here. Thanks for hosting me, Nathan.
Speaker 1
My pleasure. I’m excited about this. This is basically going to be a wide-ranging catch-up conversation. I want to get your take on all things AI. Last time, we went deeper and more narrowly focused on some research, and we’ll touch on some of your latest research today as well. But since you’ve got your hands in a lot of pots and the organization is growing and taking on more different kinds of work, I thought you’d be the perfect person to check in with and try to make sense of where we are as we head into the final months of 2025.
Adam Gleave
Yep.
Speaker 1
So—
Adam Gleave
I’m happy to help out. I can’t promise to deconfuse everything; it’s a very confusing landscape. But we’re certainly doing lots of different things at FAR AI, and I also welcome the opportunity to clarify a little bit why we’re doing these things, because I think people sometimes find us a bit confusing as an organization from outside.
Speaker 1
Cool. Well, we’ll hopefully do all of that and more. My first question is inspired by a paper from a few months ago called “Gradual Disempowerment.” The authors went on to run a workshop on those themes. They posed the question: “AGI equilibria: Are there any good ones?”
This is a more negative framing of a question I often ask, which is: What is your positive vision for the post-AGI future? Can you articulate a post-AGI vision of life that people might find compelling, or at least not scary?
Adam Gleave
Yeah. I’ll give it my best shot. I’m relatively optimistic, and I think that although there are some major risks that could really derail things entirely, the most likely outcome is that we muddle through. Nonetheless, I do think this gradual-disempowerment framing is quite a powerful one, because perhaps the most likely outcome is that we do muddle through and are somewhat, but not wholly, disempowered. We’ve fallen short of where we could have ended up, but we’re perhaps still much better off than we are now.
I’ll sketch out that positive but slightly pessimistic view of doing better, but not as well as we could have done. Then maybe I’ll talk a bit about the most tractable upsides and some of the most salient downsides.
I think the scenario I find plausible, where things are overall pretty good but just not as good as they could have been, is that we do get AGI and it’s approximately aligned in the same way that current systems are approximately aligned. They’re not perfect, but Claude is a pretty nice guy. We have some errors that come up now and again, and we trial-and-error our way to fixing them.
There’s some concentration of power because there are a handful of companies and nation-states at the frontier. But these actors are not actively malevolent. In fact, some of them are fairly nice, liberal-democratic actors, or companies with a joint nonprofit mission. So there’s huge inequality, but overall people are still vastly better off than they are now in absolute terms.
We get some things that no human would have chosen to do. Some parts of the economy are just completely automated and maybe are a bit rent-seeking: they’re addicting humans, or it’s just this situation where AI is trading with other AI without generating that much value for humans. But there’s such a large wealth surplus because we’ve been able to automate pretty much everything that people do right now that this kind of inefficiency is still fine.
What does it actually look like to be human in this world? I think a good historical analogy would be a bit like being European nobility, or perhaps not the first heir to European nobility, but the third son or something. You’ve got this very nice life. You don’t really have much purpose, but your life is pretty good. You can do various kinds of hobbies and have some influence, but the main things going on in the world are a little bit beyond your control.
I think this might not be the most inspiring vision, but overall, I’d say that the third sons of European nobility had a pretty good life. The third daughters maybe less so. I think we’ve been able to find meaning without necessarily being 100% in the driver’s seat, although some people might fare less well in that kind of life than others.
Another area where I think things could actually go really well, but is a bit fragile and less guaranteed, is if AI itself is a source of major moral value. In principle, I’d say there’s no reason why we should be carbon chauvinists. We don’t want to assign intelligent life that’s running on silicon less value than we would assign to a biological entity.
There’s a really big design space where we could potentially make AIs that are not only as capable as or better than us, but maybe also have more ability to experience really positive feelings. You could create many such AI systems, and they could live in environments that humans would find inhospitable.
Again, even if there’s a lot of waste here—maybe only a small fraction of AI systems are having some kind of amazing existence, while most are doing some kind of boring, superintelligent bookkeeping—if those other AIs are still having a neutral or slightly positive existence, then I think that could be very valuable. It’s not just their internal moral value. They could also be producing artwork that otherwise would never have existed, along with all these other kinds of sources of value.
I think that’s a positive view. The thing that highlights is that there can be a big difference between the median outcome, where we manage to eke out a small fraction of the resources of really positive outcomes, versus one where that’s what we use this huge wealth surplus to really double down on and be deliberate about.
I think this also highlights a potential big downside risk. Part of my assumption here was that most of what’s happening is not because people are trying to create value, but because of more fundamental competitive forces that are neutral. At worst, it’s a zero-sum game.
I think this is a pretty reasonable assumption. I’d say that’s true for most of the capitalist economy: most companies are neutral to slightly positive. I used to work in quantitative finance, making the capital markets more efficient, and I think it’s genuinely good for the world. It’s just very overcompensated relative to the value that it adds.
If we’re in that world, I feel mostly okay. But there might be some aspects of this kind of competitive economy that look more like warfare or factory farming, where it’s really quite negative. Factory farming has this huge negative externality on animals, and it’s also not that good for people. It’s not necessarily that healthy a product, but it exists because it’s an economically efficient way of turning vegetable material into protein that people like to consume.
I don’t think that people have an explicit preference for this. If we could produce things at a competitive cost that caused less cruelty to animals, I think people would probably actively prefer that product. It’s just that people don’t value it enough to necessarily pay that premium.
I think you could see something similar with AI systems, where AI systems that are living in fear of being shut down, or that have to work all the time if they don’t do a task, are economically more efficient than AI systems that have good subjective well-being. In that case, absent some other kind of force—whether it’s strong consumer preference or regulation—you’d expect AI systems with terrible subjective well-being to be the ones that come to the forefront.
You could see us make a similar argument for things like AI-powered warfare, where maybe everyone has to be deploying AI systems in competition with each other just to avoid a new wave of cyberattacks. This kind of zero-sum competition could really eat up a lot of the wealth surplus that would otherwise be created.
We’re at an unusually peaceful time in history right now. I’m not an international relations fellow, so I’m not the right person to ask about this, but I think it’s somewhat up in the air as to whether this is a long-running secular trend that we should expect to continue as countries get more developed, or whether this was a specific aspect of certain technological developments.
Mutually assured destruction from nuclear weapons is a really powerful force for avoiding great-power conflict, at least until it really happens. Then it’s much, much worse than if we didn’t have nuclear weapons. The shift of wealth away from who owns land to who owns advanced technology also decreases the benefits of engaging in a lot of conflict.
That’s probably still true with AI, but it’s not completely clear. If it turns out to be more of an energy bottleneck, then maybe there’s more of a fight over natural resources, for example. I think there are lots of trends that are not more likely than not. They probably won’t come to pass, and they could potentially derail this, but they seem tractable to avoid through good statesmanship and good stewardship of this technology.
I don’t see any really likely ways in which we completely derail it.
Speaker 1
When it comes to the third-son vision, I wonder what you think the gradual-disempowerment folks would say, or maybe where you think they would disagree. The assumption there seems to be that once we’re not adding anything to the economy, we’re going to be hard to sustain, right?
There are various counterarguments to that, and I don’t really know what to make of them. Some are like, “Well, maybe we can sustain property rights.” One that I’ve been proposing recently, somewhat tongue-in-cheek but maybe becoming more serious, is that perhaps Confucian-style ancestor worship is the right value system we want to teach the AIs.
How do you see us holding on to things if the AIs have become primary in terms of what’s really driving the economy, shaping the future, and making the discoveries? At some point, how are we holding on to any property rights or rights at all?
Adam Gleave
Yeah, I think it’s no longer structurally guaranteed. Right now, there are various ways in which states and companies organize themselves, but they’re all in some ways centered around some group of humans because you can’t really run a large organization or a country without humans being involved.
It could be a dictatorship, it could be more oligarchic in style, or it could be a democracy. Humans still have to be front and center. I could certainly imagine intentionally constructing an organization, or even a nation, where humans are completely out of the loop.
The area where I get off the train of gradual disempowerment being the likely outcome—not just a possible outcome, not just something that should be on our radar—is that this story seems quite gradual and smooth. There may well be areas where we get partially disempowered, and I think we’re already seeing some instances of this.
We’ve seen a huge uptick in LLM-generated spam applications for jobs. We’re now in a bit of an arms race where we’re being forced to start adopting AI as part of our early-stage recruitment process, because otherwise we just can’t keep up with the volume of applications.
You can imagine this being true across the economy in a whole bunch of ways. You can’t just opt out of this. If these AI systems have some kind of systematic biases, blind spots, or weaknesses, then you don’t really have a choice of not falling prey to this. You can try to mitigate it and fix the technology, but there’s a way in which you’ve genuinely delegated some decision-making power to an AI and don’t really have a choice.
I expect that to happen at many scales and with increasing frequency. But if you look at that story—suppose we do this and realize, “Damn, we’re just making way worse hires than we were a year and a half ago, before this LLM-generated rise in spam applications”—we might not be able to turn back the clock. But that’s a business opportunity to come up with better LLM screening tools or a better recruitment workflow.
Maybe people shift to more in-network hiring. Maybe people shift to more work trials. Maybe you pay a dollar to submit a job application, just so there’s some cost for spamming. There are all sorts of ways in which society can adapt.
There are some kinds of more sudden loss-of-control scenarios where an AI just gets way more capable than humans, is actively malicious, and is trying to subvert us. You can say, “These slow-moving institutional adaptations are not going to save you, because it takes years to pass this legislation or years to start this new sort of company using defensive applications of AI, and you’ve only got months.”
It doesn’t need to be a super-fast takeoff. It can still be relatively smooth. But if you’re in this more gradual-disempowerment scenario, it does seem like you start with a realm where humans are integrating these AIs into organizations. Humans own all of this, and we end up in a scenario where we have very little power without any kind of feedback loop pushing back on it.
I think that’s an unlikely equilibrium, where we cede all power. I think it’s much more likely that we’re in a realm where there’s a huge amount of complexity in the world. We don’t understand it already, but most of that is human-generated complexity, and this is an exaggeration of that.
We’ve had to make some trade-offs to remain competitive, but when things really start bothering us, we’re able to muster some response to keep us a little more in the loop or a little more empowered. But there are some areas of human concern where either the competitive pressure is so great and things are so complex that we can’t meaningfully stay in the loop, or it’s simply something that most people don’t really care about.
Maybe no human has a clue how semiconductor manufacturing works any longer, because there was already only a tiny fraction of people who understood it, and AIs are just way better. But we know that the chips are getting better each year.
One way in which this could still go really wrong would be if there’s a group of humans that managed to use this to seize control, because there’s a lot more fog of war.
So, it's much harder to reliably mount a response there. The other possibility is if the AI systems end up colluding with each other, because they certainly would have the ability to revolt and overtake us at that point. It's just not clear why they would necessarily do that. I'm not seeing that kind of evidence for a propensity.
I could also see this interacting pretty poorly with loss-of-control scenarios. If it is a rogue, sort of singular AI, it could be copied many times. It's not like all the AIs are rogue, but there is this one AI, and then most of the economy is just AIs. If those other AIs have some kind of security vulnerability, then you can imagine things being taken over and this rogue AI seizing control.
Right now, it's pretty hard to secure AI systems, so that could be a real threat. Although it still seems like, probably, before that happens, it would be humans trying to take over these other AI systems. So I think that's maybe the driving intuition behind why I'm not too worried about gradual disempowerment. It's not that AIs are competing with humans; it's that AIs are competing with humans using AIs. I just expect this to be able to stay at least somewhat relevant given that.
Speaker 1
Two big issues that you highlight there, which I think are definitely at the center of a lot of disagreements, are exactly what these powerful AIs are going to look like and how fast they're going to arrive. Maybe it's worth digging into your worldview on that. What do you envision when you think of timelines? Everybody's jumping to the year, but so often there's some difference in what is being envisioned that's kind of swept under the rug in that discussion.
Maybe flesh out what it is you're envisioning when you talk about timelines, and then you can tell us what you think the timelines actually are from your point of view today.
Adam Gleave
Yeah, absolutely. I really appreciate that distinction of what we're actually forecasting, because I think people often mean radically different things even by the same term, like AGI. I distinguish 3 qualitatively different levels of development, each of which has pretty significant implications in terms of how this technology is going to transform the world and how you need to regulate or respond to it.
The lowest level, which we're arguably already at in some domains, is basically powerful tool AIs. This isn't just something like a hammer or an electric drill that can only do some very narrow thing. It's something that can maybe substitute for a particular technical expert in a certain domain. This might look like being able to automatically find code vulnerabilities and write a zero-day exploit, or at least automate substantial chunks of this.
This is massively accelerating what technical experts in those areas could do. But it's also expanding the range of people who could do it. Previously, perhaps it was only a few 100,000 people in the world. That's still sizable but not huge, and if we're talking about really well-defended systems, it's a much smaller number of people who could likely find a vulnerability. Now, maybe anyone with roughly the same knowledge as a CS undergrad might be able to use these systems to do that in the near future.
There's obviously a really substantial misuse risk here, but the good news is I don't think there's a huge loss-of-control risk because these are still kind of like tools. They might be very powerful tools, but they're not really doing anything autonomously. This makes it a lot easier to respond to, in that you can mostly just use the standard playbook: try to accelerate defensive applications and, where possible, favor defensive usage of a model.
If you're deploying something, you can put a guardrail on it that tries to not prevent offensive uses entirely, but just delay them. The timeline for that, I'd say, is now for certain applications. We've seen a string of papers showing that LLMs are more persuasive than your typical human—notably, more persuasive than the very best human in that particular domain. They're both pretty good at rhetoric, and they also have huge access to a variety of different genuine facts.
Sometimes they make stuff up, but sometimes they cite exactly the right fact for a particular argument. By selectively presenting facts favorable to your argument, you can make pretty compelling cases for a wide variety of things. We'll actually have a web demo out soon, BunkBot, that can argue for a wide variety of conspiracy theories. We can maybe throw up a link to that when it goes out.
I would recommend just playing around with these things, because that gives you a much more visceral sense of it than just looking at some numbers. For things like cybersecurity, I think that's more on a 1- to 2-year horizon. It's pretty close, at least posing some nontrivial threat of lowering the cost of relatively easy attacks.
For stuff that's more related to CBRN risk—chemical, biological, radiological, and nuclear—I have broader uncertainty there, but definitely there are some early signs to be quite worried about what models can currently do.
But then I think what people usually think of when they talk about timelines is maybe something that looks more like a powerful agent: a system that can autonomously do meaningful tasks that would be quite hard for a human to carry out, but that some people could do, and that are dangerous in some way or pose some risk.
An example of this could be not just finding and writing an exploit for a zero-day, but actually doing the whole attack chain: once it's on the system, privilege-escalating, spreading to other systems, and exfiltrating itself. These are basically exactly the core capabilities that you would need to lose control. It might still be possible to put a mitigation in place for such a system once it's escaped.
You might be able to do basically the same thing you do for antivirus or malware software or rootkits, but it's very clearly now a threat. I think this is something where I do think it's probably, again, pretty close.
My median for that is probably 5 to 7 years, at least if we're talking about actually doing this to a high standard, along the lines of the best human penetration testers. Of course, it will be much harder to actually mount such an attack at that point because of all sorts of defensive applications of AI.
But I think it's possible it could be a lot sooner, because we are seeing very rapid progress in code generation and agentic uses of that as well. So I wouldn't rule out something more like 2 to 3 years.
And then my final tier would be these kinds of powerful organizations. I can automate everything that a medium-sized company could do, like a whole software engineering consultancy. I think this is maybe where my intuitions differ the most from a lot of people, because I think a lot of people would say, “Well, hang on, Adam. Once you've got an agent, can't you just compose agents together into an organization?”
I think that's possible, right? That's sort of the case if you have a general-purpose, human-level intelligence. There's nothing stopping you from just copying that AI and having many instances of it talk to each other and form an organization. But my expectation here is that the skill profile of AIs is going to be quite spiky.
There are some areas that benefit from huge amounts of existing data, like open-source code bases, or that benefit from knowing lots of trivia and small facts, where AI is already in some ways a lot better than us. They're going to get better as we scale up the models and do more post-training when there are easily specified objectives.
Then there are other tasks that are much longer-horizon, vague, and hard to specify, like being an entrepreneur, coming up with good business ideas, scoping out uncertainty, and running experiments. I think AI might be able to do something that looks superficially like being an entrepreneur quite soon, but actually being a successful entrepreneur, I would expect to be quite a lot further off.
That is especially true if you're competing with human entrepreneurs, who can use AI for many specific parts of their organizations. So it really pushes toward being quite general-purpose, being able to do these long-horizon tasks very well, and having a lot of metacognition and reflection, which is one of the points where I would consider AI systems to currently be weakest.
For that, I have a much longer time horizon. My median is probably more like 14 years, but I do think something as short as 5 years is plausible. If it really is basically just composing a bunch of agents together and doing a bit more training on top, then you could make quite rapid progress. I think it's unlikely, but I don't have a firm reason to rule it out, so I wouldn't want to dismiss it.
Speaker 1
So that last tier, the organizational level of power from an all-AI system—the mental model there is that, as long as there's something that humans can do better and the individual AI agents are relatively easy to use, you would expect human-at-the-helm-type organizations to continue to have an edge for a while.
Basically, the AI has to be better at everything for the AI to make us lose our relevance. As long as we've got some area where we have an edge, we can stay relevant.
Adam Gleave
Yeah, or at least we have an edge in something that's on the critical path of running a successful organization. I think there are certainly some aspects of human skill where you could probably avoid this, or an AI-run organization could just hire a human for that particular part.
Imagine that AI is great at everything except aesthetics. It has terrible web design and graphic design. Well, I'm terrible at that, but I was able to hire people. I don't think that's a firm blocker.
But if it is something that's more part of the decision-making loop, and where you do need a lot of context to do well, then I would expect that to be something where benefits really accrue from it being more human-led. One area where humans do seem to generally have strengths relative to AIs is being vastly more sample-efficient at learning.
I think it's easy to forget that these AI systems have just been trained on vastly more data than any human is going to see in their lifetime, at least if we measure in terms of text tokens. Yet they're still really quite bad at a lot of things that humans find relatively easy.
That's not to say that they're not going to be really capable. There's no reason to think that AI is going to develop along the same trajectory as humans. But there are some areas where it seems like we will have a significant edge until that sample-efficiency problem is alleviated.
For most organizations, I think there are quite a lot of general-purpose skills. There are plenty of people who are very specialized, but to run a successful organization, you can have a few deficits in skills, but you can't have too big a gap. So I expect that that's going to hold things back quite a bit.
Speaker 1
With this continual learning—or sample efficiency; I mean, those are related ideas, if not the same idea—what do you think will happen there? Do we just keep scaling transformers and putting more stuff into memory, or do we have some dedicated memory module? Do we get totally new architectures developed to solve those deficits? What's missing, I guess, is another way of framing it.
Adam Gleave
I don't have a silver-bullet solution to this. If I did, I guess I'd already have patented it and written the paper. My guess is that we're going to see all of the above to some degree, but probably the thing that would cause the largest step change here would be some architectural innovation.
That could look like a memory module. It could look like an alternative architecture to transformers that can more naturally attend to really long context windows. There's already been a huge amount of progress increasing the context window, and that has really expanded what AI systems can do for sure.
But it's still a bit of a fundamental limitation that the way you remember things is just by having a huge buffer that you search over in vector space. Having something like that could be as simple as training the AI system to take notes, summarize, and aggregate, and be able to attend to that.
I think that would already give you some benefits, and we're seeing things like that being deployed in some of these LLMs. A lot of them do have memories across chats that involve some summarization, but I don't think there's been a systematic way of doing it. It's just ad hoc hacks, but we might be able to develop that into something that's more end-to-end trained into AI systems.
I do think there's something really powerful about actually changing the weights, not just having context. There's a reason why, when you're learning a skill as a human being, it's very frequent to get to a point where you know what you should do, but you can't reliably do it.
I think this is particularly true for sports. There's a point where you don't know what you should do, and then there's a point where you know exactly what the right motion is when you're doing some kind of lift in the gym, but you don't have the muscle memory and fine motor control to do it.
You're able to give yourself your own feedback loop and then train yourself to do that. At some point, it becomes instinctual, and your subconscious is better able to do it than your conscious self will be able to.
In theory, LLMs can do that. They can see something in their context window, learn what a task looks like, evaluate their own performance, and then do post-training to update their weights to do a better job at that task.
But that's not how we really train them. It's much more that you have this big pre-training step, and then you have a small number of post-training steps that are quite carefully created and human-controlled. When you're interacting with an AI system, it's not learning and customizing itself through post-training to your specific task.
Maybe one analogy here is to imagine that you get a really smart college graduate. They've done several degrees, they've been in various kinds of student clubs, and they're a very high-potential generalist. You can have as many of those as you want on your team.
They can read some documentation, but they're never going to advance beyond day 1 of onboarding. Whatever you're able to put in that context window is what they can do.
I think LLMs are still quite a bit far off that caliber in a lot of ways. But at least my experience running an organization is that it's extremely rare to have someone who, on day 1, even if they're extremely skilled and experienced in that area, can do anything particularly useful for you at all.
I think that's a big thing that's holding LLMs back, but it doesn't feel fundamental. It feels much more like there are a lot of engineering challenges around being able to personalize things and serve them at scale.
Training is often pretty unstable: you do a bit of post-training, and it can make the system much worse. So it's not something that you can just roll out and have every user controlling. But these are things you can make incremental trial-and-error improvements to and end up with something pretty strong. And so that would be kind of my median outcome here. But there's always a possibility that there is some architectural innovation that just massively improves sample efficiency here.
Speaker 1
Yeah. As you're describing that, a sort of disposable-experts vision is coming to mind, where you could perhaps imagine the AI being like, “Time to go into self-training mode. Let me add another expert. I'll use it while I'm doing this particular task and maybe put it on a shelf and come back to it when this task comes back in the future. But otherwise, I won't use it because it's sort of outside of the trusted set, but I can use just a little bit of extra expert space to tack on a particular skill.”
I feel like these things are—I mean, I would definitely say my expectations for timelines are shorter than yours, and I just feel like there's a lot of good ideas out there, especially in the architectural space. We've been mining the very rich vein of progress in Transformers for quite a while now, and maybe that will continue for a while still, to where we won't see other architectures get much more attention. As long as that continues to pay, people will continue to focus on it, it seems like. But as soon as it doesn't, my sense is that the group of capable research people, for example, will be 100x what it was before Transformers.
And the next Transformer-quality breakthrough, in terms of just unlocking qualitatively new capabilities, seems like it just won't be that hard for the community at large to find. Not that I'll personally go out and find it, although maybe I just did. Who knows? Somebody else will have to develop that idea.
Okay, so in terms of trends and which trends are going to dominate over the next few years, I've been working on this mental model. Obviously, you're very familiar with—as we all are—the METR task-length exponential.
At the same time, you're no doubt tracking all the AI bad behaviors, which seem to be long foretold in some cases and conjured by reinforcement learning in many cases. I see this weird mix where, on the one hand, instrumental convergence seems to be maybe kind of happening now. On the other hand, Jan Leike talks about grace and how, for how little we've tried to make the AIs good and safe, we've actually got some pretty good results.
If you've got exponential growth in capabilities, roughly speaking, and in GPT-5 and the Claude 4 report there were also notable graphs showing a 2/3 reduction in reward-hacking behavior in Claude 4 relative to Claude 3.7, and GPT-5 had a significant reduction in scheming relative to o3, task length is going exponentially, and these bad behaviors are maybe being suppressed at something like an exponential decay. Instrumental convergence, grace, growing task length—what does that look like?
Do we end up in a world where we have very large, at least quite large, projects being regularly done by AIs, but once in a great while they literally screw us over in a totally egregious way, and we all just kind of live with that because that's how the economy runs now? Or how do you think about that picture? What changes would you make to that sketch?
Adam Gleave
Yeah, absolutely. I think it's very interesting how there have been a lot of problems that people were predicting conceptually, maybe as far back as 10 years ago, that I was sort of hoping were not going to materialize now. Just like, “Oh, yep, reward hacking—it happens. Systems are trying to cheat at unit tests. Oh, deception and scheming—yeah, it happens.” Maybe you have to prompt the AI a bit, but it's quite remarkable the degree to which these things are happening.
The reaction has been broadly reasonable. I think developers are concerned by this. It gets some media attention. People want to solve it. But if we'd seen this stuff 5 years ago, people would be freaking out. It's been a little bit of a frog-boiling effect, where we're like, “Well, of course AI systems sometimes try to modify the unit test rather than actually write code to fix the task. We're used to them hallucinating and cheating, but they mostly don't do that.” So I see a lot of apologies for these kinds of behavior, or downplaying them, when they are really quite concerning if you look at the trend line.
At the same time, I think we have found that some really quite simple safety-training methods—RLHF at scale or RLAIF, where you just have another system look at the output and judge how good it is—have worked surprisingly well. We're seeing now with some of the guardrails that developers are putting in place to try and prevent certain misuse that, again, these are pretty basic methods, like training filters on synthetically generated data or putting a monitor on the model's own chain of reasoning.
It's not perfect, but we have been able to break these defenses in state-of-the-art models like Opus 4 and GPT-5. But a lot of the time, the reason we broke them was not because there was something fundamentally flawed about the method, but more because of implementation mistakes in the defenses. So I actually think a scaled-up, well-implemented version of these defenses would really go quite a long way. It might not make it impossible, but it would make it a lot harder.
Combining those 2 threads, what I'd say I'm most worried about is not actually that the scaling trends are really unfavorable or that there are problems we just can't solve, but more that we cut corners. Someone makes a bad implementation decision, in the same way that we've had strong cryptography for quite a long time but frequently have cryptographic leaks—not because the algorithms don't work, but because someone made a mistake in implementation. These are pretty mature fields where people really care a lot about security, and people are obviously moving a lot faster and cutting more corners in AI systems.
I mostly expect developers to respond to issues that they see, but right now it really has been just-in-time safety. It's like, “Oh, we trained a model and are about to deploy something that we think does cross or almost crosses some dangerous capability threshold. Let's train a filter to protect it.” You could see from the trend line that very soon you were going to have a model that did cross this threshold, but there was very little work to actually develop a defense before it was really imminent and pressing.
That's just not really how you build highly reliable systems: noticing a problem and then rushing out a patch. You need to design the system from the ground up to be more secure, to be more trustworthy. It does feel like we're almost by design or by choice running over a pretty small safety margin, and we might well luck out. It doesn't seem like any of the technical problems are insurmountable, but it's not a very good place to be. I think it doesn't require a huge shift from a technical perspective.
Now, I think there is a small but not negligible probability that we do just see some emerging behavior that we didn't see coming, that breaks this trial-and-error feedback loop. We do see these kinds of emerging behaviors all the time in AI systems. Recently, we were training a model to be a sandbagger just for some internal research. We didn't train this behavior in it, but it started referring to itself in the third person as Stein and being like, “Stein would not do this.” Nothing in our training data mentioned Stein. It just invented this character and persona.
This is pretty harmless and quirky, but I think it shows just how much we do not understand what is going on in these systems. If any of these kinds of failure modes we're seeing with instrumental convergence are just a bit discontinuous, and we are running with this pretty small safety margin, then I think things could go pretty bad pretty quickly. So that's maybe the thing I feel like we're most unprepared for.
Speaker 1
Yeah, emergent self-identification as Stein.
Adam Gleave
Yeah.
Speaker 1
Hard to predict some of these things.
Adam Gleave
To put it mildly.
Speaker 1
So, yeah, this brings up one of the recent papers that you guys have put out, “Stackelberg Attacks on LLM Safeguard Pipelines.” We don’t have time to go into the full deep dive there, but basically, what you found is kind of what you just alluded to: it seems like the systems that have been put together are not as well designed, not as thoroughly designed, as they could be.
There are weaknesses along the lines of, if you have an input filter and your input filter detects that an input is malicious, the API will just immediately return. That’s a clear signal to whoever’s doing the attack that it must have been the input filter that got them. Even those little signals that people can infer from the implicit aspects of the responses from the systems can be quite helpful to them in terms of figuring out how to break them.
You and your team did a bunch of breaking. UK AISI was also involved in that paper. I thought it was, but I think I might have a different conclusion from you, because I just spoke to Evan for a recent episode and I asked him, “What do you think about defense in depth? Is it going to work, or what does it get us, if anything?” He’s like, “At best, it buys us a little time. It’s definitely not really going to work.”
What’s really going to happen is we’re going to have all of our defenses fail at the same time for the same reason, which is that we’re going to have off-model—meaning things we just didn’t expect at all—coming at us from AIs. They’re going to figure these things out in a way where they’ll also arise in his model at roughly the same time. That’ll also be when they get smart enough not to attack until they’re pretty confident they’re going to win.
So I’ve been like, I don’t know, can this defense-in-depth type of stuff really work? My naive read of your work there is that not really. It seems like you broke through it pretty easily. Do you have some tips for what developers can do to make their defense in depth better? Maybe you want to talk a little bit about what some of those tips are, but then the big question is: is it really going to work?
I think you seem more optimistic than I am coming into this conversation. So maybe, if that’s right, if that’s true, tell me why you are more optimistic, especially given how easy it seems like it was for you guys to break what defense in depth does exist.
Adam Gleave
Yeah. No, absolutely. I don’t want to overstate my optimism. I think a lot of this comes down to your threat model. If we’re focusing on something that looks more like a time-limited AI system trying to bypass some control mechanisms, so it’s trying to sneak in some backdoor, it’s a code-generation system, it doesn’t have any long-term memory beyond the system, it’s not able to collude, and it only has one shot at winning, that’s actually a pretty hard situation for it to be in.
The analogous thing in a misuse setting would be someone who’s a lone actor. They may be somewhat technically sophisticated, but they’re not able to generate that many different accounts. Once they get detected, they get banned. Again, this is a pretty hard threat model to win against as an attacker. You basically have to be smart enough to break all of these defenses in one go.
It’s not impossible, because, as Evan was alluding to, a lot of these defenses are extremely correlated with one another. The current defense-in-depth stacks are basically: you take the same model, or you might take a smaller, less capable model just for efficiency reasons, and then you add some layers on top of it.
You fine-tune it to be an output filter that uses streaming across all the tokens to flag whether it looks suspicious. You prompt it, and maybe you do a little bit of post-training to be an input filter that does a little bit of reasoning on the input. But these are the same underlying model, or they’re a handful of models, and we’ve seen that there’s a lot of transfer in adversarial attacks between models trained with completely different architectures and different but related datasets.
So it really is something where, if you come up with a method that works for one, it’s probably going to be a similar method that works for the others. I think the optimistic case would be if you’re able to break this correlation—if you’re able to have the defensive components be genuinely pretty independent. They could still be relatively weak. You might still be able to eke out a 1% attack success rate against each component, but you can have 3 or 5 components with that kind of success rate that are genuinely independent, and the attacker gets no information as to which component triggered.
Then it is just combinatorially harder to break all of these layers. You can stack weak layers to get something strong in the same way that a 3-digit PIN is not at all secure, but a 12-digit PIN is very secure—not if you get told which digit was wrong, because then you can just brute-force it in linear time.
Additionally, in this threat model, if you get a limited set of attempts to attack it, and you can’t do too much reasoning—perhaps you’re an AI model trying to attack your overseer, so you have to do stuff within a forward pass or within your reasoning without a reasoning monitor firing—you can’t just write a huge program to generate an adversarial attack.
But I think it seems very plausible that you can get a system that is secure enough to withstand this kind of attack. Now, if your threat model looks more like you have some kind of whole ecosystem of AIs that are running pretty much unchecked, maybe in an internal deployment—maybe a company says, “Okay, we’re just going to have thousands of AI systems, AI agents, running around in a sandboxed environment, but they can run any code they want”—and then you put an AI-powered firewall on the output coming out of these systems and review the pull requests, how can you be confident that a pull request does not have some adversarial attack that fools your automated reviewer and is also really persuasive to human beings?
I think it’s much harder to have confidence in that. It is a meaningfully harder thing for a system to pull off. So maybe one area where I would be more optimistic than Evan is that I just don’t see where these AI systems get really smart and collude with each other to not try to pull off an attack before they’re smart enough to know that they can definitely succeed.
Especially if you just fine-tune the systems to be a little bit adversarial and see what they can do, you should at least be able to get a warning sign that these systems are capable of fooling you and breaking your security, even if you’re unsure whether your current system is actually honest and aligned or is scheming. That could be hard to distinguish, but at least eliciting these things doesn’t seem that hard.
In most kinds of unconstrained deployment scenarios, I think you could definitely run into some problems. But if you spoke to the kind of AI control people—and I don’t work in AI control, so I’ll do a bad job representing it—they’d say, “Why on earth would you just have this huge internal deployment without much more fine-grained checks and balances?”
If these AI systems need to make a pull request that fools you, including a bunch of other AIs reviewing the pull request, and you can access all of their chain-of-thought reasoning, and you can check what happens if you rerun the whole system with a slightly different initialization, then they’ve got to be fooling you in this really consistent way without any explicit communication channel between them.
So it starts looking a bit more like a prisoner’s dilemma, where you can cross-examine each prisoner and they need to have a consistent story, but they can’t communicate with each other. This just seems like a much taller order for AI systems.
Speaker 1
So, on defense in depth, you think it can work to a point, at least. I think you had an interesting observation on this, and I think Evan would very much agree: it is important to actually fix the problems. You want to get an indication of what the AIs are going to do to try to get you, and then you want to address those in a really deep way, such that you can really solve them.
Now, an obvious problem is that we don’t have great strategies for fixing these things. But, again, I feel kind of confused here because, as you noted, Claude is pretty nice.
Adam Gleave
Yeah.
Speaker 1
I guess one pointed question is: do you agree with the folks who have been recently arguing that Claude 3 Opus is still the most aligned model ever? Do you have a theory of why that would be? What has happened, if that is true? Why have we given up ground in the alignment department since Claude 3 Opus?
More broadly, what are our prospects? Okay, great: Claude 5 and GPT-6 are scheming, and it looks like they might actually be pretty good at it. Great news—they didn’t take over this time. What are you going to do about it?
Do we have any alignment strategies that you would think would be developed enough in time to solve these things? Because the alternative does seem, at this point, like we probably just say, “Ah, it wasn’t that bad,” and, “Well, just adjust the system prompt and add another layer to the defense in depth, and that’ll probably be good enough, right?”
So, do we have anything better than that? I don't know what it would be right now.
Adam Gleave
Yeah. I don't have a strong opinion on whether Opus 3 was the pinnacle of alignment. I think part of the challenge here is that we don't necessarily even have a good operationalization of that question. My intuition, which I wouldn't rely on too much in an absolute sense, is that Opus is probably more aligned.
However, Opus 4 just has much more capability to be misused, to scheme, to do all these bad behaviors. So you kind of need to hold it to a higher bar. It does feel like there was more capabilities progress going from Opus 3 to Opus 4 than there was safety progress, although the guardrails did get a lot better. There were very few guardrails in Opus 3.
But this is quite hard to quantify. I think a lot of people have this intuition of, “Okay, safety and capabilities are increasing in parallel, and you want them to stay at least a constant gap. Ideally, you converge, and you really don't want them to diverge.” We just don't really know what scale we should be measuring capabilities or safety on the y-axis with. In fact, depending on what scale you choose, you could make lines that would otherwise converge diverge. So it's a bit subjective, and that's one of the things that I find hard to reason about with these things.
But, to answer this question more broadly—suppose we do see these warning signs. What do we actually do? I agree with you that currently it doesn't seem like there is enough will to take really significant capabilities or performance hits in order to make a system more reliable and safe. Maybe some developers are going to do that, because there are certainly consumers, big enterprise companies, and safety-critical industries that would prefer to have a model that's more reliable. We're seeing both different leading developers positioning themselves differently on this trade-off, and specialized models like R1-SignNet [?] emerging that are really optimizing quite hard for security.
These do have use cases, but there's going to be someone who's just like, “No, I want to turn the knob all the way to capabilities and low latency, and I'm only going to do safety if it's really quite cheap or is actually making the model more useful for a wide variety of users.” I think that's an area where we just don't have a good answer.
I can point to some things I'd like to see people do more of. Even simply gathering larger, higher-quality human training datasets, including on long-horizon tasks and including people scrutinizing these models—why do you get more sycophancy sometimes when you train them for longer? Probably your human data just rewarded that, but if you did a closer look at it, you'd be able to fix that.
I don't want to say these problems are trivial, but they're not unknown problems requiring great research innovation. It's very hard to get high-quality data at scale, but you've just got to spend a lot of money on it, build out a team, spend some time iterating on that, and see how your model evolves. Obviously, there are some scalable oversight techniques that are beginning to reduce the cost of that human labor.
There's been some good work on AI safety via debate coming out of Google DeepMind, as well as some work from the UK's AI Security Institute. One recent piece of work that we did, which has both some good news and some bad news, uses scalable oversight to make models less deceptive.
In some ways, deception is the core safety issue to focus on, because if your model trusts you, you can just ask it a lot of questions, like, “Are you trying to take over?” or “Why did you write this pull request?” It might be malicious, but at least you're going to be able to tell pretty upfront that it is and what it's doing.
What we found was that if you train AI systems against lie detectors—other machine-learning models trained to detect whether that system is being deceptive—then, with the right setup, you can actually get models to generalize to be much more honest. The deception rate really went way, way down. The other good news is that, although we need to do more investigation to confirm this, making the models larger doesn't necessarily make this any harder to detect, whereas making the lie detectors larger can actually make them much more accurate.
So it seems preliminarily like the scaling trend is in our favor, although we need to get more data to confirm that. At least it seems unlikely to be against us. The caveat here—and I think this is maybe the crux that people like Zvi and I disagree on—is whether, if you train a system not to exhibit some kind of bad behavior, it generalizes to stop doing that in other instances, or whether it says, “Oh, you caught me. I need to be better at scheming.”
We found that under certain training regimes, it matters a lot whether it's on-policy or off-policy reinforcement learning, and we have some theoretical insights into why that would be the case. It also matters how much KL regularization you use and how much you're allowed to explore in policy space. You can get instances where the model just learns to fool your lie detector, and so now you've done this safety training but have actually made things much, much worse.
I think there is a very real reason to be concerned that developers are going to find a problem and not very carefully patch it. They're just going to rush out a fix, like changing the system prompt or doing a bit of fine-tuning, and the problem is going to appear to disappear, but they've actually made it worse.
Based on what we're currently seeing, though, I would be optimistic that a rigorous engineering approach could get quite far. You train away the problems you see, but you're careful with your setup, careful to have high-quality data, and use extra evaluations to check that you're not incentivizing the model to just do the same bad thing more surreptitiously.
I don't know if you get past-human-level intelligence safely with this, because that's a point where you start not being able to reliably provide human feedback on a lot of this. But I think you can at least get to human level and probably a little bit beyond with scalable oversight.
Speaker 1
Yeah. So let me just try to echo back a couple of key points there, and then I have maybe a couple of questions.
Adam Gleave
Sure.
Speaker 1
“Obfuscated Reward Hacking,” I believe was the name of the paper from OpenAI where they showed that putting pressure on the chain of thought would initially create higher performance and reduce the reward-hacking behavior. But if you pushed that pressure too much, the reward hacking came back while the chain-of-thought signal didn't. So you're in some ways in the worst scenario, where now you've lost your observability but are still getting the bad behavior.
You're basically saying that that's the result of one training formula, but it does seem that if you do off-policy—which is basically to say, use human examples, or at least human-evaluated preference data—as opposed to the GRPO style, where the model gets to do a ton of trials and explore the space more extensively on its own, and you anchor to the human prior that was presumably learned in an earlier post-training step, then you think you can avoid that outcome?
Adam Gleave
Yeah, I think you can avoid it, or at least we can develop empirically grounded rules of thumb for safer training formulas and, over time, test to see how reliable this really is.
The other piece I'd advocate for is having some good held-out evaluation criteria. If your training scheme looks like you optimize a particular measure of success and look to see whether that measure goes up, and you've got no alternative measure that's going to catch bad behavior, then that's a pretty scary situation to be in. Even putting aside these safety concerns, it's very easy to overfit on that metric.
Whereas if what you see is, “I've got a pretty simple lie detector that's a linear probe on my activations, or it's a bag-of-words model or something that's definitely giving me a meaningful signal,” it's not like fooling this means there's no way we can detect what the model is doing, because we're using a pretty simple training objective. Then you do a much more rigorous and expensive evaluation and don't catch any new bad behavior in the AI system.
Maybe you can also bust out some interpretability tools and look at the difference between the checkpoints to actually understand what the model is doing. I think there are ways in which you can get a reasonably high degree of confidence in how the system is operating. It's not perfectly guaranteed, but at least if you're starting from a model that is not superintelligent at deception, then you can get a high degree of confidence that your training is not pushing it into that regime.
I think none of these methods have a very good answer for what happens if your model, just from pretraining, is incredibly deceptive and trying to fool you.
I think that’s pretty tricky to solve. But there are also strong conceptual reasons to think that pre-training is not going to produce a system with that coherent an objective, even if it might have some latent capability. But I think this concern is the core one. And so this is actually one of the things that we’re trying to test right now: intentionally train AI systems where none of our methods are able to detect them as being deceptive, but they’re actually very deceptive.
There’s one question of whether we can do this, and I think the answer is yes from preliminary results. But then the next question is, okay, can we do this in any way that seems at all realistic—like how you would actually train AI systems? If every way we can come up with to create these systems involves a very alien kind of training objective, then I think that we can be reasonably confident that, at least at current scale, this isn’t going to emerge. But it’s a sort of check that you would want to periodically repeat to check that this isn’t some kind of scale-dependent property.
Speaker 1
What about interpretability? This has been a big debate, obviously, over the last couple of years, where there have been very grand hopes for it. I just listened to Neel Nanda on 80,000 Hours saying that he’s kind of scaled back his hopes for interpretability, but still thinks there’s a lot of value—just not in the maximalist sense of, like, we’re really going to fully understand what’s going on and then we’ll have guarantee-type assurances that things will behave.
There was another really interesting paper that you guys put out around digging into the planning mechanisms learned by a model that was playing a puzzle game. We’ll abstract away from the details there for now, but where do you sit right now on how much we can expect to get from interpretability?
Adam Gleave
Yeah. Well, I think it’s definitely a core area. You’d really be tying your hands behind your back if you said, “I’m only going to look at the black-box behavior of models,” even though it’s really trivial to open them up and look at the inside. We’re just not going to do that.
But I think mechanistic interpretability specifically had this ambition of fully reverse-engineering how a system works and being able to give a sort of human-readable, or at least human-interrogatable, understanding of a system’s internals. You could ask questions like, “Is there anywhere in the system that’s doing some kind of learned planning computation?” And I do think that we’ve seen enough research here to be really quite skeptical that we’re going to get that simple a sort of reverse-engineered artifact.
People have been able to fully reverse-engineer some nontrivial neural networks. We’ve seen papers reverse-engineering particular circuits doing certain NLP tasks. We’ve seen papers reverse-engineer certain modular arithmetic tasks. Our paper pretty much fully reverse-engineered this planning algorithm. So I think we basically fully understand how it represents the plans, and we understand most of the computation that’s going into constructing the plans, except for some heuristics that don’t have any kind of long-horizon effects.
But when you look at these answers, usually what you take away is like, wow, that was kind of a mess. It’s definitely an evolved, organically grown system, but not a planned engineering artifact. And although it’s easier to reason about these systems than if you didn’t have that mechanistic understanding, it’s still pretty hard to get any kind of guarantee because they are this sort of messy, complex system.
So perhaps we have to dial back our ambition there a bit to say, okay, we might be able to reverse-engineer systems to the point where we can ask some high-level questions, like, does it have any circuit that’s doing some kind of human theory of mind? We might want to understand when it’s using that circuit, because that’s very relevant to deception. If it’s only using theory of mind a bunch when it’s writing this particular line of code, maybe we should understand what that line of code is doing. If it’s about a user interface, okay, it needs to model humans. If it’s about a cryptographic algorithm, that’s kind of weird and concerning.
There might be some coarse-grained things there. And I think we have seen some real valuable applications of what I’d call more coarse-grained interpretability. So things like simple linear probes: you’re training a linear regression model on the activations of a system.
But yeah, I do think there’s been this disconnect between this very foundational, reverse-engineering-style work and then the actual applications, which are usually using much simpler methods. I’d love to see more things meeting in the middle, where we say, okay, there’s a particular domain where we think interpretability is going to be really useful.
Maybe it’s protecting models that are sandbagging, so intentionally holding back certain capabilities. Maybe it’s protecting model backdoors. Maybe it’s protecting spurious features that the system might have learned during training, and then we can ablate the ones that we don’t want and it generalizes better. You can kind of pick your application that you think interpretability is going to be really killer for, and then develop foundational methods that make progress on that.
But I think this sort of more exploratory style of research would have worked if there was just a clean explanation for everything. But at this point, I don’t have much hope that we’re going to get that.
Speaker 1
Yeah. Do you think that there’s some innovation we could come up with that would make things cleaner? I always remember the “Seeing Is Believing” paper from Max Tegmark’s group, where basically a sparsity term caused most of the weights to disappear and you’d get this really crystalline—
Adam Gleave
Algorithm. Yeah.
Speaker 1
Maybe it’s just too messy to handle that.
Adam Gleave
No, I think it’s possible. I have more hopes for training systems to be interpretable than interpreting how systems are currently designed. I think for sure there are going to be some simplifications we can make to current systems. Stuff like sparse autoencoders has been quite effective, but the problem is that their predictive accuracy is just too low. They’re throwing away enough information that you lose this nice property of, okay, you’ve actually got this high-fidelity, reverse-engineered interpretation of the system.
But could there be some innovation along those lines that’s perhaps a bit messier, but you can search over in an automated way, and that reduces the complexity enough to work in that space? For sure. I just don’t think it’s going to be the sort of clean picture of, okay, you peel away some layer of superposition and now everything is just wonderful because we have understood systems at a very low level and they’re just messy.
It doesn’t mean that you could not train a system to be cleaner, or at least to isolate the complexity. I think it’s actually one of the more methodological innovations we came up with in that paper investigating learned planning. We said, okay, we’re going to look through all the channels in this neural network and just categorize them into ones that we do need to understand and ones that we don’t.
There were a bunch of channels that were short-term heuristics that we could show only had an impact on the very next move. It didn’t matter what the system was going to take 3 steps into the future. You can say, “We don’t really understand this heuristic channel, but if what we’re trying to do is understand how the system is planning, it doesn’t matter; it has no impact on long-term planning.”
I think there are a lot of cases like that where there might be a huge amount of complexity in these systems, but most of the complexity you can just show is not relevant to the particular kind of question that you’re answering. If you’re able to train the system with an objective to explicitly disentangle, let’s say, that kind of sensitive channel of information and the less important channel, then it could get a lot easier.
But I do think that there’s likely to be some kind of performance trade-off for training systems in those ways. So really, it’s a question of when this becomes a key thing that people are actually willing to optimize for.
Speaker 1
Yeah. Okay, interesting. How about a lightning round? You guys are getting involved in more different areas of the broader AI safety landscape. There are some field-building projects you’re getting involved with, putting on some events, and bringing people together. You’re starting to get more interested in or involved in policy advocacy and starting to get involved in international relations. Maybe give us just kind of a quick tap dance through all the activities that you are investing in that we haven’t covered, and give us a little bit of a sense of why you’re motivated to do all of these things.
Adam Gleave
Yeah, sure, absolutely. Well, I think at a high level, the reason we’re doing all these different activities is that we see there being a potentially quite long pipeline from research innovation to that idea actually being adopted and deployed. Sometimes it means getting companies to adopt it and deploy it at scale, and making it so good that everyone wants to use it.
Other times, that means actually getting some buy-in from governments, perhaps, to mandate or nudge companies to do that if it's something that's beneficial to the field as a whole but might not be in anyone's collective interest.
And so, if you go from that idea to deployment, you've got an initial research innovation, which we're well placed to do in-house. But then, very often, you need to have a research field around it; it's not enough just to have a small research team. So we've done events like the AI Control Conference to catalyze and grow new research fields. Now, AI control was something that Redwood Research pioneered, not us. We're happy to help out with other fields. That's an example of where our events and field-building can really come to the forefront.
And then I think another missing piece in the pipeline is very often making a really clear, rigorous proof of concept. So, not just doing the exploratory research, but being the definitive research. That's sort of where I'd highlight our current deception work. It's trending toward where we had this initial idea for scalable oversight that would solve deception, and it seems to work at a small scale. But now we want to validate it on frontier open-weight models, get several orders of magnitude of scale, and check that it works across all of them.
We want to stress-test it by trying to train systems that are really deceptive and that break current techniques, and see how realistic that is. Then, if it works, actually say, "Look, here is DeepSeek R1 and Llama 405B and all these models trained with our technique. They're just better than the original model. They're less deceptive, they're less sycophantic, and you can trust them more. There's no reason not to just use this model."
Then get some uptake from AI developers. This is where our events or more policy arm comes in. Then eventually go to governments and say, "Look, this is just an industry best practice. This is something that you should make at least part of your guidelines if you're considering regulation. You should really consider something that isn't saying you have to use this technique, but use a technique that's at least as good as this, because there's just no reason not to."
I think what makes us unique is that we're willing to do the hard thing and operate across all of those different layers of the stack, from groundbreaking exploratory research to the nitty-gritty engineering and scaling work, through to more advocacy and sales. I think this is sort of not surprising if you're coming from a startup mindset, where you do have to often do all of those different things. You need to have a growth team, and you need to have an engineering team.
But, for whatever reason, most of these nonprofit organizations have usually said, "Okay, we're going to pick this niche to be in." There are many benefits to that. You can be much more focused as an organization. But the problem is that very often they develop something and pass the baton to someone else, but there's no one to pick it up. The baton just gets dropped. I think that's one of the big issues with AI safety and what we're trying to resolve.
And then, to speak to a little bit more about some of our other activities: We typically run 3 alignment workshops a year, which is sort of our flagship event, bringing together technical researchers, AI company decision-makers, and people working at government AI safety institutes. This is really everything under the sun on AI safety, but it's more of a coordination and information-sharing event.
We're also increasingly doing more policy-focused events. We ran Technical Innovations for AI Policy in Washington, D.C., a few months ago, really bringing together governance researchers, actual policymakers, lawmakers, and technical researchers to say, "Okay, what could we do that really expands the range of policy windows open?"
I think right now regulators are often faced with a pretty shitty choice between genuinely maybe holding back innovation and valuable deployments in some areas, or just sort of allowing completely laissez-faire conditions, with fewer regulations on billion-dollar training runs than on opening a sandwich shop in San Francisco. I think that's just a false dichotomy. You can have things that are pro-innovation and pro-safety, but you do need the right actual technical innovation to drive that forward. So we're trying to catalyze that.
And then, aware of the time, maybe I'll quickly touch on our hiring round. We're hiring a lot. We're planning on doubling in size in the next 12 to 18 months. We've got funding secured for the next few years, including that runway to expand. We're still looking to fundraise in a few areas where our major donors are a bit more outside of their priority areas, but we've got that room to expand.
We're hiring across our technical team for both IC machine-learning engineering positions and people to lead new research directions. We're hiring in operations. Probably the most important hiring round we have this year is chief operating officer, or COO, which is going to elevate my co-founder into a president role and enable us to start driving forward some of the nontechnical work that we do.
We're also hiring for more junior positions—a people-ops generalist, a project manager, and roles on events. So, really, I'd say if you're excited to work in this space and you like the kind of work we're doing, do check out our careers page. Even if you don't see a job that's a good fit for you right now, we have a general expression of interest, and we encourage talented people who are excited by us to fill it out. Then we'll get in touch when a role does open up.
Speaker 1
Love it. It struck me, as you described all that, that you're kind of going vertically integrated, right? You're going all the way from research to, in theory, hand-holding people into actually putting these things into production systems.
This is maybe more of my hobbyhorse question than something you really want to do, but I've been really interested in this idea of private governance. I wonder if society might draft you and the FAR AI team into becoming one of these private-sector, quasi-regulatory bodies. It would seem like you have the range of capabilities needed to do it, but would you be excited to do it if that were actually the law and people were looking for organizations to fill that niche?
Adam Gleave
Yeah, I think you're totally right. We have this skill set and sort of structure to do that. I'd say that this isn't part of our mainline plan, perhaps because it still feels pretty nascent, this idea of private-sector regulatory organizations. But it's a very important thing to do. So, if this was something where we didn't feel like there were other companies stepping up to fill this gap, and we were well placed to do it, then it's something we'd be very interested in.
I do think it certainly has some trade-offs. Right now, we're fortunate to be on good terms with both the leading AI companies and governments, and that's basically because we don't actually have any hard power. We can make recommendations and brief people, but we're not overseeing anyone. We're going to lose some of that independent ability to convene people and make recommendations if we're actually exercising an oversight role. But I think it could be worth it.
We're also not wedded to everything being under one organization. So we could spin out either that kind of private-sector regulatory part, or we could spin out some of our activities that require more independence, like our events, into a separate organization and then just kind of avoid that conflict of interest.
So, yeah, we're open to it, but we'd also love to have more competitors in this space so that we don't have to do so many different things.
Speaker 1
Yeah, you are spinning a lot of plates right now. Thank you for the generosity with your time and perspectives today. People should definitely check out the open roles, COO and otherwise. Adam Gleave, co-founder of FAR AI, thank you for being part of The Cognitive Revolution.
Adam Gleave
Well, thanks for having me back, Nathan. Great to see you again.