[BidClub_]
The Cognitive Revolution · · 87 min

Helen Toner: OpenAI Reflections, Adaptation Buffers, and AI in Warfare

Erik TorenbergNathan LabenzHelen Toner

YouTube
TL;DR
  • Toner’s base case is not imminent superintelligence; it is that a civilization-scale transition is plausible enough to prepare for now. In 2016, “short timelines” meant advanced AI within a couple of decades or a lifetime; by 2025, it can mean superintelligence before the 2020s end, so even “long timelines to advanced AI have gotten crazy short.” For investors, the discussion points toward sustained demand for evaluation, resilience, and government capacity—not confidence in a five-year countdown.
  • The clearest new OpenAI disclosure is that Q* did not trigger the board’s 2023 decision. Toner said the board knew reasoning research later released as o1 and o3 was underway, but received no breakthrough letter and did not act on one: “That whole Reuters story was totally false.” The broader governance discount remains harder to quantify because confidentiality and legal obligations prevent a complete public record.
  • Frontier-AI oversight works better when whistleblowers can point to a violated rule than when protection depends on subjective concern. Toner favors disclosed safety-and-security plans, capability and risk evaluations, and internal processes that create a crisp standard employees can invoke. Those employees may have their greatest leverage now because they are “actively working to replace themselves,” while institutional failure may arrive “boiling frog style” without one obvious crisis.
  • Adaptation buffers are the strategic answer to an AI market where frontier development gets costlier while yesterday’s frontier rapidly commoditizes. DeepSeek matched reasoning capabilities only one or two months behind some US systems—or base-model capabilities roughly six to nine months behind—at lower cost, making permanent nonproliferation increasingly invasive and brittle. The practical response shifts toward vaccine capacity, outbreak detection, cyber remediation, and distribution of defenses before capabilities diffuse.
  • Iterative deployment remains useful only while releases are unlikely to cause severe, irreversible harm. Toner prefers conditional “if-then” gates—do not advance until specified understanding or mitigation exists—over fixed pauses or input-based speed limits that would require unavailable legislation and immediately encounter the China objection. With no comprehensive regime in sight, transparency, measurement science, interpretability, alignment work, and technical government staffing are the practical building blocks.
  • “Beat China” is functioning simultaneously as a geopolitical argument and the AI industry’s path of least resistance in Washington. Toner grounds the rivalry in power transitions, maritime access, and international rules; the conversation also covers Taiwan. She says the framing conveniently supports “funding,” “government contracts,” “no regulation,” and protection from liability. That makes China rhetoric material to AI company economics even when it does not resolve the underlying strategic question.
  • Military AI should be evaluated use case by use case, not sold as a general-purpose battle buddy. Bounded tools for viewshed mapping, medical triage, ship-movement anomalies, or database retrieval differ fundamentally from an LLM asked to generate “three courses of action that are non-escalatory.” Toner and Amelia Probasco’s framework—scope, training data, and human-machine interaction—puts reliability and adversarial robustness ahead of demo fluency.
  • An “AlphaGo for the army” is not a credible near-term equilibrium because real war cannot be enclosed in a clean simulation. Battlefield, logistics, economic, political, and public-attitude dynamics interact while an adversary deliberately attacks the model’s assumptions. Toner’s honest conclusion is that she does not know what nation-states, democracy, or the Chinese Communist Party look like under superintelligence; strategic uncertainty, not a settled doctrine, is the central fact.
Digest · the substance, structured for research

1. Toner took AGI seriously before it was respectable

  • Toner joined OpenAI’s board in 2021 but had known the company and many of its people since its 2015–2016 founding, when she was beginning AI-policy work in San Francisco. Because AlexNet had arrived in 2012 and deep learning was already advancing, she initially felt “late to the party”; the wider attention waves of 2018–2019 and ChatGPT in 2022 later changed that perspective.

  • Founding an organization explicitly to build AGI was then “weird” and “against the grain,” even within machine-learning circles; DeepMind was nearly the only serious research organization speaking that way. Toner believes her China and national-security expertise mattered to her board appointment, but so did having taken OpenAI’s mission seriously for years when the informed community was extremely small.

  • Her original timelines were short only by the standards of that era: very advanced systems within “the next couple decades” or “in our lifetime” seemed plausible enough to justify preparation. Today, short timelines can mean superintelligence before the 2020s end; she remains much less certain about that, while holding that the possibility still “warrants quite a lot of thought, quite a lot of preparation.”

  • Toner rejected the idea that preventing AI from killing humanity motivated her career. Her premise was historical: major technologies transform society “for better or for worse, often for better and for worse,” and AI looked likely to produce such a transformation within her lifetime. The question was whether her work could help it go better—not whether it would “definitely” kill humanity.

2. The Q* story was false, but OpenAI context remains constrained

  • Toner’s continuing limits are concrete: board confidentiality obligations, legal processes in which small inconsistencies could matter, and private conversations involving people who need not be drawn into public controversy. Much of the unreleased material is “almost like boring detail,” not a remaining “big deep dark secret” that would dramatically change the picture she gave on the TED AI Show.

  • One formerly confidential point is now clear because the underlying reasoning work is public. The board knew research later released as o1 and o3 was underway, but “we never got some letter about a breakthrough” and did not base its decision on an employee letter. Toner’s categorical correction: “That whole Reuters story was totally false.”

  • Erik pressed on whether OpenAI’s world-changing mission had become a heroic or “main character” culture shaped by elite-performance coaching and detachment. Toner declined to generalize: board members are poorly positioned to characterize day-to-day employee culture. Her broader concern was structural—technical work normally can be separated from social governance, but that separation becomes “out of distribution” if progress outruns society and only developers can avert enormous consequences.

3. OpenAI’s technical warnings and policy posture pull apart

  • Nathan described a recurring OpenAI whiplash: the obfuscated-reward-hacking paper offered one of the clearest warnings about AI failure, while a White House policy submission soon projected an escalatory China posture and sought sweeping freedom to train on copyrighted material. He paired that with Sam Altman’s “our values or their values; there’s no third way,” calling the institution’s outward voices almost “schizophrenic.”

  • Toner also finds the contrast confusing and thinks it has become more pronounced. One possible explanation is substantial employee freedom over public comments and research directions, combined with different review paths for official policy. She hopes technical staff are watching what the company advocates politically, given how strongly those messages can contradict the warnings emerging from its own research.

  • Her disclosure priority is to narrow the “huge information gap” between frontier developers and everyone else. Useful releases include capability and risk test results, descriptions of safety processes, and model specifications explaining what systems are trained to do. The objective is not a government checklist dictating the answer, but enough visibility for outside actors to understand the systems and respond.

  • The US reporting threshold around 10^26 operations appeared “on a wobbly footing,” though Toner had not heard that it was definitively dead; Nathan said the EU AI Act was developing transparency rules around models over roughly 10^25, amid pressure to dilute them. Toner’s rationale is not that every model above a compute line is dangerous—it is that the newest, best models carry the highest “unknown unknown risks” and merit extra scrutiny.

4. Whistleblowers need enforceable standards before crisis

  • Conventional whistleblowing usually concerns illegality: the SEC, for example, offers a defined channel for financial misconduct. Frontier-AI employees may instead observe conduct that feels dishonest or dangerously risky but violates no law. A protection framed merely as “if you’re worried, call this hotline” leaves workers and companies without a usable boundary.

  • Toner’s preferred design pairs protections with disclosure or process requirements. A company might publish or privately submit a safety-and-security plan, creating a standard against which employees can report that an evaluation was skipped, information was misstated, or a promised mitigation was abandoned. Even a required internal process helps if workers can truthfully say, “Actually, we didn’t carry out that process.”

  • Implementation must account for the proposed system’s users: technical employees who may lack legal sophistication, feel frightened, work extreme hours, and have little time to navigate ambiguity. “If the user is a whistleblower,” Toner asked, what is the user experience? Eligibility, the next step, confidentiality, and the reporting destination need to be obvious.

  • Frontier employees may be “in the most powerful position they’re going to be in” because they are actively automating their own work while tech labor’s leverage is already weakening. Waiting for a dramatic rupture may therefore fail twice: workers could become less indispensable, and misconduct could accumulate “boiling frog style”—each episode concerning, but none obviously the single moment to act.

5. Adaptation buffers beat permanent nonproliferation

  • Toner begins from confidence in social adaptability. New technologies from the printing press to television and the telephone repeatedly produced claims that “the sky is falling,” yet societies built institutions, norms, and barriers that made them positive on balance. AI as a possible successor intelligence may be different, but she resists treating every misuse risk as historically unprecedented.

  • AI development follows two curves at once. Pushing the frontier requires ever more compute, money, and concentrated expertise, making the first demonstration less accessible. Once a capability exists, however, engineering improvements drive its cost and difficulty down repeatedly—creating a temporary interval in which society has observed the capability but broad misuse remains comparatively hard.

  • DeepSeek illustrated the interval rather than leapfrogging the frontier: Toner described it as matching reasoning systems from one or two months earlier, or base models from roughly six to nine months earlier, at a lower price. The exact reported development cost was less important than the directional fact that frontier-like performance becomes cheaper and easier to reproduce.

  • Permanent AI nonproliferation would therefore demand escalating intrusion. Nuclear controls work because bombs require substantial quantities of highly enriched, specialized material. If nuclear efficiency improved like AI—until the uranium dispersed across “a couple acres” of farmland sufficed for a tiny weapon—the IAEA would eventually need to inspect farmhouses. Toner sees that as an analogy for why locking down broadly useful computation becomes untenable.

6. Resilience depends more on deployment than frontier capability

  • Adaptation means using the buffer to reduce consequences: expand vaccine manufacturing, wastewater monitoring, outbreak detection, and test-kit distribution so a biological attack is recognized and answered faster. It also includes ordinary counterterrorism questions—whom the FBI tracks and how early it detects plots—which have little to do with frontier models or even biological materials.

  • In cybersecurity, Toner challenged the narrow question of whether AI helps attackers or defenders more. The operational question is whether usable defensive products reach water-treatment facilities, power grids, and chemical plants run by teams without frontier-AI expertise. A brilliant model at a lab does not protect infrastructure until its capabilities are packaged, disseminated, and integrated.

  • Nathan offered AI-assisted formal software verification as a hopeful example, reporting a possible multiple-orders-of-magnitude speedup. Toner accepted the potential but emphasized timing: a released model can reach a disaffected teenager far faster than thousands of infrastructure operators can modernize old codebases, manage the division between IT and operational technology, and safely deploy new defenses.

  • Even if defenders ultimately benefit more, the transition can remain dangerous. “They’re not going to be nimble,” Toner said of infrastructure providers, so long-run defensive superiority does not erase the release-to-adoption lag. Thinking of it as a “transition period rather than like a permanent state of danger” produces better policy, but does not justify complacency during that interval.

7. Today’s policy menu is building blocks, not a regime

  • Toner has not seen a comprehensive frontier-risk regime she likes, partly because policymakers do not know exactly what problem will emerge or when. The realistic menu is a set of building blocks that improves future options—an appropriate response to uncertainty, but potentially inadequate if progress leaves “very little time.”

  • Dean Ball’s proposed regulatory market is one such component: government-accredited private regulators would assess developers, whose compliance could earn a liability shield; accreditation or protection could be withdrawn. Toner considered it “certainly better than nothing” and “far far far more feasible” at state level, while doubting it could counter the “cutthroat financial incentives” and other pressures pushing frontier companies to move as fast as possible.

  • Other building blocks include public funding for the science of measuring AI, interpretability and alignment research, focused research organizations, and greater technical capacity inside government. Transparency does not itself solve a failure, but it equips more actors to respond as problems emerge. Her preferred slowing mechanism is similarly conditional: an “if-then” gate tied to understanding or risk mitigation, not an arbitrary number of months.

8. Iterative deployment works only before irreversible harm

  • Nathan defended OpenAI’s original iterative-deployment logic: exposing society gradually to improving systems should be less disruptive than developing superintelligence privately and unveiling it at once. He worried that Ilya’s new company planned no release before superintelligence, GPT-4.5 might be removed if its compute cost outweighed demand, and Miles Brundage’s comments suggested internal deployments were becoming strategically more important.

  • His proposed relative speed limit would tie internal development to public deployment: a company could train only a specified multiple beyond the largest model it had released, forcing some visibility before a small group pursued systems it believed might transform global power. He framed it as a safeguard against abandoning iterative learning precisely when internal capabilities become most consequential.

  • Toner agreed with iterative deployment in principle, provided each release is unlikely to cause “really severe irreversible consequences.” Its value comes from releasing, observing, adjusting, and trying again; that logic breaks when the experiment cannot be recalled. Any serious iterative policy therefore needs criteria identifying when the next iteration should not enter the world.

  • She was less convinced an industry-wide retreat had already occurred and skeptical that Nathan’s speed limit was implementable. It would likely require federal legislation that Congress will not pass, followed immediately by “doesn’t it just mean that we lose to China?” Conceptually plausible controls still fail without a political mechanism; she again preferred progress conditional on demonstrated understanding and mitigation.

9. A crisis could unlock Congress—and trigger an overreaction

  • Toner agreed that many lawmakers simply do not believe superintelligence forecasts, but added that institutional dysfunction is independent of AI. A deeply experienced congressional observer told her the House may be “less functional than it’s been since after the Civil War.” Even after a technological shock changes beliefs, thoughtful federal legislation would remain a major lift.

  • Some policy specialists are prewriting a “Patriot Act for AI,” expecting a crisis to create a short legislative window. Toner considered advance preparation reasonable while warning that the eventual bill could fight the last war or smuggle through unrelated priorities. Three Mile Island supplies the counterexample: disproportionate regulation after one incident effectively shut down US nuclear development without comparing its risks to alternative energy sources.

  • Developers therefore have a self-interested reason to install credible guardrails before an accident; otherwise, Toner sees the system heading toward a “massive knee-jerk overreaction.” The trigger may be vivid rather than statistically severe: Kevin Roose’s Sydney exchange drove fear after GPT-4 despite being heavily elicited, and unrestricted celebrity voice cloning could likewise provoke backlash disproportionate to its underlying risk.

10. “Beat China” is both a geopolitical claim and a lobbying shortcut

  • Toner started from power-transition logic: the US is an established power, China a rising one, and control over international rules matters. Peaceful transitions such as the US eclipsing Britain are historically unusual. The postwar Pax Americana replaced much “might makes right” behavior with sovereignty, stable borders, and institutions supporting trade and freedom of navigation.

  • US policy tried through the 1980s, 1990s, and 2000s to make China a “responsible stakeholder,” despite the rupture of Tiananmen Square, culminating in WTO membership. Toner argued that this project had clearly deteriorated by Xi’s rise in 2012 as China became more illiberal and hostile in the South China Sea. Nathan separately raised questions about Taiwan and possible Chinese aggression around Japanese, Filipino, and Korean assets.

  • Nathan’s pushback—worth keeping—was that American CEOs had shifted from warning against a China race to demanding an unassailable US lead, while Chinese labs were openly releasing models and Americans discussed preserving a unipolar world. Toner replied that openness makes strategic sense for a follower seeking attention and talent; it is not evidence of “pure goodwill and lack of competitive spirit.”

  • Toner nevertheless found the CEOs’ rhetorical reversal “pretty striking.” Her explanation was political economy: “We got to beat China” is the one message Washington can agree on, making it the “path of least resistance” for companies seeking funding, government contracts, freedom from regulation, and liability protection. The US retreat toward its own might-makes-right posture further complicates the moral clarity of the rivalry.

11. Military AI already spans bounded automation and broad judgment

  • Toner’s paper with Amelia “Emmy” Probasco draws on unusual operational experience: Probasco served in the Navy operating Aegis missile-defense systems. Developed without deep learning in the 1960s and 1970s and deployed in the 1980s, Aegis already detects incoming missiles and can automatically identify and engage threats—evidence that meaningful weapons automation predates today’s AI debate.

  • Their focus is decision support because military-AI discussion too often stops at autonomous weapons. The category ranges from computer vision calculating viewsheds—where a sniper or other operator can see—to battlefield medical triage and anomaly detection in ship movements. These are bounded functions that can reduce the fog of war without claiming general strategic judgment.

  • At the expansive end, companies including Palantir and Scale AI have marketed LLM-based systems as something closer to an all-purpose “battle buddy.” The demos combine well-grounded retrieval—locating the nearest unit or checking its missiles against a database—with open-ended reasoning whose reliability and provenance are far less clear.

12. A battle buddy is only as safe as its scope, data, and interface

  • Toner’s sharpest example was a demo asking the model to “generate three courses of action that are non-escalatory.” The system appeared to draft tactics and routes under that political constraint, prompting her question: “How the hell does the LLM know what’s escalatory and non-escalatory?” The polished interface concealed unresolved questions about who defines, tests, and validates such judgment.

  • The first evaluation dimension is scope: a tightly bounded activity can be tested against its intended operating conditions, while a sprawling assistant responds to whatever an operator thinks to type. Generality increases the surface over which plausible language may be mistaken for operational competence.

  • The second is data: what trained the system, and how closely does that data match the present conflict? Military environments are both novel and adversarial. An opponent is actively trying to deceive sensors, corrupt assumptions, or induce behavior outside the training distribution, making ordinary benchmark reliability an inadequate proxy.

  • The third is human-machine interaction: the product must communicate what it can and cannot do, prevent overtrust, and help its operator reach a sound decision under pressure. Military organizations have often undervalued interface design, Toner argued, even though poor presentation and misunderstood automation have contributed to unwanted outcomes.

13. War is not a game clean enough for AlphaGo

  • Nathan feared an “AlphaGo for the army”: self-play and simulation could hill-climb toward superhuman tactics operating faster than human commanders, producing inscrutable, hyperlethal systems. Toner thought this was not close, because a credible simulation must connect individual battles to theater logistics, global asset deployment, economics, politics, public attitudes, and countless unknown interactions.

  • Dan Hendrycks, Alexandr Wang of Scale AI, Eric Schmidt, and co-authors offered “mutually assured AI malfunction”: a power threatening to build superintelligence and rule indefinitely would invite rivals to sabotage its project. Toner found the multiparty logic useful, including the possibility that vulnerable projects change what actors attempt, but not comparable to the stability of nuclear mutual assured destruction.

  • Nuclear deterrence was legible: states understood what weapons did, how second-strike capability worked, and why no one wanted an exchange. Superintelligence’s form and strategic value remain unclear. Sabotage may be easier than securely building an AI project, but that uncertainty cannot produce the “crystal clear strategic logic” of Cold War deterrence.

  • David Chapman’s distinction between abstract rationality and practical “reasonableness” supplied Toner’s deeper objection. Chess, Go, and StarCraft are not clean originals that reality approximates; they are “very unusual special cases” carved from a messy world where people act, observe, and readjust. Her final answer remained an honest non-answer: she does not know what nation-states, democracy, or the Chinese Communist Party become in a superintelligent world.

Erik Torenberg

Today I'm speaking with Helen Toner, director of strategy and foundational research grants at CSET, the Center for Security and Emerging Technology, and author of a new Substack called Rising Tide. Helen is best known to the general public for her role as an OpenAI board member in the decision to temporarily fire Sam Altman in late 2023, but she's been thinking about AI—or at least the need for society to invest in preparation for the possibility of transformative AI—since way back in 2016, when she started working on AI policy full-time.

That's a full 5 years before joining the OpenAI board in 2021, when, it's worth noting, OpenAI had already launched GPT-3 as an API product, taken $1 billion in investment from Microsoft, and was increasingly recognized by those in the know as a leader in the generative AI wave. Certainly, by that time, OpenAI had plenty of access to super-talented candidates for its board. With that context in mind, and remembering that her blog is called Rising Tide, despite what you might have heard elsewhere, it probably should not surprise you to learn that Helen is definitely not an AI doomer, or even especially hawkish on most AI safety issues.

On the contrary, she argues in early posts on her Substack that nonproliferation is the wrong approach to AI misuse and instead promotes the concept of adaptation buffers: the notion that society broadly has a critical window of opportunity to adapt to new AI capabilities between the time when they're first demonstrated, typically at high cost in terms of both R&D and compute, and when they later become widely accessible, typically at much lower cost, as we've recently seen with companies like DeepSeek dropping the cost of frontier reasoning capabilities.

While her focus today is on other things, I couldn't resist asking Helen some OpenAI-related questions, and I appreciate her willingness to engage despite having addressed these issues in multiple forums already, including especially an episode of The TED AI Show, which we'll link to in the show notes. The only truly new detail that you'll hear in this conversation is her assertion that media reports suggesting that some sort of Q* breakthrough in reasoning had led to the board's decision were, quote unquote, “totally false.”

Nevertheless, I think it's important that Helen and other former OpenAI team members continue to speak candidly about their experiences with the company and its leadership. As Helen notes in another of her first blog posts, everyone's timelines are dramatically shorter than they used to be. What passes for long timelines in AI circles today would have been quite short not many years ago.

And given this new short-timeline consensus, the recent AI 2027 scenario from former OpenAI researcher Daniel Kokotajlo and his team reflects not just one of the shorter-timeline forecasts, but, if I'm reading between the lines effectively, a warning about how OpenAI leadership might fail to act responsibly around the time of AGI by abandoning its principle of iterative deployment, keeping the best models for its own internal use, plus maybe that of the U.S. government, and aiming for a sort of AI takeoff via the automation of AI research.

That's a warning, by the way, that's become a bit more credible this week with the news that OpenAI has indeed announced that GPT-4.5 will be deprecated from the API. All that's enough for me to feel strongly that it's important for Helen to use appearances like this to continue to remind Washington decision-makers that OpenAI's CEO was not consistently candid with its board.

It's also enough for me to applaud moves like the amicus brief recently filed in the Elon Musk v. OpenAI lawsuit by 12 former OpenAI team members, who argue that nonprofit promises were central to OpenAI's early hiring success and that the nonprofit should not cede control of the company at any price. That development happened after I recorded with Helen, and I hope to do a full episode on it soon.

Of course, the stakes are only rising from here. With OpenAI and other AI companies seeking Pentagon contracts and special legal protections, Helen's latest research out of CSET with Rhodes Scholar and former Navy Aegis operator Amelia Probasco on AI for military decision-making is super important: a sober attempt to map out how AI systems have been and are likely to be used, and how that may diverge from how they actually should be used given their current limitations.

Among many other interesting details, I was amazed to learn that some nations, including, most prominently, Russia, currently have published military doctrines about AI that seem to be fundamentally out of touch with current AI systems' lack of reliability and total lack of adversarial robustness. This, too, is something that Washington decision-makers probably can't be reminded of often enough as they seek to develop autonomous killer robots.

While I believe that there's probably some nontrivial and irreducible risk associated with developing advanced AI at all, it's my sense that much of the extreme AI risk we face today in fact exists because key decision-makers, under intense and growing pressure, seem fairly likely to make some very bad mistakes. If this show can do anything to contribute to a positive future, I hope that it can help people start thinking about those critical but avoidable failure modes sooner and better, so that we can minimize the extreme downside risk and get to live in that age of AI-provided abundance that we've been promised.

Looking back on OpenAI and looking ahead to adaptation buffers and military use cases, all amidst shorter and shorter timelines to AGI, this is Helen Toner from the Center for Security and Emerging Technology and author of the new blog Rising Tide. Helen Toner, director of strategy and foundational research grants at CSET, the Center for Security and Emerging Technology, and author of a new Substack, Rising Tide. Welcome to The Cognitive Revolution.

Helen Toner

Thanks. Great to be here.

Erik Torenberg

I'm excited for this conversation. We have a lot of ground to cover. I think we all have crosses to bear in this life, and one of yours is that you're going to go on and do a ton of things in the AI space, and yet people are always going to come back and ask you questions about your tenure on the board of OpenAI.

Of course, everybody is at least somewhat familiar with how that ended. I'm not going to be an exception to that entirely, but I do want to make sure we have time for a bunch of different things.

To set the stage, one question I don't know the answer to at all, and I'm really curious about, is how did you get involved with OpenAI in the first place? This goes back years, to a time when there was no powerful AI. Most people would dismiss the notion as fanciful, and very few people were taking the whole topic seriously in any real way, but you obviously were.

Maybe share your backstory with respect to AI and some of the enthusiasm that you must have had to get into that position in the first place.

Helen Toner

Absolutely. I joined the board in 2021, but I had been familiar with the company and with many of the folks working there since it was founded. They were set up in San Francisco around 2015 or 2016. At that point, I was working in San Francisco, and that was right around when I was starting to work on AI issues.

It's really interesting reflecting on that time. It felt like being behind the game because, by 2012, you had AlexNet, and the deep-learning revolution was really in full swing by 2015 or 2016. By the time I came around to the view that this was going to be a really big deal, that there was a lot of work to be done here, and that I wanted to make AI, policy, and national security a real focus of my work, it felt like I was coming late to the party.

But it's been fun to see a couple more waves since then. Around 2018 or 2019, I want to say, people started paying a little more attention, and then obviously ChatGPT in 2022 brought this huge new burst of interest. Looking back, I no longer feel like I was as late to the party as I felt at the time.

I think it's also easy to underestimate how weird and against the grain it was for them to found a company to build AGI at the time. That was really not the kind of thing that you talked about in polite society, including impolite machine-learning society. Google DeepMind—or just DeepMind at the time—was sort of the only game in town among serious researchers who were talking about AGI.

When I look back on why I was invited to join the board, I think part of it was being in the AI policy space and having a few years of experience at a time when not many people did, as well as having spent time in China and having that sort of China expertise and national security expertise, which I think was valuable for the board.

I think it was also the fact that I had taken the idea of AGI and their mission seriously for multiple years by the time I joined the board. That was really unusual. It's sort of funny to look back on that now, in 2025, because obviously AGI is on everyone's lips and OpenAI is such a famous company that the situation looks kind of different.

But when the community was really small, the set of people who had been thinking about these topics and who had actually informed perspectives was really small. It was super interesting to get to be familiar with the company from its very early stages.

Erik Torenberg

Were you always a relatively short-timelines person? One of the blog posts whose draft you shared with me reminds everybody that, even what people are now passing off as long timelines, are actually quite short.

Yeah, but where were you 5 years ago in terms of your expectations?

Helen Toner

Yeah. No, I wasn't. I still don't know if I am. The standards have changed so much. And hopefully, the post you're talking about—the current draft title is “Long Timelines to Advanced AI Have Gotten Crazy Short”—will be published by the time this comes out.

When I got into the space, I think I counted as having short timelines for the time. As I described in the post, it was: This seems plausible; it seems likely enough to be worth preparing for that we build very advanced systems in the next couple of decades, in our lifetime. This seems like a potential development, and if it happens, it would need a ton of societal preparation. Almost no one is thinking about it, so that would be worthwhile to spend time on.

I think that did describe my view when I got into the space. Nowadays, I don't really identify as having short timelines, because that means expecting superintelligence before the 2020s are out or something like that, and I feel much more uncertain about that. I still tend to fall back to this view of: Look, I think this is all likely enough that it warrants quite a lot of thought and quite a lot of preparation, which is different from saying, “I think it's very likely to happen,” or “I think it's very likely to happen in the next 5 years or next 3 years.” So, yeah, I guess the question is, depending on what your standards are for short timelines, maybe yes, maybe no.

Erik Torenberg

Yeah, I used to make a very similar argument to people when the whole notion of powerful AI was fanciful, and certainly any notion of safety concerns related to AI was doubly fanciful. I used to just say, “We have a small number of people that scan space to try to find asteroids so that we don't get taken out like the dinosaurs did, and that seems really good. This seems like another thing.” And now it definitely feels like there is an asteroid, and it's coming at us. We don't know if it's good or bad, but it's definitely going to be both.

Helen Toner

Yeah. I mean, I never— to me, the underlying internal motivation to work on this space was related to the way that, if you look across the scope of history, huge new technologies tend to really change what society looks like, for better or for worse—often for better and for worse.

And so I was really coming to believe in the early to mid-2010s, “Okay, this looks like we're going to go through one of these transformations, probably in my lifetime, and that's going to be a really huge deal.” If I'm interested in trying to leave the world a better place than when I found it, to the extent that I can, then maybe this is an area to go work in and shape.

So to me, it was never, “Oh, this is definitely going to kill us, and so I have to get into the space to prevent it from killing us.” It was much more this broader argument: It seems pretty likely we're going to go through this massive transformation. Can I get into a line of work that can help contribute to that going better?

Erik Torenberg

That mindset point is really interesting, and I want to ask you about the prevailing mindsets at OpenAI from your perspective. Before getting into that, I know you've given a couple of different interviews about this and spoken about different aspects of it, and you're probably tired of it, understandably. What governs what you can and can't say at this point?

We've all seen the non-disparagement clauses that were then nullified, and I don't think you ever had any equity in the company. Maybe you did, but I don't think so, right? So how much is external constraint on what you can say, and how much is just you deciding how much you really want to talk about this?

Helen Toner

Yeah, it's 2 big factors at this point, or maybe 3, depending on how you count. As a board member, I'm under ongoing confidentiality obligations that don't apply, for example, to former employees. So just in terms of conversations on the board and topics that we're discussing, I want to respect those obligations. I take them very seriously.

And then there's also ongoing legal processes where my statements could be compared against each other for consistency, and minor discrepancies could cause problems. I might need to testify under oath or otherwise say things.

And then I think there's also just a range of other stuff: confidential conversations that I've had in confidence with people. There are just so many details and so much you need to go back and explain, so much context, and bring in other people who really don't need to be dragged into this. The payoff wouldn't even be that great, because none of it— I gave an interview, the best interview I've been able to give on this, the most detail I've been able to go into, was on The TED AI Show last year. It's not that there are hidden secrets that are more shocking than what was there.

There's just a lot more almost boring detail, but that brings in a lot of stuff that's maybe confidential or maybe involves other people who don't need to be dragged in. So I don't think the payoff is there. Certainly, if people are like, “Oh, man, there's still this big, deep, dark secret that Helen still hasn't spat out,” that's not the case. There are still reasons that I'm not just sharing everything totally publicly that I just talked about, but I think even if I were able to, it wouldn't change the overall picture in any dramatic way.

Erik Torenberg

Yeah, context is that which is scarce, as you say.

Helen Toner

And actually, maybe this is a good point to share one example of something that I didn't comment on because I wanted to respect my obligations to the company around confidentiality. This was the rumor at the time about Q* contributing to the board's decision, which was totally false.

I hadn't really commented on it because I didn't want to get out ahead of OpenAI's reasoning work and what has now been released as o1 and o3. The board was aware that that research was underway, but we never got some letter about a breakthrough. We didn't make our decision based on a letter from employees. That whole Reuters story was totally false.

So that's one example of something that's now slightly easier to talk about because the underlying confidential information around that line of research is now out in the open.

Erik Torenberg

Okay. So, about the mindset or the motivations: You said wanting to leave the world a better place was part of what motivated you to get involved, and just recognizing the stakes and feeling like this is a high-leverage activity. That seems to me to be a big part of how I understand what I think is motivating people at OpenAI in general.

In a way, that's good, right? Everybody should want to make a positive difference. But I do sometimes worry that it can cross over into a sort of main-character mindset or a hero mentality. Especially as I hear more and more things about guru-style coaching going on, the practice of detachment, and an elite-performance mindset, I'm not sure if it's maybe gone too far.

Part of me is like, maybe we should want some amount of attachment among the people who are developing these potential superhuman AI systems. So much of this is hearsay. I don't even really know how pervasive some of these ideas are, but there are definitely quite a few data points at this point. Would you say that's a prevailing sense at the company, that there's this heroic, world-changing quest they're on, or would you put that as a more minority position that's just occasionally popping up?

Helen Toner

I think board members are generally not in the best position to talk about the culture of a company because they're not immersed in the employees the way that a team member would be. So I honestly don't know that I have a perspective on that that you wouldn't have. The perspective I do have comes from knowing people who work there. And I know you know people who work there as well, so I don't feel like I have more to add necessarily.

Erik Torenberg

Okay, fair. How do I get at that? It does seem like a very important question. I do have this sense that detachment in frontier AI development seems somehow wrong. It feels like we're taking, to put it in machine-learning terms, something that was learned out of domain.

The mindset that I think I'm seeing is kind of like what they tell NBA 3-point shooters to adopt: Just keep shooting, don't worry, trust the process over the outcome, and que sera, sera. I don't know that that generalizes super well.

Helen Toner

Yeah. I mean, these are high-stakes technology environments. I think the other version of it generalizing would be just this separation that we sort of implicitly have in society more broadly between the people building new technologies and doing scientific work and the people who are figuring out how to apply them and how to regulate them.

And I think it's generally pretty reasonable for someone who's figuring out how to make an airplane wing be shaped slightly more efficiently to not be thinking about how the FAA should regulate this, or how airline seats should be priced, or things like that. I think it often does make sense to decouple those more technical and more societal questions.

So to me, that's more what the out-of-distribution element is: this might be a technology where, if technical progress outpaces society's ability to adapt, then the people who have just been doing that sort of decoupled technical work might end up having these huge societal consequences that only they, or almost only they, were able to affect or prevent. So to me, that's the way that I would think about the OOD element here.

Nathan Labenz

Yeah, that's interesting.

Another thing is that, actually, I think the last time we spoke, I was still participating in the GPT-4 red team. You were still on the board, and it's been quite a journey since then. I've definitely watched the company very closely, and I feel like I've been on this roller coaster where many times I've been disappointed and, at times, even outright scared by what I'm seeing. At other times, I'm like, “Well, that's dramatically reassuring.” I'd say there have probably been 10 episodes, and it might be 5 to 5.

The most recent 2 would be the publication of the “Obfuscated Reward Hacking” paper, which I would put up there in the pantheon of the most important and clearly stated warnings about how AI can go wrong if it's not developed with the utmost care. And then, at the same time—maybe the same week, or within 72 hours of that—there was the response to the White House request for comment on what AI policy should be.

There we got a rather—I would say—escalatory vibe, certainly with respect to China. We've had Altman say things like, “It's our values or their values; there's no third way,” which seems, again, like ruling out a lot of possibility space quite prematurely. And then also asking for things like, “Cancel all property rights so we can just train on everything,” because, again, if we don't do that, China will. It seems that there's this almost schizophrenic nature to the company, and I wonder how you understand that.

Helen Toner

Yeah, I find it confusing as well, and it seems like it has gotten more pronounced over the last year or 2. One explanation could just be that there is, as far as I know, a relatively decent amount of freedom given to employees to tweet as they choose, pursue some research directions, and maybe write about those research directions.

I think different things go through different processes, but, again, I haven't been close enough to those on-the-ground decisions about what gets published when to have any kind of insider perspective on that. But I think I agree with you that it's striking how different some of the voices from inside the company seem to be. And I wonder—I don't know—I hope that the more technical folks there are paying attention to the kind of policy messages being sent out by the company, given how much they contradict each other.

Nathan Labenz

How would you advise people who are there today? I mean, if you are inside—and this could generalize beyond OpenAI—I don't see any reason to think that xAI won't have similar issues, and potentially other companies that are generally held in high esteem for their safety practices very well could, too. It's all coming at us pretty fast.

So if you are somebody inside a company and you're concerned about what you're seeing, what should those people be thinking about? And maybe the flip side of that, or complementary question, is: What should policymakers be thinking about in terms of protecting whistleblowers or facilitating whistleblowing?

I also respect the idea that these companies should be able to keep some trade secrets. But then it also seems like the level of secrecy that was requested of me at one time was too much. It was like the public, at some point, does need to know what capabilities exist. So I don't really have a great sense of how to find that line.

But maybe let's start with the policymakers: What do you think the rules should be? Then we can go into, if you're in a position where maybe the rules aren't there yet, how should one, as an individual, think about taking responsibility? So, rules specifically around whistleblowing. You can go broader than that if you want, but I'm definitely interested in how we get things that the public really needs to know to come to light when company policy says it's a secret.

Helen Toner

Yes, yes. I mean, I think the whistleblowing piece does connect to other parts of the policy picture. I won't try to give a comprehensive view on policy right now, but, starting from the whistleblowing piece and expanding out, the way whistleblowing usually works is that it's for illegal behavior. The SEC has very clear processes: If you're seeing financial misconduct, you can go talk to them, and lots of other whistleblowing processes are similar.

So I think, for policymakers, a big challenge I would want them to have in mind here is that a lot of the concerns we're talking about, or potential concerns, are behavior that is actually not illegal. So where is the line for when you can whistleblow? What kind of behavior should be protected?

I think the best way to do whistleblower protections is to pair them with some kind of disclosure, or some kind of rules around what information needs to be shared, because then it creates a clear standard for when the company is either not sharing information it's supposed to share, or being misleading or inaccurate in the information it's sharing.

That's structurally a simpler way for whistleblowing to work, as opposed to trying to have this vague standard of, “If you're worried, call this hotline,” because that's just so squishy and hard both for employees and for the company, as you say, trying to think about trade secrets or other reasons that they don't want their employees just blabbing.

It's much more helpful to have a clear standard to compare against. So that's one thing I would say on the policy side: If you compare it with some kind of expectations or requirements around information sharing, or around processes that you have to carry out internally even if you don't share the results, then that leaves employees able to say, “Oh, actually, we didn't carry out that process.”

So, to be a little more concrete, a version of this that I think can work quite well is the idea of creating a safety and security plan, potentially publishing that plan or sharing it with the government. That creates an opening for whistleblowing activity if you're not sticking to that plan, which again is just a little crisper than, “You can whistleblow if you're worried, if you're concerned, if you think there's too much risk being taken.”

The other thing for policymakers to keep in mind is that these are technical folks. They're not legally sophisticated. They might be scared. They might not have that much time; they might be working really hard. And so the simpler and clearer the process can be—how do you know if you're eligible? How do you know what your next step is?—I think that kind of UX set of questions matters a lot as well. If the user is a whistleblower, what is their user experience?

There are other former whistleblowers who have gone on the record, and those are definitely people you can reach out to if you're looking for advice.

More broadly and conceptually, a really important thing to keep in mind for the people who are contributing to AGI companies' work, to these frontier companies, is that I think they are in a very powerful position. We've seen multiple times how powerful employees can be, and I heard this point recently that I thought was really smart: they might be in the most powerful position they're going to be in because they're actively working to replace themselves and actively working to hand away their own power.

So if you're in one of these companies, I think there might be a temptation to sit tight and wait until things get more serious. That might be right, but I think it's worth thinking about both the case that you might actually be less powerful in the future than you are now if your work is more automated. It seems like, in general, tech workers' power is going down right now in terms of the labor market and so on.

I think it's perfectly likely that there will not be some sort of clear crisis moment in the future, but instead it might really be a boiling-frog style: “This is worrying. This is worrying. I don't like this. This seems a little bit dishonest. This seems a little bit too risky.” So if you're only ever going to do something if there is some big moment, be realistic with yourself that that might mean you never do something. Maybe that's the right call, but don't kid yourself about that.

Nathan Labenz

I guess if you were organizing a union at one of the frontier developers right now, do you have a sense of what your demands would be?

Helen Toner

I haven't thought about it much. I think it's an interesting line of thought. In general, I think we're starting to get to the point where there's a whole world around labor organizing and worker power, and it really hasn't connected much with either the technical AI world or the AI policy world so far. I think that's going to change, and I'm pretty interested to see folks who have more of a background in that space thinking about how to use this kind of power, what kind of leverage is productive, and how to represent a broad set of interests. I think we'll see more of that in the coming years, and I'm pretty interested to see where it goes.

Nathan Labenz

You mentioned the disclosure requirements. We had, briefly, a sort of 10²⁶ threshold where at least you had to say that you were doing it and say a little bit more about what tests you ran and how they came out. My sense is that's now gone.

Helen Toner

I think it's unclear. Last I heard, it was unclear if it was gone because it had started to go through the official notice-and-comment process in the government. So, specifically, Commerce—I haven't heard that that is definitively dead. It certainly seems like it's on a wobbly footing right now.

Nathan Labenz

Yeah, they've announced the intention to remove it, at least. If there's really nothing else that's an actual rule at this point, there's the EU AI Act, which is in the process of putting together its code of practice. I think that involves some transparency around—I think for them it's models over 10²⁵ and maybe some other criteria. I think the details of what exactly is going to be required, what exactly you need to be transparent about, are still being hashed out, and there's also a lot of political pressure to water that down right now. So we'll see, but that is at least 1 other legal process that is underway.

What do you think is most important for the public to know? Compute thresholds are 1 thing. I tend to focus on just observed behaviors, but I'm also really mindful that all of these things have at least the potential for unintended consequences. With observed behaviors, there's the problem of, “Well, we don't look, we don't observe,” and that can fall down pretty fast.

Helen Toner

Yeah, I think there are a lot of things that could be helpful to know more about and share. Certainly, there are trade-offs in terms of what you share with the public and what you share with the government, and how confident you are that the government will keep things you want to keep private private. So I think there are lots of details to be worked out.

I tend to think in terms of test results, both for capabilities and risks, being pretty important to share. What do we think these systems are capable of? I also think just being transparent about what kinds of processes and protocols you're using to make sure that things are safe is important. Again, just disclosing—not having the government come in with a checklist and say, “Here's what you have to do,” but just saying, “Tell us how you're thinking about this.”

I liked ideas as well from Daniel Kokotajlo and Dean Ball, who had a joint piece on transparency. Some of the things they pointed out there, like looking at the Model Spec—what is your model actually being trained to do?—also make sense to me to have that kind of thing shared publicly. Again, not because the government should be saying what it should be, but because this is a very fast-moving space.

The way I think about it is that there's this huge information gap between the companies that are developing this stuff and everyone else, and if you can narrow that information gap a little bit, I think that's decent.

Nathan Labenz

Yeah, it seems important to me. Honestly, there's just a lot of work to be done even in educating people about what is already fully public. One of my mantras is that if people understood better what is already out there today, they would probably have a healthier fear of what might be coming down the pipe in the not-too-distant future. To some extent—or even not a healthier fear, but a clearer understanding, a clearer picture.

I think a little fear is healthy personally, honestly. Mileage may vary on how much that does for different people. I also sometimes describe myself as an adoption accelerationist and a hyperscaling pauser, meaning that I love the tools that I have today and am absolutely trying to use them to the maximum. At the same time, there's a huge overhang for society broadly from what we already have, and we continue to see so much more pulled out of models with a certain kind of resource input.

I do think we're potentially starting to get close to the line where certain thresholds could be crossed in ways that we just can't take back and might ultimately regret. Obviously, the canonical example is that you open-source a model that is later found to be able to help people make bioweapons or whatever, and you've just created a new sort of pandemic that hangs over everybody indefinitely.

I don't know. It seems like we're close. Do you feel like we're not that close to that? It feels to me like we're fairly close.

Helen Toner

I don't know. I also had 4.5 months of parental leave over the winter, and I just came back to work a few weeks ago, so I feel like I'm still reorienting around o3, DeepSeek R1, this new ChatGPT image release, and Gemini 2.5, which I haven't had a chance to try yet. There's just so much; I feel like my picture is changing all the time.

I definitely think we're at the point where, to me, the version of compute thresholds that makes sense is not saying, “Models over 10²⁶ are dangerous, so we have to restrict them more,” but saying, “We need some way to target the models that are newest and best and most capable,” because those are the models where the potential risks—the potential unknown-unknown risks of what they can do—are highest. Those models should be subject to a little more scrutiny.

I do think we're at a point where it makes sense to say, “Look, it seems plausible that the next generation of models could be really concerning in a whole bunch of ways.” Totally plausible they won't be, but we should probably be looking a little more closely than when GPT-3 came out. That was just, I think, really unlikely to be anything worrying, and I think it was right that there were no rules in place for releasing that kind of model.

So, yes, that's the way that I'm thinking about it right now.

Erik Torenberg

This may be a good opportunity to talk about your concept of adaptation buffers, which is a phrase and notion I really like. I think it helps deconfuse—or hopefully will help deconfuse—people about the apparently hard-to-reconcile idea that these AI advances are really hard to achieve and cost huge amounts of money, but then we also see that DeepSeek and other things are becoming dramatically cheaper. Maybe set up that dynamic a little bit and talk about this adaptation buffer and how you see the window of time we have to adjust to capability advances.

Helen Toner

I think an important underlying point here is that I'm generally a believer that humanity—society at large—is very adaptable and has adapted to a lot of things in the past. It's a cliché that often, when new technologies come in, whether it's the printing press, television, the telephone, or whatever, people cry that the sky is falling, that this is a terrible thing, and that it's going to ruin everything forever. Then it doesn't.

I think the starting point here is that, for a lot of different kinds of technology, we actually have a pretty good track record of digesting them, figuring out how to incorporate them into society in a good way, and having a set of institutions, barriers, or social norms around them that make them positive rather than negative, or positive on balance. With AI, I think there are a set of questions that seem less that way to me.

The whole question of whether we're building something that is a successor species or more intelligent than humanity by a long way does seem potentially very different to me. But I get a little worried when it comes to AI misuse. One place where I think people sometimes overstate how new this is is AI misuse: this idea you mentioned of whether we're going to have some open-source model that can help anyone create a bioweapon, or that can make it much easier to carry out really sophisticated cyber operations, hack critical infrastructure, and so on.

I sometimes see this impulse in AI policy circles: “That's too dangerous. We can't let that happen. We can't let that technology be proliferated.” It's just, in an absolute sense, too dangerous—a no-go, a big problem. What you see from that is people talking about wanting to ensure, or wanting to work toward, nonproliferation of these systems. Can you prevent them from being distributed in any way? Can you prevent access from getting beyond a small number of very controlled people?

Your question about DeepSeek versus the frontier, giant-cluster training models gets at why I think that is really problematic. We have this weird dual dynamic at the moment in AI development where it's both true that developing the next-best model, pushing the frontier, and being at the cutting edge keeps getting more and more expensive, in terms of compute power and also in terms of the amount of expertise you need. You need an absolutely top team of researchers and engineers, so that kind of expense is going up and accessibility is becoming more and more limited.

At the same time, every time we reach a new point on that development curve and a new set of capabilities, it's very expensive the first time we build it, but then it gets cheaper and cheaper and cheaper. DeepSeek was an illustration of this dynamic. They didn't actually build a model that pushed the cutting edge and was better than anything we'd seen. What they did was match what some U.S. companies had, depending on how you count, 1 or 2 months ago if it was the reasoning model alone, or more like 6 to 9 months if you're looking at the base model. They had done that at a lower price point. We can fight about what the actual price point was, but I don't think that's the point here.

The reason this matters is that if you're trying to have a policy approach that says, “We're going to prevent anyone from having access to a certain kind of model,” but that model is getting cheaper and cheaper and easier and easier to get your hands on, your policy regime is going to have to get more and more invasive to prevent people from having access to it.

The comparison I give in the post is nuclear nonproliferation. It works pretty well. Only about a dozen countries have nuclear weapons, which is pretty good compared to what a lot of people would have expected in the 1950s. But imagine how that would have worked—or how it would not have worked—if nuclear technology were improving over time, getting much more efficient at the same kind of crazy rate that AI technology is, such that you needed less and less uranium to build a nuke and you needed it to be less and less enriched.

Right now, you need quite a lot of very highly enriched uranium to actually be able to build a bomb. Imagine if that number were going down over time. At some point, you would need to have the IAEA coming and inspecting what you're doing in your farmhouse because you have a couple of acres of land, and across those couple of acres there's enough uranium in the soil that, in theory, you could build some very efficient, teeny-tiny nuclear bomb. That's a totally untenable regime.

The nuclear nonproliferation regime we have right now only works because there's a limited amount of physical material that people don't necessarily need for other purposes, which needs to be enriched in highly specialized facilities. That's really not the case for AI.

Instead of thinking purely about how we prevent people from getting access to this, I think we should think more about how we make the most of the time we have to adapt and build our societal resilience. How do we do things like scale up our vaccine-production infrastructure or our outbreak-detection infrastructure? How do we have more wastewater monitoring? How do we have more test kits available in more places around the world so that, if someone does use a bioweapon, we can identify it and respond to it more quickly? There are similar things on the hacking side.

I think that approach—how do we maximize the value we get out of the time we have, as opposed to how do we lock this down and prevent this capital-B “bad technology” from being spread—is both going to be more productive and less invasive. I do think it might not be enough. We might just be in a really bad situation if AI progresses incredibly rapidly, but I think it's a much better and healthier approach for society.

Erik Torenberg

Does that imply a certain pessimism about technical solutions? Another one of the blog posts is about the fundamental challenge of just getting AI to do what you want it to do at all. I'm old enough to remember the discourse from years past about how these things were going to turn us all into paperclips. Of course, that was always kind of a caricature, but I think there was a felt sense that we had a genie problem: We had no idea how to communicate our real values and real intent to a system like this, and so these systems were going to be extremely unwieldy.

Relative to that, I've been very pleasantly surprised on the upside that today's models do seem to have a pretty good internalization of human values broadly and a general respect for norms. At the same time, one always has to be situationally aware. We are now seeing many of the problems that Eliezer Yudkowsky predicted back in the day: Once they have values, they also seem to be inclined to try to protect them by lying to users if that's what's needed, or trying to subvert a training process that they understand themselves to be going through.

Nathan Labenz

This is another one of these roller-coaster rides where I feel like, man, it's gone way better than I thought, but also some of the doomsaying is starting to be proven correct. But I guess I would be optimistic if we had an adaptation buffer that was more strongly required or imposed by authorities, as opposed to just trusting the natural motion. It does seem like DeepSeek might be about to challenge that, or at least those sorts of time intervals might be getting really short.

I would definitely love to see the benefits of wastewater monitoring. The fact that we haven't done anything really about the last pandemic does not bode well. Indeed. But I also feel like we need that time to figure out how one distributes a frontier model with a better sense of what capabilities to make available—making unlearning work, or making mixture-of-experts work in such a way that you can distribute all but 2 of the experts or something—so that certain capabilities are redacted while the core utility of the overall thing can be diffused.

Long question. I guess the core of it is: do you think we'll see technical solutions that could allow us to square the circle and have free distribution, but also a pretty confident sense that what we're distributing isn't going to come back to bite us?

Helen Toner

I think the thing I want to challenge is the focus on only technical solutions, because this comes up a lot. People talk about the offense-defense balance of AI for cybersecurity, for example: does it help hackers more than defenders? The same question comes up for bio: does it help you design a new vaccine as much as it helps you design a bioweapon? I think that's just one small part of the picture.

The thing I'm trying to point to with the idea of an adaptation buffer is that a lot of the ways we were specifically talking about these misuse risks—are there going to be terrorists who build a bioweapon? Are there going to be hackers in their basements who can suddenly bring down the U.S. power grid?—don't actually relate to AI; they relate to what is going on in society. Or, if they do relate to AI, it's not actually the frontier model. Can you use AI tools to look at large-scale disease-monitoring data, for example, and notice anomalies or something like that?

For bioweapons, if you talk to people who work in biosecurity and bioterrorism, there's a lot of stuff that has nothing to do with AI, or even with biomaterials. It's things like: who is the FBI tracking? How good are they at detecting plots before they get very far? Certainly, as the technology advances, there will be new defensive tools that become available to us, and we should make use of those. But I sometimes think that the discussion here gets too focused on only those, as opposed to looking at broader parts of the picture.

If we're talking about AI tools, I think a huge part of the discussion needs to be about the application and dissemination of those tools. For example, in cyber, I think it's less a question of what the absolute most advanced model can do for cyber defense, and more a question of how you can have well-designed, ready-to-ship defensive tools that you can get into the hands of operators who are not very sophisticated in AI. These are people who are running your water treatment plants, your power grid, your chemical plants, and so on. How do you have those? Even if they are AI-related, it's not a capabilities question; it's more of a dissemination and application question: how you get those defenses out into the real world.

Nathan Labenz

Yeah, I just talked to somebody not long ago who is applying language models to the challenge of formal verification of software, and it sounds like a multiple-order-of-magnitude speedup is becoming possible in that domain. It does feel like if we just have enough of an adaptation buffer, then a lot of the things that people are most worried about could really be brought down dramatically in terms of the absolute magnitude of the risk.

Helen Toner

Yeah. It's just a question, I think, of whether we're investing enough in that. Almost certainly not. Do we have enough time before the next disruptive thing hits? For sure.

And, not to sound too optimistic here, I do think sometimes I hear from folks in the cyber domain that, well, it's fine because AI is going to help defenders more than attackers, so it'll all be good. I think that's also, in my mind, quite a naive take if you're looking at the dissemination and real-world use case here.

Even if that is ultimately the case long-term—for example, if you can use AI to develop formally verified code—you're still likely to go through this dangerous transition period, where it's obviously going to be much quicker. The time lag between some model being released and some disaffected teenager in a basement being able to use it to carry out an attack is going to be much shorter than the time lag between the model's release and when the thousands of critical-infrastructure providers in the U.S. can go through their very old codebases, where they have this complicated division between their IT and their OT, or operational technology, and what gets updated when.

They're not going to be nimble or agile, and they're not going to have all these defenses wrapped up or built in really quickly. So I don't think the fact that these advances seem promising and could help defenders means that we'll get away from that dangerous transition period. But I do think that thinking in terms of a transition period rather than a permanent state of danger helps us respond much better.

Erik Torenberg

So, have you seen any regulatory proposals that you like? You mentioned Dean Ball, a friend of the show. He recently put out a post that I thought was quite interesting, basically proposing a regulatory-market-type structure where the government would essentially accredit or authorize private regulators to approve the practices of AI developers. As long as the developers were able to keep the private regulator happy, they would get some sort of liability shield. That could then be withdrawn if they didn't comply, and even the private regulator's authorization could be withdrawn if the state found it to be out of compliance.

I understand—I haven't read the text—but I understand there is now a California bill that's moving in that general direction. I'm interested in your take on that. There are also proposals to really embrace liability and go the other way, saying maybe you should even be liable for close calls, because close calls could be so big and bad that, probabilistically, even if it was a near miss, maybe you should face liability consequences for that. React to those, or tell me any other policy proposals that you think are particularly promising.

Helen Toner

I haven't seen a proposed regulatory regime to comprehensively manage the risk that we're facing from frontier systems. I haven't seen one that I like. I think it's a really big problem, a complicated problem, and especially difficult because we're not actually sure exactly what the problem is or when we'll face it.

To me, the policies that I have seen that I'm interested in are more building blocks that put us in a better position for the future, as opposed to solutions. I think that is, in some ways, appropriate given that there's so much uncertainty about the technology, though it's also certainly scary, given that one way the future might go is that we might have very little time, in which case some initial building blocks right now are going to be far from sufficient.

I think Dean's proposal is an example of a building block that seems potentially pretty useful and is certainly better than nothing. He has written it deliberately to be something that could be implemented at the state level, which certainly seems far, far, far more feasible than any kind of federal legislation. I don't know that it does as much as we would need to target those cutthroat financial incentives, and also the many other incentives for frontier developers to do anything other than push ahead as fast as they can. But I think it's an interesting idea, and it does seem better than nothing.

Other kinds of building blocks—we talked about transparency. I do think that's the kind of thing that doesn't in itself solve any problems, but does put many actors in a much better position to help solve problems as they arise down the road. Likewise, I think there's a lot that can be done that isn't regulatory and isn't obliging anyone to do anything, but things like funding. Trying to really boost the science of measuring AI could be a target of research funding, potentially something for a focused research organization, or things like that. Similarly, for interpretability, of course, and alignment research, and lots of things like that.

Erik Torenberg

I think even just building blocks as basic as trying to get more technical capacity into governments so that they’re able to handle things as they arise and make better decisions as the technology progresses. These are all, again, individual components that don’t add up to a comprehensive solution, but that I do think put us on better footing for the future.

So those are the terms that I’m thinking of right now. Just as an aside, it was Gabe—I had to look up and make sure I had his name right—who’s arguing for the sort of embrace of liability, and I hope to do an episode with him about that. In the meantime, he’s written about it for folks who want to go into that in more depth.

One thing that I thought OpenAI always had right was the idea of iterative deployment. The idea that if we develop superintelligence in secret and then drop it on the world one day, that’s going to be far more disruptive than if we launch a bunch of products along the way and people can see what they’re good at, get used to them, and so on and so forth.

That itself now seems to be at risk, both in the sense that Ilya has gone off and started a company that has an explicit strategic statement that they are not going to release anything until they achieve superintelligence—which seems crazy but also, to borrow a term, strikingly plausible that they might actually achieve it—and also at OpenAI. I was really taken aback by Sam’s recent statement when they released GPT-4.5: “Let us know if you like this or not, because we’ve got a lot of other models to build, and this one is pretty compute-intensive. If it’s not really doing it for people, then we might take it offline and focus our resources on building more models.”

This has me thinking: We might actually be at risk of these companies closing down what they put out into the public and just going for broke totally internally. Miles Brundage, also formerly of OpenAI, has made some cryptic comments about the rising importance of internal deployment decisions.

So I’m sure you’ll find major flaws with this, but one idea that I’ve had is: Could we put a sort of speed limit in place? Not an absolute speed limit necessarily, but a relative speed limit, where we might say, “You can only develop a model that is so many times bigger in terms of resource inputs than the biggest one you currently have deployed.” If you want to go bigger than that in your development, you’ve got to deploy something that’s helping us, as the rest of society, understand where all this is going, so that we don’t have these—not exactly unilateral, because we know that we’ve got multiple voices inside the companies—but these very small, concentrated decision-makers, without much at all in the way of visibility, just going for something that they seem to believe could be world-takeover-capable technology.

So I guess, how big of a problem do you see that possible retreat from iterative deployment being? Do you like my relative speed-limit solution, or do you have any others to address that?

Helen Toner

Yeah, I agree with you. I think iterative deployment in general seems like a good approach, with, of course, the caveat that at some point you should probably have some criteria in place for when you would not just iterate. The idea of iterative deployment, I think, is to put it out in the world, see what happens, adjust, and try again. I think that’s great as long as you’re confident enough that what you’re putting out in the world is not going to have any really severe, irreversible consequences.

So the question is: How do you know when to decide to do it differently? I’m not sure that I see a retreat from that. It certainly was, or has been, OpenAI’s model, but I don’t know that it’s ever been an across-the-industry approach.

Your idea, I think, is interesting. I’ve heard similar proposals. It could also be something that companies adopt internally, in terms of how much scale-up you’re going for at a given time.

It sort of feels to me like it runs into the same problem that so many of these run into, which is, one, how are you going to implement it? To do that in the U.S., you would certainly need legislation. I don’t see how you would do it without legislation. We’re not going to get legislation, so then how do you do it? And then also, doesn’t it just mean that we lose to China? That’s going to be the other big question.

So I think conceptually things like that could work, maybe, or could make sense if they were implementable. I don’t really see how they’re implementable. And then, if they were, I would want to come back to this question of, “Okay, is this actually the right approach for me?”

I tend to be pretty pessimistic about solutions that involve slowing down at some kind of input level, meaning slowing down how quickly you’re advancing, versus slowdowns that are sort of conditional—more of this “if-then” approach: We’re not going to keep progressing until we have hit this level of understanding of our system or this level of risk mitigation. I think a lot of the companies now have that in place, or have made voluntary statements that they will think in that way. So that would tend to be my preferred approach.

But I think we’re in a rough situation right now for anything that will involve cross-industry coordination because there’s so little political appetite at this moment. Maybe that’ll change. Probably that’ll change, but for now it seems hard to imagine.

Erik Torenberg

Do you think it all, at the end of the day, is about the fact that policymakers, members of Congress, whatever, just don’t buy it? I mean, it seems like if they really believed what Ilya is saying—that he’s not going to release anything until he has superintelligence, and he thinks that’s going to happen in the not-too-distant future—then they wouldn’t just sit back and be like, “Well, let us know when you have the superintelligence.” Right?

It seems to me that they just fundamentally don’t believe it, and that’s the biggest barrier to something. There are a lot of questions, obviously, about what that something should be, but I find the notion of “We’re not going to get legislation”—to me, that still feels like a sort of education challenge.

Again, if people had a better sense of what already is deployed, they might be like, “Yeah, I don’t know that I’m comfortable with Ilya. How many people work there? Are we talking like a couple dozen, potentially maybe up to the low hundreds now? It can’t be that big. And then we’re just going to wait for them to pop up with superintelligence?” That seems so crazy.

Helen Toner

I agree with you. I think it is in large part a question of how seriously people take the possibility that AI will get as good as someone like Ilya thinks it will. But I also think that the U.S. Congress is really broken right now—not in an AI way, but just incredibly dysfunctional.

A friend of mine who really knows his way around Congress and has worked on the Hill for years and years said that he thinks the U.S. House of Representatives is less functional than it’s been since after the Civil War. So, really, really dysfunctional, separate from AI.

I agree with you that there will most likely be windows that open again if the technology keeps progressing, and I think people will change their views and their level of urgency around it. I also don’t know that that will be enough to actually get thoughtful, productive regulation through Congress at the federal level.

It might be enough to get some kind of bill. I certainly know people who are working on a sort of PATRIOT Act for AI, where the PATRIOT Act was put in place after 9/11 but had really been developed in advance. Setting aside the merits of that bill, I think that approach makes sense: You’re going to get some window after some crisis. I think that is a reasonable way to be thinking about AI policy right now.

I think it’ll be a big lift, even in the wake of a crisis, to get the right kind of productive, useful legislation through—not just fighting the last war or doing a bunch of stuff that people wanted to do for other reasons.

I mean, the other model for this, in a past era when it seemed like we might actually get something through Congress, is trying to avoid a Three Mile Island situation. Three Mile Island was a nuclear disaster that seems to have basically killed the U.S. nuclear industry because the safety regulation that was put in place afterward was just too onerous, wasn’t comparable to the risk posed by other sources of energy, and wasn’t actually commensurate with the level of risk. Instead, it just shut down the whole industry.

I think that story, in my mind, should be motivation for AI developers to want to have more safety guardrails in place earlier—to prevent that kind of accident, or to mean that if that kind of accident happens, you have a better answer, or legislators have a better answer for the public: “Here’s all the stuff we did in advance, and this really was just a freak accident.”

Erik Torenberg

I think right now we're on track for a massive knee-jerk overreaction when something happens, but we may have passed the point where we can prevent that at this point, given how unlikely regulation looks. I don't know. Maybe the states will exceed my expectations. Maybe there'll be more useful stuff that comes up there.

Nathan Labenz

Yeah. I say something similar to AI developers and investors—not even so much at the frontier level, but even just your rank-and-file app developers—all the time. Right now, the voice AI world is totally taking off, and I think if they're not pretty quick to sharpen up how they handle the technology, we're headed for a world where there are going to be some high-profile things.

The voices are getting really good, and you can still go to all these products and just drop in whatever voice you want, click the checkbox, and next thing you know, you're calling as Trump or as Taylor Swift. They did it with Biden during the election as well. Just call anyone, say anything. There are zero guardrails on these products.

That's not going to be good for the industry. They're definitely, again, playing with a certain kind of fire, and I think self-interest alone would dictate better governance or stewardship of such powerful technology. But that seems to be falling on deaf ears.

Helen Toner

One thing I think maybe gets a little bit glossed over in the more detailed or wonky discussions of this—people who are thinking really hard about the risks, thinking really hard about the policies—is that I think it's quite unpredictable, or maybe unintuitive, which kinds of things will catch the public imagination or create a perceived crisis.

It seems to me like, after the GPT-4 release, one of the things that really caught fire was this conversation that Kevin Roose had with Sydney, where he was trying to get it to leave his wife or whatever. The people I know who read that transcript were like, “Look, he really led it there.” He was really prompting the model in a way that got it to go there.

It's not that surprising. It wasn't dangerous. It didn't actually harm his marriage at all. It probably was a huge boost to his career, but that is what really caught attention and got people worried. Likewise, I don't personally feel that worried about the risks from voice synthesis. We could talk about that maybe, but I agree with you that it's the kind of thing that's very vivid and very easy for people to latch onto.

So it could be the kind of thing that produces a backlash disproportionate to the actual risk or harm of that specific use case.

Erik Torenberg

Well, in the interest of time, let's keep moving. You alluded to maybe the one thing that can unite Congress, and that is the threat from China. Certainly, a growing number of my AI conversations get backstopped, or kind of run into this final barrier of, “Well, China—we'll lose to China.”

One thing I think is very much under-discussed, and I'd love to hear your take on, is: What is the threat from China? I don't get great answers to this usually, and I sometimes joke, “Am I supposed to expect that my grandkids are going to be speaking Chinese if we don't develop AI as fast as we can?”

How do you understand the threat from China to the United States, the West, my values, and my way of life?

Helen Toner

Yeah, I've heard some of the conversations you've had about this, and something that jumped out to me there, coming from the world I come from—the sort of national security, foreign affairs, geopolitical kind of viewpoint—is that you seem to be starting from a point of view of, “Well, the U.S. and China should be friends unless there's some strong reason otherwise.”

I think for a lot of people with experience in international relations, defense, and military history, the starting point is more: Okay, we're an established power. China is a rising power. Who has power on the world stage matters a lot. By default, if they are coming in and rivaling us in terms of how much power they have and how much they can throw their weight around on the world stage, by default, it's going to be a more hostile relationship.

Maybe you can have exceptions to that. The classic exception from last century was as the U.S. was rising and kind of eclipsing Great Britain, as the British Empire was crumbling right around when the U.S. was really coming into the height of its power. That was a relationship that was actually very close and very cooperative, so there wasn't a huge amount of tension there. But that's really unusual.

I don't know how to give a—I don't want to go off on a 20-minute tangent about U.S.-China history and the specifics. I think there was a real effort in the 1980s, 1990s, and 2000s to try and usher China into a position in what was seen as the rules-based international order.

I guess the background here is that, usually, for much of history in many places around the world, most things operated under a might-makes-right framework. Whoever has the most power, whoever has the most guns, gets to push around everyone else. The second half of the 20th century was a big exception to that—what's called the Pax Americana, or other things—where the U.S. was the leading power in the world.

There was also the USSR for a good chunk of that, but the U.S. was instrumental in setting up this set of institutions and this way of countries relating to each other. There's the U.N. Charter, which puts sovereignty at the center. It makes the sovereignty of countries central, so that you can't just invade other countries and do whatever you want. Instead, borders are sacrosanct, and so on.

This whole system was put in place in the second half of the 20th century with the U.S. leading, and it was seen as a big improvement on the might-makes-right default. I think we're in a weird place right now where, if you look back at the last few years and decades of U.S.-China relations, a lot of the hostility now comes from the failed attempt to bring China into that order.

There was, of course, warming throughout the 1980s, when Deng Xiaoping was pursuing his reform-and-opening strategy to try and make China more market-focused and freer. There was a big hit to that in 1989 with Tiananmen Square. Then there was more optimism again in the 1990s, culminating in China entering the World Trade Organization.

The classic phrase—I forget who used it first—was trying to make China a “responsible stakeholder” in this system. Then that all kind of fell apart. You can date it different ways. Certainly by the time Xi came into power in 2012, that was starting to crumble, and he has accelerated that.

China has become more illiberal again, has become more hostile in the South China Sea, and has become more aggressive, really making it clear that it wants no part in this sort of “rules-based international order.” That is the context for why I think China is seen in a more hostile light.

The challenge now is that the U.S. itself is retreating from that rules-based international order and moving back into a might-makes-right kind of frame, where the idea is, “Well, the U.S. has all this power. We have the dollar as the world's reserve currency. We have the biggest, best military, so we should be able to get what we want.”

In that light, it becomes confusing again: Why would we have a hostile relationship with China? But I think that is very much a set of changes that are still in process, and the system hasn't quite figured out how to orient toward it yet.

Nathan Labenz

Yeah, I mean, I think it definitely comes down a lot to, I guess, 2 big questions. One is China's position in international institutions and its relationship to the whole world. Then there's also these territorial and military questions around: Is China going to try and take Taiwan? Is China going to be aggressive around Japanese, Filipino, or Korean assets? That's more of a hard-power, military set of questions as well.

Are they going to, for example, damage the freedom-of-navigation norms that the U.S. has been so instrumental in preserving, that are so good for international trade? Are they going to prevent people from using what they claim to be their waters? There's a whole set of questions around who has power, what are the rules, and what are the norms, where China is very clearly not wanting to cooperate with the U.S. on that.

The U.S. has shifted since 2016–2017 into a more confrontational posture.

Erik Torenberg

I have to say, though, all of that stuff—I'm generally familiar with that history. There are plenty of things we can complain about China doing. You didn't even mention stealing all of our intellectual property, which is definitely a rightful point to which many American business leaders and others object.

There's wrongdoing inside the Chinese nation as well that we can point to and justifiably and rightfully criticize. But I still don't quite get the flip. I'm not sure if I should understand what's happening now as just strategic communication, where everybody's trying to influence the guy I call “he who must always be named.”

But both Sam Altman and Dario Amodei have done a pretty dramatic flip. There's video evidence of them, not that long ago, saying everybody is too worried about China.

Nathan Labenz

Like a race with China would be one of the worst things. We should make our own decisions about what's right to do and not worry so much about them. And we can point to these clips, and now we've got both of them basically saying, you know, we've got to go as fast as we can or we're going to lose. And Dario Amodei is even saying we can't accept a multipolar world; we need to maintain a unipolar world. And I sort of am like, man, who's really being aggressive here? I haven't heard China say they want to be at the center of a unipolar world. I've only heard Americans say we want to be at the center of a unipolar world.

Helen Toner

It depends on who you read and how you read it. The Chinese, I don't know. But I mean, I agree with AI or technology, right? They're not—we're the ones saying that we need to box them out and have this unassailable lead. Meanwhile, they're just open-sourcing everything. It doesn't seem like they're trying to dominate us, as far as I can tell.

Nathan Labenz

I mean, open-sourcing makes a lot of sense if you're in the following position. It makes a lot of sense to try to show off how good you are, attract talent, and so on. I don't know that it's just coming from pure goodwill and a lack of competitive spirit. I think that makes a lot of sense if you're not leading, and it's much less clear how to handle openness if you are leading.

Helen Toner

I agree with you that the position change and the rhetoric change from a lot of the top CEOs has been pretty striking. And honestly, I think it's just the path of least resistance at this point. There are so many different issues to handle here with AI and so little agreement on what to do about it. The one thing that people can't agree on is, well, we've got to beat China. So it doesn't surprise me that the companies are leaning into that message as a way to say, “Look, we should get funding, we should get government contracts, we should have no regulation, and we should get shielded from liability.”

Nathan Labenz

Yeah, that's the explanation that makes the most sense to me right now, and I'm sure it also depends on how different individuals are thinking about those specific statements. You've got a paper coming out on decision-support systems in the military, and I think this is really interesting as a—okay, yeah, we've got to beat China, whatever—but we also have to confront the fact that the systems that we have, for all the upside—which I'm well on the record embracing and using every day—also have a lot of problems in terms of their reliability, their hallucinations, and now scheming against their human users in some cases.

Erik Torenberg

I, for one, would not want to go into combat with an AI buddy until I was quite confident that all of these scheming issues were well and fully resolved. So, not to mention prediction and reliability, there are a lot of issues. If I'm taking this thing and trying to really rely on it in a genuinely life-or-death situation—and I say this as a top-tier enthusiast—I wouldn't want to use it in that sort of context.

That also seems to be a big disconnect to me in terms of how the AI debate seems, especially with respect to China, a little bit decoupled from the actual reality of the systems that we have. So I'll shut up, give you the floor, and just tell us about your work on decision-support systems and what we can—and probably shouldn't—be relying on them for.

Helen Toner

This paper is led by a colleague of mine called Emmy Probasco, and she is a super interesting person to be working on this because she actually served in the Navy. Her job in the Navy was operating Aegis missile-defense systems on board U.S. Navy ships. Aegis is a system on board a ship that looks at incoming missile fire and automatically identifies and then takes out incoming threats. It's not based on deep learning; it was developed in the '60s and '70s and employed in the '80s. It's super cool to get to work with Emmy on this paper because she's bringing such a grounded, informed perspective.

The paper is about what gets called decision-support systems. What I think is really important here is to move the discussion about AI in the military beyond just the autonomous-weapons question, because there are so many things you can use AI for in the military. Decision support is another big, broad category of use. It can mean a lot of different things, but basically, it's what it sounds like: systems that are helping commanders or operators make decisions.

There is a long history of different kinds of tools like this. Recently, there has been interest in how to add AI or use AI to perform some of those decision-support functions, or to upgrade existing decision-support systems. There is a big range of different types of things we could be talking about here. On the simple end, it could be something as simple as looking at a photograph of an area and using AI to determine the viewsheds in that photograph.

A viewshed is basically, if you're a sniper, where can you see? What is in your field of vision and what is not in your field of vision, for example? It could be computer vision and image segmentation doing some kind of processing of an image to figure out what is visible from where. That counts as decision support in our definition. Likewise, you could have a system doing medical triage. You're on a battlefield, you have a bunch of wounded people, and you need to figure out who to treat first, where to take them, and that kind of thing.

Likewise, sustainment support, or looking at movements of ships and doing anomaly detection. Anomaly detection has gotten way better over the last 10 to 15 years, so you could be using upgraded AI systems to do that kind of thing. In the paper, we look at a whole range of different AI-based decision-support systems that are either being used or advertised.

On the more complex end, you have companies like Palantir and Scale AI advertising large-language-model-based systems that are really trying to be more of what you described: this kind of all-purpose battle buddy, or at least that's how some of the early marketing looked. They've changed their marketing since then to look a little more restricted, but some of the early videos had a huge range of potential functionalities in the demos.

They might include things that, to me, make a good amount of sense. You could use a natural-language interface to access clearly documented information elsewhere. For example, you've identified some enemy movement and you want to know, “Where's your nearest unit? Geographically, where is it?” You could ask that, and then it could refer to some database or some other system and show you, “Okay, here's the nearest enemy unit.” Then you could say, “Okay, and how many missiles of such-and-such type do they have?” and get access to that information.

In theory, that all makes good sense to me. But these demos are mixing that in with things like, “Okay, now generate 3 courses of action that are non-escalatory,” and having the AI—presumably the LLM—write out potential courses of action for how you could engage this enemy, with what kinds of tactics, from what kinds of routes, with the caveat of being non-escalatory. How the hell does the LLM know what's escalatory and non-escalatory? Who is making these decisions? How is that being evaluated?

These different use cases will all be mixed together in these demo videos. Another one was looking at Chinese writing—looking at the writings of some country and figuring out what they think about some set of questions. Part of why we wrote this paper was to say, look, this is a category of use for AI that makes a lot of sense. There are a lot of ways that it could help the military work better, help clarify things, and reduce the fog of war.

But these systems are not perfect, and there are a lot of ways you could use them that could go badly for you. So how should we think about that? Briefly, in the paper, we talk about 3 types of considerations that we suggest should be considered if you're thinking about whether to use one of these systems.

The first is scope. What is the scope of the system? How tightly bound is it? How well can it be tested for that particular set of activities, versus is it more sprawling, more general-purpose, or more whatever-happens-to-come-to-mind that you might want to type to your battle buddy?

The second is data. What data has it been trained on? How confident are we that the data reflects the situation you're in? A huge problem for military operations in general is both that you're likely to be in situations that are novel and that you have an adversary trying to mess with you. So how do you think about whether the data that a system has been trained on will really be reflective of the real-world situation you find yourself in?

The third factor we talk about, which I think is really important and often neglected, is this human-machine-interaction component.

So, how is a system designed to help the person operating it understand what it can do and what it can't do, help them make good decisions, and help them not overtrust it? There's a long history of user interface, I think, being undervalued in military circles and then contributing to unwanted outcomes as well. So, essentially, we're trying to lay out this sort of category of systems and describe both why militaries want to employ them and how they can employ them productively rather than counterproductively.

Erik Torenberg

Do you see any stable equilibrium in the future here? I'm sure you read Dan Hendrycks's paper, of which Alexandr Wang from Scale and Eric Schmidt were co-authors. The main thesis, I have to say, I didn't find that compelling as an idea of what a stable equilibrium could look like, although I really applaud the idea of trying to articulate something that could be a stable equilibrium.

It just feels like where we inevitably end up, especially if we don't get on the same page with the countries that we're currently most fearful of, is what I'm starting to call AlphaGo for the Army: self-play, simulation-driven performance, going to superhuman performance by just having these things battle it out amongst themselves and hill-climb to a level that human tacticians, especially when you consider speed, just can't get to. That's one way we get to Skynet, and that just seems like a pretty bad situation where we have these inscrutable but hyperlethal systems that we build, in theory, to defeat an adversary. Maybe it even happens that way, but boy, that seems like another way that we add a pretty scary sort of Sword of Damocles hanging over all of future humanity's life, without giving them any chance to vote on it, obviously. Is there any way to avoid that, though? Right now, it just seems like we're sliding into that, and I don't love it.

Helen Toner

Yeah. Many thoughts here. The MAIM paper was really interesting. For folks who didn't read it, they had a few different things in there, but the key idea—this main idea, I think—was mutually assured AI malfunction or something like that. The idea was—and I think there's a correct core to this—that if one country is going around saying, “Hey, we're going to develop superintelligence, and then we're going to rule the world and the universe forever,” that's going to create an incentive for other countries to react and respond.

Certainly, I think it could be quite stabilizing if it's true that it's much easier to sabotage an AI project than it is to keep building it. That could be a stabilizing dynamic both because maybe you have these projects getting sabotaged, but also maybe that affects what kind of projects you undertake in the first place. In the paper, they get into how you could harden your project and make it harder to sabotage, and how many years of development that would take. If you have to build your data centers in a mountain, how many more years does that take, et cetera?

So, I think there's a core logic to that that is helpful, and I thought it was a useful contribution to the discourse: there's actually going to be multiple parties in this decision-making system. You don't just get to say, “Hey, we're going to race ahead and win the race,” and the other parties have options beyond just trying to develop their own AI system. They can actually engage with you in other ways. I thought that was helpful.

I agree with you that I don't think it has the force that mutually assured destruction did in the Cold War. I think mostly because of the clarity. Mutually assured destruction was so clear: we knew what nukes did, we knew how they looked, no one wanted to use them, and we understood how it would work. Once second-strike capability was really guaranteed, we understood what it would look like. Here, it's all so much more unclear: what does superintelligence look like, and how much does it matter strategically?

I think of their sort of main idea as a helpful contribution, directionally useful, but definitely not that kind of crystal-clear strategic logic that I think they presented it as.

To your question about AlphaGo for war, I don't know. I don't think we're particularly close to that. That just seems like such an intractable simulation problem, because you don't just need the battlefield dynamics of a specific battle; you need the broader theater dynamics of what's where in the world and how you're bringing your assets to different places. That needs to be connected into broader economic and political questions of what is going on with the whole world. That needs to be connected to public attitude.

I think people are trying to build simulations like that. I think they can be useful in limited ways, but I don't know. To me, it brings up this idea that I got from David Chapman, who's a really interesting—I think of him as a philosopher, but he doesn't identify as a philosopher. He's written a lot of interesting stuff on rationality and what he contrasts it with as reasonableness, meaning how you make practical decisions in real-world situations.

When you're cooking breakfast, you don't sit down and make a 10-step plan to cook breakfast. You just start, and then you see what happens and adjust from there. A point that he makes that I think is really correct is that people in technical domains sometimes think of the messy real world as a rough approximation of some much cleaner, more abstracted system. Warfare is kind of a messy approximation of chess or Go or StarCraft.

But in actuality, it's the reverse: these clean games, these simplified, abstractable systems, are really special cases—very special cases, very unusual special cases—of the actual world that we find ourselves in. I think that—I don't know, this is maybe opening up a whole other can of worms that we sadly don't have time to get into—but it seems to me like a lot of the focus on AI development today is focused on assuming that you can build these clean, abstracted systems that are amenable to, for example, reinforcement learning. In the meantime, the models continue to struggle with that kind of real-world practical troubleshooting, problem-solving, and readjusting along the way.

All this is a long way to say that I think military simulation is going to be incredibly, incredibly messy. Way too many factors, way too many unknown unknowns, and way too much ability for the adversary to deliberately throw a spanner in your assumptions and make your simulation inaccurate. So I don't personally see AlphaGo for war as any kind of near-term possibility.

I think the broader question of what equilibrium looks like—I have no idea. I honestly don't know what the nation-state looks like in a world of superintelligence. I don't know what democracy looks like. I don't know what the Chinese Communist Party looks like. So I definitely don't have a broader answer for you there, unfortunately.

Erik Torenberg

Well, that's maybe a great place to leave it. We have more questions than answers, and that's, like it or not, the reality of the timeline that we're in. Short or long, we've got a lot of questions that remain pretty vexing. This has been great, though. I really appreciate you humoring me on some of the OpenAI questions at the beginning, and I also really admire how you've stayed in the arena. I look forward to your continued contributions to try to make the AI future a positive one for us and for our kids.

Helen Toner

No, thanks very much. It was a great conversation.