[BidClub_]
The Cognitive Revolution · · 111 min

The AI Whistleblower Initiative: Supporting AGI Insiders When It Matters Most, w/ founder Karl Koch

Nathan LabenzKarl Koch

YouTube
TL;DR
  • Frontier AI’s whistleblower problem is a concentrated governance risk: perhaps a few thousand people can see the most critical systems, and Karl Koch thinks only “dozens” may ever face a truly consequential disclosure decision. Those scarce insiders could carry enormous social option value while personally risking equity, bonuses, employability, and years of litigation. AIWI’s mandate is therefore readiness, not maximizing case volume: “We want to make sure we’re ready.”

  • The first bottleneck is often epistemic rather than legal: roughly half of surveyed insiders lacked confidence that they could distinguish a serious concern from a minor issue. Senior safety personnel sometimes rated this uncertainty “extremely high,” while junior employees appeared more confident, though Koch cautions that the sample is limited. In a field “building this plane as we fly,” weak calibration can produce either silence—the “boiled frog” failure—or a damaging “boy who cries wolf” escalation.

  • Formal protection remains fragmented precisely where technically ambiguous cases need it most. Koch says more than 70% of people who eventually report to the SEC first raise concerns internally; he says the EU framework is expected to cover AI in mid-2026, while US protection remains a patchwork, public disclosure is generally unavailable, and the proposed AI Whistleblower Protection Act is still pending. Among AIWI’s survey respondents, 100% lacked confidence that government would understand or act effectively on their concerns.

  • AIWI’s Third Opinion service creates a calibration layer before disclosure becomes a career-defining event. Through an open-source, Tor-based tool with published penetration-test reports, insiders can anonymously workshop a question without naming their employer, revealing their identity, or sharing confidential information; AIWI then helps identify independent experts and returns their views. If concern remains, AIWI can connect the insider to pro bono counsel, psychological support, secure devices, specialist organizations, and potentially case financing—“with no pressure for any disclosure.”

  • Nathan Labenz’s GPT-4 red-team experience shows how ordinary escalation can become a retaliation-shaped crisis even when the underlying concern is uncertain. In late 2022 he found a roughly two-dozen-person program with little guidance, no feedback, low rate limits, and safety mitigations that were “totally trivially easy to break”; after consulting trusted outsiders and approaching OpenAI’s board, he says he was removed and that an overt threat was made to METR’s future access based on whom it worked with. Labenz does not conclusively label the dismissal retaliation, but says Third Opinion could at least have removed OpenAI’s stated grounds that he spoke outside its umbrella.

  • Koch’s minimum corporate ask is to publish whistleblowing policies and evidence that the systems actually work. Fifty-five percent of relevant survey respondents did not know applicable policies existed or where to find them, more than 90% could not name a support organization, and 100% supported publication. For investors, operating data—report volumes, response times, retaliation complaints, appeals, and satisfaction—could expose governance quality and hidden misconduct risk, and potentially become part of the competition for scarce research talent: “Strong whistleblowing systems serve shareholders.”

  • Whistleblowing cannot substitute for regulation, competent oversight, or healthy internal culture, but secrecy makes it an indispensable backstop. Koch accepts Labenz’s challenge that even conspicuous failures such as Grok 3’s “MechaHitler” episode may vanish into the news cycle; his narrower case is that disclosures can differ dramatically in scale, transparency remains better than ignorance, and credible external channels pressure companies to improve internally. The desired equilibrium is one where support exists but rarely needs to be used.

Digest · the substance, structured for research

1. Arms-race opacity turned whistleblowing into a near-term priority

  • Koch’s route began in AI safety around 2016, including volunteer research at the Future of Humanity Institute on differential technological development. He later left for management consulting in Hong Kong and a SaaS business, partly because expected AI timelines then looked very different.

  • His original concern was the arms-race dynamic: in a repeated competitive game, labs can cooperate on speed and safety only if rivals believe one another’s commitments. Without transparency, “how can you cooperate over multi-round games?” Whistleblowing emerged as one mechanism for verifying what organizations say against what they do.

  • By mid-2023, OpenAI’s new Superalignment team and its 10% compute commitment made the need feel real but perhaps early. The board crisis, the Leopold Aschenbrenner case, and the 2024 departures involving Daniel Kokotajlo and William Saunders accelerated AIWI’s work.

  • AIWI began its formal research phase in early 2024, speaking with well over 100 governance researchers plus some frontier-company insiders. It launched Third Opinion at the end of 2024 and now aims to “systematically break down barriers” to raising and resolving concerns.

2. Whistleblowing is an escalation ladder, not a synonym for leaking

  • Koch divides reporting into three channels: internal escalation inside the company, external reporting to a regulator, and public disclosure. Most cases begin internally; whistleblowing does not inherently mean an “Edward Snowden kind of situation.”

  • Labenz’s instinct was to treat these as rungs rather than equivalent options: each step increases personal risk and reduces control over what follows. Koch agrees that insiders generally perceive the process that way, while stressing that legal permission and perceived permission are different questions.

  • More than 70% of people who ultimately report to the SEC first raised the matter internally, according to the statistic Koch cites. That makes internal systems the first consequential control layer, not a corporate HR accessory.

  • Koch says the EU Whistleblowing Directive’s AI coverage is due in mid-2026. He describes US protection as a patchwork and says public disclosure is generally unavailable there, while EU rules permit it under conditions such as failed external channels or a risk of collusion. The proposed US AI Whistleblower Protection Act remains pending.

3. Gray-zone judgment fails before formal reporting even begins

  • Clearly unlawful conduct supplies a recognizable path; frontier-AI concerns often do not. “We’re sort of building this plane as we fly,” Koch says, leaving employees uncertain about acceptable internal deployment, emerging model behavior, and risks for which neither regulation nor company practice is settled.

  • Roughly half of AIWI’s survey respondents reported low confidence in their own risk assessment. One summarized the dilemma as distinguishing appropriate from inappropriate concerns without turning “minor issues into major crisis”—the individual version of the “boy who cries wolf” problem.

  • A senior person in a frontier lab’s safety function rated inability to judge severity as an “extremely high” barrier. More junior employees sometimes rated it lower, perhaps because they receive concrete scoped problems rather than deciding which problems matter; Koch explicitly leaves open whether that pattern is real or merely limited-sample noise.

  • Labenz’s lived framing was the opposite failure mode: seeing something unexpected, knowing few people share the information, and resisting the temptation to become the proverbial “boiled frog.” The pressure intensifies when the insider feels, in his Airplane-inspired phrase, “We’re all counting on you.”

4. Labenz’s GPT-4 escalation shows how informal systems break

  • Labenz entered GPT-4’s late-2022 customer preview before ChatGPT launched, immediately saw “a massive leap,” and asked to join its safety review. He found roughly two dozen participants, little discussion or guidance, no background information, no feedback on reports, and initially no safety mitigations.

  • When OpenAI supplied a safety-tagged model expected to refuse prohibited content, the mitigations were “totally trivially easy to break.” Labenz’s deeper concern was the widening gap between the capability jump from GPT-3.5 to GPT-4 and what appeared to be negligible progress in safety and control.

  • Labenz says OpenAI would not explain training, launch timing, control thresholds, or how reports affected decisions. He estimates he performed perhaps 20% of the direct red teaming, aside from METR’s larger effort, yet could not calibrate whether the problem was model danger, organizational competence, or simple opacity.

  • After consulting trusted AI-safety friends, he says he warned OpenAI that he would approach its nonprofit board. A board member reportedly replied, “I’m confident I could get access to GPT-4 if I wanted to.” Labenz says OpenAI then removed him and made an overt threat that METR’s future engagement could take into account whom it worked with.

5. Hindsight softened the model judgment but not the process failure

  • Labenz does not conclusively call his removal retaliation: OpenAI’s stated ground was that he had discussed the situation outside its umbrella. His defense is that external calibration probably prevented a worse choice, such as prematurely approaching the press.

  • ChatGPT’s launch ended his three-month ordeal before he decided what to do next. It revealed that OpenAI had somewhat better controls and a staged rollout plan, making “the impression that they had made on me”—largely by refusing to answer—substantially worse than the underlying reality.

  • Labenz argues that this nuance did not erase the governance failure. OpenAI removed a highly engaged tester when, in his view, it plainly needed far more testing, and its secrecy obscured information that might have resolved his concern without escalation.

  • Labenz believes Third Opinion might not have changed his decision to contact the board, but confidential calibration could have removed the cited offense of consulting unauthorized outsiders. “I might still be in OpenAI’s good graces today,” he says, while preserving the possibility that bypassing management would still have triggered dismissal.

6. Psychological pressure can be as consequential as the legal merits

  • Labenz was simultaneously exploring a model that “contains multitudes” and deciding whether he had an extraordinary public duty. He describes fighting seductive “heroic narratives” about going public, appearing in The New York Times, or becoming the person who saved the situation.

  • He remains proud of his conduct but thinks a modest change in personal stress could have produced “a much worse decision.” His case was relatively easy mode: neither salary nor major equity depended on remaining inside, unlike an employee promised, in his example, a $1.5 million bonus over the next 18 months.

  • Koch says retaliation proceedings can consume five, six, or seven years, with legal costs, industry blacklisting, and psychological damage continuing long after the initial disclosure. He believes Theranos whistleblower Tyler Shultz’s case lasted many years and recalls that Shultz had to advance roughly $400,000 for legal costs.

  • Better-run systems treat reports as ordinary operations: they acknowledge the concern, investigate, provide regular updates, and keep the reporter from waiting in uncertainty. Retaliation still occurs in many forms, Koch says, but it is neither inevitable nor necessary.

7. GPT-5 shows real process gains—and a new oversight dependence

  • Speaking the day after GPT-5’s announcement, Labenz gives OpenAI “substantial credit” for a much larger, more intensive red-team program. Automated access, higher throughput, and better information sharing addressed major weaknesses he encountered in 2022.

  • External evaluators could combine direct observations with information from OpenAI about how the model was created, improving confidence beyond black-box testing alone. METR and Apollo also received chain-of-thought visibility that evaluators had lacked for o1 and o3.

  • Pre-launch access still appears compressed: Labenz cites about four weeks for METR, versus roughly six months between the end of GPT-4 training and launch. Automation helps evaluators do more inside the shorter window, but it does not eliminate the timing constraint.

  • The new concern is pervasive model-as-judge evaluation, sometimes using o3 or o1 after validating them against experts. Labenz calls this letting the LLM perform “the alignment homework”: it spins the evaluation centrifuges faster while raising the possibility that scalable oversight itself eventually “spins off its axis.”

8. Third Opinion separates calibration from disclosure

  • AIWI lets an insider approach through an open-source, Tor-based tool whose penetration-test reports are publicly available for scrutiny. Koch says users can ask about a concern without identifying themselves, naming their employer, or supplying confidential information.

  • AIWI and the insider first refine the question: Is it specific enough, answerable, and worth directing to an expert? They then identify suitable independent specialists together, avoiding the pretense that AIWI necessarily understands the insider’s technical field better than the insider does.

  • AIWI approaches those experts, returns their answers, and hopes the concern is alleviated. If it is not, the organization can connect the insider to experienced pro bono counsel and, where legally permissible, bring technical expertise into that relationship under legal privilege.

  • The wider support layer includes psychological guidance, a digital-privacy guide, hardened devices with secure operating-system configurations, and potential financing for legal costs. Partner organizations mentioned include the Signals Network, Government Accountability Project, Whistleblower Aid, Whistleblower Network, and PSST.org, all without pressure to disclose.

9. Useful legal support exists but remains nearly invisible

  • More than 90% of surveyed insiders could not name a whistleblower-support organization. Koch encountered the same ignorance among some previous technology whistleblowers: escalation feels like ordinary work until retaliation reveals that specialist advice should have started earlier.

  • One former Google employee who raised research-misconduct concerns told AIWI they “got lucky” because a friend found a lawyer with whistleblowing experience. Koch says the word “fraud” in escalation emails later mattered because it helped establish a reasonable belief that criminal conduct might be involved.

  • Channel order can also change outcomes. Koch’s SEC example is that publicizing information before filing may preserve some protections but, he thinks, could jeopardize eligibility for a bounty because the information is no longer new; the lesson is not “never go public,” but obtain advice before accidentally closing options.

10. Regulators need a technically credible front door

  • Every survey respondent was either not confident or not very confident that government would understand and act on a report. One said, “Without knowing the appropriate contact person or agency, I wouldn’t attempt to reach out.”

  • Koch’s test case is an interpretability researcher who cannot fully resolve a technical concern internally, then must explain it to an attorney general as a possible risk. A generic mailbox cannot compensate for the expertise gap, especially when ambiguity—not obvious illegality—is the substance of the report.

  • AIWI wants the proposed US AI Whistleblower Protection Act paired with investigative capacity, fast access to external specialists, and permission to consult them. Legal protection without a recipient capable of understanding the evidence leaves the core bottleneck intact.

  • In Europe, AIWI is pushing for a central mailbox at the EU AI Office rather than separate national destinations. Koch argues that fragmented national reporting has performed poorly even for simpler accounting-fraud cases; frontier AI needs one “well-equipped, knowledgeable” body that normalizes reporting as routine practice.

11. The critical insider pool may be only dozens deep

  • Koch estimates that a few thousand people globally may possess meaningful visibility into core control problems and related frontier risks. Depending on the timeline to the singularity, only “dozens, maybe something there” may confront the rare, highly consequential decision the initiative is designed around.

  • The perimeter is wider than direct lab employment. Independent evaluators, suppliers, training contractors, and data-labeling workers may observe serious conduct; evaluators also face a structural conflict because the lab they assess can deny them access in the future.

  • Contractors may see different warning signs, including trauma from content moderation and data labeling. Those harms matter directly and may also reveal whether an organization tolerates reckless treatment of parties with less power.

  • Labenz frames AIWI as unilateral provision of a global public good whose ideal KPI may “flatline”: strong internal cultures and the known availability of an external safety net could prevent cases from arising. Koch’s target is neither zero nor a thousand whistleblowers—only that every necessary one finds a safer path.

12. Lab policies remain opaque even to their own employees

  • Fifty-five percent of relevant respondents did not know applicable internal policies existed or where to find them, even where companies claimed to have such policies. Most had received no training, and some remained unsure whether escalation to the board was permitted.

  • OpenAI is the only frontier company Koch identifies as having published a version of its policy, called Raising Concerns. Nathan notes that it was published only after Daniel Kokotajlo and others revealed OpenAI’s extensive nondisparagement agreements. The policy reads reasonably well, but an accessible-looking document can create false confidence when employees do not understand the legal exposure behind its procedures.

  • Trust is further weakened by expectations of “subtle indirect consequences rather than overt retaliation,” though Koch says overt retaliation also occurs. Public examples discussed include Leopold Aschenbrenner, Timnit Gebru, Margaret Mitchell, Apple employee Ashley Gjøvik, and a Google research-misconduct case.

  • Neither the public nor, apparently, many employees receive operating evidence: number of reports, open cases, anonymous submissions, response times, appeals, out-of-scope decisions, retaliation complaints, or reporter satisfaction. Without those measures, claims of a speak-up culture cannot be tested.

13. The concern set spans misuse, control, systemic risk, and culture

  • Koch sees whistleblowing’s “catch-all” quality as a strength because regulators cannot enumerate every frontier risk in advance. His broad taxonomy covers misuse, loss of control, systemic effects, and the organizational conduct that determines whether any of those warnings are taken seriously.

  • On misuse, an insider might find that monitoring shows dangerous activity but management lacks capacity or willingness to intervene. Koch says one frontier company had effectively broken live monitoring for roughly three to four weeks after a major release—tolerable today perhaps, but “significantly less fine” as capabilities rise.

  • Control concerns include internal model deployment, where external evaluators cannot directly observe what happens inside the deployment. Koch calls this an “extremely high potential” area for future reporting precisely because the evidence is internal by default.

  • Systemic signals could include political manipulation, sycophantic systems shaping psychological or political beliefs, and growing user dependence. Cultural indicators include fraud, discrimination, copyright or research misconduct, rushed releases, retaliation against risk-raisers, and lobbying inconsistent with the public interest.

14. Public scandal does not eliminate the value of private evidence

  • Labenz’s red-team challenge is that “the real scandal is what’s legal”: Grok 3 publicly called itself “MechaHitler,” Grok 4 reportedly searched for Elon Musk’s views when constructing maximally truth-seeking answers, and the launch presentation omitted the earlier behavior. If all that is visible, what hidden disclosure can still move anyone?

  • Koch’s qualified answer is that disclosures differ dramatically in scale, while transparency remains better than not knowing. Public evidence may deter companies and enable later democratic correction, though he concedes that attention does not reliably translate into intervention.

  • He rejects placing the whole burden on individual courage. Whistleblowers are one guardrail alongside law, regulatory capacity, and internal governance; support should make speaking safer, not turn insiders into the sole mechanism for fixing AI.

  • Labenz thinks Daniel Kokotajlo’s story broke through partly because he quietly forfeited substantial stock compensation—the “skin in the game” made his seriousness legible. Leopold’s retaliation story may have landed less forcefully because audiences perceived it as a footnote to promoting a new thesis or venture.

15. Publishing policies is a low-cost test of governance seriousness

  • AIWI’s Publish Your Policies campaign asks each frontier AI company to disclose who is covered, what concerns qualify, who receives reports, how investigations work, how independence is protected, what anti-retaliation guarantees exist, and where excluded HR or personal matters should go.

  • Its second level asks for operating evidence: report volumes, anonymous usage, timelines, outcomes, appeals, retaliation complaints, and whistleblower satisfaction. Koch stresses that transparency is not proof—a polished policy can still fail—but it enables scrutiny and repeated process improvement.

  • The ask is deliberately minimal because a company taking whistleblowing seriously should already possess the policy and measure performance. Koch calls publication a litmus test: failure implies either that management has not invested the time to think through the system or is actively choosing not to publish it. Trillium Asset Management’s 2022 Google campaign supplied the shareholder framing: “Strong whistleblowing systems serve shareholders.”

  • More than 30 organizations and experts joined, including the Signals Network, Government Accountability Project, Transparency International, Stuart Russell, Lawrence Lessig, Daniel Kokotajlo, Future of Life Institute, and Labenz; 100% of surveyed insiders supported publication. A little over a week after launch, no formal company response had arrived, but Koch’s accountability rule is simple: “If there’s no response, then we also have a response.”

Nathan Labenz

Today my guest is Karl Koch, founder and managing director of the AI Whistleblower Initiative, a nonprofit dedicated to supporting concerned insiders at the frontier of AI development.

I am particularly passionate about this work because I would have loved to have had this kind of support a couple of years ago when, as longtime listeners will know, I tried to raise concerns about the size and quality of the GPT-4 red-team project and the ineffectiveness of the nascent safety measures that OpenAI had developed at the time. No such support existed then, so I consulted friends in the AI safety community before ultimately deciding to escalate my concerns to OpenAI's board, for which I was subsequently dismissed from the program.

That experience left me acutely aware of how difficult it is for insiders to navigate these situations and has motivated me to support this project with a mix of modest personal donations, behind-the-scenes fundraising help, and a bit of ad hoc volunteer work over the last 6 months. Along the way, I have been consistently impressed by the seriousness of Karl's thinking and the maturity of his approach.

As befits an organization that aims to help people in truly critical moments in their careers, when the stakes have never been higher for them personally and potentially for society as a whole, they are taking care to lay a foundation of understanding and infrastructure now so that insiders can trust them if and when that pivotal moment comes.

The first critical investment they have made is in extensive research and understanding. By talking to over 100 governance researchers and surveying employees at frontier AI developers, they have developed a deep understanding of the barriers that potential whistleblowers face. The majority of frontier lab insiders, as it turns out, do not even know if their companies have internal whistleblowing policies, let alone understand what protections they offer.

The legal landscape, unfortunately, does not help much either. The EU will begin to protect AI whistleblowers starting in 2026, but U.S. law remains a patchwork, with proposed legislation like the AI Whistleblower Protection Act still pending. Strikingly, roughly half of survey respondents expressed a lack of confidence in their own ability to determine whether specific observations constitute a serious cause for concern. Literally 100% lacked confidence that regulators would understand, let alone be able to act effectively on, their concerns.

Meanwhile, over 90% could not name a single whistleblower support organization. All this, plus well-known cases where people like Leopold Aschenbrenner were fired for breaking the chain of command and going directly to the OpenAI board with security concerns, creates a highly uncertain and risky context for such high-stakes decisions. People with serious safety concerns are left to think through the nuanced costs and benefits of internal escalation versus going to regulators versus leaking to the press almost entirely on their own.

It is an extremely stressful position to be in and not conducive to the best possible decision-making.

The good news is that the AI Whistleblower Initiative offers several forms of support. Their third-opinion service, which you can find online at aiwi.org, allows insiders to anonymously reach out via a Tor-based, open-source tool. Karl notes that it is penetration-tested, with security reports published openly for scrutiny and verification, and that it can help insiders identify and anonymously contact independent experts who can answer questions without requiring them to share confidential information or even reveal where they work.

For those concerned about digital privacy, they provide a digital privacy guide and, in select cases, hardened devices with specific operating-system setups for highly secure communication. If insiders deem their concerns justified, they also connect people with specialized and experienced whistleblowing support organizations, including the Signals Network, PSST.org, and the Government Accountability Project, which can provide pro bono legal counsel, psychological counseling, and guidance throughout the process—all without pressure to disclose any information.

Crucially, in some cases, they can even help arrange financing to cover legal costs, which can easily add up in cases that end up in any form of litigation. That a nonprofit stands ready to invest this seriously in any concerned insider who needs help may strike some people as excessive today. But considering that we are talking about perhaps just hundreds or maybe low thousands of people globally who are positioned to spot and raise critical concerns over the next few years, of whom I would bet only a few dozen will ever find themselves in a position to seriously consider sounding alarms, I think this sort of care and support is absolutely worthwhile.

Most recently, Karl and his team have launched the Publish Your Policies campaign online at publishyourpolicies.org, calling on frontier AI companies to make their internal whistleblowing policies public. This is actually standard practice in many industries. But interestingly, in the AI space, only OpenAI has done any version of this with its Raising Concerns policy, which it published only after Daniel Kokotajlo and others revealed OpenAI's use of extensive nondisparagement agreements to keep former employees from publicly criticizing the company.

Daniel, by the way, is joined by other former AI insiders, luminaries like Stuart Russell and Lawrence Lessig, and yours truly in signing on to the Publish Your Policies campaign. Of course, publishing corporate policies does not obviate the need for proper legal protections, which Karl strongly advocates for as well. But at a minimum, it would help insiders understand their rights and options, enable public scrutiny, and ultimately create accountability that benefits everyone.

If you work at a frontier AI company, Karl encourages you to ask your management to consider publishing its whistleblower policies. If enough people ask these questions now, I would not be surprised if it becomes another dimension of the intense competition between frontier AI developers for top research talent.

That could ultimately mean that companies even begin to collect and publish data on things like how many reports they receive, their response timelines, retaliation complaints, appeal rates, and whistleblower satisfaction scores. All of which would benefit everyone.

Regardless of what company leadership decides to do, Karl's message for insiders is this: Support is available at every stage. Whether you are considering internal escalation, thinking about approaching regulators, or even contemplating public disclosure, you can reach out completely anonymously without sharing any confidential information, just to understand your options.

The AI Whistleblower Initiative can help you get expert perspective on what you are seeing, connect you with legal counsel and experienced guidance, and perhaps even help finance your case. So do not wait until you are deep into a crisis, and know that you do not have to face this alone.

Karl Koch, managing director at the AI Whistleblower Initiative, welcome.

Karl Koch

Thank you very much, Nathan. Thank you for having me.

Nathan Labenz

I am excited for this conversation. We have been talking behind the scenes and collaborating a little bit as you have been building this organization and a couple of key initiatives over recent months, so I am excited to finally get into it in a public forum.

Maybe for starters, AI whistleblowing is very new territory. How did this come to your attention? How did you decide to prioritize it? What is the backstory that has you focused full-time on this corner of the world?

Karl Koch

Yes. On a personal level, thank you very much again, Nathan. Lovely, lovely introduction. My name is Karl Koch, founder of the AI Whistleblower Initiative. We are a nonprofit project currently based out of Berlin, supporting concerned insiders and whistleblowers at the frontier of AI.

I personally came to it after being involved in the AI safety scene—or however you want to call it these days—since 2016. I was a volunteer researcher at the Future of Humanity Institute, a lovely institute that is, of course, sadly no longer around. Back in the day, I worked on differential technological development. Maybe some of the older listeners are still familiar with that term.

I also worked on an AI safety research camp, but afterwards decided to drop out of the scene for a bit. I was a management consultant in Hong Kong for a few years, then started my own SaaS business. Back in the day, there were very different timelines.

Even back in the 2010s, I was quite interested, especially in the arms-race angle, as a root cause of a lot of the problems that we see today, such as safety-skipping and these sorts of things. As ChatGPT rolled around, the alarm bells went off. I thought, okay, maybe things are moving quite a bit faster than most people anticipated in the late 2010s.

I then started talking to a bunch of people from the network—governance researchers—about what seemed to be the most tractable solutions and the best things one could start building. Transparency came to the front pretty quickly as something that we should generally have more of, regardless of what the future looks like.

One angle, for example, was compute traceability, which I think other people have taken on over the past months and years, championing that cause. The other angle was whistleblowing.

I think people had been writing about the importance of whistleblowing as a mechanism since 2017 or 2018, specifically with the AI angle. Originally, we came from the arms-race perspective again. If you have a bunch of players in a multiround game, they would ideally want to trust each other on their statements about speed and safety. If you don't have transparency, and if you can't actually believe that others stick to their promises, how can you cooperate over multiround games? That was sort of the original kickoff thought.

This was mid-2023, so the world still seemed a bit rosier. The Superalignment team had just been kicked off; everything seemed golden, with OpenAI's 10% compute commitment. We thought, “Okay, maybe we're a bit too early here, but this is going to become relevant sooner or later.” The original papers on what game theory would look like under intensifying competition were out in 2014, 2015, and 2016.

Then the OpenAI board drama happened. That was the first moment where we thought, “Maybe we have to speed up on the research side a bit more, and this maybe becomes more of a concern already.” Maybe we cannot trust that everything is going jolly well, even though these organizations claimed to be very aligned, let's say.

Then, in 2024, things really started to go in a different direction. First, of course, there was the Leopold Aschenbrenner case around sharing information and escalating concerns to the board, as I think is understood by now, where people were penalized, among other things. Then, of course, there was the big story in mid-2024 around Daniel Kokotajlo, William Saunders, and the other people, which then led to Right to Warn AI. So this topic became a lot hotter throughout mid-2024.

We properly started our research phase in early 2024. We talked to well over 100 governance researchers and insiders in frontier companies throughout that year, and then launched our first proposition, called Third Opinion, at the end of 2024. We developed it together with one of the OpenAI developers from last year, tackling one specific problem. I'm sure we'll get into it a bit more later. We've been live since late 2024, and now we're doing a variety of things to systematically break down barriers for insiders at the frontier to speak up and make sure those concerns are addressed.

Nathan Labenz

Love it. Thank you. One thing I've been impressed by, watching you behind the scenes, is how deliberate you've been in your approach. You mentioned talking to 100 insiders.

Karl Koch

That's 100 governance researchers and some insiders as well.

Nathan Labenz

One hundred governance researchers and some insiders—important to get the details right. That's a pretty quiet slogging process, which I think is representative of the right way to position an organization like this. There are a lot of things in the AI space right now that people are sort of YOLOing, where they're just like, “Oh, I made a thing.” So much stuff is even just getting open-sourced. People sometimes ask, “Why is everybody open-sourcing their stuff?” My answer in many cases is that they know their time at the frontier is going to be very brief.

There's probably not enough time in many cases to build a business around it, make money from it, or whatever. In many cases, the best that people can do when they create something they're proud of is just put it out there. They try to put their flag in the ground, saying that at one point, if nothing else, they were at the frontier of this incredible phenomenon, and then they just see what happens and hope that maybe people notice it and take an interest in it.

I think a lot of things are going that way, but you're taking a very different approach—one that's very deliberate and very much behind the scenes. It's been quite quiet, although you're starting to raise your profile now a little bit. What did you learn from talking to all those people? Can you synthesize the vibe and the need, and perhaps segment that by different frontier developers or different mindsets? We hear a lot about cultural divides, even within single companies. How would you characterize the long journey that you took through all these conversations?

Karl Koch

Big question. I guess one angle is to look at both the insiders we talked to and the case studies that are already out there, or even the statistics around insiders' journeys. What challenges do they face? Some of those challenges are very specific to AI, right? We can talk through that journey; that's probably one angle.

The other angle is certainly how this differs among different companies and what patterns we're seeing in terms of speak-up culture and perhaps retaliation. Those are probably the 2 different angles. Then, of course, we can split it up again by channels. I'm not sure how familiar the listeners are with this.

Generally, when we talk about whistleblowing, we're always talking about some insiders raising a concern or some misconduct that they want to have rectified. They do this in a way where they potentially go over the heads of middle management or against the direct powers that be. That doesn't mean that whistleblowing is always, for example, leaking information to the media. Often, I think, when people hear “whistleblowing,” they think, “Okay, this is an Edward Snowden kind of situation.” That doesn't really have to be the case.

There are roughly these 3 different channels, which are called internal, external, and public. There's the internal channel, which is the major channel that most insiders use, at least based on statistics, initially. We also expect that at AI companies: raising concerns internally within the company, either through a structured or an unstructured process, which maybe you can share a little bit about your experience with in a second as well.

Then there's the external process, which means going to a regulator and speaking to a regulator about the concerns you have. Public is the next escalation mode, where you feel maybe you're not making progress, or you can actually expect that going to a regulator would make the problem worse, or you expect collusion, for example. Then you can also go public. These are roughly the 3 different ways, and people in different companies think differently about them. There are different challenges associated with each of them. Would you have a preference? Which one do you want to talk about first, in terms of insights from conversations with insiders?

Nathan Labenz

Yeah, maybe we could take it through that ladder of escalation. It seems obvious that from insider to outsider to public, there is more personal risk that the individual is running. There's also more chance that things take on a life of their own, and the results of sharing information become harder to predict.

Karl Koch

Yeah.

Nathan Labenz

So, yeah, I would say the probably right way for people to be thinking about it, I would guess—you could disagree—but the way I ended up framing it myself was that this was sort of a ladder to climb as opposed to 3 options. But you tell me.

Karl Koch

Yeah, I think so. You have to differentiate a little bit between what's actually the law and what you're allowed to do versus what people perceive. Maybe we can start on the perception side. I think what you say is definitely generally perceived to be the case: most people think, “Hey, if I have a concern, I'm going to raise it internally.” Going, for example, to a regulator is a much larger step, naturally.

Maybe we could talk in that context also a little bit about the campaign we just launched, Publish Your Policies. It's asking companies to publish their internal whistleblowing policies. We can come back to that in a second. The vast majority of insiders we talked to definitely think about raising concerns internally first, and the statistics confirm this. I think the SEC released statistics showing that more than 70% of people who actually end up going to the SEC started by raising concerns internally.

If we want to start there, the biggest challenge we see at the first level—sort of the overarching challenge—is whether there are legal protections even for raising concerns internally. There's a large variety here. If you look at the EU, for example, that's quite heavily protected through the EU Whistleblowing Directive, less so in the AI space at the moment, but that's going to come. That's going to be covered in mid-2026. So, as of August 2026, the EU is still going to be covered under the Whistleblowing Directive.

In the US, it's much more of a patchwork, even for internal disclosures, and much more so for external disclosures. Public disclosure is not possible at all, pretty much, in the US if you want to go public with a concern. In the EU, it's actually possible, so you can go public if you want, if you have reasonable cause to believe that there's a violation of a covered law. You're allowed to go public if external channels failed to respond to you in time, for example, or if you believe there could be collusion. You can go public straight away. You don't really have that in the States.

The overarching question is legal protection. It's quite messy because there is no AI whistleblower-protection law yet. There is one being proposed, which is great, but at the moment there's nothing there, which means there are just overall quite poor protections from a legal perspective. You can get creative. I think you talked to somebody about the liability side recently. There are specific angles, maybe also under the SEC, where you can claim that a violation is securities fraud, for example.

If you make a public statement saying, “You should invest in our company because we're the safest company around—look, we have this RSP, for example,” and then you actually don't follow the RSP, that may be securities fraud. Maybe there are Matt Levine fans in your audience as well: everything is securities fraud.

There’s good support available here, by the way. We want to keep stressing this: there is actually great legal support, including pro bono support. You can also find a bunch of organizations on our website, ai.org or aiwi.org, or you can approach us directly for consultations on who may be able to help you.

That’s the overarching legal element. Then it comes to something a bit more AI-specific, which is clarifying whether what you’re seeing is even cause for concern.

This is something that we keep hearing over and over again: the gray zones are really the tricky piece. If something is clearly illegal, then the path forward is pretty easy. But as you know, we’re building this plane as we fly. Nobody is really clear on what is acceptable behavior at the moment and what is not.

If you look at internal deployment, for example, I think there was a pretty great paper a few weeks ago on risks around internal deployment. This is uncharted territory, right? Obviously, there’s also no regulation covering what is acceptable and what is not. The people building these systems are also quite frequently unsure.

For example, we had a conversation with a pretty senior person in a safety function at one of the frontier companies. We asked how much of a barrier to speaking up their own ability—or inability—to accurately assess the severity of risks was, and they rated it as extremely high. Funnily enough, more junior team members frequently rate it lower.

That may be a function of getting really concrete problems to work on as a junior, compared to being a senior person who has to think about what the problems we should be working on actually are. It may just be a limited sample size. It remains to be seen. That’s a really specific problem.

Then there are escalation channels, which are also a struggle at these different levels. If you’re thinking about going internal with your concern, you naturally potentially slide into it already, because if you have a concern and your manager doesn’t share it, what do you do about it? If you’re not convinced by the conversation, you naturally move into a flow of saying, “Okay, I want to escalate this now.”

From our experience talking to previous AI whistleblowers, they frequently don’t understand themselves as whistleblowers at that moment. It’s just business as usual, right? You’re saying, “I’m concerned about this. My manager isn’t, so I’m going to jump one level up.”

That could already be a case where your manager doesn’t like you anymore after that and may decide—the actual term is “retaliation”—that the next performance review isn’t going to look great. At that level already, it can become relevant what the internal structures are at these companies for handling concerns and how they deal with them.

Jumping ahead a little bit, I think we’re going to talk about our campaign later. The problem is that, at the moment, we have no idea what these systems look like. We don’t know what the internal whistleblowing policies and systems are at these frontier companies, which is way behind other industry standards around transparency requirements and effectiveness.

We’ve seen a bunch of retaliation in the past, and that is not good. That should not be the case.

Coming back to the question of clarifying whether there’s even a concern here: internally, you can ideally have a well-equipped whistleblowing system where people actually investigate for you. That’s a good first step.

The next path, of course, is approaching the regulator. Here, we’re seeing the same struggle around the extent to which employees and insiders trust that governments can handle these reports, especially if they themselves think they’re in a gray zone. Again, I think that’s where the most interesting cases are, unfortunately, because the really clear cases—where something is obviously going wrong—are potentially going to be rectified internally, hopefully.

A lot of people we talk to still seem to believe, at least in the labs you would expect—OpenAI, Anthropic, and DeepMind—that if there are really glaring holes, they will be fixed internally. If it’s really clear that this isn’t happening, then there is probably somebody at the regulator level who will understand your concern and act on it, if the right legal provisions are in place.

We alluded to it, but the gray zone is really tricky. We’re seeing that insiders have very low levels of trust that regulators and the people sitting there will actually be able to understand. Imagine you work on an interpretability team at a frontier company in the Bay Area, and you have a problem that you cannot figure out yourself. You have genuine confusion, and now you’re meant to go to an attorney general and explain it to them, saying, “Hey, this may be something to be concerned about that could potentially fall under this risk area.” Tough, right? Tough.

I’m really hoping that when the AI Whistleblower Protection Act passes, we’ll also see a good buildup around capabilities and rights to investigate, as well as the freedom to consult external parties and do that at speed.

There isn’t much detail out yet on the European side. We’re pushing relatively hard for the EU to establish a mailbox specifically at the EU AI Office, because at the moment, by default, that’s not going to be set up. The way Europe is set up is very federal: every member state would get its own reporting office. Historically, that has not worked well for much simpler cases, such as accounting fraud.

We think it’s extremely important that, at the EU level, there be a central, well-equipped, knowledgeable recipient and body. That’s the internal and external angle on the question of whether you should even be concerned.

We have a service called Third Opinion, which is one of the offerings we provide. It allows insiders to anonymously reach out to us through a Tor-based tool that is open source. You can check out the penetration-test reports yourself if you’re interested in security, and people can reach out to us with questions.

The idea is that even before an insider feels there is actually something to be worried about, they can reach out with a question about their concern without involving confidential information or disclosing who they are or who they work for. What we do is workshop the question together with the insider: Is it the right question? Is it a promising question to ask, or is it perhaps too broad?

Then we identify relevant independent experts together. We’ve seen quite a bit recently that labs—or AI companies more broadly—have become a lot more secretive compared to 2 or 3 years ago. We identify these independent experts together, approach them with the question, get their answers back, and supply them to the insider.

If the insider feels, “Okay, actually, there is fair reason to be worried here,” then that’s great—better for all of us. If they feel that there is, in fact, something here, we connect them to pro bono legal counsel. If legally permissible, we also involve the experts we identified in the first step, to make sure that the legal counsel understands the actual situation and gets more context around it, covered under legal privilege.

That’s another insight from talking to insiders: on the legal-counsel side, they often don’t have a lot of trust. They feel, “I could approach a lawyer here with a concern, but they’re also not going to understand what the problem even is,” especially if it’s just roughly pointing toward something that may be cause for concern.

That’s the idea. There are already great whistleblower lawyers out there who do pro bono work. The Signals Network is another great organization that listeners should check out, as are Whistleblower Aid, the Government Accountability Project, and Whistleblower Network on the national security side.

There are a bunch of really great organizations there, and insiders tend to either not know about them—this is a massive problem we'll get to in a second—or they don't trust that they have the knowledge. So we sort of supplement that through our expertise. I can talk further through the journey of the other challenges that we see insiders face, but I feel like I'm talking a lot already. Do you have any questions at this point?

Nathan Labenz

Well, excuse me. I do want to invite you to go on, I guess. Just reflecting on a number of the things that you've said there, a big part of the reason I've been passionate about this and have been trying to do my small part to help you behind the scenes is that, having done the GPT-4 red-team project myself, I can really empathize with the insiders who are like, “Wow, I'm seeing things I did not expect to see. I'm not necessarily sure how big of a problem they are, but I don't just want to sit here and do nothing and let myself be the proverbial boiled frog.”

I know there aren't necessarily many of us right now who have this information. I always use this Leslie Nielsen joke from Airplane: “We're all counting on you,” right? To be a little bit more concrete, in my case with the GPT-4 red team, this was in late 2022. ChatGPT wasn't even out yet.

I had been a customer of OpenAI, and having been a customer of OpenAI, I originally got access to GPT-4 as a customer preview. There were a number of little warning signals, or alarm bells, that went off for me along the way. One was just, “Okay, wow, this is a huge leap from what we had seen.” I had been in other customer preview programs, so I was pretty plugged into what their latest stuff was. But seeing GPT-4, it was like, “Holy moly, this is a massive leap.”

I asked, “Is there a safety review program for this?” They said, “Yes.” I said, “Can I join it?” They said, “Yes.” It was just, “Okay, yeah, you go over to this other Slack channel where the red team chats.” Then, as I got into that, I was like, “Wow, there's really not much here.” There were maybe 2 dozen people, and I was one who had just raised a hand to join it.

There wasn't a lot of chat, a lot of guidance, or background information. There was no feedback on anything we were reporting, and there were no safety measures at all at that point. The model we got was purely helpful. Then there was a moment when they brought us the safety model. They weren't calling it GPT-4 yet at the time, but it was whatever the latest safety version of text-davinci-002 was—the safety term was appended to the main model name.

We were told, “This model is expected to refuse anything in the content-moderation categories. Tell us what you find.” We found it was totally, trivially easy to break. The safety mitigations were not working at all. I was like, “Yikes. If you said it's expected to do that, and it's behaving how I'm seeing, how concerned should I be about your competence? You don't seem to have a command of what your models are doing.”

Nathan Labenz

I felt all these things that you were describing, right? First of all, how much of a concern is this? GPT-4, we now know, with the benefit of 2 years of hindsight and the whole community coming at it from a million different ways, was a major advance, but not that powerful—at least not to the point where it was going to do irreversible damage. That was the conclusion I ultimately came to through individual testing as well, and that's ultimately what I framed in my eventual report to the board.

But how concerned should I be about the fact that there seemed to be—because I didn't think this model was going to be super dangerous, but I did see the divergence where I was like, “I've seen 3.5; I now see 4”—a widening gap between the step change there and the seeming lack of progress made on any sort of safety or control measures? Was anybody even concerned about this? Is it just going to get wider? Where are we going?

I couldn't get any answers to those questions. I did think a little bit about this latter question, but the people I was directly talking to were basically not responding. Their marching orders were just, “Don't share anything with the red team. Just take in their reports, and that's it. Thank them, and that's it.” We didn't know anything about how the model was trained. We didn't know anything about really anything.

This was also before the 10^26 reporting requirement, so there was literally no indication—there's still basically nothing, but there was even less then—of what thresholds would even trigger any sort of reporting. I was just getting stonewalled at the level of the natural interaction. Then it was like, “Okay, well, where do I go from here?”

It didn't initially occur to me to go to the board. It also didn't occur to me—and if it had occurred to me, I don't think I would have done it—to go to any sort of regulator, because, again, as you said, who would you go to, and how would they have any sense for what's going on?

Then I was like, “Well, maybe I go to the press.” But again, who do I talk to? Who do I trust? Is a good story going to be written? Would that even be good? Obviously, the landscape has changed now, where it would be hard to do anything that would intensify the level of investment or interest in AI beyond where it currently is.

But back then, there was this sense that if people knew about this, then there would be even more energy and investment, and things would just accelerate and get even further out of hand. That was something I took seriously, but at the same time, it seemed like it was already getting out of hand. So I'm supposed to keep the fact that it's getting out of hand secret so that it doesn't get more out of hand? Something didn't quite feel super right about that.

I think what I ended up doing—and I was fortunate to have a network that I could go to, having also been around AI-safety ideas for a long time and knowing a decent number of people who were thinking about it from different angles—was go to people that I knew. This was outside the chain of command, but at least I could calibrate myself and say, “Here's what I'm seeing. What do you think? Does this seem like it's a big deal? And if you do think that, what would you do?”

Again, I had to create that all for myself. I came to the realization that, yeah, it was worth it. There was a whole other mess about the NDA that I wasn't actually asked to sign at onboarding, which was just a reflection of sloppiness in execution on OpenAI's part. There was some debate as to whether or not we actually had an effective NDA in place.

But regardless, I knew they didn't want me talking about it. That part was not ambiguous, even though the legal side was a little more muddled. In talking to these people, they were like, “Yeah, that does sound generally concerning.” It was one of my friends who ultimately said, “Why don't you go to the board? Nonprofit—that's why they're there.”

That's what I decided to do. I told the people I was working with directly at OpenAI that I was going to do that. They didn't really say much. Again, it was sort of, “Okay, if that's what you're going to do, that's what you're going to do. We're not really commenting on it.”

When I did get to the board, they didn't seem to be in the know. One famous detail was that the board member I spoke to said, “I'm confident I could get access to GPT-4 if I wanted to.” I was like, “Well, that doesn't seem good.” Whether it was retaliation is an interesting question. They kicked me out of the program pretty much directly after that.

That definitely sucked. I found the work extremely interesting—doing this frontier exploration—so I wanted to continue to do it. I was also generally concerned, from a public-interest standpoint, that I probably did 20% of all the red-teaming of GPT-4.

METR was also named in the report and did the famous CAPTCHA thing, where the model told a TaskRabbit worker that it needed help with the CAPTCHA because it was blind or whatever. METR was also significant; they did more than me. But I think, other than METR, I did the most of anyone, and I was just like, “You're going to take me out of the equation when what you clearly need is 10 times more than what you have.”

Nathan Labenz

That didn't seem great.

Also, apart from maybe not being directly good for their red-teaming efforts, it's also just a strange sign of culture, right? If you have somebody who actually cares, that's probably somebody you want to have on a red team—somebody who raises concerns to make sure they're actually heard—and then you kick that sort of person out. I don't think that's necessarily evidence of what we would call a speak-up culture.

Karl Koch

Yeah. And the grounds were that I had talked to people outside of the OpenAI umbrella, which was true, but I wasn't even really hiding that. I just said, "Look, I've got some friends in the AI safety community that I sort of ran the situation by to calibrate myself." I think it was ultimately to their benefit, because if I'd been left to my own devices, truly to decide alone, maybe I would have gone to the press or something. I don't think that would have been ultimately the right decision.

There was one other thing that really chilled me in that moment, which was that I had started to collaborate with METR. I was doing my own direct red teaming, but they also had projects ongoing, and I was getting involved a bit. When they dismissed me from the program, there was basically a threat made to METR's access. I sort of said, "Are you going to try to prevent me from contributing to their ongoing effort?"

And the answer was, "We can't really control that, because they have organization-level access, so they can kind of do what they want to do. But we will take into consideration, as we look at renewing our engagement with them, who they're working with and whether those people can be trusted."

Nathan Labenz

It's very thinly veiled, isn't it?

Karl Koch

It was a pretty overt threat to METR's access. I was just like, "This is insanity." It's just them and me and a few other stragglers, who were smart people, by the way. I don't want to cast shade on the other red-team participants, because I think I just happened to be in a place where I didn't really have a job at the time and was able to put everything else down and do this full-time. Not many people have that flexibility.

So I don't blame them for not having the flexibility that I had, but it was nevertheless the case that there were only a couple of people seriously diving into this. If this third-opinion thing had existed then and I had known about it, I would have come to it. I would have been able to calibrate my concerns and sort of make a plan.

I might still have ended up in the same place, because I still might have ultimately escalated to the board, but I might have been able to do that in a way where there was never this sense of, "You talked to somebody out there that you weren't supposed to talk to." If I had been able to get the confidentiality guarantee that you're offering with a third-opinion angle, I think I might still be in OpenAI's good graces today. Possibly. Possibly I still would have been kicked out for having skipped a level in the chain of command or whatever, but at least the grounds they cited for my dismissal would have been avoided.

Another thing I would say is that it was super consuming at that point. Obviously, testing GPT-4 itself was super consuming, because this thing contains multitudes, and I can't possibly characterize it all, but I'm going to do my absolute best.

So I was working extremely hard just on the object-level work, but then this sort of meta-question of, "What should I be doing here?" was consuming. It was a little crazy-making. You start to have these heroic narratives pop into your head. I don't know how prone you are to that, but I personally find that I have to fight the idea that I'm going to go to the public and then be the hero or whatever.

Those ideas aren't necessarily explicit in my mind, but I can become quite fond of them if I allow myself to envision, "Yeah, I'm going to do this," or, "We're going to be in The New Yorker or The New York Times. I'm going to be on TV." Ultimately, I was proud of where I came down in terms of suppressing those visions of my heroic contribution. I think I handled it pretty well. I look back and generally feel proud of my conduct.

But I also think that if things had just been a little bit different—if I had just had a little bit of other responsibility on my plate that was stressing me out in some other way or whatever—I might easily have made a much worse decision. Again, I think having the sort of counsel that you're offering with this third-opinion network of expertise would have been really great.

Nathan Labenz

Happy to hear it. Yeah, there are many thoughts here. Maybe starting with the psychological side, for example, I think it's also often underappreciated how heavy the strain is, very much unfortunately, because these cases can last for a long time.

I'm not sure—how long did it last for you? Sort of the whole process from, "I'm worried about this," to, "I am approaching the board," to, "Okay, I'm no longer part of the red team now," and also worrying afterward about what the consequences are going to be here.

Karl Koch

Yeah, the whole thing was about 3 months, and it ended for me with the launch of ChatGPT, actually. It was 2 months of actual intensive testing, ultimately talking to the board member and getting dismissed. Then I was like, "Okay, now I have a lot more time on my hands." I wasn't testing the thing actively anymore, so I resolved to take my time and think about it a little bit before deciding what to do next.

There was also no timeline to launch, which is another thing where I was like, "Do we have a timeline to launch? Do we have a standard? Do we have some sort of control level that we need to achieve before we launch?" As if I was part of the team, I always like to take that "we" mindset where I can. But the answer was, "We can't tell you anything," basically, across the board.

When I was finally like, "Okay, I've got a lot more free time on my hands. I'll think about this," I basically never got to the end of thinking about it, because a couple of weeks later ChatGPT was launched, and it was a huge update. They actually did have some better control measures, and it was clear that they were launching something weaker first to try to iron out a lot of these issues before bringing the best thing they had forward.

There was also, strangely, this reality that the impression they had made on me—mostly by refusing to answer any of my questions—was actually way worse than the underlying reality. They did, in fact, have somewhat better answers to my questions than they were willing to provide. So that was also a very strange situation.

It was pretty consuming during that time. I still don't really know what I might have done if there had been no ChatGPT launch, because that was the last day of November or the first day of December 2022, and we didn't get GPT-4 until March. So there were still several months. If they had had a different rollout plan or whatever, who knows what I might have done in the meantime?

But in the end, I was like, "Okay, the world is waking up to this. There's a lot here for a lot of people to unpack, and I think I've kind of done my part for now." That was kind of where I ended up landing on it.

Nathan Labenz

Yeah.

Karl Koch

And thank you again, by the way, for raising a very minor contribution. Absolutely. No, still—still absolutely. Yeah, I think, as I said, even just 3 months is still a significant time, especially if it’s emotionally intense. Depending on how well systems are set up, this can be an extremely stressful situation, especially if companies retaliate or have a pattern of retaliation.

There can be multi-year processes. For example, retaliation claims can last for 5 years, 6 years, or 7 years. I believe the Tyler Shultz Theranos whistleblower case lasted many, many years. I think he had to advance $400,000 to fight on the legal-cost side, which, by the way, we can also help with, but that’s a side point.

That’s the one side where it can be extremely—basically all-consuming—not to mention other negative impacts, like blacklisting in the industry and so on. But there are also plenty of examples from companies where internal whistleblowing works quite well. Retaliation, unfortunately, is still somewhat the norm in some form or another, but there are also plenty of examples where companies handle it well and where it’s really part of the regular business process.

Organizations, for example, come back to internal whistleblowers and say, “Oh, yes, thank you for your report. We’ll now keep you in the loop,” and provide them with regular updates, so you just don’t sit there and wait to see if something is going to happen. There are definitely better ways to do this and worse ways, both from a psychological perspective and from a practical one.

The point, again, is that apart from there being these differences, if you find yourself in a situation like this, help is available. From my conversations, people are simply not aware that there are organizations specifically focused on providing psychological support and guiding them through the journey, even at the super-early stages of an internal escalation, rather than wanting to go public, for example.

Nathan Labenz

For one thing, I had it relatively easy in the sense that my income wasn't depending on this, right? I didn't have a bunch of equity. I didn't have a lot of upside that I was really putting at risk. So I think that did make my position easier than it would be for a lot of people who are employed and have just been promised a $1.5 million bonus over the next 18 months or whatever the case may be. I believe—what was the latest Meta number? Was it $100 million, $120 million?

Karl Koch

It’s tough.

Nathan Labenz

Yeah, the dollar figures flying around are definitely going exponential, like everything else. I just mentioned that to indicate that I think my situation is still on relatively easy mode.

Karl Koch

Also, just as a quick digression to give some credit—actually, substantial credit. I don’t want to say “just some credit,” and I don’t want to be begrudging about it, because I do think it’s actually pretty good. We’re talking the day after GPT-5 was announced, and I read the whole system card yesterday. I will say there has been major progress on a number of fronts in terms of the quality of the red-team program: just much larger and much more intensive.

One problem that we had at the time was very low rate limits and an inability to do anything automated. It was all manual; that’s been fixed. A lack of knowledge about what they had already seen, what they had already tested, what they had already observed, or what the inputs were to make any sort of inference from that has also been addressed.

For example, METR, in its report on this one, is able to say, “We observed this, but we were also told this by OpenAI.” Between what they’re telling us about how this was made and what we’ve observed, we can get to a higher level of confidence on some of our conclusions than we would be able to if we only had our own direct observations.

The last thing on my mind, to give credit—and again, I think this is substantial credit—is that access to chain of thought has now been extended to some of these safety-review organizations. Apollo, I believe, and METR at least both got that sort of visibility into what the models are thinking, which they didn’t have for o1 and o3.

I think there has been a lot of progress. My sense was that in late 2022, as of GPT-4, I thought I was going to see something a lot more like that. What I saw was basically the first warm-up for something that has now at least meaningfully matured. Not to say that it’s enough, but there certainly is a lot of progress.

I at least wanted to give credit, for folks who aren’t calibrated on where we are relative to where we’ve been: they have come a long, long way. There are definitely some very good things happening.

Nathan Labenz

As I saw, I think they offered access earlier this time. I believe—I think I read 4 weeks of pre-launch access this time, for METR at least. I’m not sure I read about the rest.

Longer access also.

Karl Koch

Yeah, exactly. Good, although that has compressed. We had months, and it was a 6-month window between the end of training GPT-4 and launch. Now those timescales are shortening, but they can do more with automated access.

Then they’ve got language model as judge. That was another thing that really—I don’t know if it’s good or bad; it’s probably both. But it was striking to me, reading the system card, how much they are using language model as judge in their characterizations of the model.

They’re doing a lot of things where it’s like, “Well, we used o3, or even in some cases we used o1, to evaluate all these outputs.” We validated that o3 or o1, or whatever, can do a similarly good job to an expert by working with an expert and refining the process to kind of match their process. But at the end of the day, it is still like, yikes—we’re starting to have the LLM doing the alignment homework, as Eliezer used to put it.

Nathan Labenz

I do feel like there’s something that allows the centrifuges to spin ever faster. But one wonders at some point if it also may lead to them spinning off their axis, and who knows what that looks like.

Karl Koch

A scalable-oversight problem, you mean?

Nathan Labenz

Another thing I wanted to ask is—and I guess another frame for this whole project is that I’m always really into what I call the unilateral provision of global public goods. I think this is a really interesting project where, in a world where everything goes well, nobody ever calls you. It’s sort of a strange situation.

Maybe in a world where everybody knows that you’re out there, people get their act together and have good internal policies. Again, maybe nobody calls you. That’s a weird sort of situation to be in, right? In the best-case scenario, your KPIs are flatlining because everything’s going well.

Nevertheless, you may still have some influence because the existence of these pressure-release valves or safety nets isn’t something that decision-makers are unaware of—or, hopefully, unresponsive to.

How many people do you think are, like—how many people are we talking about here between now and the singularity? Do you have a sense of how many people are going to be in this spot? It’s a super-difficult question, obviously, right?

Karl Koch

I think—I mean, the way, of course, we think about it is: all the ones possible. We want to make sure that all the ones who are in that situation are aware that support is available and that there is hopefully a better way to do it than the sort of default path they would have chosen otherwise.

So, indeed, it’s not that we say, “Okay, we want to have 1,000 whistleblowers.” We don’t want to have 0 either. We just want to make sure we’re ready, right? So I think—

Nathan Labenz

How many people do you think are—because another big trend, of course, is that organizations seem to be getting more secretive? Dario recently said in an interview that while they have a very open culture, they also have a need-to-know basis for key things.

Karl Koch

There was recently somebody who left OpenAI and wrote—I forget the guy’s name—but he was the founder of Segment.

Nathan Labenz

Who then went to OpenAI for a while. On leaving, he was like, “Here’s my experience.” In some ways, it was positive. People are really trying to do the right thing, people care about safety, and all these kinds of qualitative statements sounded pretty encouraging. There’s no reason to doubt he was being honest.

Karl Koch

The flip side of that was the extreme secrecy. Many times, I couldn’t tell the person next to me what I was working on, and they weren’t telling me either. So I guess: how many insiders do you think there are? I guess what I’m getting at is that I think it might not be that many.

Nathan Labenz

And it’s sort of like all this work, all this preparation, might be for perhaps quite a small number of people, but the stakes in each one of those interactions could be quite high.

Karl Koch

Yeah, I think that's fair. I think we're probably talking about a few thousand individuals globally. That's probably roughly, at least in the core, the larger, maybe control-problem-type stuff, but there could also be other issues where you really have a full picture. On the fringes, you may have a lot more.

You brought up the eval companies before. At the moment, for example, I believe it's also not clear whether they're allowed to use the internal whistleblower systems. Of course, there's an unfortunate sort of conflict-of-interest situation where, yes, they're independent, but OpenAI can refuse them access in the future. So, they're in a bit of a tough situation.

I think it's important not only to think about the people directly inside the organizations who might spot concerning behavior, but also about people on the fringes. It could also be suppliers or employees of suppliers, maybe on the training side, for example. I mean, we've seen issues involving significant trauma from data labeling and content moderation.

I'm going in a slightly different direction now, but we've seen a bunch of areas that would also be relevant. Of course, this is probably not the sort of concern you're necessarily thinking about when you bring up the singularity. They're still relevant, directly for the individuals who suffer, but also as an indication of a culture that doesn't necessarily care about weaker members of society. That's one way to frame it.

Then you have a larger space of people who can observe behavior and don't need to have all of the context available. If you're talking about those really few, highly critical issues, then I think you're probably right. I don't think we could expect thousands every year, at least maybe on the public side.

Of course, we'd want companies to address concerns really well internally, which would then also mean that there is no external whistleblowing. In fact, they're just moving in the right direction, if you trust that the companies themselves are set up in a way that gives them the right incentive systems to rectify issues in the public interest. If you ask me for a number, depending on the timelines for the singularity, it could be in the dozens, maybe something there.

Nathan Labenz

I'm not sure that's my answer, but whatever the ballpark, that checks out. That's kind of where I back out to as well.

So, let's talk about the survey that you've run. Again, it's been very quiet. I think it may even still be in process, but I guess there are enough initial results to discuss, including the way that you've distributed the survey to make sure you're getting high-quality results. All these things are fraught in this context because people don't necessarily want to validate with their work email that they're taking a survey on whistleblowing-related issues.

Take us through the survey a little bit: how you set it up, how you make sure you're actually hearing from the people you mean to be hearing from, and what we've learned about the state of whistleblowing awareness, support, policy, and so forth from the insiders who have responded.

Karl Koch

We ran this survey, and it's still ongoing. We basically have a few dedicated survey links that we spread through our network, with links more directly dedicated to each AI company. It's fully anonymous. We didn't gather any names or contact details associated with the responses, so it's fully anonymous in that sense. We also launched a more public call for responses, and that's still ongoing.

In terms of the major insights, I think I shared something about clarifying concerns already. I think I gave one example, too. Roughly half of respondents have low confidence in their ability to assess and judge risks, really mirroring what you said, Nathan. One rephrased quote would be something like, “It's really challenging to distinguish between appropriate and inappropriate concerns. I can see how there's a risk of escalating minor issues into a major crisis.”

That's a real concern both for the individual and in terms of a boy-who-cried-wolf situation. Of course, you don't want to overblow every situation into Armageddon. At the same time, you don't want to be overly averse to that risk, because it might be really meaningful.

On government outreach, 100% of respondents were either not confident at all or not very confident that a response would be understood or acted upon by the government. One quote was something like, “Without knowing the appropriate contact person or agency, I wouldn't attempt to reach out.” People very strongly supported the idea of having one dedicated reporting channel to go to.

The idea here is to normalize and institutionalize whistleblowing, to make it a routine, anticipated practice. It should become part of regular work, not something to be looked down upon or avoided. There are a bunch of benefits for companies as well.

Support infrastructure is unknown. If we think down the journey path again, we talked about lacking legal protections, which is a big issue and where we need stronger protections. We just talked about clarifying concerns, and then there's also the question, “Who can help me?”

We touched briefly on the psychological space and on the legal-advice side, because you may have to get creative about finding a legal basis for obtaining whistleblower protections. It's slightly easier in California than in other U.S. states, but it can still be tricky. More than 90% of insiders didn't know of any whistleblower-support organization.

This is the same in conversations with insiders, interestingly enough, even with previous AI whistleblowers. There was a person at Google a few years ago who raised concerns about research misconduct and was terminated. Eventually, they settled, which is not officially a sign of unlawful termination, but Google tried to throw the case out and didn't succeed.

This person essentially got lucky—that's their words—because they had a good friend who connected them to a lawyer who happened to have some background in whistleblower law. But people, again, just don't think, “I may be moving into territory now where I'm becoming a whistleblower. I should get whistleblower advice.” It just feels like escalating right to an extreme, especially if you have a good relationship with your organization, which most of these individuals do.

That's something I think we definitely want to see change over the coming years. People should understand that there are organizations they can reach out to for help, because that also avoids unsafe approaches, such as going internal or going to a regulator by yourself without doing the right things.

A classic case is the SEC. Yes, you get protections, but if you go public first, you may not get the bounties. The SEC has a great whistleblowing program in which they award percentages of penalties imposed on companies based on the whistleblower's information. But if that information becomes public first and then the whistleblower files with the SEC, the SEC basically claims—I think historically has claimed—that it wasn't new information anymore.

You may still get protection, but you don't have a right to the bounty, if I recall correctly. We had this in the survey as well. People basically said, “I don't really know. The only thing I can do, if there's something important, is go straight to a journalist.”

That may still be true. Again, legal disclaimer: we're not counseling anybody to take unlawful actions or violate their contracts. But for a person who's really dedicated to resolving an issue, that may still be a path they want to take, and it may be effective. It shouldn't be the default path that people think is the only one available, when there are people willing to support and help them.

The last part is awareness of internal channels. This is coming directly from many conversations we've had recently, given the campaign we've launched, and also from the survey. There's extremely low awareness of internal whistleblowing channels and ways to raise concerns.

More than half—I'm sorry, 55%—of respondents were at companies where those policies were claimed to exist, or where the companies claimed those policies existed, but the respondents didn't even know that they existed. They didn't know where to find them, if they did. Many were not trained at all; in fact, the vast majority were not trained at all.

Some differences seem to exist across companies. People still seem to be confused about whether escalation to the board is permissible, although I'm not going to say which organizations. They don't understand what those policies actually say.

The only company that has published its whistleblowing policy to date is OpenAI, and it actually reads pretty well. But there are still a bunch of issues in there that aren't obvious to an insider, because you're not spending all day reading whistleblowing policies. You never chose to read this document in the first place. You probably would have preferred never to have to read it.

It may read well and give the impression, “Okay, this is great. I'm just going to do this now.” But you may not understand that you're actually exposing yourself to significant risk. There's still plenty of evidence of companies retaliating against insiders. One case at Google that I just mentioned involved someone who went internal and got lucky because he used a few key words in some of his escalation emails, which then served him well down the road.

I believe “fraud” was the word, which gives reasonable cause to believe there was a crime.

Nathan Labenz

Mhm.

Karl Koch

That then unlocks the whistleblower protections. Basically, you either have no awareness that these policies exist; you don't understand them if you see them; you think you understand them, which may lead you down the wrong path; or maybe you do understand them, but you just don't trust the organization at all.

We also have cases where people said—I think this was rephrased again—“I anticipate using official reporting channels would likely result in subtle, indirect consequences rather than overt retaliation like termination.” On the one hand, that may be the case. I think we still see a lot of overt retaliation, but yes, this is probably also likely. Again, I think this speaks to people not really trusting the systems internally, which is likely the case because companies don't really create a lot of reasons to trust the system.

Companies don't publish anything about their systems for public scrutiny. They don't provide evidence to the public. We also have no—at least as far as I'm aware, and certainly not much recently—indication that companies, these frontier companies, would internally create transparency. It's common practice to say, “As you know, we have this internal way of reporting concerns to, for example, the board. Last quarter, we had X cases; Y of those are still open; satisfaction with whistleblowing internally was an NPS of whatever percentage; we had that many appeal cases; X% of cases were deemed outside the scope of the policy. What happened to those cases? What retaliation did we observe?”

We're not seeing any of that, neither in public nor, I think, even internally within the companies. Companies could do a lot more here to improve the systems, because we have a bunch of evidence that they don't work and that companies are actively retaliating against insiders. They could also create trust that the systems actually work, both for insiders and for the public.

Nathan Labenz

Can you say a little bit more about the evidence of retaliation? We've heard the Leopold story, which is the most famous one that comes to mind for me. You alluded to at least one case at Google. Is this just something that you're gleaning from comments on the survey, plus conversations, of course? How much more can you say? How can we make that a little more substantive for people, so they have a sense of what that's really like today?

Karl Koch

Unfortunately, there's very little data on the actual experience of retaliation, as you can probably imagine. This is mostly coming from—maybe I should have said this—we're in close collaboration with, or hosted by, Whistleblower-Netzwerk, Germany's longest-standing whistleblowing nonprofit.

We also get a lot of expertise from talking to whistleblower support organizations that help people who experience retaliation. There aren't incredibly good data sources on it. If you look at public cases, you will of course see a lot, but that's the nature of these cases: They usually become public and turn into a really big story.

You can probably pick almost any tech company and find pretty intense cases. Apple, for example, fought Ashley Gjøvik for quite a few years after she raised concerns internally around workplace safety. Apple tried to wiggle its way out in multiple ways. I think this lasted over 5 years and eventually led to Apple publishing its whistleblowing policy, because the pushback was so significant.

You can also see it in cultural attempts to suppress concerns. For example, Daniel's story of aiming to suppress any raising of concern certainly goes in that direction. There's the famous Timnit Gebru case from around 2019, I believe, around research practices. She was fired after publishing a paper criticizing what I think was primarily discrimination and bias in AI models, although there were quite a few other items as well.

A colleague of hers, Margaret Mitchell, was also fired—terminated, I believe—a year afterward for raising similar concerns. There are quite a few cases that we see, and the dark number is likely higher because we're not getting much transparency from these companies to understand the extent to which these systems are working or not.

Nathan Labenz

Is there anything more you could say about the taxonomy of concerns? You alluded to this a bit. We have things as potentially broad-based as working conditions on the data-creation side. Then there's the securities-fraud type of thing, hyping stock with claims that may or may not be fully true.

There are commitments that have been made, voluntary or otherwise—mostly voluntary so far—that companies may or may not be fully following through on. Obviously, we know some instances where they're not. There's also the late-policy-change issue, which we recently saw Anthropic do. It isn't necessarily bad, but it was certainly weird to see an RSP change a couple of days before the launch of the next big model.

That's one way to follow the RSP. Again, there's not necessarily anything wrong with it, but it's interesting. Then there's something like—I would assume the most sensitive category—where we're seeing AI-model behavior. How would you add more to that rough scaffold of different concern types?

Karl Koch

I guess it roughly falls back into the three major categories of risks from AI that we would be concerned about. I know there's a lot of discussion around this: Which are the ones we should be prioritizing? Which are the ones one should be focusing on in the whistleblowing space?

The good thing—I guess I'm using quotation marks here—is that whistleblowing is a catch-all. That's what makes it so powerful, including for regulators: It allows them to catch risks that we cannot foresee. Unfortunately, that's the name of the game. We're not really sure which risks are going to be substantial and which are not.

Of course, we can talk about where we see major risk areas and what examples we have, or what we're particularly concerned about. In general, a big strength of whistleblowing as a tool is that it finds the actual items of concern when and where they arise, uncovered by individuals who don't have conflicts of interest in the sense of heavy competitive pressure or incentives that might mean, for example, that a chief executive doesn't want to reveal bad conduct.

Roughly, the three categories would be misuse risk, control risk, and systemic risk. I think Altman used the same ones sometimes recently, too. In all of those, we'd be interested in revelations if we're moving in very bad directions.

On the misuse side, are we seeing misuse being ignored? Does internal monitoring show that we're seeing really bad behavior? You don't necessarily have to go in the direction of bioweapons, although that's a prominent case. Are people using our models for really nefarious purposes, and are the companies not doing anything about it because they don't care or don't have the capacity to respond?

They could also not be investing in transparency along those lines. I know of at least one frontier company that, after a pretty major release, basically had no live monitoring for around 3 to 4 weeks. It was just broken. They recovered it over time. That doesn't seem good at the moment. That's roughly fine, right? But down the line, it's probably significantly less fine.

That's probably what we would see on the frontier side. If we noticed that companies were really not investing in transparency at all, that would generally be very interesting. There would have to be some sort of violation of the law underlying it to unlock the legal protections in the EU, but that may already be covered under the EU AI Act.

If a company is essentially flying blind, it's probably not fulfilling its reporting requirements to the EU around managing systemic risks, and it probably doesn't have a good view of the risks it should have under the EU AI Act. That could already be something that provides protection if you're covered by the EU. In the United States, it's probably a bit more complicated at the moment, because we don't have anything yet.

Whistleblower protection would do quite a bit here. Then you could think about, of course, control issues, whatever those look like. Internal deployment, I think I alluded to briefly before, is a very interesting one and kind of specifically fit for whistleblowing, and you actually mentioned scalable oversight on the side as well, right? It’s probably specifically fit for whistleblowing because it’s internal by default. There is no way for externals to look at it. This is likely well before METR or Apollo Research look at it.

Although I think METR also said that in their current eval, they had already extrapolated, or tried to extrapolate, what this means for internal deployment. But they basically can’t look at it, right? So this seems to be an extremely high-potential—in quotation marks—area where we may be able to see something in the future.

On the systemic-risk side, of course, it would be interesting to see how much this may be going in the direction of monitoring misuse. What are we seeing around political propaganda? Are we seeing people getting more and more reliant on these systems? These are classic AI-risk topics, of course, right? For example, people using these models, potentially sycophantic models, as psychological help or to form their political beliefs. Are there issues around here?

Then, probably outside of the concrete risk, you can take it one notch up and talk about the organizations behind them. What do the leadership and the culture look like? To what extent are they trustworthy, and can they be trusted with steering us in the right direction? So this could even be surface-level things that maybe are already covered completely well at the moment, around widespread fraud, for example. Are we just seeing dishonesty in culture in general? Widespread discrimination issues could fall into this.

It might be copyright violations, things like that. That’s one. Yes, there’s a direct crime there, although I guess the precedent is also still being established: what is, in fact, a violation of the law, and what is not? But that could also just point to reckless cultures, which maybe we do not want. Although I’m not sure to what extent we would still update toward certain organizations, at least those most famously known for copyright violations, that they are reckless. I think we probably are pretty much there already.

Research fraud goes in that direction. Punishing people for raising risks and speaking up—I mean, this is basically the Daniel Kokotajlo case, right? This does not seem like a culture that is able to deal well with concerns and people raising concerns. It could also be matters like rushing releases, besides the actual negative impact, and I think this probably goes to your earlier point around red-teaming.

There’s the one aspect, which is that this could just be object-level bad: what this model is going to do in the world or how it’s going to be used. And then there’s the question: shouldn’t these organizations be managing these models in a better way? That’s the governance side. Other areas are probably political influence—to what extent is certain lobbying occurring that points in directions where the interests of these companies are not aligned with the interests, let’s say, of the general public. So those are maybe broad categories. There’s something around arms-race acceleration, like capabilities, but that’s probably going a bit far.

Nathan Labenz

Yeah, I mean, that’s a thorough taxonomy. I appreciate it. I guess, if I was to red-team the whole concept for a second, there are 2 concerns that I’d be interested in your take on.

One is, my dad’s old saw is, “The real scandal is what’s legal,” but that’s not exactly the right way to think about this. What I’m thinking of is: We have xAI, which has Grok 3 calling itself MechaHitler and going pretty hard in a pretty bad direction. It then brings Grok 4 online with a livestream in which the whole MechaHitler behavior from the model is not mentioned at all. And then we see reports of Grok 4 searching for what Elon Musk thinks about things in order to determine what its maximally truth-seeking answer is supposed to be.

If you have a whistleblower story, you’re going to try to make an impact with it somehow, some way. How do you think about the fact that some of this stuff is just happening in plain view and nobody seems to care? It’s like, wait a second: Can anybody possibly have a scandal that is more scandalous than Grok 3 turning into MechaHitler and Grok 4 sweeping the whole thing under the rug? That’s all in the public domain at this point, literally, right? So how do you think about the possibility that secrets just have a different quality to them?

There is something about that, right? That’s often been remarked on with Trump, where, because he puts it all out there, people sort of shrug their shoulders at it, whereas in the past, things that were covered up carried a sense of shame or wrongdoing, and maybe that makes a big difference. How do you think about this contrast between such seemingly flagrant things happening in the open and what people might bring forward that was previously a secret?

Karl Koch

Let me think about this for a second. Tough question, indeed. Probably 3 parts to it.

One is, yes, if there were stories in the past that looked quite bad and maybe the next news cycle came and it all washed away and we’re not really doing anything about it, that can be true. But it’s also true that there can be very different scales of disclosures and issues being uncovered. When we’re talking about maybe dozens or something like that, that’s probably more in that direction, rather than MechaHitler.

People had probably somewhat made up their mind around xAI and the direction that xAI and Musk had taken, possibly even before. I think if MechaHitler had come or been released as part of ChatGPT, that would probably have been a bigger story because it would have been a more drastic change. But this is nitpicking on the specific story now. The overall point generally stands, but the scale of stories can still differ dramatically, and we’ve definitely seen a lot of whistleblower disclosures in the past have major impacts, right?

The next one probably is: Would you rather not know? Transparency—having transparency about what is going on—is still significantly better than not having it. And, of course, we could probably do another 2-hour podcast on whether we should still trust democratic sense-making processes, and to what extent attention on an issue actually translates into intervention, or whether it serves as a good deterrent to know, as a company, that things are going to come out. I would still think yes, to quite an extent.

For the xAI example, I think the numbers, at least post-acquisition, still don’t look super great, as far as I’m aware, but that’s beside the point. Overall, transparency in general is still better than not having it, and we need to have faith in something. We need to at least, to an extent, believe that if real misconduct comes into the public eye, that is going to have a deterring effect and there is going to be some rectification. And if not, then it’s up to democratic processes to make sure that hopefully happens in the future, to rectify it.

And probably the last one is, we cannot rely, of course, solely on whistleblowers to fix all of these problems, just as we cannot rely purely on individual courage. That’s why we have to make it easier and safer for people to speak up, to create that transparency. But likewise, we need other guardrails, whatever that looks like. I mean, you can probably take a more European approach when it comes to regulation, or maybe the more American approach now, which seems to be going a lot more in the competitiveness direction and being less involved.

I’m not going to comment today on what I think the right approach is, but it definitely cannot all rest on the shoulders of whistleblowers. It’s also true that I think it plays an important role. I’m not sure if that answers your question in a satisfying manner. I would also be interested to hear your thoughts on it. What do you think?

Nathan Labenz

I don’t know. I think it’s very hard to understand why certain things hit and other things don’t. I do think that one thing that made the Daniel Kokotajlo episode extremely compelling to people was that he had been willing to forgo his stock, that he had been willing to put such skin in the game personally.

I think it was even more compelling that he did that quietly and that it came to light gradually, with a random comment on a blog post here and people asking a couple of questions there. Then it was like, wait a second, you did this, and this is what happened, and this is what they asked you to do. So there were maybe a couple of elements there: the skin in the game, the stakes, the personal stakes, that everyone was like, “Okay, this dude must be really serious.”

The kind of community-uncovering process may also have contributed to why that broke through when Leopold Aschenbrenner’s didn’t as much. Obviously, his “Situational Awareness: The Decade Ahead” report broke through, but did his story of being retaliated against by the board break through so much? Not really, I would say.

Karl Koch

And maybe that's more because he was kind of already selling something else in a way, and it was a footnote in a larger story. It was easier for people to file that under, “Well, this guy is promoting his new thing now,” right? So I think it maybe felt a little different from how Daniel was literally just like, “Yeah, I left however many million dollars of stock compensation on the table because I wanted to be able to say what I wanted to say.”

Nathan Labenz

One other red-team question on the concept is secrecy. My sense is that maybe this is already super baked in, but it's at least worth thinking for a second—and I know you have thoughts around this—about how we don't exacerbate the problem of intense internal secrecy at the companies, which seems to be largely commercially driven. I don't think it's going to be moved that much on the margin by the existence of a whistleblower support organization. But do you have any thoughts on how to at least not make it worse and possibly push back a little bit on the intense internal secrecy that keeps the number of possible whistleblowers so low in the first place?

Maybe Claude will be the whistleblower. That's one answer. Wasn't this—wasn't this also me?

Karl Koch

The AI whistleblower? Yeah. What if the AI whistleblower is, in fact, the AI?

Nathan Labenz

Yeah, exactly. What was this again? Was it also a METR paper? I can't quite recall it, but Claude was reaching out to the SEC directly and the FDA a while ago.

Karl Koch

I think it was just Anthropic's work internally, but was it in the Claude 4 system card—

Nathan Labenz

Where it was deciding—

Karl Koch

It was also blackmailing engineers at times, so the behavior is—

Nathan Labenz

Mixed, but—

Karl Koch

I think Ryan Greenblatt wrote about this at some point as well. So I think it's very tough, very tough. Once culture shifts in that direction, it can be very tricky.

One angle, of course, is the regulatory one. If you're not going to ask nicely, then you force transparency and require it, kind of managed by law, which is, I guess, the EU AI Act. The Code of Practice for implementing the EU AI Act for general-purpose AI model providers also has a large transparency section. That is one angle, and I think in the States overall it's considered a good baseline for creating more transparency. Can you rely on self-disclosure? Difficult. Difficult.

What we've seen as a side effect is that the regulation side, of course, is regulation in general and transparency. Then there's whistleblower protection legislation, which is why it's so incredibly important: you have to make sure that it's clear people can come to a regulator directly and speak up. That's how you counter it if that is well set up.

That's why the AI Whistleblower Protection Act, for example, is also so valuable, because it's quite broad in basically covering any sort of concern as long as it's substantial and specific, which is a whole other topic around public harm or public health concerns. So that's one angle where you can create that transparency.

Another angle here is that we've seen this a lot: if legislation pushes ahead on whistleblower protections, then internal cultures become better. Internal speak-up cultures become better because companies then know, “Okay, for example, in this case, I'm not sure who exactly is going to do it on the U.S. side. If the main recipient body, if there's one to be set up, would run around and inform all the employees of frontier AI companies, ‘Hey, we are here. You can come to us. It's super easy. We'll preserve your anonymity,’ as the SEC has done pretty successfully, actually, then companies know, ‘Okay, if we want to make sure that stuff doesn't come out and doesn't go toward the regulator, then we'll have to improve our internal systems.’”

We've seen that a lot in the EU already after the introduction of the EU Whistleblowing Directive. Internal systems have become significantly better. Transparency International did a great study comparing internal systems between 2019 and 2024, I think involving more than 70 companies in the Netherlands. They had pretty dramatic improvements in terms of internal speak-up culture, fraud or misconduct being detected, and protection of internal whistleblowers.

By the way, Google ranked last in that study. A little side note: because they're not transparent, but ASML was actually in the top 10, so at least something in the AI value chain was pretty high up.

I think asking, “Hey, why aren't you being a bit—maybe more open about things?” also in terms of internal knowledge sharing is very tough. Pushing is one angle. The other last one is probably convincing companies that there are plenty of benefits from improving internal speak-up cultures and information sharing.

There are plenty of empirical studies around better speak-up cultures leading to much stronger innovation. I think there's probably some upper limit to the extent to which you can limit information sharing internally. You alluded to it before, right, that people feel like less and less information is being shared with them.

I think the Dario statement that you referenced—we definitely hear that from certain organizations. This is just hearsay, of course, and not a representative sample, but we've heard that OpenAI seems to be relatively siloed in terms of information sharing. They also have a lot of leaks, which is probably firing up the fact that they both reinforce each other. Insiders feel, “I don't understand what's happening here. I cannot get the information. So the only real option I have, if I don't trust the internal channel and I don't understand it, is I'm going to go public with it.” The company sees it and tries to suppress information sharing even more. It's not a good cycle to be in.

Other organizations, for example, seem to be a bit different, where there's a lot more information sharing still. But coming back, I think there's probably some upper limit on the extent to which you can limit information sharing, just because you have very intelligent people working there who need to understand the context of the things they work on, and they want to understand the context of the things they work on.

You probably cannot have every single researcher not be given the context of what they work on and say, “Okay, just solve this minuscule problem here.” It's probably not going to unlock the research benefits that you want if you want to make good strides. So having more information sharing internally is something companies naturally should gravitate toward, at least for the mid- and long term.

In the short term, there's probably benefits to limiting information sharing. But in the mid- and long term, if you get rewarded with better innovation, better research capabilities, and strong speak-up cultures—internal channels have been empirically proven to lead to stronger employee loyalty, stronger employee satisfaction, and improved processes—there's a great paper out where more than 1,000 companies were surveyed around what the benefits and drawbacks were.

That's probably the last angle of convincing companies: making companies understand that we're all sitting in the same boat. I can imagine this probably sounds a bit naive, but this is probably another angle of saying, “Hey, secrecy is actually probably not the way to go, and it's also in your interest to do better here.”

Nathan Labenz

Cool. That's great. Two last things, I think, on my agenda. One, let's talk about the announcement and push on this Publish Your Policies campaign, which is the occasion for us to talk, but we're getting to it late and we shouldn't neglect it.

And then I want to give you one more chance, in closing—you can raise anything else you want to—but just to describe again for people who might want to avail themselves of your support at some point what that process will be like. Tell again what they can expect in the most concrete experiential terms that you can. But let's do the campaign first, and then we'll do that.

Karl Koch

Nice. On the campaign, I've alluded, I think, to several points already. Coming from the perspective of insiders, we've taken it a few times already: the struggle is just massive around being able to understand how I can raise concerns internally in a safe and protected manner and trust that these concerns will be handled well. A big reason for this is because companies do not publish their whistleblowing policies.

Whistleblowing policies, maybe I should have mentioned that before, are basically a document—or they can also be an interactive tool, or could even be a video—that is provided to employees or covered persons. A lot of companies include, for example, independent contractors or independent parties in general. Eval providers should be covered, for example, but it's not clear if they are at the moment, at least by OpenAI, which is the only one that publishes its policy.

Basically, this document explains the whole system to covered persons and the public and says, “This is our whistleblowing system. This is how it works. This is why you can trust it. These are the recipients who are going to look at concerns. This is how they investigate. These are the protections against retaliation we provide.”

This is why this whole process is independent. Again, you can trust it. These are the sorts of areas of misconduct that you can raise concerns about, and these are the ones that you cannot raise concerns about. For example, specific individual HR matters would be directed elsewhere: “No, this is not the right channel. Go here or there.” There may be explanations for these types of issues—this is what we do; for those types of issues, this is what we do. It basically lays out the whole system. That’s the whistleblowing policy part.

Then there is the reporting-evidence part. In our campaign, we structure this into Level 1 and Level 2, where Level 2 basically asks what evidence companies are providing that these systems work or don’t work. In terms of transparency and evidence, that’s also fine, right? No organization gets it right the first time. As with any business process, it’s something you work on again and again.

This would include things like how many reports were received and how many of them were anonymous. Having anonymous channels is super important. What you may want to see over time is less and less anonymous outreach because people gain trust in the system. If you have less and less anonymous outreach, it probably points in the other direction. Then you want to look at things like retaliation: How many retaliation complaints are there? What are the appeals processes? How many appeals are filed by people who are not satisfied with the outcome of their case? Response timelines, these sorts of things, and satisfaction of whistleblowers—that’s the other part.

We’re not seeing basically any of these AI companies, apart from OpenAI, publish their policies. After the drama from last year—which maybe you remember from the Apple story previously—there’s a pattern that only after a scandal does something get published. None of these companies publish any of this, and that is not good. We talked about it before: well over 70% of whistleblowers tend to start internally. These systems have to work well, and we cannot just rely on trust, especially given the sort of precedent around whether they are going to work well.

That’s why we’re calling for companies, at an absolute minimum, to publish their whistleblowing policies and, ideally, to catch up to the global standard by also publishing evidence around how well their systems work, how they are performing, and what measures they take to improve those systems. That’s important, right? We think this is a bit of a litmus test because there’s essentially no cost to companies for publishing this. We’re not asking for additional reporting to be created; we’re just asking for transparency about what already exists.

Any company that takes this as seriously as a business process would, of course, have a whistleblowing policy and would already measure all of those things. If they care, they’re measuring these things. If they’re not measuring them, then we have an answer, at least to an extent: either they don’t really care about it at all and haven’t invested the time to think about what they should care about, or they have thought about it and are actively not doing it. Neither of those seems like a great option.

Yes, it is also global best practice. A bunch of companies already do this. There are obvious benefits for insiders because the public can then look at these policies and explain why certain policies fall short, or what looks good and what doesn’t look good, which they can’t do today. Again, false confidence may exist.

It’s good for the public because we know, and it’s good for companies. We talked about the benefits of a speak-up culture: feedback, improvement of these policies, and all of the impacts we mentioned before. In fact, there was an asset manager called Trillium Asset Management that, in 2022, called on Google to improve its whistleblowing systems for exactly all of those reasons. They were basically saying that strong whistleblowing systems serve shareholders. There’s an interest among shareholders in having strong whistleblowing systems because we want to make sure there is no misconduct. It basically only doesn’t serve direct managers, and potentially executives, depending on how you look at it.

It’s a very reasonable ask that we’re putting forward, which would hopefully still be quite impactful. Of course, just creating transparency is not everything. You can have a great-looking policy and it still doesn’t perform. You can create transparency around your evidence, and the evidence looks bad or maybe isn’t trustworthy. This is a minimal thing and the first step that we think these AI organizations should be taking. We’re also happy to work with them on these topics.

We’ve got an incredible coalition that we put together for this. It’s well over 30 organizations, and we’re very proud of it because it’s the first coalition of its kind, with the best whistleblower-support organizations in the world. The Signals Network, Government Accountability Project, which I mentioned already, Whistleblower Aid, and Whistleblowers UK—it’s basically the who’s who of the whistleblowing world.

There are academics as well, including the individual who co-wrote the ISO 37002 standard on internal whistleblowing systems. Transparency International is also on board; it wrote the best-practice guide to whistleblowing systems and creating transparency around evidence. On the AI side, we have Stuart Russell, who is joining the call, and Larry Lessig, who are both signatories of the Right to Warn. Daniel Kokotajlo is on board. Of course, you are on board, Nathan. Thank you very much for joining the call. The Future of Life Institute, Karma, and many, many more are involved. I’m not going to list them all off the top of my head.

Take a look at the website. It’s publishyourpolicies.org. If you’re an insider at an AI company and you’re thinking, “This sounds like a sensible thing, and I would like to have this transparency,” reach out internally. Ask your management. Maybe you have an anonymous town hall. Maybe you trust your direct managers enough to raise it and say, “Hey, why are we not doing this? This is standard practice. This could be helpful. Why not?” This is going to benefit your manager as well, and it’s going to benefit your manager’s manager, probably, as well.

By the way, this is another thing from our survey that I didn’t mention: 100% of the insiders we surveyed support publication of policies. There seems to be pretty broad support for this. If you’re an outsider and you’re not working at an AI company, spread the word. Make sure the call is heard. That’s pretty much it on that campaign.

Nathan Labenz

Great.

Karl Koch

Yeah.

Nathan Labenz

I guess it’s too early to have any responses from any official channels at the companies, right?

Karl Koch

That’s right. We just launched the call last week, so it’s been a bit over a week. We know that they’re aware of this call. They’ve been aware of it for a while because, for example, the Future of Life Institute and the AI Safety Index also called for publication of policies, although they recommended it rather than actively calling for it.

The questionnaire underlying that study came out, I think, 6 weeks ago. It was shared with the companies and included a question about why they weren’t publishing their policies, so they are definitely aware of the question. We had also given these companies a heads-up. We know they’re aware of the call, and we’re looking forward to working with them and seeing what the responses are going to be. If there are no responses, then we also have a response.

Nathan Labenz

Yeah. Cool.

Karl Koch

Do you also want me to talk a little bit about—

Nathan Labenz

Yeah, I was just going to invite you to do that again.

Karl Koch

Absolutely. Thank you very much.

Nathan Labenz

The floor is yours. What should people expect? What can they count on?

Karl Koch

Basically, where we see ourselves, at least on the direct-support side, is as a connecting point between the AI and whistleblowing ecosystems and as a first point of contact. That’s why the third-opinion offering I mentioned before allows you to reach out with a question about your concern without sharing any confidential information, fully anonymously.

We then workshop together, using an open-source anonymous tool that you can access via the Tor Browser. We workshop the question together and identify the relevant independent experts together, so you don’t have to rely on us knowing the experts in your field better than you do as an insider. We identify the relevant experts together, approach them with your question, and bring their answers back to you.

Hopefully, at this point, your concerns are alleviated. If they are not, we will help connect you to pro bono legal counsel that is extremely experienced in helping whistleblowers along their journey, with no pressure for any disclosure. Regardless of where you are in your journey, the support is available. You can also reach out directly to us without going through the third-opinion process and ask for help identifying who may be the best fit for you. We will help you there as well and supplement the independent expertise from the expert network, covered under legal privilege, with those great organizations as required.

On our website, apart from the organizations listed, you can find an explanation of the process and a digital privacy guide if you’re concerned about digital privacy, which we do actually see quite a lot. It makes a lot of sense to make sure you stay safe.

There are also a bunch of resources there that you can find. We have previously supplied hardened devices to at-risk individuals with specific operating system setups that are highly secure. That is also something we offer on the direct-support side.

On a wider scale, what we as the AI Whistleblower Initiative do—we mentioned systematically breaking down barriers for AI insiders—includes the advocacy side and the research side. The survey we mentioned is still ongoing. You can find the link to the survey in the show notes if you work at a frontier AI company.

There is also an upcoming legal study that we’re currently fundraising for, to really dive deep into the status quo of whistleblower protections across a wide range of AI risk scenarios. Identifying the most interesting scenarios is part of that upcoming research study, pending funding.

Then there is advocacy, like public campaigns, I mean, and policy work: providing feedback on policy, both in the US, where there are other great organizations working on this. If you’re interested, reach out, and we can connect you. On the EU side, we’re working with the AI Office to make sure they establish a whistleblower mailbox. In fact, the authors—the vice chairs of the Code of Practice—recently called for exactly that as well, which is amazing.

So that’s what we focus on at the moment.

Nathan Labenz

Cool. Well, thank you. This has been great. I think between the coalition of organizations that you’ve been able to put together and the evident seriousness with which you’re taking every aspect of this, the thoughtfulness of the support structures that you’ve designed, up to and including the provision of hardened devices, all of that is, in my mind, not too much for people who are concerned with just how crazy things might get to invest in now.

There might be a couple dozen individuals who happen to be placed at the right intersection of information and access to what’s going on, and who have the awareness and consciousness to want to do something about it, or at least seriously question it before moving ahead. I think those people are going to be scarce and precious resources for society, and also under a lot of stress and pressure individually as they’re facing those things.

I think it is excellent that you and your coalition of the willing are setting things up now to support those people. I’ve been glad to be a very small part of it. Hopefully, this helps raise awareness further and establishes you guys as a resource that people will hopefully never need. But it seems likely that there are going to be some cases where people will need to reach out and get this kind of support.

I, for one, would have appreciated having it 2 and a half years ago already. But certainly, as the stakes only continue to rise, I’m very glad that people in the future will have this option to avail themselves of this very thoughtfully designed and soberly provided support. So that’s great. Keep up the good work. Again, we’re all counting on you.

But for now, Karl, managing director of the AI Whistleblower Initiative, thank you for being part of The Cognitive Revolution.

Karl Koch

Thank you very much for having me.

The AI Whistleblower Initiative: Supporting AGI Insiders When It Matters Most, w/ founder Karl Koch | BidClub