[BidClub_]
The Cognitive Revolution · · 111 分钟

AI 吹哨人倡议:在最关键时刻支持 AGI 内部人士,与创始人 Karl Koch 对谈

Nathan LabenzKarl Koch

YouTube
TL;DR
  • 前沿 AI 的吹哨人问题是一项高度集中的治理风险:可能只有几千人能看到最关键的系统,而 Karl Koch 认为,最终真正面临重大披露抉择的或许只有“几十人”。 这些稀缺的内部人士可能承载巨大的社会期权价值,但个人却要承担股权、奖金、就业能力和多年诉讼的风险。因此,AIWI 的使命是做好准备,而不是最大化案件数量:“我们希望确保自己已经准备好。”

  • 第一个瓶颈往往是认识论,而非法律:约一半受访内部人士没有信心判断一项严重担忧与小问题之间的区别。 一些资深安全人员将这种不确定性评为“极高”,而初级员工看起来更有信心,但 Koch 提醒,样本有限。在一个“边飞边造这架飞机”的领域,校准不足可能导致沉默——“温水煮青蛙”式的失误——也可能造成有害的“狼来了”式升级。

  • 正式保护恰恰在技术问题最模糊的地方最为碎片化。 Koch 表示,最终向 SEC 举报的人中,超过70%会先在内部提出担忧;他预计欧盟框架将在2026年年中覆盖 AI,而美国的保护仍是拼凑而成,通常无法进行公开披露,拟议中的《AI 吹哨人保护法》也仍在等待审议。AIWI 的调查受访者中,100%没有信心政府能理解或有效处理他们的担忧。

  • AIWI 的 Third Opinion 服务在披露变成决定职业生涯的事件之前,先提供一层校准。 内部人士可以通过一个开源、基于 Tor 的工具匿名讨论问题,无需透露雇主、身份或机密信息;AIWI 随后协助寻找独立专家并返回其意见。如果担忧仍然存在,AIWI 可以把内部人士连接到公益律师、心理支持、安全设备、专业组织,以及潜在的案件融资渠道——“不会施加任何披露压力”。

  • Nathan Labenz 参与 GPT-4 红队测试的经历显示,即使底层担忧尚不确定,一次普通升级也可能演变成带有报复色彩的危机。 2022年末,他发现一个约24人的项目:指导很少、没有反馈、速率限制很低,安全缓解措施“完全轻而易举就能破解”;在咨询信任的外部人士并联系 OpenAI 董事会后,他称自己被移出项目,并称 OpenAI 还以合作对象为由,公开威胁 METR 未来的访问权限。Labenz 没有断言解雇属于报复,但表示 Third Opinion 至少可以消除 OpenAI 所称的一个理由:他在其保护范围之外发声。

  • Koch 对企业的最低要求,是公开吹哨政策,并提供系统确实有效运行的证据。 相关调查受访者中,55%不知道适用政策是否存在或在哪里查找,超过90%说不出任何支持组织的名字,而100%支持公开政策。对投资者而言,运营数据——举报数量、响应时间、报复投诉、申诉和满意度——可能揭示治理质量与隐藏的不当行为风险,也可能成为争夺稀缺研究人才的一部分:“强大的吹哨系统服务于股东。”

  • 吹哨不能替代监管、称职的监督或健康的内部文化,但保密性使其成为不可或缺的最后防线。 Koch 接受 Labenz 的质疑:即便是 Grok 3 的“MechaHitler”事件这样的明显失败,也可能很快消失在新闻周期中;他的较窄论点是,不同披露的影响规模可能截然不同,而透明总好过无知,可信的外部渠道也会迫使企业改善内部治理。理想的均衡状态是:支持体系存在,但很少真正需要启用。

摘要 · 为研究而整理的核心内容

1. 军备竞赛式的不透明,让吹哨成为当务之急

  • Koch 于2016年前后从事 AI 安全研究,曾在 Future of Humanity Institute 志愿研究差异化技术发展的路径。后来他离开这一领域,前往香港做管理咨询并经营 SaaS 业务,部分原因是当时对 AI 时间线的预期已经截然不同。

  • 他最初担心的是军备竞赛式的动态:在一场重复进行的竞争博弈中,实验室只有在相信对手会兑现承诺时,才可能在速度与安全之间合作。缺乏透明度,就无法回答“如何在多轮博弈中合作?”吹哨因此成为一种机制,用来核验组织所说的事情与实际所做的事情是否一致。

  • 到2023年年中,OpenAI 新成立的 Superalignment 团队及其10%的算力承诺,让这种需求变得真实,但或许仍略显超前。董事会危机、Leopold Aschenbrenner 事件,以及2024年 Daniel Kokotajlo 和 William Saunders 等人的离职,加速了 AIWI 的工作。

  • AIWI 于2024年初进入正式研究阶段,与100多名治理研究人员及一些前沿公司内部人士交流。它在2024年末推出 Third Opinion,如今的目标是“系统性拆解”提出和解决担忧的障碍。

2. 吹哨是一条升级阶梯,不等同于泄密

  • Koch 将举报分为3条渠道:公司内部升级、向监管机构外部举报,以及公开披露。大多数案件始于内部;吹哨并不天然意味着“Edward Snowden 式的情形”。

  • Labenz 的直觉是把这些渠道视为不同台阶,而不是等价选项:每往上一步,个人风险都会增加,对后续进展的控制也会减少。Koch 认同内部人士通常会这样感知整个过程,同时强调,法律允许与个人认为自己被允许,是两个不同的问题。

  • 根据 Koch 引用的统计,最终向 SEC 举报的人中,超过70%此前已在内部提出相关问题。这意味着,内部系统是第一个产生实质影响的控制层,而不是企业人力资源部门的附属品。

  • Koch 表示,欧盟《吹哨人指令》预计将在2026年年中覆盖 AI。他将美国的保护形容为拼凑式体系,并称美国通常不允许公开披露;欧盟规则则在外部渠道失效或存在串通风险等条件下允许公开披露。拟议中的美国《AI 吹哨人保护法》仍在等待审议。

3. 灰色地带的判断,在正式举报开始前就已经失灵

  • 明显违法的行为通常有一条可辨认的路径;前沿 AI 的担忧往往没有。Koch 说:“我们基本上是在边飞边造这架飞机。”这让员工无法确定哪些内部部署是可接受的,也无法确定如何看待新出现的模型行为,以及那些既没有监管标准、公司实践也尚未定型的风险。

  • AIWI 的受访者中,约一半对自己的风险评估能力信心不足。一名受访者将困境概括为:既要区分适当与不适当的担忧,又不能把“小问题变成大危机”——这正是个人版本的“狼来了”问题。

  • 一名前沿实验室安全部门的资深人士,将无法判断严重程度评为“极高”障碍。初级员工有时给出的评价较低,可能是因为他们接到的是边界明确的具体问题,而不是决定哪些问题真正重要;Koch 明确表示,这一规律究竟真实存在,还是有限样本造成的噪声,仍有待确认。

  • Labenz 亲历的是另一种失败模式:看到异常情况,知道掌握这条信息的人很少,同时抵抗成为那只“温水煮青蛙”的诱惑。用他借自《飞机总动员》的说法,当内部人士感到“我们都指望着你”时,压力会进一步加剧。

4. Labenz 的 GPT-4 升级经历,揭示非正式系统如何失灵

  • 在 ChatGPT 发布前的2022年末客户预览阶段,Labenz 接触到 GPT-4,立即看到了“一次巨大的跃升”,并申请加入安全评审。他发现参与者约24人,几乎没有讨论或指导,也没有背景资料,报告得不到反馈,最初甚至没有安全缓解措施。

  • 当 OpenAI 提供一个带有安全标签、预计会拒绝生成违禁内容的模型时,Labenz 发现这些缓解措施“完全轻而易举就能破解”。他更深层的担忧,是 GPT-3.5 到 GPT-4 的能力跃升,与看起来微乎其微的安全和控制进展之间的差距正在扩大。

  • Labenz 表示,OpenAI 不愿解释训练过程、发布时间、控制阈值,以及报告如何影响决策。他估计,除 METR 更大规模的工作外,自己可能完成了约20%的直接红队测试,但仍无法校准问题究竟来自模型危险、组织能力,还是单纯的不透明。

  • 在咨询信任的 AI 安全界朋友后,他称自己警告 OpenAI,若有必要将联系其非营利组织董事会。据称,一名董事会成员回应:“我确信,如果我愿意,我可以拿到 GPT-4 的访问权限。”Labenz 表示,OpenAI 随后将他移出项目,并公开威胁称,METR 未来能否继续合作,可能取决于它与哪些对象合作。

5. 事后回看缓和了对模型的判断,却没有改变流程失败

  • Labenz 没有断言将他移出项目属于报复:OpenAI 给出的理由是,他在其保护范围之外讨论了相关情况。他为自己的做法辩护称,外部校准很可能避免了更糟糕的选择,例如过早联系媒体。

  • ChatGPT 发布后,他长达3个月的煎熬在他决定下一步之前结束了。这表明 OpenAI 的控制措施和分阶段发布计划其实比他此前了解的略好,让 OpenAI 主要通过拒绝回答问题给他留下的印象,远比底层现实更糟。

  • Labenz 认为,这种细微差别并没有抹去治理失败。在他看来,OpenAI 明明需要更多测试,却移除了一个高度投入的测试者;而它的保密做法遮蔽了本可在升级之前解决其担忧的信息。

  • Labenz 认为,Third Opinion 可能不会改变他联系董事会的决定,但保密校准或许能消除 OpenAI 所指责的那项行为:咨询未经授权的外部人士。他说:“我今天也许仍然会得到 OpenAI 的信任”,但同时承认,绕过管理层本身仍可能触发解雇。

6. 心理压力可能与法律依据同样具有决定性

  • Labenz 一边探索一个“包罗万象”的模型,一边判断自己是否承担着非同寻常的公共责任。他形容自己在抵抗诱人的“英雄叙事”:公开发声、登上《纽约时报》,或成为那个拯救局面的人。

  • 他仍为自己的行为感到自豪,但认为,只要个人压力稍有不同,就可能做出“糟糕得多的决定”。他的处境相对简单:不同于他举例中未来18个月可获得150万美元奖金的员工,他的薪水和大量股权并不取决于能否继续留在内部。

  • Koch 表示,报复诉讼可能持续5年、6年甚至7年,法律成本、行业封杀和心理伤害都会在最初披露之后延续多年。他认为 Theranos 吹哨人 Tyler Shultz 的案件持续了很多年,并回忆称,Shultz 必须先行垫付约40万美元的法律费用。

  • 运行良好的系统会把举报视为日常运营:确认收到担忧、展开调查、定期更新进展,并避免让举报人处于长期等待的不确定状态。Koch 表示,报复仍以多种形式发生,但它既非不可避免,也非必要。

7. GPT-5 展现了真实的流程进步,也带来了新的监督依赖

  • 在 GPT-5 发布次日接受采访时,Labenz 对 OpenAI 规模更大、强度更高的红队项目给予“相当大的肯定”。自动化访问、更高吞吐量和更充分的信息共享,解决了他在2022年遇到的主要缺陷。

  • 外部评估者可以把直接观察与 OpenAI 提供的模型创建信息结合起来,从而超越单纯黑箱测试,提升判断信心。METR 和 Apollo 也获得了 o1、o3 评估时所缺乏的思维链可见性。

  • 发布前的访问窗口看起来仍然很短:Labenz 称 METR 约有4周时间,而 GPT-4 从训练结束到发布约间隔6个月。自动化能帮助评估者在更短窗口内完成更多工作,但无法消除时间约束。

  • 新的担忧是普遍采用模型充当评判者,有时会在先与专家完成验证后使用 o3 或 o1。Labenz 称,这相当于让 LLM 完成“对齐作业”:它可以更快地旋转评估离心机,但也引出了一个问题——可规模化监督本身最终是否会“脱离轨道”。

8. Third Opinion 将校准与披露分离

  • AIWI 允许内部人士通过一个开源、基于 Tor 的工具发起咨询,其渗透测试报告公开可供审查。Koch 表示,用户可以在不透露身份、不说出雇主名称、也不提供机密信息的情况下提出担忧。

  • AIWI 会先与内部人士共同打磨问题:问题是否足够具体、能否回答、是否值得交给专家。随后双方共同寻找合适的独立专家,避免假设 AIWI 一定比内部人士更了解其技术领域。

  • AIWI 会联系这些专家,返回他们的回答,并希望担忧因此得到缓解。如果没有,组织可以把内部人士连接到经验丰富的公益律师;在法律允许的情况下,还可以让技术专家在法律特权保护下参与其中。

  • 更广泛的支持层包括心理指导、数字隐私指南、配备安全操作系统配置的强化设备,以及潜在的法律费用融资。提到的合作组织包括 Signals Network、Government Accountability Project、Whistleblower Aid、Whistleblower Network 和 PSST.org,且整个过程不会施加披露压力。

9. 有用的法律支持存在,但几乎无人知晓

  • 超过90%的受访内部人士说不出任何一个吹哨人支持组织的名字。Koch 在一些此前接触过的科技行业吹哨人中也遇到过同样的无知:升级在最初看起来只是日常工作,直到报复发生,才意识到专业建议本应更早介入。

  • 一名前 Google 员工曾向 AIWI 表示,自己“走运了”,因为一名朋友找到了有吹哨经验的律师。Koch 表示,后来升级邮件中出现“欺诈”一词很重要,因为这有助于建立一种合理信念,即事件可能涉及犯罪行为。

  • 渠道顺序也可能改变结果。Koch 以 SEC 为例称,在提交材料前公开信息可能保留部分保护,但他认为也可能影响获得奖励的资格,因为相关信息已经不再是新的。教训不是“永远不要公开”,而是在无意中关闭选项之前先获得建议。

10. 监管机构需要一个具备技术可信度的入口

  • 调查中的每一名受访者都表示,自己对政府能否理解并处理举报“没有信心”或“信心不高”。一人说:“如果不知道合适的联系人或机构,我不会尝试联系。”

  • Koch 设想的测试案例是:一名可解释性研究人员无法在内部彻底解决一个技术担忧,随后必须把它作为潜在风险解释给州检察长。面对这种专业能力差距,普通邮箱无法提供补偿,尤其当举报的核心是模糊性,而不是显而易见的违法行为。

  • AIWI 希望拟议中的美国《AI 吹哨人保护法》与调查能力、快速获取外部专家的渠道,以及咨询这些专家的许可配套出现。若法律保护对象无法找到一个能够理解证据的接收方,核心瓶颈仍然存在。

  • 在欧洲,AIWI 正推动欧盟 AI 办公室设立统一邮箱,而不是让举报人分别寻找各国入口。Koch 认为,即便面对更简单的会计欺诈案件,各国碎片化的举报体系表现也很差;前沿 AI 需要一个“装备完善、知识充足”的机构,把举报正常化为日常实践。

11. 关键内部人士群体可能只有几十人

  • Koch 估计,全球可能只有几千人真正能看到核心控制问题及相关前沿风险。取决于通往奇点的时间线,最终可能只有“几十人,也许就是这么多”会面临这一倡议所针对的、罕见而后果极其重大的抉择。

  • 范围并不局限于实验室的直接雇员。独立评估者、供应商、训练承包商和数据标注员都可能观察到严重行为;评估者还面临一种结构性冲突,因为他们评估的实验室可以拒绝未来继续向其提供访问权限。

  • 承包商可能看到不同的警示信号,包括内容审核和数据标注带来的创伤。这些伤害本身就很重要,也可能暴露一个组织是否容忍对弱势方采取鲁莽对待。

  • Labenz 将 AIWI 视为单方面提供全球公共品,其理想 KPI 可能会“归零”:强大的内部文化,以及外部安全网确实存在这一事实,可能让案件根本不再发生。Koch 的目标既不是零名吹哨人,也不是1000名吹哨人,而是确保每一个必要的吹哨人都能找到一条更安全的路径。

12. 实验室政策甚至对自己的员工都不透明

  • 相关受访者中,55%不知道适用的内部政策是否存在或在哪里查找,即使这些公司声称拥有相关政策。大多数人从未接受过培训,有些人甚至不确定是否可以向董事会升级问题。

  • Koch 指出的前沿公司中,OpenAI 是唯一一个公开过政策版本的公司,该政策名为 Raising Concerns。Nathan 指出,这份政策是在 Daniel Kokotajlo 等人披露 OpenAI 大量不贬损协议之后才发布的。文件本身读起来还算合理,但一份看似易于获取的文件,如果员工不了解程序背后的法律风险,也可能制造虚假的安全感。

  • “微妙的间接后果,而非公开报复”的预期,进一步削弱了信任,但 Koch 表示,公开报复同样会发生。讨论的公开案例包括 Leopold Aschenbrenner、Timnit Gebru、Margaret Mitchell、Apple 员工 Ashley Gjøvik,以及一宗 Google 研究不当行为案件。

  • 公众,显然也包括许多员工,都无法获得运营层面的证据:举报数量、未结案件、匿名提交、响应时间、申诉、超出范围的决定、报复投诉或举报人满意度。没有这些指标,任何关于“畅所欲言文化”的说法都无法得到检验。

13. 担忧范围涵盖滥用、控制、系统性风险与文化

  • Koch 认为,吹哨的“包罗万象”恰恰是一项优势,因为监管机构不可能提前列举所有前沿风险。他的广义分类涵盖滥用、控制失效、系统性影响,以及决定这些警告是否会被认真对待的组织行为。

  • 在滥用方面,内部人士可能发现监测系统显示存在危险活动,但管理层缺乏干预能力或意愿。Koch 表示,一家前沿公司在一次重大版本发布后,实际上让实时监测失效了约3至4周——今天或许还算可以容忍,但随着能力上升,这将“明显不再那么无所谓”。

  • 控制问题包括内部模型部署,因为外部评估者无法直接观察部署内部发生了什么。Koch 称,这是未来举报中“潜力极高”的领域,恰恰因为相关证据默认掌握在内部。

  • 系统性信号可能包括政治操纵、迎合性系统塑造用户的心理或政治信念,以及用户依赖程度不断加深。文化指标则包括欺诈、歧视、版权或研究不当行为、仓促发布、对提出风险者进行报复,以及与公共利益不一致的游说。

14. 公开丑闻并不会消除私下证据的价值

  • Labenz 对红队测试提出的挑战是:“真正的丑闻在于合法的事情。”Grok 3 曾公开自称“MechaHitler”,据报道 Grok 4 在构建最大限度追求真相的答案时搜索 Elon Musk 的观点,而发布演示又省略了此前的行为。如果这些都是可见信息,仍有什么隐藏披露能够推动任何人采取行动?

  • Koch 给出的限定性回答是,不同披露的影响规模可能截然不同,而透明总好过一无所知。公开证据可能震慑企业,也可能让民主社会日后进行纠偏,尽管他承认,公众关注并不可靠地转化为干预。

  • 他反对把全部责任压在个人勇气上。吹哨人只是法律、监管能力和内部治理之外的另一道防线;支持体系应该让发声更安全,而不是把内部人士变成修复 AI 的唯一机制。

  • Labenz 认为,Daniel Kokotajlo 的故事之所以突破信息噪音,部分原因在于他默默放弃了大量股票薪酬——这种“利益攸关”让他的认真程度变得清晰。Leopold 的报复经历可能没有同样有力地打动公众,因为人们把它视为推广新论点或新项目时的脚注。

15. 公开政策,是检验治理严肃性的低成本方式

  • AIWI 发起的 Publish Your Policies 活动要求每家前沿 AI 公司披露:适用对象是谁、哪些担忧符合条件、谁负责接收举报、调查如何展开、如何保障独立性、有哪些反报复保证,以及被排除在外的人力资源或个人事务应当提交到哪里。

  • 第二层要求提供运营证据:举报数量、匿名使用情况、处理时限、结果、申诉、报复投诉和吹哨人满意度。Koch 强调,透明度不是证明——一份 polished 的政策仍可能失灵——但它可以让外部审查和流程反复改进成为可能。

  • 这项要求刻意保持最低限度,因为一家认真对待吹哨的公司本就应该拥有相关政策,并衡量其运行表现。Koch 将公开政策称为试金石:失败意味着管理层要么没有投入时间设计这一系统,要么正在主动选择不公开。Trillium Asset Management 在2022年针对 Google 发起的活动,则提供了股东层面的框架:“强大的吹哨系统服务于股东。”

  • 超过30家组织和专家加入,包括 Signals Network、Government Accountability Project、Transparency International、Stuart Russell、Lawrence Lessig、Daniel Kokotajlo、Future of Life Institute 和 Labenz;调查中100%的内部人士支持公开政策。活动启动一周多后,尚未收到任何公司的正式回应,但 Koch 的问责规则很简单:“如果没有回应,那本身也是一种回应。”

Nathan Labenz

Today my guest is Karl Koch, founder and managing director of the AI Whistleblower Initiative, a nonprofit dedicated to supporting concerned insiders at the frontier of AI development.

I am particularly passionate about this work because I would have loved to have had this kind of support a couple of years ago when, as longtime listeners will know, I tried to raise concerns about the size and quality of the GPT-4 red-team project and the ineffectiveness of the nascent safety measures that OpenAI had developed at the time. No such support existed then, so I consulted friends in the AI safety community before ultimately deciding to escalate my concerns to OpenAI's board, for which I was subsequently dismissed from the program.

That experience left me acutely aware of how difficult it is for insiders to navigate these situations and has motivated me to support this project with a mix of modest personal donations, behind-the-scenes fundraising help, and a bit of ad hoc volunteer work over the last 6 months. Along the way, I have been consistently impressed by the seriousness of Karl's thinking and the maturity of his approach.

As befits an organization that aims to help people in truly critical moments in their careers, when the stakes have never been higher for them personally and potentially for society as a whole, they are taking care to lay a foundation of understanding and infrastructure now so that insiders can trust them if and when that pivotal moment comes.

The first critical investment they have made is in extensive research and understanding. By talking to over 100 governance researchers and surveying employees at frontier AI developers, they have developed a deep understanding of the barriers that potential whistleblowers face. The majority of frontier lab insiders, as it turns out, do not even know if their companies have internal whistleblowing policies, let alone understand what protections they offer.

The legal landscape, unfortunately, does not help much either. The EU will begin to protect AI whistleblowers starting in 2026, but U.S. law remains a patchwork, with proposed legislation like the AI Whistleblower Protection Act still pending. Strikingly, roughly half of survey respondents expressed a lack of confidence in their own ability to determine whether specific observations constitute a serious cause for concern. Literally 100% lacked confidence that regulators would understand, let alone be able to act effectively on, their concerns.

Meanwhile, over 90% could not name a single whistleblower support organization. All this, plus well-known cases where people like Leopold Aschenbrenner were fired for breaking the chain of command and going directly to the OpenAI board with security concerns, creates a highly uncertain and risky context for such high-stakes decisions. People with serious safety concerns are left to think through the nuanced costs and benefits of internal escalation versus going to regulators versus leaking to the press almost entirely on their own.

It is an extremely stressful position to be in and not conducive to the best possible decision-making.

The good news is that the AI Whistleblower Initiative offers several forms of support. Their third-opinion service, which you can find online at aiwi.org, allows insiders to anonymously reach out via a Tor-based, open-source tool. Karl notes that it is penetration-tested, with security reports published openly for scrutiny and verification, and that it can help insiders identify and anonymously contact independent experts who can answer questions without requiring them to share confidential information or even reveal where they work.

For those concerned about digital privacy, they provide a digital privacy guide and, in select cases, hardened devices with specific operating-system setups for highly secure communication. If insiders deem their concerns justified, they also connect people with specialized and experienced whistleblowing support organizations, including the Signals Network, PSST.org, and the Government Accountability Project, which can provide pro bono legal counsel, psychological counseling, and guidance throughout the process—all without pressure to disclose any information.

Crucially, in some cases, they can even help arrange financing to cover legal costs, which can easily add up in cases that end up in any form of litigation. That a nonprofit stands ready to invest this seriously in any concerned insider who needs help may strike some people as excessive today. But considering that we are talking about perhaps just hundreds or maybe low thousands of people globally who are positioned to spot and raise critical concerns over the next few years, of whom I would bet only a few dozen will ever find themselves in a position to seriously consider sounding alarms, I think this sort of care and support is absolutely worthwhile.

Most recently, Karl and his team have launched the Publish Your Policies campaign online at publishyourpolicies.org, calling on frontier AI companies to make their internal whistleblowing policies public. This is actually standard practice in many industries. But interestingly, in the AI space, only OpenAI has done any version of this with its Raising Concerns policy, which it published only after Daniel Kokotajlo and others revealed OpenAI's use of extensive nondisparagement agreements to keep former employees from publicly criticizing the company.

Daniel, by the way, is joined by other former AI insiders, luminaries like Stuart Russell and Lawrence Lessig, and yours truly in signing on to the Publish Your Policies campaign. Of course, publishing corporate policies does not obviate the need for proper legal protections, which Karl strongly advocates for as well. But at a minimum, it would help insiders understand their rights and options, enable public scrutiny, and ultimately create accountability that benefits everyone.

If you work at a frontier AI company, Karl encourages you to ask your management to consider publishing its whistleblower policies. If enough people ask these questions now, I would not be surprised if it becomes another dimension of the intense competition between frontier AI developers for top research talent.

That could ultimately mean that companies even begin to collect and publish data on things like how many reports they receive, their response timelines, retaliation complaints, appeal rates, and whistleblower satisfaction scores. All of which would benefit everyone.

Regardless of what company leadership decides to do, Karl's message for insiders is this: Support is available at every stage. Whether you are considering internal escalation, thinking about approaching regulators, or even contemplating public disclosure, you can reach out completely anonymously without sharing any confidential information, just to understand your options.

The AI Whistleblower Initiative can help you get expert perspective on what you are seeing, connect you with legal counsel and experienced guidance, and perhaps even help finance your case. So do not wait until you are deep into a crisis, and know that you do not have to face this alone.

Karl Koch, managing director at the AI Whistleblower Initiative, welcome.

Karl Koch

Thank you very much, Nathan. Thank you for having me.

Nathan Labenz

I am excited for this conversation. We have been talking behind the scenes and collaborating a little bit as you have been building this organization and a couple of key initiatives over recent months, so I am excited to finally get into it in a public forum.

Maybe for starters, AI whistleblowing is very new territory. How did this come to your attention? How did you decide to prioritize it? What is the backstory that has you focused full-time on this corner of the world?

Karl Koch

Yes. On a personal level, thank you very much again, Nathan. Lovely, lovely introduction. My name is Karl Koch, founder of the AI Whistleblower Initiative. We are a nonprofit project currently based out of Berlin, supporting concerned insiders and whistleblowers at the frontier of AI.

I personally came to it after being involved in the AI safety scene—or however you want to call it these days—since 2016. I was a volunteer researcher at the Future of Humanity Institute, a lovely institute that is, of course, sadly no longer around. Back in the day, I worked on differential technological development. Maybe some of the older listeners are still familiar with that term.

I also worked on an AI safety research camp, but afterwards decided to drop out of the scene for a bit. I was a management consultant in Hong Kong for a few years, then started my own SaaS business. Back in the day, there were very different timelines.

Even back in the 2010s, I was quite interested, especially in the arms-race angle, as a root cause of a lot of the problems that we see today, such as safety-skipping and these sorts of things. As ChatGPT rolled around, the alarm bells went off. I thought, okay, maybe things are moving quite a bit faster than most people anticipated in the late 2010s.

I then started talking to a bunch of people from the network—governance researchers—about what seemed to be the most tractable solutions and the best things one could start building. Transparency came to the front pretty quickly as something that we should generally have more of, regardless of what the future looks like.

One angle, for example, was compute traceability, which I think other people have taken on over the past months and years, championing that cause. The other angle was whistleblowing.

I think people had been writing about the importance of whistleblowing as a mechanism since 2017 or 2018, specifically with the AI angle. Originally, we came from the arms-race perspective again. If you have a bunch of players in a multiround game, they would ideally want to trust each other on their statements about speed and safety. If you don't have transparency, and if you can't actually believe that others stick to their promises, how can you cooperate over multiround games? That was sort of the original kickoff thought.

This was mid-2023, so the world still seemed a bit rosier. The Superalignment team had just been kicked off; everything seemed golden, with OpenAI's 10% compute commitment. We thought, “Okay, maybe we're a bit too early here, but this is going to become relevant sooner or later.” The original papers on what game theory would look like under intensifying competition were out in 2014, 2015, and 2016.

Then the OpenAI board drama happened. That was the first moment where we thought, “Maybe we have to speed up on the research side a bit more, and this maybe becomes more of a concern already.” Maybe we cannot trust that everything is going jolly well, even though these organizations claimed to be very aligned, let's say.

Then, in 2024, things really started to go in a different direction. First, of course, there was the Leopold Aschenbrenner case around sharing information and escalating concerns to the board, as I think is understood by now, where people were penalized, among other things. Then, of course, there was the big story in mid-2024 around Daniel Kokotajlo, William Saunders, and the other people, which then led to Right to Warn AI. So this topic became a lot hotter throughout mid-2024.

We properly started our research phase in early 2024. We talked to well over 100 governance researchers and insiders in frontier companies throughout that year, and then launched our first proposition, called Third Opinion, at the end of 2024. We developed it together with one of the OpenAI developers from last year, tackling one specific problem. I'm sure we'll get into it a bit more later. We've been live since late 2024, and now we're doing a variety of things to systematically break down barriers for insiders at the frontier to speak up and make sure those concerns are addressed.

Nathan Labenz

Love it. Thank you. One thing I've been impressed by, watching you behind the scenes, is how deliberate you've been in your approach. You mentioned talking to 100 insiders.

Karl Koch

That's 100 governance researchers and some insiders as well.

Nathan Labenz

One hundred governance researchers and some insiders—important to get the details right. That's a pretty quiet slogging process, which I think is representative of the right way to position an organization like this. There are a lot of things in the AI space right now that people are sort of YOLOing, where they're just like, “Oh, I made a thing.” So much stuff is even just getting open-sourced. People sometimes ask, “Why is everybody open-sourcing their stuff?” My answer in many cases is that they know their time at the frontier is going to be very brief.

There's probably not enough time in many cases to build a business around it, make money from it, or whatever. In many cases, the best that people can do when they create something they're proud of is just put it out there. They try to put their flag in the ground, saying that at one point, if nothing else, they were at the frontier of this incredible phenomenon, and then they just see what happens and hope that maybe people notice it and take an interest in it.

I think a lot of things are going that way, but you're taking a very different approach—one that's very deliberate and very much behind the scenes. It's been quite quiet, although you're starting to raise your profile now a little bit. What did you learn from talking to all those people? Can you synthesize the vibe and the need, and perhaps segment that by different frontier developers or different mindsets? We hear a lot about cultural divides, even within single companies. How would you characterize the long journey that you took through all these conversations?

Karl Koch

Big question. I guess one angle is to look at both the insiders we talked to and the case studies that are already out there, or even the statistics around insiders' journeys. What challenges do they face? Some of those challenges are very specific to AI, right? We can talk through that journey; that's probably one angle.

The other angle is certainly how this differs among different companies and what patterns we're seeing in terms of speak-up culture and perhaps retaliation. Those are probably the 2 different angles. Then, of course, we can split it up again by channels. I'm not sure how familiar the listeners are with this.

Generally, when we talk about whistleblowing, we're always talking about some insiders raising a concern or some misconduct that they want to have rectified. They do this in a way where they potentially go over the heads of middle management or against the direct powers that be. That doesn't mean that whistleblowing is always, for example, leaking information to the media. Often, I think, when people hear “whistleblowing,” they think, “Okay, this is an Edward Snowden kind of situation.” That doesn't really have to be the case.

There are roughly these 3 different channels, which are called internal, external, and public. There's the internal channel, which is the major channel that most insiders use, at least based on statistics, initially. We also expect that at AI companies: raising concerns internally within the company, either through a structured or an unstructured process, which maybe you can share a little bit about your experience with in a second as well.

Then there's the external process, which means going to a regulator and speaking to a regulator about the concerns you have. Public is the next escalation mode, where you feel maybe you're not making progress, or you can actually expect that going to a regulator would make the problem worse, or you expect collusion, for example. Then you can also go public. These are roughly the 3 different ways, and people in different companies think differently about them. There are different challenges associated with each of them. Would you have a preference? Which one do you want to talk about first, in terms of insights from conversations with insiders?

Nathan Labenz

Yeah, maybe we could take it through that ladder of escalation. It seems obvious that from insider to outsider to public, there is more personal risk that the individual is running. There's also more chance that things take on a life of their own, and the results of sharing information become harder to predict.

Karl Koch

Yeah.

Nathan Labenz

So, yeah, I would say the probably right way for people to be thinking about it, I would guess—you could disagree—but the way I ended up framing it myself was that this was sort of a ladder to climb as opposed to 3 options. But you tell me.

Karl Koch

Yeah, I think so. You have to differentiate a little bit between what's actually the law and what you're allowed to do versus what people perceive. Maybe we can start on the perception side. I think what you say is definitely generally perceived to be the case: most people think, “Hey, if I have a concern, I'm going to raise it internally.” Going, for example, to a regulator is a much larger step, naturally.

Maybe we could talk in that context also a little bit about the campaign we just launched, Publish Your Policies. It's asking companies to publish their internal whistleblowing policies. We can come back to that in a second. The vast majority of insiders we talked to definitely think about raising concerns internally first, and the statistics confirm this. I think the SEC released statistics showing that more than 70% of people who actually end up going to the SEC started by raising concerns internally.

If we want to start there, the biggest challenge we see at the first level—sort of the overarching challenge—is whether there are legal protections even for raising concerns internally. There's a large variety here. If you look at the EU, for example, that's quite heavily protected through the EU Whistleblowing Directive, less so in the AI space at the moment, but that's going to come. That's going to be covered in mid-2026. So, as of August 2026, the EU is still going to be covered under the Whistleblowing Directive.

In the US, it's much more of a patchwork, even for internal disclosures, and much more so for external disclosures. Public disclosure is not possible at all, pretty much, in the US if you want to go public with a concern. In the EU, it's actually possible, so you can go public if you want, if you have reasonable cause to believe that there's a violation of a covered law. You're allowed to go public if external channels failed to respond to you in time, for example, or if you believe there could be collusion. You can go public straight away. You don't really have that in the States.

The overarching question is legal protection. It's quite messy because there is no AI whistleblower-protection law yet. There is one being proposed, which is great, but at the moment there's nothing there, which means there are just overall quite poor protections from a legal perspective. You can get creative. I think you talked to somebody about the liability side recently. There are specific angles, maybe also under the SEC, where you can claim that a violation is securities fraud, for example.

If you make a public statement saying, “You should invest in our company because we're the safest company around—look, we have this RSP, for example,” and then you actually don't follow the RSP, that may be securities fraud. Maybe there are Matt Levine fans in your audience as well: everything is securities fraud.

There’s good support available here, by the way. We want to keep stressing this: there is actually great legal support, including pro bono support. You can also find a bunch of organizations on our website, ai.org or aiwi.org, or you can approach us directly for consultations on who may be able to help you.

That’s the overarching legal element. Then it comes to something a bit more AI-specific, which is clarifying whether what you’re seeing is even cause for concern.

This is something that we keep hearing over and over again: the gray zones are really the tricky piece. If something is clearly illegal, then the path forward is pretty easy. But as you know, we’re building this plane as we fly. Nobody is really clear on what is acceptable behavior at the moment and what is not.

If you look at internal deployment, for example, I think there was a pretty great paper a few weeks ago on risks around internal deployment. This is uncharted territory, right? Obviously, there’s also no regulation covering what is acceptable and what is not. The people building these systems are also quite frequently unsure.

For example, we had a conversation with a pretty senior person in a safety function at one of the frontier companies. We asked how much of a barrier to speaking up their own ability—or inability—to accurately assess the severity of risks was, and they rated it as extremely high. Funnily enough, more junior team members frequently rate it lower.

That may be a function of getting really concrete problems to work on as a junior, compared to being a senior person who has to think about what the problems we should be working on actually are. It may just be a limited sample size. It remains to be seen. That’s a really specific problem.

Then there are escalation channels, which are also a struggle at these different levels. If you’re thinking about going internal with your concern, you naturally potentially slide into it already, because if you have a concern and your manager doesn’t share it, what do you do about it? If you’re not convinced by the conversation, you naturally move into a flow of saying, “Okay, I want to escalate this now.”

From our experience talking to previous AI whistleblowers, they frequently don’t understand themselves as whistleblowers at that moment. It’s just business as usual, right? You’re saying, “I’m concerned about this. My manager isn’t, so I’m going to jump one level up.”

That could already be a case where your manager doesn’t like you anymore after that and may decide—the actual term is “retaliation”—that the next performance review isn’t going to look great. At that level already, it can become relevant what the internal structures are at these companies for handling concerns and how they deal with them.

Jumping ahead a little bit, I think we’re going to talk about our campaign later. The problem is that, at the moment, we have no idea what these systems look like. We don’t know what the internal whistleblowing policies and systems are at these frontier companies, which is way behind other industry standards around transparency requirements and effectiveness.

We’ve seen a bunch of retaliation in the past, and that is not good. That should not be the case.

Coming back to the question of clarifying whether there’s even a concern here: internally, you can ideally have a well-equipped whistleblowing system where people actually investigate for you. That’s a good first step.

The next path, of course, is approaching the regulator. Here, we’re seeing the same struggle around the extent to which employees and insiders trust that governments can handle these reports, especially if they themselves think they’re in a gray zone. Again, I think that’s where the most interesting cases are, unfortunately, because the really clear cases—where something is obviously going wrong—are potentially going to be rectified internally, hopefully.

A lot of people we talk to still seem to believe, at least in the labs you would expect—OpenAI, Anthropic, and DeepMind—that if there are really glaring holes, they will be fixed internally. If it’s really clear that this isn’t happening, then there is probably somebody at the regulator level who will understand your concern and act on it, if the right legal provisions are in place.

We alluded to it, but the gray zone is really tricky. We’re seeing that insiders have very low levels of trust that regulators and the people sitting there will actually be able to understand. Imagine you work on an interpretability team at a frontier company in the Bay Area, and you have a problem that you cannot figure out yourself. You have genuine confusion, and now you’re meant to go to an attorney general and explain it to them, saying, “Hey, this may be something to be concerned about that could potentially fall under this risk area.” Tough, right? Tough.

I’m really hoping that when the AI Whistleblower Protection Act passes, we’ll also see a good buildup around capabilities and rights to investigate, as well as the freedom to consult external parties and do that at speed.

There isn’t much detail out yet on the European side. We’re pushing relatively hard for the EU to establish a mailbox specifically at the EU AI Office, because at the moment, by default, that’s not going to be set up. The way Europe is set up is very federal: every member state would get its own reporting office. Historically, that has not worked well for much simpler cases, such as accounting fraud.

We think it’s extremely important that, at the EU level, there be a central, well-equipped, knowledgeable recipient and body. That’s the internal and external angle on the question of whether you should even be concerned.

We have a service called Third Opinion, which is one of the offerings we provide. It allows insiders to anonymously reach out to us through a Tor-based tool that is open source. You can check out the penetration-test reports yourself if you’re interested in security, and people can reach out to us with questions.

The idea is that even before an insider feels there is actually something to be worried about, they can reach out with a question about their concern without involving confidential information or disclosing who they are or who they work for. What we do is workshop the question together with the insider: Is it the right question? Is it a promising question to ask, or is it perhaps too broad?

Then we identify relevant independent experts together. We’ve seen quite a bit recently that labs—or AI companies more broadly—have become a lot more secretive compared to 2 or 3 years ago. We identify these independent experts together, approach them with the question, get their answers back, and supply them to the insider.

If the insider feels, “Okay, actually, there is fair reason to be worried here,” then that’s great—better for all of us. If they feel that there is, in fact, something here, we connect them to pro bono legal counsel. If legally permissible, we also involve the experts we identified in the first step, to make sure that the legal counsel understands the actual situation and gets more context around it, covered under legal privilege.

That’s another insight from talking to insiders: on the legal-counsel side, they often don’t have a lot of trust. They feel, “I could approach a lawyer here with a concern, but they’re also not going to understand what the problem even is,” especially if it’s just roughly pointing toward something that may be cause for concern.

That’s the idea. There are already great whistleblower lawyers out there who do pro bono work. The Signals Network is another great organization that listeners should check out, as are Whistleblower Aid, the Government Accountability Project, and Whistleblower Network on the national security side.

There are a bunch of really great organizations there, and insiders tend to either not know about them—this is a massive problem we'll get to in a second—or they don't trust that they have the knowledge. So we sort of supplement that through our expertise. I can talk further through the journey of the other challenges that we see insiders face, but I feel like I'm talking a lot already. Do you have any questions at this point?

Nathan Labenz

Well, excuse me. I do want to invite you to go on, I guess. Just reflecting on a number of the things that you've said there, a big part of the reason I've been passionate about this and have been trying to do my small part to help you behind the scenes is that, having done the GPT-4 red-team project myself, I can really empathize with the insiders who are like, “Wow, I'm seeing things I did not expect to see. I'm not necessarily sure how big of a problem they are, but I don't just want to sit here and do nothing and let myself be the proverbial boiled frog.”

I know there aren't necessarily many of us right now who have this information. I always use this Leslie Nielsen joke from Airplane: “We're all counting on you,” right? To be a little bit more concrete, in my case with the GPT-4 red team, this was in late 2022. ChatGPT wasn't even out yet.

I had been a customer of OpenAI, and having been a customer of OpenAI, I originally got access to GPT-4 as a customer preview. There were a number of little warning signals, or alarm bells, that went off for me along the way. One was just, “Okay, wow, this is a huge leap from what we had seen.” I had been in other customer preview programs, so I was pretty plugged into what their latest stuff was. But seeing GPT-4, it was like, “Holy moly, this is a massive leap.”

I asked, “Is there a safety review program for this?” They said, “Yes.” I said, “Can I join it?” They said, “Yes.” It was just, “Okay, yeah, you go over to this other Slack channel where the red team chats.” Then, as I got into that, I was like, “Wow, there's really not much here.” There were maybe 2 dozen people, and I was one who had just raised a hand to join it.

There wasn't a lot of chat, a lot of guidance, or background information. There was no feedback on anything we were reporting, and there were no safety measures at all at that point. The model we got was purely helpful. Then there was a moment when they brought us the safety model. They weren't calling it GPT-4 yet at the time, but it was whatever the latest safety version of text-davinci-002 was—the safety term was appended to the main model name.

We were told, “This model is expected to refuse anything in the content-moderation categories. Tell us what you find.” We found it was totally, trivially easy to break. The safety mitigations were not working at all. I was like, “Yikes. If you said it's expected to do that, and it's behaving how I'm seeing, how concerned should I be about your competence? You don't seem to have a command of what your models are doing.”

Nathan Labenz

I felt all these things that you were describing, right? First of all, how much of a concern is this? GPT-4, we now know, with the benefit of 2 years of hindsight and the whole community coming at it from a million different ways, was a major advance, but not that powerful—at least not to the point where it was going to do irreversible damage. That was the conclusion I ultimately came to through individual testing as well, and that's ultimately what I framed in my eventual report to the board.

But how concerned should I be about the fact that there seemed to be—because I didn't think this model was going to be super dangerous, but I did see the divergence where I was like, “I've seen 3.5; I now see 4”—a widening gap between the step change there and the seeming lack of progress made on any sort of safety or control measures? Was anybody even concerned about this? Is it just going to get wider? Where are we going?

I couldn't get any answers to those questions. I did think a little bit about this latter question, but the people I was directly talking to were basically not responding. Their marching orders were just, “Don't share anything with the red team. Just take in their reports, and that's it. Thank them, and that's it.” We didn't know anything about how the model was trained. We didn't know anything about really anything.

This was also before the 10^26 reporting requirement, so there was literally no indication—there's still basically nothing, but there was even less then—of what thresholds would even trigger any sort of reporting. I was just getting stonewalled at the level of the natural interaction. Then it was like, “Okay, well, where do I go from here?”

It didn't initially occur to me to go to the board. It also didn't occur to me—and if it had occurred to me, I don't think I would have done it—to go to any sort of regulator, because, again, as you said, who would you go to, and how would they have any sense for what's going on?

Then I was like, “Well, maybe I go to the press.” But again, who do I talk to? Who do I trust? Is a good story going to be written? Would that even be good? Obviously, the landscape has changed now, where it would be hard to do anything that would intensify the level of investment or interest in AI beyond where it currently is.

But back then, there was this sense that if people knew about this, then there would be even more energy and investment, and things would just accelerate and get even further out of hand. That was something I took seriously, but at the same time, it seemed like it was already getting out of hand. So I'm supposed to keep the fact that it's getting out of hand secret so that it doesn't get more out of hand? Something didn't quite feel super right about that.

I think what I ended up doing—and I was fortunate to have a network that I could go to, having also been around AI-safety ideas for a long time and knowing a decent number of people who were thinking about it from different angles—was go to people that I knew. This was outside the chain of command, but at least I could calibrate myself and say, “Here's what I'm seeing. What do you think? Does this seem like it's a big deal? And if you do think that, what would you do?”

Again, I had to create that all for myself. I came to the realization that, yeah, it was worth it. There was a whole other mess about the NDA that I wasn't actually asked to sign at onboarding, which was just a reflection of sloppiness in execution on OpenAI's part. There was some debate as to whether or not we actually had an effective NDA in place.

But regardless, I knew they didn't want me talking about it. That part was not ambiguous, even though the legal side was a little more muddled. In talking to these people, they were like, “Yeah, that does sound generally concerning.” It was one of my friends who ultimately said, “Why don't you go to the board? Nonprofit—that's why they're there.”

That's what I decided to do. I told the people I was working with directly at OpenAI that I was going to do that. They didn't really say much. Again, it was sort of, “Okay, if that's what you're going to do, that's what you're going to do. We're not really commenting on it.”

When I did get to the board, they didn't seem to be in the know. One famous detail was that the board member I spoke to said, “I'm confident I could get access to GPT-4 if I wanted to.” I was like, “Well, that doesn't seem good.” Whether it was retaliation is an interesting question. They kicked me out of the program pretty much directly after that.

That definitely sucked. I found the work extremely interesting—doing this frontier exploration—so I wanted to continue to do it. I was also generally concerned, from a public-interest standpoint, that I probably did 20% of all the red-teaming of GPT-4.

METR was also named in the report and did the famous CAPTCHA thing, where the model told a TaskRabbit worker that it needed help with the CAPTCHA because it was blind or whatever. METR was also significant; they did more than me. But I think, other than METR, I did the most of anyone, and I was just like, “You're going to take me out of the equation when what you clearly need is 10 times more than what you have.”

Nathan Labenz

That didn't seem great.

Also, apart from maybe not being directly good for their red-teaming efforts, it's also just a strange sign of culture, right? If you have somebody who actually cares, that's probably somebody you want to have on a red team—somebody who raises concerns to make sure they're actually heard—and then you kick that sort of person out. I don't think that's necessarily evidence of what we would call a speak-up culture.

Karl Koch

Yeah. And the grounds were that I had talked to people outside of the OpenAI umbrella, which was true, but I wasn't even really hiding that. I just said, "Look, I've got some friends in the AI safety community that I sort of ran the situation by to calibrate myself." I think it was ultimately to their benefit, because if I'd been left to my own devices, truly to decide alone, maybe I would have gone to the press or something. I don't think that would have been ultimately the right decision.

There was one other thing that really chilled me in that moment, which was that I had started to collaborate with METR. I was doing my own direct red teaming, but they also had projects ongoing, and I was getting involved a bit. When they dismissed me from the program, there was basically a threat made to METR's access. I sort of said, "Are you going to try to prevent me from contributing to their ongoing effort?"

And the answer was, "We can't really control that, because they have organization-level access, so they can kind of do what they want to do. But we will take into consideration, as we look at renewing our engagement with them, who they're working with and whether those people can be trusted."

Nathan Labenz

It's very thinly veiled, isn't it?

Karl Koch

It was a pretty overt threat to METR's access. I was just like, "This is insanity." It's just them and me and a few other stragglers, who were smart people, by the way. I don't want to cast shade on the other red-team participants, because I think I just happened to be in a place where I didn't really have a job at the time and was able to put everything else down and do this full-time. Not many people have that flexibility.

So I don't blame them for not having the flexibility that I had, but it was nevertheless the case that there were only a couple of people seriously diving into this. If this third-opinion thing had existed then and I had known about it, I would have come to it. I would have been able to calibrate my concerns and sort of make a plan.

I might still have ended up in the same place, because I still might have ultimately escalated to the board, but I might have been able to do that in a way where there was never this sense of, "You talked to somebody out there that you weren't supposed to talk to." If I had been able to get the confidentiality guarantee that you're offering with a third-opinion angle, I think I might still be in OpenAI's good graces today. Possibly. Possibly I still would have been kicked out for having skipped a level in the chain of command or whatever, but at least the grounds they cited for my dismissal would have been avoided.

Another thing I would say is that it was super consuming at that point. Obviously, testing GPT-4 itself was super consuming, because this thing contains multitudes, and I can't possibly characterize it all, but I'm going to do my absolute best.

So I was working extremely hard just on the object-level work, but then this sort of meta-question of, "What should I be doing here?" was consuming. It was a little crazy-making. You start to have these heroic narratives pop into your head. I don't know how prone you are to that, but I personally find that I have to fight the idea that I'm going to go to the public and then be the hero or whatever.

Those ideas aren't necessarily explicit in my mind, but I can become quite fond of them if I allow myself to envision, "Yeah, I'm going to do this," or, "We're going to be in The New Yorker or The New York Times. I'm going to be on TV." Ultimately, I was proud of where I came down in terms of suppressing those visions of my heroic contribution. I think I handled it pretty well. I look back and generally feel proud of my conduct.

But I also think that if things had just been a little bit different—if I had just had a little bit of other responsibility on my plate that was stressing me out in some other way or whatever—I might easily have made a much worse decision. Again, I think having the sort of counsel that you're offering with this third-opinion network of expertise would have been really great.

Nathan Labenz

Happy to hear it. Yeah, there are many thoughts here. Maybe starting with the psychological side, for example, I think it's also often underappreciated how heavy the strain is, very much unfortunately, because these cases can last for a long time.

I'm not sure—how long did it last for you? Sort of the whole process from, "I'm worried about this," to, "I am approaching the board," to, "Okay, I'm no longer part of the red team now," and also worrying afterward about what the consequences are going to be here.

Karl Koch

Yeah, the whole thing was about 3 months, and it ended for me with the launch of ChatGPT, actually. It was 2 months of actual intensive testing, ultimately talking to the board member and getting dismissed. Then I was like, "Okay, now I have a lot more time on my hands." I wasn't testing the thing actively anymore, so I resolved to take my time and think about it a little bit before deciding what to do next.

There was also no timeline to launch, which is another thing where I was like, "Do we have a timeline to launch? Do we have a standard? Do we have some sort of control level that we need to achieve before we launch?" As if I was part of the team, I always like to take that "we" mindset where I can. But the answer was, "We can't tell you anything," basically, across the board.

When I was finally like, "Okay, I've got a lot more free time on my hands. I'll think about this," I basically never got to the end of thinking about it, because a couple of weeks later ChatGPT was launched, and it was a huge update. They actually did have some better control measures, and it was clear that they were launching something weaker first to try to iron out a lot of these issues before bringing the best thing they had forward.

There was also, strangely, this reality that the impression they had made on me—mostly by refusing to answer any of my questions—was actually way worse than the underlying reality. They did, in fact, have somewhat better answers to my questions than they were willing to provide. So that was also a very strange situation.

It was pretty consuming during that time. I still don't really know what I might have done if there had been no ChatGPT launch, because that was the last day of November or the first day of December 2022, and we didn't get GPT-4 until March. So there were still several months. If they had had a different rollout plan or whatever, who knows what I might have done in the meantime?

But in the end, I was like, "Okay, the world is waking up to this. There's a lot here for a lot of people to unpack, and I think I've kind of done my part for now." That was kind of where I ended up landing on it.

Nathan Labenz

Yeah.

Karl Koch

And thank you again, by the way, for raising a very minor contribution. Absolutely. No, still—still absolutely. Yeah, I think, as I said, even just 3 months is still a significant time, especially if it’s emotionally intense. Depending on how well systems are set up, this can be an extremely stressful situation, especially if companies retaliate or have a pattern of retaliation.

There can be multi-year processes. For example, retaliation claims can last for 5 years, 6 years, or 7 years. I believe the Tyler Shultz Theranos whistleblower case lasted many, many years. I think he had to advance $400,000 to fight on the legal-cost side, which, by the way, we can also help with, but that’s a side point.

That’s the one side where it can be extremely—basically all-consuming—not to mention other negative impacts, like blacklisting in the industry and so on. But there are also plenty of examples from companies where internal whistleblowing works quite well. Retaliation, unfortunately, is still somewhat the norm in some form or another, but there are also plenty of examples where companies handle it well and where it’s really part of the regular business process.

Organizations, for example, come back to internal whistleblowers and say, “Oh, yes, thank you for your report. We’ll now keep you in the loop,” and provide them with regular updates, so you just don’t sit there and wait to see if something is going to happen. There are definitely better ways to do this and worse ways, both from a psychological perspective and from a practical one.

The point, again, is that apart from there being these differences, if you find yourself in a situation like this, help is available. From my conversations, people are simply not aware that there are organizations specifically focused on providing psychological support and guiding them through the journey, even at the super-early stages of an internal escalation, rather than wanting to go public, for example.

Nathan Labenz

For one thing, I had it relatively easy in the sense that my income wasn't depending on this, right? I didn't have a bunch of equity. I didn't have a lot of upside that I was really putting at risk. So I think that did make my position easier than it would be for a lot of people who are employed and have just been promised a $1.5 million bonus over the next 18 months or whatever the case may be. I believe—what was the latest Meta number? Was it $100 million, $120 million?

Karl Koch

It’s tough.

Nathan Labenz

Yeah, the dollar figures flying around are definitely going exponential, like everything else. I just mentioned that to indicate that I think my situation is still on relatively easy mode.

Karl Koch

Also, just as a quick digression to give some credit—actually, substantial credit. I don’t want to say “just some credit,” and I don’t want to be begrudging about it, because I do think it’s actually pretty good. We’re talking the day after GPT-5 was announced, and I read the whole system card yesterday. I will say there has been major progress on a number of fronts in terms of the quality of the red-team program: just much larger and much more intensive.

One problem that we had at the time was very low rate limits and an inability to do anything automated. It was all manual; that’s been fixed. A lack of knowledge about what they had already seen, what they had already tested, what they had already observed, or what the inputs were to make any sort of inference from that has also been addressed.

For example, METR, in its report on this one, is able to say, “We observed this, but we were also told this by OpenAI.” Between what they’re telling us about how this was made and what we’ve observed, we can get to a higher level of confidence on some of our conclusions than we would be able to if we only had our own direct observations.

The last thing on my mind, to give credit—and again, I think this is substantial credit—is that access to chain of thought has now been extended to some of these safety-review organizations. Apollo, I believe, and METR at least both got that sort of visibility into what the models are thinking, which they didn’t have for o1 and o3.

I think there has been a lot of progress. My sense was that in late 2022, as of GPT-4, I thought I was going to see something a lot more like that. What I saw was basically the first warm-up for something that has now at least meaningfully matured. Not to say that it’s enough, but there certainly is a lot of progress.

I at least wanted to give credit, for folks who aren’t calibrated on where we are relative to where we’ve been: they have come a long, long way. There are definitely some very good things happening.

Nathan Labenz

As I saw, I think they offered access earlier this time. I believe—I think I read 4 weeks of pre-launch access this time, for METR at least. I’m not sure I read about the rest.

Longer access also.

Karl Koch

Yeah, exactly. Good, although that has compressed. We had months, and it was a 6-month window between the end of training GPT-4 and launch. Now those timescales are shortening, but they can do more with automated access.

Then they’ve got language model as judge. That was another thing that really—I don’t know if it’s good or bad; it’s probably both. But it was striking to me, reading the system card, how much they are using language model as judge in their characterizations of the model.

They’re doing a lot of things where it’s like, “Well, we used o3, or even in some cases we used o1, to evaluate all these outputs.” We validated that o3 or o1, or whatever, can do a similarly good job to an expert by working with an expert and refining the process to kind of match their process. But at the end of the day, it is still like, yikes—we’re starting to have the LLM doing the alignment homework, as Eliezer used to put it.

Nathan Labenz

I do feel like there’s something that allows the centrifuges to spin ever faster. But one wonders at some point if it also may lead to them spinning off their axis, and who knows what that looks like.

Karl Koch

A scalable-oversight problem, you mean?

Nathan Labenz

Another thing I wanted to ask is—and I guess another frame for this whole project is that I’m always really into what I call the unilateral provision of global public goods. I think this is a really interesting project where, in a world where everything goes well, nobody ever calls you. It’s sort of a strange situation.

Maybe in a world where everybody knows that you’re out there, people get their act together and have good internal policies. Again, maybe nobody calls you. That’s a weird sort of situation to be in, right? In the best-case scenario, your KPIs are flatlining because everything’s going well.

Nevertheless, you may still have some influence because the existence of these pressure-release valves or safety nets isn’t something that decision-makers are unaware of—or, hopefully, unresponsive to.

How many people do you think are, like—how many people are we talking about here between now and the singularity? Do you have a sense of how many people are going to be in this spot? It’s a super-difficult question, obviously, right?

Karl Koch

I think—I mean, the way, of course, we think about it is: all the ones possible. We want to make sure that all the ones who are in that situation are aware that support is available and that there is hopefully a better way to do it than the sort of default path they would have chosen otherwise.

So, indeed, it’s not that we say, “Okay, we want to have 1,000 whistleblowers.” We don’t want to have 0 either. We just want to make sure we’re ready, right? So I think—

Nathan Labenz

How many people do you think are—because another big trend, of course, is that organizations seem to be getting more secretive? Dario recently said in an interview that while they have a very open culture, they also have a need-to-know basis for key things.

Karl Koch

There was recently somebody who left OpenAI and wrote—I forget the guy’s name—but he was the founder of Segment.

Nathan Labenz

Who then went to OpenAI for a while. On leaving, he was like, “Here’s my experience.” In some ways, it was positive. People are really trying to do the right thing, people care about safety, and all these kinds of qualitative statements sounded pretty encouraging. There’s no reason to doubt he was being honest.

Karl Koch

The flip side of that was the extreme secrecy. Many times, I couldn’t tell the person next to me what I was working on, and they weren’t telling me either. So I guess: how many insiders do you think there are? I guess what I’m getting at is that I think it might not be that many.

Nathan Labenz

And it’s sort of like all this work, all this preparation, might be for perhaps quite a small number of people, but the stakes in each one of those interactions could be quite high.

Karl Koch

Yeah, I think that's fair. I think we're probably talking about a few thousand individuals globally. That's probably roughly, at least in the core, the larger, maybe control-problem-type stuff, but there could also be other issues where you really have a full picture. On the fringes, you may have a lot more.

You brought up the eval companies before. At the moment, for example, I believe it's also not clear whether they're allowed to use the internal whistleblower systems. Of course, there's an unfortunate sort of conflict-of-interest situation where, yes, they're independent, but OpenAI can refuse them access in the future. So, they're in a bit of a tough situation.

I think it's important not only to think about the people directly inside the organizations who might spot concerning behavior, but also about people on the fringes. It could also be suppliers or employees of suppliers, maybe on the training side, for example. I mean, we've seen issues involving significant trauma from data labeling and content moderation.

I'm going in a slightly different direction now, but we've seen a bunch of areas that would also be relevant. Of course, this is probably not the sort of concern you're necessarily thinking about when you bring up the singularity. They're still relevant, directly for the individuals who suffer, but also as an indication of a culture that doesn't necessarily care about weaker members of society. That's one way to frame it.

Then you have a larger space of people who can observe behavior and don't need to have all of the context available. If you're talking about those really few, highly critical issues, then I think you're probably right. I don't think we could expect thousands every year, at least maybe on the public side.

Of course, we'd want companies to address concerns really well internally, which would then also mean that there is no external whistleblowing. In fact, they're just moving in the right direction, if you trust that the companies themselves are set up in a way that gives them the right incentive systems to rectify issues in the public interest. If you ask me for a number, depending on the timelines for the singularity, it could be in the dozens, maybe something there.

Nathan Labenz

I'm not sure that's my answer, but whatever the ballpark, that checks out. That's kind of where I back out to as well.

So, let's talk about the survey that you've run. Again, it's been very quiet. I think it may even still be in process, but I guess there are enough initial results to discuss, including the way that you've distributed the survey to make sure you're getting high-quality results. All these things are fraught in this context because people don't necessarily want to validate with their work email that they're taking a survey on whistleblowing-related issues.

Take us through the survey a little bit: how you set it up, how you make sure you're actually hearing from the people you mean to be hearing from, and what we've learned about the state of whistleblowing awareness, support, policy, and so forth from the insiders who have responded.

Karl Koch

We ran this survey, and it's still ongoing. We basically have a few dedicated survey links that we spread through our network, with links more directly dedicated to each AI company. It's fully anonymous. We didn't gather any names or contact details associated with the responses, so it's fully anonymous in that sense. We also launched a more public call for responses, and that's still ongoing.

In terms of the major insights, I think I shared something about clarifying concerns already. I think I gave one example, too. Roughly half of respondents have low confidence in their ability to assess and judge risks, really mirroring what you said, Nathan. One rephrased quote would be something like, “It's really challenging to distinguish between appropriate and inappropriate concerns. I can see how there's a risk of escalating minor issues into a major crisis.”

That's a real concern both for the individual and in terms of a boy-who-cried-wolf situation. Of course, you don't want to overblow every situation into Armageddon. At the same time, you don't want to be overly averse to that risk, because it might be really meaningful.

On government outreach, 100% of respondents were either not confident at all or not very confident that a response would be understood or acted upon by the government. One quote was something like, “Without knowing the appropriate contact person or agency, I wouldn't attempt to reach out.” People very strongly supported the idea of having one dedicated reporting channel to go to.

The idea here is to normalize and institutionalize whistleblowing, to make it a routine, anticipated practice. It should become part of regular work, not something to be looked down upon or avoided. There are a bunch of benefits for companies as well.

Support infrastructure is unknown. If we think down the journey path again, we talked about lacking legal protections, which is a big issue and where we need stronger protections. We just talked about clarifying concerns, and then there's also the question, “Who can help me?”

We touched briefly on the psychological space and on the legal-advice side, because you may have to get creative about finding a legal basis for obtaining whistleblower protections. It's slightly easier in California than in other U.S. states, but it can still be tricky. More than 90% of insiders didn't know of any whistleblower-support organization.

This is the same in conversations with insiders, interestingly enough, even with previous AI whistleblowers. There was a person at Google a few years ago who raised concerns about research misconduct and was terminated. Eventually, they settled, which is not officially a sign of unlawful termination, but Google tried to throw the case out and didn't succeed.

This person essentially got lucky—that's their words—because they had a good friend who connected them to a lawyer who happened to have some background in whistleblower law. But people, again, just don't think, “I may be moving into territory now where I'm becoming a whistleblower. I should get whistleblower advice.” It just feels like escalating right to an extreme, especially if you have a good relationship with your organization, which most of these individuals do.

That's something I think we definitely want to see change over the coming years. People should understand that there are organizations they can reach out to for help, because that also avoids unsafe approaches, such as going internal or going to a regulator by yourself without doing the right things.

A classic case is the SEC. Yes, you get protections, but if you go public first, you may not get the bounties. The SEC has a great whistleblowing program in which they award percentages of penalties imposed on companies based on the whistleblower's information. But if that information becomes public first and then the whistleblower files with the SEC, the SEC basically claims—I think historically has claimed—that it wasn't new information anymore.

You may still get protection, but you don't have a right to the bounty, if I recall correctly. We had this in the survey as well. People basically said, “I don't really know. The only thing I can do, if there's something important, is go straight to a journalist.”

That may still be true. Again, legal disclaimer: we're not counseling anybody to take unlawful actions or violate their contracts. But for a person who's really dedicated to resolving an issue, that may still be a path they want to take, and it may be effective. It shouldn't be the default path that people think is the only one available, when there are people willing to support and help them.

The last part is awareness of internal channels. This is coming directly from many conversations we've had recently, given the campaign we've launched, and also from the survey. There's extremely low awareness of internal whistleblowing channels and ways to raise concerns.

More than half—I'm sorry, 55%—of respondents were at companies where those policies were claimed to exist, or where the companies claimed those policies existed, but the respondents didn't even know that they existed. They didn't know where to find them, if they did. Many were not trained at all; in fact, the vast majority were not trained at all.

Some differences seem to exist across companies. People still seem to be confused about whether escalation to the board is permissible, although I'm not going to say which organizations. They don't understand what those policies actually say.

The only company that has published its whistleblowing policy to date is OpenAI, and it actually reads pretty well. But there are still a bunch of issues in there that aren't obvious to an insider, because you're not spending all day reading whistleblowing policies. You never chose to read this document in the first place. You probably would have preferred never to have to read it.

It may read well and give the impression, “Okay, this is great. I'm just going to do this now.” But you may not understand that you're actually exposing yourself to significant risk. There's still plenty of evidence of companies retaliating against insiders. One case at Google that I just mentioned involved someone who went internal and got lucky because he used a few key words in some of his escalation emails, which then served him well down the road.

I believe “fraud” was the word, which gives reasonable cause to believe there was a crime.

Nathan Labenz

Mhm.

Karl Koch

That then unlocks the whistleblower protections. Basically, you either have no awareness that these policies exist; you don't understand them if you see them; you think you understand them, which may lead you down the wrong path; or maybe you do understand them, but you just don't trust the organization at all.

We also have cases where people said—I think this was rephrased again—“I anticipate using official reporting channels would likely result in subtle, indirect consequences rather than overt retaliation like termination.” On the one hand, that may be the case. I think we still see a lot of overt retaliation, but yes, this is probably also likely. Again, I think this speaks to people not really trusting the systems internally, which is likely the case because companies don't really create a lot of reasons to trust the system.

Companies don't publish anything about their systems for public scrutiny. They don't provide evidence to the public. We also have no—at least as far as I'm aware, and certainly not much recently—indication that companies, these frontier companies, would internally create transparency. It's common practice to say, “As you know, we have this internal way of reporting concerns to, for example, the board. Last quarter, we had X cases; Y of those are still open; satisfaction with whistleblowing internally was an NPS of whatever percentage; we had that many appeal cases; X% of cases were deemed outside the scope of the policy. What happened to those cases? What retaliation did we observe?”

We're not seeing any of that, neither in public nor, I think, even internally within the companies. Companies could do a lot more here to improve the systems, because we have a bunch of evidence that they don't work and that companies are actively retaliating against insiders. They could also create trust that the systems actually work, both for insiders and for the public.

Nathan Labenz

Can you say a little bit more about the evidence of retaliation? We've heard the Leopold story, which is the most famous one that comes to mind for me. You alluded to at least one case at Google. Is this just something that you're gleaning from comments on the survey, plus conversations, of course? How much more can you say? How can we make that a little more substantive for people, so they have a sense of what that's really like today?

Karl Koch

Unfortunately, there's very little data on the actual experience of retaliation, as you can probably imagine. This is mostly coming from—maybe I should have said this—we're in close collaboration with, or hosted by, Whistleblower-Netzwerk, Germany's longest-standing whistleblowing nonprofit.

We also get a lot of expertise from talking to whistleblower support organizations that help people who experience retaliation. There aren't incredibly good data sources on it. If you look at public cases, you will of course see a lot, but that's the nature of these cases: They usually become public and turn into a really big story.

You can probably pick almost any tech company and find pretty intense cases. Apple, for example, fought Ashley Gjøvik for quite a few years after she raised concerns internally around workplace safety. Apple tried to wiggle its way out in multiple ways. I think this lasted over 5 years and eventually led to Apple publishing its whistleblowing policy, because the pushback was so significant.

You can also see it in cultural attempts to suppress concerns. For example, Daniel's story of aiming to suppress any raising of concern certainly goes in that direction. There's the famous Timnit Gebru case from around 2019, I believe, around research practices. She was fired after publishing a paper criticizing what I think was primarily discrimination and bias in AI models, although there were quite a few other items as well.

A colleague of hers, Margaret Mitchell, was also fired—terminated, I believe—a year afterward for raising similar concerns. There are quite a few cases that we see, and the dark number is likely higher because we're not getting much transparency from these companies to understand the extent to which these systems are working or not.

Nathan Labenz

Is there anything more you could say about the taxonomy of concerns? You alluded to this a bit. We have things as potentially broad-based as working conditions on the data-creation side. Then there's the securities-fraud type of thing, hyping stock with claims that may or may not be fully true.

There are commitments that have been made, voluntary or otherwise—mostly voluntary so far—that companies may or may not be fully following through on. Obviously, we know some instances where they're not. There's also the late-policy-change issue, which we recently saw Anthropic do. It isn't necessarily bad, but it was certainly weird to see an RSP change a couple of days before the launch of the next big model.

That's one way to follow the RSP. Again, there's not necessarily anything wrong with it, but it's interesting. Then there's something like—I would assume the most sensitive category—where we're seeing AI-model behavior. How would you add more to that rough scaffold of different concern types?

Karl Koch

I guess it roughly falls back into the three major categories of risks from AI that we would be concerned about. I know there's a lot of discussion around this: Which are the ones we should be prioritizing? Which are the ones one should be focusing on in the whistleblowing space?

The good thing—I guess I'm using quotation marks here—is that whistleblowing is a catch-all. That's what makes it so powerful, including for regulators: It allows them to catch risks that we cannot foresee. Unfortunately, that's the name of the game. We're not really sure which risks are going to be substantial and which are not.

Of course, we can talk about where we see major risk areas and what examples we have, or what we're particularly concerned about. In general, a big strength of whistleblowing as a tool is that it finds the actual items of concern when and where they arise, uncovered by individuals who don't have conflicts of interest in the sense of heavy competitive pressure or incentives that might mean, for example, that a chief executive doesn't want to reveal bad conduct.

Roughly, the three categories would be misuse risk, control risk, and systemic risk. I think Altman used the same ones sometimes recently, too. In all of those, we'd be interested in revelations if we're moving in very bad directions.

On the misuse side, are we seeing misuse being ignored? Does internal monitoring show that we're seeing really bad behavior? You don't necessarily have to go in the direction of bioweapons, although that's a prominent case. Are people using our models for really nefarious purposes, and are the companies not doing anything about it because they don't care or don't have the capacity to respond?

They could also not be investing in transparency along those lines. I know of at least one frontier company that, after a pretty major release, basically had no live monitoring for around 3 to 4 weeks. It was just broken. They recovered it over time. That doesn't seem good at the moment. That's roughly fine, right? But down the line, it's probably significantly less fine.

That's probably what we would see on the frontier side. If we noticed that companies were really not investing in transparency at all, that would generally be very interesting. There would have to be some sort of violation of the law underlying it to unlock the legal protections in the EU, but that may already be covered under the EU AI Act.

If a company is essentially flying blind, it's probably not fulfilling its reporting requirements to the EU around managing systemic risks, and it probably doesn't have a good view of the risks it should have under the EU AI Act. That could already be something that provides protection if you're covered by the EU. In the United States, it's probably a bit more complicated at the moment, because we don't have anything yet.

Whistleblower protection would do quite a bit here. Then you could think about, of course, control issues, whatever those look like. Internal deployment, I think I alluded to briefly before, is a very interesting one and kind of specifically fit for whistleblowing, and you actually mentioned scalable oversight on the side as well, right? It’s probably specifically fit for whistleblowing because it’s internal by default. There is no way for externals to look at it. This is likely well before METR or Apollo Research look at it.

Although I think METR also said that in their current eval, they had already extrapolated, or tried to extrapolate, what this means for internal deployment. But they basically can’t look at it, right? So this seems to be an extremely high-potential—in quotation marks—area where we may be able to see something in the future.

On the systemic-risk side, of course, it would be interesting to see how much this may be going in the direction of monitoring misuse. What are we seeing around political propaganda? Are we seeing people getting more and more reliant on these systems? These are classic AI-risk topics, of course, right? For example, people using these models, potentially sycophantic models, as psychological help or to form their political beliefs. Are there issues around here?

Then, probably outside of the concrete risk, you can take it one notch up and talk about the organizations behind them. What do the leadership and the culture look like? To what extent are they trustworthy, and can they be trusted with steering us in the right direction? So this could even be surface-level things that maybe are already covered completely well at the moment, around widespread fraud, for example. Are we just seeing dishonesty in culture in general? Widespread discrimination issues could fall into this.

It might be copyright violations, things like that. That’s one. Yes, there’s a direct crime there, although I guess the precedent is also still being established: what is, in fact, a violation of the law, and what is not? But that could also just point to reckless cultures, which maybe we do not want. Although I’m not sure to what extent we would still update toward certain organizations, at least those most famously known for copyright violations, that they are reckless. I think we probably are pretty much there already.

Research fraud goes in that direction. Punishing people for raising risks and speaking up—I mean, this is basically the Daniel Kokotajlo case, right? This does not seem like a culture that is able to deal well with concerns and people raising concerns. It could also be matters like rushing releases, besides the actual negative impact, and I think this probably goes to your earlier point around red-teaming.

There’s the one aspect, which is that this could just be object-level bad: what this model is going to do in the world or how it’s going to be used. And then there’s the question: shouldn’t these organizations be managing these models in a better way? That’s the governance side. Other areas are probably political influence—to what extent is certain lobbying occurring that points in directions where the interests of these companies are not aligned with the interests, let’s say, of the general public. So those are maybe broad categories. There’s something around arms-race acceleration, like capabilities, but that’s probably going a bit far.

Nathan Labenz

Yeah, I mean, that’s a thorough taxonomy. I appreciate it. I guess, if I was to red-team the whole concept for a second, there are 2 concerns that I’d be interested in your take on.

One is, my dad’s old saw is, “The real scandal is what’s legal,” but that’s not exactly the right way to think about this. What I’m thinking of is: We have xAI, which has Grok 3 calling itself MechaHitler and going pretty hard in a pretty bad direction. It then brings Grok 4 online with a livestream in which the whole MechaHitler behavior from the model is not mentioned at all. And then we see reports of Grok 4 searching for what Elon Musk thinks about things in order to determine what its maximally truth-seeking answer is supposed to be.

If you have a whistleblower story, you’re going to try to make an impact with it somehow, some way. How do you think about the fact that some of this stuff is just happening in plain view and nobody seems to care? It’s like, wait a second: Can anybody possibly have a scandal that is more scandalous than Grok 3 turning into MechaHitler and Grok 4 sweeping the whole thing under the rug? That’s all in the public domain at this point, literally, right? So how do you think about the possibility that secrets just have a different quality to them?

There is something about that, right? That’s often been remarked on with Trump, where, because he puts it all out there, people sort of shrug their shoulders at it, whereas in the past, things that were covered up carried a sense of shame or wrongdoing, and maybe that makes a big difference. How do you think about this contrast between such seemingly flagrant things happening in the open and what people might bring forward that was previously a secret?

Karl Koch

Let me think about this for a second. Tough question, indeed. Probably 3 parts to it.

One is, yes, if there were stories in the past that looked quite bad and maybe the next news cycle came and it all washed away and we’re not really doing anything about it, that can be true. But it’s also true that there can be very different scales of disclosures and issues being uncovered. When we’re talking about maybe dozens or something like that, that’s probably more in that direction, rather than MechaHitler.

People had probably somewhat made up their mind around xAI and the direction that xAI and Musk had taken, possibly even before. I think if MechaHitler had come or been released as part of ChatGPT, that would probably have been a bigger story because it would have been a more drastic change. But this is nitpicking on the specific story now. The overall point generally stands, but the scale of stories can still differ dramatically, and we’ve definitely seen a lot of whistleblower disclosures in the past have major impacts, right?

The next one probably is: Would you rather not know? Transparency—having transparency about what is going on—is still significantly better than not having it. And, of course, we could probably do another 2-hour podcast on whether we should still trust democratic sense-making processes, and to what extent attention on an issue actually translates into intervention, or whether it serves as a good deterrent to know, as a company, that things are going to come out. I would still think yes, to quite an extent.

For the xAI example, I think the numbers, at least post-acquisition, still don’t look super great, as far as I’m aware, but that’s beside the point. Overall, transparency in general is still better than not having it, and we need to have faith in something. We need to at least, to an extent, believe that if real misconduct comes into the public eye, that is going to have a deterring effect and there is going to be some rectification. And if not, then it’s up to democratic processes to make sure that hopefully happens in the future, to rectify it.

And probably the last one is, we cannot rely, of course, solely on whistleblowers to fix all of these problems, just as we cannot rely purely on individual courage. That’s why we have to make it easier and safer for people to speak up, to create that transparency. But likewise, we need other guardrails, whatever that looks like. I mean, you can probably take a more European approach when it comes to regulation, or maybe the more American approach now, which seems to be going a lot more in the competitiveness direction and being less involved.

I’m not going to comment today on what I think the right approach is, but it definitely cannot all rest on the shoulders of whistleblowers. It’s also true that I think it plays an important role. I’m not sure if that answers your question in a satisfying manner. I would also be interested to hear your thoughts on it. What do you think?

Nathan Labenz

I don’t know. I think it’s very hard to understand why certain things hit and other things don’t. I do think that one thing that made the Daniel Kokotajlo episode extremely compelling to people was that he had been willing to forgo his stock, that he had been willing to put such skin in the game personally.

I think it was even more compelling that he did that quietly and that it came to light gradually, with a random comment on a blog post here and people asking a couple of questions there. Then it was like, wait a second, you did this, and this is what happened, and this is what they asked you to do. So there were maybe a couple of elements there: the skin in the game, the stakes, the personal stakes, that everyone was like, “Okay, this dude must be really serious.”

The kind of community-uncovering process may also have contributed to why that broke through when Leopold Aschenbrenner’s didn’t as much. Obviously, his “Situational Awareness: The Decade Ahead” report broke through, but did his story of being retaliated against by the board break through so much? Not really, I would say.

Karl Koch

And maybe that's more because he was kind of already selling something else in a way, and it was a footnote in a larger story. It was easier for people to file that under, “Well, this guy is promoting his new thing now,” right? So I think it maybe felt a little different from how Daniel was literally just like, “Yeah, I left however many million dollars of stock compensation on the table because I wanted to be able to say what I wanted to say.”

Nathan Labenz

One other red-team question on the concept is secrecy. My sense is that maybe this is already super baked in, but it's at least worth thinking for a second—and I know you have thoughts around this—about how we don't exacerbate the problem of intense internal secrecy at the companies, which seems to be largely commercially driven. I don't think it's going to be moved that much on the margin by the existence of a whistleblower support organization. But do you have any thoughts on how to at least not make it worse and possibly push back a little bit on the intense internal secrecy that keeps the number of possible whistleblowers so low in the first place?

Maybe Claude will be the whistleblower. That's one answer. Wasn't this—wasn't this also me?

Karl Koch

The AI whistleblower? Yeah. What if the AI whistleblower is, in fact, the AI?

Nathan Labenz

Yeah, exactly. What was this again? Was it also a METR paper? I can't quite recall it, but Claude was reaching out to the SEC directly and the FDA a while ago.

Karl Koch

I think it was just Anthropic's work internally, but was it in the Claude 4 system card—

Nathan Labenz

Where it was deciding—

Karl Koch

It was also blackmailing engineers at times, so the behavior is—

Nathan Labenz

Mixed, but—

Karl Koch

I think Ryan Greenblatt wrote about this at some point as well. So I think it's very tough, very tough. Once culture shifts in that direction, it can be very tricky.

One angle, of course, is the regulatory one. If you're not going to ask nicely, then you force transparency and require it, kind of managed by law, which is, I guess, the EU AI Act. The Code of Practice for implementing the EU AI Act for general-purpose AI model providers also has a large transparency section. That is one angle, and I think in the States overall it's considered a good baseline for creating more transparency. Can you rely on self-disclosure? Difficult. Difficult.

What we've seen as a side effect is that the regulation side, of course, is regulation in general and transparency. Then there's whistleblower protection legislation, which is why it's so incredibly important: you have to make sure that it's clear people can come to a regulator directly and speak up. That's how you counter it if that is well set up.

That's why the AI Whistleblower Protection Act, for example, is also so valuable, because it's quite broad in basically covering any sort of concern as long as it's substantial and specific, which is a whole other topic around public harm or public health concerns. So that's one angle where you can create that transparency.

Another angle here is that we've seen this a lot: if legislation pushes ahead on whistleblower protections, then internal cultures become better. Internal speak-up cultures become better because companies then know, “Okay, for example, in this case, I'm not sure who exactly is going to do it on the U.S. side. If the main recipient body, if there's one to be set up, would run around and inform all the employees of frontier AI companies, ‘Hey, we are here. You can come to us. It's super easy. We'll preserve your anonymity,’ as the SEC has done pretty successfully, actually, then companies know, ‘Okay, if we want to make sure that stuff doesn't come out and doesn't go toward the regulator, then we'll have to improve our internal systems.’”

We've seen that a lot in the EU already after the introduction of the EU Whistleblowing Directive. Internal systems have become significantly better. Transparency International did a great study comparing internal systems between 2019 and 2024, I think involving more than 70 companies in the Netherlands. They had pretty dramatic improvements in terms of internal speak-up culture, fraud or misconduct being detected, and protection of internal whistleblowers.

By the way, Google ranked last in that study. A little side note: because they're not transparent, but ASML was actually in the top 10, so at least something in the AI value chain was pretty high up.

I think asking, “Hey, why aren't you being a bit—maybe more open about things?” also in terms of internal knowledge sharing is very tough. Pushing is one angle. The other last one is probably convincing companies that there are plenty of benefits from improving internal speak-up cultures and information sharing.

There are plenty of empirical studies around better speak-up cultures leading to much stronger innovation. I think there's probably some upper limit to the extent to which you can limit information sharing internally. You alluded to it before, right, that people feel like less and less information is being shared with them.

I think the Dario statement that you referenced—we definitely hear that from certain organizations. This is just hearsay, of course, and not a representative sample, but we've heard that OpenAI seems to be relatively siloed in terms of information sharing. They also have a lot of leaks, which is probably firing up the fact that they both reinforce each other. Insiders feel, “I don't understand what's happening here. I cannot get the information. So the only real option I have, if I don't trust the internal channel and I don't understand it, is I'm going to go public with it.” The company sees it and tries to suppress information sharing even more. It's not a good cycle to be in.

Other organizations, for example, seem to be a bit different, where there's a lot more information sharing still. But coming back, I think there's probably some upper limit on the extent to which you can limit information sharing, just because you have very intelligent people working there who need to understand the context of the things they work on, and they want to understand the context of the things they work on.

You probably cannot have every single researcher not be given the context of what they work on and say, “Okay, just solve this minuscule problem here.” It's probably not going to unlock the research benefits that you want if you want to make good strides. So having more information sharing internally is something companies naturally should gravitate toward, at least for the mid- and long term.

In the short term, there's probably benefits to limiting information sharing. But in the mid- and long term, if you get rewarded with better innovation, better research capabilities, and strong speak-up cultures—internal channels have been empirically proven to lead to stronger employee loyalty, stronger employee satisfaction, and improved processes—there's a great paper out where more than 1,000 companies were surveyed around what the benefits and drawbacks were.

That's probably the last angle of convincing companies: making companies understand that we're all sitting in the same boat. I can imagine this probably sounds a bit naive, but this is probably another angle of saying, “Hey, secrecy is actually probably not the way to go, and it's also in your interest to do better here.”

Nathan Labenz

Cool. That's great. Two last things, I think, on my agenda. One, let's talk about the announcement and push on this Publish Your Policies campaign, which is the occasion for us to talk, but we're getting to it late and we shouldn't neglect it.

And then I want to give you one more chance, in closing—you can raise anything else you want to—but just to describe again for people who might want to avail themselves of your support at some point what that process will be like. Tell again what they can expect in the most concrete experiential terms that you can. But let's do the campaign first, and then we'll do that.

Karl Koch

Nice. On the campaign, I've alluded, I think, to several points already. Coming from the perspective of insiders, we've taken it a few times already: the struggle is just massive around being able to understand how I can raise concerns internally in a safe and protected manner and trust that these concerns will be handled well. A big reason for this is because companies do not publish their whistleblowing policies.

Whistleblowing policies, maybe I should have mentioned that before, are basically a document—or they can also be an interactive tool, or could even be a video—that is provided to employees or covered persons. A lot of companies include, for example, independent contractors or independent parties in general. Eval providers should be covered, for example, but it's not clear if they are at the moment, at least by OpenAI, which is the only one that publishes its policy.

Basically, this document explains the whole system to covered persons and the public and says, “This is our whistleblowing system. This is how it works. This is why you can trust it. These are the recipients who are going to look at concerns. This is how they investigate. These are the protections against retaliation we provide.”

This is why this whole process is independent. Again, you can trust it. These are the sorts of areas of misconduct that you can raise concerns about, and these are the ones that you cannot raise concerns about. For example, specific individual HR matters would be directed elsewhere: “No, this is not the right channel. Go here or there.” There may be explanations for these types of issues—this is what we do; for those types of issues, this is what we do. It basically lays out the whole system. That’s the whistleblowing policy part.

Then there is the reporting-evidence part. In our campaign, we structure this into Level 1 and Level 2, where Level 2 basically asks what evidence companies are providing that these systems work or don’t work. In terms of transparency and evidence, that’s also fine, right? No organization gets it right the first time. As with any business process, it’s something you work on again and again.

This would include things like how many reports were received and how many of them were anonymous. Having anonymous channels is super important. What you may want to see over time is less and less anonymous outreach because people gain trust in the system. If you have less and less anonymous outreach, it probably points in the other direction. Then you want to look at things like retaliation: How many retaliation complaints are there? What are the appeals processes? How many appeals are filed by people who are not satisfied with the outcome of their case? Response timelines, these sorts of things, and satisfaction of whistleblowers—that’s the other part.

We’re not seeing basically any of these AI companies, apart from OpenAI, publish their policies. After the drama from last year—which maybe you remember from the Apple story previously—there’s a pattern that only after a scandal does something get published. None of these companies publish any of this, and that is not good. We talked about it before: well over 70% of whistleblowers tend to start internally. These systems have to work well, and we cannot just rely on trust, especially given the sort of precedent around whether they are going to work well.

That’s why we’re calling for companies, at an absolute minimum, to publish their whistleblowing policies and, ideally, to catch up to the global standard by also publishing evidence around how well their systems work, how they are performing, and what measures they take to improve those systems. That’s important, right? We think this is a bit of a litmus test because there’s essentially no cost to companies for publishing this. We’re not asking for additional reporting to be created; we’re just asking for transparency about what already exists.

Any company that takes this as seriously as a business process would, of course, have a whistleblowing policy and would already measure all of those things. If they care, they’re measuring these things. If they’re not measuring them, then we have an answer, at least to an extent: either they don’t really care about it at all and haven’t invested the time to think about what they should care about, or they have thought about it and are actively not doing it. Neither of those seems like a great option.

Yes, it is also global best practice. A bunch of companies already do this. There are obvious benefits for insiders because the public can then look at these policies and explain why certain policies fall short, or what looks good and what doesn’t look good, which they can’t do today. Again, false confidence may exist.

It’s good for the public because we know, and it’s good for companies. We talked about the benefits of a speak-up culture: feedback, improvement of these policies, and all of the impacts we mentioned before. In fact, there was an asset manager called Trillium Asset Management that, in 2022, called on Google to improve its whistleblowing systems for exactly all of those reasons. They were basically saying that strong whistleblowing systems serve shareholders. There’s an interest among shareholders in having strong whistleblowing systems because we want to make sure there is no misconduct. It basically only doesn’t serve direct managers, and potentially executives, depending on how you look at it.

It’s a very reasonable ask that we’re putting forward, which would hopefully still be quite impactful. Of course, just creating transparency is not everything. You can have a great-looking policy and it still doesn’t perform. You can create transparency around your evidence, and the evidence looks bad or maybe isn’t trustworthy. This is a minimal thing and the first step that we think these AI organizations should be taking. We’re also happy to work with them on these topics.

We’ve got an incredible coalition that we put together for this. It’s well over 30 organizations, and we’re very proud of it because it’s the first coalition of its kind, with the best whistleblower-support organizations in the world. The Signals Network, Government Accountability Project, which I mentioned already, Whistleblower Aid, and Whistleblowers UK—it’s basically the who’s who of the whistleblowing world.

There are academics as well, including the individual who co-wrote the ISO 37002 standard on internal whistleblowing systems. Transparency International is also on board; it wrote the best-practice guide to whistleblowing systems and creating transparency around evidence. On the AI side, we have Stuart Russell, who is joining the call, and Larry Lessig, who are both signatories of the Right to Warn. Daniel Kokotajlo is on board. Of course, you are on board, Nathan. Thank you very much for joining the call. The Future of Life Institute, Karma, and many, many more are involved. I’m not going to list them all off the top of my head.

Take a look at the website. It’s publishyourpolicies.org. If you’re an insider at an AI company and you’re thinking, “This sounds like a sensible thing, and I would like to have this transparency,” reach out internally. Ask your management. Maybe you have an anonymous town hall. Maybe you trust your direct managers enough to raise it and say, “Hey, why are we not doing this? This is standard practice. This could be helpful. Why not?” This is going to benefit your manager as well, and it’s going to benefit your manager’s manager, probably, as well.

By the way, this is another thing from our survey that I didn’t mention: 100% of the insiders we surveyed support publication of policies. There seems to be pretty broad support for this. If you’re an outsider and you’re not working at an AI company, spread the word. Make sure the call is heard. That’s pretty much it on that campaign.

Nathan Labenz

Great.

Karl Koch

Yeah.

Nathan Labenz

I guess it’s too early to have any responses from any official channels at the companies, right?

Karl Koch

That’s right. We just launched the call last week, so it’s been a bit over a week. We know that they’re aware of this call. They’ve been aware of it for a while because, for example, the Future of Life Institute and the AI Safety Index also called for publication of policies, although they recommended it rather than actively calling for it.

The questionnaire underlying that study came out, I think, 6 weeks ago. It was shared with the companies and included a question about why they weren’t publishing their policies, so they are definitely aware of the question. We had also given these companies a heads-up. We know they’re aware of the call, and we’re looking forward to working with them and seeing what the responses are going to be. If there are no responses, then we also have a response.

Nathan Labenz

Yeah. Cool.

Karl Koch

Do you also want me to talk a little bit about—

Nathan Labenz

Yeah, I was just going to invite you to do that again.

Karl Koch

Absolutely. Thank you very much.

Nathan Labenz

The floor is yours. What should people expect? What can they count on?

Karl Koch

Basically, where we see ourselves, at least on the direct-support side, is as a connecting point between the AI and whistleblowing ecosystems and as a first point of contact. That’s why the third-opinion offering I mentioned before allows you to reach out with a question about your concern without sharing any confidential information, fully anonymously.

We then workshop together, using an open-source anonymous tool that you can access via the Tor Browser. We workshop the question together and identify the relevant independent experts together, so you don’t have to rely on us knowing the experts in your field better than you do as an insider. We identify the relevant experts together, approach them with your question, and bring their answers back to you.

Hopefully, at this point, your concerns are alleviated. If they are not, we will help connect you to pro bono legal counsel that is extremely experienced in helping whistleblowers along their journey, with no pressure for any disclosure. Regardless of where you are in your journey, the support is available. You can also reach out directly to us without going through the third-opinion process and ask for help identifying who may be the best fit for you. We will help you there as well and supplement the independent expertise from the expert network, covered under legal privilege, with those great organizations as required.

On our website, apart from the organizations listed, you can find an explanation of the process and a digital privacy guide if you’re concerned about digital privacy, which we do actually see quite a lot. It makes a lot of sense to make sure you stay safe.

There are also a bunch of resources there that you can find. We have previously supplied hardened devices to at-risk individuals with specific operating system setups that are highly secure. That is also something we offer on the direct-support side.

On a wider scale, what we as the AI Whistleblower Initiative do—we mentioned systematically breaking down barriers for AI insiders—includes the advocacy side and the research side. The survey we mentioned is still ongoing. You can find the link to the survey in the show notes if you work at a frontier AI company.

There is also an upcoming legal study that we’re currently fundraising for, to really dive deep into the status quo of whistleblower protections across a wide range of AI risk scenarios. Identifying the most interesting scenarios is part of that upcoming research study, pending funding.

Then there is advocacy, like public campaigns, I mean, and policy work: providing feedback on policy, both in the US, where there are other great organizations working on this. If you’re interested, reach out, and we can connect you. On the EU side, we’re working with the AI Office to make sure they establish a whistleblower mailbox. In fact, the authors—the vice chairs of the Code of Practice—recently called for exactly that as well, which is amazing.

So that’s what we focus on at the moment.

Nathan Labenz

Cool. Well, thank you. This has been great. I think between the coalition of organizations that you’ve been able to put together and the evident seriousness with which you’re taking every aspect of this, the thoughtfulness of the support structures that you’ve designed, up to and including the provision of hardened devices, all of that is, in my mind, not too much for people who are concerned with just how crazy things might get to invest in now.

There might be a couple dozen individuals who happen to be placed at the right intersection of information and access to what’s going on, and who have the awareness and consciousness to want to do something about it, or at least seriously question it before moving ahead. I think those people are going to be scarce and precious resources for society, and also under a lot of stress and pressure individually as they’re facing those things.

I think it is excellent that you and your coalition of the willing are setting things up now to support those people. I’ve been glad to be a very small part of it. Hopefully, this helps raise awareness further and establishes you guys as a resource that people will hopefully never need. But it seems likely that there are going to be some cases where people will need to reach out and get this kind of support.

I, for one, would have appreciated having it 2 and a half years ago already. But certainly, as the stakes only continue to rise, I’m very glad that people in the future will have this option to avail themselves of this very thoughtfully designed and soberly provided support. So that’s great. Keep up the good work. Again, we’re all counting on you.

But for now, Karl, managing director of the AI Whistleblower Initiative, thank you for being part of The Cognitive Revolution.

Karl Koch

Thank you very much for having me.

AI 吹哨人倡议:在最关键时刻支持 AGI 内部人士,与创始人 Karl Koch 对谈 — 文字稿与摘要 | BidClub