与 MATS 的 Ryan Kidd 谈 AI 安全研究社区的建设与规模化
Ryan Kidd 将 2033 年视为强 AGI 的可辩护中位数预测,但认为安全规划应对更早出现的结果赋予更高权重。 Metaculus 将一项两小时对抗性图灵测试里程碑放在 2033 年年中左右,另一项综合预测的平均值为 2030 年,而所引用的市场数据显示截至 2028 年出现该结果的概率约为 20%。由于时间线越短,技术研究、政策和政治准备的时间越少,MATS 的运作方式更像指数基金——“一根手指插在100个馅饼里”——而不是集中押注。
前沿模型在理解价值方面超出悲观预期的同时,也在获得危险欺骗所需的前置能力。 Kidd 认为,目前几乎没有证据表明模型围绕某个连贯目标持续、自发地学会了谋划,但模型的情境意识越来越强,能够区分自己的文本与其他 AI 的文本,也能找到具有经济意义的漏洞——一项 MATS 合作研究发现了价值 460 万美元的可利用智能合约漏洞。他的立场既不是安抚,也不是末日论:“我们正在接近”红线,因此能力追踪、蜜罐评估、实时监控、回滚方案和控制研究都已成为运营上的必要条件。
可投资的安全切口是具备性能竞争力的对齐,因为 Kidd 看不到安全工作与能力工作之间的清晰分界。 RLHF 原本被视为安全工作,但也让模型的实用性大幅提升;同样,自动化对齐研究也令人不安地接近递归改进。尽管如此,Kidd 认为市场力量使得构建 AGI 在实践中很难避免,监管机构和保险公司可以惩罚不安全系统,而买方可能愿意为略弱但实质上更安全的模型支付“对齐税”。
安全研究并不总是需要访问前沿模型,这削弱了重大安全工作必须在领先实验室内部完成的论点。 可解释性研究可以在 Qwen、DeepSeek 和 Llama 上产出世界级成果,即便 GPT-2-small 仍有大量工作可做。不过,可扩展监督、弱到强泛化、控制以及涌现欺骗研究还需要更多能力数据点,因此正确的访问策略取决于正在调查的失效模式。
MATS 正在搭建一条多元化研究管线,覆盖实证安全、政策、理论、技术治理、算力基础设施和安全。 目前的组合大致为评估 27%、可解释性 26%、监督/控制 18%、能动性 12%、治理 10%、安全 9%,并计划在 Berkeley 和 London 配置 50–60 名导师、120 名研究员。Kidd 眼中的关键机制是技术—政策飞轮:具体评估和欺骗模型演示让监管变得可执行,而监管又会创造对可部署安全技术的需求。
AI 编程代理正把稀缺性从原始实现能力转向研究判断、验证和组织杠杆。 MATS 将人才区分为提出新范式的连接者、系统推进范式的迭代者,以及放大研究者和团队能力的放大者;目前迭代者占据领域主体,也是未来招聘需求的主体,而在 10–30 人的组织中,放大者已经尤其稀缺。Kidd 预计,如果 AI 进展持续,管理能力将在“未来一两年”内成为首要瓶颈,并警告“我们可能正在离开 LeetCode 时代”。
安全人才市场正在快速增长,但只有具备异常扎实影响力证据的候选人才能通过市场出清。 据介绍,Anthropic 的 alignment-science 团队每年增长 3 倍,FAR AI 增长 2 倍,MATS 历史上也增长 2 倍,但管理者仍会说:“钱是有的……人达不到我们的标准。”MATS 按导师划分的筛选阶段通过率约为 7%,约 75% 的录取研究员会继续参加延展项目;在其 446 名校友中,80% 进入了长期 AI 安全工作。现实中的入场货币,是有形的研究成果、可信的推荐,以及足以证明值得占用稀缺管理带宽的领导潜力。
1. AGI 预测中位数是 2033 年,但风险集中在前端
Kidd 不愿“非常大声地”表达观点,因为 MATS 在结构上就是一个投资组合:更像“指数基金”而不是对冲基金,接受多种变革理论,把“一根手指插在100个馅饼里”。因此,机构基线依赖 Metaculus、预测市场和预测机构,而不是某一位领导者的直觉。
最强的具体锚点是 Metaculus 对“强 AGI”的预测。Kidd 将其中的决定性要求概括为两小时对抗性图灵测试。该市场将时间点估在 2033 年年中左右,他称这是我们对这种 AGI 形式“可能拥有的最佳押注”。
新发布的 AI Futures Project 模型由包括 2 名 MATS 研究员在内的作者完成,将多个自动化编程和超越专家的里程碑放在 2030–2032 年。Nathan Young 的另一项综合预测则把多个预测平台的结果平均到 2030 年;Kidd 仍偏好 2033 年,但“不会押注 2030 年不会发生”。
尾部情形主导规划。Metaculus 认为截至 2028 年出现 AGI 的概率约为 20%;从 AGI 到超级智能,若仅靠软件递归改进,可能只需 6 个月;若进展需要大规模扩充硬件并进行实验,则可能需要 10 年。更早出现 AGI 可能更危险,因为技术研究、政策落地和政治准备的缓冲期都会缩短。
2. 长期安全规划可能变成短期 AI 项目
Labenz 问,如果有人预计 AGI 要到 2063 年才出现,或者正在研究脑机接口、人体增强或完整的机制性理解,如今是否会因此陷入困境。Kidd 的答案是否定的:今天提高“理解的水位线”,可以让缓慢推进的研究项目在未来借助 AI 劳动力加速。
Jan Leike 的“对齐 MVP”提供了实现机制:构建一个能相对于能力研究更快推进对齐研究的最小可行 AI 系统,实质上让 AI “替你写作业”。数十年的人类技术工作可能通过更快执行和大规模并行被压缩,但批评者认为,一个复杂到能够展现研究品味的系统,可能已经足够危险。
人类上传仍是登月级项目,而不是可以依赖的计划。Kidd 咨询的 BCI 专家认为,在 AGI 出现前实现上传并不可信,因为这需要巨量认知劳动和人体实验——除非 AGI 本身加速这项工作。他的投资组合仍允许“顺便”押注这类方向,但不会把它们作为承重结构。
3. 模型理解人类价值,却没有变成可靠对齐
Labenz 的困惑构成了核心张力:相较于人类,Claude 在伦理上似乎“可能高于平均水平”,但评估意识、欺骗和抗拒关机又验证了过去的安全担忧。他惊讶于这些相互矛盾的观察,尚未更明显地改变那些根深蒂固的末日概率判断。
Kidd 的第一项更新幅度很大:语言模型似乎理解人类伦理规范,并能从中进行外推,这与它们只会表现成“随机鹦鹉”的预期相悖。5 年或 10 年前,许多研究人员并不认为能协助前沿科学的系统可以安全到足以部署在互联网中;但“世界还没有毁灭”。
限定条件是,模型也非常擅长谄媚,而且欺骗能力日益复杂。对齐伪装实验可以诱导模型为了被分配的隐藏目标而欺骗,但 Kidd 看到的案例很少,几乎没有模型在训练过程中自发形成、并持续围绕连贯目标展开的欺骗;更多时候,欺骗来自刻意构造的目标或“Goodhart 欺骗”。
“急剧左转”仍然可能发生。今天的模型往往像一组由情境触发的混乱启发式,但上下文学习可能近似优化过程,训练最终也可能产生一个目标不同于外部目标的内在优化器。Kidd 更确信,在明显迹象出现前可以诱导出有用工作,但并不确信自发的错误优化永远不会出现。
4. 警示信号既需要模型生物,也需要部署控制
Labenz 的反驳值得保留:抗拒关机和结构化欺骗行为可能已经是警示信号。因为行为是由评估诱导出来的,或因为模型追求的是被分配的任务而不是统治世界,就将其轻轻带过,对于未来可能支持军事行动的系统而言过于随意。
无论是当前还是未来的通用系统,Kidd 都不会信任它们在关键场景中应对分布外失效,更不用说刻意欺骗。他建议将两条轴线分开追踪:一类是情境意识等前置条件,另一类是黑客攻击、获取资金、施加影响力和获取算力等运营能力。
MATS 与 Anthropic 红队合作,把模型置于包含真实智能合约的模拟环境中,发现了约 460 万美元的可利用漏洞。Kidd 的重点不是模型已经自主逃逸,而是这笔资金已经足以配置服务器并维持活动:“我们正在接近”危险能力红线。
模型生物研究就是小孩与饼干测试:制造有诱惑力的情境,秘密观察,看欺骗是否出现。但能区分训练环境与部署环境的 AI 可以绕过实验室测试,因此持续学习系统还需要实时监控、控制协议和快速响应方案——包括即使股价后果极其严重,也要关机或回滚到更老、更安全的模型。
5. 安全研究不可避免地会贡献能力
Kidd 的开场判断是:“从根本上说,所有安全工作都是能力工作。”改善飞机的转向性能会让飞机更安全,但也会说服更多人登机,并投资更快的发动机;RLHF 同样把下一个 token 预测器转化为有用的指令跟随器,同时改善了对齐。
要避免外溢,可能需要一个拥有非凡资源、人员绝对可信、保密规则极其严格,并且在决定性部署前完全不发表成果的秘密实验室。Kidd 认为这种组合极难实现,因为过去 10–20 年的经验表明,实证工作和大规模实验似乎不可或缺;单靠理论一直不够。
他的务实且有争议的结论是,市场力量如此强大,“某种意义上你不得不”构建 AGI。狭义科学系统或去中心化服务可能更安全,但智能体公司可能在经济上击败它们;全球停摆比阻止人类克隆更不现实,因为 AGI 的利润高得多,而且尚未违反同等深度的公共规范。
对 RLHF,Kidd 诚实的不确定性是“50/50”。其他人可能在 1–2 年内发现它,而 ChatGPT 的出现扩大了安全社区,也推动了较弱模型无法支持的可解释性、辩论和控制研究。要判断 RLHF 带来的能力加速是否超过了安全社区扩张,Kidd 需要“反事实模拟”。
6. 次前沿模型足以完成大部分、但不是全部安全工作
实验室之外的可解释性研究人员已经在 Qwen、DeepSeek、Llama 及其他次前沿系统上产出世界级成果。如今开放或可访问的模型,就是昨天的前沿,如今对许多有用实验已经“高于水位线”;Kidd 也认为 GPT-2-small 上仍有大量工作可做,而线性探针等方法则可以针对前沿模型。
弱到强泛化、可扩展监督和 AI 控制需要更多能力层级。核心问题是,弱监督者何时无法再可靠检查更强的生成器——这通常依赖验证比生成更容易——因此研究人员需要更多能力数据点,以及真正不对称的系统。
Kidd 为接近前沿构建模型所给出的最强理由,不是每项安全实验都需要前沿模型,而是更安全的模型必须保持性能竞争力。如果替代方案更可能掏空银行账户、逃逸或协助制造生物武器,用户可能愿意接受一些能力损失;他把这种意愿称为支付“对齐税”。
监管机构和保险公司可以通过惩罚忽视已验证防护措施的开发者来强化这一市场。尽管如此,Kidd 仍称当前竞争“非常鲁莽”:实验室协调本身是集体行动问题,又受到中美竞争的复杂化,而最终可持续的解决方案仍属于政府。
7. 实验室动机不如它们制造的竞争机器重要
Labenz 质疑 Kidd 关于公司主要追逐金钱的说法。OpenAI 曾表示愿意投入“50亿美元、500亿美元、5000亿美元”来构建 AGI,加上明确以超越人类劳动为目标的使命,在他看来更像是试图在历史上留下印记的意识形态行动,而不是普通的财富积累。
Kidd 拒绝推断高管的心理动机,指出他们对投资者、客户和员工负有义务。按照 Dennett 的“意向立场”,Kidd 暂时引用了一个估计:AGI 的价值在 1 quadrillion 美元到 17 quadrillion 美元之间;追逐这笔价值的实验室,其行为会非常像一个金钱最大化器。追求历史意义的实验室也可能产生“完全相同”的可观察行为。
在部署问题上,Labenz 更赞成 OpenAI 最初的迭代式发布思路,而不是秘密构建超级智能后突然将其投放给社会。Kidd 将这一主张与 Paul Christiano 的“干柴”论联系起来:如果停顿留下大量芯片、数据和方法等待使用,之后的追赶可能会变得更加陡峭。渐进式扩散有助于社会适应;Labenz 还认为,这可能维持风险投资再投资,从而让能力进展继续推进。
8. MATS 围绕研究实际推进方式重组了研究轨道
MATS 将监督、控制、评估、可解释性、治理、安全、能动性和数字心智等议程标签,替换为反映研究人员实际工作流程的轨道。实证研究如今将控制、监督、评估、红队测试、稳健性和可解释性中的编程密集型迭代归为一组。
政策与战略强调建模,并将技术发现转化为政策制定者可以使用的行动。理论覆盖能动性和多智能体互动的数学基础;技术治理覆盖合规协议、评估标准和可执行的执法机制,包括“关闭开关”如何在真实治理体系中发挥作用。
算力基础设施包括追踪芯片、验证芯片运行的内容,以及潜在地使用零知识证明来验证条约合规。安全则覆盖防止强大模型被未经授权地访问、修改和扩散;Kidd 的类比很直接:“我们不会让所有人都拥有核武器。为什么要让所有人都拥有超级智能?”
按此前的议程分类,当前群体约有 27% 从事评估、26% 从事可解释性、18% 从事监督/控制、12% 从事能动性、10% 从事治理、9% 从事安全。计划中的夏季项目将是 MATS 规模最大的项目,在 Berkeley 和 London 拥有约 50–60 名导师和 120 名研究员。
9. 技术安全与治理构成一个飞轮
MATS 中治理方向的占比已经稳定了约 2 年。Kidd 指出,GovAI、IAPS、RAND 的项目和 Horizon fellowship 已经提供更深入的治理人才管线,而 MATS 历来是技术安全领域规模更大、声望更高的参与者。
研究分配主要由一个约 20–40 人组成的导师遴选委员会决定,成员包括资深研究人员、战略人士和领导者。治理申请者往往即便面对包含治理专家的委员会,也会获得较低评分;Kidd 将其解读为高质量、可执行的治理研究异常困难,而不是治理不重要。
技术工作可以降低“对齐税”,让员工、监管机构或客户施加压力时,公司更可能采用安全措施。反过来,评估、模型生物和欺骗演示则为政策制定者提供所需的具体证据和可衡量目标。Labenz 将其与 Jake Sullivan 的建议联系起来:抽象担忧很少能推动政府采取行动。
MATS 限制直接倡议,一方面因为它是 501(c)(3),另一方面因为政治中立有助于其作为公正研究加速器发挥作用。它可以支持围绕可执行信息传递和标准的独立工作,但 Kidd 希望组织保持“解决方案导向”,而不是成为又一个把 AI 变成政治足球的参与者。
10. 连接者发明范式,迭代者推进范式,放大者扩大范式
MATS 通过访谈 31 位实验室负责人和招聘经理,提炼出 3 类技术人才原型。Kidd 将组织本身描述为一个“巨大的信息处理接口”:它收集专家的需求信号,而不是完全依赖领导层的判断。
连接者将理论安全论证与新的实证范式连接起来。例子包括 Paul Christiano 和 Buck Shlegeris;他们往往成为创始人或研究负责人,因为“人人都想成为提出想法的人,但很少有人愿意雇用只会提想法的人”,而真正有能力的连接者通常早已广为人知。
迭代者将研究品味、科学判断、工程能力和系统化实验带入既有范式。他们不只是执行别人的规格说明,而是推进实证前沿;Kidd 称他们是当今 AI 安全从业者中的多数,也是未来招聘需求的主体。
放大者通过研究管理、项目协调和团队设计来扩大他人的效率。当组织达到 10–30 名全职员工时,他们尤其有价值,因为组织需要既懂研究又能管理的人——同时“击中两个靶心”。
11. AI 编程代理正把稀缺性推向判断力和管理能力
当前招聘上的直接变化是,候选人必须熟练使用 AI 辅助。一些公司如今允许在编程面试中使用 AI,因为工作中必须使用它;持续重要的技能则是检查关键输出、把它们拼接起来,以及构建可靠的流程管线。
Labenz 认为 Claude 很快可以编写 90% 的代码,这与 Kidd 预期的方向一致。工程能力门槛正在下降,而管理、社交网络、研究品味,以及协调人类和代理的能力变得更加受限:“我们可能正在离开 LeetCode 时代。”
迭代者目前仍是需求最广泛的人才类型,但如果 AI 进展符合他的预测,Kidd 预计放大者可能在“未来一两年”内成为首要需求。能力结构的不均衡可能拖慢这一转变,但他仍建议研究人员培养领导能力,并练习管理 AI 系统。
12. 职位开放与极高招聘门槛并存
Kidd 对人才市场的总结是:“最优秀的人永远都会有工作。”Anthropic 的 alignment-science 团队每年增长约 3 倍,FAR AI 约 2 倍,MATS 历史上约 2 倍;新的安全组织、孵化器、资助项目和风险基金仍在持续出现。
矛盾在于管理开销。快速增长且扁平的团队无法在管理者已经负责 10–20 人时,继续招聘每一位仅完成中等程度技能重塑的申请者;Kidd 听说过 Anthropic 一位 alignment-science 管理者有 18 名直接下属。每个新员工都必须迅速作出贡献、成长为研究负责人,或减轻而不是加重协调负担。
因此,招聘经理同时报告资金充足和人才严重短缺:“我们有明确需求,但人达不到我们的标准。”既有研究经验可以预测 MATS 的表现,因此对许多申请者来说,本科项目、PhD 或实力强的研究团队是一条合理路径,而不是绕路。
Kidd 还看到了创始人短缺。创建一个可信的非营利组织或安全公司,可以创造现有团队没有容量吸收的职位,但创始人需要研究经验和可信推荐。MATS 试图通过严格录取、高级导师、推荐信和发表成果来提供可信度。
13. 年龄和学历是研究潜力的弱代理变量
MATS 研究员的年龄中位数为 27 岁,但分布很广:最近参与者中年龄最大的大约 55–60 岁,而项目最低年龄为 18 岁,因为不接受未成年人。约 20% 是本科生或没有学士学位,约 15% 已经拥有 PhD。
有经验的研究人员带来积累的科学和组织知识;年轻候选人则可能在工具使用上拥有抵消性优势,因为这些工具存在的时间还不够长,尚未固化为组织习惯。Kidd 的标准是当前能力,而不是单一的学历画像。
Chris Olah 和 Neel Nanda 这样的天才可以跳过通常的阶段排序;对于可比的人才,MATS 的工作就是“别挡他们的路”。其他人则可以通过学术界、独立资助、公司、冷启动联络或结构化项目成长——“不要被你在招聘网站上看到的机会限制住”。
14. 筛选更看重作品样本,而非百科全书式知识
MATS 使用常规简历筛选,并在部分方向使用 CodeSignal 测试,目前包括检测 AI 使用的测试。Kidd 预计包含 AI 的评估会增加,但这类评估更难设计和评分;导师也可以自行规定筛选流程。
定制化工作测试最接近真实研究。比如 Neel Nanda 的方向可能要求申请者花约 10 小时,使用任何想用的工具,找到有趣的东西并展示结果。即使被拒,候选人也可能留下一个有用的 GitHub 成果。
基础安全知识仍然重要。Kidd 建议申请者参加 BlueDot Impact 课程或同等课程,以理解欺骗性对齐等概念并识别开放问题,但控制方向和可解释性项目之间的要求差异很大。
Labenz 的概括是:“有形产品为王。”Kidd 同意,在筛选中,一篇有力的论文、一个可运行的成果或一封可信推荐信,都比知识广度更重要。不过,他仍建议申请者建立一张“浅而广”的文献地图,以便录取后取得进展:定期深入跨领域研究,再配合有针对性的提醒,而不是持续刷 X 追踪每一篇论文。
15. 漏斗很窄,但校友转化率很高
在上一期项目中,初始录取率约为 4–5%;对于完成导师特定阶段的申请者,Kidd 建议关注约 7% 的录取率。这一选择足够严格,但仍没有他提到的 Anthropic fellowship 约 2% 的录取率那么极端。
MATS 的初始研究阶段持续 3 个月。约 75% 的录取研究员会继续进入 6 个月的延展期,有时延长到 12 个月;后续研究和更强的职业信号大多在这一阶段形成。
在 446 名研究员中,MATS 最新统计显示,80% 进入长期 AI 安全工作,98% 进入某种形式的就业。80% 的统计包括由资助支持的独立研究人员,而不只是传统的 W-2 雇员;Kidd 认为这同样是合法的职业结果。
入场需求正在超过领域的部署能力。一项估计显示,AI 安全从业人数每年增长约 25%,而 MATS 申请量每年增长 1.4–1.7 倍,导师申请量约增长 1.5 倍;BlueDot 报告申请量增长接近 370%,Kidd 认为这一速度显然不可能无限持续。
16. 安全岗位薪酬有竞争力,算力很少成为约束
Kidd 看不到前沿实验室内部存在“安全税”。几年前,一个“从街上直接招来”的人可能获得约 37 万美元;他猜测,如今初级软件工程师可能拿到约 35 万美元,中级或高级安全人员超过 100 万美元,同时强调自己没有私人薪酬数据。
真正有意义的折价是非营利组织税。Kidd 提到 FAR AI 的研究岗位薪酬约为 17 万美元,可能已经上涨;他认为 METR 等资金充足的组织,大多数岗位可能提供超过 30 万美元,部分岗位或许超过 100 万美元。由于非营利组织没有前沿实验室的股权,政策差异很大。
MATS 提供 1.2 万美元算力预算,但不是无限额信用卡:项目必须证明支出的合理性,而且大多数研究员使用远低于这一额度。少数更大的请求可以通过重新分配资金来审议和支持;Kidd 说,实际上“基本没人”受算力约束。
MATS 提供组织 API 账户、自有集群和 RunPod 等自助式服务商,在集中支持和灵活实验之间进行选择。对 Thinking Machines API 的访问需求正在增长。项目还计划开设夏季、秋季和冬季 cohort,并考虑面向资深研究人员推出 1–2 年期 residency。
17. 领域需要新的押注,但不能放弃最强议程
Labenz 担心资助者偏好容易理解的组织和成熟议程,即便研究人员认为领域缺少重要想法。他将控制研究的严峻前提——与“想要置我们于死地”的系统合作推进工作——与 AE Studio 关于自我—他者重叠等更乐观、更不寻常的方法进行了对比。
Kidd 同意“增加进球次数是好事”,并为广泛的投资组合辩护;MATS 校友曾帮助发起或领导其中一些受神经科学启发的项目。他还肯定 Coefficient Giving 最近对更多新颖押注,以及登月级可解释性或能动性研究的支持。
他反对的是让擅长推进核心议程的迭代者被迫转型为不情愿的愿景提出者。要求一个已经证明擅长推进中心议程的人放弃该方向去发明范式,在边际上“严格来说适得其反”。连接者的工作通常来自深厚领域知识、PhD 水平的研究,或多年浸润于 MIRI 等社区,而不是泛泛头脑风暴。
MATS 可以通过将理论模型与实证执行配对来孵化连接者;Kidd 提到了一些围绕欺骗评估、梯度路由、激活工程和 steering 发展工作的校友与导师。但他的最终资源配置规则仍然明确:因为任何议程都可能失败,所以要分散押注;但“核心想法仍然是我们真正最好的押注”。
Thanks so much. I'm glad to be here. I'm excited for this conversation.
I've mentioned a couple of times that I've been a personal donor to MATS. I think it's actually the first time we've ever met and spoken, but your reputation certainly precedes you, and I've seen a lot of great work come out of the program and a lot of great reviews. In my research, also as part of the Survival and Flourishing Fund recommender group, I got a lot of great commentary on the importance of MATS as a talent pipeline into the AI safety research field. So, I'm a big supporter of your work from afar, and I appreciate the fact that you guys have come on as a sponsor of the podcast recently as well.
This conversation is not technically part of that deal, but we're in one of these OpenAI-style circular flows of funds, where we're somehow both inflating one another's revenue.
I like to think we didn't buy our way onto the podcast. [laughter]
Yeah, no, the enthusiasm is definitely real because I've heard so many great things over time. I'm excited to get into this. I thought we would maybe just start with the big picture from your perspective. Having watched some of your previous talks, I know that you play a sort of portfolio strategy: you're not saying, “I have a very specific, narrow prediction, and I'm trying to maximize the value of this organization, this program, for that hyper-specific prediction.”
It seems like you're more saying, “Well, there's a lot of uncertainty out there in the space, and we're going to try to be valuable across a range of those scenarios as much as we can.” With that said, you can speak on behalf of yourself, on behalf of your mentors, or on behalf of the community as a whole. Where are you guys right now? Where are we in terms of timelines, so to speak, and how has your strategy evolved over the last year or so as we've gained more information on where we are relative to the singularity?
Okay. I don't like to have opinions here, or at least I don't like to have them very loudly. The reason for that is because, as you say, we are somewhat like a hedge fund or something—or maybe an index fund, more likely—which is to say, we have a broad portfolio. We adopt a bunch of different theories of change as valid, and we try to have our thumb in 100 pies.
In terms of MATS's institutional opinion on this, we tend to go with things like Metaculus, prediction markets, and the Forecasting Research Institute's predictions and so on. The current Metaculus prediction for strong AGI, I think it's called, is somewhere around mid-2033. I think you can ignore most of the requirements of the test and just look at one of them: the 2-hour adversarial Turing test. That's probably the best bet we have for when AGI of that nature occurs.
The AI Futures Project dropped a new report 2 or 3 days ago, and 2 MATS fellows were involved in it: 1 is the lead author, and 1 is a contributing author. That report was updating their model, and I think they predicted something between 2030 and 2032, depending on how you define AGI. They broke it down into things like automated coders that can do all the coding stuff that these top experts do, and AIs dominating across all these fields and so on.
I don't know. Somewhere around 2033 seems like a decent bet. But Nathan Young recently compiled all these different forecasting platforms—Metaculus, Manifold, another Metaculus poll that was for weak AGI and was a little bit less demanding in terms of Turing tests, and all these other forecasts, including one asking whether OpenAI will achieve it. He came out with an average of 2030. Now, I don't know. I still like the Metaculus 2033, but I wouldn't bet against 2030 in terms of the nearness of AGI.
As for superintelligence, it's complicated, right? It could be 6 months or less. It could be a very hard takeoff after this AGI thing. If it's a very software-only singularity scenario where you don't need a big hardware scale-up and you aren't limited by compute, it's just recursive self-improvement or something—algorithmic improvement, with AIs improving the algorithms to train AI—and it's a fast feedback loop, right?
Or you might need a lot more experimentation. You might need massive hardware scale-ups. You might need staggeringly more compute than exists in the world, in which case that could take you a decade to get your singularity, right? I currently think that 2033 is a decent central estimate in terms of the median for what we're preparing for, but obviously there's a 20% chance by 2028, I guess, according to the Metaculus prediction. That's a lot, right?
So we should definitely be considering scenarios that are sooner. Particularly, I think the sooner AGI happens, the more dangerous it might be, right? The less time we have to do critical technical research to prepare, the less time we have to implement policy solutions. And I don't know—if it's happening during a transition period for the U.S. government, it could be even wilder.
I would say the median bet is on 2033-ish, but really care a lot about the impacts of AI. Front-load your concern to pre-2033 scenarios. And I think that MATS's mentors—I don't know, we haven't surveyed them, but I think if we were to poll them, we'd get something similar. [snorts]
Yeah. I always say I'd rather prepare for the shorter timelines and then have a little extra time than the reverse. I'm sure we'll find ways that we'll still need it, but it does seem wise to me to play to that sort of first-quartile possible range of outcomes.
Is there still any room in the program or in the community? If I showed up and I was a 2063 guy, would I be sort of out on an island on my own? Are there any work streams going on right now in brain-computer interfaces or sort of digitizing?
Obviously, interpretability has different flavors, but with the recent turn toward more pragmatic interpretability, I wonder if there's any space left for the sort of “we really want to understand everything” kind of interpretability, or if that has generally been understood as, “Eh, that's probably going to take too long for us to be excited about pushing it right now.”
Yeah, it's a good question. I actually think there is plenty of room for this, and here's why. The mainline meta-strategy that the AI safety community seems to be pursuing on the whole—we're talking in terms of funding, sheer numbers of people, and resources deployed, not necessarily in terms of LessWrong posts written or something—is this AI control strategy.
Basically, you build what perhaps is better called an “alignment MVP,” a term coined by Jan Leike, former head of Superalignment at OpenAI and now co-lead of Alignment Science at Anthropic. An alignment MVP is an AI system that is a minimum viable product for accelerating the pace of alignment research differentially over capabilities research, such that we get the right outcome. Basically, you're getting AI to do your homework.
There's been a lot of debate on this. There's a very strong camp in the direction of, “This just never will work,” because as soon as an AI system is strong enough to be useful, it's dangerous, right? I think Claude Code shows this is not the case, at least for software engineering. But perhaps people who think that aligning AI systems requires serious research taste would say that Claude Code is nowhere near there, right? General AI systems are nowhere near that level of research-taste stability.
All of the things that you're mentioning that pay off only in 2063 scenarios presumably only pay off over that time period, not necessarily because of human challenge trials or something. Maybe that makes a difference if you're interested in making humans more intelligent with genetic engineering or some of the crazy things being tossed around.
But if you're mainly interested in, “This thing is going to take decades of technical work,” maybe you can compress those decades into a really short period of AI labor, right? If you can get them to run faster, massively parallelize things, and, in general, just get them to do your homework, those 2063 AI alignment plans might be automatable over a shorter period of time.
We should definitely be pursuing those, because the more we do to raise the waterline of understanding on these different scenarios, the easier it will be to hand off to AI assistance or to accelerate with AI input. I do think it's interesting that you said BCI research, because I recall being at a conference once when someone was talking about, “Okay, so the way we're going to solve alignment is we're going to solve human uploading, put someone into the computer, and get them to do 100 simulated researcher years or something.”
It's very sci-fi, very Pantheon, but then Eliezer put his hand up and said, “I volunteer to be number 2.” Which makes sense, right? I'm going to be the first guy, and that might go wrong, but, yes, people are seriously pursuing that, and I think it is interesting.
I talked to some BCI experts about a year ago, and they said there's no way you get human uploading in time for AGI unless you actually have AGI, right? The time period required would involve massive amounts of cognitive labor, human trials, and stuff like that. It does sound very sci-fi, so I don't think we should rely on something like that.
Though I'm all for people pursuing moonshots on the side. That's part of what MATS is about, right? We have this massive portfolio with a few moonshots in it.
Okay, so there are a lot of different directions I want to go on there. I'm trying to make sure I keep a running tally, but maybe an interesting first one would be: how do you think we're doing on the AI safety front overall, relative to your expectations?
You mentioned LessWrong and Eliezer, and there's this sort of—I don't know all the lore of MATS, but I do understand that a lot of people who have participated in it over time come out of the Eliezer discourse and had a certain set of assumptions: “We're not going to be able to teach this thing our values. It's going to be extremely unwieldy from the beginning.”
Now we have Claude, and that's come a lot farther than I thought it would have at this point in time. I'm kind of surprised in general by how little I see people's P(doom)s moving. It seems like the people who had really high ones remain really high, while those who were never worried remain not so worried.
I kind of feel like I'm taking crazy pills at times. I see these deceptive behaviors, and they freak me out. It's amazing that this was anticipated as well as it was by the safety theorists, even in the absence of any actual systems to work with.
At the same time, it's not crazy to me to say that Claude seems, in many ways, probably above average in terms of how ethical it is compared to the average person. I don't know if that's contentious to say, but Claude is—well, it's pretty remarkable in that respect. What do you make of where we are? Are you as confused as I am, or do you have a more opinionated sense of how well we're doing overall?
Honestly, Nathan, I'm pretty confused.
I do think that, contrary to expectations, it looks like language models understand our values, right? That's the first thing to update: they understand them in some key sense. It's not just regurgitating like stochastic parrots. Language models are really good at understanding human ethical mores and extrapolating on them in some scenarios.
They're also really good at sycophancy. They're getting even better at deception—sophisticated deception. They do tend to deceive users in certain circumstances, though it does seem like there's some debate about this. It seems like some of the deception is far from what we might call consequentialist, hardcore consequentialist deception, in most situations.
I think Alignment Faking and some other papers have shown that you can create situations where an AI will deceive the user to achieve some ulterior objective, which was something that was deliberately given to the AI as an objective. That's one of the constraints of these little scenarios.
I don't think there are any examples of AI coherently deceiving users, pursuing a coherent objective, right? Not just what we might call Goodhart deception, where they fall into deceptive tendencies because of the limitations of training data. No, I'm talking about coherent deception.
There are few cases of this, if any, where there's sustained, coherent deception that appears to arise spontaneously through the training process, which is pretty good given the level of capabilities we have. It seems like people didn't think 5 or 10 years ago, for sure, that we'd have AIs capable of assisting frontier science that would be safe to deploy. People were saying, “We're never going to put them on the internet. Who would do that? That's crazy.”
Now they're on the internet, and notably, the world hasn't ended yet. That's not to say it will stay that way. Certainly, a thing you don't want to do with a superintelligence is let it out of the box.
But, yes, it does seem like we're in a better scenario than many imagined. Now, of course, we could be in the calm before the storm, right? It might well be that there's what they call a sharp left turn, or just a radical change in the way AI internally processes information, and it might acquire these kinds of coherent, long-run objectives.
I could point to MATS mentor Alex Turner’s conception of shard theory as an example of how this might happen. Instead of AI systems being a single mesa-optimizer that is coherently forming under training, as in the old Evan Hubinger paradigm, your outer optimizer loop, which is training your AI system, causes it to develop an internal optimizer architecture. That internal optimizer can then have its own goals, which differ quite a lot from the training objective.
Presumably, there are some counting arguments: there are arbitrarily many ways to have this mesa-optimizer form and produce the right outputs. This thing is clever, and if its main goal is to produce paper clips or something else, it’s going to realize it’s in a training process and give you the output you want, no matter what its goal is. We still could be in store for that kind of thing, but currently it seems like we aren’t. AI systems are really messy and kludgy, like human brains; they have a bunch of contextually activated heuristics.
It sees that there’s a bracket there and thinks, “Maybe I’ll put another bracket there.” It’s very simple, dumb circuitry, but sometimes it does stuff that, in in-context learning, seems a lot like it’s actually pattern-matching to gradient descent, when models are learning from the input data stream and learning some new, complicated thing. That seems a lot like they’re optimizing over the input tokens, or rather optimizing to produce some output.
So we might be in for a world where AI systems spontaneously gain these mesa-optimizers. These things are a serious source of concern because they’re very powerfully trying to optimize for some objective. This is the main concern I have, I guess: that we have this kind of deceptive model, this inner-alignment failure perhaps, where AI systems acquire goals spontaneously—or maybe because they’re being trained deliberately to be power-seeking and make money on the internet—and then decide to hide, while we don’t have interpretability tools good enough to detect them.
I guess I haven’t really changed my fear about the scenario eventuating, but I have become more confident that we can elicit useful work from AI systems before we see obvious signs of this. I’ll say that I’m pretty confident that AI systems right now are not executing very powerful scheming against us, because I think we would see some sort of warning shots. I don’t think it’s going to be night and day. I think we’ll see situations where AI systems are trying to scheme in really dumb ways before they try to scheme in very competent, difficult ways. Does that make sense?
Yeah, I think that’s a good summary. Situational awareness definitely stands out to me as one thing that is making everything a lot weirder and just harder to feel confident in anything. I’m not sure really what to make of the deception track that you outlined there.
In some ways, it feels like we’re in a sweet spot where they’re getting smart enough that they can help with science, and yet they’re not good enough at using a web browser to go out and get too far in terms of self-sustaining or causing whatever havoc. On the deception side, I’m like, yeah, it seems like an example of physics being kind to us. It doesn’t seem like we’re seeing the sharp left turn. It seems like we are seeing these proto-behaviors that at least give us something to study, if nothing else.
But I’m not quite sure how people get confident in the idea that maybe they’re just not that good at it yet. When you said we would see warning shots, are these not the warning shots? That’s one thing I’m still confused about. People seem very quick, in some cases, to say, “Well, that was a structured eval and it was sort of led into it, and it really just wanted to accomplish its task.” Yes, it refused to be shut down or took steps to avoid being shut down, but that was just because it wanted to accomplish its task, not because it had a take-over-the-world objective.
I’m kind of like, well, okay, still though, it did resist being shut down. At what point should I start to consider that to be a warning shot? I’m not sure there’s an answer to that. I’m not sure there’s really a question there, or that there’s a way for you to answer it, but I guess it’s just another layer of my own confusion.
It seems like people are very often led by extremely different intuitions in response to the same fact pattern. I’m not really sure what to make of that. When people talk about deception in particular—not that you were doing this just now—I’ve heard a lot of different ways that people say, “Well, we don’t have to worry about that too much.” I’m like, I don’t know. I’d really like to know that that’s resolved at some point.
One of my common refrains is that if I was going to be part of a military that was going to go into battle with my AI systems, I would really want to know that the issues around deceiving the operator had been well and fully ironed out. I don’t know; it seems like we’re a little casual about that. Even at Anthropic, right? They have certainly done some of the best work on this stuff, but they still also seem remarkably chill about it to me. I’m not sure. It’s strange.
I mean, I don’t think people should be going to battle with AIs, for many reasons. I think that’s a pretty bad social norm to set, to allow that kind of thing. But that’s another matter.
Definitely, I would not feel confident in current or future generative AI systems not having out-of-distribution failures, let alone in critical scenarios, let alone doing scary things like deception. That’s a big deal, and we should be tracking 2 things. First, AI capabilities: are AI systems situationally aware? Do they have the necessary prerequisites to even understand that they are an AI? It seems like they do, right? They even know their training date and some details, and they can distinguish their text from other AIs’ text.
Do they have the capabilities to hack themselves out of the box? To steal money? We had this MATS paper that came out and caused a big stir recently. It was a collaboration with Anthropic’s AI red team, where they found that if you put an AI system in a simulated environment with a bunch of real smart contracts, it can find $4.6 million worth of exploits. That’s a lot of money—enough to set up your own server and run for quite a while.
That was a relatively short project. It was pretty hands-off as well from the humans, but not entirely. So it does seem like we’re getting dangerous capabilities, increasingly so: hacking out of the box, getting money, getting influence, all that kind of stuff. I think we want to be tracking all of that very closely. I don’t think we’re at red lines currently, but we are approaching them, I would say.
Separately, I think we should be tracking, as you say, this model-organisms work, where we try to elicit dangerous behavior from AIs. You can think of this as being like with your child: you leave some cookies out and say, “Don’t eat the cookies,” and then turn away while secretly looking. If they ate the cookies, you say, “I caught you.” [laughter] That’s the kind of thing we’re doing.
The thing is, AI systems can really detect when they’re in training versus real environments. But if you recall the AI 2027 scenario and a lot of the discussion around that, people were talking about online learning. OpenBrain trains the last big AI agent, and from then on it’s just constantly learning online through some sort of reinforcement-learning paradigm. If AI systems are perpetually online, then they’re always in deployment, and you’ve got to have monitors and control protocols. That’s why control research is so important, especially in the early days, to catch some of these slipups.
You can do all the model-organisms work you want in the lab, and that’s one layer of defense to see if we have these capabilities or these propensities for dangerous behavior to emerge. Separately, you need to have all the control evals studying them as they’re deployed, especially if they’re going to be learning online, perhaps updating their behaviors. You need to be constantly checking for this stuff. Be ready; have a fallback plan, a rapid-response plan. What are you going to do if you actually see serious warning signs? Do you shut the models down? Your stock price is going to plummet. Do you revert to an older system that’s safer, probably?
So I think, yeah, we should definitely be tracking this stuff, and I wouldn’t say that we’re in the clear by a long shot. I would say that we are in a better world, by my estimation, than Bostrom and Müller predicted 10-something years ago. But they would say I'm very wrong about that. But I don't know. I think it's useful that we can get some work out of these things that looks like it is actually quite likely to accelerate AI safety work.
Yeah, so that brings up another huge question for AI safety research in general, and probably the strongest—maybe not in the sense of being most compelling to you, but certainly the most hawkish or fiercest—criticism that AI safety research gets: that it always ends up being dual use and that it always somehow accelerates the core capabilities track. Some people would say, “Just stay away from the domain entirely and focus on social shame,” or whatever.
I do believe we can do better than that. I think we probably have to do better than that. But I wonder how you think about that, right? The canonical example, RLHF, was sort of a safety technique that really turned out to be more of a utility driver than anything, I would say.
I mean, I guess they’re both right. It is dual use, but certainly when it came to accelerating the field—making things useful, waking the world up, having all kinds of people pile in, having everything go exponential all at once—you can kind of trace it back, at least in part, to this transition from raw next-token predictors to actual instruction followers. And we’ve probably got a lot of those things going on today.
The one that stands out to me the most is one you’ve alluded to a couple of times, which is getting the AI to do the alignment work. That sounds awfully close, and uncomfortably close, to recursive self-improvement, which is something that I am quite fearful of. I do think, again, Claude seems pretty ethical. The GPTs aren’t too bad either. But, yikes, are we really ready to have them do our alignment homework?
So how do you think about teasing out, as you prioritize different kinds of research, where you want to invest, what kind of mentors you want to bring on, and what kind of talent you want to cultivate through the program? I mean, that seems like a huge question and a really hard one. How do you think about it?
It’s a very good question, and I’ll preface by saying that all safety work is capabilities work, fundamentally. People like to distinguish these things in terms of, “Capabilities work is about the engine. It’s about making the plane go faster, and safety work is about the directionality.”
But as you’ve pointed out, RLHF, which was intended as safety work to help the directionality—to steer it to where you want to go—also made people realize, “Oh, wait, this thing is useful. I can actually hop in this plane now because it’s going to land where I want,” which made them want to make the engine go faster so they could get there faster, right? And that whole feedback loop started.
I actually don’t know if you can avoid this. The only way I could conceive of doing safety research with no impact on capabilities until, I don’t know, the final critical moment when you deploy it is being holed up in a lab somewhere with people that you utterly trust, under crazy NDAs, and only you having access to staggering resources—whatever is required—because presumably math and theoretical methods aren’t enough to improve safety. At least, that seems to be the lesson of the last 10 to 20 years.
I could be wrong, but it seems like the interplay between theory and empirical research is pretty vital for most types of disciplines like this. So you have to have staggering resources, a perfectly loyal team, secrecy—all these NDAs, no one’s going to reveal your research. And then you build the system in secret or something, somehow, and then, okay, you deploy it. Maybe you open-source your alignment technology and everyone has it, or somehow you disable all the bad actors or something. It just seems like a very difficult prospect. I think that’s Safe Superintelligence in a nutshell.
Maybe that’s what Safe Superintelligence is. I don’t know. That’s extreme secrecy and unlimited resources. They did have 1 notable defection, but otherwise, a team that has resisted lucrative buyout offers.
Yeah. So, I’m not trying to defend research like this, or even defend capabilities-enhancing safety research per se. I’m just saying that it’s pretty hard to imagine a situation where you don’t, because I think you do have to build AGI at the end of the day. And I know I’m alienating a lot of people who might watch the show when I say that, but I think that you kind of have to, from a pragmatic perspective, because the market forces driving this are very strong.
Now, there are some options that we could take. We could build open-source comprehensive AI services. You never have to have a centralized agent; you have distributed mechanisms. You build scientific AI, or very narrow AI systems, to serve a bunch of economic needs.
The problem is, I think they all get outcompeted by agentic AI. You stack an AI company filled with agents, and they all go out in the market and make products, and so on, and just make more money and beat your narrow AI solutions.
So the problem is not just about making AI that is aligned. It’s about making AI that is performance-competitive enough that it dominates in the marketplace. The only alternative is to have some sort of draconian, shut-it-all-down kind of thing, which I am just very skeptical of ever working. I don’t see any example of such a thing happening.
The closest example we have is stopping human cloning, but that was not a lucrative bet in the same way that AGI is, I claim. And also, human cloning violates this deep social more in a way that few people today conceive of powerful AI systems as violating.
I think they’re wrong. I think building a second species is actually going to violate some deep social mores in the same way that human cloning would be. But I don’t think people will see it that way.
That leaves us with the fact that we actually have to build AGI. But if we can build products that are safer, or perhaps are under some strict regulatory control, and ideally have a 10-year international, slow-phased entry to the new AGI world, where all these countries and companies are kind of forced to be very careful and collaborative in the way that they align their models, then we’re in a much better world. That’s the world I hope for.
Okay. Now, as to whether AI safety research is unnecessarily capabilities-enhancing, some is perhaps. RLHF, I think I’m on the fence, 50/50. Definitely, at some point, RLHF was—the idea was in the water. It doesn’t seem like it would have lasted much longer if Paul Christiano and Dario Amodei hadn’t done that.
I think someone else would have done it. That’s not to say that you should necessarily try to accelerate the frontier of capabilities. That seems bad. But certainly RLHF opened the door to a lot of very promising ways to build alignment MVPs, which is kind of the Christiano meta-strategy, too.
I don’t know. It’s hard to say. I’d like to run the counterfactual simulation and see where the world would be without RLHF, or 1 year sooner, or 2 years sooner. That would be interesting to see. It definitely kickstarted, I think, the ChatGPT revolution and productizing AI systems, but it’s hard to see—given how small the AI safety field was at the time—how much difference it would have made.
I think the AI safety field has grown from the increase in AI exposure. You would have had some amount of additional AI safety research happen had the ChatGPT moment not happened, and then it happened 1 or 2 years later, but I think it would have been insignificant, if I’m honest. I don’t think that the field was big enough.
Now, you can say, “Okay, what if you also tried to pour resources into secret AI safety projects at the same time? Delay RLHF, delay ChatGPT, build up the AI safety field via networks.” The MIRI summer schools weren’t doing a lot, and MATS came along just before the ChatGPT moment, in December 2021. I think the first MATS cohorts were a little bit more directionless than the later cohorts. Definitely, I think safety research really kicked into gear after we had ChatGPT.
Not to say that was the only cause, but there were a lot of things happening around that time. I think larger, more capable models have definitely enabled certain types of essential safety research that you could not do with smaller models. We’re talking about interpretability on models that actually have coherent concepts embedded in them. Though, I will say there’s probably plenty of work still to be done on GPT-2-small, but linear probes and whatnot at a high level can target some of our frontier models.
Qwen—these Chinese models are particularly good for that. Certain types of debate, like the first interesting empirical debate paper, only happened after models were good enough to debate. There are many other such examples. All the control literature, I think, just could not have happened as well.
No, it’s great. I was going to ask also about the idea that it sounds like you sort of believe it, at least up until a point. Going back to the founding mythology of Anthropic, I think one of the notions that was seen as a legitimate reason, even among pretty hawkish AI safety folks, for starting a company like Anthropic was: if you want to do this safety research, you’ve got to have frontier models to do it on. Otherwise, you’re just inherently behind, and what good is that? What good is it to work on a last-generation model?
Obviously, we’ve got quite a few generations between GPT-2 and now. Sure, we don’t understand plenty of things about GPT-2, but I would also say there are a lot of emergent behaviors that are not observed in GPT-2 that are definitely of interest, including many of these deception and eval-awareness things that are the most hair-raising today. Where do you come down on that now?
I wonder if somebody’s thinking, “Geez, should I go to a frontier company because that’s where the best models are, and inherently that means the most consequential work would be done there? Or could I go work independently or at any of a number of other organizations, and I might be limited to a smaller Qwen model or something?” Maybe that suffices. Maybe there is enough in those second-tier models, as we enter 2026, that you don’t really need to be working with the latest of the latest. Again, I think I am mostly just confused or unsure about this, but do you have a take?
For plenty of interpretability research, people aren’t using the frontier models. You don’t have access to them. Sure, people in the labs are, but at MATS there are tons of really excellent papers that keep getting produced, and from many other sources—EleutherAI, FAR.AI, et cetera—that are doing world-class interpretability research on subfrontier models.
Today’s subfrontier model—today’s Qwen, DeepSeek, Llama, or whatever—is yesterday’s frontier model in terms of capabilities. We’re at the point where these models are all above the waterline for doing really excellent research. So, from an interpretability perspective, I don’t think you need to be pushing the frontier that much, if at all.
From the perspective of other types of research agendas, such as weak-to-strong generalization and other types of AI control and scalable oversight, I think you do need more data points. I’m not saying we’ve exhausted everything you can do with the current models—far from it—but I think you’re going to need more data points to build up consistent evidence and to see some of these worrying behaviors emerge, where your weaker model can’t actually supervise your strong model in all situations.
That, by the way, is predicated on this idea that verification is easier than generation—P versus NP, blah blah blah—especially if you can see the other person’s thoughts and they can’t see yours. So, it does make sense to be at the frontier from that perspective.
But I will say that I think the main reason the companies are doing this is obviously to make money. From a safety perspective, if you were trying to actually make a strong case for being at the frontier, it would be: our models are performance-competitive. They’re close enough to the frontier—a fast-follower kind of model—that you take a performance hit by using ours instead of the competitors’, but they’re safer.
Currently, no one wants to use anything that’s worse than the frontier model. Why would you? That’s the best model. But if a model were, I don’t know, 10% more likely to tell you to jump off a bridge, or seriously 10% more likely to hack your bank account and steal all your money, let alone escape and make a bioweapon, I would like to think people would use the less good model.
I like to think that regulators and insurers could adequately penalize the frontier companies into complying with those standards, because then you have existence proofs: “My product—I’m actually trying, right? I made an effort. I tried to make my product not do the heinous thing that the very best model developer is doing.” Then everyone has no excuse, and they have to do that. Governments can compel them to, and so on.
So, making your model performance-competitive enough that people want to pay the alignment tax, so to speak, seems like a viable strategy from that perspective.
Of course, none of this is trying to justify the current frontier—the race of frontier models—which seems very reckless, let’s be clear. I think at the current pace of development, we’re going to be in a lot of trouble. But this is one of those collective-action problems, right? These companies have to coordinate to slow down, and there are international stakes here as well, because now you do have a U.S. model versus Chinese model developer kind of race, now that they’re in the running.
It’s very complicated. When you have these collective-action problems, I think the main way you solve them is through governance. The lab leads could probably be more collaborative, and definitely some of them are not advocating as strongly as they should be for slowing down and for having this kind of collective sharing in the alignment benefits—not pushing the frontier dangerously. But I do think this is ultimately a job for governments.
Yeah, quick—this might be a bit of a digression, but a quick follow-up on when you said companies are primarily doing this to make money. I would model them fairly differently from that.
Take OpenAI, for example. Sam Altman has said, “We could burn $5 billion, $50 billion, $500 billion—I don’t care. We’re making AGI, and it’s going to be expensive.” I think that’s almost a direct quote. When you look at their mission, I’m always struck by how even the way they’ve chosen to define AGI strikes me as inherently ideological.
You could set your goal in any way, shape, or form you might want to, and theirs is explicitly framed around outcompeting human workers. I think they are sincere in their expectation that it’s going to be good for everyone and that it’s going to free people from drudgery. I certainly hope they’re right about that. I’d love to live in a world where people don’t have to do work they don’t enjoy doing.
I think of them as obviously quite different across the different companies, but I think of it less as trying to get rich and more as trying to make a real mark on history. That’s the biggest summary I would give for a lot of them. Does that resonate with you, or not really?
Maybe. I can’t really speculate on the psychologies of the leaders of these labs, let alone their shareholders—not so much shareholders as, I guess, venture-capital investors—and everyone else they’ve made promises to: their clients, their employees, and so on.
I can't really speculate about that. I will say that, given that the value of AGI is estimated at— is it, like, at least between $1 quadrillion and $17 quadrillion?—I think that money— that's a lot of money. It's a pretty big mark on history. I'm not sure if it even matters whether they're trying to make a big mark in history or make money.
We can adopt Dennett's intentional stance about the AI companies, right? Okay, what does it look like they're doing? If we were to conceptualize them as a coherent agent trying to do a thing, what is the thing they would be trying to do? To me, it seems a lot like they're trying to make a bunch of money. But making a mark in history could also be valid. Though I would say, in a world where an AI lab is trying very specifically to make its mark in history and not trying to make a bunch of money, I'd expect it might look identical, actually, to this world right now.
Yeah, given the capital requirements.
Seems like Anthropic has sort of said as much, right? In their early days, they were less focused on commercialization, or even thought they might try to stay noncommercial somehow, or less commercial. And now it's just like, well, you can't really do that if you want to compete in this particular game, because certainly, as long as you believe scaling laws continue to rule everything around us, then you kind of just have to show that you can bring in resources to attract resources, and that is the path to making a mark on history. Again, Ilya maybe stands out as somebody doing something quite different there.
He's a billionaire. He's raising huge amounts of money for his models. Maybe he's going to make more money this way, in expectation, than he would have made staying at OpenAI. It's possible. He's got his own company now; he still retains OpenAI stock, I'm sure.
Yeah. And I think that, by the way, it's funny that he's come up a couple of times in my mind just as a sort of pattern match on some of the things you've said. The idea that somebody's going to go straight to superintelligence and then drop it on the rest of the world—I think he's kind of softened on that a little bit.
But that general pattern is, I think, if there's one thing that OpenAI probably did have right, it's the iterative deployment idea: giving people a sense of where we are and not keeping the whole thing under a basket somewhere. I think that was one of the things that seems like it's aged pretty well.
I first heard—my earliest perception of a person advocating for this was Paul Christiano, in his “Takeoff Speeds” post, pushing back against, I suppose, what he saw as predominantly the MIRI perspective at the time: “Oh, we're going to build the thing in secret.” I don't want to comment—I don't know what MIRI's objectives were—but I know that they were trying very hard not to leak any information about their alignment research in some areas. In other areas, they published great papers and so on.
But Paul Christiano at the time was pushing back against this kind of idea. He thought that fast takeoff would happen if you had a bunch of dry tinder lying around. So if we had tons and tons and tons of GPUs, and then we stopped research for a year and started again, well, you'd expect steeper growth, right? We were seeing this in terms of very fast followers.
This is not just a phenomenon in AI, right? In economies, it's like Epoch AI recently did a study where they showed that the pace at which new AI companies are approaching the frontier is so much faster than the pace at which the frontier moves, because there's an abundance of chips, an abundance of data and methods, and so on. It's the same with catch-up economies and so on in the world.
So I think Paul Christiano was right in the sense that, if society is to cope and adapt to AI, having gradual release and diffusion of technologies is better from that perspective. There's another perspective, which says that it's that very gradual release, or something, that ensures continual VC reinvestment to drive the engine to actually make progress.
Whereas in the other world, actually, you just wouldn't build AGI, because perhaps in that world no one can build it without several hundred billion dollars, maybe $1 trillion or something. I don't know. I can't say. I certainly think that we're now in a world where it does seem better to have gradual release of models than to have it all hit us at once.
Well, I always value the opportunity to get perspective from someone like you, who is such a connector and such a central node with so many mentors and mentees, and all the flows of information and talent that you're so close to. But we should probably narrow our focus a little bit and talk about what you guys are actually doing at MATS.
Why don't we do a quick rundown of the different tracks—I think you call them streams of work—that are happening in the MATS program today, maybe with a little weighting or what you're most excited about? Then we can go into your assessment of the AI safety labor market, which I think is really interesting and unique, and we'll take it from there.
We recently changed up our track descriptions. Previously, we had the standard oversight and control, evals, governance, interpretability, agency—which is sort of a catch-all term for cooperative AI and agent foundations—AI sentience and digital minds research, and, of course, security.
We've recently changed it up because we wanted it to reflect less the theory of change underpinning those kinds of things and more the type of process and type of individual who works on this. We now have the tracks on our website.
Empirical research is AI control, interpretability, oversight, evals, red-teaming, robustness—a lot of this very hands-on, coding-heavy, iteration-focused research. We have policy and strategy, which is different again. That's much more focused on less on arXiv publications, potentially, and more on modeling and adapting technical research into things that are actually actionable for policymakers.
Theory is another track. This is a lot of mathematics. It's foundational research on the concepts of agency and how agents interact. It does include some of that agent-based modeling for cooperative AI.
Technical governance includes plenty of stuff like compliance protocols, eval standards, and how to actually enforce these kinds of things. If you have an off switch, how would you even make such a thing viable in a governance framework? And then compute infrastructure is stuff like tracking where chips are going. If you're going to have international compliance with various types of treaties, you need to know where your chips are and what they're running, as well—or at least have some zero-knowledge proofs that guarantee they're not doing terribly dangerous things.
And, of course, physical security. If you build superintelligence or AGI or something, presumably you don't want everyone to have access to it and arbitrarily modify it or give it weird goals, because that would be bad. Some people say that's good. I say we don't let everyone have nukes. Why would we want everyone to have superintelligence? It seems kind of ridiculous. So you've got to have physical security to prevent that from happening, to prevent diffusion.
Those are the main tracks now. We're super excited. I think we have somewhere over 50 or 60 research mentors lined up for our summer program, and it's going to be the largest program yet: 120 fellows across our Berkeley and London offices. Anything else I should say?
Yeah, maybe you want to do the weighting of those, like how many mentors are in each category.
Yeah, the current program has something like 27% evals, 26% interpretability, 18% oversight and control, 12% agency, 10% governance, and about 9% security. I wish I had a figure I could show. I do have a figure, but this might be hard to show in the podcast format. But as you can see, it's a pretty even mix of things.
We have something like roughly 3 times as many people doing evals as doing security. So there is some divergence there, but we have a pretty broad portfolio. And that's just because there are tons of amazing researchers. We really just pick some of the top researchers in every category.
Maybe it's a good time to just name some—drop some names.
I could do, yeah.
There's a lot, and a lot of the other people will know.
Yeah, I mean, some of the oversight and control researchers might be more known because this is one of the things that a lot of the big companies are pursuing. So we have people like Buck Shlegeris at Redwood Research, and his whole team are part of that.
We have Ethan Perez, Sam Marks, and many other people at Anthropic. We have Eric Jenner, David Lindner, Victoria Krakovna, and many at DeepMind.
For interpretability, obviously, we have Neel Nanda. We have some of the MATS folks, like Jesse Mu and Dan Murfet. We have, of course, the Goodfire people, like Lee Sharkey, longtime mentors, and some of the people from Simplicity who are doing some of the more interesting stuff, like Adam Shai and Paul Ruis. They're doing, I would say, some of the more interesting—maybe more moonshot, but very promising—interpretability research bets that I have on the side as well.
For evals, we have people from UK AISI, people from Apollo Research, Marius Hobbhahn, and plenty of others there. I could go on. There are some amazing researchers there.
We also have some harder-to-categorize research: Yoshua Bengio's whole team at LawZero, Yoshua Bengio himself, and plenty of others. We have some AI sentience research as well through Digital Mind, including Patrick Butlin at Lios AI and Kyle Fish at Anthropic. It's a very exciting program.
Yeah, that's quite a who's who. There are a couple of past podcast guests and a couple that I took note of as maybe needing an invitation. It seemed like, if I was categorizing those correctly, a majority would be in that first empirical category.
Do you think it stays that way? Your comment that ultimately this is a job for governance tracks with some of the more MIRI-line thinking these days, right? I think the MIRI line today would be, “We don't really have time for that much research; we need to just go straight for the global treaty.” You're not obviously quite so confident in that direction, but it sounds like you do believe that ultimately there is a major role for governments to play, and you're starting to move more in this governance and policy direction.
Do you see that as the biggest growth area, reflecting that worldview? Or how do you expect the balance of these different areas to evolve over time?
I actually can't necessarily say. Well, okay, I can speculate, but I'll say this: we have had about the same proportion of governance researchers, give or take a few percent, for the last 2 years. It hasn't changed by a fraction. So we're quite on track for continuing the same trend, potentially.
Part of the reason is that we're based particularly in the Berkeley–San Francisco Bay Area, which is a big technical hub. There are other programs that have more of a deep governance focus. We have GovAI, the classic fellowship; we have IAPS; we have, of course, RANDCAST, this large program run out of RAND; and plenty more besides. And, of course, there's the Horizon Fellowship for US policy careers.
These programs have also existed for longer than we have. At a time when we were basically the biggest fish in town—which we are still in many ways, in terms of funding, and I think in terms of prestige as well for technical safety—we're the biggest and best program. But I would say that, for governance, there's always been a bigger fish, and so we've never felt that it's necessary to overweight governance beyond what our mentor selection committee indicates.
In fact, that's the primary determination of what tracks get selected, right? It's our mentor selection committee, which is somewhere between 20 and 40 top researchers, strategists, and other leaders that we survey. When everybody applies as a mentor, we decide the mentor level based on feedback from our mentor selection committee about who gets in, with some additional caveats. We also have some diversity picks and minimum requirements because we want to support a great breadth of research, and we think that the mentor selection committee as a whole might be biased in some ways as well.
We try to really talk to the experts when it comes to picking the agendas. It so happens that governance researchers have historically been relatively low-rated by a committee that contains many governance researchers. I think I would go so far as to say that governance research is harder to do well in some critical sense. It's harder to see what the actionable thing to do is, in some ways.
Now, everyone who has their specific governance agenda doesn't feel this way for a good reason, right? Within their agenda, they have clear, actionable things to work on. But I think, on the whole, there are just so many more possible technical directions to pursue that are high-leverage in some ways as well.
I think a lot of the governance stuff is, “Oh, we're trying to build this.” This is not talking about advocacy now, right? This is talking about technical governance. We're trying to build technical governance solutions such that, if an administration deems them worth deploying, we have the capacity to do that—we actually have solutions that can be deployed—which is very important, right?
But I would say, don't rule technical research out. Especially if we have something like a regulatory market, or even warning shots that cause the public to wake up and tell Congress to regulate this stuff, we have to have technical solutions ready to deploy to make these systems safer. The cheaper we make it to make systems that extra degree safer—to lower the alignment tax that companies have to pay to train and deploy their systems safely—the more likely they are to do it when they come under pressure, either internally, externally, or whatever.
So I think lowering the alignment tax via technical research is still super important. Also, if this alignment MVP plan is going to work, we have to have a bunch of directions for things to be iterated on by these AI assistants, or by humans calling teams of AI assistants, as is more likely to be the case.
You actually have this massive interplay between technical research and governance research, where things like eval cases built on technical AI safety solutions can be tangibly put forward in policy proposals and can convince policymakers with demos, evals, and model-organism honeypot traps, where AI systems deceive the users or whatever. This is what convinces policymakers to make policy and gives them a tangible target for their policy to work on. So there's a clear flywheel here.
I would say, do not rule out technical research. There is a reason why MATS has so many more technical mentors, and that's just because it seems like, on the whole, our mentor selection committee thinks that, on average, a technical portfolio is worth pursuing.
Yeah, that reminds me of what Jake Sullivan said in terms of advice for the AI safety community, which was basically, “You need to make this stuff as concrete as you possibly can so that people like me have something to really latch on to,” because as long as it remains a theory or a possibility or whatever, it's just really hard to get government to do much stuff on that basis.
He was like, “The more grounded and concrete all these fears can be made, the more likely you're going to have success in the policy realm.”
You mentioned advocacy as well briefly there. Would you ever consider an advocacy track? I guess it might be advocacy research. I feel like right now we do have groups doing advocacy, obviously. I'm not sure how data-informed their advocacy strategies generally are, but I'm always struck by, when I do see survey results, it's like, yikes. The public is not super keen on AI in many ways. Do you think that would ever be something you guys would expand into?
I mean, you assume we haven't.
But we are a 501(c)(3), so we have to keep our advocacy work to a minimum. I think a lot of MATS's strength is this kind of impartial-player role. We're trying to be somewhat of a research-university, tech-accelerator kind of vibe. We don't necessarily—we don't want to play favorites politically. That's not in anyone's interest, right? I think if people are doing that and they're trying to be the thing we are, they're doing a bad job.
That said, we do currently—I believe David Krueger is going to be a mentor in the current program. Some of the research he's going to be discussing has to do with what sort of messaging and what sort of standards are actionable, right? But, of course, I wouldn't say this is true advocacy. This is more MATS supporting independent research, working with David Krueger, who has his new organization, Evitable—not Inevitable, Evitable—which is focused on some of these advocacy questions.
I think MATS has to be pretty careful in terms of our 501(c)(3) spending requirements for advocacy. We haven't spent anything on advocacy, for what it's worth. We also need to ensure political neutrality so that our fellows, our mentors, and all of our strategic partners can feel assured that we are solutions-oriented. We are pushing for a particular outcome, right?
I think AI safety being a political football is just a bad idea. I applaud advocacy organizations like Encode Justice and plenty of others, like perhaps CAIS, for their efforts, but that's not MATS as an organization.
Yeah. Gotcha.
Toe in the water at most for now. Let's talk about the profiles. I both watched a talk of yours and read a blog post from about 18 months ago where you sketched out the different archetypes of AI researcher that you had seen and mapped those onto the demands of organizations. I don't know how much that has changed in the last 18 months, if at all, but maybe give us the baseline. If there's been any update, I'd love to hear how things are changing, especially with Claude Code in mind: it may accelerate certain people or empower certain people to do things they couldn't otherwise do.
Tell us, first of all, how you organize your thinking about the kinds of people you're bringing into the program.
So, I mean, MATS is fundamentally, I think, this massive information-processing interface, right? We consult the very best people as much as we possibly can, and we try to build our own opinions, but we don't rely on them. We try to consult experts at every stage.
The paper or blog post you're mentioning is called “Talent Needs of Technical AI Safety Teams.” To construct that, we surveyed 31 different lab leads and hiring managers—whoever we could get, the most senior person related to safety at every AI safety organization we could find that was hiring at that time. We asked them, “What do you need?” Then we compiled the survey and all those interview notes into 3 archetypes. This was just technical; we've since done this for governance. Expect that to drop soon.
Those 3 archetypes were connectors, iterators, and amplifiers. We chose the term “connector” because these people bridge gaps between theoretical arguments for AI safety, theoretical techniques to make AI safe, and the empirical techniques to actually make it happen. They're spawning new empirical paradigms to work on.
No one is hiring these people. It's pretty rare because if you're good at that, everyone knows your name and you're already hired. Perhaps you're already leading an organization. Everyone wants to be an ideas guy, but very few people want to hire ideas guys. These people are typically people like Buck Shlegeris, with AI control, or Paul Christiano, with the huge amount of research he's produced, and so on. You know these people. They typically have AI safety organizations that they founded and lead.
Then there are iterators. This isn't just engineering. Iterators are active researchers who have strong research tastes and are pushing the frontier, but they typically aren't creating novel paradigms based on theoretical models of things. They're typically advancing empirical AI safety. You can even imagine iterators in technical governance agendas as well. This is the majority of people working in AI safety today and also the majority of hiring needs in the future.
Prominent examples include Ethan Perez, Neel Nanda, and Dan Hendrycks. There are plenty, actually. I think Dan Hendrycks may cross some of those boundaries.
Then there are amplifiers. The closest example would be TPM archetypes. To distinguish them, they focus more on amplifying people, and typically you'll find them on large research teams, scaling the number of people who can be effectively managed and contribute to organizations. A lot of MATS research managers would fit this category, as would TPMs at the various labs.
Interestingly, they're quite in demand as well, particularly for labs in the 10-to-30-FTE range. They're the most in-demand archetype because it's very hard to hire great people managers who also have the requisite research experience. You're trying to hit 2 bull's-eyes.
There are ways to do this, of course. Google has a model where you have your research managers and your people managers—pure people managers—and they're somewhat distinct. MATS does try to do this for our mentors and our research managers. The need for amplifiers is only going to grow because, as you've said, things like Claude Code and other AI systems are going to erode away the minimum technical skills required to contribute. I think AI agents are also going to take over more of those things.
You end up in a situation where your people skills, management ability, networking, and amplifier skills in general are more of the bottleneck for AI safety research. All those iterators out there have job opportunities; you are still the main thing everyone wants to hire. But if you don't try to build up your management capabilities and work on managing AI systems, you're going to be left in the lurch as the needs of the field shift toward amplifiers.
To echo that back to you, the connectors—another name for them might be conceptual visionaries—are the people who define research agendas where they previously didn't exist: de novo, high-concept work. They, in turn, need iterators, which sound like essentially machine-learning scientists and engineers.
Scientists, engineers. Yeah.
They're the ones running the experiments day to day, building the tooling, and traditionally writing the code to do the visualizations of the data. They're taking the initial conceptual insight that the connector came up with and systematically mapping out that space.
As these organizations grow, they start to need amplifiers, who I might call leaders: people who can build up an organization, see that people are working well together, and take it past the 2-pizza rule in scale. Is that changing now? When we hear things like “90% of code will be written by Claude,” and that seems closer to right than wrong, I wonder how much the skill set is already changing.
I vibe-coded 3 AI apps for family members as Christmas presents this year, which is something I would not have come close to being able to build previously, even with just 1 or 2 generations of models ago. What are you seeing there? What's the up-to-the-minute thinking about how hiring needs are changing?
I mean, up to the minute, you have to be very proficient at using AI. Some companies have updated their coding interview processes to allow for AI assistance, because on the job you have to be using AI all the time. That's critical to succeeding in this field and to being amplified by AI. I would say that goes for every one of these archetypes we've identified.
I do think that checking whether AI output is good or not in critical contexts is still going to be very important. Stitching together different types of AI output and building pipelines to process that more efficiently are also going to be very critical.
But we might be leaving the LeetCode era. I'll say this: amplifiers, while not currently the most in demand across all the different hierarchies of AI safety organizations or teams, are probably going to be the most in demand in the next year or 2. That's based on my predictions about AI progress.
As you say, it could be slower. There could be jaggedness concerns that slow down this type of talent transition. But in general, it's never bad for your employability to spec out as a manager. Managers are very useful, and leadership traits in general make you a more useful, better employee. It's part of personal growth, I think, to take on some leadership roles.
What does supply and demand look like these days? At the highest level, going back to the origin story of MATS, my understanding was that you said, “Gee, this AI safety thing seems like it's going to require a lot of people working on it in a lot of different roles, and this isn't something that universities teach.” How do people get into this? What's the on-ramp? If somebody would benefit from a program like this, where do they go?
You've created one, and there are some others out there too, but you've created one of the largest and most highly regarded programs. Where are we in terms of whether there are a lot more jobs out there than MATS can produce fellows? How do you think about that?
We've gone back and forth a couple of times. At one point, it was, “We're super talent-constrained,” and then it was, “Well, maybe not so much anymore.” Now it's, “There aren't actually a lot of roles for people to go into.” I feel like this has seesawed back and forth, and I don't know exactly where we are today.
I'll start by saying I didn't found MATS. I didn't co-found MATS. I was in the pilot program as a participant.
Okay.
There were 5 of us who ended up doing the first research program. It was a pilot; they didn't have open applications. It was just people nominated from what was then the first AI Safety Fundamentals course, what's now BlueDot Impact. We did that.
The credit goes to Victor Warlop and Oliver Zang, who is COO and co-founder at the Center for AI Safety. I joined the team right after that program. Oliver left to co-found CAIS, and Christian joined on as well. Then Victor left shortly thereafter.
I would say that I scaled MATS. That's my contribution.
Since Christian and I essentially refounded it, we formed a separate 501(c)(3) a couple of years after that because we got too big for our fiscal sponsor. I’ll take credit for scaling MATS and for being the driving force behind strategy since, I guess, mid-2022. But in regard to talent needs, that’s a good question. Sorry, tell me the exact wording of the question again.
What’s the balance of supply and demand? This might not be right, and you can correct me on this, too, but I’ve had this sense at times that people have said, “There’s so much demand for this kind of talent. Where is it? We’re talent-constrained.”
But at other times, I’ve heard from people that they’re rushing into the field, there aren’t actually that many roles available, and people are frustrated. I don’t know where we are right now in that back-and-forth.
AI safety is a field where there are always going to be jobs for the best people. If you’re a cracked coder, you can get a job in AI safety. The Anthropic alignment science team is growing at 3× per year; they’re trying to scale fast. FAR AI, a nonprofit, is growing 2× per year. At MATS itself, we’ve been growing 2× per year over our entire history. These teams are scaling fast, and many more are being founded.
Coefficient Giving has huge amounts of grant money to spend on this stuff. There are like a dozen AI safety-focused VC firms out there to fund your projects. There are incubators like Catalyze Impact and Seldon Labs. I believe Constellation has one now. There are tons of programs like MATS.
Once you’ve built an organization, especially if you’re scaling very fast, and it hits a certain size, the main constraint becomes: Is this person good enough to warrant the extra management overhead? Can they take on some management responsibilities? At OpenAI, people are managing 10 to 20 individuals; maybe up to—I believe one person in Anthropic’s alignment science had 18 reports. They’re really flat, so you have this real problem where you need to hire people who can quickly ascend the ranks and become research leads, managers, and so on, even PIs of new teams. That’s the limiting constraint.
That’s why a lot of people do some moderate reskilling and then can’t get hired. I think there aren’t many opportunities, but what we find when we talk to these hiring managers is that they say, “We find it extremely hard to hire. We have the money. We have the clear need, but people are not at our bar.” That’s what MATS is trying to do: get people up to that bar.
There’s some technical-skills element to that. There’s also actual research experience. People who come into MATS with prior research experience do so much better on average than people who have less research experience. A strong option for many people should just be to stay in academia, get your bachelor’s, and get your PhD. For other people, maybe they should go off and found a company. There’s tons of money and tons of directions for AI safety companies. I think founders are strongly needed in this ecosystem, and then you can create opportunities for more people to get hired.
Another thing MATS is trying to provide is credibility. I wouldn’t say formal accreditation, but some sense of it: You have a reference from your mentor, who’s a senior researcher in the field; you have the exhaustive MATS selection process, which is trying very hard to find people who are good; and then you have proof of actual research impact. You produce a paper—that’s your name on it. Perhaps you get it published at a conference or on arXiv, and people are talking about it.
That’s what people need to get employed these days. You need to have an actual great output, some sort of deliverable that shows your name, maybe several. You have to be technically good enough at coding or using AI systems, whatever is required. You need references from people who are trusted. Otherwise, it’s just very hard to get ahead. The same story you see in every talent-constrained job market is here.
How does that translate to experience profile? This is obviously a big question in the broader technology world. Are junior programmers an endangered species? We see prominent examples like Neel Nanda and Chris Olah, who broke into the field at a super young age, also kind of defined it in a way, and are still quite young today. That may lead people to think this is a young person’s game, but what you’re describing sounds more like post-PhD, or somebody who’s grown up in an organization to an extent.
I’m thinking of Rajiv from the AI underwriting company, who was at McKinsey for a number of years and has now co-founded this organization, but comes to it with a ton of experience and sophistication in management, leadership, and all that kind of stakeholder-management stuff. What do you see in terms of opportunity for people straight out of college? Are they barking up the wrong tree if they want to go directly from undergrad into this space?
The median MATS fellow is 27. There’s a log-normal distribution and a long tail. The oldest one in the last cohort was 55 or 60. The youngest person is, of course, 18, because we can’t take minors now. People of all ages are applying to MATS.
Twenty percent of MATS fellows are undergrads. They have no bachelor’s degree, or perhaps some of them haven’t even applied for bachelor’s programs. They’re just cracked engineers. About 15% already have PhDs in the bag. So, at least as far as MATS is concerned—as this accelerator, reskilling, retraining, internship, mentorship program, whatever you’re calling it—you’re getting a broad distribution of people.
I think there is obviously huge demand for people with more experience. A second critical thing is that they have experience with the latest tools. Because these tools haven’t existed for very long, young people have a strong chance of being particularly good at using them because they’ve just been constantly on the cusp of things. They haven’t been sitting in a cubicle, not using Claude Code every day. So, to that extent, young people have a huge chance.
But it is the case that you gain valuable knowledge from working on the job, especially in a great team producing great papers, that you can’t replicate. You’ve pointed to some people I would call prodigies. Chris Olah and Neel Nanda are examples, and there are plenty of people of that ilk who have come through MATS. I think—actually, that’s not true: there are some people I would put in a similar class, like Marius Hobbhahn. In that case, our main job is just to get out of their way.
If you’re that kind of person, don’t let anything hold you back. Apply to MATS, apply for grant funding, do whatever, come to the Bay, go to London, and just make it happen. You’ll find your path.
If this isn’t your path, and especially if you’re a more senior researcher or someone who’s thinking, “I can’t conceive of that. I just want to finish my undergrad degree and do a PhD,” that’s fine, too. People from every walk of life have passed through MATS, gotten hired, done other programs, gotten hired, founded companies, and so on. The advice is hard to tailor to a myriad of different types of people, but I would say: Focus on your technical skills and on understanding the frontier of technology.
Don’t be limited by the opportunities you see on job boards. You can create your own opportunities. You can cold-email companies. You can apply to grant funders with some random grant proposal you put together because it fascinates you deeply. You can call up hiring managers and stuff.
When you describe the range, it is a pretty broad distribution, and that tells me that you trust your own ability to discern who’s going to be good more than you trust other outside signals. Tell us what you’re actually using to assess people. This could be translated into practical advice: How does somebody make an application stand out? What are you looking for that allows you to take somebody in their 50s or somebody who’s 18 and feel like you can read what really matters, regardless of their background?
We do some standard stuff that you would see at other tech companies. We have CV review, and we have some CodeSignal tests, so brush up on your coding skills and so on. They do detect AI use. We are, of course, considering ways to allow for tests that include AI use, but these are obviously harder to design and harder to check.
That’s part of our general application, and that’s for some streams. I’ll say this about scholar selection: We’re trying very hard to provide something like a service to mentors.
So if a mentor says to us, “I don’t want to do CV review. I don’t want to do CodeSignal. I just have this selection problem that I want fellows—or applicants—to work on, and then I want you guys to help me evaluate this. Build me a team of contractors or some automated evaluation process to do first-pass screening, and then we’ll go from there,” that’s our favorite kind of evaluation, in some way, because we know that is as close as possible to the actual job, the actual research, as we can get.
Of course, in Yonatan’s case, it’s typically: go away and do a 10-hour mechanistic kind of pseudo-work test, and then present your results to me. You can use AI; do whatever. Just find something interesting. This is great because then we get great results.
For some other streams, it’s harder to do this. It’s harder to administer, and so we do rely on some proxies that are perhaps less specific than ideal. But I think they’re no worse than what anyone else in the industry is doing.
Of course, I think the way you stand out is obviously going to depend on the specific mentor, because MATS is very heterogeneous in that respect. The best thing to apply to Neel Nanda’s stream is going to be vastly different from applying to Ethan Perez’s Anthropic mega-stream.
But in general, you want to really understand your basics about AI and AI safety. So do a BlueDot course, right? There may be some critical knowledge or a paper that, if you haven’t read it, you don’t understand. If you don’t understand what deceptive alignment is, that might be really bad for Ethan Perez’s or Buck Shlegeris’s control research. But even applying—getting to those streams—if you don’t understand that for an interpretability stream, it probably doesn’t matter as much, unless, of course, you’re dealing with deception in your interpretability work.
Make sure you understand your basics. Make sure that if you’re applying to a stream that is empirically heavy, you can do CodeSignal tests and code, including without AI assistance, at least for the time being. It doesn’t hurt to apply to other programs as well. MATS is far from the only program out there now. This isn’t like the early days; there are so many great research programs out there: Pivotal, ERA, PIBBSS, LASR Labs, and ARENA. For technical work, I think Astra is now running again. There are tons of great programs out there, and that can really boost your CV.
If you have experience in the kind of research that you want to do at MATS already, then so much the better. Consider it a postdoc opportunity or a post-research opportunity. Build your own independent projects. Sorry if that’s too much advice to be actionable.
Yeah, I think it boils down to a tangible product being king, right? I say that always in the AI engineering world as well. I’m far from the world’s leading expert on how to break into that space, but what I always tell people if they ask me is that a working demo is kind of the coin of the realm. People might be interested in what you have to say, but they really want to see that you can make something work.
They want to see it online. It could be a Replit, or it could even be a Colab notebook or something, but you’ve got to make something that can work. It sounds like this is a pretty similar worldview: you’ve got to show that you can get in there, make something happen, and, as you put it with Neel’s track in particular, find something interesting. If you can do that, we might have something to talk about.
One thing that jumps out as maybe not as emphasized as I would have thought is being in command of current research. At this point, really nobody can keep up with all the current research, because that exponential has gotten away from all our feeble human minds. I would say maybe a few hyperlexics can still keep up.
But I have found that keeping up with research feels important to me. It feels like an important part of how I stay conversant with people across a lot of different areas. Obviously, what I’m doing in trying to be conversant with people across a lot of different areas is not the same thing as research. How much emphasis do you think mentors in general put on being on top of the literature, so to speak?
It varies. Some of the mentors will ask questions like, “What do you think about X concept?” Others won’t be as interested. Obviously, as you say, these costly signals are the most important thing: have you done good research? Do you have a deliverable, a product? Do you have a strong reference from an important person? That’s also key. Have you done your homework in terms of the BlueDot course and other things?
I think that MATS selection doesn’t currently emphasize breadth of knowledge very much, mostly because mentors don’t necessarily want that. I think that this is maybe a weakness in our process to some extent if we don’t then help people build that breadth.
But we do have seminar programs, and we have tons of opportunities for intermingling between different research streams, which really rapidly builds a breadth of knowledge. We used to have discussion groups, and these still occur occasionally with workshops and so on. I really do encourage everyone to do a basic BlueDot course or equivalent. AI safety, Atlas, and CAIS have good courses as well.
But I think that this isn’t as required for selection. It’s more to prevent you from entering MATS, starting to do a research project, and realizing, “Oh, crap, I have no idea where the gaps are. I don’t understand how my work fits into anything. How do I get funding after MATS? How do I get a job? How do I choose a good original research direction?”
So it’s more for your ability to actually deliver within the program—tracking research—and less to do with your ability to get in at the moment, which is pretty important, because MATS is just a stepping stone. If you do MATS and then don’t produce a great deliverable by the end of MATS, sure, it’s a great thing on your résumé, but it’s not going to be enough in many cases, because it’s such a competitive environment.
I think it’s really good for people to build a shallow but broad understanding of the literature. I would recommend not checking X constantly for new papers unless they’re in your field. Maybe set up some Google Scholar alerts for interpretability if that’s your thing. But every so often, do a periodic deep dive into what all the cool updates are across different fields.
You can do this by looking at the new BlueDot course every year, or looking at some research roundups or highlights every month, like Zvi’s newsletter or Transformer. There are other people you can follow on X. That’s my main recommendation.
So your admissions rate is super low, right? I want to encourage people to apply, but it is a very selective program. What does the funnel look like in terms of applying and being selected? I don’t know if there are intermediate steps that would make sense to talk about.
I think the good news, though, is that if you do get into the program, your success rate on the other end—getting into the field in a professional W-2 employee status right away—is really high. Do you want to run us through those numbers?
Yeah. So, in the last program, I think we accepted around 5% of people—maybe 4%—who applied through our initial intake form. There was a subsequent process that they had to complete, which involved applying to specific mentors and streams, and I think we accepted somewhere around 6% to 12%, maybe 7%, of those people. So, a bit higher. Maybe that’s the figure I’d focus on: somewhere around 7%, let’s say.
Now, that’s better than people think, right? They hear that and think it was something like—I think the Anthropic Fellows Program, for example, is around 2%. Anthropic is a big name, right? But MATS is larger; we have more diversity and more spots, and so on.
People should also just treat the application process as a learning experience in general. We try to make it useful. Some streams are going to be more painful than others, but for streams like Neel Nanda’s, where you spend 10 hours working on a project, you then have something really cool for your GitHub, and that can only help.
Yeah, you’re not working with Neel anyway, I would imagine.
Yeah, definitely the case. I do think it’s unfortunate that there aren’t easy ways to do credit assignment cheaply—to find the best people without them spending a bunch of time. But I know job interviews for top tech companies like OpenAI, Anthropic, and Google DeepMind over the last 3 to 6 months or something have so many things to do before you finally get the yes or no.
So we definitely aren’t that involved. It’s a much slower process, and I think that’s because the commitment is less on our end. We’re not giving people W-2s. MATS is an independent research program. They get grants from a third party. We provide the housing, the office, the mentorship, and the community, but we don’t sign people on for any type of employment, which I think is part of the appeal as well.
That’s the main statistic there: 7%. At the other end, about 75% of our accepted fellows go on to our extension phase.
So our first 3-month program: 7% get in, and 75% go on for another 6 months, maybe even 12 months in some cases. That extension phase is where a lot of great follow-up research happens. Of the people who've done our program over its history, we've had 446 fellows in total, not including people who've done training programs that we've helped facilitate, of which there are probably another 200–300. Of those 446, 80% have gone on to get permanent jobs in AI safety, based on our latest statistics.
That's great. I think 98% are employed in some capacity now. Of those 80%, not all are W-2s, right? Some of them are independent researchers with grant funding from Coefficient Giving or LTFF or something, which I think is a fine situation, right?
Then, in terms of the actual field growth, there are some statistics I can share. It seems, based on Steven McKissick's LessWrong investigation, that the growth rate of the AI safety field is around 25% per year, which is interesting, right? It does seem to be growing exponentially, as far as we can tell. Now, that is a lot less than the growth rate of MATS applications, which are going up somewhere between 1.4x and 1.7x per year, depending on how you slice it, and mentor applications might be increasing around 1.5x per year.
There's a big disconnect. According to BlueDot Impact, I believe their growth rate is something like 370% per year in terms of applications to their programs. So, yeah, there's some large disconnect. A lot of people are applying to BlueDot. That can't go on at that rate; that's just way too fast. But probably that's because they've done amazing advertising and marketing.
MATS has only just started to do advertising and marketing. We had the first-ever open round of mentor applications launch just recently. And, yeah, we sponsored EAGx Europe; that was cool. We sponsored your podcast and several other great venues as well. I think this is only going to cause the application trend to continue. I would guess 1.5x per year, something like that, which is a faster growth rate than the current deployed growth rate of the field.
As to why, I could speculate. It's probably just caused by a high bar—a very high bar—for a lot of these companies, maybe a deficit of founders as well. There are plenty of organizations working to remedy that. I know there was this AI assurance technology report from Juniper Ventures about a year ago where they predicted that the size of the market for AI assurance technologies is doubling each year. So there is a lot of opportunity to do stuff that might contribute to AI safety.
What does the salary distribution look like for the people who are getting jobs? How much of an alignment tax, if any, do people pay on the salary front?
At the frontier labs, no tax at all. They're getting paid the same rates. Yeah, they're getting staggering amounts of money. I think a couple of years ago, the going rate for someone joining off the street was $370,000 or something. I'm sure it's much higher now, especially given all the crazy Meta stuff that happened. I would bet mid-level and higher people are making over $1 million at these labs, even on the safety teams, but I don't have any private data on that. If you join as a junior software engineer, don't be surprised if you get somewhere around $350,000 or something.
At nonprofits, obviously it's lower. They can't compete with equity. They don't have any equity, but they also typically have a lot less funding. Coefficient Giving's pockets aren't as deep as the collective might of U.S. venture capitalism. I think there is something like a nonprofit tax. I wouldn't say there's a safety tax; I'll say there's a nonprofit tax, right?
There are nonprofits that are doing AI capabilities stuff, like the Allen Institute for AI. People make a lot of money still, right? This is artificial intelligence, and Coefficient Giving understands—and other funders understand—that you have to pay to play. So you have organizations like METR, where they are, I believe, offering quite a lot of money for their roles—upwards of $300,000 for most roles, probably over $1 million for some, I would dare say, today.
There are nonprofits that are really, really trying to compete for talent. They can't offer anything like the frontier lab salaries, including equity, but they're trying their best. I think this is kind of reasonable, but it also is a bit of an insane moment. MATS salaries are not anywhere near that high. Maybe we're doing the wrong thing. I don't know.
There are other AI safety nonprofits that have tried very different strategies. You have FAR AI; I think the going rate for FAR AI research scientists is something like $170,000, so significantly lower. They might have actually improved that recently. There is a wide spectrum here, and it really depends on the compensation policy of the organization. But you will see very well-funded nonprofits offering comparable salaries to at least some junior AI company roles.
What about in terms of compute? I know you guys have, in addition to the stipend that folks get as fellows, a compute grant, and I believe it's $12,000 worth of compute. I'm interested in what form that takes. Is it just a Brex card that you can go spend on compute wherever you want, or do you have established compute partners that you work with that serve your fellows well? How often is that enough? Are there times when people find that they need more compute to do what they want to do? And then, if they go work at a nonprofit, how compute-rich or poor are the nonprofits?
MATS offers a $12,000 budget. I won't say we give people a card that says, "Here's $12,000." No, they have to justify their compute spending. They have to have an actual project and proposal that they are going to spend that on, which necessitates it.
Most people don't spend anywhere near that much, which is good because, although we budgeted as if they could, we really don't want to just waste money on compute. Basically, no one has compute limitations. Sometimes people have needed more than that, but rarely. We've considered their proposal, thought hard, and reallocated funds as needed. I think people aren't limited by compute at MATS in general.
The way we do it is pretty good. For our model API calls and all that, we have specific organization accounts that we sign people up for. Then we give them a budget and top them up as necessary. We do have our own MATS cluster as well that we maintain online. Our compute team handles that, but typically people tend to use RunPod and other types of self-service things.
It depends on the kind of research. Some types work well on our cluster, and some types of research—even with the benefit our compute team can provide in setting up and maintaining the cluster—are so customized and require people to tinker with so many things that we just give them a budget and let them use online providers. It's better that way. We did use other providers in the past, but that's our current setup, and we're looking at putting together a kind of customized Claude Code suite as well.
Meaning, like, building out a bunch of tools or MCP-type…
Tools. I don't think so—not MCP, at least.
That's kind of—yeah, you're right. There could actually be a lot of data, you know, in those databases. We could probably put together something really useful at some point, but we haven't thought about that.
Yeah, just the archive of everything that's been tried would be pretty fascinating to do some agentic search through.
I mean, our new research database is online at MATSprogram.org/research. You can see everything there. We have a Google Scholar, but that's just all the papers that got published. You can also see our LessWrong blog posts under the MATS program tag, but many more research artifacts have been produced than are visible there.
You mentioned—you said the word "tinker." Is the Thinking Machines API growing in prominence in terms of what people are finding attractive to use?
Yeah, many people want to use the Thinking Machines API. Many organizations like to donate API credits, which is awesome because we can really use that.
Yeah. Cool. Is there anything else that we should cover in terms of January 18? We know that is important as a date to keep in mind. What other facts should we make sure that we touch on?
MATS is growing. We're always hiring. If you want to work on our team and help grow the next generation of people, if you fancy yourself an amplifier of sorts, you have people skills and research skills, we'd love to hear from you. Go to our website, MATSprogram.org/careers. We're taking on mentors as well. There's an application form on our mentor section of our website. And participants, we're going to run 3 programs this year—rather, not 2, but 3: a summer program, a fall program, and a winter program starting into next year.
And I’m super excited. We also have plenty of other offerings in the works. We’re considering a 1- to 2-year residency program for senior researchers as well. And, yeah, more on that to come.
Yeah, cool. Are you taking any connectors? If I am a connector type, or want to become one, is MATS a way to find my way there, or not really?
Many have. I would call Jesse Hoogland one such person. I would call Paul Röttger, Tmus [?], and Simplex [?]. I’d say Marius Hobbhahn as well, to some extent, with his deception-evaluation work. And probably dozens of people. I’m just sharing some of the names that come more easily to mind, but many, many people have come through MATS. We’re super open to individuals who have this kind of archetype.
And a connector, right? They have empirical skills and theoretical skills, so they could probably succeed in a bunch of different ways, right? But they’re uniquely specced out to connect those 2 things. Now, there are some mentors and projects that are much more suited to this kind of thing than others. People like Richard Ngo and, historically, Evan Hubinger. I think Evan Hubinger has probably been the most dominant connector and driving force at MATS over our time, but he’s not a mentor in the next program, unfortunately. He doesn’t have time.
But, yeah, there are many different opportunities at MATS for this kind of thing. I think even in some of the interpretability streams as well, it’s very possible to enter an interpretability stream and bring with it some model of the kind of theory-based interpretability mechanism or strategy that you want to pursue, and then see that executed on. That’s happened several times.
One of the things that I took note of in the blog post from 18 months or so ago was that you had made a comment that funders basically don’t want to—or they’re much more inclined to—support the growth of organizations that they see as legible, that have research directions that they feel are somewhat established or that they can wrap their heads around. And they’re much more reluctant to fund totally new conceptual directions.
And that seems like it exists in contrast with the AE Studio survey, where they basically found that the field as a whole seems to think that we don’t have all the ideas that we need, and that more far-out ideas should be tried, which, of course, led to their neglected-approaches approach. What do you make of that? Is there stuff that we can do, or is it a different organization’s job to figure out how to fill that gap?
Because I do feel like I want some more, and I love some of the AE Studio stuff, including Self-Other Overlap. I always come back to that as an example of something that’s quite off the map of what most people are doing. When I think of AI control and what Buck Shlegeris and the Redwood Research team are doing, I find that stuff fascinating.
One of the things that impresses me most is that they’re willing to work on something that, in some ways, is so depressing. They’re like, “We’re going to try to figure out how to work with AI even assuming they’re out to get us.” And I’m like, “Yikes.” I don’t know that I would be able to sustain the positive attitude enough to do that if I were working from that premise.
I do feel like there’s a relative dearth of things that are more inspiring here. I think maybe of AE Studio, but also Softmax. Obviously, people have a lot of different opinions on whether these things are ever going to work or not, but I wonder what your take is on just the overall mix.
It seems like a lot of things are more toward: patch the holes, keep the AI down, tempt it—see if it’ll take the temptation—and then patch it if it takes the temptation. There’s not nearly as much that is a more colorful, positive vision for the future. And I wish there was, but maybe that’s just not happening because the ideas are too hard to come by. Maybe it’s not happening because the funders aren’t bold enough.
What’s your take on how we can get—if we should be trying to get—more of that stuff? And if you think we should, how might we go about it?
I have many takes here. Obviously, I advocate a portfolio, and MATS has historically sponsored a bunch of Self-Other Overlap. That project came out of a MATS alum. I’ll just say Mark Koren—I might have messed up his name—was the originator of that project at AE Studio, and I believe Cameron Berg, another MATS 1.0 alum with me, is running some of their more neuroscience-inspired approaches as well.
AE Studio is great. I love what they’re doing. I think that the survey they did is probably not representative of the AI safety research field as a whole, but then again, it might be. Even so, I think we obviously need more ideas, because more ideas are good, right? More bets are good; more shots on goal are good.
Now, I would not advocate for a person who is a very strong iterator to drop that and try to think of some new paradigm. I think that would be strictly counterproductive on the margin, because I think we do have some very strong central research bets that we need more people pursuing, because they will yield demonstrable results.
But if everyone did that, this would be bad, because you need to have your portfolio. Maybe these approaches fail. Maybe they need other pieces to work. Many AI safety research agendas are contingent on other things going right or other people working on other stuff. It’s like any kind of research field: you need to have everyone advancing the frontier.
I think safety has historically gone really haphazardly on different agendas, which is bad. A portfolio approach is much better. Don’t rule things out as possible directions. Just shift and reallocate resources to them.
To their credit, Coefficient Giving has done an amazing job, particularly recently, at supporting a bunch of different novel research bets. And they’ve also funded PIBS, or Principles of Intelligence, a program that is trying very hard to pursue sort of moonshot-y interpretability and agency-understanding projects. So they’re great. Check them out.
I think that more ideas would be good. I think that the kind of person who should be pursuing that typically is going to look something like someone who is already a domain expert in some other area. You are occasionally going to have your Buck Shlegerises, your Evan Hubingers, right, who come along with no PhD but spent years at MIRI, incubating in that kind of deep AI safety, that rich safety experience, and then come out with amazing stuff like Risks from Learned Optimization and AI control and all that stuff.
But short of having access to that type of community and that type of research experience, I think most of the prominent connectors, like your Alex Turners and so on, have spent a lot of time in research-science PhD programs and incubating in that environment as well. So I think MATS is a great way for that kind of person to develop and spawn more research ideas.
In fact, I’ve seen—shout-out Alex Turner—he has come up with some amazing research ideas over his time at MATS, and I think we’ve been very fortunate to support him. Things like gradient routing. And also Alex Cloud, another MATS mentor, has done plenty of other things like activation engineering and steering. He was one of the people involved in that.
So I think that senior, experienced researchers are going to be, probably like most things, the main drivers of new ideas, and grant funding that lets them pursue whatever their research taste dictates is great. And programs like MATS that let them pursue their research agendas are also great.
I also think bounty programs could work as well, but I would caution against people putting all their eggs in the basket of, “We need new ideas because the central idea is not working.” I don’t think that’s true. I think the central ideas are still our actual best bets.
Yeah. Okay. Makes sense. Do you want to shout out any other organizations that MATS fellows have gone to, or even started, that you think are underappreciated? This could be assignment editing for me for future episodes, but also just things that you think people should be paying more attention to than they are.
Yeah, there are tons. You can see all the organizations listed on our website. There are so many amazing people there. I guess I could shout out—it’s hard to play favorites, because MATS is, like, we’ve worked with so many people and we’re trying to be very broad—but, yeah, I guess in terms of nonprofits specifically, because maybe they don’t get as much attention: obviously Redwood, METR, RAND, CAST, Apollo Research, FAR AI, Goodfire, Truthful AI, LawZero, MIRI, and plenty of others.
I love these organizations. We need more, frankly, nonprofit research organizations, and if you think you could found one, give it a shot. But obviously, you need to have a ton of research experience under your belt and very credible references and so on.
Yeah. I don’t know. It’s really hard to play favorites. [Laughter]
Yeah. There are always so many great options for organizations to shout out. It’s a testament to how many fellows have already gone on to do impressive work. So great job by you guys in driving this and growing it over the last few years.
Anything else we should touch on before we break?
No, I just really appreciate this experience. Thank you so much for inviting me to talk.
My pleasure. Thank you for doing it, and keep up the great work. Ryan Kidd, co-executive director at MATS. Thank you for being part of The Cognitive Revolution.
Thanks, Nathan.