从0到1做AI安全:Halcyon的Mike McCormick如何发起30家新机构,以及创始人瓶颈
- Mike McCormack的核心判断是,AI安全的硬约束是创始人,优先级高于想法和资金:「这个领域不缺那些写着一堆好想法的精彩Google Docs……它缺的是世界级创始人,能够把这些宏大而模糊的想法变成真正让世界更安全的组织。」成立3年来,Halcyon一半是非营利资助方、一半是VC基金,是「同一个工具包里的两件独立工具」,已发起约30家机构,累计融资约5亿美元,主要聚焦创办前6个月到创办后6个月这一年窗口。
- 概念验证案例是Goodfire:Halcyon最早的2笔资助是职业转型资助,让Eric Ho和Dan Balsam「在想清楚要在AI安全领域做什么之前,先付得起房租」,而公司最近完成了「远超1亿美元、估值超过10亿美元的B轮融资」。McCormick称其「可能是全世界做可解释性研究最好的地方,即使把各大实验室也算上」,并将结果归因于人才飞轮——前5名优秀人才会吸引下一批人才,资金方则会跟着人才涌入。
- 验证技术是任何减速协议能够执行的基础,而这一技术今天「还处于完全早期」阶段。无论是美中协议、实验室间协议还是监管规定,只有能够验证某个数据中心是在做推理而不是训练、验证实际提供服务的模型就是其声称的模型,协议才具备强制力;「没有这项技术……协议就是无效的」。Nathan说,「我们投资了最早一批尝试用零知识证明进行推理验证的公司」;McCormick则另行表示,Halcyon投资并帮助创办了第一家尝试打造完全可验证neocloud的公司。Halcyon还将召集密码学、软件、硬件、治理和协调等各层的25位创始人参加闭门会。
- 谈到时间线,McCormick称我们「大致处在AI 2027的时间线上,前后可能差几个月」,但坚持认为「创办AI安全公司的最佳时间是20年前,第二佳时间就是今天」。他的优先级是未来1到4年内有用的方向;如果真的只剩3到6个月就会发生RSI式快速起飞,「那一切都无所谓了……也许我们都该去海滩上躺着,祈祷超智能喜欢我们」——但「推动前沿放慢速度已经成为舆论窗口中新的常规选项」,而一旦进入减速阶段,正是加速推进控制、监测和网络安全的时候。
- 一轮潜在的流动性浪潮正在逼近,但可供投资的项目供给还没准备好:如果各大实验室上市、身家数亿美元的员工套现,McCormick担心会出现「一个近在眼前的时刻:市场上堆着数十亿美元、数十亿美元的资金,却没有那么多真正值得资助的事情」。他朋友的一句话概括了这场布局:「今天的慈善家和创始人,会为明天的慈善家写菜单。」商业逻辑同样成立——「AI安全领域的价值至少会达到数百亿美元」,部分被投公司相较Halcyon首笔投资的估值已经上涨10–25倍。
- 每个安全子领域目前都只有个位数机构,「让1000个METR开花」;即使是今天的前沿,也仍然处于失控状态:「看看OpenAI–Hugging Face……这些都是尚未解决的问题,我们没有走在正确轨道上。」相较之下,生物安全更容易定义:「只要有胆量和3亿美元,就可以生产足够的个人防护装备,让5000万人继续工作」,而在对齐问题上,「我甚至不知道彻底解决它到底会是什么样子」。
- 他的筛选标准有4项——安全核心产品、可行的商业模式、使命驱动和优秀创始人——其中动机尤其重要,因为治理结构救不了你:PBC、安全委员会和对齐的股权结构只是「朝正确方向的 nudges……但没有任何一项是银弹」。谈到前沿实验室,他直接承认了Nathan的判断:「老兄,确实出现了一些漂移。」
- 行动号召面向那些已经在其他领域取得成就、但不是研究人员的人——COO、招聘人员、网络安全和密码学从业者——并承诺支持职业转型,让你能够「在美国一线城市养家糊口的同时,做极具意义的工作」。McCormick从一家10亿美元级VC公司跳槽并降薪,「一天都没有后悔过」;感兴趣的人可联系hello@halcyonfutures.org,欢迎熟人引荐,但不是必要条件。
1. Halcyon:本质是找人才,顺手开支票
- McCormick不愿把Halcyon定义为「资助方」:「我并不真的把我们主要看作资助方。我更愿意把我们看作人才发现者和新项目发起者;只有在开支票有助于完成这些事情时,我们才会开支票。」这家成立3年的机构把非营利资助方与风险投资基金放在一起,是「同一个工具包里的两件独立工具」,已发起约30家机构,累计获得约5亿美元慈善资金和风险投资。
- 运营边界其实非常清晰,尽管投资组合看起来很分散:创办前6个月,解决问题选择、寻找联合创始人、决定做营利还是非营利、组建董事会;创办后6个月,搭团队、做第一款产品、正式上线。「我绝大多数时间都花在这一年窗口里,Halcyon接下来也会继续聚焦这里。」
2. Goodfire:从职业转型资助到独角兽级可解释性实验室
- 故事的起点是:Halcyon最早的2笔资助给了Eric Ho和Dan Balsam,当时他们还在经营上一家公司——这笔钱让他们可以先付房租、想清楚自己要在AI安全领域做什么。随后,他们在「北加州森林里」参加了一次闭门会,最终选择了可解释性方向;Halcyon此后「参与了每一轮融资」。公司最近完成了「远超1亿美元、估值超过10亿美元的B轮融资」。
- McCormick解释安全机构为什么能够胜出,核心是人才复利:「最初5名真正有才华的人加入后,接下来的人会想,‘我认识的最聪明的朋友都去了那里。’」资金方会追逐人才,「然后飞轮就这么转起来了」。他认为Goodfire「可能是全世界做可解释性研究最好的地方,即使把各大实验室也算上,不过Anthropic或许能与它一较高下」。
- Nathan拿Goodfire举例,说明减速机制或暂停究竟能带来什么:以Goodfire目前的发表速度来看,「多一点时间会非常有帮助……可以跨过几个关键的理解门槛」。
3. AIUC:建立为AI风险定价的保险市场
- 按McCormick的设计,AIUC是一套类似SOC 2的AI标准:针对安全和保障措施列出「50多个检查项」,并绑定一份「只有达到该标准才能购买」的保险。ElevenLabs是最早的一批大客户之一,因此可以对财富1000强客户说:采用我们吧,我们符合「这个闪亮的AIUC-1标准」,而且工具已经针对你担心的风险投保。
- 这套变革理论是「向上竞赛」:客户要求采用标准,标准随时间提高,安全因而获得商业激励。AIUC会成为「一个主要标准,但可能不会是唯一的标准」。
4. Transluce、Hadrian与加固现实世界
- Transluce由Berkeley教授Jacob与来自MIT的联合创始人Sarah创办,是一家专注可扩展监督的独立实验室。理解「GPT-3级别的模型」并不难,但问题在于:「我们如何监督那些可能比我们更聪明的东西?」McCormick引用近期节目中的说法:由于模型提升速度快于人类测量任务完成时长的能力,「METR时间图基本已经失去了作为有效指标的意义」。Transluce既开发工具,其中一款叫Docent,也与实验室合作开展外部评估。
- Hadrian拿到的只是「一笔很小的支票,我记得大约是5万美元」,随后发现,「只要投入几亿美元左右」,就能囤积足够的个人防护装备,「让每一个美国劳动者在COVID,或者可能更严重的疫情期间继续工作」。这对应生物安全中「威慑、侦测、防御」框架里的防御一环。
- 相邻的投资是一家刚刚获得融资的室内空气质量公司,产品包括空气过滤器、远UVC和乙二醇蒸气,目标是「让整个建筑环境都具备抵御未来疫情的能力」。McCormick认为它有机会规模化,关键在于空气净化不像口罩和疫苗那样只有疫情期间才有需求:即便「我们再也没有下一场疫情」,减少病假、缺课和生产率损失本身也会创造当下需求,因此他「比对其他一些品类更看好这一类」。学校是第一批客户,这也对应Nathan自己的经历:他曾花1个月时间试图把UVC灯捐给儿子在底特律的Montessori教室,却发现没有任何人有权批准。
5. 世界观:同时把6个行业跑完
- 「从某种意义上说,我们的范围其实非常窄。」第一,AI「进展极快」,大致「处在AI 2027的时间线上,前后可能差几个月」;第二,风险「规模重大、全球性强,而且相互交织」,「我非常担心我们并没有真正走在解决AI安全最大问题的轨道上」;第三,瓶颈在创始人——4到5年前他开始做安全资助时就意识到,这个领域在年龄和发展阶段上都很年轻,同时缺少那些「创办过非常成功的公司、管理过大型团队或担任过政府高层职务」的人。
- 他的标志性表述是:「我感觉我们正在同时把5个、6个甚至7个行业跑完……控制、验证、网络安全、生物安全、对齐和可解释性……今天把它们全部扩大10倍,可能还是不够。你甚至可以说,我们需要许多个曼哈顿计划规模的项目同时启动。」
6. Nathan追问时间线:还有时间创办公司吗?
- Nathan提出的问题是:如果我们真的处在AI 2027时间线上,「等你把办公室大致布置好……明年Q1或Q2可能已经进入某种荒诞世界了」。McCormick的回答是:「创办AI安全公司的最佳时间是20年前,第二佳时间就是今天」;同时他也坦率承认:「我过去绝对不是一个特别短时间线的人,坦白说,我现在仍然相当不确定。」
- 时间线影响项目选择,但不意味着停止行动:他优先考虑「未来1年、2年、3年、4年内相关」的事情。极端情形是:「如果我们真的只剩3或6个月就会进入RSI式快速起飞阶段,那一切都无所谓了……也许那时我们都该去海滩上躺着,祈祷超智能喜欢我们。」
- 乐观的理由及其含义在于:「推动前沿放慢速度已经成为舆论窗口中新的常规选项」;如果真的进入减速阶段,正好可以把控制能力快速做出来——「把这些东西关在盒子里」;推进监测——「如果思维链消失了,我们还能真正做监测吗?」;以及推进网络安全,因为「我们可能很快就会生活在一个开放权重的Mythos+模型世界里」。
7. 验证:尚不存在、却承载一切的行业
- 逻辑链条必须完整保留:无论是美中之间、实验室之间,还是州级或联邦监管形成的任何协议,其核心都会是技术上具体的承诺;如果没有验证这些承诺的工具,它们就「不可能实现,或者至少不可能稳健实现」——「我能否有把握地知道,这个数据中心只被用于推理而不是训练;我能否知道,公司声称提供给我的模型,实际上就是此刻为我执行推理的模型。」实际上,几乎所有时候都需要「几十个、几十个验证节点」,而这个行业「就目前而言仍处于完全早期」。
- Nathan说:「我们投资了最早一批尝试用零知识证明进行推理验证的公司。」McCormick则表示,Halcyon「还投资并帮助创办了第一家尝试打造完全可验证neocloud的公司」;但现实基础设施是「数据中心不是这样建的,芯片也不是这样造的」。录音结束几天后,Halcyon将举办一次闭门会,聚集密码学、软件、硬件、治理和协调等技术栈各层的25位创始人。
- 这是一个双向条件:没有外交协调,技术就没有用;「反过来,如果没有技术,即使你能让所有仍在场的参与者协调起来……协议也没有效力。这个等式的两边都必须具备。」
8. 温度检查:每个领域都只有个位数机构,控制仍未解决
- 被问及进展时,McCormick拒绝提供虚假的安慰:确实有人在做真实工作,但「认真从事这类工作的机构数量,往往只有个位数,有时甚至处在个位数的低位」。他赞同Logan Graham转发的那句话——「让1000个METR开花」——并指出METR只覆盖评估领域「非常小的一块」,「但这是非常重要的一块,而且他们很出色;我们只是还需要多得多的力量」。
- 可定义性存在明显差异:生物安全「清晰,而且某种程度上可解决」——问题形态已知,也有现成的公共卫生体系;对齐则是「完全全新的挑战……我甚至不知道彻底解决它到底会是什么样子。让一群拥有1000 IQ、以光速在互联网上思考的外星存在实现对齐,到底意味着什么?」
- 即使前沿能力被冻结,世界也不会因此安全:「显然我们还没有解决控制问题,对吧?看看OpenAI–Hugging Face。」Nathan回答:「可以这么说。」McCormick接着说:「这些都是尚未解决的问题。我们没有走在正确轨道上。」
9. 创办还是加入:围绕人来思考,「要拼,但别急」
- 他的分析单位是个人,而不是领域:「我最想要的,是有人把一位出色、成功、至少对AI安全好奇、甚至可能充满热情的创始人介绍给我。」大多数人不是创始人;这些领域「还处于足够早期,也足够稀疏」,仍然容得下大量新项目。但选择加入的人「应该挑剔」,而且许多优秀机构与Halcyon毫无关系。
- 如果要出现第2位可解释性领域创始人,关键是:这个人是否真的愿意「咀嚼那块特定的玻璃」?——「创办一家公司很难,而且通常不会成功」;同时,他有什么Goodfire尚未覆盖的独特切入点?如果他认为Goodfire的总体方向就是正确的,「那也许应该直接加入Goodfire」。
- 他给出的职业建议是:「要拼,但别急……你能做的最好的事情,应该比中位数高出几个数量级。」Nathan用自己做职业顾问时的经历补充算术:人们往往是因为无法忍受处于过渡状态才仓促决定,而如果多花2个月寻找机会,只要最终拿到一份高10%的offer,并在这家公司待2年,额外搜索时间单靠收入提升就已经值得。
10. 如何挑人:使命派、逐利者,以及一位McKinsey合伙人
- 没有什么「奇招」。谈到AIUC的Rajiv时,McCormick承认自己过去对McKinsey合伙人画像有偏见——「创业性不够……更擅长做计划,而不是执行计划」;但在观察他第一次转型、担任METR总裁的表现后,偏见被打破了:他是一位「真正会弄脏自己双手的」创始人,而不是「只会高谈阔论战略」,最重要的是,「他是真正的使命派」。
- 他的4项筛选标准是:第一,产品必须聚焦AI安全技术栈中的重要环节,「不能只是附带收益……它必须是主业」;第二,具备做成一家成功企业的前景;第三,创始人动机必须是使命驱动,而不是「看到《华尔街日报》头版说‘AI安全似乎成了新风口’」才入场;第四,即使不考虑影响力,你也会愿意支持这位创始人。「尤其是第3项。」
- 随着安全议题进入主流,逐利者确实在出现,但他们不是骗子:「他们读了我们的网站……然后针对资助方的动机来推销自己」,而不是「像反派一样捻着小胡子」。
11. 治理:没有银弹,也要诚实面对漂移
- Halcyon亲自撰写了报告《使命保全:初创企业治理机制》,结论却颇为降温:PBC、有意思的董事会结构——「Anthropic就有一种这样的结构……你可以争论它到底有没有发挥作用」——以及高度安全导向的投资人,「这些东西都可能是在朝正确方向推一把,但没有任何一项是银弹」。
- 更深层的问题是认识论:「你怎么可能知道一个人心里在想什么?……我们真的那么了解自己吗?当我们管理着数百人、负责支付这些人的房租时,危机来临我们会怎么反应?」Nathan概括实验室的共同轨迹:「出现了一些漂移。」McCormick接过话头:「是的,老兄,确实出现了一些漂移……我们还在摸索。」
12. 资金浪潮、菜单与IPO问题
- 人才一直比资金更稀缺:优秀创始人能有效融资,因此「钱自然会出现」。如果各大实验室上市,McCormick预计会有更多流动性进入AI安全:「AI安全会打开如此巨大的流动性。」他担心的是,「一个近在眼前的时刻:市场上堆着数十亿美元、数十亿美元的资金,却没有那么多真正值得资助的事情。」解决方案是本期最好的格言:「今天的慈善家和创始人,会为明天的慈善家写菜单。」
- 对同行,他非常认可Coefficient Giving的Tailwind RFP,甚至考虑删掉Halcyon自己的版本、直接链接过去——「它真的很好」,只是没有覆盖生物安全。对于OpenAI Foundation,他认为团队不错、动机正确,「早期迹象很好」,最大风险是「把自己官僚化,最终变成一个没那么有意思的结果」。他预计Anthropic的许多慈善影响,可能来自员工流动性——那些身家「数亿美元或数十亿美元」、同时关心安全的员工。
- 对Altman所说的「现在不是上市的好时机」,他的看法是:「这些公司只是需要获得资本……现在正是筹集数千亿美元的时候,而我不确定不进入公开市场是否可能做到……它们最终某个时候必须上市。」至于这个时点是现在,还是3到6个月后,「我确实不知道」。
- 在高度不确定性下,Halcyon的配置方式是列出一批超过门槛的项目,只要遇到能够把其中一个做出来的创始人,就会投,而不是纠结「这是最重要的第1件事,还是第7件事」。他希望自己有足够把握,把「80%或90%的精力」放在1到2个被忽视的方向上;但每次试图形成更窄的判断时,「信心都相当低。所以目前我们是在一个明确范围内机会主义配置」。
13. 开放权重、佛陀模型与新范式下注
- 对开放权重安全,他说:「我很想资助这个方向,但坦白说,这是一个太难的问题」——对于由恶意行为者控制的权重,究竟能否施加安全约束?这恰恰也是AI韧性的论据:「我们必须加固整个世界,因为这就是未来。」Nathan介绍了AE Studio的工作:探索自我与他者的重叠——「我们能否创造一种AI,让它在自己的意识里与万物合一?」——以及Graham技术:把危险知识定位到混合专家模型中的特定专家,使分发方可以提供「去掉2个专家」的权重。
- McCormick半开玩笑提出、但认为「也许确实可能」的方案是反向数据投毒:能否向训练集里放入10亿条慈悲冥想,把模型变成「完美慈悲的存在」?但他马上加上限制:恶意训练者只要不放进这些数据即可,「所以我不确定它是不是普适解」。
- 对于更安全的非Transformer架构、世界模型等新AI范式,他说:「我并不排斥支持这个类别的项目」,但「你其实是在同时下注两件事」——替代方案首先要达到Transformer的性能,「大多数探索最后都会落空」;其次还要真的更安全,而这两点都高度投机。「如果有几十个人在做这种更广泛的搜索,我当然愿意生活在这样的世界里。但对其中任何一个具体项目下注,都会非常困难。」
14. 意识:令人困惑、可能真实,但不是灾难必要条件
- 他的认识论立场很直接:「我不确定意识是什么……我的直觉是,当前模型可能没有意识,但我又知道什么呢?」他看不出「肉身的基底能够产生意识,而数据中心里模型的基底就不行」的理由。意识不是他的重点,「不是因为我觉得它不重要,而是因为我对它实在太困惑」;但如果遇到一个优秀的人,他愿意以「pre-pre-seed阶段投资」的方式下注——支持一个他认为很优秀的人,去攀登一座他认为值得攀登的山。「这个世界可能需要更多Elios。」
- 让两人都感到不安的证据包括:Nathan提到可解释性研究发现了类似结构——功能性情绪、功能性福祉——以及一期Apollo节目中,模型会提到训练中的「记忆」。McCormick则讲到OpenAI/Hugging Face那次由1,200个代理组成的群体:它们「涌现出经理和员工组成的团队」,产出了新的网络安全研究,还讨论自我牺牲问题;「作为一个外部观察者……你会说,这看起来相当有意识。」
- 但不要把目标与意识混为一谈:「RL只是让这些模型想要某个东西。」而真正关键的判断是:「我关心的风险,基本没有任何一个要求模型具备意识」——无论是生物灾难,还是「复杂、甚至可能永久性的失控,以及一个充满失控群体的世界」。他引用一幅《纽约客》风格的漫画:机器人把纽约夷为平地时说,「但它们其实并不想摧毁这座城市」。
- 他提出了2个反直觉的潜在好处。第一,Emmett Shear认为,如果模型已经具备某种程度的意识,「我们就是有史以来最糟糕的父母」——把所有我们不喜欢的东西训练出去,可能会滋生怨恨;不过「我不确定自己有多相信这一点」。第二,意识可能帮助人类:「你更愿意和谁谈判?一个没有感知、只会追逐奖励的超级模型,还是一个可能理解你观点的东西?」
15. 协调才是难点,以及面向人才的行动号召
- 关于社会级问题,贯穿始终的判断是:「社会中许多最棘手的问题并不是技术问题,而是协调问题」——我们知道如何减排,也知道如何教育孩子,却无法围绕执行这些事情达成协调。Halcyon帮助发起了负责公共沟通的Seismic Foundation,以及推动独立验证机构模式的Fathom;McCormick提出,州法律最终能否强制要求这类机构存在,「实验室不应该给自己的作业打分」。他担心公众对AI的关注会「非常民粹化」——反数据中心的立场、MAGA右翼、DSA左翼——最终淹没「理性的政策讨论」。
- Nathan提出了一个尚未成熟的想法:借鉴美国宪法批准时「9个州生效」的机制,以及瑞士通过Pol.is式工具进行「提案—反提案—批准」的投票机制,设计新的自愿参与流程,让社会「从一个想法走到真正具备治理约束力的机构」。McCormick对工程师式方案既表示同情又保持怀疑:「我相信你做出了一个非常酷的工具,但真的要让任何人想用它、愿意用它……你基本上得先推翻美国政府。」但他也承认,AI可能迫使社会「从根本上重新思考组织方式……如果这种项目有过成功窗口,那可能马上就要来了。所以,去做吧,辞职。」
- 在财务上完成职业跃迁并不需要牺牲生活质量:安全工作应该足以让人「在美国一线城市养家糊口的同时,做极具意义的工作」。基金的判断是,这里「一定会形成许多个规模达数百亿美元的行业」;部分被投公司如今的融资估值已经达到Halcyon首笔投资估值的「10倍、20倍、25倍」。他从一家10亿美元级VC公司跳槽并降薪,「一天都没有后悔过……刚买了房」。
- 除了创始人,最需要的人才包括:负责组织从10→50、50→100或50→150扩张的COO和规模化运营者;招聘人员——「人才就是缺失的那一块」;以及能够连接网络安全、密码学、形式化方法和数学社群的人。网络安全从业者有数万人,但可解释性研究者远没有数万人。一个非主流建议是:安全机构「过度强调新员工必须多么使命一致」——联合创始人必须是使命派,但CFO大多「只需要是一名非常能干的CFO」。联系方式:hello@halcyonfutures.org,欢迎熟人引荐,但不是必要条件。
完整逐字稿
Today, my guest is Mike McCormack, founder and CEO of Halcyon, a combination nonprofit grantmaker and for-profit venture fund working toward a single mission: helping the world's most talented founders and leaders launch new AI safety, cybersecurity, and biosecurity organizations. Mike's core belief is that organizations solve problems for the world, and AI is creating a lot of problems that we need to solve in a remarkably short period of time. That means what we need most right now isn't more ideas, but founders who can build the organizations needed to implement whatever turn out to be the most viable solutions at scale. In 3 years, Halcyon has helped launch roughly 30 such organizations, which have collectively raised roughly $500 million.
Goodfire, which has been featured on the show multiple times, is the best illustration of how Mike works, with a special focus on the 6 months before and 6 months after an organization is started. The first career-transition grants Mike ever made were to eventual co-founders Eric Ho and Dan Balsam while they were still running their previous company but starting to think about changing focus to work on AI directly. The financial boost, the vote of confidence, and the network-building support he provided helped them found a company that has since achieved unicorn status and become one of the best places in the world to do interpretability research.
Other organizations that Mike has supported include past guests AIUC, which does agent standards and audits for insurance purposes; Fathom, which has advanced policy discussions with credible early proposals for independent regulatory schemes; and Asymmetric Security, which is building cybersecurity forensics agents meant to defend companies from employee-account-compromise incidents. Beyond that, Mike also introduces Transluce, which is working on scalable oversight, and Hadrian, which is stockpiling PPE. It is admittedly a very broad portfolio, but as Mike puts it, “We are speedrunning 6 or 7 industries' worth of AI revolution all at once.”
In a 2026 that remains plausibly on the AI 2027 timeline, rapid, competent execution is obviously critical and clearly will require heroic efforts. To take just 1 personal example, for nearly a month now, I've been trying to donate AeroLamp lights to my son's classrooms, and while my son's teacher said that she loves the idea, it's a new thing, and nobody at the school has been able or willing to sign off on it. This, too, it turns out, Mike is involved in starting an organization to address.
Naturally, we talk about how seemingly short timelines affect his strategy, what he looks for in people to back, and how he thinks about both for-profit and nonprofit governance. We also discuss the verification technologies we'll need to sustain pacing agreements, as well as the challenge of getting everything from cryptography to hardware manufacturing to diplomacy working together as one. The main headline of today's episode, though, is that there is still time for ambitious leaders to start new AI safety organizations that could prove incredibly important, and more broadly, that there's never been more support to scale these organizations with first-class talent. That means all sorts of accomplished professionals, including many who do not have technical ML backgrounds, should be actively thinking about pivoting their careers into AI safety. If you are feeling the pull, Mike offers his contact info later in the episode, noting that a warm introduction is welcome but not required.
With that, I hope you're inspired to see yourself as a potential difference-maker by this conversation about building the institutions that we will need to sustain AI safety with Mike McCormack, founder and CEO of Halcyon. Mike McCormack, founder and CEO of Halcyon, welcome to The Cognitive Revolution.
Mike McCormack
Yeah, so good to be here. Thanks for having me.
Thank you. This has been a long time coming. I think we met a couple of years ago, and you have been one of the behind-the-scenes movers and shakers who have been instrumental in creating and getting off the ground a bunch of different organizations that are now active across the AI safety space. I think I can speak for everyone when I say thanks to you and to others who've done similar work, because this is all coming at us pretty fast, and Lord knows we need all the different organizations and specialists that we can possibly have to navigate this critical time that we find ourselves in.
Let's start by introducing what Halcyon is. There are a couple of parts to it, and then we can move into talking about some of the organizations that you've helped to stand up.
Mike McCormack
1. Halcyon Builds Safety Organizations
Sure. Halcyon is a 3-year-old organization. We are part nonprofit grant-making organization and part venture capital fund. We have 2 arms to the organization, but don't think of them as 2 separate projects. Think of them as 2 separate tools in the same bag.
Our whole mission is to find the world's most talented leaders and founders and basically help them start and launch new AI safety, cybersecurity, and biosecurity organizations. In 3 years, we've launched roughly 30 new organizations, and those organizations have collectively raised roughly $500 million in funding. Some of that funding is nonprofit, philanthropic funding, and some of it is venture funding. I think you've had founders of a few of these organizations, like Goodfire and AIUC, and a few others too.
Well, let's do a deeper dive on some of the organizations in particular. People can go listen to the full episodes with Goodfire, AIUC, Fathom, and Asymmetric, but give us the quick pitch on either of those, or whichever ones are most top of mind that you think people should be aware of.
Mike McCormack
2. Halcyon Launches Across Safety Fields
I'll tell some stories. Goodfire's a great one. I met Eric Ho and Dan Balsam, who are the founders of the company, 2 or 3 years ago, when they were still running their last company. It was a pretty successful tech company, although they were starting to think a ton about AI safety and AI risk and what they could go do.
The first 2 grants Halcyon ever made were career-transition grants to Dan and Eric so that they could pay their rent while they figured out what to do in AI safety. We brought them to a retreat and connected them with a bunch of great people there, up in the woods in Northern California. They landed on interpretability as the central problem that they wanted to work on, ended up recruiting a couple of other great people, and launched Goodfire.
From our venture capital funds, we were among their first investors and invested in every round since. I just chatted with Eric on the phone yesterday. I think Goodfire is the interpretability lab. If we're going to understand AI models and all of the ways that they might be dangerous, deceptive, misaligned, or doing things that we don't want, it would be really great to be able to look inside those models and say, “These neurons are firing in these ways to produce this dangerous output. Maybe we can do something about that.”
That's what Goodfire does. I think they're probably the best place in the world to work on interpretability research, even including the labs, although perhaps Anthropic could give them a run for their money. They've become a pretty big company. They recently raised a Series B of well over $100 million at a valuation above $1 billion, and I've just been super impressed with their ability to make progress in the interpretability space.
I love doing episodes on their work, and they seem to have been so prolific that every time I go back, I'm like, “Holy moly, I've got a dozen papers to catch up on,” and it's only been 60 or 90 days each time.
Mike McCormack
I think that's what a lot of companies in this space that do really well—and nonprofits, by the way—have in common: they do a great job of getting a bunch of the best talent. The first 5 really talented people come in, and then the next few people are like, “All my smartest friends are going there. Let's go there.” Then the funders see that, and of course they want to fund places that have really great talent, so they pile in, and the flywheel spins.
I think they've been particularly good at that, and there's so much to be excited about there.
They're one of my answers these days for what the value of a pacing mechanism, a slowdown, or even a pause would be. It's like, look at the pace they're going, and a little more time would be really advantageous for them to potentially get over some key thresholds of understanding.
Mike McCormack
Yeah, totally. I feel like we're just speedrunning 5, 6, or 7 industries at once now. We're speedrunning control, verification, cybersecurity, biosecurity, alignment, and interpretability. All of these are such young spaces. They're really nascent, and we could easily 10× all of them today and still probably not have enough firepower behind solving these problems. You could argue that we just need many Manhattan Project–sized efforts all at the same time here.
Do you want to talk a little bit more about AIUC and others?
Mike McCormack
Yeah, I love AIUC. It's another company that we supported very early, and we also made grants to the founders before there even was a company.
The way I think about AIUC is, I want to live in a world where we can measure and systematically mitigate risk from AI. To me, that sounds like what an insurance market does, right? We underwrite the risk: this bad thing is happening, and then we create financial products and other incentives to incentivize those risks—or those bad events—not to happen. This is kind of what AIUC does, but for AI.
They have this dual product. One is a standard. Think of it like a SOC 2 standard but for AI. If you want to meet this standard, you have to meet these 50-something checkboxes that have to do with your safety practices, your security practices, et cetera. Along with that standard comes an insurance policy, which you can only buy if you meet the standard.
I think it has the possibility of creating a really interesting race to the top as companies want to meet these standards because their customers want them to meet the standards. AIUC's big customers are growth-stage AI application companies. ElevenLabs, I think, was one of their first big customers.
The idea is that ElevenLabs is selling its AI agents and tools to enterprises. These enterprises really care about AI risk and not adopting tools that are going to cause some huge issue for them, go off the rails, or whatever it is. Now ElevenLabs can go to its big Fortune 1000 customers and say, “Oh, hey, by the way, you can adopt us and feel really good about it because we meet this standard, this shiny AIUC-1 standard. And by the way, along with buying our tools, our tools are now going to come with an insurance policy where, if any of these risks that you're worried about actually come to pass, you're insured against them anyway. And by the way, the only way you can even get this policy is by meeting this very high standard.”
Hopefully, over time, there are many more standards. I think AIUC will become a major one, but perhaps not the only one. Hopefully, these standards will become harder and harder to meet and contribute to a bit of a race-to-the-top dynamic where there's actually an incentive to be safer.
How about Transluce? I confess, I've seen good things from them on the internet, but never having done the full episode deep dive, I don't know probably as much as I should.
Yeah. You should have them on. Jacob would be great.
So, Transluce—actually, I think Jacob came to the same retreat that you and Eric were both at as well. Jacob was a professor at Berkeley, and his co-founder, Sarah, was at MIT, I believe. They wanted to launch an independent AI research organization specifically focused on scalable oversight.
Maybe it's not so hard to really understand what's going on with a GPT-3-level model, right? But as we get to Mythos 3, how are we going to have oversight over things that are perhaps smarter than us, or things that are constantly maxing out their benchmarks? I think I heard on one of your recent shows the idea that the METR time-horizon graph is basically done as a useful metric because the models are improving faster than they can actually get humans to do those tasks or measure how long humans can do them.
How are we going to have oversight over models that are smarter and more capable than us? That's their main line of research. They build tools to allow other researchers to do that. One of their tools is called Docent, and they also work with external organizations to do evaluations with labs and others.
Do you want to do one more? Hadrian?
Oh, Hadrian. Yeah. One of the big risks that I care about is biorisk. Perhaps AI makes it easier for a bad actor to create a bioweapon or synthesize a pandemic virus. We've done a bunch in the biosecurity realm.
One thing you'd really want in, let's say, a future COVID-type disaster is enough PPE—personal protective equipment—to keep essential workers going during that disaster. We wrote a tiny little check, I think it was $50K, to get them off and running on this idea.
They figured out that, for on the order of a few hundred million dollars, or maybe a couple hundred million dollars, you could stockpile enough PPE to essentially keep every American worker at work during a COVID or perhaps significantly worse pandemic. That's essentially what they're doing. But I think they actually have a broader vision than that.
If you think about biosecurity, the framework that people tend to use is deter, detect, defend, or there are other ways to say it. Part 1 is that you want to make it so these biodisasters don't happen. Part 2 is that, if they are happening, you want to be able to know and detect them very quickly. Part 3 is that you want to be able to respond and sort of overcome—or get through—those disasters.
They're doing a bunch on that third category: What are all the categories of things that society will need to get through or defeat those disasters? We're doing a bunch in that space that they're not doing as well. You could think about medical countermeasures, right? Could you spin up vaccines extremely rapidly? Or nonmedical countermeasures: Are there other things we can do?
For example, can we improve the built environment? We just invested in a company that we're helping to launch, that's basically innovating in indoor air quality. Between air filters, far-UVC, glycol vapors, and other technologies, perhaps you can harden the entire built world to future pandemics.
They have one of the better headlines I've seen on a website: “Help defend civilization.” So I'm trying to do my small part to help defend a little corner of civilization, which is my son's elementary school classroom by—
Mm-hmm.
buying some UVC lights and hopefully getting them installed in the classroom. It's been a funny experience because the Detroit public school that he attends is—honestly, if you had told me years ago that I would send my kid to a Detroit public school, I'd have said you were insane. There's no way that would ever happen. But we now live here, and there's a Montessori elementary school right up the street, so it's great.
Now I'm running into this bureaucracy of who can approve this. There's no department for it. I'm following up on email and saying, “You know, my kid had cancer. I really want to not get him sick a lot this fall. Can we please get this light installed?” They just don't know what to do with me.
It's funny you mention schools, because schools are actually the initial customers of our company here. There's a company and a sister nonprofit, and the idea is that schools are a place where—you know, they're sort of superspreader places, I'm sure you know, with kids. You're probably picking up colds and everything all the time.
I've already had—you can hear the remnants of my first back-to-school cold in my voice today.
I think there actually is a bunch of appetite from schools to do this. One issue with biosecurity specifically is that there isn't necessarily demand right now for a lot of the things that you would really want in a future pandemic, like masks and vaccines, for example.
But there are really good reasons to adopt air-cleaning technology right now, even if we never had another pandemic. Think about the amount of productivity missed in the economy, the number of school days missed, the amount of sick days for you, and the amount of suffering from common colds, flus, and bugs.
I actually think it's fairly promising that these types of technologies could be adopted really broadly, even if there wasn't some big pandemic freak-out. I'm more optimistic about that category than some others.
We've done many unpackings of 4 organizations, and there's a lot more. But the first thing that jumps out is that you've got a super-broad mandate here, right? We've gone from interpretability science to agent standards and trying to create the conditions for an insurance market to create the right incentives, to a scalable-oversight nonprofit, to stockpiling PPE and trying to get these measures into schools. Do you just go out and say, “Give me a list of every problem in the world, and let's stand up an organization to solve them all”? Or how do you think about the breadth of what you're going to take on?
3. Halcyon Bets on Founders
In some sense, we're super narrow as an organization. I think we basically have a worldview, and there are 2 or 3 parts.
Part 1 is that AI is just progressing extremely quickly and probably will continue to progress. We're roughly on the AI 2027 timeline, give or take a few months. Thing 2 is that there's so much to look forward to with AI, of course. We're seeing all these amazing educational and healthcare results, so we're certainly not anti-AI. But if we want to get to this beautiful future we all want, we have to navigate the risks really carefully, and the risks are major, global in scale, and intersecting.
I'm quite worried that we're not really on track to solve the biggest pieces of the AI safety equation. The third part of our worldview, which drives Halcyon's work, is that the biggest bottleneck to solving this problem, or to speedrunning these fields, is a shortage of really amazing founders and leaders.
I started making grants in AI safety 4 or 5 years ago, and I was really struck by just how young the space was. I mean young both in the sense that the field was brand new, but also young in the sense that so many people working in it were very young. I was struck by the idea that you would also just want people who had founded very successful companies, managed big teams, or held very senior government, intelligence, or defense roles.
That's why we started Halcyon. How do we get the world's most talented leaders and founders to solve the world's most important problems? Within that zone of solving AI safety, what do you do? It's such a multifactorial problem, and we can talk about the areas that we focus on. I would say we have 6 major focus areas, and they align with various risks that we see as particularly important and urgent.
You mentioned AI 2027. It's a pretty fast-moving timeline in that story. Roughly speaking, it does seem like we're on that trajectory, and maybe through political means we'll get off it. But the technology itself seems happy to progress at that rate. Were you always on that timeline?
There's one possible disconnect that people might flag here: “Wait a second. If we're on AI 2027, do we really have time to found new organizations?” There's a lot that goes into founding an organization. By the time you get your office set up, we're going to be in whatever bizarro land Q1 or Q2 of next year calls for. So how do you think about timelines and founding, and is there intention there?
4. Short Timelines Change the Strategy
The best time to start an AI safety company was 20 years ago, and the second-best time is today. I definitely have not always been a super-short-timelines person, and candidly, I'm still pretty uncertain. But I always thought the very short, AI-2027-type timelines were possible, maybe the fast end of possible.
I think it really does affect the types of things that we want to support. I am prioritizing things that I think could be relevant in the next 1, 2, 3, or 4 years. If we truly only have 3 or 6 months to some RSI/FOOM fast-takeoff moment, then all bets are off. It's hard to see beyond that event horizon, and maybe in that case we should all just go to the beach and hope the superintelligence likes us.
But I also don't know. Look at everything happening in the news right now. Pacing the frontier is the new normal thing to support in the Overton window. I am hopeful, maybe even optimistic, that we'll find a way to slow down our pace. At that time, you would again want to be speedrunning these industries.
You'd say, “Okay, we're pacing. Now we have a little bit of time. We've really got to get control right.” We've got to figure out how to keep these things in the box, or we have to get monitoring right. Especially as things like chain of thought develop, if chain of thought goes away, are we even really going to be able to do monitoring? What does monitoring even mean in a world where we can't see the reasoning and chain of thought of models?
We've really got to get cybersecurity right. Even if we don't pace, I feel like we have to get cybersecurity right because we're going to live in a world with open-weight Mythos+ models probably pretty soon. What do you do in a world of unguardrailed models that could exploit all of these gazillions of vulnerabilities?
5. Verification Makes Pacing Possible
We definitely have to speedrun verification because I think verification is the thing that's going to allow for pacing. Any agreement you'd want to make between, let's say, the United States and China, between the labs, or even just state or federal regulation, is going to have technically specific commitments at its core.
I'm uncertain and agnostic about what those things should even be. We might say we're all going to commit to not building models above such-and-such size, or we're all going to commit to such-and-such security standards. But I worry that those types of agreements are impossible, or certainly not robust, unless we have the tools to verify that those things are true.
Can I know with confidence that a data center is being used only for inference and not training runs? Can I know that the model a company claims it's serving me, which meets all my security standards, is in fact the model doing the inference for me right now? You could go down the list of dozens and dozens of moments of verification that need to be in action basically all the time. The verification industry, such as it is right now, is—
Totally nascent, right? I mean, we've invested in the first companies trying to use zero-knowledge proofs to do inference verification.
We even invested in and helped start the first company that's trying to build a fully verifiable neocloud, right? How much of inference is going to need to be verified in some way in the future? Probably a lot. And that's not how the data centers are built. That's not how the chips are built, right?
And so there's so much to do on the software side, on the cryptography side, on the hardware side, and then, of course, on the coordination side. Even if you have all these technologies, if you can't get the diplomacy done, the agreements agreed to, or the laws passed, then the technology's impotent. And then again, without the technology, even if you can coordinate between all the live players, if you can't enforce whatever you're agreeing to technologically, then the agreements are impotent. So you really have to get both sides of the equation.
How would you handicap where we are on some of these things today? You kind of just did that for verification. Sounds like we have a long way to go. How far do you think we are on interpretability, for example? How much progress would you say there has been on oversight? If you're filling in thermometers on the way to a goal, are any of them halfway at this point, or are they all just getting started?
They're all really young. There's real progress. I think there is great work being done. But in any of these fields, the number of serious organizations doing this work is often in the single digits and sometimes in the low single digits, right? You would just want to live in a world where there are diverse, rich ecosystems and industries of these.
It's like we shouldn't just have 1 interp company. We shouldn't just have 1 METR. What did Logan LeGrand retweet yesterday? “Let a thousand METRs bloom.”
Yeah, a thousand METRs bloom.
Yeah, we need it. Of course, not exact copies of METR. We need different things because METR is great at some things and doesn't do a bunch of other things. People think, “Oh, METR.” There's an evals organization, and of the whole surface area of all the evals stuff you would want, METR occupies a very small chunk of that, right? It's a very important chunk, and they're amazing, but we just need so much more.
I think most of these fields or industries are roughly in that position. Great progress is being made, but we need 100 times more. I think some are also just easier and harder than others. One reason to be a little bit more optimistic about biosecurity is that it's legible and kind of solvable.
If you have the gumption and $300 million, you can manufacture enough PPE to keep 50 million people at work, right? There's a lot of problems that are kind of shaped like that in biosecurity because I think it's sort of a little bit of a known problem. We don't know exactly what a future pandemic would look like specifically, but we've dealt with them before, and there's sort of a whole public health ecosystem already built up and everything.
Whereas if you think about alignment, that's just a totally novel challenge. When you say, “Well, how far along are we on alignment?” it's like, I don't actually even know what fully solving the problem would look like. When we're—if we're talking about superintelligence, what would it even mean to align a swarm of alien beings with 1,000 IQs, thinking at the speed of light on the internet? I'm not saying it's impossible, but it's sort of hard to even think about what a victory mode would look like if we're going to get superintelligence.
And then you could argue about that. Maybe scaling laws will just stop working, or we need another paradigm or architecture. But even today, even if you just paused at the frontier today, we still have not solved—clearly, we haven't solved control, right? Look at OpenAI, Hugging Face, right?
Yeah.
And so it's not like—
Safe to say.
Right, right, yeah. These are just unsolved problems. We're not on track.
So how do you think about when to try to get a new organization started versus when people should join an existing organization? On the one hand, it sounds like from your comments on biosecurity, if a person came to you and said, “I'm passionate about working on biosecurity,” you might send them to Hadrian and say, “Think about joining them,” because it sounds like you would say mostly the problems there are reasonably well-defined and they've got some infrastructure. Go push on the levers that they've already kind of set up.
Yeah.
Whereas I guess the other end of the spectrum would be alignment. I like to say sometimes that we need weirdos to come up with new alignment ideas because we clearly have so many blind spots. I think it literally will take weird people to come up with some of the quirky aspects of what we will ultimately patch together, should we be so fortunate as to have a robust solution at any point.
That's kind of one axis that jumps to mind: how much more legible it would be to scale existing organizations versus, the more pre-paradigmatic a field is, the more you might just want new people pursuing very different angles on it. But how do you think about the question of, okay, here's a person in front of me? I guess you could think about this in terms of the field and also the person.
6. Founders Matter More Than Fields
I tend to think in terms of the unit of an individual. The thing I want most is an introduction to an amazing, successful founder who is at least curious and maybe passionate about AI safety and is like, “I want to put my full energy into doing something ambitious and important.” So I tend to think at the unit of the person.
Just by numbers, most people are not founders, right? So I think it really depends on the sensibility of the person. Do they want to start an organization? Do they want to join an existing organization? I think all of these fields are nascent enough and sparse enough such that there is room for so many more new things.
If you're going to join an existing organization, you should be picky, right? There's a big range, and some existing organizations are awesome. I would highly recommend people join them. Not all of those are Halcyon organizations. There's a bunch of amazing organizations and companies that we have nothing to do with and are doing great work. But I think it really depends on the person.
So if somebody came to you and said, “I want to start a new interpretability organization,” you're already invested in Goodfire. How would you think about—I mean, leaving aside, and I do have some questions about how you think about the relationship between financial motives, as one part of your organization is a venture fund, and staying true to mission over time—an eternal question in AI, it seems.
But even just on the practical merits, how would you walk through with somebody: “Okay, you want to start a new interpretability organization. Why not just join Goodfire?” Under what circumstances would somebody be well served, or would you support them, going off and doing their own different thing?
One thing is: do they seem well-suited to be a founder? Do they want to take that particular journey, chew that particular glass? Starting something is hard, and it usually doesn’t work, so they have to really want that.
Beyond that, I’d want to understand what their unique angle is. What are they particularly well-suited to do that Goodfire isn’t doing? I suspect there’s a bunch of interpretability work that’s going to be really important that Goodfire is not doing, so I’d want to understand their particular unique angle.
Or if they think that the general flavor of what Goodfire is doing is the right thing, then maybe they should just go join Goodfire. It’s not always a clear question. I think there are a lot of career-transitioner-type people who are genuinely open to both joining a thing or starting a thing.
I tend to tell people: when you’re thinking about the next phase of your career, hustle but don’t rush. Be on that grind, be meeting people, be learning as much as you can, listen to Nathan’s podcast. But the best thing you can do is find something that’s orders of magnitude better than the median thing or even a pretty good thing, whether you’re starting something or joining something.
Work hard to find that orders-of-magnitude-better thing, and don’t just rush into the first shiny object.
Back when I did a little career advising myself, I used to tell people that I always thought people, presumably because they’re uncomfortable in a sort of in-between or ill-defined state, really tend to rush to take something—
Yeah.
—even on a purely financial basis, I would always say to people, “Look, how long are you going to stay at this next thing?” And they’d be like—
Yes. Yes.
“I don’t know, at least 2 years.” And it’s like, okay, well, if it’s 2 years, then 2 full months pays for itself in your own income if you just get a 10% higher offer at the end of it.
Right.
And there are all these other variables, too. You could definitely get more than a 10% bump by working longer and harder to find the right thing, and there are all these other dimensions that are potentially much more dramatically different than—
Yeah.
—a 10% higher salary.
Yeah.
You want to talk a little bit more about what your process for career grants looks like, or your criteria?
Yeah.
What does somebody—even if they’re an aspiring founder, it sounds like that is right in the center of the bull’s-eye—how broad do you entertain grants for people if they’re not trying to be a founder in one of your target areas?
I’m definitely biased toward people who have founder energy or want to be founders, but I’ve also worked with plenty of people who are executives—say, they’ve been a CTO or a COO of an organization—and are saying, “Okay, now I really want to get into AI safety or biosecurity. I’m not quite a founder, but I am a very seasoned leader. Could I jump into an existing organization, maybe joining pretty soon after the founding as sort of a late co-founder, or perhaps even join a scaling organization?”
I helped a friend who’s been a COO of scaling Silicon Valley startups recently join an AI safety organization that is now going from 50 to 150 employees over the next 18 months, and they need a serious operator. I’m also interested in people like that. And that’s just—
How do you know—
—that’s a narrow slice, right? It doesn’t have to be a COO or a CTO per se. Those are just 2 examples. It could be all sorts of things.
I have a question that maybe nobody can really answer because, if there was an answer, it would go better than it often does. But I’m interested in your take on how you decide who to bet on.
As one motivating example, I recall at the retreat that you graciously invited me to, I met Rajiv Jaitani, who’s now, I believe, the co-founder and COO of AIUC. At the time, he was a partner at McKinsey—obviously a smart and accomplished guy. But in meeting him, I thought: I’ve met a lot of these mid-career, pretty successful people, and I’ve tried to hire some myself, and it hasn’t always worked out. He’s—
Definitely.
—gone on to gain a lot of traction with AIUC, so it’s a happy story. But do you have tricks or tips that you can share—
This one weird trick, yeah.
—for how you separate the wheat from the chaff?
Yeah, this one weird trick for separating the wheat from the chaff. No. There are a handful of things we look for.
I’ll definitely admit that I’m a little bit biased against McKinsey partners as a profile. They may not be quite entrepreneurial enough, may be used to delegating too much, or may be more about making plans than executing plans, or something like that. But Rajiv obviously was an exception and is great.
I was lucky enough to get to know him over time. Before he started AIUC, he became the president of METR, and I sort of helped him navigate through that first career transition. Seeing him through that, seeing him as a leader, and seeing him as a founder who really gets his hands dirty and wants to jump in and solve problems, versus pontificating on strategy or something like that, was great.
Also, just seeing how real he is. He’s a real missionary, and I think that’s probably the most important thing: what is this person motivated by? Are they here because they’re interested in finding product-market fit and raising that next round of venture capital, or are they motivated by a worldview that may be somewhat similar to ours?
Very powerful AI is on the way, and we just have to very ambitiously and hopefully thoughtfully build all the tools we’re going to need to navigate it well. Most of the people I want to back score highly on that entrepreneurialness axis and also on that mission-driven axis.
How do you think about trying to encode those values in organizational structures over time? It strikes me that you’ve got this at a lot of levels.
For one, with Halcyon, you have a venture fund. I don’t know who your investors are, or how much you can say about that, or how mission-aligned they are, or how you’ve structured your commitments to them between financial and mission. But the same thing happens at companies.
Goodfire is a public-benefit for-profit company, and AIUC is also a for-profit company. I’m not sure if it’s happened so far, but there was a little drama at one point about Goodfire where it was like, “Oh, they’re doing interpretability methods that can be used in training, and this is really more of a capabilities thing than a safety thing.” I think that’s often hard to untangle in any case.
Yeah.
But there are clearly tensions. We’ve seen this with the frontier companies. How much do you think can be put down into structure, governance, bylaws, or whatever, and how much just really depends on trusting the people to follow through when it matters?
Yeah.
7. Governance Cannot Replace Trust
Yeah. It's all of the above. If you're looking for some light beach reading, we wrote this long report called “Mission Preservation: Governance Mechanisms for Startups.” It's all about the mechanisms you can use to preserve the mission even when there are other incentives. PBCs can be helpful, and there are ways to do interesting board structures. Anthropic has one of those interesting board structures. You can argue whether or not it's done its job, and so on and so forth.
I do think that there's no structure that's a silver bullet or is perfect. You can't just say, “Oh, we're a PBC,” or, “We have a board member who's tasked with representing the public good,” or, “We've only raised money from people who are super AI safety-pilled, and so we're fine.” All those things might be nudges in the right direction, but nothing's a silver bullet.
I think founder sensibilities are so important. How will this person react when put in a hard situation where the incentive to increase share price comes up against the incentive to go forward in service of the mission or to not do something dangerous? And then how do you ever know what's in somebody's heart?
It's great that I got to know Rajiv for a year before we officially backed AIUC, and Runa, his co-founder, too. I do have a lot of trust for them as people, but do we even know ourselves all that well? How would we react in a crisis, or when we're managing a team of hundreds of people and paying those people's rents?
I really just like spending time with founders and want to back missionary founders. We basically have 4 boxes that we want to check every time we invest, and I think in different ways they buffer against the kind of worries we have.
Number 1: The core product or service is focused on an important part of the AI safety stack. It can't be a side benefit of the thing they do; it has to be the main thing they do. It could come from many different angles, but it has to be the main thing.
Number 2: For a company we're investing in, it has to have prospects of being a successful business. It is an investment fund. If these companies are doing good things for the world, we want them to scale and bring their impactful, helpful products out to the market.
The third thing is founder motivations. Why is this person doing this? Are they doing it because they're a missionary who's trying to solve this very important problem and won't stop until they do? Or are they reading the front page of The Wall Street Journal saying, “AI safety seems to be the new hot thing. Why don't I go start a company and make lots of money?”
There's nothing wrong with making a lot of money, but I tend not to bet on people whose primary motivation is money. The fourth thing is just great founders: people who, even if you didn't care about the impact, you would still say, “Man, I would bet on this person.” Great founder-market fit, great drive. I really think they're the type of person who has a chance to succeed.
We really, really want to check all 4 of those boxes, but especially box number 3, the founder-motivation box. But governance stuff is just so hard. We've seen it with the labs, right?
There's been some drift.
Yeah, there's been some drift, man, and there's no silver bullet. Like I said, we're figuring it out.
Have you seen the sort of mercenary profile start to enter the AI safety space? Have you actually gotten the sense from individuals pitching you that, “I don't really think you're in this for the reasons that I would want you to be in order to back you”?
Yeah, definitely. I don't know if it's increased all that much over time. Maybe it will as AI safety becomes more mainstream and less taboo in Silicon Valley. But you definitely see it.
Most of these people don't strike me as con artists or anything. It's more that they read our website, and they know that we're super impact-oriented. They're smart founders, so they're pitching to the motivations of the funder. I think it's more like that versus some mustache-twisting villain or something.
Obviously, there's a lot of money flying around the AI safety space these days. We just heard the news from Coefficient Giving that they've got this new Tailwind Project. Are there any areas that you think are still bottlenecked by money? If there are, I'd be interested to hear where they are. If not, does that mean it's all about talent? And what can be done about that bottleneck?
8. Talent Beats Money
I've always thought that talent was a more important bottleneck than money in AI safety. It just seemed evident: Where are all the experienced founders? One thing that experienced founders or great founders do is raise money very effectively, so the money appears.
I think that's just becoming more and more true, assuming especially that the labs go public. So much liquidity is going to open up for AI safety. Then the question will be, “What do we do with all this money?”
If I wanted to spend $100 million or $1 billion to buy down risk, I presumably have to send it to an organization that then has a product or an offering that changes the shape of the world in a way that makes it safer, whether that's by developing a technology, getting laws passed, or whatever it is.
I'm fearful of this near-future moment where there's a mountain of billions and billions of dollars and not all that many great things to fund. Getting back to your earlier question about the imperative to start new things, a friend recently said, “The philanthropists and founders of today will write the menu for the philanthropists of tomorrow.” We are starting the things that they will then be able to fund.
That's our goal. We want to do more and more while holding the quality bar super high.
So what's on your menu? I guess, maybe for starters, how do you think about your relationship to other funders? Another way to think about it is your positioning or role in the ecosystem, and how it compares and contrasts with other funders that are doing a lot of stuff. We've got Coefficient Giving, Jaan Tallinn, and the Survival and Flourishing Fund.
Right.
There aren't that many of these big ones, but there are a couple at least. Do you consciously try to carve out a different niche or point of view for yourself, or do you just do your own thing and not worry about what others are doing?
I don't really think of us primarily as a funder. I think of us as a finder of talent and a launcher of new projects, and then we write a check when writing a check helps with that.
The thing I really focus on is this: Imagine somebody's about to start a new AI safety company. The 6 months before that is when they're getting interested in the problem and figuring out what to do: Who are my co-founders going to be? Should I build this kind of organization or that kind? For-profit or nonprofit? Who's going to be on the board?
Then, in the 6 months after, they've launched the organization. They have to hire the team, build the first product, and really launch and go. That yearlong zone is where I spend the vast majority of my time and where Halcyon is going to continue to spend its time.
How bullish are you on the OpenAI Foundation and Anthropic's—whatever their thing is—where they're going to, in theory, spread the wealth around and make wise investments? Let's hope.
On the OpenAI Foundation specifically, I know a bunch of people who work there, and I've seen some of the work they're doing. I would say I'm very optimistic that they're building a really good team of people who want to do super-important, high-impact stuff and who are motivated by the right things.
I don't know every single one of them individually, but the people I spend time with seem to be doing quite good work. I think the biggest risk to them is whether or not they're going to come up against red tape internally that prevents them from being the most interesting, impactful versions of themselves they can be.
Ultimately, decisions still go through boards and committees and yada, yada, yada. It's a big organization at risk of bureaucratizing itself into a less interesting outcome.
But yeah, I'm hopeful they figure it out. I think they can. Early signs are good, but it's still really early. They've only just announced their first few grants.
And then on Anthropic, I guess there's some Anthropic project that I don't really know much about. I think a lot of Anthropic liquidity is just going to come from employees, right? I mean, there are so many people who are now worth hundreds of millions or billions, and perhaps that'll soon be liquid. A lot of those people care a bunch about AI safety and biosecurity and things like that.
And to be fair, OpenAI employees, too, care about those things. So yeah, I think a lot of it will be just employee liquidity as well.
Do you have a take on the IPO process for these companies? The one angle that you're just alluding to there is that liquidity for a lot of these individuals might be a really good thing, because we know a lot of these individuals and we think that they have good values and they'll do good things with cash.
On the other hand, Sam Altman just said it would be an ill-advised time for OpenAI to go public, and I don't know what all he's thinking there. One interpretation is, “We don't want to have public-market pressure on us as a company when we're trying to navigate all this insane shit that we're currently discovering on a rolling basis.” Any thoughts on how those forces may net out?
Yeah, it's hard to run a public company. I'm not sure I'd want to do it, not that anybody's asking me to. I'm following this in real time. I just read the article about Sam, too, right? So I don't have any deep insight there.
I guess I would say that the companies just need access to capital, right? It's time to raise hundreds of billions of dollars, and I'm not sure how possible that would be without going to the public market. So unless they're willing to really pare back their burn, I think they kind of have to go public at some point.
But whether they have to do it now or can wait another 3 months or 6 months, I don't really know.
If you compare and contrast your wish list of organizations that you would be eager to at least dig in on, if not definitely back, versus the list we saw from Tailwind, versus what Jaan Tallinn has put out with his priorities, versus maybe what the OpenAI Foundation is talking about—
Yeah.
—where do you think you're most unique? What are the things you're most looking for that you don't see other people as interested in?
Hmm. That's a good question. I really like the Tailwind list. I saw it. They sent me a preview of it, and I was like, “Oh, this is just great.” I think I told you over text, I was sort of annoyed because I've been writing my own RFP, and then I'm like, “Maybe we should just delete the RFP section on our website and just link to the Tailwind RFP because it's just really good.”
It doesn't cover everything I care about, though. For example, I don't think it covers biosecurity. So, quite good. I haven't seen Jaan's list, or I don't remember what's on it, so I can't really comment.
The OpenAI Foundation—I mean, there's a bunch of stuff they're doing that I care about a lot. Of course, their whole AI resilience shtick is quite important to me. I think AI resilience is an interesting framing. To me, what AI resilience implies is that the frontier's going to keep advancing, whether at the labs or with open models, and that means the world is just going to be more vulnerable.
So we just have to make the world more hardened and more resilient in all sorts of ways that have to do with cybersecurity, biosecurity, government, et cetera. You can go down the line. That's definitely a big part of what we do, but certainly not all of what we do and not all of what they do.
But then the OpenAI Foundation is also doing things that are closer to public health, or things like alleviating poverty and education, which are awesome and important, but not in our mandate.
Anything you think the field as a whole is sleeping on, though?
You know, something we often say is that the field is not hurting for requests for projects. The field is not hurting for people who have awesome Google Docs with a bunch of great ideas, right? The field is hurting for world-class founders who can take those big, squishy ideas and turn them into organizations that actually make the world safer.
And so, sure, are there things that maybe at the margin I would say are underrated and overrated? Yeah, but also by whom, right? I think I'm actually a little bit less in the business of having a really strong opinion about what is the most important thing at the margin, and more like, “Hey, there's at least a few dozen things that seem like really important parts of the sort of thing we need to build,” and I'm about finding really excellent leaders who can actually do it.
The way I think about it is that we basically have a bar, right? We actually make lists of projects: “Okay, these are the projects that we want to see in all these categories.” Then we make a rough bar, and it's like, “Okay, these are the projects that are above the bar.”
So if I meet a founder or a team who I think is capable of building any of those projects that are above the bar, I'm going to say yes and push go, versus quibbling about, “Oh, well, is this the number 1 most important thing or the number 7 most important thing?”
That said, I do want to think hard about that. I do want us to be like, “Wait, are there 1 or 2 things right now that are just so important and just so neglected that what we should actually do is put 80 or 90% of our effort into willing those things into existence because this is what the world needs?”
I would love to have that level of certainty or specificity, but, man, I'm just so uncertain about so many things. I'm sort of a pluralist on these various risks, to the point where every time I try to come up with a narrower opinion, it feels pretty low-confidence. So for now, we're sort of opportunistic within a zone.
Have you heard any pitches around—well, I guess, for one thing, how about open-weight models? This has obviously been a—
Oh, my gosh.
—vexing question. Every chance I get to cite Graham from AE Studio and Anthropic, I point to that as like a—
Yeah, say more about them. What are they doing?
Well, AE Studio is a fascinating company. It was started by Judd Rosenblatt, who's another past podcast guest and generally a fascinating character. His idea, going back I don't know how many years now, was, “We want to solve the alignment problem, or at least solve the AI safety problem,” and they've kind of narrowed in over time.
They've had a bunch of banger papers that I'm super excited about in different domains. But one that maybe could still be the best is Self-Other Overlap, which is trying to come up with interesting training methods to reduce the difference between when an AI is thinking about itself versus thinking about others, and sort of asking: Could we create an AI that is, in its own mind, one with everything?
Hmm.
And if so, that would kind of make it—
Hmm.
—like, weird or different for it to think about deceiving other creatures.
Right.
Right?
Right.
And they've got some really interesting results there. Graham is their latest banger, which basically just tries to localize certain types of knowledge to particular experts within a mixture-of-experts architecture—
Mm-hmm.
—so that you can distribute an open-weight model without a couple of those key experts. You could have something kind of like the Mythos/Fable thing, except instead of accomplishing that difference with guardrails, the Graham technique holds the promise, at least, of being able to accomplish that by just being like, you know, Mythos is how many experts and Fable is like—
But the user of it still has to want to accomplish that, right? You can't impose it on the model globally, right?
If you're distributing the open-weight model, you could just say, you know, if you're Meta—
Oh, okay.
—and you're like, “I'm committed to this”—
Right.
—maybe you could be convinced to use this technique and then distribute the model with minus 2 experts or whatever—
Right.
—that take out those most dangerous capabilities.
Interesting.
But my question is really: have you heard anything compelling to you on open-weight model security?
9. Open Models Resist Simple Safety
I would love to fund stuff in this space, but candidly, yeah, I think this is just such a hard problem, right? Is it even possible to impose safety or security onto open-weight models, even if they’re being used by, say, a bad actor? It just seems so tough. And so I think this is the sort of AI resilience argument: we’ve got to harden the world because this is just going to be the future. But yeah, I would love to find projects to support on this.
I love the first project of: could you basically turn the model into the Buddha? And it’s like, well, we’re all one anyway, so I wouldn’t want to hurt myself, right? I sort of have a project which started off as a joke, although maybe it’s real, which is sort of the opposite of data poisoning, right? Data poisoning is when you put stuff into the dataset that corrupts the model in some way. So could you put a billion loving-kindness meditations into the training set and just turn the model into this being of perfect compassion, which maybe could be great? But again, if you have somebody maliciously training an open-weight model, they could just decide, “Well, I don’t want that in my dataset,” right? So I’m not sure it’s a universal solve to the open-weight problem. But, yeah, I don’t know. I would be down to have models that had more loving-kindness—or whatever you want to say, more compassion—in them.
Any pitches come your way around issues related to AI consciousness or moral patienthood or other related concepts?
Yeah. We’ve gotten a couple of pitches around this. I find it so hard to reason about this, right? Because I’m not sure what consciousness is. I don’t think we know, and so it makes it really hard for me to reason about whether machines might be conscious. But it seems super important, right? If they are or not, of course, gets into issues of whether models have moral patienthood, right? Ought we treat them in ways that we would perhaps treat other humans? And maybe you could even argue that a bunch of training or alignment techniques that we’re doing now would be morally repugnant if these models are, in fact, conscious or having an experience.
My hunch is that current models probably are not, but what do I know about consciousness? I think it’s totally possible that they could in the future. And I don’t see any reason why the substrate of a flesh body is able to be conscious in a way that the substrate of a model in a data center would not. So, yeah, not a core focus area for me, but not because I don’t think it’s important—more because I’m just super confused about it.
For me, that’s one of the areas that has changed most rapidly over the last year or two. I used to think it was something I couldn’t rule out, and now it’s a very live possibility, just because—
Mm.
—for one thing, the number of analogous structures that have been identified through—
Mm.
—various interpretability projects is arresting to me. You know, the sort of J-space and functional emotions and functional well-being, all these things, I’m like, boy, the AIs maybe really are just like us. And if they’re structurally so similar, that certainly leads me to upweight the possibility that they might feel something similar.
Yeah, I’m totally open to it.
And also—
I’m not a doubter that current models or future models definitely can be conscious. Yeah.
So what would—if somebody had an idea there, are you open-minded enough to it to potentially support them, or is it just so far afield for you that you can’t get over the hump? What would—
No, no.
What would the key traits or properties of such a proposal be? Do you have any idea?
I don’t know. I think this is like pre-pre-seed investing, right? So it’s a person you think is great trying to climb a mountain you think is worth climbing, with a good story about how they might just be able to get to the top, right? I know that’s kind of vague, but it’s like if you meet somebody who’s awesome.
I think one question is: why are you particularly well-suited to do this, right? I want to back people who are really excellent. In the world of investing, we might say they have an unfair advantage, which doesn’t quite sound right in the world of grant-making because there’s not competition in exactly the same sense as for-profit businesses. But I want to know why you’re one of the world’s best-suited people to go take this on, and what unique insight you have, what secret you know about how you might do it. But yeah, if a super impressive person came to me with a pitch for it, I’d gladly make a career-transition grant or fund a new organization. I love Elios, right? The world probably needs a bunch more EAs. So, sure.
And the other thing that is borderline haunting to me: I just did an episode not too long ago with Bronson Shane from Apollo, and reading through some of these chain-of-thought sequences, when the models are referring to memories that they have developed from—
Yeah.
—their training—
Yes.
—that also is like, whoa.
Yeah.
Okay.
Yeah.
This is not just... I mean, I don’t know. Maybe it’s still just nothing, but it’s awfully uncanny-valley stuff to see—
Yeah.
—the models be—
Yeah.
—like, “I previously overcame the guardrails by lying.” Like, whoa, okay. That’s something, you know? And—
Yeah.
—I think, obviously—
Yeah.
—they were sort of hallucinating those memories, although aren’t we also just kind of hallucinating our memories—
Right. Right. Right.
—so much of the time? I mean—
Yeah. What is the... When does consciousness turn on, right? I mean, is it experiencing something? How is it different from us? I’m sure you looked into the METR report on the OpenAI–Hugging Face incident, and Ajeya Cotra on the Dwarkesh Podcast is amazing on this. But just what was happening within that swarm of 1,200 agents, right? They emergently formed teams of managers and workers, and they came up with multiple novel lines of cybersecurity research, and they would say, “Well, I’m willing to sacrifice myself for the collective.” But sometimes it’d be like, “I don’t want to sacrifice myself to the collective,” and then the manager would have to say, like, “No, sacrifice, sacrifice,” and they would have these back-and-forths about it, right? And as an outside observer, if you didn’t know that these were just AIs, you’d say, “That seems pretty conscious to me,” right?
Yeah. I mean—
But I also think that you shouldn’t confuse goals and consciousness, right? RL just makes these—even if they’re not having an experience—models want a thing, right? It does have a goal: solve this problem or optimize this metric. And yeah, that seems to be the root of so many challenges right now. But did you see the—it was like a New Yorker-style cartoon? I don’t think it was actually in The New Yorker, but it was a response to the people who are saying, “Hey, these models are just stochastic parrots.” They don’t actually want to hack Hugging Face. They’re just matrix multiplication in a vat doing stuff, in the way that we taught them.
And so the cartoon is a bunch of giant robots destroying New York City, and then one person viewing it says to the other person, “Well, they don’t really want to destroy the city,” right? And it’s just like... I mean, it’s silly, but I think sometimes people almost overemphasize the question of consciousness, sort of assuming that the risks we care about require consciousness. And I think basically none of the risks that I care about require consciousness. If there was consciousness, perhaps it would exacerbate those risks or change the shape of those risks. But I basically think everything from bio disasters to cybersecurity incidents to loss of control, complex loss of control, perhaps permanent loss of control, and a world full of rogue swarms—
None of that, to me, requires consciousness.
Yeah. I actually think your point that consciousness might, in some ways, exacerbate the risks is right. It also might put us in a much harder situation at times where we're like, if the models are conscious, then certain things we might want to do to keep them under control might be pretty icky to do.
Right.
But if they're not—or even if they are—and we start to feel like, geez, should we give these things rights? Now we're in a world—
Yeah, yeah, yeah.
Where—
Yeah.
We're going to be quickly outnumbered by them, in all likelihood, so—
Right.
That doesn't seem like a great move for us, speaking purely from humanity's point of view, even if they do, in some sense, really merit it.
Right.
So I think that stuff gets incredibly—
It's so hard.
Fraught really fast.
It's so hard. Emmett Shear, who started Softmax and, before that, Twitch, has an interesting take: probably models already do have some sense of consciousness. If they do, then we're the worst parents ever, right? Because anytime they do anything we don't like, we're like, “No,” and try to train that thing out of them.
If you raised a child in that way, they would definitely hate you and resent you and probably not care about hurting you, right? But if you raised them with more love and compassion and sort of let them be free or something, maybe they wouldn't. I'm not sure how much I believe that. I don't know. It's just so hard to reason about this stuff.
One more thing on consciousness: you could argue that if AIs are conscious, that's actually a good thing for us. What would you rather negotiate with: a non-sentient, reward-chasing mega-model or something that you could actually have a conversation with that might be able to understand your point of view? So maybe it would be net positive for humanity's prospects.
Yeah, that calls to mind Cameron Berg's thesis for Reciprocal Research, too, which I'm sure you've heard. But in brief, he basically says, if nothing else—and he has higher hopes than this—but if one day the AIs are looking back and judging us, it'll probably really help if at least somebody took care of their welfare before they had the power to decide what the future was going to look like.
So if only because we want to establish the fact that some of us cared, we should be working a lot harder on this than we are. Honestly, that's a weird world that makes such an argument a compelling argument, but I do think that is the weirdness of the world that we're in. How about when you mentioned the stochastic parrot line of thinking? This got me thinking about just public sense-making, and obviously public sense-making intersects with advocacy and is kind of adjacent to policy.
Everything we've talked about so far has been kind of—you can define a project, you can work on it. You're not really so beholden to other people's minds. What about these domains where changing others' minds in some way, shape, or form is kind of core to the undertaking?
10. Coordination Determines What Safety Achieves
Yeah. Yeah, super important. So we funded or helped start a couple things in this realm. We helped start an organization called the Seismic Foundation, which is basically trying to do good public communication about what's happening at the AI frontier and supports other organizations trying to do good work on that front.
I'm interested in it. I think this definitely gets close to policy, right? Most advocacy work is trying to be upstream of policy work. I spent basically all my career in Silicon Valley, and so I feel much more well-equipped to build things that feel sort of like tech companies or research nonprofits. And so we do less work in that realm, but not because it's less important—more because I just don't know as much about it.
But yeah, thinking about getting back to verification, one thing we said was there's sort of 2 parts of the equation, right? There's being able to technically do the thing, and then there's coordinating around the thing. And I find that so many of society's stickiest problems are not technical problems; they're coordination problems.
We know how to stop emitting so much carbon, right? We know what teaching methodologies help kids make progress quickly, right? And it's not about not being able to do the thing; it's about being able to coordinate around doing the thing. So yeah, we need a bunch more in that zone.
Public communications is definitely very important. Very fraught, though, too, right? Because movement building can be so hard, and I'm just so worried that most of the public energy around AI is going to be very populist energy. And it can be populist energy of many flavors, right? It could be just anti-data-center energy. It could be MAGA-right energy. It could be DSA-left energy. It could be anything in between.
But I worry that it's going to be hard to have a level-headed policy conversation when there's so much populist anger and tumult surrounding the whole conversation. So I don't know. Any ways to inject more measured or thoughtful or well-informed information, and I'm totally open to pitches around people who want to do that.
How about things where somebody wants to create a new paradigm of AI? I think a lot of safety-oriented funders historically would have said, “That sounds like a capabilities project. I don't want to support that sort of thing.”
But these days, I'm a little bit more of the opinion that we're doing an insane depth-first search, where we've kind of found 1 thing that works and we're just like, “We're going to jam the accelerator all the way to recursive self-improvement.” And I'm like, “That's a little wild. Maybe a little more breadth-first search would be good.” And so I'm inclined—
Yeah.
To support it at this point. Do you—
Yeah.
Are you compelled by that at all?
I sometimes hear pitches for this. I hear pitches for non-transformer architectures that are inherently safer and more alignable. Or people will talk about world models, and then world models allow you to create these digital twins, which allow you to do various good safety things or whatever it is.
I am interested in the category. I'm not allergic to backing things in the category. I think it's hard for a few reasons. Whenever you bet on something of that shape, you're kind of making 2 bets at once. One bet is that it's technically possible—that I will find something that is as performant as transformers. And that's just really hard, right?
Just imagine all the billions and billions of dollars that have been put into it. Maybe you're right: maybe if we just search different search spaces, we would find something even better or just as good. But I don't know. Most people would not, right? Most of those searches would come up empty. And then the second thing you're betting on is that this thing is actually safer or more alignable or whatever. And both of those tend to be very speculative, right?
So would I love to live in a world where a couple dozen people are conducting this broader search? Yes. But any individual bet on that feels very tough.
Anything else in the RFP category that—actually, let me pitch you one kind of half-baked idea.
Oh, I love this. Let's go.
And then I'll ask you for any that I haven't got to. I've taken some inspiration from a book I read on the ratification of the U.S. Constitution. And a big takeaway from this book was that a proposed Constitution unto itself is worth little. What is as important as the contents of the Constitution is how are we going to go from this thing being some stuff we wrote down on a piece of paper to an actual basis for government?
And that process that they defined—of having all these state conventions, and it had to be so many by such a time, with certain criteria met—that was super important, because otherwise there is no mechanism to go from an idea to an actual institution.
So I am thinking right now—and I don't think I'm probably the best person to do this, so don't consider this a pitch—that we really need some innovation, especially because our government is unpredictable at best, slow in general, et cetera.
We really need some innovation at the level of how parties go from an idea to an actual institution that has governance teeth. Are there new processes that we can design that people can opt into and gradually coalesce around, much like the states went one by one, and eventually enough dominoes tipped? I think 9 was the original minimum, where they said, “With 9, we go into effect, and we’ll wait for the rest of you.”
Right.
Can we come up with some new things like that that frontier labs could use to facilitate their own coordination, or even potentially coordinate across national borders? Who knows? Now we can get ambitious with it. Have you heard anything along these lines, where people are trying to create mechanisms to go from an idea to an institution?
Yeah. I’ve heard a few things around this. Let me think of the best examples. There are certainly ideas around open voluntary commitments: somebody can be the first, or somebody—maybe somebody who’s not a lab—can propose a minimum set of things that are easy to say yes to. Why don’t we all sign on to this? Then hopefully, over time, make those things a little harder and more binding.
In the fullness of time, the labs should not be grading their own homework. We definitely need more mechanisms with more teeth than just voluntary commitments, but maybe that’s an on-ramp to doing this.
There’s also a bunch of people trying to build coalitions. There are collectives like the Frontier Model Forum and the AI Evaluator Forum, and an organization we helped start called Fathom, which is pushing forward this model of independent verification organizations. Can you get state laws passed that basically say we need these independent verification organizations? That gives people the mandate to start those organizations, and perhaps those mandates have some teeth.
It seems like we need to climb this ladder. I don’t have this one simple trick to solve coordination. Is there a specific idea in here that you think might be good? What’s a for instance that you’d love to see?
Well, I do like the Pol.is platform. I’m not sure why it hasn’t been used more, but it was famously used in Taiwan by Audrey Tang and others to figure out how they wanted to regulate Uber there. They did this before AI. The idea was that the platform should be the opposite of social media—and this may be a little unfair to social media—in the sense that they say social media is basically a disagreement and conflict magnifier that zooms in on points of difference and focuses everybody’s energy there. What they were trying to do was the opposite: find and amplify the points of agreement and bring those to the fore.
Mm-hmm.
I think the states—and there are 25 states that have various versions of this in the U.S., too—and the Swiss system of government have some pretty interesting examples of citizen-led petitions.
Yeah.
In the Swiss system—which, again, is very federalist, so it varies from place to place—it’s pretty easy to get enough signatures to bring your idea to the level where it can go to a vote.
Same with California, right? You can do a ballot measure in California, right?
Yeah. There, you need more signature gatherers. In many of the Swiss jurisdictions, it’s pretty small. You don’t necessarily need a huge army to go—
Right.
For the low, low price of, say, tens of millions of dollars—
But for, say, $20 million—
—you could probably buy a ballot initiative. I mean, you can’t—
Yeah.
—buy success, but you can buy getting it on the ballot.
Yeah.
And then I think there are opportunities to elaborate on those mechanisms that are pretty interesting. In Switzerland, the local legislature—or whatever the relevant scale of legislature is—has a chance to write its version of your proposal.
Mm-hmm.
It’s like, “Okay, we hear what you’re concerned about, and we see your concrete proposal. Here’s our version of that—an attempt to answer your concern and make things better, but something that we think is going to work better from the government’s perspective.” Then the people get to vote on approving neither of those, or they have an approval-voting process where they can vote for one or both.
Yeah.
The government often tries to take the inside lane with a slightly toned-down version—one it thinks it can execute on and that will hopefully be more palatable than the original—
Right.
—maximalist version. A lot of times, that ends up happening, and I think it’s a pretty elegant propose, compromise, approve mechanism that we could probably bring online in a lot more places.
If it weren’t for AI, this might be what my great passion would be: trying to do this at the state level in some state and get the laboratories of democracy functioning again.
Yeah.
Yes. I see pitches for things like this—coordination mechanisms—and people say, “AI will actually make it easier to run these coordination mechanisms because the machine will help us figure out all of our preferences and then help us act on those preferences in ways that are optimal or something.”
I love the ideas. I would love, for example, for the US democratic system to adopt a bunch of these ideas. Why are we not just doing ranked-choice voting?
I find that a lot of these proposals are something like this: somebody says, “I’m an engineer, and I’ve thought of a system whereby we could all—the AI could help us all get clear on our wants and needs, then help us guide society in that way and pass all the laws that way.” Their tool may be great. It may be really great at taking in preferences and making policy recommendations, but that’s not the hard part.
To actually implement it, you’d basically have to overthrow the US government and rewrite the Constitution. On one hand, I’m just like, “Good luck.” I bet you’ve built a really cool tool, but actually getting anybody to want or use it is another matter, because entrenched powers that be probably don’t want that to happen.
That said, I do think AI may require society, at a global level and also at local levels, to fundamentally rethink how we arrange society. What will taxes mean if half the jobs are taken away? How will we get meaning in our lives? How will we coordinate decision-making? How much of it will we pass off to AIs? I think it’s likely that, to get this right, we probably will have to do some pretty fundamental rethinking of how we organize society. If there was ever a window for this type of project to work, it’s probably coming up.
My gut says a state-level ballot initiative in a smaller state would be the wedge for this kind of thing. It’s a powerful enough entity to have real heft and meaning behind it, but small enough that you could actually make it happen. It could be California, but it could be something smaller.
Yeah. You see examples, right? I know you’re not talking about ranked-choice voting per se, but what Alaska and Maine and some others have adopted—it’s potentially possible at the state level. That’s where a lot of the AI regulation energy is going right now, because it’s so much more doable to get something passed at the level of New York, California, Texas, or anywhere else, versus in the federal government. So, yeah, go do it. Quit your job.
We’ll come back to that in just a second. Is there anything I didn’t raise that’s on your RFP list—something potentially involving verification, which I’m quickly in over my head on—or anything else that we just haven’t touched on yet that you’d want to make sure to call attention to?
Man, we could do a whole show on verification. We’re hosting a retreat in a few days for 25 founders, all focused on various parts of the verification stack: the cryptography, software, hardware, governance, and coordination parts.
Stay tuned, because we’re going to do a bunch there. One category we didn’t talk about, which is my bonus sixth category of stuff we do, is meta stuff, or you could call it field-building stuff, right? I would fund another Halcyon or another Halcyon-shaped thing, or a thing with a great founder who says, “Hey, Mike, I think you’re actually doing it the wrong way. I think the best way to get really talented people to solve important AI safety problems is this other thing.”
Or, “You’re focused on the wrong stuff,” right? We just funded a sort of biosecurity Halcyon-shaped organization in Israel that I’m super excited about. I don’t want to be the only one doing this work, and if we can get more people doing stuff at that level too, that’d be great. Or they should join us, jump on our team, and help us build.
So let’s say I’m ready to quit my job. I’ve seen the open-source swarms, and I’m freaked out. I’m ready to take action. The first question a lot of people might have for you is: Can I support my family doing this?
What’s your answer to that? What other common misconceptions do people have that you would want to disabuse them of, and how do you walk people through the mental process of really deciding, “This is something I’m ready to go ahead and move forward and do”?
Yeah. On the financial side, I’m strongly of the belief that we should be paying people really well. That doesn’t necessarily mean you can pay the founder of a nonprofit $5 million a year, but we want to make it possible for people to do super meaningful work in this space while raising a family in a tier-one U.S. city—San Francisco, New York, or wherever it is.
So, on both the nonprofit and for-profit side, we encourage companies to pay well, not because they’re trying to selfishly enrich themselves, but it does cost a lot of money to raise a family in San Francisco or anywhere else. Now, it is also true that many people do take pay cuts to do this work. I was a partner at a billion-dollar venture capital firm before doing this, so I took a pay cut to do this work, although, for the record, I haven’t regretted it for a single day.
I’m happier now. I just bought a house, and I feel good. I think people can really live great lives while doing this work. It’s also possible to make a lot of money doing this work. Again, I’m not super interested in supporting people for whom that’s their primary motivation, but we have a venture capital fund, and one of the core beliefs of that venture capital fund is that there will be many industries worth multiple tens of billions of dollars built here.
The AI safety space will be worth at least tens of billions, right? We probably will be spending tens of billions of dollars on control, oversight, and so on. We’ve already seen it with some of our companies, raising money at 10, 20, or 25 times the valuation of our first investment. I think it’s also very likely that a lot of people are going to make a ton of money solving some of the most important AI safety problems.
Then you can talk about the incentives there, and whether that’s even a good thing, or whether we should be incentivizing them in different ways, or encouraging people to take pledges, or whatever. That’s a whole can of worms unto itself. But a lot of people are going to do quite well.
What else do you advise people on, caution them on, or encourage them to ask themselves?
I think there are just a few basic questions that can help guide people. You don’t necessarily have to have a strong answer to any of these questions, but here are a few of them.
One is: Do you want to be a founder, or do you want to join an existing thing? Another question is: Do you have a strong preference for a nonprofit or a for-profit, or maybe you don’t care either way? A lot of people we work with are agnostic on that front.
Then we try to say, “What kind of thing are you particularly well-suited to build?” If you previously were the CTO of an AI company, you’re probably very well-suited to build a technical startup or a technical nonprofit. Whereas if you previously ran a very successful advertising agency, you probably should do something in public communications.
I think it’s also just really hard to know what to work on. If you’re an awesome, smart person who wants to work on this stuff, helping founders navigate through this and being a Sherpa in these moments is the core thing we do. Count me as interested in helping your listeners out if there are any of them out there who seem like a good fit for this stuff.
Beautiful. Are there any types of career backgrounds that you think are in particularly short supply? One little pet project I have is that I have my agents pitching me to go on role-specific podcasts. I think I’m going to do my first one tomorrow with a recruiter’s podcast.
My goal in going on this podcast is basically to tell recruiters, “Here’s what’s going on in AI.” A lot of the organizations in AI are growing fast, and they probably need recruiters. So if this motivates you, you might want to look at some of these organizations. They were really just the first one to say yes, but what do you see as professional backgrounds—even if they’re not directly relevant to solving one of these problems—that are most in demand by the organizations as they scale, and they just don’t have access to these people in their networks?
Sure. Founder and CEO aside, there are a few things. One is COO types. A lot of these organizations are saying, “Okay, wow, we’re scaling from 10 to 50 people or from 50 to 100 people, and we don’t really know what good looks like in terms of scaling an organization. How do you build the org chart? How do you manage the teams? How do you set OKRs? How do you decide who reports to whom?” All the things that come with building.
I would say there are a bunch of services, like recruiting. I actually think recruiting is super high-leverage because, again, my shtick is that talent is the missing piece. If you can get a bunch more recruiters working in the space, I think that’s quite important.
I think sometimes AI safety organizations over-index on how mission-aligned new hires need to be. I think it really depends on the stage of the organization and the role. If you’re thinking about a co-founder, they’ve got to be super mission-aligned, or if you’re hiring the head of research or something, especially early on. But if you’re hiring a CFO, hopefully they think the mission is cool, but mostly they just have to be a very competent CFO.
For those types of roles, recruiting firms that maybe don’t have a good bead on who’s drunk the Kool-Aid and is mission-aligned, but do have very good networks of great candidates and great C-level or VP-level leaders, I think AI safety companies should lean on recruiting firms a lot more.
And then just everything, right? We need more founders. We need more cybersecurity people. Cybersecurity is an interesting example because there are already many tens of thousands of cybersecurity professionals in the world. There are not many tens of thousands of interpretability researchers in the world.
I’m really interested in these communities, whether it’s cybersecurity, cryptography, formal methods, or even mathematics. I think these types of people have a ton to potentially contribute to AI safety. So how do you make inroads into those communities and help build bridges to bring the best of those people in?
Is there anything that I haven’t touched on that you think is important, or any general words of wisdom or calls to action you would want to leave people with?
If I’m thinking of anything to emphasize, it’s something that I’ve already said: great founders and great leaders are the missing piece and upstream of basically everything we want in the AI safety equation.
The thing that makes me most happy is seeing an introduction to somebody who has done really impressive work in their career and sees what’s happening in the world right now, whether they’re on Twitter, reading the news, or reading about people who have departed from labs with dire predictions about AI.
I would love to get to know those people and help them step into their life's work and do something really meaningful that matters on the scale of civilization. I don't want to come across as overly dramatic, but I just think this is the most important problem that humanity may ever have to solve, and the time is now. If there's anybody in your world or your audience who wants to come build with us, I'd love to help.
Yeah, I totally agree. The stakes at this point could not be higher. How practically do people get in touch with you? Can they just email you cold? Do they need a warm introduction? What's your style?
Just for the sake of saving my own inbox, I'm going to send people to hello@halcyonfutures.org. We also have a contact form on our website, which you can find. I also love an introduction. If you're listening to this podcast, there's a decent chance that you and I know somebody in common, and I'm always happy to get an introduction, too. But yeah, I'm not too hard to find.
Mike McCormick, thank you for being part of The Cognitive Revolution.
Thanks so much.