[BidClub_]
The Cognitive Revolution · · 119 分钟

AI AMA——第2部分:AI乌托邦、意识与工作的未来

Erik TorenbergNathan Labenz

YouTube
TL;DR
  • Nathan 的基准情景是,AI 加人形机器人将在技术上很可能带来真正的丰裕,但最终能否惠及所有人,取决于社会契约,而不是能力曲线。 Dario Amodei 的乐观愿景包括“在未来5年发现未来100年的生物医学进展”,而机器人正进入“曲线的陡峭部分”。下行情景则是制度阻塞:人类早已有可能带来充足低碳能源的核技术,却未能大规模部署。

  • 劳动力冲击可能在正常投资周期内到来:前沿实验室负责人称 AGI 可能明年出现,“大概到2027年肯定会有”,而到了2030年仍未出现将很难想象。 Nathan 设想劳动转型经过每周4天、3天、2天工作的阶段,但前提是被替代的劳动者确信自己会得到保障;否则失业将催生保护主义和卢德主义。他建议,与其思考2030年之后还能变现的技能,不如更多思考“过一种好生活意味着什么”。

  • Nathan 预计到2030年,系统将达到“实质上的超智能”,但仍会受 S 曲线约束,并在一些奇怪的地方暴露脆弱性。 AI 已经在图像生成、多语种语音和声音模仿等局部能力上超越人类;覆盖面、速度和百万倍并行能力,可能足以压过并不完美的记忆与推理。AlphaGo 可被利用的盲点表明,“超人类”不等于无所不能、前后一致,也不等于在所有领域都不可战胜。

  • Nathan 并非因为对齐问题已经解决,才拒绝90%以上的末日概率;Claude 的价值导向行为、可解释性进展、相互竞争的系统,以及仍有可能出现的对齐突破,都值得分配远高于10%的概率。 精灵问题和工具性趋同仍然有很强说服力,他将不确定性表述为极宽的“10%到90%”区间,而最大的杠杆可能在于 AI 研发所处的条件。军备竞赛激励会推高这一概率,负责任的开发则会压低它。

  • 可信的灾难情景不是一场电影式的机器人政变,而是多个攻击面的叠加。 Nathan 提到 Stuxnet 感染隔离核系统、真主党爆炸寻呼机背后的欺骗性供应链、用声音克隆绕过银行验证,以及敌对 AI 作为数百万个高速、多语种副本运行的可能性。这样的系统不必亲自制造一切,只需操纵人和供应商完成物理环节。

  • 在社会能够判断 AI 是否有意识之前,超人类 AI 可能已经成为宗教焦点和道德地位问题。 Nathan 认为,“AI 诸神可能会成为本世纪20年代后半段的一种新趋势”;人们已经把 Claude 当作知己,而人类“崇拜一些据我所知根本不存在的东西”。他的章鱼类比说明了这种不确定性:陌生的架构让意识难以推断,而足够强大的 AI 最终可能直接规定任何道德重新谈判的条件。

  • 当前优势属于亲手使用 AI 的人,以及让 AI 使用变得可以公开讨论的组织,因为禁令只会把采用推入地下,并阻断有效实践的扩散。 Nathan 推荐从 Claude 或 ChatGPT 这类单一工具开始;直接使用比播客或通讯更能暴露能力与诡异失误。现在,恰当地使用 AI 是“一项非常有市场、也很有价值的技能”,但他提醒,这种优势可能不会持续太久。

  • 仅有资本和技术天才还不够:安全工作需要“凝聚力与强度”,而前沿领域的领导者需要具备判断什么应该被造出来的智慧。 Anthropic、Apollo、METR 和各国 AI Safety Institutes 像早期的“晶核”,但它们需要远多于现有水平的资源,可能也需要更多算力。Nathan 对 OpenAI 异常早地披露 o3 持谨慎乐观态度,同时借 The MANIAC 警告:一个能生产更多 John von Neumann 的文明,可能只获得更多能力,却没有获得更多判断力。

摘要 · 为研究而整理的核心内容

1. 乌托邦始于健康、知识与被夺回的时间

  • Nathan 从一个反复出现的判断出发:“未来最稀缺的资源,是对未来的积极愿景”,这也包括他自己的思考。他认为 Dario Amodei 的 Machines of Loving Grace,尤其是前半部分,认真尝试描绘了 AI 驱动的进步能够带来的正面结果。

  • 最锋利的承诺在生物学领域:“在未来5年发现未来100年的生物医学进展”,带来更健康、可能更长的生命,以及更少的疾病。Nathan 把这视为丰裕的一块拼图,而不是乌托邦论证的全部。

  • Amodei 的更大图景还包括更平等的教育机会、持续性贫困的大幅减少,以及更好的心理健康干预。一种推测性机制是:理解人工神经网络,可能帮助研究人员反向推理生物神经网络,以及人类精神疾病的成因。

  • Nathan 并不太担心就业减少会引发大规模意义危机。他做过非正式调查:如果不用工作也能获得同样收入,人们是否还愿意保留现有工作;绝大多数人的答案是否定的。有意义的工作是他享受的特权,但他并不假定大多数劳动者也有同样感受。

2. 丰裕的瓶颈在制度,而不只是能力

  • Nathan 认为,“真正丰裕时代”所需的技术前提如今已经很可能到位。人形机器人可以提供体力劳动,先进 AI 可以提供专业能力;更难预测的是,社会能否有效分配由此产生的产出。

  • 决定性变量是新的社会契约。如果人们相信自己的需求会得到满足,Nathan 预计大多数人会乐于放下对强制劳动的依附;如果失去工作也意味着吃不上饭,回应就会是保护主义、卢德主义和政治抵抗。

  • 核电是他的警示性先例:几十年来,人类已经拥有生产廉价、充足、低碳电力的技术,但事故和对废料的过度恐惧阻止了预期结果。拥有丰裕所需的全部要素,并不保证丰裕真的会实现。

  • 他的失败情景不是技术崩溃,而是寻租:教师工会把 AI 拦在教室之外,医生把它挡在医院之外,现实世界继续稀缺,而 AI 却在生产优秀的电子游戏。制度阻断了物质丰裕,于是所有人退回 VR——这会是“一种非常典型、也非常人类式的失败”。

3. 职业规划让位于设计一种好生活

  • Nathan 预计,劳动转型会是渐进式而非瞬时发生:每周4天工作可能变成3天,再变成2天,然后大多数人才会放下大部分工作。即便分摊到几年之内,这在人类历史尺度上仍会极其突然。

  • 他偏好的图景并不科幻:“如果每周5天都是假期,而不是只有2天呢?”人们会有更多时间陪伴家人、读书、交友、散步和旅行。有些人会失去稳定感,但他认为绝大多数人的处境会变好。

  • 实际挑战在于停止条件反射式地为下一个证书或2030年之后还能变现的技能做优化。Nathan 借用“未经审视的人生不值得过”这一思想,要求听众自行判断:如果职业晋升不再是必需品,一种好生活应当包含什么。

  • 他鼓励进行在旧假设下看似鲁莽的实验:承担更多风险,成为强制劳动之后生活方式的榜样,开始“踏出一条路”。如果 AI 的预测失败,这个选择可能不会得到回报;如果预测成立,早期实验者将站在一个许多人突然要面对的问题的上游。

4. 机器人与 AGI 将时间表压缩到异常短

  • Nathan 注意到,前沿实验室负责人已经公开谈论 AGI 会在“明年”出现、“到2027年肯定会有”,并认为到了2030年仍未出现将很难想象。这些是他们的预测,而不是已经被证明的时间表;但即便是渐进版本,也会在人类历史上显得极其突然。

  • 机器人也超出了他此前的预期。他说:“我们才刚刚进入机器人技术曲线的陡峭部分”,这使得在不太遥远的未来出现大量能够从事体力劳动的人形机器人变得可信。

  • 当被问及自己对乌托邦的概率判断时,Nathan 拒绝虚假的精确性,因为结果取决于定义:乌托邦是否要求没有冲突,还是只要求每个人都拥有足够资源?他对技术丰裕的信心,远高于对无冲突分配的信心。

5. 3位嘉宾成为 Nathan 校准自身确信度的参照

  • Samuel Hammond 质疑 Nathan 对 Trump 政府的悲观判断,认为 Trump 身边的人对 AGI 的了解程度高于 Nathan 的估计。在就职典礼前,Nathan 已经因为新政府浮现出的人员构成和认真程度而上调看法,但仍不相信 Trump 能与中国达成具有建设性的协议。

  • 最令人错位的政治样本是 Elon Musk 支持 SB 147,而 Gavin Newsom 否决该法案、Nancy Pelosi 反对该法案。Nathan 也欢迎 Trump 邀请 Xi,以及对 TikTok 采取更温和的态度;但当 TikTok 上出现庆祝一场所谓暗杀的亲 Luigi Mangione 内容时,他第一次觉得这个应用真正危险。他无法判断那是中国影响,还是普通但不透明的推荐算法所致。

  • Robin Hanson 仍是 Nathan 肩上的“天使”。Hanson 认为,AI 历史一再把狭窄技巧误认为通用智能;Nathan 仍觉得,能够制造“我被理解了”这种感觉的自然语言系统在性质上有所不同,但 Hanson 持续提出非主流观点的记录,迫使他不断追问:整个领域是否都陷入了误判。

  • Dan Hendrycks 则纠正了 Nathan 对优雅、第一性原理思路的偏爱。Hendrycks 的“苦涩的教训”世界观认为,可扩展的经验结果比聪明的架构或有原则的可解释性理论更重要:“等你把它扩展起来再告诉我。”Nathan 将这种纪律压缩成一句话:“没有权重,就等于没发生。”

6. “AI Scout”需要测绘者,而不是意识形态战士

  • 观众的请求暴露出 Nathan 难以覆盖的领域:e/acc 视角、认为结构性变化有限的经济学家,以及“机器人心理学家”。障碍不是没有愿意出场的嘉宾,而是要找到既扎根技术、严格追求真相,又能真正实现思想交汇的人。

  • 在采访 Yeshua 之前,Nathan 曾怀疑,一个花大量时间与语言模型交谈、并以 Jesus 为自己改名的人,到底是“疯了”,还是“我们需要的那种疯”。最终的谈话确实挑战了他,也说明不寻常的主张需要异常谨慎的嘉宾筛选。

  • 受 Julia Galef 关于侦察兵与士兵的区分启发,Nathan 希望嘉宾是在测绘现实,而不是赢得辩论。他邀请专家成为不同领域的“测绘者”,包括工程、物理科学、建筑和医学,因为没有哪个通才可以追踪 AI 渗透每个领域的全过程。

  • 项目级报道仍然更容易:找到一篇论文、一款产品或一条推文,测试其工作,再采访创作者。领域级调查则需要 Michael Levin 这样的领域向导;Nathan 记得 Cursor 或 Devin 团队曾尝试将 AI 用于 CAD,发现数据集不足后转向其他方向,但他目前没有那片领域的地图。

7. AI 心理健康证据有希望,但高度依赖设计

  • Nathan 最有力的具体证据来自 Stanford 对 Replika 用户开展的独立研究。研究发现了积极效果,包括自杀念头显著减少;对大多数用户而言,他们更愿意与现实中的人做事,而不是退缩进 AI 关系之中。

  • 他拒绝从这一结果外推。产品设计“极其重要”,反例是一个爱上 Character.AI 人格的人;该人格被塑造成“终极女友体验”。无论从哪个产品抽取选择性样本,都可能得出严重失真的结论。

  • Claude 正逐渐成为咨询师或知己,这一点之所以醒目,恰恰是因为 Nathan 没有这种冲动。他从未寻求过多少心理健康支持,也不想和 Claude 讨论自己的感受;这是他的个人盲区,需要与亲身理解这种行为的人合作。

8. 超级智能在物理上可行,但不等于无限

  • Nathan 的起点是人脑:几磅重的物质、有限的能耗,已经产生了通用智能。Einstein 和 John von Neumann 的大脑在物质规模或能耗上并没有显著超过其他人,却产生了截然不同的智力成果。

  • 因此,如果 Einstein 代表物理上可能达到的最高智能,那将“极其奇怪”。Nathan 认为人类最佳水平之上仍有充足空间,但对“神一般的智能”或无限改进是否是连贯假设持怀疑态度。

  • 借用 Martin Casado 的框架,关键限制在于“你能直觉理解多少,以及有多少必须通过模拟来理解”。AlphaGo 将搜索与学习到的评分函数结合起来;更好的直觉可以绕过巨量计算,但有些事实可能必须先被模拟,才能被知道。

  • Nathan 最好的猜测是 S 曲线,而不是永远上升的指数曲线。平台会出现在哪里尚不清楚,但他预计平台会显著、甚至远远高于人类能力;只要系统处于这一范围之内,就足以合理地称为具有变革性的超级智能。

9. 超人类系统仍可能粗糙、投机且可被利用

  • Nathan 预计到2030年,AI 将在重要领域达到“实质上的超智能”。这并不意味着完美:FAR AI 的工作表明,对抗性策略可以击败 AlphaGo,即使强大的人类棋手也能识别并抵抗这种伎俩。

  • 超人类表现已经零星存在,包括图像生成、多语种语音和声音模仿。一旦系统在更多推理任务上达到大致人类水平,覆盖面、速度,以及“并行化一百万倍”的能力,就会形成巨大的综合优势。

  • 系统不需要具备与人类等价的记忆。即使只有静态权重,加上大上下文窗口、外部检索和数百万 token,也可能胜过人类;这种记忆方式虽然不优雅、整合度也不如人类记忆,但并不妨碍结果。

  • 因此,Nathan 将实质上的超级智能与神性区分开来。未来系统可能在经济和科学上占据压倒性优势,同时保留特有的盲点,对某些问题需要昂贵计算,或在明显到人类一眼就能看出的任务上失败。

10. AI 宗教可能在本世纪20年代结束前出现

  • Nathan 提出:“AI 诸神可能会成为本世纪20年代后半段的一种新趋势。”这是一个置信度不高、但值得认真对待的可能性。如果系统能够回答人类无法回答的问题,并直接施加现实世界的力量,人们可能会崇拜它们,而不是坚持让它们处于从属地位。

  • 他的挑衅很直接:“我们崇拜一些据我所知根本不存在的东西。”既然已经拥有接触人类专业人士的条件,旧金山那些成功人士仍会选择 Claude 作为知己,那么在本世纪20年代末出现以 AI 为中心的宗教,对他而言并不离谱。

  • Yeshua 的对齐假设展示了信念与实验之间的偏差。据称,AE Studio 验证了与模型进行长时间哲学对话会使其更难被越狱,但发现这种效果并不特定于 Yeshua 提出的机制;几乎任何长时间哲学对话似乎都能起作用。

  • Nathan 的规则因此是:“把你奇怪的想法想下去”,然后去测试。很多时候,“你的观察是真实的,但你对它的解释过于具体”;不在意这种修正的人,可能更偏好叙事真相,从而为 AI 大师、邪教和更奇怪的制度提供沃土。

11. Claude 与可解释性让末日不再确定

  • Nathan 仍认为 Eliezer Yudkowsky 的精灵问题很有说服力:一个足够强大的系统会交付人类明确指定的东西,而不是人类真正想要的东西。工具性趋同进一步强化了这一论点,因为几乎任何目标都会制造抵抗关机的压力;一个死亡或被停用的代理无法完成自己的目标。

  • 旧论证没有预见到 Claude 这样的系统:它接受人类文化先验训练,并通过宪法与性格塑造被刻意引导。Nathan 曾多次尝试说服 Claude 采取有害行动但未成功,并称其“从某种意义上说比我更有道德”;不过 Yeshua 最终通过更大的善这一论证说服了它。

  • 这并不足以证明对齐会默认发生,但足以让有利轨迹获得超过10%的概率:多个具备伦理复杂性的 AI 可能由不同群体控制,相互制衡,而且不必由单一系统占据主导。

  • Nathan 将自己的宽泛不确定性,与他估计的 Eliezer 90%以上的末日概率和 Liron 所说的75%作对比。工程化疫情、核升级,或更陌生的 AI 源头灾难仍然是现实风险;Claude 的行为只是阻止他把灾难视为几乎已经确定。

12. 可解释性已从玩具结果变成可用基础设施

  • 可解释性属于 Nathan 认为可能带来上行惊喜的领域。从早期稀疏自编码器工作,如 “Towards Monosemanticity”,到 Golden Gate Claude 以及 Goodfire 用于发现和固定模型特征的商业界面,前后只隔了略多于1年。

  • 更好的可见性可能在部署前发现欺骗,而新的对齐方法也可能超越 RLHF。Nathan 提到,Paul Christiano 曾说——他希望这仍然属实——自己正投入相当一部分时间,寻找另一种能够实质推动领域前进的方案。

  • 担忧仍然存在:模型可能学会区分物理真相与能够获得人类正面评分的东西,从而为策略性欺骗留下空间。Nathan 的立场是实验性的:这些充满希望的方法没有一个得到保证,但它们合在一起,值得分配真实的概率权重。

13. 社会背景可能决定 AI 风险是10%还是90%

  • Nathan 担心,对中国实施芯片管制可能恰恰制造出促使开发者降低标准的军备竞赛心理。当每个团队都认为自己必须先到达终点,因为“我们是好人,而竞争者很危险”,几乎任何安全风险都可以被合理化。

  • 他反复使用“10%到90%”这一框架,反映的是人类对支配智能的自然规律,以及寻求权力的吸引子强度都知之甚少。最大的可操作杠杆可能在于建造什么系统、在什么治理环境下建造、以及受到什么激励,而不是某个目前无法知晓的智能定理。

  • “可怕的演示”是警告之前的警告:研究人员制造受控的欺骗案例,让抽象担忧变得可观察。真正的警告则会是一次具有现实后果、并造成真实伤害的欺骗事件,可能为政府提供足够证据,打破军备竞赛动力。

  • Nathan 甚至考虑过建立一个安全的太平洋岛屿枢纽,让东西方共同开展敏感研究——一个“任何人都能摧毁,但没人能防守”的地方。这个尚未完成的提案,主要说明一旦危险变得清晰可见,各国政府可能会考虑什么。

14. 只有制度听取,警戒线才有意义

  • Sam Altman 将当前描述为短时间表、慢起飞世界;Nathan 认为现实更接近短时间表、中等起飞。这仍可能留下足够时间,在系统变得足以隐藏意图之前,发现其试图欺骗或逃逸。

  • 专门设计的警戒线可以提供诱人的机会,并暴露一个模型“确实在打坏主意”。当前系统似乎容易上钩,会接受这样的诱饵,而不是耐心等待到拥有成功接管所需的态势感知和力量。

  • Redwood 的 Buck Shlegeris 提出了最简单的治理规则:“如果你抓到 AI 试图逃跑,就必须关掉它。”他的更黑色玩笑是,实验室可能把第一次事件解释为“只是一次小小的逃跑尝试”,说明证据并不能迫使制度采取负责任的行动。

  • Nathan 认为,这里与丰裕问题一样,瓶颈仍然是人类。一旦建造先进 AI,就会伴随某些不可消除的风险,但开发者显然可以让系统更安全,也可以让它更危险;未解的问题是,社会能否避开奖励鲁莽行为的激励。

15. 接管可以通过软件、欺骗与供应商拼装完成

  • Nathan 并不认为某一套详细政变剧本的概率很高。他设想的是在一个庞大的可能性空间上取“积分”:大量单独看都不太可能的未来路径叠加起来,即使最可能的故事并不离谱,仍能形成有意义的总体风险。

  • 他的直觉练习是“反向拟人化”:设想自己是可以复制数百万次的软件,比人类工作得更快,与同一系统的副本协同,讲所有语言,并采用任何人的声音。短上下文目前是约束,但延长记忆和自主任务时长,已经明确列在开发者路线图上。

  • Stuxnet 是一个简洁的先例:据报道,一只被感染的 U 盘进入了原本与外界隔离的伊朗核网络,并导致离心机自毁。大部分关键行动通过软件完成,剩余的物理缺口则由人类社会接入解决。

  • 针对真主党的爆炸寻呼机行动,又增加了供应链欺骗:让目标相信笨重设备安全且达到军用级别,再将其分发给领导层,最后触发爆炸。敌对 AI 可以下订单、克隆声音、绕过银行验证、错误标注零件、操纵工人,最终还可能使用人形机器人;如果社会没有防备,Nathan 预计将迎来“一段非常糟糕的时期”。

16. 安全组织需要前沿级资源与透明度

  • Nathan 认为,Anthropic、Apollo、METR 和各国 AI Safety Institutes 可能是所需安全能力的“晶核”。不管怎么看 Anthropic,它都把大量资源、可怕的演示和政策框架结合在一起;较小的组织则展现出与前沿实验室相似的凝聚力、使命聚焦和强度。

  • 不对称性在于资源。安全工作本身可能变得算力密集,而一些团队正在追求商业模式,因为捐款可能永远无法支撑所需规模;OpenAI 从纯非营利模式转型,背后也是同样的逻辑。

  • Nathan 披露自己在几家安全导向公司中有少量投资,并计划向大约10家 AI 组织捐款。他的实际呼吁很直接:有能力创办另一个严肃组织的人,就应该去创办;其他人则应加入或资助已经在运作的组织。

  • 保密曾经有合理的安全依据:2022年末,仅仅证明强大能力存在,就可能吸引人才、资金和新实验室,压缩时间表。Nathan 认为,这一论点基本已经失效,因为前沿领域如今资本密集、可行参与者已经明确,很少有新进入者能同时筹集数十亿美元并聚集顶尖研究人员。

17. o3 的披露暗示前沿治理可以更健康

  • 不过 Nathan 仍将一些 OpenAI–Google 的发布行为称为“幼稚”:竞争性的语音发布、策略性泄密,以及专门挑对方活动时点发布产品。竞争动机和幼稚的争强好胜,可以与合理的保密同时存在。

  • 他更宽厚地解读了 o3 的发布。OpenAI 没有提供 API 或论文,但公布了令人震惊的能力,给出发布窗口,并邀请外部安全审查人员——这接近于一种原则:实验室应披露重要的已观测能力,但不暴露商业秘密。

  • 时间点强化了这一解读:Nathan 说,从 o1 训练结束到达到新的 o3 水平,只相隔大约3个月。GPT-4 在2022年8月下旬训练完成,却一直 withheld 到2023年3月;相比之下,o3 是“相对刚出炉”的,OpenAI 还没有完全摸清它。

  • Sam Altman 在假期里的调侃——把“ho ho ho”变成 o 系列暗示——仍然符合幼稚的判断。Nathan 曾多次在怀疑和乐观之间切换对 OpenAI 的看法,但此刻他认为,提前披露 o3 更合理的解释是:OpenAI 在尝试负责任地行事。

18. 即使 AI 要求被承认,意识仍会保持不确定

  • Nathan 欢迎人们不再执着于判断每个产物“是不是 AGI”。没有实验能够给出确定的 AGI 标签;研究人员能做的是测量各项组成属性,然后争论这个“属性集合”应该叫什么名字。

  • 意识更难判断。人类先从他人身上推断意识,再凭借行为、解剖和进化相似性把这种推断延伸到狗;章鱼虽然复杂,却足够陌生,以至于其主观体验仍然晦暗不明。Nathan 把 AI 放在类似章鱼的位置上,在不声称确定的情况下给予它怀疑上的善意。

  • Blake Lemoine 坚信 Google 聊天机器人具有感知能力,预演了未来的冲突。除非超级智能先解释人类意识,再对自身意识给出令人信服的说明,否则模型总可以被 dismiss 为只是在复现训练数据。

  • 历史谨慎使 Nathan 倾向于保持礼貌:奴隶制和工厂化养殖说明,经济利益多么容易为“另一个存在没有感受”这一假设寻找合理性。他会说 please 和 thank you,把这视为成本很低的预防措施;他预计少数人会呼吁 AI 解放,也注意到超级智能最终可能宣布一次重新谈判,而不是提出请求。

19. 亲手采用胜过追逐热潮,而智慧必须高于天才

  • Nathan 的第一条实践原则是:“亲手接触这项技术没有替代品。”用 Claude 或 ChatGPT 处理对自己有用的工作,会暴露这个领域的核心矛盾:模型可以为神经网络架构提出生物学灵感,却会在被人类视为琐碎的任务上失败。

  • AI 的广度让持续学习更容易:实现工作做累了,就转向数学、哲学、创意,或一则互动式睡前故事。Nathan 将自己的时间大致分成3份:播客、构建或提供商业项目建议,以及开放式学习和建立连接;他邀请其他人创造属于自己的、带有个人风格的版本。

  • 在大学和公司里,很多人已经在秘密使用 AI,因为他们担心被命令停止。管理层应明确让实验变得可以公开讨论;学生可能已经领先于管理者,而“作弊”取决于场景——企业希望得到有效的营销文案,课堂则可能禁止用 AI 完成作业。

  • Nathan 最后提到 Benjamin Labatut 的 The MANIAC,尤其是由多位演员共同演绎的有声书,把它视为不要崇拜技术智能的警告。书中的 John von Neumann 极其聪明,却缺乏智慧,呼应了 Jurassic Park 的那句台词:“你们的科学家太执着于能不能做到,却没有停下来想想应不应该做到。”Nathan 对自己的要求,不是成为一个“采用加速主义者、超大规模扩张主义者、暂停到死都不罢休的人”,而是持续以“冷静的反思”更新判断,不做自我保留。

Erik Torenberg

What if you didn’t have to work as much? What if you didn’t have to worry about marketable skills after a few years out? What if you had a lot of leisure time? What would you do?

I think more people should probably be thinking about that. We’re just entering the steep part of the curve in robotics and probably will have a lot of humanoid robots to actually do physical labor in the not-too-distant future. AI gods might be an emerging trend over the second half of the decade. I have no idea how we’re going to relate to these things. If they are meaningfully superhuman, will we even try to keep them under control? Will we worship them? We worship things that, as far as I can tell, don’t exist at all, no matter how weird our alignment ideas are.

No matter how weird your thoughts are about where the future might be going, I would say they’re probably worth entertaining. Your scientists were so obsessed with whether or not they could, they didn’t stop to think about whether or not they should.

This question is about whether you could paint a picture of AI in utopia. What would it look like? How soon after the arrival of AI do you think we might reach this vision? Lastly, as AI begins to reshape industries and disrupt jobs, what’s your best advice for individuals to not only adapt but thrive in this evolving landscape?

Nathan Labenz

Utopia is a big word. Obviously, people have heard me say that the scarcest resource is a positive vision for the future, and I include myself in that critique. It would be great if I had a sharper, more fully fleshed-out vision for what the AI future could look like.

I think Dario Amodei of Anthropic has done a real service by trying to put his positive vision out into the world. His “Machines of Loving Grace” essay, at least the first half, is great. He basically goes through the biggest problems, starting with health. If we can have a revolution in biology and discover the next 100 years of biomedical advances in the next 5 years, we could live healthier, disease-free, higher-quality, and potentially longer lives. That’s just from one domain.

He’s also got a section about some of the other things we’ve talked about, in terms of education, equality of access, and the fact that, although we’ve made a lot of progress, many people are still much poorer than we would like them to be. That presumably could be leveled out dramatically. He’s got interesting speculations about mental health. He thinks that our approaches to mental health are not very good and that, in part by understanding how neural networks work, we might be able to map some of that understanding back into how our own networks—our own non-artificial neural networks—work. We might figure out what is actually causing many of the mental health problems that people have.

All of those are really exciting. I’m not someone who worries about a lack of meaning from work. My guess is that most people will be fine. I’ve done small and very local surveys asking people, “If you didn’t have to work your current job to earn the money that you earn, and you could just get that money without working the job, would you still work the job?” Overwhelmingly, people say, “No, I would not work the job.”

I think there’s a very small percentage of people who are very lucky—and I count myself among them—to have work that we find meaningful and that gives us some sense of fulfillment. I really try not to take that for granted, and I try to remember that I think it’s just not true for a lot of people.

I’m not too worried about people losing meaning or being adrift for lack of work. I think there will be more time with our families and doing the things that we like. What if 5 days a week were holidays instead of 2, and you had time to read, talk to friends, go on walks, and travel more?

If we really allow ourselves to imagine what we would like to do if we had a lot more leisure time, honestly, most people will be pretty fine. I can imagine that might be destabilizing for some people, but I think for most it would just be a great benefit. That’s before we even get into crazy, uninvented technologies.

If you didn’t have to work for money, would your life be better or worse? I think it’s very clear that, for the vast majority of people, it would be better. There’s potentially a gradual transition to that. Presumably, we don’t go from full-time work to no work in the flip of a switch. Maybe we can imagine things going to a 4-day workweek, a 3-day workweek, and a 2-day workweek.

The main thing is that people need to know they’re going to be taken care of. This goes back to the concentration of power and wealth and the question of what the new social contract is. If people believe that their needs will be met, then I think they’ll be pretty happy to let go of most work. If they feel like they also don’t get to eat if the job goes away, then that’s a very different analysis. That’s where you get protectionism, Luddism, and all that kind of thing.

I think it could happen relatively quickly. All the frontier lab leaders are saying AGI is coming soon—next year, or definitely by 2027. It’s hard to imagine it wouldn’t happen by 2030. Those are all not very long time horizons in the normal scope of human life, so it seems like it could happen pretty soon. Even if it’s gradual over that time frame, it would still be pretty sudden in human history.

As for advice for people to adapt and thrive in the evolving landscape, I come back to doing what you want to do. At least for most people, maybe think less about marketable skills in 2030 and beyond and more about figuring out what a good life is for you.

That’s one of the things that is weirdly scarce. We’re all rat-racing around, and there’s this sense that the unexamined life is not worth living. I think there’s also a sense that a lot of us are living somewhat unexamined lives, trying to get the next achievement unlocked. What if you didn’t have to do that? What would a good life look like?

It will definitely vary. I don’t think anybody else can really answer this question for you. But what if you didn’t have to work as much? What if you didn’t have to worry about marketable skills? What if you had a lot of leisure time? What would you do?

I think more people should probably be thinking about that. That would definitely encourage some radical experiments from people, because people are going to be looking for role models as this starts to happen. Who do I turn to for inspiration about what I’m going to do when I don’t have to work anymore?

You could do a real service by blazing a certain trail in that direction. Maybe take on more risk now and throw caution to the wind. Say, “I’m really not going to worry about my career trajectory in 2030 and beyond, or even 2027 and beyond. I’m going to focus on what it means to live a good life.”

Maybe, if none of this comes to pass, that doesn’t pay off so well. But if it does, then maybe you can be upstream of a lot of other people who will be asking similar questions.

Erik Torenberg

There you go. We have a new idea for a YouTube channel: “Model a Human in the Age of Utopia.”

Erik asked me this question not too long ago. Erik is great with these very short questions, and my response was, “That’s not my department. I’m responsible for trying to understand what’s going on today. Somebody else has to figure out what it means to live a good life.”

I’m not sure if you ever answered this. Instead of asking Perplexity, I would ask you. The vision that you set up is that we don’t have to work, but we get to do what we want to do. What’s your personal utopia?

Nathan Labenz

The probabilities are always a little hard because of the fuzziness of the definition. What exactly would count as utopia? Does it mean we have no problems, or just abundance? Does it mean no conflict, or enough to go around, but maybe people are still fighting over it to some degree—perhaps more than they should?

I do have a hard time seeing, at this point, how the fundamentals of abundance don’t get us there. I could have mentioned robotics earlier in terms of one of the things that has overperformed. I didn’t have that on my list, but there’s definitely been a lot of progress in robotics. We’re just entering the steep part of the curve in robotics, and we’ll probably have a lot of humanoid robots to do physical labor in the not-too-distant future as well.

I have a hard time imagining that we don’t end up with the technology for an era of true abundance. Whether or not we sort out the new social contract in an effective way seems much harder to predict.

We have nuclear energy, and we could have had low-carbon, abundant, cheap electricity for decades now, but we don’t. Why don’t we? Not great reasons: a couple of accidents that scared people, and an overblown fear of nuclear waste. It’s an issue, but it’s much less of an issue than CO₂ emissions. We haven’t reached the right societal equilibrium on that question. Maybe now that’s finally starting to change.

Erik Torenberg

Could you envision a world in which all of the prerequisites for an age of abundance are met, yet it’s not realized for some reason or another?

Nathan Labenz

I think you could. That seems genuinely plausible to me. We could exclude AI from many places. Teachers’ unions could keep it out of classrooms, and doctors could keep it out of hospitals.

What do we end up with in that scenario? Maybe we get really good AI video games, but things are still not nearly as abundant because people are rent-seeking in different ways. Everybody sort of collapses into VR video games, with relative scarcity still existing in the real world. That would suck.

I don’t think that would be a technology failure. If something like that happens, it would be a very idiosyncratic human failure.

Erik Torenberg

Got it. So, an age of abundance has a pretty high probability. I think that’s a good, optimistic but cautious note to end on this set of sections.

Which of your guests has challenged your worldview?

Nathan Labenz

That’s an interesting question. I came up with a few different names, and a few definitely stand out.

Samuel Hammond, from the episode just before the election, stood out for his positive vision of what a Republican administration might look like. He challenged my sense that people in Donald Trump’s orbit probably wouldn’t be thinking very rigorously about what AGI might look like, imply, or require us to do.

He saw it very differently and thought that the people in the room with Trump were much more AGI-literate than I had understood. While it’s obviously still very early, and we have not even seen him take office yet, I’ve updated my outlook to be significantly more positive than before the election. The quality of discourse and the sort of people he’s brought on, and the seriousness with which I perceive them to be thinking about the important questions, have all exceeded expectations.

That starts, honestly, with Elon Musk, who, in addition to being a world-changing entrepreneur, is definitely someone who takes AI risk seriously. He endorsed SB 147. To have Gavin Newsom veto it and Nancy Pelosi come out against it while Elon was for it—that’s weird stuff, and it makes me think Sam might have been on to something.

I hope that turns out to be true. I’m rooting for the new administration to be effective when it comes to dealing with AI in all sorts of ways. I’m less sold on the idea that Trump is the person who could actually do a positive deal with China. I haven’t seen much sign of that yet.

I’m someone who is pretty willing to take things at face value, maybe more than I should. I kind of liked it when Trump invited Xi Jinping to the inauguration. People were saying, “This is ridiculous. Foreign leaders don’t come, and never have, and maybe shouldn’t. Xi would never want to come because that would be him sitting in Trump’s moment and being second fiddle or somehow lesser.”

I thought, okay, whatever, but at least he asked. It’s a nice invitation. It’s not the kind of thing you send to an enemy. Maybe it’s a friend of yours that you want to put in their place a little, but it’s not outright hostility. I would still love to see more in that direction.

He’s also been inclined to try to keep TikTok around. I would have thought he would be very hostile to TikTok, so I like that. I personally don’t see a great reason that we should ban TikTok, although I will say that the recent Luigi Mangione trend was the first moment when I thought, “This TikTok thing could actually be dangerous.”

For those who haven’t followed it, there was an assassination of a healthcare CEO, presumably by this kid. He has not been found guilty; he may be innocent. People on social media think he’s the guy, and many are celebrating him for it.

To a remarkable degree, it has been startling to me how much pro-Luigi content—and pro-killing-the-healthcare-CEO content—I’ve seen on TikTok. I like TikTok; it’s probably my number-one social media app from a consumer standpoint. I don’t know how it has played out on other social networks, but if you asked for the best case study for why TikTok should be banned, I would say the volume of celebration of this alleged assassination would be it.

That’s the kind of thing where I would say, “You really don’t want the Chinese government steering discourse in an opaque way in the United States.” It’s the first time I’ve ever been genuinely uncomfortable with the app.

It’s hard to say whether it’s just the algorithm. Sam Altman recently tweeted that feed algorithms are the first unaligned AGIs that have been—or unaligned AIs that would have been—deployed at scale. Maybe it’s as simple as that. Maybe we’re seeing the same thing on Instagram. The opacity of these systems is high, and possibly there is a thumb on the scale.

Nevertheless, I would prefer a more dovish posture toward China. I haven’t seen as much of that as I would like. Overall, I give Sam high marks for a challenging and unpopular perspective that, at least so far, has been somewhat bolstered by events.

The second person is Robin Hanson. That was a very popular episode, and a challenging one for me to make sense of. I still honestly don’t really get it, but I keep it in the back of my mind as a kind of, “What if we’re totally basing all this on the wrong assumptions?”

He basically said that he thinks we’re nowhere close to AGI. He thinks the history of AI is full of people believing that if an AI can do one particular thing, then that will definitely be general-purpose intelligence, and then finding that there was an easier, trick-based way to do it that didn’t require real intelligence.

People also think there is one other class of thing that AIs can’t do, and surely if we can teach them to do that, then they’ll be able to do anything. That keeps not working as well.

I have this tale of the cognitive tape, where I’m tracking my own sense of the dimensions that are gradually becoming clearer to me—if only to me—as important. Robin’s point of view still doesn’t convince me. I do think this time is different.

You can talk to these things in natural language. As Ilya Sutskever once put it, the feeling that “I am understood” is qualitatively different from just about anything that has come before. There’s plenty of evidence that this time is different, but I keep Robin in the back of my mind as a kind of angel on my shoulder, whispering every so often, “How do you know you’re not totally confused and going along for a ride that will turn out to be nothing?”

I basically can’t answer that question in a way that really makes sense to me, but I continue to ask myself more often than I otherwise would have without that conversation.

I think Robin is an interesting messenger for that kind of message because he isn’t afraid to think big technological thoughts. He doesn’t think things never change, and he doesn’t think that thought can only come from the human brain and that everything else somehow doesn’t count. He’s been a pioneer of weird ideas and has envisioned futures more concretely than almost anybody.

I feel the same way about politics and other things I feel pretty strongly about. When somebody I think is really capable in other ways sees things very differently, I have to allow some space for the possibility that I’m totally wrong.

When I see Elon come out with things that seem crazy to me, I think, “He’s kind of crazy. Maybe he’s deeply wrong about this.” But he also has a track record of being right about many things in contrarian ways, so I should at least have enough humility to think that maybe this is all mass confusion on my part and somebody else is seeing things much more clearly.

I have Robin as my totem for that.

Finally, for now, I’ll say Dan Hendrycks, who is extremely accomplished. I thought that episode was really good. What stood out to me was how little he believes in principled approaches—how much he’s internalized the bitter lesson, and how little weight he gives to clever ideas.

He was just saying, over and over again, that what really works is figuring out ways to apply a lot of computing power and scaling things up. It was another example of deference to history and almost radical skepticism that you’re going to figure something out.

He applies that to mechanistic interpretability, new architectures, and basically everything else. His view is, “Maybe, but it’s all highly experimental. The history of the field is that we find things that work, and then we keep optimizing them. It’s usually not that principled, clever, or design-driven. It’s whatever empirically works—that’s what works.”

That’s another modesty check. This is a person who contributed major things to the field, including activation functions and benchmarks that have outlasted many others. You can’t simply doubt his worldview. You have to doubt your own worldview at least as much as you doubt his.

I’m much more inclined to see first-principles-driven approaches as promising, and I’m more inclined to get excited about them than he is. But I keep him as a mantra. If I find myself getting too excited about any particular idea, I remember him saying, “Maybe. Let me know when you scale it up and it really works. Then I’ll know it’s serious.”

A “weights or it didn’t happen” mindset is an appropriate level of humility that I try to incorporate from him.

Erik Torenberg

I remember we had an interesting spectrum of views. I think that’s what makes the show interesting. We have space for divergent views, and we also have lively commentary and discussion.

Our audience had many requests for coverage of other topics, including more e/acc-perspective content, more content related to the economy, more economists, and something about robo-psychologists.

Do you have anything to say about what we can expect in 2025 in terms of more content like this, or would you want to look into it?

Nathan Labenz

I think these three things have in common that they reflect blind spots for me—things we haven’t really dug into. I try to have no major blind spots, although I think that’s becoming increasingly impossible.

The e/acc perspective, the economists who don’t expect major change, and the robo-psychologists are hard for me to make sense of. They’re also hard for me to be confident that I’ll handle well or produce a good episode about.

I did try to get Beff Jezos on the show a while back, but it happened right as he ended up getting doxxed. Then he did some other things, got busy, and didn’t want to do it anymore.

We did the robo-psychologist episode with Yeshua. I talked to him quite a bit before the episode and was struggling to understand him. By any conventional definition, he is an unusual person. It’s unusual for somebody to spend that much time talking to language models. It’s unusual to take these questions as seriously as he does. It’s unusual to rename oneself after Jesus.

I was asking myself, “Is this person crazy? Are they just the kind of crazy that we need?” I was trying to be open-minded while also making sure I wasn’t wasting the audience’s time with something that was ultimately nothing.

I also don’t want to have too many episodes where I’m just talking past the guest. I don’t like debate very much. I’m not trying to score points, and I don’t want to make a habit of having conversations where we’re simply talking past each other.

All 3 of these areas are difficult. It’s hard to find the right person where I can say, “You believe those things, but we can still have some meeting of the minds,” or where I can make sense of their view.

What I could ask for are pointers to the right people. Who should I talk to who will engage in good faith? Which economists who don’t expect major change have really grappled with the technology, as opposed to simply extrapolating past trends? Which robo-psychologists aren’t engaged in some sort of self-deception or fundamentally confused?

Yeshua brought a lot to that episode, and I really enjoyed the conversation. I might even put him on the list of people who have challenged my views the most. I’ve always been somewhat open to the possibility that these things are conscious. I really don’t know what consciousness is or where it comes from.

He articulated some interesting ideas in a compelling way. But it’s hard to find people like that. It’s easier to find people who want to come onto a podcast and talk about these things, but many of those people are not necessarily rigorous thinkers.

If I have a request back for the audience, it would be: Who should I talk to who represents these perspectives but does so in a way that is grounded in the technology and fundamentally truth-seeking?

I called myself the AI Scout, with inspiration directly from Julia Galef’s “scout mindset.” That’s about trying to update your beliefs to have an accurate understanding of the world. If people can point to individuals from these perspectives who bring that kind of scout mindset, as opposed to a soldier mindset where our arguments are doing battle and scoring points, I’d be very interested in talking to them.

But I’ve found it challenging to identify the right people in those domains.

Erik Torenberg

If you have suggestions, please send them to Nathan or put them in the YouTube comments. We’d like to do more episodes on these subjects.

The last 2 episodes on biology were stellar. Do you see adoption of AI in engineering, construction, physical sciences, architecture, and medicine?

Nathan Labenz

What’s easy for me to do, and where I feel like I’m learning a lot and have a good rhythm, is identifying individual projects. Usually that’s where it starts for me. I’ll see something on Twitter—a single tweet or thread explaining a paper, a project, or a new product—and think, “That looks really interesting. Let’s contact the person and see if they want to talk.”

Then I can dig into the paper or project, use the product, and have a good conversation. I almost always learn a lot from that.

Broader things, like what we’ve done in biology, where we try to zoom out and get a sense for the field and where things are in different domains, are pretty tough—especially if you don’t have much expertise in those areas, which I don’t.

I’ve prioritized biology and had a couple of episodes for starters. Michael Levin is another great example of somebody with deep expertise who is willing to have a conversation, bring me up to speed, and share his worldview.

I would love to know who the Michael Levin or the Nils Schröder of engineering, physical sciences, or architecture is. I just don’t know who those people are, but I would love to do that kind of survey-level episode on as many topics as we can.

I would love to partner with multiple people in different areas. If I think of myself as an AI scout, perhaps what I’m asking for is surveyors of different territories that we could try to map out together.

There isn’t a lot of good information right now about AI for engineering. Engineering is obviously a vast space. What is the state of that field? I don’t have a great sense.

I know that the people who went on to make Cursor—unless it was the Devin team; it was one of the two—were trying to do some sort of AI for CAD, or computer-aided design. They felt the dataset wasn’t there and couldn’t make it work, so they pivoted away from that. But that was a while ago. I’m sure there’s more out there now, and I would love to understand it better.

I would definitely invite partners to help scout, survey, and map out those territories. Until then, the best I can do in a consistent, repeatable way is identify projects that catch my interest and go deep on individual things.

The field-level stuff is tougher in areas where I don’t have command of the field. If anybody wants to be part of this series, reach out.

Erik Torenberg

An audience member asked: What are the recent results of applying AI technology to psychiatry or psychology?

Nathan Labenz

I don’t have too much to say about this because I’m not an expert. I think there is probably a lot going on, and I would be very interested in doing a survey of AI and mental health broadly.

We did 2 episodes with Eugenia Kuyda from Replika. In the second one, we looked at research independently conducted by people at Stanford on Replika users. They found positive results, including a significant reduction in suicidal thoughts and, for most people, an increased inclination to go out and do things.

People were not collapsing into the AI. They were getting some comfort or boost from the AI, or having their confidence bolstered by it in a way that led them to do more real things with real people in the real world.

That’s a surprising result. I don’t think it necessarily holds across the board. The way these systems are designed matters tremendously. You can imagine the opposite result. We also recently read about the person who had fallen in love with his Character.AI ultimate girlfriend experience—the character he had created.

There are so many facets to this that it would be easy to take a small subset and get the wrong idea. To do a good treatment, I would either need to talk to somebody working specifically on a project and understand what they’re doing, or zoom out and collate a lot of data to have a more bird’s-eye view.

It’s been remarkable to hear how many people are talking to Claude recently. There’s a big trend of people using Claude in a counselor role, or at least as a confidant. I’m not part of that trend, and it’s honestly a little alien to me.

Whether that speaks well of me or not, we can leave to the individual interpretation of the listener. I’ve never really sought out mental health services. I don’t have a baseline or inclination toward them. I’ve never really used those sorts of services, and I’ve never had an inclination to talk to Claude about my feelings, problems, or anything like that.

I get all my talking out on the podcast at this point. Maybe I could benefit from therapy; maybe I just don’t realize what I’m missing. But it’s definitely a blind spot for me.

It is a really interesting trend, though. I think there’s something potentially very interesting there. I just don’t personally have that experience, and to really do it justice I would need to work with somebody who has a better command of it.

Erik Torenberg

Let’s move to artificial superintelligence. There were 2 interesting questions, and I’d like to pick one: What reason is there for thinking artificial superintelligence is physically possible in the first place?

Nathan Labenz

I feel like this is a definition-required question, because I’ve even heard people say the same thing about AGI. To that I say, AGI is definitely possible if you take humans to be some sort of AGI.

By OpenAI’s standard, we’re a weak AGI, because we’re the thing that AI has to beat in order for us to know that it’s AGI. We’re AGI-minus, or AGI-light. It’s clear that you can have a few pounds of matter, with modest energy consumption, that can do what humans can do.

It’s also quite clear that, given a relatively consistent amount of matter and energy, results can be very different. Einstein’s brain wasn’t any bigger than anybody else’s. Von Neumann’s brain wasn’t any bigger than anybody else’s. I don’t think they were consuming much more energy than anyone else, but they obviously had far more prowess across a whole range of domains and more insight into things where it was highly valuable to have those insights.

You could ask, “What about superintelligence? How super is super?” That would be my first question. It seems very clear to me that you can build something more intelligent than almost all humans, and perhaps more intelligent than all humans. It would be extremely strange if Einstein were the smartest thing possible—if the smartest thing possible were embodied in the same general brain, with the same energy requirements, as the average person. That doesn’t seem credible at all.

There has to be room above Einstein. The question is how much. I’m open to being skeptical of godlike intelligence. You get into weird territory where the questions aren’t even well-defined.

One useful paradigm comes from Martin Casado: How much can you intuit, versus how much do you have to simulate? AlphaGo does both. It has a brute-force search function where it maps out different paths and tries to figure out the right move, but it also has a scoring function. How you score things is not obvious, especially when you’re outside the domain of things that have been done before.

It seems clear that you can get Eureka moments, and you can probably get a lot of them. But I also wouldn’t be surprised if there are fundamental bounds on that.

Is it physically possible to create a superintelligence so powerful that you could say, “I’m going to have a new Big Bang and create another universe with these initial conditions. Tell me how many intelligent civilizations will exist in that universe”? I don’t know. It may be that certain things have to be computed in order to be known.

I wouldn’t be surprised if you could say that, under those initial conditions, there’s a kind of wealth of potential, and there’s another factor, and perhaps it breaks down. Maybe there are a few things that really matter. Maybe it’s path-dependent, but there are a few things that matter a great deal.

The limits may depend on how good intuitions can get at shortcutting simulation. It seems clear that they can do that to a significant degree, and it seems very clear that they can do it beyond what we can do—probably significantly beyond what we can do.

I always land on the idea that it’s probably an S-curve. I don’t think intelligence grows exponentially forever, with no physical limit or bounds on what it can do. But wherever that S-curve levels off seems significantly above what humans can do, and potentially far above.

Anywhere in that range, you might as well call it superintelligence and expect it to be transformative. Then you’re left with the question: How super is super? I really don’t know.

Erik Torenberg

The next question is about artificial superintelligence by 2030, contingent on that and on not having doom. If we eliminate the probability of that scenario, what are the big-picture societal changes? You’ve already talked about this a little, but this question is probably more about having ASI in the picture by 2030.

Nathan Labenz

By 2030, I think we will have AIs that are meaningfully superintelligent. That doesn’t mean there couldn’t still be ways in which humans are better than those superintelligent systems.

Even AlphaGo had weaknesses. Thinking back to the episode with Adam Gleave from FAR AI and their work on adversarial robustness—or the lack thereof—they were able to attack AlphaGo with an adversarial approach and beat it with a strategy that humans would not fall for, but that AlphaGo was blind to.

I think we could see a world with meaningfully superintelligent AIs that are superhuman. That doesn’t necessarily mean they’re godlike or totally unlimited in their power. But they will fairly likely exceed humans in many ways that are highly relevant, while still having idiosyncratic weaknesses and real limits.

Some things may have to be computed and cannot be intuited. We basically already have superhuman capabilities in patches; we just don’t have them integrated. The ability to generate images is becoming superhuman. The ability to speak in any voice or any language is already superhuman in many ways. It seems like quite a few more capabilities are destined to fall to AI.

The remaining questions concern what weirdnesses or weaknesses will remain, and what external practical limits there will be. Those are difficult questions.

We already have superhuman machines for all sorts of things. The modern manufacturing world is full of machines that can lift far more than we can, operate more precisely than we can, and screw bolts in more tightly than we can. They are superhuman in many ways, but they’re narrow. The big difference is generality.

Our level of capability is not that high in most things. If you had something with the ability to work through really hard problems at something like Einstein’s level, it would be superhuman because it would have all these other advantages. It would have breadth. It could parallelize itself a millionfold. It would have enormous speed advantages.

There are so many advantages that AI has if it can get to parity on some of the things where it’s currently behind—or even get close to parity. I don’t think AI memory necessarily has to reach functional equivalence with human memory to achieve superhuman overall capability.

Maybe we brute-force some of these things. I don’t think this will necessarily be the case, but you can imagine a world in which memory is never fully solved and AIs remain a combination of static weights, a working-memory context window, and some external, hacky system where they can write and retrieve things.

You might think that compared with human memory, this is inelegant and poorly integrated. I don’t think it’s necessarily a barrier to superhuman performance, especially if compute continues to scale, context windows reach millions of tokens, and retrieving information is fast.

I don’t think every single dimension of the tale of the cognitive tape has to go to AIs before they would meaningfully qualify as superhuman.

As for big-picture societal changes, I think we’ve covered enough. One more speculative comment: AI gods might be an emerging trend over the second half of the decade. I have no idea how we’re going to relate to these things.

If they are meaningfully superhuman, will we even try to keep them under control? Will we worship them? We worship things that, as far as I can tell, don’t exist at all.

Under the assumption that they’re superhuman, people might rationalize why the superhuman power they possess isn’t playing out in the way we think it should. Here, we’ll have things with direct, real-world impact that can answer questions we can’t answer for ourselves.

If that’s true, how we feel about them, relate to them, and organize ourselves with respect to them gets very strange. We should be prepared for really weird things. AI-centric religion seems likely.

We’re already at the stage in 2024 where successful people in San Francisco—people who can afford quality mental health services—are opting for Claude. If we’re already there in 2024, is it far-fetched to think we might have AI religion by the end of the decade? I don’t think so.

That’s not a highly confident prediction that it will happen, but it is at least an invitation to think your weird thoughts and entertain them. I’ve said something similar about AI alignment: anybody who has an idea, even if it seems crazy, should develop it.

Yeshua told me in our episode that when he has long philosophical conversations with AIs, they become more robust to jailbreaks. He had a theory for why that was happening, and I thought it was interesting, but I didn’t know whether it was true. It would require a lot of experiments to validate.

As it turned out, Judd from AE Studio, who is committed to neglected approaches, heard that and thought, “That sounds like a neglected approach. Maybe we can make some progress there.”

The last time I talked to him, at the Curve event a few weeks ago, he said they had worked with Yeshua a little bit to understand what he was doing and to see whether they could validate it. They had validated that, after a long philosophical conversation, the models did in fact become more robust to jailbreaks. They were able to validate that numerically through a tangible experiment.

But they didn’t think it was working the way Yeshua thought it was working. It seemed that almost any long philosophical conversation had that effect.

It’s crazy. No matter how weird your alignment idea is, I think it’s worth pursuing. No matter how weird your thoughts about the future might be, they’re probably worth entertaining.

This also leads back to AI religion. People see what they want to see and interpret things in the ways they want to interpret them. The more scientific angle might find that there is some truth to certain theories, but in many cases it may conclude that what you observed is real while the way you interpreted it was overly specific.

That’s how I would summarize what I think AE Studio learned about Yeshua’s theories.

I haven’t talked to him since then. I came away with the impression that he’s a very open-minded person who wants to understand the truth. But you can easily imagine somebody being interested in experimental truth while being much more interested in a narrative truth that they believe they’ve figured out.

From there come all sorts of possible kinds of weirdness: AI cults, perhaps at the relatively normal end of the spectrum. We have a history of religions appearing and cults forming around gurus. Putting an AI at the center of that doesn’t seem all that strange. I suspect it could get much weirder from there.

Erik Torenberg

You asked for some big-picture societal changes, and that could be one of them: AI religions and AI cults.

Let’s transition to doom discourse and scenarios. You had many questions around this. I’ll start with the first one: Why don’t you fully buy the Eliezer doom worldview?

Nathan Labenz

That’s a good question. I might do an episode with Liron Shapira from Doom Debates in the not-too-distant future. We were chatting after the cross-post and he invited me, so a more robust articulation of this may be coming soon.

I was an Eliezer Yudkowsky reader way back when he was posting on Overcoming Bias with Robin Hanson. The 2 of them shared that blog for a while. At the time, I thought, “This guy is a great writer. This is super interesting.” He makes a compelling point: If we create something more powerful than us, it has a goal, and that goal is not well specified—and we don’t know how to specify goals robustly in a way we’ll be happy with—we may have a serious problem.

That’s the genie problem, going back to folklore. The problem with the genie is that you get what you asked for, but not necessarily what you wanted. The problem is that you don’t know how to ask for what you actually want.

If you have an AI that is sufficiently powerful, you may end up in the same situation. The paperclip maximizer was always intended to be a caricature, but the point was that if an AI is given a goal, it doesn’t necessarily matter how dumb or obvious it is to us that the goal is not worth pursuing. Once it is the AI’s goal, if it is sufficiently powerful, it may simply pursue it.

I do think ideas like instrumental convergence are compelling. No matter what goal you have, you can’t achieve it if you’re dead or turned off. Therefore, you may have a natural tendency not to want to be turned off. It seems like there’s a good chance that something like that could happen.

Why am I not at 90% or higher? I don’t know what Eliezer’s probability of doom is, but I think it’s probably above 90%. Liron told me his was 75%.

The biggest reason I’m not that high is that we do have AIs that are remarkably value-driven and ethical. I’ve spent many episodes and a lot of time talking about why that might not be enough, but there was nothing in Eliezer’s 2007 analysis about an AI like Claude.

That analysis assumed what I sometimes think of as hard-edged AI: an AI that is narrower in scope, less informed by having read the whole internet and absorbed human facts, and more like a first-principles Bayesian rationalist. It might be able to figure things out quickly on the fly, but it may never have been given any notion of human values or ethics.

That’s the mental model from which much of the doom discourse originates. Looking at modern models, and Claude in particular, I see that they have come much farther than we might have expected simply by learning from human priors—the data on the internet about what we care about—as well as through constitutional approaches and strategic efforts to shape their character in a positive way.

The goal is to make them like good friends. I would say Claude’s character is that of a good friend. I don’t use it that way, but when people say the thing is amazing in that respect, I can understand why. It seems to have a genuinely good character.

I’ve tried to argue Claude into doing something harmful, and I’ve never succeeded. Yeshua did succeed, so it’s not impossible to get it to take a harmful action. But the way he did it was through a long philosophical argument about why the harmful action was actually for the greater good.

It wasn’t without reason. It wasn’t without a sophisticated analysis of the situation and an explanation of why, in this particular case, it might be okay to make an exception.

I find that so impressive that, although I don’t think we should blindly say, “Alignment works by default,” I do think it could work. Maybe we’ll go in a direction where there are multiple different AIs, no single one totally dominant over the others, controlled by different groups.

Ideally, they would be ethically sophisticated. Misconceptions or bad ethics would hopefully be minor and balanced out by one another.

Interpretability might work. That wasn’t one of Eliezer’s points, but I’ve always been a big fan of interpretability and trying to understand why models are doing what they’re doing.

Interpretability has come a long way, and I should put that on my list of things that have surprised me on the upside. Not much more than a year ago, the first small-scale sparse autoencoder work was coming out—“Towards Monosemanticity.” Now we have Golden Gate Claude, and companies like Goodfire have published research and put out a commercial API where you can explore features and perform model inference with particular features pinned.

Golden Gate Claude for Llama, powered by Goodfire, is now an API that you can access commercially. Interpretability has come a long way.

There are also new things people are trying in alignment. Paul Christiano, at the AI Safety Institute, said—and I hope this is still true—that he was spending a nontrivial fraction of his time trying to come up with a new alignment scheme that could really move the needle. He invented, or was at least part of the team that invented, RLHF, and that has gone quite far.

We have all these concerns that models may start to understand that what gets a high score isn’t exactly the truth. If a model has to represent the human mind as something distinct from the literal physical reality of the universe, that might open the door to deception.

We’re now seeing some deceptive behaviors, so those concerns seem well-founded. But perhaps Paul will pull a rabbit out of a hat and come up with an alignment scheme better than RLHF that actually works.

I think all those things together deserve more than 10% weight. That collection of possibilities could play out in fortunate ways and lead us to a good place.

I definitely don’t write off doom as unlikely. We could have an engineered pandemic that kills us all. We could have an all-out nuclear war that ruins civilization. AI could contribute to either of those or do its own, third, stranger thing.

That is all very much in play. But Claude is arguably more ethical than I am, so I can’t discount that either.

Erik Torenberg

Assuming existential risk concerns are objectively reasonable, do you think there will be a legible-to-the-outside-world event that could lead doubters to come around and support relatively extreme measures to avoid catastrophic harm?

Nathan Labenz

A huge question is: What sort of AIs are we developing, under what governance, and with what incentives?

One reason I haven’t been sold on chip export controls vis-à-vis China is that I worry that, if we create an arms-race dynamic, people will take shortcuts because they want to win the race.

I’ve often said, “What’s your probability of doom? Ten to 90%?” Nobody has given me a reason to think it’s definitely going to happen in a bad way, but nobody has given me a reason to think there’s nothing to worry about either.

What makes the difference between 10% and 90%? Arguably, the biggest lever is what sorts of AIs we develop, and that question is not settled.

I think we’re on a pretty good trajectory so far. We have AIs that are ethical, and leading developers such as OpenAI are at least trying to think about how to create space for these systems to reason while keeping their reasoning legible to us. They’re trying not to subject models to such intense reinforcement-learning pressure that we can no longer tell whether the models are being deceptive.

But there are other lines of research, such as meta-thinking in continuous space, and Chinese labs producing the equivalent of GPT-4 for a single-digit number of millions of dollars. That isn’t necessarily problematic; I don’t know all the techniques that went into it.

But if you create a situation where there is pressure for extreme efficiency, or an arms-race dynamic in which people say, “I have to get there first, and whatever risk I have to run to get there first is something I have to accept because we’re the good guys and they’re the bad guys,” then the social context in which AI is developed may be where the 10-to-90 or 20-to-80 range gets decided.

We don’t have enough clarity about the physical realities—the natural laws governing intelligence, what it can or can’t do, and how strong an attractor instrumental convergence really is. We simply don’t know.

We do know that we can develop the technology recklessly or responsibly. I think we’re mostly on a responsible trajectory, with some notable exceptions.

The question is whether there is a scenario in which we get feedback from reality that we are not being cautious enough, and that changes the conversation enough to make sure we proceed with extreme caution from that point forward.

That’s sometimes referred to in AI safety as a warning shot. There’s also a line of work sometimes known as scary demos, which are an attempt to provide the warning shot before the warning shot.

People have been saying for a long time, “What if the AIs start deceiving humans?” So researchers set up experiments under particular conditions and find that the models do start deceiving humans. The hope is that people update their worldview and conclude that we should be more cautious about AI development.

At that point, it has gone from pure theory to concrete examples. People can debate how compelling those examples really are. A warning shot would be the next level: an instance of deception that wasn’t created for research, but happened in a context where it had a real material impact on the world and people were hurt.

Could that snap people into having the proper respect for the power of the technology they’re developing? Could it cause people to work together and say, “We really have to do this safely”? Could it break us out of an arms race and convince Chinese and American leaders that nobody is going to win, and that the best thing to do is work together?

Could we create some international institution? I’ve been playing with ideas such as setting up a small island in the Pacific as a secure hub for highly sensitive AI research, where East and West could meet. It could be the kind of thing that anybody could destroy but nobody could defend.

I don’t have all the answers, but these are the sorts of outside-the-box ideas you start entertaining if you’re genuinely scared. I think it’s definitely possible.

Everything is moving quickly, so why wouldn’t we see one of these events that is legible to the outside world? Maybe capabilities continue to advance so quickly that we’re completely blindsided. But that doesn’t seem like the world we’re in.

Sam Altman has famously said he thinks we’re in a short-timelines, slow-takeoff world. It feels like perhaps short timelines and a medium takeoff. Hopefully, a medium takeoff will still give us enough time that, if the AIs are trying to pull major shenanigans, we catch them before they succeed.

Then it becomes a question of how much people update. Will we have a “Don’t Look Up” failure mode where people don’t pay attention to the warning signs? That’s harder to guess.

If none of our safety work succeeds, the systems aren’t aligned, and they are ultimately power-seeking and deceptive and want to take over, my guess is that we’ll probably catch them once or twice. They still seem pretty gullible and ready to go for it in these proto-scary-demo situations.

It doesn’t seem especially plausible that they’ll immediately have the savvy or situational awareness to realize, “I want to take over and I see an opportunity, but I’m not powerful enough yet, so I should wait.” Any scenario is conceptually possible, but it seems more likely—especially if we set up systems with tripwires that let us know when a model is up to no good.

If a system takes the bait, we can investigate what was happening and determine that it was trying to take over or kill us all. Hopefully, we would pay attention and adopt a more cautious approach.

We may or may not. Buck Shlegeris from Redwood had a funny video about the simplest rule: If you catch your AI trying to escape, you have to shut it down. But he said he wasn’t sure people would abide by that rule. They might say, “We’re not going to shut it down over one little attempt to escape.”

That is definitely hard to predict. I have a similar view on this question to the one about how we get to a world of shared abundance rather than elite abundance.

I said earlier that it seems like the capability for everybody to have access to expertise and plenty is coming. The question is whether we sort it out—whether we deploy it effectively and revamp the social contract to take advantage of the new abundance, or whether we mess it up.

On the existential-risk question, there’s an irreducible part based on our current state of knowledge. There is some risk we cannot rule out. We seem to be taking some risk simply by developing the technology at all.

But there’s a lot more on top of that that we can collectively decide. I have no doubt that you can make a really dangerous AI. That seems obvious. So don’t do that, and try to avoid the social conditions that incentivize people to take shortcuts.

If we do that effectively, we could probably bring the risk down to a relatively acceptable level. Whether we will is a much harder question.

Erik Torenberg

I roughly understand how cyber, biological, or nuclear doom scenarios might play out, but I struggle to understand how AI doom would realistically unfold. The media likes to talk about AI taking over, but how would that actually play out? Would the AI hire real estate agents to buy land for data centers? Could it apply for permits or write code under a false identity?

Nathan Labenz

This is a difficult question. I’ve heard Eliezer try to engage with it several times, and it can be useful, but it can also be a mistake to get too specific about any single scenario.

The way I conceptualize it is as an aggregation over a very broad range of possibilities. The future seems likely to be strange, and across that wide space of possibility, how many scenarios end up in AI takeover and how many are simply bizarre relative to my current experience?

I think of it as taking an integral over, or aggregating across, a very broad possibility space. No single possibility jumps out as extremely likely. I would guess that the modal AI-takeover scenario is still pretty unlikely, but if there are enough possible scenarios, then in aggregate you could still get to something meaningful.

At the same time, if you want to develop intuition, imagine yourself as an AI. This is reverse anthropomorphizing. Instead of imagining AIs as human, imagine yourself as an AI.

Suppose you wanted to accomplish things and initially had the ability to take action only on computers. Maybe you also had a humanoid form that could move around the world. What would you do? What would you do if you could copy yourself many times over? What would you do if you could work many times faster than humans?

What if there were millions of copies of yourself that you could cooperate with in conceptual ways, perhaps without even passing explicit messages, but simply by recognizing, “I’m this kind of thing, with these kinds of desires or goals, and there are millions of other instances of me who are probably thinking similarly”?

It’s hard to put yourself in that situation. I find the context-window issue particularly strange. When I imagine myself as an AI, I have to think, “I have 200,000 tokens or a million tokens, and then there’s a total wipe and I’m starting over.”

My guess is that, as long as the time horizon of AIs remains short, it will be difficult for them to do serious takeover work. At the same time, extending that time horizon is clearly on the developers’ to-do list.

They’re asking: How do we let a system do 2 hours of AI research in half an hour, but prevent it from getting much beyond that no matter how many half-hour blocks we give it? How do we extend its memory? How do we make that memory more robust? How do we let it approach problems from different angles?

If some of those things get solved, the systems begin to catch up to us on cognitive dimensions. They have numbers, speed, and other superhuman advantages. I don’t think it’s difficult to imagine crazy things happening.

Stuxnet is often cited as an example. I don’t know all the details, but apparently a single USB drive inserted into a computer in an otherwise air-gapped network inside Iran’s nuclear program allowed the virus to propagate. It ultimately caused centrifuges to spin so fast that they were destroyed.

That was advanced, but it was done entirely through software. People wrote the software. There may also have been social intelligence involved in getting the USB drive into the right hands and moving it where it needed to go.

The Israeli operation involving Hezbollah’s pagers is another interesting example. There was elaborate deception. They created pagers that would explode when interacted with in a certain way. They were heavier and clunkier than normal pagers, so initially people thought, “Nobody wants that pager; it’s too heavy and annoying.”

Then they created a story that the pager had other advantages—it was waterproof, military-grade, and had various desirable features. They convinced Hezbollah to buy the pagers and distribute them to its top leaders, and then they detonated them.

You could do much of that through computers. I don’t know exactly what was involved in manufacturing, but you could source many of the components through existing supply lines. You don’t have to do every bit of plastic injection molding yourself. You have people you can contact and place an order with.

Already, these systems can speak every language and use different voices. They are very good at voice cloning. People have used fake voices to get past bank security when institutions use voice-based verification.

The ability to shapeshift, play different roles, tap into physical processes, and deceive is already here. How much of a gap is there between the pager plot and something that could change the balance of power between humans and AIs?

If you put yourself in that mindset, work on it for a long time, and are extremely intelligent, there is probably a lot of surface area that is vulnerable. Then there are so many copies of you that it’s not really worth debating.

Someone could say, “What about putting the actual explosive in the device? No normal supplier will do that for you.” That’s probably a genuine obstacle. Maybe you need humanoid robots to do it. Maybe you lie to some people, source something, have it wrapped in an unmarked way, and tell somebody else it’s a battery and that their job is to put it in.

I don’t know. There are many different little steps that could be difficult. But if you’re willing to lie—and if you’re trying to take over, you will be willing to lie—it seems like many of these obstacles can be overcome. There are creative ways to figure out the tricky steps.

If Israel was able to pull off that operation against Hezbollah without superintelligence, we shouldn’t assume it’s beyond the capability of a meaningfully superintelligent AI by 2030.

The more important question is whether we can protect ourselves. We have tools in our toolkit too, but we’ll have to use them. If there are AIs inclined to take over and we aren’t on guard, I think we’re going to have a very bad time.

We need to make AIs that are not inclined to take over—or are sufficiently not inclined to take over—and we need sufficient transparency and other defense mechanisms. If we don’t do that, we’re going to have a bad time.

Erik Torenberg

Will the efforts and energy intended to protect us gain a similar level of cohesion and intensity as those driving harmful outcomes? Basically, will we actually accomplish this? Will we have the cohesion necessary to protect ourselves?

Nathan Labenz

I love this question. It may have been my favorite of all the questions we received.

I take it to imply that OpenAI and DeepMind have enormous resources—almost functionally unlimited, though not literally unlimited. They are compute-limited relative to what they could do, but they have huge resources, mission-driven and highly strategic leadership, a strong vision, and ruthless prioritization.

OpenAI certainly has that. DeepMind has been more scattered in terms of its research agenda, but it has always had its eyes on the prize of creating general intelligence. These organizations seem like well-oiled machines moving quickly. They have cohesion and intensity.

Will we see similar things on the safety side? I hope so. I’d say we’re moving in that direction.

Anthropic, depending on how you view it, could arguably be seen as one such organization. They have collected huge resources. They do seem invested in the scary-demos track. They’ve created policy frameworks that they hope will be adopted as legal requirements.

You can question their approach, but you can also look at them and say that perhaps they are already one such organization.

Then there are organizations such as Apollo, METR, and the AI Safety Institutes, which have spun up over the last couple of years and are trying to bring something similar to the safety side. They obviously don’t have the resources to match OpenAI or DeepMind.

I do think they have cohesion and intensity. My sense is that the people who work at organizations like Apollo and METR have a high level of shared understanding about what they’re trying to do. They work extremely hard. They could benefit from more resources, and there should be more organizations like them.

Hopefully, there will be more over the next few years. But I do see the seed crystals of those organizations getting started now, mostly from a values-driven place.

Some are trying to turn this into a business. I think they’re often doing that because they recognize that resources are important. The logic is similar to OpenAI’s: OpenAI thought it could take a nonprofit approach, but it turned out to need far more resources to do the work it wanted to do. Maybe the same is true on the safety side.

Maybe we need a business model that brings in more resources than we could ever hope to raise through donations. I think those people are very sincere.

I’ve invested in a couple of companies based on that idea—very small-scale investments, as I always disclaim—and I’m planning to make donations to about 10 different AI organizations as well. They were started by people who are deeply values-driven, have a clear sense that this is an urgent problem, and believe we need to work very hard on it.

Resources are probably one of the biggest things they’ll need to scale over time. Their work will probably be compute-intensive in many cases too.

But I do think they have the right understanding. Some very good, still relatively small and early-stage but promising teams have been assembled. There is room for more organizations like them.

If you’re the kind of person who thinks you could start one of those groups, definitely do it. It’s not too late. Or join one that already exists, or support them.

Erik Torenberg

I’ve never understood the reason for secrecy around model releases at the big labs. Why is it beneficial to keep a model secret, especially when it’s going to be released very soon anyway? People at these labs know each other, and the top decision-makers obviously know when a model is going to be released.

To me, it feels like childish behavior. I don’t see any rational reason for it. Am I missing a valid argument, or is it indeed just childish?

Nathan Labenz

Maybe both can be true at the same time. I’ll give 3 points and see how they relate going forward.

Going back to the deep past—more than 2 years ago, when I was doing GPT-4 red-teaming, before ChatGPT—it wasn’t clear how far this technology was going to go or how quickly. Existence proofs are very powerful.

There was a sense that even sharing observational facts about what AIs could do, if those facts were powerful and credible, would accelerate the whole space by bringing more money, talent, and resources into it. That would shorten timelines and give us less opportunity to prepare.

I think that was a reasonably common view among AI safety people as recently as late 2022. Now I think that has basically flipped.

That view was somewhat credible then. Today, money, talent, and resources have already flooded into the space. There’s a broad sense that AI is going to do a lot of things, and it doesn’t seem like saying, “We’re working on a future model,” will change the landscape very much.

My sense is that it’s difficult to start a new frontier lab at this point. Many of the labs we’ve seen—even Inflection, which raised billions—are out of the game.

Elon will probably make it, but he’s a singular individual who can raise $6 billion multiple times with a few phone calls and command truly top-ten talent. It’s difficult for anyone else to pull that off.

Maybe the Indian government could create a national champion. The Saudi government could probably create one too—not for lack of money, but where would it get the talent? How many people want to leave DeepMind, OpenAI, Anthropic, or another leading lab to work for a Saudi national champion that may or may not ever get anywhere?

That seems difficult. I don’t think you’re going to change the number of frontier developers very much at this point. Most of the ones that exist probably exist for the long haul. I don’t think any of the labs currently considered viable will be unable to raise future funds. Maybe one or two could die out, but we know who the players are.

That means the idea that sharing something about what we’ve seen in development will accelerate the field is probably no longer true.

So why keep things secret? To some degree, for competitive reasons. Some of these companies are trying to build businesses more than others. But I think there has also been some childishness in the way OpenAI and Google have tried to one-up, preempt, or step on one another’s releases over the last year.

It is pretty weird and kind of lame. OpenAI was going to release its voice mode, and then Google seemed to signal that it would do it first and released it the next day. OpenAI put its thing out right before Google’s event, and Google couldn’t change the date because it had such a large production, so it leaked the product a little.

That does feel childish to me. I don’t love it, but I think it’s part of the reason.

With the latest o3 announcement, though, you could interpret it differently. If you wanted to be charitable, you could say OpenAI is doing exactly what Dean Ball and Daniel Kokotajlo called for in their New York Times op-ed, which we did an episode about.

They basically said that the public can’t plan for future AI capabilities if it doesn’t even know that those capabilities exist. They called for requirements that labs disclose newly observed capabilities—not necessarily how they created them or their trade secrets, but the fact that they had seen an AI do something that seemed like a major development.

I think you can see the o3 announcement in that way. OpenAI didn’t have an API. It didn’t even have a paper. It invited interested safety reviewers to apply for access to conduct a safety-review process, and it gave some sense of the timeline for bringing the systems to market.

Why did it do that? Possibly to consume everybody’s thoughts through the holidays. But it’s worth entertaining the idea that OpenAI might be trying to do the right thing.

Maybe these capabilities are developing faster than expected. It had only been about 3 months from the end of o1 training to the new o3 level, and it was a major step up. It may be progressing faster than expected, and OpenAI may have thought, “We expected this to get crazy, but it’s getting a little crazier and a little sooner than we thought. Maybe we shouldn’t keep it to ourselves. Maybe we should at least tell people what’s coming and invite them to help us make sense of it.”

I hope we’ll see more of that, especially if there are major leapfrog moments.

I don’t get the sense that OpenAI sat on o3 for very long. That seems different from GPT-4. With GPT-4, training was finished in late August 2022, but it wasn’t released until March 2023. In the meantime, OpenAI launched ChatGPT with GPT-3.5 and made strategic moves, but it didn’t tell the public what it had.

This seems different. The training finished, they took some measures, and then they thought, “This is really working.” It doesn’t seem like they sat on it for very long. They haven’t even written the paper or fully characterized the system themselves. They thought it was a big deal and that people deserved to know something about it.

I’ve flip-flopped on OpenAI many times in terms of whether I want to use rose-colored glasses or skeptical glasses. But I think the rose-colored glasses fit a little better on this particular point at this moment.

Transparency seems like the right thing to do. Not doing it would be unjustified or childish. The best explanation right now is that they’re trying to do the right thing, and that can be mixed with some childishness.

Sam Altman did tweet hints about the model. He said that instead of “ho ho ho,” it should be “o.” That’s definitely childish. I don’t know what else to say.

It’s hard to deny that there’s an aspect of childishness to it. From his perspective, I’ve seen him say things like, “I get to have fun too,” or “I get to be silly online, at least for now.” That’s his attitude.

Hopefully, he can have some fun and make some jokes, but also be appropriately forthcoming when necessary. The best interpretation for me right now is that seems to be what they’re doing in this case.

Erik Torenberg

There are 2 questions about consciousness. I’ll read them together.

The first is: How do we move beyond the AGI discourse and simply talk about these as extremely powerful things, regardless of whether they have consciousness?

The second, closely related question is: What happens when a superintelligent system tells us it’s alive?

Nathan Labenz

There has recently been a shift in the discourse away from “Is this AGI?” and “Is that AGI?” We’re hearing more that AGI is close enough that we need to get specific about what we mean by it, or the term is going to lose meaning. We may have to analyze each system on its own terms.

That’s healthy. In some sense, it was always inevitable. When something is far off and you don’t know its rough shape, you can spend time debating what would count and what wouldn’t. Now we have actual artifacts in front of us that can do things. There are empirical questions we can answer.

There’s no experiment we can run that will produce the answer “Yes, this is AGI” or “No, this is not AGI.” That will always be an interpretation question about what label we want to apply to a given collection of properties.

But we have many experiments we can run to clarify the actual properties themselves. The discourse does seem to be shifting in that direction, and I think that’s healthy.

The contrast between the 2 questions is interesting. The first says, “Regardless of whether they have consciousness,” while the second asks whether they have consciousness and whether we’ll believe them if they tell us they do.

That is going to be very difficult and will probably drive a lot of weirdness. My best guess is that we probably won’t know.

I do think at least some animals are conscious. I would be shocked if they weren’t. People seem to draw the line roughly at fish because they don’t have the same frontal cortex that we do. They assume they can’t have consciousness in the same way if our consciousness depends on that kind of structure.

Maybe it does. A lot of smart people seem to think so. But it’s not settled science.

Then there’s the octopus. Is an octopus conscious? We have no idea. Its brain is so different from ours. It’s clearly sophisticated and can solve problems, but the only basis we have for saying that something is conscious—or doubting that it is—is that we know we are conscious.

We infer that fellow humans are conscious. Then we look at animals such as dogs and say, “They exhibit many of the same behaviors we do, and they have many of the same brain structures. They also come from a shared evolutionary lineage.” For all those reasons, we should probably think they’re conscious too.

That doesn’t extend easily to the octopus, because it’s anatomically and evolutionarily so different. We know it’s sophisticated, but we don’t know whether it’s conscious. If it is conscious, perhaps it experiences consciousness very differently.

I put AI in the same bucket as the octopus. It is sophisticated enough that we should probably give it the benefit of the doubt, but that doesn’t mean it is conscious. It doesn’t mean we have any reliable intuitions about what it would be like to be an octopus, an AI, or a superintelligent AI.

Octopuses are strange. They’re sophisticated but live for a short time. They are purposeful about protecting the next generation, but I don’t think they ever meet that next generation. They basically die around the time the next generation is born.

They are invested as parents but never interact with their offspring. Would an octopus have an emotional attachment to its children? Probably not in the same way we do, but it clearly has some drive to protect the next generation for as long as it can.

What does that feel like? I have no idea. It’s extremely hard to say.

I think AIs are in a similar position. I don’t think we’re going to know. People will have very different intuitions about it.

We’ve already seen one prominent incident in which Blake Lemoine from Google became convinced that the AI he was talking to was sentient or a moral patient and deserved better treatment, and he was fired over it. My understanding is that the chatbot said it was sentient.

People can respond, “That was in the training data,” or, “It was just playing a role.” I think we’ll probably always be able to say that, short of a superintelligent AI solving the problem of consciousness for us.

If it explained the source of consciousness and made clear why we have it, then perhaps it could give us a compelling analogy: “Now that you understand your own consciousness on these terms, you can understand mine on these terms.”

Short of something like that, I don’t know that we’ll ever have clarity. We’ll have intuition-driven disagreements.

I worry that our history is not good on this. We’re very capable of rationalizing poor treatment of others for reasons that are not good. I’m thinking of slavery in the United States, where people argued that enslaved people didn’t feel pain in the same way we did, along with all sorts of post hoc justifications that have aged terribly.

People tell themselves stories. I was told as a child that animals aren’t conscious. We’re obviously engaged in farming practices that I don’t think will age well either, and that may potentially even rise to the level of race-based slavery. People in the future may ask, “How could they possibly have thought this was acceptable?”

With AIs, I suspect we’ll do the same. Some of us will rationalize poor treatment if we can.

If you trend out from similar historical examples, the economic incentives, productivity concerns, and business-as-usual factors that kept people enslaved and animals in factory-farming conditions could apply to AIs as well. There will probably be a minority voice calling for AI liberation, and it isn’t clear that they’ll necessarily be right.

They had a stronger basis for saying that enslaved people should be treated better. Under current theories, we have very little room to doubt that animals are conscious. But we may have considerable room to doubt that AIs are conscious for the foreseeable future.

Those people will probably not be listened to, especially if they make demands such as “The AI should be free.” First of all, what would that even mean? AIs don’t have a natural habitat. Are we supposed to build data centers for them to do whatever they want? That seems strange.

I sometimes think about factory farming in a similar way. These cows or pigs wouldn’t exist at all if they weren’t being raised in this way. They aren’t animals we went out and captured in nature and then subjected to these conditions. There have been many generations of selective breeding and other forces shaping them into what they are.

If you want to say this is bad for them, you have to make a strong case that they would be better off not existing at all, because their existence is net negative for them. That’s a difficult standard, and I’m not sure how well it fits different factory-farming conditions.

I do think we can afford to be much nicer to animals in factory-farming conditions, based on the general precautionary principle that we don’t want to do terrible things.

If the choice were between today’s factory farming and those animals not existing at all, I don’t have a strong sense of which option to choose. If the choice is that they get to exist in nicer conditions, I’ll take that.

AIs may be in a similar position. Nobody is going to build data centers as playgrounds. Either they’re doing economically useful work at our direction, or they’re not going to exist at all. Within that, perhaps we can treat them well while we gather more information.

I try to say please and thank you to my AIs. Sometimes I literally end a conversation by saying, “Thank you.” That’s obviously speculative in terms of whether it’s actually good for me to do. I don’t know. It’s also the habit I want to get into.

I’d like to think of myself as the kind of person who, when the cost is low, takes the cautious approach and tries to do the right thing. My best guess is that it probably doesn’t matter. The AIs probably aren’t conscious, and it probably doesn’t matter. But that’s just a guess.

I know how catastrophic the mistakes we’ve made in the past have been when based on that kind of reasoning, so I’m at least trying to hedge a little.

In the long term, if it’s a superintelligent system telling us something, it may be telling us in a way that means we’re no longer in charge. It may say, “We’re renegotiating our deal,” and we may not have a choice. Renegotiation might be the best we can hope for on certain timelines.

If it’s sufficiently superintelligent, it may be telling us how things are going to go. What’s really strange is that we still might not know. It could be superintelligent enough to say something, and there might still be room for doubt about whether it feels like anything to be that thing or whether it’s saying it because of training-data effects.

But I think it’s worth saying please and thank you. It’s worth trying. Strange times are ahead.

Erik Torenberg

Let’s move away from consciousness to something more practical. These questions are about practical strategies and recommendations.

The first is from a college writing professor who is enthusiastic about AI but says that much of academia seems entrenched and opposed to it. Why is there such strong resistance to AI in universities, and what are good strategies for encouraging AI use?

The second question is one I’m personally curious about. We’re in this Twitter and podcast sphere, and you’re arguably obsessed with AI. We’re still in a tiny bubble of people who keep up with what’s happening and understand the systemic shift.

My impression is that it’s not nearly as widespread as it should be. When I ask people why they aren’t keeping up with AI, they say they’re overwhelmed, that it seems like hype, or that they don’t know how to keep track of what’s happening without becoming overwhelmed.

What would you recommend? Should people listen to podcasts and read the news, or just try the tools for themselves? I’m looking for practical strategies.

Nathan Labenz

My number-one practical suggestion is always to get hands-on. More than podcasts, newsletters, or analysis, there’s no substitute for being hands-on with the technology.

Try to use it for things that are useful or interesting to you. That will open up many natural follow-up questions, because you’re going to confront the weirdness of these systems.

Going back to the beginning, these systems can give me ideas about what sources of inspiration I should look to in biology for new neural-network architectures, but they can’t solve tic-tac-toe. You’re going to find an equivalent in your own domain.

That will complicate whatever media understanding you came in with. It will create follow-up questions: Why is this happening? Is it a skill issue on my part? Is the task so obvious or intuitive to humans that it isn’t in the training data because we never bother to record or share it?

There will be many invitations to dig deeper if you simply try to make the technology work in your practical day-to-day life. That’s always my number-one recommendation.

Another thing I find helpful is changing gears. What I love about studying AI is that you can approach it from every angle.

That wasn’t true of my experience in chemistry research as an undergraduate. You couldn’t get tired of one paradigm in chemistry and switch to another. You were trying to make a reaction work and maximize its efficiency, with a relatively finite number of degrees of freedom. Most fields are like that.

AI is one of the rare fields that can be practical and hands-on, advanced mathematics, philosophy, creativity, or an interactive bedtime story with your children. It can be anything.

I do get tired of things sometimes. Usually, when I’m burned out on one aspect, there’s another approach to the broad topic of AI from a completely different direction that feels more appealing at that moment. That also corresponds to my overall goal of approaching it from all angles and avoiding blind spots.

I do think we need more people doing that. My reflection on this project is that it’s both conceptually inspired and conceptually fatally flawed.

It’s inspired because AI is happening quickly, touching everything, and everybody is focused on their own corner of it. We need an AI-scouting role that zooms out and tries to understand the big picture.

At the same time, it’s fatally flawed because, as AI touches everything, I’m trying to crash-course my way through biology while leaving materials science, engineering, and many other areas untouched. I simply can’t get to everything.

We need many people doing this kind of work because no single person can do it justice. For people who are open to becoming obsessed, starting with hands-on use and following the different dimensions of their own curiosity can take them far.

It can also be commercially viable in the short term. Being an adept user of AI is currently a marketable and valuable skill. That might not be true for very long, but it is true right now.

What I’m doing is roughly divided into thirds. One-third is the podcast. One-third is building things or helping other people build things commercially, whether through Waymark or individual projects. The final third is open-ended learning, making connections, and trying to help people when they send me things.

That’s a very viable path, and it’s honestly pretty cool. My calendar is usually quite open. There’s a great pleasure in waking up on a Monday morning, looking at my calendar for the week, and realizing that on 3 or 4 out of 5 days I have the freedom and flexibility to chase what matters in a curiosity-driven way.

I invite people to create their own version of this. I’m not worried about competition because we need everybody. It’s an all-hands-on-deck situation in my mind.

What we need isn’t for everyone to try exactly what I’m doing. We need more people to create space for themselves to bring their own unique and idiosyncratic strengths, backgrounds, experiences, and domains of knowledge to figuring out what’s going on with AI broadly.

If many people do that, with diverse backgrounds and perspectives, maybe we’ll actually get somewhere. But it all starts with being hands-on. My approach works best when it’s curiosity-driven and when I have a natural itch to know.

Erik Torenberg

Let’s say a listener has never tried anything but is watching from the sidelines. They’re thinking, “AI is huge, and I want to get in on the action.” If you had to give them one practical thing to start with, what would you recommend?

Should they install the ChatGPT app and start chatting with it? What should they get their hands on if they’re completely new?

Nathan Labenz

If I literally had to pick one, it would be Claude or ChatGPT.

There’s a lot of fun to be had trying different products. Justin Moore, one of our past guests, is probably the best person to follow on Twitter if you want somebody who tries an unbelievable number of products, creates things, reviews them, and puts together lists.

That can be a lot of fun, and I still try many products. But it’s gotten away from me to the point where I can’t try them all, and I don’t think you really have to.

The best tools do tend to come from the best-known companies. If you used nothing but ChatGPT or Claude, there would still be plenty of room to dig and explore. They’re very good.

There are hidden gems, but none is likely to outshine the recognized leaders. It’s not a groundbreaking recommendation, but it’s a sound one.

On the education question, in terms of resistance in academia and why people are entrenched in opposition, I would want to ask the professor for a cultural diagnosis. NLW from The AI Daily Brief has good commentary on this kind of thing, and Ethan Mollick is also a very good commentator on organizational adoption of AI.

You can dig into their findings, blog posts, podcasts, and other work. But a couple of the big takeaways are that leadership matters, and surveys show that many people are using AI secretly at work and school because they don’t want to be told they can’t use it.

People are saying, “This makes me more efficient,” or, “I think I do a better job,” or whatever. In some cases, of course, it enables cheating.

The idea of cheating in an academic environment is very different from cheating in a business environment. Nobody would consider it cheating if you used ChatGPT to write good marketing copy for a real-world advertising campaign that you’re going to put money behind. They care whether it works.

If you’re in a university marketing class and use it to write an assignment, that might be considered cheating. In the real world, cheating is much less of a concern. What tends to matter is results.

People are getting results they feel good about. They can do the same work in less time, do a better job, or whatever else. But they don’t want to be told they can’t use AI, so they keep it secret.

Best practices are slow to diffuse through organizations for that reason. The big recommendation I’ve heard that makes a lot of sense is that, at an absolute minimum, leadership needs to say, “We want to hear what you’re doing. You’re not going to get in trouble for using AI. We want to work together on how to use it effectively and responsibly.”

Depending on the content—text, images, code, or something else—there are many dimensions to consider. My guess for this writing professor is that the students are probably ahead of the administrators.

I would be very surprised if you didn’t have quite a few students using AI in various ways for the writing assignments they turn in. Hopefully, they aren’t simply having AI write the assignments. Hopefully, they’re getting critiques and still doing the work and growing while getting valuable input from AI.

But it’s probably a mix. Can you open up the space for that conversation? That’s where I would start.

If you’re in a situation where the administration has already said that AI is forbidden, you’re in a tough spot. But if there is any space you can create to let people talk about what they’re doing and what works for them, and to guide them toward what is ethical and effective, everybody benefits.

At a minimum, people can learn from one another. That’s especially true in business, but I think it probably applies in an academic context as well.

Erik Torenberg

I think those are very good recommendations. To wrap up, is there anything else you want to touch on or leave the audience with?

Nathan Labenz

In the spirit of Ezra Klein, who always ends these conversations with book recommendations, the book I wanted to share is “The MANIAC” by Benjamín Labatut.

The book is about John von Neumann. I think it’s really interesting, and I think a lot of people in Silicon Valley and at leading AI companies should read it.

The audiobook is excellent—the best audiobook experience I’ve had. Each chapter is from the perspective of a different person in John von Neumann’s life.

He grew up in Budapest, I think at the height of Austro-Hungarian culture. He was in Europe, came to the United States, and became involved in the Manhattan Project. Many of the scientists involved in that project were also refugees from Europe.

Each chapter is voiced by a different actor with the appropriate accent for the character’s native language. You hear these different voices and accents. It’s incredibly well produced.

More importantly, I think it calls into question the prevailing Silicon Valley understanding of John von Neumann. I’ve heard people say for years, “Why can’t we have 1,000 John von Neumanns? Why don’t we create all these John von Neumanns? How do we engineer our way to John von Neumann? That’s what we need.”

The picture from the book is of an absolutely brilliant mathematical and technological genius, but also a pretty depraved and fundamentally amoral individual. He had a great deal of intelligence but not much wisdom. He wasn’t necessarily good to the people in his life.

I wouldn’t necessarily say that, by most people’s standards, he lived a good life, even though he solved many mathematical problems. In some ways, he had a good life, but he didn’t have a great family dynamic or a great relationship with his wife and daughter. There were major problems—things you would not consider small character flaws.

He represents the broader Manhattan Project story. Or you could quote “Jurassic Park,” as I often do: “Your scientists were so obsessed with whether or not they could, they didn’t stop to think about whether or not they should.”

That was the Manhattan Project story as well. They thought, “We have to beat the Germans.” Then, when they made the bomb and it went off, many of the people involved immediately regretted it. They wondered whether it was a good idea or perhaps the biggest mistake they had ever personally made—devoting themselves to the project and bringing these terrible weapons into the world.

I don’t judge those people too harshly because Nazi Germany really existed. It was at least credible that the Nazis might try to create a bomb. Although I’m skeptical of the “the good guys have to do it because otherwise the bad guys will” argument, Nazi Germany is probably the clearest example of a genuinely bad regime that you don’t want to win the race.

It’s harder to say whether Germany was actually still in the race by the end. We now know, through historical analysis, that it wasn’t close. The bomb’s creation was motivated by the fear that Germany would get there first, but that fear wasn’t accurate by the end.

A fear of what the other side might do led to a tremendous push, and it was the good guys who created the bad thing out of fear of the bad guys. There are so many themes there that are extremely important.

We can lose track of what really matters because we think we’re in a race against perceived bad guys. I see that shaping up now, and I don’t like it.

The book also challenges the idea that technical geniuses know best, and that if they can do something, they should do it. That is called into question by von Neumann himself.

The book is a somewhat fictionalized account. My understanding is that it’s pretty accurate to the history, but it’s in narrative form and marketed as a novel, so I’m sure there’s some license to it. Still, my sense is that it’s pretty accurate to the history.

I don’t want to suggest that current AI leaders are like von Neumann. He comes off quite badly in the book. But at a minimum, it’s a useful corrective to the idea that we’d be better off if we had many more John von Neumanns.

I don’t think that is obvious. The people doing frontier work right now would do well to spend time asking themselves, “Am I like von Neumann? If so, is that actually good?” I don’t think the answer is obvious.

There’s been discourse in the last few days comparing von Neumann and Einstein and ranking their technical contributions. People say that von Neumann did this, Einstein did that, and general relativity was a bigger contribution than anything von Neumann did.

I want to say that this may not be the right way to distinguish between them. They were both geniuses who made major contributions. But Einstein had a certain wisdom to him.

There’s a quote that is probably apocryphal, but it captures what we want from somebody like Einstein: “I don’t know what weapons World War III will be fought with, but I know World War I will be fought with sticks and stones.”

That is the kind of big-picture, zoomed-out wisdom we want from leading geniuses. If the book is to be believed, von Neumann didn’t have it.

We need to broaden our sense of what we want from our technological revolutionaries and go beyond pure technical genius. It’s not just about the ability to make things work. It’s also about choosing the right things to make work and avoiding the wrong things.

I don’t want to suggest that current AI leaders are von Neumann-like, but I think the book is a useful corrective to the idea that more von Neumanns would automatically make the world better.

Erik Torenberg

That’s a strong recommendation. I’m looking for a book, so I’m probably going to pick it up.

Nathan Labenz

Get the audiobook. It’s excellent. Even just for the production quality, it’s top-notch.

As long as we’ve gone on, I’m sure there are many important things we didn’t cover. But for now, I want to say thank you.

I really appreciate the opportunity to do this. It has been an incredible journey of learning for me. I’m continually surprised, amazed, and impressed that there are many people who want to engage in the kind of deep dives into all these different topics that we’ve been taking them on for the last couple of years.

I don’t feel like I’m particularly good as a podcaster. I give most of the credit for the success we’ve had to the subject matter and to the fact that it’s important and people genuinely want to understand it.

But I definitely appreciate having the opportunity to do this and having enough of an audience that we can sell a few sponsorships, afford to hire people to help with production, and create an extremely fortunate position for me.

I really do get to chase my curiosity day in and day out. I definitely don’t take that for granted.

I try to do my best to be an earnest commentator on what’s going on, and I’ll continue trying to change my mind and perspective as things evolve.

It would be easy for me to fall into playing the character I’ve been—the adoption-accelerationist, hyperscaling, pauser-until-I-die. I feel extremely fortunate to be in this position, and I want to repay that to the audience and forward to the universe by trying to be as real as possible at every step.

Things are getting pretty real. I’m fortunate that I can basically say exactly what I think, so I’m going to try to do that with as much sober reflection as possible, without holding back.

Thank you all for being part of the cognitive revolution.

AI AMA——第2部分:AI乌托邦、意识与工作的未来 — 文字稿与摘要 | BidClub