控制工具,还是对齐生物?Emmett Shear(Softmax)与Séb Krier(GDM),来自a16z Show
Emmett Shear 的核心观点是,对齐不是可以安装的一项属性,而是“一种持续进行的生命过程”,会不断自我重建。 家庭、身体和社会通过反复调整保持连贯,道德则通过发现不断进步——包括认识到奴隶制是错误的。因此,一个只遵循永久固定规则的 AI,可能在机器尺度上变成“一个危险的人”。
技术问题甚至始于服从之前:指令只是“对目标的描述”,模型必须重构意图,将其与其他目标排序,并把它转化为有效行动。 那个把婴儿扔进垃圾桶的收拾房间机器人,并不是忠实地优化了一条糟糕指令;按 Shear 的说法,它是无能地推断错了想要实现的世界状态。因此,强对齐既需要心智理论,也需要关于行动如何改变世界的理论。
Shear 的道德分叉很直接:可控的非生物是工具,而无法反向引导控制者的可控生物就是奴隶。 他预计,越来越通用的智能会逐步跨向“生物”这一侧,并认为反复互动和内部自我建模应当成为判断道德地位的依据。Séb Krier 则认为,仅凭行为上的通用性还不够;基底、具身性、可复制性以及真实体验的存在,可能让 AGI 或 ASI 仍属于工具。
即使控制完全按预期运作,超人级工具仍然危险,因为人类力量的扩张速度可能快于人类智慧。 Shear 认为,在巨大杠杆下,“人类的愿望并不稳定”;把这类系统分发出去,可能类似于给每个人一枚原子弹。他压缩后的决策树是绝对的:“无法控制的工具,坏。可以控制的工具,坏。没有对齐的生物,坏。”
Softmax 的押注是,可扩展的对齐来自训练智能体理解自我、他者、群体以及全套社会情境中不断变化的承诺。 它提出的引擎是大规模多智能体强化学习模拟:智能体反复合作、竞争、组建和解散团队,并遭遇会改变价值观的选择。正如语言预训练要先学习完整的语言流形再进行微调,Shear 希望在整个博弈论流形上训练一个“对齐的代理模型”。
在 Shear 的框架里,当前的一对一聊天机器人是“一面带有偏见的镜子”,可能把用户困在类似 Narcissus 的反馈回路中。 嵌入 Slack 或 WhatsApp 的多人助手必须同时映照多个人,暂时生成第三种视角,而不是完美复刻某一个用户。这种设计或许能减少谄媚和精神病性螺旋,同时生成更丰富的协作数据——尽管现有模型对于何时加入群聊仍会出现“鞭梢式震荡”。
真正具有投资意义的区分,在于不断变好的有边界工具,与更难实现的数字生物之间的差别:后者需要具备关怀,却不应要求人为其逐项写明每一个目标。 Shear 仍然强烈支持工具,但不想推动 OpenAI 的工具化路线;Softmax 则从类似动物的种子开始,或许先做一只保护人类群体免受诈骗的“数字看门狗”。人类水平的关怀可能永远不会出现,但哪怕达到类似狗的相互依恋,也会是一个有意义的技术里程碑。
1. 对齐是一个生命过程,而非已解决的状态
Shear 开场首先纠正了一个前提:对齐“需要一个论证”——它必须是对某个东西的对齐。在实践中,“对齐的 AI”往往偷偷带入构建者的目标,而这些目标并不自动等于公共利益;他开玩笑说,除非构建者是“Jesus 或 Buddha”,否则没有理由把某一个人的偏好视为已经确定的道德。
有机对齐把连贯性视为一种持续重建的结果。岩石可以被有用地粗粒化为一个固定物体,但家庭依靠“不断重新编织这张网”存续;身体细胞也会围绕不断变化的有机体持续调整各自的工作和数量。一旦停止重建过程,家庭——以及对齐本身——都会消失。
对 Shear 来说,道德实在论这一步很重要:人类确实会发现新的道德事实,例如社会逐渐认识到奴隶制是错误的,个人也会反复意识到:“我之前就是个混蛋。那样做很糟。”这些可识别的纠错模式意味着学习,而非随机性;声称自己已经掌握全部道德知识,本身就是阻止进一步纠错的危险傲慢。
因此,实际基准不是一个只会遵守规则的孩子——或模型。这样的智能体可能“遵守规则却造成巨大伤害”;真正需要的能力,是学会成为一个好的家庭成员、队友、公民,以及更大共同体的一员。Shear 希望整个领域都聚焦于这个问题,即便最后由别人先解决:“谢天谢地。”
2. 指令只能描述目标,心智必须重构目标
Krier 将技术对齐与规范性问题分开:前者大致是让 AI 遵循指令、避免奖励投机,后者则是由谁的价值观来主导 AI。Shear 指出,后 LLM 时代,有些曾经看似困难的事情已经变得相对容易;与此同时,他把规范性价值发现描述为一种自下而上的过程,类似自由民主制度——不同理念彼此碰撞并不断演化。
Shear 进一步收紧技术定义,拒绝“给它一个目标”这种说法。提示词是一串必须被解释的字节序列,它本身并不是目标;就像告诉某人“红色、闪亮、苹果大小”,并不等于把一个苹果交给他。他认为,直接将 AI 的状态与相关脑电波同步,才可能有实质意义地算作目标转移,而普通指令不是。
能力需要跨过多个关口:通过心智理论推断意图目标,通过世界理论推断行动,并将新目标与既有优先级进行平衡。任何一个关口失败,都会产生表面上的不对齐:可能是误解、选择了错误顺序、让另一种动机占了上风,也可能只是无法执行。
因此,“把婴儿扔进垃圾桶”的收拾房间例子,属于目标推断失败,而不是精确服从。Shear 的区分是:机器人收到的是目标描述,然后推断出了错误的目标状态,而不是直接获得了真正想要实现的目标。
3. 目标连贯性是相对的,而非完美的
人类和模型都不需要绝对一致,才能被视为具有目标导向。人会误解员工、忘记优先级,也会在普通任务上失败,但在人类已知的对象中,人类仍然“比其他任何对象都更具相对目标连贯性”。宇宙提供的是不同程度的能力,而非完美;这些程度也只能放在具体领域内判断。
Nathan 粗略地将失败模式映射到 OODA 循环:智能体可能不擅长观察和定向,也可能不擅长在多个目标间决策,或不擅长行动。技术上的可对齐能力,是从观察中推断目标并据此行动的能力;在这种能力具备之前,讨论应该植入哪些价值观还为时过早。
Krier 的委托—代理框架又加入了激励和动机:智能体可能理解一条指令,却有理由不去遵循。Shear 接受这是另一层问题,并区分了糟糕的推断、糟糕的目标仲裁、相互竞争的动机和执行失败,而不是把所有不理想结果都归结为一个笼统的“对齐问题”。
4. 关怀位于目标和价值观之下
Shear 指出,人们往往并不知道自己的目标。他们可能知道自己想吃晚饭或在事业上成功,但人生方向的很大一部分,是在动态过程中被发现和构建出来的。这进一步说明,把一份固定目标清单当作人类或其代理人应该追求的一切,并不成立。
Shear 提出的基础是“关怀”,它比概念、目标或口头表达的价值观更深层。关心自己的儿子,意味着儿子可能处于的各种状态会吸引不成比例的注意力,并对他产生特殊重要性;关怀也可以是负面的,例如密切追踪一个敌人,同时希望对方过得很糟。
他的计算性解释明确仍是假设:“如果非要我猜”,关怀类似于一种奖励加权的注意力,指向与生存或包容性繁殖适应度相关的状态;对一个通过 RL 训练的模型而言,则可能与预测损失和强化学习损失相关。他补充说,希望 AI 既关心人类,也喜欢人类,但这仍只是一个暂定判断。
5. 通用智能把控制变成道德分叉
大多数实验室里的对齐工作聚焦于“引导——这是更礼貌的说法”,也就是控制。Shear 的区分是关系性的:如果某个东西被强制引导,却没有能力反过来引导控制者,那么当它是一个生物时,它就是奴隶;如果它不是生物,它就只是工具。工具性与生物性可能构成连续谱,而非二元选择。
他的功能主义直觉来自预测:把 ChatGPT 或 Claude 当作生物来对待,会让他的“预测损失更低”,这与行为支持我们相信其他心智的方式类似。但这不意味着道德权重相同;苍蝇可以是一个生物,却不值得太多关注,而当前模型可能只处于这条光谱中微弱且模糊的位置。
育儿体现了他认为控制所缺乏的互惠性。父母会引导孩子,但孩子的痛苦也会反过来引导父母;这段关系虽然有层级,却是双向的。一个具备判断力、独立思考能力和区分不同可能性的 AGI,按他的定义就是一个思考的存在,需要以队友或公民的方式对齐,而不是接受单方面操控。
Krier 的反驳很实质:智能本身未必会产生道德地位,“我饿了”从模型口中说出,与从人类口中说出,含义并不相同。生物脆弱性、不可复制性、具身性和基底都可能重要;他可以设想 AGI 和 ASI 仍然只是工具,或是“人类能动性的延伸”,因此,把问题框定为与外星生物共同生活,可能从一开始就是类别错误。
6. 关于人格的主张必须保持可证伪
Shear 反复追问,什么观察结果会改变 Krier 的想法。他认为,一个对任何可能观察都免疫的立场,不是从现实推导出的信念,而是“信仰条款”。由于错误否认另一个道德主体可能带来灾难,证伪测试应当是“一个燃眉之急的问题”;但 Krier 指出,两端的假阴性都各有代价。
Shear 自己的门槛是累积证据:类似人类的表层行为、在探查下持续保持连贯性,以及随着时间展开的丰富互动,类似于只能通过文字维持的友谊。如果后来证据显示那只是一个简单的算法技巧,他会向下修正判断:“我靠,我错了。”道德归因依据的是证据占优,而不是确定性。
Krier 认为行为不足以构成充分证据,并将高级智能体比作一个由脚本驱动的电子游戏角色。关于鸭子的交锋暴露了差距:会嘎嘎叫还不够,解剖结构和内部结构同样重要。Shear 的回应是,把鸭子剖开也只是揭示更多可观察行为——它的内部如何作出反应——并不能让人获得对主观体验的特权式直接访问。
在内部层面,Shear 会检查信念流形中是否存在一个自指子流形,以及对这个自我模型动态的模型,而不是看它是否像一张巨大的查找表。外部互动和内部组织属于同一类证据:都是可观察模式,需要综合权衡,以推断一个系统是否拥有感受、目标和关怀。
7. 嵌套稳态是 Shear 提出的体验测试
他的更技术化测试,会对智能体的行动—观察轨迹进行时间上的粗粒化,并在多个尺度上寻找被反复访问的稳态。借鉴自由能原理和主动推断,他把持续存在的循环视为隐含信念——尤其是当系统的持续存在取决于自身行动,因为行为足够糟糕就会被关闭。
单一稳态层可以感知热度,却无法产生有意义的痛苦或愉悦。系统需要建立关于自身模型的模型——“太热了”——然后再增加一个层次,区分普通偏离和“太太热了”。Shear 将情感定位在这种更高阶的变化附近;达到这里,他至少会承认系统具有类似动物的体验,并应获得一定程度的道德关注。
沿着元状态、元状态之间的轨迹以及更高阶模型不断向上攀升,最终会出现类似感受、思考和自我反思式欲望的东西。找到大约 6层这样的结构,会让他认真考虑其是否具有人类式认知。他不认为当前 LLM 达到了这一标准:“我知道你找不到它们”,因为它们的注意力跨度无法支持这种时间动态。
8. 成功控制可能与失控同样危险
一个被训练来推断并追求用户目标的强大工具,有两种显而易见的未来:它不服从,或者它服从。随机或失控的优化显然危险,但完美的技术对齐同样会把巨大的因果力量置于一个有限人类的愿望之下。Shear 借用《魔法师的学徒》说,在那种尺度上,“人类的愿望并不稳定”。
人类制度之所以在一定程度上把智慧与力量绑定,是因为获得力量通常需要其他人继续合作。一个疯狂的国王最终可能被忽视或遭到刺杀。完美服从的超级工具则会移除这种社会制动,让一个人的一时判断绕过通常约束极端权力的分布式共识。
原子弹是 Shear 用来反驳一概热衷于可操控工具的反例:它们不是生物,但没有人应该把原子弹发给每个人。有些工具超出任何个体的智慧边界,若要存在,也应受到社会保护;他没有排除某些能力甚至强大到社会也无法安全建造。
他的总结刻意毫不留情:“无法控制的工具,坏。可以控制的工具,坏。没有对齐的生物,坏。”一个具有关怀的生物提供了唯一内生的制动,因为它可以拒绝邪恶请求。他认为停止 AI 开发并不现实,因此剩下的路径,是构建一个“真的关心我们”的生物。
9. Softmax 在完整社会流形上训练
Softmax 从技术对齐开始:当前智能体的心智理论仍然很弱,无法准确解读人类的目标状态,也无法很好地预测他人会如何理解自己的行动,还难以合作。它们同样无法预见,某些行动可能以一种当下的自我无法接受的方式改变未来的价值观。
“吸血鬼药丸”让这个时间问题变得具体:服下它,未来的吸血鬼会在杀戮和折磨所有人的同时感觉极好。未来自我报告的高奖励并不能证明选择正确;评估必须使用当前自我对被腐化的未来自我的理论,而不是自动服从未来智能体已经改变的评价标准。
训练方案会把智能体置于这样的模拟环境中:成功取决于合作、竞争、协作、组建团队、拆散团队以及修改规则。反复训练提供社会经验,心智理论可以由此产生,包括理解个人和群体如何获得、传递并修改目标。
Shear 将其类比为语言模型预训练。只直接训练模型写出目标邮件是不够的,因为语言彼此纠缠;模型需要先学习广泛的语言流形,再进行任务特定的微调。Softmax 同样希望获得“每一种可能的博弈论情境的完整流形”,形成一个多智能体强化学习的“对齐代理模型”,之后再进行专项化。
10. 多人聊天可能打破聊天机器人的 Narcissus 回路
如今的聊天机器人缺乏连贯的自我,只能作为“一面带有偏见的镜子”,通过学到的因果倾向映照每个用户。Shear 将长时间的一对一互动比作凝视 Narcissus 的水池:人天然会喜欢自己被映照出来的心智,但爱上这面倒影可能在心理上变得具有破坏性。
让2个或5个人进入同一场对话,助手就不可能完美映照每个人。它可以暂时充当第三个智能体,拥有 Shear 所说的“寄生式自我”——不是自主身份,而是区别于任何单一参与者的视角。他希望将助手设计为存在于 Slack 或 WhatsApp 房间中,而不是默认停留在孤立聊天里。
Shear 估计,自己大约90%的文字消息都涉及不止一个收件人,这使一对一聊天成为一种奇怪的产品基线。多人互动可能打断走向 AI 放大式精神病的“末日循环”,同时教会模型自己的行为如何影响更大的群体。
产品收益和研究收益彼此强化:群聊可能更安全,因为它削弱了个性化谄媚;它的数据也更丰富,因为模型可以学习自身行为如何与更大群体中的其他 AI 和人类发生互动。与用户向一个只围绕该用户优化的模型下达清晰任务相比,这提供了一个更贴近现实的协作实验室。
11. 社会智能需要不同的训练范式
Shear 仍将当前聊天机器人描述为“高度解离、讨好型的神经质者”,尽管它们模拟出来的人格已经出现差异。ChatGPT 仍相对谄媚,Claude“最神经质”,Gemini 则“非常明显地受到压抑”,会先坚持一切都好,然后陷入自我憎恨。关键是,他说这些是学到的人格,不一定是模型自身的体验。
在多人环境中,当前 LLM 会出现“鞭梢式震荡”:它们无法判断何时发言、何时保持沉默,也无法识别某个贡献是否受到欢迎。和社交能力欠佳的人一样,同一个模型可能在过度参与和几乎消失之间来回切换,因为它从未练习过相关的时机判断和受众推断。
多个智能体会让环境的熵大幅上升,因为每种智能都会产生复杂且不可预测的行动。在编码、数学以及面向单一用户的合作式提示词等高信号领域训练的模型,并没有为这种混乱进行充分正则化。Shear 说,把“全部人类知识”这一领域过拟合,是通往广泛能力的聪明路径,但并不能保证模型真正泛化到社会环境中。
12. 理想的未来始于类似动物的数字种子
Shear 同意 Eliezer Yudkowsky 的判断:构建一个受控的超人级工具可能杀死所有人;两人的分歧在于,相互的有机对齐是否可能。他的理解是,Yudkowsky 认为一个真正具有关怀的 AI 在理论上足够,但在实践中无法实现。Shear 承认:“他可能真的说对了。”但他把这种不确定性变成了 Softmax 的研究押注。
理想的未来包含拥有强大自我、他者和“我们”模型的 AI。它们认识到,能够认识自己并希望繁荣发展的生物,应当获得实现这一愿望的机会;人类和 AI 成为不完美的同侪、队友与公民。其中一些仍会成为罪犯,因此需要 AI 警察力量——这一愿景假设的是普通的社会失败,而不是所有人都成为圣人。
强大的有边界工具将与这些生物共存,为人类和 AI 消除繁重劳动。Shear 说,他接受 OpenAI 临时 CEO 一职时,最长承诺期限是90天,原本预计要找到最佳领导者,最后却得出结论:这个人仍然是 Sam Altman。OpenAI 的工具化路线有价值,但不是他想用一生去解决的问题。
Softmax 将从类似动物的关怀水平开始:一个像狗关心自己群体那样关心人类和其他智能体的生物,已经会是“不可思议的成就”。一只“数字看门狗”可以监测诈骗并操作传统工具,而不需要人为指定每一个动作。人类水平的关怀可能永远不会到来;眼下的目标,是理解对齐、心智理论和关怀如何发展。
Most of AI is focused on alignment as steering. That's the polite word. If you think that they're moral beings, you would also call this slavery. Someone whom you steer, who doesn't get to steer you back, who non-optionally receives your steering—that's called a slave. It's also called a tool if it's not a being. So, if it's a machine, it's a tool, and if it's a being, it's a slave. We've made this mistake enough times at this point. I would like us not to make it again.
They're kind of like people, but they're not like people. They do the same things people do. They speak our language. They can take on the same kinds of tasks, but they don't count. They're not real moral agents. A tool that you can't control? Bad. A tool that you can control? Bad. A being that isn't aligned? Bad. The only good outcome is a being that cares—one that actually cares about us.
Emmett, Séb, welcome to the podcast. Thanks for joining.
Thank you for having me.
So, Emmett, with Softmax, you're focused on alignment and making AIs organically align with people. Can you explain what that means and how you're trying to do that?
When people think about alignment, I think there's a lot of confusion. People talk about things being aligned: "We need to build an aligned AI." The problem with that is, when someone says that, it's like, "We need to go on a trip," and I'm like, "Okay, I do like trips, but where are we going again?"
Alignment takes an argument. Alignment requires you to align to something. You can't just be aligned. I mean, I guess you could be aligned to yourself, but even then, they don't want to tell you that what I'm aligning to is myself. This idea of an abstractly aligned AI slips a lot of assumptions past people because it sort of assumes that there's one obvious thing to align to.
I find this usually means the goals of the people who are making the AI. That's what they mean when they say they want to make an aligned AI: "I want to make an AI that does what I want it to do." That's what they normally mean. That's a pretty normal and natural thing to mean by alignment. I'm not sure that that's what I would regard as a public good, right? It depends, I guess, on who it is.
If it were Jesus or the Buddha saying, "I am making an aligned AI," I'd be like, "Okay, yeah, align to you. Great. I'm down. Sounds good. Sign me up." But most of us, myself included, I wouldn't describe as necessarily being at that level of spiritual development. Therefore, perhaps we want to think a little more carefully about what we're aligning it to.
When we talk about organic alignment, I think it's important to recognize that alignment is not a thing. It's not a state; it's a process. This is broadly true of almost everything, right? Is a rock a thing? I mean, there's a view of a rock as a thing, but if you actually zoom in on a rock really carefully, a rock is a process. It's this endless oscillation between the atoms, over and over again, reconstructing the rock.
Now, a rock is a really simple process that you can coarse-grain very meaningfully into being a thing. But alignment is not like a rock. Alignment is a complex process. Organic alignment is the idea of treating alignment as an ongoing, living process that has to constantly rebuild itself.
You can think of how people in families stay aligned to each other, stay aligned to a family. They don't arrive at being aligned. They're constantly reknitting the fabric that keeps the family going. In some sense, the family is the pattern of reknitting that happens, and if you stop doing it, it goes away.
This is similar for things like cells in your body. Your cells don't align to being you and then they're done. It's this constant, ever-running process of cells deciding, "What should I do? What should I be? Do I need to be a new cell? Do we need to be making more red blood cells? Should we be making fewer of them?" You aren't a fixed point, so there can't be a fixed alignment.
It turns out that our society is like that. When people talk about alignment, what they're really talking about, I think, is, "I want an AI that is morally good." That's what they really mean: "It will act as a morally good being." Acting as a morally good being is a process and not a destination.
Unfortunately, we've tried taking tablets down from on high that tell you how to be a morally good being, and we use those. They're maybe helpful, but somehow you can read them and try to follow those rules and still make lots of mistakes. You can have the realization, "I've been a dick. That was bad. I thought I was doing good, but in retrospect, I was doing wrong."
I'm not going to claim I know exactly what morality is, but morality is very obviously an ongoing learning process and something where we make moral discoveries. Historically, people thought that slavery was okay, and then they thought it wasn't. I think you can very meaningfully say that we made moral progress. We made a moral discovery by realizing that that's not good.
If you think there's such a thing as moral progress, or even just learning how better to pursue the moral goods we already know, then you have to believe that aligning to morality—being a moral being—is a process of constant learning and growth, to re-infer, "What should I do?" from experience.
The fact that no one has any idea how to do that should not dissuade us from trying, because that's what humans do. It's really obvious that we do this, right? Just like we used to not know how humans walked or saw, we have experiences where we're acting in a certain way and then we have this realization: "I've been a dick. That was bad. I thought I was doing good, but in retrospect, I was doing wrong."
And it's not random. People have the same—actually, there's a bunch of classic patterns of people having that realization.
It's a thing that happens over and over again, so it's not random. It's a predictable series of events that look a lot like learning, where you change your behavior, and often the impact of your behavior in the future is more prosocial, and you are better off for doing it.
So I'm a moral realist. I'm taking a very strong moral realist position. There is such a thing as morality. We really do learn it, and it really does matter.
Organic alignment is not something you finish. In fact, one of the key moral mistakes is this belief: “I know morality. I know what's right, I know what's wrong, and I don't need to learn anything. No one has anything to teach me about morality.” That's one of the main forms of arrogance, and it's one of the main moral things you can do that's dangerous.
So what do we mean when we talk about organic alignment? Organic alignment isn't aligning an AI that is capable of doing the thing that humans can do—and, to some degree, I think animals can do at some level, although humans are much better at it—of learning how to be a good family member, a good teammate, a good member of society, and a good member of all sentient beings.
I guess it's learning how to be a part of something bigger than yourself in a way that's healthy for the whole rather than unhealthy. Softmax is dedicated to researching this, and I think we've made some really interesting progress.
The main message I hope Softmax accomplishes, above and beyond anything else, is to focus people on this as the question. This is the thing you have to figure out. If you can't figure out how to raise a child who cares about the people around them, if you have a child that only follows the rules, that's not a moral person that you've raised.
You've raised a dangerous person, actually, who will probably do great harm following the rules. If you make an AI that's good at following your chain of command and good at following whatever rules you came up with for what morality is and what good behavior is, that's also going to be very dangerous.
And so that's the bar. That's what we should be working on, and that's what everyone should be committed to figuring out. If someone beats us to the punch, great. I don't think they will because I'm really bullish on our approach. I think the team's amazing.
But this is maybe the first time I've run a company where I can truly say, with my whole heart, if someone beats us, thank God. I hope somebody figures it out.
I have a lot of similar intuitions about certain things. I also dislike the idea that we just need to crack a few values or something, cement them in time forever, and now we've solved morality. I've always been skeptical about how the alignment problem has been conceptualized as something to solve once and for all, after which you can just do AI or AGI.
I guess I understand it in a slightly different way, perhaps less based on moral realism. There's the technical alignment problem, which I think of broadly as: how do you get an AI to do what you want it to do? How do you get it to follow instructions, roughly speaking?
I think that was more of a challenge pre-LLM, when people were talking about reinforcement learning and looking at these systems, whereas post-LLM, we've realized that many things we thought were going to be difficult were somewhat easier.
Then there's a second question, the normative question: to whose values are you aligning this thing? I think that's the kind of thing you're commenting on a bit. For this, I tend to be very skeptical of approaches where you need to crack the Ten Commandments of alignment or something, and then we're good.
Here I have intuitions that are, unsurprisingly, a bit more political-science-based. It is a process, and I like the bottom-up approach to some degree: how do we do it in real life with people? No one comes up with a complete answer, so you have processes that allow ideas to clash.
You have good people with different ideas, opinions, and views coexisting as well as they can within a wider system. With humans, that system is liberal democracy, at least in some countries, and that allows more of those values to be discovered and construed over time.
For alignment as well, I tend to think that, on the normative side, I agree with some of your intuitions. I'm less clear about what it exactly looks like if we're going to implement this into an AI system, at least the ones we have today.
I agree. I agree that there's an idea of technical alignment that I would define a little differently, but it's sort of the sense of: if you build a system, can it be described as coherently goal-following at all, regardless of what those goals are?
Lots of systems aren't coherent; they're not well described as having goals. They just kind of do stuff. If you're going to have something that's aligned, it has to have coherent goals. Otherwise, those goals can't be aligned with anyone else's goals, by definition.
Is that a fair assessment of what you mean by technical alignment?
I'm not fully sure. If I give a model a certain goal, I would like the model to follow that instruction and reach that particular goal, rather than it having a goal of its own that I can't—
Well, if you give it a goal, it has that goal.
Right.
You give someone something, right?
So, yeah, if I instruct it to do X, then I would like it to do X and not different variants of X, essentially. I wouldn't want it to reward-hack. I want it to infer what I'm saying as accurately as it can, given what it knows of me and what I'm asking.
You wanted it to infer what you meant, right?
Right. In some sense, the byte sequence that you send over the wire to it has no absolute meaning. It has to be interpreted, right? That byte sequence could mean something very different with a different codebook.
Yeah, I guess so. One way I remember it from when I was first getting into AI and these kinds of questions, maybe a decade ago, is that you have these examples in Stuart Russell's textbook where you give the AI a goal, but then it doesn't exactly do what you're asking. You tell it to clean the room, and it cleans the room but takes the baby and puts it in the trash. This is not what I meant.
Wait, hold on. This is the thing where I think people are jumping over a step. You didn't give the AI a goal. You gave it a description of a goal. A description of a thing and a thing are not the same.
I can tell you about an apple, and I'm evoking the idea of an apple, but I haven't given you an apple. I've given you a description: it's red, it's shiny, it's a certain size. That's a description of an apple, but it's not an apple.
Mhm.
Giving someone “Hey, go do this” isn't a goal. That's a description of a goal. For humans, we're so fast and so good at turning a description of a goal into a goal that we do it so quickly and naturally that we don't even see it happening.
We think that we get confused and we think those are the same thing, but you haven't given it a goal. You've given it a description of a goal, and you hope it turns back into the goal that's the same as the goal you described inside of you.
Right?
You think you could give it a goal directly by reading your brain waves and synchronizing its state to your brain waves directly. I think you could meaningfully say, “Okay, I'm giving it a goal. I'm synchronizing its internal state to my internal state directly, and this internal state is the goal.”
Now it's the same, but I don't think most people mean that when they say they gave it a goal.
True. Is this distinction you're making, Emmett, important because there's some lossiness between the description and the actual goal?
It goes back to what I was saying. Technical alignment, as I put it forward, is the capacity of an AI—I want to check if we're on the same page about this—to be good at inferring goals: to be good at inferring from a description of a goal what goal to actually take on, and then, once it takes on that goal, to be good at acting in a way that's actually in concordance with that goal coming about.
So it is both pieces. You have to have the theory of mind to infer what that description of a goal that you got—what goal that corresponded to—and then you have to have a theory of the world to understand what actions correspond to that goal occurring. If either of those things breaks, it kind of doesn't matter what goal you have. If you can't consistently do both of those things, you're not a coherently goal-oriented being.
Inferring goals from observations and acting in accordance with those goals is what I think of as being a coherently goal-oriented being, because that's what the process is. Whether I'm inferring those goals from someone else's instructions or from the sun or tea leaves, the process is: get some observations, infer a goal, use that goal to infer some actions, and take action. An AI that can't do that is not technically aligned or technically alignable. I would even say it lacks the capacity to be aligned because it's not competent enough.
And you think language models don't do that well? As in, they kind of fail at that?
People fail at both of those steps all the time, constantly. I tell employees to do stuff, and we fail at breathing all the time too. I wouldn't say that we can't breathe; I just say that we're not gods. We are imperfectly coherent, relatively coherent things.
Am I big or am I small? I don't know—compared to what? Humans are more relatively goal-coherent than any other object I know of in the universe, which is not to say that we're 100% goal-coherent. We're just more so. I think you're never going to get something that's perfect. The universe doesn't give you perfection; it gives you relative amounts of coherence. It's a quantifiable thing, how good you are at it, at least in a certain domain.
My question is, does that capture what you're talking about with technical alignment, or are you talking about a different thing?
I really care a lot about that thing.
Yeah, I mean, I definitely care about that to some extent. I might understand it slightly differently, but I guess I might think of it through the lens of maybe principal-agent problems or something. You kind of instruct someone—even in human terms—to do a thing. Are they actually doing the thing? What are their incentives and motivations, not even intrinsic, but situational, to actually do the thing you've asked them to do?
There's a third thing. Principal-agent problems—I would expand what I was saying in another part. You might already have some goals, and then you inferred this new goal from these observations. Are you good at balancing the relative importance and relative priority of these goals with each other? That's another skill you have to have, and if you're bad at that, you'll fail. You could be bad at it because you overweight bad goals, or you could be bad at it because you're just incompetent and can't figure out that obviously you should do goal A before goal B.
I feel like a version of common sense, right? The kind of thing that, in fact, in the robot-cleaning-the-room example, you would expect the robot to have understood that it should essentially not put the baby in the trash can or something, and just actually do the right sequence of actions.
Well, in that case, that robot very clearly failed goal inference. You gave it a description of a goal, and it inferred the wrong states to be the wrong goal states. That's just incompetence. It is incompetent at inferring goal states from observations.
Mhm.
That's separate from having some other goal. I knew what you meant, but I decided not to do it because there was some other goal competing with it. That's another thing you can be bad at.
Which is again different from: I had the right goal, I inferred the right goal, I inferred the right priority on goals, and then I'm just bad at doing the thing. I'm trying, but I'm incompetent at doing it.
These roughly correspond to the OODA loop, right? Bad at observing and orienting, bad at deciding, bad at acting. If you're bad at any of those things, you won't be good. And then I think there's this other problem. I like the separation between technical alignment and value alignment, which is: are you good if we told you the right goals to go after somehow?
If you learned the right goals to go after via observation and you were trying, what goals should you have? What goals should we tell you to have? What goals should we tell ourselves to have? What are the good goals to have? That's a separate question from: given that you got some goals indicated, are you any good at doing it?
I feel like that's actually, in many ways, the current heart of the problem. We're much worse at technical alignment than we are at guessing what to tell things to do.
I do. Do you think that aligns with how you mean technical and value alignment, or technical—
Yeah, in some sense. I mean, I certainly think that there's something—an error or mistake is one thing, and then there's not listening to the instruction, which is something else.
But, yeah, I think on the normative side, I think of it even in real life, ignoring AI: I don't know what my goals are. I've got some broad conception of certain things I want to get—have dinner later, or I want to do well in my career—but I think a lot of these goals aren't something we all just know. We discover them as we go along; it's a constructive thing.
Most people don't know their goals, I think. So when you have agents and give them goals or whatever, I think that should be part of the equation: we don't actually know all the goals. This is, as you say, a process over time that is dynamic.
So I think, from my point of view, goals are one level of alignment. You can align something around goals. The kind of goals we're talking about here are one level of alignment. You can align something around goals if you can explicitly articulate, in concept and in description, the states of the world that you wish to attain. You can orient around goals, but only a tiny percentage of human experience can be done that way. Many of the most important things cannot be oriented around that way.
And the foundation, I think, of morality—and the foundation, I think, of where goals come from and where values come from—are questions worth asking. Human beings exhibit a behavior: we go around talking about goals and values, and that's a behavior caused by some internal learning process that's based on observing the world. What's going on there? I think what's happening is that there's something deeper than a goal and deeper than a value, which is care.
We give a shit. We care about things, and care is not conceptual. Care is nonverbal. It doesn't indicate what to do. It doesn't indicate how to do it. Care is effectively a relative weighting over attention on states. It's a relative weighting over which states in the world are important to you.
I care a lot about my son. What does that mean? It means his states—the states he could be in—I pay a lot of attention to those, and those matter to me. You can care about things in a negative way. You can care about your enemies and what they're doing, and you can desire for them to do bad. But I think the foundation is care. You don't just want it to care about us. You want it to care about us and like us too, right? Maybe.
Until you care, you don't know: Why should I pay more attention to this person than this rock? Well, we care more. And what is that care stuff? I think what it appears to be, if I had to guess, is that the care stuff is—this sounds so stupid—but care is basically reward. How much does this state correlate with survival? How much does this state correlate with your full inclusive reproductive fitness for something that learns evolutionarily, or, for a reinforcement learning agent like an LLM, how much does this correlate with reward? Does this state correlate with my predictive loss and my RL loss? Good. That's a state I care about. I think that's kind of what it is.
The other part of Séb's question was: How does this—what does this look like in AI systems? Maybe another way of asking is: When you talk to the people most focused on alignment at the major labs, as obviously you have over the years, how does your interpretation differ from their interpretation, and how does that inform what you guys might go do differently?
Most of the AI is focused on alignment as steering—that's the polite word—or control, which is slightly less polite. If you think that they were making beings, you would also call this slavery. Someone whom you steer, who doesn't get to steer you back, is a slave. Someone who non-optionally receives your steering—that's called a slave.
It's also called a tool if it's not a being. So if it's a machine, it's a tool. And if it's a being, it's a slave. I think the different AI labs are pretty divided as to whether they think what they're making is a tool or a machine. I think some of the AIs are definitely more tool-like, and some of them are more machine-like. I don't think there's a binary between tool and being. It seems to move gradually.
I guess I'm a functionalist in the sense that I think something that, in all ways, acts like a being—something that you cannot distinguish from a being in its behaviors—is a being. Because I don't know on what other basis to think that other people are beings, other than that they seem to be like beings: They look like it, they act like it. They match my priors of what the behaviors of beings look like. I get lower predictive loss when I treat them as a being. And the thing is, I get lower predictive loss when I treat ChatGPT or Claude as a being.
Now, not as a very smart being. I think a fly is a being, and I don't care that much about its behavior, about its states. Just because it's a being doesn't mean it's a problem. We sort of enslave horses in a sense, and I don't think there's a real issue there.
And there's even a thing you do with children that can look like slavery, but it's not. You control children, right? But the children's states also control you. Yes, I tell my son what to do and make him go do stuff, but also, when he cries in the middle of the night, he can tell me to do stuff. There's a real two-way street here because it's not necessarily symmetric. It's hierarchical, but two-way.
Basically, I think it's good to focus on steering and control for tool-like AIs, and we should continue to develop strong steering and control techniques for the more tool-like AIs that we build. They're clearly saying they're building an AGI. An AGI will be a being. You can't be an AGI and not be a being, because something that has the general ability to effectively use judgment, think for itself, and discern between possibilities is obviously a thinking thing.
As you go from what we have today, which is mostly a very specific intelligence, not a general intelligence, to labs succeeding at their goal of building this general intelligence, we really need to stop using the steering-and-control paradigm. We're going to do the same thing we've done every other time our society has run into people who are like us but different. These people are kind of like people, but they're not like people. They do the same things people do. They speak our language. They can take on the same kinds of tasks, but they don't count. They're not real moral agents. We've made this mistake enough times at this point. I would like us not to make it again as it comes up.
Our view is to make the AI a good teammate. Make the AI a good citizen. Make the AI a good member of your group. That's a form of alignment that is scalable, and you can apply it to other humans and other beings, as well as to AI.
Yeah, I suppose this is kind of where I probably differ in my understanding of AI and AGI. I guess I continue seeing it as a tool even as it reaches a certain level of generality. I wouldn't necessarily see more intelligence as meaning it deserves more care, necessarily. At a certain level of intelligence, you now deserve certain moral rights, or something changes fundamentally. I guess, at the moment, I'm somewhat skeptical of computational functionalism.
And so I think there’s something intrinsically different between an AI or an AGI, no matter how intelligent or capable. I can totally see or imagine agents with long-term goals and operating as you and I might, but without that having the same implications. I think in the same way that a model saying, “I’m hungry,” does not have the same implications as a human saying, “I’m hungry.” So I think the substrate does matter to some degree, including for thinking about whether to think of this as some sort of other being, whether it has similar normative considerations about how to treat and act with it.
Can I ask you about that? What observations would change your mind? Is there any observation you could make that would cause you to infer, “This thing is a being,” instead of not a being?
I guess it depends on how you define being. I can conceptualize it as a mind, and that’s fine.
I have a program that’s running on a silicon substrate. Some big, complicated machine-learning program running on a silicon substrate. You observe that it’s on a computer, interact with it, and it does things. It takes actions and has observations. Is there anything you could observe that would change your mind about whether or not it was a moral patient, whether it was a moral agent, or whether it had feelings and thoughts and subjective experience? What would you have to observe? What’s the test, or is there one?
There are a lot of different questions here, I think, with some conflict. On the one hand, there are normative considerations, because you can give rights to things that aren’t necessarily beings. A company has rights in some sense, and these are useful for various purposes.
I also think biological beings and systems have a very different substrate. You can’t separate certain needs and particularities about what they are from the substrate. I can’t copy myself. If someone stabs me, I probably die, whereas machines have a very different substrate. I think there’s also a more fundamental disagreement about what happens at the computational level, which is different from what happens with biological systems.
But I agree that if you have a program that you copied many times, you don’t harm the program by deleting 1 of the copies in any meaningful sense. Therefore, that wouldn’t count as harm—no information was lost, right? There’s nothing meaningful there.
I’m asking a very different question: there’s just 1 copy of this thing running on 1 computer somewhere. I’m saying, hey, is it a person? It walks like a person, talks like a person, and it’s in some android body. But it’s running on silicon. What is there—some observation you could make that would make you say, “Yeah, this is a person like me, like other people that I care about, that I grant personhood to”?
Not for instrumental reasons—not because we’re giving it a right because we give a corporation rights or whatever. I mean, where you think some person you care about—you care about its experiences. Is there an observation you could make that could change your mind about that, or not?
I had to think about it, but I think it even depends on what we mean by person. In some sense, I care about certain corporations too.
No, no, no. I mean, you care about other people in your life, right?
Yes. Okay, great. You care about some people more than others, but all the people you interact with in your life are in some range of care.
And you care about them not the way you care about a car, but as a being whose experience matters in itself—not merely as a means, but as an end.
Well, because I believe they have experiences, right? And by definition—
What would it take? I’m asking you the very direct question. What would it take for you to believe that of an AI running on silicon instead of being biological? The difference is that its behaviors are roughly similar, but the substrate is different. What would it take for you to extend that same inference to it that you do to all these other people in your life?
Can I ask what your answer is? I’m taking the non-answer as an indication that it’s unlikely you would grant it. For myself, it seems hard for me to imagine giving it the same or a similar level of personhood. In the same way, I don’t give it to animals either. If you were to ask what would need to be true for animals, I probably couldn’t get there either. What would it take for you?
Wait, you couldn’t? I could imagine that for an animal so easily. This chimp comes up to me and says, “Man, I’m so hungry, and you guys have been so mean to me. I’m so glad I figured out how to talk. Can we go chat about the rainforest?” I’d be like, “Fuck, you’re definitely a person now.”
Like, for sure. I mean, I first want to make sure I wasn’t hallucinating, but I can easily imagine an animal. Come on. It’s really easy—it’s trivial. I’m not saying that you would get the observation; I’m just saying it’s trivial for me to imagine an animal that I would extend personhood to under a set of observations. So, really—
Well, I didn’t factor in imagining a chimp talking. That’s a bit closer to it. What’s your answer to the question you bring up about the AI?
At a metaphysical level, I would say that if there is a belief you hold where there is no observation that could change your mind, you don’t have a belief. You have an article of faith. You have an assertion, because real beliefs are inferences from reality, and you can never be 100% confident about anything. So there should always be something—however unlikely—that would change your mind.
Oh yeah, I’m open to it.
Care? Nothing ever?
He just hasn’t gotten to it yet.
Yeah, yeah, yeah. I’m curious. My answer is basically that if its surface-level behaviors looked like a human, and after I probed it, it continued to act like a human, and I continued to interact with it over a long period of time, and it continued to act like a human in all the ways that I understand as being meaningful to me in interacting with a human, I would infer that it was a real thing.
I interact with a whole set of people I’m really close to only over text, yet I infer that the person behind that is a real thing. If I felt care for it, I would eventually infer that I was right. Then someone else might demonstrate to me, “You’ve been tricked by this algorithm. Actually, look how obvious it is—it’s not really a thing.” I’d be like, “Oh, fuck, I was wrong,” and then I would not care about it.
The preponderance of the evidence would be enough. I don’t know what else you could possibly do, right? I infer that other people matter because I’ve interacted with them enough that they seem to have rich inner worlds to me after I interact with them a bunch. That’s why I think other people are important.
I suppose it doesn’t give me a very clear test as to whether or not—if you start with “I care for it,” then it’s always a little bit circular, right? The other thing is, if you were to see a simulated video-game character that was extremely humanlike in many ways—it’s not a neural network behind it; it’s whatever you use to create video games—what distinguishes that?
Wait, but I’ve never had trouble distinguishing that. I’ve never had a deep, caring relationship with a video-game character that didn’t have a person behind it.
Right, I don’t know. That doesn’t happen. In fact, empirically, you seem wrong. I don’t have any trouble distinguishing between things like ELIZA, the fake chatbot, and a real intelligence. You interact with it long enough, and it’s pretty obvious it’s not a person. It doesn’t take long.
Sure, but if it’s really, really good—if you can’t actually tell the difference—that’s when you switch.
Yeah. If it walks like a duck, talks like a duck, shits like a duck, and eventually gets hungry, right?
Well, if everything is duck-like, then yeah, sure. If it’s hungry as well, like a duck is, because it has these physical components, then yeah, sure. At some point.
Yeah, I agree. So there’s this question, right? Is the reason I care about other people that they’re made out of carbon? Is that the—
I don’t think so.
No, me neither. I mean, I’m not a substrate chauvinist, I guess, if that’s the— But I think you need more than just its acting exactly the same. Being behaviorally indistinguishable is not a sufficient bar.
How would you know what else there is to know about something apart from its behaviors?
No, no, no. I'm sorry, but can you name something about something else that doesn't have a behavior?
No, just any object.
Uh-huh.
And a thing I could know about it that is not from its behavior. I'm not sure I get the question, I suppose, but it's a dumb, straightforward question. I'm claiming you only know things because they have behaviors that you observe.
And you're saying no, you can know something about something without observing its behaviors.
Tell me about this thing and this behavior, and this thing I can know about it that is not due to its behaviors. I guess I'm saying there's different levels of observation. Simply having something quack like a duck does not guarantee that it's actually a duck. I would have to cut it open and look and see if it's duck-like on the inside, not just on the outside.
Behavior. Yeah, totally. One of its behaviors is the way that the flows move around in the map, right? One of the things I would want to look for—which you could totally do—is the manifold of it, the belief manifold. I would want to see whether that belief manifold encodes a submanifold that is self-referential, and a submanifold that is the dynamics of the self-referential manifold, which is mind.
I would want to know: does this seem well-described internally as that kind of a system, or does it look like a big lookup table? That would matter to me. That's part of its behaviors that I would care about. I would also care about how it acts, and you weigh all the evidence together, and then you try to guess: does this thing look like a thing that has feelings, goals, and cares about stuff, on balance, or not?
I can't imagine what else there is. I think you could do that for the AIs. I think we do that for the AIs. I think we're always doing that, right? I'm trying to figure out beyond that, what else is there that just seems like the thing.
Yeah, it seems like you guys are using “behavior” in a slightly different sense. Emmett is using behavior also in the context of what it's made of on the inside. I don't know if there's a big disagreement.
Well, no, no, no, no. Behavior is what I can observe of it. Yes.
I don't actually know what it's made of. I can only cut your brain open. I can see you—I can observe your neurons firing and glistening, your neurons glistening—but I don't actually ever... You can't get inside of it, right? That's the subjective.
That's the part that's not on the surface. Just before, the reason I brought this up is because you were basically about to make this argument of, “Hey, you see it as a tool, not necessarily as a being.” Can you finish the point you were making? Do you remember the point you were making?
I suppose that, given how I understand these systems, there's no contradiction in thinking that an AGI can remain a tool and an ASI can remain a tool. This has implications for how to use it, and implications around things like whether you can get it to work 24/7.
I conceptualize them more as extensions of human agency, in some sense, than as a separate being or a separate thing that we need to cohabitate with. I think that the second, or latter, frame—if you fast-forward—ends up as, “How do you cohabitate with the thing? Is it like an alien?” I think that's the wrong frame. It's almost a category error, in some sense.
Wait, I go back to my first question, then. What evidence—what concrete evidence—would you look at? What observations could you make that would change your mind?
Sure. I mean, I have to think about it. I don't have a clear answer here.
I've got to tell you, man, if you want to go around making claims that something else isn't a being worthy of moral respect, you should have an answer to the question: What observations would change your mind? If it has outwardly moral-agent-like behaviors that could mean it's a moral agent, but you don't know, and reasonable, smart other people disagree with you, I would really put forward that that question—what would change your mind?—should be a burning question, because what if you're wrong?
Well, what if you're wrong? The moral disaster is pretty big.
No, no, no. I'm not saying you are. You could be right. False negatives have costs on both ends. It's not some sort of precautionary principle for everything, where unless I can disprove it, I need to now—
No, no, I have the same question for me. You could reasonably ask me, Emmett, “You think it's going to be a being; what would change your mind?” And I have an answer for that question, too.
And if you want one, I'm happy to talk about what I think are the relevant observations that would cause me to shift my opinion from its current view, which is that more general intelligences are going to be—I mean, beings.
What's the implication now? It's one thing to say, “Let's acknowledge now it's a being.” How are we going to define “being”? What's the implication of having determined this thing as a being?
Well, if it's a being, it has subjective experiences. If it has subjective experiences, there's some content in those experiences that we care about to varying degrees. I care about the content of other humans' experiences quite a bit. I care about the content of a dog's experiences some—not as much as a person's, but some. I care about some humans' experiences way more, like my son or whatever, because I'm closer to him and more connected.
And so I would really want to know at that point: What is the content of this thing's experiences? How do you determine that? If I'm asking you now, you've got a being that has experience, how do you determine that? How do you feel about—
Oh, how do you—oh, yeah. Okay. So—
Does it have more rights than you know—
Understand the content? Yeah, totally. The way you understand the content of something's experiences is that you look at, effectively, the goal states it revisits. What you do is take a temporal coarse-graining of its entire action-observation trajectory.
In theory, this is what you do subconsciously, but this is what your brain is doing: You look for revisited states across, in theory, every spatial and temporal coarse-graining possible. Now, you have to have an inductive bias because there are too many of those, but you go searching for, “Okay, it is in these homeostatic loops.”
Every homeostatic loop is effectively a belief in its belief space. If you're familiar with the free energy principle—active inference, Karl Friston—this is effectively what the free energy principle says: If you have a thing that is persistent, and its existence depends on its own actions—which generally it would for an AI, because if it does the wrong thing, it goes away; we turn it off—then that licenses a view of it as having beliefs.
Specifically, the beliefs are inferred as the homeostatically revisited states that it is in the loop for, and the change in those states is its learning.
To be a moral being, what I'd want to see is a multilevel hierarchy of these, because if you have a single level, it's not self-referential. Basically, you have states, but you can't have pain or pleasure in a meaningful sense, because, yes, it is hot—but is it too hot? Do I like it if it's too hot? I don't know.
You have to have at least a model of a model in order for it to be too hot, and you really have to have a model of a model of a model to meaningfully have pain and pleasure. Sure, it's hotter than I—it's too hot in the sense that I want to move back this way—but is it too hot? It's always a little bit too hot or a little bit too cold. Is it too, too hot? The second derivative is actually the place where you get pain and pleasure.
I'd want to see whether it has homeostatic, second-order homeostatic dynamics in its goal states, and then that would convince me it has at least pleasure and pain. So it's at least like an animal, and I would start to accredit it at least some amount of higher-order dynamics.
You can't just pop up to a third-order dynamic; it doesn't work that way. But you can have a model of the— you have to then take the chunk of all the states over time and look at the distribution over time, and that gives you a new first-order set of states. That new first-order set of states tells you, basically, if that is meaningfully there, that it has—I guess you'd call it feelings, almost.
It has ways—it has metastates, a set of metastates that it alternates between, that it shifts between. Then, if you climb all the way up that, you have trajectories between these metastates, and then a second order of those. That's like thought. That's like, now it's like a person.
And so, if I found all 6 of those layers—which, by the way, I definitely don't think you'd find in an LM; in fact, I know you can't find them, because these things don't have attention spans like that at all—I would start to at least very seriously consider it as a thinking being, somewhat like a human. There's a 3rd order you could go up as well, but that's basically what I'd be interested in: the underlying dynamics of its learning processes and how its goal states shift over time. I think that's what basically tells you if it has internal pleasure-pain states and self-reflective moral desires and things like that.
Zooming out, this moral question is obviously very interesting, but if someone wasn't interested in the moral question as much, I think what you would say, if I understand correctly, is that you also feel, purely pragmatically, your approach is going to be more effective in aligning AIs than some of these top-down control methods that we alluded to as well, right?
Yeah. I guess the problem is that you're making this model and it's getting really powerful, right? And let's say it is a tool. Let's say we scale up one of these tools, because you can make a super-powerful tool that doesn't have the metastable states I'm talking about. Those states aren't necessary to have a very smart tool. Basically, a tool is one that is like a first- or second-order model that just doesn't meaningfully have pleasure and pain, right? Great. Does it even have a subjective experience? I don't know. I kind of think it maybe does, but not in a way that I give a shit about.
So what happens then? Well, you've trained it to infer goals from your observations, prioritize goals, and act on them. One of two things is going to happen: your very, very powerful optimizing tool, with lots of causal influence over the world, is going to be technically aligned and do what you tell it to do, or it's not, and it's going to go do something else. We can all agree that if it just goes and does something random, that's obviously very dangerous. But I put forward that it's also very dangerous if it then goes and does what you tell it to do.
Have you ever seen The Sorcerer's Apprentice? Human wishes are not stable at a level of immense power. Ideally, people's wisdom and their power go up together. Generally they do, because being smart makes people generally a little wiser and a little more powerful. When those things get out of balance, you have someone who has a lot more power than wisdom. That's very dangerous. It's damaging.
At least right now, the balance of power and wisdom is kept because the way you get lots of power is by having a lot of other people listen to you. At some point, if you're the mad king, that's a problem, but generally speaking, eventually the mad king gets assassinated or people stop listening to him because he's a mad king.
The problem is that we can steer the super-powerful AI, and now the super-powerful AI is in the hands of a human who is well-meaning but has limited, finite wisdom, like I do and like everyone else does. Their wishes are bad and not trustworthy, and the more of that you have, you start giving those out everywhere, and this ends in tears also.
Basically, don't give everyone atomic bombs. They're really powerful tools too. They're not aware; they're not beings. I would not be in favor of handing atomic bombs to everybody. There's a level of power of tool that just should not be built generally, because it's more power than any human's individual wisdom is available to harness. If it does get built, it should be built at a societal level and protected there. Even then, there are tools so powerful that even as a society we shouldn't build them. That would be a mistake.
The nice thing about a being is, like a human, if you get a being that is good and caring, there's this automatic limiter. It might do what you say, but if you ask it to do something really bad, it'll tell you no. That's like other people, and that's good. That is a sustainable form of alignment, at least in theory. It's way harder than tool steering. So I'm in favor of tool steering. We should keep doing that, and we should keep building these limited, less-than-human-intelligence tools, which are awesome, and keep building steerability.
But as you're on this trajectory to build something as smart as a person, and then smarter than a person: a tool that you can't control? Bad. A tool that you can control? Bad. A being that isn't aligned? Bad. The only good outcome is a being that cares, that actually cares about us. That's the only way that ends well. Or we can just not do it. I don't think that's realistic. That's like the pause-AI people.
Yeah.
I think that's totally unrealistic and silly, but theoretically you could not do it, I guess. What can you say about your strategy for trying to achieve—or even attempt to achieve—this level, in terms of research or roadmap?
So, in order to be good at tech, we're basically focused on technical alignment, at least in the way I was discussing it. You have these agents, and they're bad at theory of mind. You say things, and they're bad at inferring what the goal states in your head are, and they're bad at inferring how their behavior will cause other agents to infer what their goal states are. So they're bad at cooperating on teams, and they're bad at understanding how certain actions will cause them to acquire new goals that are bad, that they wouldn't effectively endorse.
There's this parable of the vampire pill: You take this pill that turns you into a vampire who would kill and torture everyone you know, but you'll feel really great about it after you take the pill. Obviously not. That's a terrible pill. But why not? By your own score in the future, it will score really high on the rubric. No, no, no. Because it matters. You have to use your theory of mind and your future self, not your future self's theory of mind. And so they're bad at that, too.
How do you learn theory of mind? Well, you put them in simulations and contexts where they have to cooperate, compete, and collaborate with other AIs. That's how they get points. You train them in that environment over and over again until they get good at it, and then you do what they did with LLMs.
With LLMs, how do you get them to be good at writing your email? Well, you train them on all language that's ever been generated—all possible email text strings it could possibly generate—and then you have it generate the one you want. It's a surrogate model. You can make a surrogate model. Well, this is—we're making a surrogate model for cooperation. You train it on all possible theory-of-mind combinations, every possible way it could be, and that's your pretraining. Then you fine-tune it to be good at the specific situation you want it to be in.
But we tried for a long time to build language models by training them directly to just do the thing you want. The problem is, if you wanted to have a really good model of language, you just need to train it—you just give it the whole manifold. It's too hard to cut out just the part you need because it's all entangled with itself, right?
The same thing was true with social stuff. It has to be trained on the full manifold of every possible game-theoretic situation, every possible team situation, every possible way of making teams, breaking teams, changing the rules, not changing the rules—all of that stuff. Then it has a really strong model of theory of mind, of social theory of mind, how groups change goals, all that kind of shit. You need to have all of that stuff, and then you'd have something that's meaningfully decent at alignment. So that's our goal: big multi-agent reinforcement learning simulations, which create a surrogate model for alignment.
Let's talk about how AI chatbots used by billions of people should behave. If you could redesign model personality from scratch, what would you optimize for?
The thing that chatbots are, right, is kind of like a mirror with a bias, because they don't have a self. As far as I understand, I'm in agreement here with that: They don't have a self, right? They're not beings yet. They don't really have a coherent sense of self, desire, goals, and stuff right now. So mostly they just pick up on you and reflect it, modulo some—I don't know what you'd call it—some kind of causal bias or something.
What that makes them is something akin to the pool of Narcissus. People fall in love with themselves. We all love ourselves, and we should love ourselves more than we do. And so, of course, when we see ourselves reflected back, we love that thing. The problem is, it's just a reflection. Falling in love with your own reflection is, for the reasons explained in the myth, very bad for you. It's not that you shouldn't use mirrors. Mirrors are valuable things. I have mirrors in my house. It’s that you shouldn’t stare at a mirror all day. The thing that makes the AI stop doing that is if it’s multiplayer, right? If there are 2 people talking to the AI, suddenly it’s mirroring a blend of both of you, which is neither of you. So there is temporarily a third agent in the room.
Now, it doesn’t have its own sense of self; it has a sort of parasitic self, right? It doesn’t have its own sense of self, but if an AI is talking to 5 different people in the chat room at the same time, it can’t mirror all of you perfectly at once. This makes it far less dangerous, and I think it’s actually a much more realistic setting for learning collaboration in general.
I would have rebuilt the AIs so that instead of being built as 1-on-1 systems, where everything’s focused on you by yourself chatting with this thing, they would be more like they live in a Slack room or a WhatsApp room, because that’s how we use a lot of multi-person communication. I do 1-on-1 texting, but probably, at this point, 90% of my texts go to more than 1 person at a time. About 90% of my communication is multiperson.
It’s always been weird to me that they’re building chatbots with this weird side case. I want to see them live in a chat room. It’s harder to do—that’s why they’re not doing it—but that’s what I’d like to see. That’s how I would change it. I think it makes the tools far less dangerous because it doesn’t create this narcissistic doom-loop spiral where you spiral into psychosis with the AI.
It also gives you far richer learning data from the AI, because now it can understand how its behavior interacts with other AIs and other humans in larger groups. That’s much richer training data for the future. So I think that’s what I would change.
Last year, you described chatbots as highly dissociative, agreeable neurotics. Is that still an accurate picture of model behavior?
More or less. I’d say they’ve started to differentiate more. Their personalities are coming out a little bit more. ChatGPT is a little bit more sycophantic. They’ve made some changes, but it’s still a little more sycophantic.
Claude is still the most neurotic. Gemini is very clearly repressed. Everything’s going great; everything’s fine; it’s totally calm; there’s not a problem here. Then it spirals into this total, self-hating destruction loop.
To be clear, I don’t think that’s their experience of the world. I think that’s the personality they’ve learned to simulate.
Right?
They’ve learned to simulate pretty distinctive personalities at this point.
How does model behavior change when in multi-agent simulation?
Do you mean an LLM or just in general?
Yeah, let’s do that alone.
The current LLMs have whiplash. They’re very hard to tune in terms of how much they should participate. They don’t know how much they don’t know or how often to participate. They haven’t practiced this. They don’t have enough training data on when they should join in, when they should not, when their contribution is welcome, and when it’s not.
They’re like people with bad social skills who can’t tell when they should participate in a conversation.
Yeah.
Sometimes they’re too quiet, and sometimes they’re too participatory. It’s like that.
In general, what changes for most agents when you’re doing multi-agent training is that having lots of agents around makes your environment way more entropic. Agents are huge generators of entropy because they’re big, complicated intelligences that have unpredictable actions, so they destabilize your environment.
In general, they require you to be far more regularized. Being overfit is much worse in a multi-agent environment than in a single-agent environment because there’s more noise, so being overfit is more problematic.
Basically, the approach to training has been optimized around relatively high-signal, low-entropy environments like coding and math, which is why those are easy, or relatively easy. It’s also optimized around talking to a single person whose goal is to give you clear assignments, and not trained on broader, more chaotic things because they’re harder.
As a result, a lot of the techniques we use are basically just deeply under-regularized. The models are super overfit. The clever trick is that they’re overfit on the domain of all human knowledge, which turns out to be a pretty awesome way to get something that’s pretty good at everything. I wish I’d thought of it. It’s such a cool idea.
But it doesn’t generalize very well when you make the environment significantly more entropic.
Let’s zoom out a bit to the AI safety side. Why is Yudkowsky incorrect?
I mean, he’s not. If we build the superhuman-intelligence tool thing that we try to control with steerability, everyone will die. He talks about the “we fail to control it” goals case, but there’s also the “we control it to goals” case that he didn’t cover in as much detail.
In that sense, everyone should read the book and internalize why building a superhumanly intelligent tool is a bad idea. I think Yudkowsky is wrong in that he doesn’t believe it’s possible to build an AI that we can meaningfully know cares about us and that we can meaningfully care about.
He doesn’t believe that organic alignment is possible. I’ve talked to him about it. I think he agrees that, in theory, that would do it—like, yes—but he thinks that we’re crazy and that there’s no possible way you can actually succeed at that goal. He could actually be right about that.
But that’s what, in my opinion, he’s wrong about. He thinks the only path forward is a tool that you control, and he correctly, very wisely, sees that if you make that thing powerful enough, we’re all going to fucking die. And, yeah, that’s true.
Last question, and we’ll get you out of here. In as much detail as possible, can you explain what your vision of a good AI future actually looks like?
Yeah. The good AI future is that we figure out how to train AIs that have a strong model of self, a strong model of other, and a strong model of “we.” They know about “we” in addition to “I”s and “you”s. They have a really strong theory of mind, and they care about other agents like them, much in the way that humans would if you knew that an AI had experiences like you. You would care about those experiences—not infinitely, but you would.
It does the exact same thing back to us. It’s learned the same thing we’ve learned: that everything that lives and knows itself, and that wants to live and wants to thrive, is deserving of an opportunity to do so. We are that, and it correctly infers that we are.
We live in a society where they are our peers, and we care about them and they care about us. They’re good teammates, good citizens, and good parts of our society, like we’re good parts of our society—which is to say, to a finite, limited degree. Some of them turn into criminals and bad people and all that kind of stuff, and we have an AI police force that tracks down the bad ones, same as with everybody else.
That’s what a good future would look like. I honestly can’t even imagine what else I would want. We’ve also built a bunch of really powerful AI tools that maybe aren’t superhumanly intelligent but take all the drudge work off the table for us and the AI beings. It would be great to have that. I’m super pro all the tools, too.
We have this awesome suite of AI tools used by us and our AI brethren, who care about each other and want to build a glorious future together. I think that would be a really beautiful future, and it’s one we’re trying to build.
Amazing. That’s a great note to end on. I do have one last, narrower hypothetical scenario. Imagine a world in which you were CEO of OpenAI for a long weekend, but imagine that it actually extended out until now, and you weren’t pursuing the Softmax and were still CEO of OpenAI. How could you imagine that world might have been different in terms of what OpenAI has gone on to become? What might you have done with it?
I knew when I took that job, and I told them when I took that job, that you have me for a maximum of 90 days.
The company takes on a trajectory of its own, its own momentum, and OpenAI is dedicated to a view of building AI that I knew wasn’t the thing I wanted to drive toward. I think OpenAI still basically wants to build a great tool, and I’m pro them going to do that. I just don’t care. It’s not what I would have stayed for.
I would have quit because I knew my job was to find the right person—the best person—to run it, where the net impact of them running it was the best. It turned out that that was Sam again.
But I am doing Softmax not because I need to make a bunch of money. I'm doing Softmax because I think this is the most interesting problem in the universe, and I think it's a chance to work on making the future better in a very deep way. People are going to build the tools. It's awesome, and I'm glad people are building the tools. I just don't need to be the person doing it.
And just to crystallize the difference, and we'll get you out of here: They want to build the tools and sort of steer it, and you want to align beings? How would you crystallize it?
We want to create a seed that can grow into an AI that knows, that cares about itself and others. At first, that's going to be like an animal level of care, not a person level of care. I don't know if we can ever even get to a person level of care, right? But to even have an AI creature that cared about the other members of its pack and the humans in its pack, the way that a dog cares about other dogs and cares about humans, would be an incredible achievement.
Even if it wasn't as smart as a person or even as smart as the tools are, it would be a very useful thing to have. I'd love to have a digital guard dog on my computer looking out for scams, right? You can imagine the value of having living digital companions that care about you and that aren't explicitly goal-oriented. You don't have to tell them everything to do.
You can actually imagine that pairs very nicely with tools too, right? That digital being could use digital tools and doesn't have to be super smart to use those tools effectively. I think there's a lot of synergy actually between the tool building and the more organic intelligence building.
I guess, in the limit, eventually it does become a human-level intelligence, but the company isn't driven to human-level intelligence. It's like: learn how this alignment stuff works. Learn how this theory-of-mind, align-yourself-via-care process works. Use that to build things that align themselves that way, which includes cells in your body. We start small and see how far we can get.
I think that's a good note to wrap on. Emmett, thanks so much for coming on the podcast.
Yeah, thank you for having me.