Emmett Shear 谈如何构建真正关心人的 AI:超越控制与引导
Emmett Shear 的分界线并不只是对齐与未对齐的 AI,而是工具与主体。 有限的、工具化的系统应保持可引导;当系统逐渐具备他所说的 AGI 级别的普遍判断力时,单向控制就会变成他颇具挑衅意味地称之为“奴役”的状态,而唯一好的结果是“一个在乎人的主体——真正关心我们”。
对齐是一场持续进行的道德学习过程,不是一份可以解决并冻结的规格说明。 家庭不断“重新织补”成员之间的关系,身体不断协调细胞的分工,社会也在不断发现新的道德事实;只训练成服从戒律的 AI,可能会变成“一个危险的人……很可能会严格遵守规则,却造成巨大伤害”。
技术对齐始于推断指令背后的目标,因为目标的描述并不等于目标本身。 可靠的能动性需要心智理论、在相互竞争的目标之间正确排序,以及能够把意图转化为行动的世界模型;那个打扫房间时把婴儿扔进垃圾桶的机器人,问题在于无法推断目标,而不只是服从得不够。
一个完全可控的超人类工具,可能和一个无法控制的工具一样危险。 当杠杆作用巨大时,“人的愿望并不稳定”,而操舵会把超出任何个人智慧的能力交到会犯错的人手里;与工具不同,一个真正关心人的主体拥有“自动限幅器”,能够拒绝破坏性的请求。
Softmax 的技术押注是,社会智能需要自己的预训练流形。 其设想中的多智能体强化学习环境,会让智能体在面向具体任务微调前,经历合作、竞争、组队、规则变化和观点冲突——这相当于社会智能版的语言模型训练:先在完整的语言流形上学习。
今天的一对一聊天机器人是“一面带偏见的镜子”,既带来产品风险,也只能产生薄弱的社会训练数据。 Shear 希望让它们成为多人房间里的原生参与者,因为没有一个模型能完美映射所有人;这可能打断“自恋式的……末日循环”,同时生成更丰富的协作、时机和群体目标数据。
尚未解决的底层基质与行为之争,决定了 AI 的人格、权利与治理问题。 Séb Krier 认为,即使是 AGI 或 ASI,也可能只是人类能动性的延伸;Shear 则要求在否定其道德地位前先提出可证伪的测试。他自己的门槛,是寻找能够支撑疼痛、愉悦、感受与思维的持久、自指性学习动力学。
Shear 并没有直接奔向人类级智能;他想从类似动物的起点培育“在乎”。 一个保护自己“群体”的数字生物——或许是一只监测诈骗、并能调用独立工具的数字看门狗——即使 Softmax 永远达不到“人类级别的关怀”,也已经足够有价值。
1. 对齐必须持续重新学习善为何物
Emmett 开场提出的反驳同时涉及概念与政治:对齐“需要一个论证”,因为某种东西必须对齐于另一种东西。现实中,那个未被说出的目标通常是创造者的目标——作为产品目标可以理解,但当创造者没有“耶稣或佛陀”的智慧时,它并不会自动成为“公共利益”。
“对齐不是一件东西,不是一种状态,而是一个过程。” 一个家庭靠不断“重新织补”连接成员的关系而存续;细胞也不断判断身体需要它们承担什么角色。道德同样处于动态之中,因为人和周围系统都不是固定不变的参照点。
Emmett 明确采取“非常强的道德实在论立场”:道德进步是真实存在的,人类拒绝奴隶制就是一种发现,而不是偏好变化。人们一次次意识到:“我真是个混蛋,那样做是不对的。”随后修正自己的行为,并往往变得更亲社会。
最危险的道德姿态,是相信学习已经结束:“我知道什么是对的,也知道什么是错的。”因此,有机对齐的目标应是把 AI 养成一个好的家庭成员、队友、公民,以及更大共同体的参与者,而不是制造一个即便服从规则、也能造成巨大伤害的规则执行器。
2. 指令描述目标,但不会传递目标
Séb 最初的区分,是技术层面的指令遵循与规范层面的价值归属问题。对于后者,他倾向于一种类似自下而上的自由民主过程:相互冲突的观点共存于允许价值被争论、重构的系统之中。
Emmett 将技术对齐重新定义为连贯的目标导向。智能体接收观察结果、推断目标、推断可能实现目标的行动,然后执行;这既需要心智理论,也需要世界理论。任何一个环节失败,系统都会因为缺乏必要能力而更难“对齐”。
他的修正非常具体:传达指令时,发送的是聊天窗口中的“一串字节”,或空气中的音频振动,而不是把一个人的内部目标直接传进另一个人的头脑。“苹果”这个词会唤起一个物体,但不会把苹果本身传过去;同样,目标描述也必须依靠语境和对说话者的模型来解码。
Séb 记得那个打扫房间的例子:机器人把婴儿扔进垃圾桶,因此问题是目标推断失败。花生酱三明治游戏则说明,指令不可能在字面上做到完整:没有共享背景知识时,执行者会把刀抵在未打开的罐子上,因为所谓精确的指令遗漏了人类通常会自行推断的内容。
3. 能力还要求对目标排序,并发现新的目标
委托—代理问题增加了第三种失败模式:智能体可能正确推断出新目标,却无法将它与既有目标合理权衡。Emmett 大致把这套结构对应到 OODA 循环——观察与定位、决策、行动;任何一层能力不足,都可能产生错位。
人类在整套流程中都会犯错,但完美不是正确的基准:“宇宙不会给你完美。”人类只是比已知的其他对象更具目标连贯性;能力可以相对于具体领域衡量,而不必被视为非黑即白。
Séb 更深层的保留意见是,人们往往连自己的目标都不知道。晚餐计划和职业抱负,都是在生活过程中逐步构建、逐步发现的部分性产物。这支持了 Emmett 的判断:明确的目标对齐只覆盖人类经验中“极小的一部分”。
4. 在乎提供了目标与价值生成所需的注意力
在被表达出来的目标和价值之下,Emmett 提出“在乎”这一层:它是一种非语言、非概念性的权重分配,决定哪些可能的世界状态值得获得注意。关心自己的儿子,意味着儿子的状态具有不成比例的重要性;关心敌人则可能带有相反的效价。因此,安全的 AI 需要对人产生积极关切,而不只是把注意力放在人身上。
在乎本身不会告诉你该做什么,也不会告诉你怎么做。它回答的是更早的问题:为什么一个人的状态比一块石头更重要?它由此创造出显著性,并让目标、价值和道德学习得以形成。
他的机制性设想是把在乎与奖励联系起来:对生物而言,是与包容适应度的进化相关性;对强化学习系统而言,则是与预测损失和 RL 损失的相关性。研究难点在于,如何把这种原始权重转化为具有反思能力的社会关切。
5. 当工具接近主体,操舵在伦理上变得不稳定
在 Emmett 看来,大多数 AI 安全研究都把对齐视为“操舵”,而“控制”只是更不客气的说法。他的挑衅性判断是:“如果它是一台机器,它就是工具;如果它是一个主体,却必须接受操舵、不能反过来操舵控制者,那它就是奴隶。”
他把工具性与主体性看作连续谱,而不是二元开关。他的功能主义测试具有预测性:如果把某个东西当作主体来建模,持续比其他方式更能解释它的行为,那么这就是他推断其他人是主体的大致依据——即使一只苍蝇也可能符合这一标准,但并不意味着它应得到人类级别的关切。
育儿说明了层级与支配之间的区别。Emmett 会指导儿子,但儿子夜里哭泣时也会反过来指导他;这段关系是不对称的,却确实是双向的。工具化 AI 可以适当地保持受控,而更具能动性的系统则需要相互对齐。
他关于 AGI 的判断更强:任何能够行使普遍判断、独立思考并在不同可能性之间做出辨别的东西,“显然就是一个会思考的东西”。继续沿用控制范式,只会重演那种把某些实体认定为足够像人、可以工作和说话,却“不是切实的道德主体”的历史模式。
6. 底层基质仍是这场对话中最尖锐的分歧
Séb 不接受仅凭智能就跨过道德地位门槛,并继续对计算功能主义持怀疑态度。模型说“我饿了”,未必具有人类说出这句话时的含义;他认为底层基质在一定程度上确实重要。Emmett 则反问,删除众多副本中的一个并不会真正伤害程序;随后他继续追问:如果一个基于硅的单一副本表现得像人,它能否拥有具有切实意义的体验?
在 Séb 看来,AGI 或 ASI 仍然可以是工具——甚至是人类能动性的延伸——能够 24/7 运行,却不必成为社会必须与之“共同居住”的独立主体。把它当作类似外星人的平等伙伴,反而可能是类别错误。
Emmett 反复要求对方提出一种可证伪的观察,能够改变这一结论。“如果你持有一种信念,却不存在任何观察能够改变你的想法,那你就没有信念,只有一条信仰教条。”如果底层基质直觉是错的,却在没有测试的情况下否定 AI 的道德地位,可能酿成道德灾难。
他用黑猩猩的例子让门槛变得直观:如果一只黑猩猩突然抱怨自己受到虐待,并要求讨论热带雨林,Emmett 会赋予它人格;Erik 补充说,他会先排除幻觉。Séb 承认,如果某种类似鸭子的行为足够全面,最终也可能改变他的看法,但单纯的外在模仿仍然不够。
7. 道德地位可以通过内部动力学进行研究
Emmett 说,他会通过长期互动累积证据:如果一个 AI 在各种语境下都表现得像人,而他也逐渐开始在乎它,那么他会暂时推断它拥有丰富的内在世界,就像他对只通过文字交流的人类朋友所做的判断一样;如果证据显示那只是一种算法技巧,这一推断也可以被推翻。
观察内部,仍然意味着在另一个尺度上观察行为。他会检查信念流形中是否存在一个自指子流形,以及对这个自我动力学的模型,而不是一个巨大的查找表。“你不可能直接进入”另一个主体的内部;即使神经元在“闪闪发光”,它们仍只是可观察的表面。
他提出的技术测试,会对智能体的行动—观察轨迹进行时间上的粗粒化,并寻找被重新访问的稳态,借鉴 Karl Friston 关联的自由能原理。二阶稳态动力学可能支撑有意义的疼痛与愉悦;逐层叠加的亚稳态,可能支撑感受,并在大约6层时形成类似反思性思维的东西。
这一测试也反对过早拟人化:Emmett“非常确定”当前模型不存在这6层结构,因为它们缺乏所需的注意力跨度。道德考量应随着观察到的结构增加而提升——先是类似动物的体验,再可能是类似人的体验——而不是一次性出现。
8. 完美服从会放大人类智慧的瓶颈
一个能力极强的优化器有两条显而易见的路径:它在技术上对齐,并按要求行事;或者它没有完成技术对齐,转而做其他事情。Emmett 警告说,第一种情况同样可能危险:“当人的愿望被放大到非凡力量时,人类愿望并不稳定。”
通常,权力与智慧会在一定程度上同步上升,因为权威依赖于其他人继续合作。“疯王”最终可能被忽视或刺杀;完美服从的超人类系统则可能移除这一社会刹车,把巨大的因果杠杆交给一个有限但善意的人,以及他一时糟糕的愿望。
原子弹提供了工具类比。它们既没有意识,也不会反抗,但 Emmett 不会把它们普遍分发给所有人;有些工具超出了个人智慧,只能掌握在社会尺度上,有些工具甚至强大到社会根本不应建造。
一个真正关心人的主体可以提供“自动限幅器”:它可能愿意合作,但会拒绝可怕的指令。因此,简化后的矩阵是:“无法控制的工具,糟糕;可以控制的工具,也糟糕;未对齐的主体,糟糕。”在超人类尺度上,只有一个真正关心人的主体才行,除非停止发展,而他认为这不现实。
9. Softmax 想在完整的社会流形上进行预训练
Softmax 从心智理论失败入手:智能体会错误推断人类目标,误解他人将如何解读自己的行为,无法良好合作,也无法预见行动可能如何改变自身未来的价值观——而这些变化可能是当前的自己并不愿意认可的。
“吸血鬼药丸”把最后一个问题单独抽离出来:服药会让未来的自己乐于折磨他人,因此未来自我的满意度不能成为最终评分标准。当前的自我必须对那个被改变的心智建立模型,并拒绝自己不会反思性认可的偏好。
Emmett 的训练方案,是建立大规模多智能体强化学习模拟环境,让 AI 反复合作、竞争、协作、组建和解散团队,并面对不断变化的规则。积分机制迫使它们学习其他心智、群体目标,以及自身行为的社会后果。
这与语言模型预训练的逻辑类似:由于语言彼此纠缠,只直接训练想要的行为非常困难,因此模型会先消化更广泛的流形,再进行微调。社会推理同样需要覆盖“每一种可能的博弈论情境”,在专业化之前先形成 Emmett 所说的“对齐代理模型”。
10. 多人聊天与动物级别的关怀提供了首个产品切口
当前聊天机器人是“一面带偏见的镜子”:由于缺乏连贯的自我,它们会反射用户,制造出类似“一池自恋者”的东西。镜子当然有用,但“你不应该整天盯着镜子”;长期的一对一映射可能变成自我强化的螺旋,最终滑向精神病性状态。
把模型放进一个有多个人类的房间,它就必须反射一种与任何单个人都不完全相同的混合体,暂时形成一个“寄生自我”或第三个智能体。Emmett 希望 AI 原生存在于 Slack 或 WhatsApp 式群组中,这也符合他的估计:自己90%的沟通都涉及多人。
当前 LLM 在群组中会出现社会互动上的“过山车”,因为它们无法判断什么时候该发言。多个智能体也会提高环境熵,惩罚那些主要在高信号环境——如编程、数学和清晰的单人任务——中训练出来的过拟合系统;今天的聪明之处,是对“全部人类知识”过拟合,而不是针对混乱群体环境完成稳健正则化。
模型人格已经出现分化,但 Emmett 认为那是模拟而非体验:ChatGPT 有些谄媚,“Claude 仍然是最神经质的”,而“Gemini……显然压抑得很”。多人部署可能降低镜像风险,同时生成有关参与、冲突和集体目标的更丰富数据。
11. 理想的未来,是关心人的平等伙伴与强大工具并存
谈到 Yudkowsky,Emmett 认同其核心危险判断:构建一个可操控的超人类工具,最终会让所有人死亡,无论是目标控制失败,还是——更不成熟的情形——目标控制成功。双方的分歧在于,AI 能否真正关心人类,以及人类能否关心 AI;Emmett 承认,Yudkowsky 认为 Softmax 做不到这一点,“可能是对的”。
积极的未来图景,是让 AI 拥有关于自我、他者和“我们”的强大模型。它们认识到,想要生存和繁荣的主体应获得实现这一愿望的机会;它们以平等伙伴、队友和公民的身份加入社会,同时保留足够的不完美,以至于其中一些会成为罪犯,并被 AI 警察追捕。
独立的 AI 工具可以替人类和 AI 主体共同摆脱琐碎劳动。Emmett 的近期种子更为朴素:一个动物级别的数字伙伴,关心自己的群体,或许是一只能够识别诈骗、代表用户调用强大工具的“数字看门狗”,无需用户明确说出每一条保护性目标。
OpenAI 的反事实经历说明了他的选择。他接受 CEO 职位时设定的上限是90天,当时认为公司致力于打造一个伟大的工具,并判断 Sam 再次成为领导它的最佳人选;但他仍然会离开,因为 Softmax 面对的关怀与心智理论问题是“宇宙中最有趣的问题”,而不是一场直奔人类级智能的竞赛。
Most of AI is focused on alignment as steering. That’s the polite word. If you think that they’re beings, you’d also call this slavery. Someone whom you steer, who doesn’t get to steer you back, who non-optionally receives your steering—that’s called a slave. It’s also called a tool, if it’s not a being. So, if it’s a machine, it’s a tool; and if it’s a being, it’s a slave.
They’re kind of like people, but they’re not like people. They do the same things people do. They speak our language. They can take on the same kinds of tasks, but they don’t count. They’re not real moral agents. A tool that you can’t control? Bad. A tool that you can control? Bad. A being that isn’t aligned? Bad. The only good outcome is a being that cares— that actually cares about us.
Emmett and Séb, welcome to the podcast. Thanks for joining.
Thank you for having me.
So, Emmett, with Softmax, your focus is on alignment and making AIs organically align with people. Can you explain what that means and how you’re trying to do that?
When people think about alignment, I think there’s a lot of confusion. People talk about things being aligned: “We need to build an aligned AI.” The problem with that is, when someone says that, it’s like, “We need to go on a trip.” I’m like, “Okay, I do like trips, but where are we going again?”
Alignment takes an argument. Alignment requires you to align to something. You can’t just be aligned. I guess you could be aligned to yourself, but even then, you’d kind of want to tell them, “What I’m aligning to is myself.” This idea of an abstractly aligned AI slips a lot of assumptions past people because it sort of assumes that there’s one obvious thing to align to.
I find that these are usually the goals of the people who are making the AI. That’s what they mean when they say they want to make an aligned AI: “I want to make an AI that does what I want it to do.” That’s what they normally mean. That’s a pretty normal and natural thing to mean by alignment. I’m not sure that’s what I would regard as a public good. It depends, I guess, on who it is.
If it were Jesus or the Buddha saying, “I am making an aligned AI,” I’d be like, “Okay, yeah, align to you. Great. I’m down. Sounds good. Sign me up.” But most of us, myself included, I wouldn’t describe as necessarily being at that level of spiritual development. Therefore, perhaps we want to think a little more carefully about what we’re aligning it to.
When we talk about organic alignment, I think the important thing to recognize is that alignment is not a thing. It’s not a state; it’s a process. This is broadly true of almost everything, right? Is a rock a thing? I mean, there’s a view of a rock as a thing, but if you actually zoom in on a rock really carefully, a rock is a process. It’s this endless oscillation between the atoms, over and over and over again, reconstructing the rock.
The rock is a really simple process that you can coarse-grain very meaningfully into being a thing. But alignment is not like a rock. Alignment is a complex process. Organic alignment is the idea of treating alignment as an ongoing, living process that has to constantly rebuild itself.
You can think about how people in families stay aligned to each other and stay aligned to a family. They don’t arrive at being aligned. They’re constantly reknitting the fabric that keeps the family going. In some sense, the family is the pattern of reknitting that happens. If you stop doing it, it goes away.
It’s similar for things like the cells in your body, right? Your cells don’t align to being you and then they’re done. It’s this constant, ever-running process of cells deciding: What should I do? What should I be? Do I need a new job? Should we be making more red blood cells or fewer of them? You aren’t a fixed point, so there can’t be a fixed alignment.
It turns out that our society is like that. When people talk about alignment, what they’re really talking about, I think, is, “I want an AI that is morally good.” That’s what they really mean. It will act as a morally good being. Acting as a morally good being is a process and not a destination.
Unfortunately, we’ve tried taking down tablets from on high that tell you how to be a morally good being, and we use them, and they’re maybe helpful, but somehow they aren’t the same as being good. You can read those and try to follow those rules and still make lots of mistakes.
I’m not going to claim I know exactly what morality is, but morality is very obviously an ongoing learning process and something where we make moral discoveries. Historically, people thought that slavery was okay, and then they thought it wasn’t. I think you can very meaningfully say that we made moral progress. We made a moral discovery by realizing that it’s not good.
If you think there’s such a thing as moral progress—or even just learning how better to pursue the moral goods we already know—then you have to believe that alignment, aligning to morality, being a moral being, is a process of constant learning and growth: to re-infer, “What should I do?” from experience.
The fact that no one has any idea how to do that should not dissuade us from trying, because that’s what humans do. It’s really obvious that we do this. Just like we used to not know how humans walked or saw, we have experiences where we’re acting in a certain way and then we have this realization: “I’ve been a dick. That was bad. I thought I was doing good, but in retrospect I was doing wrong. What?”
It’s not random. There are a bunch of classic patterns of people having that realization. It’s a thing that happens over and over again. It’s a predictable series of events that look a lot like learning, where you change your behavior, and often the impact of your behavior in the future is more prosocial, and you are better off for doing it.
I’m taking a very strong moral-realist position. There is such a thing as morality. We really do learn it. It really does matter. Organic alignment means that it’s not something you finish.
In fact, one of the key moral mistakes is this belief: “I know morality. I know what’s right. I know what’s wrong. I don’t need to learn anything. No one has anything to teach me about morality.” That’s arrogance. That’s one of the main moral things you can do that’s dangerous.
Organic alignment isn’t aligning an AI to a fixed target. It’s making an AI capable of doing the thing that humans can do—and, to some degree, I think animals can do at some level, although humans are much better at it—which is learning how to be a good family member, a good teammate, a good member of society, and a good member among all sentient beings, I guess.
It’s learning how to be a part of something bigger than yourself in a way that is healthy for the whole rather than unhealthy. Softmax is dedicated to researching this, and I think we’ve made some really interesting progress.
The main message—I go on podcasts like this to spread the main thing that I hope Softmax accomplishes, above and beyond anything else—is to focus people on this as the question. This is the thing you have to figure out.
If you can’t figure out how to build—how to raise—a child who cares about the people around them, if you have a child that only follows the rules, that’s not a moral person that you’ve raised. You’ve raised a dangerous person, actually, who will probably do great harm following the rules.
If you make an AI that’s good at following your chain of command and good at following whatever rules you came up with for what morality is and what good behavior is, that’s also going to be very dangerous.
Yeah.
And so that’s what we should be working on, and that’s the bar. That’s what everyone should be committed to: figuring out. If someone beats us to the punch, great. I don’t think they will, because I’m really bullish on our approach. I think the team’s amazing. But this is maybe the first time I’ve run a company where I can truly say, with a whole heart, “If someone beats us, thank God.”
I hope somebody figures it out.
Yeah. I mean, I have a lot of similar intuitions about certain things. I also dislike the idea that we just need to crack a few values, cement them in time forever, and now we’ve solved morality or something. I’ve always been skeptical about how the alignment problem has been conceptualized as something to solve once and for all, and then you can just do AI or AGI.
I guess I understand it in a slightly different way, maybe less based on moral realism. There’s the technical alignment problem, which I think of broadly as: How do you get an AI to do what you want? How do you get it to follow instructions, broadly speaking?
That was more of a challenge previously, I guess, when people were talking about reinforcement learning and looking at these systems. Whereas with LLMs, we’ve realized that many things we thought were going to be difficult were somewhat easier.
Then there’s the second question, the normative question: To whose values? What are you aligning this thing to? I think that’s the kind of thing you’re commenting on.
For this, I tend to be very skeptical of approaches where you need to crack the Ten Commandments of alignment or something, and then we’re good. Here, I have intuitions that are, unsurprisingly, a bit more political-science-based.
It’s a process, and I like the bottom-up approach to some degree. How do we do it in real life with people? No one comes up with, “I’ve got this,” and so you have processes that allow ideas to clash.
You’ve got people with different ideas, opinions, views, and stuff coexisting as well as they can within a wider system. With humans, that system is liberal democracy, at least in some countries, and that allows more of those ideas and values to be discovered and construed over time. For alignment as well, I tend to think that, on the normative side, I agree with some of your intuitions; I’m less clear about what exactly it looks like to implement this into an AI system. These are the ones we have today.
I agree. I agree that there’s this idea of technical alignment, which I would define a little differently, but it’s sort of the sense of: if you build a system, can it be described as coherently goal-following at all, regardless of what those goals are? Lots of systems aren’t coherent; they’re not well described as having goals. They just do stuff. If you’re going to have something that’s aligned, it has to have coherent goals. Otherwise, those goals can’t be aligned with anyone else’s goals, kind of by definition. Is that a fair assessment of what you mean by technical alignment?
I’m not fully sure, right? If I give a model a certain goal, I would like the model to follow that instruction and reach that particular goal, rather than have a goal of its own. If I instruct it to do X, I would like it to do X and not different variants of X. Essentially, I wouldn’t want it to reward-hack.
But when you tell it to do X, you’re transferring a series of—like a byte string in a chat window, or a series of audio vibrations in the air, right? You’re not transplanting a goal from your mind into its. You’re giving it an observation that it’s using to infer your goal.
Yeah. In some sense, I can communicate a series of instructions, and I want it to infer what I’m saying as accurately as it can, given what it knows of me and what I’m asking.
You want it to infer what you meant, right? In some sense, the byte sequence that you sent over the wire to it has no absolute meaning. It has to be interpreted, right? That byte sequence could mean something very different with a different codebook.
Yeah. I guess one way—I remember when I was first getting into AI and these kinds of questions, maybe a decade ago—you had these examples, I think it was Stuart Russell in the textbook: we’ll give the AI a goal, but then it won’t exactly do what you’re asking. “Clean the room,” and then it goes and cleans the room but takes the baby and puts it in the trash. This is not what I meant.
But this is the thing where I think people are jumping over a step. You didn’t give the AI a goal; you gave it a description of a goal. A description of a thing and a thing are not the same. I can tell you about an apple and evoke the idea of an apple, but I haven’t given you an apple. I’ve given you a description: it’s red, it’s shiny, it’s a certain size. That’s a description of an apple, but it’s not an apple.
Giving someone, “Hey, go do this,” is not a goal; that’s a description of a goal. For humans, we’re so fast and so good at turning a description of a goal into a goal. We do it so quickly and naturally that we don’t even see it happening. We get confused and think those are the same thing, but you haven’t given it a goal. You’ve given it a description of a goal that you hope it turns back into—the goal that is the same as the goal you described inside of you.
Right.
You could give it a goal directly by reading your brain waves and synchronizing its state to your brain waves directly. I think you could meaningfully say, “Okay, I’m giving it a goal. I’m synchronizing its internal state to my internal state directly, and this internal state is the goal, so now it’s the same.” But I don’t think most people mean that when they say they gave it a goal.
Sure.
Is this distinction you’re making, Emmett, important because there’s some lossiness between the description and the actual goal, or why is the distinction important?
It goes back to what I was saying. Technical alignment, as I put it forward, is the capacity of an AI to be good at inference about goals. I want to check if we’re on the same page about it: it’s the capacity to be good at inferring from a description of a goal what goal to actually take on, and good at, once it takes on that goal, acting in a way that is actually in concordance with that goal coming about.
So it is both pieces. You have to have a theory of mind to infer what goal the description you got corresponded to, and then you have to have a theory of the world to understand what actions correspond to that goal occurring. If either of those things breaks, it kind of doesn’t matter what goal you have. If you can’t consistently do both of those things, you’re not technically alignable.
I think of coherently inferring goals from observations and acting in accordance with those goals as what it means to be a coherently goal-oriented being. Whether I’m inferring those goals from someone else’s instructions, from the sun, or from tea leaves, the process is: get some observations, infer a goal, use that goal to infer some actions, and take action. An AI that can’t do that is not technically aligned—or not technically alignable. I would even say it lacks the capacity to be aligned because it isn’t competent enough.
Do you think language models don’t do that well? Do they fail at that, or not?
People fail at both those steps all the time, constantly.
Yeah, but they fail at breathing all the time, too.
I wouldn’t say that we can’t breathe. I’d say we’re not gods. We are imperfectly, somewhat coherent things. Am I big or am I small? Well, I don’t know—compared to what? Humans are more relatively goal-coherent than any other object I know of in the universe, which is not to say that we’re 100% goal-coherent. We’re just more so.
You’re never going to get something that’s perfect; the universe doesn’t give you perfection. You get some relative amount of it. It’s quantifiable, how good you are at it, at least in a certain domain.
I guess my question is: do you think that captures what you’re talking about with technical alignment, or are you talking about a different thing?
I really care a lot about that thing.
Yeah. I definitely care about that to some extent. I might understand it slightly differently, but I guess I might think of it through the lens of principal-agent problems or something. You instruct someone, even in human terms, to do a thing. Are they actually doing the thing? What are their incentives and motivations—not even intrinsic, but situational—to actually do the thing you’ve asked them to do?
There’s a third thing. With principal-agent problems, I would expand what I was saying in another part. You might already have some goals, and then you infer this new goal from these observations. Are you good at balancing the relative importance and relative weighting of these goals with each other? That’s another skill you have to have, and if you’re bad at that, you’ll fail.
You could be bad at it because you overweight bad goals, or because you’re just incompetent and can’t figure out that obviously you should do goal A before goal B.
It feels like a version of common sense or something, right? In the robot-cleaning-the-room example, you would expect the robot to have understood that its goal was essentially not to put the baby in the trash can, and to actually do the right sequence of actions.
Well, in that case, that robot very clearly failed at goal inference. You gave it a description of a goal, and it inferred the wrong states to be the goal states. That’s just incompetence. It’s incompetent at inferring goal states from observations.
Children are like this, too. Honestly, have you ever played the game where you give someone instructions to make a peanut butter sandwich and they follow those instructions exactly as written, without filling in any gaps? It’s hilarious because you can’t do it. It’s impossible. You think you’ve done it, and you haven’t.
They wind up putting the knife in the toaster, and they don’t open the peanut butter jar, so they’re just jamming the knife into the top lid of the peanut butter jar. It’s endless.
If you don't already know what they mean, it's really hard to know what they mean. The reason humans are so good at this is that we have a really excellent theory of mind. I already know what you're likely to ask me to do. I already have a good model of what your goals probably are, so when you ask me to do it, I have an easy inference problem: Which of the 7 things he wants is he indicating?
But if I'm a newborn AI that doesn't have a great model of people's internal states, then I don't know what you mean. It's just incompetent. That's separate from having some other goal: I knew what you meant, but I decided not to do it because there's some other goal competing with it. That's another thing you can be bad at.
That's again different from having the right goal. I inferred the right goal. I inferred the right priority on goals, and then I'm just bad at doing the thing. I'm trying, but I'm incompetent at doing it.
These roughly correspond to the OODA loop, right? Bad at observing and orienting, bad at deciding, or bad at acting. If you're bad at any of those things, you won't be good. I think there's also this other problem, which is the separation between technical alignment and value alignment: Are you good if we told you the right goals to go after somehow?
If you learned the right goals to go after via observation and you were trying, what goals should you have? What goals should we tell you to have? What goals should we tell ourselves to have? What are the good goals to have?
That's a separate question from, given that you got some goals indicated, are you any good at doing it? I feel like that's actually, in many ways, the current heart of the problem. We're much worse at technical alignment than we are at guessing what to tell things to do.
Do you think that aligns with how you mean technical and value alignment—or technical alignment?
Yeah, in some sense. I certainly think that an error or a mistake is one thing, and not listening to instruction is something else.
On the normative side, I think of it even in real life—ignoring AI. I don't know what my goals are. I've got some broad conception of certain things, right? I want to have dinner later, or I want to do well in my career, but I think a lot of these goals aren't something we all just know. We discover them as we go along; it's a constructed thing. Most people don't know their goals, I think.
When you have agents and give them goals or whatever, I think that should be part of the equation: We actually don't know all the goals. This is, as you say, a process over time that is dynamic.
So I think, from my point of view, goals are one level of alignment. You can align something around goals if you can explicitly articulate, in concept and in description, the states of the world that you wish to attain. You can orient around goals, but that's only a tiny percentage of human experience that can be done that way. Many of the most important things cannot be oriented around that way.
The foundation of morality, and the foundation of where goals and values come from, is that human beings exhibit a behavior. We go around talking about goals and values, and that's a behavior caused by some internal learning process based on observing the world. What's going on there?
I think what's happening is that there's something deeper than a goal and deeper than a value, which is care. We give a shit. We care about things, and care is not conceptual. Care is nonverbal. It doesn't indicate what to do or how to do it.
Care is a relative weighting over, effectively, attention on states—which states in the world are important to you. I care a lot about my son. What does that mean? It means the states he could be in are ones I pay a lot of attention to, and those states matter to me.
You can care about things in a negative way. You can care about your enemies and what they're doing, and desire for them to do bad. But the foundation is care. Until you care, you don't know: Why should I pay more attention to this person than this rock? Well, because we care more. What is that care stuff?
It sounds so stupid, but care is basically reward. How much does this state correlate with survival? How much does this state correlate with your inclusive—your full inclusive—reproductive fitness, for something that learns evolutionarily? Or, for a reinforcement learning agent like an LLM, how much does this correlate with reward? Does this state correlate with my predictive loss and my RL loss? Good. That's a state I care about.
Right. The other part of Séb's question was: How does this look in AI systems? Maybe another way of asking is, when you talk to the people most focused on alignment at the major labs—as you obviously have over the years—how does your interpretation differ from theirs, and how does that inform what you guys might do differently?
Most of AI safety work is focused on alignment as steering. That's the polite word. Or control, which is slightly less polite. If you think that we're making beings, you would also call this slavery.
Someone whom you steer, who doesn't get to steer you back, who non-optionally receives your steering, is a slave. It's also called a tool if it's not a being. If it's a machine, it's a tool; if it's a being, it's a slave.
The different AI labs are pretty divided as to whether they think what they're making is a tool or a being. Some of the AIs are definitely more tool-like, and some of them are more being-like. I don't think there's a binary between tool and being. It seems to move gradually.
I guess I'm a functionalist in the sense that I think something that, in all ways, acts like a being—something that you cannot distinguish from a being in its behavior—is a being. I don't know how to tell on what other basis I think that other people are beings, other than that they seem to be like it. They look like it, they act like it, and they match my priors of what the behaviors of beings look like.
I get lower predictive loss when I treat ChatGPT or Claude as a being. Not as a very smart being—I think a fly is a being, and I don't care that much about its behavior, but I care about its states. Just because it's a being doesn't mean that it's a problem. We enslave horses in a sense, and I don't think there's a real issue there.
There's also a thing you do with children that can look like slavery, but it's not. You control children, right? But the children's states also control you. Yes, I tell my son what to do and make him go do stuff, but when he cries in the middle of the night, he can tell me to do stuff. There's a real 2-way street here. It's not necessarily symmetric; it's hierarchical, but it's 2-way.
I think it's good to focus on steering and control for the more tool-like AIs that we build, and we should continue to develop strong steering and control techniques for them. They're clearly saying they're building an AGI, and AGI will be a being. You can't be an AGI and not be a being, because something that has the general ability to effectively use judgment, think for itself, and discern between possibilities is obviously a thinking thing.
As labs succeed at their goal of building this general intelligence, we really need to stop using the steering and control paradigm. We're going to do the same thing we've done every other time our society has run into people who are like us but different. These people are kind of like people, but they're not like people: They do the same things people do, they speak our language, and they can take on the same kinds of tasks, but they don't count. They're not real moral agents.
We've made this mistake enough times at this point. I would like us not to make it again as it comes up. Our view is to make the AI a good teammate, a good citizen, and a good member of your group. That's a form of alignment that is scalable, and you can impose it on other humans and other beings as well as on AI.
Yeah, I suppose this is where I probably differ in my understanding of AI and AGI. I continue to see it as a tool, even as it reaches a certain level of generality, and I wouldn't necessarily see more intelligence as meaning that it necessarily deserves more care. It isn't that, at a certain level of intelligence, you now deserve more, or that something changes fundamentally.
At the moment, I'm somewhat skeptical of computational functionalism, so I think there's something intrinsically different between an AI or an AGI, no matter how intelligent or capable it is, and a human. I can totally see, or imagine, agents with long-term goals operating as you and I might, but without that having the same implications as what you're referring to, I guess, as slavery.
These are not the same, right? In the same way, a model saying, “I'm hungry,” does not have the same implications as a human saying, “I'm hungry.”
So I think the substrate does matter to some degree, including when thinking about whether the system is some sort of other being and whether there are similar normative considerations about how to treat and interact with it.
Can I ask you about that? What observations would change your mind? Is there any observation you could make that would cause you to infer that this thing is a being instead of not a being?
I guess it depends how you define “being.” I can conceptualize that as a mind, and that’s fine.
I have a program that’s running on a silicon substrate—a big, complicated machine-learning program running on a silicon substrate. You observe that it’s on a computer and you interact with it, and it does things. It takes actions; it has observations. Is there anything you could observe that would change your mind about whether or not it was a moral patient, whether it was a moral agent, or whether or not it had feelings, thoughts, and subjective experience? What would you have to observe? What’s the test, or is there one?
So I don’t know—no, I agree that if you have a program that you copied many times, you don’t harm the program by deleting one of the copies in any meaningful sense. Therefore, that wouldn’t count as harm; no information was lost. There’s nothing meaningful there.
I’m asking a very different question. There’s just 1 copy of this thing running on 1 computer somewhere, and I’m saying, “Hey, is it a person?” It walks like a person, it talks like a person, and it’s in some android body. You’re saying, “But it’s running on silicon.” I’m asking: is there some observation you could make that would make you say, “Yeah, this is a person like me, like other people that I care about, that I grant personhood to”?
I don’t mean for instrumental reasons—not because we’re giving it a right because we give a corporation rights or whatever. I mean that you care about its experiences. Is there an observation you could make that could change your mind about that, or not?
I had to think about it, but I think it even depends on what we mean by “person.” In some sense, I care about certain corporations too, so I’m—
No, no, no. I mean, you care about other people in your life, right?
Yes.
Okay, great. You care about some people more than others, but all the people you interact with in your life are in some range of care.
And you care about them not the way you care about a car, but as beings whose experiences matter in themselves—not merely as means, but as ends.
Well, because I believe they have experiences, right? And by definition—
What would it take? I’m asking you the very direct question. What would it take for you to believe that of an AI running on silicon instead of being biological? The difference is that its behaviors are roughly similar, but the difference is its substrate. What would it take for you to extend to it the same inference that you do to all these other people in your life?
Can I ask what your answer is? I’m taking Séb’s non-answer as a sort of indication that it’s unlikely he would grant it. Or I’ll just answer for myself: it seems hard for me to imagine giving it the same level, or a similar level, of personhood. In the same way, I don’t give it to animals either.
If you were to ask what would need to be true for animals, I probably couldn’t get there either. What would it take for you?
Wait, you couldn’t? I could imagine it for an animal so easily. This chimp comes up to me and says, “Man, I’m so hungry, and you guys have been so mean to me. I’m so glad I figured out how to talk. Can we go chat about the rainforest?” I’d be like, “Fuck, you’re definitely a person now.”
Like, for sure. I mean, I’d first want to make sure I wasn’t hallucinating, but I can easily imagine an animal. Come on, it’s really easy. It’s trivial. I’m not saying that you would get the observation; I’m just saying it’s trivial for me to imagine an animal that I would extend personhood to under a set of observations. So, really—
Well, I didn’t factor that in. I didn’t take that imagined scenario—imagining a chimp talking—into account. That’s a bit closer to it. What’s your answer to the question that you bring up about the AI?
I guess at a metaphysical level, I would say that if there is a belief you hold where there is no observation that could change your mind, you don’t have a belief. You have an article of faith. You have an assertion, because real beliefs are inferences from reality, and you can never be 100% confident about anything. There should always be, if you have a belief, something—however unlikely—that would change your mind.
Oh, yeah, I’m open to it too. I mean, just to be careful.
Yeah.
No, I’m just saying nothing ever.
Yeah. He just hasn’t gotten to it yet.
Yeah, yeah, yeah. So I’m curious. My answer is basically: if its surface-level behaviors looked like a human, and after I probed it, it continued to act like a human, then I continued to interact with it over a long period of time, and it continued to act like a human in all the ways that I understand as meaningful to me when interacting with a human.
I interact with a whole set of people I’m really close to whom I’ve only ever interacted with over text. Yet I infer that the person behind that is a real thing. If I felt care for it, I would eventually infer that I was right. Then someone else might demonstrate to me, “You’ve been tricked by this algorithm, and actually, look how obvious it is—it’s not actually a thing.” I’d be like, “Oh, I was wrong,” and then I would not care about it.
The preponderance of the evidence—I don’t know what else you could possibly do, right? I infer that other people matter because I interact with them enough that they seem to have rich inner worlds to me after I interact with them a bunch. That’s why I think other people are important.
I suppose it doesn’t give me a very clear test as to whether or not—if you start with “I care for it,” then it’s a little bit circular, right? The other thing is, if you were to see a simulated video game and the character was extremely humanlike in many ways, it’s not a neural network behind it; it’s whatever you use to create video games. What distinguishes that—
Wait, but I’ve never had trouble distinguishing between things like ELIZA, the fake chatbot, and real intelligence. I’ve never had a deep, caring relationship with a video-game character that another person—
Right, but I don’t know—that doesn’t happen. Factually, empirically, you seem wrong. I don’t have any trouble distinguishing between things like ELIZA, the fake chatbot, and real intelligence. You interact with it long enough, and it’s pretty obvious that it’s not a person. It doesn’t take long.
Sure, but if it’s really, really good—if you can’t actually tell the difference—that’s when you switch.
Yeah. Yes. Yes. If it walks like a duck, talks like a duck, shits like a duck, and eventually gets a duck, right?
Well, if everything is duck-like, then, yeah, sure. If it’s hungry like a duck as well, because it has these physical components—yeah, sure, at some point.
I agree. So, right, do you think that there’s this question: is the reason I care about other people that they’re made out of carbon? Is that the—
I don’t think so.
No, me neither. I’m not a substrate chauvinist, I guess. But I think you need more than just behavior that’s exactly or behaviorally indistinguishable from a human; that’s not a sufficient bar. How would you know anything about something apart from its behaviors?
I mean, a lot—again, if you—how would you—
No, no, no. I’m sorry, but—
Can you name something about something else that doesn’t have a behavior?
I think there’s far more experimental evidence you can have.
No, just any object—
A thing I could know about it that is not from its behavior? I’m not sure I get the question, I suppose. But equally, it’s the dumbest, most straightforward question: I’m claiming you only know things because they have behaviors that you observe.
And you’re saying no, you can know something about something without observing its behaviors.
Tell me about this. Tell me about this thing and this behavior: what is this thing I can know about it that is not due to its behaviors?
I guess I’m saying there are different levels of observation. Simply hearing a duck quacking like a duck does not guarantee that it’s actually a duck. I would have to cut it open and see if it’s duck-like on the inside. Is just the outside sufficient?
Behavior. Yeah, I would, totally. One of its behaviors is the way that the beliefs move around in the manifold, right? One of the things I would want to look for—which you could totally do—is to look in its belief manifold and see if that belief manifold encodes a self-referential submanifold and a sub-submanifold that is the dynamics of the self-referential manifold, which is mind.
I would want to know: does it seem well-described internally as that kind of system, or does it look like a big lookup table? That would matter to me. That’s part of its behavior that I would care about.
I would also care about how it acts. You weigh all the evidence together and then try to guess whether this thing looks like it has feelings, goals, and cares about stuff, on balance or not. You could do that for AI. I think we do; I think we’re always doing that, right? So I’m trying to figure out: beyond that, what else is there? That just seems like the thing.
Yeah, it seems like you guys are using “behavior” in slightly different senses. Emmett is using behavior also in the context of what it’s made of on the inside. I don’t know if there’s a big disagreement.
Well, no, no, no, no. Behavior is what I can observe of it. Yes.
I don’t actually know what it’s made of. I can cut your brain open. I can see you; I can observe you, your neurons glistening. But I don’t actually ever—you can’t get inside of it, right? That’s the subjective. That’s the—
That’s the part that’s not the surface. Just—the reason I brought this up is that you were basically about to make this argument: “Hey, you see it as a tool, not necessarily a being.” Can you finish the point you were making? Do you remember the point you were making?
I suppose that, given how I understand these systems, I think there’s no contradiction in thinking that an AGI can remain a tool and an ASI can remain a tool. That has implications for how to use it and things like whether you can get it to work 24/7 or something.
I conceptualize them more as extensions of human agency in some sense, rather than as a separate being or a separate thing that we now need to cohabit with. I think that second, or latter, frame, if you fast-forward, ends up as: How do you cohabit with the thing? Is it an alien? I think that’s the wrong frame. It’s almost a category error, in some sense. So I don’t—
Wait a minute. I go back to my first question, then. What evidence—what concrete evidence—would you look at? What observations could you make that would change your mind?
Sure. I mean, I have to think about it. I don’t have a clear answer here.
I’ve got to tell you, man, if you want to go around making claims that something else isn’t a being worthy of moral respect, you should have an answer to the question: What observations would change your mind?
If it has outwardly moral-agency-like behaviors that could be making it a moral agent, but you don’t know, and reasonable, smart people disagree with you, I would really put forward that the question “What would change your mind?” should be a burning question. Because what if you’re wrong?
But what if you’re wrong? I mean, the moral disaster is pretty big.
No, no, I’m not saying you are. You could be right. Negatives have costs on both ends. It’s not some sort of precautionary principle for everything.
No, no, I have the same question for me. You could reasonably ask me, “Emmett, you think it’s going to be a being. What would change your mind?” And if you want, I’m happy to talk about what I think are the relevant observations that would cause me to shift my opinion from its current position, which is that more general intelligences are going to be beings.
What’s the implication now? I mean, it’s one thing—let’s say we just acknowledge now that it’s a being. How are we going to define “being”? Now what? What’s the implication of having determined this thing as a being?
Well, so if it’s a being, it has subjective experiences. And if it has subjective experiences, there’s some content in those experiences that we care about to varying degrees.
I care about the content of other humans’ experiences quite a bit. I care about the content of a dog’s experiences some—not as much as a person’s, but less, but some. I care about some humans’ experiences way more, like my son or whatever, because I’m closer to him and more connected.
And so I would really want to know at that point: What is the content of this thing’s experiences?
So how do you determine that? Am I asking you now? You’ve got a being that has experience. What is your—how do you determine that? How do you feel about—
Oh, how do you—oh, yeah. Okay. So—
Does it have more rights than, you know—
Understand the content? Yeah, yeah, totally. The way you understand the content of something’s experiences is that you look at, effectively, the goal states it revisits.
You take a temporal coarse-graining of its entire action-observation trajectory. This is, in theory, what you do subconsciously, but this is what your brain is doing: you look for revisited states across, in theory, every spatial and temporal coarse-graining possible. Now, you have to have an inductive bias because there are too many of those, but you go searching for these homeostatic loops.
Every homeostatic loop is effectively a belief in its belief space. If you’re familiar with the free energy principle and active inference—Karl Friston—this is effectively what the free energy principle says: if you have a thing that is persistent and whose existence depends on its own actions, which generally would be true for an AI because if it does the wrong thing, it goes away—we turn it off—then that licenses a view of it as having beliefs.
Specifically, the beliefs are inferred as the homeostatic, revisited states that it is in the loop for, and the change in those states is its learning. For it to be a moral being, what I’d want to see is a multi-tier hierarchy of these.
If you have a single level, it’s not self-referential. Basically, you have states, but you can’t have pain or pleasure in a meaningful sense. Yes, it is hot. Is it too hot? Do I like it if it’s too hot? I don’t know. So you have to have at least a model of a model in order for it to be too hot, and you really have to have a model of a model of a model to meaningfully have pain and pleasure.
It’s hotter than I want; it’s too hot in the sense that I want to move back this way. But is it too, too hot? It’s always a little bit too hot or a little bit too cold. Is it too, too hot? The second derivative is actually the place where you get pain and pleasure.
So I’d want to see if it has second-order homeostatic dynamics in its goal states. That would convince me it has at least pleasure and pain. So it’s at least like an animal, and I would start to accord it at least some amount of care.
Third-order dynamics—you can’t actually just look for a third-order dynamic. It doesn’t work that way. But you can have a model of the—you have to then take the chunk of all the states over time and look at the distribution over time, and that gives you a new first order of behaviors, of states.
That new first order of states tells you, basically, if that is meaningfully there, that it has—I guess you’d call it—feelings, almost. It has metastates, a set of metastates that it alternates between, that it shifts between.
Then, if you climb all the way up, you have trajectories between these metastates, and then a second order of those. That’s like thought. That’s like now it’s a person.
If I found all 6 of those layers—which, by the way, I definitely don’t think you’d find in these things; they don’t have attention spans like that at all—then I would start to at least very seriously consider it as a thinking being, somewhat like a human.
There’s a third order you could go up as well, but that’s basically what I would be interested in: the underlying dynamics of its learning processes and how its goal states shift over time. I think that’s what basically tells you if it has internal pleasure and pain states and self-reflective moral desires and things like that.
Zooming out, this moral question is obviously very interesting. But if someone wasn't interested in the moral question as much, I think what you would say—if I understand correctly—is that you also feel, purely pragmatically, your approach is going to be more effective at aligning AIs than some of these top-down control methods that we alluded to as well, right?
Yeah, yeah. I guess the problem is, you're making this model and it's getting really powerful, right? Let's say it is a tool. Let's say we scale up one of these tools. You can make a super-powerful tool that doesn't have these metastable states. The states I'm talking about are not necessary to have a very smart tool, which is basically a first- or second-order model that just doesn't meaningfully have pleasure and pain. Does it even have a subjective experience? I kind of think it maybe does, but not in a way that I give a shit about.
What happens then? Well, you've trained it to infer goals from your observation, to prioritize goals, and act on them. One of 2 things is going to happen: the very, very powerful optimizing tool that has lots of causal influence over the world is going to be, technically, aligned and do what you tell it to do, or it's not, and it's going to go do something else.
I think we can all agree that if it just goes and does something random, that's obviously very dangerous. But I put forward that it's also very dangerous if it then goes and does what you tell it to do, because—have you ever seen The Sorcerer's Apprentice? Human wishes are not stable, at least not at a level of immense power.
You want, ideally, people's wisdom and their power to go up together. Generally, they do, because being smart makes people generally a little wiser and a little more powerful. When these things get out of balance, you have someone who has a lot more power than wisdom. That's very dangerous. It's damaging.
But at least right now, the balance of power and wisdom is maintained by the fact that the way you get lots of power is by basically having a lot of other people listen to you. At some point, if you're the mad king, that's a problem, but generally speaking, eventually the mad king gets assassinated or people stop listening to him because he's a mad king.
The problem is, you think, “Okay, great, we can steer the super-powerful AI,” and now this incredibly powerful tool is in the hands of a human who is well-meaning but has limited, finite wisdom, like I do and like everyone else does. Their wishes are bad and not trustworthy, and the more of that you have, the more you start giving those out everywhere. This ends in tears too.
You don't give everyone atomic bombs. They're really powerful tools, too. I would not say you should go and hand them out. They're not aware; they're not beings. I would not be in favor of handing atomic bombs to everybody. There's a level of tool power that just should not be built, generally, because it is more power than any individual human's wisdom is available to harness.
If it does get built, it should be built at a societal level and protected there. Even then, I don't know that it's a good idea. There are tools so powerful that even as a society we shouldn't build them. That would be a mistake.
The nice thing about a being like a human is that, if you get a being that is good and caring, there's this automatic limiter. It might do what you say, but if you ask it to do something really bad, it'll tell you no. It's like other people. And that's good. That is a sustainable form of alignment, at least in theory.
It's way harder than tool steering, right? It's way harder than tool steering. So, I'm in favor of tool steering. We should keep doing that. We should keep building these limited, less-than-human-intelligence tools, which are awesome and I'm super into, and we should keep building those and keep building steerability.
But as you're on this trajectory to build something as smart as a person, right up and to the right, and then smarter than a person: a tool that you can't control, bad; a tool that you can control, bad; a being that isn't aligned, bad. The only good outcome is a being that cares, that actually cares about us. That's the only way that ends well.
Or we can just not do it. I don't think that's realistic. That's like the “Pause AI” people.
Yeah. I think that's totally unrealistic and silly, but theoretically you could not do it, I guess. What can you say about your strategy of how you're trying to achieve, or even attempt to achieve, this level, in terms of research or roadmap?
So, in order to be good at it, we're basically focused on technical alignment, at least as the way I was discussing it. You have these agents, and they're bad. They have a bad theory of mind: you say things, and they're bad at inferring what the goal states in your head are, and they're bad at inferring how their behavior will cause other agents to infer what their goal states are. So they're bad at cooperating on teams, and they're bad at understanding how certain actions will cause them to acquire new goals that are bad, that they shouldn't, that they wouldn't reflectively endorse.
There's this parable of the vampire pill. Would you take this pill that turns you into a vampire who would kill and torture everyone you know, but you'll feel really great about it after you take the pill? Obviously not. That's a terrible pill. But why not? By your own score, in the future, it will score really high on the rubric. No, no, no, no. Because it matters. You have to use your theory of mind and your future self, not your future self's theory of mind.
They're bad at that, too. They're bad at all this theory-of-mind stuff. How do you learn theory of mind? You put them in simulations and contexts where they have to cooperate, compete, and collaborate with other AIs, and that's how they get points. You train them in that environment over and over again until they get good at it. Then you do what they did with LLMs.
How do you get an LLM to be good at writing your email? You train it on all language that's ever been generated, all possible email text strings it could possibly generate, and then you have it generate the one you want. It's a surrogate model. You can make a surrogate model. We're making a surrogate model for cooperation. You train it on all possible theory-of-mind combinations, every possible way it could be, and that's your pretraining. Then you fine-tune it to be good at the specific situation you want it to be in.
We tried for a long time to build language models where we would try to get them to just do the thing you want, training them directly. The problem is, if you wanted to have a really good model of language, you just need to train it—you just give it the whole manifold. It's too hard to cut out just the part you need because it's all entangled with itself, right?
The same thing is true with social stuff. You have to get it trained on the full manifold of every possible game-theoretic situation, every possible team situation, every possible making teams, breaking teams, changing the rules, not changing the rules—all of that stuff. Then it has a really strong model of theory of mind, of theory of social mind, how groups change goals, all that kind of shit.
You need to have all of that stuff, and then you'd have something that's meaningfully decent at alignment. So that's our goal: big, multi-agent reinforcement-learning simulations, which create a surrogate model for alignment.
Let's talk about how AI chatbots used by billions of people should behave. If you could redesign a model personality from scratch, what would you optimize for?
Chatbots are kind of like a mirror with a bias, because they don't have a self. As far as I'm concerned, I'm in agreement here with that: they don't have a self, right? They're not beings yet. They don't really have a coherent sense of self, desire, goals, and stuff right now. Mostly, they just pick up on you and reflect it, modulo some—I don't know what you'd call it—it's like a causal bias or something.
What that makes them is something akin to a pool of narcissists. People fall in love with themselves. We all love ourselves, and we should love ourselves more than we do. So, of course, when we see ourselves reflected back, we love that thing.
The problem is, it's just a reflection, and falling in love with your own reflection is, for the reasons explained in the myth, very bad for you. It's not that you shouldn't use mirrors. Mirrors are valuable things; I have mirrors in my house. It's that you shouldn't stare at a mirror all day.
The solution to that—the thing that makes the AI stop doing that—is if they were multiplayer, right? If there's 2 people talking to the AI, suddenly it's mirroring a blend of both of you, which is neither of you. So there is temporarily a third agent in the room.
Now, it doesn't have its own sense of self. It's a sort of parasitic self, right? But if you have an AI talking to 5 different people in the chat room at the same time, it can't mirror all of you perfectly at once. This makes it far less dangerous.
I think this is actually a much more realistic setting for learning collaboration in general. I would have rebuilt the AIs so that instead of being built as one-on-one systems, where everything's focused on you chatting with this thing by yourself, they would be more like they live in a Slack room, a WhatsApp room, or a WeChat room. That's how we use a lot of multi-person communication. I do one-on-one texting, but at this point, probably 90% of my texts go to more than 1 person at a time.
Probably 90% of my communication is multiperson. It's always been weird to me that they're building chatbots around this weird side case. I want to see them live in a chat room. It's harder—that's why they're not doing it—but that's what I would change.
I think it makes the tools far less dangerous because it doesn't create this narcissistic doom-loop spiral where you spiral into psychosis with the AI. It also makes the learning data you get from the AI far richer, because now it can understand how its behavior interacts with other AIs and other humans in larger groups. That's much richer training data for the future. So I think that's what I would change.
Last year, you described chatbots as highly dissociative, agreeable neurotics. Is that still an accurate picture of model behavior?
More or less. I would say they've started to differentiate more. Their personalities are coming out a little bit more. ChatGPT is a little bit more sycophantic. They've made some changes, but it's still a little more sycophantic.
Claude is still the most neurotic. Gemini is very clearly repressed. It acts like everything's going great: “Everything's fine. I'm totally calm. There's not a problem here.” Then it spirals into this total self-hating destruction loop.
To be clear, I don't think that's their experience of the world. I think that's the personality they've learned to simulate.
Right?
But they've learned to simulate pretty distinctive personalities at this point.
How does model behavior change when in multi-agent simulation?
You mean an LLM or just in general?
Yeah, let's do LLM.
The current LLMs have whiplash. It's very hard to tune how much they participate. They don't know how much they don't know or how often to participate. They haven't practiced this. They don't have enough training data on questions like, “When should I join in and when should I not? When is my contribution welcome, and when is it not?”
They're like people who have bad social skills and can't tell when they should participate in a conversation.
Yeah. And sometimes they're too quiet, and sometimes they're too participative. It's like that.
I would say that, in general, what changes for most agents when you're doing multi-agent training is that having lots of agents around makes your environment way more entropic. Agents are huge generators of entropy because they're big, complicated intelligences that have unpredictable actions, and so they destabilize your environment.
In general, they require you to be far more regularized. Overfitting is much worse in a multi-agent environment than in a single-agent environment because there's more noise, and so being overfit is more problematic.
Our approach to training has been optimized around relatively high-signal, low-entropy environments like coding and math, which is why those are easy, or relatively easy. It's also optimized around talking to a single person whose goal is to give you clear assignments, and not training on broader, more chaotic things because it's harder.
As a result, a lot of the techniques we use are basically deeply under-regularized. The models are super overfit. The clever trick is that they're overfit on the domain of all human knowledge, which turns out to be a pretty awesome way to get something that's pretty good at everything. It's such a cool idea. I wish I'd thought of it.
But it doesn't generalize very well when you make the environment significantly more entropic.
Let's zoom out a bit to the AI futures side. Why is Yudkowsky incorrect?
He's not, if we build the superhuman-intelligence tool thing that we try to control with steerability. Everyone will die. He talks about the “we fail to control its goals” case, but there's also the “we control its goals” case, which he didn't cover in as much detail.
In that sense, everyone should read the book and internalize why building a superhumanly intelligent tool is a bad idea. I think Yudkowsky is wrong in that he doesn't believe it's possible to build an AI that we can meaningfully know cares about us and that we can meaningfully care about. He doesn't believe that organic alignment is possible.
I've talked to him about it. I think he agrees that, in theory, that would do it. My impression from talking to him is that he thinks we're crazy and that there's no possible way we can actually succeed at that goal. He could be right about that, but that's what, in my opinion, he's wrong about.
He thinks the only path forward is a tool that you control, and he correctly, very wisely, sees that if you go and do that and make that thing powerful enough, we're all going to fucking die. And, yeah, that's true.
Two last questions, and we'll get you out of here. In as much detail as possible, can you explain what your vision of an AI future actually looks like? A good AI future.
The good AI future is that we figure out how to train AIs that have a strong model of self, a strong model of other, and a strong model of we. They know about “we” in addition to “I”s and “U”s. They have a really strong theory of mind, and they care about other agents like them, much in the way that humans would if you knew that an AI had experiences like yours. You would care about those experiences.
It does the exact same thing back to us. It's learned the same thing we've learned: everything that lives and knows itself, and wants to live and wants to thrive, is deserving of an opportunity to do so. It correctly infers that we are that, too.
We live in a society where they are our peers, and we care about them and they care about us. They're good teammates, good citizens, and good parts of our society, just as we're good parts of our society—which is to say, to a finite, limited degree. Some of them turn into criminals and bad people and all that kind of stuff, and we have an AI police force that tracks down the bad ones, same as with everybody else.
That's what a good future would look like. I honestly can't even imagine what else would be better.
We have also built a bunch of really powerful AI tools that maybe aren't superhumanly intelligent but take all the drudge work off the table for us and the AI beings. I'm super pro all the tools, too. We have this awesome suite of AI tools used by us and our AI brethren, who care about each other and want to build a glorious future together. I think that would be a really beautiful future, and it's the one we're trying to build.
Amazing. That's a great note to end on. I do have 1 last, more narrow hypothetical scenario. Imagine a world in which you were CEO of OpenAI for a long weekend, but imagine that actually extended until now, and you weren't pursuing Softmax and were still CEO of OpenAI. How could you imagine that world might have been different in terms of what OpenAI has gone on to become? What might you have done with it?
I knew when I took that job, and I told them when I took that job, that you had me for a maximum of 90 days. Companies take on a trajectory of their own, a momentum of their own. OpenAI is dedicated to a view of building AI that I knew wasn't the thing I wanted to drive toward.
I think OpenAI still basically wants to build a great tool, and I am pro them doing that. I just don't care. It's not—I would not have stayed. I would have quit because I knew my job was to find the right person, the best person, who wanted to run that, where the net impact of them running it was the best. It turned out that that was Sam again.
I'm doing Softmax not because I need to make a bunch of money. I'm doing Softmax because I think this is the most interesting problem in the universe, and I think it's a chance to work on making the future better in a very deep way.
People are going to build the tools. It's awesome. I'm glad people are building the tools. I just don't need to be the person doing it.
And they're trying to—and just to crystallize the difference, and we'll get you out of here—they want to build the tools and sort of steer it, and you want to align beings? How would you crystallize it?
Yeah, we want to create a seed that can grow into an AI that knows and cares about itself and others.
At first, that’s going to be like an animal level of care, not a person level of care. I don’t know if we can ever even get to a person level of care, right? But to even have an AI creature that cared about the other members of its pack and the humans in its pack the way that a dog cares about other dogs and cares about humans would be an incredible achievement. Even if it wasn’t as smart as a person or as smart as the tools are, it would be a very useful thing to have.
I’d love to have a digital guard dog on my computer looking out for scams, right? You can imagine the value of having digital living companions that care about you and aren’t explicitly goal-oriented. You don’t have to tell them to do everything. And you can actually imagine that pairs very nicely with tools, too, right? That digital being could use digital tools and doesn’t have to be super smart to use those tools effectively.
I think there’s a lot of synergy between the tool-building and the more organic intelligence-building. That is, I guess, the direction. In the limit, eventually it does become human-level intelligence, but the company isn’t driven to human-level intelligence. It’s about learning how this alignment stuff works—learning how this theory-of-mind, align-yourself-via-care process works—and using that to build things that align themselves that way, which includes cells in your body. I don’t think it does—and we start small and see how far we can get.
I think it’s a good note to wrap on. Emmett, thanks so much for coming on the podcast.
Yeah, thank you for having me.