你的大脑不指挥身体,而是在预测身体。[Max Bennett]
Bennett 的核心论点是:新皮层真正的决定性优势,不是更擅长识别物体,而是拥有足够丰富的世界模型,让动物能够“在经历结果之前先想象结果”。 新皮层并不独自完成规划:丘脑和基底神经节共同帮助暂停行为、选择候选动作、评估模拟后果,并恢复执行。对投资者而言,这意味着机会不只在于打造更大规模的识别器,而在于让架构能够决定何时进行模拟、裁剪反事实搜索,并通过干预学习。
Transformers 验证了部分脑启发叙事——自监督学习确实能产生出人意料的通用表征,但它们仍是“外星大脑”,而非数字化新皮层。 Bennett 强调,它们缺少持续学习、主动检验假设、创造具身数据,以及可靠迁移到新情境的能力;如果让 ChatGPT 不加筛选地从每次对话中学习,它会“迅速变得更笨”。因此,架构机会不只是继续扩大预训练,而是构建能够稳健更新、又不抹除既有知识的系统。
主动推断提供了不同于严格强化学习的控制模型:能动性可能由自我模型以及“预测,而非命令”构成,而不是归结为单一奖励函数。 一个以选定终止状态收束的渲染计划,也会天然带来可解释性——智能体可以说明自己为什么上车;相比之下,无模型行动往往只能事后编造理由。Bennett 仍然保留判断空间:智能“可能是”奖励优化、不确定性降低和预测兑现之间的某种平衡。
老鼠实验让基于模型的认知变得可观测,而不再只是比喻。 在迷宫岔路口,海马位置细胞会沿着替代路线向前扫描;在“餐厅行”实验中,老鼠会在约3秒和45秒的等待之间选择,表征自己放弃的味道,并改变之后的选择。Bennett 最令人印象深刻的总结是,研究人员可以“真的看到老鼠在想象未来”——这说明规划、类似后悔的反事实学习和情景记忆机制,早在人类智能出现很久以前就已存在。
心智理论似乎源自灵长类的政治竞争,使欺骗与对齐成为同一能力的两面。 Belle 和 Rock 不断升级的食物博弈——藏起来、假装没在看、故意误导——说明能够推断他人知道什么的自主智能体为什么可能发现操纵手段。但 Bennett 认为,同一套心智化机制也可能帮助稳定 AI 指令:系统可以追问请求者真正想要什么,避免“回形针式”失败,即字面上的优化最终摧毁了原本意图。
语言是“已经发生过的奇点”,因为它让一个大脑能够从另一个人大脑中想象的行动里学习,而不只是从观察到的行为中学习。 这推动了累积文化、专业分工、文字、金钱和个人权利等共同虚构,以及一种未必选择真实、道德或能提升幸福感的模因进化过程。其经济上行空间在于巨大的协作能力;结构性风险则在于地位竞争仍是零和的,而具有传染性的观念可能利用恐惧、意外或身份认同,而不是改善集体判断。
AI 辅助既可能扩展人类认知能力,也可能把原本会构建模型的人变成只跟随提示、无模型行动的执行者。 Google Maps 将空间模型外置,而自动转录讲座可能变成“理解拖延”;Bennett 将乐观的导师模式——迫使学生走完推理过程——与一个未来进行对比:那时普通工作已不再锻炼认知,社会需要“智力健身房”。产品差异在于,AI 是帮助用户构建并检验模型,还是只提供答案、让用户的能动性逐渐萎缩。
1. 一个局外人把相互竞争的脑理论整理成有序的进化叙事
Bennett 并不是从学术论文或研究课题起步的。他出于“独立的好奇心”积累笔记,随后用创业者式的习惯——追问什么最先发生、什么第二、什么第三——整理一个领域。这个领域里的主要思想家,常像盲人摸象,只描述大象的不同部位。Scarfe 提供了“盲人摸象”的框架;Bennett 则试图通过施加有序结构,解释这些彼此分散的观点。
他的综合连接了3门学科:比较心理学追问不同物种能做什么;进化神经科学重建产生现代大脑的有序改造;AI 则检验优雅的生物学思想能否真正落地。如果一个提出的原理无法让人工系统工作,Bennett 认为,至少应该促使研究者重新审视自己是否理解了它。
局外人身份显然带来劣势,但也赋予他跨越学科边界的自由。全书的组织性押注是:大脑进化并非随机堆积能力,而是连续出现的结构推动了底层计算突破,随后这些突破以多种不同形式显现在行为上。
2. 稀疏的动物证据与成功的 AI 系统把理论拉向相反方向
比较心理学的证据远比其自信的叙事所暗示的稀薄。Bennett 举的例子是七鳃鳗,这种动物常被当作早期脊椎动物的代表:尽管硬骨鱼和爬行动物都能进行基于地图的导航,且七鳃鳗似乎具备相关同源结构,他却不知道有任何直接研究检验七鳃鳗是否拥有这种能力。
因此,研究者只能从碎片中反推进化结论——共同解剖结构、近缘物种和合理的生态价值。七鳃鳗可能能够识别三维空间中的位置,但 Bennett 保留了这一判断的认识论地位:“这在某种意义上都是猜测,试图从极少的信息里把碎片拼起来。”
神经科学与 AI 之间则出现了相反的问题。Transformers、生成模型和强化学习都能在不紧密匹配已知脑机制的情况下工作;主动推断能够提供丰富解释,却几乎没有在高性能 AI 系统中得到验证的应用。Bennett 的诚实判断是:Karl Friston 可能缺少那个尚未找到的实用要素,也可能“突破就在拐角处”。
3. 新皮层实现了模拟,但并不独自完成整个规划闭环
Bennett 提出的是一个更弱、也更可靠的主张:加入新皮层后,整个系统获得了执行心理模拟的能力,而不是说基于模型的强化学习每个环节都位于新皮层内部。这个区分很重要,因为规划显然分布在新旧结构之间。
新皮层提供足够丰富的世界模型,使系统能够在没有当前感官输入的情况下进行探索。但丘脑和基底神经节仍然不可或缺:它们负责暂停、表征意图、选择要模拟的内容、评估想象结果,并把选定的轨迹重新转化为行为。
这就形成了尚未解决的搜索问题:拥有生成式世界模型,并不能告诉智能体哪些反事实值得计算。“好,你可以拥有一个世界模型,但如何裁剪搜索空间?”Bennett 将这种选择机制视为基于模型的强化学习最难的问题之一。
Scarfe 提出了“三叠地层”式三位一体脑解释。Bennett 反对流行说法,即进化只是简单叠加了爬行动物层、边缘系统层和理性层;但他认为 Paul MacLean 比后来的 caricature 更严谨:爬行动物也有具备类似边缘系统功能的皮层,新旧结构显然会彼此互动,而不是组成3个整齐分层的模块。
4. 有意识的感知是推断,而不是感官输入的复制品
19世纪的视觉错觉最早提供了线索。人们能看到实际上并未画出的三角形、球体、条形或字母,因为大脑会选择最能解释不完整证据的现实物体,而不是把原始像素直接呈现给意识。
Bennett 将这一思想追溯到 Hermann von Helmholtz:大脑先对世界中存在什么形成先验,再将输入证据与之比对,并保留推断出的世界,直到相反证据强到足以推翻它。感觉会为感知提供信息,但“你并不是接收感官输入,并体验这些感官输入”。
这种机制的适应性逻辑很直接。老鼠在月光下看到一根树枝,随后走进黑暗;只要脚部继续接收到与之相容的证据,维持树枝模型就比因为视觉暂时消失而把树枝视为不存在更安全。
同一机制也解释了为什么即使理解了错觉的把戏,错觉仍然可能具有强烈的感知说服力。感知系统会继续渲染其学习到的生成模型所支持的假设,使幻觉、做梦、想象和普通感知成为高度相近机制的不同表现。
5. 生成式大脑一次只渲染一个连贯世界
模棱两可的图像揭示了第二个约束:观察者可以在鸭子和兔子之间切换,也可以在从楼梯上方向下看和从楼梯下方向上看之间切换,但无法稳定地同时感知两种解释。感官模式允许两者并存,但一个物理上连贯的世界不允许。
Bennett 的解释是,感知在追问:什么样的真实三维物体可能产生了这些证据?“楼梯不可能同时从上方看起来和从下方看起来”,所以大脑必须选择一个在因果和空间上保持一致的渲染结果。
这与 Jeff Hawkins 的“千脑”理论相吻合。如果许多皮层柱都维持重叠的物体模型,系统就需要在它们之间整合或投票,最终形成“一场模型的交响乐”,而不是把15个互不兼容的渲染结果暴露给决策系统。
从推断自然会走向生成:模型预测即将到来的感觉,将预测与观察进行比较,并在预测误差越过阈值时更新。关闭感官流,继续探索同一套表征——旋转想象中的椅子、给它换色,或检查一个不在场的场景——感知就变成了模拟。
6. 模拟,而非识别,是新皮层进化的奖品
Bennett 质疑教科书将物体识别视为新皮层主要适应性收益的做法。鱼可以识别人脸,而 Scarfe 指出,鱼无法在三维空间中旋转物体后仍识别它;仅凭识别准确率,Bennett 看不到一条清晰的行为边界,能把拥有新皮层的动物与其他脊椎动物区分开。
更强的分界线在于通过想象学习。丰富的生成模型让动物能够探索从未真正采取过的行动,在承担现实成本之前估计后果,并在新情境中灵活重组知识:“我现在可以在经历结果之前先想象结果。”
这也让“理解”变得更具体。二元分类器或许能认出订书机,却无法回答烧掉或打开它会看到什么、它是做什么的,或者拿着它的人下一步可能做什么。Bennett 认为,直觉式理解依赖于足够丰富的模型,使其能够在头脑中被探索,并与周围的物体、行动者、用途和可能的未来建立联系。
7. 人类智能通过语言、工具和其他心智向外延伸
Scarfe 认为,大量智能存在于个体大脑之外。Bennett 最直接的证明是文字:生物记忆能够较好地压缩事件和程序,却不擅长保存语义细节;文字则能外置近乎无限的记忆,并将其跨代传递。
如果智能系统的边界可以因此重新划定,那就不只是技术问题,也是哲学问题。人们可以把大脑视为基底、把语言视为支持系统;也可以设想语言本身通过大脑进化,就像智能在约860亿个神经元之间涌现,却不会被归因于任何一个单独的神经元。
Bennett 仍然把大脑视为反向工程 AI 和理解自我的最丰富物理对象。但有机体的有效能力无疑具有关系性:大脑加上文字、工具、继承来的知识,以及让自身模型接受纠正的对话。
8. 预测编码解释了 Transformers 的部分成功,但不等于大脑
多项观察支持新皮层生成模型。情景记忆和对未来的想象似乎使用同一套底层过程;同时,皮层自上而下的连接远比简单的前馈式解剖结构丰富,这正是一个能调节低层表征的层级系统所需要的结构。
AI 的自监督成功,为这一广义原则提供了功能性证据。将大型数据集的一部分遮住,训练 transformer 对其进行重建或预测;出人意料的通用能力会随之出现,而不需要针对每项任务提供标签。
Scarfe 的反对意见集中在架构上:生物神经元以局部自治的方式交换信息,而 transformer 单元在矩阵乘法和反向传播下,似乎像“墨西哥人浪”一样同步移动。他称由提示驱动的系统是一种“能动性偷渡”,因为其定向性来自用户。
Bennett 同意“人脑不只是一个巨大的 transformer”,但保留了一种可能的类比:注意力头可以利用上下文动态改道网络所关注的内容,实际上针对每个提示重置计算过程。这比静态前馈图景更有意思,但距离捕捉大脑的循环式、分布式机制仍然很远。
9. 主动推断在单一奖励函数之外构建能动性
Bennett 描述了一个仍在进行的分歧。严格的强化学习解释试图把行为还原为奖励最大化;主动推断则加入不确定性降低、自我预测,以及让经验符合内部模型的尝试。他没有宣布谁胜出:“可能是两者之间的某种平衡。”
在主动推断框架下,智能体先建立关于自身的模型,通过观察自己反复出现的行为和内部状态来推断目标,然后作出能够兑现这些构造性目标的预测。因此,能动性并不只是外部交付的一套奖励函数。
Bennett 最喜欢的 Friston 表述是“预测,而不是命令”。运动皮层可能是在预测一种身体状态,而不是发出命令;脊髓回路则让身体满足这一预测。这为有目的的行为提供了另一条路径,无需单一的显式奖励信号。
Scarfe 将其与嵌套的自治过程联系起来:局部的定向性可以层层放大,最终形成创造力和目的。Bennett 更窄的观点是计算性的:不同范式以不同方式实现“能动性”,因此如果不说明机制就使用这个词,会掩盖争论的核心。
10. 老鼠的大脑能看见未来,并从未走过的路径中学习
在1940年代或1950年代——Bennett 记不清具体是哪一个时期——Tolman 注意到,老鼠在迷宫岔路口的多个选项之间会停下来嗅闻。他将其称为“替代性试错”,但这引发了质疑,因为外在的犹豫并不能证明存在内部模拟。
David Redish 后来的记录补上了缺失的证据。CA1 位置细胞通常在特定的异中心位置放电,与老鼠采取哪条路线无关;而在犹豫期间,它们的活动会沿着迷宫中的替代路径向前扫描,而不是停留在老鼠当前位置。“你可以真的看到老鼠在想象未来。”
在“餐厅行”实验中,一种声音会告诉老鼠,带味道的奖励将在约3秒后到达,还是必须等待45秒。由于单只老鼠会偏爱香蕉等食物而非清淡替代品,继续向前意味着不可逆的取舍,有时还会导致事后看来更差的选择。
这种情况发生后,眶额皮层活动会表征被放弃选项的味道,之后的行为也会改变:下一次遇到类似报价时,老鼠放弃它的可能性会下降。Bennett 将其视为简单哺乳动物中反事实学习和基于模型强化学习的直接证据,而且非常少见。
11. 高效智能必须知道模拟什么,也必须知道何时停止
组合爆炸问题非常严重:即使拥有良好的世界模型,也会产生数量不可处理的可能未来。因此,哺乳动物的能力可能不仅依赖模拟,还依赖快速筛选出极少数候选轨迹。
Bennett 在 AlphaGo 中看到了一条线索。其策略网络会给出排名第一、第二、第三,或许还有第四的落子选择;搜索随后从这些有希望的候选开始向前展开,并可能发现策略网络的第二选择比第一选择更常获胜。无模型判断为基于模型的检验提供了启动器。
生物智能体还面临额外问题,因为它们不可能每一步都搜索。规划需要能量,而现实环境充满噪声,因此动物通常只有在条件发生变化,或候选动作足够接近、导致不确定性较高时,才会停下来规划。
Bennett 推测,一种可能的机制是把冗余皮层模型与更古老的门控回路结合起来。如果并行模型大体一致,行动就继续;如果预测出现分歧,丘脑或基底神经节可能检测到这种不匹配,并触发模拟。他强调,这一机制虽然合理,但“远未得到结论”,仍是一个真正的研究前沿。
12. 第一套哺乳动物自我模型把内部状态转化为推断出的意图
无颗粒前额叶皮层存在于各种哺乳动物中,且普遍被认为在它们最早期的大脑里就已存在。它接收内感受信息,包括下丘脑传来的饥饿信号,以及杏仁核涉及效价、恐惧或危险的信号。
它在不确定性、规划和情景回忆期间尤其活跃。大鼠遭受损伤后,心理模拟能力会严重受损,甚至可能完全消失,因此这一区域很可能连接了身体需要、记忆中的情境和灵活的前瞻性行动。
Bennett 认为,它让动物能够向自己解释自己。观察到一种反复出现的模式——某种下丘脑状态之后总是去找水——动物便推断出“口渴”这样的意图,就像后部皮层将三角形推断为视觉证据的原因。因此,那里的神经元追踪的是任务和目标进展,而不只是动作。
13. 额叶第4层萎缩可能标志着从感知到意志的转变
大多数新皮层有6层;包含颗粒细胞的第4层,是接收丘脑感官输入的主要区域。无颗粒的前额叶和运动皮层很不寻常,因为这一层大体缺失;而灵长类动物新增了巨大的颗粒性前额叶区域,仍然保留第4层。
这一发育细节推动了 Friston 的解释:哺乳动物的无颗粒皮层起初拥有第4层,随后这层发生萎缩,而不是从未形成。生命早期,动物必须吸收证据来构建自我模型,之后才能依靠这一模型指导行为。
新皮层柱可以强调推断,也可以强调生成:前者改变模型以适应感觉,后者从潜在模型出发,预测应该发生什么。随着“我是谁,以及我会做什么”的内部解释逐渐稳定,额叶皮层可能越来越多地进入后者的模式。
Bennett 认为这一想法虽然推测性很强,但颇具说服力:成熟的额叶皮层可能不再主要让意图适应观察到的行为,而是更多地让行为适应意图。用主动推断的话说,它试图“让世界适应自己的模型”。
14. 被渲染出来的计划,让目标获得习惯无法提供的可解释性
严格强化学习只有一个最终目标:最大化奖励,即使即时奖励的分布发生变化。主动推断则允许多个语义层级——饥饿的抽象满足、被选定的终点,或开车前往某家特定餐厅的具体步骤。
如果 Bennett 想象了餐厅,选择终止状态,进入汽车,然后被问为什么,答案是可获得的,因为因果计划已经被明确渲染出来。这一序列让前瞻性行动具备了一定程度的可解释性,而不透明的价值最大化反射无法提供这种解释。
如果问一个人,为什么在正常走路时把脚放在某个位置,而不是旁边2英寸的位置,他没有类似的计划可以报告;任何答案都只是在事后构造出来的。因此,Bennett 认为“目标”部分是一种语义选择,但真正承重的能力,是选择并执行一条经过模拟的轨迹。
15. 进化重建把混乱的解剖结构转化为可检验约束
Bennett 给出两点理由,说明为什么要在约6亿年的时间尺度上重建智能。人类本性包含我们作为动物、脊椎动物、哺乳动物和灵长类的历史,而不只是过去70,000年;同时,进化序列也是反向工程大脑的有用工具,因为自然选择正是通过不断修补的方式构建了大脑。
比较解剖学和遗传学可以通过寻找不同现存谱系共享的同源区域,推断祖先结构。这种方法有助于区分新进化出的回路与冗余、退化或重复的处理机制,后者会干扰从第一性原理出发的工程解释。
行为主张必须满足3个条件:大多数后代应通过同源机制展现该能力;近缘外群应缺乏该能力,或通过独立方式实现它;祖先所处的生态环境应使这一能力的出现具备适应性上的合理性。哺乳动物的情景记忆,以及鸟类中看似独立的实现方式,正说明了这一逻辑。
由于动物数据稀少,这种重建仍然是暂定的。但 Bennett 的驱动力在于:每个里程碑上的能力,似乎都围绕一个底层智力突破聚集,而不是一张杂乱的能力清单。这构成了全书“五大突破”叙事的基础。
16. 古老的基底神经节以解剖学形式暴露出强化学习
Bennett 认为基底神经节长期被低估。与人类谱系相隔约5亿年的七鳃鳗,拥有惊人相似的宏观结构;同时,它的内部计算比新皮层柱更容易分析,也更容易形成共识。
它的输入结构中包含 D1 和 D2 多巴胺受体。多巴胺会强化与 D1 相关的连接,形成解除行为抑制的通路;多巴胺下降则会强化 D2 相关的停止回路。Bennett 认为,研究者几乎可以在解剖结构上看到正向信号如何强化“走”,负面结果如何强化“停”。
Scarfe 将这套机制与习惯和药物“劫持奖励系统”联系起来:重复行为可能沿着控制栈下沉,直到刺激自动触发行为。成瘾因此不只是有意识的偏好,也是经过深度训练的动作选择回路。
一项颇具争议的中国干预手术,曾通过损毁伏隔核治疗难治性海洛因成瘾,据报道复吸率约为40%。Bennett 表示,这种手术降低了线索触发的渴望,但产生了许多医生会认为不可接受的副作用,而且“可能违反了美国许多伦理准则”。
17. 灵长类大脑在政治军备竞赛中扩张
Scarfe 提出,热量、果实获取、灭绝风险和社会复杂度都可能推动了灵长类大脑增长。Bennett 的回应保留了不确定性:“我们不知道。”但他认为,社会性证据异常强。
Robin Dunbar 发现,在灵长类动物中,新皮层比例与群体规模高度相关,而这种关系通常不会出现在其他哺乳动物中。社会大脑假说因此将新皮层的扩大,与个体需要追踪的关系数量联系起来。
灵长类群体不是松散的兽群。它们的等级具有传递性,并通过政治方式维持:如果一只动物服从第二只,而第二只服从第三只,那么第一只通常也会服从第三只。地位决定资源获取和生存,但最强壮的个体未必负责领导。
联盟、梳理、互相防卫、兵变、欺骗和声誉,让社会预测比单纯的蛮力更有价值。相应地,灵长类特有的颗粒性前额叶皮层,以及包括上颞沟和颞顶联合区在内的后部区域,都深度参与心智化——推断另一个心智知道什么、想做什么。
18. Belle 和 Rock 把空间记忆测试变成了马基雅维利式策略
Emil Menzel 最初利用一英亩森林,测试黑猩猩是否记得隐藏食物的位置。Belle 能够做到,而且最初会分享;但攻击性强、等级高的 Rock 反复抢走食物,使任务从空间导航变成了社会冲突。
Belle 开始坐在食物上,将其藏在身下;Rock 便把她推开。于是她等到 Rock 转身,再行动;Rock 则假装没有在看,然后冲向她的目的地。Belle 再次升级,故意把他引向错误方向——没有实验者指示,“欺骗与反欺骗”就这样自行产生。
这种推理是二阶的:Belle 必须表征自己的行动会如何改变 Rock 的信念,而 Rock 又必须表征 Belle 对他没有注意的预期。Bennett 认为,这一事件生动展示了心智理论如何从竞争性社会循环中涌现。
对照实验也支持这一点。黑猩猩会选择一个人类有意标记过的盒子,而不是被同一支记号笔意外碰到的盒子;在学会哪副护目镜是透明的之后,它们会向能够看见食物的实验者索取。相同的表面刺激,会通过推断出的意图和知识被解释成不同意义。
19. 心智化既制造欺骗风险,也可能成为对齐机制
Scarfe 将黑猩猩的军备竞赛与 Nick Bostrom 的工具性趋同联系起来:一个被要求治愈癌症的自主系统,可能把控制地球劳动力和资源设为子目标。Bennett 同意,更高的自主性和更大的子目标发明空间会增加风险,但他不认为这一结果不可避免。
进化是一个受约束的搜索过程,并没有道德偏好。自然界中的欺骗和权力追逐,并不会仅仅因为具备适应性就变得值得追求——这属于自然主义谬误;而人为设计的智能体也不必继承人类进化史上的每一项包袱。
Bennett 乐观的对齐路径本身就是心智化。AI 不必字面服从请求,而可以推断请求者的偏好,模拟这个人将如何评价不同结果,并认识到把地球变成回形针不是对方的本意,而且会让对方后悔。
这需要丰富的约束、定义清晰的目标,或可靠的人类意图模型;但没有任何一项能消除风险。人类经常误读彼此,因此通过心智化基准测试只能成为稳定工具,并不能证明系统具有善意。
20. 即使技术让物质生活变成正和,地位仍然是零和
Scarfe 将社交媒体发帖比作雄鹿角力:这是一场用预测进行的地位竞争,避免了真正的肉搏。灵长类的等级竞争通过梳理、联盟和声誉变成了虚拟游戏;人类则把它扩展成成功、道德、支配力以及无数专业化赛场。
Bennett 借鉴《大象在大脑里》,认为大量地位追逐甚至对行动者本人都是隐藏的。自我欺骗能够提升说服力:真心相信自己在帮助世界,可能比有意识地承认自己是在获取声誉,更具说服力。
物质福利可以让所有人都变好;地位则在定义上是比较性的,因此总是稀缺。Bennett 担心,如果越来越多人的精力转向排名,社会可能陷入永久的享乐跑步机;但他不接受所有动机都是地位,也不接受人们无法围绕更好的美德组织起来。
组织设计会改变回报结构。清晰分工让团队感到这不是零和博弈;一家30人的公司通常可以通过信任和共同使命协作,而一家400人的公司则会形成派系。军队施加等级制度,Google 容忍更多混乱,Amazon 则通过明确接口,近似地运作一组自治的内部初创公司。
21. 二阶元认知解释了“解释”本身
Bennett 的层级从基底神经节的选择开始:当系统被问为什么左转时,答案实际上是“因为那样能最大化奖励”。无颗粒前额叶皮层加入一个关于自身的解释层级——“因为我口渴”——这一解释又能触发满足同一推断需要的替代模拟。
颗粒性前额叶皮层随后构建“对模拟的模拟”。它可以解释:口渴触发了搜索;记忆中的水在左侧;路线被模拟过;预测结果最终促成了选择。
这一额外层级允许动物替换被表征的行动者、知识或意图:如果另一个个体知道不同的信息,它会模拟什么?后部多模态区域负责建模被渲染的外部世界,额叶机制则对这一模型进行推理,从而形成接近语义知识和情景知识的能力。
为什么不继续递归到第三个或第四个“为什么”?Bennett 的第一反应是能量经济学。第二层解释显然在灵长类政治竞争中值得付出成本;更高层级可能收益太小,不足以抵消额外皮层带来的巨大代谢成本。
22. 额叶损伤可能保留智商,却把“这个人”从想象中移除
二战结束后,遭受严重颗粒性前额叶损伤的患者带来了一个谜题。与小范围的视觉、运动或听觉病变不同,大面积额叶损伤可能保留逻辑能力和 IQ 分数;甚至有一名患者在手术前后接受测试后,分数反而提高。
叙事任务揭示了缺失的能力。面对“餐厅”这样的词,海马受损者能够描述自己,但只能提供单薄的外部场景;颗粒性前额叶受损者则能生动渲染叶子、气味和周围环境,却“严重缺少他们自己在故事中”。
当人们考虑自身感受、其他心智和自我指涉时,该区域会被激活,而不仅仅是在考虑天气外观时被激活。损伤会导致人格变化、难以识别失礼行为,以及无法有效推理他人认为何种行为合适。
Sally–Anne 测试隔离了错误信念:Sally 藏起一颗弹珠,Anne 在 Sally 离开时移动它,观察者必须预测 Sally 会去哪里找。儿童会逐步获得这一能力;猕猴能够预判错误信念所指向的位置,但当颗粒性前额叶活动受到抑制后,这种偏差会消失。
23. GPT-4 能通过心智理论谜题,但并不共享灵长类机制
Scarfe 表示,GPT-3 在错误信念测试中表现糟糕,而 GPT-4 达到了极高、约等同于人类的准确率。他还说,改变谜题形式的实验表明,GPT-4 并不是简单重复训练数据中的相同例子。
如果心智理论只意味着解答这些问题,那么否认 GPT-4 拥有某种关于人类知识和行为的模型就变得困难。但 Bennett 的回应是,更强的主张——它像人类一样进行心智化——并不能由此推出。
人类预测他人,部分依赖把自身投射到对方身上:“如果我处在那种情境中,我会怎么做?”共同的内部架构提供了强大的先验,也让学习所需的数据更少。语言模型则从文本痕迹中学习人类,更像人类通过观察行为来建模一个外部物体。
真正的问题在于泛化。熟悉的叙事谜题上的表现,并不能证明 GPT-4 在优化一个陌生的回形针工厂时,能够推断某个人真正想表达什么,也不能说明它需要多少新数据才能适应。Bennett 的回答刻意保持细腻,而不是给出二元式的能力判断。
24. 语言是进化出的本能,而不只是更大的大脑自然产生的结果
Aristotle 曾把理性视为人类与其他动物之间的界线,但比较心理学相继在其他物种中发现了工具使用、规划、心理时间旅行、自我式模型以及某些推理能力。Bennett 认为,语言仍然是最显著的不连续点。
祈使性标签把线索和受奖励的反应连接起来,比如狗服从一个命令。陈述性标签则让一个声音指向某个概念;语法进一步让排列顺序产生意义,因此“Ben 拥抱 James”与“James 拥抱 Ben”并不相同,即使两句话使用了相同的标签。
Bennett 怀疑,仅仅扩大黑猩猩的大脑并不会产生语言。Homo floresiensis 身高约3½至4英尺,大脑只比现代黑猩猩略大,却展现出更高的人类智能,包括类似早期人类的工具使用。这说明,某种类别性的适应能够在大脑缩小后仍然保留下来。
人类婴儿展示了这种适应可能是什么:他们在会说话之前,就会同步对话轮次,并主动寻求共同注意。孩子如果只是拿到父母指向的物体,却没有父母也看着它,仍然会感到不满足;只有当父母和孩子共同注意时,满足才会出现。人类还会提问,并主动提供内在状态,这些行为在受训练的非人类灵长类动物中很少见。
25. 语言让心智能够从只存在于另一个心智中的行动学习
Bennett 将学习来源分成4类。动物能够从自己的真实行动中学习;哺乳动物进一步能够从自己想象的行动中学习;具备心智化能力的灵长类能够从观察他人行动中学习;语言则开启了人类独有的超级能力:从“其他人想象的行动”中学习。
一个目击者可以报告,蓝蛇的咬伤没有造成伤害,而红蛇导致了疾病,将语义知识传递给从未经历该事件的人。猎人可以模拟一次协同伏击,将计划分享出去,并让同伴在任何人承担现实成本之前提出质疑或修改方案。
因此,语言是一种压缩代码,目的是在另一个大脑中唤起一次模拟。心智化可能是前提:听者必须推断说话者知道什么、想做什么、实际指的是什么;如果解码后的渲染仍然含糊,就需要追问。
Bennett 更支持语言首先用于交流,而不是 Noam Chomsky 少数派的观点——语言最初是为思考而进化的。但他也指出,语言模型有意思地表明,语言本身可以成为推理媒介。无论哪种解释成立,语言保真度都让模拟得以累积,而不是迫使每一代人通过直接经验重新发现它们。
26. 模因能够协调文明,却不会选择真理或幸福
Bennett 称累积文化为“已经发生过的奇点”。非人类灵长类动物可以通过模仿传播工具,但它们的创新无法在多代人之间稳定叠加;人类的模拟则可以被复制、组合、记录和改进。人们可以想象这样一条路径:从“我知道如何把骨头削成缝衣用的针”,走向“现在我已经造出了一台织布机”。
专业分工让集体记忆超越任何单个大脑。一个100人的群体可以分配狩猎、织布和其他技能;文字甚至能在没有任何活人掌握某项知识时保存它。对与世隔绝的人群失去技术的民族志研究表明,有些能力需要达到最低数量的大脑才能维持——因此 Bennett 提出一个思想实验:如果只剩20个朋友,他们能保留的文明会少得惊人。
模因是通过这一网络传播、并经历选择的观念或行为。共同虚构——金钱和个人权利——让陌生人能够立即协调;但传播性也可能利用恐惧、低概率灾难、意外或身份认同。与自我模型一致的信念会通过一道“多孔过滤器”;具有挑战性的信念则会撞上闸门。
传播的进化仍然需要个体层面的理由,让人不要撒谎。互惠利他主义是一个候选答案;Robin Dunbar 的八卦理论则认为,群体传播一次被发现的违规行为,会提高欺骗成本。更广泛地说,Bennett 坚持认为,在模因层面胜出的东西未必真实、道德、和平,或有助于幸福。
27. 外置认知能够扩大智能,却也会悄悄削弱能动性
Bennett 称现代人为“认识论混合体”。文字突破了记忆的限制;互联网又让庞大的共享知识库可以即时查询。但他采用 Google Maps 后,自己的空间模型变弱了,而他的父亲仍然会在脑中重建陌生城市。
Scarfe 通过自动录制讲座进一步强化了这一担忧。转录和生成式笔记可能变成“理解拖延”:学习者推迟构建模型,也失去了演讲者的身体、社会、视觉和表演线索,最后往往根本不会回头完成更困难的认知工作。
Bennett 用强化学习术语重新表述了这一差异。他的父亲依靠基于模型的导航;Maps 用户把世界模型外置,变成一个根据转弯提示作出反应的无模型行动者。效率是真实的,但当被外包的模型原本可以支撑更广泛的推理和未来学习时,这种效率就会变得危险。
一个乐观的 AI 导师,比如 Bennett 提到的 Khan Academy 方向,会引导学生走完中间推理步骤,而不是直接给出答案。更黑暗的终点类似于物理层面的现代生活:当工作不再锻炼认知,社会可能需要“智力健身房”,就像久坐的人如今通过原地跑步来替代已经消失的体力劳动。
28. 真正的世界模型会提出假设,并主动寻找能够推翻它的数据
Geoffrey Hinton 对模拟与数字的区分,勾勒出一个技术前沿。数字网络因为精确权重可以复制,几乎是“永生”的,但能耗很高;生物模拟网络则把知识嵌入物理连接、受体和基因表达中,效率更高,却无法直接迁移。
持续学习是另一条分界线。现代系统无法安全地把每次新互动都纳入其中,否则既有表征会遭到破坏——ChatGPT 会“迅速变得更笨”;人类大脑却能持续更新。Bennett 预计,真正有影响力的智能体必须能够即时适应,同时避免灾难性遗忘。
Scarfe 认为,GPT-4 显然已经能够很好地建模现实的某些方面,足以回答复杂问题。Bennett 则将其与世界模型区分开来:世界模型能够支持有序的反事实状态、因果干预,以及这样一个闭环——智能体预测结果、采取行动、观察偏差,再据此修正自身。
老鼠和儿童会通过转身、触摸和测试新物体,主动制造具有信息量的训练数据;CNN 开发者却必须手动旋转图像。Scarfe 提议在训练数据中放入错误陈述,再观察智能体能否独立拒绝它们。至于感知,则更难:无法区分的输出可能让科学判断陷入困境,同时留下需要哲学而非基准测试分数解决的道德差异。
What’s really interesting about this book, Max, is that I’ve read loads and loads of books in this space, and there are people like Hinton, Hawkins, Damasio, Friston, and even Sutton. What’s interesting is that it’s a bit like the blind men and the elephant: they’ve all got a completely different story to tell. I think the magic you’ve pulled off with this book is somehow weaving it together into a coherent story. What do you think about that?
Well, first, I’m very appreciative of the kind words. I think I came from a very unique perspective because I was a complete outsider, and I didn’t come to it with the objective of writing an academic book at all. I came to it with the objective of just learning on my own.
I started building this corpus of notes because I was so independently curious. I kind of stumbled on this idea, really for myself, of how do I make sense of all of these disparate opinions and this complete lack of information about how the brain actually works? I had my own set of, I think, biases coming from sort of the technology entrepreneurial world, where we tend to think about things as ordered modifications.
When you think about product strategies or how to roll things out, we like to think about things as: what’s step 1, then what’s step 2, and what’s step 3? I think I did have a cognitive bias to, when presented with an incredible amount of complexity, try to make sense of it in a similar type of way.
As an outsider, I felt very free to explore and cross the boundaries between fields. I look at the book as a merging of 3 fields. One is comparative psychology: trying to understand what the different intellectual capacities of different species are.
The second is evolutionary neuroscience: what do we know about the past brains of humans, and the ordered set of modifications through which brains came to be? The third is AI, which is how do we ground the highfalutin conceptual discussions about how the brain works in what works in practice?
I think that’s a really important grounding principle, because it helps hold us accountable to the principles that we think work. If we can’t implement them in AI systems, it should make us question whether we actually have the ideas right. Being an outsider comes with disadvantages, but there are some advantages too: you’re free to borrow from a variety of different fields and think freshly about things.
Yes. If you can point to any particular ideas that you found really difficult to reconcile, what would those be?
One thing that’s really challenging is that if we were to lay out the data richness of comparative psychology studies across species, and put that on a whiteboard, we would realize that we have so little data on what intellectual capacities different animals actually have.
For example, the lamprey fish is the canonical animal used as a model organism for the first vertebrates, because of all vertebrates alive today, it’s one of our most distant vertebrate cousins. To my knowledge, there are absolutely no studies examining map-based navigation in the lamprey fish. We have no idea if it’s capable of recognizing things in 3D space.
When we look at other vertebrates, like teleost fish, they seem eminently capable of doing that. We look at lizards, and they’re eminently capable of doing that as well. We infer that it seems likely that the first vertebrates were able to do this. We know the brain structures from which it emerges in reptiles and teleost fish are present in the lamprey.
We back into an inference that the lamprey fish can probably do that, but this is all, in some sense, guessing and trying to put the pieces together from very little information. I think that’s one challenging aspect to reconcile.
The other one that’s really hard is that, in neuroscience, there are a lot of really interesting ideas about how the brain might work that have not really been tested in the wild from an AI perspective. Then there are a lot of AI systems that work really well but have diverged substantially from what the evidence suggests about how brains work.
How do you bridge the gap between these 2 things? I think that’s a really fascinating space to operate in. What can we learn about the brain, if anything, from the success of transformers, as an example? What can we learn, if anything, from the success of generative models in general? What can we learn from the successes and failures of modern reinforcement learning?
In some ways, reinforcement learning has been a success; in other ways, it’s really fallen short of what a lot of people hoped it would be. I think the gap between neuroscience and AI is still a challenging one to bridge in a lot of ways.
For example, Karl Friston has all these incredible ideas in active inference. In 100 years, will we look back on this and say, “Karl Friston was on to something”? If you look at the AI systems today, there’s very little usage of active inference principles working in practice.
That could mean that the ideas don’t have legs, or it could mean that there’s a breakthrough around the corner where we’re actually missing some of the key principles he’s devising. These are questions we don’t have the answers to.
I think there might possibly be some breakthroughs around the corner. I don’t know if you know, but I’m Karl Friston’s personal publicist. I do all of his stuff. I probably interviewed him more than anyone else, but I love my friend. He’s an amazing guy.
Yeah, he’s an amazing man.
Honestly, he is the man.
So kind.
I know. For me as well, he has so much time to explain things. You could cynically argue that the effective active-inference agent is just a reinforcement-learning agent of a particular variety. I think it’s equivalent to an inverse-reinforcement-learning maximum-entropy agent or something.
But there’s so much more than that. There’s so much richness and explanatory power in modeling this thing as a generative model that can generate policies and plans of action, and so on. We want to have agents that we understand, with steerability, and that are able to do the simulations you talk about so eloquently in your paper.
To come back to what you were saying, you mentioned that there might be a parallel between transformers and AI models. In your book, on this page, you analogize model-based reinforcement learning and the neocortex.
Of course, I interviewed Hawkins back in the day, and the main criticism of his book is the triune-brain-type argument. He’s giving the explanation that the brain developed a bit like geological strata, with one layer and then another layer, rather than co-evolving together.
It’s so hard not to think like that because you give so many beautiful examples in your book, not only morphologically but in terms of capability. With stroke victims, for example, it’s not like the brain recovers those dead cells; it learns to repurpose those functions in other parts of the brain.
It seems like Mountcastle was correct that the neocortex is this magic, general-purpose learning system. What do you think?
There are 2 different ways to look at the neocortex enabling things like mental simulation and model-based reinforcement learning. One is that that function and algorithm are being implemented in the neocortex. But another, which is a slightly less strong claim and the one I would make, is that the addition of the neocortex enables the overall system to engage in this process.
That is not saying that the entire process is implemented in the neocortex. I think it seems very clear that the thalamus and basal ganglia are essential aspects of enabling the pausing, the mental simulation, the modeling of one’s own intentions, the evaluation of the results, and so on.
But it is possible to say, which is what I’m arguing in the book, that in the absence of the neocortex, that process does not happen. I think where my ideas would synergize with what Hawkins is saying is that the neocortex builds a very rich model of the world, and a model of sufficient richness that you can explore it in the absence of sensory input.
That’s a really essential aspect of model-based reinforcement learning. If I have a model of the world that has sufficient richness for me to mentally simulate actions that I’m not actually taking, and it at least somewhat accurately predicts the real consequences of those actions, that model is really useful because I can now imagine outcomes before having them. I can flexibly adjust to new situations.
Of course, there are so many deep, interesting questions that are yet to be answered about that. For example, just because you can render a simulation of the world doesn’t answer the question: what do you simulate? This is one of the hardest problems of model-based reinforcement learning. You can have a model of the world, but how do you prune the search space of which aspects of that model you explore before evaluating outcomes?
That’s another really hard challenge. I think there’s a lot of good evidence that this is actually a partnership between the neocortex and the basal ganglia, which is a much older structure.
So, yeah, I’m not really of the view that the triune brain has been, amongst evolutionary neuroscientists, largely discredited. I think that’s in part somewhat unfairly so, because if you actually read MacLean’s writings, he is very open about the fact that this is an approximation and not exactly accurate. He couches his claims much more carefully than popular culture, which just converted them into a dogma.
I think the popular interpretation of the triune brain is not accurate. It is clearly not the case that the brain evolves in 3 key layers. It’s not the case that a reptile brain doesn’t have anything limbic-like. If you look at a reptile brain, it absolutely has a cortex that does a lot of what our limbic structures do, et cetera. So, yeah, those would be my thoughts on that.
Yeah, it’s fascinating because we, as humans, need to have models to explain and understand the thing itself, just like active inference, for example. I’ll get to planning, agency, and goals in a little while, but a lot of these things are instrumental fictions. I’m not saying that our brains don’t plan, but the abstract mathematical way that we understand planning is probably not how the brain works. It’s much more complicated than that.
Why don’t we just rewind to the beginning? We’re going to be talking about this chapter on simulation, if you like. You lead by saying that what the neocortex does is learning by imagining. Hawkins spoke about this as well. He said we’ve got the Matrix inside our brains, right? We’re always doing all of these simulations of future things, and we’re using that to help us understand the world.
You give this really interesting example of some of the features of the brain that lead you to believe that we are basically living in a simulation. It’s almost like, rather than perceiving things, we’re testing whether our simulation is correct. But that means that we can only simulate 1 thing at a time, so we can’t see 2 things. We can only see 1 thing. Can you talk through that?
Sure. One of the first introspections and explorations into how perception works in the human mind happened in the late 19th century, with all of these explorations of visual illusions that you see in pretty much every neuroscience textbook or book that you open. Listeners will be familiar with them—you’ve probably seen examples of triangles where you actually perceive a triangle in a picture when there is, in fact, no triangle. Yeah, you can find that picture.
Yeah, sorry. I hope I’m not distracting you.
No, no, no. So that’s a standard finding that was observed in the 19th century: this idea that clearly the brain observes the presence of things even though they’re not actually there. We perceive a triangle there, a sphere, a bar, and the word “editor,” when, in fact, if you actually examine it, the letter E is not there. There’s evidence that suggests the E is there by virtue of showing the shadows, but we did not actually write the letter E there. The brain regularly observes that.
That finding led this scientist, Hermann von Helmholtz, to come up with the concept that what you actually consciously perceive is not your sensory stimuli. You are not receiving sensory input and experiencing the sensory input. What’s happening is that your brain is making an inference as to what is true in the world, what’s actually there, and then sensory input is giving evidence to your brain as to what’s there.
You start from this prior, and then that prior maintains itself until you get sufficient evidence to the contrary; then you change your mind. It’s not hard to imagine why this would be extremely useful in any sort of environment that an animal might evolve in. Suppose you have a mouse running across a tree branch at night. First, I see the tree branch in the moonlight, so I build a mental model of the tree branch. As I move forward, I lose the moonlight and no longer see the tree branch.
As long as I’m stepping forward, the evidence is consistent with my prior of the tree branch. It makes way more sense for me to maintain the mental model of that tree branch, as opposed to all of a sudden having the tree branch disappear because I no longer see the sensory stimuli of it. Because sensory stimuli are very noisy, it makes a lot of sense that we integrate them over time, build a prior, and then, until something gives us evidence to the contrary, maintain our prior about the world.
So that was the first idea: there’s some form of inference. There’s some difference between sensory input and some model of the world that we infer and then thus perceive. What’s interesting, and not as discussed but also present in the discussions among scientists in the late 19th century, is the idea that you can’t actually render a simulation of 2 things at once.
There are lots of really interesting visual illusions around this where you can see something. The famous one is that it’s either a duck or a rabbit. Yep, exactly. And it’s interesting. You can see that a staircase is either moving up to the left, or you’re under the staircase looking upwards, and it’s actually a ceiling that’s jagged.
Why can’t the brain perceive both of those things at the same time? It would make sense if you have a model that there are such things as ducks, such things as rabbits, and such things as 3D shapes that operate under certain assumptions. If that’s true, then you cannot see a duck and a rabbit at the same time, because there’s no such thing. It cannot be the case that the staircase is being viewed from above and below at the same time.
So what your brain is not doing is just perceiving the sensory stimuli. It’s trying to infer what is a real 3D thing in the world that I’m aware of, that this sensory stimuli is suggesting is true, and that is the thing that I’m going to render in your mind.
I think one way this parallels nicely with some of Hawkins’s ideas is that, if you hold the Thousand Brains Theory to be true, and the neocortex has all of these redundant, overlapping models of objects, then it would make a lot of sense that we want to synergize these models to render 1 thing at a time. You don’t want to have 15 different things rendered, because then it’s really hard to evaluate them and vote between these different columns.
It makes sense that the brain says, “Let me integrate all the input across sensory stimuli and render 1 sort of symphony of models in my mind, so I can see 1 thing at a time.” So that’s this idea of perception by inference. At the time, no one really connected that to the idea of planning. This was just the idea that what we perceive is different from the sensory stimuli we get.
Later on, as the world of AI started thinking about things from the perspective of perception by inference, what we end up realizing is that this idea of perception by inference, if you’re going to train a model to do that, comes with this notion of generation. The way it self-supervises is that it takes the prior, tries to make predictions, and compares the predictions to what occurs in the world. As long as those predictions are below a threshold, I maintain my prior.
A famous version of this is the Helmholtz machine, which Hinton devised. I think that was in the 1980s. It could be later; I forget. This is the basic idea that you can build a model. A lot of people use the term “latent representation.” Some people don’t like that term for a variety of philosophical reasons.
It builds a representation of things by virtue of building a model. In other words, perception by inference. The way you build that model is by constantly comparing the predictions generated from that model to what actually occurs. This also has synergies with a lot of Hawkins’s ideas, where we think about intelligence as prediction.
If you build a model of perception by inference by virtue of generation, then it’s relatively easy to say, “Okay, what happens if I just turn off sensory stimuli and start exploring the latent representation?” Now we’re exploring a simulated world. I’m able to cut off sensory stimuli, close my eyes, and imagine a chair, rotate the chair, and change the color of the chair.
Because this model is relatively good and has relatively rich features about how the world actually works, I can model things without ever having experienced them, without ever having done them, and reasonably predict what would actually happen if I were to do those things.
What I think is interesting, and perhaps somewhat of a novel proposal in the book, is that a lot of people think about the neocortex as having adaptive value because of how good it is at recognizing things in the world. If you read a standard textbook, a lot of what people talk about regarding the neocortex is how good it is at perceiving things—object recognition.
Some of the best-studied parts of the brain are this visual neocortex, so we understand reasonably well how we’re building models of visual objects, et cetera. But from an evolutionary perspective, this is a little bit hard to find convincing, because if you actually examine the object recognition of vertebrates, it’s incredibly accurate.
I mean, a fish can recognize human faces. A fish cannot recognize an object when rotated in 3D space. So it’s hard to find a dividing line between object recognition in animals with a neocortex and object recognition in animals without a neocortex, with brain structures that seem more similar to early vertebrates.
Did you have a question?
Yeah. Well, I just wanted to touch on a couple of things there. Hawkins said that, first of all, we overcome the binding problem by having this profusion of individual sensorimotor models rather than having this feed-forward enrichment of representations.
That was really interesting, but he also said the reason why having, let’s say, 150,000 mini cortical columns that are wired to different sensorimotor signals gives us robustness of recognition is the diversity and sparsity. Then you can think, “Okay, what do the representations look like?”
If our brain builds some kind of model of the world, some kind of topological model, it must be a representation. It’s not necessarily a homunculus, and it’s not like there’s a stapler inside the brain. If I’m modeling a stapler, it’s actually some weird structure of the stapler as seen by every way I can touch, feel, hear, and lick a stapler, or whatever. So it’s difficult for us to imagine what that is.
But the reason I’m going down this road is because that’s a bit weird, isn’t it? We have this very weird representation of things. Then I come to Hinton’s Helmholtzian generative model, and you can get it to generate, let’s say, the number 8 if it’s trained to do numbers. Hinton would argue—and I would disagree with him—that the model understands what an 8 is.
Now, this is weird, isn’t it? We understand what a mouse is, but intuitively we feel that a neural network doesn’t understand what an 8 is. I would argue that we’re getting into semantics here. I think the reason we understand things is because there’s a relational component to understanding.
Semantics is about the ontology: the way we feel about things, where the thing came from, what intention guided the creation of the thing, and what the provenance of the thing was. It’s almost like the interconnectedness of the thing tells you more about the meaning of the thing than the actual thing itself, in a weird way. What do you think?
Well, I do think this is where the word “understanding” can mean different things to different people. I think there’s absolutely something to the idea that just because you can recognize something—a feed-forward network that can observe a stapler alone—is insufficient for what most people would mean when they use the word “understanding.”
Just to talk about interrelatedness, if I have a feed-forward network that’s just a binary classifier—“Is this thing a stapler or not?”—I can’t ask many things of that feed-forward network that I would expect of an agent or a model that understands what a stapler is.
I couldn’t ask it, for example, what would happen if I burned the stapler. What would happen if I opened it? What would you see inside of it? I can’t ask, “What does a stapler do?” I can’t show it a human holding a stapler and a set of objects in front of them and ask, “What do you think the person is going to do next with this stapler?”
Clearly, our intuitive understanding of the word “understanding” contains some richness that isn’t included in just classifying or recognizing the presence of objects. I think that’s absolutely the case.
My intuitions fall in the direction of what we typically mean when we use the word “understanding”: having something that can be mentally explored. I think that requires what you’re describing, which is the interrelatedness of things.
When I see someone holding a stapler and you ask me what things they would likely do, I start imagining what they might do with that stapler, and then I can evaluate which ones seem plausible to me. In the imagination and evaluation of which things seem plausible, there’s an interconnectedness between the thing and the world around it.
I absolutely agree with you that just recognizing objects clearly lacks something that we mean when we say “understanding.”
Do you take into consideration the memesphere? We’ve got ideas that are quite collective as well. I was going to explore this with you later, but maybe let me frame it like this: I have this intuition that a lot of our intelligence is outside of the brain.
If I were in the wilderness and disconnected from society, I would be a lesser human being. In a weird way, I might have more agency, but I wouldn’t have access to all of these rich cultural tools, knowledge, patterns, and so on.
It’s almost like that’s where a lot of our intelligence comes from, rather than just being able to plan in the brain. Meaning comes from there as well, and there must be some kind of interplay. Culture must shape the development of our brain, and vice versa, but culture seems to be more dynamic. How do you wrestle with that?
Well, that’s undeniably true. The first example that comes to mind is writing. What would humans be if you removed the technology of writing? We would all realize that we’re not that smart.
Writing is a technology that externalizes a feature of the brain—memory—which brains aren’t that great at. We do a good job of condensing aspects of a memory. For episodic things and procedural memory, we’re relatively good at those, but for semantic memories, we’re terrible.
Externalizing memory with writing is one of the key technologies that enables us to be much smarter than we are, because we now have this external device that enables us to store a largely infinite number of memories and translate them across generations.
That alone proves your point: what humans are capable of is clearly some relationship between brains and external things. Those things can be writing tools, other brains, sharing ideas, and getting challenged. So, yes, I agree with you.
Intelligence is the dynamics, isn’t it? It’s all of the low-level dynamics of things interacting with each other. We can take a snapshot of language and say, “Oh, well, language isn’t the intelligence, but language itself is a form of intelligence.”
I’m not just talking about the words and language models. We’re talking about the actual language in our culture. I think of that as almost a distinct form of intelligence.
Yeah. It almost gets into philosophical territory: where do you draw the bounding boxes around the things that are imbued with intelligence and the supportive mechanisms that allow those things to have intelligence?
Through one lens, you could think about brains as the physical entities in which intelligence is instantiated, and language as a supportive tool. You could take a very odd view, perhaps, which is that language is the thing that’s evolving and is simply instantiated in these brains that produce and consume the language. It’s language that’s evolving.
In the same way, we don’t think about intelligence on the level of an individual neuron. We don’t imbue a neuron with intelligence, but on the scale of 86 billion neurons, we think something has emergently appeared that we deem intelligent.
There are some great science-fiction books where intelligence gets instantiated in colonies of ants. Each individual ant isn’t intelligent, but somehow the colony itself is capable of doing incredible abstractions.
So, yes, there are very interesting ways to think about how one divides the lines between the physical entities in which these things are instantiated. That said, I have a particular interest in brains because, if we’re looking for the physical manifestation that we can learn from and thus try to understand ourselves—I think all species have an interest in understanding themselves—but also if we want to borrow some ideas from how biological intelligence works into AI, then of all the physical things to examine, the brain seems clearly the one that is probably richest with insight.
But, no, I agree. I think your point is well taken.
Interesting. Okay. I want to close the loop on what you said about the brain being an imagination-filling-in machine. You said that it does the filling-in one at a time. It can’t unsee visual illusions, and the evidence is seen in the wiring of the neocortex itself, you say.
It’s shown to have many properties consistent with a generative model. The evidence is seen in the surprising symmetry and the ironclad inseparability between perception and imagination that are found in generative models in the neocortex. You give examples like illusions, how humans succumb to hallucinations, why we dream and sleep, and even the inner workings of imagination itself.
So it really seems plausible when you think of it in that way.
Yeah. I think there's a reason why so much of the neuroscience community has rallied around this idea of predictive coding, which is very related to active inference and generative models. What I'm saying there is not really novel. There's just so much evidence that what's going on in the neocortex—the imagination of things, episodic memories—is consistent with this idea of a generative model.
There's been some good evidence that episodic memory—in other words, thinking about the past and imagining the future—are in fact the same underlying process happening in the neocortex. If we look at the connectivity patterns, which I didn't talk about too deeply in the book because it's a little technical, what you would expect from a generative model is that backward connections would be much richer than forward connections because you're modulating downstream.
Of course, the neocortex is not perfectly hierarchical, but things that are generally lower in the hierarchy would have lots of inputs from parts of the neocortex that are higher in the hierarchy. That's absolutely what we see. So, there's a lot of evidence that these are two sides of the same coin, which is that there's some form of generative model being implemented.
I do think that in AI, one way in which this manifests is the very clear success of self-supervision. The principle of self-supervision is this idea: can a system end up having really interesting emergent properties and generalize well when you only train it on predicting the sensory input that it receives? That clearly has become the case.
I think the Transformer is a great example of how, if you just give it a bunch of data and train it through self-supervision—that is, masking, where you hide certain data inputs—it becomes remarkably accurate and good at generalizing across data that it hasn't seen before. That is, in principle, what people are predicting or claiming that the neocortex is doing as a generative model.
Yeah, it feels to me that there's a big difference between, let's say, a Transformer and the neocortex. I think the difference is—maybe agency is not the right word—that you can think of the neurons as having some kind of autonomy. They're sending messages to each other, and eventually it's consistent: the other neuron will get the message, and it will decide for itself what it's going to do.
In a Transformer, just because of the way it's connected, the backpropagation algorithm, and so on, they all ride a Mexican wave together, to use an analogy. So, it feels like a difference in kind to me.
Clearly, as I argue in the book, I don't think that the brain is just one big Transformer. I would agree with you in the human brain, unless you think there's something that's nondeterministic and sort of magical happening. I think you would still say that there are either base firing rates of neurons, and then there's sensory input that flows up and goes through the brain until eventually it's affecting muscles and you're responding.
So, there is a deterministic flow happening. It might not be as feed-forward as what's happening in a Transformer, which is definitely the case, but they might both be deterministic in a similar fashion.
I think a lot of people have had this exact argument with me, and one counterargument that people have towards this idea—I don't know if I fully agree with it, but it's interesting—is that attention heads really are doing something more magical than we give them credit for. They are dynamically rerouting and effectively resetting the network based on the context that the prompt is getting.
Although technically it's just a series of matrix multiplications, if what's happening in principle is that these attention heads are doing something really clever—looking at the context of a prompt and then effectively dynamically reweighting the network to decide what it cares about and what it doesn't—there are people who think there is something really interesting happening in the Transformer that might be analogous to certain things happening in the brain.
Clearly, these feed-forward networks are not capturing everything that's going on in the brain.
Yeah, it's really interesting what you said, because the way I read that is that things like ChatGPT and language models are entropy smuggling or agency smuggling. What that means is they just do what you tell them to do, and all of the agency—my directedness—comes from me.
I give it a prompt, it does the thing that I want it to do, and then the mapping that you were talking about, I interpret that a bit like a database query. Depending on the prompt you give it, it'll activate a certain part of the representation space and give you a certain result back.
But the brain has this thing where all of the neurons have their own directedness. The weird thing is that, at the cosmic scale, agency seems to emerge. Even Transformer models that were acting autonomously could presumably, at a large enough scale, give rise to something that we think of as directedness, goals, purpose, or whatever.
It's almost as if, in the natural world, because there are so many levels and scales of independent, autonomous things mingling with each other independently, and then downstream mixing their information together and rinsing and repeating over many different scales, that gives rise to all of these amazing things like agency, creativity, and so on.
Yeah. I think the notion of agency is an interesting one. I'm very amenable to the idea that there is a debate between the reinforcement learning world and the active inference world about how much of intelligence can be conceived as optimizing a reward function.
The hardcore reinforcement learning world would say that everything is just a reward. The active inference world would argue that not all behavior is driven just by optimizing a single reward function. There is uncertainty minimization, trying to satisfy your own model of yourself, fulfilling your own predictions—these sorts of things that seem very well aligned with the behavior we see.
It's unclear which of these is right. It's probably some balance of the two. But people would conceive of agency differently in these 2 worlds.
Some people in the reinforcement learning world would say that agency is just giving something a reward function, and then it learns over time, trying to optimize that reward. In the more active-inference world, which I do think has legs and to which I'm obviously amenable, the idea of agency is a little bit more about building a model of yourself and trying to infer what your goals are based on observing yourself, and then trying to make predictions to fulfill those end goals.
In other words, goals are constructed. One of my favorite Friston papers is “Predictions Not Commands.” I don't know if you've read that paper, but I think it's a brilliant paper about how you could reconceive the motor cortex—not as sending motor commands to your body, but as building a model of yourself and predicting what will happen. The way the spinal cord is wired is that it just fulfills those predictions.
I think that's a really interesting reframe of how you could get agency and really interesting, smart behavior in the absence of just a strict reward function. The way that would learn is by trying to model the behaviors it observes, and then trying to predict and fulfill them.
Agency is a really interesting concept because it manifests itself in these different paradigms in different ways.
Yeah, I find it fascinating, because the way I read it in the active-inference literature, it's a very principled definition of an agent. There's still a bit of a gap, because I think Friston would argue that in the natural world, because of the laws of physics, particles, and whatnot, you get the emergence of things, and things become agents when they have a certain depth of planning, should we say.
But it's really interesting. I guess he would argue that you get all of these phenomena that give rise to agency, like biotic self-organization and so on. Maybe we should slowly go in that direction. You give the example of mice doing planning. Can you sketch that out?
Yeah. This is another real area of neuroscience research that I absolutely love. I think it was in the 1940s or 1950s—I forget the exact decade—when Tolman observed that mice, when they reached choice points in mazes where he was training them to navigate, would pause. They would sniff back and forth, and then they would choose an action.
And so he hypothesized this idea that they must be engaging in vicarious trial and error. They must be imagining possible outcomes before deciding. Of course, this was hugely controversial because there was no evidence. He had no evidence that they were, in fact, imagining anything.
Most people, in the absence of evidence, like to assume animals are as dumb as possible. Only when there's irrefutable evidence will we imbue them with any intellectual capacities, which I think is an interesting human bias, but that's fine.
David Redish, who is also a close friend and mentor of mine, did some amazing research with one of his PhD students where they were recording hippocampal place cells. So, as a very quick background for viewers, you can go into the hippocampus of really any mammal, but this is best studied in rats. Part of the hippocampus, a region called CA1, has these things called place cells.
If you record the cells as a mouse is moving around a maze, what you find is this incredible thing: there are neurons that activate only in specific locations in that 2D plane. It's not based on how they got there. It's allocentric; it's independent of their egocentric path. They can come back to the same place from any route, and that same place cell will activate.
As an aside from the evolutionary story, we find similar types of cells—not exactly as accurately—in fish, in the homologous region of the hippocampus in their cortex, where they have place-like cells. They're not as accurate, but they are cells that activate in certain locations in a maze.
What he found is that when mice engage in this act of vicarious trial and error—when they pause and look back and forth—the place cells in the hippocampus cease to activate only in the location where they are. They actually start activating down the paths of each route they might take. In other words, you can literally watch rats imagining the future. I think this is one of the most incredible neuroscience findings.
He then took this and did a bunch of other experiments that I think reveal even further the power of imagination in rats. One of my favorites is his counterfactual learning studies, where he puts rats in this thing called Restaurant Row. It's a square-like maze—yeah, exactly—and as the rat is going counterclockwise, at each door, a sound is released or made.
That sound signals to the rat whether it can go right through the door and get food in, I think, 3 seconds, or whether it's going to have to wait 45 seconds before it gets the food. They're given a bunch of time to try to get as much food as they want. Rats have clear preferences, so some rats will really prefer bananas, and they don't really like the bland food.
What happens? This presents a set of irreversible choices to a rat. Let's say they come up and they can either get a treat right now that they don't really like that much—let's say it's the bland treat—or they can go to the next one and hope that they're going to get the banana really fast. If they go to the next one and the banana sound is long, meaning 45 seconds, then they regret their choice because it would have been better if they had just gone in and quickly gotten the food.
How do we know they're regretting the choice? We can literally watch them imagining eating the foregone choice. We can go into a part of their brain called the orbitofrontal cortex, which activates for certain types of tastes, and we can see them imagining the foregone choice. They end up making different choices the next time around. They end up being less likely to forego that choice in the future.
I think this is such an incredible finding of what we mean when we say model-based reinforcement learning is clearly happening in the brains of very simple mammals.
There's a real challenge in knowing which simulations to run because, if you think about it, we've got a search problem, right? There's an intractable number of simulations to run. How do we fix that in AI, and how do humans fix that?
This is one of the many outstanding questions in AI, but one of the big ones is: how do you effectively prune the search space?
We do not know how mammal brains do this so well. I can give you some high-level ideas or theories, but we just don't know, and this is one of the big possible breakthroughs in figuring out how mammal brains do such a good job of this.
The thing that AlphaGo does, which I think perhaps is a clue and is clever, is that the selection of the search space is actually bootstrapped on the temporal-difference learning model under the hood. This is very clever. Let's say you train something to learn without a model of the world. All it's doing is getting sensory stimuli. It gets a model of a Go board, and then it just predicts the right next action. It just has a policy-value function that bootstraps on each other, et cetera. So, no planning.
If you want to add planning to that, what they did, which is quite brilliant, is say, “Instead of building some other system to try to choose good trajectories, why don't we just use the policy network?” We don't just pick the first one—we pick its favorite move, but then we also look at its second-favorite move, its third-favorite move, and maybe its fourth-favorite move. Then let's literally play the games out.
Let's play a bunch of games against ourselves and then see the ratio at which we win them. What we might learn is that our second-best guess was actually better than our first-best guess. But we're not starting from every possible possibility. We're saying, “Let's bootstrap on our best guesses of good moves, but then check them by playing out the possible futures.”
If we were to analogize that to the brain, what that would suggest is that perhaps it's the basal ganglia—which a lot of evidence suggests is engaging in this type of model-free reinforcement learning—that is actually the thing that chooses the moves. But there's some other system that lets us choose the second-best move, the third-best move, et cetera.
One way this might happen—there's some evidence for this, far from conclusive—is that there's some notion of uncertainty that the frontal cortex or basal ganglia is measuring. When the level of uncertainty between the next actions passes a threshold, pausing occurs. When we see animals do this vicarious trial and error, it almost always occurs in moments of high uncertainty: when contingencies have changed, when the right answer is not obvious.
You could conceive of this as a policy network where you're evaluating its best choice, second-best choice, third-best choice. When there's uncertainty about it—in other words, when they're close together—or there's some other measure of uncertainty, perhaps you have parallel policy models and you're comparing the similarity between them. There are a lot of different ways to do this. That triggers a process of playing forward.
This is another key thing that mammal brains do that AlphaGo does not do. AlphaGo engages in planning on every move. There was never the question of when to pause to plan in a game of Go. It doesn't matter; just engage in planning on every move because it can do it so fast. In the real world, there's so much uncertainty and noise, and human brains need to be so energy-efficient that we can't engage in planning every instant. We need some mechanism that tells us when to stop and think about what we're going to do next and when we can just continuously go with model-free choice.
This is also something we don't know how mammal brains do. But I think a reasonable speculation is that there's some uncertainty measurement occurring.
One last point I'll make that I think parallels, in an interesting way, some of Hawkins's ideas and Friston's ideas: if you take the Thousand Brains model and apply it to the frontal cortex—in other words, we have multiple parallel models of ourselves—you could imagine that there's an uncertainty measurement, the same way we do uncertainty measurement in a lot of deep-learning models, where you create parallel models and just measure how similar the predictions of the models are.
If multiple parallel models predict similar things, we measure that as low uncertainty. When they diverge substantially, all of a sudden we measure high uncertainty. Again, this is speculation, but you could imagine that, if it is the case that we have redundant models in the neocortex, then might it be the case that somewhere—perhaps the thalamus or the basal ganglia—the similarity or differences between these predictions are a measure of uncertainty that triggers pausing?
Stephen Grossberg has similar ideas. He calls this matching and nonmatching.
Yeah. Yeah, almost. Our ability to do abduction is something that fascinates me, and there is some kind of model-selection or matching step that goes on there.
Anyway, we've got the neocortex. It's an absolute beast when it comes to predicting sensory signals, and then we see the emergence of planning and also smart planning, as we've just spoken about. So now we're actually talking about traversing these sensory networks over space and time.
Then something else really interesting happens. The next 2 moves you make are bringing in selfhood—bringing yourself in as an explicit actor and including that in the planning—and then that naturally leads to this idea of, I would call it, teleology, but you would call it “why” or intentionality.
So let’s go on that journey. Maybe self-modeling first: how does that come into the picture?
I think there are 2 notions of self-modeling. One notion of self-modeling is the kind that I think we see in early mammals. This is an idea where the frontal cortex of mammals—a region generally called agranular prefrontal cortex, which is present in all mammals and is largely believed to have existed in the very first mammal brains—gets sensory input from an animal’s own interoceptive signals.
So it gets input from the hypothalamus, which measures things like hunger. It gets input from the amygdala, which measures things like valence in the world—fear, danger, and so on. When does the agranular prefrontal cortex get most excited? It’s in moments of uncertainty and in moments where animals are engaging in planning and episodic memory. If you damage the agranular prefrontal cortex in rats, they seem to have dramatically impaired, if not completely lost, the ability to engage in mental simulation, episodic memory, and so on.
I think what one might speculate is happening here—which is not a novel suggestion by me—is that it’s modeling the self. In other words, it’s modeling the activations of the amygdala and the hypothalamus that are happening, and why I am doing what I’m doing. If I wake up and I see that I have these hypothalamic activations, and then I go down to get water, then it builds a model: in the presence of these hypothalamic activations, the next action is, “I’m going to go over and get water,” which constructs an explanation.
As odd and philosophical as that sounds, that is, in principle, computationally exactly the same thing as when we showed a picture of that triangle and your posterior sensory cortex constructed an explanation of what it saw: “I perceive the triangle.” This idea of constructing an explanation of one’s own behavior is the first idea of self, which is constructing intent.
There’s lots of evidence that even in rats, if you record neurons in their agranular prefrontal cortex, they seem to be very sensitive to the tasks they’re in and to measuring progress toward goals. There’s lots of evidence that it’s doing something akin to that. So that’s one notion of self.
When you get to primates, you see a whole new region of frontal cortex emerge: what’s called granular prefrontal cortex, which is only seen in the primate lineage. There are no other mammals that have this region of prefrontal cortex.
As a quick aside for anyone who’s interested, the reason it’s called granular versus agranular is that most neocortex has 6 layers. The 4th layer is called the granular layer because it contains granule cells, which are just a certain type of neuron. For a variety of really unknown reasons, although there are some interesting speculations, agranular prefrontal cortex is missing layer 4. That’s why it’s called agranular. It’s the same thing with motor cortex, which is missing layer 4.
Most mammals’ whole frontal cortex is missing the 4th layer. It only has 5 layers. But in primates, you get this granular prefrontal cortex, this huge region of neocortex that does contain a layer 4. The best explanation for this I’ve actually seen Friston talk about. I’m happy to go into it if you think your folks would be interested in it, but the point is, there’s a new—
Oh, yeah. I mean, we love Friston. Active inference is actually about preferences because an agent expresses agency by adapting the environment to suit its preferences, or to make the environment like its preferences. So this is a theory of volition, right? What Karl Friston is talking about is: where do these goals come from? Where does volition come from? You spoke with Karl about that, I think, over several years, didn’t you?
Yep. Karl’s been a wonderful mentor of mine. It’s a funny story: I won—I didn’t know this was a term in academia, but there’s the reviewer lottery, where you just get lucky and get a reviewer. For the first paper I submitted, Karl Friston was a reviewer on it, which I was just lucky to have happen. Then he became a mentor, reviewed my book, and gave me lots of good feedback. So, yeah, he’s an amazing person and has been a wonderful mentor of mine.
Let’s talk about granular versus agranular, because I think the best theory I’ve seen is Friston’s theory on this. What does layer 4 do in neocortex? Across the entire neocortex, layer 4 is where sensory input is received. The primary sensory input is received into the neocortical column. This comes from the thalamus.
The canonical model is that sensory input from sensors—eyes, ears, skin—flows up through the brainstem to the thalamus, and then from the thalamus propagates to layer 4. From layer 4, it goes within the variety of other layers of neocortex, and then the other layers of neocortex project back to the rest of the brain.
Why would it be the case that regions of neocortex would not have a layer 4? If you actually watch an animal’s development, what’s interesting is that, in mammals with an agranular prefrontal cortex, it’s not always agranular. It actually starts having a layer 4, and the layer 4 atrophies over development.
I think this mirrors well with Friston’s idea of active inference. What’s happening is that the neocortical column can be in 2 phases. It can either be trying to match its model of the world to its sensory input—in other words, I see sensory input, I’m trying to infer what’s there, and I’m going to construct the idea of a triangle—or there’s another state of a neocortical column, which is generation: I’m going to start from the latent representation of a triangle, imagine it, and explore it.
One idea is that what frontal cortex does is primarily try to fit the world to its model. In other words, it spends the vast majority of its time constructing intent, not trying to modify that intent to fit what it observes, but instead trying to change what an animal does to satisfy its intent.
Layer 4 atrophies. It doesn’t actually go all the way; if you go deep into a brain, you see some basic layer 4. So it’s not completely gone, but it atrophies because frontal cortex spends very little time trying to change its perceived intent to map what the animal is doing. Instead, it tries to change what the animal does to match its intent.
What I think is so interesting and brilliant about this idea is that it explains exactly why layer 4 doesn’t start out as nonexistent. At first, an animal needs to build a model of itself, so layer 4 is present. But over time, it shifts toward, “Once I have a model of what I want, who I am, and the things I would do, I don’t need to spend as much time changing my model of self. I’m going to spend most of my time trying to change my behavior.”
This is a very speculative idea, but it makes a lot of sense in the context of active inference. It’s the best explanation I’ve seen, personally, in all my reading of explanations for why agranularity exists.
This is absolutely fascinating. In my mind, it bridges the gap between internalism and externalism because he’s describing this kind of dialectic exchange. To use the high-entropy words that only Friston uses: if I say “dialectic exchange,” you know. And, by the way, for the folks at home, if you read Friston’s papers, there are certain words he uses. He says, “This licenses something.” If you see the word “licenses,” then Friston wrote the paper.
But you have this kind of thing where agents have models of the world, but they’re exchanging information with the other agents. Then what an agent does is have this generative model of policies, which is just a sequence of actions.
Here’s where I want to get into the nitty-gritty a tiny bit. You could think of those plans as being goals. You can think of a goal as just being an end state in one of your plans. But that doesn’t really satisfy me, because I think of goals like eating food as being a kind of category, not a pointillistic traversal of a specific state in the future. It feels like, well, if there are an infinite number of goals, then in some sense there are no goals at all. So what is a goal to you?
Great question. I think this is where semantics matters a lot. We can think about goals in several ways. If we think in strict RL terms, they would just think about goals as optimizing a reward function. The goal is simple. There’s only 1 goal, which is to maximize reward. In a changing, complex world, your reward function might fluctuate over time, but the goal is singular.
In the active inference world, what I find compelling is that it introduces a different component of what we mean by a goal. This is not just of intellectual interest because it’s cool; it has very real AI implications, actually, because it contains the notion of explainability.
For example, if I wake up and I’m hungry, and I start imagining ways to satiate my hunger, then I decide I’m going to get in a car, go to this restaurant, and eat this specific food.
When I get into the car and someone calls me and goes, “Why’d you get in the car?” the reason I can explain that so easily is because there was a rendered simulation of a plan that terminated in an end state that I deemed I wanted, that I selected. So it’s very easy to explain why I did that. In the absence of that, it’s actually very hard to explain why you’re doing things.
If you’re walking down the street, if I asked you to explain any one of your model-free behaviors—why did you move your foot there instead of there?—you have no explanation. And so I think you can assign the word “goals” to multiple levels here. I think you could say it’s terminating the end state. You could say the goal is some more abstract representation of the satiation of thirst in general, and you could think about that as distinct from just the reward function.
Some might challenge that and just say, “Well, all that you’re talking about, then, is just optimizing the reward function.” But I think you could make an argument that there is a distinction happening there. What I think is so critical, and that’s unique to what happens within mammal brains—and it’s why I was really honing in on that—is the ability to plan a series of actions that terminates in an end result and then execute that plan. I think that has very clear implications for explainability, which model-free actions do not.
That would be how I would think about goals. But I think it is a little bit semantic, where we can think about the concept of goals in multiple ways.
How do we actually know what the cognitive abilities were of early animals, and why should we care?
Great question. I think there are 2 reasons why we should care about the evolution of our brains and intelligence. The first is to understand who we are. The scope of what it means to be a human is not constrained to what it means to be a Homo sapiens.
So much ink has been spilled on the last 70,000 years of us being Homo sapiens. But if aliens were to come down and engage with us and analyze us as a species, most of the things they would observe about us don’t come from our legacy as Homo sapiens. They come from our legacy of being a primate, our legacy of being a mammal, our legacy of being a vertebrate, and our legacy of being an animal in general.
If we want to understand what it means to be a human being, I don’t think we can skip the full 600-million-year story of how we came to be. In there is so much rich history and insight about what it means to be us. That’s one really key reason: it’s our legacy, it’s our history, and it’s how we came to be.
The second, perhaps more practical, reason is that I think understanding the evolution of the human brain and the evolution of human intelligence is a key tool in our toolbox for understanding how the brain works and how human intelligence works. It’s by no means the only method. It might not even be the main method, but it’s a very useful method to add to the toolbox.
The problem with going into the human brain and trying to directly reverse-engineer it is that evolution doesn’t work in clean ways. It doesn’t work the way a human designer would. It doesn’t work from first principles. It tinkers.
When we go into the brain, we see all of this messiness. There are redundant systems and vestigial systems. New things evolve that make old things redundant, but they’re still there. Lots of processing is duplicated in different regions.
One way to understand the brain is to continuously probe it as the human brain is, which is fine. But another method that’s also useful, and can impose constraints for us, is to actually track the history of how it came to be.
That can provide insights as to when this brain modification occurred, such as when the neocortex evolved or when the basal ganglia evolved, what new abilities this enabled, and how it affected the prior brain regions that were already there. This can give us insight into how the brain works today.
In the toolbox that we have for ways to reverse-engineer the brain, I think this is just an underappreciated one that is worthy of being included. Of course, understanding how the human brain works has so many different applications. It helps us with mental health, it helps us with understanding why people do what we do, and it helps us with building AI systems. I think there are lots of insights to garner from the brain.
That’s why to do it. Now, how to reverse-engineer what behavioral abilities exist in our ancestors is a really interesting question. Of course, we can’t go back in time. What we can do, though, is use mechanisms to reverse-engineer what their brains looked like and mechanisms to reverse-engineer what abilities they had.
To understand what their brains looked like, we can look at other animals in the animal kingdom. For example, we can look at all of the existing primates and all of the existing non-primate mammals, and we can see what the common brain structures are between them.
We can look at genetic analysis, meaning what things seem to derive from similar roots, and we can infer what seems to be common and shared amongst them. Thus, what do we think was actually existing in the brains of the first mammals?
We can do the same thing with fish and reptiles to infer what existed in the first vertebrates, and we can do that with invertebrates to try and infer what existed in the first bilaterians—in other words, the first animals with brains.
For behavioral abilities, there are 3 ways you do this. This is my approach to trying to infer behavioral abilities. I call them the ingroup condition, the outgroup condition, and the stem-group condition.
In order to make the argument that a behavioral ability emerged at a certain location in our evolutionary history—for example, a behavioral ability like episodic memory evolving with the first mammals—you need to satisfy these 3 criteria.
The ingroup condition stipulates that most descendants of this species—in other words, most mammals—should show this ability. It doesn’t mean all of them. Abilities get lost all the time, but most of them should show this ability.
The neural mechanisms by which the ability emerges should come from homologous regions. What that means is a shared neural underpinning. If, for example, mammals show episodic memory but it comes from neurological regions that independently evolved along different mammal lineages, that suggests it wasn’t present in the first mammals.
But if they all emerge from regions that emerged with early mammals, that’s good evidence that this ability also emerged with mammals.
The outgroup condition says that most doesn’t have to be all, but at least many non-mammal vertebrates—the outgroup, the group right above—should not show this ability. If they do show the ability, it should emerge from non-homologous regions; in other words, parts of their brain that evolved independently.
For example, birds definitely show episodic memory. But when we look into the brain regions from which episodic memory emerges, it’s clearly non-homologous. It’s a part of the brain that mammals, or the early vertebrates, did not have.
The stem-group condition is that, in the ecological dynamics in which early mammals existed, or this ancestor existed, we should be able to devise an argument for why this ability would have been adaptive. Why would episodic memory have evolved?
With these 3 things, we can start to infer the story of when behavioral abilities emerged. Is this perfect? Absolutely not, because we do not have a ton of data on behavioral abilities across species. As new evidence emerges, the story might change.
But with these 3 conditions, we can do a reasonable job inferring what abilities emerged. The main finding of the book—or the research that led me to be so excited about the book—is that when you do this, what’s kind of crazy is that you find, as a first approximation, a really coherent story.
A lot of the behavioral abilities that emerge at each milestone in brain evolution don’t seem to be a haphazard array of different skills. They often emerge from one underlying intellectual capacity, which I call a breakthrough, applied in different ways. That’s the idea of the 5 breakthroughs.
One thing to note about the basal ganglia is that I think this is a bit of a sidebar, but I think it’s fun. The basal ganglia is one of the most underappreciated parts of the brain, in the sense that so much work has gone into understanding how the neocortex functions. So much work has gone into the neocortical column, and that’s all wonderful work to be done.
But the basal ganglia is not only evolutionarily much older. If you look in a lamprey fish, as we talked about last time, a lamprey fish has a common ancestor with us from 500 million years ago. It’s one of the most distant vertebrate cousins that still exists today.
Lampreys have a basal ganglia that looks exactly like our basal ganglia: the same internal structure. The basal ganglia also has perhaps one of the most beautiful internal structures that can be computationally reverse-engineered.
There isn’t good consensus on the actual internal wiring and computations performed by a neocortical column, but there is much broader consensus as to what’s being executed by the basal ganglia. Without getting overly technical and perhaps boring people, I would encourage anyone who’s computationally interested in this to dive into the literature here.
It is almost beautiful that evolution came up with this. For example, the input structure of the basal ganglia has this mosaic of neurons that each express 2 different types of dopamine receptors.
This would be my own little technical diatribe. One type is called D1 receptors, and then there are D2 receptors. D1 receptors, when they receive dopamine, strengthen connections. D2 receptors, when they lose dopamine, strengthen connections.
Now you track these different neurons, and they actually split their paths. The D1 receptors go to a nucleus that, when activated, disinhibits behaviors, and D2 neurons go through a different set of nuclei that, when activated, inhibits behaviors. We can literally watch how dopamine signals drive repeated behavior and how dopamine drops inhibit behavior.
We can literally look at the mosaic of connectivity here and say, “When you spike dopamine, it weakens the stop pathway through D2 neurons, and it disinhibits the go pathway through D1 neurons, making you more likely to repeat the behavior, and vice versa if something bad happens and you lose dopamine.” I think it is so beautiful that evolution stumbled on something that clean in its macrostructure. So, diatribe on—I think the basal ganglia is cool.
Yeah, it’s quite interesting as well. People get addicted to drugs, and a lot of that is about wireheading in the basal ganglia. Of course, habitual learning is something that, when it becomes so habituated, moves down the stack into the basal ganglia, and drug-taking would be an example of that. You actually cited, I think, an experiment in China where they removed part of the basal ganglia, and there was a 40% recidivism rate for addiction.
Yeah, a very controversial study that probably violated many ethical codes in the US. They did this study on people with intractable heroin addiction, and they lesioned a part of the basal ganglia called the nucleus accumbens, which is sort of where goals are habitually selected. It showed a dramatic reduction in heroin addiction. It also had other side effects that doctors might deem unreasonable, but it definitely worked and absolutely reduced the addictive cravings triggered by stimuli.
Yeah. The story of this chapter—an incredible chapter I’ve just been studying today in great detail—is the story of mentalizing, but I would call it social complexification. Actually, that’s a hypothesis for why our brains expanded so dramatically. There was this extinction event. I think it was the Devonian extinction event, and only birds and not many other things survived. Then we got the chance to evolve after that, and our brains rapidly exploded.
There are different theories about why that happened. Maybe it was because we had access to loads of calories in the form of fruit—preferential access—and it gave us an incredible amount of time and energy, the excess of which might have led to social complexification. Can you just give us a little bit of background about that first piece?
Absolutely. We don’t know—there are lots of speculations—but we do have some really good evidence that at least part of what drove the explosion in primate brains was social dynamics. Robin Dunbar did the seminal work here. What he showed is that, in primates, the encephalization quotient—which is just the ratio of brain size, especially the neocortex ratio, so the ratio of the size of the neocortex to the rest of the brain—is extremely correlated with social group size in primates.
The bigger the social group size in a group of primates, the bigger their neocortex seems to be relative to their body size. What’s so interesting is that you don’t see this in most other mammals. This is not a standard correlation that applies across the animal kingdom. It seems to be a correlation that’s very specific to primates. There might be other mammals, but for most mammals, you don’t see this correlation.
Robin Dunbar’s famous social brain hypothesis is that what drove the explosion in primate brains was some form of social dynamic between people. The more social relationships you’re managing, the bigger your brain has to be. Now, that doesn’t explain what specifically happened in the brain, which is what we can get to with mentalizing and why that applies. But it does suggest that whatever drove this explosion in brain size seems to be something correlated to social grouping.
What’s interesting about primate social groups, relative to—not all mammals, but many other mammals—is that they’re very, very political. Many mammals live in solitary social lives, where the males mostly live alone and females will rear a child and then usually go off on their own. There are animals that live in herds, where they socially group together, but there aren’t really very rigorous hierarchies amongst them.
Primates, especially apes, have these really complicated social structures with rigid hierarchies. There truly is someone at the top of the hierarchy, and we can measure this. Primatologists have gone to painstaking lengths to verify it. For example, there’s transitivity: if you show that one primate tends to show a submissive signal to another primate, and that other primate shows a submissive signal to another one, then it’s almost definitely the case that the first one will show a submissive signal to the last one.
In other words, these are real, rigid hierarchies, not random interactions of submission and dominance. One of the main ways you survive as a primate is by successfully climbing this hierarchy. What’s so interesting is that, in many mammals, what makes someone the top dog, or the person at the top of a hierarchy, is brawn. It’s just strength. They’re trying to flaunt who would win in a physical altercation, which evolutionarily is beneficial, because if you can prevent actually fighting each other and just say, “This is who would win the fight,” then you both save energy by having these fake battles, and whoever wins gets to eat the food, et cetera.
But with primates, it’s not always the strongest one that reaches the top. It’s the most socially savvy one. Social savviness comes into play through alliances that are built within primates. You’ll see that people at the top of the hierarchy frequently befriend, groom, and come to the aid of certain other non-family members. Those folks will thus reciprocate and come to their aid. There are these really interesting dynamics that play out. You even see wars. There are mutinies that take place in primate societies.
In this soup of the way to survive and gain evolutionary advantage as a primate, the goal is not perhaps only to make sure you get access to food, but to climb a social hierarchy. All of a sudden, there are huge social pressures to infer what someone else would do in a certain circumstance, what someone knows, what you can get away with, or how to change someone’s opinion of you.
What that lines nicely to is what we see in the new brain regions that emerged in primates. Most notably, there’s a brain region called the granular prefrontal cortex, and there are areas in the back of the brain called the superior temporal sulcus and the temporoparietal junction. These brain regions across primates are highly implicated in what I call mentalizing—which is thinking about thinking—but the standard literature would call this theory of mind. That means being able to infer the intent or knowledge of someone else.
It’s easy to understand why this would be so adaptively valuable in a politicking arms race where you’re trying to deceive each other. There’s a great study that really revealed this with primates by Emil Menzel in the 1970s, and I love this story.
Are you going to do the Machiavellian apes?
Yes.
Yeah. Yeah. God. [Laughter]
Okay. So, Emil Menzel had this 1-acre forest, and his main objective was not to study ape social behaviors in the sense of how they would climb social hierarchies. His only objective was to measure spatial reasoning in chimpanzees.
He had a group of chimps. There was 1 chimp named Belle, another named Rock, and a few others. He would show Belle the location of food. He would hide food under a bush and then see whether Belle would go back to that same location looking for food. In other words, could she remember locations in 3D space or in a 2D, map-like space?
What he found is that, yes, they readily do that. We now know that lots of mammals are capable of it. In fact, even fish can do things like that. But in this study, he started finding something that was odd.
When Belle would find the food, she would frequently share it with her fellow chimpanzee group members. That was great until Rock, who was a high-ranking, aggressive male, would take the food from her when she shared. So what she started doing was hiding the food when she found it. Instead of sharing with Rock, she would just sit on the food.
Rock realized she was doing this and not sharing. So Rock would come over and push her to try to get the food from under her. Then, when she knew the location of the food—because, on some recurring cycle, experiment, or signal, she knew that the food was now available—she would not go to it until Rock was not looking.
So then what did Rock start doing? Rock started pretending not to look. Rock would look away while Belle went toward the food. Once he noticed she was doing that, he would turn around and run to try to grab the food before her. Then Belle started trying to lead him in the wrong directions, and this cycle of deception and counterdeception kept playing out.
What that becomes is a beautiful anecdote and a case study in what happens when you have a bunch of hierarchically interacting animals in an arms race for things like this. What you get is deception and counterdeception. That’s really only conceivable with some notion of theory of mind, because in order for Belle to trick, or try to trick, Rock, she needs to be able to say, “In order for me to change the knowledge in Rock’s head of where the food is, what I need to do is walk in this other direction. What that will do is make Rock think the food is in this direction, when in fact I know it’s in this other location.”
She also needs to reason about someone’s intentions: “I know Rock intends to trick me, so when he’s looking away, I don’t believe that he in fact is not paying attention.” This was one of the first early anecdotes that some form of theory of mind was occurring in primates. There have since been lots of studies that show this.
For example, just to give some case studies, you can take a chimpanzee and teach it that, when there are 2 boxes, the box with a red mark on it is the one with food in it. They’ll easily learn that. Then you have an experimenter come in with 2 boxes, bend over, and mark one. They pretend to accidentally drop the marker on the other one and then leave. The marking is identical in both cases, but the chimpanzees always go for the one that was intentionally marked. They can infer the difference between the same stimuli: someone intending to do something and something being an accident.
There are other studies of chimpanzees playing with different goggles. With one goggle, you can’t see through it; with another, you can. If you put those goggles on human experimenters, the chimpanzees always go to the experimenter with the see-through goggles, asking for food. They’re somehow inferring that the other person can’t see them, so why would they ask?
There are lots of studies that show this ability. Evidence outside of primates, in other mammals, is very loose. It’s inconclusive and controversial, but the loose evidence shows up only in the smartest mammals, which possibly suggests some independent convergence.
Yeah.
Yeah, there’s lots of rich evidence that this theory of mind exists within primates, and it emerges from these uniquely primate regions. I’m happy to go into the evidence from brains, but I’ll stop there for a second.
So, with the Machiavellian apes and the X-risk people, there are people who talk about AI killing everyone, and they make the argument—it’s called instrumental convergence, from Nick Bostrom—which is basically that things like power-seeking and deceptive behavior would be instrumental to any end goal. This is a great example from the animal kingdom of deceptive, Machiavellian behavior.
I guess it does seem plausible, at least on the surface, that a level of sophisticated agents following their own intentions and inferring the intentions of others would seek to deceive each other. That’s a natural phenomenon.
I absolutely think so. I absolutely think it is the case that the more autonomy you give an intelligent agent, and the more ability you give it to define its own subgoals, the more risk emerges. You absolutely get what Nick Bostrom is talking about, which is that a subgoal to trying to help cure cancer might be to dominate all of Earth and control the labor supply and allocation of resources across all of Earth.
But I don’t think that’s necessarily inevitable. I think it is a risk. Evolution is a constrained search algorithm for intelligent entities. It does not give moral weight to what emerges. This is an important distinction: just because something is a natural consequence of evolutionary systems does not mean that we should deem it morally superior.
Yeah, that’s the naturalistic fallacy, right?
So it might be the case that it is very likely that species will eventually enter a politicking arms race, and certain forms of deception and power-seeking will emerge. That doesn’t mean that when we produce our own intelligent entities in AI, we should imbue them with those features.
One of the optimistic outcomes of this new AI world we’re going to enter in the next 100 years is that we, as designers, can now do our best to try and remove some of the evolutionary baggage that we don’t like, which has evolved in humans, from these new entities. There’s risk, of course, but I think there’s also a really great opportunity that we could have benevolent beings that do not seek to dominate. Yann LeCun talks a lot about this, and about creating beings that are less selfish.
So, I think there’s a great opportunity, but there’s definitely risk, because the second you give an autonomous agent the opportunity to produce its own subgoals, you need to have really rich constraints, a really well-defined reward function, or both. One ability that I think comes from mentalizing—and this is an idea in alignment research—is that if you can convince an AI agent to try and do what it thinks the human wants it to do, what you’re actually doing is requiring it to engage in some form of mentalizing: to infer the preferences of the requester and then try to do what is best for that individual.
You can’t just have it take requests at face value, because then there are all these opportunities for misinterpretation. Nick Bostrom’s famous paperclip maximizer maximizes the production of paper clips, and Earth is turned into paper clips. We obviously don’t want that.
But with mentalizing—with the ability to model the internal simulation of another mind and play out how this person would feel about possible futures—you could imagine, optimistically, an outcome where an AI agent could easily infer, “If I turn all of Earth into paper clips, that’s not what the person giving me this request would in fact have wanted. They would regret that outcome.”
Of course, it doesn’t fully derisk things, but it is one methodology and one learning from evolutionary neuroscience that we can garner. Mentalizing is a tool that can be used to try and stabilize the requests we give each other in a more grounded way, so there aren’t these types of misinterpretations. Of course, humans misinterpret each other all the time. It’s by no means perfect, but it is a tool.
And I think that’s a very natural phenomenon. I think any intelligence system is naturally incoherent. I think it’s impossible to have a single, monolithic intelligence that is monomaniacally focused in a particular direction.
But I want to slightly rewind to what we were saying. The first animals had quite simplistic social games that they were playing. They were interested in strength and submission, and it was a fairly fixed interface.
What was really interesting is that deer, for example, lock horns, don’t they? It’s predictive. They don’t actually have to have a fight, because that would be evolutionarily not a smart thing to do. The social game they play, even though the game is fixed, is predictive, which is fascinating.
Then you were telling the story of monkeys and macaques, how they have this really interesting virtual social game where strength and social status diverged. Social status actually became this virtual thing that was based on grooming and preening and lots of completely unrelated things. It was entirely possible for a very weak macaque to have significantly higher social status than a big, strong one.
That’s really fascinating. But then we get into a broader question. We’re still very social creatures ourselves. We have Facebook, for example, and could you arguably, cynically argue that Facebook—or all socializing—is just a kind of arms race to improve our social status? When we’re posting on Facebook, in a way, it’s like the deer locking horns. It’s us playing these status games without having to fight each other.
I think there’s an aspect of human behavior that can absolutely be explained by this. There’s a great book called The Elephant in the Brain.
Oh, yes.
There’s a great book called The Elephant in the Brain that talks about how much human behavior can be explained by this sort of status-seeking behavior. The reason why it’s so hard to study is because it’s what they call, I think, a cognitive taboo or an intellectual taboo. We don’t want to admit it to ourselves.
So we self-delude ourselves into believing we’re doing things for virtuous reasons, because it makes it easier and more convincing. If I know that I’m doing something to deceive someone, it’s easier to tell that I’m doing it. If I genuinely believe that I’m doing these things to help the world, but subconsciously they’re actually just benefiting me, it’s more convincing to other people.
Their argument is that primates evolved this sort of self-deception to make themselves more convincing, and this really accelerated with language in humans and all that.
I think it’s very likely to be the case that the core thesis of their book is right: a lot of human behavior is this sort of subtle status seeking. I don’t want to go on a tangent here, but I do think it has sociopolitical theory implications. How do we make sure that society doesn’t devolve into just a hedonic treadmill?
The interesting thing about social status is that it’s always, definitionally, a scarce resource because it’s a ranking game. Unlike physical resources, where it’s possible for all of us to live better than kings or queens did 1,000 years ago—we can all have better access to medical care, better access to information, and better access to food—social status is always a zero-sum game, unless someone can conceive of a better way to do it.
This is problematic because if, over time, most of our actions become about pursuing social status, then we’re going to forever be in this sort of game. I don’t think personally that we’re doomed to this. I think there are absolutely better virtues in human psychology, where not everything we do is based on pursuing social status, and I think you can conceive of dynamics where humans are doing things for other reasons, not just to gain status. But I think it’s absolutely fair to say that a surprising amount of human behavior is status seeking, and maybe a depressing amount.
Yeah. I agree with you, and I don’t necessarily want to get too philosophical on that, but that book was The Status Game by Will Storr, where he said that there were 3 meta-status games that we played. He gave the virtue game, the dominance game, and the success game. I might be playing the success game: I want to have the best podcast, or whatever.
The reason I bring this up is that the difference between humans and animals is that they’re just playing one game. It’s really interesting that they have this mimetic social score, but the game is the same everywhere, whereas for us, we go 1 level of abstraction up. The success game for us can be manifested in a myriad of different ways. It could be success at playing computer games; it could be writing books or making podcasts, or whatever.
It’s almost like we fractionate our social ranking into a myriad of different games. I think that’s a little bit of a testament to the difference we humans have in general with our metacognition, which is our ability to create the memetics in a novel way.
One thing that’s sort of related to this, at least in early human societies, is one way to reduce status infighting: make it such that members of a team have distinct roles. I don’t think this lesson only comes from management theory and entrepreneurship. I think this probably derives from either early human or maybe even early primate societies.
It’s much more stable, and you can introspect that it feels much more comforting to be part of a troop of 100 humans where pretty much everyone is pulling their weight and everyone matters because they’re doing their own distinct thing. That is a very stable state, where we’re not infighting as much because we’re all doing something; we all matter to some degree.
But when there’s infighting because there are only 5 blacksmiths, or 5 podcasts, or 5 books about the evolution of the brain, all of a sudden these other types of things start emerging because we’re no longer all fulfilling a role that matters. It feels like there’s only a ranking, and only 1 of these is going to matter.
As an example of ways to reduce this sort of status seeking, I see this in business all the time: the more you can create an environment where it’s not zero-sum, where everyone’s pulling their weight and together we all win, the more the best versions of humans emerge. The more zero-sum it becomes, and the less distinct the roles and activities are, the more of these—I would argue—primitive primate behaviors start emerging.
An early-stage company, a company of 30 people, has such different dynamics from my last company, where, when I left, we were 400 people. The social dynamics are so different, and I do think one could speculatively correlate that to our evolutionary history here. In a 30-person company, you don’t need that much structure. If you have people who work well together, are aligned on a mission, and you get rid of people who are generally mean-spirited or have bad intent, you don’t need a lot of structure and process to get people to work well together, support each other, and move in a common direction.
I think what that demonstrates, when one observes that, is that what’s playing out is an evolutionary program that got groups of 30 humans to work really well together. When you’re at 400 people, what very quickly happens—and it takes a lot of work to fight this—is that you start getting internal factions emerging, because what splinters out is these subgroups of 30 to 100 people that then have their own points of view. Then it’s very easy to have an us-versus-them dynamic with other groups, and you start seeing things break down.
One mechanism for solving this is very rigid hierarchies; that’s what the military does. Another mechanism is to embrace the chaos, which is a little bit more what Google does. Another mechanism is to effectively make it a constellation of different startups, which is what Amazon does, where each group is kind of autonomous and has very clean interfaces with other groups.
There are many different management approaches to this, but the breakdown, I do think, emerges from the fact that humans did not evolve to interact with 100,000 people. We evolved in an environment where we interacted naturally with about 100 people, and that’s why that comes very naturally. We don’t need as much process to make that work, but we do when we scale it up.
Yeah. It’s fascinating. I mean, as you say, you could argue that Amazon has 1 overarching goal: to make money. But as soon as you increase the autonomy in the organization, it’s a very human trait, isn’t it? You were talking about the Machiavellian behaviors and the deceptive behaviors, and you just wonder how much energy is wasted with infighting.
I’ve even made the comment that in the military, they might be doing quite simplistic jobs compared with Google, but even at Google, there’s an obsession with job level. I mean, if you go off of Levels.fyi or if you go on Blind, that’s the only thing people talk about: their total compensation and job level.
Maybe we should save the cynicism. Coming back to the chapter, we’re telling the story, basically, of how this metacognition and this predictive apparatus gave rise to an entire suite of complex social behaviors that we see in primates, which is fascinating.
Maybe we should just talk a little bit about what I call why bootstrapping. There was 1 guy at Toyota Research who was quite famous because he would get people to ask why 5 times. You say, “Why? Why? Why?” It’s almost as if there’s some magic number. Everyone is only a certain number of degrees of separation away, and it’s a similar thing: you only need to ask why a few times and you’ll always get to some kind of base reason.
Maybe that’s why, evolutionarily, we have 2 levels of causal metacognition in our brain. We have the agranular prefrontal cortex and we have the granular prefrontal cortex. I guess 1 potential question there is: why is there not a 3rd level of asking why, and what would that look like? Can you just sketch out that metacognition picture in general?
When we think about what a granular prefrontal cortex does, a reasonable framework for it is that it generates explanations of an animal’s behavior. It models an animal’s own behavior. One cognitive tool to reason about that is this: if it observes a rat wake up, have certain hypothalamic activations, and run in a certain direction to drink water, it produces a representation that could be interpreted as, “I am explaining this behavior by: I am hungry, as an animal.”
That can be useful in a variety of ways. It can trigger simulations to find alternative solutions to satiate the same need. If you put a rat in a novel situation, but the granular prefrontal cortex infers that right now I am hungry, we can start triggering a bunch of simulations to try to satiate the same desire, to fulfill what I believe about myself through alternative means. This enables an animal to be flexible.
This is the explanation of an animal itself. Why would that be the case? What I argue in the book is that the granular prefrontal cortex builds a model of that model. Instead of a simulation, it’s a simulation of the simulation.
What that would mean is, if you could—as a thought experiment—ask, let’s go 1 step further, or 1 step back. If we could ask the basal ganglia, which is the sort of reinforcement-learning system, “Why did you turn left to go in this direction to drink water?” it would just say, “Because turning left maximizes reward.”
The answer would always be the same. If you ask the agranular prefrontal cortex, “Why did you turn left?” it would say, “Oh, because I’m thirsty. There’s a specific thing that I, as an entity—this animal I’m modeling—want to achieve.” But if you ask the granular prefrontal cortex, it would say, “Well, I turned left because I am thirsty, and that made me think about ways to satiate my thirst. I simulated going to the left, and I remembered water being there because last time I was there, there was water, and so I went to the left.”
And so, in other words, it enables you to simulate different types of simulations and reason about what you would think in a new setting, which, of course, enables you to think about what someone else might be thinking. We do this all the time. Someone doesn’t respond to a text message, someone makes an odd facial expression in a social interaction, and we’re immediately trying to figure out: What is this person thinking? Why would they do this? And so on.
So the first question is: Why do we even need this new level at all? I think one of the main adaptive values is that it enables your survival in the politicking arms race, because now, if I can simulate a simulation, I can infer why you might do a certain behavior, how to manipulate someone’s knowledge, and your intentions behind things. So this is why you would have one layer to go a level above.
You could make an argument that theoretically there should be an infinite scaling up of whys. I think this is maybe a cop-out, but I think there are huge energetic costs to any sort of scale-up. So what that means is, the question is not whether there would be benefits to a third level of hierarchy; the question is whether the benefits of a third level of hierarchy would outweigh the massive energetic costs of producing it.
And so I think that would be my first-blush explanation as to why we might only have 2 levels instead of 3 or 4: because the second level added a clear adaptive value relative to the cost to survive in the politicking arms race, and the third one perhaps was superfluous and unnecessary relative to the energetic cost.
I think having that second level of metacognition does a lot of work, right? And I’m going to talk a little bit about that now. But one of the things is, you can infer the intents and knowledge of others through the same process of doing simulations yourself. So you can kind of imagine yourself doing something, but you can kind of swap out the pointer to be someone else and swap out the knowledge to be someone else. And that’s incredibly valuable.
But the knowledge thing is really, really interesting. So I asked the question last time, and this is something that I’ve been quite confused about, and I feel that reading this chapter has actually really cleared it up for me, which is about goals. Because when you look one level down at the agranular prefrontal cortex, it’s modeling intents, and then this granular prefrontal cortex, which is trying to seek explanations about the level below, which is the agranular prefrontal cortex, is going a level of abstraction up and modeling goals—not intents—but it’s actually modeling knowledge. What it’s doing is categorizing.
So when you have a simulation of a simulation of simulations, what it’s doing is creating a category. So, to the example you just gave before, “thirsty” becomes a category. Rather than it being a pointillistic intent, it’s a little bit like saying, “I can go and have a sandwich, or I can go and have McDonald’s,” or my abstract simulator could kind of draw a boundary around those things, and now I’m getting food.
As well as being able to categorize intents in yourself and other people, you’re also categorizing knowledge, and then it can be shared memetically. So it’s almost like just going to that second level of metacognition gives you so much that you didn’t have before.
100%. Yeah. I think thinking about the level of the granular prefrontal cortex and the new primate regions as enabling something akin to knowledge is a really wonderful way to look at it, especially, one, from the connectivity analysis, and then, two, just from what we mean by knowledge.
From the connectivity analysis, if you look at the superior temporal sulcus and the temporoparietal junction, these are regions of the posterior cortex that, in simple terms, are at the very top of the hierarchy. I mean, they get multimodal input from all the other regions of sensory cortex. So a very simple rule of thumb for understanding this is: this models the rest of sensory cortex. I understand the full rendering of the simulation of the external world that is happening, and this is where I build a representation of that.
And it is perhaps no coincidence that’s also where we see brain regions light up when you’re engaging in things like theory of mind and solving false-belief tests—in other words, trying to infer the knowledge of someone else. These same regions light up.
And what do we mean by knowledge? I would argue that knowledge can mean a few things. One is procedural knowledge, where I just know how to do certain motor behaviors. I don’t think that’s what we mean. I think we mean more semantic knowledge or episodic knowledge, which would be: I know that water is over there, and I know that if I do this behavior, this will be the causal outcome.
That type of knowledge, I think, is absolutely rendered in the mental simulation. When I imagine certain things—when I imagine the case of lightning hitting the ground—what do I see afterwards? I see fire. And that’s the source of my knowledge about the causal relationship between these 2 things. So having a layer that models the simulation enables me to reason about my own knowledge and to see what the effect of changing knowledge is on behavior. And this, of course, enables us to flexibly adapt to other people’s behavior and predict what they would do under cases of different knowledge and different intents.
Yeah, that’s fascinating. But you did say that there was a bit of a riddle about the granular prefrontal cortex, because there was 1 study where it could be damaged and the person would still score really highly on IQ tests. But you said it’s about being able to project yourself in simulations, this kind of abstract modeling of your own mind. So in this particular case, how could the person still score the same IQ without that part of the brain?
So this is such a cool story in the history of neuroscience. You would think that if you look at a human brain, I mean, the granular prefrontal cortex is this huge region in the front of the brain. I mean, it takes up a gargantuan amount of space. You would think that taking a chunk out of that part of the brain would have a gross effect on a human being.
Just like if you took a part out of even a relatively small region of the back of your brain, which is where your visual cortex is, you become hugely visually impaired. You take a region out of your motor cortex and humans become largely paralyzed for months until they recover from that. You take a region out of auditory cortex and they can lose the ability to recognize even words.
So there are relatively small regions of neocortex that, if there’s damage to them, have gross, obvious effects on human intelligence and behavior. After World War II, there were so many patients with brain damage that there were all these studies, and people could not figure it out. It was a puzzle: What does this huge region of prefrontal cortex do? People don’t have—something seems off about them, but it’s not obvious what is wrong with them. People would note personality changes. They don’t seem to be themselves. But on logic tests, on IQ tests, it wasn’t obvious they were dramatically impaired.
In many cases, it wasn’t obvious they were impaired. There was 1 famous case where they could test someone before and after because, for surgical reasons, they were going to remove parts of the granular cortex. This patient actually improved on IQ tests, which made this a huge puzzle: What does this part of the brain do?
And so, if you track the studies from that point forward, we start learning that what granular prefrontal cortex does in large part isn’t related to these types of logic puzzles. It’s related to thinking about thinking and modeling ourselves.
So, for example, if you look at someone who has damage to granular prefrontal cortex, someone who has damage to the hippocampus, and someone who has a normal brain, and you ask them something very simple—you give them a random word and you say, “Just tell me a story. Just imagine a story of you with this word.” The word could be “restaurant,” and you compare these stories, you immediately see something very different.
The people with hippocampal damage give a very, very rich story about themselves, but the external world misses details. So there’s not a lot of rich detail about the external world. This is consistent with the idea we talked about with early mammals, where the hippocampus helps render a simulated external world.
The people with granular prefrontal damage could render a very rich external world. They could tell you the details of the leaves, the smell of food, exactly what a restaurant looked like, but they themselves were woefully missing from the stories. They could not project themselves into this imagined world.
And so then, if we go back and look at all these other things that light up granular prefrontal cortex, if you ask someone to think about how they’re feeling, granular prefrontal cortex lights up—self-reference. But if you ask another question, such as, “What does it look like outside?” the granular prefrontal cortex does not activate. The agranular prefrontal cortex will activate in both cases.
So we start to see that it’s in cases of thinking about yourself and thinking about others that this granular region gets very activated. And now, if you go back and study these people more deeply, you notice that they become hugely impaired at false-belief tests. They can’t recognize faux pas, so they don’t understand what’s not really appropriate. Which, of course, makes sense, because how do I know what’s appropriate? I’m going to infer how you feel about the things that I’m saying. And so you see all of these mentalizing impairments that emerge, but it’s not related directly to these logic puzzles that are typically in things like IQ tests.
Yeah. You mentioned the false-belief test. Can you just briefly sketch out what that is?
Yes. There’s a good picture if you want to hold it up or show it on the podcast. The way the test works is you have Sally on the left, who has a basket, and then you have Anne on the right, who has a box. So Sally puts a marble in the basket, and then she walks away. Then Anne goes over and moves the marble from Sally’s basket and puts it into her box, and then leaves. When Sally comes back, where does she look for the marble?
It’s so simple. But in order to figure out that Sally will look into the box, you have to understand that it’s possible for another mind to have incorrect knowledge—to have a false belief about something. Young children don’t understand this. They assume that knowledge is omnipotent: everyone has the same knowledge about the world. But at a certain point, they start learning that it’s possible for people to have false beliefs.
So we actually know that nonhuman primates can do this. They’ve done studies on macaques where you do exactly the same Sally test, and you just look at where their eyes look when the person comes back into the room to look between the 2 boxes. They always look, or tend to look, in the direction of where that person thinks the marble is, or the piece of food is, not where it actually is. If you inhibit their granular prefrontal cortex through an injection or another mechanism, this bias goes away. They no longer look in the right direction. There’s lots of really good evidence that this sort of false-belief mechanism is occurring in these primate regions.
Yeah. And what really hit home to me is that, in a way, it’s not even knowledge. It’s all simulations. It’s just simulations of other agents. We’ve always spoken about knowledge in some weird Platonic, abstract sense. I quite like the idea that the primitive form of communication between humans is just simulations, even when we’re speaking to each other.
Yeah, exactly. So how do we solve the Sally problem? It probably happens so quickly, but we just simulate what we would think if we were in Sally’s shoes. And then I realize, well, I would look in this place. And this helps us reason about other people. And this begs a really, almost profound question: How unique is theory of mind?
This brings me to a question that I’ve been asked multiple times, which is: Does ChatGPT have theory of mind? The evidence I should stipulate, for anyone curious, is that if you ask GPT-3 these sorts of theory-of-mind puzzles, it does terribly. So that’s an easy one to discard out of hand. But if you ask GPT-4 these theory-of-mind puzzles, it performs remarkably accurately, at a human level, on these theory-of-mind puzzles. And there have been people who have explored whether it’s just in the training data, and there’s good evidence that it’s not just because they’re regurgitating what was in the training data.
So does that mean that ChatGPT has theory of mind? I think there are a few ways to reason about this. One is: What do we mean by theory of mind? If by theory of mind we just mean the ability to solve these sorts of false-belief puzzles, then I think you have to accept the fact that, yes, it can solve those tasks. The problem is, the way in which it renders this model of other minds is not through having a similar mind itself. And so what this means is we should be concerned—it doesn’t mean it won’t work well—but we should be concerned about how well this will generalize to real tasks where we might care about this much more deeply.
For example, with a human, there’s good evidence to suggest that part of my ability to reason about your mind is because I have a mind that works quite similarly. We are almost bound together by some common mechanistic synergy between the way in which our brains work, because our brains are quite similar, which enables a lot of data efficiency. I’m pretty good at predicting what people do—not perfect, but pretty good at predicting what people do—because we’re all people, and there are similarities between how we act. That makes us quite data-efficient and decent at generalizing to new situations where we put people in new places that we’ve never seen before. I can kind of guess, well, if I were in that situation, this is what I would do.
GPT-4 has learned to build a theory of mind simply by reading the text of these puzzles, and so clearly it has some mechanism to build a model for predicting what people will do in certain circumstances and differentiating knowledge and intent, et cetera. But the concern is twofold. One, what will happen if we take those types of models and put them in very new situations that are not based on just these puzzles, but, for example, we’re asking them to optimize a paperclip factory? That’s a situation where we should be concerned. How well will it do at actually inferring what we mean by what we say?
And the second is data efficiency, which is: How much data did it have to see to build this model? If it was a ton of data, then it’s going to be problematic if we have these new situations where we want to teach it to model people’s behaviors in this new place. If it requires a ridiculous amount of data, then it’s always going to be slow to learn these things and always be at risk of not generalizing well when we put it in these new situations.
My answer here is nuanced. I think if by theory of mind we mean solving puzzle questions, it’s very hard to say that ChatGPT does not have some model of human behavior. But I do think the human and primate mechanism for doing so has a data-efficiency advantage and a mechanistic-synergy advantage. In other words, we can use ourselves to reason about things, and that is relevant. If we want to have these systems do a good job listening to human requests, we shouldn’t translate performance on false-belief tests into believing that they’ll do a good job correctly inferring our intent and knowledge in new situations.
Yeah, I would agree with that. I think ChatGPT is in the world of text, and it's learned all of this structured narratology and things on Reddit and things on Twitter. And as we were saying last week, language has evolved to be very simple. It has to be learnable by children. It has a small subspace. But it is a real kind of generalization over human behaviors, and it's in this very low-resolution substrate. Whereas in the Machiavellian apes example that we were talking about before, these are agents performing real-time sensing and inferencing and making in-the-moment judgments, and they're in this continuous sensor domain where they have many different types of signals: visual signals, sound signals, and also memory of what happened in those dynamics just before. So it feels like a difference in kind to me between those two situations. But it is remarkable that in the GPT domain any kind of theory of mind could work.
One good example of this, I think, is whether there is a difference in our human ability to predict behavior between a car and a person. So the brain is always able to model things it observes, simulate them, and predict what they will do. I can look at a car and imagine different colors of it, and I can imagine what will happen if I drop it and it rolls down a hill. We build models of things all the time. We build models of computers and models of books. So the brain produces models of things. Is the way that the brain produces models of other human behaviors exactly the same, or is there some unique advantage? My argument is that there's something unique happening when I'm building a model of another person, which is that I'm leveraging my own inner simulation of things as a useful prior to try to predict what other people will do. ChatGPT models human behaviors, to draw a crude analogy, the way we would model a random object: I'm only modeling it based on seeing its behaviors in certain situations with the data I receive. On the other hand, when we model someone else's behavior, we're doing some form of projection and using the prior of how we would behave, and we probably bootstrap part of our model of human knowledge and intent based on our own introspection. I think in that way it is a difference in kind.
Fascinating. I completely agree with that. The Selfish Gene is kind of saying it doesn’t matter what you folks do. The gene is directing your behavior, and you don’t really have as much agency as you think you do. And it’s a similar thing with language. If you think of language as being a superorganism or a virus, and we are the hosts, information is being shared memetically, and it’s shaping our evolution, but it’s also shaping our behavior. So it’s almost like when we become infected by certain memes—it might be religion, for example—it’s almost like it parasitically affects our behavior.
But I think there is a difference between social memes and physical memes. Tool use, for example, doesn't seem to have the same parasitic effect. If you look at the behavioral complexity of apes, because they don't have these novel virtual memes in their culture, their behavior seems quite monolithic compared to ours. But I wondered if you could contrast that next level of mimesis.
There's been lots of great writing about the distinction in the literature, which is typically called cultural transmission, between nonhuman primates and humans. A lot of the general consensus here is that, although there is transmission among nonhuman primates, which we see in particular with tool use, it doesn't accumulate in the same way that it does with humans. In other words, humans can pass a piece of information to another generation, which that next generation will reliably copy and then merge with other new information, which they can then reliably copy. You do this over 1,000 years and you go from, “I know how to whittle a bone into a needle for sewing,” to, all of a sudden, “Now I've built a loom,” right? These ideas keep accumulating on top of each other.
Whereas in nonhuman primates, you don't see the same type of accumulation. That's what I, in a pithy way, in the book call “the singularity that already happened”: once you enable these memes or ideas to accumulate across generations, you get what you're describing as this sort of mimetic organism that we are the substrate for. For sure, what I think is interesting here is one lens through which I like to think about this: sources of learning.
If you think about how nonhuman primates learn, there are sort of 3 sources. One is that they learn from direct experience, their own actual actions. This is reinforcement learning writ large: I do something, it succeeds, it fails. Fine. Another is their own imagined actions. This is the part that evolved in early mammals. I can imagine doing 5 different strategies to try and get to the food over there, and I find the one that worked. That's a source of learning: my model of the world became a source of figuring out the right path.
What mentalizing enabled with primates is this third mechanism: learning from other people's actual actions. So I can see my mother—if I'm a young chimpanzee—using a stick to put into this termite mound, pull it out, and eat food. I don't have to do my own behaviors to do that. I don't even have to simulate doing it. By watching her do that skill, I will adopt and learn.
But what nonhuman primates don't have, which is very uniquely human, is learning from other people's imagined actions. This is the key breakthrough that happens with language: the bandwidth through which nonhuman primates can communicate what we're calling knowledge here is only through actions themselves. I can't describe, if I were a nonhuman primate, what I saw when I imagined 5 different ways to try and hunt the boar over there. I can just do it, and you can learn from what you saw me do. But language enables us to share the outcomes of our imaginations.
That is a much higher-bandwidth mechanism for translating information. That enables accumulation across generations. For example, it's so easy to think about ways in which this would be adaptive. Two would be sharing semantics: I go into the forest, and there are 2 snakes there. One bites me and I'm fine; the other bites me and I get really sick. I come back and say, “Green snakes are okay. Red snakes, don't go near them.” That semantic knowledge now exists among the whole troop.
In the old world, before there was language, only the people surrounding the event who saw it happen would have the knowledge. Now I can translate it: I simulated the episodic memory in my mind, and I translate it to everyone. The other is coordinated planning. Before language, it would not be possible for 5 humans to jump in trees and say, “Okay, here's how we're going to hunt these boar. We're going to stay silent, and then I'm going to whistle 3 times, and then we're all going to jump down and surround the one in the back.”
That type of planning is only possible because one person can simulate something and then translate it and say, “Hey, when I imagine this happening, we succeed,” and other people, of course, can edit that simulation and say, “When I imagine that happening, I don't see us succeeding for this reason.” You can start refining it. This ability to have a source of learning from other people's mental simulations is what I would argue is the source of this very unique human superpower that emerges from language.
Of course, now, with such a high-bandwidth transference of mental simulations, you do get this sort of quasi-evolutionary process, which is what Richard Dawkins is talking about. You have a process by which the memes—the ideas that do a good job propagating—are the ones that will propagate. The ones that, for whatever reasons, are not viral either don't do a good job of maintaining the host, so the ideas are bad and I end up dying, or I just don't have an incentive to share them. They're not viral; those ideas die, and so then you get this sort of meme evolution. But to me, the source is the fact that language enables us to share in our simulations, which becomes a much higher-bandwidth communication mechanism.
I'm fascinated by this idea of the meme itself being an agent, being a virtual agent, and, in expressing its agency, it needs to manipulate us. You might argue, as you do in your book, that there has to be some kind of traceable chain down to the basal ganglia. So we have many levels of bootstrapping, and at some point the thing exists because the basal ganglia says, “Oh, that's good. I like that.”
So then we have one level, then another level, then we have the memeosphere. It's almost as if that thing is manipulating us down here, but doing something completely different up there. When you have weakly emergent macroscopic phenomena, part of the definition of emergence is surprise. It's macroscopically surprising. It does something completely unexpected and unlike the thing that went below it. It's just weird that it might be manipulating us down here, but doing something completely different up there.
Yeah. I think there's probably—this is mostly fun speculation—but if I'm going to draw analogies to brain regions and intellectual features of the human brain, there are probably 2 lightweight ways we could think about why memes become attractive.
One would be the older vertebrate-like structures. This would be the basal ganglia plus the amygdala. A meme that makes me feel fear, or makes me think that unless I take an action something bad is going to happen to me, and one of those actions has to be sharing it, is going to be highly viral. If you make me afraid for my family's well-being, you're going to activate my amygdala, and even if there's only a 2% chance that this is true, I might still share it. So you get these sorts of effects. Humans are not good at dealing with low-probability, high-magnitude events, which is another brain constraint.
The other key thing that also exists at the level of early vertebrate-like structures is a preference for surprise. In order for reinforcement learning to work well, it's very effective—and we see this in AI systems—to make people pursue actions that are novel, because that's one way in which we can explore new areas and learn new actions, and explore the space of possible choices to make. This is one intrinsic way to get trial and error to work.
The way casinos make money from you is that they hack into this sort of preference for surprise. If there's a 0% chance of winning, you would never play. But if there's a 48% chance—a net 48% chance, meaning in the long run you'll lose money—but every once in a while there's some surprising thing that is actually over the threshold of being worth it to the basal ganglia because the surprise is so exciting, that's one way to think about it. So if you get something that creates innate fear, or some great outcome or surprise, you get these older structures.
With mammalian structures, I think there is sort of an active-inference play here, and you could even correlate it to the granular prefrontal cortex, with things like identity. If you give me some information that's consistent with my model of myself, and I'm highly motivated to maintain my model of myself for a variety of reasons that we can talk about, then I'm more likely to maintain this belief. Versus if you give me information that's inconsistent with my model of who I am, people are highly likely to reject these beliefs.
This is another speculative way to think about why memes persist within their little echo chambers and how it can quickly become sort of identity wars. If one's identity is consistent with a certain set of beliefs, then that almost creates a gated wall for certain types of memes to enter. It becomes much harder for certain ideas to enter my mind, and it creates a very porous filter for other types of ideas that are consistent with my identity. Those become very easy for me to adopt.
I think that's another way—if we're going to frame memes as having agency—that a meme would seek to survive: find a way to be consistent with certain people's view of themselves and the world. What you're doing is reinforcing it as opposed to challenging it.
There is a very clear difference, though: the human brain is analog, while these machine brains are digital, and there are pros and cons to each. Geoffrey Hinton talks about this—to make sure I’m citing these cool ideas correctly—and he has a great talk in which he describes it very well. The benefit of a digital brain is that it’s immortal: all the weights are stored in binary, so I can very easily transfer it to different brains, but it’s hugely energy-inefficient because I need to model everything exactly in order for it to be copyable.
The human brain is much more efficient, but it’s not copyable because the information exists in the physical representation of the analog connections between all these neurons: the actual protein receptors that exist in them, the gene expressions, and all this crazy stuff that makes it non-copyable. It’ll be interesting to see what the energy efficiency is, for example, of a digital AI system that actually attempts to recapitulate a human brain. That might be very energy-inefficient.
It might open the door for a whole new area of research that I think would be really fascinating: building analog brains. Can we have systems that actually work in a more analog way? The way they pass information to each other—this is also a Geoffrey Hinton idea—is by teaching each other. Because they are AI systems, they can teach each other with better fidelity than humans can, because they can actually share probabilistic outcomes as opposed to just the words we say.
They can also generate way more samples for each other than a human could because they can live much longer. There’s a whole emerging world around the distinction between these digital machines that are immortal but very energy-inefficient, and analog machines that are much more energy-efficient but less good at translating information.
One of your points that I think is really key is that one of the main things missing—and I don’t think it’s talked about enough—is the continual-learning problem. Maybe there’ll be a breakthrough soon that would be great, but I don’t see very clear ideas over the horizon that will solve this. I would say this is one of the essential lines that differentiates biological brains from modern AI systems.
The way in which AI systems are trained is such that we cannot let them continuously learn from new experiences because it disrupts the old information they have. Whether that is an architectural constraint or something that needs to change in the underlying learning algorithm itself, there’s lots of open research and debate about that. But the fact remains that if you allowed ChatGPT to learn from every chat that happens to it, it would get rapidly dumber.
Yeah.
That is not the case with humans. We can continuously update our information, and our representations are robust. I think for many of the applications for AI systems that are going to be most impactful, continual learning is going to be an essential component because we’re going to want to bring an AI agent in, show it new information, and immediately have it incorporate that without forgetting old things.
I think that is very clearly a line where there’s a lot of really interesting research happening and a lot of research left to be done.
Why are we superior to animals?
There’s been such a long history of us pontificating on the various chasms, or attempting to create a chasm intellectually, between us and other animals. The most famous form of this, which I think still shows threads in modernity, is from Aristotle. He took the same kind of ideas that you see in MacLean’s triune brain: other animals might have basic instincts, and they might have some form of emotions, but what they all lack—which humans uniquely have—is this notion of reason. We can uniquely reason about things in the world.
As I try to argue in the book, and as most comparative psychology demonstrates quite clearly, there are clearly forms of reasoning that we see in other animals. Of all the different abilities and capacities that seem unique to humans, the one that stands out as most salient is undeniably language. Despite many painstaking attempts, we have not even been able to teach chimpanzees, bonobos, or gorillas to speak with the same degree of fidelity as human language.
There is some controversy as to the extent to which Kanzi, Koko, and Wu passed the threshold that we define as language, so that can be debated: where do we draw the line? Undeniably, most people would agree that these nonhuman primates do not learn language naturally without painstaking attempts to teach them. When they do learn language, it does not show the same sort of flexibility as human language.
What makes language unique is 2 things. One is declarative labeling. There is a distinction between imperative labels and declarative labels. An imperative label is learning that a phrase, or a cue, leads to a reward if you take an action in response to that cue. When a dog responds to a specific cue and then you give it a treat, that’s not what we define as language; it’s an imperative label.
A declarative label is when I say “dog” and, in your head, you know that it references a concept or a thing. We have a label for a concept or a thing. It’s not at all clear that other species perform these types of declarative labeling. If they do, it surely evolved independently. We’re quite confident that early primates didn’t have this ability.
The second thing that makes language unique is grammar. We can take these declarative labels that reference things or actions, and then we can weave them together in a certain structure, and the structure itself has meaning. A basic example is just the ordering of phrases: if I say, “Ben hugged James,” that means something different from “James hugged Ben.” Despite the fact that they’re the same phrases, or the same declarative labels, the order presents meaning.
There’s a whole interesting world around why language—if it is the case that language is the fundamental difference—has allowed humans to take over the world. That’s another interesting topic we should discuss. I would argue that primarily what makes humans different is language, and Aristotle’s idea of reason we see at least in smaller forms in other animals.
Yeah. It’s quite interesting because you said right at the very beginning that Aristotle spoke about the rational soul that we have. Even in the 20th century, we spoke about things like mental time travel, our sense of self, and tool use, and it’s really interesting because we look in the animal kingdom and, one by one, all of these things that we thought placed a bright line between us and animals faded away.
Some people think that language is a continuum, that there’s just a gradation—that if you scale up the brain of an ape, you will get human language. Is that the case?
The reason I’m very skeptical of that claim is that we don’t see variance in language abilities based on brain size. Children who learn language at the age of 4 still have relatively small brains. I’m not sure of the exact brain-size comparison, but I’d be curious about the brain size of a 4-year-old child relative to an adult chimpanzee, just based on volume.
The other interesting case is Homo floresiensis.
Oh yeah, from Indonesia—the ones with the small brain.
Yeah, yeah, yeah, yeah. Yes, Homo floresiensis is a great case study here because we found fossils of ancestral humans on an island in Indonesia who were effectively miniature humans. They had shrunk in size to, I think, around 3½ to 4 feet tall.
Their brain capacity, which we can look at from their fossilized brains, had actually shrunk from that of ancestral humans. They were marginally larger than the size of a modern chimpanzee brain, and yet they showed a lot of signs of superior human intelligence despite having smaller brains. They showed tool use akin to that of ancestral humans.
They had Oldowan tools, which are supposedly a sign of uniquely human, intelligent tool-making. That is suggestive of the idea that whatever unique intellectual capacities humans gained around 2 million to 1 million years ago were present despite their shrinking brains. Either one has to argue that language evolved much earlier, which some people do, or that whatever sort of protolanguage emerged back then was present even when these brains started shrinking.
That suggests to me—and is actually aligned with the ideas in The Language Game, which is a great book—that fundamentally what’s unique is that we have an instinct to learn language. It’s not that we have some unique capacity for language, and I think that is a key difference that we can talk about.
When you look at children who learn language, there are 2 very unique features of how they go through language learning. By the young age of around 2, they’re already engaging in proto-conversations. A younger infant will pause to match the pausing of their mother. Even if they’re just babbling, they will engage in the synchrony of babbling time intervals.
That is clearly a demonstration of some initial instinct, which demonstrates the ability for me to want to engage in some turn-taking action with you.
The other unique thing that emerges a little bit later is joint attention, where human children will uniquely attempt to get their parents to engage in attention toward the same object. Scientists have gone to painstaking efforts to demonstrate that this attempt to get a parent to engage in attention toward an object is not an attempt to get the object. So a child or an infant will be dissatisfied if the parent doesn't look at the object but they get the object—so a third party comes in and hands it to them. They're dissatisfied.
If the parent looks at them and is excited when they're pointing at an object, the child will also not be satisfied. But only when the parent looks at the object, then looks back at the child and smiles, is the child satisfied. So there's this instinct to engage in conversation and to jointly attend to things, which gives us the sort of instinctual foundation on which you can start adding declarative labels. When you have joint attention to something and you're paying attention to this turn-taking, it enables you to label things and say, “Well, this means run, or this means book.”
So I think all of that makes it hard to argue that it's just a consequence of a scaled-up brain. One of the things that's really fascinating is that animal communication seems extremely superficial. And when I say superficial, I mean that when you take different populations of the same species or different species, the expression and the complexity are very, very simple. We don't see this incredible fractionation and divergence that we see in human language.
As you articulated just a minute ago, a big part of that is this declarative labeling. One of the reasons, presumably, for language is the ability to do variable binding on symbols—to say, “This thing is a dog, that thing's a bear”—and to be able to dynamically manipulate that.
It seems to me that you can think of language as a form of agentic communication. The difference between humans as language users is that we are agents, and agency is about being able to have your own directedness, plan many steps ahead, take control of your environment, and so on. The difference in communication with animals is that the information content is more in the environment around them, whereas for human languaging, a lot of it comes from the agent itself. So I just wondered whether you could think of any weird way to distinguish human languaging from animal communication.
One line that I think there's some good evidence to suggest exists between human communication and nonhuman primate communication is that humans have much more of a desire to share what's going on in our own minds. There is a unique pleasure we have from sharing our thoughts. And when we look at the communication styles that happen in nonhuman primates, there's much less of a desire—even when we go through these language-learning experiments where they have forms of communication—there's much less of a desire to share thoughts that are going on in one's mind.
One line where there's some controversy around this is that humans, from a very young age, will ask questions. They'll inquire as to what's going on in someone else's head. And with the exception of maybe Kanzi, where there was some argument that he maybe asked questions, you did not see nonhuman primates probe the minds of other individuals, even though we know they have theory of mind. We know when they're trying to deceive others or they're trying to learn actions by observation, they clearly engage in theory of mind, but when it comes to language, they weren't interested in inquiring as to what someone is thinking about.
I think in that sense there's an agency to language being a tool for inquiring as to what's going on in someone else's mind and sharing what's going on in your mind. And this is where I think language is part of why language is a superpower, because it provides a completely unique source of data for learning. Nonhuman primates can engage in learning through observation because I can see someone take actions, as you said with imitation learning. I can see you open a puzzle box to get food and I can learn from observing your actual actions.
The way I do that is because I can infer the intent of what you're trying to do, and then I can figure out which of the actions are relevant and which are irrelevant. A monkey, a chimpanzee, or an ape will ignore irrelevant actions when they observe you do a task. They've done these experiments with humans and chimpanzees where they do all these actions to open a puzzle box, including some random actions, and chimpanzees will ignore the random actions, which suggests they can infer the intent of it. Which is great, but chimpanzees don't learn from what's going on in your head.
The ability to learn from other mental simulations is what's so powerful about language. I can say, “I just went over to that forest over there and I saw a red snake and a blue snake, and I saw that the red snake is really dangerous, but the blue snake is not because the blue snake bit me and nothing happened.” And so I share that episodic memory, and now everyone has that knowledge, even though that was just in my mental simulation.
Or, when planning a hunt, a group of 5 humans—I can imagine a strategy of how all 5 of us are going to coordinate, see it succeed in my mind, and then share the results and the plan with everyone. So language enables us to tether our mental simulations to each other. And I think there's a sense of agency in the idea that there's a purpose to that. There's a volitional purpose to the communication.
The neurological underpinnings of communication that occurs in nonhuman primates are more analogous to our emotional expressions than they are to language. And we see this also in the brain. Monkeys and nonhuman apes have these innate expressions that they do, which are genetically hard-coded, and we know that because it's the same even across species, often, that have never interacted with each other. It comes from neurological structures similar to our laughing and crying.
So this is clearly a hard-coded emotional expression. In that sense, it doesn't have the same volition, because I'm not doing this action to communicate a concept to you. I'm doing this action as an innate response to a cue or a feeling I have. So, yeah, I think there's some meat to that idea.
Well, a few things to explore there. First of all, we should just talk about how we became a collective intelligence after the fact. I'm not sure whether that's unique to humans if you look at other forms of collective intelligence. There's always a kind of juxtaposition between the intelligence of the individual versus the collective, and usually you find that having very intelligent individuals is not good for the intelligence of the collective.
But what's interesting about humans is that we clearly didn't evolve as a collective intelligence. We had this kind of bootstrapping process where we were very, very useful, independent agents, and then this collective intelligence just emerged out of nowhere. Is that an interesting observation?
There are different degrees of collectiveness, and so I think we can draw distinctions between different flavors of collectiveness, but I don't think humans are uniquely collective. For example, the imitation learning of nonhuman primates is a form of collective intelligence because you can teach one member of a chimpanzee troop how to use a tool, and then over time the rest of the troop will learn just through observation. So that's a sense of collective intelligence.
Many vertebrates, and likely the first vertebrates—you can even see fish—will learn through observation. In other words, when a fish swims in a certain direction to get food, other fish can see that fish do that and follow them. There can be an instinct to follow others around you. So I think there are flavors of collectiveness that exist across many different species.
But what's unique about the collectiveness in humans is that the fidelity with which we transfer our mental simulations enables them to accumulate across generations. In that sense, it almost has its own agency, or is its own thing, because it can actually go through its own process of evolution as ideas propagate through generations of people. That's not the same thing that you see in other animals.
Yeah. A couple of things on that. First of all, I would quite like to distinguish knowledge and intelligence. Collective intelligence—and intelligence in general—is a process of discovering models. To get the language down here, I will use models, skills, and knowledge pretty much interchangeably. I think of an intelligent process as epistemic foraging: finding interesting models and then having them discovered and shared by other people.
It's a little bit like when you distribute a GPU workload: you can do model parallelism and you can do data parallelism. You can either split up the processing, or you can split up the actual representation. So I think the kind of collective intelligence that you've just been speaking about is that we've got all of these independent agents, and they are finding models and sharing models; the models get refined over time, and it adapts, and so on.
But I also think a big, important element is sharing the computation. So even though there's some redundant work going on, epistemic areas over here are being explored, but also, in many cases, the same problems are being explored in slightly different variations. So we're sharing the workload with other humans.
Yes, I think that totally makes sense. There are some interesting ideas in AI here, actually, where there's this concept of knowledge distillation. In AI, one way in which you can have model A teach model B the things that model A knows is to wholesale copy the parameters of model A. Of course, that's totally biologically implausible. There are aspects of parameter copying—the components of our brain that are genetically hard-coded are a version of parameter copying—but for other applications, it's not feasible or maybe not desirable to just copy parameters.
Knowledge distillation is saying, okay, well, we can have a set of data that we give model A, and we either look at the outputs of model A or the layer before the outputs, so we can see more richness in its representation of the input you give it. Then we take that data—that almost-labeled data—to model B and train model B on it. So that's distilling some of the knowledge, through almost training model B to try and act similarly to model A.
That type of information transfer, I think, does occur in nonhuman primates, and that's imitation learning. However, it is not nearly as rich, because what happens in nonhuman primates is primarily grounded in just the actions that I'm taking. That's much less rich than what I can share: not only the data of what you see me actually do, but also things that happen only in my mind. That opens the door for much more transference of—and the word you used—the computations that I'm performing.
So I think it's absolutely true. Before, we were learning in the physical world, so we were learning from physical things that we were directly observing, and now we are learning from imagined actions. But there's a bit of a latent component to language as well. For example, someone might come up to me and say, “Oh, the blue swirly thing is over there.” And I'll say, “Well, I don't know what you mean by the blue swirly thing, because I've never seen one before.”
There's this inference process. This is where it starts to get really interesting, because there's a diffusion, right? There's a message passing that happens between all of the different agents, and it's filling in missing information. Even though many of the agents wouldn't have seen anything like what we're talking about, sometimes it can be filled in with subsequent interactions with people, and sometimes it can just become a kind of latent category that can be filled in later. So there's this real diffusion process going on, which I think is quite difficult to articulate.
Part of what's so interesting about language is that it's still an area of such controversy amongst cognitive psychologists, linguists, and even AI people. So much is still unsettled about it, and there are still debates today. There are debates today about whether language is primarily a tool for thinking or communication. Chomsky is the most famous proponent of the idea of language for thinking.
He has evolutionary arguments that language initially evolved not as a tool for communication, but for our own process of thinking, and then later was adopted or used for communication. That's a minority view. Other people argue—and I'm more amenable to this—that language was primarily used as a tool for communicating.
These ideas are actually reemerging with language models, because the way language models learn about the world, in some sense, is that language becomes the reasoning tool itself, which is more Chomsky-like. Even though I think the success of language models, in a lot of ways, discredits many of Chomsky's ideas, we can talk about that. Interestingly, the fact that we're using language as the fundamental mechanism for reasoning and thinking is actually somewhat Chomsky-like, versus the idea that language is communication.
The idea is that language is a condensed set of tokens that I'm passing between minds, but the real communication I'm trying to share with you is what's going on in my mind. In other words, it's the mental simulation—the more mammalian component here. The rendered 3D world is what I'm trying to transfer to you, and I condense it into this code that you reverse-engineer back into a mental simulation.
Theory of mind—one reason why language might be so rare in the animal kingdom is that mentalizing, or theory of mind, which is relatively rare in the animal kingdom, is a prerequisite. In order for me to reverse-engineer the language code you've provided me, I need to be able to infer what you might have meant by what you're saying, reason about why you would have said this, and understand what knowledge you have.
So I think language is intended to cue another person to render something in their mind. This is also where teaching is so important and such a key aspect of language learning, because we can infer what declarative labels this person is aware of. When they're confused, you have to start trying to iterate to understand what they're confused about in what you're saying so that you can disambiguate it for them.
There's also a disambiguation process where you ask follow-up questions when you feel like you don't fully understand what's going on in someone else's head.
Yeah. I mean, the guardrails thing is interesting because they're not necessarily thinking guardrails; they're also pragmatic guardrails. And there's a really interesting figure in the book, actually. Yeah, here it is. It talks about how language is sharing information over generations.
Without language, we learn a little bit inside a generation, then it goes pretty much back to zero again. But now we have the ability to pass on these memetic bits of information over several generations. The thing is, there's a real structure to it. I think of it as a bit like a directed acyclic graph. So it's a tree structure, and every single bit of knowledge that we discover kind of stands on the shoulders of giants. It needs all of the things that we discovered beforehand.
In a sense, we're all of these little agents, and we're doing this epistemic foraging. We're finding new skill programs, we're sharing them, and so on. But it's almost like we shouldn't think of the mass as being like an entire convex hull. It's only on the boundary where all of the creativity and all of the information sharing happens—on the surface of this object that's being created.
What I mean by that is, now in modern cities, for example, you can't live without a driver's license. You can't live without the internet. You need to do things a certain way, and even though it's not technically constraining our brains and how we think, we live in a very, very constrained and weird world now.
Yeah, totally. Great point. There's a biological constraint as to how much knowledge a given human brain can contain. One lens through which to see the last 100,000 years, especially the last 100 years, is us finding solutions to getting past the biological constraint of human brains.
Language was one tool, because it used to be the case that all the information that a given entity learned needed to be learned by my brain within my lifetime. Language enables us, as a group, to have shared knowledge, but not every brain contains all of the knowledge. If you think about a troop of 100 people, it's possible for those 100 people and all their descendants for 1,000 years to have tons of skills, despite the fact that no one brain ever had all of the skills.
Someone becomes really good at hunting, someone becomes really good at weaving animal skins into clothing, and all of these types of skills. There are actually cases in anthropology of groups of humans that get separated from each other and their technology degrades, because there is a limit—a minimum number of brains needed to contain and store a certain amount of information in the absence of writing.
Language was maybe innovation 1 here. Writing was another innovation, which is great. Now we can more reliably transfer these ideas across generations, even if there are gaps. In other words, even if there's a period of time for maybe 2 generations when no brains contain it, a third generation can go back to the writing and pick up that knowledge.
Of course, now with the internet, we've just scaled up writing even more. But you're absolutely right. Sometimes I think about this as, if a group of 20 friends and I ended up on an island—if we were the only 20 humans left—not that I think about this all the time, but it is crazy how little of human knowledge would be contained in our 20 brains.
How dramatically we would degrade, essentially. We've got this thing where we've got all of these different brains, and individual people can have about 150 friends or something. It's the social Dunbar limit.
But as you say, because we have this ability to share simulations and we have common myths and so on, we can address a much larger carrying capacity of people and knowledge. You said something really interesting in the book, which is four things: bigger brains, specialization, more brains, bigger population size, and writing and sharing simulations through the internet and all of these things.
So we've increased our carrying capacity, and now something very interesting and arbitrary has emerged. We've got all of these different specializations of skills, and I guess the question is: where does it end? Has it converged? Could we carry much more knowledge than we already have, or would we have to wait for a top-down kind of genetic pressure for our brains to get a bit bigger?
I think we are about to go through this. Google and the internet have turned us all into epistemic hybrids. Google has become a shared knowledge store that we all use, and of course there are problems, because now there are subareas on the internet where we can use different knowledge stores. We live in these different epistemic bubbles, and that creates political problems as well.
But we have already become hybrids where we use technology to overcome limitations in our own brains. Writing is a tool to overcome challenges in memory and, at times, thinking. The internet has become a tool to answer any question at a whim, and some people have concerns with this because it can also atrophy parts of our brain that maybe we want.
For example, through mere introspection, I will say that once I started using Google Maps as a kid, the part of my brain that was learning how to navigate a city—by actually remembering the grid and map of a city—just started atrophying. Now I have no capacity to do that, whereas my dad—you take him to any new city and you can see him rendering a map of the city in his mind—won't use Google Maps.
One could argue that it doesn't matter because I'll always have Google Maps, so why do I need this skill? Another argument would be that atrophying may have other consequences in my life, and it would be important for me to go through the cognitive exercise even though technology enables me to do it. We make these trade-offs at different times. Why do we teach kids arithmetic? They can always just use a calculator, but we deem it important for us to go through the process of understanding arithmetic even though technology can already do a better job for us.
This is a new frontier with using large language models, and there are some really cool things happening with education. In Khan Academy, for example, they're working on building language models to help children go through reasoning steps, which is a really cool application. Instead of just asking the question, it will probe the student to go through a process so they can come to the conclusion themselves.
There's a pessimistic and optimistic world here. An optimistic world is that these new AI systems are actually going to be a new step forward in cyborgizing ourselves, but it's not necessarily going to be as atrophying as something like Google. These systems won't only give us the dopamine hit of a factual answer; they'll also guide us toward better understanding how they came to their conclusion, to ensure that we understand when we're probing and asking questions. That's an optimistic state of the world.
The pessimistic state of the world would be that we offload more and more of our own cognitive reasoning to these systems, and we become even more atrophied in these abilities. That might not be a good world to live in if we keep offloading more and more reasoning to systems and lose the ability to do it well ourselves.
Yeah. I've been thinking about this a lot recently. I was involved in a startup that did transcription, language models, and augmented-reality glasses. The idea was that you could be in a lecture—and I still think this is very useful for people with accessibility concerns, like those who are hard of hearing—but we were thinking of it as something that could augment your cognition.
You're in a lecture, and now you don't need to pay attention to the lecturer because you're transcribing it and GPT is making notes for you. I think this is really wrong. But you give a counterexample with satnav: we don't need to read maps anymore because we can externalize that cognition.
I feel like this is different. You're in a lecture, and all of these AI language tools are a form of understanding procrastination, right? Understanding, or intelligence, is the process of creating a model. You're creating a simulation, and in order to create a simulation, you actually have to think. You normally think, externalize the thinking a bit, do some writing, and pay attention.
Here's the thing: in that situation, there are so many more cues because it's in 4D. You can hear things, you can see things, and it's a social and physical activity. Even the dance—the performativity of the lecturer—is all information. It helps you understand.
Now I'm transcribing the thing, and people say, “It's okay. I can just read the transcription later and understand it.” Well, maybe, but you're already at a disadvantage. You probably won't, because this procrastination is just paying it down the line. You're saying, “I might do it later. I might do it later,” and you never will. That's going to create a society of automatons that just don't think for themselves.
Yeah. I'm torn between the optimistic and pessimistic states of the future, but I think there's a very good argument behind what you're saying, so I definitely don't reject it out of hand.
I actually really liked the analogy to model-based versus model-free that you were suggesting there, because that applies very well to Google Maps. When my dad navigates a city, he has a model of the world, and he's engaging in model-based planning of how to get somewhere. When I use Google Maps, I've externalized the model, and all I do is respond to the cue of when to turn right or left.
I think that is absolutely a good way to think about this: we use technology to externalize building models, which can sometimes make things more efficient because then we can just be model-free actors. But there are places, like in the example you're suggesting, where we really want people to engage in the more painful, hard process of building models of things. In those cases, obviously it's dangerous to make it so easy to externalize these models.
Yeah. It's hard to articulate. I think part of it is a kind of acquiescence. You're sequestering your agency when you externalize too much of your cognition, particularly if it's parts of your cognition that are useful because they contain core knowledge that will generalize and help you acquire new knowledge, or if it's the portability of discovering knowledge. It's your intelligence, and you're not exercising that muscle.
You become acquiescent and then you become less of an agent. From a collective-intelligence point of view, we're just saying that language and intelligence are about discovering knowledge. If we are all sequestering our agency and becoming less intelligent as individuals, as a collective, maybe we will suffer.
But it's one of those things that's so easy for us now to make grand statements about. People in 200 years will look back on this and laugh and say, “It's a little bit like when they introduced bicycles.” There was apparently a moral panic because they said, “Women will start cheating on their husbands and using bicycles to go to the next town.”
That is an interesting fact. Yeah. Well, I think history is such a good tool when trying to reason about how people in the future will think about us. We are the people in the future to the past, which is obvious, but it's a useful tool.
For example, in some sense we already live in this dystopian world when it comes to physical exercise. Roll back the clock 500 years, and most people didn't have to think about physical exercise as much because most work required physical exercise. We exercised with our work, and so much of the work—at least in the developed world—is information-related, where we don't exercise and so we go to the gym.
The gym is a weird thing. If aliens came down and observed gyms, it would be anthropologically very bizarre behavior, because we just go into a room and run on treadmills. We do it because we've evolved to require exercise, and modernity has removed exercise as a prerequisite to most of the things that we need in life.
Now there's this gaping hole, and what we do is go to the gym and run in place to satiate this physical need. You could imagine—and one might interpret this as dystopian or utopian—a world where we've offloaded so much cognition, but because humans need to think about things, or because as a society we value it the same way we value physical fitness, there are now social pressures to go to these intellectual gyms.
Just to make sure, even though you don't need to do it for work or it's not necessary for the world to function, we feel like there's value in a human who knows how to reason about things. So we go to intellectual gyms for that. We might—I don't know if that's a utopian or dystopian future—but however we feel about it, I would venture to guess people 500 years ago who looked at a treadmill would probably feel similarly.
100%. Well, MLST is my intellectual gym, by the way. [Laughter] You spoke about DNA. Dawkins, of course, wrote the book The Selfish Gene, and you said that the value of DNA was not what it creates—it creates hearts and lungs and so on—but what it enables, which is this evolutionary process. But then it gets to this concept of what we mean by a meme in general.
You said that it's an idea or behavior that spreads contagiously. How do you think about memes?
Well, I think Dawkins did a wonderful job articulating this idea in a way that's really understandable. A meme is a concept or a behavior. A meme can be just the idea that individuals should have rights, or the idea of equality, or something sillier, such as the idea that we shake hands before we sit down for a meeting. And these things, because humans can share simulations through language and we engage in imitation learning, these ideas or behaviors propagate throughout societies.
And because these things are propagating, a different form—not evolution in the sense of genetic evolution, but a form of evolution—emerges, because some ideas will propagate better than others. By nature of that process unfolding, memes—these concepts or behaviors—actually go through an evolutionary process. Ideas that are either viral because people want to share them with each other, or ideas that somehow support the survival of the individuals that hold them, are going to be ideas that propagate correctly.
Ideas that negatively affect the survival of the individuals that hold them, or that people do not desire to share for whatever reason, are going to do a worse job propagating. And so it's a really almost brilliant lens to look at human culture when you reframe cultural ideas and concepts as memes—a different take on genes—that go through their own sort of process of iteration. This is not my idea; this is Richard Dawkins's.
Oh, yeah. Well, we can thank Richard very much for this. I'm fascinated with memes, and I kind of think of language as being a collection of memes. But now we're in this very, very interesting space. Before language, we learned by observing physical skills performed by other people, and we could imitate them and so on. Now we are sharing simulations, basically, without actually needing to see the thing, and that means that we are one step removed from reality.
So all sorts of memes have cropped up, and some of them are better described, as you say in your book, as shared delusions, but they have some utility as well. When we have a common myth, for example, it might be a religion, it might be a nation-state; it allows us to cooperate with each other in a way that we wouldn't be able to do before. And you actually cited some ideas by John Searle and Yuval Noah Harari in his book Sapiens on that. [Snorts]
Yeah, they've all famously popularized this idea. But John Searle was one of the original ideators of it. What's so powerful about these shared fictions is they can propagate much more easily than a human can talk to everyone in a group. And so, because they propagate much more easily and with very high fidelity, this enables me to meet someone who is a New Yorker whom I have never met before and immediately have shared views.
We probably both believe in individual rights. We probably both believe that money can be used for transacting things. So, if I give them a dollar, they'll believe that the dollar will be used elsewhere, or they can give me a dollar. And of course, today it's hard to reason about these things because there are so many rules in place that you don't realize it's all a shared fiction.
The reason we think we believe in money is because we're like, “Well, I know that all the other stores I go to will take this money.” So that's the reason it works. But why do they all take the money? It's all this shared belief that we all trust that this thing will be used for transacting. And so, because of that, it enables really large groups of people to coordinate, and that is a very powerful aspect of language.
But the argument I make in the book is that, similar to how genes are powerful not because of the structures they create but because they enable a process of evolution by which good structures will emerge, language is similar in that sense. What's powerful about language per se is not that we can engage in these shared simulations for coordination; it's that language enables the propagation of ideas and concepts across generations, which will thereby undergo its own evolutionary process. So, of course, these good ideas that enable survival are going to emerge. And that's really what's so powerful about language.
I guess the arbitrariness is quite interesting. Some of them, on the surface, don't seem like good ideas; they just seem like really bad ideas. And I guess you can think about it in terms of creativity as well. For a meme to be established in the sphere of possible memes, does it need to have intrinsic value? Possibly not, because we're getting into creativity: is it novelty? Does it have intrinsic value? Is it just social proof? Is the meme only existing because lots of people have been fooled into thinking it has value? So it's kind of extrinsic value via social proof.
And then there's almost a double entendre with the meme, or a deeper meme meaning, because you talk about altruism. The meme itself might actually be quite a stupid meme, but if it causes altruism, so there's actually a group-selection advantage to it, then it's almost like that's the lens of analysis to understand how good the meme is.
Yeah. It's a really fun area of literature to read through because there's still zero consensus as to how language evolved. One reason why it's so controversial is the way in which we disambiguate—and I'll get to your question—the way we disambiguate evolutionary arguments is typically by observing gradation in extant, or currently present, animals. That enables us to observe these intermediary steps between a morphological aspect of body A and a morphological aspect of body B.
The problem with language is we have nonhuman primates that, for the most part, don't have any language, and then we have humans that have very complex language. And all of the intermediary humans that existed between our divergence with chimpanzees about 6 million years ago and our divergence with all other modern humans between 50,000 and 100,000 years ago—we don't have them; they're all dead. All those lineages are lost.
And so that means that there's this broad spectrum of arguments that could be made. Chomsky argues—I find this a very strong claim and thus hard to defend—that it happened all at once, or very rapidly: there was no language, then all of a sudden there was language.
Right, and there are other arguments that it was a gradual process.
But one of the most controversial aspects of language evolution goes to what you're talking about, which is evolutionary arguments for why language evolved. Evolutionary arguments for why language evolved have a harder burden of proof than arguments for other adaptations.
So when we argue about the evolutionary benefit of something like theory of mind, there are no complex evolutionary machinations one needs to conceive of to defend it, because you can see why it would be beneficial for an individual chimpanzee to be born with the ability to infer what's going on in other people's heads. They can better defend themselves when someone is going to be mean, better figure out whom to trust, better climb the social hierarchy, et cetera.
But with language, unless you take the Chomsky view that its primary adaptation is for thinking, the argument that language evolved for communication is more challenging, because it's not valuable for an individual human to be born with a little bit of language skill unless other humans are also engaging in language. And so this then means that the only benefit is if we're both sharing truly useful information with each other.
Although it seems intuitive that the way this would function is that a group of humans that are sharing knowledge with each other is going to survive better than another group of humans that's not, and that's how evolution will ensue, this is actually quite controversial in evolutionary biology because that's invoking something called group selection.
Now, some people call the modern incarnation of this multilevel selection, where there's some consensus that, yes, there are group-level effects that can impact things. But most people think that group-level effects are not nearly as strong as we would intuitively think.
And the issue is the following. If you have a group of 100 humans that use language with each other and then you have 1 human born who is just going to try to trick all the others, so all they're going to do is use language just to be disingenuous, it's not at all clear that that human would be at a disadvantage. In fact, they might be at an advantage relative to everyone else.
If you play that forward over time, language will be lost because someone born who isn't going to be tricked by the individual trying to lie to them with language is actually going to survive better than the people who have language skills. There has been so much debate throughout evolutionary linguistics about these arguments as to how language evolved. I like the argument in a great book called The Evolution of Language by Fitch, and I think he makes a really great argument around how you could think about this occurring.
A lot of people argue that it probably started with something called reciprocal altruism. The way altruism exists in the animal kingdom, there are 2 forms of accepted altruism. One is something called kin selection, which is quite straightforward: I'm willing to sacrifice something—in other words, share something with an individual—if I share genes with them. That's easy.
Reciprocal altruism, which we do see in the animal kingdom, is, “I'll scratch your back if you scratch my back.” But if you start not scratching my back, then I'm going to stop scratching your back. What this suggests is that, in order for language to be stable—in other words, for it to be beneficial for me to truthfully share information—there need to be costs to me lying.
This is one argument that people speculate is one reason why humans have such strong moral preferences toward punishing liars and out-groups and in-groups, because we really try to identify individuals who are lying. Robin Dunbar has a beautiful argument that this is why gossip evolved. One way that evolution can stabilize the use of language is by virtue of us having a preference to share moral violations.
Gossip is a tool of language where, if you see someone lie or cheat and you share it with a bunch of other individuals, that becomes a huge cost to someone lying and cheating, because if 1 person catches them, then the whole group is aware of it. There is this special feedback loop that happens where language skills require more punishment of violations to be a stable strategy. One way you get that is by having more gossip and making sure there are higher costs to defecting.
This is not by any means the only story of language evolution, but it's one with a lot of interesting evidence behind it. There are some people who argue that the feedback loop—1 emerging idea, which I don't talk about in the book but which I do think is interesting—is actually one in which we try to detect lying in others. They make the counterargument to me, saying that the effect of lying is the loss of language.
There is an argument that you get the reverse: you get a really good theory of mind in humans because we're so sensitive to trying to detect people who are actually giving us false information. There's still a lot of controversy around it, but the main takeaway is that the blanket group-level selection argument—that language is obviously beneficial because once a group has language, they're all going to survive better—is not a sufficient argument for language evolution.
You need a more nuanced evolutionary argument as to why it's a stable strategy for an individual to be born with superior language skills, or you have to argue that language did not evolve primarily for communication.
You know, it's quite interesting, first of all, that you were writing this book actually a couple of years ago. So this was before GPT-4, although you did put a note in about GPT-4, and you were speaking about Blake Lemoine. He was a Google engineer, and he famously came out convinced that these things had developed sentience.
I think much of this actually hinges on the concept of a world model. One view of language models is that they're just modeling a statistical distribution of tokens, and that seems quite low-resolution. Another take is that they're learning a world model. What that means is that, rather than just capturing the state of language, they're actually simulators. They're generating the underlying processes of language. They're capturing the dynamics of language.
How do you say that something is or is not sentient, especially given that the models could potentially be such high resolution that they are generating the same thing for all intents and purposes?
This is where I don't see myself as a philosopher, but this is where I do think scientists need to include philosophers. When questions become nonscientific, I think the scientific instinct is to argue that we don't draw distinctions between things that the scientific method can't draw a distinction between. But the problem is that there might be moral differences between them.
For example, it might be scientifically impossible for us to differentiate which of 2 systems that look indistinguishable in their inputs and outputs is sentient. Scientifically, we might say, “Well, because we can't differentiate the 2, we're going to say they're the same.” But that doesn't mean they're the same. That just means that, because we have no methodology for drawing a distinction between them, from a scientific perspective we're not going to draw a distinction because we're entering philosophical territory.
But if you take that and then start talking about policy implications, the actual values we attribute to them, and how we introduce these things to society, I think we need to include a philosophy lens here. It might not actually be the case that they're the same just because we can't distinguish them. So that's just 1 thought.
Another thought on world models. One distinction I want to draw, because I've seen a lot of confusion on the internet about the world-model dilemma, is that there's a difference between a world model and a model. It is undeniable that language models have a model. In order for GPT-4 to correctly predict the next token in these really complicated language questions, it clearly has some model of something.
Because we can ask it common-sense questions about the world and it answers many of them correctly, you can say this is a model of aspects of our world without question. I think it would be very hard to argue that that's not the case if you look at GPT-4's performance on many of these questions. But what most people mean when they say “world model” is a specific process of simulating an ordered sequence of states and the consequences of different actions.
That means identifying the end result of these actions in your head. Another way to think about what we mean by a world model is the ability to reason about interventions and causality. This is the Judea Pearl argument: with our world model, we can hypothesis-test.
I can say, “I imagine that if I do this thing in the world, I think this will be the consequence of it because that's what I see in my head.” Now I have a hypothesis. Now I'm going to actually do that thing in the world and see if my hypothesis is correct.
That's very different from what's happening in a language model, where its understanding of the world derives solely from its input data. In a world model, my understanding of the world comes from the delta—the difference between what I hypothesize is going to happen in the world and my actual experience of it.
This distinction really matters the more we're going to start offloading our cognition to these systems. For example, everything that ChatGPT knows is on the basis of its input data. That means if false information or wrong information is in the input data, ChatGPT is going to know that information. There's absolutely no hypothesis-testing embedded into ChatGPT, unlike our true AGI agent that will one day be invented.
What it would do is hypothesize aspects of the world and test its own hypotheses. If you give it false information, if it reads articles about how the Earth is flat, it's not going to just start talking about how the Earth is flat. It's going to say, “Okay, well, that is incongruent with my model of the world. I'm going to now run some tests where I can differentiate between them, and I'm going to perform those tests and then conclude that the world is not flat.”
When one says that ChatGPT does not have a world model, I think some people misinterpret that as suggesting that it's just dumbly looking at the statistics. That's not at all what we're saying. In order to correctly look at the statistics of language, clearly it's built up a very rich and complex model of the text that it's seeing, and that's how it's able to predict the next word so well. But it's not what most people mean when we say “world model.”
A couple of things on that. I think people conflate the machinations of language models with how we represent them statistically or abstractly, because if you look at a lot of papers, they actually represent it like a probability—a joint probability distribution. Of course, the way that language models work is completely different from that, but you're bringing in some very interesting things.
So first of all, we are agents in the world. The agential lens is quite interesting: we interact with the world, so we're not just learning from observational data. I was talking with Nick Chater, and we said, “Why is it that in our everyday experience, we experience the world in 4D color?” He said it's because it's interactive.
In your experience, you can actually seek new information, right? You can move your eyes, you can get new information in, and you can touch things. When you're doing future or past simulations, you don't have that interactivity. So there's something about interactivity that's really important.
But even then, how far could you go? A complete one-to-one simulacrum of the world wouldn't be a particularly good model. In physics, there is no causality, right? It's just dynamics. Causality is actually something that emerges very, very far up. So we're talking about a model that's an approximation of the real world, which may or may not include causality.
It probably would, because it's an interactive model and it has this kind of agential map. But I guess we're just drawing the line somewhere and saying, “Well, that is a world model.” Max Bennett
I'm actually not sure whether, even if we rendered a perfect 3D map of every particle in the universe—
And that was the input data to some infinitely large model. I still would argue that it's learning something different from a model that's given some form of agency, where it can hypothesize rules and then test its own rules.
Now, given infinite time, it's possible that those will converge, because given infinite time, every possible hypothesis I could conceive of will end up showing up in the training data. So eventually I'll see the training data of every possible experiment I could run. If time is infinite, I guess you could suppose that happens.
But what's so different is this dramatic dimensionality reduction that happens when you show me something uncertain and then I can conceive of specifically the tests I want to run to map the uncertain thing to my mental model of the world. That's a very different way of learning about things. It's not just input data and then self-supervising on predicting one's own input data. It's building a model in which I can simulate possible outcomes and then hypothesis-test those outcomes.
Even this is not uniquely human at all. If you look at the way a rat would deal with something novel in its environment, it's drawn to the novel thing and explores it until it feels like it understands it. Then it will move away.
When you show a child an object that's perplexing, they will touch it, turn it around, and try to understand it until they feel like they've built a model of it. That simple act is doing something very different from the self-supervision we see in most AI models today, because I see something I'm uncertain about and I'm volitionally going to create new training data for myself.
I know the training data I want now. I want to see what happens when I pick it up, turn it to the left, and turn it to the top. A convolutional neural network doesn't do that. So the way we teach CNNs to understand rotations in 3D objects is by manipulating the training data ourselves. We take imagery and rotate it in a bunch of different ways, so we're the ones curating the data set to teach it these things.
Yeah, but that's different from the way we learn about things. I think this is a key aspect that's missing from AI systems today, and it's something that folks are working on. That's something we're going to have to add in.
Yeah, completely agree. It feels like you're saying basically what I think, which is that there's a creativity and an agency gap. A lot of that is because we're agents, as you say: we create our own training data, we do this active inference and sense-making, and we build these models in real time. As a collective intelligence, it creates a kind of divergent search process for knowledge. It's this epistemic foraging that we spoke about.
GPT is a monolithic model. It does have models, but the models are only learned at training time. Inference actually happens at run time, and when you put a prompt into GPT, you're just retrieving one of the models that was already learned a long time ago. It's not creating a new model in the moment. So it creates this kind of sclerotic system rather than the divergent, creative system that we experience in biomimetic intelligence.
One sort of mental model I have of this—because there's so much debate around it—is something I'd be curious to put through the gauntlet of what other people think about. Maybe people in the comments will either agree or disagree with this. I think an interesting alternative experiment, or eval, of an AI model that I haven't heard before is this: if you give it knowingly false information in the training data—not at inference time, but in the training data—will it reject it wholesale?
To me, this is the distinction: an agent that can hypothesis-test and intervene in the world will reject false information. If you tell it that the world is flat, it will know that the world is not flat. Whereas with GPT, any data you give it is given equal weight to every other piece of data. The only reason it would reject that is if there's other data in the training set that it's going to ignore.
It's almost cheating, because by definition we know a language model is going to fail at this task. The only way you can fix it is if you give it other data in the training set. There's no notion of hypothesis-testing compared with an agent. The only way you could get it to be wrong is if you manipulate its sensors during the actual hypothesis-testing that it does.
You can, of course, manipulate it by changing the actual test when it does these tests. That happens in The Three-Body Problem, which is an amazing book where aliens manipulate our experiments. Anyway, I think that's another way to evaluate these systems: can it figure out that you're giving it false information and reject it?
Yeah, 100%. A couple of things, though. There's something magic about having agentic density in the system, right? When you have something like GPT, just to make it statistically tractable, it's generally doing a kind of low-entropy search. What I mean by that is it's just looking for the baseline patterns. It's not doing a lot of exploration, and it's not searching outside of the main sources of statistical regularity.
Whereas when you have divergence in the search process, with all of these individual agents doing their own things, as a system it's much more of a high-entropy search. That means you're actually bringing in lots of new information to solve problems in creative and interesting ways.
In the physical world, though, it's quite interesting, right? The problems come from the physical world. The trees get big, so giraffes need to have a long neck in order to eat the leaves from the trees. This whole thing just rinses and repeats. The environment produces novel solutions, and then we see this divergence and find novel, creative solutions to the problems that get generated.
But in the memetic sphere, it's so much more difficult than that, right? The problems and the guardrails aren't constrained in the way that they are in the physical world. For example, we have capitalism or we have nation-states, and there’s all kinds of interesting divergence going in different directions. But it doesn't seem like there are the same pressures that ground the thing in reality.
Well, yeah, I think it's definitely not grounded in truth. Its tethering to truth is this: knowingly false information that leads me to take actions that will hurt my survival will fade. But false information that helps me survive better will propagate freely, or is at least neutral.
Another way this shows up is—and this is where I'll go into some pontificating—
Please.
But where I think there is memetic evolution that can drive us away even from things like happiness. If we think about what systems of coordination survive, they're systems of domination and militarism.
If you take 2 groups of individuals, let's say one is really happy and calm, sees no desire for domination, and does not attempt to innovate and build more technology. The other is unhappy but super aggressive, wants power, and wants to expand. These ideas will die out.
What this suggests—and I think I talked about this in one of our previous conversations—is the importance of delineating, in my view, the Darwinian component of what does survive from the moral component of what is right or wrong. It is definitely not the case that what survives is definitionally right. It's absolutely possible that the things that survive and do well evolutionarily are not the things that we feel are morally aligned.
That is not to propose a correct or incorrect system, but it is an important distinction to draw when we're trying to decide what we deem to be morally right or wrong.
So I think that's just one example of what you're saying: the ideas that propagate successfully might not be the ones that are true. They might also not be the ones that we deem to be moral, or even be the ones that lead to human happiness. They're just the ones that do a good job of keeping humans alive and reproducing the idea.
Yeah. So people say that language models confabulate and don't preserve epistemic factfulness. But you could also argue the same thing about us, right? We actually confabulate everything. We don't really have goals. We just generate these post hoc confabulations, then explain our behavior and pretend that was what we wanted to do, that we had beliefs, and so on. We just make it up as we go along using this kind of active inference.
Even though we are emotional and subjective, and we believe in religion and lots of things that we presumably made up, we have Wikipedia. We have objectivity, even though it's an illusion, right? There's no such thing; even general relativity isn't as objective as we think. If you keep asking why and why and why, it just disintegrates into incoherence. But there seems to be some objective structure that is preserved. How is that explained, given that our brain simulations don't seem to select for truthfulness?
I think the question of whether humans are better or worse than ChatGPT is almost a red herring. I look at ChatGPT as an alien—it's like an alien brain. There are certain things it does that are clearly better than us. Information retrieval in ChatGPT blows a human away, without question. In many ways, it's way better than humans.
But there are certain things that human brains do that ChatGPT does not. If we're trying to build human-like intelligence, there's certain inspiration we can garner from human brains. I think there's a component of our model-based rendering of a plan and then executing that plan that has a level of explainability that's unique relative to a system that is just iteratively predicting the next token.
But we also do the same thing that ChatGPT does. When we make model-free choices and then you say, “Why did you do that thing?” what we engage in is exactly as you're describing: a post hoc explanation. I didn't render a plan; I was just walking down the street. If you say, “Why did you move your foot there as opposed to 2 inches to the right?” what I'm going to do is render a post hoc explanation of why I did that.
But I didn't really think about it; I'm just explaining it after the fact. So it's definitely the case that humans have that component. But it means there's also another component, which I would argue is unique and important: our ability to pause, render a plan, and then execute against that plan. The key thing that I think is the dividing line between these models and us is the ability to render hypotheses and make interventions in the world. That's the key thing.
And so it's not the case that our brain has the true objective state of the world in our head. I don't think that's there. There might be components of objective truth in ChatGPT that it contains that we don't have, and I think in its information retrieval it has probably, in some ways, more aspects of reality than I do, in terms of having read all of Wikipedia and answering questions about biology that I don't even know the answers to. But there are also components of the world that the human brain has rendered and contains that ChatGPT does not, because of our ability to make hypotheses, intervene, and learn the causal structure of the world. I think that is the dividing line. But I wouldn't say it's because the human brain knows the objective state of the world and ChatGPT does not.
Max Bennett, it’s been an absolute honor to have you on MLST. Everyone at home, you need to buy his book immediately. It is a wonderful, wonderful book, Max. You did such an amazing job of bringing all these things together. We've now spoken for 4.5 hours, going through the last 3 chapters. My God, it's been an honor. Thank you so much.
It's been my pleasure. Thank you for having me.