[BidClub_]
Machine Learning Street Talk · · 82 分钟

Karl Friston——为什么智能不能无限做大(金发姑娘原则)

Tim ScarfeKeithKarl Friston

播客
TL;DR
  • Friston 的核心判断是,智能存在于一个“金发姑娘区间”:系统太小或太大,智能都可能消失。 在微观尺度,随机性会压过形成主体性的循环结构;到了行星或天文尺度,平均化又会抹平组成部分的复杂性,只剩没有规划能力的运动。因此,星际文明可能需要联邦式组织,而不是一个统一智能体:进化“无需具备智能,也可以很美”。

  • 大型组织不会因为把一群聪明的部分简单加总起来,就自动变得更聪明。 企业只有建立能够平衡耗散与循环、保守动力的结构,才能保有大尺度智能——这相当于让组织运行在“混沌边缘”。Friston 对机构的更尖锐警告是,全球化和组织可能会“大到伤害自身”,因为它们逐渐失去规划能力。

  • 机器意识在原理上可行,但 Friston 怀疑标准的 von Neumann 架构能否高效、甚至能否实现机器意识。 有意识的人工系统需要主体性、足够广的反事实空间与时间深度,还要具身为一台“兽性机器”,并依赖具体基底、具备“会死亡”的计算属性。他押注的硬件方向是存内计算、忆阻器和神经形态光子学——不一定是脉冲神经网络——因为记忆基底必须参与计算并实现自组织。

  • 主体性与意识彼此正交,但递归因果结构既被用于解释主体性,也出现在部分意识理论中。 Friston 认为,主体性源于系统无法直接观测自身行动,因此只能把行动推断为感觉的原因,并由此在“内部拥有一个未来”,从中选择路径。意识可能还需要额外的精度控制、注意力和元认知——“观察自己正在观察”——而不只是拥有贝叶斯信念。

  • 大约20年后,Friston 说自己仍未遇到自由能原理决定性的“那就行不通了”时刻。 其核心几乎简约到极致:分割会产生条件概率分布,而这些分布的动力学遵循最小作用量原理。真正困难的是,如何传达一个“几乎同义反复般简单、但后果深远”的理论;Friston 仍然矛盾,因为少量“魔法与神秘主义”也会推动探索。

  • 这期节目同时拒绝大脑沙文主义和不加区分的泛智能论。 Mike Levin 的基础认知研究把病毒、黏菌和 xenobots 是否具备智能视为一个经验问题;Keith 则认为,病毒本质上是一台把自身机制外包给宿主的“捕鼠器”,在整个生命周期内都不学习。真正有用的单位或许是群落或生态系统,但 Friston 仍要求其具备正确的因果结构、反事实深度和合适尺度。

  • 对自主系统而言,静态感知远远不够,因为有意义的边界是历史,而不只是像素。 马尔可夫毯的发现必须识别那些能够随时间持续存在并反复出现的实体;单靠图像分割无法产生理解。因此,结构学习和动态世界模型是机器人与自动驾驶的核心,而 Friston 也承认,远距离作用仍让他缺少在相互竞争的因果结构之间做选择的“好数学”。

摘要 · 为研究而整理的核心内容

1. 20年时间强化了 Friston 的信念,却没有终结争论

  • Keith 将自由能原理追溯到约2005年,并要求对这20年做一次复盘。Friston 拒绝提前庆功:一边发展理论一边评估理论,就像一边做饭一边点评这顿饭;但每当某个现象“相当顺滑地嵌入”框架,他都会获得一阵多巴胺刺激,而截至目前,他仍未遇到决定性的“那就行不通了”时刻。

  • Friston 失望的地方在于传播。他对人们称自由能原理“出了名地难”感到矛盾:一些“魔法与神秘主义”确实能让人觉得这是值得攻克的挑战,但这个框架本应“几乎同义反复般简单”,而他怀疑自己的解释还没有让它保持足够直观。

  • Keith 将其与概率论相比:和规则很容易写下来,但推论意义深远,而且经常被误用。Friston 认同这一点,因为条件概率正是框架的核心:先对状态进行分割,再让一组状态以另一组为条件,最后为这些条件密度的动力学描述一个最小作用量原理——位于“热力学之前”,也位于“量子力学之前”。

  • Friston 主要把生成式 AI 和大语言模型视为理解自然智能的工程工具。这种理解可以改善生活、支持可持续发展,并通过计算精神病学服务心理健康;他援引 Feynman 的话——“我无法创造的东西,我就不理解”——并补充了一个令人不安的推论:真正的理解可能要求人类创造出“会受苦的东西”。

2. 马尔可夫毯把缺失的因果连接变成自然类

  • Friston 从外部状态、内部状态、感觉状态和主动状态出发。不同的禁止性因果影响会产生一套“自然类”本体论:既没有内部状态、也没有主动状态的实体,就是因果黑洞,典型地处于惰性状态;由于它从不反作用于环境,也几乎不可见。

  • 在细胞这类普通的接触边界系统中,感觉状态构成外层,主动状态位于其下,内部状态处在中心。这种排列同时施加了两个方向的条件独立:外部状态无法直接触达行动,内部状态也无法直接造成感觉。

  • 层级生物体有所不同,因为它们的主动状态会与内部状态隔离。从内部看,那些不可见的行动现在会和外部世界一起显现为感觉的原因;因此,生物体不仅要推断环境,还要推断自身。正是这种递归,让这些系统成为“奇怪的东西”。

  • 远距离作用让图景变得更复杂。细菌生活在接触局部的世界里,但视觉、听觉、电磁辐射和磁场——比如节目讨论的指南针——创造了非局部因果可能性。Friston 坦承,自己的结构学习仍存在空白:他不知道有什么好的数学方法,能够判断应如何在相互竞争的因果组织之间做出最佳区分。

3. 自我建模让主体拥有一个私有的未来

  • Keith 的框架是,计算性的奇异循环打开了完全不同的一类行为。Friston 表示认同:一旦系统把自己建模为环境的一个原因,它就必须推断自己正在做什么,以及这些行动将产生什么结果;这构成了“真正意义上的主体性”,而不只是被动调节。

  • 因为行动的后果尚未到来,自我建模天然指向未来。从恒温器、病毒或单个细胞转变为主体,发生在系统拥有“内部的未来”、能够表征分叉的可能路径并在其中做出选择之时。Friston 认为,这可能是真正主体性、甚至某种基本意识的必要条件,但并不是两者的定义。

  • Tim 追问,主体性与现象体验之间似乎存在混淆:更深的反身层级或许会强化主体性,但为什么它们会产生意识?Friston 完全承认这一点,并认同 Anil Seth 的区分:智能、主体性与意识彼此正交,“具备主体性并不意味着具备意识”。

4. 意识可能源自精度、可访问性和递归注意力

  • 自由能框架对意识提供的是双重面向一元论,而不是直接的意识解答。同一脑基底同时承载热力学信息几何和表征几何;在后者中,内部状态编码关于世界的贝叶斯信念,其中也包括行动。物理动力学和推断,是对同一过程的两种读法。

  • 仅仅编码后验信念可能还不够。恒温器可以被解释为持有关于温度的信念,但这并不意味着它有意识。因此,更强的理论要求信念在动力学上被“点燃”——具备足够精度,能够影响层级或分布式系统中的其他更新过程。

  • Friston 提到一种与 Mark Solms 相关的“被感受到的不确定性”理论:意识追踪的是对置信度、精度或不确定性的表征,而不只是内容本身。他谨慎回忆说,有观点认为,只要关闭一个约3立方毫米的脑干区域,就能关闭意识;该区域包含上行系统的起源,而这些系统编码并表征精度信号。

  • 更具要求的理论还加入控制和自我识别:先注意,再识别自己正在注意,然后再识别这种识别。Chris Fields 的“内屏”假说把嵌套的马尔可夫毯重新表述为全息屏幕,并提出一个不可约的内屏:它只能通过作用于大脑其他部分来认识自己。Friston 认为,这一设想与高阶思维理论和精度门控的全局神经元工作空间理论相容。

5. 机器意识需要长时间范围和会死亡的硬件

  • Keith 认为,意识需要最低限度的复杂性和因果结构,并把它与 Game of Life 中的滑翔机相比:低于一定像素数量,滑翔机就无法存在。Friston 表示认同,但认为边界在技术上很模糊——就像追问多少粒粮食才算一堆——并提名时间深度作为最合理的维度。

  • 反事实广度,是可用未来路径的数量;反事实或时间深度,则是这些路径延伸的距离。恒温器只有微分方程中隐含的即时未来。具有人类式主体性的机器,需要一个能够模拟自身行动后果的世界模型,同时具备足够宽的备选路径和明显更长的时间范围。

  • 对于机器最终能否理解世界或获得意识,Friston 的回答是“原则上可以”,但仅靠数学模拟可能不够。沿着 Seth 的“兽性机器”框架和 Geoffrey Hinton 的会死亡计算,他主张具身性与基底依赖:计算必须依赖具体物理基底,而不能只是独立于基底、在其上运行。

  • 因此,他怀疑 von Neumann 系统——处理器读取和写入分离的存储器——因为这种整体结构让记忆难以自组织。Friston 认为,有意识的机器必须在热力学和信息层面都遵循最小作用量路径,否则它无法在非平凡的时间长度内保持意识。他转而指向存内计算、忆阻器和神经形态光子学,同时明确表示并不要求脉冲神经网络。

6. 基础认知是经验问题,但病毒暴露了边界难题

  • Tim 追问,植物和病毒是否意味着“泛智能”,并挑战大脑被赋予特殊地位的做法。Friston 将问题引向 Mike Levin 和 Chris Fields 的基础认知研究计划:智能可能广泛存在,但必须通过实验将其揭示出来,而不能根据外观预设其存在。

  • 按照许多适应性行为标准,Friston 说,黏菌、病毒和 xenobots 看起来都“极其聪明”。他支持这种反对沙文主义的挑战,但坚持认为病毒不具备奇异事物那种表现性主体性;某些因果结构是分类意义上存在或不存在的,即使更广义的智能门槛仍然模糊。

  • Keith 的反驳聚焦于生物学外包。病毒通过辐射或转录错误发生突变,而不是主动修改自己;它像一台注入遗传物质的“捕鼠器”,把转录和其余工作都交给宿主细胞。在自身生命周期内,病毒既不学习,也不适应,因此与其称之为“智能”,不如把它视为普通的展开动力学。

  • 尺度可以在一定程度上调和两种观点:单个病毒可能并不值得关注,但群落或生态系统可能才是相关单位。Friston 同样把自然选择称为贝叶斯模型选择——偏好模型证据最高的表型——但不把适应性等同于规划。Keith 更严格的规则是:智能应归属于足够解释现象的最小边界,而不是归属于所有包含它的系统。

7. 智能会在金发姑娘尺度以下和以上消失

  • Friston 主张,应在每个候选对象的马尔可夫毯尺度上进行评估:其中是否拥有足够的机制,能够形成具备有意义反事实深度和广度的世界模型?病毒可能太小;生物圈可能太大,因为粗粒度观测会把组成部分丰富的因果组织平均掉。

  • 物理机制始于对动力学进行 Helmholtz 分解,将其拆分为耗散部分和保守部分。耗散性波动提供随机性和开放性;保守性或无散流则在状态之间循环,支持惯例、生物节律、繁殖和 Poincaré recurrence——即反复回到接近起点的位置,使系统得以维持可识别的身份。

  • Friston 认为,在极小尺度,量子随机性占据主导,循环性的保守组织不足以维持系统结构。人类尺度的生物体将一个不断迁移、持续变化的世界,与对特征状态的稳定重访结合起来。到了天文尺度,平均化会移除波动,留下以保守性 Newton 运动为主的系统:行星会运行轨道,但运动本身没有规划的证据。

  • 这就是处在“混沌边缘”的金发姑娘区间:完全有序和完全噪声都没有取胜。Friston 说,他看不出月亮、天气或进化在思考自己的未来。不过,当他把论证推向最终尺度边界时,又留下保留意见:“我不确定这是否正确。”

8. 粗粒化创造新的智能,却没有凭空增加新物质

  • Tim 追问,Gaia 看似没有智能,是否只是观察者能力有限的结果。人类通过忽略细节来抽象,通过扭曲细节来理想化;如果智能意味着“用更少做更多”,而涌现意味着“更多就是不同”,那么更高尺度的系统可能确实发生了重新组织,而不只是其组成部分的模糊视图。

  • Friston 给出的数学答案是重整化群:该算子递归地降低维度,把细尺度状态归并为适合更高尺度的变量。更高层级可以呈现全新的自证或智能动力学,但这些变量仍然是细尺度活动的函数——这构成了真实的涌现,却没有引入无法解释的新实体。

  • Keith 认为、Friston 也同意,大型组织如企业的智能需要刻意安装的组织结构。企业的粗粒度指标是生存——它是否能以可识别的形式从“Series A 走到不知道第几轮”——但规模扩大后,丰富递归组织的空间会减少,规划也会越来越困难。

  • Keith 将这一限制延伸到星际文明:最终,光速会阻止跨越光年的集中式协调。Friston 说,物理学论证似乎指向一个金发姑娘区间,但随后重新定义了这种失望:联邦式生态系统依然可以很美。“进化是一件美丽的事”,而且“它无需具备智能,也可以很美”。

9. 形态承载模型,DNA 约束策略

  • Tim 以延展认知为挑战,追问思维究竟在哪里结束:头脑内部、手机里、植物的形态之中,还是整个双向因果过程?Friston 回答说,讨论必须先确定一个马尔可夫边界,但每个有边界的事物都会受到更高尺度的环境限定;即便是没有智能的病毒,也需要一个有利于其持续存在的宿主世界。

  • 对植物而言,形态并不只是内部模型所表征的对象。物理结构会参数化条件分布,使“基底”本身成为生成模型的一部分。按照良好调节器的思想,生物体必须承载其环境的因果组织;因此,一个无尺度的世界应当对应生物体中的无尺度层级结构。

  • Friston 最喜欢的例子,是 David Attenborough 将植物画面加速10倍或100倍后的影像。根和茎会朝特定方向移动,植物会相互争夺阳光,有些还会捕食昆虫;当它们被加速到人类的时间尺度后,行为看起来极其像动物,很难再将其简单斥为没有智能,无论它们是否有意识。

  • Tim 起初提出,DNA 或许可以被视为一种类似软件的模型,随后 Keith 将问题进一步明确:DNA 指定的是一套策略——每个细胞在面对感觉状态时应如何行动——而不是一份完整的身体蓝图。Friston 表示认同:遗传代码提供变化缓慢的结构先验和预期约束,但每个生物体仍必须学习具体细节,比如该向哪里生长;同一套代码可以分化出根、树皮和器官。

10. 实际的智能发现必须沿着时间追踪动力学

  • Tim 转述 Maxwell Ramstead 的压缩表述:无法合并的有边界开放系统,会通过交换信息并实现同步。Friston 认可这一点——松散耦合系统会收敛到一个混沌同步流形,而自由能最小化可以被描述为广义同步,不必引入 Bayes、预测加工或自证。

  • Chris Fields 的量子版本,则用纠缠和幺正性来描述对应关系。Friston 开玩笑说,在共同相处“超过5分钟”后,参与者就会完全纠缠在一起:也许严格来说并不是一个系统,但 Fields 可能会这么说;至少,它们的动力学会占据同一个流形,变得越来越难以区分。

  • 对一台配备摄像头的机器人而言,图像分割只是迈出的第一步,即承认环境中存在某些东西。Tim 认为,理解需要历史,而不是快照;Friston 也认同,因为马尔可夫毯的发现必须利用动力学。如果每次只能看到一个宇宙状态,那么支撑概率分布的重复实现,就只能跨越过去和未来的时间产生。

  • 因此,自动驾驶系统需要在时间维度上保存推断出的马尔可夫毯,并区分持续存在的事物与水、雾这类无定形物质。Josh Tenenbaum 研究的以对象为中心的 Newton 先验是一种选择,也可以使用更宽泛的模型;但 Friston 的边界表述令人印象深刻:“我会放松要求,但不会放松到重整化群之外。”

Tim Scarfe

Keith has been talking about this moment where we have a glass of sherry with Professor Friston for about the last 4 years.

Keith

Right.

Tim Scarfe

And our dream has been realized today.

Keith

Absolutely. Such a pleasure to meet you. It's absolutely been a pleasure.

Karl Friston

Cheers.

Keith

Cheers.

Karl Friston

You mentioned consciousness, which you shouldn't really do with me, but if we stay here for long enough—more than 5 minutes—we will ultimately become completely entangled. We will not become one. Well, actually, Chris Fields thinks you would become one. Is evolution an intelligent process? It's certainly a free energy-minimizing process. So it's just Bayesian model selection.

Keith

Is the planet intelligent?

Karl Friston

Yeah.

Keith

Yeah.

Karl Friston

I don't see the weather planning. I don't see evolution planning. It doesn't think about its future.

Tim Scarfe

Life is the intensive property of matter, and intelligence is the extensive property.

Karl Friston

There's a large ensemble of universes. Well, of course, for the purpose of this argument, there's only 1 state of the universe at any one time. I'm not sure this is correct. I don't want to disappoint you.

Keith

Don't worry, it's well beyond my lifetime when that will happen.

Karl Friston

If you're going to make a difference, it's really about understanding how to use all the marvelous engineering that we've witnessed in terms of generative AI and large language models and the like—

Keith

Mm.

Karl Friston

—in the service of understanding the principles that underwrite natural intelligence—

Keith

Mm.

Karl Friston

—and deploying that understanding either to make life better—

Keith

Mm.

Karl Friston

—or to underwrite sustainability in a slightly more political way. Or to deploy it—which is why I got into this game—in the context of mental well-being and computational psychiatry. To quote Feynman, “That which I cannot create, I do not understand.”

Keith

Yeah.

Karl Friston

That means that we have to be able to create—

Keith

Yeah.

Karl Friston

—things that suffer.

Keith

Yeah.

Karl Friston

Which means that we need to understand the principles of natural intelligence.

Tim Scarfe

I don't know if you saw, I interviewed Yoshua Bengio.

Karl Friston

All right.

Tim Scarfe

I used the term “epistemic foraging” a few times, and he said, “I love that term. I love that term. Where did it come from?” Professor Friston?

Keith

Absolutely. Epistemic foraging is what we do. It's what we do.

Tim Scarfe

Yeah. To epistemic foraging.

Karl Friston

To epistemic foraging.

Tim Scarfe

It's just wonderful.

Karl Friston

There are certain causal forces in our universe that are at a distance, usually mediated by electromagnetic radiation or magnetic fields, in this particular instance. That really complicates the way in which we construe all the cause-and-effect structures that we have to model in our brain. If I was a simple little virus or a little bacterium, I wouldn't have to worry about electric fields, seeing things, or hearing anything. My world would just be that physical world that I could touch, literally. It would just be my next-door neighbors.

That leads to a very particular kind of Markov blanket structure and ecosystem of Markov blankets that precludes, or certainly does not license, some of the deep structures that we were talking about earlier in terms of strange things. The strangeness of a beautiful sort—beautiful loops, to use one of my colleagues' notions—rests upon action at a distance of a very non-spoofy sort that is nicely exemplified by this compass. At the moment, I don't know of any good maths that allows you to work out what is the best thing to do to disambiguate between this structure and that structure.

Tim Scarfe

Mm. Yeah.

Karl Friston

So if you can do that, that'd be good.

Tim Scarfe

Structural learning.

Karl Friston

Structural learning.

Tim Scarfe

To structural learning.

Karl Friston

Structural learning.

Keith

Somebody out there, help us out.

Tim Scarfe

Professor Friston, it is absolutely amazing to have you back on MLST. As you know, Keith and I are incredibly fond of you. You've been a big hero of ours for many years. I think this is the 4th time—

Keith

That you've been on the show?

Tim Scarfe

Yeah. I believe so.

Karl Friston

Yeah.

Tim Scarfe

Welcome, Professor Friston.

Karl Friston

I've lost count. But it's lovely to see you both in person, and congratulations on your recent marriage.

Tim Scarfe

Oh, thank you very much. Thank you very much.

1. The Free Energy Retrospective

Keith

So, we were just thinking about this. We've done these interviews. The free energy principle originated around about 2005, right? I wanted to get a bit of a retrospective from you on, let's say, what have been the successes so far of the free energy principle. How has it been going? What could have gone better over the last 20-odd years? Really, what's your take on the progress, I guess, of the free energy principle?

Karl Friston

It's a difficult question to answer, in the sense that when you're doing it, you're in the middle of it, and it's very difficult to assess. If you asked me if I'd made a meal and then asked me, “How's that meal going? Did you really enjoy it?”—when you're doing it, you're in the middle of it, and it's very difficult to assess.

So it's been a journey. I do catch myself from time to time thinking, “This is the right way to think about things.” Whenever you come across a phenomenon or a problem or an application that seems to just slip in quite neatly to the overall theoretical framework, every time that happens, I get a little buzz of dopamine, and that consolidates and comes with, “Yeah, I'm on the right track. I haven't wasted so much time.” There have been times in my life when I've realized I'm on the wrong track, so I'm used to that.

But for the particular application of the free energy principle, I have yet to have that: “Oh, well, that doesn't work, then. Think about it; do something else.”

What could have gone better? I'm ambivalent about this when I read that the free energy principle is notoriously difficult to understand. Now, half of me thinks, “Good,” because it's always important to have a slight degree of magic and mysticism to engage people. If people don't think there's a challenge, or that they've accomplished something by understanding it, then they're not going to be motivated to inquire, think about it, or indeed debate it.

On the other hand, the free energy principle is not meant to be complicated or difficult to understand. It's actually almost tautologically simple. Being able to communicate the free energy principle in a way that people find it a useful and obvious tool or method to apply could have, I think, gone better. That may be because I'm not very good at keeping things simple or intuitive.

Keith

We can attest to that.

Karl Friston

Yeah.

Keith

One thing that's interesting is I feel the same way about probability theory. Conditional probability theory is very easy to write down. You end up with the 2 fundamental rules, the sum rule and the product rule, but the consequences of it are very profound, and probability is notoriously misunderstood and misapplied. There are many errors of reasoning that happen when people try to reason through statistical correlations or whatever. So I think there are things like that in life that are very simple and yet very hard to understand the full extent of their impact, right?

Karl Friston

Yes.

I mean, it's interesting you picked up on conditional probabilities. That is the heart of the free energy principle at many different levels.

The classical formulation of the free energy principle starts off with a partition of different states of being. That partition suddenly means you've got the probability distribution or density over one set of states of being that can now be conditioned upon another.

The whole free energy principle is basically a principle of least action pertaining to density dynamics: the dynamics or evolution not of densities, but of conditional densities. That's it.

Keith

Mm-hmm.

Karl Friston

You know, this is before thermodynamics. It's before quantum mechanics. It's just about conditional probability distributions. So it's interesting you picked up on that as something that is so simple, yet so absolutely powerful.

2. Strange Things Become Agents

Tim Scarfe

Professor Friston, we were reading your paper earlier. It's from 2023, “Path integrals, particular kinds, and strange things.” It introduced a categorization of particles, from inert to active and ordinary to strange. The strange categorization was particularly interesting because you were giving an account of how these particular particles could give rise to phenomenal, conscious states or even agentic states.

We felt that this was almost an account of pan-agentialism. Historically, we've done interviews about the free energy principle, and we've spoken about self-organization and emergence. Agency and phenomenal states are things that emerge when you have this temporal and counterfactual depth. Maybe we're misunderstanding, but it feels like this is an account of low-level agentic potential and phenomenal potential.

Karl Friston

To set the scene, just to come back to that partition that we were talking about, there are many ways of arranging the conditional independencies to disconnect various partitions: external states, internal states, sensory states, and active states. The combinations of disallowed influences give rise to an ontology of different kinds of things, and you could call them natural kinds.

I've been told I shouldn't use that, but I like “natural kinds”: the things in nature that prescribe themselves just by being special instances of a lack of causal influence.

If we start off with the simplest case, where there are no internal states and no active states, what are we talking about? We're talking about some kind of causal black hole, something that is quintessentially inert. You can never see it because you can only see the active states that act back on the environment in which it is embedded.

But then we get to more interesting things that have a full complement of internal and external states, and a bidirectional coupling between the two, which is mediated by the sensory states and the active states. If you disallow action at a distance—in other words, I have to be, in some metric sense, next to you in order for my active states to become your sensory states, and your sensory states to become my active states—then you have a very simple kind of Markov blanket.

That would be fine for describing active matter, for example, or any medium where I have to touch to feel, to influence, or to sense. Interestingly, this kind of particle has its active states on the inside, which I had to think about. It can be no other way if you just try to arrange all the dependencies.

So what would that look like? It would look like something like a cell that had, on the outside, sensory states supported by a layer of active states that surrounded the internal state. You're complying with the laws of the Markov blanket of thingness: you need to have that conditional independence to separate the thing from everything else, or the self from the non-self.

Everything's fine because the active states are hiding behind the sensory states, so the external states can't influence the active state. So that's tick one. That's what we need.

But we also have to ensure that the internal states, symmetrically, cannot act to influence the sensory states, and that's fine because they're hiding behind the active states. So you've got a nice mathematical image of a cell, if you like.

But that doesn't work for things like you and me. Things like you and me have a hierarchical structure. What that basically means is that the active states which, in very simple organisms, were immediately juxtaposed to the internal states now become sequestered.

Now, effectively, the active states are no longer seen by the internal states, and then something quite remarkable happens mathematically. It's really simple, and it's just another aspect of conditional probability distributions. Because you can't see your active states, you can now only see your sensory states. It looks as if, from the point of view of the internal states, the active states have now become causes of sensory states.

So now my world is caused not just by the external world, by the environmental states, by my heat path, by my external milieu, but also by my own actions. Now I'm inferring the causes of my sensorium, where I'm actually a cause, so I'm inferring myself.

There's this beautiful recursion which licenses this strange-loop analogy, with a bit of poetic license. Another way of putting that is that, for these simple structures—for these natural kinds that would be, say, single-celled organisms—the internal states can be read as modeling the causes of their sensations, which just are the external states.

They have direct access to the active state, so there's no conditional independence that licenses the notion of description in terms of inference or sense-making. Whereas now, the Bayesian mechanics that attends the internal activity, the internal machinations, covers both my action and the things that I'm acting upon.

Of course, this is a nice metaphor for planning as inference: I'm now thinking about and trying to infer what I'm actually doing. That, to my mind, produces a very unique and special kind of thing—things like you and me, basically—which are pretty unique when we look at all the different kinds of things that could be around.

Keith

Well, it's another example of how, as soon as you had this recursion—as soon as you had these strange loops, especially for systems that are computational.

I forget who it was that said, “Life is the computational phase of matter.” When you're at this level of complexity, where you're having some kind of information processing, and you introduce this strange loop, that unlocks a completely different category of behavior, a different tier of computation.

Karl Friston

Absolutely.

Tim Scarfe

It was David Krakauer, and he said, “Life is the intensive property of matter, and intelligence is the extensive property.”

Keith

No, this is a different one, though. It may have been Wolfram who said, “Life is a computation.” I don't remember. I'd have to look it up. But, yeah, it unlocks a completely different category of behavior.

Karl Friston

Yeah. Well, you alluded to that earlier in terms of phenomenology, and you mentioned consciousness, which you shouldn't really do with me. We'll read that as an elementary kind of sentience.

I think that's really important: that you do unlock or manifest, or now at least have a mathematical calculus that allows you to talk about things that model themselves. As soon as you're modeling yourself, you become an agent, or you have an authentic kind of agency.

In order to model myself, or to model the me as a cause of my environment, I have to infer what I'm doing, in particular, the consequences of what I'm doing. Because the consequences have not yet appeared, this has a quintessentially future-pointing aspect.

So now we're moving from a thermostat, a virus, or a single-celled organism to things that actually have a future on the inside—their private future that just exists for them. Of course, if you can get suitably far into the future, where you have a divergence of particular paths into the future, just as a nod to the path-integral formulation, then you have to select one.

So now you've really unlocked, again, a calculus of selection of your paths into the future, which I think is not a definition of agency and certainly not consciousness, but would arguably be necessary to support something that had true agency and possibly even consciousness of an elemental kind.

3. Consciousness Beyond Agency

Tim Scarfe

Could we press on consciousness a bit? The account that you just gave is a beautiful account of agency, and it's very plausible. As you increase this reflexive layering, this recursion, you get these multiple paths into the future, and you become the cause of your own actions. So you become more of an agent. It's almost an account of the strength of agency, with more and more layers.

But in the abstract of this paper, we were quite struck by the language that conflated consciousness with agency. Intuitively, I feel that phenomenal experience is quite orthogonal to agency. Why were they convolved together?

Karl Friston

I'm desperately trying to remember how I referred to consciousness. You've got to be very careful when using that word, depending on who you're talking to.

Tim Scarfe

Yes, the C-word.

Karl Friston

The C-word. Oh, naughty.

I think people like Anil Seth make a very similar point quite earnestly: intelligence and consciousness are completely orthogonal, and you're making agency and consciousness completely orthogonal. I would agree entirely. Consciousness isn't implied by being agentic, by having agency.

There are many stories you can tell here, and my mind goes to the people who might be watching this to make sure I don't offend anybody by not mentioning them. From the point of view of the free-energy principle, the way that you'd look at consciousness is in terms of a dual-aspect monism, which you get for free from the treatment of the density dynamics in exactly the way we were talking about before.

Now my internal brain states have a thermodynamics. That can be written in terms of an information geometry that itself is just predicated on conditional probability distributions. But they also have an information geometry that inherits from the fact that my internal brain states represent the outside world, including my own actions. So there's both an information geometry and, by implication, a thermodynamics of my brain activity, which supervenes on exactly the same substrate as does the information geometry, which is representational—the inference, the Bayesian mechanics side of things. That gives you license to talk about a sort of dual-aspect monism.

But what does it really mean to be conscious? Some people might believe that it's just having a posterior belief, having a Bayesian belief, which is literally just a conditional probability distribution that is encoded or parameterized by some physical state of being. In this instance, under the free-energy principle, it's the internal states of being.

Other people say, “Well, no, just being able to read a thermostat, for example, as having posterior beliefs about the temperature of the external world does not license you to ascribe consciousness to this thing.” So the next step would be: it has to be ignited in some way. It has to be realized. It has to be emergent.

I'm sort of paraphrasing Jakob Powrie here and, to a certain extent, Mark Solms. They emphasize that it is when these beliefs become sufficiently precise in a dynamical way that they are realized and influence other aspects of belief updating in a hierarchical or distributed system. Mark Soames would call this felt uncertainty. He would emphasize that the feeling part of consciousness is mediated by representations not of the content, but of the precision, uncertainty, or confidence with which these beliefs are currently in operation.

He will bring to the table all sorts of very compelling neurobiological evidence as to why this is. You can switch off consciousness literally by, I think he says, a 3-millimeter cube of the brainstem containing the cells of origin of those ascending systems that encode and represent—not the content, but the confidence, uncertainty, or precision of these conditional probability distributions.

Jakob, I think, would say that to be aware is to equip all those messages that are providing evidence for your current explanation with precision, which has physiological and bioelectric mechanisms behind it. Other people go further.

People in the world of phenomenology, of the kind that Thomas Metzinger pursues, would say that it is necessary to have control over the precision and, furthermore, to be aware that you are rendering that which was once transparent opaque. You actually have to recognize that you're attending.

So, again, you not only have this very simple recursion of modeling and self-modeling in terms of planning as inference, but now you've got this hierarchical recursion where you're recognizing that you're attending to something. If you're somebody like Lars Sunved Smith, you'd say that to actually be self-conscious, I now have to recognize that I was recognizing that I was attending to something. You get layer upon layer upon layer.

To join the dots with Chris Fields, who has brought to the table the inner-screen hypothesis—have you come across this yet?

Tim Scarfe

No. No, tell us about that.

Karl Friston

Strictly speaking, you should get Chris to tell you about that. But I'll give you a quick preamble.

Chris Fields is the theoretician who has formulated the quantum information-theoretic version of the free-energy principle. For him, the Markov blanket that separates self from non-self becomes a holographic screen upon which classical information is written and from which it is read by some internal bulk and some external bulk.

In the sense that strange loops and strange recursions are only allowable when you have this hierarchical structure on the inside, you've now got lots of inner holographic screens, inner Markov blankets. Markov blankets, in their pragmatic sense, just define things like hierarchies, for example, and they define the architecture of any network, factor graph, or computer. When you've got lots of them, you've effectively got lots of inner screens.

My reading of that idea is that there is 1 irreducible inner screen—1 inner screen that within it has no other screens. This is an interesting screen, a Markov blanket from the classical perspective, because the only way that the internal states of this irreducible Markov blanket can know themselves is by acting on the exterior. Of course, the exterior is the rest of your brain. So you've got this metacognitive, self-recursive aspect to the inner-screen hypothesis.

That notion has a lot of mileage in relation to other people's theories of consciousness. We're talking about higher-order thought theory, so the very notion of higher-order thought and its implicit appeal to metacognition, meta-metacognition, and meta-meta-metacognition is again this notion of looking at oneself and inferring oneself, but on the inside. Looking at the looking at the looking. This is, of course, a natural consequence of having this hierarchical structure.

You could also argue that it's completely compatible with global neuronal workspace theories: there is this precision-dependent ignition of certain sources or messages—sufficient statistics of conditional probability distributions—that gain access to all of these inner screens in a dynamic way because you're controlling access by attending, by encoding and representing, and by optimizing the precision, uncertainty, or confidence.

So you've got this notion of ignition and penetration through to the global workspace, which we could read as this sort of irreducible inner screen: the core, the deepest part of your sense-making, your brain. I think this notion is not only consistent with the conditional independencies that are implicit in the free-energy principle as applied to strange things that have this hierarchical and recursive aspect, but also relates comfortably to extant theories, all coming from different perspectives in different parts of the elephant, as it were.

It's a comfortable accommodation of things that we presume, or things that people have brought to the table in order to explain consciousness from their direction of theorizing.

Keith

It seems that, if we split the camps—or, let's say, the thought groups—that thought about this, almost all of what you just talked about now are accounts in which there is a minimal complexity required for there to be consciousness. Then we can talk about what that level is.

Well, it requires this or this or this. But then there are also people who say, “No, there’s no minimal complexity.” Electrons have some small amount of consciousness. I fall much more into the category of people who say there is a minimum—you have to reach a certain something.

I don’t really know where that boundary is. Maybe it’s one of the things we’ve mentioned. Maybe we’ll figure out a more elegant categorization, but it’s just like needing a certain number of pixels in the Game of Life before you can have a glider. You can’t do it with 3. You need some minimum number, and so there’s some minimum amount of brain material. Three cubic millimeters is still a lot of neurons. I don’t know how many are in there, but hundreds, thousands, or whatever it is.

You need some level of complexity before you have consciousness, or before you have recursive agency, intelligence, or whatever. Would you agree with that? There is some minimum. We don’t know where it is yet; it’s hard to place, but there is a minimum?

Karl Friston

No, I would agree entirely.

Keith

Okay. So there is a minimum, and not only that, it has to have a certain causal structure, right? We can debate that, and that’s really in line with Searle in the Chinese room argument, which is: a dictionary doesn’t have understanding because it doesn’t have the right causal structure. You have to have a certain causal structure or a certain minimum complexity, and then you reach this—whatever it is, whether we’re talking about consciousness, understanding, agency, or all of these things.

4. Building Conscious Machines

So I guess my question to you is: will we be able to build machines based on our current computer architectures someday, whether it’s 100 years from now or 200 years from now? It doesn’t matter. In principle, can we build machines that have understanding, consciousness, and all these capabilities?

Karl Friston

Yes, that’s a good question, and the answer, I think, is yes. Can I come back to qualify that answer by reference to that wonderful example about what I would read as vagueness in a technical sense? Vagueness: how many grains of sand constitute a pile?

Keith

Right. Yeah.

Karl Friston

It’s not well defined. It is in philosophy, but not in mathematics. I think that’s absolutely the right way to think about these bright lines. If I had to commit to the dimension—the number of grains of sand that you get before you have consciousness, or there is a pile—I would say it’s something you actually referred to earlier on. I think it’s the depth of your future, or your future in your head.

If you’re talking now about an algorithm or some artificial intelligence that is equipped with a generative or world model of the consequences of its actions, there will be a time horizon associated with that component of its generative model. I think it’s the depth of that time horizon—the thing that Anil Seth would refer to, not as counterfactual breadth, which is the number of divergent paths one could take into the future, the options that you select among, but counterfactual depth.

Keith

Okay.

Karl Friston

The temporal depth. That means that you can be panpsychic. You can say that a thermostat has a notion of the future, in the sense that it operates through path-integral control and differential equations. As soon as you put a differential equation in play, you’ve got an instantaneous future because you’ve got some gradient with respect to time.

That’s not, though, the kind of depth that you and I enjoy. I would imagine it is really just the depth. Maxwell Ramstead talks about this as merely reflexive active inference, with very myopic, very short-term self-models, right through to fully—well, to be conscious in the way that we’ve been talking about, I think you’d need to have a long depth.

So what does that mean for building AGI, conscious artifacts, or machine consciousness? First of all, they have to be agentic, because we’ve just said that having a world model of the consequences of your actions would be necessary to be an agent. More than that, you’d have to look quite a long way into the future.

Your generative model, your world model, would have to go quite a long way into the future. I repeat: it would have to have both counterfactual breadth and counterfactual depth at hand. Would that be sufficient? I’m now remembering that I forgot to mention Anil Seth in the list of people not to upset when reviewing theories of consciousness, so this is an opportunity just to say that there are people out there who would say, “Well, okay, you can write down the maths of all this, and you can write down in silico hypotheses, or indeed simulate global neuronal workspace theories and try to produce things that look as if they have consciousness.”

That’s not going to work unless you actually embody it, unless you are, in Anil’s words, a beast machine. I think that coheres with the argument for mortal computation. So when you ask the question, “Can we build it on our computer architectures?” I would have to ask you: do you mean a von Neumann architecture, or do you mean a memory-processing or in-memory-processing architecture?

I would take that as synonymous with a neuromorphic architecture—not spiking neural networks. You don’t need those, but you do need processing in memory to be mortal. You need that substrate dependence, read in terms of Geoffrey Hinton’s definition of mortal computation and Alex’s subsequent elaborations of that. I think I would subscribe to that, largely to keep Anil happy.

I don’t think you can do this on a von Neumann architecture, because the Markov blankets of a von Neumann architecture, where you’re reading and writing from memory, make it very difficult for the memory to self-organize.

Tim Scarfe

I see.

Karl Friston

People have written about this philosophically. I think Vanya Weiss has written a paper about this, and I think Anil Seth speaks to this argument in a recent Behavioral and Brain Sciences paper. It may not be possible, and certainly, from the point of view of efficiency, it’s highly unlikely that von Neumann architectures are the kind of things that would conform to—

Let me reverse. To be is to pursue a path of least action in accordance with the free energy principle. To be conscious does not excuse you from that. If you want a conscious machine, you have to have a machine that pursues a path of least action.

Via the Janiskee quality, that has to be true both thermodynamically and informationally, in terms of the conditional probability distributions. If you don’t pursue that path of least action, you can’t be conscious for a nontrivial amount of time. I can see you want to ask a question, so I’ll interject.

Tim Scarfe

Well, only to say that Maxwell pointed me to Anil’s work, where he defends biological naturalism. Of course, even Searle said that the biological substrate is an existence proof. It is not saying that it has to be biology. But Keith and I certainly agree that there’s something about the substrate that is very important.

5. Intelligence Beyond Brains

I wanted to talk a little bit about viruses. As you know, I’m a bit of an externalist. I’ve never been able to completely pin you down, Professor Friston, because there have been so many interpretations of the free energy principle that lean internalist and externalist, and even the hybrid version, which Maxwell also wrote a paper about.

But I’m fascinated by this idea of diverse intelligences. For example, could a virus be intelligent? I spoke with David Krakauer, and he said intelligent things do inference, have representations, and are adaptable.

And we're a little bit chauvinistic about our brains, aren't we? Our brains seem to have a privileged status. But what say you of viruses? Or, actually, we read your paper with Calvo, “Predicting Green: Really Radical (Plant) Predictive Processing.” In that paper, you gave a beautiful account of how even plants could be doing inferencing.

And that, to me, seems incredible because I'm amenable to the idea. But it seems to me intuitively that plants are not as sophisticated as we are. So it comes back to this line that Keith was talking about before: if you go too far down the stack, it's an account of—we'll call it pan-intelligence—where you have a tiny little bit of intelligence even if you go all the way down the stack.

Karl Friston

Have you spoken to Mike Levin?

Tim Scarfe

Yes, relatively recently, about a year ago.

Karl Friston

Right.

Tim Scarfe

Yes.

Karl Friston

I mean, it would be nice to revisit him. That is exactly his big question at the moment. And, of course, he has, as an intellectual accomplice, Chris Fields with him as well. The two of them are pursuing this notion of basal cognition: that there is intelligence everywhere. It's just a question of how we conceive of it and how we test for it.

And indeed, Mike's argument is that you have to design the right experiments. It's an empirical question: is this virus intelligent or not? Well, you have to design the right experiments to disclose or evince intelligent behavior, and Mike would claim that, by many metrics and yardsticks that we use to measure adaptive intelligent behavior, slime molds, viruses, and xenobots are incredibly intelligent.

I think he's fighting against the chauvinism that you mentioned: that intelligent things are just properties of creatures like you and me and our brains, and that this kind of cognitive capacity and competence can be found everywhere. So I'm very sympathetic to that. But it does tread on the toes of the vague argument.

I certainly don't think that viruses have the same expressive kind of agency that we were talking about. And indeed, the counterargument to the vague notion of intelligence and/or consciousness is hitting you in the face when you read that Strange Things paper. I'm talking about categorically different natural kinds that do and do not have, for example, a causal power or influence of active states on internal states. You're either one of these, or you're one of these.

Keith

I'm not so sympathetic to the view that a virus is intelligent, right? Because I think this falls into the category of things like when you're in a biology class and you learn how to define life, and then somebody says, “What about fire?” It seems to meet all these criteria. It grows, it can expand, and it uses resources.

For example, when a virus “mutates,” the virus doesn't mutate itself. It is mutated by a gamma ray hitting it, or by a transcription error, or whatever. Oil, vinegar, and baking soda undergoing a reaction is not intelligent, right? I think there's some level of complexity—whether it's recursion or these other types of causal structures—that has to be there before I'm even interested in talking about intelligence or entertaining intelligence. Other things are just dynamics that unfold in certain ways: chemical reactions and that sort of thing.

6. Intelligence Has a Goldilocks Scale

Karl Friston

And I think that, at that point, is a nice opportunity just to introduce the notion of scale-freeness, or scale invariance.

Keith

Okay.

Karl Friston

I can see you could also take this toward collective intelligence, federated learning, and federated inference. It's not the single cell; it's the single cell with its neighbors, and its neighbors' neighbors' neighbors' neighbors. What one's looking for is some conservation of intelligent dynamics that is preserved over different scales.

The single virus is probably not interesting. It's probably the colony that is interesting. I thought that was a nice point at which to introduce the notion of that kind of scale invariance. And, of course, as a mathematician, you would be looking at the renormalization group to see how that unpacked mathematically.

Interestingly, it also speaks to this notion of a mutation. Is evolution an intelligent process? It's certainly adaptive. It certainly has the level of complexity that you would require in order to pass those vague thresholds. But is evolution in and of itself an intelligent process? It's certainly a free-energy-minimizing process.

It's just Bayesian model selection. Natural selection just is selecting those things that have the highest model evidence, or marginal likelihood, of being that phenotype in this kind of environment.

Keith

Well, it's like, is the planet intelligent? It certainly contains 8 billion or so of us, so does that count? I mean—

Karl Friston

Yeah. Why not?

Keith

Well, I would say—and we talked about this way back when we talked about whether a flotilla is an agent, or whether it's the pilot who's controlling the convoy or whatever—I would say that if there's a more minimal boundary that contains the intelligent entity, then the larger one is not. The planet, for example: since I can draw a smaller boundary, which is an individual person, the individual person is intelligent. Anything that contains that person is just more stuff.

Tim Scarfe

But Keith, what is your— The issue with viruses: is it because the individual virus is inert? Whereas a cell, for example, you can partition down to the individual cells, and the cell still has some degree of agentic property, as Professor Friston was describing. Is it that you see them as inert and being carried by something else?

Keith

Well, I think it's that, too: too much of their causal structure and machinery is basically outsourced to other things.

Tim Scarfe

Yeah.

Keith

They're kind of just these little particles that go around, and if they happen to stick on a cell, they're like a mousetrap that activates and just injects some DNA. But the cell does all the rest of the work, right? It does the transcription. It has all that machinery.

By themselves, a virus doesn't change over the course of its lifetime. Other than this simple mousetrap kind of activation, it does nothing. It doesn't have any machinery to learn and adapt by itself over the course of the lifetime of a virus.

I might be more amenable to somebody saying that an ecosystem of viruses or something has some degree of intelligence. But it still doesn't have the type of processing and causal structure, and the minimal complexity, that, at least for me, is the useful concept of intelligence.

Karl Friston

So I think we're coming back now to the number of inner screens, the counterfactual depth, and the definition of agency in the sense of Strange Things that we're talking about for this particular scale. If we just read the scale as, if you like, the size of the Markov blanket, then at this scale, for this kind of thing, if there is a sufficient degree of complexity—read specifically, though, in this instance, as the counterfactual depth and breadth of your world model about the consequences of your action, that particular part of your implicit generative model—that would happily accommodate the fact that when you get something of the size or scale of a virus, there just isn't the machinery or the space to entertain that in any nontrivial way.

Keith

Right.

Karl Friston

One could also argue—and I got a sense that you were arguing yourself toward this—that if you go too big, you also lose that.

Keith

Mm-hmm.

Karl Friston

Coming back to this wonderful question: we have 8 billion intelligent—really intelligent, I sound like Trump there, didn't I?—entities constituting our biosphere. But is the biosphere, from the point of view of, say, the Carr hypothesis, intelligent in and of itself? I would say no. It's simply because, at that scale, all the complexity of the constituent elements disappears.

Keith

Yes.

Karl Friston

It would be a little bit like saying, “Let's just take it to the limit. Let's just take it to the astronomical limit, the motion of heavenly bodies.” The motion of heavenly bodies is completely described by the position of the planets and the Moon, right, and the position of the Earth.

At that scale, you've averaged away the fact that the Earth contains a biosphere, the biosphere contains human beings, human beings contain cells, and the cells may or may not contain viruses, depending upon who you've been exposed to. So there will be no intelligence at that level.

So it’s perfectly possible to have intelligence at a particular scale that disappears when you get too big and when you get too small. Another way of looking at that is going right back to things like Prigogine and dissipative structures. You’re talking about complexity. If we just think about how people have tried to understand complex systems and self-organizing systems that are open, we’re not talking about 20th-century physics and equilibrium physics. We’re talking about the physics of nonequilibrium and things that are open, in open exchange with each other.

And, of course, you get to the notion of dissipative structures. What does that tell you mathematically? Well, what’s not dissipative? That’s a schoolboy question. Can you remember? You’ve probably forgotten. So what is not dissipative is conservative.

Keith

Oh, got you. Sure.

Karl Friston

Yeah, I’m sure I’m teasing. Another gift of Helmholtz, of course, is that any dynamics can be partitioned into a dissipative part and a conservative part. The dissipative part is that which rests upon very fast, random, complicated fluctuations, of the kind you might find in, say, quantum mechanics or thermodynamics, whereas the conservative part does not rely upon that and just goes round in circles, basically. It’s literally called solenoidal flow, or conservative flow.

So that basically means that dissipative structures have to have an admixture of this circular aspect, this conservative, classical aspect—life cycles, reproduction, oscillations—plus the random dissipative part that you’ll find in things like thermodynamics and quantum mechanics. That is definitional of dissipative structures.

As you get too small, everything becomes quantum and random. It all becomes probabilistic. There is no conservative stuff other than a Schrödinger potential, at which point I think you would find it very difficult to find something that was intelligent in the sense we’re talking about, because we have to have this recurrence, this solenoidal aspect, in order to revisit the states, so you know a virus is a virus, in order for there to be a nonequilibrium steady-state distributional solution to that.

But as you get bigger and bigger and bigger, you get to viruses, and they’re still not quite complex enough to be intelligent in the way that we mean. Then you get to our size, and that’s perfect because we have both this dissipative aspect—we deal with a random world, with an itinerant world—and yet we keep revisiting states of being. So we have this sort of conservative, biomimetic kind of self-organization.

But then you get bigger and bigger and bigger, and you get to the level of the biosphere or the size of the Moon or the Sun. Of course, by averaging, all the random fluctuations go away. So you’re just left with the solenoidal part. You’re just left with the solenoidal motion, the Newtonian motion of classical mechanics.

Keith

I want to jump in here because one thing that really fascinates me about this is that it’s possible for us to construct large-scale things that are intelligent, like a corporation. A corporation is seen as a superintelligence. But we’re only able to do that by putting in structure that maintains this balance of dissipative and conservative flow. We have to put in organizational structure so that it continues to function at these larger scales, right? Isn’t that pretty interesting?

It’s again this yin-and-yang thing that we run into all the time, where it’s the balance between 2 forces: either between complete order and complete noise, or between dissipation and conservation. You have to be almost on the edge of chaos, and it has to have a certain causal structure in order for it to be intelligent.

Karl Friston

No, well, I agree entirely. That sort of Goldilocks regime, where you are on the edge of chaos, I think is quite specific to a particular scale. Certainly, you could invoke a sort of strong anthropic principle here and say that the kind of intelligence that we will recognize has to be at our scale.

But I think there’s something more fundamental than that. I think that the very existence, if you subscribe as an externalist to quantum physics and Newtonian physics or Lagrangian classical mechanics, means that there is a Goldilocks regime. It means that we can only exist at this scale with, as you say, this sort of yin and yang, this admixture of dissipative dynamics and conservative dynamics.

Just to reinforce this, conservative dynamics is absolutely essential because it is that which causes this Poincaré recurrence. It defines these strange attractors, which is the other way of licensing the notion of strange things. You always come back to somewhere near where you started, like Red Queen dynamics in theoretical biology.

So you have to have the conservative, circular motion just to have a routine, have biorhythms, replicate, and reproduce at many different levels and at many different scales. But it’s always remarkable in the face of a dissipative, itinerant world. Everything is changing all the time, and yet we somehow resist that change by being at the edge of chaos.

You’re more anxious than I am to say something. Carry on.

Keith

Please go on, Professor Friston.

Karl Friston

No, no, no.

Keith

I’m just so fascinated by your observation that with Gaia theory, for example, when we zoom out, the apparent phenomenon seems less intelligent. As you said, maybe intelligence is just what we recognize. When we spoke with Wolfram that time, he was talking about how we’re computationally bounded as observers.

We do this thing called abstraction, where we ignore details, or even idealization, where we deliberately distort the truth. David Krakauer said that intelligence is doing more with less, and emergence is “more is different.” The Earth isn’t doing less; we just don’t see what it’s doing. But if emergence is “more is different,” then that licenses a fundamental reorganization of the underlying substrate, which means it isn’t actually doing the thing at the lower level anymore.

It’s a fundamental coarse-graining, and it’s actually changing at the higher level. So which is it?

Karl Friston

Well, I didn’t realize there was a choice there because I agree with everything you’ve said. But I like the notion of coarse-graining because, of course, that is exactly what you get from the renormalization-group treatment of these things. You can simulate free-energy-minimizing processes that are running at different scales, and then you have to ask the deep questions of how you couple between the scales. There you get to supervenience and emergence, mathematically so defined.

But it all boils down to exactly where you started and where you ended up, which is a coarse-graining in the right kind of way. To actually write down a renormalization group, you have to have an RG operator. What is an RG operator? It has 2 parts to it. It has a sort of dimension reduction—the R part, if you like—and the grouping operator.

They basically do the right kind of coarse-graining. They reduce dimensionality, and they group together in the right way those states of being at the higher scale, and so on, ad infinitum. So I think the notion of coarse-graining is absolutely essential here. On that level, you could argue both ways.

I’m getting a sense that your question is something like: Is there a true emergentism here or not? Again, I don’t have philosophical training, so I’m not sure I can really answer that. But certainly, the intelligent dynamics, read as a self-evidencing perspective on a free-energy-minimizing process, are recapitulated at a completely different scale in a different way at each scale through the recursive application of RG operators.

So yes, there is something brand new going on. And yet, in the spirit of Haken synergetics, these RG operators are just taking functions of stuff that’s happening at the finer scale. So you’re not inventing anything new. It’s just predicated on, emerging from, or supervening on finer-scale stuff. But what emerges at the higher scale has all the attributes of self-evidencing, and you could argue for consciousness or intelligence at some level.

But coming back to the point about building bigger organizations: If the argument is that, as you get bigger and bigger and bigger, you have to be effectively conservative in order to exist, then a company, for example—what is its metric of goodness? It’s how long it survives. Can you get from Series A to Series whatever in some recognizable form?

So as you get bigger and bigger and bigger, I think there’s less opportunity for this complex, hierarchical recursive structure within the scale. I don’t see the Moon thinking or planning.

I don't see the weather planning. I don't see evolution planning. It doesn't think about its future. It's too big. I would imagine that you could apply the same arguments to globalization and institutions that get too big for their own good, because they can't plan anymore. So we're coming back to the pilot, who's really the intelligent person, not the flotilla.

Keith

Mm-hmm.

Karl Friston
Keith

Right. Yeah. In a way, it's slightly depressing to me because it means we can't build an intelligent intergalactic civilization. It's almost like, beyond a certain scale, we just have to leave it up to distributed, emergent processes that may or may not end up doing something intelligent as a whole. At some point, you've got the speed of light, and that's basically going to be a barrier to organization across light-years, right? So maybe there is an ultimate Goldilocks limit to intelligence. It can't get too big.

Karl Friston

That's certainly what the arguments from the physics part suggest: there is a Goldilocks of scale, or zone in scale space. I'm not sure this is correct, and I don't want to disappoint you, so—

Keith

Don't worry. It's well beyond my lifetime when that will happen.

Karl Friston

There's a certain beauty, though, in federating and distributing and thinking in terms of ecosystems of things at a particular scale. Evolution is a beautiful thing. It doesn't have to be intelligent to be beautiful.

Keith

Yeah, yeah.

Karl Friston
Keith

Correct. Correct. Well, good point. It could be beautiful but not intelligent. We could have a beautiful, distributed civilization. That's fine. I'm happy with that.

7. Drawing the Cognitive Boundary

Tim Scarfe

On the internalism and externalism debate, I'm a huge fan of Andy Clark, as you know, and he famously argued in his 1999 paper with David Chalmers, that the phone extends your mind. Even when we were talking about the plant example before, when we were talking about plants doing modeling, there's an enactive interpretation there. Keith and I were arguing about this, and he doesn't agree with me.

Is the model the physical morphology of the plant? It fits the environment like a key in the lock. Or is the model the software? Is it the DNA of the plant? It's fascinating, isn't it, that you can think of cognition as not being entirely inside our heads, but as this unfurling, causal, bidirectional process of so many things around us.

So how do we actually draw boundaries around things where we say, “Okay, well, here's the principle: modeling and cognition are in here. This is where bona fide beliefs are happening. This is where bona fide thinking is happening”? Is it even possible to make that distinction?

Karl Friston

I think it is. I'm answering it in an almost trivial way: if you want to talk about something, you have to be able to define its Markov blanket.

I think your question, though, speaks to something that we've just been covering, which is the separation of scales. Everything, literally, from the point of view of the free energy principle, that is equipped with a Markov blanket is contextualized by a scale above, and the same rules apply to the scale above as to the thing itself. So you have to have a context in which everything is operating.

The virus may not be intelligent, but it certainly has to be living in a world that is conducive to its existence. That world, and the states of that particular world at the scale above—for example, the host cell and the host organism—have to comply and be intelligent in some sense in order for the virus to be there, even if the virus itself is not intelligent.

But we should come back to extended cognition and Andy Clark. To the plant, you made an interesting point: is it in the DNA, or is it in the morphology, the phototaxics, and everything else that plants possess and that we infer goes on inside?

First of all, I think you're absolutely right to say that it is in the morphology. In a sense, that's what I was getting at when talking about the importance of mortal computation for machine consciousness. To realize Bayesian mechanics or self-evidencing, you physically have to parameterize your conditional probability distributions—your Bayesian beliefs about the environment with which you are coupled. The substrate is the parameterization.

By definition, it's substrate-dependent computation, which means the structural form—the morphology—is the structure of the generative model, in the spirit of structural learning. That is very much in the spirit of the good regulator theorem: to be in my world, I have to physically encode and embody a model of the structure of my world.

If you think your world is scale-free, that means my brain must have some scale-free hierarchical aspect, and indeed it does. You could actually write that down and model it with a renormalization group. That's just a reflection of the fact that I am immersed in a world that has some scale invariance in it. So that is certainly true for the plant.

Just one little aside here: that paper was written before David Attenborough's series on the life of plants, where he was able to speed things up by a factor of 10 or 100.

Tim Scarfe

Yes.

Karl Friston

Of course, if you look at plants doing things—sending their little roots and shoots off in a particular direction, or eating insects or whatever—if you speed it up, these things are very animalistic. They look like they could be intelligent, irrespective of whether they're conscious or not. They certainly start to look much more like you and me when you speed things up.

Their morphology matters more than one could possibly imagine, because this is how you become a good regulator. You model your environment. You install that cause-and-effect structure into your computer architecture, which is, again, the argument against von Neumann architectures. That's why one might look to all the processing and memory in neuromorphic photonics, possibly quantum computation. Quantum computation has gone off the boil recently, but all that processing and memory stuff—memristors, for example—I think that's where the answer will be.

Is it in the DNA? Does the DNA cause the structure? So—

Tim Scarfe

Keep your comment about gene expression to make that point.

Karl Friston

You could, if you wanted to simulate these kinds of things, read the DNA as the code that specifies the priors that specify the structure. If you think of DNA as prescribing the structure of your generative model, your world model, or your factor graph, that will be fit for purpose and is learnable.

You start off with, say, DNA. DNA does not tell you what particular kind of plant you're going to be. It's not going to tell you how to engage in your phototaxis and point towards the sun, or compete with other plants that are trying to deny you sunlight. That is something that you have to update and learn during your particular lifetime. But you're equipped with the basic structure—the prior on the structure of your generative model.

And of course, that is most gracefully accommodated, I think, again with respect to the renormalization group. We’re just talking about 2 scales. So you’ve got a slow scale, where you’ve got the viral mutation you mentioned before. The viral DNA—should it have RNA or DNA?—is changing very, very slowly, on a timescale that is greater than or equivalent to the lifespan of any given virus.

That’s, if you like, specifying the initial conditions—the structure for the specification of a particular instance of a virus that lives—

Keith

Well, maybe just jump in for a minute, because I think it’s not even specifying the structure. This is what you pointed out before about the difference between, say, active inference and modeling: DNA is specifying a policy.

Karl Friston

Yes.

Keith

It’s effectively just specifying a policy that every single cell in the organism follows. This policy is not instructions for a structure, but instructions on what to do in response to a certain set of sensory states, right? So it’s almost the policy that the DNA specifies, and this policy applies to every cell. Then, somehow, miraculously, it works out to create a structure that is fit for purpose in its particular environment.

Karl Friston

Yes. And of course, for that specification to work, it has to change at the same rate at which the environment changes, which is very, very slowly. I think that’s absolutely right.

Practically, that hits you in the face when it comes to thinking about how you commit to one policy or another policy. In my world, that would be the expected free energy. The expected free energy comes in 2 parts. It has a sort of expected information gain and the expected cost, or constraints, or utility. That is the point at which you need your DNA. You need to know what it is to be the kind of thing that I am. What do I not do, and what do I do?

Keith

The fascinating thing in organisms is that this is different even on a cell-by-cell level, right? Because morphologically, the same code evolves into all the different organs and parts, roots, bark, and whatever else, which is pretty fascinating.

Karl Friston

Mm-hmm.

Keith

It’s—

Karl Friston

And, yeah, again, it’s exactly the kind of fascinating issue that preoccupies Mike Levin.

Keith

Yeah.

Tim Scarfe

I spoke with our mutual friend Maxwell Ramstead the other day, and he actually gave a wonderful description of the free energy principle. What you want in a good pitch is for it to be replicable and easily understandable, and I’ll try and recapitulate it to you to test how replicable it was.

He said, “Second law of thermodynamics: closed systems. What if we have open systems with boundaries?” He was saying that when things can’t merge together, they instead share information with each other. So there’s this informational synchrony, and that was his way of describing the free energy principle.

First of all, is that a good description? Number 2, where do the boundaries come from? We want to have practical implementations of the free energy principle, and of course we can implement it with computer games or contrived environments where the boundaries are clear. But what if I want to build a robot, and I have a camera, and I see all of these pixels on the screen? I need to divide that scene up into objects so I can start applying the free energy principle. How do we do that?

Karl Friston

Right. 5 minutes, 5 questions. Good.

Keith

It’s got to be practical. That’s why we’re only giving you 5 minutes.

Karl Friston

Right. First of all, the notion of being separate but part of a universe through synchronization, I think, is absolutely correct, and it could be unpacked in 2 ways.

One, you could say that to find a free energy-minimizing solution to your dynamics, namely the path of least action, is just to evince a generalized synchrony between the inside and the outside. Generalized synchrony is also known as synchronization of chaos. Basically, 2 systems that are loosely coupled, in the sense of dynamical systems, will ultimately converge on a synchronization manifold, and they will show chaotic dynamics on that manifold.

That manifold is a synchronization manifold, and being on that manifold is the minimization of free energy. So you can talk about the free energy principle without mentioning self-evidencing, without mentioning Bayes, without mentioning predictive processing, or even extended cognition. You can just talk about it as a variational principle specifying a variational bound on the Lagrangian that specifies generalized synchrony, or synchronization of chaos.

What you’re talking about is exactly this sort of “separate but the same.” I think it’s quite nicely, although I don’t fully understand it, recapitulated in Chris Fields’s quantum treatment. What he would talk about there is not the synchronization between the inside and the outside, or me and you and everything else like me and you, but entanglement.

The principle of unitarity is this—well, you can read the free energy principle as just the principle of unitarity, which means that if we stay here for long enough, for more than 5 minutes, we will ultimately become completely entangled. Classically, that means we will be engaged in a generalized synchrony. There will literally be a synchronization of our itinerant dynamics. We will not become one. Well, actually, Chris Field thinks you would become one.

We will be indistinguishable because we’re just playing out on exactly the same synchronization manifold. So I think that’s absolutely right.

Your last question—and you had about 3 in between—was how you would practically use the free energy principle, and the notion of thingness inherited from Markov blankets, if you’re building robots. I think you’re talking here about physics-discovery or Markov-blanket-discovery algorithms, which, of course, Maxwell and Jeff Becker are furiously working away on.

Tim Scarfe

Yeah.

Assuming you didn’t know anything about that stuff, how would you approach that problem? From first principles, if you had a robot and a camera, how would you partition the scene into boundaries?

Karl Friston

You would be appealing to the notion of Markov boundaries. You can turn it on its head and look at image segmentation as what has emerged in computer vision as a way of identifying things. What you’re doing is committing to an implicit world model, or generative model, for your robot. You’re assuming that it is going to be operating as a good regulator, as a good model of its environment, in an environment that is composed of things.

Tim Scarfe

Yes. But I think my intuition is that we’ve been talking a lot about understanding and creativity recently, and I intuitively feel that understanding is about knowing the history of something. It’s about knowing how you got there. It’s not about knowing the state.

So it’s a little bit different from segmenting an image. To understand the history and the dynamics of a system seems to be much, much more powerful.

Karl Friston

Absolutely. Which is why Markov-blanket discovery is not image segmentation. To do Markov-blanket discovery, you have to look at the dynamics and the history.

Tim Scarfe

Yes.

Karl Friston

There’s quite a fundamental point to be made here. We’ve been talking from the beginning right through to the end about conditional probability distributions. Probability distributions over what? There’s only 1 state of the universe at any one time. There are not many worlds. There’s not an ensemble of universes.

For the purposes of this argument, there’s only 1 state of the universe at any one time. So you can’t have a probability distribution unless you invoke some sort of ensemble or canonical-ensemble assumption, as they do in thermodynamics, where there’s some exchangeability. You can swap around universes and have a distribution.

Keith

Well, you can have an epistemological probability distribution, right?

Karl Friston

I would argue, though, that what you’re implicitly doing is actually doing that over time. There’s a history, a past, and a future. That epistemological notion really has to refer to what is the background over which these multiple realizations occur, from which you can select to build a probability distribution.

Of course, from the point of view of the free energy principle and also building robots, this has to be time. So, by definition, it has to be in the dynamics and the history—not just the short-term dynamics, but recurrence: the characteristic states that I keep returning to as a robot, as a good robot.

So, yeah, absolutely. Segmentation algorithms are a good start, but they’re going to be completely useless when it comes to autonomous vehicles, for example, unless you’ve built in the fact that there is conservation of this particular Markov blanket over time, which, of course, they don’t.

It’s probably better to start with the generative model: this 2-dimensional RGB feed has been generated by Markov blankets that, by definition, persist over time, because time is the support of the conditional distributions that define the existence of the Markov blanket.

Then it’s a question of what kind of priors you put on these things. Are they things? Are they stuff? What’s the Markov blanket of water? Or perhaps fog, if we’re doing autonomous vehicles.

To what extent would I now apply my priors to these things? You can go right through to object-centric priors, of the kind—physics engine-based stuff—that, say, Josh Tenenbaum pursues, assuming that you’ve got some Newtonian behavior. Or you could be much more relaxed. I would relax, but not beyond the renormalization group.

Tim Scarfe

Professor Friston, it’s been an absolute honor. Thank you so much for joining us today. It’s been amazing.

Karl Friston

I really enjoyed it. It’s lovely to speak to you both again.

Keith

Yeah. It’s really been an honor for me because this is the first time I’ve met you in person. I’ve been thinking about you for many years, so it’s been a pleasure to meet you, and thank you for coming to the studio today.

Karl Friston

Well, thank you for coming to England.

Keith

Absolutely.