[BidClub_]
Machine Learning Street Talk · · 82 min

Karl Friston - Why Intelligence Can't Get Too Large (Goldilocks principle)

Tim ScarfeKeithKarl Friston

Podcast
TL;DR
  • Friston’s core call is that intelligence occupies a Goldilocks zone of scale: it can disappear when a system becomes either too small or too large. At tiny scales, randomness overwhelms the recurrent structure needed for agency; at planetary or astronomical scales, averaging erases constituent complexity, leaving motion without planning. An intergalactic civilization may therefore need federation rather than a single intelligence: evolution “doesn’t have to be intelligent to be beautiful.”

  • Bigger organizations do not become smarter merely by aggregating smart parts. Corporations can preserve large-scale intelligence only by imposing structures that balance dissipation with recurrent, conservative dynamics—the organizational equivalent of operating “on the edge of chaos.” Friston’s sharper institutional warning is that globalization and organizations may become “too big for their own good” because they lose the ability to plan.

  • Machine consciousness is possible in principle, but Friston doubts standard von Neumann architecture can deliver it efficiently—or perhaps at all. A conscious artifact would need agency, substantial counterfactual breadth and temporal depth, embodiment as a “beast machine,” and substrate-dependent or “mortal” computation. His directional hardware bet favors processing-in-memory, memristors and neuromorphic photonics—not necessarily spiking neural networks—because the memory substrate must participate and self-organize.

  • Agency and consciousness are orthogonal, although recursive causal structure is used to explain agency and some consciousness proposals. Friston locates agency in a system’s inability to directly observe its own actions, so it infers them as causes of sensation and acquires a “future on the inside” from which it selects paths. Consciousness may require additional precision control, attention and metacognition—“looking at the looking at the looking”—rather than merely having Bayesian beliefs.

  • After roughly 20 years, Friston says he has not encountered the decisive “that doesn’t work then” moment for the free energy principle. Its core is almost minimal: partitions create conditional probability distributions, whose dynamics follow a principle of least action. The difficulty is communicating something “almost tautologically simple” whose consequences are profound; Friston remains ambivalent because a little “magic and mysticism” also motivates inquiry.

  • The episode rejects both brain chauvinism and indiscriminate pan-intelligence. Mike Levin’s basal-cognition program treats intelligence in viruses, slime molds and xenobots as an empirical question, while Keith argues that a virus is essentially a “mousetrap” outsourcing its machinery to a host and doing no lifetime learning. The useful unit may instead be a colony or ecosystem—but Friston still requires the right causal structure, counterfactual depth and scale.

  • For autonomous systems, static perception is insufficient because meaningful boundaries are histories, not just pixels. Markov-blanket discovery must identify entities that persist and recur over time; image segmentation alone cannot supply understanding. That makes structural learning and dynamic world models central to robotics and autonomous vehicles, while Friston concedes that action at a distance still leaves him without “good maths” for choosing between competing causal structures.

Digest · the substance, structured for research

1. Twenty years have strengthened, not settled, Friston’s conviction

  • Keith dates the free energy principle to around 2005 and asks for a twenty-year audit. Friston resists a victory lap: assessing it while developing it is like reviewing a meal while cooking, but every phenomenon that “slip[s] in quite neatly” produces a dopamine hit—and he has yet to reach the decisive “that doesn’t work then” moment.

  • Friston’s disappointment is communicative. He is ambivalent when the principle is called notoriously difficult: some “magic and mysticism” gives people a challenge worth mastering, yet the framework is meant to be “almost tautologically simple,” and he suspects his own explanations have not kept it sufficiently intuitive.

  • Keith compares it with probability theory: the sum and product rules are easy to write down, yet their consequences are profound and routinely misapplied. Friston agrees because conditional probability is the framework’s heart: partition states, condition one set upon another, then describe a principle of least action for the dynamics of those conditional densities—“before thermodynamics” and “before quantum mechanics.”

  • Friston sees generative AI and large language models primarily as engineering resources for understanding natural intelligence. That understanding could make life better, support sustainability and serve mental wellbeing through computational psychiatry; invoking Feynman—“That which I cannot create, I do not understand”—he adds the uncomfortable implication that genuine understanding may require creating “things that suffer.”

2. Markov blankets turn missing causal links into natural kinds

  • Friston begins with external, internal, sensory and active states. Different forbidden causal influences produce an ontology of “natural kinds”: an entity with neither internal nor active states is a causal black hole, quintessentially inert and effectively invisible because it never acts back upon its environment.

  • In an ordinary touch-bound system such as a cell, sensory states form the outside layer, active states sit beneath them, and internal states occupy the center. The arrangement enforces both directions of conditional independence: external states cannot directly reach action, while internal states cannot directly cause sensation.

  • Hierarchical organisms are different because their active states become sequestered from internal states. From the inside, unseen actions now appear as causes of sensation alongside the external world; the organism must therefore infer not only its environment but itself. That recursion is what makes these systems “strange things.”

  • Action at a distance complicates the picture. A bacterium inhabits a touch-local world, but vision, hearing, electromagnetic radiation and magnetic fields—as illustrated by the compass discussed—create nonlocal causal possibilities. Friston’s candid gap is structural learning: he knows of no good mathematics for deciding how best to disambiguate between competing causal organizations.

3. Self-modeling gives an agent a private future

  • Keith’s framing is that computational strange loops unlock an entirely different category of behavior. Friston agrees: once a system models itself as a cause of its environment, it must infer what it is doing and what those actions will produce; that is an “authentic kind of agency,” not merely reactive regulation.

  • Because an action’s consequences have not yet arrived, self-modeling is intrinsically future-pointing. The transition from thermostat, virus or single cell to agent occurs when a system has a “future on the inside,” represents divergent possible paths, and selects among them—arguably necessary, Friston says, for true agency and possibly elemental consciousness, though not a definition of either.

  • Tim presses the apparent conflation of agency with phenomenal experience: deeper reflexive layers plausibly strengthen agency, but why should they create consciousness? Friston concedes the point completely, aligning with Anil Seth’s distinction: intelligence, agency and consciousness are orthogonal, and “consciousness is not implied by being agentic.”

4. Consciousness may arise from precision, access and recursive attention

  • The free-energy treatment offers dual-aspect monism rather than a direct solution to consciousness. The same brain substrate carries a thermodynamic information geometry and a representational geometry in which internal states encode Bayesian beliefs about the world, including action; physical dynamics and inference are two readings of one process.

  • Merely encoding a posterior belief may not suffice. A thermostat can be interpreted as holding beliefs about temperature without thereby being conscious, so stronger accounts require beliefs to become dynamically “ignited”—sufficiently precise to affect updating elsewhere in a hierarchical or distributed system.

  • Friston invokes a “felt uncertainty” account associated in his remarks with Mark Solms: consciousness tracks representations of confidence, precision or uncertainty, not simply content. He cautiously recalls the claim that consciousness can be switched off through a brain-stem region roughly three millimeters cubed, containing the origins of ascending systems that encode and represent those precision signals.

  • More demanding theories add control and self-recognition: attend, recognize that you are attending, then recognize that recognition. Chris Fields’s inner-screen hypothesis recasts nested Markov blankets as holographic screens and posits an irreducible inner screen that knows itself only by acting on the rest of the brain—compatible, Friston thinks, with higher-order thought and precision-gated global neuronal workspace theories.

5. Machine consciousness requires long horizons and mortal hardware

  • Keith argues for a minimum complexity and causal structure, comparing consciousness with a Game of Life glider that cannot exist below some number of pixels. Friston agrees, while treating the boundary as technically vague—like asking how many grains make a pile—and nominates temporal depth as the most plausible dimension.

  • Counterfactual breadth is the number of available future paths; counterfactual or temporal depth is how far those paths extend. A thermostat has only the instantaneous future implicit in a differential equation. A machine with human-like agency would need a world model of its actions’ consequences with both wide alternatives and a substantially longer horizon.

  • Friston’s answer to whether machines could eventually understand or become conscious is “in principle, yes,” but mathematical simulation alone may not suffice. Following Seth’s “beast machine” framing and Geoffrey Hinton’s mortal computation, he argues for embodiment and substrate dependence: the computation must depend on its physical substrate rather than merely being implemented independently of it.

  • He therefore doubts von Neumann systems, where processors read and write separate memory, because that blanket structure makes memory difficult to self-organize. Friston says a conscious machine would have to pursue a path of least action thermodynamically and informationally; otherwise it could not be conscious for a nontrivial amount of time. He points instead toward processing-in-memory, memristors and neuromorphic photonics, while explicitly saying spiking neural networks are not required.

6. Basal cognition is empirical, but viruses expose the boundary problem

  • Tim asks whether plants and viruses imply “pan-intelligence,” challenging the privileged status given to brains. Friston directs the question toward Mike Levin and Chris Fields’s basal-cognition program: intelligence may be widespread, but experiments must be designed to reveal it rather than assumed from appearance.

  • By many adaptive-behavior yardsticks, Friston says, slime molds, viruses and xenobots can look “incredibly intelligent.” He is sympathetic to that attack on chauvinism, yet insists that a virus lacks the expressive agency of a strange thing; some causal structures are categorically present or absent, even if the broader intelligence threshold remains vague.

  • Keith’s pushback is biological outsourcing. A virus is mutated by radiation or transcription error rather than modifying itself, behaves like a “mousetrap” that injects genetic material, and leaves transcription and the rest of the work to the host cell. Over its own lifetime it neither learns nor adapts, making “intelligence” less useful than ordinary unfolding dynamics.

  • Scale partly reconciles them: the individual virus may be uninteresting while a colony or ecosystem becomes the relevant unit. Friston similarly calls natural selection Bayesian model selection—favoring phenotypes with the greatest model evidence—without equating adaptiveness with planning. Keith’s stricter rule assigns intelligence to the smallest sufficient boundary, not everything containing it.

7. Intelligence vanishes below and above a Goldilocks scale

  • Friston proposes evaluating each candidate at the scale of its Markov blanket: does it contain enough machinery for a world model with meaningful counterfactual depth and breadth? A virus may be too small; a biosphere may be too large because coarse observation averages away the rich causal organization of its constituents.

  • The physical mechanism begins with the Helmholtz decomposition of dynamics into dissipative and conservative parts. Dissipative fluctuations supply randomness and openness; conservative or solenoidal flows cycle through states, supporting routines, biorhythms, reproduction and Poincaré recurrence—the repeated return near a starting point that lets a system maintain a recognizable identity.

  • At very small scales, Friston argues, quantum randomness dominates and recurrent conservative organization becomes insufficient. Human-scale organisms mix an itinerant, changing world with stable revisitation of characteristic states. At astronomical scales, averaging removes the fluctuations, leaving predominantly conservative Newtonian motion: planets orbit, but the motion contains no evidence of planning.

  • This is the Goldilocks regime “on the edge of chaos,” where neither complete order nor complete noise wins. Friston says he does not see the Moon, weather or evolution thinking about their futures. When extending the argument toward an ultimate scale limit, however, he hedges: “I’m not sure this is correct.”

8. Coarse-graining creates new intelligence without inventing new matter

  • Tim asks whether the apparent unintelligence of Gaia merely reflects observer limitations. Humans abstract by ignoring detail and idealize by distorting it; if intelligence is “doing more with less” and emergence means “more is different,” then a higher-scale system may be genuinely reorganized rather than merely a blurry view of its components.

  • Friston’s mathematical answer is the renormalization group: an operator reduces dimensions and groups fine-scale states into the appropriate higher-scale variables, recursively. Higher levels can display brand-new self-evidencing or intelligent dynamics, yet those variables remain functions of finer-scale activity—real emergence without adding an unexplained substance.

  • Keith argues—and Friston agrees—that large-scale intelligence such as a corporation requires deliberately installed organizational structure. The corporation’s coarse metric is survival—whether it progresses from “Series A to Series whatever” in recognizable form—but increasing size leaves less room for rich recursive organization, making planning progressively harder.

  • Keith extends the limit to intergalactic civilization, where the speed of light eventually blocks centralized coordination across light-years. Friston says the physics arguments suggest a Goldilocks zone, then reframes the disappointment: federated ecosystems can still be beautiful. “Evolution is a beautiful thing,” and “it doesn’t have to be intelligent to be beautiful.”

9. Morphology embodies the model while DNA constrains the policy

  • Tim’s extended-cognition challenge asks where thinking ends: inside the head, in a phone, across a plant’s morphology, or throughout a bidirectional causal process? Friston answers that discussion requires a Markov boundary, but every bounded thing is contextualized by the scale above; even an unintelligent virus requires a host world conducive to its persistence.

  • For plants, morphology is not merely what an internal model represents. Physical structure parameterizes conditional distributions, making “the substrate” part of the generative model itself. Under the good-regulator idea, an organism must embody the causal organization of its environment; a scale-free world should therefore be matched by scale-free hierarchy in the organism.

  • Friston’s best specimen is David Attenborough’s plant footage accelerated by a factor of 10 or 100. Roots and shoots move in particular directions, plants compete with other plants for sunlight and some consume insects; sped into a human temporal register, their behavior looks strikingly animal-like and difficult to dismiss as unintelligent, irrespective of consciousness.

  • Tim initially raises DNA as a possible software-like model, then Keith sharpens the point: DNA specifies a policy—what each cell does given sensory states—not a completed body plan. Friston agrees that inherited code supplies slowly changing structural priors and expected constraints, while each organism must learn particulars such as where to grow; the same code can differentiate into roots, bark and organs.

10. Practical intelligence discovery must follow dynamics through time

  • Tim relays Maxwell Ramstead’s compressed account: bounded open systems that cannot merge instead exchange information and synchronize. Friston endorses it—loosely coupled systems converge on a chaotic synchronization manifold, and free-energy minimization can be described as generalized synchrony without invoking Bayes, predictive processing or self-evidencing.

  • Chris Fields’s quantum version frames the counterpart in terms of entanglement and unitarity. Friston jokes that after “more than five minutes” together, the participants will become completely entangled: perhaps not literally one system, though Fields might say so, but increasingly indistinguishable as their dynamics occupy the same manifold.

  • For a camera-equipped robot, image segmentation is only a starting commitment that the environment contains things. Tim argues that understanding requires history, not a snapshot; Friston agrees because Markov-blanket discovery must use dynamics. With only one universe-state at a time, the repeated realizations supporting a probability distribution arise across past and future time.

  • Autonomous vehicles therefore need to conserve inferred blankets across time and distinguish persistent things from amorphous stuff such as water or fog. Object-centric Newtonian priors of the kind Josh Tenenbaum studies are one option; broader models are possible, but Friston’s boundary is memorable: “I would relax, but not beyond the renormalization group.”

Tim Scarfe

Keith has been talking about this moment where we have a glass of sherry with Professor Friston for about the last 4 years.

Keith

Right.

Tim Scarfe

And our dream has been realized today.

Keith

Absolutely. Such a pleasure to meet you. It's absolutely been a pleasure.

Karl Friston

Cheers.

Keith

Cheers.

Karl Friston

You mentioned consciousness, which you shouldn't really do with me, but if we stay here for long enough—more than 5 minutes—we will ultimately become completely entangled. We will not become one. Well, actually, Chris Fields thinks you would become one. Is evolution an intelligent process? It's certainly a free energy-minimizing process. So it's just Bayesian model selection.

Keith

Is the planet intelligent?

Karl Friston

Yeah.

Keith

Yeah.

Karl Friston

I don't see the weather planning. I don't see evolution planning. It doesn't think about its future.

Tim Scarfe

Life is the intensive property of matter, and intelligence is the extensive property.

Karl Friston

There's a large ensemble of universes. Well, of course, for the purpose of this argument, there's only 1 state of the universe at any one time. I'm not sure this is correct. I don't want to disappoint you.

Keith

Don't worry, it's well beyond my lifetime when that will happen.

Karl Friston

If you're going to make a difference, it's really about understanding how to use all the marvelous engineering that we've witnessed in terms of generative AI and large language models and the like—

Keith

Mm.

Karl Friston

—in the service of understanding the principles that underwrite natural intelligence—

Keith

Mm.

Karl Friston

—and deploying that understanding either to make life better—

Keith

Mm.

Karl Friston

—or to underwrite sustainability in a slightly more political way. Or to deploy it—which is why I got into this game—in the context of mental well-being and computational psychiatry. To quote Feynman, “That which I cannot create, I do not understand.”

Keith

Yeah.

Karl Friston

That means that we have to be able to create—

Keith

Yeah.

Karl Friston

—things that suffer.

Keith

Yeah.

Karl Friston

Which means that we need to understand the principles of natural intelligence.

Tim Scarfe

I don't know if you saw, I interviewed Yoshua Bengio.

Karl Friston

All right.

Tim Scarfe

I used the term “epistemic foraging” a few times, and he said, “I love that term. I love that term. Where did it come from?” Professor Friston?

Keith

Absolutely. Epistemic foraging is what we do. It's what we do.

Tim Scarfe

Yeah. To epistemic foraging.

Karl Friston

To epistemic foraging.

Tim Scarfe

It's just wonderful.

Karl Friston

There are certain causal forces in our universe that are at a distance, usually mediated by electromagnetic radiation or magnetic fields, in this particular instance. That really complicates the way in which we construe all the cause-and-effect structures that we have to model in our brain. If I was a simple little virus or a little bacterium, I wouldn't have to worry about electric fields, seeing things, or hearing anything. My world would just be that physical world that I could touch, literally. It would just be my next-door neighbors.

That leads to a very particular kind of Markov blanket structure and ecosystem of Markov blankets that precludes, or certainly does not license, some of the deep structures that we were talking about earlier in terms of strange things. The strangeness of a beautiful sort—beautiful loops, to use one of my colleagues' notions—rests upon action at a distance of a very non-spoofy sort that is nicely exemplified by this compass. At the moment, I don't know of any good maths that allows you to work out what is the best thing to do to disambiguate between this structure and that structure.

Tim Scarfe

Mm. Yeah.

Karl Friston

So if you can do that, that'd be good.

Tim Scarfe

Structural learning.

Karl Friston

Structural learning.

Tim Scarfe

To structural learning.

Karl Friston

Structural learning.

Keith

Somebody out there, help us out.

Tim Scarfe

Professor Friston, it is absolutely amazing to have you back on MLST. As you know, Keith and I are incredibly fond of you. You've been a big hero of ours for many years. I think this is the 4th time—

Keith

That you've been on the show?

Tim Scarfe

Yeah. I believe so.

Karl Friston

Yeah.

Tim Scarfe

Welcome, Professor Friston.

Karl Friston

I've lost count. But it's lovely to see you both in person, and congratulations on your recent marriage.

Tim Scarfe

Oh, thank you very much. Thank you very much.

1. The Free Energy Retrospective

Keith

So, we were just thinking about this. We've done these interviews. The free energy principle originated around about 2005, right? I wanted to get a bit of a retrospective from you on, let's say, what have been the successes so far of the free energy principle. How has it been going? What could have gone better over the last 20-odd years? Really, what's your take on the progress, I guess, of the free energy principle?

Karl Friston

It's a difficult question to answer, in the sense that when you're doing it, you're in the middle of it, and it's very difficult to assess. If you asked me if I'd made a meal and then asked me, “How's that meal going? Did you really enjoy it?”—when you're doing it, you're in the middle of it, and it's very difficult to assess.

So it's been a journey. I do catch myself from time to time thinking, “This is the right way to think about things.” Whenever you come across a phenomenon or a problem or an application that seems to just slip in quite neatly to the overall theoretical framework, every time that happens, I get a little buzz of dopamine, and that consolidates and comes with, “Yeah, I'm on the right track. I haven't wasted so much time.” There have been times in my life when I've realized I'm on the wrong track, so I'm used to that.

But for the particular application of the free energy principle, I have yet to have that: “Oh, well, that doesn't work, then. Think about it; do something else.”

What could have gone better? I'm ambivalent about this when I read that the free energy principle is notoriously difficult to understand. Now, half of me thinks, “Good,” because it's always important to have a slight degree of magic and mysticism to engage people. If people don't think there's a challenge, or that they've accomplished something by understanding it, then they're not going to be motivated to inquire, think about it, or indeed debate it.

On the other hand, the free energy principle is not meant to be complicated or difficult to understand. It's actually almost tautologically simple. Being able to communicate the free energy principle in a way that people find it a useful and obvious tool or method to apply could have, I think, gone better. That may be because I'm not very good at keeping things simple or intuitive.

Keith

We can attest to that.

Karl Friston

Yeah.

Keith

One thing that's interesting is I feel the same way about probability theory. Conditional probability theory is very easy to write down. You end up with the 2 fundamental rules, the sum rule and the product rule, but the consequences of it are very profound, and probability is notoriously misunderstood and misapplied. There are many errors of reasoning that happen when people try to reason through statistical correlations or whatever. So I think there are things like that in life that are very simple and yet very hard to understand the full extent of their impact, right?

Karl Friston

Yes.

I mean, it's interesting you picked up on conditional probabilities. That is the heart of the free energy principle at many different levels.

The classical formulation of the free energy principle starts off with a partition of different states of being. That partition suddenly means you've got the probability distribution or density over one set of states of being that can now be conditioned upon another.

The whole free energy principle is basically a principle of least action pertaining to density dynamics: the dynamics or evolution not of densities, but of conditional densities. That's it.

Keith

Mm-hmm.

Karl Friston

You know, this is before thermodynamics. It's before quantum mechanics. It's just about conditional probability distributions. So it's interesting you picked up on that as something that is so simple, yet so absolutely powerful.

2. Strange Things Become Agents

Tim Scarfe

Professor Friston, we were reading your paper earlier. It's from 2023, “Path integrals, particular kinds, and strange things.” It introduced a categorization of particles, from inert to active and ordinary to strange. The strange categorization was particularly interesting because you were giving an account of how these particular particles could give rise to phenomenal, conscious states or even agentic states.

We felt that this was almost an account of pan-agentialism. Historically, we've done interviews about the free energy principle, and we've spoken about self-organization and emergence. Agency and phenomenal states are things that emerge when you have this temporal and counterfactual depth. Maybe we're misunderstanding, but it feels like this is an account of low-level agentic potential and phenomenal potential.

Karl Friston

To set the scene, just to come back to that partition that we were talking about, there are many ways of arranging the conditional independencies to disconnect various partitions: external states, internal states, sensory states, and active states. The combinations of disallowed influences give rise to an ontology of different kinds of things, and you could call them natural kinds.

I've been told I shouldn't use that, but I like “natural kinds”: the things in nature that prescribe themselves just by being special instances of a lack of causal influence.

If we start off with the simplest case, where there are no internal states and no active states, what are we talking about? We're talking about some kind of causal black hole, something that is quintessentially inert. You can never see it because you can only see the active states that act back on the environment in which it is embedded.

But then we get to more interesting things that have a full complement of internal and external states, and a bidirectional coupling between the two, which is mediated by the sensory states and the active states. If you disallow action at a distance—in other words, I have to be, in some metric sense, next to you in order for my active states to become your sensory states, and your sensory states to become my active states—then you have a very simple kind of Markov blanket.

That would be fine for describing active matter, for example, or any medium where I have to touch to feel, to influence, or to sense. Interestingly, this kind of particle has its active states on the inside, which I had to think about. It can be no other way if you just try to arrange all the dependencies.

So what would that look like? It would look like something like a cell that had, on the outside, sensory states supported by a layer of active states that surrounded the internal state. You're complying with the laws of the Markov blanket of thingness: you need to have that conditional independence to separate the thing from everything else, or the self from the non-self.

Everything's fine because the active states are hiding behind the sensory states, so the external states can't influence the active state. So that's tick one. That's what we need.

But we also have to ensure that the internal states, symmetrically, cannot act to influence the sensory states, and that's fine because they're hiding behind the active states. So you've got a nice mathematical image of a cell, if you like.

But that doesn't work for things like you and me. Things like you and me have a hierarchical structure. What that basically means is that the active states which, in very simple organisms, were immediately juxtaposed to the internal states now become sequestered.

Now, effectively, the active states are no longer seen by the internal states, and then something quite remarkable happens mathematically. It's really simple, and it's just another aspect of conditional probability distributions. Because you can't see your active states, you can now only see your sensory states. It looks as if, from the point of view of the internal states, the active states have now become causes of sensory states.

So now my world is caused not just by the external world, by the environmental states, by my heat path, by my external milieu, but also by my own actions. Now I'm inferring the causes of my sensorium, where I'm actually a cause, so I'm inferring myself.

There's this beautiful recursion which licenses this strange-loop analogy, with a bit of poetic license. Another way of putting that is that, for these simple structures—for these natural kinds that would be, say, single-celled organisms—the internal states can be read as modeling the causes of their sensations, which just are the external states.

They have direct access to the active state, so there's no conditional independence that licenses the notion of description in terms of inference or sense-making. Whereas now, the Bayesian mechanics that attends the internal activity, the internal machinations, covers both my action and the things that I'm acting upon.

Of course, this is a nice metaphor for planning as inference: I'm now thinking about and trying to infer what I'm actually doing. That, to my mind, produces a very unique and special kind of thing—things like you and me, basically—which are pretty unique when we look at all the different kinds of things that could be around.

Keith

Well, it's another example of how, as soon as you had this recursion—as soon as you had these strange loops, especially for systems that are computational.

I forget who it was that said, “Life is the computational phase of matter.” When you're at this level of complexity, where you're having some kind of information processing, and you introduce this strange loop, that unlocks a completely different category of behavior, a different tier of computation.

Karl Friston

Absolutely.

Tim Scarfe

It was David Krakauer, and he said, “Life is the intensive property of matter, and intelligence is the extensive property.”

Keith

No, this is a different one, though. It may have been Wolfram who said, “Life is a computation.” I don't remember. I'd have to look it up. But, yeah, it unlocks a completely different category of behavior.

Karl Friston

Yeah. Well, you alluded to that earlier in terms of phenomenology, and you mentioned consciousness, which you shouldn't really do with me. We'll read that as an elementary kind of sentience.

I think that's really important: that you do unlock or manifest, or now at least have a mathematical calculus that allows you to talk about things that model themselves. As soon as you're modeling yourself, you become an agent, or you have an authentic kind of agency.

In order to model myself, or to model the me as a cause of my environment, I have to infer what I'm doing, in particular, the consequences of what I'm doing. Because the consequences have not yet appeared, this has a quintessentially future-pointing aspect.

So now we're moving from a thermostat, a virus, or a single-celled organism to things that actually have a future on the inside—their private future that just exists for them. Of course, if you can get suitably far into the future, where you have a divergence of particular paths into the future, just as a nod to the path-integral formulation, then you have to select one.

So now you've really unlocked, again, a calculus of selection of your paths into the future, which I think is not a definition of agency and certainly not consciousness, but would arguably be necessary to support something that had true agency and possibly even consciousness of an elemental kind.

3. Consciousness Beyond Agency

Tim Scarfe

Could we press on consciousness a bit? The account that you just gave is a beautiful account of agency, and it's very plausible. As you increase this reflexive layering, this recursion, you get these multiple paths into the future, and you become the cause of your own actions. So you become more of an agent. It's almost an account of the strength of agency, with more and more layers.

But in the abstract of this paper, we were quite struck by the language that conflated consciousness with agency. Intuitively, I feel that phenomenal experience is quite orthogonal to agency. Why were they convolved together?

Karl Friston

I'm desperately trying to remember how I referred to consciousness. You've got to be very careful when using that word, depending on who you're talking to.

Tim Scarfe

Yes, the C-word.

Karl Friston

The C-word. Oh, naughty.

I think people like Anil Seth make a very similar point quite earnestly: intelligence and consciousness are completely orthogonal, and you're making agency and consciousness completely orthogonal. I would agree entirely. Consciousness isn't implied by being agentic, by having agency.

There are many stories you can tell here, and my mind goes to the people who might be watching this to make sure I don't offend anybody by not mentioning them. From the point of view of the free-energy principle, the way that you'd look at consciousness is in terms of a dual-aspect monism, which you get for free from the treatment of the density dynamics in exactly the way we were talking about before.

Now my internal brain states have a thermodynamics. That can be written in terms of an information geometry that itself is just predicated on conditional probability distributions. But they also have an information geometry that inherits from the fact that my internal brain states represent the outside world, including my own actions. So there's both an information geometry and, by implication, a thermodynamics of my brain activity, which supervenes on exactly the same substrate as does the information geometry, which is representational—the inference, the Bayesian mechanics side of things. That gives you license to talk about a sort of dual-aspect monism.

But what does it really mean to be conscious? Some people might believe that it's just having a posterior belief, having a Bayesian belief, which is literally just a conditional probability distribution that is encoded or parameterized by some physical state of being. In this instance, under the free-energy principle, it's the internal states of being.

Other people say, “Well, no, just being able to read a thermostat, for example, as having posterior beliefs about the temperature of the external world does not license you to ascribe consciousness to this thing.” So the next step would be: it has to be ignited in some way. It has to be realized. It has to be emergent.

I'm sort of paraphrasing Jakob Powrie here and, to a certain extent, Mark Solms. They emphasize that it is when these beliefs become sufficiently precise in a dynamical way that they are realized and influence other aspects of belief updating in a hierarchical or distributed system. Mark Soames would call this felt uncertainty. He would emphasize that the feeling part of consciousness is mediated by representations not of the content, but of the precision, uncertainty, or confidence with which these beliefs are currently in operation.

He will bring to the table all sorts of very compelling neurobiological evidence as to why this is. You can switch off consciousness literally by, I think he says, a 3-millimeter cube of the brainstem containing the cells of origin of those ascending systems that encode and represent—not the content, but the confidence, uncertainty, or precision of these conditional probability distributions.

Jakob, I think, would say that to be aware is to equip all those messages that are providing evidence for your current explanation with precision, which has physiological and bioelectric mechanisms behind it. Other people go further.

People in the world of phenomenology, of the kind that Thomas Metzinger pursues, would say that it is necessary to have control over the precision and, furthermore, to be aware that you are rendering that which was once transparent opaque. You actually have to recognize that you're attending.

So, again, you not only have this very simple recursion of modeling and self-modeling in terms of planning as inference, but now you've got this hierarchical recursion where you're recognizing that you're attending to something. If you're somebody like Lars Sunved Smith, you'd say that to actually be self-conscious, I now have to recognize that I was recognizing that I was attending to something. You get layer upon layer upon layer.

To join the dots with Chris Fields, who has brought to the table the inner-screen hypothesis—have you come across this yet?

Tim Scarfe

No. No, tell us about that.

Karl Friston

Strictly speaking, you should get Chris to tell you about that. But I'll give you a quick preamble.

Chris Fields is the theoretician who has formulated the quantum information-theoretic version of the free-energy principle. For him, the Markov blanket that separates self from non-self becomes a holographic screen upon which classical information is written and from which it is read by some internal bulk and some external bulk.

In the sense that strange loops and strange recursions are only allowable when you have this hierarchical structure on the inside, you've now got lots of inner holographic screens, inner Markov blankets. Markov blankets, in their pragmatic sense, just define things like hierarchies, for example, and they define the architecture of any network, factor graph, or computer. When you've got lots of them, you've effectively got lots of inner screens.

My reading of that idea is that there is 1 irreducible inner screen—1 inner screen that within it has no other screens. This is an interesting screen, a Markov blanket from the classical perspective, because the only way that the internal states of this irreducible Markov blanket can know themselves is by acting on the exterior. Of course, the exterior is the rest of your brain. So you've got this metacognitive, self-recursive aspect to the inner-screen hypothesis.

That notion has a lot of mileage in relation to other people's theories of consciousness. We're talking about higher-order thought theory, so the very notion of higher-order thought and its implicit appeal to metacognition, meta-metacognition, and meta-meta-metacognition is again this notion of looking at oneself and inferring oneself, but on the inside. Looking at the looking at the looking. This is, of course, a natural consequence of having this hierarchical structure.

You could also argue that it's completely compatible with global neuronal workspace theories: there is this precision-dependent ignition of certain sources or messages—sufficient statistics of conditional probability distributions—that gain access to all of these inner screens in a dynamic way because you're controlling access by attending, by encoding and representing, and by optimizing the precision, uncertainty, or confidence.

So you've got this notion of ignition and penetration through to the global workspace, which we could read as this sort of irreducible inner screen: the core, the deepest part of your sense-making, your brain. I think this notion is not only consistent with the conditional independencies that are implicit in the free-energy principle as applied to strange things that have this hierarchical and recursive aspect, but also relates comfortably to extant theories, all coming from different perspectives in different parts of the elephant, as it were.

It's a comfortable accommodation of things that we presume, or things that people have brought to the table in order to explain consciousness from their direction of theorizing.

Keith

It seems that, if we split the camps—or, let's say, the thought groups—that thought about this, almost all of what you just talked about now are accounts in which there is a minimal complexity required for there to be consciousness. Then we can talk about what that level is.

Well, it requires this or this or this. But then there are also people who say, “No, there’s no minimal complexity.” Electrons have some small amount of consciousness. I fall much more into the category of people who say there is a minimum—you have to reach a certain something.

I don’t really know where that boundary is. Maybe it’s one of the things we’ve mentioned. Maybe we’ll figure out a more elegant categorization, but it’s just like needing a certain number of pixels in the Game of Life before you can have a glider. You can’t do it with 3. You need some minimum number, and so there’s some minimum amount of brain material. Three cubic millimeters is still a lot of neurons. I don’t know how many are in there, but hundreds, thousands, or whatever it is.

You need some level of complexity before you have consciousness, or before you have recursive agency, intelligence, or whatever. Would you agree with that? There is some minimum. We don’t know where it is yet; it’s hard to place, but there is a minimum?

Karl Friston

No, I would agree entirely.

Keith

Okay. So there is a minimum, and not only that, it has to have a certain causal structure, right? We can debate that, and that’s really in line with Searle in the Chinese room argument, which is: a dictionary doesn’t have understanding because it doesn’t have the right causal structure. You have to have a certain causal structure or a certain minimum complexity, and then you reach this—whatever it is, whether we’re talking about consciousness, understanding, agency, or all of these things.

4. Building Conscious Machines

So I guess my question to you is: will we be able to build machines based on our current computer architectures someday, whether it’s 100 years from now or 200 years from now? It doesn’t matter. In principle, can we build machines that have understanding, consciousness, and all these capabilities?

Karl Friston

Yes, that’s a good question, and the answer, I think, is yes. Can I come back to qualify that answer by reference to that wonderful example about what I would read as vagueness in a technical sense? Vagueness: how many grains of sand constitute a pile?

Keith

Right. Yeah.

Karl Friston

It’s not well defined. It is in philosophy, but not in mathematics. I think that’s absolutely the right way to think about these bright lines. If I had to commit to the dimension—the number of grains of sand that you get before you have consciousness, or there is a pile—I would say it’s something you actually referred to earlier on. I think it’s the depth of your future, or your future in your head.

If you’re talking now about an algorithm or some artificial intelligence that is equipped with a generative or world model of the consequences of its actions, there will be a time horizon associated with that component of its generative model. I think it’s the depth of that time horizon—the thing that Anil Seth would refer to, not as counterfactual breadth, which is the number of divergent paths one could take into the future, the options that you select among, but counterfactual depth.

Keith

Okay.

Karl Friston

The temporal depth. That means that you can be panpsychic. You can say that a thermostat has a notion of the future, in the sense that it operates through path-integral control and differential equations. As soon as you put a differential equation in play, you’ve got an instantaneous future because you’ve got some gradient with respect to time.

That’s not, though, the kind of depth that you and I enjoy. I would imagine it is really just the depth. Maxwell Ramstead talks about this as merely reflexive active inference, with very myopic, very short-term self-models, right through to fully—well, to be conscious in the way that we’ve been talking about, I think you’d need to have a long depth.

So what does that mean for building AGI, conscious artifacts, or machine consciousness? First of all, they have to be agentic, because we’ve just said that having a world model of the consequences of your actions would be necessary to be an agent. More than that, you’d have to look quite a long way into the future.

Your generative model, your world model, would have to go quite a long way into the future. I repeat: it would have to have both counterfactual breadth and counterfactual depth at hand. Would that be sufficient? I’m now remembering that I forgot to mention Anil Seth in the list of people not to upset when reviewing theories of consciousness, so this is an opportunity just to say that there are people out there who would say, “Well, okay, you can write down the maths of all this, and you can write down in silico hypotheses, or indeed simulate global neuronal workspace theories and try to produce things that look as if they have consciousness.”

That’s not going to work unless you actually embody it, unless you are, in Anil’s words, a beast machine. I think that coheres with the argument for mortal computation. So when you ask the question, “Can we build it on our computer architectures?” I would have to ask you: do you mean a von Neumann architecture, or do you mean a memory-processing or in-memory-processing architecture?

I would take that as synonymous with a neuromorphic architecture—not spiking neural networks. You don’t need those, but you do need processing in memory to be mortal. You need that substrate dependence, read in terms of Geoffrey Hinton’s definition of mortal computation and Alex’s subsequent elaborations of that. I think I would subscribe to that, largely to keep Anil happy.

I don’t think you can do this on a von Neumann architecture, because the Markov blankets of a von Neumann architecture, where you’re reading and writing from memory, make it very difficult for the memory to self-organize.

Tim Scarfe

I see.

Karl Friston

People have written about this philosophically. I think Vanya Weiss has written a paper about this, and I think Anil Seth speaks to this argument in a recent Behavioral and Brain Sciences paper. It may not be possible, and certainly, from the point of view of efficiency, it’s highly unlikely that von Neumann architectures are the kind of things that would conform to—

Let me reverse. To be is to pursue a path of least action in accordance with the free energy principle. To be conscious does not excuse you from that. If you want a conscious machine, you have to have a machine that pursues a path of least action.

Via the Janiskee quality, that has to be true both thermodynamically and informationally, in terms of the conditional probability distributions. If you don’t pursue that path of least action, you can’t be conscious for a nontrivial amount of time. I can see you want to ask a question, so I’ll interject.

Tim Scarfe

Well, only to say that Maxwell pointed me to Anil’s work, where he defends biological naturalism. Of course, even Searle said that the biological substrate is an existence proof. It is not saying that it has to be biology. But Keith and I certainly agree that there’s something about the substrate that is very important.

5. Intelligence Beyond Brains

I wanted to talk a little bit about viruses. As you know, I’m a bit of an externalist. I’ve never been able to completely pin you down, Professor Friston, because there have been so many interpretations of the free energy principle that lean internalist and externalist, and even the hybrid version, which Maxwell also wrote a paper about.

But I’m fascinated by this idea of diverse intelligences. For example, could a virus be intelligent? I spoke with David Krakauer, and he said intelligent things do inference, have representations, and are adaptable.

And we're a little bit chauvinistic about our brains, aren't we? Our brains seem to have a privileged status. But what say you of viruses? Or, actually, we read your paper with Calvo, “Predicting Green: Really Radical (Plant) Predictive Processing.” In that paper, you gave a beautiful account of how even plants could be doing inferencing.

And that, to me, seems incredible because I'm amenable to the idea. But it seems to me intuitively that plants are not as sophisticated as we are. So it comes back to this line that Keith was talking about before: if you go too far down the stack, it's an account of—we'll call it pan-intelligence—where you have a tiny little bit of intelligence even if you go all the way down the stack.

Karl Friston

Have you spoken to Mike Levin?

Tim Scarfe

Yes, relatively recently, about a year ago.

Karl Friston

Right.

Tim Scarfe

Yes.

Karl Friston

I mean, it would be nice to revisit him. That is exactly his big question at the moment. And, of course, he has, as an intellectual accomplice, Chris Fields with him as well. The two of them are pursuing this notion of basal cognition: that there is intelligence everywhere. It's just a question of how we conceive of it and how we test for it.

And indeed, Mike's argument is that you have to design the right experiments. It's an empirical question: is this virus intelligent or not? Well, you have to design the right experiments to disclose or evince intelligent behavior, and Mike would claim that, by many metrics and yardsticks that we use to measure adaptive intelligent behavior, slime molds, viruses, and xenobots are incredibly intelligent.

I think he's fighting against the chauvinism that you mentioned: that intelligent things are just properties of creatures like you and me and our brains, and that this kind of cognitive capacity and competence can be found everywhere. So I'm very sympathetic to that. But it does tread on the toes of the vague argument.

I certainly don't think that viruses have the same expressive kind of agency that we were talking about. And indeed, the counterargument to the vague notion of intelligence and/or consciousness is hitting you in the face when you read that Strange Things paper. I'm talking about categorically different natural kinds that do and do not have, for example, a causal power or influence of active states on internal states. You're either one of these, or you're one of these.

Keith

I'm not so sympathetic to the view that a virus is intelligent, right? Because I think this falls into the category of things like when you're in a biology class and you learn how to define life, and then somebody says, “What about fire?” It seems to meet all these criteria. It grows, it can expand, and it uses resources.

For example, when a virus “mutates,” the virus doesn't mutate itself. It is mutated by a gamma ray hitting it, or by a transcription error, or whatever. Oil, vinegar, and baking soda undergoing a reaction is not intelligent, right? I think there's some level of complexity—whether it's recursion or these other types of causal structures—that has to be there before I'm even interested in talking about intelligence or entertaining intelligence. Other things are just dynamics that unfold in certain ways: chemical reactions and that sort of thing.

6. Intelligence Has a Goldilocks Scale

Karl Friston

And I think that, at that point, is a nice opportunity just to introduce the notion of scale-freeness, or scale invariance.

Keith

Okay.

Karl Friston

I can see you could also take this toward collective intelligence, federated learning, and federated inference. It's not the single cell; it's the single cell with its neighbors, and its neighbors' neighbors' neighbors' neighbors. What one's looking for is some conservation of intelligent dynamics that is preserved over different scales.

The single virus is probably not interesting. It's probably the colony that is interesting. I thought that was a nice point at which to introduce the notion of that kind of scale invariance. And, of course, as a mathematician, you would be looking at the renormalization group to see how that unpacked mathematically.

Interestingly, it also speaks to this notion of a mutation. Is evolution an intelligent process? It's certainly adaptive. It certainly has the level of complexity that you would require in order to pass those vague thresholds. But is evolution in and of itself an intelligent process? It's certainly a free-energy-minimizing process.

It's just Bayesian model selection. Natural selection just is selecting those things that have the highest model evidence, or marginal likelihood, of being that phenotype in this kind of environment.

Keith

Well, it's like, is the planet intelligent? It certainly contains 8 billion or so of us, so does that count? I mean—

Karl Friston

Yeah. Why not?

Keith

Well, I would say—and we talked about this way back when we talked about whether a flotilla is an agent, or whether it's the pilot who's controlling the convoy or whatever—I would say that if there's a more minimal boundary that contains the intelligent entity, then the larger one is not. The planet, for example: since I can draw a smaller boundary, which is an individual person, the individual person is intelligent. Anything that contains that person is just more stuff.

Tim Scarfe

But Keith, what is your— The issue with viruses: is it because the individual virus is inert? Whereas a cell, for example, you can partition down to the individual cells, and the cell still has some degree of agentic property, as Professor Friston was describing. Is it that you see them as inert and being carried by something else?

Keith

Well, I think it's that, too: too much of their causal structure and machinery is basically outsourced to other things.

Tim Scarfe

Yeah.

Keith

They're kind of just these little particles that go around, and if they happen to stick on a cell, they're like a mousetrap that activates and just injects some DNA. But the cell does all the rest of the work, right? It does the transcription. It has all that machinery.

By themselves, a virus doesn't change over the course of its lifetime. Other than this simple mousetrap kind of activation, it does nothing. It doesn't have any machinery to learn and adapt by itself over the course of the lifetime of a virus.

I might be more amenable to somebody saying that an ecosystem of viruses or something has some degree of intelligence. But it still doesn't have the type of processing and causal structure, and the minimal complexity, that, at least for me, is the useful concept of intelligence.

Karl Friston

So I think we're coming back now to the number of inner screens, the counterfactual depth, and the definition of agency in the sense of Strange Things that we're talking about for this particular scale. If we just read the scale as, if you like, the size of the Markov blanket, then at this scale, for this kind of thing, if there is a sufficient degree of complexity—read specifically, though, in this instance, as the counterfactual depth and breadth of your world model about the consequences of your action, that particular part of your implicit generative model—that would happily accommodate the fact that when you get something of the size or scale of a virus, there just isn't the machinery or the space to entertain that in any nontrivial way.

Keith

Right.

Karl Friston

One could also argue—and I got a sense that you were arguing yourself toward this—that if you go too big, you also lose that.

Keith

Mm-hmm.

Karl Friston

Coming back to this wonderful question: we have 8 billion intelligent—really intelligent, I sound like Trump there, didn't I?—entities constituting our biosphere. But is the biosphere, from the point of view of, say, the Carr hypothesis, intelligent in and of itself? I would say no. It's simply because, at that scale, all the complexity of the constituent elements disappears.

Keith

Yes.

Karl Friston

It would be a little bit like saying, “Let's just take it to the limit. Let's just take it to the astronomical limit, the motion of heavenly bodies.” The motion of heavenly bodies is completely described by the position of the planets and the Moon, right, and the position of the Earth.

At that scale, you've averaged away the fact that the Earth contains a biosphere, the biosphere contains human beings, human beings contain cells, and the cells may or may not contain viruses, depending upon who you've been exposed to. So there will be no intelligence at that level.

So it’s perfectly possible to have intelligence at a particular scale that disappears when you get too big and when you get too small. Another way of looking at that is going right back to things like Prigogine and dissipative structures. You’re talking about complexity. If we just think about how people have tried to understand complex systems and self-organizing systems that are open, we’re not talking about 20th-century physics and equilibrium physics. We’re talking about the physics of nonequilibrium and things that are open, in open exchange with each other.

And, of course, you get to the notion of dissipative structures. What does that tell you mathematically? Well, what’s not dissipative? That’s a schoolboy question. Can you remember? You’ve probably forgotten. So what is not dissipative is conservative.

Keith

Oh, got you. Sure.

Karl Friston

Yeah, I’m sure I’m teasing. Another gift of Helmholtz, of course, is that any dynamics can be partitioned into a dissipative part and a conservative part. The dissipative part is that which rests upon very fast, random, complicated fluctuations, of the kind you might find in, say, quantum mechanics or thermodynamics, whereas the conservative part does not rely upon that and just goes round in circles, basically. It’s literally called solenoidal flow, or conservative flow.

So that basically means that dissipative structures have to have an admixture of this circular aspect, this conservative, classical aspect—life cycles, reproduction, oscillations—plus the random dissipative part that you’ll find in things like thermodynamics and quantum mechanics. That is definitional of dissipative structures.

As you get too small, everything becomes quantum and random. It all becomes probabilistic. There is no conservative stuff other than a Schrödinger potential, at which point I think you would find it very difficult to find something that was intelligent in the sense we’re talking about, because we have to have this recurrence, this solenoidal aspect, in order to revisit the states, so you know a virus is a virus, in order for there to be a nonequilibrium steady-state distributional solution to that.

But as you get bigger and bigger and bigger, you get to viruses, and they’re still not quite complex enough to be intelligent in the way that we mean. Then you get to our size, and that’s perfect because we have both this dissipative aspect—we deal with a random world, with an itinerant world—and yet we keep revisiting states of being. So we have this sort of conservative, biomimetic kind of self-organization.

But then you get bigger and bigger and bigger, and you get to the level of the biosphere or the size of the Moon or the Sun. Of course, by averaging, all the random fluctuations go away. So you’re just left with the solenoidal part. You’re just left with the solenoidal motion, the Newtonian motion of classical mechanics.

Keith

I want to jump in here because one thing that really fascinates me about this is that it’s possible for us to construct large-scale things that are intelligent, like a corporation. A corporation is seen as a superintelligence. But we’re only able to do that by putting in structure that maintains this balance of dissipative and conservative flow. We have to put in organizational structure so that it continues to function at these larger scales, right? Isn’t that pretty interesting?

It’s again this yin-and-yang thing that we run into all the time, where it’s the balance between 2 forces: either between complete order and complete noise, or between dissipation and conservation. You have to be almost on the edge of chaos, and it has to have a certain causal structure in order for it to be intelligent.

Karl Friston

No, well, I agree entirely. That sort of Goldilocks regime, where you are on the edge of chaos, I think is quite specific to a particular scale. Certainly, you could invoke a sort of strong anthropic principle here and say that the kind of intelligence that we will recognize has to be at our scale.

But I think there’s something more fundamental than that. I think that the very existence, if you subscribe as an externalist to quantum physics and Newtonian physics or Lagrangian classical mechanics, means that there is a Goldilocks regime. It means that we can only exist at this scale with, as you say, this sort of yin and yang, this admixture of dissipative dynamics and conservative dynamics.

Just to reinforce this, conservative dynamics is absolutely essential because it is that which causes this Poincaré recurrence. It defines these strange attractors, which is the other way of licensing the notion of strange things. You always come back to somewhere near where you started, like Red Queen dynamics in theoretical biology.

So you have to have the conservative, circular motion just to have a routine, have biorhythms, replicate, and reproduce at many different levels and at many different scales. But it’s always remarkable in the face of a dissipative, itinerant world. Everything is changing all the time, and yet we somehow resist that change by being at the edge of chaos.

You’re more anxious than I am to say something. Carry on.

Keith

Please go on, Professor Friston.

Karl Friston

No, no, no.

Keith

I’m just so fascinated by your observation that with Gaia theory, for example, when we zoom out, the apparent phenomenon seems less intelligent. As you said, maybe intelligence is just what we recognize. When we spoke with Wolfram that time, he was talking about how we’re computationally bounded as observers.

We do this thing called abstraction, where we ignore details, or even idealization, where we deliberately distort the truth. David Krakauer said that intelligence is doing more with less, and emergence is “more is different.” The Earth isn’t doing less; we just don’t see what it’s doing. But if emergence is “more is different,” then that licenses a fundamental reorganization of the underlying substrate, which means it isn’t actually doing the thing at the lower level anymore.

It’s a fundamental coarse-graining, and it’s actually changing at the higher level. So which is it?

Karl Friston

Well, I didn’t realize there was a choice there because I agree with everything you’ve said. But I like the notion of coarse-graining because, of course, that is exactly what you get from the renormalization-group treatment of these things. You can simulate free-energy-minimizing processes that are running at different scales, and then you have to ask the deep questions of how you couple between the scales. There you get to supervenience and emergence, mathematically so defined.

But it all boils down to exactly where you started and where you ended up, which is a coarse-graining in the right kind of way. To actually write down a renormalization group, you have to have an RG operator. What is an RG operator? It has 2 parts to it. It has a sort of dimension reduction—the R part, if you like—and the grouping operator.

They basically do the right kind of coarse-graining. They reduce dimensionality, and they group together in the right way those states of being at the higher scale, and so on, ad infinitum. So I think the notion of coarse-graining is absolutely essential here. On that level, you could argue both ways.

I’m getting a sense that your question is something like: Is there a true emergentism here or not? Again, I don’t have philosophical training, so I’m not sure I can really answer that. But certainly, the intelligent dynamics, read as a self-evidencing perspective on a free-energy-minimizing process, are recapitulated at a completely different scale in a different way at each scale through the recursive application of RG operators.

So yes, there is something brand new going on. And yet, in the spirit of Haken synergetics, these RG operators are just taking functions of stuff that’s happening at the finer scale. So you’re not inventing anything new. It’s just predicated on, emerging from, or supervening on finer-scale stuff. But what emerges at the higher scale has all the attributes of self-evidencing, and you could argue for consciousness or intelligence at some level.

But coming back to the point about building bigger organizations: If the argument is that, as you get bigger and bigger and bigger, you have to be effectively conservative in order to exist, then a company, for example—what is its metric of goodness? It’s how long it survives. Can you get from Series A to Series whatever in some recognizable form?

So as you get bigger and bigger and bigger, I think there’s less opportunity for this complex, hierarchical recursive structure within the scale. I don’t see the Moon thinking or planning.

I don't see the weather planning. I don't see evolution planning. It doesn't think about its future. It's too big. I would imagine that you could apply the same arguments to globalization and institutions that get too big for their own good, because they can't plan anymore. So we're coming back to the pilot, who's really the intelligent person, not the flotilla.

Keith

Mm-hmm.

Karl Friston
Keith

Right. Yeah. In a way, it's slightly depressing to me because it means we can't build an intelligent intergalactic civilization. It's almost like, beyond a certain scale, we just have to leave it up to distributed, emergent processes that may or may not end up doing something intelligent as a whole. At some point, you've got the speed of light, and that's basically going to be a barrier to organization across light-years, right? So maybe there is an ultimate Goldilocks limit to intelligence. It can't get too big.

Karl Friston

That's certainly what the arguments from the physics part suggest: there is a Goldilocks of scale, or zone in scale space. I'm not sure this is correct, and I don't want to disappoint you, so—

Keith

Don't worry. It's well beyond my lifetime when that will happen.

Karl Friston

There's a certain beauty, though, in federating and distributing and thinking in terms of ecosystems of things at a particular scale. Evolution is a beautiful thing. It doesn't have to be intelligent to be beautiful.

Keith

Yeah, yeah.

Karl Friston
Keith

Correct. Correct. Well, good point. It could be beautiful but not intelligent. We could have a beautiful, distributed civilization. That's fine. I'm happy with that.

7. Drawing the Cognitive Boundary

Tim Scarfe

On the internalism and externalism debate, I'm a huge fan of Andy Clark, as you know, and he famously argued in his 1999 paper with David Chalmers, that the phone extends your mind. Even when we were talking about the plant example before, when we were talking about plants doing modeling, there's an enactive interpretation there. Keith and I were arguing about this, and he doesn't agree with me.

Is the model the physical morphology of the plant? It fits the environment like a key in the lock. Or is the model the software? Is it the DNA of the plant? It's fascinating, isn't it, that you can think of cognition as not being entirely inside our heads, but as this unfurling, causal, bidirectional process of so many things around us.

So how do we actually draw boundaries around things where we say, “Okay, well, here's the principle: modeling and cognition are in here. This is where bona fide beliefs are happening. This is where bona fide thinking is happening”? Is it even possible to make that distinction?

Karl Friston

I think it is. I'm answering it in an almost trivial way: if you want to talk about something, you have to be able to define its Markov blanket.

I think your question, though, speaks to something that we've just been covering, which is the separation of scales. Everything, literally, from the point of view of the free energy principle, that is equipped with a Markov blanket is contextualized by a scale above, and the same rules apply to the scale above as to the thing itself. So you have to have a context in which everything is operating.

The virus may not be intelligent, but it certainly has to be living in a world that is conducive to its existence. That world, and the states of that particular world at the scale above—for example, the host cell and the host organism—have to comply and be intelligent in some sense in order for the virus to be there, even if the virus itself is not intelligent.

But we should come back to extended cognition and Andy Clark. To the plant, you made an interesting point: is it in the DNA, or is it in the morphology, the phototaxics, and everything else that plants possess and that we infer goes on inside?

First of all, I think you're absolutely right to say that it is in the morphology. In a sense, that's what I was getting at when talking about the importance of mortal computation for machine consciousness. To realize Bayesian mechanics or self-evidencing, you physically have to parameterize your conditional probability distributions—your Bayesian beliefs about the environment with which you are coupled. The substrate is the parameterization.

By definition, it's substrate-dependent computation, which means the structural form—the morphology—is the structure of the generative model, in the spirit of structural learning. That is very much in the spirit of the good regulator theorem: to be in my world, I have to physically encode and embody a model of the structure of my world.

If you think your world is scale-free, that means my brain must have some scale-free hierarchical aspect, and indeed it does. You could actually write that down and model it with a renormalization group. That's just a reflection of the fact that I am immersed in a world that has some scale invariance in it. So that is certainly true for the plant.

Just one little aside here: that paper was written before David Attenborough's series on the life of plants, where he was able to speed things up by a factor of 10 or 100.

Tim Scarfe

Yes.

Karl Friston

Of course, if you look at plants doing things—sending their little roots and shoots off in a particular direction, or eating insects or whatever—if you speed it up, these things are very animalistic. They look like they could be intelligent, irrespective of whether they're conscious or not. They certainly start to look much more like you and me when you speed things up.

Their morphology matters more than one could possibly imagine, because this is how you become a good regulator. You model your environment. You install that cause-and-effect structure into your computer architecture, which is, again, the argument against von Neumann architectures. That's why one might look to all the processing and memory in neuromorphic photonics, possibly quantum computation. Quantum computation has gone off the boil recently, but all that processing and memory stuff—memristors, for example—I think that's where the answer will be.

Is it in the DNA? Does the DNA cause the structure? So—

Tim Scarfe

Keep your comment about gene expression to make that point.

Karl Friston

You could, if you wanted to simulate these kinds of things, read the DNA as the code that specifies the priors that specify the structure. If you think of DNA as prescribing the structure of your generative model, your world model, or your factor graph, that will be fit for purpose and is learnable.

You start off with, say, DNA. DNA does not tell you what particular kind of plant you're going to be. It's not going to tell you how to engage in your phototaxis and point towards the sun, or compete with other plants that are trying to deny you sunlight. That is something that you have to update and learn during your particular lifetime. But you're equipped with the basic structure—the prior on the structure of your generative model.

And of course, that is most gracefully accommodated, I think, again with respect to the renormalization group. We’re just talking about 2 scales. So you’ve got a slow scale, where you’ve got the viral mutation you mentioned before. The viral DNA—should it have RNA or DNA?—is changing very, very slowly, on a timescale that is greater than or equivalent to the lifespan of any given virus.

That’s, if you like, specifying the initial conditions—the structure for the specification of a particular instance of a virus that lives—

Keith

Well, maybe just jump in for a minute, because I think it’s not even specifying the structure. This is what you pointed out before about the difference between, say, active inference and modeling: DNA is specifying a policy.

Karl Friston

Yes.

Keith

It’s effectively just specifying a policy that every single cell in the organism follows. This policy is not instructions for a structure, but instructions on what to do in response to a certain set of sensory states, right? So it’s almost the policy that the DNA specifies, and this policy applies to every cell. Then, somehow, miraculously, it works out to create a structure that is fit for purpose in its particular environment.

Karl Friston

Yes. And of course, for that specification to work, it has to change at the same rate at which the environment changes, which is very, very slowly. I think that’s absolutely right.

Practically, that hits you in the face when it comes to thinking about how you commit to one policy or another policy. In my world, that would be the expected free energy. The expected free energy comes in 2 parts. It has a sort of expected information gain and the expected cost, or constraints, or utility. That is the point at which you need your DNA. You need to know what it is to be the kind of thing that I am. What do I not do, and what do I do?

Keith

The fascinating thing in organisms is that this is different even on a cell-by-cell level, right? Because morphologically, the same code evolves into all the different organs and parts, roots, bark, and whatever else, which is pretty fascinating.

Karl Friston

Mm-hmm.

Keith

It’s—

Karl Friston

And, yeah, again, it’s exactly the kind of fascinating issue that preoccupies Mike Levin.

Keith

Yeah.

Tim Scarfe

I spoke with our mutual friend Maxwell Ramstead the other day, and he actually gave a wonderful description of the free energy principle. What you want in a good pitch is for it to be replicable and easily understandable, and I’ll try and recapitulate it to you to test how replicable it was.

He said, “Second law of thermodynamics: closed systems. What if we have open systems with boundaries?” He was saying that when things can’t merge together, they instead share information with each other. So there’s this informational synchrony, and that was his way of describing the free energy principle.

First of all, is that a good description? Number 2, where do the boundaries come from? We want to have practical implementations of the free energy principle, and of course we can implement it with computer games or contrived environments where the boundaries are clear. But what if I want to build a robot, and I have a camera, and I see all of these pixels on the screen? I need to divide that scene up into objects so I can start applying the free energy principle. How do we do that?

Karl Friston

Right. 5 minutes, 5 questions. Good.

Keith

It’s got to be practical. That’s why we’re only giving you 5 minutes.

Karl Friston

Right. First of all, the notion of being separate but part of a universe through synchronization, I think, is absolutely correct, and it could be unpacked in 2 ways.

One, you could say that to find a free energy-minimizing solution to your dynamics, namely the path of least action, is just to evince a generalized synchrony between the inside and the outside. Generalized synchrony is also known as synchronization of chaos. Basically, 2 systems that are loosely coupled, in the sense of dynamical systems, will ultimately converge on a synchronization manifold, and they will show chaotic dynamics on that manifold.

That manifold is a synchronization manifold, and being on that manifold is the minimization of free energy. So you can talk about the free energy principle without mentioning self-evidencing, without mentioning Bayes, without mentioning predictive processing, or even extended cognition. You can just talk about it as a variational principle specifying a variational bound on the Lagrangian that specifies generalized synchrony, or synchronization of chaos.

What you’re talking about is exactly this sort of “separate but the same.” I think it’s quite nicely, although I don’t fully understand it, recapitulated in Chris Fields’s quantum treatment. What he would talk about there is not the synchronization between the inside and the outside, or me and you and everything else like me and you, but entanglement.

The principle of unitarity is this—well, you can read the free energy principle as just the principle of unitarity, which means that if we stay here for long enough, for more than 5 minutes, we will ultimately become completely entangled. Classically, that means we will be engaged in a generalized synchrony. There will literally be a synchronization of our itinerant dynamics. We will not become one. Well, actually, Chris Field thinks you would become one.

We will be indistinguishable because we’re just playing out on exactly the same synchronization manifold. So I think that’s absolutely right.

Your last question—and you had about 3 in between—was how you would practically use the free energy principle, and the notion of thingness inherited from Markov blankets, if you’re building robots. I think you’re talking here about physics-discovery or Markov-blanket-discovery algorithms, which, of course, Maxwell and Jeff Becker are furiously working away on.

Tim Scarfe

Yeah.

Assuming you didn’t know anything about that stuff, how would you approach that problem? From first principles, if you had a robot and a camera, how would you partition the scene into boundaries?

Karl Friston

You would be appealing to the notion of Markov boundaries. You can turn it on its head and look at image segmentation as what has emerged in computer vision as a way of identifying things. What you’re doing is committing to an implicit world model, or generative model, for your robot. You’re assuming that it is going to be operating as a good regulator, as a good model of its environment, in an environment that is composed of things.

Tim Scarfe

Yes. But I think my intuition is that we’ve been talking a lot about understanding and creativity recently, and I intuitively feel that understanding is about knowing the history of something. It’s about knowing how you got there. It’s not about knowing the state.

So it’s a little bit different from segmenting an image. To understand the history and the dynamics of a system seems to be much, much more powerful.

Karl Friston

Absolutely. Which is why Markov-blanket discovery is not image segmentation. To do Markov-blanket discovery, you have to look at the dynamics and the history.

Tim Scarfe

Yes.

Karl Friston

There’s quite a fundamental point to be made here. We’ve been talking from the beginning right through to the end about conditional probability distributions. Probability distributions over what? There’s only 1 state of the universe at any one time. There are not many worlds. There’s not an ensemble of universes.

For the purposes of this argument, there’s only 1 state of the universe at any one time. So you can’t have a probability distribution unless you invoke some sort of ensemble or canonical-ensemble assumption, as they do in thermodynamics, where there’s some exchangeability. You can swap around universes and have a distribution.

Keith

Well, you can have an epistemological probability distribution, right?

Karl Friston

I would argue, though, that what you’re implicitly doing is actually doing that over time. There’s a history, a past, and a future. That epistemological notion really has to refer to what is the background over which these multiple realizations occur, from which you can select to build a probability distribution.

Of course, from the point of view of the free energy principle and also building robots, this has to be time. So, by definition, it has to be in the dynamics and the history—not just the short-term dynamics, but recurrence: the characteristic states that I keep returning to as a robot, as a good robot.

So, yeah, absolutely. Segmentation algorithms are a good start, but they’re going to be completely useless when it comes to autonomous vehicles, for example, unless you’ve built in the fact that there is conservation of this particular Markov blanket over time, which, of course, they don’t.

It’s probably better to start with the generative model: this 2-dimensional RGB feed has been generated by Markov blankets that, by definition, persist over time, because time is the support of the conditional distributions that define the existence of the Markov blanket.

Then it’s a question of what kind of priors you put on these things. Are they things? Are they stuff? What’s the Markov blanket of water? Or perhaps fog, if we’re doing autonomous vehicles.

To what extent would I now apply my priors to these things? You can go right through to object-centric priors, of the kind—physics engine-based stuff—that, say, Josh Tenenbaum pursues, assuming that you’ve got some Newtonian behavior. Or you could be much more relaxed. I would relax, but not beyond the renormalization group.

Tim Scarfe

Professor Friston, it’s been an absolute honor. Thank you so much for joining us today. It’s been amazing.

Karl Friston

I really enjoyed it. It’s lovely to speak to you both again.

Keith

Yeah. It’s really been an honor for me because this is the first time I’ve met you in person. I’ve been thinking about you for many years, so it’s been a pleasure to meet you, and thank you for coming to the studio today.

Karl Friston

Well, thank you for coming to England.

Keith

Absolutely.

Karl Friston - Why Intelligence Can't Get Too Large (Goldilocks principle) | BidClub