Tim Scarfe
What’s really interesting about this book, Max, is that I’ve read loads and loads of books in this space, and there are people like Hinton, Hawkins, Damasio, Friston, and even Sutton. What’s interesting is that it’s a bit like the blind men and the elephant: they’ve all got a completely different story to tell. I think the magic you’ve pulled off with this book is somehow weaving it together into a coherent story. What do you think about that?
Max Bennett
Well, first, I’m very appreciative of the kind words. I think I came from a very unique perspective because I was a complete outsider, and I didn’t come to it with the objective of writing an academic book at all. I came to it with the objective of just learning on my own.
I started building this corpus of notes because I was so independently curious. I kind of stumbled on this idea, really for myself, of how do I make sense of all of these disparate opinions and this complete lack of information about how the brain actually works? I had my own set of, I think, biases coming from sort of the technology entrepreneurial world, where we tend to think about things as ordered modifications.
When you think about product strategies or how to roll things out, we like to think about things as: what’s step 1, then what’s step 2, and what’s step 3? I think I did have a cognitive bias to, when presented with an incredible amount of complexity, try to make sense of it in a similar type of way.
As an outsider, I felt very free to explore and cross the boundaries between fields. I look at the book as a merging of 3 fields. One is comparative psychology: trying to understand what the different intellectual capacities of different species are.
The second is evolutionary neuroscience: what do we know about the past brains of humans, and the ordered set of modifications through which brains came to be? The third is AI, which is how do we ground the highfalutin conceptual discussions about how the brain works in what works in practice?
I think that’s a really important grounding principle, because it helps hold us accountable to the principles that we think work. If we can’t implement them in AI systems, it should make us question whether we actually have the ideas right. Being an outsider comes with disadvantages, but there are some advantages too: you’re free to borrow from a variety of different fields and think freshly about things.
Tim Scarfe
Yes. If you can point to any particular ideas that you found really difficult to reconcile, what would those be?
Max Bennett
One thing that’s really challenging is that if we were to lay out the data richness of comparative psychology studies across species, and put that on a whiteboard, we would realize that we have so little data on what intellectual capacities different animals actually have.
For example, the lamprey fish is the canonical animal used as a model organism for the first vertebrates, because of all vertebrates alive today, it’s one of our most distant vertebrate cousins. To my knowledge, there are absolutely no studies examining map-based navigation in the lamprey fish. We have no idea if it’s capable of recognizing things in 3D space.
When we look at other vertebrates, like teleost fish, they seem eminently capable of doing that. We look at lizards, and they’re eminently capable of doing that as well. We infer that it seems likely that the first vertebrates were able to do this. We know the brain structures from which it emerges in reptiles and teleost fish are present in the lamprey.
We back into an inference that the lamprey fish can probably do that, but this is all, in some sense, guessing and trying to put the pieces together from very little information. I think that’s one challenging aspect to reconcile.
The other one that’s really hard is that, in neuroscience, there are a lot of really interesting ideas about how the brain might work that have not really been tested in the wild from an AI perspective. Then there are a lot of AI systems that work really well but have diverged substantially from what the evidence suggests about how brains work.
How do you bridge the gap between these 2 things? I think that’s a really fascinating space to operate in. What can we learn about the brain, if anything, from the success of transformers, as an example? What can we learn, if anything, from the success of generative models in general? What can we learn from the successes and failures of modern reinforcement learning?
In some ways, reinforcement learning has been a success; in other ways, it’s really fallen short of what a lot of people hoped it would be. I think the gap between neuroscience and AI is still a challenging one to bridge in a lot of ways.
For example, Karl Friston has all these incredible ideas in active inference. In 100 years, will we look back on this and say, “Karl Friston was on to something”? If you look at the AI systems today, there’s very little usage of active inference principles working in practice.
That could mean that the ideas don’t have legs, or it could mean that there’s a breakthrough around the corner where we’re actually missing some of the key principles he’s devising. These are questions we don’t have the answers to.
Tim Scarfe
I think there might possibly be some breakthroughs around the corner. I don’t know if you know, but I’m Karl Friston’s personal publicist. I do all of his stuff. I probably interviewed him more than anyone else, but I love my friend. He’s an amazing guy.
Max Bennett
Yeah, he’s an amazing man.
Tim Scarfe
Honestly, he is the man.
Max Bennett
So kind.
Tim Scarfe
I know. For me as well, he has so much time to explain things. You could cynically argue that the effective active-inference agent is just a reinforcement-learning agent of a particular variety. I think it’s equivalent to an inverse-reinforcement-learning maximum-entropy agent or something.
But there’s so much more than that. There’s so much richness and explanatory power in modeling this thing as a generative model that can generate policies and plans of action, and so on. We want to have agents that we understand, with steerability, and that are able to do the simulations you talk about so eloquently in your paper.
To come back to what you were saying, you mentioned that there might be a parallel between transformers and AI models. In your book, on this page, you analogize model-based reinforcement learning and the neocortex.
Of course, I interviewed Hawkins back in the day, and the main criticism of his book is the triune-brain-type argument. He’s giving the explanation that the brain developed a bit like geological strata, with one layer and then another layer, rather than co-evolving together.
It’s so hard not to think like that because you give so many beautiful examples in your book, not only morphologically but in terms of capability. With stroke victims, for example, it’s not like the brain recovers those dead cells; it learns to repurpose those functions in other parts of the brain.
It seems like Mountcastle was correct that the neocortex is this magic, general-purpose learning system. What do you think?
Max Bennett
There are 2 different ways to look at the neocortex enabling things like mental simulation and model-based reinforcement learning. One is that that function and algorithm are being implemented in the neocortex. But another, which is a slightly less strong claim and the one I would make, is that the addition of the neocortex enables the overall system to engage in this process.
That is not saying that the entire process is implemented in the neocortex. I think it seems very clear that the thalamus and basal ganglia are essential aspects of enabling the pausing, the mental simulation, the modeling of one’s own intentions, the evaluation of the results, and so on.
But it is possible to say, which is what I’m arguing in the book, that in the absence of the neocortex, that process does not happen. I think where my ideas would synergize with what Hawkins is saying is that the neocortex builds a very rich model of the world, and a model of sufficient richness that you can explore it in the absence of sensory input.
That’s a really essential aspect of model-based reinforcement learning. If I have a model of the world that has sufficient richness for me to mentally simulate actions that I’m not actually taking, and it at least somewhat accurately predicts the real consequences of those actions, that model is really useful because I can now imagine outcomes before having them. I can flexibly adjust to new situations.
Of course, there are so many deep, interesting questions that are yet to be answered about that. For example, just because you can render a simulation of the world doesn’t answer the question: what do you simulate? This is one of the hardest problems of model-based reinforcement learning. You can have a model of the world, but how do you prune the search space of which aspects of that model you explore before evaluating outcomes?
That’s another really hard challenge. I think there’s a lot of good evidence that this is actually a partnership between the neocortex and the basal ganglia, which is a much older structure.
So, yeah, I’m not really of the view that the triune brain has been, amongst evolutionary neuroscientists, largely discredited. I think that’s in part somewhat unfairly so, because if you actually read MacLean’s writings, he is very open about the fact that this is an approximation and not exactly accurate. He couches his claims much more carefully than popular culture, which just converted them into a dogma.
I think the popular interpretation of the triune brain is not accurate. It is clearly not the case that the brain evolves in 3 key layers. It’s not the case that a reptile brain doesn’t have anything limbic-like. If you look at a reptile brain, it absolutely has a cortex that does a lot of what our limbic structures do, et cetera. So, yeah, those would be my thoughts on that.
Tim Scarfe
Yeah, it’s fascinating because we, as humans, need to have models to explain and understand the thing itself, just like active inference, for example. I’ll get to planning, agency, and goals in a little while, but a lot of these things are instrumental fictions. I’m not saying that our brains don’t plan, but the abstract mathematical way that we understand planning is probably not how the brain works. It’s much more complicated than that.
Why don’t we just rewind to the beginning? We’re going to be talking about this chapter on simulation, if you like. You lead by saying that what the neocortex does is learning by imagining. Hawkins spoke about this as well. He said we’ve got the Matrix inside our brains, right? We’re always doing all of these simulations of future things, and we’re using that to help us understand the world.
You give this really interesting example of some of the features of the brain that lead you to believe that we are basically living in a simulation. It’s almost like, rather than perceiving things, we’re testing whether our simulation is correct. But that means that we can only simulate 1 thing at a time, so we can’t see 2 things. We can only see 1 thing. Can you talk through that?
Max Bennett
Sure. One of the first introspections and explorations into how perception works in the human mind happened in the late 19th century, with all of these explorations of visual illusions that you see in pretty much every neuroscience textbook or book that you open. Listeners will be familiar with them—you’ve probably seen examples of triangles where you actually perceive a triangle in a picture when there is, in fact, no triangle. Yeah, you can find that picture.
Tim Scarfe
Yeah, sorry. I hope I’m not distracting you.
Max Bennett
No, no, no. So that’s a standard finding that was observed in the 19th century: this idea that clearly the brain observes the presence of things even though they’re not actually there. We perceive a triangle there, a sphere, a bar, and the word “editor,” when, in fact, if you actually examine it, the letter E is not there. There’s evidence that suggests the E is there by virtue of showing the shadows, but we did not actually write the letter E there. The brain regularly observes that.
That finding led this scientist, Hermann von Helmholtz, to come up with the concept that what you actually consciously perceive is not your sensory stimuli. You are not receiving sensory input and experiencing the sensory input. What’s happening is that your brain is making an inference as to what is true in the world, what’s actually there, and then sensory input is giving evidence to your brain as to what’s there.
You start from this prior, and then that prior maintains itself until you get sufficient evidence to the contrary; then you change your mind. It’s not hard to imagine why this would be extremely useful in any sort of environment that an animal might evolve in. Suppose you have a mouse running across a tree branch at night. First, I see the tree branch in the moonlight, so I build a mental model of the tree branch. As I move forward, I lose the moonlight and no longer see the tree branch.
As long as I’m stepping forward, the evidence is consistent with my prior of the tree branch. It makes way more sense for me to maintain the mental model of that tree branch, as opposed to all of a sudden having the tree branch disappear because I no longer see the sensory stimuli of it. Because sensory stimuli are very noisy, it makes a lot of sense that we integrate them over time, build a prior, and then, until something gives us evidence to the contrary, maintain our prior about the world.
So that was the first idea: there’s some form of inference. There’s some difference between sensory input and some model of the world that we infer and then thus perceive. What’s interesting, and not as discussed but also present in the discussions among scientists in the late 19th century, is the idea that you can’t actually render a simulation of 2 things at once.
There are lots of really interesting visual illusions around this where you can see something. The famous one is that it’s either a duck or a rabbit. Yep, exactly. And it’s interesting. You can see that a staircase is either moving up to the left, or you’re under the staircase looking upwards, and it’s actually a ceiling that’s jagged.
Why can’t the brain perceive both of those things at the same time? It would make sense if you have a model that there are such things as ducks, such things as rabbits, and such things as 3D shapes that operate under certain assumptions. If that’s true, then you cannot see a duck and a rabbit at the same time, because there’s no such thing. It cannot be the case that the staircase is being viewed from above and below at the same time.
So what your brain is not doing is just perceiving the sensory stimuli. It’s trying to infer what is a real 3D thing in the world that I’m aware of, that this sensory stimuli is suggesting is true, and that is the thing that I’m going to render in your mind.
I think one way this parallels nicely with some of Hawkins’s ideas is that, if you hold the Thousand Brains Theory to be true, and the neocortex has all of these redundant, overlapping models of objects, then it would make a lot of sense that we want to synergize these models to render 1 thing at a time. You don’t want to have 15 different things rendered, because then it’s really hard to evaluate them and vote between these different columns.
It makes sense that the brain says, “Let me integrate all the input across sensory stimuli and render 1 sort of symphony of models in my mind, so I can see 1 thing at a time.” So that’s this idea of perception by inference. At the time, no one really connected that to the idea of planning. This was just the idea that what we perceive is different from the sensory stimuli we get.
Later on, as the world of AI started thinking about things from the perspective of perception by inference, what we end up realizing is that this idea of perception by inference, if you’re going to train a model to do that, comes with this notion of generation. The way it self-supervises is that it takes the prior, tries to make predictions, and compares the predictions to what occurs in the world. As long as those predictions are below a threshold, I maintain my prior.
A famous version of this is the Helmholtz machine, which Hinton devised. I think that was in the 1980s. It could be later; I forget. This is the basic idea that you can build a model. A lot of people use the term “latent representation.” Some people don’t like that term for a variety of philosophical reasons.
It builds a representation of things by virtue of building a model. In other words, perception by inference. The way you build that model is by constantly comparing the predictions generated from that model to what actually occurs. This also has synergies with a lot of Hawkins’s ideas, where we think about intelligence as prediction.
If you build a model of perception by inference by virtue of generation, then it’s relatively easy to say, “Okay, what happens if I just turn off sensory stimuli and start exploring the latent representation?” Now we’re exploring a simulated world. I’m able to cut off sensory stimuli, close my eyes, and imagine a chair, rotate the chair, and change the color of the chair.
Because this model is relatively good and has relatively rich features about how the world actually works, I can model things without ever having experienced them, without ever having done them, and reasonably predict what would actually happen if I were to do those things.
What I think is interesting, and perhaps somewhat of a novel proposal in the book, is that a lot of people think about the neocortex as having adaptive value because of how good it is at recognizing things in the world. If you read a standard textbook, a lot of what people talk about regarding the neocortex is how good it is at perceiving things—object recognition.
Some of the best-studied parts of the brain are this visual neocortex, so we understand reasonably well how we’re building models of visual objects, et cetera. But from an evolutionary perspective, this is a little bit hard to find convincing, because if you actually examine the object recognition of vertebrates, it’s incredibly accurate.
Tim Scarfe
I mean, a fish can recognize human faces. A fish cannot recognize an object when rotated in 3D space. So it’s hard to find a dividing line between object recognition in animals with a neocortex and object recognition in animals without a neocortex, with brain structures that seem more similar to early vertebrates.
Max Bennett
Did you have a question?
Tim Scarfe
Yeah. Well, I just wanted to touch on a couple of things there. Hawkins said that, first of all, we overcome the binding problem by having this profusion of individual sensorimotor models rather than having this feed-forward enrichment of representations.
That was really interesting, but he also said the reason why having, let’s say, 150,000 mini cortical columns that are wired to different sensorimotor signals gives us robustness of recognition is the diversity and sparsity. Then you can think, “Okay, what do the representations look like?”
If our brain builds some kind of model of the world, some kind of topological model, it must be a representation. It’s not necessarily a homunculus, and it’s not like there’s a stapler inside the brain. If I’m modeling a stapler, it’s actually some weird structure of the stapler as seen by every way I can touch, feel, hear, and lick a stapler, or whatever. So it’s difficult for us to imagine what that is.
But the reason I’m going down this road is because that’s a bit weird, isn’t it? We have this very weird representation of things. Then I come to Hinton’s Helmholtzian generative model, and you can get it to generate, let’s say, the number 8 if it’s trained to do numbers. Hinton would argue—and I would disagree with him—that the model understands what an 8 is.
Now, this is weird, isn’t it? We understand what a mouse is, but intuitively we feel that a neural network doesn’t understand what an 8 is. I would argue that we’re getting into semantics here. I think the reason we understand things is because there’s a relational component to understanding.
Semantics is about the ontology: the way we feel about things, where the thing came from, what intention guided the creation of the thing, and what the provenance of the thing was. It’s almost like the interconnectedness of the thing tells you more about the meaning of the thing than the actual thing itself, in a weird way. What do you think?
Max Bennett
Well, I do think this is where the word “understanding” can mean different things to different people. I think there’s absolutely something to the idea that just because you can recognize something—a feed-forward network that can observe a stapler alone—is insufficient for what most people would mean when they use the word “understanding.”
Just to talk about interrelatedness, if I have a feed-forward network that’s just a binary classifier—“Is this thing a stapler or not?”—I can’t ask many things of that feed-forward network that I would expect of an agent or a model that understands what a stapler is.
I couldn’t ask it, for example, what would happen if I burned the stapler. What would happen if I opened it? What would you see inside of it? I can’t ask, “What does a stapler do?” I can’t show it a human holding a stapler and a set of objects in front of them and ask, “What do you think the person is going to do next with this stapler?”
Clearly, our intuitive understanding of the word “understanding” contains some richness that isn’t included in just classifying or recognizing the presence of objects. I think that’s absolutely the case.
My intuitions fall in the direction of what we typically mean when we use the word “understanding”: having something that can be mentally explored. I think that requires what you’re describing, which is the interrelatedness of things.
When I see someone holding a stapler and you ask me what things they would likely do, I start imagining what they might do with that stapler, and then I can evaluate which ones seem plausible to me. In the imagination and evaluation of which things seem plausible, there’s an interconnectedness between the thing and the world around it.
I absolutely agree with you that just recognizing objects clearly lacks something that we mean when we say “understanding.”
Tim Scarfe
Do you take into consideration the memesphere? We’ve got ideas that are quite collective as well. I was going to explore this with you later, but maybe let me frame it like this: I have this intuition that a lot of our intelligence is outside of the brain.
If I were in the wilderness and disconnected from society, I would be a lesser human being. In a weird way, I might have more agency, but I wouldn’t have access to all of these rich cultural tools, knowledge, patterns, and so on.
It’s almost like that’s where a lot of our intelligence comes from, rather than just being able to plan in the brain. Meaning comes from there as well, and there must be some kind of interplay. Culture must shape the development of our brain, and vice versa, but culture seems to be more dynamic. How do you wrestle with that?
Max Bennett
Well, that’s undeniably true. The first example that comes to mind is writing. What would humans be if you removed the technology of writing? We would all realize that we’re not that smart.
Writing is a technology that externalizes a feature of the brain—memory—which brains aren’t that great at. We do a good job of condensing aspects of a memory. For episodic things and procedural memory, we’re relatively good at those, but for semantic memories, we’re terrible.
Externalizing memory with writing is one of the key technologies that enables us to be much smarter than we are, because we now have this external device that enables us to store a largely infinite number of memories and translate them across generations.
That alone proves your point: what humans are capable of is clearly some relationship between brains and external things. Those things can be writing tools, other brains, sharing ideas, and getting challenged. So, yes, I agree with you.
Tim Scarfe
Intelligence is the dynamics, isn’t it? It’s all of the low-level dynamics of things interacting with each other. We can take a snapshot of language and say, “Oh, well, language isn’t the intelligence, but language itself is a form of intelligence.”
I’m not just talking about the words and language models. We’re talking about the actual language in our culture. I think of that as almost a distinct form of intelligence.
Max Bennett
Yeah. It almost gets into philosophical territory: where do you draw the bounding boxes around the things that are imbued with intelligence and the supportive mechanisms that allow those things to have intelligence?
Through one lens, you could think about brains as the physical entities in which intelligence is instantiated, and language as a supportive tool. You could take a very odd view, perhaps, which is that language is the thing that’s evolving and is simply instantiated in these brains that produce and consume the language. It’s language that’s evolving.
In the same way, we don’t think about intelligence on the level of an individual neuron. We don’t imbue a neuron with intelligence, but on the scale of 86 billion neurons, we think something has emergently appeared that we deem intelligent.
There are some great science-fiction books where intelligence gets instantiated in colonies of ants. Each individual ant isn’t intelligent, but somehow the colony itself is capable of doing incredible abstractions.
So, yes, there are very interesting ways to think about how one divides the lines between the physical entities in which these things are instantiated. That said, I have a particular interest in brains because, if we’re looking for the physical manifestation that we can learn from and thus try to understand ourselves—I think all species have an interest in understanding themselves—but also if we want to borrow some ideas from how biological intelligence works into AI, then of all the physical things to examine, the brain seems clearly the one that is probably richest with insight.
But, no, I agree. I think your point is well taken.
Tim Scarfe
Interesting. Okay. I want to close the loop on what you said about the brain being an imagination-filling-in machine. You said that it does the filling-in one at a time. It can’t unsee visual illusions, and the evidence is seen in the wiring of the neocortex itself, you say.
It’s shown to have many properties consistent with a generative model. The evidence is seen in the surprising symmetry and the ironclad inseparability between perception and imagination that are found in generative models in the neocortex. You give examples like illusions, how humans succumb to hallucinations, why we dream and sleep, and even the inner workings of imagination itself.
So it really seems plausible when you think of it in that way.
Max Bennett
Yeah. I think there's a reason why so much of the neuroscience community has rallied around this idea of predictive coding, which is very related to active inference and generative models. What I'm saying there is not really novel. There's just so much evidence that what's going on in the neocortex—the imagination of things, episodic memories—is consistent with this idea of a generative model.
There's been some good evidence that episodic memory—in other words, thinking about the past and imagining the future—are in fact the same underlying process happening in the neocortex. If we look at the connectivity patterns, which I didn't talk about too deeply in the book because it's a little technical, what you would expect from a generative model is that backward connections would be much richer than forward connections because you're modulating downstream.
Of course, the neocortex is not perfectly hierarchical, but things that are generally lower in the hierarchy would have lots of inputs from parts of the neocortex that are higher in the hierarchy. That's absolutely what we see. So, there's a lot of evidence that these are two sides of the same coin, which is that there's some form of generative model being implemented.
I do think that in AI, one way in which this manifests is the very clear success of self-supervision. The principle of self-supervision is this idea: can a system end up having really interesting emergent properties and generalize well when you only train it on predicting the sensory input that it receives? That clearly has become the case.
I think the Transformer is a great example of how, if you just give it a bunch of data and train it through self-supervision—that is, masking, where you hide certain data inputs—it becomes remarkably accurate and good at generalizing across data that it hasn't seen before. That is, in principle, what people are predicting or claiming that the neocortex is doing as a generative model.
Tim Scarfe
Yeah, it feels to me that there's a big difference between, let's say, a Transformer and the neocortex. I think the difference is—maybe agency is not the right word—that you can think of the neurons as having some kind of autonomy. They're sending messages to each other, and eventually it's consistent: the other neuron will get the message, and it will decide for itself what it's going to do.
In a Transformer, just because of the way it's connected, the backpropagation algorithm, and so on, they all ride a Mexican wave together, to use an analogy. So, it feels like a difference in kind to me.
Max Bennett
Clearly, as I argue in the book, I don't think that the brain is just one big Transformer. I would agree with you in the human brain, unless you think there's something that's nondeterministic and sort of magical happening. I think you would still say that there are either base firing rates of neurons, and then there's sensory input that flows up and goes through the brain until eventually it's affecting muscles and you're responding.
So, there is a deterministic flow happening. It might not be as feed-forward as what's happening in a Transformer, which is definitely the case, but they might both be deterministic in a similar fashion.
I think a lot of people have had this exact argument with me, and one counterargument that people have towards this idea—I don't know if I fully agree with it, but it's interesting—is that attention heads really are doing something more magical than we give them credit for. They are dynamically rerouting and effectively resetting the network based on the context that the prompt is getting.
Although technically it's just a series of matrix multiplications, if what's happening in principle is that these attention heads are doing something really clever—looking at the context of a prompt and then effectively dynamically reweighting the network to decide what it cares about and what it doesn't—there are people who think there is something really interesting happening in the Transformer that might be analogous to certain things happening in the brain.
Clearly, these feed-forward networks are not capturing everything that's going on in the brain.
Tim Scarfe
Yeah, it's really interesting what you said, because the way I read that is that things like ChatGPT and language models are entropy smuggling or agency smuggling. What that means is they just do what you tell them to do, and all of the agency—my directedness—comes from me.
I give it a prompt, it does the thing that I want it to do, and then the mapping that you were talking about, I interpret that a bit like a database query. Depending on the prompt you give it, it'll activate a certain part of the representation space and give you a certain result back.
But the brain has this thing where all of the neurons have their own directedness. The weird thing is that, at the cosmic scale, agency seems to emerge. Even Transformer models that were acting autonomously could presumably, at a large enough scale, give rise to something that we think of as directedness, goals, purpose, or whatever.
It's almost as if, in the natural world, because there are so many levels and scales of independent, autonomous things mingling with each other independently, and then downstream mixing their information together and rinsing and repeating over many different scales, that gives rise to all of these amazing things like agency, creativity, and so on.
Max Bennett
Yeah. I think the notion of agency is an interesting one. I'm very amenable to the idea that there is a debate between the reinforcement learning world and the active inference world about how much of intelligence can be conceived as optimizing a reward function.
The hardcore reinforcement learning world would say that everything is just a reward. The active inference world would argue that not all behavior is driven just by optimizing a single reward function. There is uncertainty minimization, trying to satisfy your own model of yourself, fulfilling your own predictions—these sorts of things that seem very well aligned with the behavior we see.
It's unclear which of these is right. It's probably some balance of the two. But people would conceive of agency differently in these 2 worlds.
Some people in the reinforcement learning world would say that agency is just giving something a reward function, and then it learns over time, trying to optimize that reward. In the more active-inference world, which I do think has legs and to which I'm obviously amenable, the idea of agency is a little bit more about building a model of yourself and trying to infer what your goals are based on observing yourself, and then trying to make predictions to fulfill those end goals.
In other words, goals are constructed. One of my favorite Friston papers is “Predictions Not Commands.” I don't know if you've read that paper, but I think it's a brilliant paper about how you could reconceive the motor cortex—not as sending motor commands to your body, but as building a model of yourself and predicting what will happen. The way the spinal cord is wired is that it just fulfills those predictions.
I think that's a really interesting reframe of how you could get agency and really interesting, smart behavior in the absence of just a strict reward function. The way that would learn is by trying to model the behaviors it observes, and then trying to predict and fulfill them.
Agency is a really interesting concept because it manifests itself in these different paradigms in different ways.
Tim Scarfe
Yeah, I find it fascinating, because the way I read it in the active-inference literature, it's a very principled definition of an agent. There's still a bit of a gap, because I think Friston would argue that in the natural world, because of the laws of physics, particles, and whatnot, you get the emergence of things, and things become agents when they have a certain depth of planning, should we say.
But it's really interesting. I guess he would argue that you get all of these phenomena that give rise to agency, like biotic self-organization and so on. Maybe we should slowly go in that direction. You give the example of mice doing planning. Can you sketch that out?
Max Bennett
Yeah. This is another real area of neuroscience research that I absolutely love. I think it was in the 1940s or 1950s—I forget the exact decade—when Tolman observed that mice, when they reached choice points in mazes where he was training them to navigate, would pause. They would sniff back and forth, and then they would choose an action.
And so he hypothesized this idea that they must be engaging in vicarious trial and error. They must be imagining possible outcomes before deciding. Of course, this was hugely controversial because there was no evidence. He had no evidence that they were, in fact, imagining anything.
Most people, in the absence of evidence, like to assume animals are as dumb as possible. Only when there's irrefutable evidence will we imbue them with any intellectual capacities, which I think is an interesting human bias, but that's fine.
David Redish, who is also a close friend and mentor of mine, did some amazing research with one of his PhD students where they were recording hippocampal place cells. So, as a very quick background for viewers, you can go into the hippocampus of really any mammal, but this is best studied in rats. Part of the hippocampus, a region called CA1, has these things called place cells.
If you record the cells as a mouse is moving around a maze, what you find is this incredible thing: there are neurons that activate only in specific locations in that 2D plane. It's not based on how they got there. It's allocentric; it's independent of their egocentric path. They can come back to the same place from any route, and that same place cell will activate.
As an aside from the evolutionary story, we find similar types of cells—not exactly as accurately—in fish, in the homologous region of the hippocampus in their cortex, where they have place-like cells. They're not as accurate, but they are cells that activate in certain locations in a maze.
What he found is that when mice engage in this act of vicarious trial and error—when they pause and look back and forth—the place cells in the hippocampus cease to activate only in the location where they are. They actually start activating down the paths of each route they might take. In other words, you can literally watch rats imagining the future. I think this is one of the most incredible neuroscience findings.
He then took this and did a bunch of other experiments that I think reveal even further the power of imagination in rats. One of my favorites is his counterfactual learning studies, where he puts rats in this thing called Restaurant Row. It's a square-like maze—yeah, exactly—and as the rat is going counterclockwise, at each door, a sound is released or made.
That sound signals to the rat whether it can go right through the door and get food in, I think, 3 seconds, or whether it's going to have to wait 45 seconds before it gets the food. They're given a bunch of time to try to get as much food as they want. Rats have clear preferences, so some rats will really prefer bananas, and they don't really like the bland food.
What happens? This presents a set of irreversible choices to a rat. Let's say they come up and they can either get a treat right now that they don't really like that much—let's say it's the bland treat—or they can go to the next one and hope that they're going to get the banana really fast. If they go to the next one and the banana sound is long, meaning 45 seconds, then they regret their choice because it would have been better if they had just gone in and quickly gotten the food.
How do we know they're regretting the choice? We can literally watch them imagining eating the foregone choice. We can go into a part of their brain called the orbitofrontal cortex, which activates for certain types of tastes, and we can see them imagining the foregone choice. They end up making different choices the next time around. They end up being less likely to forego that choice in the future.
I think this is such an incredible finding of what we mean when we say model-based reinforcement learning is clearly happening in the brains of very simple mammals.
Tim Scarfe
There's a real challenge in knowing which simulations to run because, if you think about it, we've got a search problem, right? There's an intractable number of simulations to run. How do we fix that in AI, and how do humans fix that?
Max Bennett
This is one of the many outstanding questions in AI, but one of the big ones is: how do you effectively prune the search space?
We do not know how mammal brains do this so well. I can give you some high-level ideas or theories, but we just don't know, and this is one of the big possible breakthroughs in figuring out how mammal brains do such a good job of this.
The thing that AlphaGo does, which I think perhaps is a clue and is clever, is that the selection of the search space is actually bootstrapped on the temporal-difference learning model under the hood. This is very clever. Let's say you train something to learn without a model of the world. All it's doing is getting sensory stimuli. It gets a model of a Go board, and then it just predicts the right next action. It just has a policy-value function that bootstraps on each other, et cetera. So, no planning.
If you want to add planning to that, what they did, which is quite brilliant, is say, “Instead of building some other system to try to choose good trajectories, why don't we just use the policy network?” We don't just pick the first one—we pick its favorite move, but then we also look at its second-favorite move, its third-favorite move, and maybe its fourth-favorite move. Then let's literally play the games out.
Let's play a bunch of games against ourselves and then see the ratio at which we win them. What we might learn is that our second-best guess was actually better than our first-best guess. But we're not starting from every possible possibility. We're saying, “Let's bootstrap on our best guesses of good moves, but then check them by playing out the possible futures.”
If we were to analogize that to the brain, what that would suggest is that perhaps it's the basal ganglia—which a lot of evidence suggests is engaging in this type of model-free reinforcement learning—that is actually the thing that chooses the moves. But there's some other system that lets us choose the second-best move, the third-best move, et cetera.
One way this might happen—there's some evidence for this, far from conclusive—is that there's some notion of uncertainty that the frontal cortex or basal ganglia is measuring. When the level of uncertainty between the next actions passes a threshold, pausing occurs. When we see animals do this vicarious trial and error, it almost always occurs in moments of high uncertainty: when contingencies have changed, when the right answer is not obvious.
You could conceive of this as a policy network where you're evaluating its best choice, second-best choice, third-best choice. When there's uncertainty about it—in other words, when they're close together—or there's some other measure of uncertainty, perhaps you have parallel policy models and you're comparing the similarity between them. There are a lot of different ways to do this. That triggers a process of playing forward.
This is another key thing that mammal brains do that AlphaGo does not do. AlphaGo engages in planning on every move. There was never the question of when to pause to plan in a game of Go. It doesn't matter; just engage in planning on every move because it can do it so fast. In the real world, there's so much uncertainty and noise, and human brains need to be so energy-efficient that we can't engage in planning every instant. We need some mechanism that tells us when to stop and think about what we're going to do next and when we can just continuously go with model-free choice.
This is also something we don't know how mammal brains do. But I think a reasonable speculation is that there's some uncertainty measurement occurring.
One last point I'll make that I think parallels, in an interesting way, some of Hawkins's ideas and Friston's ideas: if you take the Thousand Brains model and apply it to the frontal cortex—in other words, we have multiple parallel models of ourselves—you could imagine that there's an uncertainty measurement, the same way we do uncertainty measurement in a lot of deep-learning models, where you create parallel models and just measure how similar the predictions of the models are.
If multiple parallel models predict similar things, we measure that as low uncertainty. When they diverge substantially, all of a sudden we measure high uncertainty. Again, this is speculation, but you could imagine that, if it is the case that we have redundant models in the neocortex, then might it be the case that somewhere—perhaps the thalamus or the basal ganglia—the similarity or differences between these predictions are a measure of uncertainty that triggers pausing?
Stephen Grossberg has similar ideas. He calls this matching and nonmatching.
Tim Scarfe
Yeah. Yeah, almost. Our ability to do abduction is something that fascinates me, and there is some kind of model-selection or matching step that goes on there.
Anyway, we've got the neocortex. It's an absolute beast when it comes to predicting sensory signals, and then we see the emergence of planning and also smart planning, as we've just spoken about. So now we're actually talking about traversing these sensory networks over space and time.
Then something else really interesting happens. The next 2 moves you make are bringing in selfhood—bringing yourself in as an explicit actor and including that in the planning—and then that naturally leads to this idea of, I would call it, teleology, but you would call it “why” or intentionality.
So let’s go on that journey. Maybe self-modeling first: how does that come into the picture?
Max Bennett
I think there are 2 notions of self-modeling. One notion of self-modeling is the kind that I think we see in early mammals. This is an idea where the frontal cortex of mammals—a region generally called agranular prefrontal cortex, which is present in all mammals and is largely believed to have existed in the very first mammal brains—gets sensory input from an animal’s own interoceptive signals.
So it gets input from the hypothalamus, which measures things like hunger. It gets input from the amygdala, which measures things like valence in the world—fear, danger, and so on. When does the agranular prefrontal cortex get most excited? It’s in moments of uncertainty and in moments where animals are engaging in planning and episodic memory. If you damage the agranular prefrontal cortex in rats, they seem to have dramatically impaired, if not completely lost, the ability to engage in mental simulation, episodic memory, and so on.
I think what one might speculate is happening here—which is not a novel suggestion by me—is that it’s modeling the self. In other words, it’s modeling the activations of the amygdala and the hypothalamus that are happening, and why I am doing what I’m doing. If I wake up and I see that I have these hypothalamic activations, and then I go down to get water, then it builds a model: in the presence of these hypothalamic activations, the next action is, “I’m going to go over and get water,” which constructs an explanation.
As odd and philosophical as that sounds, that is, in principle, computationally exactly the same thing as when we showed a picture of that triangle and your posterior sensory cortex constructed an explanation of what it saw: “I perceive the triangle.” This idea of constructing an explanation of one’s own behavior is the first idea of self, which is constructing intent.
There’s lots of evidence that even in rats, if you record neurons in their agranular prefrontal cortex, they seem to be very sensitive to the tasks they’re in and to measuring progress toward goals. There’s lots of evidence that it’s doing something akin to that. So that’s one notion of self.
When you get to primates, you see a whole new region of frontal cortex emerge: what’s called granular prefrontal cortex, which is only seen in the primate lineage. There are no other mammals that have this region of prefrontal cortex.
As a quick aside for anyone who’s interested, the reason it’s called granular versus agranular is that most neocortex has 6 layers. The 4th layer is called the granular layer because it contains granule cells, which are just a certain type of neuron. For a variety of really unknown reasons, although there are some interesting speculations, agranular prefrontal cortex is missing layer 4. That’s why it’s called agranular. It’s the same thing with motor cortex, which is missing layer 4.
Most mammals’ whole frontal cortex is missing the 4th layer. It only has 5 layers. But in primates, you get this granular prefrontal cortex, this huge region of neocortex that does contain a layer 4. The best explanation for this I’ve actually seen Friston talk about. I’m happy to go into it if you think your folks would be interested in it, but the point is, there’s a new—
Tim Scarfe
Oh, yeah. I mean, we love Friston. Active inference is actually about preferences because an agent expresses agency by adapting the environment to suit its preferences, or to make the environment like its preferences. So this is a theory of volition, right? What Karl Friston is talking about is: where do these goals come from? Where does volition come from? You spoke with Karl about that, I think, over several years, didn’t you?
Max Bennett
Yep. Karl’s been a wonderful mentor of mine. It’s a funny story: I won—I didn’t know this was a term in academia, but there’s the reviewer lottery, where you just get lucky and get a reviewer. For the first paper I submitted, Karl Friston was a reviewer on it, which I was just lucky to have happen. Then he became a mentor, reviewed my book, and gave me lots of good feedback. So, yeah, he’s an amazing person and has been a wonderful mentor of mine.
Let’s talk about granular versus agranular, because I think the best theory I’ve seen is Friston’s theory on this. What does layer 4 do in neocortex? Across the entire neocortex, layer 4 is where sensory input is received. The primary sensory input is received into the neocortical column. This comes from the thalamus.
The canonical model is that sensory input from sensors—eyes, ears, skin—flows up through the brainstem to the thalamus, and then from the thalamus propagates to layer 4. From layer 4, it goes within the variety of other layers of neocortex, and then the other layers of neocortex project back to the rest of the brain.
Why would it be the case that regions of neocortex would not have a layer 4? If you actually watch an animal’s development, what’s interesting is that, in mammals with an agranular prefrontal cortex, it’s not always agranular. It actually starts having a layer 4, and the layer 4 atrophies over development.
I think this mirrors well with Friston’s idea of active inference. What’s happening is that the neocortical column can be in 2 phases. It can either be trying to match its model of the world to its sensory input—in other words, I see sensory input, I’m trying to infer what’s there, and I’m going to construct the idea of a triangle—or there’s another state of a neocortical column, which is generation: I’m going to start from the latent representation of a triangle, imagine it, and explore it.
One idea is that what frontal cortex does is primarily try to fit the world to its model. In other words, it spends the vast majority of its time constructing intent, not trying to modify that intent to fit what it observes, but instead trying to change what an animal does to satisfy its intent.
Layer 4 atrophies. It doesn’t actually go all the way; if you go deep into a brain, you see some basic layer 4. So it’s not completely gone, but it atrophies because frontal cortex spends very little time trying to change its perceived intent to map what the animal is doing. Instead, it tries to change what the animal does to match its intent.
What I think is so interesting and brilliant about this idea is that it explains exactly why layer 4 doesn’t start out as nonexistent. At first, an animal needs to build a model of itself, so layer 4 is present. But over time, it shifts toward, “Once I have a model of what I want, who I am, and the things I would do, I don’t need to spend as much time changing my model of self. I’m going to spend most of my time trying to change my behavior.”
This is a very speculative idea, but it makes a lot of sense in the context of active inference. It’s the best explanation I’ve seen, personally, in all my reading of explanations for why agranularity exists.
Tim Scarfe
This is absolutely fascinating. In my mind, it bridges the gap between internalism and externalism because he’s describing this kind of dialectic exchange. To use the high-entropy words that only Friston uses: if I say “dialectic exchange,” you know. And, by the way, for the folks at home, if you read Friston’s papers, there are certain words he uses. He says, “This licenses something.” If you see the word “licenses,” then Friston wrote the paper.
But you have this kind of thing where agents have models of the world, but they’re exchanging information with the other agents. Then what an agent does is have this generative model of policies, which is just a sequence of actions.
Here’s where I want to get into the nitty-gritty a tiny bit. You could think of those plans as being goals. You can think of a goal as just being an end state in one of your plans. But that doesn’t really satisfy me, because I think of goals like eating food as being a kind of category, not a pointillistic traversal of a specific state in the future. It feels like, well, if there are an infinite number of goals, then in some sense there are no goals at all. So what is a goal to you?
Max Bennett
Great question. I think this is where semantics matters a lot. We can think about goals in several ways. If we think in strict RL terms, they would just think about goals as optimizing a reward function. The goal is simple. There’s only 1 goal, which is to maximize reward. In a changing, complex world, your reward function might fluctuate over time, but the goal is singular.
In the active inference world, what I find compelling is that it introduces a different component of what we mean by a goal. This is not just of intellectual interest because it’s cool; it has very real AI implications, actually, because it contains the notion of explainability.
For example, if I wake up and I’m hungry, and I start imagining ways to satiate my hunger, then I decide I’m going to get in a car, go to this restaurant, and eat this specific food.
When I get into the car and someone calls me and goes, “Why’d you get in the car?” the reason I can explain that so easily is because there was a rendered simulation of a plan that terminated in an end state that I deemed I wanted, that I selected. So it’s very easy to explain why I did that. In the absence of that, it’s actually very hard to explain why you’re doing things.
If you’re walking down the street, if I asked you to explain any one of your model-free behaviors—why did you move your foot there instead of there?—you have no explanation. And so I think you can assign the word “goals” to multiple levels here. I think you could say it’s terminating the end state. You could say the goal is some more abstract representation of the satiation of thirst in general, and you could think about that as distinct from just the reward function.
Some might challenge that and just say, “Well, all that you’re talking about, then, is just optimizing the reward function.” But I think you could make an argument that there is a distinction happening there. What I think is so critical, and that’s unique to what happens within mammal brains—and it’s why I was really honing in on that—is the ability to plan a series of actions that terminates in an end result and then execute that plan. I think that has very clear implications for explainability, which model-free actions do not.
That would be how I would think about goals. But I think it is a little bit semantic, where we can think about the concept of goals in multiple ways.
Tim Scarfe
How do we actually know what the cognitive abilities were of early animals, and why should we care?
Max Bennett
Great question. I think there are 2 reasons why we should care about the evolution of our brains and intelligence. The first is to understand who we are. The scope of what it means to be a human is not constrained to what it means to be a Homo sapiens.
So much ink has been spilled on the last 70,000 years of us being Homo sapiens. But if aliens were to come down and engage with us and analyze us as a species, most of the things they would observe about us don’t come from our legacy as Homo sapiens. They come from our legacy of being a primate, our legacy of being a mammal, our legacy of being a vertebrate, and our legacy of being an animal in general.
If we want to understand what it means to be a human being, I don’t think we can skip the full 600-million-year story of how we came to be. In there is so much rich history and insight about what it means to be us. That’s one really key reason: it’s our legacy, it’s our history, and it’s how we came to be.
The second, perhaps more practical, reason is that I think understanding the evolution of the human brain and the evolution of human intelligence is a key tool in our toolbox for understanding how the brain works and how human intelligence works. It’s by no means the only method. It might not even be the main method, but it’s a very useful method to add to the toolbox.
The problem with going into the human brain and trying to directly reverse-engineer it is that evolution doesn’t work in clean ways. It doesn’t work the way a human designer would. It doesn’t work from first principles. It tinkers.
When we go into the brain, we see all of this messiness. There are redundant systems and vestigial systems. New things evolve that make old things redundant, but they’re still there. Lots of processing is duplicated in different regions.
One way to understand the brain is to continuously probe it as the human brain is, which is fine. But another method that’s also useful, and can impose constraints for us, is to actually track the history of how it came to be.
That can provide insights as to when this brain modification occurred, such as when the neocortex evolved or when the basal ganglia evolved, what new abilities this enabled, and how it affected the prior brain regions that were already there. This can give us insight into how the brain works today.
In the toolbox that we have for ways to reverse-engineer the brain, I think this is just an underappreciated one that is worthy of being included. Of course, understanding how the human brain works has so many different applications. It helps us with mental health, it helps us with understanding why people do what we do, and it helps us with building AI systems. I think there are lots of insights to garner from the brain.
That’s why to do it. Now, how to reverse-engineer what behavioral abilities exist in our ancestors is a really interesting question. Of course, we can’t go back in time. What we can do, though, is use mechanisms to reverse-engineer what their brains looked like and mechanisms to reverse-engineer what abilities they had.
To understand what their brains looked like, we can look at other animals in the animal kingdom. For example, we can look at all of the existing primates and all of the existing non-primate mammals, and we can see what the common brain structures are between them.
We can look at genetic analysis, meaning what things seem to derive from similar roots, and we can infer what seems to be common and shared amongst them. Thus, what do we think was actually existing in the brains of the first mammals?
We can do the same thing with fish and reptiles to infer what existed in the first vertebrates, and we can do that with invertebrates to try and infer what existed in the first bilaterians—in other words, the first animals with brains.
For behavioral abilities, there are 3 ways you do this. This is my approach to trying to infer behavioral abilities. I call them the ingroup condition, the outgroup condition, and the stem-group condition.
In order to make the argument that a behavioral ability emerged at a certain location in our evolutionary history—for example, a behavioral ability like episodic memory evolving with the first mammals—you need to satisfy these 3 criteria.
The ingroup condition stipulates that most descendants of this species—in other words, most mammals—should show this ability. It doesn’t mean all of them. Abilities get lost all the time, but most of them should show this ability.
The neural mechanisms by which the ability emerges should come from homologous regions. What that means is a shared neural underpinning. If, for example, mammals show episodic memory but it comes from neurological regions that independently evolved along different mammal lineages, that suggests it wasn’t present in the first mammals.
But if they all emerge from regions that emerged with early mammals, that’s good evidence that this ability also emerged with mammals.
The outgroup condition says that most doesn’t have to be all, but at least many non-mammal vertebrates—the outgroup, the group right above—should not show this ability. If they do show the ability, it should emerge from non-homologous regions; in other words, parts of their brain that evolved independently.
For example, birds definitely show episodic memory. But when we look into the brain regions from which episodic memory emerges, it’s clearly non-homologous. It’s a part of the brain that mammals, or the early vertebrates, did not have.
The stem-group condition is that, in the ecological dynamics in which early mammals existed, or this ancestor existed, we should be able to devise an argument for why this ability would have been adaptive. Why would episodic memory have evolved?
With these 3 things, we can start to infer the story of when behavioral abilities emerged. Is this perfect? Absolutely not, because we do not have a ton of data on behavioral abilities across species. As new evidence emerges, the story might change.
But with these 3 conditions, we can do a reasonable job inferring what abilities emerged. The main finding of the book—or the research that led me to be so excited about the book—is that when you do this, what’s kind of crazy is that you find, as a first approximation, a really coherent story.
A lot of the behavioral abilities that emerge at each milestone in brain evolution don’t seem to be a haphazard array of different skills. They often emerge from one underlying intellectual capacity, which I call a breakthrough, applied in different ways. That’s the idea of the 5 breakthroughs.
One thing to note about the basal ganglia is that I think this is a bit of a sidebar, but I think it’s fun. The basal ganglia is one of the most underappreciated parts of the brain, in the sense that so much work has gone into understanding how the neocortex functions. So much work has gone into the neocortical column, and that’s all wonderful work to be done.
But the basal ganglia is not only evolutionarily much older. If you look in a lamprey fish, as we talked about last time, a lamprey fish has a common ancestor with us from 500 million years ago. It’s one of the most distant vertebrate cousins that still exists today.
Lampreys have a basal ganglia that looks exactly like our basal ganglia: the same internal structure. The basal ganglia also has perhaps one of the most beautiful internal structures that can be computationally reverse-engineered.
There isn’t good consensus on the actual internal wiring and computations performed by a neocortical column, but there is much broader consensus as to what’s being executed by the basal ganglia. Without getting overly technical and perhaps boring people, I would encourage anyone who’s computationally interested in this to dive into the literature here.
It is almost beautiful that evolution came up with this. For example, the input structure of the basal ganglia has this mosaic of neurons that each express 2 different types of dopamine receptors.
This would be my own little technical diatribe. One type is called D1 receptors, and then there are D2 receptors. D1 receptors, when they receive dopamine, strengthen connections. D2 receptors, when they lose dopamine, strengthen connections.
Now you track these different neurons, and they actually split their paths. The D1 receptors go to a nucleus that, when activated, disinhibits behaviors, and D2 neurons go through a different set of nuclei that, when activated, inhibits behaviors. We can literally watch how dopamine signals drive repeated behavior and how dopamine drops inhibit behavior.
We can literally look at the mosaic of connectivity here and say, “When you spike dopamine, it weakens the stop pathway through D2 neurons, and it disinhibits the go pathway through D1 neurons, making you more likely to repeat the behavior, and vice versa if something bad happens and you lose dopamine.” I think it is so beautiful that evolution stumbled on something that clean in its macrostructure. So, diatribe on—I think the basal ganglia is cool.
Tim Scarfe
Yeah, it’s quite interesting as well. People get addicted to drugs, and a lot of that is about wireheading in the basal ganglia. Of course, habitual learning is something that, when it becomes so habituated, moves down the stack into the basal ganglia, and drug-taking would be an example of that. You actually cited, I think, an experiment in China where they removed part of the basal ganglia, and there was a 40% recidivism rate for addiction.
Max Bennett
Yeah, a very controversial study that probably violated many ethical codes in the US. They did this study on people with intractable heroin addiction, and they lesioned a part of the basal ganglia called the nucleus accumbens, which is sort of where goals are habitually selected. It showed a dramatic reduction in heroin addiction. It also had other side effects that doctors might deem unreasonable, but it definitely worked and absolutely reduced the addictive cravings triggered by stimuli.
Tim Scarfe
Yeah. The story of this chapter—an incredible chapter I’ve just been studying today in great detail—is the story of mentalizing, but I would call it social complexification. Actually, that’s a hypothesis for why our brains expanded so dramatically. There was this extinction event. I think it was the Devonian extinction event, and only birds and not many other things survived. Then we got the chance to evolve after that, and our brains rapidly exploded.
There are different theories about why that happened. Maybe it was because we had access to loads of calories in the form of fruit—preferential access—and it gave us an incredible amount of time and energy, the excess of which might have led to social complexification. Can you just give us a little bit of background about that first piece?
Max Bennett
Absolutely. We don’t know—there are lots of speculations—but we do have some really good evidence that at least part of what drove the explosion in primate brains was social dynamics. Robin Dunbar did the seminal work here. What he showed is that, in primates, the encephalization quotient—which is just the ratio of brain size, especially the neocortex ratio, so the ratio of the size of the neocortex to the rest of the brain—is extremely correlated with social group size in primates.
The bigger the social group size in a group of primates, the bigger their neocortex seems to be relative to their body size. What’s so interesting is that you don’t see this in most other mammals. This is not a standard correlation that applies across the animal kingdom. It seems to be a correlation that’s very specific to primates. There might be other mammals, but for most mammals, you don’t see this correlation.
Robin Dunbar’s famous social brain hypothesis is that what drove the explosion in primate brains was some form of social dynamic between people. The more social relationships you’re managing, the bigger your brain has to be. Now, that doesn’t explain what specifically happened in the brain, which is what we can get to with mentalizing and why that applies. But it does suggest that whatever drove this explosion in brain size seems to be something correlated to social grouping.
What’s interesting about primate social groups, relative to—not all mammals, but many other mammals—is that they’re very, very political. Many mammals live in solitary social lives, where the males mostly live alone and females will rear a child and then usually go off on their own. There are animals that live in herds, where they socially group together, but there aren’t really very rigorous hierarchies amongst them.
Primates, especially apes, have these really complicated social structures with rigid hierarchies. There truly is someone at the top of the hierarchy, and we can measure this. Primatologists have gone to painstaking lengths to verify it. For example, there’s transitivity: if you show that one primate tends to show a submissive signal to another primate, and that other primate shows a submissive signal to another one, then it’s almost definitely the case that the first one will show a submissive signal to the last one.
In other words, these are real, rigid hierarchies, not random interactions of submission and dominance. One of the main ways you survive as a primate is by successfully climbing this hierarchy. What’s so interesting is that, in many mammals, what makes someone the top dog, or the person at the top of a hierarchy, is brawn. It’s just strength. They’re trying to flaunt who would win in a physical altercation, which evolutionarily is beneficial, because if you can prevent actually fighting each other and just say, “This is who would win the fight,” then you both save energy by having these fake battles, and whoever wins gets to eat the food, et cetera.
But with primates, it’s not always the strongest one that reaches the top. It’s the most socially savvy one. Social savviness comes into play through alliances that are built within primates. You’ll see that people at the top of the hierarchy frequently befriend, groom, and come to the aid of certain other non-family members. Those folks will thus reciprocate and come to their aid. There are these really interesting dynamics that play out. You even see wars. There are mutinies that take place in primate societies.
In this soup of the way to survive and gain evolutionary advantage as a primate, the goal is not perhaps only to make sure you get access to food, but to climb a social hierarchy. All of a sudden, there are huge social pressures to infer what someone else would do in a certain circumstance, what someone knows, what you can get away with, or how to change someone’s opinion of you.
What that lines nicely to is what we see in the new brain regions that emerged in primates. Most notably, there’s a brain region called the granular prefrontal cortex, and there are areas in the back of the brain called the superior temporal sulcus and the temporoparietal junction. These brain regions across primates are highly implicated in what I call mentalizing—which is thinking about thinking—but the standard literature would call this theory of mind. That means being able to infer the intent or knowledge of someone else.
It’s easy to understand why this would be so adaptively valuable in a politicking arms race where you’re trying to deceive each other. There’s a great study that really revealed this with primates by Emil Menzel in the 1970s, and I love this story.
Tim Scarfe
Are you going to do the Machiavellian apes?
Max Bennett
Yes.
Tim Scarfe
Yeah. Yeah. God. [Laughter]
Max Bennett
Okay. So, Emil Menzel had this 1-acre forest, and his main objective was not to study ape social behaviors in the sense of how they would climb social hierarchies. His only objective was to measure spatial reasoning in chimpanzees.
He had a group of chimps. There was 1 chimp named Belle, another named Rock, and a few others. He would show Belle the location of food. He would hide food under a bush and then see whether Belle would go back to that same location looking for food. In other words, could she remember locations in 3D space or in a 2D, map-like space?
What he found is that, yes, they readily do that. We now know that lots of mammals are capable of it. In fact, even fish can do things like that. But in this study, he started finding something that was odd.
When Belle would find the food, she would frequently share it with her fellow chimpanzee group members. That was great until Rock, who was a high-ranking, aggressive male, would take the food from her when she shared. So what she started doing was hiding the food when she found it. Instead of sharing with Rock, she would just sit on the food.
Rock realized she was doing this and not sharing. So Rock would come over and push her to try to get the food from under her. Then, when she knew the location of the food—because, on some recurring cycle, experiment, or signal, she knew that the food was now available—she would not go to it until Rock was not looking.
So then what did Rock start doing? Rock started pretending not to look. Rock would look away while Belle went toward the food. Once he noticed she was doing that, he would turn around and run to try to grab the food before her. Then Belle started trying to lead him in the wrong directions, and this cycle of deception and counterdeception kept playing out.
What that becomes is a beautiful anecdote and a case study in what happens when you have a bunch of hierarchically interacting animals in an arms race for things like this. What you get is deception and counterdeception. That’s really only conceivable with some notion of theory of mind, because in order for Belle to trick, or try to trick, Rock, she needs to be able to say, “In order for me to change the knowledge in Rock’s head of where the food is, what I need to do is walk in this other direction. What that will do is make Rock think the food is in this direction, when in fact I know it’s in this other location.”
She also needs to reason about someone’s intentions: “I know Rock intends to trick me, so when he’s looking away, I don’t believe that he in fact is not paying attention.” This was one of the first early anecdotes that some form of theory of mind was occurring in primates. There have since been lots of studies that show this.
For example, just to give some case studies, you can take a chimpanzee and teach it that, when there are 2 boxes, the box with a red mark on it is the one with food in it. They’ll easily learn that. Then you have an experimenter come in with 2 boxes, bend over, and mark one. They pretend to accidentally drop the marker on the other one and then leave. The marking is identical in both cases, but the chimpanzees always go for the one that was intentionally marked. They can infer the difference between the same stimuli: someone intending to do something and something being an accident.
There are other studies of chimpanzees playing with different goggles. With one goggle, you can’t see through it; with another, you can. If you put those goggles on human experimenters, the chimpanzees always go to the experimenter with the see-through goggles, asking for food. They’re somehow inferring that the other person can’t see them, so why would they ask?
There are lots of studies that show this ability. Evidence outside of primates, in other mammals, is very loose. It’s inconclusive and controversial, but the loose evidence shows up only in the smartest mammals, which possibly suggests some independent convergence.
Tim Scarfe
Yeah.
Max Bennett
Yeah, there’s lots of rich evidence that this theory of mind exists within primates, and it emerges from these uniquely primate regions. I’m happy to go into the evidence from brains, but I’ll stop there for a second.
Tim Scarfe
So, with the Machiavellian apes and the X-risk people, there are people who talk about AI killing everyone, and they make the argument—it’s called instrumental convergence, from Nick Bostrom—which is basically that things like power-seeking and deceptive behavior would be instrumental to any end goal. This is a great example from the animal kingdom of deceptive, Machiavellian behavior.
I guess it does seem plausible, at least on the surface, that a level of sophisticated agents following their own intentions and inferring the intentions of others would seek to deceive each other. That’s a natural phenomenon.
Max Bennett
I absolutely think so. I absolutely think it is the case that the more autonomy you give an intelligent agent, and the more ability you give it to define its own subgoals, the more risk emerges. You absolutely get what Nick Bostrom is talking about, which is that a subgoal to trying to help cure cancer might be to dominate all of Earth and control the labor supply and allocation of resources across all of Earth.
But I don’t think that’s necessarily inevitable. I think it is a risk. Evolution is a constrained search algorithm for intelligent entities. It does not give moral weight to what emerges. This is an important distinction: just because something is a natural consequence of evolutionary systems does not mean that we should deem it morally superior.
Tim Scarfe
Yeah, that’s the naturalistic fallacy, right?
Max Bennett
So it might be the case that it is very likely that species will eventually enter a politicking arms race, and certain forms of deception and power-seeking will emerge. That doesn’t mean that when we produce our own intelligent entities in AI, we should imbue them with those features.
One of the optimistic outcomes of this new AI world we’re going to enter in the next 100 years is that we, as designers, can now do our best to try and remove some of the evolutionary baggage that we don’t like, which has evolved in humans, from these new entities. There’s risk, of course, but I think there’s also a really great opportunity that we could have benevolent beings that do not seek to dominate. Yann LeCun talks a lot about this, and about creating beings that are less selfish.
So, I think there’s a great opportunity, but there’s definitely risk, because the second you give an autonomous agent the opportunity to produce its own subgoals, you need to have really rich constraints, a really well-defined reward function, or both. One ability that I think comes from mentalizing—and this is an idea in alignment research—is that if you can convince an AI agent to try and do what it thinks the human wants it to do, what you’re actually doing is requiring it to engage in some form of mentalizing: to infer the preferences of the requester and then try to do what is best for that individual.
You can’t just have it take requests at face value, because then there are all these opportunities for misinterpretation. Nick Bostrom’s famous paperclip maximizer maximizes the production of paper clips, and Earth is turned into paper clips. We obviously don’t want that.
But with mentalizing—with the ability to model the internal simulation of another mind and play out how this person would feel about possible futures—you could imagine, optimistically, an outcome where an AI agent could easily infer, “If I turn all of Earth into paper clips, that’s not what the person giving me this request would in fact have wanted. They would regret that outcome.”
Of course, it doesn’t fully derisk things, but it is one methodology and one learning from evolutionary neuroscience that we can garner. Mentalizing is a tool that can be used to try and stabilize the requests we give each other in a more grounded way, so there aren’t these types of misinterpretations. Of course, humans misinterpret each other all the time. It’s by no means perfect, but it is a tool.
Tim Scarfe
And I think that’s a very natural phenomenon. I think any intelligence system is naturally incoherent. I think it’s impossible to have a single, monolithic intelligence that is monomaniacally focused in a particular direction.
But I want to slightly rewind to what we were saying. The first animals had quite simplistic social games that they were playing. They were interested in strength and submission, and it was a fairly fixed interface.
What was really interesting is that deer, for example, lock horns, don’t they? It’s predictive. They don’t actually have to have a fight, because that would be evolutionarily not a smart thing to do. The social game they play, even though the game is fixed, is predictive, which is fascinating.
Then you were telling the story of monkeys and macaques, how they have this really interesting virtual social game where strength and social status diverged. Social status actually became this virtual thing that was based on grooming and preening and lots of completely unrelated things. It was entirely possible for a very weak macaque to have significantly higher social status than a big, strong one.
That’s really fascinating. But then we get into a broader question. We’re still very social creatures ourselves. We have Facebook, for example, and could you arguably, cynically argue that Facebook—or all socializing—is just a kind of arms race to improve our social status? When we’re posting on Facebook, in a way, it’s like the deer locking horns. It’s us playing these status games without having to fight each other.
Max Bennett
I think there’s an aspect of human behavior that can absolutely be explained by this. There’s a great book called The Elephant in the Brain.
Tim Scarfe
Oh, yes.
Max Bennett
There’s a great book called The Elephant in the Brain that talks about how much human behavior can be explained by this sort of status-seeking behavior. The reason why it’s so hard to study is because it’s what they call, I think, a cognitive taboo or an intellectual taboo. We don’t want to admit it to ourselves.
So we self-delude ourselves into believing we’re doing things for virtuous reasons, because it makes it easier and more convincing. If I know that I’m doing something to deceive someone, it’s easier to tell that I’m doing it. If I genuinely believe that I’m doing these things to help the world, but subconsciously they’re actually just benefiting me, it’s more convincing to other people.
Their argument is that primates evolved this sort of self-deception to make themselves more convincing, and this really accelerated with language in humans and all that.
I think it’s very likely to be the case that the core thesis of their book is right: a lot of human behavior is this sort of subtle status seeking. I don’t want to go on a tangent here, but I do think it has sociopolitical theory implications. How do we make sure that society doesn’t devolve into just a hedonic treadmill?
The interesting thing about social status is that it’s always, definitionally, a scarce resource because it’s a ranking game. Unlike physical resources, where it’s possible for all of us to live better than kings or queens did 1,000 years ago—we can all have better access to medical care, better access to information, and better access to food—social status is always a zero-sum game, unless someone can conceive of a better way to do it.
This is problematic because if, over time, most of our actions become about pursuing social status, then we’re going to forever be in this sort of game. I don’t think personally that we’re doomed to this. I think there are absolutely better virtues in human psychology, where not everything we do is based on pursuing social status, and I think you can conceive of dynamics where humans are doing things for other reasons, not just to gain status. But I think it’s absolutely fair to say that a surprising amount of human behavior is status seeking, and maybe a depressing amount.
Tim Scarfe
Yeah. I agree with you, and I don’t necessarily want to get too philosophical on that, but that book was The Status Game by Will Storr, where he said that there were 3 meta-status games that we played. He gave the virtue game, the dominance game, and the success game. I might be playing the success game: I want to have the best podcast, or whatever.
The reason I bring this up is that the difference between humans and animals is that they’re just playing one game. It’s really interesting that they have this mimetic social score, but the game is the same everywhere, whereas for us, we go 1 level of abstraction up. The success game for us can be manifested in a myriad of different ways. It could be success at playing computer games; it could be writing books or making podcasts, or whatever.
It’s almost like we fractionate our social ranking into a myriad of different games. I think that’s a little bit of a testament to the difference we humans have in general with our metacognition, which is our ability to create the memetics in a novel way.
Max Bennett
One thing that’s sort of related to this, at least in early human societies, is one way to reduce status infighting: make it such that members of a team have distinct roles. I don’t think this lesson only comes from management theory and entrepreneurship. I think this probably derives from either early human or maybe even early primate societies.
It’s much more stable, and you can introspect that it feels much more comforting to be part of a troop of 100 humans where pretty much everyone is pulling their weight and everyone matters because they’re doing their own distinct thing. That is a very stable state, where we’re not infighting as much because we’re all doing something; we all matter to some degree.
But when there’s infighting because there are only 5 blacksmiths, or 5 podcasts, or 5 books about the evolution of the brain, all of a sudden these other types of things start emerging because we’re no longer all fulfilling a role that matters. It feels like there’s only a ranking, and only 1 of these is going to matter.
As an example of ways to reduce this sort of status seeking, I see this in business all the time: the more you can create an environment where it’s not zero-sum, where everyone’s pulling their weight and together we all win, the more the best versions of humans emerge. The more zero-sum it becomes, and the less distinct the roles and activities are, the more of these—I would argue—primitive primate behaviors start emerging.
An early-stage company, a company of 30 people, has such different dynamics from my last company, where, when I left, we were 400 people. The social dynamics are so different, and I do think one could speculatively correlate that to our evolutionary history here. In a 30-person company, you don’t need that much structure. If you have people who work well together, are aligned on a mission, and you get rid of people who are generally mean-spirited or have bad intent, you don’t need a lot of structure and process to get people to work well together, support each other, and move in a common direction.
I think what that demonstrates, when one observes that, is that what’s playing out is an evolutionary program that got groups of 30 humans to work really well together. When you’re at 400 people, what very quickly happens—and it takes a lot of work to fight this—is that you start getting internal factions emerging, because what splinters out is these subgroups of 30 to 100 people that then have their own points of view. Then it’s very easy to have an us-versus-them dynamic with other groups, and you start seeing things break down.
One mechanism for solving this is very rigid hierarchies; that’s what the military does. Another mechanism is to embrace the chaos, which is a little bit more what Google does. Another mechanism is to effectively make it a constellation of different startups, which is what Amazon does, where each group is kind of autonomous and has very clean interfaces with other groups.
There are many different management approaches to this, but the breakdown, I do think, emerges from the fact that humans did not evolve to interact with 100,000 people. We evolved in an environment where we interacted naturally with about 100 people, and that’s why that comes very naturally. We don’t need as much process to make that work, but we do when we scale it up.
Tim Scarfe
Yeah. It’s fascinating. I mean, as you say, you could argue that Amazon has 1 overarching goal: to make money. But as soon as you increase the autonomy in the organization, it’s a very human trait, isn’t it? You were talking about the Machiavellian behaviors and the deceptive behaviors, and you just wonder how much energy is wasted with infighting.
I’ve even made the comment that in the military, they might be doing quite simplistic jobs compared with Google, but even at Google, there’s an obsession with job level. I mean, if you go off of Levels.fyi or if you go on Blind, that’s the only thing people talk about: their total compensation and job level.
Maybe we should save the cynicism. Coming back to the chapter, we’re telling the story, basically, of how this metacognition and this predictive apparatus gave rise to an entire suite of complex social behaviors that we see in primates, which is fascinating.
Maybe we should just talk a little bit about what I call why bootstrapping. There was 1 guy at Toyota Research who was quite famous because he would get people to ask why 5 times. You say, “Why? Why? Why?” It’s almost as if there’s some magic number. Everyone is only a certain number of degrees of separation away, and it’s a similar thing: you only need to ask why a few times and you’ll always get to some kind of base reason.
Maybe that’s why, evolutionarily, we have 2 levels of causal metacognition in our brain. We have the agranular prefrontal cortex and we have the granular prefrontal cortex. I guess 1 potential question there is: why is there not a 3rd level of asking why, and what would that look like? Can you just sketch out that metacognition picture in general?
Max Bennett
When we think about what a granular prefrontal cortex does, a reasonable framework for it is that it generates explanations of an animal’s behavior. It models an animal’s own behavior. One cognitive tool to reason about that is this: if it observes a rat wake up, have certain hypothalamic activations, and run in a certain direction to drink water, it produces a representation that could be interpreted as, “I am explaining this behavior by: I am hungry, as an animal.”
That can be useful in a variety of ways. It can trigger simulations to find alternative solutions to satiate the same need. If you put a rat in a novel situation, but the granular prefrontal cortex infers that right now I am hungry, we can start triggering a bunch of simulations to try to satiate the same desire, to fulfill what I believe about myself through alternative means. This enables an animal to be flexible.
This is the explanation of an animal itself. Why would that be the case? What I argue in the book is that the granular prefrontal cortex builds a model of that model. Instead of a simulation, it’s a simulation of the simulation.
What that would mean is, if you could—as a thought experiment—ask, let’s go 1 step further, or 1 step back. If we could ask the basal ganglia, which is the sort of reinforcement-learning system, “Why did you turn left to go in this direction to drink water?” it would just say, “Because turning left maximizes reward.”
The answer would always be the same. If you ask the agranular prefrontal cortex, “Why did you turn left?” it would say, “Oh, because I’m thirsty. There’s a specific thing that I, as an entity—this animal I’m modeling—want to achieve.” But if you ask the granular prefrontal cortex, it would say, “Well, I turned left because I am thirsty, and that made me think about ways to satiate my thirst. I simulated going to the left, and I remembered water being there because last time I was there, there was water, and so I went to the left.”
And so, in other words, it enables you to simulate different types of simulations and reason about what you would think in a new setting, which, of course, enables you to think about what someone else might be thinking. We do this all the time. Someone doesn’t respond to a text message, someone makes an odd facial expression in a social interaction, and we’re immediately trying to figure out: What is this person thinking? Why would they do this? And so on.
So the first question is: Why do we even need this new level at all? I think one of the main adaptive values is that it enables your survival in the politicking arms race, because now, if I can simulate a simulation, I can infer why you might do a certain behavior, how to manipulate someone’s knowledge, and your intentions behind things. So this is why you would have one layer to go a level above.
You could make an argument that theoretically there should be an infinite scaling up of whys. I think this is maybe a cop-out, but I think there are huge energetic costs to any sort of scale-up. So what that means is, the question is not whether there would be benefits to a third level of hierarchy; the question is whether the benefits of a third level of hierarchy would outweigh the massive energetic costs of producing it.
And so I think that would be my first-blush explanation as to why we might only have 2 levels instead of 3 or 4: because the second level added a clear adaptive value relative to the cost to survive in the politicking arms race, and the third one perhaps was superfluous and unnecessary relative to the energetic cost.
Tim Scarfe
I think having that second level of metacognition does a lot of work, right? And I’m going to talk a little bit about that now. But one of the things is, you can infer the intents and knowledge of others through the same process of doing simulations yourself. So you can kind of imagine yourself doing something, but you can kind of swap out the pointer to be someone else and swap out the knowledge to be someone else. And that’s incredibly valuable.
But the knowledge thing is really, really interesting. So I asked the question last time, and this is something that I’ve been quite confused about, and I feel that reading this chapter has actually really cleared it up for me, which is about goals. Because when you look one level down at the agranular prefrontal cortex, it’s modeling intents, and then this granular prefrontal cortex, which is trying to seek explanations about the level below, which is the agranular prefrontal cortex, is going a level of abstraction up and modeling goals—not intents—but it’s actually modeling knowledge. What it’s doing is categorizing.
So when you have a simulation of a simulation of simulations, what it’s doing is creating a category. So, to the example you just gave before, “thirsty” becomes a category. Rather than it being a pointillistic intent, it’s a little bit like saying, “I can go and have a sandwich, or I can go and have McDonald’s,” or my abstract simulator could kind of draw a boundary around those things, and now I’m getting food.
As well as being able to categorize intents in yourself and other people, you’re also categorizing knowledge, and then it can be shared memetically. So it’s almost like just going to that second level of metacognition gives you so much that you didn’t have before.
Max Bennett
100%. Yeah. I think thinking about the level of the granular prefrontal cortex and the new primate regions as enabling something akin to knowledge is a really wonderful way to look at it, especially, one, from the connectivity analysis, and then, two, just from what we mean by knowledge.
From the connectivity analysis, if you look at the superior temporal sulcus and the temporoparietal junction, these are regions of the posterior cortex that, in simple terms, are at the very top of the hierarchy. I mean, they get multimodal input from all the other regions of sensory cortex. So a very simple rule of thumb for understanding this is: this models the rest of sensory cortex. I understand the full rendering of the simulation of the external world that is happening, and this is where I build a representation of that.
And it is perhaps no coincidence that’s also where we see brain regions light up when you’re engaging in things like theory of mind and solving false-belief tests—in other words, trying to infer the knowledge of someone else. These same regions light up.
And what do we mean by knowledge? I would argue that knowledge can mean a few things. One is procedural knowledge, where I just know how to do certain motor behaviors. I don’t think that’s what we mean. I think we mean more semantic knowledge or episodic knowledge, which would be: I know that water is over there, and I know that if I do this behavior, this will be the causal outcome.
That type of knowledge, I think, is absolutely rendered in the mental simulation. When I imagine certain things—when I imagine the case of lightning hitting the ground—what do I see afterwards? I see fire. And that’s the source of my knowledge about the causal relationship between these 2 things. So having a layer that models the simulation enables me to reason about my own knowledge and to see what the effect of changing knowledge is on behavior. And this, of course, enables us to flexibly adapt to other people’s behavior and predict what they would do under cases of different knowledge and different intents.
Tim Scarfe
Yeah, that’s fascinating. But you did say that there was a bit of a riddle about the granular prefrontal cortex, because there was 1 study where it could be damaged and the person would still score really highly on IQ tests. But you said it’s about being able to project yourself in simulations, this kind of abstract modeling of your own mind. So in this particular case, how could the person still score the same IQ without that part of the brain?
Max Bennett
So this is such a cool story in the history of neuroscience. You would think that if you look at a human brain, I mean, the granular prefrontal cortex is this huge region in the front of the brain. I mean, it takes up a gargantuan amount of space. You would think that taking a chunk out of that part of the brain would have a gross effect on a human being.
Just like if you took a part out of even a relatively small region of the back of your brain, which is where your visual cortex is, you become hugely visually impaired. You take a region out of your motor cortex and humans become largely paralyzed for months until they recover from that. You take a region out of auditory cortex and they can lose the ability to recognize even words.
So there are relatively small regions of neocortex that, if there’s damage to them, have gross, obvious effects on human intelligence and behavior. After World War II, there were so many patients with brain damage that there were all these studies, and people could not figure it out. It was a puzzle: What does this huge region of prefrontal cortex do? People don’t have—something seems off about them, but it’s not obvious what is wrong with them. People would note personality changes. They don’t seem to be themselves. But on logic tests, on IQ tests, it wasn’t obvious they were dramatically impaired.
In many cases, it wasn’t obvious they were impaired. There was 1 famous case where they could test someone before and after because, for surgical reasons, they were going to remove parts of the granular cortex. This patient actually improved on IQ tests, which made this a huge puzzle: What does this part of the brain do?
And so, if you track the studies from that point forward, we start learning that what granular prefrontal cortex does in large part isn’t related to these types of logic puzzles. It’s related to thinking about thinking and modeling ourselves.
So, for example, if you look at someone who has damage to granular prefrontal cortex, someone who has damage to the hippocampus, and someone who has a normal brain, and you ask them something very simple—you give them a random word and you say, “Just tell me a story. Just imagine a story of you with this word.” The word could be “restaurant,” and you compare these stories, you immediately see something very different.
The people with hippocampal damage give a very, very rich story about themselves, but the external world misses details. So there’s not a lot of rich detail about the external world. This is consistent with the idea we talked about with early mammals, where the hippocampus helps render a simulated external world.
The people with granular prefrontal damage could render a very rich external world. They could tell you the details of the leaves, the smell of food, exactly what a restaurant looked like, but they themselves were woefully missing from the stories. They could not project themselves into this imagined world.
And so then, if we go back and look at all these other things that light up granular prefrontal cortex, if you ask someone to think about how they’re feeling, granular prefrontal cortex lights up—self-reference. But if you ask another question, such as, “What does it look like outside?” the granular prefrontal cortex does not activate. The agranular prefrontal cortex will activate in both cases.
So we start to see that it’s in cases of thinking about yourself and thinking about others that this granular region gets very activated. And now, if you go back and study these people more deeply, you notice that they become hugely impaired at false-belief tests. They can’t recognize faux pas, so they don’t understand what’s not really appropriate. Which, of course, makes sense, because how do I know what’s appropriate? I’m going to infer how you feel about the things that I’m saying. And so you see all of these mentalizing impairments that emerge, but it’s not related directly to these logic puzzles that are typically in things like IQ tests.
Tim Scarfe
Yeah. You mentioned the false-belief test. Can you just briefly sketch out what that is?
Max Bennett
Yes. There’s a good picture if you want to hold it up or show it on the podcast. The way the test works is you have Sally on the left, who has a basket, and then you have Anne on the right, who has a box. So Sally puts a marble in the basket, and then she walks away. Then Anne goes over and moves the marble from Sally’s basket and puts it into her box, and then leaves. When Sally comes back, where does she look for the marble?
It’s so simple. But in order to figure out that Sally will look into the box, you have to understand that it’s possible for another mind to have incorrect knowledge—to have a false belief about something. Young children don’t understand this. They assume that knowledge is omnipotent: everyone has the same knowledge about the world. But at a certain point, they start learning that it’s possible for people to have false beliefs.
So we actually know that nonhuman primates can do this. They’ve done studies on macaques where you do exactly the same Sally test, and you just look at where their eyes look when the person comes back into the room to look between the 2 boxes. They always look, or tend to look, in the direction of where that person thinks the marble is, or the piece of food is, not where it actually is. If you inhibit their granular prefrontal cortex through an injection or another mechanism, this bias goes away. They no longer look in the right direction. There’s lots of really good evidence that this sort of false-belief mechanism is occurring in these primate regions.
Tim Scarfe
Yeah. And what really hit home to me is that, in a way, it’s not even knowledge. It’s all simulations. It’s just simulations of other agents. We’ve always spoken about knowledge in some weird Platonic, abstract sense. I quite like the idea that the primitive form of communication between humans is just simulations, even when we’re speaking to each other.
Yeah, exactly. So how do we solve the Sally problem? It probably happens so quickly, but we just simulate what we would think if we were in Sally’s shoes. And then I realize, well, I would look in this place. And this helps us reason about other people. And this begs a really, almost profound question: How unique is theory of mind?
This brings me to a question that I’ve been asked multiple times, which is: Does ChatGPT have theory of mind? The evidence I should stipulate, for anyone curious, is that if you ask GPT-3 these sorts of theory-of-mind puzzles, it does terribly. So that’s an easy one to discard out of hand. But if you ask GPT-4 these theory-of-mind puzzles, it performs remarkably accurately, at a human level, on these theory-of-mind puzzles. And there have been people who have explored whether it’s just in the training data, and there’s good evidence that it’s not just because they’re regurgitating what was in the training data.
So does that mean that ChatGPT has theory of mind? I think there are a few ways to reason about this. One is: What do we mean by theory of mind? If by theory of mind we just mean the ability to solve these sorts of false-belief puzzles, then I think you have to accept the fact that, yes, it can solve those tasks. The problem is, the way in which it renders this model of other minds is not through having a similar mind itself. And so what this means is we should be concerned—it doesn’t mean it won’t work well—but we should be concerned about how well this will generalize to real tasks where we might care about this much more deeply.
For example, with a human, there’s good evidence to suggest that part of my ability to reason about your mind is because I have a mind that works quite similarly. We are almost bound together by some common mechanistic synergy between the way in which our brains work, because our brains are quite similar, which enables a lot of data efficiency. I’m pretty good at predicting what people do—not perfect, but pretty good at predicting what people do—because we’re all people, and there are similarities between how we act. That makes us quite data-efficient and decent at generalizing to new situations where we put people in new places that we’ve never seen before. I can kind of guess, well, if I were in that situation, this is what I would do.
GPT-4 has learned to build a theory of mind simply by reading the text of these puzzles, and so clearly it has some mechanism to build a model for predicting what people will do in certain circumstances and differentiating knowledge and intent, et cetera. But the concern is twofold. One, what will happen if we take those types of models and put them in very new situations that are not based on just these puzzles, but, for example, we’re asking them to optimize a paperclip factory? That’s a situation where we should be concerned. How well will it do at actually inferring what we mean by what we say?
And the second is data efficiency, which is: How much data did it have to see to build this model? If it was a ton of data, then it’s going to be problematic if we have these new situations where we want to teach it to model people’s behaviors in this new place. If it requires a ridiculous amount of data, then it’s always going to be slow to learn these things and always be at risk of not generalizing well when we put it in these new situations.
My answer here is nuanced. I think if by theory of mind we mean solving puzzle questions, it’s very hard to say that ChatGPT does not have some model of human behavior. But I do think the human and primate mechanism for doing so has a data-efficiency advantage and a mechanistic-synergy advantage. In other words, we can use ourselves to reason about things, and that is relevant. If we want to have these systems do a good job listening to human requests, we shouldn’t translate performance on false-belief tests into believing that they’ll do a good job correctly inferring our intent and knowledge in new situations.
Tim Scarfe
Yeah, I would agree with that. I think ChatGPT is in the world of text, and it's learned all of this structured narratology and things on Reddit and things on Twitter. And as we were saying last week, language has evolved to be very simple. It has to be learnable by children. It has a small subspace. But it is a real kind of generalization over human behaviors, and it's in this very low-resolution substrate. Whereas in the Machiavellian apes example that we were talking about before, these are agents performing real-time sensing and inferencing and making in-the-moment judgments, and they're in this continuous sensor domain where they have many different types of signals: visual signals, sound signals, and also memory of what happened in those dynamics just before. So it feels like a difference in kind to me between those two situations. But it is remarkable that in the GPT domain any kind of theory of mind could work.
Max Bennett
One good example of this, I think, is whether there is a difference in our human ability to predict behavior between a car and a person. So the brain is always able to model things it observes, simulate them, and predict what they will do. I can look at a car and imagine different colors of it, and I can imagine what will happen if I drop it and it rolls down a hill. We build models of things all the time. We build models of computers and models of books. So the brain produces models of things. Is the way that the brain produces models of other human behaviors exactly the same, or is there some unique advantage? My argument is that there's something unique happening when I'm building a model of another person, which is that I'm leveraging my own inner simulation of things as a useful prior to try to predict what other people will do. ChatGPT models human behaviors, to draw a crude analogy, the way we would model a random object: I'm only modeling it based on seeing its behaviors in certain situations with the data I receive. On the other hand, when we model someone else's behavior, we're doing some form of projection and using the prior of how we would behave, and we probably bootstrap part of our model of human knowledge and intent based on our own introspection. I think in that way it is a difference in kind.
Tim Scarfe
Fascinating. I completely agree with that. The Selfish Gene is kind of saying it doesn’t matter what you folks do. The gene is directing your behavior, and you don’t really have as much agency as you think you do. And it’s a similar thing with language. If you think of language as being a superorganism or a virus, and we are the hosts, information is being shared memetically, and it’s shaping our evolution, but it’s also shaping our behavior. So it’s almost like when we become infected by certain memes—it might be religion, for example—it’s almost like it parasitically affects our behavior.
Tim Scarfe
But I think there is a difference between social memes and physical memes. Tool use, for example, doesn't seem to have the same parasitic effect. If you look at the behavioral complexity of apes, because they don't have these novel virtual memes in their culture, their behavior seems quite monolithic compared to ours. But I wondered if you could contrast that next level of mimesis.
Max Bennett
There's been lots of great writing about the distinction in the literature, which is typically called cultural transmission, between nonhuman primates and humans. A lot of the general consensus here is that, although there is transmission among nonhuman primates, which we see in particular with tool use, it doesn't accumulate in the same way that it does with humans. In other words, humans can pass a piece of information to another generation, which that next generation will reliably copy and then merge with other new information, which they can then reliably copy. You do this over 1,000 years and you go from, “I know how to whittle a bone into a needle for sewing,” to, all of a sudden, “Now I've built a loom,” right? These ideas keep accumulating on top of each other.
Whereas in nonhuman primates, you don't see the same type of accumulation. That's what I, in a pithy way, in the book call “the singularity that already happened”: once you enable these memes or ideas to accumulate across generations, you get what you're describing as this sort of mimetic organism that we are the substrate for. For sure, what I think is interesting here is one lens through which I like to think about this: sources of learning.
If you think about how nonhuman primates learn, there are sort of 3 sources. One is that they learn from direct experience, their own actual actions. This is reinforcement learning writ large: I do something, it succeeds, it fails. Fine. Another is their own imagined actions. This is the part that evolved in early mammals. I can imagine doing 5 different strategies to try and get to the food over there, and I find the one that worked. That's a source of learning: my model of the world became a source of figuring out the right path.
What mentalizing enabled with primates is this third mechanism: learning from other people's actual actions. So I can see my mother—if I'm a young chimpanzee—using a stick to put into this termite mound, pull it out, and eat food. I don't have to do my own behaviors to do that. I don't even have to simulate doing it. By watching her do that skill, I will adopt and learn.
But what nonhuman primates don't have, which is very uniquely human, is learning from other people's imagined actions. This is the key breakthrough that happens with language: the bandwidth through which nonhuman primates can communicate what we're calling knowledge here is only through actions themselves. I can't describe, if I were a nonhuman primate, what I saw when I imagined 5 different ways to try and hunt the boar over there. I can just do it, and you can learn from what you saw me do. But language enables us to share the outcomes of our imaginations.
That is a much higher-bandwidth mechanism for translating information. That enables accumulation across generations. For example, it's so easy to think about ways in which this would be adaptive. Two would be sharing semantics: I go into the forest, and there are 2 snakes there. One bites me and I'm fine; the other bites me and I get really sick. I come back and say, “Green snakes are okay. Red snakes, don't go near them.” That semantic knowledge now exists among the whole troop.
In the old world, before there was language, only the people surrounding the event who saw it happen would have the knowledge. Now I can translate it: I simulated the episodic memory in my mind, and I translate it to everyone. The other is coordinated planning. Before language, it would not be possible for 5 humans to jump in trees and say, “Okay, here's how we're going to hunt these boar. We're going to stay silent, and then I'm going to whistle 3 times, and then we're all going to jump down and surround the one in the back.”
That type of planning is only possible because one person can simulate something and then translate it and say, “Hey, when I imagine this happening, we succeed,” and other people, of course, can edit that simulation and say, “When I imagine that happening, I don't see us succeeding for this reason.” You can start refining it. This ability to have a source of learning from other people's mental simulations is what I would argue is the source of this very unique human superpower that emerges from language.
Of course, now, with such a high-bandwidth transference of mental simulations, you do get this sort of quasi-evolutionary process, which is what Richard Dawkins is talking about. You have a process by which the memes—the ideas that do a good job propagating—are the ones that will propagate. The ones that, for whatever reasons, are not viral either don't do a good job of maintaining the host, so the ideas are bad and I end up dying, or I just don't have an incentive to share them. They're not viral; those ideas die, and so then you get this sort of meme evolution. But to me, the source is the fact that language enables us to share in our simulations, which becomes a much higher-bandwidth communication mechanism.
Tim Scarfe
I'm fascinated by this idea of the meme itself being an agent, being a virtual agent, and, in expressing its agency, it needs to manipulate us. You might argue, as you do in your book, that there has to be some kind of traceable chain down to the basal ganglia. So we have many levels of bootstrapping, and at some point the thing exists because the basal ganglia says, “Oh, that's good. I like that.”
So then we have one level, then another level, then we have the memeosphere. It's almost as if that thing is manipulating us down here, but doing something completely different up there. When you have weakly emergent macroscopic phenomena, part of the definition of emergence is surprise. It's macroscopically surprising. It does something completely unexpected and unlike the thing that went below it. It's just weird that it might be manipulating us down here, but doing something completely different up there.
Max Bennett
Yeah. I think there's probably—this is mostly fun speculation—but if I'm going to draw analogies to brain regions and intellectual features of the human brain, there are probably 2 lightweight ways we could think about why memes become attractive.
One would be the older vertebrate-like structures. This would be the basal ganglia plus the amygdala. A meme that makes me feel fear, or makes me think that unless I take an action something bad is going to happen to me, and one of those actions has to be sharing it, is going to be highly viral. If you make me afraid for my family's well-being, you're going to activate my amygdala, and even if there's only a 2% chance that this is true, I might still share it. So you get these sorts of effects. Humans are not good at dealing with low-probability, high-magnitude events, which is another brain constraint.
The other key thing that also exists at the level of early vertebrate-like structures is a preference for surprise. In order for reinforcement learning to work well, it's very effective—and we see this in AI systems—to make people pursue actions that are novel, because that's one way in which we can explore new areas and learn new actions, and explore the space of possible choices to make. This is one intrinsic way to get trial and error to work.
The way casinos make money from you is that they hack into this sort of preference for surprise. If there's a 0% chance of winning, you would never play. But if there's a 48% chance—a net 48% chance, meaning in the long run you'll lose money—but every once in a while there's some surprising thing that is actually over the threshold of being worth it to the basal ganglia because the surprise is so exciting, that's one way to think about it. So if you get something that creates innate fear, or some great outcome or surprise, you get these older structures.
With mammalian structures, I think there is sort of an active-inference play here, and you could even correlate it to the granular prefrontal cortex, with things like identity. If you give me some information that's consistent with my model of myself, and I'm highly motivated to maintain my model of myself for a variety of reasons that we can talk about, then I'm more likely to maintain this belief. Versus if you give me information that's inconsistent with my model of who I am, people are highly likely to reject these beliefs.
This is another speculative way to think about why memes persist within their little echo chambers and how it can quickly become sort of identity wars. If one's identity is consistent with a certain set of beliefs, then that almost creates a gated wall for certain types of memes to enter. It becomes much harder for certain ideas to enter my mind, and it creates a very porous filter for other types of ideas that are consistent with my identity. Those become very easy for me to adopt.
I think that's another way—if we're going to frame memes as having agency—that a meme would seek to survive: find a way to be consistent with certain people's view of themselves and the world. What you're doing is reinforcing it as opposed to challenging it.
There is a very clear difference, though: the human brain is analog, while these machine brains are digital, and there are pros and cons to each. Geoffrey Hinton talks about this—to make sure I’m citing these cool ideas correctly—and he has a great talk in which he describes it very well. The benefit of a digital brain is that it’s immortal: all the weights are stored in binary, so I can very easily transfer it to different brains, but it’s hugely energy-inefficient because I need to model everything exactly in order for it to be copyable.
The human brain is much more efficient, but it’s not copyable because the information exists in the physical representation of the analog connections between all these neurons: the actual protein receptors that exist in them, the gene expressions, and all this crazy stuff that makes it non-copyable. It’ll be interesting to see what the energy efficiency is, for example, of a digital AI system that actually attempts to recapitulate a human brain. That might be very energy-inefficient.
It might open the door for a whole new area of research that I think would be really fascinating: building analog brains. Can we have systems that actually work in a more analog way? The way they pass information to each other—this is also a Geoffrey Hinton idea—is by teaching each other. Because they are AI systems, they can teach each other with better fidelity than humans can, because they can actually share probabilistic outcomes as opposed to just the words we say.
They can also generate way more samples for each other than a human could because they can live much longer. There’s a whole emerging world around the distinction between these digital machines that are immortal but very energy-inefficient, and analog machines that are much more energy-efficient but less good at translating information.
One of your points that I think is really key is that one of the main things missing—and I don’t think it’s talked about enough—is the continual-learning problem. Maybe there’ll be a breakthrough soon that would be great, but I don’t see very clear ideas over the horizon that will solve this. I would say this is one of the essential lines that differentiates biological brains from modern AI systems.
The way in which AI systems are trained is such that we cannot let them continuously learn from new experiences because it disrupts the old information they have. Whether that is an architectural constraint or something that needs to change in the underlying learning algorithm itself, there’s lots of open research and debate about that. But the fact remains that if you allowed ChatGPT to learn from every chat that happens to it, it would get rapidly dumber.
Tim Scarfe
Yeah.
Max Bennett
That is not the case with humans. We can continuously update our information, and our representations are robust. I think for many of the applications for AI systems that are going to be most impactful, continual learning is going to be an essential component because we’re going to want to bring an AI agent in, show it new information, and immediately have it incorporate that without forgetting old things.
I think that is very clearly a line where there’s a lot of really interesting research happening and a lot of research left to be done.
Tim Scarfe
Why are we superior to animals?
Max Bennett
There’s been such a long history of us pontificating on the various chasms, or attempting to create a chasm intellectually, between us and other animals. The most famous form of this, which I think still shows threads in modernity, is from Aristotle. He took the same kind of ideas that you see in MacLean’s triune brain: other animals might have basic instincts, and they might have some form of emotions, but what they all lack—which humans uniquely have—is this notion of reason. We can uniquely reason about things in the world.
As I try to argue in the book, and as most comparative psychology demonstrates quite clearly, there are clearly forms of reasoning that we see in other animals. Of all the different abilities and capacities that seem unique to humans, the one that stands out as most salient is undeniably language. Despite many painstaking attempts, we have not even been able to teach chimpanzees, bonobos, or gorillas to speak with the same degree of fidelity as human language.
There is some controversy as to the extent to which Kanzi, Koko, and Wu passed the threshold that we define as language, so that can be debated: where do we draw the line? Undeniably, most people would agree that these nonhuman primates do not learn language naturally without painstaking attempts to teach them. When they do learn language, it does not show the same sort of flexibility as human language.
What makes language unique is 2 things. One is declarative labeling. There is a distinction between imperative labels and declarative labels. An imperative label is learning that a phrase, or a cue, leads to a reward if you take an action in response to that cue. When a dog responds to a specific cue and then you give it a treat, that’s not what we define as language; it’s an imperative label.
A declarative label is when I say “dog” and, in your head, you know that it references a concept or a thing. We have a label for a concept or a thing. It’s not at all clear that other species perform these types of declarative labeling. If they do, it surely evolved independently. We’re quite confident that early primates didn’t have this ability.
The second thing that makes language unique is grammar. We can take these declarative labels that reference things or actions, and then we can weave them together in a certain structure, and the structure itself has meaning. A basic example is just the ordering of phrases: if I say, “Ben hugged James,” that means something different from “James hugged Ben.” Despite the fact that they’re the same phrases, or the same declarative labels, the order presents meaning.
There’s a whole interesting world around why language—if it is the case that language is the fundamental difference—has allowed humans to take over the world. That’s another interesting topic we should discuss. I would argue that primarily what makes humans different is language, and Aristotle’s idea of reason we see at least in smaller forms in other animals.
Tim Scarfe
Yeah. It’s quite interesting because you said right at the very beginning that Aristotle spoke about the rational soul that we have. Even in the 20th century, we spoke about things like mental time travel, our sense of self, and tool use, and it’s really interesting because we look in the animal kingdom and, one by one, all of these things that we thought placed a bright line between us and animals faded away.
Some people think that language is a continuum, that there’s just a gradation—that if you scale up the brain of an ape, you will get human language. Is that the case?
Max Bennett
The reason I’m very skeptical of that claim is that we don’t see variance in language abilities based on brain size. Children who learn language at the age of 4 still have relatively small brains. I’m not sure of the exact brain-size comparison, but I’d be curious about the brain size of a 4-year-old child relative to an adult chimpanzee, just based on volume.
The other interesting case is Homo floresiensis.
Tim Scarfe
Oh yeah, from Indonesia—the ones with the small brain.
Max Bennett
Yeah, yeah, yeah, yeah. Yes, Homo floresiensis is a great case study here because we found fossils of ancestral humans on an island in Indonesia who were effectively miniature humans. They had shrunk in size to, I think, around 3½ to 4 feet tall.
Their brain capacity, which we can look at from their fossilized brains, had actually shrunk from that of ancestral humans. They were marginally larger than the size of a modern chimpanzee brain, and yet they showed a lot of signs of superior human intelligence despite having smaller brains. They showed tool use akin to that of ancestral humans.
They had Oldowan tools, which are supposedly a sign of uniquely human, intelligent tool-making. That is suggestive of the idea that whatever unique intellectual capacities humans gained around 2 million to 1 million years ago were present despite their shrinking brains. Either one has to argue that language evolved much earlier, which some people do, or that whatever sort of protolanguage emerged back then was present even when these brains started shrinking.
That suggests to me—and is actually aligned with the ideas in The Language Game, which is a great book—that fundamentally what’s unique is that we have an instinct to learn language. It’s not that we have some unique capacity for language, and I think that is a key difference that we can talk about.
When you look at children who learn language, there are 2 very unique features of how they go through language learning. By the young age of around 2, they’re already engaging in proto-conversations. A younger infant will pause to match the pausing of their mother. Even if they’re just babbling, they will engage in the synchrony of babbling time intervals.
That is clearly a demonstration of some initial instinct, which demonstrates the ability for me to want to engage in some turn-taking action with you.
The other unique thing that emerges a little bit later is joint attention, where human children will uniquely attempt to get their parents to engage in attention toward the same object. Scientists have gone to painstaking efforts to demonstrate that this attempt to get a parent to engage in attention toward an object is not an attempt to get the object. So a child or an infant will be dissatisfied if the parent doesn't look at the object but they get the object—so a third party comes in and hands it to them. They're dissatisfied.
If the parent looks at them and is excited when they're pointing at an object, the child will also not be satisfied. But only when the parent looks at the object, then looks back at the child and smiles, is the child satisfied. So there's this instinct to engage in conversation and to jointly attend to things, which gives us the sort of instinctual foundation on which you can start adding declarative labels. When you have joint attention to something and you're paying attention to this turn-taking, it enables you to label things and say, “Well, this means run, or this means book.”
So I think all of that makes it hard to argue that it's just a consequence of a scaled-up brain. One of the things that's really fascinating is that animal communication seems extremely superficial. And when I say superficial, I mean that when you take different populations of the same species or different species, the expression and the complexity are very, very simple. We don't see this incredible fractionation and divergence that we see in human language.
Tim Scarfe
As you articulated just a minute ago, a big part of that is this declarative labeling. One of the reasons, presumably, for language is the ability to do variable binding on symbols—to say, “This thing is a dog, that thing's a bear”—and to be able to dynamically manipulate that.
It seems to me that you can think of language as a form of agentic communication. The difference between humans as language users is that we are agents, and agency is about being able to have your own directedness, plan many steps ahead, take control of your environment, and so on. The difference in communication with animals is that the information content is more in the environment around them, whereas for human languaging, a lot of it comes from the agent itself. So I just wondered whether you could think of any weird way to distinguish human languaging from animal communication.
Max Bennett
One line that I think there's some good evidence to suggest exists between human communication and nonhuman primate communication is that humans have much more of a desire to share what's going on in our own minds. There is a unique pleasure we have from sharing our thoughts. And when we look at the communication styles that happen in nonhuman primates, there's much less of a desire—even when we go through these language-learning experiments where they have forms of communication—there's much less of a desire to share thoughts that are going on in one's mind.
One line where there's some controversy around this is that humans, from a very young age, will ask questions. They'll inquire as to what's going on in someone else's head. And with the exception of maybe Kanzi, where there was some argument that he maybe asked questions, you did not see nonhuman primates probe the minds of other individuals, even though we know they have theory of mind. We know when they're trying to deceive others or they're trying to learn actions by observation, they clearly engage in theory of mind, but when it comes to language, they weren't interested in inquiring as to what someone is thinking about.
I think in that sense there's an agency to language being a tool for inquiring as to what's going on in someone else's mind and sharing what's going on in your mind. And this is where I think language is part of why language is a superpower, because it provides a completely unique source of data for learning. Nonhuman primates can engage in learning through observation because I can see someone take actions, as you said with imitation learning. I can see you open a puzzle box to get food and I can learn from observing your actual actions.
The way I do that is because I can infer the intent of what you're trying to do, and then I can figure out which of the actions are relevant and which are irrelevant. A monkey, a chimpanzee, or an ape will ignore irrelevant actions when they observe you do a task. They've done these experiments with humans and chimpanzees where they do all these actions to open a puzzle box, including some random actions, and chimpanzees will ignore the random actions, which suggests they can infer the intent of it. Which is great, but chimpanzees don't learn from what's going on in your head.
The ability to learn from other mental simulations is what's so powerful about language. I can say, “I just went over to that forest over there and I saw a red snake and a blue snake, and I saw that the red snake is really dangerous, but the blue snake is not because the blue snake bit me and nothing happened.” And so I share that episodic memory, and now everyone has that knowledge, even though that was just in my mental simulation.
Or, when planning a hunt, a group of 5 humans—I can imagine a strategy of how all 5 of us are going to coordinate, see it succeed in my mind, and then share the results and the plan with everyone. So language enables us to tether our mental simulations to each other. And I think there's a sense of agency in the idea that there's a purpose to that. There's a volitional purpose to the communication.
The neurological underpinnings of communication that occurs in nonhuman primates are more analogous to our emotional expressions than they are to language. And we see this also in the brain. Monkeys and nonhuman apes have these innate expressions that they do, which are genetically hard-coded, and we know that because it's the same even across species, often, that have never interacted with each other. It comes from neurological structures similar to our laughing and crying.
So this is clearly a hard-coded emotional expression. In that sense, it doesn't have the same volition, because I'm not doing this action to communicate a concept to you. I'm doing this action as an innate response to a cue or a feeling I have. So, yeah, I think there's some meat to that idea.
Tim Scarfe
Well, a few things to explore there. First of all, we should just talk about how we became a collective intelligence after the fact. I'm not sure whether that's unique to humans if you look at other forms of collective intelligence. There's always a kind of juxtaposition between the intelligence of the individual versus the collective, and usually you find that having very intelligent individuals is not good for the intelligence of the collective.
But what's interesting about humans is that we clearly didn't evolve as a collective intelligence. We had this kind of bootstrapping process where we were very, very useful, independent agents, and then this collective intelligence just emerged out of nowhere. Is that an interesting observation?
Max Bennett
There are different degrees of collectiveness, and so I think we can draw distinctions between different flavors of collectiveness, but I don't think humans are uniquely collective. For example, the imitation learning of nonhuman primates is a form of collective intelligence because you can teach one member of a chimpanzee troop how to use a tool, and then over time the rest of the troop will learn just through observation. So that's a sense of collective intelligence.
Many vertebrates, and likely the first vertebrates—you can even see fish—will learn through observation. In other words, when a fish swims in a certain direction to get food, other fish can see that fish do that and follow them. There can be an instinct to follow others around you. So I think there are flavors of collectiveness that exist across many different species.
But what's unique about the collectiveness in humans is that the fidelity with which we transfer our mental simulations enables them to accumulate across generations. In that sense, it almost has its own agency, or is its own thing, because it can actually go through its own process of evolution as ideas propagate through generations of people. That's not the same thing that you see in other animals.
Tim Scarfe
Yeah. A couple of things on that. First of all, I would quite like to distinguish knowledge and intelligence. Collective intelligence—and intelligence in general—is a process of discovering models. To get the language down here, I will use models, skills, and knowledge pretty much interchangeably. I think of an intelligent process as epistemic foraging: finding interesting models and then having them discovered and shared by other people.
It's a little bit like when you distribute a GPU workload: you can do model parallelism and you can do data parallelism. You can either split up the processing, or you can split up the actual representation. So I think the kind of collective intelligence that you've just been speaking about is that we've got all of these independent agents, and they are finding models and sharing models; the models get refined over time, and it adapts, and so on.
But I also think a big, important element is sharing the computation. So even though there's some redundant work going on, epistemic areas over here are being explored, but also, in many cases, the same problems are being explored in slightly different variations. So we're sharing the workload with other humans.
Max Bennett
Yes, I think that totally makes sense. There are some interesting ideas in AI here, actually, where there's this concept of knowledge distillation. In AI, one way in which you can have model A teach model B the things that model A knows is to wholesale copy the parameters of model A. Of course, that's totally biologically implausible. There are aspects of parameter copying—the components of our brain that are genetically hard-coded are a version of parameter copying—but for other applications, it's not feasible or maybe not desirable to just copy parameters.
Knowledge distillation is saying, okay, well, we can have a set of data that we give model A, and we either look at the outputs of model A or the layer before the outputs, so we can see more richness in its representation of the input you give it. Then we take that data—that almost-labeled data—to model B and train model B on it. So that's distilling some of the knowledge, through almost training model B to try and act similarly to model A.
That type of information transfer, I think, does occur in nonhuman primates, and that's imitation learning. However, it is not nearly as rich, because what happens in nonhuman primates is primarily grounded in just the actions that I'm taking. That's much less rich than what I can share: not only the data of what you see me actually do, but also things that happen only in my mind. That opens the door for much more transference of—and the word you used—the computations that I'm performing.
Tim Scarfe
So I think it's absolutely true. Before, we were learning in the physical world, so we were learning from physical things that we were directly observing, and now we are learning from imagined actions. But there's a bit of a latent component to language as well. For example, someone might come up to me and say, “Oh, the blue swirly thing is over there.” And I'll say, “Well, I don't know what you mean by the blue swirly thing, because I've never seen one before.”
There's this inference process. This is where it starts to get really interesting, because there's a diffusion, right? There's a message passing that happens between all of the different agents, and it's filling in missing information. Even though many of the agents wouldn't have seen anything like what we're talking about, sometimes it can be filled in with subsequent interactions with people, and sometimes it can just become a kind of latent category that can be filled in later. So there's this real diffusion process going on, which I think is quite difficult to articulate.
Max Bennett
Part of what's so interesting about language is that it's still an area of such controversy amongst cognitive psychologists, linguists, and even AI people. So much is still unsettled about it, and there are still debates today. There are debates today about whether language is primarily a tool for thinking or communication. Chomsky is the most famous proponent of the idea of language for thinking.
He has evolutionary arguments that language initially evolved not as a tool for communication, but for our own process of thinking, and then later was adopted or used for communication. That's a minority view. Other people argue—and I'm more amenable to this—that language was primarily used as a tool for communicating.
These ideas are actually reemerging with language models, because the way language models learn about the world, in some sense, is that language becomes the reasoning tool itself, which is more Chomsky-like. Even though I think the success of language models, in a lot of ways, discredits many of Chomsky's ideas, we can talk about that. Interestingly, the fact that we're using language as the fundamental mechanism for reasoning and thinking is actually somewhat Chomsky-like, versus the idea that language is communication.
The idea is that language is a condensed set of tokens that I'm passing between minds, but the real communication I'm trying to share with you is what's going on in my mind. In other words, it's the mental simulation—the more mammalian component here. The rendered 3D world is what I'm trying to transfer to you, and I condense it into this code that you reverse-engineer back into a mental simulation.
Theory of mind—one reason why language might be so rare in the animal kingdom is that mentalizing, or theory of mind, which is relatively rare in the animal kingdom, is a prerequisite. In order for me to reverse-engineer the language code you've provided me, I need to be able to infer what you might have meant by what you're saying, reason about why you would have said this, and understand what knowledge you have.
So I think language is intended to cue another person to render something in their mind. This is also where teaching is so important and such a key aspect of language learning, because we can infer what declarative labels this person is aware of. When they're confused, you have to start trying to iterate to understand what they're confused about in what you're saying so that you can disambiguate it for them.
There's also a disambiguation process where you ask follow-up questions when you feel like you don't fully understand what's going on in someone else's head.
Tim Scarfe
Yeah. I mean, the guardrails thing is interesting because they're not necessarily thinking guardrails; they're also pragmatic guardrails. And there's a really interesting figure in the book, actually. Yeah, here it is. It talks about how language is sharing information over generations.
Without language, we learn a little bit inside a generation, then it goes pretty much back to zero again. But now we have the ability to pass on these memetic bits of information over several generations. The thing is, there's a real structure to it. I think of it as a bit like a directed acyclic graph. So it's a tree structure, and every single bit of knowledge that we discover kind of stands on the shoulders of giants. It needs all of the things that we discovered beforehand.
In a sense, we're all of these little agents, and we're doing this epistemic foraging. We're finding new skill programs, we're sharing them, and so on. But it's almost like we shouldn't think of the mass as being like an entire convex hull. It's only on the boundary where all of the creativity and all of the information sharing happens—on the surface of this object that's being created.
What I mean by that is, now in modern cities, for example, you can't live without a driver's license. You can't live without the internet. You need to do things a certain way, and even though it's not technically constraining our brains and how we think, we live in a very, very constrained and weird world now.
Max Bennett
Yeah, totally. Great point. There's a biological constraint as to how much knowledge a given human brain can contain. One lens through which to see the last 100,000 years, especially the last 100 years, is us finding solutions to getting past the biological constraint of human brains.
Language was one tool, because it used to be the case that all the information that a given entity learned needed to be learned by my brain within my lifetime. Language enables us, as a group, to have shared knowledge, but not every brain contains all of the knowledge. If you think about a troop of 100 people, it's possible for those 100 people and all their descendants for 1,000 years to have tons of skills, despite the fact that no one brain ever had all of the skills.
Someone becomes really good at hunting, someone becomes really good at weaving animal skins into clothing, and all of these types of skills. There are actually cases in anthropology of groups of humans that get separated from each other and their technology degrades, because there is a limit—a minimum number of brains needed to contain and store a certain amount of information in the absence of writing.
Language was maybe innovation 1 here. Writing was another innovation, which is great. Now we can more reliably transfer these ideas across generations, even if there are gaps. In other words, even if there's a period of time for maybe 2 generations when no brains contain it, a third generation can go back to the writing and pick up that knowledge.
Of course, now with the internet, we've just scaled up writing even more. But you're absolutely right. Sometimes I think about this as, if a group of 20 friends and I ended up on an island—if we were the only 20 humans left—not that I think about this all the time, but it is crazy how little of human knowledge would be contained in our 20 brains.
How dramatically we would degrade, essentially. We've got this thing where we've got all of these different brains, and individual people can have about 150 friends or something. It's the social Dunbar limit.
Tim Scarfe
But as you say, because we have this ability to share simulations and we have common myths and so on, we can address a much larger carrying capacity of people and knowledge. You said something really interesting in the book, which is four things: bigger brains, specialization, more brains, bigger population size, and writing and sharing simulations through the internet and all of these things.
So we've increased our carrying capacity, and now something very interesting and arbitrary has emerged. We've got all of these different specializations of skills, and I guess the question is: where does it end? Has it converged? Could we carry much more knowledge than we already have, or would we have to wait for a top-down kind of genetic pressure for our brains to get a bit bigger?
Max Bennett
I think we are about to go through this. Google and the internet have turned us all into epistemic hybrids. Google has become a shared knowledge store that we all use, and of course there are problems, because now there are subareas on the internet where we can use different knowledge stores. We live in these different epistemic bubbles, and that creates political problems as well.
But we have already become hybrids where we use technology to overcome limitations in our own brains. Writing is a tool to overcome challenges in memory and, at times, thinking. The internet has become a tool to answer any question at a whim, and some people have concerns with this because it can also atrophy parts of our brain that maybe we want.
For example, through mere introspection, I will say that once I started using Google Maps as a kid, the part of my brain that was learning how to navigate a city—by actually remembering the grid and map of a city—just started atrophying. Now I have no capacity to do that, whereas my dad—you take him to any new city and you can see him rendering a map of the city in his mind—won't use Google Maps.
One could argue that it doesn't matter because I'll always have Google Maps, so why do I need this skill? Another argument would be that atrophying may have other consequences in my life, and it would be important for me to go through the cognitive exercise even though technology enables me to do it. We make these trade-offs at different times. Why do we teach kids arithmetic? They can always just use a calculator, but we deem it important for us to go through the process of understanding arithmetic even though technology can already do a better job for us.
This is a new frontier with using large language models, and there are some really cool things happening with education. In Khan Academy, for example, they're working on building language models to help children go through reasoning steps, which is a really cool application. Instead of just asking the question, it will probe the student to go through a process so they can come to the conclusion themselves.
There's a pessimistic and optimistic world here. An optimistic world is that these new AI systems are actually going to be a new step forward in cyborgizing ourselves, but it's not necessarily going to be as atrophying as something like Google. These systems won't only give us the dopamine hit of a factual answer; they'll also guide us toward better understanding how they came to their conclusion, to ensure that we understand when we're probing and asking questions. That's an optimistic state of the world.
The pessimistic state of the world would be that we offload more and more of our own cognitive reasoning to these systems, and we become even more atrophied in these abilities. That might not be a good world to live in if we keep offloading more and more reasoning to systems and lose the ability to do it well ourselves.
Tim Scarfe
Yeah. I've been thinking about this a lot recently. I was involved in a startup that did transcription, language models, and augmented-reality glasses. The idea was that you could be in a lecture—and I still think this is very useful for people with accessibility concerns, like those who are hard of hearing—but we were thinking of it as something that could augment your cognition.
You're in a lecture, and now you don't need to pay attention to the lecturer because you're transcribing it and GPT is making notes for you. I think this is really wrong. But you give a counterexample with satnav: we don't need to read maps anymore because we can externalize that cognition.
I feel like this is different. You're in a lecture, and all of these AI language tools are a form of understanding procrastination, right? Understanding, or intelligence, is the process of creating a model. You're creating a simulation, and in order to create a simulation, you actually have to think. You normally think, externalize the thinking a bit, do some writing, and pay attention.
Here's the thing: in that situation, there are so many more cues because it's in 4D. You can hear things, you can see things, and it's a social and physical activity. Even the dance—the performativity of the lecturer—is all information. It helps you understand.
Now I'm transcribing the thing, and people say, “It's okay. I can just read the transcription later and understand it.” Well, maybe, but you're already at a disadvantage. You probably won't, because this procrastination is just paying it down the line. You're saying, “I might do it later. I might do it later,” and you never will. That's going to create a society of automatons that just don't think for themselves.
Max Bennett
Yeah. I'm torn between the optimistic and pessimistic states of the future, but I think there's a very good argument behind what you're saying, so I definitely don't reject it out of hand.
I actually really liked the analogy to model-based versus model-free that you were suggesting there, because that applies very well to Google Maps. When my dad navigates a city, he has a model of the world, and he's engaging in model-based planning of how to get somewhere. When I use Google Maps, I've externalized the model, and all I do is respond to the cue of when to turn right or left.
I think that is absolutely a good way to think about this: we use technology to externalize building models, which can sometimes make things more efficient because then we can just be model-free actors. But there are places, like in the example you're suggesting, where we really want people to engage in the more painful, hard process of building models of things. In those cases, obviously it's dangerous to make it so easy to externalize these models.
Tim Scarfe
Yeah. It's hard to articulate. I think part of it is a kind of acquiescence. You're sequestering your agency when you externalize too much of your cognition, particularly if it's parts of your cognition that are useful because they contain core knowledge that will generalize and help you acquire new knowledge, or if it's the portability of discovering knowledge. It's your intelligence, and you're not exercising that muscle.
You become acquiescent and then you become less of an agent. From a collective-intelligence point of view, we're just saying that language and intelligence are about discovering knowledge. If we are all sequestering our agency and becoming less intelligent as individuals, as a collective, maybe we will suffer.
But it's one of those things that's so easy for us now to make grand statements about. People in 200 years will look back on this and laugh and say, “It's a little bit like when they introduced bicycles.” There was apparently a moral panic because they said, “Women will start cheating on their husbands and using bicycles to go to the next town.”
Max Bennett
That is an interesting fact. Yeah. Well, I think history is such a good tool when trying to reason about how people in the future will think about us. We are the people in the future to the past, which is obvious, but it's a useful tool.
For example, in some sense we already live in this dystopian world when it comes to physical exercise. Roll back the clock 500 years, and most people didn't have to think about physical exercise as much because most work required physical exercise. We exercised with our work, and so much of the work—at least in the developed world—is information-related, where we don't exercise and so we go to the gym.
The gym is a weird thing. If aliens came down and observed gyms, it would be anthropologically very bizarre behavior, because we just go into a room and run on treadmills. We do it because we've evolved to require exercise, and modernity has removed exercise as a prerequisite to most of the things that we need in life.
Now there's this gaping hole, and what we do is go to the gym and run in place to satiate this physical need. You could imagine—and one might interpret this as dystopian or utopian—a world where we've offloaded so much cognition, but because humans need to think about things, or because as a society we value it the same way we value physical fitness, there are now social pressures to go to these intellectual gyms.
Just to make sure, even though you don't need to do it for work or it's not necessary for the world to function, we feel like there's value in a human who knows how to reason about things. So we go to intellectual gyms for that. We might—I don't know if that's a utopian or dystopian future—but however we feel about it, I would venture to guess people 500 years ago who looked at a treadmill would probably feel similarly.
Tim Scarfe
100%. Well, MLST is my intellectual gym, by the way. [Laughter] You spoke about DNA. Dawkins, of course, wrote the book The Selfish Gene, and you said that the value of DNA was not what it creates—it creates hearts and lungs and so on—but what it enables, which is this evolutionary process. But then it gets to this concept of what we mean by a meme in general.
You said that it's an idea or behavior that spreads contagiously. How do you think about memes?
Max Bennett
Well, I think Dawkins did a wonderful job articulating this idea in a way that's really understandable. A meme is a concept or a behavior. A meme can be just the idea that individuals should have rights, or the idea of equality, or something sillier, such as the idea that we shake hands before we sit down for a meeting. And these things, because humans can share simulations through language and we engage in imitation learning, these ideas or behaviors propagate throughout societies.
And because these things are propagating, a different form—not evolution in the sense of genetic evolution, but a form of evolution—emerges, because some ideas will propagate better than others. By nature of that process unfolding, memes—these concepts or behaviors—actually go through an evolutionary process. Ideas that are either viral because people want to share them with each other, or ideas that somehow support the survival of the individuals that hold them, are going to be ideas that propagate correctly.
Ideas that negatively affect the survival of the individuals that hold them, or that people do not desire to share for whatever reason, are going to do a worse job propagating. And so it's a really almost brilliant lens to look at human culture when you reframe cultural ideas and concepts as memes—a different take on genes—that go through their own sort of process of iteration. This is not my idea; this is Richard Dawkins's.
Tim Scarfe
Oh, yeah. Well, we can thank Richard very much for this. I'm fascinated with memes, and I kind of think of language as being a collection of memes. But now we're in this very, very interesting space. Before language, we learned by observing physical skills performed by other people, and we could imitate them and so on. Now we are sharing simulations, basically, without actually needing to see the thing, and that means that we are one step removed from reality.
So all sorts of memes have cropped up, and some of them are better described, as you say in your book, as shared delusions, but they have some utility as well. When we have a common myth, for example, it might be a religion, it might be a nation-state; it allows us to cooperate with each other in a way that we wouldn't be able to do before. And you actually cited some ideas by John Searle and Yuval Noah Harari in his book Sapiens on that. [Snorts]
Max Bennett
Yeah, they've all famously popularized this idea. But John Searle was one of the original ideators of it. What's so powerful about these shared fictions is they can propagate much more easily than a human can talk to everyone in a group. And so, because they propagate much more easily and with very high fidelity, this enables me to meet someone who is a New Yorker whom I have never met before and immediately have shared views.
We probably both believe in individual rights. We probably both believe that money can be used for transacting things. So, if I give them a dollar, they'll believe that the dollar will be used elsewhere, or they can give me a dollar. And of course, today it's hard to reason about these things because there are so many rules in place that you don't realize it's all a shared fiction.
The reason we think we believe in money is because we're like, “Well, I know that all the other stores I go to will take this money.” So that's the reason it works. But why do they all take the money? It's all this shared belief that we all trust that this thing will be used for transacting. And so, because of that, it enables really large groups of people to coordinate, and that is a very powerful aspect of language.
But the argument I make in the book is that, similar to how genes are powerful not because of the structures they create but because they enable a process of evolution by which good structures will emerge, language is similar in that sense. What's powerful about language per se is not that we can engage in these shared simulations for coordination; it's that language enables the propagation of ideas and concepts across generations, which will thereby undergo its own evolutionary process. So, of course, these good ideas that enable survival are going to emerge. And that's really what's so powerful about language.
Tim Scarfe
I guess the arbitrariness is quite interesting. Some of them, on the surface, don't seem like good ideas; they just seem like really bad ideas. And I guess you can think about it in terms of creativity as well. For a meme to be established in the sphere of possible memes, does it need to have intrinsic value? Possibly not, because we're getting into creativity: is it novelty? Does it have intrinsic value? Is it just social proof? Is the meme only existing because lots of people have been fooled into thinking it has value? So it's kind of extrinsic value via social proof.
And then there's almost a double entendre with the meme, or a deeper meme meaning, because you talk about altruism. The meme itself might actually be quite a stupid meme, but if it causes altruism, so there's actually a group-selection advantage to it, then it's almost like that's the lens of analysis to understand how good the meme is.
Max Bennett
Yeah. It's a really fun area of literature to read through because there's still zero consensus as to how language evolved. One reason why it's so controversial is the way in which we disambiguate—and I'll get to your question—the way we disambiguate evolutionary arguments is typically by observing gradation in extant, or currently present, animals. That enables us to observe these intermediary steps between a morphological aspect of body A and a morphological aspect of body B.
The problem with language is we have nonhuman primates that, for the most part, don't have any language, and then we have humans that have very complex language. And all of the intermediary humans that existed between our divergence with chimpanzees about 6 million years ago and our divergence with all other modern humans between 50,000 and 100,000 years ago—we don't have them; they're all dead. All those lineages are lost.
And so that means that there's this broad spectrum of arguments that could be made. Chomsky argues—I find this a very strong claim and thus hard to defend—that it happened all at once, or very rapidly: there was no language, then all of a sudden there was language.
Tim Scarfe
Right, and there are other arguments that it was a gradual process.
Max Bennett
But one of the most controversial aspects of language evolution goes to what you're talking about, which is evolutionary arguments for why language evolved. Evolutionary arguments for why language evolved have a harder burden of proof than arguments for other adaptations.
So when we argue about the evolutionary benefit of something like theory of mind, there are no complex evolutionary machinations one needs to conceive of to defend it, because you can see why it would be beneficial for an individual chimpanzee to be born with the ability to infer what's going on in other people's heads. They can better defend themselves when someone is going to be mean, better figure out whom to trust, better climb the social hierarchy, et cetera.
But with language, unless you take the Chomsky view that its primary adaptation is for thinking, the argument that language evolved for communication is more challenging, because it's not valuable for an individual human to be born with a little bit of language skill unless other humans are also engaging in language. And so this then means that the only benefit is if we're both sharing truly useful information with each other.
Although it seems intuitive that the way this would function is that a group of humans that are sharing knowledge with each other is going to survive better than another group of humans that's not, and that's how evolution will ensue, this is actually quite controversial in evolutionary biology because that's invoking something called group selection.
Now, some people call the modern incarnation of this multilevel selection, where there's some consensus that, yes, there are group-level effects that can impact things. But most people think that group-level effects are not nearly as strong as we would intuitively think.
And the issue is the following. If you have a group of 100 humans that use language with each other and then you have 1 human born who is just going to try to trick all the others, so all they're going to do is use language just to be disingenuous, it's not at all clear that that human would be at a disadvantage. In fact, they might be at an advantage relative to everyone else.
If you play that forward over time, language will be lost because someone born who isn't going to be tricked by the individual trying to lie to them with language is actually going to survive better than the people who have language skills. There has been so much debate throughout evolutionary linguistics about these arguments as to how language evolved. I like the argument in a great book called The Evolution of Language by Fitch, and I think he makes a really great argument around how you could think about this occurring.
A lot of people argue that it probably started with something called reciprocal altruism. The way altruism exists in the animal kingdom, there are 2 forms of accepted altruism. One is something called kin selection, which is quite straightforward: I'm willing to sacrifice something—in other words, share something with an individual—if I share genes with them. That's easy.
Reciprocal altruism, which we do see in the animal kingdom, is, “I'll scratch your back if you scratch my back.” But if you start not scratching my back, then I'm going to stop scratching your back. What this suggests is that, in order for language to be stable—in other words, for it to be beneficial for me to truthfully share information—there need to be costs to me lying.
This is one argument that people speculate is one reason why humans have such strong moral preferences toward punishing liars and out-groups and in-groups, because we really try to identify individuals who are lying. Robin Dunbar has a beautiful argument that this is why gossip evolved. One way that evolution can stabilize the use of language is by virtue of us having a preference to share moral violations.
Gossip is a tool of language where, if you see someone lie or cheat and you share it with a bunch of other individuals, that becomes a huge cost to someone lying and cheating, because if 1 person catches them, then the whole group is aware of it. There is this special feedback loop that happens where language skills require more punishment of violations to be a stable strategy. One way you get that is by having more gossip and making sure there are higher costs to defecting.
This is not by any means the only story of language evolution, but it's one with a lot of interesting evidence behind it. There are some people who argue that the feedback loop—1 emerging idea, which I don't talk about in the book but which I do think is interesting—is actually one in which we try to detect lying in others. They make the counterargument to me, saying that the effect of lying is the loss of language.
There is an argument that you get the reverse: you get a really good theory of mind in humans because we're so sensitive to trying to detect people who are actually giving us false information. There's still a lot of controversy around it, but the main takeaway is that the blanket group-level selection argument—that language is obviously beneficial because once a group has language, they're all going to survive better—is not a sufficient argument for language evolution.
You need a more nuanced evolutionary argument as to why it's a stable strategy for an individual to be born with superior language skills, or you have to argue that language did not evolve primarily for communication.
Tim Scarfe
You know, it's quite interesting, first of all, that you were writing this book actually a couple of years ago. So this was before GPT-4, although you did put a note in about GPT-4, and you were speaking about Blake Lemoine. He was a Google engineer, and he famously came out convinced that these things had developed sentience.
I think much of this actually hinges on the concept of a world model. One view of language models is that they're just modeling a statistical distribution of tokens, and that seems quite low-resolution. Another take is that they're learning a world model. What that means is that, rather than just capturing the state of language, they're actually simulators. They're generating the underlying processes of language. They're capturing the dynamics of language.
How do you say that something is or is not sentient, especially given that the models could potentially be such high resolution that they are generating the same thing for all intents and purposes?
Max Bennett
This is where I don't see myself as a philosopher, but this is where I do think scientists need to include philosophers. When questions become nonscientific, I think the scientific instinct is to argue that we don't draw distinctions between things that the scientific method can't draw a distinction between. But the problem is that there might be moral differences between them.
For example, it might be scientifically impossible for us to differentiate which of 2 systems that look indistinguishable in their inputs and outputs is sentient. Scientifically, we might say, “Well, because we can't differentiate the 2, we're going to say they're the same.” But that doesn't mean they're the same. That just means that, because we have no methodology for drawing a distinction between them, from a scientific perspective we're not going to draw a distinction because we're entering philosophical territory.
But if you take that and then start talking about policy implications, the actual values we attribute to them, and how we introduce these things to society, I think we need to include a philosophy lens here. It might not actually be the case that they're the same just because we can't distinguish them. So that's just 1 thought.
Tim Scarfe
Another thought on world models. One distinction I want to draw, because I've seen a lot of confusion on the internet about the world-model dilemma, is that there's a difference between a world model and a model. It is undeniable that language models have a model. In order for GPT-4 to correctly predict the next token in these really complicated language questions, it clearly has some model of something.
Because we can ask it common-sense questions about the world and it answers many of them correctly, you can say this is a model of aspects of our world without question. I think it would be very hard to argue that that's not the case if you look at GPT-4's performance on many of these questions. But what most people mean when they say “world model” is a specific process of simulating an ordered sequence of states and the consequences of different actions.
Max Bennett
That means identifying the end result of these actions in your head. Another way to think about what we mean by a world model is the ability to reason about interventions and causality. This is the Judea Pearl argument: with our world model, we can hypothesis-test.
I can say, “I imagine that if I do this thing in the world, I think this will be the consequence of it because that's what I see in my head.” Now I have a hypothesis. Now I'm going to actually do that thing in the world and see if my hypothesis is correct.
That's very different from what's happening in a language model, where its understanding of the world derives solely from its input data. In a world model, my understanding of the world comes from the delta—the difference between what I hypothesize is going to happen in the world and my actual experience of it.
This distinction really matters the more we're going to start offloading our cognition to these systems. For example, everything that ChatGPT knows is on the basis of its input data. That means if false information or wrong information is in the input data, ChatGPT is going to know that information. There's absolutely no hypothesis-testing embedded into ChatGPT, unlike our true AGI agent that will one day be invented.
What it would do is hypothesize aspects of the world and test its own hypotheses. If you give it false information, if it reads articles about how the Earth is flat, it's not going to just start talking about how the Earth is flat. It's going to say, “Okay, well, that is incongruent with my model of the world. I'm going to now run some tests where I can differentiate between them, and I'm going to perform those tests and then conclude that the world is not flat.”
When one says that ChatGPT does not have a world model, I think some people misinterpret that as suggesting that it's just dumbly looking at the statistics. That's not at all what we're saying. In order to correctly look at the statistics of language, clearly it's built up a very rich and complex model of the text that it's seeing, and that's how it's able to predict the next word so well. But it's not what most people mean when we say “world model.”
Tim Scarfe
A couple of things on that. I think people conflate the machinations of language models with how we represent them statistically or abstractly, because if you look at a lot of papers, they actually represent it like a probability—a joint probability distribution. Of course, the way that language models work is completely different from that, but you're bringing in some very interesting things.
So first of all, we are agents in the world. The agential lens is quite interesting: we interact with the world, so we're not just learning from observational data. I was talking with Nick Chater, and we said, “Why is it that in our everyday experience, we experience the world in 4D color?” He said it's because it's interactive.
In your experience, you can actually seek new information, right? You can move your eyes, you can get new information in, and you can touch things. When you're doing future or past simulations, you don't have that interactivity. So there's something about interactivity that's really important.
But even then, how far could you go? A complete one-to-one simulacrum of the world wouldn't be a particularly good model. In physics, there is no causality, right? It's just dynamics. Causality is actually something that emerges very, very far up. So we're talking about a model that's an approximation of the real world, which may or may not include causality.
It probably would, because it's an interactive model and it has this kind of agential map. But I guess we're just drawing the line somewhere and saying, “Well, that is a world model.” Max Bennett
I'm actually not sure whether, even if we rendered a perfect 3D map of every particle in the universe—
And that was the input data to some infinitely large model. I still would argue that it's learning something different from a model that's given some form of agency, where it can hypothesize rules and then test its own rules.
Now, given infinite time, it's possible that those will converge, because given infinite time, every possible hypothesis I could conceive of will end up showing up in the training data. So eventually I'll see the training data of every possible experiment I could run. If time is infinite, I guess you could suppose that happens.
But what's so different is this dramatic dimensionality reduction that happens when you show me something uncertain and then I can conceive of specifically the tests I want to run to map the uncertain thing to my mental model of the world. That's a very different way of learning about things. It's not just input data and then self-supervising on predicting one's own input data. It's building a model in which I can simulate possible outcomes and then hypothesis-test those outcomes.
Even this is not uniquely human at all. If you look at the way a rat would deal with something novel in its environment, it's drawn to the novel thing and explores it until it feels like it understands it. Then it will move away.
When you show a child an object that's perplexing, they will touch it, turn it around, and try to understand it until they feel like they've built a model of it. That simple act is doing something very different from the self-supervision we see in most AI models today, because I see something I'm uncertain about and I'm volitionally going to create new training data for myself.
I know the training data I want now. I want to see what happens when I pick it up, turn it to the left, and turn it to the top. A convolutional neural network doesn't do that. So the way we teach CNNs to understand rotations in 3D objects is by manipulating the training data ourselves. We take imagery and rotate it in a bunch of different ways, so we're the ones curating the data set to teach it these things.
Max Bennett
Yeah, but that's different from the way we learn about things. I think this is a key aspect that's missing from AI systems today, and it's something that folks are working on. That's something we're going to have to add in.
Tim Scarfe
Yeah, completely agree. It feels like you're saying basically what I think, which is that there's a creativity and an agency gap. A lot of that is because we're agents, as you say: we create our own training data, we do this active inference and sense-making, and we build these models in real time. As a collective intelligence, it creates a kind of divergent search process for knowledge. It's this epistemic foraging that we spoke about.
GPT is a monolithic model. It does have models, but the models are only learned at training time. Inference actually happens at run time, and when you put a prompt into GPT, you're just retrieving one of the models that was already learned a long time ago. It's not creating a new model in the moment. So it creates this kind of sclerotic system rather than the divergent, creative system that we experience in biomimetic intelligence.
One sort of mental model I have of this—because there's so much debate around it—is something I'd be curious to put through the gauntlet of what other people think about. Maybe people in the comments will either agree or disagree with this. I think an interesting alternative experiment, or eval, of an AI model that I haven't heard before is this: if you give it knowingly false information in the training data—not at inference time, but in the training data—will it reject it wholesale?
To me, this is the distinction: an agent that can hypothesis-test and intervene in the world will reject false information. If you tell it that the world is flat, it will know that the world is not flat. Whereas with GPT, any data you give it is given equal weight to every other piece of data. The only reason it would reject that is if there's other data in the training set that it's going to ignore.
It's almost cheating, because by definition we know a language model is going to fail at this task. The only way you can fix it is if you give it other data in the training set. There's no notion of hypothesis-testing compared with an agent. The only way you could get it to be wrong is if you manipulate its sensors during the actual hypothesis-testing that it does.
You can, of course, manipulate it by changing the actual test when it does these tests. That happens in The Three-Body Problem, which is an amazing book where aliens manipulate our experiments. Anyway, I think that's another way to evaluate these systems: can it figure out that you're giving it false information and reject it?
Max Bennett
Yeah, 100%. A couple of things, though. There's something magic about having agentic density in the system, right? When you have something like GPT, just to make it statistically tractable, it's generally doing a kind of low-entropy search. What I mean by that is it's just looking for the baseline patterns. It's not doing a lot of exploration, and it's not searching outside of the main sources of statistical regularity.
Whereas when you have divergence in the search process, with all of these individual agents doing their own things, as a system it's much more of a high-entropy search. That means you're actually bringing in lots of new information to solve problems in creative and interesting ways.
In the physical world, though, it's quite interesting, right? The problems come from the physical world. The trees get big, so giraffes need to have a long neck in order to eat the leaves from the trees. This whole thing just rinses and repeats. The environment produces novel solutions, and then we see this divergence and find novel, creative solutions to the problems that get generated.
But in the memetic sphere, it's so much more difficult than that, right? The problems and the guardrails aren't constrained in the way that they are in the physical world. For example, we have capitalism or we have nation-states, and there’s all kinds of interesting divergence going in different directions. But it doesn't seem like there are the same pressures that ground the thing in reality.
Tim Scarfe
Well, yeah, I think it's definitely not grounded in truth. Its tethering to truth is this: knowingly false information that leads me to take actions that will hurt my survival will fade. But false information that helps me survive better will propagate freely, or is at least neutral.
Another way this shows up is—and this is where I'll go into some pontificating—
Max Bennett
Please.
Tim Scarfe
But where I think there is memetic evolution that can drive us away even from things like happiness. If we think about what systems of coordination survive, they're systems of domination and militarism.
If you take 2 groups of individuals, let's say one is really happy and calm, sees no desire for domination, and does not attempt to innovate and build more technology. The other is unhappy but super aggressive, wants power, and wants to expand. These ideas will die out.
What this suggests—and I think I talked about this in one of our previous conversations—is the importance of delineating, in my view, the Darwinian component of what does survive from the moral component of what is right or wrong. It is definitely not the case that what survives is definitionally right. It's absolutely possible that the things that survive and do well evolutionarily are not the things that we feel are morally aligned.
That is not to propose a correct or incorrect system, but it is an important distinction to draw when we're trying to decide what we deem to be morally right or wrong.
Max Bennett
So I think that's just one example of what you're saying: the ideas that propagate successfully might not be the ones that are true. They might also not be the ones that we deem to be moral, or even be the ones that lead to human happiness. They're just the ones that do a good job of keeping humans alive and reproducing the idea.
Tim Scarfe
Yeah. So people say that language models confabulate and don't preserve epistemic factfulness. But you could also argue the same thing about us, right? We actually confabulate everything. We don't really have goals. We just generate these post hoc confabulations, then explain our behavior and pretend that was what we wanted to do, that we had beliefs, and so on. We just make it up as we go along using this kind of active inference.
Even though we are emotional and subjective, and we believe in religion and lots of things that we presumably made up, we have Wikipedia. We have objectivity, even though it's an illusion, right? There's no such thing; even general relativity isn't as objective as we think. If you keep asking why and why and why, it just disintegrates into incoherence. But there seems to be some objective structure that is preserved. How is that explained, given that our brain simulations don't seem to select for truthfulness?
Max Bennett
I think the question of whether humans are better or worse than ChatGPT is almost a red herring. I look at ChatGPT as an alien—it's like an alien brain. There are certain things it does that are clearly better than us. Information retrieval in ChatGPT blows a human away, without question. In many ways, it's way better than humans.
But there are certain things that human brains do that ChatGPT does not. If we're trying to build human-like intelligence, there's certain inspiration we can garner from human brains. I think there's a component of our model-based rendering of a plan and then executing that plan that has a level of explainability that's unique relative to a system that is just iteratively predicting the next token.
But we also do the same thing that ChatGPT does. When we make model-free choices and then you say, “Why did you do that thing?” what we engage in is exactly as you're describing: a post hoc explanation. I didn't render a plan; I was just walking down the street. If you say, “Why did you move your foot there as opposed to 2 inches to the right?” what I'm going to do is render a post hoc explanation of why I did that.
But I didn't really think about it; I'm just explaining it after the fact. So it's definitely the case that humans have that component. But it means there's also another component, which I would argue is unique and important: our ability to pause, render a plan, and then execute against that plan. The key thing that I think is the dividing line between these models and us is the ability to render hypotheses and make interventions in the world. That's the key thing.
And so it's not the case that our brain has the true objective state of the world in our head. I don't think that's there. There might be components of objective truth in ChatGPT that it contains that we don't have, and I think in its information retrieval it has probably, in some ways, more aspects of reality than I do, in terms of having read all of Wikipedia and answering questions about biology that I don't even know the answers to. But there are also components of the world that the human brain has rendered and contains that ChatGPT does not, because of our ability to make hypotheses, intervene, and learn the causal structure of the world. I think that is the dividing line. But I wouldn't say it's because the human brain knows the objective state of the world and ChatGPT does not.
Tim Scarfe
Max Bennett, it’s been an absolute honor to have you on MLST. Everyone at home, you need to buy his book immediately. It is a wonderful, wonderful book, Max. You did such an amazing job of bringing all these things together. We've now spoken for 4.5 hours, going through the last 3 chapters. My God, it's been an honor. Thank you so much.
Max Bennett
It's been my pleasure. Thank you for having me.