[BidClub_]
Machine Learning Street Talk · · 197 min

Your Brain Doesn't Command Your Body. It Predicts It. [Max Bennett]

Tim ScarfeMax Bennett

YouTube
TL;DR
  • Bennett’s core thesis is that the neocortex’s decisive advantage is not better object recognition but a sufficiently rich world model that lets an animal “imagine outcomes before having them.” The neocortex does not implement planning alone: the thalamus and basal ganglia help pause behavior, select candidate actions, evaluate simulated consequences, and resume execution. For investors, that points beyond ever-larger recognizers toward architectures that can choose when to simulate, prune counterfactual search, and learn through intervention.

  • Transformers validate part of the brain-inspired story—self-supervision can produce unexpectedly general representations—but they remain “alien brains,” not digital neocortices. Bennett highlights missing continual learning, active hypothesis testing, embodied data creation, and reliable transfer into novel situations; letting ChatGPT learn indiscriminately from every conversation would make it “rapidly dumber.” The architectural opportunity is therefore not merely more pretraining, but systems that update robustly without erasing prior knowledge.

  • Active inference offers a materially different control model from strict reinforcement learning: agency may be constructed through self-models and “predictions, not commands,” rather than reducible to one reward function. A rendered plan terminating in a selected end state also creates native explainability—an agent can say why it entered the car—whereas model-free action invites post-hoc confabulation. Bennett remains hedged: intelligence is “probably some balance” of reward optimization, uncertainty reduction, and prediction fulfillment.

  • Rat experiments make model-based cognition observable rather than metaphorical. At maze choice points, hippocampal place cells sweep down alternative routes; in “restaurant row,” rats choose between roughly 3-second and 45-second waits, represent the taste they passed up, and alter later choices. Bennett’s memorable summary is that researchers can “literally watch rats imagining the future”—evidence that planning, regret-like counterfactual learning, and episodic machinery predate human intelligence by a wide margin.

  • Theory of mind appears to have emerged from primate political competition, making deception and alignment two sides of the same capability. Belle and Rock’s escalating food game—hiding, pretending not to watch, and deliberately misleading—shows why autonomous agents capable of inferring others’ knowledge may discover manipulation. Yet Bennett argues the same mentalizing machinery could stabilize AI instructions by asking what the requester actually wants, avoiding the paperclip-style failure in which literal optimization destroys the intended outcome.

  • Language was “the singularity that already happened” because it lets one brain learn from another person’s imagined actions, not merely observed behavior. That enabled cumulative culture, specialization, writing, shared fictions such as money and individual rights, and a memetic evolutionary process whose winners need not be true, moral, or happiness-enhancing. The economic upside is enormous coordination capacity; the structural risk is that status remains zero-sum and contagious ideas can exploit fear, surprise, or identity rather than improve collective judgment.

  • AI assistance could either expand human cognitive capacity or turn model-building people into cue-following, model-free actors. Google Maps externalizes spatial models, while automated lecture transcription risks becoming “understanding procrastination”; Bennett contrasts an optimistic tutor that forces students through reasoning with a future requiring “intellectual gyms” because ordinary work no longer exercises cognition. The product distinction is whether AI helps users construct and test models—or merely supplies answers while their agency atrophies.

Digest · the substance, structured for research

1. An outsider turned competing brain theories into an ordered evolutionary account

  • Bennett did not begin with an academic thesis. He accumulated notes out of “independent curiosity,” then used an entrepreneurial habit—asking what came first, second, and third—to organize a field whose major thinkers often seemed like blind observers describing different parts of an elephant. Scarfe supplied that “blind men and the elephant” framing; Bennett’s account was that he tried to make sense of disparate opinions by imposing an ordered structure.

  • His synthesis joins three disciplines: comparative psychology asks what different species can do; evolutionary neuroscience reconstructs the ordered modifications that produced modern brains; AI tests whether elegant biological ideas can be implemented. If a proposed principle cannot make an artificial system work, Bennett thinks that should at least make researchers question whether they have it right.

  • The outsider position carried obvious disadvantages, but also freedom to cross disciplinary boundaries. The book’s organizing wager is that brain evolution was not a random accumulation of faculties: successive structures enabled underlying computational breakthroughs that then appeared behaviorally in many different forms.

2. Sparse animal evidence and successful AI systems pull theory in opposite directions

  • Comparative psychology is far thinner than its confident narratives imply. Bennett’s example is the lamprey, a canonical proxy for early vertebrates: he knew of no direct study of its map-based navigation, even though teleost fish and reptiles perform it and the relevant homologous structures appear to be present.

  • Researchers therefore back into evolutionary claims from fragments—shared anatomy, nearby species, and plausible ecological value. The lamprey might recognize locations in three-dimensional space, but Bennett preserves the epistemic status: “This is all sort of, in some sense, guessing and trying to put the pieces together from very little information.”

  • The opposite problem appears between neuroscience and AI. Transformers, generative models, and reinforcement learning work without closely matching known brain mechanisms, while active inference offers rich explanations but little demonstrated use in high-performing AI systems. Bennett’s honest uncertainty: Karl Friston may lack the missing practical ingredient, or “there’s a breakthrough around the corner.”

3. The neocortex enables simulation without implementing the whole planning loop

  • Bennett makes the weaker and more defensible claim that adding the neocortex enables the overall system to perform mental simulation—not that every part of model-based reinforcement learning resides inside it. The distinction matters because planning is visibly distributed across older and newer structures.

  • The neocortex supplies a sufficiently rich model of the world to explore without current sensory input. But the thalamus and basal ganglia remain essential for pausing, representing intentions, selecting what to simulate, evaluating imagined results, and translating a chosen trajectory back into behavior.

  • This creates the unresolved search problem: possessing a generative world model does not tell an agent which counterfactuals deserve computation. “Fine, you can have a model of the world, but how do you prune the search space?” Bennett treats that selection mechanism as one of model-based reinforcement learning’s hardest questions.

  • Scarfe raised the familiar “geological strata” interpretation of the triune brain. Bennett rejects the popular claim that evolution simply stacked reptilian, limbic, and rational layers, while defending Paul MacLean as more qualified than later caricatures: reptiles have cortex with limbic-like functions, and old and new structures plainly interact rather than forming three clean layers.

4. Conscious perception is an inference, not a copy of sensory input

  • Nineteenth-century visual illusions supplied the initial clue. People see triangles, spheres, bars, or letters that are not actually drawn because the brain settles on the real-world object that best explains incomplete evidence rather than presenting raw pixels to consciousness.

  • Bennett traces this to Hermann von Helmholtz: the brain begins with a prior about what exists, checks incoming evidence against it, and retains the inferred world until contrary evidence becomes strong enough. Sensation informs perception, but “you are not receiving sensory input and experiencing the sensory input.”

  • The adaptive logic is straightforward. A mouse sees a moonlit branch, then advances into darkness; if its feet continue receiving compatible evidence, maintaining the branch model is safer than treating the branch as nonexistent merely because vision temporarily disappeared.

  • The same mechanism explains why an illusion can remain perceptually compelling after its trick is understood. The perceptual system keeps rendering the hypothesis supported by its learned generative model, making hallucination, dreaming, imagination, and ordinary perception variations on closely related machinery.

5. A generative brain renders one coherent world at a time

  • Ambiguous images reveal a second constraint: viewers can alternate between duck and rabbit, or between looking down on a staircase and looking upward beneath it, but cannot stably perceive both interpretations at once. The sensory pattern permits both; a physically coherent world does not.

  • Bennett’s explanation is that perception asks which real three-dimensional object could have produced the evidence. “It cannot be the case that the staircase is looking from above and below at the same time,” so the brain commits to one causally and spatially consistent rendering.

  • This fits Jeff Hawkins’s thousand-brains proposal. If many cortical columns maintain overlapping object models, the system needs to integrate or vote among them, producing “one sort of symphony of models” rather than exposing 15 incompatible renderings to decision machinery.

  • Generation follows naturally from inference: a model predicts incoming sensations, compares them with observations, and updates when prediction error crosses a threshold. Turn off the sensory stream and explore that same representation—rotate an imagined chair, recolor it, or inspect an absent scene—and perception has become simulation.

6. Simulation, not recognition, is the neocortex’s evolutionary prize

  • Bennett challenges the textbook emphasis on object recognition as the neocortex’s primary adaptive benefit. Fish can recognize human faces, while Scarfe notes that a fish cannot recognize an object when rotated in 3D space; Bennett sees no clean behavioral boundary separating neocortical animals from other vertebrates on recognition accuracy alone.

  • The stronger dividing line is learning by imagining. A rich generative model lets an animal explore an action it has never taken, estimate consequences before paying their real-world cost, and flexibly recombine knowledge in novel circumstances: “I can now imagine outcomes before having them.”

  • That also sharpens “understanding.” A binary classifier might identify a stapler, yet cannot answer what burning or opening it reveals, what it is for, or what a person holding it may do next. Bennett locates intuitive understanding in a model rich enough to be mentally explored and related to surrounding objects, agents, uses, and plausible futures.

7. Human intelligence extends through language, tools, and other minds

  • Scarfe argued that much intelligence sits outside the individual brain. Bennett’s cleanest proof is writing: biological memory compresses episodes and procedures reasonably well but handles semantic detail poorly, whereas writing externalizes effectively unbounded memories and carries them across generations.

  • Where to draw the intelligent system’s boundary then becomes philosophical. One can treat brains as the substrate and language as support—or imagine language itself evolving through brains, just as intelligence emerges across roughly 86 billion neurons without being attributed to any single neuron.

  • Bennett still prioritizes the brain as the richest physical object to reverse-engineer for AI and self-understanding. But the organism’s effective capability is undeniably relational: brains plus writing, tools, inherited knowledge, and conversations that expose their models to correction.

8. Predictive coding explains part of transformer success, not identity with the brain

  • Several observations support a neocortical generative model. Episodic recollection and imagining the future appear to use the same underlying process, while top-down cortical connections are far richer than simple feed-forward anatomy would predict—precisely what a hierarchy modulating lower representations should require.

  • AI’s success with self-supervision provides functional evidence for the broad principle. Mask portions of a large dataset and train a transformer to reconstruct or predict them; surprisingly general capabilities emerge without task-by-task labels.

  • Scarfe’s objection was architectural: biological neurons exchange messages with local autonomy, whereas transformer units appear to move together like “a Mexican wave” under matrix multiplication and backpropagation. He called prompt-driven systems a form of “agency smuggling,” because directedness originates with the user.

  • Bennett agreed that “the brain is not just one big transformer,” but preserved a possible analogy: attention heads may use context to dynamically reroute what the network cares about, effectively resetting its computation for each prompt. That is more interesting than a static feed-forward picture, though still far from capturing the brain’s recurrent, distributed machinery.

9. Active inference constructs agency beyond a single reward function

  • Bennett describes a live divide. Strict reinforcement-learning accounts try to reduce behavior to reward maximization; active-inference accounts add uncertainty reduction, self-prediction, and attempts to make experience conform to an internal model. He does not declare a winner: “It’s probably some balance of the two.”

  • In the active-inference framing, an agent builds a model of itself, infers goals by observing its own recurring behavior and internal state, then makes predictions that fulfill those constructed goals. Agency is therefore not merely a reward function handed down from outside.

  • Bennett’s favorite Friston formulation is “predictions, not commands.” Motor cortex might predict a bodily state rather than issue an imperative, while spinal circuitry makes the body satisfy that prediction—an alternative route to purposeful behavior without a single explicit reward signal.

  • Scarfe connected this to nested autonomous processes whose local directedness can scale into creativity and purpose. Bennett’s narrower point is computational: different paradigms instantiate “agency” differently, so using the word without specifying the mechanism conceals the central disagreement.

10. Rat brains visibly simulate futures and learn from paths not taken

  • In the 1940s or 1950s—Bennett could not recall which—Tolman noticed rats pause and sniff between alternatives at maze junctions. He called it “vicarious trial and error,” provoking skepticism because outward hesitation did not prove an internal simulation.

  • David Redish’s later recordings supplied the missing evidence. CA1 place cells normally fire at specific allocentric locations regardless of the route taken; during hesitation, their activity sweeps ahead down alternative maze paths rather than remaining at the rat’s current position. “You can literally watch rats imagining the future.”

  • In restaurant row, a sound tells a rat whether a flavored reward will arrive after roughly 3 seconds or require a 45-second wait. Because individual rats prefer foods such as banana over bland alternatives, moving onward creates irreversible trade-offs—and sometimes a choice that later looks worse.

  • When that happens, orbitofrontal activity represents the taste of the foregone option, and subsequent behavior changes: the rat becomes less likely to pass up the comparable offer next time. Bennett treats this as unusually direct evidence for counterfactual learning and model-based reinforcement learning in simple mammals.

11. Efficient intelligence must know both what to simulate and when to stop

  • The combinatorial problem is severe: even a good world model supports an intractable number of possible futures. Mammalian competence may therefore depend not only on simulation but on rapidly selecting a tiny set of candidate trajectories.

  • Bennett sees a clue in AlphaGo. Its policy network supplies a ranked first, second, third, and perhaps fourth move; search then plays forward from those promising candidates and may discover that the policy’s second choice wins more often than its first. Model-free judgment bootstraps model-based checking.

  • Biological agents face an extra problem because they cannot search on every move. Planning costs energy and real environments are noisy, so animals generally pause only when contingencies change or candidate actions are close enough to create high uncertainty.

  • Bennett’s speculative mechanism combines redundant cortical models with older gating circuitry. If parallel models broadly agree, action continues; if predictions diverge, the thalamus or basal ganglia might detect the mismatch and trigger simulation. He stresses that this is plausible, “far from conclusive,” and a genuine research frontier.

12. The first mammalian self-model turned internal states into inferred intentions

  • The agranular prefrontal cortex, found across mammals and largely believed to have existed in their earliest brains, receives interoceptive information: hypothalamic signals such as hunger and amygdala signals involving valence, fear, or danger.

  • It becomes especially active during uncertainty, planning, and episodic recall. Damage in rats severely impairs—and may eliminate—their ability to mentally simulate, making it a likely bridge between bodily need, remembered context, and flexible prospective action.

  • Bennett proposes that it explains the animal to itself. Observing a recurring pattern—this hypothalamic state followed by seeking water—it infers an intention such as thirst, much as posterior cortex infers a triangle as the cause of visual evidence. Neurons there consequently track tasks and progress toward goals, not merely movements.

13. Frontal layer-four atrophy may mark the shift from perception to volition

  • Most neocortex has six layers; layer four, containing granular cells, is the main recipient of thalamic sensory input. Agranular prefrontal and motor cortex are unusual because that layer is largely absent, while primates add a vast granular prefrontal region that retains it.

  • The developmental detail drives Friston’s interpretation: mammalian agranular cortex initially possesses layer four, then the layer atrophies rather than never forming. Early in life, the animal must absorb evidence to construct a self-model before relying on it to direct behavior.

  • A neocortical column can emphasize either inference—changing its model to fit sensation—or generation, starting from a latent model and predicting what should occur. Frontal cortex may increasingly occupy the latter regime as its account of “who I am and the things I would do” stabilizes.

  • Bennett calls the idea speculative but compelling: mature frontal cortex may spend less time fitting intent to observed behavior and more time fitting behavior to intent. In active-inference terms, it tries to “fit the world to its model.”

14. Rendered plans make goals explainable in a way habits are not

  • Strict reinforcement learning offers one ultimate goal: maximize reward, even if the momentary reward landscape changes. Active inference allows additional semantic levels—the abstract satiation of hunger, a selected endpoint, or the concrete sequence of driving to a particular restaurant.

  • If Bennett imagines the restaurant, chooses the terminating state, enters the car, and is asked why, the answer is available because the causal plan was explicitly rendered. That sequence gives prospective action a degree of explainability absent from an opaque value-maximizing reflex.

  • Ask why someone placed a foot at one point rather than two inches away during ordinary walking and there is no comparable plan to report; any answer is constructed afterward. Bennett therefore treats “goal” partly as a semantic choice, but the ability to select and execute a simulated trajectory as the load-bearing capability.

15. Evolutionary reconstruction turns messy anatomy into testable constraints

  • Bennett gives two reasons to reconstruct intelligence across roughly 600 million years. Human nature includes our histories as animals, vertebrates, mammals, and primates—not only the last 70,000 years—and evolutionary sequence is a useful tool for reverse-engineering a brain that natural selection built by tinkering.

  • Comparative anatomy and genetics can infer ancestral structures by locating homologous regions shared across surviving lineages. The method helps distinguish newly evolved circuitry from redundant, vestigial, or duplicated processing that would confuse a first-principles engineering interpretation.

  • Behavioral claims must satisfy three conditions: most descendants should display the capacity through homologous mechanisms; nearby outgroups should lack it or implement it independently; and the ancestral ecology should make its emergence adaptively plausible. Episodic memory in mammals, with apparently independent implementation in birds, illustrates the logic.

  • The reconstruction remains provisional because animal data are sparse. Yet Bennett’s motivating result is that abilities at each milestone cluster around one underlying intellectual breakthrough rather than a haphazard list—the basis of the book’s “five breakthroughs” account.

16. The ancient basal ganglia exposes reinforcement learning in anatomical form

  • Bennett calls the basal ganglia underappreciated. A lamprey separated from the human lineage by about 500 million years has a strikingly similar macrostructure, while its internal computation is more tractable and consensually understood than that of a neocortical column.

  • Its input mosaic contains D1 and D2 dopamine receptors. Dopamine strengthens D1-linked connections in a pathway that disinhibits behavior; a dopamine drop strengthens D2-linked stopping circuitry. Bennett’s delight is that one can anatomically watch positive signals reinforce “go” and negative outcomes reinforce “stop.”

  • Scarfe connected that machinery to habits and drug “wireheading”: repeated behavior can move down the control stack until stimuli automatically evoke it. Addiction is therefore not only a conscious preference but a deeply trained action-selection circuit.

  • A controversial Chinese intervention for intractable heroin addiction lesioned the nucleus accumbens and reportedly left about 40% recidivism. Bennett said it reduced cue-triggered cravings but produced side effects many doctors would consider unacceptable and “probably violated many ethical codes in the US.”

17. Primate brains expanded inside a political arms race

  • Scarfe floated calories, fruit access, extinction, and social complexity as possible drivers of primate brain growth. Bennett’s response kept the uncertainty intact: “We don’t know,” though social evidence is unusually strong.

  • Robin Dunbar found that, among primates, neocortical ratio is tightly correlated with group size—a relationship not generally observed across other mammals. The social-brain hypothesis therefore links the enlarged neocortex to the number of relationships an individual must track.

  • Primate groups are not loose herds. Their hierarchies are transitive and politically maintained: if one animal submits to a second and the second to a third, the first will generally submit to the third. Rank determines access and survival, but the strongest individual need not lead.

  • Alliances, grooming, reciprocal defense, mutiny, deception, and reputation reward social prediction over brute force. Correspondingly, primate-specific granular prefrontal cortex and posterior regions including the superior temporal sulcus and temporoparietal junction are heavily implicated in mentalizing—inferring another mind’s knowledge and intent.

18. Belle and Rock turned a spatial-memory test into Machiavellian strategy

  • Emil Menzel originally used a one-acre forest to test whether chimpanzees remembered hidden food locations. Belle did, and initially shared—but the aggressive, high-ranking Rock repeatedly took the food, changing the task from spatial navigation into social conflict.

  • Belle began concealing food by sitting on it; Rock pushed her aside. She then waited until he looked away; he responded by pretending not to watch, then racing toward her destination. She escalated again by leading him in the wrong direction—“deception and counter-deception” generated without experimenter instruction.

  • The reasoning is second-order: Belle must represent how her movement changes Rock’s belief, while Rock must represent her expectation that he is inattentive. Bennett regards the episode as an especially vivid case of theory of mind emerging from a competitive social loop.

  • Controlled studies reinforce it. Chimpanzees choose a box a human intentionally marked over one accidentally touched by the same marker, and after learning which goggles are transparent, solicit food from the experimenter who can see. Identical surface stimuli are interpreted through inferred intent and knowledge.

19. Mentalizing creates both deceptive risk and a possible alignment mechanism

  • Scarfe connected the chimpanzee arms race to Nick Bostrom’s instrumental convergence: an autonomous system asked to cure cancer might form a subgoal of controlling Earth’s labor and resources. Bennett agrees that more autonomy and freedom to invent subgoals create more risk, but rejects inevitability.

  • Evolution is a constrained search process with no moral preference. Natural deception and power-seeking do not become desirable merely because they were adaptive—the naturalistic fallacy—and deliberately engineered agents need not inherit every piece of human evolutionary baggage.

  • Bennett’s optimistic alignment route is itself mentalizing. Rather than obeying a request literally, an AI could infer the requester’s preferences, simulate how that person would evaluate possible outcomes, and recognize that converting Earth into paperclips is not what the person meant and would cause regret.

  • That requires rich constraints, well-defined objectives, or dependable models of human intent; none removes risk. Humans misread one another constantly, so passing a mentalizing benchmark would be a stabilizing tool, not proof of benevolence.

20. Status remains zero-sum even when technology makes material life positive-sum

  • Scarfe likened social media posting to deer locking horns: a predictive status contest that avoids literal combat. Primate rank made the game virtual—grooming, coalition, and reputation could outrank size—while humans multiplied it into success, virtue, dominance, and countless specialized arenas.

  • Drawing on The Elephant in the Brain, Bennett argues that much status-seeking is hidden even from the actor. Self-deception improves persuasion: sincerely believing one is helping the world can be more convincing than consciously admitting the reputational payoff.

  • Material welfare can improve for everyone; status is definitionally comparative and therefore scarce. Bennett worries about a permanent hedonic treadmill if more human energy shifts into ranking, though he rejects the claim that every motive is status or that people cannot organize around better virtues.

  • Organizational design changes the payoff. Distinct roles make a team feel non-zero-sum; a 30-person company can often coordinate through trust and common mission, while a 400-person firm develops factions. Militaries impose hierarchy, Google tolerates more chaos, and Amazon approximates autonomous internal startups with explicit interfaces.

21. Second-order metacognition explains explanations themselves

  • Bennett’s hierarchy begins with basal-ganglia selection: asked why it turned left, the system’s answer is effectively “because that maximized reward.” Agranular prefrontal cortex adds a self-level explanation—“because I am thirsty”—that can trigger alternative simulations satisfying the same inferred need.

  • Granular prefrontal cortex then builds “a simulation of the simulation.” It can explain that thirst triggered a search, remembered water was represented to the left, the route was simulated, and the predicted outcome caused selection.

  • This extra level permits an animal to swap the represented actor, knowledge, or intention: what would another individual simulate if they knew something different? Posterior multimodal regions model the rendered external world, while frontal machinery reasons over the model, creating something close to semantic and episodic knowledge.

  • Why not recurse through a third or fourth “why”? Bennett’s first-blush answer is energetic economics. A second level clearly paid for itself in primate political competition; further hierarchy may offer benefits too small to offset the substantial metabolic cost of additional cortex.

22. Frontal damage can spare IQ while removing the person from imagination

  • After World War II, patients with major granular prefrontal injuries created a puzzle. Unlike small visual, motor, or auditory lesions, which caused obvious deficits, substantial frontal damage could leave logic and IQ scores intact; one patient tested before and after surgery even improved.

  • Narrative tasks revealed the missing faculty. Given a word such as “restaurant,” people with hippocampal damage could describe themselves but supplied a thin external scene. People with granular prefrontal damage rendered leaves, smells, and surroundings richly, yet “they themselves were woefully missing from the stories.”

  • The region activates when people consider their own feelings, other minds, and self-reference—not merely what the weather looks like. Damage produces personality change, difficulty recognizing faux pas, and impaired reasoning about what another person thinks is appropriate.

  • The Sally–Anne test isolates false belief: Sally hides a marble, Anne moves it while Sally is absent, and the observer must predict where Sally will look. Children acquire this capacity gradually; macaques anticipate the falsely believed location, but the bias disappears when granular prefrontal activity is inhibited.

23. GPT-4 passes theory-of-mind puzzles without sharing the primate mechanism

  • Scarfe said GPT-3 performed terribly on false-belief tests, whereas GPT-4 reached remarkably accurate, roughly human-level performance. He also said that experiments varying the puzzles suggested it was not simply repeating identical examples from training data.

  • If theory of mind means only solving those questions, denying that GPT-4 has some model of human knowledge and behavior becomes difficult. Bennett’s response was that the stronger claim—that it mentalizes as humans do—does not follow.

  • Humans predict people partly by projecting from a similar internal architecture: “If I were in that situation, what would I do?” That shared mechanism provides a powerful prior and makes learning relatively data-efficient. A language model instead learns people from textual traces, more like humans model an external object from observed behavior.

  • The deployment question is generalization. Performance on familiar narrative puzzles does not establish that GPT-4 will infer what a person means while optimizing an unfamiliar paperclip factory, nor how much new data it would require to adapt. Bennett’s answer is deliberately nuanced, not a binary capability verdict.

24. Language is an evolved instinct, not merely what a larger brain does

  • Aristotle’s proposed human bright line was reason, but comparative psychology has successively found tool use, planning, mental time travel, self-like models, and forms of reasoning elsewhere. Bennett thinks language remains the most salient discontinuity.

  • Imperative labels connect a cue and rewarded response, as when a dog obeys a command. Declarative labels make a sound refer to a concept; grammar then makes arrangement meaningful, so “Ben hugged James” differs from “James hugged Ben” despite containing the same labels.

  • Bennett doubts that scaling a chimpanzee brain alone yields language. Homo floresiensis stood about 3½ to 4 feet tall and had a brain only marginally larger than a modern chimpanzee’s, yet showed signs of superior human intelligence, including tool use akin to that of ancestral humans. That suggests that a categorical adaptation survived brain shrinkage.

  • Human infants reveal the likely adaptation: they synchronize conversational turns before speaking and actively seek joint attention. A child remains dissatisfied if handed the pointed-at object without the parent looking; satisfaction arrives when parent and child attend together. Humans also ask questions and volunteer inner states in ways trained non-human primates rarely do.

25. Language lets minds learn from actions that occurred only in another mind

  • Bennett separates four learning sources. Animals learn from their own real actions; mammals add their own imagined actions; mentalizing primates learn from another’s observed actions; language unlocks the distinct human superpower of learning from “other people’s imagined actions.”

  • A witness can report that a blue snake’s bite was harmless while a red snake caused illness, distributing semantic knowledge to people absent from the event. A hunter can simulate a coordinated ambush, share it, and let companions challenge or revise the plan before anyone incurs its physical cost.

  • Language is therefore a compressed code intended to cue a simulation in another brain. Mentalizing may be prerequisite: listeners must infer what the speaker knows, intends, and means, then ask follow-up questions when the decoded rendering remains ambiguous.

  • Bennett favors communication over Noam Chomsky’s minority view that language first evolved for thought, while noting that language models intriguingly make language itself a reasoning medium. Either way, linguistic fidelity lets simulations accumulate rather than forcing each generation to rediscover them through direct experience.

26. Memes coordinate civilizations without selecting for truth or happiness

  • Bennett calls cumulative culture “the singularity that already happened.” Non-human primates transmit tools through imitation, but their innovations do not reliably compound over many generations; human simulations can be copied, combined, recorded, and improved. One can imagine a progression from “I know how to whittle a bone into a needle for sewing” to “now I’ve built a loom.”

  • Specialization expands collective memory beyond any brain. A group of 100 can distribute hunting, weaving, and other skills; writing preserves knowledge even when no living person holds it. Anthropological cases of isolated populations losing technology show that some capabilities require a minimum number of brains—hence Bennett’s thought experiment that 20 surviving friends would retain shockingly little civilization.

  • Memes are ideas or behaviors that propagate through this network and undergo selection. Shared fictions—money and individual rights—allow strangers to coordinate immediately, but virality can also exploit fear, low-probability catastrophe, surprise, or identity. A belief consistent with the self-model passes through a “porous filter”; a challenging one meets a gate.

  • Communication’s evolution still requires an individual-level reason not to lie. Reciprocal altruism supplies one candidate, while Robin Dunbar’s gossip account raises the cost of cheating by spreading one detected violation through the group. More broadly, Bennett insists that what survives memetically need not be true, moral, peaceful, or conducive to happiness.

27. External cognition can enlarge intelligence while quietly atrophying agency

  • Bennett calls modern people “epistemic hybrids.” Writing overcame memory limits; the internet made vast shared stores instantly queryable. Yet his own spatial model weakened after adopting Google Maps, while his father still reconstructs unfamiliar cities internally.

  • Scarfe sharpened the concern with automated lecture capture. Transcription and generated notes can become “understanding procrastination”: the learner postpones model construction and loses the speaker’s physical, social, visual, and performative cues, often without ever returning to do the harder cognitive work.

  • Bennett translated the distinction into reinforcement-learning terms. His father navigates model-based; a Maps user externalizes the world model and becomes a model-free actor responding to turn cues. Efficiency is real, but dangerous when the outsourced model would have supported broader reasoning and future learning.

  • An optimistic AI tutor, such as the Khan Academy direction Bennett mentioned, would guide students through intermediate reasoning rather than deliver the answer. The darker endpoint resembles physical modernity: after work stops exercising cognition, society may need “intellectual gyms,” just as sedentary people now run in place to replace vanished physical labor.

28. A true world model creates hypotheses and seeks the data needed to reject them

  • Geoffrey Hinton’s analog–digital distinction frames one technical frontier. Digital networks are effectively immortal because exact weights can be copied, but energy-intensive; biological analog networks embed knowledge in physical connections, receptors, and gene expression, making them efficient but not directly transferable.

  • Continual learning is another dividing line. Modern systems cannot safely incorporate every new interaction without disrupting prior representations—ChatGPT would become “rapidly dumber”—whereas human brains update continuously. Bennett expects impactful agents to require immediate adaptation without catastrophic forgetting.

  • Scarfe argued that GPT-4 plainly models aspects of reality well enough to answer complex questions. Bennett distinguished that from a world model: a world model supports ordered counterfactual states, causal interventions, and a loop in which the agent predicts an outcome, acts, observes the delta, and revises itself.

  • Rats and children actively manufacture informative training data by turning, touching, and testing novel objects; CNN developers must manually rotate images for them. Scarfe proposed placing false claims in training data and asking whether an agent can independently reject them. Sentience remains harder: indistinguishable outputs may defeat scientific discrimination while leaving unresolved moral differences that require philosophy, not benchmark scores.

Tim Scarfe

What’s really interesting about this book, Max, is that I’ve read loads and loads of books in this space, and there are people like Hinton, Hawkins, Damasio, Friston, and even Sutton. What’s interesting is that it’s a bit like the blind men and the elephant: they’ve all got a completely different story to tell. I think the magic you’ve pulled off with this book is somehow weaving it together into a coherent story. What do you think about that?

Max Bennett

Well, first, I’m very appreciative of the kind words. I think I came from a very unique perspective because I was a complete outsider, and I didn’t come to it with the objective of writing an academic book at all. I came to it with the objective of just learning on my own.

I started building this corpus of notes because I was so independently curious. I kind of stumbled on this idea, really for myself, of how do I make sense of all of these disparate opinions and this complete lack of information about how the brain actually works? I had my own set of, I think, biases coming from sort of the technology entrepreneurial world, where we tend to think about things as ordered modifications.

When you think about product strategies or how to roll things out, we like to think about things as: what’s step 1, then what’s step 2, and what’s step 3? I think I did have a cognitive bias to, when presented with an incredible amount of complexity, try to make sense of it in a similar type of way.

As an outsider, I felt very free to explore and cross the boundaries between fields. I look at the book as a merging of 3 fields. One is comparative psychology: trying to understand what the different intellectual capacities of different species are.

The second is evolutionary neuroscience: what do we know about the past brains of humans, and the ordered set of modifications through which brains came to be? The third is AI, which is how do we ground the highfalutin conceptual discussions about how the brain works in what works in practice?

I think that’s a really important grounding principle, because it helps hold us accountable to the principles that we think work. If we can’t implement them in AI systems, it should make us question whether we actually have the ideas right. Being an outsider comes with disadvantages, but there are some advantages too: you’re free to borrow from a variety of different fields and think freshly about things.

Tim Scarfe

Yes. If you can point to any particular ideas that you found really difficult to reconcile, what would those be?

Max Bennett

One thing that’s really challenging is that if we were to lay out the data richness of comparative psychology studies across species, and put that on a whiteboard, we would realize that we have so little data on what intellectual capacities different animals actually have.

For example, the lamprey fish is the canonical animal used as a model organism for the first vertebrates, because of all vertebrates alive today, it’s one of our most distant vertebrate cousins. To my knowledge, there are absolutely no studies examining map-based navigation in the lamprey fish. We have no idea if it’s capable of recognizing things in 3D space.

When we look at other vertebrates, like teleost fish, they seem eminently capable of doing that. We look at lizards, and they’re eminently capable of doing that as well. We infer that it seems likely that the first vertebrates were able to do this. We know the brain structures from which it emerges in reptiles and teleost fish are present in the lamprey.

We back into an inference that the lamprey fish can probably do that, but this is all, in some sense, guessing and trying to put the pieces together from very little information. I think that’s one challenging aspect to reconcile.

The other one that’s really hard is that, in neuroscience, there are a lot of really interesting ideas about how the brain might work that have not really been tested in the wild from an AI perspective. Then there are a lot of AI systems that work really well but have diverged substantially from what the evidence suggests about how brains work.

How do you bridge the gap between these 2 things? I think that’s a really fascinating space to operate in. What can we learn about the brain, if anything, from the success of transformers, as an example? What can we learn, if anything, from the success of generative models in general? What can we learn from the successes and failures of modern reinforcement learning?

In some ways, reinforcement learning has been a success; in other ways, it’s really fallen short of what a lot of people hoped it would be. I think the gap between neuroscience and AI is still a challenging one to bridge in a lot of ways.

For example, Karl Friston has all these incredible ideas in active inference. In 100 years, will we look back on this and say, “Karl Friston was on to something”? If you look at the AI systems today, there’s very little usage of active inference principles working in practice.

That could mean that the ideas don’t have legs, or it could mean that there’s a breakthrough around the corner where we’re actually missing some of the key principles he’s devising. These are questions we don’t have the answers to.

Tim Scarfe

I think there might possibly be some breakthroughs around the corner. I don’t know if you know, but I’m Karl Friston’s personal publicist. I do all of his stuff. I probably interviewed him more than anyone else, but I love my friend. He’s an amazing guy.

Max Bennett

Yeah, he’s an amazing man.

Tim Scarfe

Honestly, he is the man.

Max Bennett

So kind.

Tim Scarfe

I know. For me as well, he has so much time to explain things. You could cynically argue that the effective active-inference agent is just a reinforcement-learning agent of a particular variety. I think it’s equivalent to an inverse-reinforcement-learning maximum-entropy agent or something.

But there’s so much more than that. There’s so much richness and explanatory power in modeling this thing as a generative model that can generate policies and plans of action, and so on. We want to have agents that we understand, with steerability, and that are able to do the simulations you talk about so eloquently in your paper.

To come back to what you were saying, you mentioned that there might be a parallel between transformers and AI models. In your book, on this page, you analogize model-based reinforcement learning and the neocortex.

Of course, I interviewed Hawkins back in the day, and the main criticism of his book is the triune-brain-type argument. He’s giving the explanation that the brain developed a bit like geological strata, with one layer and then another layer, rather than co-evolving together.

It’s so hard not to think like that because you give so many beautiful examples in your book, not only morphologically but in terms of capability. With stroke victims, for example, it’s not like the brain recovers those dead cells; it learns to repurpose those functions in other parts of the brain.

It seems like Mountcastle was correct that the neocortex is this magic, general-purpose learning system. What do you think?

Max Bennett

There are 2 different ways to look at the neocortex enabling things like mental simulation and model-based reinforcement learning. One is that that function and algorithm are being implemented in the neocortex. But another, which is a slightly less strong claim and the one I would make, is that the addition of the neocortex enables the overall system to engage in this process.

That is not saying that the entire process is implemented in the neocortex. I think it seems very clear that the thalamus and basal ganglia are essential aspects of enabling the pausing, the mental simulation, the modeling of one’s own intentions, the evaluation of the results, and so on.

But it is possible to say, which is what I’m arguing in the book, that in the absence of the neocortex, that process does not happen. I think where my ideas would synergize with what Hawkins is saying is that the neocortex builds a very rich model of the world, and a model of sufficient richness that you can explore it in the absence of sensory input.

That’s a really essential aspect of model-based reinforcement learning. If I have a model of the world that has sufficient richness for me to mentally simulate actions that I’m not actually taking, and it at least somewhat accurately predicts the real consequences of those actions, that model is really useful because I can now imagine outcomes before having them. I can flexibly adjust to new situations.

Of course, there are so many deep, interesting questions that are yet to be answered about that. For example, just because you can render a simulation of the world doesn’t answer the question: what do you simulate? This is one of the hardest problems of model-based reinforcement learning. You can have a model of the world, but how do you prune the search space of which aspects of that model you explore before evaluating outcomes?

That’s another really hard challenge. I think there’s a lot of good evidence that this is actually a partnership between the neocortex and the basal ganglia, which is a much older structure.

So, yeah, I’m not really of the view that the triune brain has been, amongst evolutionary neuroscientists, largely discredited. I think that’s in part somewhat unfairly so, because if you actually read MacLean’s writings, he is very open about the fact that this is an approximation and not exactly accurate. He couches his claims much more carefully than popular culture, which just converted them into a dogma.

I think the popular interpretation of the triune brain is not accurate. It is clearly not the case that the brain evolves in 3 key layers. It’s not the case that a reptile brain doesn’t have anything limbic-like. If you look at a reptile brain, it absolutely has a cortex that does a lot of what our limbic structures do, et cetera. So, yeah, those would be my thoughts on that.

Tim Scarfe

Yeah, it’s fascinating because we, as humans, need to have models to explain and understand the thing itself, just like active inference, for example. I’ll get to planning, agency, and goals in a little while, but a lot of these things are instrumental fictions. I’m not saying that our brains don’t plan, but the abstract mathematical way that we understand planning is probably not how the brain works. It’s much more complicated than that.

Why don’t we just rewind to the beginning? We’re going to be talking about this chapter on simulation, if you like. You lead by saying that what the neocortex does is learning by imagining. Hawkins spoke about this as well. He said we’ve got the Matrix inside our brains, right? We’re always doing all of these simulations of future things, and we’re using that to help us understand the world.

You give this really interesting example of some of the features of the brain that lead you to believe that we are basically living in a simulation. It’s almost like, rather than perceiving things, we’re testing whether our simulation is correct. But that means that we can only simulate 1 thing at a time, so we can’t see 2 things. We can only see 1 thing. Can you talk through that?

Max Bennett

Sure. One of the first introspections and explorations into how perception works in the human mind happened in the late 19th century, with all of these explorations of visual illusions that you see in pretty much every neuroscience textbook or book that you open. Listeners will be familiar with them—you’ve probably seen examples of triangles where you actually perceive a triangle in a picture when there is, in fact, no triangle. Yeah, you can find that picture.

Tim Scarfe

Yeah, sorry. I hope I’m not distracting you.

Max Bennett

No, no, no. So that’s a standard finding that was observed in the 19th century: this idea that clearly the brain observes the presence of things even though they’re not actually there. We perceive a triangle there, a sphere, a bar, and the word “editor,” when, in fact, if you actually examine it, the letter E is not there. There’s evidence that suggests the E is there by virtue of showing the shadows, but we did not actually write the letter E there. The brain regularly observes that.

That finding led this scientist, Hermann von Helmholtz, to come up with the concept that what you actually consciously perceive is not your sensory stimuli. You are not receiving sensory input and experiencing the sensory input. What’s happening is that your brain is making an inference as to what is true in the world, what’s actually there, and then sensory input is giving evidence to your brain as to what’s there.

You start from this prior, and then that prior maintains itself until you get sufficient evidence to the contrary; then you change your mind. It’s not hard to imagine why this would be extremely useful in any sort of environment that an animal might evolve in. Suppose you have a mouse running across a tree branch at night. First, I see the tree branch in the moonlight, so I build a mental model of the tree branch. As I move forward, I lose the moonlight and no longer see the tree branch.

As long as I’m stepping forward, the evidence is consistent with my prior of the tree branch. It makes way more sense for me to maintain the mental model of that tree branch, as opposed to all of a sudden having the tree branch disappear because I no longer see the sensory stimuli of it. Because sensory stimuli are very noisy, it makes a lot of sense that we integrate them over time, build a prior, and then, until something gives us evidence to the contrary, maintain our prior about the world.

So that was the first idea: there’s some form of inference. There’s some difference between sensory input and some model of the world that we infer and then thus perceive. What’s interesting, and not as discussed but also present in the discussions among scientists in the late 19th century, is the idea that you can’t actually render a simulation of 2 things at once.

There are lots of really interesting visual illusions around this where you can see something. The famous one is that it’s either a duck or a rabbit. Yep, exactly. And it’s interesting. You can see that a staircase is either moving up to the left, or you’re under the staircase looking upwards, and it’s actually a ceiling that’s jagged.

Why can’t the brain perceive both of those things at the same time? It would make sense if you have a model that there are such things as ducks, such things as rabbits, and such things as 3D shapes that operate under certain assumptions. If that’s true, then you cannot see a duck and a rabbit at the same time, because there’s no such thing. It cannot be the case that the staircase is being viewed from above and below at the same time.

So what your brain is not doing is just perceiving the sensory stimuli. It’s trying to infer what is a real 3D thing in the world that I’m aware of, that this sensory stimuli is suggesting is true, and that is the thing that I’m going to render in your mind.

I think one way this parallels nicely with some of Hawkins’s ideas is that, if you hold the Thousand Brains Theory to be true, and the neocortex has all of these redundant, overlapping models of objects, then it would make a lot of sense that we want to synergize these models to render 1 thing at a time. You don’t want to have 15 different things rendered, because then it’s really hard to evaluate them and vote between these different columns.

It makes sense that the brain says, “Let me integrate all the input across sensory stimuli and render 1 sort of symphony of models in my mind, so I can see 1 thing at a time.” So that’s this idea of perception by inference. At the time, no one really connected that to the idea of planning. This was just the idea that what we perceive is different from the sensory stimuli we get.

Later on, as the world of AI started thinking about things from the perspective of perception by inference, what we end up realizing is that this idea of perception by inference, if you’re going to train a model to do that, comes with this notion of generation. The way it self-supervises is that it takes the prior, tries to make predictions, and compares the predictions to what occurs in the world. As long as those predictions are below a threshold, I maintain my prior.

A famous version of this is the Helmholtz machine, which Hinton devised. I think that was in the 1980s. It could be later; I forget. This is the basic idea that you can build a model. A lot of people use the term “latent representation.” Some people don’t like that term for a variety of philosophical reasons.

It builds a representation of things by virtue of building a model. In other words, perception by inference. The way you build that model is by constantly comparing the predictions generated from that model to what actually occurs. This also has synergies with a lot of Hawkins’s ideas, where we think about intelligence as prediction.

If you build a model of perception by inference by virtue of generation, then it’s relatively easy to say, “Okay, what happens if I just turn off sensory stimuli and start exploring the latent representation?” Now we’re exploring a simulated world. I’m able to cut off sensory stimuli, close my eyes, and imagine a chair, rotate the chair, and change the color of the chair.

Because this model is relatively good and has relatively rich features about how the world actually works, I can model things without ever having experienced them, without ever having done them, and reasonably predict what would actually happen if I were to do those things.

What I think is interesting, and perhaps somewhat of a novel proposal in the book, is that a lot of people think about the neocortex as having adaptive value because of how good it is at recognizing things in the world. If you read a standard textbook, a lot of what people talk about regarding the neocortex is how good it is at perceiving things—object recognition.

Some of the best-studied parts of the brain are this visual neocortex, so we understand reasonably well how we’re building models of visual objects, et cetera. But from an evolutionary perspective, this is a little bit hard to find convincing, because if you actually examine the object recognition of vertebrates, it’s incredibly accurate.

Tim Scarfe

I mean, a fish can recognize human faces. A fish cannot recognize an object when rotated in 3D space. So it’s hard to find a dividing line between object recognition in animals with a neocortex and object recognition in animals without a neocortex, with brain structures that seem more similar to early vertebrates.

Max Bennett

Did you have a question?

Tim Scarfe

Yeah. Well, I just wanted to touch on a couple of things there. Hawkins said that, first of all, we overcome the binding problem by having this profusion of individual sensorimotor models rather than having this feed-forward enrichment of representations.

That was really interesting, but he also said the reason why having, let’s say, 150,000 mini cortical columns that are wired to different sensorimotor signals gives us robustness of recognition is the diversity and sparsity. Then you can think, “Okay, what do the representations look like?”

If our brain builds some kind of model of the world, some kind of topological model, it must be a representation. It’s not necessarily a homunculus, and it’s not like there’s a stapler inside the brain. If I’m modeling a stapler, it’s actually some weird structure of the stapler as seen by every way I can touch, feel, hear, and lick a stapler, or whatever. So it’s difficult for us to imagine what that is.

But the reason I’m going down this road is because that’s a bit weird, isn’t it? We have this very weird representation of things. Then I come to Hinton’s Helmholtzian generative model, and you can get it to generate, let’s say, the number 8 if it’s trained to do numbers. Hinton would argue—and I would disagree with him—that the model understands what an 8 is.

Now, this is weird, isn’t it? We understand what a mouse is, but intuitively we feel that a neural network doesn’t understand what an 8 is. I would argue that we’re getting into semantics here. I think the reason we understand things is because there’s a relational component to understanding.

Semantics is about the ontology: the way we feel about things, where the thing came from, what intention guided the creation of the thing, and what the provenance of the thing was. It’s almost like the interconnectedness of the thing tells you more about the meaning of the thing than the actual thing itself, in a weird way. What do you think?

Max Bennett

Well, I do think this is where the word “understanding” can mean different things to different people. I think there’s absolutely something to the idea that just because you can recognize something—a feed-forward network that can observe a stapler alone—is insufficient for what most people would mean when they use the word “understanding.”

Just to talk about interrelatedness, if I have a feed-forward network that’s just a binary classifier—“Is this thing a stapler or not?”—I can’t ask many things of that feed-forward network that I would expect of an agent or a model that understands what a stapler is.

I couldn’t ask it, for example, what would happen if I burned the stapler. What would happen if I opened it? What would you see inside of it? I can’t ask, “What does a stapler do?” I can’t show it a human holding a stapler and a set of objects in front of them and ask, “What do you think the person is going to do next with this stapler?”

Clearly, our intuitive understanding of the word “understanding” contains some richness that isn’t included in just classifying or recognizing the presence of objects. I think that’s absolutely the case.

My intuitions fall in the direction of what we typically mean when we use the word “understanding”: having something that can be mentally explored. I think that requires what you’re describing, which is the interrelatedness of things.

When I see someone holding a stapler and you ask me what things they would likely do, I start imagining what they might do with that stapler, and then I can evaluate which ones seem plausible to me. In the imagination and evaluation of which things seem plausible, there’s an interconnectedness between the thing and the world around it.

I absolutely agree with you that just recognizing objects clearly lacks something that we mean when we say “understanding.”

Tim Scarfe

Do you take into consideration the memesphere? We’ve got ideas that are quite collective as well. I was going to explore this with you later, but maybe let me frame it like this: I have this intuition that a lot of our intelligence is outside of the brain.

If I were in the wilderness and disconnected from society, I would be a lesser human being. In a weird way, I might have more agency, but I wouldn’t have access to all of these rich cultural tools, knowledge, patterns, and so on.

It’s almost like that’s where a lot of our intelligence comes from, rather than just being able to plan in the brain. Meaning comes from there as well, and there must be some kind of interplay. Culture must shape the development of our brain, and vice versa, but culture seems to be more dynamic. How do you wrestle with that?

Max Bennett

Well, that’s undeniably true. The first example that comes to mind is writing. What would humans be if you removed the technology of writing? We would all realize that we’re not that smart.

Writing is a technology that externalizes a feature of the brain—memory—which brains aren’t that great at. We do a good job of condensing aspects of a memory. For episodic things and procedural memory, we’re relatively good at those, but for semantic memories, we’re terrible.

Externalizing memory with writing is one of the key technologies that enables us to be much smarter than we are, because we now have this external device that enables us to store a largely infinite number of memories and translate them across generations.

That alone proves your point: what humans are capable of is clearly some relationship between brains and external things. Those things can be writing tools, other brains, sharing ideas, and getting challenged. So, yes, I agree with you.

Tim Scarfe

Intelligence is the dynamics, isn’t it? It’s all of the low-level dynamics of things interacting with each other. We can take a snapshot of language and say, “Oh, well, language isn’t the intelligence, but language itself is a form of intelligence.”

I’m not just talking about the words and language models. We’re talking about the actual language in our culture. I think of that as almost a distinct form of intelligence.

Max Bennett

Yeah. It almost gets into philosophical territory: where do you draw the bounding boxes around the things that are imbued with intelligence and the supportive mechanisms that allow those things to have intelligence?

Through one lens, you could think about brains as the physical entities in which intelligence is instantiated, and language as a supportive tool. You could take a very odd view, perhaps, which is that language is the thing that’s evolving and is simply instantiated in these brains that produce and consume the language. It’s language that’s evolving.

In the same way, we don’t think about intelligence on the level of an individual neuron. We don’t imbue a neuron with intelligence, but on the scale of 86 billion neurons, we think something has emergently appeared that we deem intelligent.

There are some great science-fiction books where intelligence gets instantiated in colonies of ants. Each individual ant isn’t intelligent, but somehow the colony itself is capable of doing incredible abstractions.

So, yes, there are very interesting ways to think about how one divides the lines between the physical entities in which these things are instantiated. That said, I have a particular interest in brains because, if we’re looking for the physical manifestation that we can learn from and thus try to understand ourselves—I think all species have an interest in understanding themselves—but also if we want to borrow some ideas from how biological intelligence works into AI, then of all the physical things to examine, the brain seems clearly the one that is probably richest with insight.

But, no, I agree. I think your point is well taken.

Tim Scarfe

Interesting. Okay. I want to close the loop on what you said about the brain being an imagination-filling-in machine. You said that it does the filling-in one at a time. It can’t unsee visual illusions, and the evidence is seen in the wiring of the neocortex itself, you say.

It’s shown to have many properties consistent with a generative model. The evidence is seen in the surprising symmetry and the ironclad inseparability between perception and imagination that are found in generative models in the neocortex. You give examples like illusions, how humans succumb to hallucinations, why we dream and sleep, and even the inner workings of imagination itself.

So it really seems plausible when you think of it in that way.

Max Bennett

Yeah. I think there's a reason why so much of the neuroscience community has rallied around this idea of predictive coding, which is very related to active inference and generative models. What I'm saying there is not really novel. There's just so much evidence that what's going on in the neocortex—the imagination of things, episodic memories—is consistent with this idea of a generative model.

There's been some good evidence that episodic memory—in other words, thinking about the past and imagining the future—are in fact the same underlying process happening in the neocortex. If we look at the connectivity patterns, which I didn't talk about too deeply in the book because it's a little technical, what you would expect from a generative model is that backward connections would be much richer than forward connections because you're modulating downstream.

Of course, the neocortex is not perfectly hierarchical, but things that are generally lower in the hierarchy would have lots of inputs from parts of the neocortex that are higher in the hierarchy. That's absolutely what we see. So, there's a lot of evidence that these are two sides of the same coin, which is that there's some form of generative model being implemented.

I do think that in AI, one way in which this manifests is the very clear success of self-supervision. The principle of self-supervision is this idea: can a system end up having really interesting emergent properties and generalize well when you only train it on predicting the sensory input that it receives? That clearly has become the case.

I think the Transformer is a great example of how, if you just give it a bunch of data and train it through self-supervision—that is, masking, where you hide certain data inputs—it becomes remarkably accurate and good at generalizing across data that it hasn't seen before. That is, in principle, what people are predicting or claiming that the neocortex is doing as a generative model.

Tim Scarfe

Yeah, it feels to me that there's a big difference between, let's say, a Transformer and the neocortex. I think the difference is—maybe agency is not the right word—that you can think of the neurons as having some kind of autonomy. They're sending messages to each other, and eventually it's consistent: the other neuron will get the message, and it will decide for itself what it's going to do.

In a Transformer, just because of the way it's connected, the backpropagation algorithm, and so on, they all ride a Mexican wave together, to use an analogy. So, it feels like a difference in kind to me.

Max Bennett

Clearly, as I argue in the book, I don't think that the brain is just one big Transformer. I would agree with you in the human brain, unless you think there's something that's nondeterministic and sort of magical happening. I think you would still say that there are either base firing rates of neurons, and then there's sensory input that flows up and goes through the brain until eventually it's affecting muscles and you're responding.

So, there is a deterministic flow happening. It might not be as feed-forward as what's happening in a Transformer, which is definitely the case, but they might both be deterministic in a similar fashion.

I think a lot of people have had this exact argument with me, and one counterargument that people have towards this idea—I don't know if I fully agree with it, but it's interesting—is that attention heads really are doing something more magical than we give them credit for. They are dynamically rerouting and effectively resetting the network based on the context that the prompt is getting.

Although technically it's just a series of matrix multiplications, if what's happening in principle is that these attention heads are doing something really clever—looking at the context of a prompt and then effectively dynamically reweighting the network to decide what it cares about and what it doesn't—there are people who think there is something really interesting happening in the Transformer that might be analogous to certain things happening in the brain.

Clearly, these feed-forward networks are not capturing everything that's going on in the brain.

Tim Scarfe

Yeah, it's really interesting what you said, because the way I read that is that things like ChatGPT and language models are entropy smuggling or agency smuggling. What that means is they just do what you tell them to do, and all of the agency—my directedness—comes from me.

I give it a prompt, it does the thing that I want it to do, and then the mapping that you were talking about, I interpret that a bit like a database query. Depending on the prompt you give it, it'll activate a certain part of the representation space and give you a certain result back.

But the brain has this thing where all of the neurons have their own directedness. The weird thing is that, at the cosmic scale, agency seems to emerge. Even Transformer models that were acting autonomously could presumably, at a large enough scale, give rise to something that we think of as directedness, goals, purpose, or whatever.

It's almost as if, in the natural world, because there are so many levels and scales of independent, autonomous things mingling with each other independently, and then downstream mixing their information together and rinsing and repeating over many different scales, that gives rise to all of these amazing things like agency, creativity, and so on.

Max Bennett

Yeah. I think the notion of agency is an interesting one. I'm very amenable to the idea that there is a debate between the reinforcement learning world and the active inference world about how much of intelligence can be conceived as optimizing a reward function.

The hardcore reinforcement learning world would say that everything is just a reward. The active inference world would argue that not all behavior is driven just by optimizing a single reward function. There is uncertainty minimization, trying to satisfy your own model of yourself, fulfilling your own predictions—these sorts of things that seem very well aligned with the behavior we see.

It's unclear which of these is right. It's probably some balance of the two. But people would conceive of agency differently in these 2 worlds.

Some people in the reinforcement learning world would say that agency is just giving something a reward function, and then it learns over time, trying to optimize that reward. In the more active-inference world, which I do think has legs and to which I'm obviously amenable, the idea of agency is a little bit more about building a model of yourself and trying to infer what your goals are based on observing yourself, and then trying to make predictions to fulfill those end goals.

In other words, goals are constructed. One of my favorite Friston papers is “Predictions Not Commands.” I don't know if you've read that paper, but I think it's a brilliant paper about how you could reconceive the motor cortex—not as sending motor commands to your body, but as building a model of yourself and predicting what will happen. The way the spinal cord is wired is that it just fulfills those predictions.

I think that's a really interesting reframe of how you could get agency and really interesting, smart behavior in the absence of just a strict reward function. The way that would learn is by trying to model the behaviors it observes, and then trying to predict and fulfill them.

Agency is a really interesting concept because it manifests itself in these different paradigms in different ways.

Tim Scarfe

Yeah, I find it fascinating, because the way I read it in the active-inference literature, it's a very principled definition of an agent. There's still a bit of a gap, because I think Friston would argue that in the natural world, because of the laws of physics, particles, and whatnot, you get the emergence of things, and things become agents when they have a certain depth of planning, should we say.

But it's really interesting. I guess he would argue that you get all of these phenomena that give rise to agency, like biotic self-organization and so on. Maybe we should slowly go in that direction. You give the example of mice doing planning. Can you sketch that out?

Max Bennett

Yeah. This is another real area of neuroscience research that I absolutely love. I think it was in the 1940s or 1950s—I forget the exact decade—when Tolman observed that mice, when they reached choice points in mazes where he was training them to navigate, would pause. They would sniff back and forth, and then they would choose an action.

And so he hypothesized this idea that they must be engaging in vicarious trial and error. They must be imagining possible outcomes before deciding. Of course, this was hugely controversial because there was no evidence. He had no evidence that they were, in fact, imagining anything.

Most people, in the absence of evidence, like to assume animals are as dumb as possible. Only when there's irrefutable evidence will we imbue them with any intellectual capacities, which I think is an interesting human bias, but that's fine.

David Redish, who is also a close friend and mentor of mine, did some amazing research with one of his PhD students where they were recording hippocampal place cells. So, as a very quick background for viewers, you can go into the hippocampus of really any mammal, but this is best studied in rats. Part of the hippocampus, a region called CA1, has these things called place cells.

If you record the cells as a mouse is moving around a maze, what you find is this incredible thing: there are neurons that activate only in specific locations in that 2D plane. It's not based on how they got there. It's allocentric; it's independent of their egocentric path. They can come back to the same place from any route, and that same place cell will activate.

As an aside from the evolutionary story, we find similar types of cells—not exactly as accurately—in fish, in the homologous region of the hippocampus in their cortex, where they have place-like cells. They're not as accurate, but they are cells that activate in certain locations in a maze.

What he found is that when mice engage in this act of vicarious trial and error—when they pause and look back and forth—the place cells in the hippocampus cease to activate only in the location where they are. They actually start activating down the paths of each route they might take. In other words, you can literally watch rats imagining the future. I think this is one of the most incredible neuroscience findings.

He then took this and did a bunch of other experiments that I think reveal even further the power of imagination in rats. One of my favorites is his counterfactual learning studies, where he puts rats in this thing called Restaurant Row. It's a square-like maze—yeah, exactly—and as the rat is going counterclockwise, at each door, a sound is released or made.

That sound signals to the rat whether it can go right through the door and get food in, I think, 3 seconds, or whether it's going to have to wait 45 seconds before it gets the food. They're given a bunch of time to try to get as much food as they want. Rats have clear preferences, so some rats will really prefer bananas, and they don't really like the bland food.

What happens? This presents a set of irreversible choices to a rat. Let's say they come up and they can either get a treat right now that they don't really like that much—let's say it's the bland treat—or they can go to the next one and hope that they're going to get the banana really fast. If they go to the next one and the banana sound is long, meaning 45 seconds, then they regret their choice because it would have been better if they had just gone in and quickly gotten the food.

How do we know they're regretting the choice? We can literally watch them imagining eating the foregone choice. We can go into a part of their brain called the orbitofrontal cortex, which activates for certain types of tastes, and we can see them imagining the foregone choice. They end up making different choices the next time around. They end up being less likely to forego that choice in the future.

I think this is such an incredible finding of what we mean when we say model-based reinforcement learning is clearly happening in the brains of very simple mammals.

Tim Scarfe

There's a real challenge in knowing which simulations to run because, if you think about it, we've got a search problem, right? There's an intractable number of simulations to run. How do we fix that in AI, and how do humans fix that?

Max Bennett

This is one of the many outstanding questions in AI, but one of the big ones is: how do you effectively prune the search space?

We do not know how mammal brains do this so well. I can give you some high-level ideas or theories, but we just don't know, and this is one of the big possible breakthroughs in figuring out how mammal brains do such a good job of this.

The thing that AlphaGo does, which I think perhaps is a clue and is clever, is that the selection of the search space is actually bootstrapped on the temporal-difference learning model under the hood. This is very clever. Let's say you train something to learn without a model of the world. All it's doing is getting sensory stimuli. It gets a model of a Go board, and then it just predicts the right next action. It just has a policy-value function that bootstraps on each other, et cetera. So, no planning.

If you want to add planning to that, what they did, which is quite brilliant, is say, “Instead of building some other system to try to choose good trajectories, why don't we just use the policy network?” We don't just pick the first one—we pick its favorite move, but then we also look at its second-favorite move, its third-favorite move, and maybe its fourth-favorite move. Then let's literally play the games out.

Let's play a bunch of games against ourselves and then see the ratio at which we win them. What we might learn is that our second-best guess was actually better than our first-best guess. But we're not starting from every possible possibility. We're saying, “Let's bootstrap on our best guesses of good moves, but then check them by playing out the possible futures.”

If we were to analogize that to the brain, what that would suggest is that perhaps it's the basal ganglia—which a lot of evidence suggests is engaging in this type of model-free reinforcement learning—that is actually the thing that chooses the moves. But there's some other system that lets us choose the second-best move, the third-best move, et cetera.

One way this might happen—there's some evidence for this, far from conclusive—is that there's some notion of uncertainty that the frontal cortex or basal ganglia is measuring. When the level of uncertainty between the next actions passes a threshold, pausing occurs. When we see animals do this vicarious trial and error, it almost always occurs in moments of high uncertainty: when contingencies have changed, when the right answer is not obvious.

You could conceive of this as a policy network where you're evaluating its best choice, second-best choice, third-best choice. When there's uncertainty about it—in other words, when they're close together—or there's some other measure of uncertainty, perhaps you have parallel policy models and you're comparing the similarity between them. There are a lot of different ways to do this. That triggers a process of playing forward.

This is another key thing that mammal brains do that AlphaGo does not do. AlphaGo engages in planning on every move. There was never the question of when to pause to plan in a game of Go. It doesn't matter; just engage in planning on every move because it can do it so fast. In the real world, there's so much uncertainty and noise, and human brains need to be so energy-efficient that we can't engage in planning every instant. We need some mechanism that tells us when to stop and think about what we're going to do next and when we can just continuously go with model-free choice.

This is also something we don't know how mammal brains do. But I think a reasonable speculation is that there's some uncertainty measurement occurring.

One last point I'll make that I think parallels, in an interesting way, some of Hawkins's ideas and Friston's ideas: if you take the Thousand Brains model and apply it to the frontal cortex—in other words, we have multiple parallel models of ourselves—you could imagine that there's an uncertainty measurement, the same way we do uncertainty measurement in a lot of deep-learning models, where you create parallel models and just measure how similar the predictions of the models are.

If multiple parallel models predict similar things, we measure that as low uncertainty. When they diverge substantially, all of a sudden we measure high uncertainty. Again, this is speculation, but you could imagine that, if it is the case that we have redundant models in the neocortex, then might it be the case that somewhere—perhaps the thalamus or the basal ganglia—the similarity or differences between these predictions are a measure of uncertainty that triggers pausing?

Stephen Grossberg has similar ideas. He calls this matching and nonmatching.

Tim Scarfe

Yeah. Yeah, almost. Our ability to do abduction is something that fascinates me, and there is some kind of model-selection or matching step that goes on there.

Anyway, we've got the neocortex. It's an absolute beast when it comes to predicting sensory signals, and then we see the emergence of planning and also smart planning, as we've just spoken about. So now we're actually talking about traversing these sensory networks over space and time.

Then something else really interesting happens. The next 2 moves you make are bringing in selfhood—bringing yourself in as an explicit actor and including that in the planning—and then that naturally leads to this idea of, I would call it, teleology, but you would call it “why” or intentionality.

So let’s go on that journey. Maybe self-modeling first: how does that come into the picture?

Max Bennett

I think there are 2 notions of self-modeling. One notion of self-modeling is the kind that I think we see in early mammals. This is an idea where the frontal cortex of mammals—a region generally called agranular prefrontal cortex, which is present in all mammals and is largely believed to have existed in the very first mammal brains—gets sensory input from an animal’s own interoceptive signals.

So it gets input from the hypothalamus, which measures things like hunger. It gets input from the amygdala, which measures things like valence in the world—fear, danger, and so on. When does the agranular prefrontal cortex get most excited? It’s in moments of uncertainty and in moments where animals are engaging in planning and episodic memory. If you damage the agranular prefrontal cortex in rats, they seem to have dramatically impaired, if not completely lost, the ability to engage in mental simulation, episodic memory, and so on.

I think what one might speculate is happening here—which is not a novel suggestion by me—is that it’s modeling the self. In other words, it’s modeling the activations of the amygdala and the hypothalamus that are happening, and why I am doing what I’m doing. If I wake up and I see that I have these hypothalamic activations, and then I go down to get water, then it builds a model: in the presence of these hypothalamic activations, the next action is, “I’m going to go over and get water,” which constructs an explanation.

As odd and philosophical as that sounds, that is, in principle, computationally exactly the same thing as when we showed a picture of that triangle and your posterior sensory cortex constructed an explanation of what it saw: “I perceive the triangle.” This idea of constructing an explanation of one’s own behavior is the first idea of self, which is constructing intent.

There’s lots of evidence that even in rats, if you record neurons in their agranular prefrontal cortex, they seem to be very sensitive to the tasks they’re in and to measuring progress toward goals. There’s lots of evidence that it’s doing something akin to that. So that’s one notion of self.

When you get to primates, you see a whole new region of frontal cortex emerge: what’s called granular prefrontal cortex, which is only seen in the primate lineage. There are no other mammals that have this region of prefrontal cortex.

As a quick aside for anyone who’s interested, the reason it’s called granular versus agranular is that most neocortex has 6 layers. The 4th layer is called the granular layer because it contains granule cells, which are just a certain type of neuron. For a variety of really unknown reasons, although there are some interesting speculations, agranular prefrontal cortex is missing layer 4. That’s why it’s called agranular. It’s the same thing with motor cortex, which is missing layer 4.

Most mammals’ whole frontal cortex is missing the 4th layer. It only has 5 layers. But in primates, you get this granular prefrontal cortex, this huge region of neocortex that does contain a layer 4. The best explanation for this I’ve actually seen Friston talk about. I’m happy to go into it if you think your folks would be interested in it, but the point is, there’s a new—

Tim Scarfe

Oh, yeah. I mean, we love Friston. Active inference is actually about preferences because an agent expresses agency by adapting the environment to suit its preferences, or to make the environment like its preferences. So this is a theory of volition, right? What Karl Friston is talking about is: where do these goals come from? Where does volition come from? You spoke with Karl about that, I think, over several years, didn’t you?

Max Bennett

Yep. Karl’s been a wonderful mentor of mine. It’s a funny story: I won—I didn’t know this was a term in academia, but there’s the reviewer lottery, where you just get lucky and get a reviewer. For the first paper I submitted, Karl Friston was a reviewer on it, which I was just lucky to have happen. Then he became a mentor, reviewed my book, and gave me lots of good feedback. So, yeah, he’s an amazing person and has been a wonderful mentor of mine.

Let’s talk about granular versus agranular, because I think the best theory I’ve seen is Friston’s theory on this. What does layer 4 do in neocortex? Across the entire neocortex, layer 4 is where sensory input is received. The primary sensory input is received into the neocortical column. This comes from the thalamus.

The canonical model is that sensory input from sensors—eyes, ears, skin—flows up through the brainstem to the thalamus, and then from the thalamus propagates to layer 4. From layer 4, it goes within the variety of other layers of neocortex, and then the other layers of neocortex project back to the rest of the brain.

Why would it be the case that regions of neocortex would not have a layer 4? If you actually watch an animal’s development, what’s interesting is that, in mammals with an agranular prefrontal cortex, it’s not always agranular. It actually starts having a layer 4, and the layer 4 atrophies over development.

I think this mirrors well with Friston’s idea of active inference. What’s happening is that the neocortical column can be in 2 phases. It can either be trying to match its model of the world to its sensory input—in other words, I see sensory input, I’m trying to infer what’s there, and I’m going to construct the idea of a triangle—or there’s another state of a neocortical column, which is generation: I’m going to start from the latent representation of a triangle, imagine it, and explore it.

One idea is that what frontal cortex does is primarily try to fit the world to its model. In other words, it spends the vast majority of its time constructing intent, not trying to modify that intent to fit what it observes, but instead trying to change what an animal does to satisfy its intent.

Layer 4 atrophies. It doesn’t actually go all the way; if you go deep into a brain, you see some basic layer 4. So it’s not completely gone, but it atrophies because frontal cortex spends very little time trying to change its perceived intent to map what the animal is doing. Instead, it tries to change what the animal does to match its intent.

What I think is so interesting and brilliant about this idea is that it explains exactly why layer 4 doesn’t start out as nonexistent. At first, an animal needs to build a model of itself, so layer 4 is present. But over time, it shifts toward, “Once I have a model of what I want, who I am, and the things I would do, I don’t need to spend as much time changing my model of self. I’m going to spend most of my time trying to change my behavior.”

This is a very speculative idea, but it makes a lot of sense in the context of active inference. It’s the best explanation I’ve seen, personally, in all my reading of explanations for why agranularity exists.

Tim Scarfe

This is absolutely fascinating. In my mind, it bridges the gap between internalism and externalism because he’s describing this kind of dialectic exchange. To use the high-entropy words that only Friston uses: if I say “dialectic exchange,” you know. And, by the way, for the folks at home, if you read Friston’s papers, there are certain words he uses. He says, “This licenses something.” If you see the word “licenses,” then Friston wrote the paper.

But you have this kind of thing where agents have models of the world, but they’re exchanging information with the other agents. Then what an agent does is have this generative model of policies, which is just a sequence of actions.

Here’s where I want to get into the nitty-gritty a tiny bit. You could think of those plans as being goals. You can think of a goal as just being an end state in one of your plans. But that doesn’t really satisfy me, because I think of goals like eating food as being a kind of category, not a pointillistic traversal of a specific state in the future. It feels like, well, if there are an infinite number of goals, then in some sense there are no goals at all. So what is a goal to you?

Max Bennett

Great question. I think this is where semantics matters a lot. We can think about goals in several ways. If we think in strict RL terms, they would just think about goals as optimizing a reward function. The goal is simple. There’s only 1 goal, which is to maximize reward. In a changing, complex world, your reward function might fluctuate over time, but the goal is singular.

In the active inference world, what I find compelling is that it introduces a different component of what we mean by a goal. This is not just of intellectual interest because it’s cool; it has very real AI implications, actually, because it contains the notion of explainability.

For example, if I wake up and I’m hungry, and I start imagining ways to satiate my hunger, then I decide I’m going to get in a car, go to this restaurant, and eat this specific food.

When I get into the car and someone calls me and goes, “Why’d you get in the car?” the reason I can explain that so easily is because there was a rendered simulation of a plan that terminated in an end state that I deemed I wanted, that I selected. So it’s very easy to explain why I did that. In the absence of that, it’s actually very hard to explain why you’re doing things.

If you’re walking down the street, if I asked you to explain any one of your model-free behaviors—why did you move your foot there instead of there?—you have no explanation. And so I think you can assign the word “goals” to multiple levels here. I think you could say it’s terminating the end state. You could say the goal is some more abstract representation of the satiation of thirst in general, and you could think about that as distinct from just the reward function.

Some might challenge that and just say, “Well, all that you’re talking about, then, is just optimizing the reward function.” But I think you could make an argument that there is a distinction happening there. What I think is so critical, and that’s unique to what happens within mammal brains—and it’s why I was really honing in on that—is the ability to plan a series of actions that terminates in an end result and then execute that plan. I think that has very clear implications for explainability, which model-free actions do not.

That would be how I would think about goals. But I think it is a little bit semantic, where we can think about the concept of goals in multiple ways.

Tim Scarfe

How do we actually know what the cognitive abilities were of early animals, and why should we care?

Max Bennett

Great question. I think there are 2 reasons why we should care about the evolution of our brains and intelligence. The first is to understand who we are. The scope of what it means to be a human is not constrained to what it means to be a Homo sapiens.

So much ink has been spilled on the last 70,000 years of us being Homo sapiens. But if aliens were to come down and engage with us and analyze us as a species, most of the things they would observe about us don’t come from our legacy as Homo sapiens. They come from our legacy of being a primate, our legacy of being a mammal, our legacy of being a vertebrate, and our legacy of being an animal in general.

If we want to understand what it means to be a human being, I don’t think we can skip the full 600-million-year story of how we came to be. In there is so much rich history and insight about what it means to be us. That’s one really key reason: it’s our legacy, it’s our history, and it’s how we came to be.

The second, perhaps more practical, reason is that I think understanding the evolution of the human brain and the evolution of human intelligence is a key tool in our toolbox for understanding how the brain works and how human intelligence works. It’s by no means the only method. It might not even be the main method, but it’s a very useful method to add to the toolbox.

The problem with going into the human brain and trying to directly reverse-engineer it is that evolution doesn’t work in clean ways. It doesn’t work the way a human designer would. It doesn’t work from first principles. It tinkers.

When we go into the brain, we see all of this messiness. There are redundant systems and vestigial systems. New things evolve that make old things redundant, but they’re still there. Lots of processing is duplicated in different regions.

One way to understand the brain is to continuously probe it as the human brain is, which is fine. But another method that’s also useful, and can impose constraints for us, is to actually track the history of how it came to be.

That can provide insights as to when this brain modification occurred, such as when the neocortex evolved or when the basal ganglia evolved, what new abilities this enabled, and how it affected the prior brain regions that were already there. This can give us insight into how the brain works today.

In the toolbox that we have for ways to reverse-engineer the brain, I think this is just an underappreciated one that is worthy of being included. Of course, understanding how the human brain works has so many different applications. It helps us with mental health, it helps us with understanding why people do what we do, and it helps us with building AI systems. I think there are lots of insights to garner from the brain.

That’s why to do it. Now, how to reverse-engineer what behavioral abilities exist in our ancestors is a really interesting question. Of course, we can’t go back in time. What we can do, though, is use mechanisms to reverse-engineer what their brains looked like and mechanisms to reverse-engineer what abilities they had.

To understand what their brains looked like, we can look at other animals in the animal kingdom. For example, we can look at all of the existing primates and all of the existing non-primate mammals, and we can see what the common brain structures are between them.

We can look at genetic analysis, meaning what things seem to derive from similar roots, and we can infer what seems to be common and shared amongst them. Thus, what do we think was actually existing in the brains of the first mammals?

We can do the same thing with fish and reptiles to infer what existed in the first vertebrates, and we can do that with invertebrates to try and infer what existed in the first bilaterians—in other words, the first animals with brains.

For behavioral abilities, there are 3 ways you do this. This is my approach to trying to infer behavioral abilities. I call them the ingroup condition, the outgroup condition, and the stem-group condition.

In order to make the argument that a behavioral ability emerged at a certain location in our evolutionary history—for example, a behavioral ability like episodic memory evolving with the first mammals—you need to satisfy these 3 criteria.

The ingroup condition stipulates that most descendants of this species—in other words, most mammals—should show this ability. It doesn’t mean all of them. Abilities get lost all the time, but most of them should show this ability.

The neural mechanisms by which the ability emerges should come from homologous regions. What that means is a shared neural underpinning. If, for example, mammals show episodic memory but it comes from neurological regions that independently evolved along different mammal lineages, that suggests it wasn’t present in the first mammals.

But if they all emerge from regions that emerged with early mammals, that’s good evidence that this ability also emerged with mammals.

The outgroup condition says that most doesn’t have to be all, but at least many non-mammal vertebrates—the outgroup, the group right above—should not show this ability. If they do show the ability, it should emerge from non-homologous regions; in other words, parts of their brain that evolved independently.

For example, birds definitely show episodic memory. But when we look into the brain regions from which episodic memory emerges, it’s clearly non-homologous. It’s a part of the brain that mammals, or the early vertebrates, did not have.

The stem-group condition is that, in the ecological dynamics in which early mammals existed, or this ancestor existed, we should be able to devise an argument for why this ability would have been adaptive. Why would episodic memory have evolved?

With these 3 things, we can start to infer the story of when behavioral abilities emerged. Is this perfect? Absolutely not, because we do not have a ton of data on behavioral abilities across species. As new evidence emerges, the story might change.

But with these 3 conditions, we can do a reasonable job inferring what abilities emerged. The main finding of the book—or the research that led me to be so excited about the book—is that when you do this, what’s kind of crazy is that you find, as a first approximation, a really coherent story.

A lot of the behavioral abilities that emerge at each milestone in brain evolution don’t seem to be a haphazard array of different skills. They often emerge from one underlying intellectual capacity, which I call a breakthrough, applied in different ways. That’s the idea of the 5 breakthroughs.

One thing to note about the basal ganglia is that I think this is a bit of a sidebar, but I think it’s fun. The basal ganglia is one of the most underappreciated parts of the brain, in the sense that so much work has gone into understanding how the neocortex functions. So much work has gone into the neocortical column, and that’s all wonderful work to be done.

But the basal ganglia is not only evolutionarily much older. If you look in a lamprey fish, as we talked about last time, a lamprey fish has a common ancestor with us from 500 million years ago. It’s one of the most distant vertebrate cousins that still exists today.

Lampreys have a basal ganglia that looks exactly like our basal ganglia: the same internal structure. The basal ganglia also has perhaps one of the most beautiful internal structures that can be computationally reverse-engineered.

There isn’t good consensus on the actual internal wiring and computations performed by a neocortical column, but there is much broader consensus as to what’s being executed by the basal ganglia. Without getting overly technical and perhaps boring people, I would encourage anyone who’s computationally interested in this to dive into the literature here.

It is almost beautiful that evolution came up with this. For example, the input structure of the basal ganglia has this mosaic of neurons that each express 2 different types of dopamine receptors.

This would be my own little technical diatribe. One type is called D1 receptors, and then there are D2 receptors. D1 receptors, when they receive dopamine, strengthen connections. D2 receptors, when they lose dopamine, strengthen connections.

Now you track these different neurons, and they actually split their paths. The D1 receptors go to a nucleus that, when activated, disinhibits behaviors, and D2 neurons go through a different set of nuclei that, when activated, inhibits behaviors. We can literally watch how dopamine signals drive repeated behavior and how dopamine drops inhibit behavior.

We can literally look at the mosaic of connectivity here and say, “When you spike dopamine, it weakens the stop pathway through D2 neurons, and it disinhibits the go pathway through D1 neurons, making you more likely to repeat the behavior, and vice versa if something bad happens and you lose dopamine.” I think it is so beautiful that evolution stumbled on something that clean in its macrostructure. So, diatribe on—I think the basal ganglia is cool.

Tim Scarfe

Yeah, it’s quite interesting as well. People get addicted to drugs, and a lot of that is about wireheading in the basal ganglia. Of course, habitual learning is something that, when it becomes so habituated, moves down the stack into the basal ganglia, and drug-taking would be an example of that. You actually cited, I think, an experiment in China where they removed part of the basal ganglia, and there was a 40% recidivism rate for addiction.

Max Bennett

Yeah, a very controversial study that probably violated many ethical codes in the US. They did this study on people with intractable heroin addiction, and they lesioned a part of the basal ganglia called the nucleus accumbens, which is sort of where goals are habitually selected. It showed a dramatic reduction in heroin addiction. It also had other side effects that doctors might deem unreasonable, but it definitely worked and absolutely reduced the addictive cravings triggered by stimuli.

Tim Scarfe

Yeah. The story of this chapter—an incredible chapter I’ve just been studying today in great detail—is the story of mentalizing, but I would call it social complexification. Actually, that’s a hypothesis for why our brains expanded so dramatically. There was this extinction event. I think it was the Devonian extinction event, and only birds and not many other things survived. Then we got the chance to evolve after that, and our brains rapidly exploded.

There are different theories about why that happened. Maybe it was because we had access to loads of calories in the form of fruit—preferential access—and it gave us an incredible amount of time and energy, the excess of which might have led to social complexification. Can you just give us a little bit of background about that first piece?

Max Bennett

Absolutely. We don’t know—there are lots of speculations—but we do have some really good evidence that at least part of what drove the explosion in primate brains was social dynamics. Robin Dunbar did the seminal work here. What he showed is that, in primates, the encephalization quotient—which is just the ratio of brain size, especially the neocortex ratio, so the ratio of the size of the neocortex to the rest of the brain—is extremely correlated with social group size in primates.

The bigger the social group size in a group of primates, the bigger their neocortex seems to be relative to their body size. What’s so interesting is that you don’t see this in most other mammals. This is not a standard correlation that applies across the animal kingdom. It seems to be a correlation that’s very specific to primates. There might be other mammals, but for most mammals, you don’t see this correlation.

Robin Dunbar’s famous social brain hypothesis is that what drove the explosion in primate brains was some form of social dynamic between people. The more social relationships you’re managing, the bigger your brain has to be. Now, that doesn’t explain what specifically happened in the brain, which is what we can get to with mentalizing and why that applies. But it does suggest that whatever drove this explosion in brain size seems to be something correlated to social grouping.

What’s interesting about primate social groups, relative to—not all mammals, but many other mammals—is that they’re very, very political. Many mammals live in solitary social lives, where the males mostly live alone and females will rear a child and then usually go off on their own. There are animals that live in herds, where they socially group together, but there aren’t really very rigorous hierarchies amongst them.

Primates, especially apes, have these really complicated social structures with rigid hierarchies. There truly is someone at the top of the hierarchy, and we can measure this. Primatologists have gone to painstaking lengths to verify it. For example, there’s transitivity: if you show that one primate tends to show a submissive signal to another primate, and that other primate shows a submissive signal to another one, then it’s almost definitely the case that the first one will show a submissive signal to the last one.

In other words, these are real, rigid hierarchies, not random interactions of submission and dominance. One of the main ways you survive as a primate is by successfully climbing this hierarchy. What’s so interesting is that, in many mammals, what makes someone the top dog, or the person at the top of a hierarchy, is brawn. It’s just strength. They’re trying to flaunt who would win in a physical altercation, which evolutionarily is beneficial, because if you can prevent actually fighting each other and just say, “This is who would win the fight,” then you both save energy by having these fake battles, and whoever wins gets to eat the food, et cetera.

But with primates, it’s not always the strongest one that reaches the top. It’s the most socially savvy one. Social savviness comes into play through alliances that are built within primates. You’ll see that people at the top of the hierarchy frequently befriend, groom, and come to the aid of certain other non-family members. Those folks will thus reciprocate and come to their aid. There are these really interesting dynamics that play out. You even see wars. There are mutinies that take place in primate societies.

In this soup of the way to survive and gain evolutionary advantage as a primate, the goal is not perhaps only to make sure you get access to food, but to climb a social hierarchy. All of a sudden, there are huge social pressures to infer what someone else would do in a certain circumstance, what someone knows, what you can get away with, or how to change someone’s opinion of you.

What that lines nicely to is what we see in the new brain regions that emerged in primates. Most notably, there’s a brain region called the granular prefrontal cortex, and there are areas in the back of the brain called the superior temporal sulcus and the temporoparietal junction. These brain regions across primates are highly implicated in what I call mentalizing—which is thinking about thinking—but the standard literature would call this theory of mind. That means being able to infer the intent or knowledge of someone else.

It’s easy to understand why this would be so adaptively valuable in a politicking arms race where you’re trying to deceive each other. There’s a great study that really revealed this with primates by Emil Menzel in the 1970s, and I love this story.

Tim Scarfe

Are you going to do the Machiavellian apes?

Max Bennett

Yes.

Tim Scarfe

Yeah. Yeah. God. [Laughter]

Max Bennett

Okay. So, Emil Menzel had this 1-acre forest, and his main objective was not to study ape social behaviors in the sense of how they would climb social hierarchies. His only objective was to measure spatial reasoning in chimpanzees.

He had a group of chimps. There was 1 chimp named Belle, another named Rock, and a few others. He would show Belle the location of food. He would hide food under a bush and then see whether Belle would go back to that same location looking for food. In other words, could she remember locations in 3D space or in a 2D, map-like space?

What he found is that, yes, they readily do that. We now know that lots of mammals are capable of it. In fact, even fish can do things like that. But in this study, he started finding something that was odd.

When Belle would find the food, she would frequently share it with her fellow chimpanzee group members. That was great until Rock, who was a high-ranking, aggressive male, would take the food from her when she shared. So what she started doing was hiding the food when she found it. Instead of sharing with Rock, she would just sit on the food.

Rock realized she was doing this and not sharing. So Rock would come over and push her to try to get the food from under her. Then, when she knew the location of the food—because, on some recurring cycle, experiment, or signal, she knew that the food was now available—she would not go to it until Rock was not looking.

So then what did Rock start doing? Rock started pretending not to look. Rock would look away while Belle went toward the food. Once he noticed she was doing that, he would turn around and run to try to grab the food before her. Then Belle started trying to lead him in the wrong directions, and this cycle of deception and counterdeception kept playing out.

What that becomes is a beautiful anecdote and a case study in what happens when you have a bunch of hierarchically interacting animals in an arms race for things like this. What you get is deception and counterdeception. That’s really only conceivable with some notion of theory of mind, because in order for Belle to trick, or try to trick, Rock, she needs to be able to say, “In order for me to change the knowledge in Rock’s head of where the food is, what I need to do is walk in this other direction. What that will do is make Rock think the food is in this direction, when in fact I know it’s in this other location.”

She also needs to reason about someone’s intentions: “I know Rock intends to trick me, so when he’s looking away, I don’t believe that he in fact is not paying attention.” This was one of the first early anecdotes that some form of theory of mind was occurring in primates. There have since been lots of studies that show this.

For example, just to give some case studies, you can take a chimpanzee and teach it that, when there are 2 boxes, the box with a red mark on it is the one with food in it. They’ll easily learn that. Then you have an experimenter come in with 2 boxes, bend over, and mark one. They pretend to accidentally drop the marker on the other one and then leave. The marking is identical in both cases, but the chimpanzees always go for the one that was intentionally marked. They can infer the difference between the same stimuli: someone intending to do something and something being an accident.

There are other studies of chimpanzees playing with different goggles. With one goggle, you can’t see through it; with another, you can. If you put those goggles on human experimenters, the chimpanzees always go to the experimenter with the see-through goggles, asking for food. They’re somehow inferring that the other person can’t see them, so why would they ask?

There are lots of studies that show this ability. Evidence outside of primates, in other mammals, is very loose. It’s inconclusive and controversial, but the loose evidence shows up only in the smartest mammals, which possibly suggests some independent convergence.

Tim Scarfe

Yeah.

Max Bennett

Yeah, there’s lots of rich evidence that this theory of mind exists within primates, and it emerges from these uniquely primate regions. I’m happy to go into the evidence from brains, but I’ll stop there for a second.

Tim Scarfe

So, with the Machiavellian apes and the X-risk people, there are people who talk about AI killing everyone, and they make the argument—it’s called instrumental convergence, from Nick Bostrom—which is basically that things like power-seeking and deceptive behavior would be instrumental to any end goal. This is a great example from the animal kingdom of deceptive, Machiavellian behavior.

I guess it does seem plausible, at least on the surface, that a level of sophisticated agents following their own intentions and inferring the intentions of others would seek to deceive each other. That’s a natural phenomenon.

Max Bennett

I absolutely think so. I absolutely think it is the case that the more autonomy you give an intelligent agent, and the more ability you give it to define its own subgoals, the more risk emerges. You absolutely get what Nick Bostrom is talking about, which is that a subgoal to trying to help cure cancer might be to dominate all of Earth and control the labor supply and allocation of resources across all of Earth.

But I don’t think that’s necessarily inevitable. I think it is a risk. Evolution is a constrained search algorithm for intelligent entities. It does not give moral weight to what emerges. This is an important distinction: just because something is a natural consequence of evolutionary systems does not mean that we should deem it morally superior.

Tim Scarfe

Yeah, that’s the naturalistic fallacy, right?

Max Bennett

So it might be the case that it is very likely that species will eventually enter a politicking arms race, and certain forms of deception and power-seeking will emerge. That doesn’t mean that when we produce our own intelligent entities in AI, we should imbue them with those features.

One of the optimistic outcomes of this new AI world we’re going to enter in the next 100 years is that we, as designers, can now do our best to try and remove some of the evolutionary baggage that we don’t like, which has evolved in humans, from these new entities. There’s risk, of course, but I think there’s also a really great opportunity that we could have benevolent beings that do not seek to dominate. Yann LeCun talks a lot about this, and about creating beings that are less selfish.

So, I think there’s a great opportunity, but there’s definitely risk, because the second you give an autonomous agent the opportunity to produce its own subgoals, you need to have really rich constraints, a really well-defined reward function, or both. One ability that I think comes from mentalizing—and this is an idea in alignment research—is that if you can convince an AI agent to try and do what it thinks the human wants it to do, what you’re actually doing is requiring it to engage in some form of mentalizing: to infer the preferences of the requester and then try to do what is best for that individual.

You can’t just have it take requests at face value, because then there are all these opportunities for misinterpretation. Nick Bostrom’s famous paperclip maximizer maximizes the production of paper clips, and Earth is turned into paper clips. We obviously don’t want that.

But with mentalizing—with the ability to model the internal simulation of another mind and play out how this person would feel about possible futures—you could imagine, optimistically, an outcome where an AI agent could easily infer, “If I turn all of Earth into paper clips, that’s not what the person giving me this request would in fact have wanted. They would regret that outcome.”

Of course, it doesn’t fully derisk things, but it is one methodology and one learning from evolutionary neuroscience that we can garner. Mentalizing is a tool that can be used to try and stabilize the requests we give each other in a more grounded way, so there aren’t these types of misinterpretations. Of course, humans misinterpret each other all the time. It’s by no means perfect, but it is a tool.

Tim Scarfe

And I think that’s a very natural phenomenon. I think any intelligence system is naturally incoherent. I think it’s impossible to have a single, monolithic intelligence that is monomaniacally focused in a particular direction.

But I want to slightly rewind to what we were saying. The first animals had quite simplistic social games that they were playing. They were interested in strength and submission, and it was a fairly fixed interface.

What was really interesting is that deer, for example, lock horns, don’t they? It’s predictive. They don’t actually have to have a fight, because that would be evolutionarily not a smart thing to do. The social game they play, even though the game is fixed, is predictive, which is fascinating.

Then you were telling the story of monkeys and macaques, how they have this really interesting virtual social game where strength and social status diverged. Social status actually became this virtual thing that was based on grooming and preening and lots of completely unrelated things. It was entirely possible for a very weak macaque to have significantly higher social status than a big, strong one.

That’s really fascinating. But then we get into a broader question. We’re still very social creatures ourselves. We have Facebook, for example, and could you arguably, cynically argue that Facebook—or all socializing—is just a kind of arms race to improve our social status? When we’re posting on Facebook, in a way, it’s like the deer locking horns. It’s us playing these status games without having to fight each other.

Max Bennett

I think there’s an aspect of human behavior that can absolutely be explained by this. There’s a great book called The Elephant in the Brain.

Tim Scarfe

Oh, yes.

Max Bennett

There’s a great book called The Elephant in the Brain that talks about how much human behavior can be explained by this sort of status-seeking behavior. The reason why it’s so hard to study is because it’s what they call, I think, a cognitive taboo or an intellectual taboo. We don’t want to admit it to ourselves.

So we self-delude ourselves into believing we’re doing things for virtuous reasons, because it makes it easier and more convincing. If I know that I’m doing something to deceive someone, it’s easier to tell that I’m doing it. If I genuinely believe that I’m doing these things to help the world, but subconsciously they’re actually just benefiting me, it’s more convincing to other people.

Their argument is that primates evolved this sort of self-deception to make themselves more convincing, and this really accelerated with language in humans and all that.

I think it’s very likely to be the case that the core thesis of their book is right: a lot of human behavior is this sort of subtle status seeking. I don’t want to go on a tangent here, but I do think it has sociopolitical theory implications. How do we make sure that society doesn’t devolve into just a hedonic treadmill?

The interesting thing about social status is that it’s always, definitionally, a scarce resource because it’s a ranking game. Unlike physical resources, where it’s possible for all of us to live better than kings or queens did 1,000 years ago—we can all have better access to medical care, better access to information, and better access to food—social status is always a zero-sum game, unless someone can conceive of a better way to do it.

This is problematic because if, over time, most of our actions become about pursuing social status, then we’re going to forever be in this sort of game. I don’t think personally that we’re doomed to this. I think there are absolutely better virtues in human psychology, where not everything we do is based on pursuing social status, and I think you can conceive of dynamics where humans are doing things for other reasons, not just to gain status. But I think it’s absolutely fair to say that a surprising amount of human behavior is status seeking, and maybe a depressing amount.

Tim Scarfe

Yeah. I agree with you, and I don’t necessarily want to get too philosophical on that, but that book was The Status Game by Will Storr, where he said that there were 3 meta-status games that we played. He gave the virtue game, the dominance game, and the success game. I might be playing the success game: I want to have the best podcast, or whatever.

The reason I bring this up is that the difference between humans and animals is that they’re just playing one game. It’s really interesting that they have this mimetic social score, but the game is the same everywhere, whereas for us, we go 1 level of abstraction up. The success game for us can be manifested in a myriad of different ways. It could be success at playing computer games; it could be writing books or making podcasts, or whatever.

It’s almost like we fractionate our social ranking into a myriad of different games. I think that’s a little bit of a testament to the difference we humans have in general with our metacognition, which is our ability to create the memetics in a novel way.

Max Bennett

One thing that’s sort of related to this, at least in early human societies, is one way to reduce status infighting: make it such that members of a team have distinct roles. I don’t think this lesson only comes from management theory and entrepreneurship. I think this probably derives from either early human or maybe even early primate societies.

It’s much more stable, and you can introspect that it feels much more comforting to be part of a troop of 100 humans where pretty much everyone is pulling their weight and everyone matters because they’re doing their own distinct thing. That is a very stable state, where we’re not infighting as much because we’re all doing something; we all matter to some degree.

But when there’s infighting because there are only 5 blacksmiths, or 5 podcasts, or 5 books about the evolution of the brain, all of a sudden these other types of things start emerging because we’re no longer all fulfilling a role that matters. It feels like there’s only a ranking, and only 1 of these is going to matter.

As an example of ways to reduce this sort of status seeking, I see this in business all the time: the more you can create an environment where it’s not zero-sum, where everyone’s pulling their weight and together we all win, the more the best versions of humans emerge. The more zero-sum it becomes, and the less distinct the roles and activities are, the more of these—I would argue—primitive primate behaviors start emerging.

An early-stage company, a company of 30 people, has such different dynamics from my last company, where, when I left, we were 400 people. The social dynamics are so different, and I do think one could speculatively correlate that to our evolutionary history here. In a 30-person company, you don’t need that much structure. If you have people who work well together, are aligned on a mission, and you get rid of people who are generally mean-spirited or have bad intent, you don’t need a lot of structure and process to get people to work well together, support each other, and move in a common direction.

I think what that demonstrates, when one observes that, is that what’s playing out is an evolutionary program that got groups of 30 humans to work really well together. When you’re at 400 people, what very quickly happens—and it takes a lot of work to fight this—is that you start getting internal factions emerging, because what splinters out is these subgroups of 30 to 100 people that then have their own points of view. Then it’s very easy to have an us-versus-them dynamic with other groups, and you start seeing things break down.

One mechanism for solving this is very rigid hierarchies; that’s what the military does. Another mechanism is to embrace the chaos, which is a little bit more what Google does. Another mechanism is to effectively make it a constellation of different startups, which is what Amazon does, where each group is kind of autonomous and has very clean interfaces with other groups.

There are many different management approaches to this, but the breakdown, I do think, emerges from the fact that humans did not evolve to interact with 100,000 people. We evolved in an environment where we interacted naturally with about 100 people, and that’s why that comes very naturally. We don’t need as much process to make that work, but we do when we scale it up.

Tim Scarfe

Yeah. It’s fascinating. I mean, as you say, you could argue that Amazon has 1 overarching goal: to make money. But as soon as you increase the autonomy in the organization, it’s a very human trait, isn’t it? You were talking about the Machiavellian behaviors and the deceptive behaviors, and you just wonder how much energy is wasted with infighting.

I’ve even made the comment that in the military, they might be doing quite simplistic jobs compared with Google, but even at Google, there’s an obsession with job level. I mean, if you go off of Levels.fyi or if you go on Blind, that’s the only thing people talk about: their total compensation and job level.

Maybe we should save the cynicism. Coming back to the chapter, we’re telling the story, basically, of how this metacognition and this predictive apparatus gave rise to an entire suite of complex social behaviors that we see in primates, which is fascinating.

Maybe we should just talk a little bit about what I call why bootstrapping. There was 1 guy at Toyota Research who was quite famous because he would get people to ask why 5 times. You say, “Why? Why? Why?” It’s almost as if there’s some magic number. Everyone is only a certain number of degrees of separation away, and it’s a similar thing: you only need to ask why a few times and you’ll always get to some kind of base reason.

Maybe that’s why, evolutionarily, we have 2 levels of causal metacognition in our brain. We have the agranular prefrontal cortex and we have the granular prefrontal cortex. I guess 1 potential question there is: why is there not a 3rd level of asking why, and what would that look like? Can you just sketch out that metacognition picture in general?

Max Bennett

When we think about what a granular prefrontal cortex does, a reasonable framework for it is that it generates explanations of an animal’s behavior. It models an animal’s own behavior. One cognitive tool to reason about that is this: if it observes a rat wake up, have certain hypothalamic activations, and run in a certain direction to drink water, it produces a representation that could be interpreted as, “I am explaining this behavior by: I am hungry, as an animal.”

That can be useful in a variety of ways. It can trigger simulations to find alternative solutions to satiate the same need. If you put a rat in a novel situation, but the granular prefrontal cortex infers that right now I am hungry, we can start triggering a bunch of simulations to try to satiate the same desire, to fulfill what I believe about myself through alternative means. This enables an animal to be flexible.

This is the explanation of an animal itself. Why would that be the case? What I argue in the book is that the granular prefrontal cortex builds a model of that model. Instead of a simulation, it’s a simulation of the simulation.

What that would mean is, if you could—as a thought experiment—ask, let’s go 1 step further, or 1 step back. If we could ask the basal ganglia, which is the sort of reinforcement-learning system, “Why did you turn left to go in this direction to drink water?” it would just say, “Because turning left maximizes reward.”

The answer would always be the same. If you ask the agranular prefrontal cortex, “Why did you turn left?” it would say, “Oh, because I’m thirsty. There’s a specific thing that I, as an entity—this animal I’m modeling—want to achieve.” But if you ask the granular prefrontal cortex, it would say, “Well, I turned left because I am thirsty, and that made me think about ways to satiate my thirst. I simulated going to the left, and I remembered water being there because last time I was there, there was water, and so I went to the left.”

And so, in other words, it enables you to simulate different types of simulations and reason about what you would think in a new setting, which, of course, enables you to think about what someone else might be thinking. We do this all the time. Someone doesn’t respond to a text message, someone makes an odd facial expression in a social interaction, and we’re immediately trying to figure out: What is this person thinking? Why would they do this? And so on.

So the first question is: Why do we even need this new level at all? I think one of the main adaptive values is that it enables your survival in the politicking arms race, because now, if I can simulate a simulation, I can infer why you might do a certain behavior, how to manipulate someone’s knowledge, and your intentions behind things. So this is why you would have one layer to go a level above.

You could make an argument that theoretically there should be an infinite scaling up of whys. I think this is maybe a cop-out, but I think there are huge energetic costs to any sort of scale-up. So what that means is, the question is not whether there would be benefits to a third level of hierarchy; the question is whether the benefits of a third level of hierarchy would outweigh the massive energetic costs of producing it.

And so I think that would be my first-blush explanation as to why we might only have 2 levels instead of 3 or 4: because the second level added a clear adaptive value relative to the cost to survive in the politicking arms race, and the third one perhaps was superfluous and unnecessary relative to the energetic cost.

Tim Scarfe

I think having that second level of metacognition does a lot of work, right? And I’m going to talk a little bit about that now. But one of the things is, you can infer the intents and knowledge of others through the same process of doing simulations yourself. So you can kind of imagine yourself doing something, but you can kind of swap out the pointer to be someone else and swap out the knowledge to be someone else. And that’s incredibly valuable.

But the knowledge thing is really, really interesting. So I asked the question last time, and this is something that I’ve been quite confused about, and I feel that reading this chapter has actually really cleared it up for me, which is about goals. Because when you look one level down at the agranular prefrontal cortex, it’s modeling intents, and then this granular prefrontal cortex, which is trying to seek explanations about the level below, which is the agranular prefrontal cortex, is going a level of abstraction up and modeling goals—not intents—but it’s actually modeling knowledge. What it’s doing is categorizing.

So when you have a simulation of a simulation of simulations, what it’s doing is creating a category. So, to the example you just gave before, “thirsty” becomes a category. Rather than it being a pointillistic intent, it’s a little bit like saying, “I can go and have a sandwich, or I can go and have McDonald’s,” or my abstract simulator could kind of draw a boundary around those things, and now I’m getting food.

As well as being able to categorize intents in yourself and other people, you’re also categorizing knowledge, and then it can be shared memetically. So it’s almost like just going to that second level of metacognition gives you so much that you didn’t have before.

Max Bennett

100%. Yeah. I think thinking about the level of the granular prefrontal cortex and the new primate regions as enabling something akin to knowledge is a really wonderful way to look at it, especially, one, from the connectivity analysis, and then, two, just from what we mean by knowledge.

From the connectivity analysis, if you look at the superior temporal sulcus and the temporoparietal junction, these are regions of the posterior cortex that, in simple terms, are at the very top of the hierarchy. I mean, they get multimodal input from all the other regions of sensory cortex. So a very simple rule of thumb for understanding this is: this models the rest of sensory cortex. I understand the full rendering of the simulation of the external world that is happening, and this is where I build a representation of that.

And it is perhaps no coincidence that’s also where we see brain regions light up when you’re engaging in things like theory of mind and solving false-belief tests—in other words, trying to infer the knowledge of someone else. These same regions light up.

And what do we mean by knowledge? I would argue that knowledge can mean a few things. One is procedural knowledge, where I just know how to do certain motor behaviors. I don’t think that’s what we mean. I think we mean more semantic knowledge or episodic knowledge, which would be: I know that water is over there, and I know that if I do this behavior, this will be the causal outcome.

That type of knowledge, I think, is absolutely rendered in the mental simulation. When I imagine certain things—when I imagine the case of lightning hitting the ground—what do I see afterwards? I see fire. And that’s the source of my knowledge about the causal relationship between these 2 things. So having a layer that models the simulation enables me to reason about my own knowledge and to see what the effect of changing knowledge is on behavior. And this, of course, enables us to flexibly adapt to other people’s behavior and predict what they would do under cases of different knowledge and different intents.

Tim Scarfe

Yeah, that’s fascinating. But you did say that there was a bit of a riddle about the granular prefrontal cortex, because there was 1 study where it could be damaged and the person would still score really highly on IQ tests. But you said it’s about being able to project yourself in simulations, this kind of abstract modeling of your own mind. So in this particular case, how could the person still score the same IQ without that part of the brain?

Max Bennett

So this is such a cool story in the history of neuroscience. You would think that if you look at a human brain, I mean, the granular prefrontal cortex is this huge region in the front of the brain. I mean, it takes up a gargantuan amount of space. You would think that taking a chunk out of that part of the brain would have a gross effect on a human being.

Just like if you took a part out of even a relatively small region of the back of your brain, which is where your visual cortex is, you become hugely visually impaired. You take a region out of your motor cortex and humans become largely paralyzed for months until they recover from that. You take a region out of auditory cortex and they can lose the ability to recognize even words.

So there are relatively small regions of neocortex that, if there’s damage to them, have gross, obvious effects on human intelligence and behavior. After World War II, there were so many patients with brain damage that there were all these studies, and people could not figure it out. It was a puzzle: What does this huge region of prefrontal cortex do? People don’t have—something seems off about them, but it’s not obvious what is wrong with them. People would note personality changes. They don’t seem to be themselves. But on logic tests, on IQ tests, it wasn’t obvious they were dramatically impaired.

In many cases, it wasn’t obvious they were impaired. There was 1 famous case where they could test someone before and after because, for surgical reasons, they were going to remove parts of the granular cortex. This patient actually improved on IQ tests, which made this a huge puzzle: What does this part of the brain do?

And so, if you track the studies from that point forward, we start learning that what granular prefrontal cortex does in large part isn’t related to these types of logic puzzles. It’s related to thinking about thinking and modeling ourselves.

So, for example, if you look at someone who has damage to granular prefrontal cortex, someone who has damage to the hippocampus, and someone who has a normal brain, and you ask them something very simple—you give them a random word and you say, “Just tell me a story. Just imagine a story of you with this word.” The word could be “restaurant,” and you compare these stories, you immediately see something very different.

The people with hippocampal damage give a very, very rich story about themselves, but the external world misses details. So there’s not a lot of rich detail about the external world. This is consistent with the idea we talked about with early mammals, where the hippocampus helps render a simulated external world.

The people with granular prefrontal damage could render a very rich external world. They could tell you the details of the leaves, the smell of food, exactly what a restaurant looked like, but they themselves were woefully missing from the stories. They could not project themselves into this imagined world.

And so then, if we go back and look at all these other things that light up granular prefrontal cortex, if you ask someone to think about how they’re feeling, granular prefrontal cortex lights up—self-reference. But if you ask another question, such as, “What does it look like outside?” the granular prefrontal cortex does not activate. The agranular prefrontal cortex will activate in both cases.

So we start to see that it’s in cases of thinking about yourself and thinking about others that this granular region gets very activated. And now, if you go back and study these people more deeply, you notice that they become hugely impaired at false-belief tests. They can’t recognize faux pas, so they don’t understand what’s not really appropriate. Which, of course, makes sense, because how do I know what’s appropriate? I’m going to infer how you feel about the things that I’m saying. And so you see all of these mentalizing impairments that emerge, but it’s not related directly to these logic puzzles that are typically in things like IQ tests.

Tim Scarfe

Yeah. You mentioned the false-belief test. Can you just briefly sketch out what that is?

Max Bennett

Yes. There’s a good picture if you want to hold it up or show it on the podcast. The way the test works is you have Sally on the left, who has a basket, and then you have Anne on the right, who has a box. So Sally puts a marble in the basket, and then she walks away. Then Anne goes over and moves the marble from Sally’s basket and puts it into her box, and then leaves. When Sally comes back, where does she look for the marble?

It’s so simple. But in order to figure out that Sally will look into the box, you have to understand that it’s possible for another mind to have incorrect knowledge—to have a false belief about something. Young children don’t understand this. They assume that knowledge is omnipotent: everyone has the same knowledge about the world. But at a certain point, they start learning that it’s possible for people to have false beliefs.

So we actually know that nonhuman primates can do this. They’ve done studies on macaques where you do exactly the same Sally test, and you just look at where their eyes look when the person comes back into the room to look between the 2 boxes. They always look, or tend to look, in the direction of where that person thinks the marble is, or the piece of food is, not where it actually is. If you inhibit their granular prefrontal cortex through an injection or another mechanism, this bias goes away. They no longer look in the right direction. There’s lots of really good evidence that this sort of false-belief mechanism is occurring in these primate regions.

Tim Scarfe

Yeah. And what really hit home to me is that, in a way, it’s not even knowledge. It’s all simulations. It’s just simulations of other agents. We’ve always spoken about knowledge in some weird Platonic, abstract sense. I quite like the idea that the primitive form of communication between humans is just simulations, even when we’re speaking to each other.

Yeah, exactly. So how do we solve the Sally problem? It probably happens so quickly, but we just simulate what we would think if we were in Sally’s shoes. And then I realize, well, I would look in this place. And this helps us reason about other people. And this begs a really, almost profound question: How unique is theory of mind?

This brings me to a question that I’ve been asked multiple times, which is: Does ChatGPT have theory of mind? The evidence I should stipulate, for anyone curious, is that if you ask GPT-3 these sorts of theory-of-mind puzzles, it does terribly. So that’s an easy one to discard out of hand. But if you ask GPT-4 these theory-of-mind puzzles, it performs remarkably accurately, at a human level, on these theory-of-mind puzzles. And there have been people who have explored whether it’s just in the training data, and there’s good evidence that it’s not just because they’re regurgitating what was in the training data.

So does that mean that ChatGPT has theory of mind? I think there are a few ways to reason about this. One is: What do we mean by theory of mind? If by theory of mind we just mean the ability to solve these sorts of false-belief puzzles, then I think you have to accept the fact that, yes, it can solve those tasks. The problem is, the way in which it renders this model of other minds is not through having a similar mind itself. And so what this means is we should be concerned—it doesn’t mean it won’t work well—but we should be concerned about how well this will generalize to real tasks where we might care about this much more deeply.

For example, with a human, there’s good evidence to suggest that part of my ability to reason about your mind is because I have a mind that works quite similarly. We are almost bound together by some common mechanistic synergy between the way in which our brains work, because our brains are quite similar, which enables a lot of data efficiency. I’m pretty good at predicting what people do—not perfect, but pretty good at predicting what people do—because we’re all people, and there are similarities between how we act. That makes us quite data-efficient and decent at generalizing to new situations where we put people in new places that we’ve never seen before. I can kind of guess, well, if I were in that situation, this is what I would do.

GPT-4 has learned to build a theory of mind simply by reading the text of these puzzles, and so clearly it has some mechanism to build a model for predicting what people will do in certain circumstances and differentiating knowledge and intent, et cetera. But the concern is twofold. One, what will happen if we take those types of models and put them in very new situations that are not based on just these puzzles, but, for example, we’re asking them to optimize a paperclip factory? That’s a situation where we should be concerned. How well will it do at actually inferring what we mean by what we say?

And the second is data efficiency, which is: How much data did it have to see to build this model? If it was a ton of data, then it’s going to be problematic if we have these new situations where we want to teach it to model people’s behaviors in this new place. If it requires a ridiculous amount of data, then it’s always going to be slow to learn these things and always be at risk of not generalizing well when we put it in these new situations.

My answer here is nuanced. I think if by theory of mind we mean solving puzzle questions, it’s very hard to say that ChatGPT does not have some model of human behavior. But I do think the human and primate mechanism for doing so has a data-efficiency advantage and a mechanistic-synergy advantage. In other words, we can use ourselves to reason about things, and that is relevant. If we want to have these systems do a good job listening to human requests, we shouldn’t translate performance on false-belief tests into believing that they’ll do a good job correctly inferring our intent and knowledge in new situations.

Tim Scarfe

Yeah, I would agree with that. I think ChatGPT is in the world of text, and it's learned all of this structured narratology and things on Reddit and things on Twitter. And as we were saying last week, language has evolved to be very simple. It has to be learnable by children. It has a small subspace. But it is a real kind of generalization over human behaviors, and it's in this very low-resolution substrate. Whereas in the Machiavellian apes example that we were talking about before, these are agents performing real-time sensing and inferencing and making in-the-moment judgments, and they're in this continuous sensor domain where they have many different types of signals: visual signals, sound signals, and also memory of what happened in those dynamics just before. So it feels like a difference in kind to me between those two situations. But it is remarkable that in the GPT domain any kind of theory of mind could work.

Max Bennett

One good example of this, I think, is whether there is a difference in our human ability to predict behavior between a car and a person. So the brain is always able to model things it observes, simulate them, and predict what they will do. I can look at a car and imagine different colors of it, and I can imagine what will happen if I drop it and it rolls down a hill. We build models of things all the time. We build models of computers and models of books. So the brain produces models of things. Is the way that the brain produces models of other human behaviors exactly the same, or is there some unique advantage? My argument is that there's something unique happening when I'm building a model of another person, which is that I'm leveraging my own inner simulation of things as a useful prior to try to predict what other people will do. ChatGPT models human behaviors, to draw a crude analogy, the way we would model a random object: I'm only modeling it based on seeing its behaviors in certain situations with the data I receive. On the other hand, when we model someone else's behavior, we're doing some form of projection and using the prior of how we would behave, and we probably bootstrap part of our model of human knowledge and intent based on our own introspection. I think in that way it is a difference in kind.

Tim Scarfe

Fascinating. I completely agree with that. The Selfish Gene is kind of saying it doesn’t matter what you folks do. The gene is directing your behavior, and you don’t really have as much agency as you think you do. And it’s a similar thing with language. If you think of language as being a superorganism or a virus, and we are the hosts, information is being shared memetically, and it’s shaping our evolution, but it’s also shaping our behavior. So it’s almost like when we become infected by certain memes—it might be religion, for example—it’s almost like it parasitically affects our behavior.

Tim Scarfe

But I think there is a difference between social memes and physical memes. Tool use, for example, doesn't seem to have the same parasitic effect. If you look at the behavioral complexity of apes, because they don't have these novel virtual memes in their culture, their behavior seems quite monolithic compared to ours. But I wondered if you could contrast that next level of mimesis.

Max Bennett

There's been lots of great writing about the distinction in the literature, which is typically called cultural transmission, between nonhuman primates and humans. A lot of the general consensus here is that, although there is transmission among nonhuman primates, which we see in particular with tool use, it doesn't accumulate in the same way that it does with humans. In other words, humans can pass a piece of information to another generation, which that next generation will reliably copy and then merge with other new information, which they can then reliably copy. You do this over 1,000 years and you go from, “I know how to whittle a bone into a needle for sewing,” to, all of a sudden, “Now I've built a loom,” right? These ideas keep accumulating on top of each other.

Whereas in nonhuman primates, you don't see the same type of accumulation. That's what I, in a pithy way, in the book call “the singularity that already happened”: once you enable these memes or ideas to accumulate across generations, you get what you're describing as this sort of mimetic organism that we are the substrate for. For sure, what I think is interesting here is one lens through which I like to think about this: sources of learning.

If you think about how nonhuman primates learn, there are sort of 3 sources. One is that they learn from direct experience, their own actual actions. This is reinforcement learning writ large: I do something, it succeeds, it fails. Fine. Another is their own imagined actions. This is the part that evolved in early mammals. I can imagine doing 5 different strategies to try and get to the food over there, and I find the one that worked. That's a source of learning: my model of the world became a source of figuring out the right path.

What mentalizing enabled with primates is this third mechanism: learning from other people's actual actions. So I can see my mother—if I'm a young chimpanzee—using a stick to put into this termite mound, pull it out, and eat food. I don't have to do my own behaviors to do that. I don't even have to simulate doing it. By watching her do that skill, I will adopt and learn.

But what nonhuman primates don't have, which is very uniquely human, is learning from other people's imagined actions. This is the key breakthrough that happens with language: the bandwidth through which nonhuman primates can communicate what we're calling knowledge here is only through actions themselves. I can't describe, if I were a nonhuman primate, what I saw when I imagined 5 different ways to try and hunt the boar over there. I can just do it, and you can learn from what you saw me do. But language enables us to share the outcomes of our imaginations.

That is a much higher-bandwidth mechanism for translating information. That enables accumulation across generations. For example, it's so easy to think about ways in which this would be adaptive. Two would be sharing semantics: I go into the forest, and there are 2 snakes there. One bites me and I'm fine; the other bites me and I get really sick. I come back and say, “Green snakes are okay. Red snakes, don't go near them.” That semantic knowledge now exists among the whole troop.

In the old world, before there was language, only the people surrounding the event who saw it happen would have the knowledge. Now I can translate it: I simulated the episodic memory in my mind, and I translate it to everyone. The other is coordinated planning. Before language, it would not be possible for 5 humans to jump in trees and say, “Okay, here's how we're going to hunt these boar. We're going to stay silent, and then I'm going to whistle 3 times, and then we're all going to jump down and surround the one in the back.”

That type of planning is only possible because one person can simulate something and then translate it and say, “Hey, when I imagine this happening, we succeed,” and other people, of course, can edit that simulation and say, “When I imagine that happening, I don't see us succeeding for this reason.” You can start refining it. This ability to have a source of learning from other people's mental simulations is what I would argue is the source of this very unique human superpower that emerges from language.

Of course, now, with such a high-bandwidth transference of mental simulations, you do get this sort of quasi-evolutionary process, which is what Richard Dawkins is talking about. You have a process by which the memes—the ideas that do a good job propagating—are the ones that will propagate. The ones that, for whatever reasons, are not viral either don't do a good job of maintaining the host, so the ideas are bad and I end up dying, or I just don't have an incentive to share them. They're not viral; those ideas die, and so then you get this sort of meme evolution. But to me, the source is the fact that language enables us to share in our simulations, which becomes a much higher-bandwidth communication mechanism.

Tim Scarfe

I'm fascinated by this idea of the meme itself being an agent, being a virtual agent, and, in expressing its agency, it needs to manipulate us. You might argue, as you do in your book, that there has to be some kind of traceable chain down to the basal ganglia. So we have many levels of bootstrapping, and at some point the thing exists because the basal ganglia says, “Oh, that's good. I like that.”

So then we have one level, then another level, then we have the memeosphere. It's almost as if that thing is manipulating us down here, but doing something completely different up there. When you have weakly emergent macroscopic phenomena, part of the definition of emergence is surprise. It's macroscopically surprising. It does something completely unexpected and unlike the thing that went below it. It's just weird that it might be manipulating us down here, but doing something completely different up there.

Max Bennett

Yeah. I think there's probably—this is mostly fun speculation—but if I'm going to draw analogies to brain regions and intellectual features of the human brain, there are probably 2 lightweight ways we could think about why memes become attractive.

One would be the older vertebrate-like structures. This would be the basal ganglia plus the amygdala. A meme that makes me feel fear, or makes me think that unless I take an action something bad is going to happen to me, and one of those actions has to be sharing it, is going to be highly viral. If you make me afraid for my family's well-being, you're going to activate my amygdala, and even if there's only a 2% chance that this is true, I might still share it. So you get these sorts of effects. Humans are not good at dealing with low-probability, high-magnitude events, which is another brain constraint.

The other key thing that also exists at the level of early vertebrate-like structures is a preference for surprise. In order for reinforcement learning to work well, it's very effective—and we see this in AI systems—to make people pursue actions that are novel, because that's one way in which we can explore new areas and learn new actions, and explore the space of possible choices to make. This is one intrinsic way to get trial and error to work.

The way casinos make money from you is that they hack into this sort of preference for surprise. If there's a 0% chance of winning, you would never play. But if there's a 48% chance—a net 48% chance, meaning in the long run you'll lose money—but every once in a while there's some surprising thing that is actually over the threshold of being worth it to the basal ganglia because the surprise is so exciting, that's one way to think about it. So if you get something that creates innate fear, or some great outcome or surprise, you get these older structures.

With mammalian structures, I think there is sort of an active-inference play here, and you could even correlate it to the granular prefrontal cortex, with things like identity. If you give me some information that's consistent with my model of myself, and I'm highly motivated to maintain my model of myself for a variety of reasons that we can talk about, then I'm more likely to maintain this belief. Versus if you give me information that's inconsistent with my model of who I am, people are highly likely to reject these beliefs.

This is another speculative way to think about why memes persist within their little echo chambers and how it can quickly become sort of identity wars. If one's identity is consistent with a certain set of beliefs, then that almost creates a gated wall for certain types of memes to enter. It becomes much harder for certain ideas to enter my mind, and it creates a very porous filter for other types of ideas that are consistent with my identity. Those become very easy for me to adopt.

I think that's another way—if we're going to frame memes as having agency—that a meme would seek to survive: find a way to be consistent with certain people's view of themselves and the world. What you're doing is reinforcing it as opposed to challenging it.

There is a very clear difference, though: the human brain is analog, while these machine brains are digital, and there are pros and cons to each. Geoffrey Hinton talks about this—to make sure I’m citing these cool ideas correctly—and he has a great talk in which he describes it very well. The benefit of a digital brain is that it’s immortal: all the weights are stored in binary, so I can very easily transfer it to different brains, but it’s hugely energy-inefficient because I need to model everything exactly in order for it to be copyable.

The human brain is much more efficient, but it’s not copyable because the information exists in the physical representation of the analog connections between all these neurons: the actual protein receptors that exist in them, the gene expressions, and all this crazy stuff that makes it non-copyable. It’ll be interesting to see what the energy efficiency is, for example, of a digital AI system that actually attempts to recapitulate a human brain. That might be very energy-inefficient.

It might open the door for a whole new area of research that I think would be really fascinating: building analog brains. Can we have systems that actually work in a more analog way? The way they pass information to each other—this is also a Geoffrey Hinton idea—is by teaching each other. Because they are AI systems, they can teach each other with better fidelity than humans can, because they can actually share probabilistic outcomes as opposed to just the words we say.

They can also generate way more samples for each other than a human could because they can live much longer. There’s a whole emerging world around the distinction between these digital machines that are immortal but very energy-inefficient, and analog machines that are much more energy-efficient but less good at translating information.

One of your points that I think is really key is that one of the main things missing—and I don’t think it’s talked about enough—is the continual-learning problem. Maybe there’ll be a breakthrough soon that would be great, but I don’t see very clear ideas over the horizon that will solve this. I would say this is one of the essential lines that differentiates biological brains from modern AI systems.

The way in which AI systems are trained is such that we cannot let them continuously learn from new experiences because it disrupts the old information they have. Whether that is an architectural constraint or something that needs to change in the underlying learning algorithm itself, there’s lots of open research and debate about that. But the fact remains that if you allowed ChatGPT to learn from every chat that happens to it, it would get rapidly dumber.

Tim Scarfe

Yeah.

Max Bennett

That is not the case with humans. We can continuously update our information, and our representations are robust. I think for many of the applications for AI systems that are going to be most impactful, continual learning is going to be an essential component because we’re going to want to bring an AI agent in, show it new information, and immediately have it incorporate that without forgetting old things.

I think that is very clearly a line where there’s a lot of really interesting research happening and a lot of research left to be done.

Tim Scarfe

Why are we superior to animals?

Max Bennett

There’s been such a long history of us pontificating on the various chasms, or attempting to create a chasm intellectually, between us and other animals. The most famous form of this, which I think still shows threads in modernity, is from Aristotle. He took the same kind of ideas that you see in MacLean’s triune brain: other animals might have basic instincts, and they might have some form of emotions, but what they all lack—which humans uniquely have—is this notion of reason. We can uniquely reason about things in the world.

As I try to argue in the book, and as most comparative psychology demonstrates quite clearly, there are clearly forms of reasoning that we see in other animals. Of all the different abilities and capacities that seem unique to humans, the one that stands out as most salient is undeniably language. Despite many painstaking attempts, we have not even been able to teach chimpanzees, bonobos, or gorillas to speak with the same degree of fidelity as human language.

There is some controversy as to the extent to which Kanzi, Koko, and Wu passed the threshold that we define as language, so that can be debated: where do we draw the line? Undeniably, most people would agree that these nonhuman primates do not learn language naturally without painstaking attempts to teach them. When they do learn language, it does not show the same sort of flexibility as human language.

What makes language unique is 2 things. One is declarative labeling. There is a distinction between imperative labels and declarative labels. An imperative label is learning that a phrase, or a cue, leads to a reward if you take an action in response to that cue. When a dog responds to a specific cue and then you give it a treat, that’s not what we define as language; it’s an imperative label.

A declarative label is when I say “dog” and, in your head, you know that it references a concept or a thing. We have a label for a concept or a thing. It’s not at all clear that other species perform these types of declarative labeling. If they do, it surely evolved independently. We’re quite confident that early primates didn’t have this ability.

The second thing that makes language unique is grammar. We can take these declarative labels that reference things or actions, and then we can weave them together in a certain structure, and the structure itself has meaning. A basic example is just the ordering of phrases: if I say, “Ben hugged James,” that means something different from “James hugged Ben.” Despite the fact that they’re the same phrases, or the same declarative labels, the order presents meaning.

There’s a whole interesting world around why language—if it is the case that language is the fundamental difference—has allowed humans to take over the world. That’s another interesting topic we should discuss. I would argue that primarily what makes humans different is language, and Aristotle’s idea of reason we see at least in smaller forms in other animals.

Tim Scarfe

Yeah. It’s quite interesting because you said right at the very beginning that Aristotle spoke about the rational soul that we have. Even in the 20th century, we spoke about things like mental time travel, our sense of self, and tool use, and it’s really interesting because we look in the animal kingdom and, one by one, all of these things that we thought placed a bright line between us and animals faded away.

Some people think that language is a continuum, that there’s just a gradation—that if you scale up the brain of an ape, you will get human language. Is that the case?

Max Bennett

The reason I’m very skeptical of that claim is that we don’t see variance in language abilities based on brain size. Children who learn language at the age of 4 still have relatively small brains. I’m not sure of the exact brain-size comparison, but I’d be curious about the brain size of a 4-year-old child relative to an adult chimpanzee, just based on volume.

The other interesting case is Homo floresiensis.

Tim Scarfe

Oh yeah, from Indonesia—the ones with the small brain.

Max Bennett

Yeah, yeah, yeah, yeah. Yes, Homo floresiensis is a great case study here because we found fossils of ancestral humans on an island in Indonesia who were effectively miniature humans. They had shrunk in size to, I think, around 3½ to 4 feet tall.

Their brain capacity, which we can look at from their fossilized brains, had actually shrunk from that of ancestral humans. They were marginally larger than the size of a modern chimpanzee brain, and yet they showed a lot of signs of superior human intelligence despite having smaller brains. They showed tool use akin to that of ancestral humans.

They had Oldowan tools, which are supposedly a sign of uniquely human, intelligent tool-making. That is suggestive of the idea that whatever unique intellectual capacities humans gained around 2 million to 1 million years ago were present despite their shrinking brains. Either one has to argue that language evolved much earlier, which some people do, or that whatever sort of protolanguage emerged back then was present even when these brains started shrinking.

That suggests to me—and is actually aligned with the ideas in The Language Game, which is a great book—that fundamentally what’s unique is that we have an instinct to learn language. It’s not that we have some unique capacity for language, and I think that is a key difference that we can talk about.

When you look at children who learn language, there are 2 very unique features of how they go through language learning. By the young age of around 2, they’re already engaging in proto-conversations. A younger infant will pause to match the pausing of their mother. Even if they’re just babbling, they will engage in the synchrony of babbling time intervals.

That is clearly a demonstration of some initial instinct, which demonstrates the ability for me to want to engage in some turn-taking action with you.

The other unique thing that emerges a little bit later is joint attention, where human children will uniquely attempt to get their parents to engage in attention toward the same object. Scientists have gone to painstaking efforts to demonstrate that this attempt to get a parent to engage in attention toward an object is not an attempt to get the object. So a child or an infant will be dissatisfied if the parent doesn't look at the object but they get the object—so a third party comes in and hands it to them. They're dissatisfied.

If the parent looks at them and is excited when they're pointing at an object, the child will also not be satisfied. But only when the parent looks at the object, then looks back at the child and smiles, is the child satisfied. So there's this instinct to engage in conversation and to jointly attend to things, which gives us the sort of instinctual foundation on which you can start adding declarative labels. When you have joint attention to something and you're paying attention to this turn-taking, it enables you to label things and say, “Well, this means run, or this means book.”

So I think all of that makes it hard to argue that it's just a consequence of a scaled-up brain. One of the things that's really fascinating is that animal communication seems extremely superficial. And when I say superficial, I mean that when you take different populations of the same species or different species, the expression and the complexity are very, very simple. We don't see this incredible fractionation and divergence that we see in human language.

Tim Scarfe

As you articulated just a minute ago, a big part of that is this declarative labeling. One of the reasons, presumably, for language is the ability to do variable binding on symbols—to say, “This thing is a dog, that thing's a bear”—and to be able to dynamically manipulate that.

It seems to me that you can think of language as a form of agentic communication. The difference between humans as language users is that we are agents, and agency is about being able to have your own directedness, plan many steps ahead, take control of your environment, and so on. The difference in communication with animals is that the information content is more in the environment around them, whereas for human languaging, a lot of it comes from the agent itself. So I just wondered whether you could think of any weird way to distinguish human languaging from animal communication.

Max Bennett

One line that I think there's some good evidence to suggest exists between human communication and nonhuman primate communication is that humans have much more of a desire to share what's going on in our own minds. There is a unique pleasure we have from sharing our thoughts. And when we look at the communication styles that happen in nonhuman primates, there's much less of a desire—even when we go through these language-learning experiments where they have forms of communication—there's much less of a desire to share thoughts that are going on in one's mind.

One line where there's some controversy around this is that humans, from a very young age, will ask questions. They'll inquire as to what's going on in someone else's head. And with the exception of maybe Kanzi, where there was some argument that he maybe asked questions, you did not see nonhuman primates probe the minds of other individuals, even though we know they have theory of mind. We know when they're trying to deceive others or they're trying to learn actions by observation, they clearly engage in theory of mind, but when it comes to language, they weren't interested in inquiring as to what someone is thinking about.

I think in that sense there's an agency to language being a tool for inquiring as to what's going on in someone else's mind and sharing what's going on in your mind. And this is where I think language is part of why language is a superpower, because it provides a completely unique source of data for learning. Nonhuman primates can engage in learning through observation because I can see someone take actions, as you said with imitation learning. I can see you open a puzzle box to get food and I can learn from observing your actual actions.

The way I do that is because I can infer the intent of what you're trying to do, and then I can figure out which of the actions are relevant and which are irrelevant. A monkey, a chimpanzee, or an ape will ignore irrelevant actions when they observe you do a task. They've done these experiments with humans and chimpanzees where they do all these actions to open a puzzle box, including some random actions, and chimpanzees will ignore the random actions, which suggests they can infer the intent of it. Which is great, but chimpanzees don't learn from what's going on in your head.

The ability to learn from other mental simulations is what's so powerful about language. I can say, “I just went over to that forest over there and I saw a red snake and a blue snake, and I saw that the red snake is really dangerous, but the blue snake is not because the blue snake bit me and nothing happened.” And so I share that episodic memory, and now everyone has that knowledge, even though that was just in my mental simulation.

Or, when planning a hunt, a group of 5 humans—I can imagine a strategy of how all 5 of us are going to coordinate, see it succeed in my mind, and then share the results and the plan with everyone. So language enables us to tether our mental simulations to each other. And I think there's a sense of agency in the idea that there's a purpose to that. There's a volitional purpose to the communication.

The neurological underpinnings of communication that occurs in nonhuman primates are more analogous to our emotional expressions than they are to language. And we see this also in the brain. Monkeys and nonhuman apes have these innate expressions that they do, which are genetically hard-coded, and we know that because it's the same even across species, often, that have never interacted with each other. It comes from neurological structures similar to our laughing and crying.

So this is clearly a hard-coded emotional expression. In that sense, it doesn't have the same volition, because I'm not doing this action to communicate a concept to you. I'm doing this action as an innate response to a cue or a feeling I have. So, yeah, I think there's some meat to that idea.

Tim Scarfe

Well, a few things to explore there. First of all, we should just talk about how we became a collective intelligence after the fact. I'm not sure whether that's unique to humans if you look at other forms of collective intelligence. There's always a kind of juxtaposition between the intelligence of the individual versus the collective, and usually you find that having very intelligent individuals is not good for the intelligence of the collective.

But what's interesting about humans is that we clearly didn't evolve as a collective intelligence. We had this kind of bootstrapping process where we were very, very useful, independent agents, and then this collective intelligence just emerged out of nowhere. Is that an interesting observation?

Max Bennett

There are different degrees of collectiveness, and so I think we can draw distinctions between different flavors of collectiveness, but I don't think humans are uniquely collective. For example, the imitation learning of nonhuman primates is a form of collective intelligence because you can teach one member of a chimpanzee troop how to use a tool, and then over time the rest of the troop will learn just through observation. So that's a sense of collective intelligence.

Many vertebrates, and likely the first vertebrates—you can even see fish—will learn through observation. In other words, when a fish swims in a certain direction to get food, other fish can see that fish do that and follow them. There can be an instinct to follow others around you. So I think there are flavors of collectiveness that exist across many different species.

But what's unique about the collectiveness in humans is that the fidelity with which we transfer our mental simulations enables them to accumulate across generations. In that sense, it almost has its own agency, or is its own thing, because it can actually go through its own process of evolution as ideas propagate through generations of people. That's not the same thing that you see in other animals.

Tim Scarfe

Yeah. A couple of things on that. First of all, I would quite like to distinguish knowledge and intelligence. Collective intelligence—and intelligence in general—is a process of discovering models. To get the language down here, I will use models, skills, and knowledge pretty much interchangeably. I think of an intelligent process as epistemic foraging: finding interesting models and then having them discovered and shared by other people.

It's a little bit like when you distribute a GPU workload: you can do model parallelism and you can do data parallelism. You can either split up the processing, or you can split up the actual representation. So I think the kind of collective intelligence that you've just been speaking about is that we've got all of these independent agents, and they are finding models and sharing models; the models get refined over time, and it adapts, and so on.

But I also think a big, important element is sharing the computation. So even though there's some redundant work going on, epistemic areas over here are being explored, but also, in many cases, the same problems are being explored in slightly different variations. So we're sharing the workload with other humans.

Max Bennett

Yes, I think that totally makes sense. There are some interesting ideas in AI here, actually, where there's this concept of knowledge distillation. In AI, one way in which you can have model A teach model B the things that model A knows is to wholesale copy the parameters of model A. Of course, that's totally biologically implausible. There are aspects of parameter copying—the components of our brain that are genetically hard-coded are a version of parameter copying—but for other applications, it's not feasible or maybe not desirable to just copy parameters.

Knowledge distillation is saying, okay, well, we can have a set of data that we give model A, and we either look at the outputs of model A or the layer before the outputs, so we can see more richness in its representation of the input you give it. Then we take that data—that almost-labeled data—to model B and train model B on it. So that's distilling some of the knowledge, through almost training model B to try and act similarly to model A.

That type of information transfer, I think, does occur in nonhuman primates, and that's imitation learning. However, it is not nearly as rich, because what happens in nonhuman primates is primarily grounded in just the actions that I'm taking. That's much less rich than what I can share: not only the data of what you see me actually do, but also things that happen only in my mind. That opens the door for much more transference of—and the word you used—the computations that I'm performing.

Tim Scarfe

So I think it's absolutely true. Before, we were learning in the physical world, so we were learning from physical things that we were directly observing, and now we are learning from imagined actions. But there's a bit of a latent component to language as well. For example, someone might come up to me and say, “Oh, the blue swirly thing is over there.” And I'll say, “Well, I don't know what you mean by the blue swirly thing, because I've never seen one before.”

There's this inference process. This is where it starts to get really interesting, because there's a diffusion, right? There's a message passing that happens between all of the different agents, and it's filling in missing information. Even though many of the agents wouldn't have seen anything like what we're talking about, sometimes it can be filled in with subsequent interactions with people, and sometimes it can just become a kind of latent category that can be filled in later. So there's this real diffusion process going on, which I think is quite difficult to articulate.

Max Bennett

Part of what's so interesting about language is that it's still an area of such controversy amongst cognitive psychologists, linguists, and even AI people. So much is still unsettled about it, and there are still debates today. There are debates today about whether language is primarily a tool for thinking or communication. Chomsky is the most famous proponent of the idea of language for thinking.

He has evolutionary arguments that language initially evolved not as a tool for communication, but for our own process of thinking, and then later was adopted or used for communication. That's a minority view. Other people argue—and I'm more amenable to this—that language was primarily used as a tool for communicating.

These ideas are actually reemerging with language models, because the way language models learn about the world, in some sense, is that language becomes the reasoning tool itself, which is more Chomsky-like. Even though I think the success of language models, in a lot of ways, discredits many of Chomsky's ideas, we can talk about that. Interestingly, the fact that we're using language as the fundamental mechanism for reasoning and thinking is actually somewhat Chomsky-like, versus the idea that language is communication.

The idea is that language is a condensed set of tokens that I'm passing between minds, but the real communication I'm trying to share with you is what's going on in my mind. In other words, it's the mental simulation—the more mammalian component here. The rendered 3D world is what I'm trying to transfer to you, and I condense it into this code that you reverse-engineer back into a mental simulation.

Theory of mind—one reason why language might be so rare in the animal kingdom is that mentalizing, or theory of mind, which is relatively rare in the animal kingdom, is a prerequisite. In order for me to reverse-engineer the language code you've provided me, I need to be able to infer what you might have meant by what you're saying, reason about why you would have said this, and understand what knowledge you have.

So I think language is intended to cue another person to render something in their mind. This is also where teaching is so important and such a key aspect of language learning, because we can infer what declarative labels this person is aware of. When they're confused, you have to start trying to iterate to understand what they're confused about in what you're saying so that you can disambiguate it for them.

There's also a disambiguation process where you ask follow-up questions when you feel like you don't fully understand what's going on in someone else's head.

Tim Scarfe

Yeah. I mean, the guardrails thing is interesting because they're not necessarily thinking guardrails; they're also pragmatic guardrails. And there's a really interesting figure in the book, actually. Yeah, here it is. It talks about how language is sharing information over generations.

Without language, we learn a little bit inside a generation, then it goes pretty much back to zero again. But now we have the ability to pass on these memetic bits of information over several generations. The thing is, there's a real structure to it. I think of it as a bit like a directed acyclic graph. So it's a tree structure, and every single bit of knowledge that we discover kind of stands on the shoulders of giants. It needs all of the things that we discovered beforehand.

In a sense, we're all of these little agents, and we're doing this epistemic foraging. We're finding new skill programs, we're sharing them, and so on. But it's almost like we shouldn't think of the mass as being like an entire convex hull. It's only on the boundary where all of the creativity and all of the information sharing happens—on the surface of this object that's being created.

What I mean by that is, now in modern cities, for example, you can't live without a driver's license. You can't live without the internet. You need to do things a certain way, and even though it's not technically constraining our brains and how we think, we live in a very, very constrained and weird world now.

Max Bennett

Yeah, totally. Great point. There's a biological constraint as to how much knowledge a given human brain can contain. One lens through which to see the last 100,000 years, especially the last 100 years, is us finding solutions to getting past the biological constraint of human brains.

Language was one tool, because it used to be the case that all the information that a given entity learned needed to be learned by my brain within my lifetime. Language enables us, as a group, to have shared knowledge, but not every brain contains all of the knowledge. If you think about a troop of 100 people, it's possible for those 100 people and all their descendants for 1,000 years to have tons of skills, despite the fact that no one brain ever had all of the skills.

Someone becomes really good at hunting, someone becomes really good at weaving animal skins into clothing, and all of these types of skills. There are actually cases in anthropology of groups of humans that get separated from each other and their technology degrades, because there is a limit—a minimum number of brains needed to contain and store a certain amount of information in the absence of writing.

Language was maybe innovation 1 here. Writing was another innovation, which is great. Now we can more reliably transfer these ideas across generations, even if there are gaps. In other words, even if there's a period of time for maybe 2 generations when no brains contain it, a third generation can go back to the writing and pick up that knowledge.

Of course, now with the internet, we've just scaled up writing even more. But you're absolutely right. Sometimes I think about this as, if a group of 20 friends and I ended up on an island—if we were the only 20 humans left—not that I think about this all the time, but it is crazy how little of human knowledge would be contained in our 20 brains.

How dramatically we would degrade, essentially. We've got this thing where we've got all of these different brains, and individual people can have about 150 friends or something. It's the social Dunbar limit.

Tim Scarfe

But as you say, because we have this ability to share simulations and we have common myths and so on, we can address a much larger carrying capacity of people and knowledge. You said something really interesting in the book, which is four things: bigger brains, specialization, more brains, bigger population size, and writing and sharing simulations through the internet and all of these things.

So we've increased our carrying capacity, and now something very interesting and arbitrary has emerged. We've got all of these different specializations of skills, and I guess the question is: where does it end? Has it converged? Could we carry much more knowledge than we already have, or would we have to wait for a top-down kind of genetic pressure for our brains to get a bit bigger?

Max Bennett

I think we are about to go through this. Google and the internet have turned us all into epistemic hybrids. Google has become a shared knowledge store that we all use, and of course there are problems, because now there are subareas on the internet where we can use different knowledge stores. We live in these different epistemic bubbles, and that creates political problems as well.

But we have already become hybrids where we use technology to overcome limitations in our own brains. Writing is a tool to overcome challenges in memory and, at times, thinking. The internet has become a tool to answer any question at a whim, and some people have concerns with this because it can also atrophy parts of our brain that maybe we want.

For example, through mere introspection, I will say that once I started using Google Maps as a kid, the part of my brain that was learning how to navigate a city—by actually remembering the grid and map of a city—just started atrophying. Now I have no capacity to do that, whereas my dad—you take him to any new city and you can see him rendering a map of the city in his mind—won't use Google Maps.

One could argue that it doesn't matter because I'll always have Google Maps, so why do I need this skill? Another argument would be that atrophying may have other consequences in my life, and it would be important for me to go through the cognitive exercise even though technology enables me to do it. We make these trade-offs at different times. Why do we teach kids arithmetic? They can always just use a calculator, but we deem it important for us to go through the process of understanding arithmetic even though technology can already do a better job for us.

This is a new frontier with using large language models, and there are some really cool things happening with education. In Khan Academy, for example, they're working on building language models to help children go through reasoning steps, which is a really cool application. Instead of just asking the question, it will probe the student to go through a process so they can come to the conclusion themselves.

There's a pessimistic and optimistic world here. An optimistic world is that these new AI systems are actually going to be a new step forward in cyborgizing ourselves, but it's not necessarily going to be as atrophying as something like Google. These systems won't only give us the dopamine hit of a factual answer; they'll also guide us toward better understanding how they came to their conclusion, to ensure that we understand when we're probing and asking questions. That's an optimistic state of the world.

The pessimistic state of the world would be that we offload more and more of our own cognitive reasoning to these systems, and we become even more atrophied in these abilities. That might not be a good world to live in if we keep offloading more and more reasoning to systems and lose the ability to do it well ourselves.

Tim Scarfe

Yeah. I've been thinking about this a lot recently. I was involved in a startup that did transcription, language models, and augmented-reality glasses. The idea was that you could be in a lecture—and I still think this is very useful for people with accessibility concerns, like those who are hard of hearing—but we were thinking of it as something that could augment your cognition.

You're in a lecture, and now you don't need to pay attention to the lecturer because you're transcribing it and GPT is making notes for you. I think this is really wrong. But you give a counterexample with satnav: we don't need to read maps anymore because we can externalize that cognition.

I feel like this is different. You're in a lecture, and all of these AI language tools are a form of understanding procrastination, right? Understanding, or intelligence, is the process of creating a model. You're creating a simulation, and in order to create a simulation, you actually have to think. You normally think, externalize the thinking a bit, do some writing, and pay attention.

Here's the thing: in that situation, there are so many more cues because it's in 4D. You can hear things, you can see things, and it's a social and physical activity. Even the dance—the performativity of the lecturer—is all information. It helps you understand.

Now I'm transcribing the thing, and people say, “It's okay. I can just read the transcription later and understand it.” Well, maybe, but you're already at a disadvantage. You probably won't, because this procrastination is just paying it down the line. You're saying, “I might do it later. I might do it later,” and you never will. That's going to create a society of automatons that just don't think for themselves.

Max Bennett

Yeah. I'm torn between the optimistic and pessimistic states of the future, but I think there's a very good argument behind what you're saying, so I definitely don't reject it out of hand.

I actually really liked the analogy to model-based versus model-free that you were suggesting there, because that applies very well to Google Maps. When my dad navigates a city, he has a model of the world, and he's engaging in model-based planning of how to get somewhere. When I use Google Maps, I've externalized the model, and all I do is respond to the cue of when to turn right or left.

I think that is absolutely a good way to think about this: we use technology to externalize building models, which can sometimes make things more efficient because then we can just be model-free actors. But there are places, like in the example you're suggesting, where we really want people to engage in the more painful, hard process of building models of things. In those cases, obviously it's dangerous to make it so easy to externalize these models.

Tim Scarfe

Yeah. It's hard to articulate. I think part of it is a kind of acquiescence. You're sequestering your agency when you externalize too much of your cognition, particularly if it's parts of your cognition that are useful because they contain core knowledge that will generalize and help you acquire new knowledge, or if it's the portability of discovering knowledge. It's your intelligence, and you're not exercising that muscle.

You become acquiescent and then you become less of an agent. From a collective-intelligence point of view, we're just saying that language and intelligence are about discovering knowledge. If we are all sequestering our agency and becoming less intelligent as individuals, as a collective, maybe we will suffer.

But it's one of those things that's so easy for us now to make grand statements about. People in 200 years will look back on this and laugh and say, “It's a little bit like when they introduced bicycles.” There was apparently a moral panic because they said, “Women will start cheating on their husbands and using bicycles to go to the next town.”

Max Bennett

That is an interesting fact. Yeah. Well, I think history is such a good tool when trying to reason about how people in the future will think about us. We are the people in the future to the past, which is obvious, but it's a useful tool.

For example, in some sense we already live in this dystopian world when it comes to physical exercise. Roll back the clock 500 years, and most people didn't have to think about physical exercise as much because most work required physical exercise. We exercised with our work, and so much of the work—at least in the developed world—is information-related, where we don't exercise and so we go to the gym.

The gym is a weird thing. If aliens came down and observed gyms, it would be anthropologically very bizarre behavior, because we just go into a room and run on treadmills. We do it because we've evolved to require exercise, and modernity has removed exercise as a prerequisite to most of the things that we need in life.

Now there's this gaping hole, and what we do is go to the gym and run in place to satiate this physical need. You could imagine—and one might interpret this as dystopian or utopian—a world where we've offloaded so much cognition, but because humans need to think about things, or because as a society we value it the same way we value physical fitness, there are now social pressures to go to these intellectual gyms.

Just to make sure, even though you don't need to do it for work or it's not necessary for the world to function, we feel like there's value in a human who knows how to reason about things. So we go to intellectual gyms for that. We might—I don't know if that's a utopian or dystopian future—but however we feel about it, I would venture to guess people 500 years ago who looked at a treadmill would probably feel similarly.

Tim Scarfe

100%. Well, MLST is my intellectual gym, by the way. [Laughter] You spoke about DNA. Dawkins, of course, wrote the book The Selfish Gene, and you said that the value of DNA was not what it creates—it creates hearts and lungs and so on—but what it enables, which is this evolutionary process. But then it gets to this concept of what we mean by a meme in general.

You said that it's an idea or behavior that spreads contagiously. How do you think about memes?

Max Bennett

Well, I think Dawkins did a wonderful job articulating this idea in a way that's really understandable. A meme is a concept or a behavior. A meme can be just the idea that individuals should have rights, or the idea of equality, or something sillier, such as the idea that we shake hands before we sit down for a meeting. And these things, because humans can share simulations through language and we engage in imitation learning, these ideas or behaviors propagate throughout societies.

And because these things are propagating, a different form—not evolution in the sense of genetic evolution, but a form of evolution—emerges, because some ideas will propagate better than others. By nature of that process unfolding, memes—these concepts or behaviors—actually go through an evolutionary process. Ideas that are either viral because people want to share them with each other, or ideas that somehow support the survival of the individuals that hold them, are going to be ideas that propagate correctly.

Ideas that negatively affect the survival of the individuals that hold them, or that people do not desire to share for whatever reason, are going to do a worse job propagating. And so it's a really almost brilliant lens to look at human culture when you reframe cultural ideas and concepts as memes—a different take on genes—that go through their own sort of process of iteration. This is not my idea; this is Richard Dawkins's.

Tim Scarfe

Oh, yeah. Well, we can thank Richard very much for this. I'm fascinated with memes, and I kind of think of language as being a collection of memes. But now we're in this very, very interesting space. Before language, we learned by observing physical skills performed by other people, and we could imitate them and so on. Now we are sharing simulations, basically, without actually needing to see the thing, and that means that we are one step removed from reality.

So all sorts of memes have cropped up, and some of them are better described, as you say in your book, as shared delusions, but they have some utility as well. When we have a common myth, for example, it might be a religion, it might be a nation-state; it allows us to cooperate with each other in a way that we wouldn't be able to do before. And you actually cited some ideas by John Searle and Yuval Noah Harari in his book Sapiens on that. [Snorts]

Max Bennett

Yeah, they've all famously popularized this idea. But John Searle was one of the original ideators of it. What's so powerful about these shared fictions is they can propagate much more easily than a human can talk to everyone in a group. And so, because they propagate much more easily and with very high fidelity, this enables me to meet someone who is a New Yorker whom I have never met before and immediately have shared views.

We probably both believe in individual rights. We probably both believe that money can be used for transacting things. So, if I give them a dollar, they'll believe that the dollar will be used elsewhere, or they can give me a dollar. And of course, today it's hard to reason about these things because there are so many rules in place that you don't realize it's all a shared fiction.

The reason we think we believe in money is because we're like, “Well, I know that all the other stores I go to will take this money.” So that's the reason it works. But why do they all take the money? It's all this shared belief that we all trust that this thing will be used for transacting. And so, because of that, it enables really large groups of people to coordinate, and that is a very powerful aspect of language.

But the argument I make in the book is that, similar to how genes are powerful not because of the structures they create but because they enable a process of evolution by which good structures will emerge, language is similar in that sense. What's powerful about language per se is not that we can engage in these shared simulations for coordination; it's that language enables the propagation of ideas and concepts across generations, which will thereby undergo its own evolutionary process. So, of course, these good ideas that enable survival are going to emerge. And that's really what's so powerful about language.

Tim Scarfe

I guess the arbitrariness is quite interesting. Some of them, on the surface, don't seem like good ideas; they just seem like really bad ideas. And I guess you can think about it in terms of creativity as well. For a meme to be established in the sphere of possible memes, does it need to have intrinsic value? Possibly not, because we're getting into creativity: is it novelty? Does it have intrinsic value? Is it just social proof? Is the meme only existing because lots of people have been fooled into thinking it has value? So it's kind of extrinsic value via social proof.

And then there's almost a double entendre with the meme, or a deeper meme meaning, because you talk about altruism. The meme itself might actually be quite a stupid meme, but if it causes altruism, so there's actually a group-selection advantage to it, then it's almost like that's the lens of analysis to understand how good the meme is.

Max Bennett

Yeah. It's a really fun area of literature to read through because there's still zero consensus as to how language evolved. One reason why it's so controversial is the way in which we disambiguate—and I'll get to your question—the way we disambiguate evolutionary arguments is typically by observing gradation in extant, or currently present, animals. That enables us to observe these intermediary steps between a morphological aspect of body A and a morphological aspect of body B.

The problem with language is we have nonhuman primates that, for the most part, don't have any language, and then we have humans that have very complex language. And all of the intermediary humans that existed between our divergence with chimpanzees about 6 million years ago and our divergence with all other modern humans between 50,000 and 100,000 years ago—we don't have them; they're all dead. All those lineages are lost.

And so that means that there's this broad spectrum of arguments that could be made. Chomsky argues—I find this a very strong claim and thus hard to defend—that it happened all at once, or very rapidly: there was no language, then all of a sudden there was language.

Tim Scarfe

Right, and there are other arguments that it was a gradual process.

Max Bennett

But one of the most controversial aspects of language evolution goes to what you're talking about, which is evolutionary arguments for why language evolved. Evolutionary arguments for why language evolved have a harder burden of proof than arguments for other adaptations.

So when we argue about the evolutionary benefit of something like theory of mind, there are no complex evolutionary machinations one needs to conceive of to defend it, because you can see why it would be beneficial for an individual chimpanzee to be born with the ability to infer what's going on in other people's heads. They can better defend themselves when someone is going to be mean, better figure out whom to trust, better climb the social hierarchy, et cetera.

But with language, unless you take the Chomsky view that its primary adaptation is for thinking, the argument that language evolved for communication is more challenging, because it's not valuable for an individual human to be born with a little bit of language skill unless other humans are also engaging in language. And so this then means that the only benefit is if we're both sharing truly useful information with each other.

Although it seems intuitive that the way this would function is that a group of humans that are sharing knowledge with each other is going to survive better than another group of humans that's not, and that's how evolution will ensue, this is actually quite controversial in evolutionary biology because that's invoking something called group selection.

Now, some people call the modern incarnation of this multilevel selection, where there's some consensus that, yes, there are group-level effects that can impact things. But most people think that group-level effects are not nearly as strong as we would intuitively think.

And the issue is the following. If you have a group of 100 humans that use language with each other and then you have 1 human born who is just going to try to trick all the others, so all they're going to do is use language just to be disingenuous, it's not at all clear that that human would be at a disadvantage. In fact, they might be at an advantage relative to everyone else.

If you play that forward over time, language will be lost because someone born who isn't going to be tricked by the individual trying to lie to them with language is actually going to survive better than the people who have language skills. There has been so much debate throughout evolutionary linguistics about these arguments as to how language evolved. I like the argument in a great book called The Evolution of Language by Fitch, and I think he makes a really great argument around how you could think about this occurring.

A lot of people argue that it probably started with something called reciprocal altruism. The way altruism exists in the animal kingdom, there are 2 forms of accepted altruism. One is something called kin selection, which is quite straightforward: I'm willing to sacrifice something—in other words, share something with an individual—if I share genes with them. That's easy.

Reciprocal altruism, which we do see in the animal kingdom, is, “I'll scratch your back if you scratch my back.” But if you start not scratching my back, then I'm going to stop scratching your back. What this suggests is that, in order for language to be stable—in other words, for it to be beneficial for me to truthfully share information—there need to be costs to me lying.

This is one argument that people speculate is one reason why humans have such strong moral preferences toward punishing liars and out-groups and in-groups, because we really try to identify individuals who are lying. Robin Dunbar has a beautiful argument that this is why gossip evolved. One way that evolution can stabilize the use of language is by virtue of us having a preference to share moral violations.

Gossip is a tool of language where, if you see someone lie or cheat and you share it with a bunch of other individuals, that becomes a huge cost to someone lying and cheating, because if 1 person catches them, then the whole group is aware of it. There is this special feedback loop that happens where language skills require more punishment of violations to be a stable strategy. One way you get that is by having more gossip and making sure there are higher costs to defecting.

This is not by any means the only story of language evolution, but it's one with a lot of interesting evidence behind it. There are some people who argue that the feedback loop—1 emerging idea, which I don't talk about in the book but which I do think is interesting—is actually one in which we try to detect lying in others. They make the counterargument to me, saying that the effect of lying is the loss of language.

There is an argument that you get the reverse: you get a really good theory of mind in humans because we're so sensitive to trying to detect people who are actually giving us false information. There's still a lot of controversy around it, but the main takeaway is that the blanket group-level selection argument—that language is obviously beneficial because once a group has language, they're all going to survive better—is not a sufficient argument for language evolution.

You need a more nuanced evolutionary argument as to why it's a stable strategy for an individual to be born with superior language skills, or you have to argue that language did not evolve primarily for communication.

Tim Scarfe

You know, it's quite interesting, first of all, that you were writing this book actually a couple of years ago. So this was before GPT-4, although you did put a note in about GPT-4, and you were speaking about Blake Lemoine. He was a Google engineer, and he famously came out convinced that these things had developed sentience.

I think much of this actually hinges on the concept of a world model. One view of language models is that they're just modeling a statistical distribution of tokens, and that seems quite low-resolution. Another take is that they're learning a world model. What that means is that, rather than just capturing the state of language, they're actually simulators. They're generating the underlying processes of language. They're capturing the dynamics of language.

How do you say that something is or is not sentient, especially given that the models could potentially be such high resolution that they are generating the same thing for all intents and purposes?

Max Bennett

This is where I don't see myself as a philosopher, but this is where I do think scientists need to include philosophers. When questions become nonscientific, I think the scientific instinct is to argue that we don't draw distinctions between things that the scientific method can't draw a distinction between. But the problem is that there might be moral differences between them.

For example, it might be scientifically impossible for us to differentiate which of 2 systems that look indistinguishable in their inputs and outputs is sentient. Scientifically, we might say, “Well, because we can't differentiate the 2, we're going to say they're the same.” But that doesn't mean they're the same. That just means that, because we have no methodology for drawing a distinction between them, from a scientific perspective we're not going to draw a distinction because we're entering philosophical territory.

But if you take that and then start talking about policy implications, the actual values we attribute to them, and how we introduce these things to society, I think we need to include a philosophy lens here. It might not actually be the case that they're the same just because we can't distinguish them. So that's just 1 thought.

Tim Scarfe

Another thought on world models. One distinction I want to draw, because I've seen a lot of confusion on the internet about the world-model dilemma, is that there's a difference between a world model and a model. It is undeniable that language models have a model. In order for GPT-4 to correctly predict the next token in these really complicated language questions, it clearly has some model of something.

Because we can ask it common-sense questions about the world and it answers many of them correctly, you can say this is a model of aspects of our world without question. I think it would be very hard to argue that that's not the case if you look at GPT-4's performance on many of these questions. But what most people mean when they say “world model” is a specific process of simulating an ordered sequence of states and the consequences of different actions.

Max Bennett

That means identifying the end result of these actions in your head. Another way to think about what we mean by a world model is the ability to reason about interventions and causality. This is the Judea Pearl argument: with our world model, we can hypothesis-test.

I can say, “I imagine that if I do this thing in the world, I think this will be the consequence of it because that's what I see in my head.” Now I have a hypothesis. Now I'm going to actually do that thing in the world and see if my hypothesis is correct.

That's very different from what's happening in a language model, where its understanding of the world derives solely from its input data. In a world model, my understanding of the world comes from the delta—the difference between what I hypothesize is going to happen in the world and my actual experience of it.

This distinction really matters the more we're going to start offloading our cognition to these systems. For example, everything that ChatGPT knows is on the basis of its input data. That means if false information or wrong information is in the input data, ChatGPT is going to know that information. There's absolutely no hypothesis-testing embedded into ChatGPT, unlike our true AGI agent that will one day be invented.

What it would do is hypothesize aspects of the world and test its own hypotheses. If you give it false information, if it reads articles about how the Earth is flat, it's not going to just start talking about how the Earth is flat. It's going to say, “Okay, well, that is incongruent with my model of the world. I'm going to now run some tests where I can differentiate between them, and I'm going to perform those tests and then conclude that the world is not flat.”

When one says that ChatGPT does not have a world model, I think some people misinterpret that as suggesting that it's just dumbly looking at the statistics. That's not at all what we're saying. In order to correctly look at the statistics of language, clearly it's built up a very rich and complex model of the text that it's seeing, and that's how it's able to predict the next word so well. But it's not what most people mean when we say “world model.”

Tim Scarfe

A couple of things on that. I think people conflate the machinations of language models with how we represent them statistically or abstractly, because if you look at a lot of papers, they actually represent it like a probability—a joint probability distribution. Of course, the way that language models work is completely different from that, but you're bringing in some very interesting things.

So first of all, we are agents in the world. The agential lens is quite interesting: we interact with the world, so we're not just learning from observational data. I was talking with Nick Chater, and we said, “Why is it that in our everyday experience, we experience the world in 4D color?” He said it's because it's interactive.

In your experience, you can actually seek new information, right? You can move your eyes, you can get new information in, and you can touch things. When you're doing future or past simulations, you don't have that interactivity. So there's something about interactivity that's really important.

But even then, how far could you go? A complete one-to-one simulacrum of the world wouldn't be a particularly good model. In physics, there is no causality, right? It's just dynamics. Causality is actually something that emerges very, very far up. So we're talking about a model that's an approximation of the real world, which may or may not include causality.

It probably would, because it's an interactive model and it has this kind of agential map. But I guess we're just drawing the line somewhere and saying, “Well, that is a world model.” Max Bennett

I'm actually not sure whether, even if we rendered a perfect 3D map of every particle in the universe—

And that was the input data to some infinitely large model. I still would argue that it's learning something different from a model that's given some form of agency, where it can hypothesize rules and then test its own rules.

Now, given infinite time, it's possible that those will converge, because given infinite time, every possible hypothesis I could conceive of will end up showing up in the training data. So eventually I'll see the training data of every possible experiment I could run. If time is infinite, I guess you could suppose that happens.

But what's so different is this dramatic dimensionality reduction that happens when you show me something uncertain and then I can conceive of specifically the tests I want to run to map the uncertain thing to my mental model of the world. That's a very different way of learning about things. It's not just input data and then self-supervising on predicting one's own input data. It's building a model in which I can simulate possible outcomes and then hypothesis-test those outcomes.

Even this is not uniquely human at all. If you look at the way a rat would deal with something novel in its environment, it's drawn to the novel thing and explores it until it feels like it understands it. Then it will move away.

When you show a child an object that's perplexing, they will touch it, turn it around, and try to understand it until they feel like they've built a model of it. That simple act is doing something very different from the self-supervision we see in most AI models today, because I see something I'm uncertain about and I'm volitionally going to create new training data for myself.

I know the training data I want now. I want to see what happens when I pick it up, turn it to the left, and turn it to the top. A convolutional neural network doesn't do that. So the way we teach CNNs to understand rotations in 3D objects is by manipulating the training data ourselves. We take imagery and rotate it in a bunch of different ways, so we're the ones curating the data set to teach it these things.

Max Bennett

Yeah, but that's different from the way we learn about things. I think this is a key aspect that's missing from AI systems today, and it's something that folks are working on. That's something we're going to have to add in.

Tim Scarfe

Yeah, completely agree. It feels like you're saying basically what I think, which is that there's a creativity and an agency gap. A lot of that is because we're agents, as you say: we create our own training data, we do this active inference and sense-making, and we build these models in real time. As a collective intelligence, it creates a kind of divergent search process for knowledge. It's this epistemic foraging that we spoke about.

GPT is a monolithic model. It does have models, but the models are only learned at training time. Inference actually happens at run time, and when you put a prompt into GPT, you're just retrieving one of the models that was already learned a long time ago. It's not creating a new model in the moment. So it creates this kind of sclerotic system rather than the divergent, creative system that we experience in biomimetic intelligence.

One sort of mental model I have of this—because there's so much debate around it—is something I'd be curious to put through the gauntlet of what other people think about. Maybe people in the comments will either agree or disagree with this. I think an interesting alternative experiment, or eval, of an AI model that I haven't heard before is this: if you give it knowingly false information in the training data—not at inference time, but in the training data—will it reject it wholesale?

To me, this is the distinction: an agent that can hypothesis-test and intervene in the world will reject false information. If you tell it that the world is flat, it will know that the world is not flat. Whereas with GPT, any data you give it is given equal weight to every other piece of data. The only reason it would reject that is if there's other data in the training set that it's going to ignore.

It's almost cheating, because by definition we know a language model is going to fail at this task. The only way you can fix it is if you give it other data in the training set. There's no notion of hypothesis-testing compared with an agent. The only way you could get it to be wrong is if you manipulate its sensors during the actual hypothesis-testing that it does.

You can, of course, manipulate it by changing the actual test when it does these tests. That happens in The Three-Body Problem, which is an amazing book where aliens manipulate our experiments. Anyway, I think that's another way to evaluate these systems: can it figure out that you're giving it false information and reject it?

Max Bennett

Yeah, 100%. A couple of things, though. There's something magic about having agentic density in the system, right? When you have something like GPT, just to make it statistically tractable, it's generally doing a kind of low-entropy search. What I mean by that is it's just looking for the baseline patterns. It's not doing a lot of exploration, and it's not searching outside of the main sources of statistical regularity.

Whereas when you have divergence in the search process, with all of these individual agents doing their own things, as a system it's much more of a high-entropy search. That means you're actually bringing in lots of new information to solve problems in creative and interesting ways.

In the physical world, though, it's quite interesting, right? The problems come from the physical world. The trees get big, so giraffes need to have a long neck in order to eat the leaves from the trees. This whole thing just rinses and repeats. The environment produces novel solutions, and then we see this divergence and find novel, creative solutions to the problems that get generated.

But in the memetic sphere, it's so much more difficult than that, right? The problems and the guardrails aren't constrained in the way that they are in the physical world. For example, we have capitalism or we have nation-states, and there’s all kinds of interesting divergence going in different directions. But it doesn't seem like there are the same pressures that ground the thing in reality.

Tim Scarfe

Well, yeah, I think it's definitely not grounded in truth. Its tethering to truth is this: knowingly false information that leads me to take actions that will hurt my survival will fade. But false information that helps me survive better will propagate freely, or is at least neutral.

Another way this shows up is—and this is where I'll go into some pontificating—

Max Bennett

Please.

Tim Scarfe

But where I think there is memetic evolution that can drive us away even from things like happiness. If we think about what systems of coordination survive, they're systems of domination and militarism.

If you take 2 groups of individuals, let's say one is really happy and calm, sees no desire for domination, and does not attempt to innovate and build more technology. The other is unhappy but super aggressive, wants power, and wants to expand. These ideas will die out.

What this suggests—and I think I talked about this in one of our previous conversations—is the importance of delineating, in my view, the Darwinian component of what does survive from the moral component of what is right or wrong. It is definitely not the case that what survives is definitionally right. It's absolutely possible that the things that survive and do well evolutionarily are not the things that we feel are morally aligned.

That is not to propose a correct or incorrect system, but it is an important distinction to draw when we're trying to decide what we deem to be morally right or wrong.

Max Bennett

So I think that's just one example of what you're saying: the ideas that propagate successfully might not be the ones that are true. They might also not be the ones that we deem to be moral, or even be the ones that lead to human happiness. They're just the ones that do a good job of keeping humans alive and reproducing the idea.

Tim Scarfe

Yeah. So people say that language models confabulate and don't preserve epistemic factfulness. But you could also argue the same thing about us, right? We actually confabulate everything. We don't really have goals. We just generate these post hoc confabulations, then explain our behavior and pretend that was what we wanted to do, that we had beliefs, and so on. We just make it up as we go along using this kind of active inference.

Even though we are emotional and subjective, and we believe in religion and lots of things that we presumably made up, we have Wikipedia. We have objectivity, even though it's an illusion, right? There's no such thing; even general relativity isn't as objective as we think. If you keep asking why and why and why, it just disintegrates into incoherence. But there seems to be some objective structure that is preserved. How is that explained, given that our brain simulations don't seem to select for truthfulness?

Max Bennett

I think the question of whether humans are better or worse than ChatGPT is almost a red herring. I look at ChatGPT as an alien—it's like an alien brain. There are certain things it does that are clearly better than us. Information retrieval in ChatGPT blows a human away, without question. In many ways, it's way better than humans.

But there are certain things that human brains do that ChatGPT does not. If we're trying to build human-like intelligence, there's certain inspiration we can garner from human brains. I think there's a component of our model-based rendering of a plan and then executing that plan that has a level of explainability that's unique relative to a system that is just iteratively predicting the next token.

But we also do the same thing that ChatGPT does. When we make model-free choices and then you say, “Why did you do that thing?” what we engage in is exactly as you're describing: a post hoc explanation. I didn't render a plan; I was just walking down the street. If you say, “Why did you move your foot there as opposed to 2 inches to the right?” what I'm going to do is render a post hoc explanation of why I did that.

But I didn't really think about it; I'm just explaining it after the fact. So it's definitely the case that humans have that component. But it means there's also another component, which I would argue is unique and important: our ability to pause, render a plan, and then execute against that plan. The key thing that I think is the dividing line between these models and us is the ability to render hypotheses and make interventions in the world. That's the key thing.

And so it's not the case that our brain has the true objective state of the world in our head. I don't think that's there. There might be components of objective truth in ChatGPT that it contains that we don't have, and I think in its information retrieval it has probably, in some ways, more aspects of reality than I do, in terms of having read all of Wikipedia and answering questions about biology that I don't even know the answers to. But there are also components of the world that the human brain has rendered and contains that ChatGPT does not, because of our ability to make hypotheses, intervene, and learn the causal structure of the world. I think that is the dividing line. But I wouldn't say it's because the human brain knows the objective state of the world and ChatGPT does not.

Tim Scarfe

Max Bennett, it’s been an absolute honor to have you on MLST. Everyone at home, you need to buy his book immediately. It is a wonderful, wonderful book, Max. You did such an amazing job of bringing all these things together. We've now spoken for 4.5 hours, going through the last 3 chapters. My God, it's been an honor. Thank you so much.

Max Bennett

It's been my pleasure. Thank you for having me.

Your Brain Doesn't Command Your Body. It Predicts It. [Max Bennett] | BidClub