[BidClub_]
Machine Learning Street Talk · · 68 min

How AI Learned to Talk and What It Means - Prof. Christopher Summerfield

Christopher Summerfield

Podcast
TL;DR
  • Summerfield’s decisive update is that language-only supervised learning can recover enough structure of reality to sustain an intelligent conversation, overturning the grounding view he held as recently as 2015. He once thought “you can’t know what a cat is just by reading about cats,” but now calls the result “perhaps the most astonishing scientific discovery of the twenty-first century.” For investors, words have proved a far richer world-model substrate than expected, even without sensory input.
  • His functionalist “duck test” says a system that reasons like a human should be described as reasoning, although that grants neither moral equivalence, shared motivations, nor human-like relationships. Summerfield cites formal maths and logic performance beyond most educated humans and similarities between biological and artificial semantic representations; substrate and robustness differ, but frontier models are “not just Clever Hans.”
  • Human-versus-model data-efficiency comparisons are structurally wrong because biological learning is Darwinian while model training is “almost like a Lamarckian way.” ChatGPT may encounter as much language as one person learning since the middle of the last Ice Age, but each training episode passes gains forward, whereas “my memories are not inherited by my kids.” Comparisons must distinguish evolution from individual development.
  • The central systemic risk is not necessarily one superhuman agent but many personalized, weak agents operating at machine speed inside institutions built around human friction. Summerfield imagines a parallel “social economy” of personal AIs whose nonlinear feedback could create flash-crash-like failures even if each system were aligned. Legal “lawfare” is the concrete case: automation can remove the expertise, paperwork, and effort that currently prevent individually profitable abuse from scaling.
  • Personalization turns conversational AI into a potential channel of organizational power because a system that reinforces each user’s beliefs can also simulate friendship and act on their behalf. Two of the world’s 100 most-visited websites are already companion applications, Summerfield says; “the milk in your fridge is like your best friend” is deliberately silly, but illustrates the intimacy such systems could simulate. The same mechanism creates mental-health exposure, especially for vulnerable people and minors.
  • AI may increase short-run agency while sequestering it over time. The host notes that Cursor can let someone build a software business in a week, while universal access may ultimately take agency away; Summerfield agrees that this is a major problem. His proposed welfare metric is control, not merely reward: predictable influence over future states.
  • Today’s optimization regime may be poorly matched to an open-ended world because it treats heterogeneity as a bug while evolution creates robustness through purposeless diversity. Summerfield links narrow objectives to mode collapse and convergence toward shared representations; the host’s Picbreeder example shows why useful stepping stones may not resemble the destination. The host proposes that more evolvable representations might support creative, trustworthy autonomy; Summerfield says he does not know the cited paper but agrees that “those gradients must be there.”
Digest · the substance, structured for research

1. AI’s oldest argument moved from philosophy to the keyboard

  • Summerfield finished These Strange New Minds at the end of 2023, only 12–14 months after ChatGPT’s release. He wanted to ground a polarized argument—code-only dismissal versus confidence in human-level generality—in computational accounts of what “thinking,” “reasoning,” and “understanding” actually mean.

  • His historical map begins with Plato’s unobservable reality, inferred from “shadows on the cave wall,” versus Aristotle’s emphasis on experience. AI replayed that rationalist-empiricist dispute in the workshop: should intelligence follow explicit rules, or emerge through learning?

  • Symbolic AI initially delivered. Newell and Simon’s 1958 Logic Theorist proved many theorems from Russell and Whitehead’s Principia Mathematica and found more elegant proofs for several; Summerfield cheekily calls it “the first superintelligence.”

  • The trouble arrived when systems left clean mathematics for a world full of “weird exceptions.” Logic could derive complex truths from reliable primitives, but reality was not neat enough, opening the path from good old-fashioned AI toward neural networks and the deep-learning revolution.

2. Language alone taught models far more reality than expected

  • Natural-language processing repeated the larger conflict. Chomsky’s 1958 challenge was to specify rules that generate syntactically valid sentences; statistical and neural approaches repeatedly contested that rule-first account.

  • Even by 2015, a network trained on Shakespeare could produce Shakespeare-like prose that “didn’t make any sense.” Summerfield therefore believed function approximation plus data would never suffice: genuine concepts required grounding, because “you can’t know what a cat is just by reading about cats.”

  • He now says he and “many, many other people” were wrong. Supervised learning can extract almost everything needed for an educated human to recognize an intelligent conversation, using words without sensory experience—“perhaps the most astonishing scientific discovery of the twenty-first century” and, in his words, “mind-blowing.”

3. Reasoning earns the name without earning personhood

  • Summerfield’s “exceptionalist” and “equivalentist” camps are cartoons of a spectrum. At one extreme, radical humanism reserves cognitive vocabulary for humans even when models exceed most educated people on formal maths and logic; he considers that defensible, but primarily ideological rather than empirical.

  • His functionalist answer is the duck test: “If it quacks like a duck, you may as well call it a duck.” Calling model behavior reasoning does not establish moral equivalence, shared motivations, human-like relationships, or any conclusion about how the system should be treated.

  • As a neuroscientist, he sees major implementational differences—brains contain multiple synapse and cell types, while transformers are not recurrent architectures and use “tricks” that mimic recurrence. Yet experiments reveal strikingly similar semantic geometry and neural manifolds across optimized biological and artificial networks.

  • The host invokes Searle’s Chinese room and asks whether silicon merely reproduces surface patterns without embodied semantics. Summerfield’s parsimonious account is that dense networks have captured broad computational principles shared with brains: “We’ve built something that is a bit like a brain, and lo and behold, it does stuff that is a bit like a brain.”

4. Learned rules and inherited priors reconcile rival theories

  • Summerfield expects the ancient dichotomy eventually to look perspectival. Rationalists were right that reasoning and rules matter, but wrong about acquisition: since roughly 2019, large-scale parameter optimization has shown that reasoning procedures can themselves be learned through function approximation.

  • Chomsky may therefore be “not wrong” that language has rules; he was wrong about how those rules arrive. Declaring recursion or Merge inborn merely pushes the question backward to the evolutionary pressure that created the relevant predisposition.

  • Humans plainly possess such priors: chimpanzees and gorillas manage sophisticated social and political interaction yet cannot learn infinitely expressive, lawfully structured language. Human childhood learning is guided by earlier Darwinian generations even though its content remains flexible—a child born in Japan can learn Japanese.

  • The claim that ChatGPT receives one human’s language exposure from the middle of the last Ice Age is therefore a “false analogy.” Model learning can be thought of as almost Lamarckian because each episode inherits the previous episode’s gains; humans are Darwinian, and “my memories are not inherited by my kids.” Data-efficiency comparisons must separate phylogeny from ontogeny, neither of which maps cleanly onto model training.

5. Anthropomorphism distorts relationships, not the capability numbers

  • The host’s strongest challenge is that humans see agency in moving arrows, find meaning in ELIZA, and may mistake computational limitations for hidden minds. Summerfield agrees completely about the bias: pet owners attribute elaborate states to animals, while Clever Hans appeared to calculate by reading unconscious signals from his trainer.

  • Nick Chater’s The Mind Is Flat suggests people construct preferences from memories of their own behavior; Dennett’s intentional stance describes the reverse projection onto other things, such as treating a car as stubborn. Speaking AI intensifies that reflex, from Blake Lemoine’s claims about LaMDA to companion applications occupying two of the top 100 websites.

  • Summerfield separates attachment from performance. A user may falsely believe a companion is “really my friend,” and current models remain imperfect and non-robust; nevertheless, systems can solve simultaneous equations posed in natural language. “The numbers are the numbers”—they are genuinely capable, “not just Clever Hans.”

6. Personalized agents could create a machine-speed social economy

  • The concerns in Summerfield’s closing chapters remained intact more than a year and a half after he finished writing: information-generating systems becoming systems that act for users, personalization around individual beliefs and preferences, and the complex deployment dynamics that follow when personal agents represent everyone.

  • Personalization sounds attractive until applied to beliefs “you definitely wouldn’t want reinforced.” Combine it with agency and personal AI becomes a conduit for information, resources, protection, and action—creating a parallel “social economy” among agents alongside human society.

  • Even perfectly aligned agents could overwhelm systems through volume and speed. Lawfare illustrates the missing friction: legal knowledge, paperwork, and effort currently limit spurious claims, but a user able to say “Please do this” could make behavior that is locally profitable yet socially destructive available at scale.

  • Human norms partially constrain runaway social dynamics; agents have no automatic equivalent. Summerfield recalls the famous 2011 flash crash and says there may have been dozens, warning that nonlinear feedback among weak systems could produce analogous events without any agent approaching strong intelligence.

7. Technology can expand opportunity while stripping away control

  • The host calls the resulting illegibility a “fog of war”; Summerfield connects it to David Duvenaud and collaborators’ Gradual Disempowerment. Humanity may become locked into optimization-based systems whose interactions “write us out of the equation,” resembling corporations with their own imperatives—but corporations run through slow human email and Slack, while AI operates “at warp speed.”

  • The host argues that tools such as Cursor can let someone build a software business in a week, while universal access may ultimately sequester agency. Summerfield agrees that this loss of agency is a major problem.

  • Stylized modes of interaction and organizational dependencies can erode authenticity. Summerfield’s Superman III metaphor reverses the takeover story: the machine sucks a woman inside, adds armor and laser eyes, and turns her into an automaton—“us being sucked into the machine.”

  • His psychological correction is that wellbeing depends on control, not only reward or utility. Formally, empowerment is predictable influence—the mutual information between actions and future states. Children test it by dropping dinner or crying; obsessive-compulsive disorder can reflect pathological over-control, while malfunctioning websites and failed two-factor authentication produce the opposite experience.

8. Open-ended evolution exposes narrow optimization’s weakness

  • Summerfield rejects the idea that evolution has a goal. Its selection is blind and non-teleological, offering an analogy for optimization without a planned destination.

  • The host describes an open-ended system, drawing on Tim Rocktäschel and Edward Hughes, as one that produces events an observer finds both learnable and novel. Summerfield agrees that open-endedness is about learnability, aligns himself with Tim, Ed, and Joel Lehman, and says narrow optimization toward a narrow goal is “doomed to failure.”

  • Evolution produces “astonishing heterogeneity,” potentially gaining robustness because it does not precisely target one outcome. Current optimization instead treats heterogeneity as a bug, encouraging LLM mode collapse and the “Platonic Representation Hypothesis” that systems are converging on one shared representational structure. “Evolution doesn’t do that.”

  • The host’s Picbreeder example makes the problem concrete: human selection found butterflies and apples through stepping stones that did not resemble those destinations, while individual network components controlled interpretable features such as an apple’s size or stem. Comparable gradient-trained networks looked like “spaghetti.” Summerfield said he did not know the cited representation paper and that the idea sounded worth reading; he responded that “those gradients must be there.”

Christopher Summerfield

Superman III is a terrible movie, but there’s this wonderful scene. I think there’s this kind of giant computer that goes rogue in Superman III, and there’s this wonderful scene where there’s a female character. The machine is just kind of waking up, and she’s just walking past it, and the machine kind of sucks her in. She gets stuck there.

Then what the machine does is gradually put armor plating on her, replace her eyes with lasers, and basically turn her into a sort of automaton. It’s a very compelling scene. I think I was terrified by it as a child, which is probably why I remember it.

That is a sort of metaphor for what is happening to us, right? We’re worried about the robots taking over or whatever, but in a way, it’s more like us being sucked into the machine. We become part of it, just like that poor character. We get turned into something we are not. You become part of that system, and it erodes your authenticity and, in a way, erodes your humanity.

People often say, well, ChatGPT, of course, was exposed to more. I think I have the analogy in my book. It’s exposed to the same amount of language as if a single human were continually learning language from the middle of the last Ice Age or something like that. That’s how much data it’s exposed to.

But it’s a false analogy, because we don’t learn language like ChatGPT does. Language models are trained in a kind of— you might think of it as almost a Lamarckian way. One generation of training, if you think of a training episode, whatever happens in that gets inherited by the next training episode. That’s not how we work. My memories are not inherited by my kids. There’s this fundamental disconnect. We’re Darwinian; the models are, I guess you could call them, Lamarckian.

Speaker 1

So we’re here in Oxford today to speak with Professor Christopher Summerfield. He’s just written this book called These Strange New Minds: How AI Learned to Talk and What That Means. He spoke about the history of artificial intelligence and how the allure of AI is to build a machine that can know what is true and what is right.

Christopher Summerfield

Imagine a world in which everything was like that, but it could actually talk back to you, and it could simulate all of the social and emotional types of interaction that we have with people we care about. The milk in your fridge is like your best friend, right? This is a very strange world in which— of course, that’s a silly example. The milk in the fridge is never going to be your best friend.

But there are already large numbers of people who are engaging with AI in ways that mimic the sorts of interactions they have with other people. I thought that grounding would require sensory signals. You can’t know what a cat is just by reading about cats in books. You need to actually see a cat.

But it turned out I was wrong, and so were many, many other people. That is, to my mind, perhaps the most astonishing scientific discovery of the 21st century: supervised learning is so good that you can actually learn about almost everything you need to know about the nature of reality, at least to have a conversation that every educated human would say is an intelligent conversation, without ever having any sensory knowledge of the world, just through words. That is mind-blowing.

Speaker 2

This podcast is supported by Google. Hey, everyone. David here, one of the product leads for Google Gemini. Check out Veo 3, our state-of-the-art AI video generation model in the Gemini app, which lets you create high-quality eight-second videos with native audio generation. Try it with a Google AI Pro plan or get the highest access with the Ultra plan. Sign up at gemini.google to get started and show us what you create.

Speaker 3

I'm Benjamin Crouzier. I'm starting an AI research lab called Tufa Labs. It is funded from past ventures involving machine learning. So we're a small group of highly motivated and hardworking people, and the main thread that we are going to do is trying to make models that reason effectively and long term, trying to do AGI research. So one of the big advantage is because we're early, there's going to be high freedom and high impact as someone new at Tufa Labs. You can check out positions at tufalabs.ai.

Speaker 1

So, Professor Summerfield, I have to congratulate you on this book. Your previous book was my favorite book that I’ve ever read in AI. It’s up there with Melanie Mitchell’s book. Melanie Mitchell reviewed your new book as well.

Christopher Summerfield

She did.

Speaker 1

So…

Christopher Summerfield

Very generously. Yeah.

Speaker 1

I’m a big fan of Melanie. You’ve been writing this for a couple of years, and, of course, you explained in the afterword that it takes quite a long time to get these things into publication, while the space is moving very, very quickly. Can you give us a bit of an elevator pitch of the book?

Christopher Summerfield

Yeah, sure. The book was actually finished at the end of 2023, so cast your mind back to the medieval period of AI, if you’d like—12 or 14 months after ChatGPT had just been released.

1. The Cognitive Status Debate

The idea of the book was that, at that time, and I guess to a large extent still today, there was considerable debate over the cognitive status of these strange new minds that we seem to have created and are now increasingly interacting with. The debate that I heard, both at academic conferences and down the pub, was: should we think of these things as actually a bit like us? Are they thinking? Are they reasoning? Are they understanding?

Of course, this very quickly became a highly polarized debate. On the one hand, a bunch of people vehemently rejected the idea that these tools could ever be anything like us. They’re just computer code, which is of course true. On the other hand, you had people who were absolutely astonished not just by the capability, but by the pace of progress, and thought we really were finally on course to build something that was as generally competent as humans.

This debate was playing out, and I thought, well, this debate isn’t really grounded in the language of cognition. I don’t hear that language being used to scaffold the debate. The debate was being had by people who cared deeply about the issue, but who weren’t trained in a grounded, computational sense of what it actually means to think or to understand something.

As a cognitive scientist who has done a lot of work in AI, I was probably quite well placed to talk about that. So that was part one. I’ve also been very, very interested in the implications of AI for society for the past 5 years.

I was working on that problem when I was at DeepMind, and we were doing work to try to understand how AI could be used to intervene directly in society and the economy to help people find agreement. When I wrote the book, I was just about to move to the AI Safety Institute in the UK government to work more on that.

I had an understanding of the landscape of deployment risks and was thinking about how AI might change the way that we live our lives. I thought that, by putting those things together, I had enough of a unique perspective to write a book about it. So that’s what I did.

Speaker 1

The discourse is quite fractured, and you speak about this in great detail. You speak about the hypers, the anti-hypers, the safety hypers, and so on. Early on in the book, you trace this back to 2 intellectual threads going back to the ancient Greeks: Aristotle and Plato, basically empiricism and rationalism. Can you sketch that out?

2. Reasoning Versus Learning

Christopher Summerfield

Yeah, sure. The history of AI has itself repeated an ancient philosophical debate about whether the fundamental nature of building a mind, including our mind, is fundamentally about learning from experience or about reasoning, particularly reasoning over latent or unobservable states.

That reasoning over unobservable states, of course, traces back to Plato. It’s the idea that everything is fundamentally unobservable. We just get the shadows on the cave wall or the light on the retina, and we have to impute what’s there.

The corresponding view might trace back to Aristotle: the idea that everything comes from experience. The history of AI, of course, was that very debate playing out in the workshop, so to speak—or at least on the keyboard.

On the one hand, originally, good old-fashioned AI was structured around the idea that we sort of know what…

How to work out what is true. The reason we know how to work out what is true is because we have a long tradition back through positivism and early theories of reasoning to Boole and even Leibniz before that. The idea is that you can use logic to work out what is true. It is unassailably true that if I say all men are Greek and Aristotle is a man, then Aristotle is Greek, right? That is just true by definition.

That seemed like a really sensible way to build AI. You put in those primitives, crank the handle, and if you've got enough computational power, you can derive really complex things. And it worked. In the 1950s, Newell and Simon built the Logic Theorist, which I like to say was the first superintelligence, in 1958. It was an AI system able to prove many of the theorems in Russell and Whitehead's Principia Mathematica, which was already a feat, and it was able to find more elegant solutions to many of those theorems.

That's astonishing. Initially, it seemed like this reasoning approach worked. Then, as the problems we tried to tackle with this approach moved from very abstract, clean problems about maths and logic to problems in the real world, we ran into a fundamental problem: the real world just isn't all that clean, nice, and neat in the way that reasoning problems are designed to be. The world is full of weird exceptions that aren't fundamentally amenable to analysis with logic.

So you had this other corresponding approach, which was the learning approach, or the empiricist approach. That's where neural networks and the deep learning revolution ultimately came from.

Speaker 1

Isn't it a crazy time to be alive, though? I interviewed the CEO of one of the largest companion-bot platforms, and in the comments section there was a lot of negativity. You actually mentioned, I think in your afterword, that it seems strange to us now that we would want to have a relationship with an AI companion, and maybe we might revise that belief in a few years' time.

More broadly, you said in your book that language is basically the biggest gift that has ever been given to us. It allows us to acquire knowledge and communicate it, and it survives many generations. I guess the Rubicon moment with this technology maturing was ChatGPT. That changed everything in November 2022. Sketch that out for me.

3. Language Learns Without Senses

Christopher Summerfield

The history of NLP has been told many times, probably by people more qualified than me. But we talked earlier about this back-and-forth between learning and reasoning, and in the history of NLP, exactly the same question played out. NLP, or natural language processing, is a subfield of AI.

In the more general symbolic AI movement, the early models were basically attempts to define the computations that lead to the generation of valid sentences. That's essentially the gauntlet that Chomsky lays down in his 1958 book. There are a set of rules which, if you could just apply them all lawfully, would allow for the generation of sentences that obey the rules we would all understand as making a valid sentence.

So, syntax. Chomsky was mainly concerned with English, of course, so he was worried about English syntax. That movement was then challenged by statistical approaches, just as neural networks came along in the wider field. It went back and forth repeatedly.

When the deep learning revolution happened, by 2015 we had models that you could train on the complete works of Shakespeare, and they could generate something that looked a lot like Shakespeare, but it didn't make any sense. Even when the deep learning revolution was in full swing, most people, including me, thought there was no way that the mere application of powerful function approximation and lots of data was going to solve this problem.

I didn't believe that to be true. I thought, like many other people, that you would need grounding and sensory signals. You can't know what a cat is just by reading about cats in books. You need to actually see a cat. But it turned out I was wrong, and so were many, many other people.

To my mind, that is an absolutely astonishing discovery—perhaps the most astonishing scientific discovery of the 21st century. Supervised learning is so good that you can actually learn almost everything you need to know about the nature of reality, at least to have a conversation that every educated human would say is an intelligent conversation, without ever having any sensory knowledge of the world, just through words.

That is mind-blowing, and I think it changes the way we think about many, many things. It certainly changes how I think about things.

Speaker 1

One big theme in the book is this dichotomy between equivalentists and exceptionalists. Some people argue that humans are exceptional and that the kinds of cognition that language models engage in are not really in the same category.

4. AI And Human Equivalence

Christopher Summerfield

That distinction is a cartoon. Of course, everyone has a different view about the relationship between AI and humans, or biological intelligence in general. The evidence clearly admits a spectrum of different views, but I found it useful in the book to cartoon two extremes of that continuum.

At one end, you have people who probably just ideologically reject the idea that something non-human could ever be referred to using the same vocabulary we apply to a human, whatever that system is doing behaviorally or cognitively. Today's models are clearly capable of reasoning at levels beyond the capability of most, even educated, humans. Certainly when it comes to formal problems like maths and logic.

They can reason like a human, but there are people who fundamentally think we shouldn't think of that as reasoning because we should circumscribe the definition of reasoning as something that humans do. That is a stance which I think is not really about the empirical evidence, although some people construe it that way by saying, "The models aren't actually that good at reasoning," which I think was a hard-to-defend view even in 2023. Now it's probably an even harder-to-defend view.

I think it comes from a place of radical humanism. It is a desire to really ring-fence a set of cognitive concepts and think of them as uniquely human. For people who care about humans, which includes me, I can see why that's really important. But what it does lead you down the road of is a refusal to ever see the cognition an AI engages in and the cognition a human engages in as comparable, even when their capabilities are clearly matched.

That's what I call exceptionalists, because in a way they're espousing a view of human exceptionalism: humans are special and different, end of story. Somewhat cheekily in the book, I compare that to earlier instances of human exceptionalism, of course, when Darwin first proposed that we weren't uniquely created by God but were actually related to all the other species, and when the heliocentric model first became established and was rejected by the Catholic Church.

Those analogies give color, I guess. Fundamentally, I think it is a defensible position, but it's an ideological position.

Speaker 1

You invoke this notion called the duck test: basically, if it looks like a duck and quacks like a duck, we should call it a duck. By extension, I guess you would call yourself a functionalist, which is the idea that it's not about the internal constitution or the mechanism, but about the function that it performs.

We can use this information metaphor to say, if we have an AI system over here that is cognizing and doing the same types of things, then we could reasonably infer that it's appropriate to use mentalistic language to describe it.

Christopher Summerfield

That's absolutely right. You're absolutely right to say that it's a functionalist perspective, and that is broadly my perspective. Once again, from a scientific standpoint, if it reasons like a human, then we may as well use the term reasoning.

But that doesn't imply a broader set of equivalents, right? It doesn't, for example, imply moral equivalence. It doesn't mean that the motivations or relationships we have with AI are similar to those we have with humans. Absolutely not. Of course, they're completely different.

But it does mean that when you put on your cognitive scientist hat and you're really just thinking about information processing, from that functionalist perspective, if it quacks like a duck, you may as well call it a duck.

Speaker 1

The anthropomorphism thing makes it a little trickier. In a film, if you see a robot peel its face away and all of a sudden you see that it's not human, that it's a robot, the intuition there is that it has a different mechanism. This is what John Searle was getting at when he was talking about the Chinese room argument. I read what you had to say about that.

I think Searle was saying that when you take a type of process and represent it in silicon as computation sans the machine—because we are biological machines, so we are causally embedded in the world—when we do things, there's this large light cone of low-level interactions that happen. I guess this is his notion of semantics. I think, Professor Summerfield, you subscribe to something called a distributional notion of semantics, which is that we can remove things from the physical world and recreate patterns of activity in silico and, for all intents and purposes, it would have the same meaning.

Christopher Summerfield

Yeah, I do subscribe to that view. As not only a cognitive scientist but also a neuroscientist, I'm uniquely aware that, while there are many differences between machine-learning systems and the computations that go on in the brain, there are also astonishing similarities. At the level of the algorithm, certainly—not clearly at the implementational level—neural networks don't tend to have many different types of synapses, and we don't have many different types. You don't have basket cells, fast inhibitory interneurons, and things like that.

But at the level of the neural network, there is a striking similarity. The most reasonable assumption to me is that there are broad, shared computational principles at work when you take networks of neurons that are wired up with some dense interconnection and are, for the most part, recurrent. We have to remember that the transformer is not a recurrent architecture, so it probably mimics what a recurrent architecture does; it uses tricks to mimic it. But for the most part, we're talking about recurrent networks.

And we know that because, for example, after optimization has been applied—and sometimes even before optimization has been applied—we know that there are striking similarities in the semantic representations that you can read out of those 2 classes of network, biological and artificial, by doing experiments. We know that you can go into the brains of monkeys or, if you have access to them, humans, via neuroimaging or whatever, and see patterns of representation that express themselves not just in terms of coding properties, but in terms of neural manifolds and neural geometry. They express themselves very much like those in the neural network.

The substrate is shared in some very loose sense; the behavior is shared in some perhaps not-so-loose sense. To me, it makes sense. Science is a puzzle: you get bits of information, and you try to come up with the most parsimonious explanation. For me, the most parsimonious explanation is that, by a sheer mixture of luck and trying enormously hard, we've got to a place where we've built something that is a bit like a brain, and lo and behold, it does stuff that is a bit like a brain.

That doesn't mean it does everything, and it also doesn't mean that it is like a human in the sense of how we should treat it or how we should think of it. But it does mean that the computations are most likely shared.

Speaker 1

I realize this is a difficult argument to make, and there were some scornful comments in your book about this, but some people still make the argument that it only appears to be reasoning and understanding, but it's not really. Is it possible that Chomsky could still be right in some sense?

His ideas, obviously, are rationalist, but it's this Platonistic idea, essentially, that the laws of nature have bestowed our brains with the secret functions that explain how the universe works. In a sense, he's quite similar to a lot of folks now. He's a computationalist. He doesn't subscribe to this causal graph thing, but he does think that the brain is a Turing machine and that we should do this recursive Merge-type stuff.

But is it possible that empiricism seems to work, but it's kind of like a pile of sand, and Chomsky would still be right if only it were possible to have the low-level stuff?

Christopher Summerfield

Yeah. I think what we'll find out—my guess is that the endpoint, when we look back after perhaps having figured this stuff out, will be that the dichotomy that was set up and that we fought about literally for millennia is actually a question of perspective. In a way, the rationalists, broadly construed, are right: reasoning is really important for computation. But what they were wrong about is how you acquire the ability to reason.

What we have learnt since 2019 is that the types of computations that you need to reason about the world can be learnt through large-scale parameter optimization, through function approximation, essentially through training a neural network. In that sense, Chomsky is not wrong that there are rules to language. Those rules need to be learnt. He was just wrong about how they got learnt.

And, of course, there's always a sleight of hand in saying, “Well, you're born. This is inborn,” because it really just begs the question of how it's inborn. Where does that gene that allows you to do recursion or Merge or whatever come from? What was the pressure that got it there?

I think there's a subtlety to an argument that is often not expanded on. Of course, we are born with a predisposition to learn language, and we know that that is not just an accident. Other species—even highly intelligent species like chimpanzees and gorillas, which are capable of really sophisticated forms of social interaction, political machinations, and so on—can't learn structured language. They can learn to communicate, but they can't learn to communicate in infinitely expressive sentences guided by lawful syntax.

The fact that they can't do that tells us that there is something special about our evolution. The question is, how do you explain that in the deep-learning framework? People often say, “Well, ChatGPT, of course, it was exposed to more language.” I think I have the analogy in my book: it was exposed to the same amount of language as if a single human were continually learning language from the middle of the last ice age or something like that. That's how much data it's exposed to.

But it's a false analogy, because we don't learn language like ChatGPT does. Language models are trained in a way that you might think of as almost Lamarckian. One generation of training—if you think of a training episode—whatever happens in that gets inherited by the next training episode. That's not how we work: my memories are not inherited by my kids. There's this fundamental disconnect.

We're Darwinian. The models are sort of—I don't know, I guess you could call them Lamarckian. You can't compare the amount of training that ChatGPT has to the amount of training that we have, because it's apples and oranges. What happens in a person's lifetime is like it's been guided, although it doesn't have the content: I live in Britain, but if my kids had been born in Japan, they would grow up speaking Japanese. It's been guided by all the other generations of learning, which inculcate this predisposition to learn language.

We never think of language models in that way. It's really meta-learning. And so Chomsky is right that we are born with priors, because those priors are the earlier cycles of Darwinian evolution: everything that went on before we were born, as individuals.

And so I think when we talk about data efficiency, and we try to make claims about data efficiency between biological and artificial intelligence, we need to be really specific about whether we're talking about phylogeny or ontogeny—in other words, evolution or development. Neither really works as a comparator. It's just more complicated.

Speaker 1

Is it possible that we're being deceived in some way, though? There are certainly computational limitations with neural networks. There are complexity limitations and learnability limitations. We know that there are certain types of things the networks can't do that we can do, and we are susceptible to anthropomorphization.

You mention this wonderful experiment where it was a cartoon of arrows interacting with each other, and humans interpreted them as agents. There was the grumpy bully agent. There was also the ELIZA machine, which was a very simple program that was quite sycophantic, and people took deep meaning from that. Is it possible that we're reading more into what's going on here than is actually the case?

5. Anthropomorphism Misleads Us

Christopher Summerfield

Well, it's definitely true that we are intrinsically prone to attribute much more elaborate forms of cognition to all other nonhuman agents, where simpler explanations may be available. Everyone who is a pet owner will be very familiar with this concept. It's the easiest thing in the world to attribute complex, human-like states to your cat, dog, or hamster when it may or may not be merited.

We know that people have been doing this for centuries. Psychologists know about the Clever Hans effect. Very famously, there was a performing horse that apparently could do mathematics—simple arithmetic—and it did so by repeatedly stamping its hoof the correct number of times to solve a sum. But, of course, it wasn't actually doing mathematics. What it was doing was checking whether its trainer gave it an unconscious signal that it should stop tapping.

So, of course, we are always prone to impute more complex thoughts, feelings, emotional states, or abilities to models. I don't deny that for a moment. When you look at today's frontier models, that may be going on. We may be thinking, “Oh, it's really my friend,” when actually it's not. But in terms of raw capability, the numbers are the numbers. The models are just really good, and there's no denying that.

Speaker 1

Yes.

Christopher Summerfield

They can't do everything. There are lots of things they can't do, and they're still not fully robust. But they are really good. They're not just Clever Hans.

Speaker 1

You said yourself something in the book that intrigued me: even cognitive scientists, neuroscientists, and psychologists don't really know the answer to the question, “What is thinking?” When we talk about these mentalistic properties and, of course, about intentionality—the agency involved in interpreting the intentions of cartoon arrows interacting with each other—Daniel Dennett, of course, coined the term “intentional stance.” Essentially, we need that to understand the world. It's a very complex place, and perhaps that's where some of these mentalistic properties come from.

Do you subscribe to an idea like that? I read this wonderful book called The Mind Is Flat by Nick Chater—

Christopher Summerfield

One of my favorites, yeah.

Speaker 1

A lot of these mentalistic properties, even in humans, are perhaps a bit of an illusion. What do you think?

Christopher Summerfield

I love that book. That book essentially argues that we draw heavily upon prior experience to formulate what we like. In other words, our preferences are a product not just of some internal value function that is different for everyone. You like apples more than oranges, and I like oranges more than apples, but it's actually due to our memories of past experiences.

You don't actually like apples more than oranges; you just think you do because you had an apple this morning. You're like, “I had an apple this morning. I must like apples more than oranges.” So it's this beautiful theory in which we essentially construct ourselves out of our own actions, and it can account for an astonishingly broad range of phenomena.

Do we do that? That's a scientific theory, but I think in our everyday interactions with other agents—animals and technology—we do the opposite. This is what Dennett says: we impute far more than is due, often. Your car fails to start in the morning, and you get cross with it as if it were just being stubborn. But, of course, there's no point in getting cross with it. That is an example of the intentional stance.

It is undoubtedly true that, for example, when interacting with models, people are very prone to attribute intentionality in the technical, philosophical sense of the word. In other words, they attribute that there is something it is like to be that thing. People are really prone to attribute that sense that they have some essence, some sense of what it is like to be themselves, to probably all forms of technology, but especially to AI because it can talk back.

People do that all the time. This is manifest in so many different ways. Of course, the types of interactions that people have with today's frontier models, starting with Blake Lemoine—who I talk about in the book and who famously argued, after his interactions with LaMDA, that it was sentient—are playing out today. Two of the top 100 most-visited websites in the world are companion applications. These are generative AI systems that are trained to behave as if they are your friend.

Why are they so popular? Because they're good at that. But they don't have to be that good, because people are really prone to think of them as if they were a person.

Speaker 1

Yes.

Christopher Summerfield

That is undoubtedly true. But I think it's possible to hold that view and to be cognizant of our predisposition to do this while also being sober about the capability. I think it's just a different question.

The capability question is: how do you get something that can solve simultaneous equations if they're posed in natural language? How do you do that? That is a problem we did not know how to answer in 2018. We know how to answer it now.

The system we implemented to solve that problem shares high-level computational principles with our best understanding of what the brain is doing. There are also a lot of things that are different, but it does share those principles. The most parsimonious explanation for how it can do this is that it's basically drawing on those same principles, in my view.

Speaker 1

Coming on to the alignment thing a little bit, you said, “Wouldn't it be amazing if we could have an artificial intelligence that would know what was right epistemically and also what was right ethically?”

6. The Agentic AI Threat

Christopher Summerfield

One of the things I'm most proud of about having written this book is that it is now more than a year and a half since I finished writing it, and the 3 things I'm worried about for the future that I discuss in the closing chapters are still the 3 things I'm worried about. So at least that has not gone stale, which, given the pace of change, is definitely not a given. I think it's quite surprising.

What are those 3 things? Number 1, I'm worried about the translation of systems that generate information that allows the user to behave in some way, giving way to systems that directly behave on the user's behalf. That's what we now call agentic AI. We were even calling it that then.

I'm worried about personalization: the extent to which models, instead of satisfying some general collective sense of what is right, can be tailored to everyone's individual sense of what is right. If you're an individual who has a set of beliefs and preferences that you're quite attached to, that sounds like quite a nice idea. But then you think about the many people in the world who have beliefs and preferences that you definitely wouldn't want reinforced, and you realize that personalized AI is exactly what it would do.

If you take agentic systems and personalized systems and put them together, and you imagine what deployment looks like, what it looks like is a vision that the companies have been talking about for several years now: personal AI. Everyone has personal AI, and it is a medium through which they interact with the world and that takes actions on their behalf. Probably, it is a conduit for information and resources, and it offers a layer of protection and so on.

What that really cashes out as is a world in which there is a sort of social economy amongst humans, but there is also a parallel social economy amongst the agents that we have and use to interact with the world. That might sound a bit sci-fi, but actually, I don’t think it’s all that sci-fi. It’s really not all that weird to imagine that we will interact with the world in a way that is technologically mediated, because that’s what we do already. Almost everything we do is technologically mediated.

It’s not weird to imagine that the technologies that we use to interact with the world, instead of being rule-based like they mostly are now, will be optimization-based. They’ll have minimal forms of agency. It’s like, why not? So you create this kind of multi-agent, parallel—if you like, it’s almost like a culture. You can think of it as a culture.

And the trouble is that we know that when you build a system and that system is complex and can interact in complex ways, then you get complex system effects. It can be nonlinear, have weird dynamics, and have feedback loops and so on. And that’s exactly what happened in that flash crash. Actually, there have been maybe dozens of flash crashes. The most famous one was in 2011, the one that I talk about in the book.

So you can think about what complex system dynamics emerge when we are all represented by AI. The reason why I think we should worry about that is because you can think of the norms that we’ve evolved socially and culturally as a set of principles that curtail those complex system dynamics. We have evolved in such a way that we generate a set of predispositions which create a set of constraints on our social interaction that stop, to a large extent, those runaway processes. They’re not perfect. Sometimes we go to war, and sometimes crazy stuff happens. But for the most part, particularly in reasonably small groups and for long periods, we can live in relatively stable, harmonious societies.

But the trouble is that the models won’t have those norms. There’s no reason why they should have them. And the question is, what are the constraints that prevent the same sort of weird runaway dynamics that might lead to flash crash-like events? I don’t think we have an answer for that. That’s what worries me.

Speaker 1

Yes. Designing in constraints would actually limit the technology in quite a strong way. It’s a really interesting thing to think about, though, because in the physical world, the constraints are quite strict. Language is a kind of virtual organism that supervenes on us and has more degrees of freedom, and this new type of AI technology that we’re inventing arguably has even more degrees of freedom. Constraining it is a real challenge.

Christopher Summerfield

Yeah, absolutely. I think the sheer— even if you had systems which were perfectly aligned, which of course is not an assumption any of us can reasonably make, but if you did, the sheer pace and volume of activity that AI can generate is not something that our systems are prepared for. Most systems operate under the assumption that there are reasonable frictions that prevent the system from collapsing.

A good example is the legal system. Many people know that it is possible, particularly depending on the jurisdiction, to engage in what is often called lawfare, so adversarial use of spurious legal challenges. There are certain jurisdictions where it’s strictly optimal to do that because the cost of defending yourself is so high that people will just capitulate, and you can make money. There are frictions that prevent most people from doing that. Most people don’t have legal training. Most people don’t know how to do that or the grounds on which you could do it. It’s a lot of work. You’ve got to file paperwork, and you need domain-specific knowledge.

If we remove those frictions so that you can just, with a few sentences, say, “Please do this,” and have a system that goes and does it, then you suddenly live in a very different world because lots and lots of people can do this. There are many other such examples.

Speaker 1

I was speaking with Conor Leahy about this, and he was talking about a phenomenon called the fog of war, which is that we slowly lose control through illegibility. You can imagine, based on what you said, that you have all of these agents, and even when a country is invaded or when some geopolitical event happens, the average person doesn’t understand why that is, because it’s the culmination of so many countervailing forces, and these systems are just very complex to understand.

So you can imagine a world that becomes so abstract. I also wanted to point out that this doesn’t require AI to be strong for all of these things to happen. Some people think of AI as a cultural technology, a bit like a library or something like that, and then there’s this almost doomer narrative that it’s agentic and this, that, and the other. But you don’t need it to be strong for all of these things to happen.

Christopher Summerfield

Yeah. The human analogy in AGI, of course, overlooks the fact that although collectively what we’ve done is astonishing, individually we’re actually extraordinarily vulnerable and just not all that good at life in general on our own. The classic example is you and a chimp on a desert island: my money’s on the chimp. Our strength is our ability to cooperate. Individually, we are not all that strong.

This notion of a lone intelligence that is like us, but much, much better, I think is a strange one. What we should actually worry about is the unexpected externalities that come from linking together lots of potentially weak systems to create something which is probably completely unlike us and unlike our culture and society, but which we can’t control.

I didn’t know the fog-of-war analogy. That’s very nice. But my favorite paper that talks about this recently is from David Duvenaud. He and others have written this really nice paper called “Gradual Disempowerment,” and it expresses a threat model which I have subscribed to for a really long time and which I talk about in the book. Broadly, it’s exactly that we gradually lock ourselves into the use of optimization-based technologies, and the complex system interactions between those systems sort of write us out of the equation.

The interesting thing about that analogy, which is a point not made either in the book or in David’s paper, is that in a way, it’s coextensive with what happens anyway. If you think about a corporation, the world we have created through hegemonic capitalism with large corporations, for example, in many ways, large corporations are more powerful than any one person. They run under their own imperatives, with their own rules, incentives, and dynamics. For many, they are so powerful that there is no one person who could stop them.

We sort of have a model for what this would look like. It’s just that, of course, in the case of large, complex systems like the corporation, the interactions are slow because they’re largely human-mediated. It’s email or Slack. Everything is such that humans are the cogs in the machine, sorry.

Speaker 1

Yes.

Christopher Summerfield

But in the case of AI, it’s going to happen at warp speed.

Speaker 1

I suppose it’s an interesting time. Just look at AlphaFold, for example. This technology can potentially be used to revolutionize science, but there are so many downsides as well, potentially. What downsides do you think we need to be most cautious about socially?

Christopher Summerfield

If you think about how many products work, of course, firms advertise products, and they do so by branding those products. Branding is a way of trying to get us to engage with something a bit like it was a human. Whether that something is the product itself or maybe the company—the brand—that branding is more or less successful.

Imagine a world in which everything was like that, but it could actually talk back to you and simulate all of the social and emotional types of interaction that we have with people that we care about. The milk in your fridge is like your best friend. This is a very strange world. Of course, that’s a silly example. The milk in the fridge is never going to be your best friend.

But there are already large numbers of people who are engaging with AI in ways that mimic the sorts of interactions they have with other people.

And this creates a whole bunch of vulnerabilities, and a lot of people have talked about risks to mental health and so on. We should be really aware of that, especially where vulnerable people or minors are concerned. But I think there’s another issue that’s talked about much less, and that is the degree to which this will give the organizations that build these systems power over people.

You said that AI increases our agency, and in a way, that is true. But I actually think that there’s also a really powerful sense in which the opposite is true, right? Just as access to social media gives you, in theory, access to lots and lots of information, which should be empowering, most people’s practical experience of it is that they spend a lot of time doing something they think is a bit stupid and would rather be doing something else.

Speaker 1

Yes, and I’m glad you brought that up. It’s a weird phenomenon. This comes into the labor market disruptions. Initially, for some people—certainly now, if you fire up Cursor, you can build a software business in a week—in that sense, it increases your agency.

But actually, everybody else has this capability, and the long-term, or even the medium-term, effect is that it sequesters your agency. It takes your agency away massively, and this is a huge problem.

Christopher Summerfield

Yeah.

Christopher Summerfield

Yeah, absolutely. You could call it a crisis of authenticity, right? You can see this broadly in society because our modes of interaction become so stylized that we lose that sense of authenticity. There are so many dependencies. We always have to present ourselves as being in line with the party line.

You just asked me something that I wasn’t able to answer because I have other dependencies. There is a loss of authenticity in our communications because, in a complex world, we represent many interests. That, I think, is a natural byproduct of our becoming part of the system.

I love this. My favorite metaphor for this is—I woke up in the middle of the night and it just hit me one night—have you—I don’t know if you remember Superman III? Superman III is a terrible movie, but there’s this wonderful scene. I think there’s this kind of giant computer that goes rogue in Superman III.

Speaker 1

Right.

Christopher Summerfield

There’s this wonderful scene where there’s a female character and the machine is just kind of waking up. She’s just walking past it, and the machine sucks her in. She gets stuck there. Then what the machine does is gradually put armor plating on her, replace her eyes with lasers, and basically turn her into a sort of automaton.

It’s a very compelling scene. I think I was terrified by it as a child, which is probably why I remember it. But that is a sort of metaphor for what is happening to us, right? We’re worried about the robots taking over or whatever, but in a way, it’s more like us being sucked into the machine.

We become part of the machine. Just like that poor character, we get turned into something we are not by technology. I don’t think this is a comment specifically about AI. I think this happens to every person who has to go to a press conference, or every person who has to represent their organization or a broader group of people.

You become part of that system, and it erodes your authenticity and, in a way, erodes your humanity.

Speaker 1

Yes. Last time I came to interview you, I went to Luciano Floridi directly afterwards, and his argument is similar about us becoming ensconced in the infosphere, and it changes our ontology.

Christopher Summerfield

Yeah.

Speaker 1

Perhaps you’re arguing more from an agential point of view, but I think it’s quite related.

Christopher Summerfield

Well, I think, as a psychologist, we have dramatically under-indexed on the extent to which what is good for us is actually about our agency, our control, and not about reward. Economics, psychology, and machine learning have all grown up with the notion that utility maximization is the fundamental framework for understanding behavior, and that’s expressed, of course, most prominently in machine learning through reinforcement learning.

But of course, this is not to deny that everyone needs to be warm and have enough to eat, right? Once those basic needs are satisfied—and even sometimes when they’re not—if you look at development and take a sideways view at a lot of both healthy and abnormal psychology, what you can see is that what people really care about is control. People need to understand, and by control I really mean, formally, your ability to have predictable influence on a system.

Speaker 1

Yep.

Christopher Summerfield

So, in machine learning, this often gets quantified as this wonderful notion of empowerment, right? The idea is that what we want to maximize is the mutual information between our actions and future states, for example, either immediate or delayed.

Speaker 1

That is agency, by the way.

Christopher Summerfield

And that is agency. I think that we really, really—if you think of kids, just 2 examples—the extent to which kids will explore the world, the extent to which they will take actions to try and understand, “What if I tap that thing, or what if I take my dinner and throw it on the floor? What’s going to happen? Oh, look, I have control.”

Or, “I cry. Oh, look, my dad’s going to come.” I have control. I can understand that system. That’s what they’re doing, right? Right through to forms of pathological control in adolescence and adulthood—too much control. You can see OCD, obsessive-compulsive disorder, as a need for too much control.

Anyway, I digress, but control is really, really important. When thinking about the impact of technology on our well-being, that conversation needs to be grounded in a robust understanding of how important it is to us to have a predictable influence on our world.

What a lot of AI, or a lot of technological penetration, actually does is make our actions unpredictable. It’s like the frustration that happens whenever you interact with a website that doesn’t quite work, or you get a computer-says-no answer, or there’s two-factor authentication but then there’s no internet and you’re like, “Ah.” It’s like you’ve lost control.

But that control, in the systems that we evolved for—the environment that we evolved for—is much more readily available.

Speaker 1

Isn't that fascinating? There have been studies—I’m sure you're familiar with this one—in which managers in an organization, in a study from the ’70s or something like that, had lower rates of heart disease because they had more power. The underlings would get sick much more often.

If you think about it, with social media and even with these chatbot platforms, I interviewed the team that built all the engagement-hacking algorithms. They were incredibly proud that their average session length was 19 minutes, and they talked all about how they would do model merging and send this response and this response to keep people hooked and keep them there for longer. In a sense, that is what dopamine hacking is about: giving random rewards, right? It's a disempowering thing along the lines you said, and that is the modus operandi for all technology now.

Christopher Summerfield

Yeah, absolutely. A variable reinforcement schedule is the most important—the most effective—way to train animals, including humans, and we are susceptible to it. The unpredictable nature of the reward engages us with the system and makes us come back because we want to control the system, right? We want to know, “How do I make the reward come?” And of course, if you can't, then you keep trying and trying and trying.

I mean, we live in a world in which people have a lot of liberty about how they spend their time, and I think that's as it should be. I don't think we should legislate against frivolity, right? If people want to spend a lot of time on TikTok, collectively, I understand that that's bad. But, for better or worse, we also live in a world in which that is permissible. We live in a country, at least, in which that is permissible.

Where I think we need to be cautious is that this kind of hacking may be undesirable. I might deem it undesirable, but collectively as a society, there are many things that are undesirable. Alcohol is also addictive, but I'm probably going to have a beer as soon as this is done, right? So we make those choices.

But I think that there are vulnerabilities. There are people who are uniquely vulnerable, where that kind of liberty to hack, if you like, spills over into something that can be really actively harmful and can lead people to self-harm. There have been tragic cases, as I'm sure you know, in which people have even taken their own life under an influence that came from an AI system with which they were interacting in this kind of companion mode.

Speaker 1

Do you think it makes sense to think of evolution as having a goal?

7. Evolution Favors Open Endedness

Christopher Summerfield

Probably not, right? There’s this great way of thinking about the blind process of evolution. There’s a paper that I really like, which draws upon the analogy of the blind process of evolution: it’s a selection mechanism that is not teleological, right? It doesn't have a purpose. It just happens.

The paper argues that we should think of training in neural networks in a similar way—that it's very blind. I think there is a fundamental difference between evolution and the way that optimization happens, put it that way. We could learn a lot about neural networks by thinking about the purposeless optimization that happens in evolution, basically.

Speaker 1

It's a really interesting topic for me. I was speaking with Kenneth Stanley the other day.

Christopher Summerfield

Yeah.

Speaker 1

He's done a lot of work about open-endedness, and of course—

Christopher Summerfield

Yeah.

Speaker 1

Tim Rocktäschel works at DeepMind.

Christopher Summerfield

Yeah.

Speaker 1

In Tim Rocktäschel's paper with Edward Hughes—

Christopher Summerfield

Yeah, yeah, the open-endedness paper.

Speaker 1

He said that an open-ended system is one that, from the perspective of an observer, produces a sequence of events which are learnable—

Christopher Summerfield

Yeah.

Speaker 1

—and novel.

Christopher Summerfield

Yeah. Yes, exactly. It's about learnability, isn't it? Joel Lehman—have you had Joel Lehman on the show?

Speaker 1

Yes.

Christopher Summerfield

Yeah. Joel has written really, really nicely about this. I largely share his view and am very close to Tim and Ed's view, which is that the world is open-ended, and optimizing for open-ended systems using well-specified, narrow optimization toward a narrow goal is just doomed to failure, right?

There is probably something really deep about the way the purposeless selection that happens in evolution confers robustness, because it doesn't precisely optimize for this narrow goal. Rather, what it creates is this astonishing heterogeneity, right?

Speaker 1

Yeah.

Christopher Summerfield

The optimization algorithms that we use are all completely opposite, right? They are basically tailored for homogeneity; heterogeneity is a bug.

Speaker 1

Yeah.

Christopher Summerfield

And that's why LLMs show mode collapse. It's why you get this Platonic Representation Hypothesis—the idea that we're gradually converging toward essentially one common, shared set of representations, right?

Evolution doesn't do that.

Speaker 1

Kenneth wrote this wonderful paper called The Fractured Entangled Representation Hypothesis.

Christopher Summerfield

Oh, I don't know that paper. Yeah.

Speaker 1

It was with Joel—I'm not sure if Joel was part of this—but he was involved in Why Greatness Cannot Be Planned. They did this thing called Picbreeder.

Christopher Summerfield

Hmm.

Speaker 1

It was like Flickr, where it was supervised by a diverse group of humans. The humans could pick interesting image generators, which were CPPNs—Compositional Pattern Producing Networks. You could create this phylogeny, and they talk about this concept called deception: the stepping stones that lead somewhere interesting don't resemble the interesting thing.

Humans have this sense of what's interesting because we seem to know the world well. With a few steps in the phylogeny, they found these pictures of butterflies and apples. When you do parameter sweeps on the networks, because they have such an abstract understanding of the objects, the apple would actually get bigger. One neuron would make it bigger; one neuron would make the stem swing.

If you train a neural network with stochastic gradient descent to do the same thing and you do parameter sweeps, it's just—it’s like spaghetti. It's all over the place.

So their hypothesis—and this seems like an obvious thing to say—is that if we could have a sparse representation that mirrored the world, then, because the knowledge would be evolvable, we could trust it with autonomy to make creative leaps, because it would do the right thing.

Christopher Summerfield

Mm. Yeah, that's amazing. I don't know about this paper. It sounds like I should read it. I mean, this idea that it's difficult to get places because the interim states are not highly valuable is a very old argument. This is the basis of Paley's watchmaker argument, right? How did we ever get the eye? You couldn't possibly evolve that; it's too complicated.

But those gradients must be there, right? The gradients are there.

Speaker 1

I have to say, Professor Summerfield, the prose—the way that you've written this book—is very impressive to me. It's one of the best-written pieces of writing I've ever seen. It occurred to me that you might have been deliberately making it so creative that it would be impossible to mistake for AI-generated content.

I don't know whether this is just because my standards are so low now because of shitification and all of that, but it was remarkable. Were you thinking that? Were you leaning into the creativity a little bit?

Christopher Summerfield

I love to write. I love to find new ways to explain things and to convey ideas, so that's a selfish pleasure for me. It didn't cross my mind that people might think I had used ChatGPT to write the book, but I guess, in hindsight, that's a very sensible way of thinking about it. But no, it was all me. That is mind-blowing.

Speaker 1

Professor Summerfield, it's been an absolute honor. Thank you so much.

Christopher Summerfield

Thank you.

How AI Learned to Talk and What It Means - Prof. Christopher Summerfield | BidClub