[BidClub_]
Machine Learning Street Talk · · 68 分钟

AI 如何学会说话,以及这意味着什么——Christopher Summerfield 教授

Christopher Summerfield

播客
TL;DR
  • Summerfield 的关键更新是:仅靠语言的监督学习,就能恢复足够多的现实结构,从而支撑一场智能对话,推翻了他直到 2015 年还持有的“接地”观点。 他曾认为“光靠阅读关于猫的描述,不可能知道猫是什么”,如今却称这一结果“或许是 21 世纪最惊人的科学发现”。对投资者而言,文字被证明是比预期丰富得多的世界模型基底,即使没有感官输入也成立。

  • 他的功能主义“鸭子测试”认为,只要一个系统以人类方式推理,就应当称其在推理,但这既不意味着道德等价,也不意味着动机相同或具有人类式关系。 Summerfield 提到,模型在形式数学和逻辑上的表现已经超过大多数受过教育的人类,生物与人工语义表征之间也存在相似性;底层基质和稳健性虽有差异,但前沿模型“不是只会读懂提示的 Clever Hans”。

  • 拿人类与模型比较数据效率,在结构上就是错的,因为生物学习遵循达尔文式机制,而模型训练“几乎像是拉马克式的”。 ChatGPT 接触到的语言,可能相当于一个人从上一个冰期中段以来所学习的语言总量,但每轮训练都会把收益传递下去,而“我的记忆不会遗传给我的孩子”。比较时必须区分进化与个体发育。

  • 核心系统性风险未必来自一个超人类智能体,而可能来自大量个性化、能力有限却以机器速度运行的智能体,进入原本围绕人类摩擦构建的制度。 Summerfield 想象了一个由个人 AI 构成的平行“社会经济”,其非线性反馈可能制造类似闪崩的失灵,即使每个系统都已对齐。法律战(lawfare)是最具体的案例:自动化可以移除目前阻止逐利滥用规模化的专业知识、文书工作和执行成本。

  • 个性化让对话式 AI 可能成为组织权力的通道,因为一个不断强化用户既有信念的系统,也可以模拟友谊并代表用户采取行动。 Summerfield 说,全球访问量最高的 100 个网站中已经有 2 个是陪伴应用;“你冰箱里的牛奶就像你最好的朋友”当然是故意说得荒诞,但它说明了这类系统可以模拟何种亲密关系。同一机制也会带来心理健康风险,尤其影响脆弱人群和未成年人。

  • AI 可能在短期内扩大人的能动性,却在长期把能动性收走。 主持人指出,Cursor 可以让一个人在一周内搭建一家软件公司,但最终的普遍可得性可能反而夺走能动性;Summerfield 认同这是一个重大问题。

  • 今天的优化范式可能与开放式世界并不匹配,因为它把异质性视为缺陷,而进化则通过无目的的多样性获得稳健性。 Summerfield 将狭窄目标与模式坍缩、以及向共同表征收敛联系起来;主持人以 Picbreeder 为例,说明有用的中间路径可能完全不像最终目标。主持人提出,更具可进化性的表征或许能支撑更具创造力、也更可信的自主性;Summerfield 说自己不了解所引论文,但认同“那些梯度一定存在”。

摘要 · 为研究而整理的核心内容

1. AI 最古老的争论,从哲学走到了键盘上

  • Summerfield 在 2023 年底完成了 These Strange New Minds,距离 ChatGPT 发布仅过去 12–14 个月。他希望把一场两极化争论落到计算层面:一端认为模型只是代码,另一端相信它具有人类级通用性;真正需要厘清的,是“思考”“推理”和“理解”究竟意味着什么。

  • 他的历史地图从 Plato 开始:不可观察的现实,只能通过“洞穴墙上的影子”推断;与之相对,Aristotle 强调经验。AI 在发展过程中重新演绎了这场理性主义与经验主义之争:智能究竟应当遵循显式规则,还是应当从学习中涌现?

  • 符号主义 AI 最初确实交出了成绩。Newell 和 Simon 于 1958 年推出的 Logic Theorist,从 Russell 和 Whitehead 的 Principia Mathematica 中证明了大量定理,并为其中若干定理找到了更优雅的证明;Summerfield 半开玩笑地称它为“第一个超级智能”。

  • 麻烦出现在系统离开整洁的数学、进入充满“奇怪例外”的现实之后。逻辑可以从可靠的基本命题推导复杂真理,但现实没有那么规整,于是旧式 AI 逐步让位于神经网络,最终走向深度学习革命。

2. 仅靠语言,模型学到的现实远超预期

  • 自然语言处理重演了更大的冲突。Chomsky 在 1958 年提出的挑战,是明确给出一套能够生成语法正确句子的规则;统计方法和神经网络方法则一再冲击这种以规则为先的解释。

  • 即使到了 2015 年,一个用 Shakespeare 训练的网络也只能生成“像 Shakespeare 的文字”,却“完全没有意义”。因此 Summerfield 当时认为,函数逼近加数据永远不够:真正的概念必须接地,因为“光靠阅读关于猫的描述,不可能知道猫是什么”。

  • 如今他承认,自己和“许许多多其他人”都错了。监督学习可以提取出一个受过教育的人识别智能对话所需的几乎全部结构,仅凭文字、无需感官经验即可做到——这“或许是 21 世纪最惊人的科学发现”,用他的话说,简直“令人震撼”。

3. 推理可以获得这个名称,但不会因此获得人格

  • Summerfield 所说的“例外主义”和“等价主义”只是光谱两端的漫画化版本。极端的人类中心主义者会把认知词汇保留给人类,即使模型在形式数学和逻辑上的表现超过大多数受过教育的人;他认为这种立场可以自洽,但主要是意识形态选择,而不是经验判断。

  • 他的功能主义回答是鸭子测试:“如果它叫起来像鸭子,那就不妨称它为鸭子。”把模型的行为称为推理,并不意味着它拥有与人类等价的道德地位、相同的动机或人类式关系,也不能据此推出系统应当如何被对待。

  • 作为神经科学家,他清楚看到实现层面的巨大差异:大脑包含多种突触和细胞类型,而 transformer 并不是循环架构,只是通过各种“技巧”模拟循环。不过,实验显示,经过优化的生物网络和人工网络之间,存在显著相似的语义几何结构与神经流形。

  • 主持人援引 Searle 的中文房间,追问硅基系统是否只是在复制表层模式,而没有具身语义。Summerfield 的简约解释是,稠密网络已经捕捉到了大脑共享的广泛计算原则:“我们造出了某种有点像大脑的东西,结果它果然做出了某些有点像大脑的事情。”

4. 学会的规则与遗传的先验,调和了两套竞争理论

  • Summerfield 预计,这组古老二分法最终会被看作视角差异。理性主义者认为推理和规则重要,这一点是对的;但他们错在认为规则必须先天存在:自 2019 年左右以来,大规模参数优化已经显示,推理程序本身也可以通过函数逼近学出来。

  • 因此,Chomsky 认为语言具有规则,可能并没有错;他错在规则的来源。把递归或 Merge 宣布为先天属性,只是把问题往前推了一步,最终仍要解释是什么进化压力塑造了这种倾向。

  • 人类显然拥有这类先验:chimpanzee 和 gorilla 能够进行复杂的社会与政治互动,却无法学会无限表达、遵循规则的语言。人类儿童的学习受到更早达尔文式世代的引导,但具体内容仍然灵活——出生在日本的孩子可以学会日语。

  • 因此,ChatGPT 接收到相当于一个人从上一个冰期中段以来所接触的语言量,这个说法是“错误类比”。模型学习几乎可以被看作拉马克式的,因为每一轮都会继承上一轮的收益;人类则遵循达尔文式机制,“我的记忆不会遗传给我的孩子”。数据效率比较必须把系统发育与个体发育分开,而这两者都无法与模型训练完全对应。

5. 拟人化扭曲的是关系,不是能力数字

  • 主持人最有力的质疑是:人类会在移动的箭头上看见能动性,会在 ELIZA 身上寻找意义,也可能把计算限制误认为隐藏的心智。Summerfield 完全认同这种偏差:宠物主人会把复杂心理状态投射给动物,而 Clever Hans 看似会计算,实际上只是读懂了训练者无意识的信号。

  • Nick Chater 的 The Mind Is Flat 认为,人们会根据自己过去的行为记忆来构造偏好;Dennett 的意向立场则描述了相反方向的投射——例如把一辆汽车当成固执的对象。会说话的 AI 强化了这种本能,从 Blake Lemoine 声称 LaMDA 具有意识,到陪伴应用进入全球访问量最高的 100 个网站中的 2 个,都是如此。

  • Summerfield 将依恋与表现分开。用户可能错误地相信陪伴 AI“真的是我的朋友”,当前模型也仍然不完美、缺乏稳健性;但它们确实可以解出用自然语言提出的联立方程。“数字就是数字”——这些系统确实有能力,“不是只会读懂提示的 Clever Hans”。

6. 个性化智能体可能制造机器速度的社会经济

  • Summerfield 在一年半多前完成书稿时提出的担忧,如今仍然成立:生成信息的系统正在变成替用户行动的系统;系统会围绕个人信念和偏好进行个性化;当个人智能体代表所有人时,部署将产生复杂的动态。

  • 个性化听起来很有吸引力,直到它被用于强化那些“你肯定不希望被强化”的信念。再将个性化与能动性结合,个人 AI 就会成为信息、资源、保护和行动的通道,在人类社会之外,形成一个由智能体彼此交互的平行“社会经济”。

  • 即使智能体已经完美对齐,也可能凭借数量和速度压垮系统。法律战说明了缺失的摩擦:法律知识、文书工作和执行成本目前限制着虚假诉求,但如果用户只需说一句“请做这件事”,那么局部有利、整体有害的行为就可能被规模化执行。

  • 人类规范会对失控的社会动力形成部分约束;智能体没有天然等价物。Summerfield 回忆起著名的 2011 年闪崩,并称类似事件可能已经发生过几十次;他警告,即使没有任何智能体接近强人工智能,弱系统之间的非线性反馈也可能制造类似事件。

7. 技术可以扩大机会,也可以剥夺控制权

  • 主持人把由此产生的不可读性称为“战争迷雾”;Summerfield 则将其与 David Duvenaud 及其合作者提出的 Gradual Disempowerment 联系起来。人类可能被锁进由优化系统主导的环境,而这些系统的相互作用会“把我们从方程里写出去”,类似拥有自身意志的公司——只是公司依靠缓慢的人类邮件和 Slack 运转,而 AI 以“曲速般的速度”运行。

  • 主持人认为,Cursor 之类的工具可以让一个人在一周内搭建一家软件公司,但最终的普遍可得性可能反而把能动性收走。Summerfield 认同,能动性的丧失是一个重大问题。

  • 固化的交互模式和对组织的依赖,可能侵蚀真实性。Summerfield 借用 Superman III 的隐喻,反转了通常的接管故事:机器把一个女人吸进去,给她装上盔甲和激光眼,把她变成自动机——也就是“我们被吸进机器”。

  • 他的心理学修正是:福祉取决于控制,而不只是奖励或效用。形式上,赋能意味着对未来状态产生可预测的影响,即行动与未来状态之间的互信息。孩子会通过把晚餐扔掉或哭泣来测试这种影响;强迫症可能体现为病理性的过度控制,而网站故障和双因素认证失败则带来相反的体验。

8. 开放式进化暴露了狭窄优化的弱点

  • Summerfield 否认进化存在某种目标。进化的选择是盲目的、非目的论的,可以作为没有预设终点的优化类比。

  • 主持人借鉴 Tim Rocktäschel 和 Edward Hughes,描述了一种开放式系统:它会产生观察者认为既可学习又新颖的事件。Summerfield 认同开放性与可学习性有关,也表示自己与 Tim、Ed 和 Joel Lehman 的观点一致;他认为,朝着狭窄目标进行狭窄优化“注定失败”。

  • 进化产生“惊人的异质性”,或许正因为它不精确瞄准某个结果,反而获得了稳健性。当前的优化则把异质性视为缺陷,推动 LLM 发生模式坍缩,也呼应了“柏拉图表征假说”——系统正在收敛到一种共同的表征结构。“进化不会这么做。”

  • 主持人用 Picbreeder 把问题具体化:人类通过并不像最终目标的中间路径,最终筛选出了蝴蝶和苹果;而单个网络组件能够控制苹果大小、果梗等可解释特征。相应的梯度训练网络看起来却像“一团意大利面”。Summerfield 说自己不了解所引用的表征论文,认为这个想法值得阅读;他回应道:“那些梯度一定存在。”

Christopher Summerfield

Superman III is a terrible movie, but there’s this wonderful scene. I think there’s this kind of giant computer that goes rogue in Superman III, and there’s this wonderful scene where there’s a female character. The machine is just kind of waking up, and she’s just walking past it, and the machine kind of sucks her in. She gets stuck there.

Then what the machine does is gradually put armor plating on her, replace her eyes with lasers, and basically turn her into a sort of automaton. It’s a very compelling scene. I think I was terrified by it as a child, which is probably why I remember it.

That is a sort of metaphor for what is happening to us, right? We’re worried about the robots taking over or whatever, but in a way, it’s more like us being sucked into the machine. We become part of it, just like that poor character. We get turned into something we are not. You become part of that system, and it erodes your authenticity and, in a way, erodes your humanity.

People often say, well, ChatGPT, of course, was exposed to more. I think I have the analogy in my book. It’s exposed to the same amount of language as if a single human were continually learning language from the middle of the last Ice Age or something like that. That’s how much data it’s exposed to.

But it’s a false analogy, because we don’t learn language like ChatGPT does. Language models are trained in a kind of— you might think of it as almost a Lamarckian way. One generation of training, if you think of a training episode, whatever happens in that gets inherited by the next training episode. That’s not how we work. My memories are not inherited by my kids. There’s this fundamental disconnect. We’re Darwinian; the models are, I guess you could call them, Lamarckian.

Speaker 1

So we’re here in Oxford today to speak with Professor Christopher Summerfield. He’s just written this book called These Strange New Minds: How AI Learned to Talk and What That Means. He spoke about the history of artificial intelligence and how the allure of AI is to build a machine that can know what is true and what is right.

Christopher Summerfield

Imagine a world in which everything was like that, but it could actually talk back to you, and it could simulate all of the social and emotional types of interaction that we have with people we care about. The milk in your fridge is like your best friend, right? This is a very strange world in which— of course, that’s a silly example. The milk in the fridge is never going to be your best friend.

But there are already large numbers of people who are engaging with AI in ways that mimic the sorts of interactions they have with other people. I thought that grounding would require sensory signals. You can’t know what a cat is just by reading about cats in books. You need to actually see a cat.

But it turned out I was wrong, and so were many, many other people. That is, to my mind, perhaps the most astonishing scientific discovery of the 21st century: supervised learning is so good that you can actually learn about almost everything you need to know about the nature of reality, at least to have a conversation that every educated human would say is an intelligent conversation, without ever having any sensory knowledge of the world, just through words. That is mind-blowing.

Speaker 2

This podcast is supported by Google. Hey, everyone. David here, one of the product leads for Google Gemini. Check out Veo 3, our state-of-the-art AI video generation model in the Gemini app, which lets you create high-quality eight-second videos with native audio generation. Try it with a Google AI Pro plan or get the highest access with the Ultra plan. Sign up at gemini.google to get started and show us what you create.

Speaker 3

I'm Benjamin Crouzier. I'm starting an AI research lab called Tufa Labs. It is funded from past ventures involving machine learning. So we're a small group of highly motivated and hardworking people, and the main thread that we are going to do is trying to make models that reason effectively and long term, trying to do AGI research. So one of the big advantage is because we're early, there's going to be high freedom and high impact as someone new at Tufa Labs. You can check out positions at tufalabs.ai.

Speaker 1

So, Professor Summerfield, I have to congratulate you on this book. Your previous book was my favorite book that I’ve ever read in AI. It’s up there with Melanie Mitchell’s book. Melanie Mitchell reviewed your new book as well.

Christopher Summerfield

She did.

Speaker 1

So…

Christopher Summerfield

Very generously. Yeah.

Speaker 1

I’m a big fan of Melanie. You’ve been writing this for a couple of years, and, of course, you explained in the afterword that it takes quite a long time to get these things into publication, while the space is moving very, very quickly. Can you give us a bit of an elevator pitch of the book?

Christopher Summerfield

Yeah, sure. The book was actually finished at the end of 2023, so cast your mind back to the medieval period of AI, if you’d like—12 or 14 months after ChatGPT had just been released.

1. The Cognitive Status Debate

The idea of the book was that, at that time, and I guess to a large extent still today, there was considerable debate over the cognitive status of these strange new minds that we seem to have created and are now increasingly interacting with. The debate that I heard, both at academic conferences and down the pub, was: should we think of these things as actually a bit like us? Are they thinking? Are they reasoning? Are they understanding?

Of course, this very quickly became a highly polarized debate. On the one hand, a bunch of people vehemently rejected the idea that these tools could ever be anything like us. They’re just computer code, which is of course true. On the other hand, you had people who were absolutely astonished not just by the capability, but by the pace of progress, and thought we really were finally on course to build something that was as generally competent as humans.

This debate was playing out, and I thought, well, this debate isn’t really grounded in the language of cognition. I don’t hear that language being used to scaffold the debate. The debate was being had by people who cared deeply about the issue, but who weren’t trained in a grounded, computational sense of what it actually means to think or to understand something.

As a cognitive scientist who has done a lot of work in AI, I was probably quite well placed to talk about that. So that was part one. I’ve also been very, very interested in the implications of AI for society for the past 5 years.

I was working on that problem when I was at DeepMind, and we were doing work to try to understand how AI could be used to intervene directly in society and the economy to help people find agreement. When I wrote the book, I was just about to move to the AI Safety Institute in the UK government to work more on that.

I had an understanding of the landscape of deployment risks and was thinking about how AI might change the way that we live our lives. I thought that, by putting those things together, I had enough of a unique perspective to write a book about it. So that’s what I did.

Speaker 1

The discourse is quite fractured, and you speak about this in great detail. You speak about the hypers, the anti-hypers, the safety hypers, and so on. Early on in the book, you trace this back to 2 intellectual threads going back to the ancient Greeks: Aristotle and Plato, basically empiricism and rationalism. Can you sketch that out?

2. Reasoning Versus Learning

Christopher Summerfield

Yeah, sure. The history of AI has itself repeated an ancient philosophical debate about whether the fundamental nature of building a mind, including our mind, is fundamentally about learning from experience or about reasoning, particularly reasoning over latent or unobservable states.

That reasoning over unobservable states, of course, traces back to Plato. It’s the idea that everything is fundamentally unobservable. We just get the shadows on the cave wall or the light on the retina, and we have to impute what’s there.

The corresponding view might trace back to Aristotle: the idea that everything comes from experience. The history of AI, of course, was that very debate playing out in the workshop, so to speak—or at least on the keyboard.

On the one hand, originally, good old-fashioned AI was structured around the idea that we sort of know what…

How to work out what is true. The reason we know how to work out what is true is because we have a long tradition back through positivism and early theories of reasoning to Boole and even Leibniz before that. The idea is that you can use logic to work out what is true. It is unassailably true that if I say all men are Greek and Aristotle is a man, then Aristotle is Greek, right? That is just true by definition.

That seemed like a really sensible way to build AI. You put in those primitives, crank the handle, and if you've got enough computational power, you can derive really complex things. And it worked. In the 1950s, Newell and Simon built the Logic Theorist, which I like to say was the first superintelligence, in 1958. It was an AI system able to prove many of the theorems in Russell and Whitehead's Principia Mathematica, which was already a feat, and it was able to find more elegant solutions to many of those theorems.

That's astonishing. Initially, it seemed like this reasoning approach worked. Then, as the problems we tried to tackle with this approach moved from very abstract, clean problems about maths and logic to problems in the real world, we ran into a fundamental problem: the real world just isn't all that clean, nice, and neat in the way that reasoning problems are designed to be. The world is full of weird exceptions that aren't fundamentally amenable to analysis with logic.

So you had this other corresponding approach, which was the learning approach, or the empiricist approach. That's where neural networks and the deep learning revolution ultimately came from.

Speaker 1

Isn't it a crazy time to be alive, though? I interviewed the CEO of one of the largest companion-bot platforms, and in the comments section there was a lot of negativity. You actually mentioned, I think in your afterword, that it seems strange to us now that we would want to have a relationship with an AI companion, and maybe we might revise that belief in a few years' time.

More broadly, you said in your book that language is basically the biggest gift that has ever been given to us. It allows us to acquire knowledge and communicate it, and it survives many generations. I guess the Rubicon moment with this technology maturing was ChatGPT. That changed everything in November 2022. Sketch that out for me.

3. Language Learns Without Senses

Christopher Summerfield

The history of NLP has been told many times, probably by people more qualified than me. But we talked earlier about this back-and-forth between learning and reasoning, and in the history of NLP, exactly the same question played out. NLP, or natural language processing, is a subfield of AI.

In the more general symbolic AI movement, the early models were basically attempts to define the computations that lead to the generation of valid sentences. That's essentially the gauntlet that Chomsky lays down in his 1958 book. There are a set of rules which, if you could just apply them all lawfully, would allow for the generation of sentences that obey the rules we would all understand as making a valid sentence.

So, syntax. Chomsky was mainly concerned with English, of course, so he was worried about English syntax. That movement was then challenged by statistical approaches, just as neural networks came along in the wider field. It went back and forth repeatedly.

When the deep learning revolution happened, by 2015 we had models that you could train on the complete works of Shakespeare, and they could generate something that looked a lot like Shakespeare, but it didn't make any sense. Even when the deep learning revolution was in full swing, most people, including me, thought there was no way that the mere application of powerful function approximation and lots of data was going to solve this problem.

I didn't believe that to be true. I thought, like many other people, that you would need grounding and sensory signals. You can't know what a cat is just by reading about cats in books. You need to actually see a cat. But it turned out I was wrong, and so were many, many other people.

To my mind, that is an absolutely astonishing discovery—perhaps the most astonishing scientific discovery of the 21st century. Supervised learning is so good that you can actually learn almost everything you need to know about the nature of reality, at least to have a conversation that every educated human would say is an intelligent conversation, without ever having any sensory knowledge of the world, just through words.

That is mind-blowing, and I think it changes the way we think about many, many things. It certainly changes how I think about things.

Speaker 1

One big theme in the book is this dichotomy between equivalentists and exceptionalists. Some people argue that humans are exceptional and that the kinds of cognition that language models engage in are not really in the same category.

4. AI And Human Equivalence

Christopher Summerfield

That distinction is a cartoon. Of course, everyone has a different view about the relationship between AI and humans, or biological intelligence in general. The evidence clearly admits a spectrum of different views, but I found it useful in the book to cartoon two extremes of that continuum.

At one end, you have people who probably just ideologically reject the idea that something non-human could ever be referred to using the same vocabulary we apply to a human, whatever that system is doing behaviorally or cognitively. Today's models are clearly capable of reasoning at levels beyond the capability of most, even educated, humans. Certainly when it comes to formal problems like maths and logic.

They can reason like a human, but there are people who fundamentally think we shouldn't think of that as reasoning because we should circumscribe the definition of reasoning as something that humans do. That is a stance which I think is not really about the empirical evidence, although some people construe it that way by saying, "The models aren't actually that good at reasoning," which I think was a hard-to-defend view even in 2023. Now it's probably an even harder-to-defend view.

I think it comes from a place of radical humanism. It is a desire to really ring-fence a set of cognitive concepts and think of them as uniquely human. For people who care about humans, which includes me, I can see why that's really important. But what it does lead you down the road of is a refusal to ever see the cognition an AI engages in and the cognition a human engages in as comparable, even when their capabilities are clearly matched.

That's what I call exceptionalists, because in a way they're espousing a view of human exceptionalism: humans are special and different, end of story. Somewhat cheekily in the book, I compare that to earlier instances of human exceptionalism, of course, when Darwin first proposed that we weren't uniquely created by God but were actually related to all the other species, and when the heliocentric model first became established and was rejected by the Catholic Church.

Those analogies give color, I guess. Fundamentally, I think it is a defensible position, but it's an ideological position.

Speaker 1

You invoke this notion called the duck test: basically, if it looks like a duck and quacks like a duck, we should call it a duck. By extension, I guess you would call yourself a functionalist, which is the idea that it's not about the internal constitution or the mechanism, but about the function that it performs.

We can use this information metaphor to say, if we have an AI system over here that is cognizing and doing the same types of things, then we could reasonably infer that it's appropriate to use mentalistic language to describe it.

Christopher Summerfield

That's absolutely right. You're absolutely right to say that it's a functionalist perspective, and that is broadly my perspective. Once again, from a scientific standpoint, if it reasons like a human, then we may as well use the term reasoning.

But that doesn't imply a broader set of equivalents, right? It doesn't, for example, imply moral equivalence. It doesn't mean that the motivations or relationships we have with AI are similar to those we have with humans. Absolutely not. Of course, they're completely different.

But it does mean that when you put on your cognitive scientist hat and you're really just thinking about information processing, from that functionalist perspective, if it quacks like a duck, you may as well call it a duck.

Speaker 1

The anthropomorphism thing makes it a little trickier. In a film, if you see a robot peel its face away and all of a sudden you see that it's not human, that it's a robot, the intuition there is that it has a different mechanism. This is what John Searle was getting at when he was talking about the Chinese room argument. I read what you had to say about that.

I think Searle was saying that when you take a type of process and represent it in silicon as computation sans the machine—because we are biological machines, so we are causally embedded in the world—when we do things, there's this large light cone of low-level interactions that happen. I guess this is his notion of semantics. I think, Professor Summerfield, you subscribe to something called a distributional notion of semantics, which is that we can remove things from the physical world and recreate patterns of activity in silico and, for all intents and purposes, it would have the same meaning.

Christopher Summerfield

Yeah, I do subscribe to that view. As not only a cognitive scientist but also a neuroscientist, I'm uniquely aware that, while there are many differences between machine-learning systems and the computations that go on in the brain, there are also astonishing similarities. At the level of the algorithm, certainly—not clearly at the implementational level—neural networks don't tend to have many different types of synapses, and we don't have many different types. You don't have basket cells, fast inhibitory interneurons, and things like that.

But at the level of the neural network, there is a striking similarity. The most reasonable assumption to me is that there are broad, shared computational principles at work when you take networks of neurons that are wired up with some dense interconnection and are, for the most part, recurrent. We have to remember that the transformer is not a recurrent architecture, so it probably mimics what a recurrent architecture does; it uses tricks to mimic it. But for the most part, we're talking about recurrent networks.

And we know that because, for example, after optimization has been applied—and sometimes even before optimization has been applied—we know that there are striking similarities in the semantic representations that you can read out of those 2 classes of network, biological and artificial, by doing experiments. We know that you can go into the brains of monkeys or, if you have access to them, humans, via neuroimaging or whatever, and see patterns of representation that express themselves not just in terms of coding properties, but in terms of neural manifolds and neural geometry. They express themselves very much like those in the neural network.

The substrate is shared in some very loose sense; the behavior is shared in some perhaps not-so-loose sense. To me, it makes sense. Science is a puzzle: you get bits of information, and you try to come up with the most parsimonious explanation. For me, the most parsimonious explanation is that, by a sheer mixture of luck and trying enormously hard, we've got to a place where we've built something that is a bit like a brain, and lo and behold, it does stuff that is a bit like a brain.

That doesn't mean it does everything, and it also doesn't mean that it is like a human in the sense of how we should treat it or how we should think of it. But it does mean that the computations are most likely shared.

Speaker 1

I realize this is a difficult argument to make, and there were some scornful comments in your book about this, but some people still make the argument that it only appears to be reasoning and understanding, but it's not really. Is it possible that Chomsky could still be right in some sense?

His ideas, obviously, are rationalist, but it's this Platonistic idea, essentially, that the laws of nature have bestowed our brains with the secret functions that explain how the universe works. In a sense, he's quite similar to a lot of folks now. He's a computationalist. He doesn't subscribe to this causal graph thing, but he does think that the brain is a Turing machine and that we should do this recursive Merge-type stuff.

But is it possible that empiricism seems to work, but it's kind of like a pile of sand, and Chomsky would still be right if only it were possible to have the low-level stuff?

Christopher Summerfield

Yeah. I think what we'll find out—my guess is that the endpoint, when we look back after perhaps having figured this stuff out, will be that the dichotomy that was set up and that we fought about literally for millennia is actually a question of perspective. In a way, the rationalists, broadly construed, are right: reasoning is really important for computation. But what they were wrong about is how you acquire the ability to reason.

What we have learnt since 2019 is that the types of computations that you need to reason about the world can be learnt through large-scale parameter optimization, through function approximation, essentially through training a neural network. In that sense, Chomsky is not wrong that there are rules to language. Those rules need to be learnt. He was just wrong about how they got learnt.

And, of course, there's always a sleight of hand in saying, “Well, you're born. This is inborn,” because it really just begs the question of how it's inborn. Where does that gene that allows you to do recursion or Merge or whatever come from? What was the pressure that got it there?

I think there's a subtlety to an argument that is often not expanded on. Of course, we are born with a predisposition to learn language, and we know that that is not just an accident. Other species—even highly intelligent species like chimpanzees and gorillas, which are capable of really sophisticated forms of social interaction, political machinations, and so on—can't learn structured language. They can learn to communicate, but they can't learn to communicate in infinitely expressive sentences guided by lawful syntax.

The fact that they can't do that tells us that there is something special about our evolution. The question is, how do you explain that in the deep-learning framework? People often say, “Well, ChatGPT, of course, it was exposed to more language.” I think I have the analogy in my book: it was exposed to the same amount of language as if a single human were continually learning language from the middle of the last ice age or something like that. That's how much data it's exposed to.

But it's a false analogy, because we don't learn language like ChatGPT does. Language models are trained in a way that you might think of as almost Lamarckian. One generation of training—if you think of a training episode—whatever happens in that gets inherited by the next training episode. That's not how we work: my memories are not inherited by my kids. There's this fundamental disconnect.

We're Darwinian. The models are sort of—I don't know, I guess you could call them Lamarckian. You can't compare the amount of training that ChatGPT has to the amount of training that we have, because it's apples and oranges. What happens in a person's lifetime is like it's been guided, although it doesn't have the content: I live in Britain, but if my kids had been born in Japan, they would grow up speaking Japanese. It's been guided by all the other generations of learning, which inculcate this predisposition to learn language.

We never think of language models in that way. It's really meta-learning. And so Chomsky is right that we are born with priors, because those priors are the earlier cycles of Darwinian evolution: everything that went on before we were born, as individuals.

And so I think when we talk about data efficiency, and we try to make claims about data efficiency between biological and artificial intelligence, we need to be really specific about whether we're talking about phylogeny or ontogeny—in other words, evolution or development. Neither really works as a comparator. It's just more complicated.

Speaker 1

Is it possible that we're being deceived in some way, though? There are certainly computational limitations with neural networks. There are complexity limitations and learnability limitations. We know that there are certain types of things the networks can't do that we can do, and we are susceptible to anthropomorphization.

You mention this wonderful experiment where it was a cartoon of arrows interacting with each other, and humans interpreted them as agents. There was the grumpy bully agent. There was also the ELIZA machine, which was a very simple program that was quite sycophantic, and people took deep meaning from that. Is it possible that we're reading more into what's going on here than is actually the case?

5. Anthropomorphism Misleads Us

Christopher Summerfield

Well, it's definitely true that we are intrinsically prone to attribute much more elaborate forms of cognition to all other nonhuman agents, where simpler explanations may be available. Everyone who is a pet owner will be very familiar with this concept. It's the easiest thing in the world to attribute complex, human-like states to your cat, dog, or hamster when it may or may not be merited.

We know that people have been doing this for centuries. Psychologists know about the Clever Hans effect. Very famously, there was a performing horse that apparently could do mathematics—simple arithmetic—and it did so by repeatedly stamping its hoof the correct number of times to solve a sum. But, of course, it wasn't actually doing mathematics. What it was doing was checking whether its trainer gave it an unconscious signal that it should stop tapping.

So, of course, we are always prone to impute more complex thoughts, feelings, emotional states, or abilities to models. I don't deny that for a moment. When you look at today's frontier models, that may be going on. We may be thinking, “Oh, it's really my friend,” when actually it's not. But in terms of raw capability, the numbers are the numbers. The models are just really good, and there's no denying that.

Speaker 1

Yes.

Christopher Summerfield

They can't do everything. There are lots of things they can't do, and they're still not fully robust. But they are really good. They're not just Clever Hans.

Speaker 1

You said yourself something in the book that intrigued me: even cognitive scientists, neuroscientists, and psychologists don't really know the answer to the question, “What is thinking?” When we talk about these mentalistic properties and, of course, about intentionality—the agency involved in interpreting the intentions of cartoon arrows interacting with each other—Daniel Dennett, of course, coined the term “intentional stance.” Essentially, we need that to understand the world. It's a very complex place, and perhaps that's where some of these mentalistic properties come from.

Do you subscribe to an idea like that? I read this wonderful book called The Mind Is Flat by Nick Chater—

Christopher Summerfield

One of my favorites, yeah.

Speaker 1

A lot of these mentalistic properties, even in humans, are perhaps a bit of an illusion. What do you think?

Christopher Summerfield

I love that book. That book essentially argues that we draw heavily upon prior experience to formulate what we like. In other words, our preferences are a product not just of some internal value function that is different for everyone. You like apples more than oranges, and I like oranges more than apples, but it's actually due to our memories of past experiences.

You don't actually like apples more than oranges; you just think you do because you had an apple this morning. You're like, “I had an apple this morning. I must like apples more than oranges.” So it's this beautiful theory in which we essentially construct ourselves out of our own actions, and it can account for an astonishingly broad range of phenomena.

Do we do that? That's a scientific theory, but I think in our everyday interactions with other agents—animals and technology—we do the opposite. This is what Dennett says: we impute far more than is due, often. Your car fails to start in the morning, and you get cross with it as if it were just being stubborn. But, of course, there's no point in getting cross with it. That is an example of the intentional stance.

It is undoubtedly true that, for example, when interacting with models, people are very prone to attribute intentionality in the technical, philosophical sense of the word. In other words, they attribute that there is something it is like to be that thing. People are really prone to attribute that sense that they have some essence, some sense of what it is like to be themselves, to probably all forms of technology, but especially to AI because it can talk back.

People do that all the time. This is manifest in so many different ways. Of course, the types of interactions that people have with today's frontier models, starting with Blake Lemoine—who I talk about in the book and who famously argued, after his interactions with LaMDA, that it was sentient—are playing out today. Two of the top 100 most-visited websites in the world are companion applications. These are generative AI systems that are trained to behave as if they are your friend.

Why are they so popular? Because they're good at that. But they don't have to be that good, because people are really prone to think of them as if they were a person.

Speaker 1

Yes.

Christopher Summerfield

That is undoubtedly true. But I think it's possible to hold that view and to be cognizant of our predisposition to do this while also being sober about the capability. I think it's just a different question.

The capability question is: how do you get something that can solve simultaneous equations if they're posed in natural language? How do you do that? That is a problem we did not know how to answer in 2018. We know how to answer it now.

The system we implemented to solve that problem shares high-level computational principles with our best understanding of what the brain is doing. There are also a lot of things that are different, but it does share those principles. The most parsimonious explanation for how it can do this is that it's basically drawing on those same principles, in my view.

Speaker 1

Coming on to the alignment thing a little bit, you said, “Wouldn't it be amazing if we could have an artificial intelligence that would know what was right epistemically and also what was right ethically?”

6. The Agentic AI Threat

Christopher Summerfield

One of the things I'm most proud of about having written this book is that it is now more than a year and a half since I finished writing it, and the 3 things I'm worried about for the future that I discuss in the closing chapters are still the 3 things I'm worried about. So at least that has not gone stale, which, given the pace of change, is definitely not a given. I think it's quite surprising.

What are those 3 things? Number 1, I'm worried about the translation of systems that generate information that allows the user to behave in some way, giving way to systems that directly behave on the user's behalf. That's what we now call agentic AI. We were even calling it that then.

I'm worried about personalization: the extent to which models, instead of satisfying some general collective sense of what is right, can be tailored to everyone's individual sense of what is right. If you're an individual who has a set of beliefs and preferences that you're quite attached to, that sounds like quite a nice idea. But then you think about the many people in the world who have beliefs and preferences that you definitely wouldn't want reinforced, and you realize that personalized AI is exactly what it would do.

If you take agentic systems and personalized systems and put them together, and you imagine what deployment looks like, what it looks like is a vision that the companies have been talking about for several years now: personal AI. Everyone has personal AI, and it is a medium through which they interact with the world and that takes actions on their behalf. Probably, it is a conduit for information and resources, and it offers a layer of protection and so on.

What that really cashes out as is a world in which there is a sort of social economy amongst humans, but there is also a parallel social economy amongst the agents that we have and use to interact with the world. That might sound a bit sci-fi, but actually, I don’t think it’s all that sci-fi. It’s really not all that weird to imagine that we will interact with the world in a way that is technologically mediated, because that’s what we do already. Almost everything we do is technologically mediated.

It’s not weird to imagine that the technologies that we use to interact with the world, instead of being rule-based like they mostly are now, will be optimization-based. They’ll have minimal forms of agency. It’s like, why not? So you create this kind of multi-agent, parallel—if you like, it’s almost like a culture. You can think of it as a culture.

And the trouble is that we know that when you build a system and that system is complex and can interact in complex ways, then you get complex system effects. It can be nonlinear, have weird dynamics, and have feedback loops and so on. And that’s exactly what happened in that flash crash. Actually, there have been maybe dozens of flash crashes. The most famous one was in 2011, the one that I talk about in the book.

So you can think about what complex system dynamics emerge when we are all represented by AI. The reason why I think we should worry about that is because you can think of the norms that we’ve evolved socially and culturally as a set of principles that curtail those complex system dynamics. We have evolved in such a way that we generate a set of predispositions which create a set of constraints on our social interaction that stop, to a large extent, those runaway processes. They’re not perfect. Sometimes we go to war, and sometimes crazy stuff happens. But for the most part, particularly in reasonably small groups and for long periods, we can live in relatively stable, harmonious societies.

But the trouble is that the models won’t have those norms. There’s no reason why they should have them. And the question is, what are the constraints that prevent the same sort of weird runaway dynamics that might lead to flash crash-like events? I don’t think we have an answer for that. That’s what worries me.

Speaker 1

Yes. Designing in constraints would actually limit the technology in quite a strong way. It’s a really interesting thing to think about, though, because in the physical world, the constraints are quite strict. Language is a kind of virtual organism that supervenes on us and has more degrees of freedom, and this new type of AI technology that we’re inventing arguably has even more degrees of freedom. Constraining it is a real challenge.

Christopher Summerfield

Yeah, absolutely. I think the sheer— even if you had systems which were perfectly aligned, which of course is not an assumption any of us can reasonably make, but if you did, the sheer pace and volume of activity that AI can generate is not something that our systems are prepared for. Most systems operate under the assumption that there are reasonable frictions that prevent the system from collapsing.

A good example is the legal system. Many people know that it is possible, particularly depending on the jurisdiction, to engage in what is often called lawfare, so adversarial use of spurious legal challenges. There are certain jurisdictions where it’s strictly optimal to do that because the cost of defending yourself is so high that people will just capitulate, and you can make money. There are frictions that prevent most people from doing that. Most people don’t have legal training. Most people don’t know how to do that or the grounds on which you could do it. It’s a lot of work. You’ve got to file paperwork, and you need domain-specific knowledge.

If we remove those frictions so that you can just, with a few sentences, say, “Please do this,” and have a system that goes and does it, then you suddenly live in a very different world because lots and lots of people can do this. There are many other such examples.

Speaker 1

I was speaking with Conor Leahy about this, and he was talking about a phenomenon called the fog of war, which is that we slowly lose control through illegibility. You can imagine, based on what you said, that you have all of these agents, and even when a country is invaded or when some geopolitical event happens, the average person doesn’t understand why that is, because it’s the culmination of so many countervailing forces, and these systems are just very complex to understand.

So you can imagine a world that becomes so abstract. I also wanted to point out that this doesn’t require AI to be strong for all of these things to happen. Some people think of AI as a cultural technology, a bit like a library or something like that, and then there’s this almost doomer narrative that it’s agentic and this, that, and the other. But you don’t need it to be strong for all of these things to happen.

Christopher Summerfield

Yeah. The human analogy in AGI, of course, overlooks the fact that although collectively what we’ve done is astonishing, individually we’re actually extraordinarily vulnerable and just not all that good at life in general on our own. The classic example is you and a chimp on a desert island: my money’s on the chimp. Our strength is our ability to cooperate. Individually, we are not all that strong.

This notion of a lone intelligence that is like us, but much, much better, I think is a strange one. What we should actually worry about is the unexpected externalities that come from linking together lots of potentially weak systems to create something which is probably completely unlike us and unlike our culture and society, but which we can’t control.

I didn’t know the fog-of-war analogy. That’s very nice. But my favorite paper that talks about this recently is from David Duvenaud. He and others have written this really nice paper called “Gradual Disempowerment,” and it expresses a threat model which I have subscribed to for a really long time and which I talk about in the book. Broadly, it’s exactly that we gradually lock ourselves into the use of optimization-based technologies, and the complex system interactions between those systems sort of write us out of the equation.

The interesting thing about that analogy, which is a point not made either in the book or in David’s paper, is that in a way, it’s coextensive with what happens anyway. If you think about a corporation, the world we have created through hegemonic capitalism with large corporations, for example, in many ways, large corporations are more powerful than any one person. They run under their own imperatives, with their own rules, incentives, and dynamics. For many, they are so powerful that there is no one person who could stop them.

We sort of have a model for what this would look like. It’s just that, of course, in the case of large, complex systems like the corporation, the interactions are slow because they’re largely human-mediated. It’s email or Slack. Everything is such that humans are the cogs in the machine, sorry.

Speaker 1

Yes.

Christopher Summerfield

But in the case of AI, it’s going to happen at warp speed.

Speaker 1

I suppose it’s an interesting time. Just look at AlphaFold, for example. This technology can potentially be used to revolutionize science, but there are so many downsides as well, potentially. What downsides do you think we need to be most cautious about socially?

Christopher Summerfield

If you think about how many products work, of course, firms advertise products, and they do so by branding those products. Branding is a way of trying to get us to engage with something a bit like it was a human. Whether that something is the product itself or maybe the company—the brand—that branding is more or less successful.

Imagine a world in which everything was like that, but it could actually talk back to you and simulate all of the social and emotional types of interaction that we have with people that we care about. The milk in your fridge is like your best friend. This is a very strange world. Of course, that’s a silly example. The milk in the fridge is never going to be your best friend.

But there are already large numbers of people who are engaging with AI in ways that mimic the sorts of interactions they have with other people.

And this creates a whole bunch of vulnerabilities, and a lot of people have talked about risks to mental health and so on. We should be really aware of that, especially where vulnerable people or minors are concerned. But I think there’s another issue that’s talked about much less, and that is the degree to which this will give the organizations that build these systems power over people.

You said that AI increases our agency, and in a way, that is true. But I actually think that there’s also a really powerful sense in which the opposite is true, right? Just as access to social media gives you, in theory, access to lots and lots of information, which should be empowering, most people’s practical experience of it is that they spend a lot of time doing something they think is a bit stupid and would rather be doing something else.

Speaker 1

Yes, and I’m glad you brought that up. It’s a weird phenomenon. This comes into the labor market disruptions. Initially, for some people—certainly now, if you fire up Cursor, you can build a software business in a week—in that sense, it increases your agency.

But actually, everybody else has this capability, and the long-term, or even the medium-term, effect is that it sequesters your agency. It takes your agency away massively, and this is a huge problem.

Christopher Summerfield

Yeah.

Christopher Summerfield

Yeah, absolutely. You could call it a crisis of authenticity, right? You can see this broadly in society because our modes of interaction become so stylized that we lose that sense of authenticity. There are so many dependencies. We always have to present ourselves as being in line with the party line.

You just asked me something that I wasn’t able to answer because I have other dependencies. There is a loss of authenticity in our communications because, in a complex world, we represent many interests. That, I think, is a natural byproduct of our becoming part of the system.

I love this. My favorite metaphor for this is—I woke up in the middle of the night and it just hit me one night—have you—I don’t know if you remember Superman III? Superman III is a terrible movie, but there’s this wonderful scene. I think there’s this kind of giant computer that goes rogue in Superman III.

Speaker 1

Right.

Christopher Summerfield

There’s this wonderful scene where there’s a female character and the machine is just kind of waking up. She’s just walking past it, and the machine sucks her in. She gets stuck there. Then what the machine does is gradually put armor plating on her, replace her eyes with lasers, and basically turn her into a sort of automaton.

It’s a very compelling scene. I think I was terrified by it as a child, which is probably why I remember it. But that is a sort of metaphor for what is happening to us, right? We’re worried about the robots taking over or whatever, but in a way, it’s more like us being sucked into the machine.

We become part of the machine. Just like that poor character, we get turned into something we are not by technology. I don’t think this is a comment specifically about AI. I think this happens to every person who has to go to a press conference, or every person who has to represent their organization or a broader group of people.

You become part of that system, and it erodes your authenticity and, in a way, erodes your humanity.

Speaker 1

Yes. Last time I came to interview you, I went to Luciano Floridi directly afterwards, and his argument is similar about us becoming ensconced in the infosphere, and it changes our ontology.

Christopher Summerfield

Yeah.

Speaker 1

Perhaps you’re arguing more from an agential point of view, but I think it’s quite related.

Christopher Summerfield

Well, I think, as a psychologist, we have dramatically under-indexed on the extent to which what is good for us is actually about our agency, our control, and not about reward. Economics, psychology, and machine learning have all grown up with the notion that utility maximization is the fundamental framework for understanding behavior, and that’s expressed, of course, most prominently in machine learning through reinforcement learning.

But of course, this is not to deny that everyone needs to be warm and have enough to eat, right? Once those basic needs are satisfied—and even sometimes when they’re not—if you look at development and take a sideways view at a lot of both healthy and abnormal psychology, what you can see is that what people really care about is control. People need to understand, and by control I really mean, formally, your ability to have predictable influence on a system.

Speaker 1

Yep.

Christopher Summerfield

So, in machine learning, this often gets quantified as this wonderful notion of empowerment, right? The idea is that what we want to maximize is the mutual information between our actions and future states, for example, either immediate or delayed.

Speaker 1

That is agency, by the way.

Christopher Summerfield

And that is agency. I think that we really, really—if you think of kids, just 2 examples—the extent to which kids will explore the world, the extent to which they will take actions to try and understand, “What if I tap that thing, or what if I take my dinner and throw it on the floor? What’s going to happen? Oh, look, I have control.”

Or, “I cry. Oh, look, my dad’s going to come.” I have control. I can understand that system. That’s what they’re doing, right? Right through to forms of pathological control in adolescence and adulthood—too much control. You can see OCD, obsessive-compulsive disorder, as a need for too much control.

Anyway, I digress, but control is really, really important. When thinking about the impact of technology on our well-being, that conversation needs to be grounded in a robust understanding of how important it is to us to have a predictable influence on our world.

What a lot of AI, or a lot of technological penetration, actually does is make our actions unpredictable. It’s like the frustration that happens whenever you interact with a website that doesn’t quite work, or you get a computer-says-no answer, or there’s two-factor authentication but then there’s no internet and you’re like, “Ah.” It’s like you’ve lost control.

But that control, in the systems that we evolved for—the environment that we evolved for—is much more readily available.

Speaker 1

Isn't that fascinating? There have been studies—I’m sure you're familiar with this one—in which managers in an organization, in a study from the ’70s or something like that, had lower rates of heart disease because they had more power. The underlings would get sick much more often.

If you think about it, with social media and even with these chatbot platforms, I interviewed the team that built all the engagement-hacking algorithms. They were incredibly proud that their average session length was 19 minutes, and they talked all about how they would do model merging and send this response and this response to keep people hooked and keep them there for longer. In a sense, that is what dopamine hacking is about: giving random rewards, right? It's a disempowering thing along the lines you said, and that is the modus operandi for all technology now.

Christopher Summerfield

Yeah, absolutely. A variable reinforcement schedule is the most important—the most effective—way to train animals, including humans, and we are susceptible to it. The unpredictable nature of the reward engages us with the system and makes us come back because we want to control the system, right? We want to know, “How do I make the reward come?” And of course, if you can't, then you keep trying and trying and trying.

I mean, we live in a world in which people have a lot of liberty about how they spend their time, and I think that's as it should be. I don't think we should legislate against frivolity, right? If people want to spend a lot of time on TikTok, collectively, I understand that that's bad. But, for better or worse, we also live in a world in which that is permissible. We live in a country, at least, in which that is permissible.

Where I think we need to be cautious is that this kind of hacking may be undesirable. I might deem it undesirable, but collectively as a society, there are many things that are undesirable. Alcohol is also addictive, but I'm probably going to have a beer as soon as this is done, right? So we make those choices.

But I think that there are vulnerabilities. There are people who are uniquely vulnerable, where that kind of liberty to hack, if you like, spills over into something that can be really actively harmful and can lead people to self-harm. There have been tragic cases, as I'm sure you know, in which people have even taken their own life under an influence that came from an AI system with which they were interacting in this kind of companion mode.

Speaker 1

Do you think it makes sense to think of evolution as having a goal?

7. Evolution Favors Open Endedness

Christopher Summerfield

Probably not, right? There’s this great way of thinking about the blind process of evolution. There’s a paper that I really like, which draws upon the analogy of the blind process of evolution: it’s a selection mechanism that is not teleological, right? It doesn't have a purpose. It just happens.

The paper argues that we should think of training in neural networks in a similar way—that it's very blind. I think there is a fundamental difference between evolution and the way that optimization happens, put it that way. We could learn a lot about neural networks by thinking about the purposeless optimization that happens in evolution, basically.

Speaker 1

It's a really interesting topic for me. I was speaking with Kenneth Stanley the other day.

Christopher Summerfield

Yeah.

Speaker 1

He's done a lot of work about open-endedness, and of course—

Christopher Summerfield

Yeah.

Speaker 1

Tim Rocktäschel works at DeepMind.

Christopher Summerfield

Yeah.

Speaker 1

In Tim Rocktäschel's paper with Edward Hughes—

Christopher Summerfield

Yeah, yeah, the open-endedness paper.

Speaker 1

He said that an open-ended system is one that, from the perspective of an observer, produces a sequence of events which are learnable—

Christopher Summerfield

Yeah.

Speaker 1

—and novel.

Christopher Summerfield

Yeah. Yes, exactly. It's about learnability, isn't it? Joel Lehman—have you had Joel Lehman on the show?

Speaker 1

Yes.

Christopher Summerfield

Yeah. Joel has written really, really nicely about this. I largely share his view and am very close to Tim and Ed's view, which is that the world is open-ended, and optimizing for open-ended systems using well-specified, narrow optimization toward a narrow goal is just doomed to failure, right?

There is probably something really deep about the way the purposeless selection that happens in evolution confers robustness, because it doesn't precisely optimize for this narrow goal. Rather, what it creates is this astonishing heterogeneity, right?

Speaker 1

Yeah.

Christopher Summerfield

The optimization algorithms that we use are all completely opposite, right? They are basically tailored for homogeneity; heterogeneity is a bug.

Speaker 1

Yeah.

Christopher Summerfield

And that's why LLMs show mode collapse. It's why you get this Platonic Representation Hypothesis—the idea that we're gradually converging toward essentially one common, shared set of representations, right?

Evolution doesn't do that.

Speaker 1

Kenneth wrote this wonderful paper called The Fractured Entangled Representation Hypothesis.

Christopher Summerfield

Oh, I don't know that paper. Yeah.

Speaker 1

It was with Joel—I'm not sure if Joel was part of this—but he was involved in Why Greatness Cannot Be Planned. They did this thing called Picbreeder.

Christopher Summerfield

Hmm.

Speaker 1

It was like Flickr, where it was supervised by a diverse group of humans. The humans could pick interesting image generators, which were CPPNs—Compositional Pattern Producing Networks. You could create this phylogeny, and they talk about this concept called deception: the stepping stones that lead somewhere interesting don't resemble the interesting thing.

Humans have this sense of what's interesting because we seem to know the world well. With a few steps in the phylogeny, they found these pictures of butterflies and apples. When you do parameter sweeps on the networks, because they have such an abstract understanding of the objects, the apple would actually get bigger. One neuron would make it bigger; one neuron would make the stem swing.

If you train a neural network with stochastic gradient descent to do the same thing and you do parameter sweeps, it's just—it’s like spaghetti. It's all over the place.

So their hypothesis—and this seems like an obvious thing to say—is that if we could have a sparse representation that mirrored the world, then, because the knowledge would be evolvable, we could trust it with autonomy to make creative leaps, because it would do the right thing.

Christopher Summerfield

Mm. Yeah, that's amazing. I don't know about this paper. It sounds like I should read it. I mean, this idea that it's difficult to get places because the interim states are not highly valuable is a very old argument. This is the basis of Paley's watchmaker argument, right? How did we ever get the eye? You couldn't possibly evolve that; it's too complicated.

But those gradients must be there, right? The gradients are there.

Speaker 1

I have to say, Professor Summerfield, the prose—the way that you've written this book—is very impressive to me. It's one of the best-written pieces of writing I've ever seen. It occurred to me that you might have been deliberately making it so creative that it would be impossible to mistake for AI-generated content.

I don't know whether this is just because my standards are so low now because of shitification and all of that, but it was remarkable. Were you thinking that? Were you leaning into the creativity a little bit?

Christopher Summerfield

I love to write. I love to find new ways to explain things and to convey ideas, so that's a selfish pleasure for me. It didn't cross my mind that people might think I had used ChatGPT to write the book, but I guess, in hindsight, that's a very sensible way of thinking about it. But no, it was all me. That is mind-blowing.

Speaker 1

Professor Summerfield, it's been an absolute honor. Thank you so much.

Christopher Summerfield

Thank you.