Google 研究员展示生命“从代码中涌现” [Blaise Agüera y Arcas]
Agüera y Arcas 的核心判断是字面意义上的:“DNA 是一段计算机程序”,因此生命是智能的一个子集,因为可遗传的自我复制需要一台通用计算机。 Von Neumann 在分子生物学为其提供组成部件之前,就已预见这套堆栈:DNA 是图灵机磁带,核糖体是构造器,DNA polymerase 是复制器。大脑和文化是更快的计算层,但细胞搭起了最底层。
他的 BFF 实验表明,目的性可以通过一次剧烈的相变,从随机代码中自发涌现。 由1,000条随机64字节磁带组成的“汤”,每条平均只有约2条指令,在数百万次两两互动中几乎始终处于惰性状态;随后熵突然坍缩,重复的复杂程序出现,逆向工程显示它们能够繁殖。某种东西之所以变得有功能,是因为它现在可以被破坏:扰乱代码,它复制自身的“目的”就会消失。
真正改变他进化论判断的意外是:当突变率从每次互动约万分之一降至0时,涌现仍在继续。 这让解释重心从突变转向合并:能够独立复制的元素可以组合成更复杂的整体,而其组装历史本身会成为信息。因此,内共生有助于解释复杂度为何会沿方向持续上升,而普通的突变—选择机制无法完整解释这一点。
这种合并逻辑既适用于线粒体进入古菌,也适用于技术和生物机制的再利用。 Agüera y Arcas 举例称,病毒原本具备的细胞膜融合能力,后来被用于胎盘形成。在 BFF 中,即便2个原始复制器究竟以 AB 还是 BA 的顺序合并,也会成为持久信息:“最终,合并树恰好成为最终基因组所编码的信息。”
在 Google,他带领的约50人 Paradigms of Intelligence 团队明确试图在利用当下成功模型之外,“重新装满桶”,补回基础性洞见。 促成这一转向的思想催化剂,是他在2020年之后得出的结论:大规模序列模型似乎已经具备一般智能。讨论重点包括组合、合作、多智能体组织,以及持久记忆等研究方向。
他的功能主义认为,基质不是生命或意识的本质:肾脏就是任何履行肾脏功能的东西,即使实现方式不同。 他强烈反驳 Anil Seth 和 John Searle 偏向基质的直觉,认为如果用功能等价的部件替换神经元,意识不会在无声无息中被“调低”。生物学接口湿润而复杂,但“多重实现性与再利用性”正是生命的本质。
他将 AI 视为人类集体智能的延伸,而不是一种截然分离的外星智能,但仍担心极化、虚假信息,以及20年后可能已不再适用的政治经济系统。 “AI 从一开始就是人类智能”,因为通用模型是在吸收人类语言后涌现的;Scarfe 更尖锐的反驳——单个产物可能超越全人类——仍基本悬置。对于当前的 transformer 系统,Agüera y Arcas 认为最大的缺口不是组合能力本身,而是“叙事记忆”,尤其是能够维持长期自我的持久记忆。
1. 大规模序列模型迫使人们重新思考智能
Agüera y Arcas 将《What Is Intelligence?》描述为一次始于2020年前后的思想冲击记录:大规模序列模型似乎具备“普遍智能”,迫使他追问,如果接受这一观察,它对进化、心智以及人类自身意味着什么。该书的开篇章节也以《What Is Life?》为名单独发布,主张生命属于更广义的智能范畴。
他目前所在的机构是 Google 的 Paradigms of Intelligence 团队,约有50人,研究重点是基础问题。团队承认当前范式有效,但希望“用新的洞见和新想法重新装满桶”,而不是只集中于利用现有模型。
Scarfe 提出 DNA 适应缓慢,而神经系统和文化以“光速”进化;Agüera y Arcas 明确承认这一点,但没有因此改变结论。DNA 只是最底层:一旦通用计算存在,生命就会“用计算机构建计算机,再用计算机构建计算机”,不断叠加细胞、神经、文化和社会层。
2. 自我复制要求生物体内部存在一台计算机
Von Neumann 的思想实验从一个乐高机器人开始:它从池塘里收集散落的零件,组装出一个后代。机器人需要一条指令带、一个读取指令带的通用构造器,以及一个把自身指令交给后代的复制器——其中也包括构造构造器和复制器的指令。
生物学后来提供了惊人的对应关系:DNA 是指令带,核糖体负责构造,DNA polymerase 复制指令带。可遗传改变之所以重要,是因为改变基因组会改变后续后代;Agüera y Arcas 的结论是绝对的:“不可能成为一个活的生物体,却不是字面意义上的一台计算机、一台通用计算机。”
元胞自动机让计算获得实体形态。传统图灵机中,磁带、读写头、规则和符号在概念上彼此分离;而在元胞自动机中,每个位置都遵循宇宙的局部“物理定律”,使机器能够打印出自身的物理机体——“一台既是笔记本电脑又是3D打印机、还能打印出另一台笔记本电脑的机器”。
Scarfe 对生化机制的质疑仍然成立:Conway’s Game of Life 是二维且确定性的,而物理生命依赖三维空间和热随机性。因此,Agüera y Arcas 引入随机图灵机,同时强调大规模并行与嵌套结构:细胞内部运行着数以10^18计的核糖体,细胞嵌套于生物体之中,人又嵌套于社会之中。
3. 随机代码跨过相变边界,涌现出目的性
BFF 从1,000条随机磁带开始,每条长64字节,使用一种由 Brainfuck 衍生出的7条指令、自修改语言。随机字节中约31/32是 no-op,因此每条磁带只剩约2条指令。每次迭代随机将2条磁带拼成一个128字节程序,运行后再将其拆开,放回这锅“汤”中。
数百万次互动几乎什么都没有发生;随后“某种看似神奇的事情发生了”。熵骤降,原本不可压缩的“汤”变得高度可压缩,复杂程序以大量副本出现。重复出现暴露了它们的功能:它们已经成为复制器,因此生命的涌现同时也是“目的性的涌现”。
目的性是操作层面的,而非神秘属性:改变一个关键字节,程序就会停止复制,说明它具备一种可以被破坏的功能。Scarfe 提出架构偏置问题,Agüera y Arcas 也承认语言会塑造最终程序;但 Z80 汇编中出现的类似涌现让他相信,这一现象本身具有普遍性。
4. 合并,而非单靠突变,为进化提供复杂度棘轮
从随机到有序的表面逆转依然符合热力学。借助 Adi Pross 的“动态动力学稳定性”概念,Agüera y Arcas 认为稳定性可以是循环的:反复制造 DNA 的脆弱 DNA,可能比只会不断侵蚀的花岗岩存在更久。从这个意义上说,“进化就是第二定律在发挥作用”。
他最初在 BFF 中以每次互动约万分之一的概率加入随机突变,假设达尔文式的“偶然与必然”驱动了结果。但当突变率降至0时,复制器仍然出现,这改变了他的看法:目的性、生命起源以及复杂度增长,都无法完全用纯粹的达尔文机制解释。
在他的框架中,内共生为后续复杂度提供了机制。他描述了一个真核生物的形成过程:线粒体进入古菌体内;由此产生的复合体比任一组成部分都更复杂,就像长矛比木棍和石尖更复杂。把既有部件组合起来,能够让进化在之后产生更精密的系统。
Scarfe 将同一逻辑延伸至信息:合并保留了谱系、复用、路径依赖,以及部件如何组装的历史。Agüera y Arcas 表示,他“完全”接受这一框架,并展示了 BFF 中合并树本身如何成为最终基因组中的信息。
5. 组合将偶然性转化为可复用结构
Scarfe 将合并与谱系、路径依赖、复用和沟通化联系起来。Agüera y Arcas 表示自己“完全”接受这一框架,同时通过 Dan Everett 对 Pirahã 的研究强烈批评 Chomsky 的语言理论。他指出,Pirahã 语言不符合 Chomsky 的要求,包括递归或中心嵌入,而且没有数字以及过去时和将来时。
他对 Chomsky 更深层的反对,是自上而下的形式主义:语法和程序路径催生了“老派 AI”,他认为这是导致多轮 AI 寒冬的一次错误起点。Scarfe 反驳称,Chomsky 的自动机和计算思想在堆栈更底层处与这一论点相似;Agüera y Arcas 则回应,Von Neumann 和早期人工生命研究者 Nils Aall Barricelli 已经具备了关键机制。
W. Brian Arthur 的灯泡例子解释了发明为何会聚集出现。一旦玻璃吹制、真空技术、灯丝和电流都已存在,多名发明者就可能独立组合出同一类产品;灯丝、灯座直径或螺纹方向等偶然决策,随后会约束之后建造的一切。
BFF 在最小尺度上暴露了这段历史:弱小的单字节复制器相遇,有时会作为一对持续存在,并在结合后表现更好。它们以 AB 还是 BA 的顺序合并,就是新增信息,因此“合并树”成为最终基因组。更高层次的感知则更平滑:即使规则被破坏,人类也能识别出古怪的自行车,这更支持神经网络和梯度下降,而不是手写的圆形检测器。
6. 功能可以跨越基质变化而延续
Agüera y Arcas 认为自己最接近功能主义者。Scarfe 以肾脏为例:物理学或许能够完整描述原子,但肾脏由过滤尿素这一功能定义,因此,只要一种截然不同的人工装置履行同一功能,它仍可被称为人工肾脏。讨论将功能视为一种关系属性,只有在由其他功能构成的生态中才具有意义。
多重实现性被视为功能的标志:不同路径都可以制造 ATP,昆虫翅膀和蝙蝠翅膀也以不同方式实现飞行。Scarfe 更强的质疑来自历史维度:替换天然肾脏或植物现在可能有效,却可能破坏生态未来的演化轨迹,因为替代物带着另一套来源进入系统。
Agüera y Arcas 将这一质疑转化为对内共生的支持:生命不断引入、再利用并平行改造那些原本服务于其他功能的机制——“没有任何智能设计者,智能设计也可以发生”。他的例子是细胞膜融合能力:它最初与病毒有关,后来被纳入胎盘形成过程。
因此,他强烈反对 Seth 和 Searle 的基质本质主义:用功能等价的神经元替换原有神经元,不会抹去意识,尽管湿润的生物学接口使这种等价远比替换软件例程困难。“多重实现性和再利用性”对他而言正是生命的本质。
7. 意识是一种合作技术
在 Agüera y Arcas 的功能主义框架下,哲学僵尸——行为与人完全一致、却“内里已死”的实体——没有看上去那么自洽。意识既不是惰性的副现象,也不是特定物质的特权;它存在于生成行为的关系之中,并在其中发挥功能。
他的 Paradigms of Intelligence 团队通过多智能体强化学习研究这一功能。合作先于内共生,而合作智能体需要心智理论:每个智能体都必须对一个包含博弈、自身内部状态、另一个智能体内部状态,以及“我的微笑意味着快乐,所以你的微笑大概率也意味着快乐”这类对应关系的世界进行推断。
对自我、他者、他者对自我的模型,以及更深层嵌套模型进行建模,会形成类似 Hofstadter 的“奇异环”。现实中的智能体受计算能力限制,其心理表征仍然像卡通画;Agüera y Arcas 说,这种递归建模最多只能达到约第6阶。
划船提供了体验层面的类比。在“同步摆动”中,8名桨手完全同步,以至于“船获得了灵魂”,当彼此分离的目的合而为一,船的速度也会提升。Scarfe 将此联系到招聘时同时寻找能动性和协同性:有时,最合理的意向边界应当包住整个协同群体,而不是其中任何一个成员。
8. 自我是一个带有“内在律师”的协商边界
连体双胞胎 Abby 和 Brittany Hensel 展示了流动的能动性:2个独立的大脑和脊髓分别控制一只手臂和一条腿,但终身的行为交叉提示让她们能够驾驶、运动、写字,并经常同步说话。她们可以作为一个协调系统行动,同时保留不同意见。
裂脑患者则呈现相反情形。外部观察者可以通过实验看到2个半球接收不同信息、控制不同的手,但患者坚持自己仍是一个人。Agüera y Arcas 拒绝接受一个唯一的特权答案:自我是关系性的,即便一只手在扣纽扣、另一只手在解纽扣,也可能只被体验为一种不便。
Petter Johansson 的选择盲实验揭示了连续性是如何被制造出来的。受试者被要求解释自己对某张脸的吸引力选择,却常常拿到自己刚刚拒绝的那张脸;他们很少察觉,并以不变的流畅度和反应延迟为其辩护。“内在律师”编造出一套叙事,甚至可以改变未来选择;大脑中分布式的各个部分彼此遮掩,因为它们“都在同一条船上”。
9. AI 延伸集体人类智能,但记忆仍是缺口
Agüera y Arcas 担心极化、虚假信息,以及政治经济制度在20年后是否仍然适用,但不担心 Eliezer Yudkowsky 讨论的那些情景。他的理由是本体论层面的:单个人类并没有显著超越其他灵长类;文明中的数百万人和数十亿人协作,才共同制造出器官移植和太空飞行。
因此,AI 更像既有集体过程的延伸:“AI 从一开始就是人类智能”,因为通用 AI 是通过在海量人类语言上训练模型实现的。Scarfe 反驳称,单个产物可能超越全人类的总和;这一问题并未得到正面解决。Agüera y Arcas 转而强调,个体经常把社会分布式知识误认为自己的知识,自行车绘画失败就是例证。
今天的模型在架构、训练、接口和模态上显然不同于大脑,但其内部表征与人类之间可以在 Brain-Score 一类测量中出现令人意外的趋同;即便只有语言输入的系统,也能复现感知结构的某些方面。模型组合出非寻常场景的能力,削弱了“它们没有组合能力”的说法。他认为最明确的缺陷是“叙事记忆”:能够维持一个自我的持久长期经验。
Scarfe 所说的“肤浅冒牌货”担忧,集中于模型在轻微逻辑变化下给出脆弱答案。Agüera y Arcas 的回应是方法论上的:先运行人类基线,因为除非刻意放慢速度并形式化推理,人类往往也会表现出同样的框架效应和逻辑错觉。Transformers 不会系统性地搜索所有图灵机程序;无论系统是大脑还是 transformer,可计算的归纳都必须依赖捷径。
The new book—it’s called What Is Intelligence? Thank you for asking. It was just published by MIT Press about 3 weeks ago, so it’s very fresh off the presses. There’s an online version of it as well that is free and very rich, full of all kinds of rich media.
Chapter 1 of that book is called “What Is Life?” “What Is Life?” is sort of the single to the album, and works as a book in its own right. That book is also for sale from MIT Press. I think I’ve explained why I think of life as being a subset of intelligence, and why the story of artificial life and abiogenesis is relevant to the story of intelligence and what it is.
The subtitle of the book is Lessons from AI About Evolution, Minds, and something something. It’s basically documenting the time since about 2020, when I got really shocked by seeing that these large sequence models seemed to be generally intelligent, and starting to think through the implications of that. What would it mean if we believe our eyes and that is what intelligence is? What does that tell us about ourselves and about the properties of intelligence more broadly? That’s the whole intellectual journey that has taken us on over the last few years.
MLST is supported by Cyber Fund. So, I’m Falm. I’m the co-founder and CEO of Prolific, and Prolific is a human data infrastructure company. So, we make it easy for people developing frontier AI models and running research to get access to trustworthy, high-quality participants for high-quality online data collection.
I am at Google. I’ve been there for about 12 years now, and I am their CTO of Technology and Society and also the founder of a new research group—well, newish; we’ve been around for a couple of years—called Paradigms of Intelligence, or PI. It’s much smaller than the previous organization that I ran at Google Research. It’s about 50 people, so enough to do some real damage.
The idea is to really focus on the fundamentals of artificial intelligence and go beyond exploiting the current models and paradigms that are working well. We believe in those, but we also think that we have to refill the bucket with new insights and new ideas as well.
I’ve just watched your talk, and you said that life and intelligence are the same thing. They are both computational. What do you mean by that?
This is a surprising claim. I know it sounds a bit odd, but what I mean by that is—well, let’s begin with life and why life is computational.
In the middle of the 20th century, John von Neumann, who was one of the founders of computer science, realized that in order for a robot paddling around on a pond to make another robot out of loose Legos that it finds floating around in the pond, just like itself, it needs to have instructions inside itself. Let’s suppose the robot is made out of Legos. Its job is to make another robot out of those loose Legos.
He imagined a tape with instructions for how to assemble a machine. The robot would also have to have a machine inside itself that would be able to walk along that tape and follow those instructions to take the loose Legos and put them together into its own form. It would also have to have a tape copier so that it could endow the offspring with that tape. The tape would have to include the instructions for the copier and the assembler—the universal constructor, as he called it.
The cool thing is that he made all of those predictions, if you like, on the basis of pure theory, before Watson and Crick and their unacknowledged collaborators had figured out the structure and function of DNA, before we knew how ribosomes worked—which are, in fact, exactly that universal constructor—and before we had discovered DNA polymerase, which is the tape copier.
All of those things have to exist in order for an organism not only to be able to reproduce itself, but to do so heritably, such that if you make a change in its genome, in what’s on the tape, then the offspring will also have that change. The kicker is that this universal constructor is a universal Turing machine. In other words, you have to have a computer inside yourself as a cell in order to make another cell. DNA is, in this very literal sense, a Turing tape.
It’s a very profound insight because, basically, he’s saying you cannot be a living organism without literally being a computer—a universal computer.
Very interesting. So, are you saying that DNA is basically a computer program?
DNA is a computer program, yes.
Very cool. Many folks in the audience would have been inspired by Conway’s Game of Life, for example, and you were talking about computational equivalence. The Game of Life, of course, is Turing-complete because it has expandable memory. Just as DNA has expandable memory, the grid size could just keep growing.
As you’re pointing to, when we see these weakly emergent behaviors, they appear very lifelike. Does that in any way downplay the biochemical and thermodynamic realities in the physical world?
Of course. A cellular automaton like the Game of Life is a lot less complex than the real world. It’s in 2 dimensions rather than 3, and there’s no thermal randomness, which turns out to be very important, actually. The fact that the computation is deterministic is also a little different from real life, where those thermal fluctuations mean that there’s always a probabilistic element to things.
You do have to extend Turing’s original ideas about computation to make a so-called stochastic Turing machine in order to really do a proper job. What von Neumann was getting at—and I’m glad you bring up cellular automata—is that they’re really a generalization of the Turing machine to the laws of physics.
Essentially, every pixel on your game board, if you like, is performing a computation. It has a state, and it’s performing some very simple computation to say what the next state is on the basis of its neighbors. The idea is that the rules determining the next state of a particular pixel are the laws of physics of that universe.
The reason von Neumann came up with this idea of cellular automata is because he wanted a system that would allow one to do computation, but in which the computation is embodied. What I mean by that is, in a Turing machine, there’s a tape and there’s a head, and then there are the symbols that are written on the tape. But the symbols are not the same stuff as the tape and the head and the instruction, or the table of rules. Those things are abstract and separate from the symbols that are written.
Whereas in a cellular automaton, the machine can literally print itself. It’s not just a laptop, if you like, but a laptop and a 3D printer in one that can print another laptop. Embodied computation is computation where the memory is written and read in atoms rather than in bits, and therefore the machine can make another of itself.
David Krakauer said to me that—we can agree that intelligence is in adaptivity, inference, and representation. Adaptivity is very important. DNA is adaptive, but it’s very slow. He was saying that the nervous system, the brain, and culture are evolution at light speed because they allow us to overcome the information-transfer bottleneck between successive generations.
Does it make sense to think of the program of intelligence at the DNA level when so much of the adaptivity seems to be happening higher up?
A lot of the adaptivity very much happens higher up. In humans, we have cultural evolution, which goes, as David says, at light speed relative to genetic evolution.
When I say that DNA is a Turing tape and that ribosomes are universal computers that construct life, that’s really only the ground level. Or maybe it’s Level 1 or Level 2. There’s physics underneath that, but there’s a Level 3, a Level 4, a Level 5, and so on. There are computers built out of computers built out of computers.
The thing about it is that once you have that ground level, then you can build as many floors above that as you like. The point is that the moment you have life—meaning that you have something that can build a copy of itself—you have a general computer, which allows you to do anything. What that means is that life, from the very beginning, is computational and can start to compute in parallel.
This gets us into symbiogenesis, which I guess we’ll cover in a bit. The fact that that ground floor is computational answers the question of why brains are computational. It’s because cells were computational from long before there were action potentials and other fast electrical processes allowing us to think.
Can you speak to this notion of recursion that you were just pointing to? Carl Friston, for example, thinks about this division in systems called a Markov blanket, like a statistical independence. We seem to observe empirically that complex, intelligent adaptive systems have nesting.
You were just speaking about having levels upon levels upon levels. What does that buy you? Is it a kind of recursion? How does that improve the sophistication of the intelligence?
Yeah.
Well, 2 things happen. One of them is things inside things inside things, and the other one is parallelism—in other words, a lot of things at the same level happening at once. They’re both important. They’re both important parts of the story.
First of all, when you have a cellular automaton like Von Neumann was imagining, that’s already massively parallel computation, because every pixel is like a little computer doing its thing. In the same way, in physical space, there can be molecules in a lot of spots, all of which are doing something. You can think of them as computational operations, perhaps, and they’re all happening at once.
In your body, you have quintillions of ribosomes, and all of those ribosomes are little, tiny universal computers working in all of your cells at once. All of your cells are working at once. But there is also nesting, because you are a person. Of course, you’re already a part of a society, which in some sense is an intelligence bigger than an individual human.
You’re made out of cells. Those cells are made out of organelles. Those organelles are made out of proteins. Those proteins are made out of molecules. That sort of nestedness is really important as well. You’re not only a lot of computers working together in parallel, but you’re also a system of computers made of computers.
One thing I find fascinating that you’ve thought a lot about, Blaise, is where the purpose comes from. Folks like Harrison, for example, have a no-nonsense physics interpretation that it’s the second law of thermodynamics at the end of the day, because there’s some notion of valence: we build these complex adaptive systems, and there needs to be something that drives them forward, something that propels them in a certain direction. Your experiments have shown that this kind of falls out of computation. What do you mean by that?
Yes. To be clear, computation and the second law very much work together here. The experiment that I did a couple of years ago that really got me started with artificial life is called BFF. It’s based on a language called Brainfuck, which is where the first B and F come from. I didn’t name it that, although I admit I enjoyed that it was called that.
This is a minimal Turing-complete language designed by a grad student. I think he was a grad student in physics, Urban Müller, in the ’90s. It’s a very minimal language. It has only 8 instructions. I only use 7 of them.
The basic setup is that I begin with a bunch of tapes of length 64, just 64-byte-long tapes, like those Turing tapes or von Neumann’s tapes. They start off filled with random bytes. There are only 7 instructions, so the great majority of those bytes—about 31 out of 32 of them—are no-ops, meaning that they don’t code for any instruction at all. They start off random and very much purposeless.
You have 1,000 of them in your soup. The procedure is really simple: it’s just plucking 2 of those tapes at random out of the soup and sticking them end to end. You make 1 tape that is 128 long, and then run it. This modification of Brainfuck is self-modifying, meaning that when you run it, it can modify values on that combined tape. Then you pull the tapes back apart and drop them back in the soup. That’s it. You just repeat that process.
If you do that a few million times, you start off with nothing much going on. Again, the huge majority of those bytes are not even instructions. There are only an average of 2 or so on each tape, so the likelihood of them doing anything is almost zero. Once in a while, you might see 1 byte somewhere change, but after a few million interactions, something apparently magical happens, which is that suddenly the entropy of the soup drops dramatically.
It goes from being incompressible because it’s all random bytes to being very highly compressible, and programs emerge on those tapes. Those programs are complex. They take some real effort to reverse-engineer, and you can see that they’re occurring in a lot of copies. That’s why it’s compressible.
The fact that they’re occurring in a lot of copies tells you what the programs are doing. They’re reproducing. They’re copying themselves. What’s so cool about this experiment is that it really shows you how life emerges from nothing. The emergence of life is, in some sense, the emergence of purpose.
What is the purpose of one of these programs? It is to reproduce. If you were to mess with one of those bytes, if you were to change it, you would in most cases break the program. When you break the program, it no longer functions to reproduce. Something that can break is something that is functional or that has purpose.
Absolutely fascinating. You said there was a phase change that was quite sudden. Would David Krakauer acknowledge that as being a form of emergence?
I think so. I’ve actually never asked David that question. We disagree on a lot of things AI-related, but I think he would acknowledge that this is a phase change and that it is an example of emergence.
I think he would, because he has a bunch of criteria, but one of them is a fundamental coarse-graining and reorganization of the microsubstrate such that the new phenomena can be described with a simple new variable. This seems to match that description.
Is it possible that there’s some kind of design bias? When we design machine-learning architectures, there’s so much information in the architecture, and in this case there’s so much information in the Brainfuck language and the terms and so on. Could that have influenced it to emerge in a certain way?
Yes. The structure of those programs does change depending on the language. We’ve tried this with other languages. We’ve tried it with Z80 assembly language, which is the assembly language of these Zilog chips that were invented sometime in the ’70s and just got discontinued last year—a very long-running microprocessor architecture.
The phenomenon is very generic. What those programs look like is shaped by the specifics of the language. But the reason those programs emerge, the reason that they develop purpose, is actually thermodynamic.
That might seem puzzling, because you would think thermodynamics is about things becoming more random, and apparently the exact opposite is happening here. You start with randomness and you get order. How could that be?
I think the answer was well characterized by a chemist—an organic chemist—Adi Pross at Ben-Gurion University of the Negev in Israel, who is now emeritus. He did a lot of work on so-called dynamic kinetic stability. The idea is that it’s an extension of the second law that says things seek their most stable state, their most stable form.
Usually, we think about those stabilities as only being fixed points, but those stabilities can be cycles too. If something dynamically makes itself, if something forms more copies of itself, that’s more stable than something that just settles.
It’s like the old joke about DNA being the stablest molecule in the universe. Obviously, DNA is fragile, but at the same time, if the DNA makes more DNA, then it will be around a long time after granite, which can only erode.
In terms of this valence question, though, does that imply to you that there is a natural drive to survive, almost? For these systems to maintain their existence, assuming that’s a primary force, they would need to have a degree of sophistication. They would need to be doing modeling; they would need to be doing sophisticated things. But is that something that just always happens? Is it a convergent property?
Yes, it is. In that sense, evolution is the second law at work. If you have a bunch of things that are not copying themselves in the BFF soup, and you have something that emerges that can copy itself, then that thing that can copy itself will write over the things that can’t copy themselves, which means it’s more fit—or more stable, if you like.
That is written into the laws of statistics in just the same way that the second law is. It’s just the kinetic, or cyclic, form of that same law rather than the steady state.
You said in your talk that merging is more important than mutation. Tell me more.
Yes. The usual thing—what we learned in school—was that Darwinian evolution consists of mutation and selection, or what Jacques Monod, the Nobel winner, called “chance and necessity.”
In other words, mutations, maybe from cosmic rays or whatever, to our DNA, sort of throw spaghetti at the wall. Whatever sticks is what remains—whatever doesn’t kill us and whatever hopefully makes us stronger.
That was my assumption as well. Starting out with these BFF experiments, I had a mutation rate where a byte could change at random with probability 1 in 10,000 or something with every interaction. Then I began playing with the mutation rate and found that this emergence of these complex programs occurred even when the mutation rate was turned down to 0.
Which is really a surprising finding. It tells you that this emergence of purpose comes about even without any random changes in the code. It’s not explainable in purely Darwinian terms.
The other things that are not explainable in purely Darwinian terms are the emergence of life in the first place. This greatly puzzled Darwin. He thought this problem of abiogenesis, or the emergence of life, was impossible to reckon with. You might as well talk about the origin of matter, is how he put it in one of his letters.
The other thing I can’t explain is the increases in complexity that occur. Why is life now more complex than bacterial life 1 billion years after it began on Earth? Why do we have human societies now? If we go back 100 million years, we had only things with much simpler brains. We had octopuses—they had pretty complex brains—but the tendency has been toward greater complexity.
There are some people who have argued against that. Famously, Stephen Jay Gould has said things like, “Everything on Earth is the same amount evolved. We’ve all been evolving for 3 billion years. It’s all equally evolved.” I think Gould was wrong when he said this. The reason being symbiogenesis: when a eukaryote is formed by a mitochondrion finding itself inside an archaeon and then becoming a eukaryote, that resulting composite organism is more complex than either of the 2 parts that made it up.
It’s the same way that a spear is more complex than a stick and a stone point. You put 2 things together, and now you have something more complex than the parts. If this idea that symbiosis, or symbiogenesis, is an essential part of evolution is correct, then you absolutely get more sophisticated things coming about later in evolution because they’re being put together from preexisting parts.
Yeah, I wanted to touch on the importance of the merge operator. We were talking about that earlier, and even Chomsky spoke about this. You could argue whether the merge operator in language evolution was the Prometheus moment, whether it was phylogenetic or ontogenetic, because you were just talking about symbiosis and merging in a physical substrate.
But it also happens in the information substrate. You get these memetic computer programs that ensconce themselves, and maybe language was that. I don’t know, but I have a theory about why merge is so important as opposed to random selection. I think creativity is about grounding. It’s about path dependence, basically.
Even the retroviruses and all of these things form a lineage. I think that if you don’t use merge, you lose the lineage. Also, something about the recursive merge operation allows you to build more complex computer programs by allowing for this kind of reuse and canalization. There’s something very natural about that.
Yeah, I completely buy everything that you’re saying. I think that’s exactly right, except that I dislike Chomsky. So, I think he’s wrong. He’s wrong about language. I’m much more of a fan of Dan Everett. I don’t know if you’re familiar with his work with the Pirahã. It’s wonderful.
He spent a long time with the Pirahã in Brazil, who are a people whose language does not obey Chomsky’s requirements for language. They don’t have recursion. They don’t have anything like center embedding. They also don’t have numbers, and they don’t have past and future tenses.
Everett wrote a great book some years ago called Don’t Sleep, There Are Snakes, which talks both about his experiences among the Pirahã and their language, and also his big fight with Chomsky over this. Chomsky’s papers are filled with theory and pseudomathematics, and have no time to give to ethnography or to actually studying any real languages. But anyway, I’m digressing.
Putting Chomsky aside, though, what you’re saying about merge—or, as I would see it, composition, functional composition—I think is absolutely fundamental. It’s how all technology is built. W. Brian Arthur has written about this and how technology evolves. Every technological invention gets invented a dozen times around the same time, as if everybody’s in telepathic communication.
The reason is that every technology has precursors. You can’t get a light bulb until you know how to blow glass, how to make a vacuum, how to draw a filament, and how to generate electric current. When all those things were there and the need for light was there, the light bulb was going to get invented. But it was invented a dozen times by different inventors with different contingent choices.
They might choose which kind of filament to use, whether it’s prongs or whether you screw in the light bulb, which way you screw it in, what the diameter is, and so on. Those decisions, as they get locked in, determine the course of everything after that which incorporates light bulbs.
So, in a way, this contingency—these choices about exactly which way those combinations go—is actually what the entire genome, or whatever it is, is made out of. In the case of BFF, the original replicators are really just single instructions that sometimes, randomly and weakly, might copy themselves. One byte moves here and there, but as those bytes get copied around, sometimes a couple of them end up together, and then they’ll copy as a group.
They’ll do better together. The contingent thing— which way they ended up getting copied, whether it was AB or BA, that they stuck together— is the information that the bigger thing is made out of.
That little extra bit, because in this case you just had single bytes, was not information to begin with. So the merger tree ends up being exactly the information that is encoded in the final genome. It’s all about the history.
Yes, absolutely fascinating. In a sense, I’m surprised you’re not a fan of Chomsky, because he was talking about automata and Turing completeness. He was the ultimate computationalist, and in a sense, what you’re describing is Chomsky’s ideas just applied lower down the stack.
That’s right. In that sense, I think he was correct, but I also think all of those ideas were already there in von Neumann in the 1950s. Even Nils Aall Barricelli, the first artificial-life researcher, worked on one of von Neumann’s machines. I think he sort of snagged time on the MANIAC to do some of his first artificial-life experiments.
They’re kind of pseudodocumented in Benjamin Labatut’s book MANIAC. It was really fun. Or no, that was in his first book, I think, When We Cease to Understand the World. But anyway, my point is that those ideas were there before Chomsky.
The thing that Chomsky really pushed, during his reign of terror over linguistics—sorry, I’m being a little bit mean—was the movement in artificial intelligence that we now call GOFAI, or good old-fashioned AI. It held that you could formalize what AI is as grammars and programs, which turned out to be wrong. That turned out to be a false start in AI, and it’s why there were so many AI winters.
There seems to be a bit of a tension, because the GOFAI folks had some very interesting ideas. I mean, I’m a big fan of Fodor and Pylyshyn, for example, and they spoke about strong compositionality. We have semantics and intentionality, and it’s possible to build these cognitive representations, but we have the issue that we can’t really design them to represent the world in a high-fidelity way, and we have semantic divergence.
Then you’re pointing to this very interesting constructive thing, and I think a constructive form of AI and compositionality solves a lot of problems because of this path-dependence problem and this canalization that we’re talking about. When you build intelligence brick by brick, you can build artifacts of incredible sophistication, but unfortunately, we can’t design the artifacts to do exactly what we want. We can gently steer them in a certain direction.
Even with Friston, I feel that even though he’s talking about the what of intelligence as prediction and adaptivity, I think the implementation matters. I think adaptivity means structure learning. I think there’s something about having a substrate which actually does this form of composition that you’re talking about that seems to be a mechanistically necessary condition for intelligence, right?
Yes. I think in many ways what we’re talking about is the tension between analog and digital ways of thinking, or bottom-up and top-down ways of thinking.
For instance, let’s talk about how you would recognize a bicycle. In the good old-fashioned AI world, you would say, “Well, you’ve got a circle detector and a line detector that will detect the lines that make up the frame of the bike,” and so on. You’ll handwrite code for all of those things.
Of course, the problem is that there are many ways of looking at a bike where you’re not going to see the wheels at once, or maybe the bike is of a weird design. There are those funny bikes that have shoes instead of wheels, and when you look at one of those in a Gestalt sort of way, you recognize a bike immediately, even if all of the rules are broken, as it were.
That’s really important, because when you’re looking as an intelligent being at the world, you have to cluster. You have to find regularities in the world whose shapes are not well defined by a set of rules. They’re not just carved up by hyperplanes; they’re blobby.
Intelligence requires methods that are very neural-net-like, that look more like continuous function approximators. That’s why gradient descent is a good idea, for instance, and learning these things via smooth functions is a good idea. Trying to encode them with rules never worked out well.
On the other hand, DNA is discrete, right? There are 4 symbols, and you order them in a certain way, and that’s it. It doesn’t mean that there’s no randomness in the way proteins are folded and so on, but composition at the level of DNA really does have to do with chopping up programs essentially made of discrete symbols, inserting bits of code, and so on.
When you’re looking from the bottom up, it’s a very, very quantized world. But when you start to look at giant, complex things like us from a high level, you have to begin from a more continuous perspective.
I think you’ve hinted that there are natural convergent patterns in computation. Can we sort of get a convex hull of your philosophy?
We could try. I hesitate to say I’m an anything-ist, but functionalist comes closest.
Functionalist.
Yeah. The reason for that is that in the old days, in the 19th century, we used to think that to be alive meant that there was some vital spirit or vital force that living things have and dead things don’t. As we started to figure out that the laws of chemistry were the same for living things and dead things, and that urea can be synthesized in a test tube and so on, those ideas really went out of fashion, and we moved into a very strict materialist kind of perspective.
Right, or everything is just physics. I mean, I was trained as a physicist. I believe in physics fully, but I also think that there is more to life, in the sense that if everything is just physics, then you have no way of saying what it means for you or me to be alive. To understand what it means to be alive, I think you have to come to grips with the idea of purpose. You have to bring teleology back into the equation.
What I mean by that is that a kidney is not just a collection of atoms. It’s an organ that performs a function, right? The function is to filter urea. If you implant an artificial kidney that works on totally different principles but also filters urea, it’s an artificial kidney. It’s still meaningful to say that.
That means there is something about the word “kidney” that means something which goes beyond the matter that the kidney is made out of. Conversely, if I come back from the future and show you an object, and you’re like, “What is that?” and I tell you it’s an artificial kidney, there’s nothing about this set of weird carbon nanotubes and so on inside that would say to you, “Kidney.” It’s just that if you happen to implant it in a body and sew it in the right way, then all of those relationships would show up in the right way for your body to persist.
This idea of things serving functions for other things, and functions only having meaning in the context of yet other functions, is ecological. I think this is really central. A rock on an inanimate planet has no function. If I break it in half, I now have 2 rocks. But a living thing has a function.
The hallmark of function is multiple realizability, just like Turing talked about for Turing machines. If you have a need to make ATP for energy inside your cells, you’re going to have multiple pathways for doing it, because sometimes the aerobic way works and sometimes you need the anaerobic way. Whenever you start to have multiple pathways—wings in insects, wings in bats—there is a function in play.
The alternative position would be essentialism. Folks like Anil Seth and John Searle think that certain types of material have a certain type of causal graph. For example, brains might give rise to consciousness, and if we simulated a brain, it wouldn’t have the same causal graph; therefore, it would be different.
But I would like to—we’ll just park that for the moment. It seems a little bit like you’re talking about this as a computer software architecture diagram. It’s like that Ship of Theseus type of thing, where we can swap things out and ask whether it’s still the same thing.
But I think path dependence is very important. The kidney evolved; it has this rich phylogeny of evolution. When you replace it with something that came from a different substrate, which has a different provenance, then it’s almost like it is a kidney now and it works now, but it breaks the ecology.
Like, imagine in an ecology if I swapped a plant out with an artificial plant and kept doing that. It might work now, but doesn't that affect its future trajectory?
Yes, it does. But that's exactly what symbiogenesis is all about.
Often, you will have a repurposing of something that was designed, if you like, by nature. One of the cool things about the BFF experiment is that it shows you how intelligent design can happen without any intelligent designer. Something that was designed for one purpose, or to serve one function, can come back around and serve another function. That brings a whole different contingent history with it.
The RSV example that I gave—the ability to fuse cell membranes together—came from a virus whose original purpose had nothing to do with building placentas, but it gets incorporated and repurposed. This is the kind of bricolage that life is made out of.
I think that kind of replacement, parallel pathing, and so on doesn't just happen when we make artificial kidneys. It's happening all the time in nature, and is the very hallmark of life. So, yes, I disagree strongly with Anil Seth and with John Searle on this point.
The brain-prosthesis experiments that you've alluded to—the idea that if you took an emulator or a simulator of a neuron and plugged it into your brain so that its inputs and outputs were connected to the other neurons, then the other neurons wouldn't know the difference. What if you did that for half of your neurons, or for all of them? Would your consciousness get dialed down even if you behaved the same way? Of course not.
For me, your consciousness is obviously a function of the relationships of all of those things with each other. It doesn't mean that it's as simple as a computer program where you can just substitute one subroutine for another. We've made computers very abstract in that way, but biology is wet and messy. The interfaces are complex and hard.
This same idea of multiple realizability and repurposability is the very stuff of life.
What is your position on consciousness? What is it? What's its purpose? Is it epiphenomenal? Can it be measured?
Yes, great question. I think that the idea of philosophical zombies, which David Chalmers has talked about—the idea that maybe something could behave just like you or me but be dead on the inside, not have any experiences, not feel anything—is actually a lot less coherent than it sounds.
I'm a functionalist about consciousness, too, and what I mean by that is twofold. One is that I don't think consciousness is some kind of epiphenomenon that we just happen to have for reasons that have nothing to do with our behavior. Nor do I think that it is somehow tied to anything about the way we're physically made. I think it is functional.
Why do we have it? In my team, Paradigms of Intelligence, we've been doing a lot of work over the last year on multi-agent reinforcement learning. The reason is that we're very interested in the precondition for symbiogenesis, which is symbiosis: cooperation.
When 2 things, or 700 things, or whatever, start to cooperate closely, that's the beginning of them really fusing together and becoming 1 thing. In order for 2 intelligent agents to cooperate, it turns out they have to have a theory of mind. They have to model each other and be able to put themselves in the place of the other.
We have a whole long theory called MUPAI about how that all works, but the CliffsNotes version of it is that it requires you to do induction over a universe that includes not only the game that we're playing, but also what is happening in your head and what is happening in my head.
In other words, you have to have a universe that includes yourself and the other, and that allows you to generalize over the class of you and me. I know that my internal state is happy when I smile, and when I see you smile, I know that you're happy on the inside as well. I can make that inference in the same way that if I see a bunch of peaches, I know that they're all the same object and I know what the backside of one will look like, and so on.
This ability to do psychological induction is really important for cooperation, and that's why we have it. One of the consequences is that we model ourselves, and we model our own models of others' models of our models, and so on. There's a kind of strange loop, as Douglas Hofstadter would have called it.
Yes, I love Douglas Hofstadter. So there's this kind of self-modeling, and then second-order self-modeling and third-order self-modeling, which could be applied to other agents. Of course, in the real world, we are computationally bounded. We can't make sense of all of the complexity. So when we do this modeling of other agents, our modeling is quite cartoonish and quite structured.
And it only goes up to sixth order as well, at most.
Oh, interesting. How does this affect—we haven't really spoken about agency yet—your ideas of purposeful behavior? Presumably, you could have a strong agent that's just doing something quite trivial. But when we have this collective intelligence and this information synchrony between agents, how does that affect your ideas of purposeful behavior?
I sometimes use the example of rowing to describe what's happening when purposes merge into a single purpose and consciousnesses merge into a single consciousness. There's this term that I learned from Daniel James Brown's book The Boys in the Boat: “swing.” That's when the 6 oarsmen—or 8 oarsmen, sorry—all achieve this kind of state where they're in perfect sync with each other.
You know it when you experience it. The boat acquires a soul, as it were. You all feel like you're pulling as 1. Boats with that property go a lot faster than boats where people haven't quite achieved that sync.
That, I think, is kind of what happens when we think of ourselves as being a self, despite the fact that our brain actually consists of a lot of parts. In the same way that the oarsmen, in some sense, began with their own purposes, their own self-models, and their own models of the other parts of the brain, through this process of subjective symbiogenesis, I guess you could call it, all of those wills become 1 and all of those selves become 1 self.
In hiring, for example, you want folks with high agency, but you also want alignment, which is the potential for this kind of synchrony. We often do a thought experiment on MLST where you can look at a boat or a flotilla of boats, and you're trying to draw a boundary. The boundary for the agent should be the minimal description. It should be where most of the agency is, where most of the planning and future modeling is happening.
Usually, it's the pilot; it's the driver of the boat. But you're talking about the situation where there is such synchrony and alignment between the agencies that almost the best intentional stance, if you like, is to draw a boundary around all of them.
I also think that there's not necessarily a single right answer. In my book, What Is Intelligence?, I talk about a few interesting cases. One of them is, for instance, the conjoined twins Abby and Brittany Hensel. I don't know if you've seen them on YouTube or on TV shows. Fascinating case.
These are 2 people who share 1 pair of arms and 1 pair of legs. Each of them controls 1 arm and 1 leg, so they're in a 2-legged race. They often speak in synchrony. They play volleyball and sports, they drive a car, and they can write emails with no problem.
They also sometimes have differences of opinion. They'll come together and apart in a remarkably fluid way, and all of that is done purely with behavioral cross-cueing, as Mike Gazzaniga would call it. Their nervous systems are separate: separate brains and separate spinal cords.
In that case, they're able to model each other extremely well because their entire lives they've been right next to each other. Another interesting case would be split-brain patients of the kind that Gazzaniga spent a lot of his career studying. Those are cases where, in adulthood, the brain is essentially cut in half.
Each hemisphere can only see the left or the right visual field, and controls 1 arm and 1 leg. The most fascinating thing about these split-brain experiments is that, from the outside point of view, it is obvious that there are 2 consciousnesses in there.
Each hemisphere is conscious of different things. You can make disjunctions between what shows up in the left and right hemispheres, and the left and right hands can be drawing different things, and so on. But if you talk to somebody who's a split-brain patient, they're always like, “Yeah, I'm still one person.” They will never admit that there are 2 people in there.
So is there somebody who is right and somebody who is wrong? No. This is entirely relational. It's a relational description. And the fact that, for them, they're the same person they always were, just occasionally something takes a little more work. Occasionally one hand will be buttoning the shirt while the other hand is unbuttoning it. It's just an inconvenience.
There are split-brain experiments as well, even just with a normal brain. And I can believe that we are sort of separately conscious in different parts of our brain. You get out of bed in the morning, and you must be a slightly different person. But we kind of gloss over that, don't we?
Absolutely. We make a narrative. The best, coolest experiments about this, I think, are the ones from Petter Johansson at the University of Lund. He's done a bunch; he was the one who discovered choice blindness. In these experiments, a subject is—I think the very first one was face choice blindness—shown 2 faces on cards and asked which one is more attractive, and you pick. Every so often, the one you're handed to explain why you thought that face was more attractive is the one you didn't pick.
So there's a kind of sleight-of-hand trick. The cool thing is that very few people notice that they're being handed the wrong face. There is no difference in the fluency or the latency of the description. You have an inner lawyer ready to spring up and justify whatever choice you made, even if it's not the choice you made. That narrative that you invent then influences your future choices.
It's as if we all make up a story about ourselves. And, of course, the reason is that we're all split-brain patients in a way. The left-hemisphere interpreter that generates the speech is likely not the same part of the brain that actually did the choosing, if you know what I mean, and yet all of those parts of your brain are invested in the idea that they're all in the same boat, that it's all one me. So they're all covering for each other.
In the same way, in a split-brain patient, if you show the non-left-brain-interpreter hemisphere “Stand up,” the person stands up, and you ask them, “Why did you stand up?” They'll say, “Oh, I was thirsty. I'm going to the kitchen for a drink of water.”
It's the same thing with artificial intelligence. It's becoming more sophisticated, and there's the social question. I suppose, actually, you can think of it as a Ship of Theseus for society. We're going to have agents embedded in society, and we're going to form a large collective intelligence. Do you worry about that future? I mean, what do you predict is going to happen?
Well, there are certainly things that I worry about. I don't want to come across as a Pollyanna. I'm worried about polarization. I'm worried about disinformation. I'm worried about our political and economic systems not necessarily being fit for purpose in the world that we'll all be living in in 20 years. But I'm certainly not concerned about a lot of the kinds of things that I hear Eliezer Yudkowsky talking about, for instance.
One of the reasons that I feel very differently is because I feel like human intelligence, in the usual sense that we think of it, is already a collective phenomenon. We're not that smart individually. We're not that much better individually than our primate cousins. It's only because we get together in large societies of millions and billions of people that we can do these amazing things, that we can transplant organs and go to space, and so on. Individually, we're just not all that.
So for me, AI is actually a part of human intelligence. It's literally already the same thing. I find it very interesting that we only achieved general AI when we began to literally train the models on reams and reams and reams of human language. So AI was human intelligence from the start.
I suppose the thesis of Eliezer is that it's possible to have artifacts which are dramatically more intelligent than we are. Maybe you think there's some kind of a limit, but do you think, in principle, that we could build artifacts which are significantly more intelligent?
Well, I think that collective humanity is already vastly more intelligent than individual humans. In that sense, and in many cases, it operates at very different time scales. For instance, I think these things are already true.
In a sense, our biggest difference is about thinking of it as an other versus already as a part of ourselves. What do we even mean by human? There was a wonderful paper from 2006 called “The Science of Cycology.” I'm not remembering her name, but she is a psychologist, and “The Science of Cycology” is spelled C-Y-C-L-O-G-Y. She asks people to draw bicycles.
First, they say, “Do you know how a bicycle works?” Everybody says, “Yeah, of course I know how a bicycle works.” “Okay, draw one.” Nobody can draw it, even if it's just looking at a sketch of a bicycle and saying, “Okay, where does the chain go?” or “Where are the pedals?” Most people don't know. They make some very fundamental error in this. It's a very funny paper, but the point is that we all have these illusions about what our own knowledge is, what our own capabilities are, and what our own intelligence is.
We already have this thing, in the sense that we identify what we think of as our intelligence with something that is actually in a bunch of other people and a bunch of other stuff around us. We do that kind of unconsciously. So for me, there's not really a discontinuity between what's already going on and AI. It's really just more of that.
Interesting. I think they would make the argument that you could build a single artifact which is more intelligent than the totality of humans. But just parking that to one side, I spoke with Judith Fan. She's a wonderful professor at Stanford, and she's done studies on drawing, comparing how humans draw to computers using CLIP models and stuff like that.
She found something fascinating. Because we have quite an abstract understanding, when we make sketches, she was grading it on the progression—progression 1, progression 2, progression 3—and we start very coarse and very abstract, while AI systems start with the edges and the details. And that, to me, indicates that AI models today don't really understand things at a very deep, abstract level like we do, perhaps because we have this compositional synthesis of knowledge that we were alluding to earlier. Do you see that as a gap?
There are a few questions, I guess, hidden in there. One of them is: Do I think of LLMs, for instance—of today's frontier models—as being less than or different from, in some basic way, our brains? What are those gaps?
First of all, they're obviously very different. Their architectures are different. They're trained in a very different way. The remarkable thing for me is actually how convergent a lot of their properties are with those of brains, despite all of that.
The fact that you find internal representations in many of them that surprisingly resemble ones you can measure in human brains—these Brain-Score-type measures from Martin Schrimpf and colleagues—or that sensory modalities in humans can be reproduced remarkably well even by models trained on pure language is really remarkable. It speaks to how much is encoded in language, how much of what is encoded in language is a reflection of the architectural properties of our brains and umwelts, and how much of that is then reconstructed, essentially, by those models.
Now, the question of what we draw first when we draw a picture and how that all works—I mean, remember that image-synthesis models like CLIP or what have you are working in pixel space to begin with. Diffusion models, by the way, work very differently from various other kinds of models. We now know that you can drive a robot with a transformer.
If you give one of those robots a paintbrush—or a pen—and you say, “Now draw,” what it will draw is going to be very different from what you get from a diffusion model that starts filling in pixels. And, for that matter, all of that is different from what happens in your own head when you're visualizing something. So I think a lot of this is not so straightforward to analyze because of all the differences in the way that AI and representation space work.
I do think that today’s models are highly compositional. Even with many of those original image-synthesis models, the fact that you could say, “A teddy bear at the bottom of the sea playing with the Speak & Spell,” or whatever, and it’ll do it tells you that they can compose. Again, are there capabilities like ours? No. There are definitely places where they’re better, places where they’re worse, and places where they have surprising gaps. So it’s different, but I wouldn’t say that there’s a fundamental lack of composition there at all.
I think, if anything, the biggest gap between transformer-based models and what we do is actually narrative memory, or being able to form long-term memories and, in that way, have a kind of persistence of a self over long periods of time. They don’t have that yet.
I’m conflicted. You are pointing to this universal representation hypothesis. I think Chris Olah popularized it with some of his visualization experiments, and it’s true: the representations are very convergent. Other things lead me to believe that the models produce these superficial impostors—that they give you exactly the right answer but for the wrong reasons. One of the hints of that is when you do variations on the input: it’s not robust.
There’s the Turing machine argument as well. These LLMs are finite-state automata, but they can access tools that are Turing-complete. So perhaps we could say the system is Turing-complete, but I don’t believe that ChatGPT is effectively searching the space of Turing machine algorithms. It hasn’t been trained to do that, but it is surprisingly robust with the ARC challenge. It can do really well, especially if you do some evolution, some refinement, and so on. So it feels like we’re knocking on the door, but there’s something missing.
I think that in many of those cases, we’re not doing a fair human comparison. This is a little bit similar to our illusions about knowing how bicycles work and so on. I hear a lot of people say things like, “Look at this case where we just flip the logic: we change it from do to don’t, and then it gets it wrong 30% more often,” and so on. My first question is always, “Have we done the human baseline?” It turns out that, surprisingly often, the human baseline shows the same property.
This doesn’t mean that humans are incapable of doing the fully robust, fully general version of these things. If you’re a logician, or if you think about it carefully, you can really write down your premises and be super robust to flipping the “nots” in the way something is formulated. But most of us don’t operate that way most of the time. We’re highly susceptible to logical illusions, cognitive illusions, et cetera, which turn out, in many cases, to be surprisingly similar to the machine case. So I’m kind of unmoved by a lot of those. I think often we’re being a little sloppy about how we do it.
It’s certainly the case that transformers aren’t searching systematically over all possible Turing machines. We don’t know how to do that. You have to take shortcuts of various kinds in order to make that whole problem of induction over programs computationally tractable, whether you’re a brain or a transformer.
Blaise, thank you so much for joining us today. It’s been an honor.
Thank you. Thank you for the really thoughtful questions.