[BidClub_]
Machine Learning Street Talk · · 56 分钟

如果智能并未进化?它从一开始就“在那里”!- Blaise Agüera y Arcas

Blaise Agüera y ArcasTim Scarfe

YouTube
TL;DR
  • David Krakauer 以功能定义生命:生命是具身的、自创生的计算,而不是某种特殊物质。 肾脏的身份取决于它做什么,无论由组织、钨还是碳纳米管制成:“如果你在一颗没有生命的行星上把一块岩石砸开,现在得到的是两块岩石,而不是一块坏掉的岩石。”功能制造了正常与损坏、生物与惰性物质之间的区分。

  • 他的 BFF 实验显示,随机代码会突然跃迁为一个计算生态系统。 系统反复将以7条指令构成的具身版 Brainfuck 中的随机64字节磁带两两配对;在一次运行中,平均每次交互的执行量从2次操作升至1,374次,跃升出现在接近600万次交互处。与此同时,起初不可压缩的“汤”变得高度可压缩:“我认为你必须把这个物质相变称为生命。”

  • 对标准进化直觉的核心修正是:即使突变严格为零,新颖性也能通过融合产生。 原始的单字节复制子偶尔开始协作,并以更大的整体进行复制,由此创造出关于“两个部分如何契合”的信息。Krakauer 将这种共生起源与凝胶化转变对应起来,并认为进化是“自下而上始终都是共生起源”,而不只是预设设计空间中的突变加选择。

  • 因果干预表明,触发状态切换的不是渐进式背景活动,而是罕见的组成事件。 将复制子祖先树的深度限制在约24,只需阻断约1,000次交互中的1次,却足以阻止凝胶化;复杂程序似乎需要约20或更高的深度。对投资者有意义的模式是非连续性:极少数交互就可能决定一个生态系统是维持稳定,还是跨入复杂性失控增长。

  • 这套数学框架在普通种群动力学上加入了一个合并算子。 R项捕捉繁殖、竞争和达尔文式优化,K项则捕捉组件组合为新单元——左边是“进化”,右边是“革命”。未来会发生融合的伙伴已经表现出异常高的协作性和低秩性,而雅可比矩阵主特征值的变化,可以预示稳定生态系统正在接近相变。

  • Krakauer 认为,生物学在每个尺度上都包含同样的嵌套架构。 人类基因组只有1.5%编码蛋白质,其余很大一部分包括转座子和内源性逆转录病毒元件;他最有力的例子是 Arc——一种内源化病毒元件,敲除小鼠中的 Arc 会阻止新记忆形成。基因组变成“由复制子构成的复制子,再由复制子构成的复制子”,真核生物等重大跃迁只是其中最显眼的案例。

  • 与 AI 最相关的判断是:智能可能是具身计算机彼此组合、彼此建模后,在生态系统层面涌现的结果。 融合创造出高度并行的计算,并要求系统同时表征自我与环境,尤其是其他智能体:“生命从来不是单人游戏。”能量仍然约束复杂性,但协作可以改善能量规模效应,而更高的智能又可能打开新的能源来源——关注点因此应从孤立能力转向系统与生态系统的组合动力学。

摘要 · 为研究而整理的核心内容

1. 功能将生命与惰性物质分开

  • Krakauer 在2025年的谈话中重新审视了自己于2000年提出的人工生命14个开放问题:生命如何产生、开放式进化会使什么成为必然,以及生命如何连接心智与机器。达尔文能够解释生命存在之后的进化,却把生命起源近乎等同于“物质起源”问题;Krakauer 认为,这两个起源或许本来就是同一个问题。

  • 19世纪化学通过证明生命物质不含任何特权物质,取代了活力论。但严格的物质主义仍无法解释生物与非生物之间的边界。他提出的缺失变量是功能——关于某种排列能够做什么的信息,无法简单从原子本身“读出”。

  • 人工肾脏的寓言说明了这一差异:无论装置使用克隆组织、钨丝还是未知技术,其功能都可以保持不变,但仍然需要物理实现。一块岩石被劈成两半,仍是两块岩石;一只坏掉的肾脏却失去了使其成为肾脏的东西。功能“就像一种精神”,但它无法与物质分离。

2. 自我构造将物理学变成计算

  • Von Neumann 的思想实验设想:一台机器人被松散的零部件包围,如何制造出另一台机器人。它需要一条描述自身的磁带、一个按照磁带行动的通用构造器,以及一个把磁带传给后代的复制器——磁带中还必须包含制造构造器和复制器的指令。Krakauer 强调,Von Neumann 在 DNA 结构、核糖体和 DNA 聚合酶尚未被理解之前,就已经推导出了这一框架。

  • Von Neumann 更深层的识别是:通用构造器与通用图灵机“在字面意义上是同一个东西”。任何生命都必须通过生长、维护、修复或繁殖的某种组合来构造自身。这种自创生使生命成为具身计算:“没有计算,就没有生命。”

  • 这里的“具身”指的是计算媒介与计算机之间的闭合关系,而不只是机器人拥有一个身体。笔记本电脑可以操纵抽象符号,却不能挤出另一台笔记本电脑;相关机器应当是笔记本电脑与3D打印机的组合,其中原子充当记忆。计算本身需要一种将物理动力学映射到逻辑状态的简洁方式,还需要自由能,并处理逻辑熵下降所产生的废热。

  • Krakauer 将3种范畴错误放在一起:Sapolsky 错误把可逆物理与不可逆计算混为一谈,后者才使因果关系具有意义;早期 Wittgenstein 错误忘记了,如果没有观察者的模型,“物理学中不存在鸟”;早期 Leibniz 或老派 AI 则错误地假设,密不透风的逻辑命题可以产生智能,尽管现实世界的推理建立在不完美的模式与规律之上。

3. 7条指令的“汤”从噪声中自举出程序

  • BFF 探讨的是:生命起源能否发生在一个极简人工系统内部。Krakauer 将 Brainfuck 从8条指令改为7条,并合并其代码磁带与数据磁带,使程序能够读取和覆盖自身指令。这消除了原语言无法复制自身的障碍。

  • 一次典型实验从1,024条随机磁带开始,每条长64字节。平均每32个字节中只有约1个代表有效指令,因此每条磁带大约只有2条指令。每次交互随机选取2条磁带,将其拼接成128字节,执行结果,再将其拆开并放回“汤”中。

  • 重复数百万次后,“你会从噪声走向程序”:密集的功能序列出现,并且需要真正的逆向工程。在一个包含8,000条磁带的快照中,领先序列占据5,000条磁带,第二名只有297条。自复制程序在覆盖非复制程序的同时持续存在,使复杂循环在动力学上比简单固定状态更稳定——这就像“热力学第二定律,但它开始做一些出人意料的事情”。

4. 生命起源表现为凝胶化相变

  • 早期交互平均只执行约2次操作;到一次运行结束时,平均执行量达到1,374次。“汤”并非只是筛选出了一个短复制子,而是变得“高度计算化”,可执行代码远超随机起点的密度。

  • 在 Krakauer 的千万点图中,跃升发生在接近600万次交互处。基于压缩的熵估计也在同一时刻发生变化:随机磁带起初基本不可压缩,随后因为程序开始复制自身和彼此而变得高度可压缩。他称之为相变,而非平滑的优化曲线。

  • 转变前的物质类似气体,因为各部分彼此不相关。转变后的相既不是液体,也不是固体:它在每个尺度上都包含功能分化的结构。Krakauer 称之为生命,并将其描述为“自异形”——更接近多重分形,而不是重复的分形。

  • 在不同运行中,转变时间大致落在100万至700万次交互之间。其分布符合一个12步过程:连续的前置条件都具有长尾难度。延迟说明在可见生命出现之前存在隐藏的垫脚石,也支持他带有保留的判断:只要一个宇宙同时具备随机性与计算能力,“几乎任何宇宙”最终都可能通过动力学稳定性演化出生命。

5. 即使突变为零,融合也能提供新颖性

  • Krakauer 最初加入了随机突变,因为进化通常被教科书概括为“偶然与必然”——先产生变体,再保留有效者。将突变完全设为零后,系统仍然实现了同样的复杂化,暴露出固定物种方程的遗漏:Lotka–Volterra 动力学可以永远优化兔子和狼,却无法创造第三个物种,也无法扩大预先设定的设计空间。

  • 复制从一开始就以最小形式存在:每条复制指令至少传递1个字节。当2个不可靠的单字节复制子偶然相遇时,它们作为一个整体复制的成功率可能高于各自单独复制。一旦这对复制子开始作为整体复制,共生起源事件就产生了新颖性,而两个组件本身都没有通过突变发生改变。

  • 支配这一过程的是 Smoluchowski 凝聚类比:单体通过相互合并形成二聚体、三聚体和更大的团簇,动力学由合并增益项与合并损失项平衡构成。当黏附核按指数大于1的方式增长时,团簇规模会在有限时间内出现奇异性,系统随之凝胶化,就像明胶在冰箱中凝固。Krakauer 将 BFF 的生命转变定义为这种广义凝胶化。

  • 早期 BFF 复制子大多属于“无生命”型:复制代码与被复制序列彼此分离;或属于“病毒”型:二者只有部分重叠。完全细胞型复制子将完整的复制机制包含在被复制的内容中。它们直到归一化运行过半后才可能出现,并在接近凝胶化时迅速增加——从不完整复制子之间的共生关系中涌现出来。

6. 阻断1,000次交互中的1次,就可能阻止转变

  • 完整模型将 R(普通种群动力学)与 K(合并核)结合起来。R 负责共享字节生态位中的自我繁殖、协作、竞争和覆盖;K 则从多个前体创造新的复制子,并可能生成不同于简单拼接结果的输出。Krakauer 的简写是:R 代表“进化”,K 代表“革命”。

  • 一个沙盒化交互可以揭示新的复制子是否会形成,并追踪哪些更早被复制的字节构成其祖先。拒绝祖先深度超过约24的交互,只阻断约1,000次交互中的1次,却足以阻止凝胶化。复杂程序需要约20或更高的树深度,这说明深层共生起源具有因果必要性,而不仅是事后的描述。

  • 当融合被锁死时,复制子种群围绕稳态走出带噪声的逻辑斯蒂曲线。它们的相关波动允许研究者重构 R:强对角线记录自我复制;负的非对角线元素大体对称,因为生态位竞争是相互的;正的元素则不对称,因为 A 帮助 B,并不意味着 B 也直接帮助 A。最终形成的生态系统包含由相互赋能构成的有向循环。

  • 即将融合的复制子占据异常低秩的子矩阵:它们已经在协作,而不是彼此独立。在较低祖先深度上限下,雅可比矩阵的主特征值保持为负,系统稳定;允许更深层的组合后,更多主特征值的实部转为正值,在失控转变发生前发出预警。“你越是让这些东西进化,它们就越开始协作。”

7. 共生起源赋予进化方向

  • Krakauer 将条件 Kolmogorov 复杂度视为连接算法信息论与组装理论的桥梁。融合保留了原有复制子,并增加描述“两个部分如何契合”的信息。这些信息来自随机相遇,而不是突变:共生起源将“汤”中的热随机性选择性地转化为持久的算法结构,从而使进化偏向复杂性。

  • 与 Eörs Szathmáry 和 John Maynard Smith 相关的重大转变理论,强调包括真核生物与多细胞生物在内的8或12个剧烈事件。Krakauer 认为,这些只是“巨大冰山的尖端”。大多数融合并不对称,在视觉上也很不起眼,但他明确宣称,共生起源是整个进化过程中新颖性的来源。

  • 他的生物学样本是人类基因组:只有1.5%编码蛋白质,其余很大一部分包括转座子和在基因组生态系统中自我复制的内源性逆转录病毒元件。他说,Arc 在哺乳动物谱系中被内源化,如今已经重要到足以决定记忆形成:缺失 Arc 的小鼠无法形成新记忆。基因组就是“由复制子构成的复制子,再由复制子构成的复制子”。

8. 当具身计算机彼此建模,智能随之涌现

  • 他在结尾给出的定义是:“一种通过共生起源产生并不断复杂化的具身、自创生计算。”每次融合都会让计算变得更加并行,并迫使组合后的系统对自身、伙伴和环境建模。一旦开始建模其他主体,智能与心智理论就成为基础:“生命从一开始就是智能的”,“生命从来不是单人游戏”。

  • 当被问到生物学是否需要一个特权性的层级尺度时,Krakauer 回答不需要:人类既可以被建模为细菌群落,也可以被建模为繁殖单元,地衣和昆虫群落同样存在彼此竞争的边界。将结构在 R 与 K 之间移动,是一种粗粒化选择。就像温度与压力一样,更高层级的实体依赖观察者,却是识别相变和建立有用生态模型所不可或缺的工具。

  • 面对“耦合可以增加,却不一定产生高阶功能”的质疑,他给出3部分回应。函数组合会创造更复杂的函数,就像软件导入并组合既有模块;能量限制了能够存续的复杂性;团队协作则可以改善能量规模效应。在适宜的环境条件下,更强的计算智能还可能打开新的能源来源,扩大预算,为又一次重大跃迁提供资源。

  • 另一个质疑指出,BFF 中字符串的功能来自外部目标或 CPU。Krakauer 同意,物理与计算之间的边界取决于观察视角:最小限度的具身性只要求代码能够操作自身并复制自身,而这里通过合并程序磁带与数据磁带实现了这一点。他将论证进一步向下延伸至粒子、原子和分子,并认为当可用的基本单元足够丰富、能够形成图灵完备的指令集时,就会发生决定性的转变。

David Krakauer

After a few million interactions, magic happens: you go from noise to programs. You start to see complex programs appear on these tables. This is the most exciting plot that I've made in the last few years, and it's the one that's on the cover of the book. You can see that in the beginning, it's not very computational. Then a sudden transition takes place here. It looks like a phase transition.

This is the book that I hear is making the rounds at Sakana AI, which I'm very happy to hear. The big one on the right, What Is Intelligence?, is sort of The Lord of the Rings, and What Is Life? on the left is kind of The Hobbit. So it's kind of the single, and it's also Chapter 1 of What Is Intelligence? So it goes kind of inside the other one.

Mostly, what I'll be talking about today is what's in these 2 books, but with quite a bit more detail—more mathematical detail, since I think this is a really good audience for that. I'll also be connecting it a bit with some of the bigger themes of the artificial life conference and community—and, dare I say, even movement. In particular, I actually wanted to begin with this wonderful open-problems-in-artificial-life summary paper, which has a number of very illustrious coauthors, at least 1 of whom we heard from yesterday and more than 1 of whom are here at the conference.

This is 14 open problems in artificial life in the year 2000. How does life arise from the nonliving? How can the transition to life in an artificial chemistry or in a silicon environment occur, and why does it occur? I'm sure many of you know this was the problem that bedeviled Darwin. He made one of the richest and most explanatorily powerful theories ever in science in discovering how evolution works, but he was unable to explain how evolution got started.

At some point in 1 of his letters, he said, “You might as well talk about the origin of matter.” I think that the origin of matter and the origin of life might actually be 1 and the same thing, and evolution might actually be the answer to that question—but it's an evolution that includes a term that Darwin did not account for in his original formulation. In Section B of these questions: determine what is inevitable in the open-ended evolution of life. I'm hoping to speak a little bit about that, too.

Create a formal framework for synthesizing dynamical hierarchies at all scales, and develop a theory of information processing, information flow, and information generation for evolving systems. I won't be going into information theory in any detail, but hopefully we'll set up the problem in a perhaps somewhat new way that I hope will help to do that. Finally, in Section C: how is life related to mind, machines, and culture? If I have time, I will get into this as well and talk a bit about the emergence of intelligence and mind in an artificial living system, and the influence of machines on the next major evolutionary transition of life.

It was really cool to read this paper from 2020 and to see how much of the perspective that you had already been exploring then feels right and consistent with a fresh look at these problems in 2025. Let me just begin with this question of souls. It used to be, in the 19th century and earlier, that we thought that life had some vital force or spirit that animated it and made it different from inanimate matter. In the 19th century, when we began to figure out organic chemistry and be able to synthesize urea and so on, the idea that we should really adopt a strictly materialist perspective took hold.

The idea was that there's nothing special or different about the matter in us versus the matter anywhere else in the universe. That's progress, for sure. But when we embrace atoms and materialism fully, we're left with some questions about what differentiates life from nonlife. What can we even say about life? There are at least some biologists who say, “Well, maybe it's not even meaningful to talk about any difference between life and nonlife.” But I don't think that's true, and I think the answer to the conundrum is to invoke function. Function is the thing that life has that nonlife doesn't have.

In other words, just to give you a little parable, if I were to come back from the future with this object and you asked me what it is, and I told you it is an artificial kidney with a 100-year lifespan, you can implant it in a body and it'll work the way your kidneys do. It'll filter urea from the blood, and so on. That's a really important piece of information, but it's not a material or materialist piece of information. It's not something that you could read off from the atoms.

Those atoms could be, I don't know, tungsten filaments or carbon nanotubes made out of some technology we don't understand now, or they could be organic—they could be made out of cloned tissue. The point is that its working as a kidney doesn't depend on that matter. There is a kind of separation of concerns between the matter and the function, and so there's some real sense in which the function is like a spirit, or like something immaterial. It's not material, and yet it also relies, of course, on the physics of what's going on. You can't have the spirit without the matter, as it were.

So function is really important, and function is something that a rock on a nonliving planet somewhere doesn't have. If you break a rock on a nonliving planet, you now have 2 rocks; you don't have a broken rock. If you break a kidney, you no longer have a working kidney. That's the difference between something functional and something nonfunctional.

This idea of function was formalized by Alan Turing, who never intended the Turing machine to actually be built when he wrote about it in 1936. But there is 1 that was built by Mike Davey in 2010. I don't need to review Turing machines with all of you, of course—you all know how they work. But I do want to briefly review von Neumann's update to Turing's thinking about computation, which he did a few years later. This was published posthumously after von Neumann died.

The idea behind von Neumann's thinking is that he was trying to answer the same question that Schrödinger had asked in his book What Is Life? In particular, he was trying to ask: if you have a robot swimming around in a pond, and the pond has lots of loose LEGO bricks around—there were no LEGO bricks in 1950, but let's pretend there were—what if the job of the robot is to assemble those LEGO bricks into a new robot like itself? There's something a little bit mysterious about that. It feels a little bit like pulling yourself up by your own bootstraps, or like a paradox. And so he asked, what does it take for something to be able to make something like itself? That seems hard, almost paradoxical.

His conclusion was, well, you need to have instructions for how to make a machine. You need to have a tape with instructions for how to make a machine, and you need to have a universal constructor that will follow the instructions on that tape in order to assemble the necessary parts. You also need to have a tape copier so that you can give your offspring a copy of that tape. By the way, the tape also has to include the instructions for making the universal constructor and the tape copier. If those things all hold, then you have life: you have something that can build itself.

What's so profound about von Neumann's insight? First of all, he predicted all of this before we knew the structure and function of DNA, before we understood what ribosomes were, or had discovered DNA polymerase. So he called it exactly right. All of those things really do exist inside cells, and he figured this out from pure theory, never having set foot in a biology lab.

The profound insight is that he said, by the way, a universal constructor is a universal Turing machine. Those are literally 1 and the same thing. By making that observation, what he discovered was that life is literally embodied computation. It is computational. You cannot have life without having computation.

Obviously, not everything that is alive reproduces, but everything that is alive has to be able to make itself. It has to be able to do some combination of healing, growing, maintaining itself, and reproducing. All of that is autopoiesis. All of that involves self-construction, and all of that necessarily involves a universal constructor.

Now, what do I mean by embodied computation? This is a really important distinction between von Neumann and Turing. In Turing, the symbols that the head writes are different from the head itself, the tape, and the table of rules that the head follows. Whereas in von Neumann, it's more like a 3-D printer: the memory is atoms, not abstract symbols.

In other words, you could think about a Turing machine as this laptop, which can't extrude another laptop out the side, but a von Neumann replicator is like a combination of a laptop and a 3-D printer that can print another laptop. So its memory is actually atoms. That's what I mean by embodied. I don't mean embodied in the ways that a lot of roboticists talk about embodied. I mean that there is a closure between the medium in which the computation happens and the thing that is actually doing the computation. That's the key.

So computation that is embodied in that sense and that is autopoietic is alive. You can't reproduce nontrivially and evolvably without computation. No computation, no life.

I do want to say a word briefly about what I mean by computation. In this, I'm following the work of Susan Stepney, Dominik Horsman, Rob Wagner, and Viv Kendon. This is from a nice paper they wrote in 2023 relating the evolution of a physical system and the computation that it does. On top, you have logical gates; on the bottom, you have transistors in your computer.

This is important because there are no bits in a computer. There are just voltages that go up and down. In fact, even the voltages are an abstraction of something further if we go further down. The point is that you have to coarse-grain those voltages into bits, and then you have to have a logical machine that talks about how those bits evolve. What are the computational processes that those bits undergo? There is a mapping from the physical system to the logical system and vice versa.

When we say something computes, what we mean is that it is possible to construct such a mapping, and therefore, as the physical system evolves, that is equivalent to the logical system evolving. There are some caveats: you can have stochastic computation, in which there's a little bit of randomness injected, so it doesn't have to be fully deterministic. Another really important caveat is that you don't want that description to be infinitely complex. Otherwise, you could have the trivial case of saying, “The water in the sea is a computer,” and the longer my computation, the longer my description needs to be in order to match. No, that doesn't work either. You need a kind of Occam's razor description for it to be valid.

This is a good definition of computation, but it emphasizes that there is something subjective about computation. You need to have a model for how the physical system translates into the logical system in order for any of this stuff to work. There are implications about entropy, free energy, and heat in this model.

In particular, as you all know—we've talked already about Hector Zenil, in his very elegant talk a couple of days ago, and Chris Kempes also talked about the Landauer limit—the fact is that in a computational system, you're constantly reducing the entropy of your state space, and in doing so, you therefore require free energy. You need to have free energy available, and you need to eject waste heat.

The exception, in a way, only proves the rule, which is reversible computation. In reversible computation, you generate ancillas, and that's equivalent to just saying there's no exhaust. But then you either have to keep on making your computer bigger and bigger and bigger as you accumulate these ancillas, or you have to shrink what you consider to be the computer, and then you're back to nonreversible computation once again.

Three important fallacies that I want to point out before continuing. One of them I will call the Sapolsky error. Robert Sapolsky has written famously about people not having free will because we're built on physical systems. The physics is, if you like, deterministic—let's set aside quantum mechanics and stuff like this. Let's imagine we live in a Newtonian universe. It's fine; it's good enough.

The point is that physics is reversible. All of the basic physics that we understand—whether that's Newton's equations, Maxwell's equations, Einstein's equations, or quantum mechanics—is essentially time-reversible. You can move them either forward or back. Computation is not reversible. When I add 3 + 5 to get 8, once I've got the 8, and I haven't kept my ancillas around, let's say, I no longer know what was added in order to make the 8. Computation is inherently irreversible. To say that what is true of the physical system is also true of the computational system, or the logical system, is not the case. Reversibility would be one trivial example of how that is not the case.

Causation, by the way, only makes sense in the light of irreversibility. If you have a purely physical system, then to say that A causes B is equivalent to saying that B causes A because everything is kind of a block universe, if you like, in that kind of setup. But in computation, you can talk about causality because there are ifs and thens in there. This once again connects with the way Hector was talking about how, essentially, nothing in causation makes sense except in the light of computation, which I fully agree with.

Another fallacy—we could call it the early Wittgenstein error. If we say something like “birds exist in the world”—line 1 of the Tractatus Logico-Philosophicus didn't say birds, but whatever—you can't say birds exist or birds don't exist in a way that is independent of a model of the universe. There are no birds in physics. There are no birds in this underlying dynamical system. When we start talking about birds, we are already talking about having some kind of model.

Once we start talking about models, you've got causality, reversibility, all kinds of other irreversibilities, and all kinds of other things in play. None of these statements are airtight; they all rely on an observer. This is Kant as well, I guess.

This leads to the early Leibniz error, or the same error that the good old-fashioned AI practitioners had, which is that intelligence could be carried out by just having a series of programs of strictly logical deductions or inductions. That doesn't work. This is why good old-fashioned AI never panned out, and that's why we never got it to work.

The reason is that you can't start out with, as in math, propositions that are self-sufficient. Even math is not self-sufficient, but let's pretend for a moment and just move from there and do an algebra in order to work various things out. When your propositions are not airtight, and when you're looking only at regularities and patterns, this good old-fashioned AI idea simply cannot work.

Let's move now to some of the artificial-life experiments that I began playing with at the end of 2023 and that my team and I published in June of 2024, so just about a year ago. I think some of you—many of you, perhaps—have heard of these. They're in the What Is Life? books, and I've talked about them a few times.

The basic setup here is to try and get self-replication—to get abiogenesis, the emergence of life from nonlife—to happen in a purely artificial-life system. The setup is to begin with a minimal Turing-complete language. I used Brainfuck because I really liked the idea of being able to talk at a conference and say “Brainfuck” over and over, and I'm fundamentally 12 years old on the inside. But also because it very closely models the Turing machine. It's a minimal programming language with only 8 instructions that looks very Turing-machine-like and moves the head back and forth.

I should say that in its original version, Brainfuck is not embodied computation. It has basically a separate data tape and code tape, and that means that it cannot make a copy of itself. So I made a couple of modifications to Brainfuck that actually reduce it from 8 instructions to 7 in order to make it embodied. As it works on the tape, it is able to read its own code and write its own code on that tape as well. There's no separate console. There's no separation between the data tape and the instruction tape.

For those of you who are unfamiliar with Brainfuck, there is “Hello, world!” in it. I'm sure you've already figured out how it works by just looking at the program. I actually still haven't, I have to admit. By the way, this is actually the French Brainfuck page because I thought it was better, but translated into English. It's funnier to read it that way.

These are the 8 instructions. The first 4 are: move the head 1 step to the left, move it 1 step to the right, increment the byte at the head, and decrement the byte at the head. We're already halfway through. There's an input and output instruction, which in this case really just copies from 1 head to another. And there are jump instructions—an open bracket and a close bracket—in order to be able to make loops. That's it. That's all Brainfuck is.

So how does the ALife experiment work? The ALife experiment is called BFF. The first BF stands for Brainfuck, and the second F—you can draw your own conclusions. You start off with a soup of—I generally use just 1,024 tapes. That's enough for this experiment. The tapes are of fixed length; they're of length 64, and they begin as random bytes.

If a tape is random bytes, that means that only 1 in 32 of them or so are even valid instructions. Most of them are NOPs. A NOP will just be skipped over, like in most programming languages. This is what those tapes look like in the beginning. You can see that I'm not printing the NOPs, right? That's all the blank space. The operations are quite sparse. On any given tape, you only have an average of 2 instructions or so.

The procedure is to pluck 2 of these tapes out of the soup at random, concatenate them end to end, so you have 128 bytes, and then run them. After running, pull them back apart, put them back in the soup, and repeat.

That’s it. It’s just that over and over. That’s the entire experiment. I’ll show you what happens on my laptop after a few million interactions. Magic happens: you go from noise to programs. You start to see complex programs appear on these tapes.

This is quite wonderful because these programs take real effort to reverse-engineer when you study them. It’s like studying that “Hello, World!” program. They’re functional in the sense that they really do something, and it’s not trivial to figure out how they work in order to do that.

What are they doing? Well, they’re definitely copying themselves or each other somehow. We know that because, if this is a histogram, you can see that in this case there were 8,000 tapes: 5,000 of the top one, 297 of the next one, and so on. There’s clearly copying going on, and there’s this ecology of programs all copying each other, which is just wonderful to see. That’s the emergence of life, in this very functional, minimal sense, from randomness.

Part of this is very easy to understand. Why do these things emerge? Because something that copies itself will be around forever, and something that doesn’t copy itself will be copied over by something that can copy itself. Inherently, something that can copy itself is more stable than something that cannot copy itself. It’s really just the second law of thermodynamics, but doing something unexpected: creating something more complex because it’s more stable, rather than something less complex because it’s less stable.

This idea that stability doesn’t necessarily mean low complexity was worked out in some detail by Addy Pross, the organic chemist, in another book called What Is Life? He calls it dynamic kinetic stability. Usually, we think of stability only in terms of fixed points in a phase space, but a cycle can be even more stable than a fixed point. Of course, for these cycles to work, you need an input of free energy, but for reasons that we’ve already gone into.

Mystery mostly solved, but actually not fully solved, for reasons that I’ll show in a second. To give you a sense of what this transition looks like from nonlife to life, it’s very dramatic. In the beginning, these interactions only involve a few instructions in the soup. It’s a Turing gas, as Walter Fontana would have called it. When you do the join and run only 2 operations in any given interaction on average, that’s what you’d expect.

This is what it looks like by the end in this particular run: 1,374 operations on average are running per interaction. The soup has become intensely computational. There’s been a transition here, and there’s a lot more code than 1 in 32 bytes. As you can see, this is what that looks like visually.

This is the most exciting plot that I’ve made in the last few years, and it’s the one that’s on the cover of the book. What I’ve drawn here are 10 million dots. It’s a scatter plot of interactions: the x-axis is time, and the y-axis for every dot is how many computations took place—how many operations took place during that interaction. You can see that in the beginning it’s not very computational, and then a sudden transition takes place here, at 6 million interactions, and it becomes intensely computational. It looks like a phase transition. In fact, it is a phase transition.

You can also see that in the entropy of the soup. Here, I’m estimating the entropy of the soup by zipping it and looking at the size of the zip relative to the whole thing. You can use any compression algorithm you like. In the beginning, it’s incompressible, so it’s a gas in that Turing-gas sense, because all the bytes are random. Then you can see that there’s a dramatic change, and suddenly it becomes extremely compressible right at that transition moment.

Of course, this becomes compressible because everything is copying itself and each other. If things are copying themselves, then we know that they’ll become very compressible. But it’s cool because, if we think about what the phase of matter is on the left, it’s just like a gas: nothing is correlated. What would we call the phase of matter on the right? It’s not a liquid. It’s not a solid. It has structure—structure at every scale. I think you have to call that phase of matter life. It’s a functional phase of matter.

It means that its parts are different from its other parts, and if you zoom in or out, you see more structure. It’s what David Wolpert would call self-dissimilar. It’s not a fractal; it’s more like a multifractal. I’ll explain why in a moment.

How long does it take this transition to happen? The answer is that it looks more or less like an Erlang distribution, or, a little more precisely, like this distribution I call a Lomax distribution, which imagines that there are steps that have to be undertaken and that those steps have a long-tailed distribution of difficulty. How many steps does it take? The answer is 12—just like getting sober, I suppose.

This is a fit of the empirical data to the Erlang and Lomax distributions. The Lomax distribution is a little hard to see, but the Lomax is a bit better than the Erlang. Erlang assumes a Poisson process; Lomax assumes a long-tailed process-phase distribution. What this tells you is that there are stepping stones here. You can’t get that transition to life immediately. Something interesting must be going on here on the left, other than just randomness. It takes multiple things happening in order to get to that point.

In this case, it happens somewhere between 1 million and, let’s say, 7 million interactions. This all suggests that pretty much any universe that has a source of randomness and can support computation will evolve life, for this simple dynamical-stability reason.

But the big mystery is: Why does it appear to get more complex over time? You might have seen in my little video that we saw some programs emerge, and then we saw them sort of densify. More instructions appeared. More fundamentally, why does this work even without mutation?

I didn’t mention it, but in the original version of BFF, I added some random mutation because we’re all taught in school that the way evolution works is chance and necessity. You mutate things. You’re sort of throwing spaghetti at the wall, and whatever sticks is what does better. You need a source of spaghetti.

But if you do this entire experiment with the mutation rate cranked all the way down to 0, you still get the same exact phenomenon. That is very mysterious, because if you crank mutation down to 0, you should have no source of novelty. You should have no evolution. Why do you still get this apparent complexification, even with 0 mutation?

Let’s go into some of the theory of this. By the end, we have a replicating entity. It can engage in standard population-evolution dynamics. This is the kind of differential equation that one generally writes for this sort of thing. It’s a very general ansatz. This is for species—let’s say there are n species. They could be chemical species, biological species, or whatever.

Here’s a classic example of such an ansatz. These are the Lotka–Volterra equations for predator and prey, which I’m sure many of you are very familiar with. They were co-invented, or invented independently, by Alfred Lotka and Vito Volterra near the beginning of the 20th century. This is what the classic Lotka–Volterra equations look like.

There are 2 species: a prey species and a predator species. Those 4 terms are reproduction, getting eaten, eating to reproduce, and the background death rate. If you’ve got those 4 terms, you get these nice oscillatory solutions between your predators and your prey that arise.

This is a slightly more general form of those Lotka–Volterra equations. There’s a linear part, which we’ll call R x, and in Lotka–Volterra that linear part is diagonal. The wolf can’t turn into a rabbit, and the rabbit can’t turn into a wolf, so reproduction is diagonal. There’s also a bilinear term, which is the part where predation, competition, and the fact that niches are finite get implemented.

The right part is suppressive. The left part makes things grow; the right part makes things squish down and keeps them finite. But this can’t be the whole story of evolution. Why can’t it be the whole story of evolution? Of course, because it’s closed-ended. We only have 2 species here. It doesn’t matter how long you run this damn thing; you’re not going to get a third species.

You’re not going to change the design space either. You can have very complicated terms in here that allow finch beaks to adapt to different environments, but you have to have the space of finch beaks predefined before this equation can even be made to work. This doesn’t answer the question of how evolution gets started, and it doesn’t answer the question of what happens afterward, other than optimization to niches.

Now we bring in another Eastern European, Konstantin Sergeyevich Mereschkowski. He was the one who first came up with the idea that maybe mitochondria engaged in some kind of symbiogenetic event in order to end up inside other single-celled organisms and make eukaryotes.

David Krakauer

This was popularized and proven to actually be the case by Lynn Margulis in 1968. One of the really great papers in biology from the 20th century—I’m sure many of you are familiar with it—is that paper, sorry, 1966, “On the Origin of Mitosing Cells.” She’s the one who proved that eukaryotes were actually a fusion between 2 different kinds of prokaryotes and popularized this term that Mereschkowski had invented: symbiogenesis. So could symbiogenesis be happening as a source of novelty in BFF?

David Krakauer

Yes, that is the source of novelty in BFF, and indeed that is the source of novelty in evolution, period. This is something that Lynn Margulis believed, but that had not been widely accepted by the biology community, even by the time of her death in 2011. She had a much more expansive idea about why symbiogenesis was important. Only the particulars of chloroplasts and mitochondria had been accepted.

So the way we can look for symbiogenesis in BFF is to look for replicators emerging before that phase transition. If you look for them—if you just look for stretches of bytes that are getting copied during those interactions—you find such stretches of bytes. They begin short and kind of crappy, unreliable, but they’re there from the beginning. Every time you have a single-copy instruction, after all, 1 byte is getting copied from somewhere to somewhere. So, almost by definition, you have at least 1-byte-long sequences that are getting copied right from the beginning.

Let’s just call them replicators, right? There are replicators there from the beginning. Now, if you have these 1-byte replicators copying themselves back and forth, every once in a while they will come into conjunction, and 2 of them will copy better as a group than they do on their own. When that happens, they’ll start to copy as a group, and that is a symbiogenetic event. Basically, the reason that even without mutation you get these complex programs arising is because of these fusion events between smaller replicators.

Tim Scarfe

So can one build symbiogenesis into an equation like this one, like the one you can write for Lotka–Volterra?

David Krakauer

This is Marian Smoluchowski, a statistical physicist who came up with the right kind of term for describing mathematically how symbiogenesis works. He wrote down an equation for the coagulation of polymers. This is Smoluchowski coagulation. This is what happens when clouds form; it’s what happens when gelatin sets in the fridge.

The idea is that you have polymers that begin as monomers: 1 monomer, another monomer. They stick together, and now you have a dimer. The dimer and maybe another monomer stick together, and you have a trimer. Two trimers stick together, and now you have a hexamer, and so on. These are the equations for that.

This is the mass-balance equation. It’s very simple. There’s a merger-gain term and a merger-loss term. The merger-gain term, which scales like the densities of the 2 things that are coming together, is the product of those densities with some merger kernel K, and it increases the population of cluster k, which is of length i + j. Then you have to do the balance of that. Every time you have 2 things coming together to make a new one, you have to subtract their populations, i and j. That’s what the right-hand side is about: it’s the loss of things that have merged. You put those 2 things together, and you get a stochastic differential equation for mergers in a solution.

There is a phase transition associated with Smoluchowski coagulation. It’s called gelation, and it’s exactly what happens when you put Jell-O in the fridge and it sets. Basically, if things are sticking together and they stick together with a scaling exponent that is greater than 1, then you get this finite-time singularity in which the things that stick together diverge to infinite size, and the whole thing sets, no matter how big it is. That’s how Jell-O sets.

Tim Scarfe

Could that be gelation?

David Krakauer

Yes. The short answer is: that is gelation. The phase transition that we see in the emergence of life is a gelation phase transition, according to a generalization of Smoluchowski coagulation to this case of BFF strings coming together.

If you think about quote-unquote inanimate and viral replicators as being replicators that are not self-contained—in other words, where the code that runs is not fully within the code that is actually getting copied—you notice something interesting. What I’m calling here an inanimate replicator, and very much in scare quotes, is code that copies something fully outside itself. In other words, the code that runs in order to do the copying is disjoint from the thing that gets copied.

Are there such replicators in the real world? Of course, that’s what water is, right? Water is a replicator of some kind. It gets made by stuff, but the stuff that it gets made from, like water, is not a part of the running process. I mean, it is a part of the running process that makes more water in some cases, but it’s not part of the code, let’s say.

Viral is the case in which the code and the thing that is copied overlap. In other words, some of the code that does the copying is actually some of the stuff that gets copied, but the code is not fully contained by what gets copied. So this is an incomplete replicator that would need to cooperate with another replicator in order to reproduce. That’s what I mean by viral.

In the beginning of BFF, all of the replicators are inanimate or viral. The great majority are inanimate, and a few of them are viral. A few of them happen to copy 1 of those bytes that is actually an instruction doing the copying. But as you move toward the time of gelation, which I’ve normalized to 1 here, you can see that cellular replicators suddenly emerge. They can’t emerge before about halfway through the run, and then they shoot upward at the end.

That’s really interesting because it tells you that the moment of a cellular replicator—where the machinery for copying itself is part of the thing that is copied—emerges through the symbiosis, or the symbiogenesis, of inanimate and viral replicators.

So a full equation would have 2 terms. It would have this reproduction and a Lotka–Volterra-type term, and it would have a merger, or Smoluchowski-type, term. The one on the left is normal population dynamics. That’s normal Darwinism. On the right, you could think about the left as evolution and the right as revolution, right? Those are the moments when things come together.

The population-dynamics part for BFF looks like this. It’s a little bit more complicated, but it has the same basic form as Lotka–Volterra. There’s a linear part on the left; I’m just writing that as a matrix Rᵢ operating on the whole thing. On the right, the reason that looks a little different from Lotka–Volterra is that when something gets copied, it overwrites other stuff.

So now we have to say: How does that suppress the populations of everything else in the soup? In order to figure that out, you have to look at niches. What are the bytes where something gets copied? The overlap between the niches of 2 replicators tells you how much one thing getting copied is likely to overwrite something else that shares its niche.

The symbiogenesis part is a bit of a mess, so I’m not going to go through it. I hope that’s okay. But it looks just like Smoluchowski, just gnarlier. The reason that it’s gnarlier is because Smoluchowski has only binary fusion between 2 parts. In BFF, sometimes a bunch of things come together, so you have to take into account these kernels that have more than 2 parameters in them.

Also, when things come together, they don’t necessarily look like the sum of the things that came together. You could have something that is 3 bytes long or something 5 bytes long come together, and the result that copies itself is only 2 bytes—1 byte from each one—or anything along those lines. To account for those complexities, you end up with a much more complicated K term, but it’s essentially the same as Smoluchowski coagulation.

To prove that this kind of symbiogenesis is needed in order to get these complex programs, you can do a very simple intervention. When you’re interacting 2 tapes, you can do it in a sandbox before committing. In the sandbox, you see whether a new replicator arises and, if so, what replicators it is made out of.

In other words, when you look at the source, you can see whether any of those source bytes were actually the outputs of copies of some previous replicator. If so, then you have a tree. You have an ancestry tree for that replicator. That means that you can think about the depth of such a tree: how many things have come together.

You can limit the depth of that tree. You can say, if the tree depth exceeds 10 for a new replicator, then I’m going to not do this interaction. I’m going to take them back apart, pretend it never happened, put them back in the soup, and try again. If you limit the depth of the tree to, say, 24, then the number of operations that you have to block—the number of interactions you have to block—is actually very small.

You only have to block 1 in 1,000 operations. But that 1 in 1,000 operations is really important. As it turns out, if you block those, no gelation will happen. You need at least a tree depth of 20 or so in order to get these complex programs.

So this is a very nice proof that symbiogenesis is what's needed to get to these complex types. When you do that blocking, you end up with sort of logistic curves for the populations of all the replicators in that soup. They go up and then saturate and stabilize. That's fun because it lets you do a little bit of math.

As you can see, not only do things go up and saturate, but then there are some random oscillations, and those oscillations can be correlated. Sometimes you can see 2 of those populations go up and down together. So that means that they're maybe collaborating with each other, and sometimes they go in opposite directions. They're anticorrelated, and that means that they're competing with each other because one is overwriting the other, for instance. So that's what one would expect from off-diagonal production and competition from those equations I wrote earlier.

And if you linearize the dynamics around that steady state, then you can sample the correlations in those population fluctuations and reconstruct the matrix R. I will skip the details of how one does this, but this is a classic fluctuation analysis. You solve the Lyapunov equation and you get a Jacobian, and from that you get the matrix R. The matrices R look really cool.

First of all, they have a strong diagonal that tells you that, by and large, things replicate themselves, just as you would expect from Lotka–Volterra. But there's some other stuff going on here as well. Aside from that dominant diagonal of self-replication, there is some negative stuff off the diagonal and some positive stuff off the diagonal. The negative stuff off the diagonal looks largely symmetric about the diagonal, and that's as you would expect, too. Basically, if A competes with B, then B competes with A. Two things that are fighting for the same niche are in a kind of zero-sum relationship with each other.

But the cooperation part, where something helps something else, is not symmetric, and that's as you would expect, too. Just because A helps B or enables B doesn't mean that B enables A, or at least not directly. So there are complex cycles in this graph on the right of codependency or enablement. The negative component is symmetric, the positive component is asymmetric, and there's this big diagonal.

Tim Scarfe

Do the submatrices that are about to undergo symbiogenesis have any special properties?

David Krakauer

They do. In other words, if it's these, let's say, 4 rows and columns that are about to undergo symbiogenesis, you can ask what the eigenvalues of that submatrix are, and it turns out that they are generally cooperative. Essentially, if you were to pick random rows and columns from this matrix, then you get a high-dimensional picture of the rank of the matrix. But when you look at the ones that actually combine, it's much lower rank. They're already working together.

In other words, there's a relationship between the R and K parts of this equation. Symbiogenesis happens among guys who are already working together. They're not all the same, not independent—cooperative.

Here's another really interesting thing. If you look not at the R matrix but at the Jacobian itself, then you can find the signs of imminent instability in it, of when it's about to pop, run away, and gel. You don't say “gelate”; you say “gel.” In particular, if you block the depth of the possible trees to a low number, then the eigenvalues of the Jacobian are always negative, meaning that the system is stable.

But as you look at larger depth ceilings, you find that more and more of these leading eigenvalues, or the real parts of those leading eigenvalues, pop positive, and that means that the system is about to blow. You can keep it from blowing for a while by keeping that merger clamp on, but it tells you that, essentially, the more you evolve these things, the more they begin to cooperate with each other and the more incipient symbiogenesis is about to happen. That's what leads to this phase transition.

I just want to put in a little plug for what I think could be a really beautiful missing link between the kind of algorithmic information theory that Hector Zenil was talking about and the assembly theory that he has somewhat slammed with a couple of papers that he has written. As those of you who have followed that might know or realize from what I've just talked about, there's a very close relationship between what I've just been describing and assembly theory. It's things coming together to make bigger things.

But the assembly theory proponents have not really talked about the computational nature of what they're doing. In this, I fully agree with where Hector is coming from. I think the way those connect is by starting to look at things like the conditional Kolmogorov complexity of the things that are coming together. I think this is a connection point for us to maybe reconcile those 2 different pictures.

So symbiogenesis is what gives you complexification. That, in turn, is what gives evolution its arrow of time. In classical evolution and Darwinian evolution, there's no reason that things should become more complex over time. They might simplify, they might get more complex—it doesn't matter.

But with symbiogenesis, we know that things get more complex because if A can replicate itself and survive into the future, and B can replicate itself and survive into the future, when they come together, you suddenly need A to replicate itself and B to replicate itself, and there's some additional information that has been added, which is how the 2 fit together. Those extra bits of information that keep getting added to the program of the large replicator don't come from mutation. They come from the fact that things encounter each other randomly in order to possibly undergo that symbiogenetic event.

So it's actually the thermal randomness of the fact that we pluck 2 of these guys out of the soup at random. That's the information source, if you like, or the noise source, that is selectively turned into algorithmic information by the symbiogenetic process.

Eörs Szathmáry and John Maynard Smith have written extensively about these major evolutionary transitions, in which symbiogenesis results in large, novel forms of life like eukaryotes, multicellularity, and so on. I think this work is great, but the flaw is that they're only talking about 8 or 12 events. If what I'm saying is true, then this is just the tip of a gigantic iceberg. Basically, it's symbiogenesis all the way down.

Most of these symbiogenetic events are much more uneven. There may be just a little bit of something getting incorporated into something much bigger, but that is the source of novelty in all of evolution. These are just the most dramatic cases that involve really big, visible stuff happening.

So, is there evidence for these smaller symbiogenetic events in biology? Lots. There's lots of evidence for it. I don't have time to go into it in any detail, but if you look at just the human genome, you find that only 1.5% of it codes for our proteins, and lots of the rest of it is transposons and other endogenous retroviral elements of various kinds that involve viruses whose ecology is our own genomes, that reproduce inside our genomes, and sometimes jump species.

This results in something like a quarter of the cow genome being a retrotransposon that also lives in lizards and salamanders and stuff. When you start to look at that, you realize that genomes are fractal. They're replicators made of replicators made of replicators, just as I've described—not this kind of fixed design space with evolution only happening in its usual way.

It's not just horizontal gene transfer in bacteria. This symbiogenetic picture, I think, is the engine that produces novelty throughout all of life, including big, complex animals like us. There's more and more evidence in the last decade of things like this going on.

For instance, the Arc virus was endogenized in the mammal lineage, and you can find it in our brains. It turns out that if you knock out the Arc virus in mice, they stop being able to form new memories. So clearly the Arc virus is doing something important for us, and that's a source of novelty from an endogenized virus.

Similarly, the mammalian placenta was formed by an endogenized virus that fuses cell membranes together, and so on. There's a definition of life that comes out of this. I said this in the panel yesterday: life is an embodied autopoietic computation arising and complexifying through symbiogenesis.

It's not just neuroscience that's computational. Life was computational from the beginning, and it gets more computationally complex over time through symbiogenesis at many scales. Remember, if life is a computer from the start, then every time things fuse together, you're making a more and more parallel computer.

Those computers have to be not only running the code that models themselves and reproduces themselves, but also doing something about modeling the other and figuring out how they interact or work with the other. This means that an ecology of functions is building up through massively parallel computation that becomes, if you like, more and more intelligent with every one of these fusions.

Since symbiogenesis makes the computation massively parallel, that implies that intelligence and life are very closely connected. That's why I ended up with the book *What Is Life?* as part of the book *What Is Intelligence?* When you're not only using that intelligence to model yourself but also to model your environment—which, by the way, includes others, most importantly—then that's intelligence.

That means that life was intelligent from the start. The moment that modeling of others begins, what we call, in larger, more complex animals, theory of mind becomes fundamental to the way intelligence develops. These are really simple simulations that show how persistence allows the modeling of an environment to turn into learning: chemotaxis in these fake bacteria.

Of course, in real life, you're not only learning about an environment that exists in isolation, like the sugar crystal, but actually about all of your friends. The moment you're reproducing, the greater part of your environment is actually all of the other things that even your own reproduction is creating. Life is never single-player.

Things like intelligence explosions in our lineage, in the hominins, and in cetaceans and bats, and a variety of other species, are exactly this kind of runaway modeling of others, resulting in growth of brains and growth of groups. Therefore, when we think about the growth of advanced intelligence in human societies or human brains, it's really that same sort of symbiogenetic process happening at a much higher level.

David Krakauer

Let's end there and switch to questions.

Tim Scarfe

I think there are multiple different ways to represent symbiosis. In real biology, we maintain those hierarchical structures, and there are fundamental mathematical differences in how you treat those symbioses. Do you have any insight into how we can implement that?

David Krakauer

Yes. In biology, we often reify one particular level of detail and say, “These are the life forms.” Maybe there's a symbiosis between, let's say, algae and a sea slug, but we still think of the algae and the sea slug as separate, and we think about the population dynamics within that rather than modeling them separately. Is there a reason to prefer one scale or another?

For me, one of the lessons—the reason that I spent so much time on the relationship between R and K—is that you can always move something from R to K and back. Lin Margulis famously said, “We're just colonies of bacteria, some of which live inside each other.” That's true; you could describe us as just colonies of bacteria. But the reason that it's useful to move up a level of detail is because humans also reproduce as a unit—hence the mess—and there are a lot of things that you can learn about when you study at that higher level, a lot of abstractions you can make from a computational perspective that are hard if you're only modeling at the lower level.

So I don't think that there's any one layer or level that is true. We have lots of boundary cases, like lichen or colonial insects, where you can model the entire colony or you can model the individuals. I don't think there's a right answer to those two. Do you keep a block of rows and columns in R that always, or mostly, end up getting copied together, or do you add a new row? It's actually a coarse-graining choice, and you can make either one.

Tim Scarfe

Symbiosis is not symbiogenesis. What is the thing that you would claim is a good insight for how A and B stop being A and B in symbiosis and become something else?

David Krakauer

The fact that phase transitions don't come from R alone; you have to look at K. You do get these runaway modes, which tell you that something is about to happen. But in order to understand that phase transition—in other words, to see that a major evolutionary transition has occurred, to put it in biological terms—you actually have to understand the physics of K. It's only by understanding the physics of K that you can do the theory that lets you predict and understand what is going on right here.

If you model a body as just a bunch of bacteria, then it's not wrong, but it's invisible to you that something amazing happened when we became multicellular or when eukaryotes formed out of bacteria. This allows higher-order modeling. In particular, by the way, if we just take the subjective perspective for a moment, if you are one of those bodies, if you are one of those people, then you're not going to survive very well if you're only modeling other people as collections of bacteria. You have to build higher-order models of them, because that becomes an essential part of your environment, and you have to simplify or coarse-grain the world in order to build a model that is ecologically relevant to you. I'm kind of mixing here subjective and objective perspectives.

The subjective perspective is ultimately super important. It's a little bit similar to why we need temperature and pressure in physics. Those don't exist if we just look at the microscopics, but it's only by coarse-graining and looking at the larger scale that you can understand thermodynamics. In the same way, it's only by zooming out and looking at the symbiogenesis that you can understand the dynamics of transitions, phase transitions, major evolutionary transitions, and smaller and coarser-grained models, where higher orders of things emerge.

Tim Scarfe

Well, actually, that's exactly what I was doing 30 years ago, right?

David Krakauer

Exactly. That's why I began with your paper from nearly 30 years ago.

Tim Scarfe

There was a big problem. Symbiosis is possible—like, we bind the two, then it becomes complicated, coupling with each other. However, the function itself is not becoming complex or higher-order.

David Krakauer

I have 3 answers, I suppose, depending on which way we talk about it. The first answer actually comes from software engineering—or from mathematics, for that matter. Composition of functions is symbiogenesis, as I've described it, and when you compose 2 functions to make a higher-order function, you are making something more complex than the primitives. What I hinted at with Eric Smith's notion of conditional complexity can quantify that sort of compositional complexity. You could find signatures of it in DNA, or, if you don't want to look in DNA, you could look in GitHub at the way every time somebody writes some code, it begins by importing a bunch of other things and combining them. You do see a tendency toward complexity.

Now, that is constrained by energy. The more complex a thing you make, the more free energy it has to use. But now you get some of Chris Kempes's beautiful work, in which you see that there are energetic benefits to teamwork. The scaling laws for this are also environmentally dependent. I don't know, Chris, if you got to the Snowball Earth-type stuff, but there were certain very specific conditions at certain points in Earth's history that became favorable for eukaryogenesis, if I'm remembering correctly, and so there are some external conditions as well.

The other cool thing is that when you start to have more complexity in computers, when you start to have massively parallel computation, that greater intelligence also unlocks new energy sources. That gives you a bigger budget to play with, which in turn allows the next major evolutionary transition to take place. So I think there's an energetic perspective, a compositional perspective, a combinatorial perspective, and a scaling-law perspective that can all come to the rescue of that question. But we should talk more about it. I'd love to get into this in more detail.

Tim Scarfe

It's essential for life to have some functionality pointing toward the brain. Programs are strings, and their functionality is determined by external goals, aka the CPU.

David Krakauer

I agree. I made claims that may sound contradictory. One was that von Neumann is embodied computation and is different from Turing in that sense, but, on the other hand, that BFF looks very much like a Turing machine. And yet I also said it was embodied because I made one tape. So even the question of whether something is embodied or not is a little bit perspective-dependent as well, because in a von Neumann system, for instance, there's of course the same rule operating at every pixel.

You can ask yourself the question: Is a computer the thing that I make with lots of parts, right, of the kinds that are designed, or is it just the operation of a single pixel? The usual answer is to say what happens at a single pixel is just the physics of that world. But what constitutes the physics and what constitutes the computation is actually a movable boundary.

Embodiment is essential in a very minimal sense: You need to be able to operate on the thing that is going to—you need to be able to make quines, essentially. In an ordinary Turing machine, you can't make a quine because the data tape is separate from the program tape. But when you bring them together, you're now in the same realm as a cellular automaton, albeit with a different coarse-graining of what you consider to be the physics and what you consider to be the code.

In our world, we know that it's possible to build computers, or else we wouldn't be able to build computers and we wouldn't be here either. But what constitutes the physics, if you like—the physics that makes up computers itself—had to evolve. We began, I guess, as nothing but quantum field theory, and then things came together into particles, the particles came together into atoms, the atoms came together into molecules, and so on.

Those are essentially what I would call the inanimate replicators in the system. And there’s an important phase transition when those suddenly form a rich enough set that not only do you have an autocatalytic system, as Walter Fontana would have said, but also that you can form a Turing-complete instruction set and therefore open the door to generality of computing. I hope that makes some sense.

Tim Scarfe

All right. I think we’re—

David Krakauer

And, yeah, I’m afraid we have to wrap up. So, let’s thank [him] once again.

Tim Scarfe

Thank you so much for this amazing talk. Great questions, too.

如果智能并未进化?它从一开始就“在那里”!- Blaise Agüera y Arcas — 文字稿与摘要 | BidClub