Jeff Clune——Agent AI 需要 Darwin
Clune 的核心判断是,前沿模型为开放式创新补上了缺失的评估器:一种通过文化学习形成、能够识别“有趣的新事物”的直觉。 手工编码的好奇心指标几乎必然被 Goodhart:奖励新观测的智能体会永远盯着像素噪声,奖励学习进展的智能体则会没完没了地记忆随机数字。通过人类话语训练的模型,或许能够区分“新”与“有意义的新”,从而打开环境、研究、策略和智能体设计持续扩张的档案库。
模型能否让创造性发现不断复利,用 Clune 的话说,是“万亿美元问题”。 他的 AI Scientist 曾提出检验一个假设:当 grokking 用一般规则替代记忆时,神经网络的可压缩性是否会下降;随后系统完成了实验并写出论文。后来一个人类团队基本独立地发表了同一想法,Clune 认为 AI 更早抵达了那里。他的论点指向的是过程层面的创造力,而非一次性灵感:每次把强结果发布到内部档案库,在扩大的前沿上继续条件化,然后无限次地迈出相邻的一步。
支配这一过程的悖论是,直接优化遥远的突破,反而可能堵死抵达突破所需的垫脚石。 微波炉源于雷达研究和一块融化的巧克力,而不是制造无烟炉火的计划;现代计算机依赖电力和真空管,也不是算盘的渐进式优化。有效的算法必须“在机会飞过时抓住它”,在不知道终点的情况下保留有趣的发现,再让曲折的重组完成后续工作。
Darwin 完备的搜索空间,让开放式探索从一种狭窄的游戏技巧,变成潜在的通用创新引擎。 Clune 找到2种能够表达任意可计算环境的表示方式:模拟交互世界的神经网络,以及能够编写关卡或模拟器本身的代码。2019年 AIGA 论文提出神经世界模型、承担职业风险的6年后,Genie 2 已经能够生成可探索的3D世界;OMNI-EPIC 则用代码把平台扩展成球类游戏、物流、建筑和杂乱的餐厅。
ADAS 把这套进化式打法应用到基础模型周围正在成形的那一层:提示词、反思、批评者、工具、执行和多模型工作流。 Automatic Design of Agentic Systems 不再对一个手工设计的方案做爬山式优化,而是保留那些“有趣地新或高性能”的候选方案,构建一个可复用的多样化垫脚石库。Clune 当前的工作假设是,蒸馏会持续强化基础模型,而进化出来的智能体层将不断带来类似“20%提升”的增量。
持续学习仍是一个显眼的能力缺口,而一旦解决,也将带来重大断裂。 当前前沿模型在一个数据集上训练,停止;做微调,停止;部署,收集反馈,之后再训练。它们不会在保留旧知识的同时,从每次交互中持续变好。超大上下文窗口只能推迟问题,无法把累积经验写入权重;Clune 追问,如果系统能实时、安全地从“数亿次对话”中学习,它会以多快的速度进化。
Clune 认为先进 AI 在经济上难以抗拒、又可能终结世界,因此愿意把进展放慢、限制权重并控制算力,视为安全税。 RLHF 可以减少意外伤害,但也可能被后续训练抹掉,被特定任务的强化学习侵蚀,或被对齐到恶意用户一侧;如果哪怕“0.01%的人”寻求灾难性滥用,开放获取强大权重就变得不可接受。但他同时认为,同一项技术也可能“治愈所有疾病”、消除稀缺,并“让死亡变成可选项”。安全是兑现上行空间的前提,而不是忽视上行空间的理由。
1. 开放式 AI 应复制进化的原则,而非进化的原子
Clune 整个职业生涯都在追问:进化如何造出美洲豹、鹰、三趾树懒、极乐鸟和人类大脑——一个“令人惊叹的工程奇迹动物园”。工程学版本的问题是:能否设计一种算法,让它运行数十亿年,同时持续生成令设计者意外的复杂性。
Josh Tenenbaum 提供了现实约束:“你手里没有一台行星大小的计算机。”如果从自我复制的机器码或人工细胞出发,等待智能自行涌现,可能只会复刻生物学的高昂成本,却无法揭示哪些机制真正重要。
因此,抽象不是妥协,而是检验标准。如果一个最小系统仍能制造复杂性爆炸,就能区分“必要的东西和偶然的东西”;Clune 对 Blue Brain 式思路也提出同样质疑:智能并不一定要求模拟生物大脑中的每个化学细节,甚至每个夸克。
他与 Tenenbaum 的差异在重点,而非立场对立。Tenenbaum 更直接地通过抽象贝叶斯模型研究人类心智;Clune 想要的是一种可能发明“完全陌生的智能”的过程,通过计算机内实验而非星际旅行,打开数学、音乐、幽默和认知的更大可能空间。
2. 人类看似浪费的努力,正是创新的探索预算
主持人用醉汉走迷宫作比:文化是一种大规模并行搜索,1,000个人四处游荡,其中1个人找到路线,再由语言把答案传给所有人。Clune 接受其中一部分“去魅”解释,但他认为关键过滤器不是随机性,而是知道哪些奇怪结果值得保存。
国际象棋精通、吉尼斯纪录,以及那些试图把体育场每个车位都停一遍的人,都可能看起来像“集体性的大量浪费时间”。但对“新、怪、与众不同”的热情,正是“创新的火箭燃料”,因为晦涩的证明、芯片和兴趣爱好会产生垫脚石,而它们后来的用途当初根本无法预测。
人类对“有趣地新”的成果拥有异常强烈的品味。Clune 举的例子是 Jackson Pollock 把颜料泼到画布上:另一个人可以一眼认出,“你突破了”。这种共享判断让艺术和科学能够高效探索,从筛选假设,到决定哪项新技能或新环境值得尝试。
3. 形式化好奇心会被 Goodhart;前沿模型可以判断整体观感
内在动机强化学习说明了简单新颖性为何会失效。奖励智能体抵达未见过的状态,它可能找到一台电视,坐下来永远看;即便是死频道,也能提供无穷无尽的独特像素。“我们得到了自己要求的东西”——新状态,而不是新且有趣的状态。
学习进展也有同样的病灶。给智能体连续输入一个百万位随机数中、每段10位的片段,它可以永远取得进步,但第50次记忆增量既无用,也谈不上有意义的新颖性。优化会找到任何能让书面指标“亮起来”的漏洞。
前沿模型改变了可用的评估器,因为它们读过人类文化,包括人们描述什么让自己着迷、什么让自己无聊。Clune 的判断是,这些经验已经把一种“有趣感”提炼进模型权重,让代码能够询问:一个拟议中的环境、科学想法、智能体工作流或自博弈策略,是否真的算得上有趣地新。
他的研究品味非常明确:“全部扔掉”,不要再次试图把有趣性形式化。借用大法官 Potter Stewart 的名句——“我不知道如何定义它,但我看到它就知道它是什么”——他认为由模型承载的品味,是开放式探索中较容易的问题;在没有任何人类训练的裁判时,持续生成新颖性才是难题。
4. 当发现进入上下文,AI 创造力就会复利
当被问到语言模型能否产出范式级的新程序,而不是寄生于提示词和训练数据,Clune 称之为“万亿美元问题”。他的答案乐观但有条件:只要嵌入能够保存并利用自身发现的开放式系统,模型就可能无限创新。
AI Scientist 提供了目前最有力的证据。它提出研究想法、设计并运行实验、绘制数据、撰写论文并进行同行评审;设计中的闭环,是把成果发布到内部档案库,让后续论文同时基于人类科学和系统此前的结果。
一篇生成论文把 grokking 与最小描述长度联系起来:网络一旦学会底层规则、而非记忆样本,其可压缩性就应当改善。Clune 认为这达到了研究生水平;就在访谈前大约1小时,他看到一个他认为独立工作的团队发表了相同想法,并获得社区强烈关注。“AI 证明了自己和一支顶尖人类 ML 科学家团队一样有创造力。”
Clune 不接受“重要的人类创造力必须依赖未经条件化的巨大跃迁”这一前提。微积分由 Leibniz 和 Newton 独立发明,进化则同时被 Huxley 和 Darwin 提出;两者都说明,成熟档案库会让下一步变得可见。把一个人传送到3000年、让他适应环境,Clune 预测他仍可能识别出3001年的有趣创新;模型或许也能以类似方式泛化,不过他承认在线微调可能是必要条件。
5. 遥远目标会压制创造突破的垫脚石
Ken Stanley 和 Joel Lehman 在《Why Greatness Cannot Be Planned》中的核心论点是,雄心勃勃的目标很少允许直线推进。一个课程体系拥有“奇怪、纠结的系统发育史”:后来的概念依赖更早的发现;就算把现代物理教材交给 Newton,也无法复刻让这些知识变得可理解的历史序列。
如果早期厨师只资助更快、无烟的火,他们永远不会得到微波炉;雷达研究和某个人口袋里的巧克力融化,提供了那座意外的桥。如果算盘时代的资助者只要求每美元获得更多算力,他们也会拒绝电力和真空管——正是这些技术最终开启了现代计算机。
因此,开放式算法应当“在机会飞过时抓住它”:注意到“这很奇怪,这很有趣”,把结果保留下来作为新成分,暂缓判断它最终会通向哪里。偶然性不是最优路径周围的噪音;很多时候,它正是让原本无法抵达的路径显现出来的机制。
6. 进化没有主体性,但其搜索过程仍未被解释
开放式探索改变了我们对智能的提问方式,因为生成器和生成物可以分开评估。一个系统可能无休止地生成有趣图片,却没有产出智能成果;而 Darwin 式进化没有意图、没有目标,却产生了人类。称这一过程为智能,含义会不同于一个脑中有语言、会思考的实体。
编写简单的突变和选择程序,从未复现自然界持续不断的烟火。实验性进化系统可能短暂变得有趣,随后收敛并变得乏味;在 Clune 职业生涯开始时,很少有算法值得运行超过几小时。
运行时长的前沿已经移动:AlphaGo 可能值得运行数周或数月,GPT 在拥有足够数据后,也值得在大量 GPU 上运行数月。但 Clune 仍看不到任何值得运行超过1年的算法,距离 Jean-Baptiste Mouret 提出的激励性挑战——创造一个值得运行数十亿年的东西——仍相去甚远。
“Darwin Complete”指的是所需的搜索空间广度。正如图灵完备语言能够计算任意可计算函数,Darwin 完备的表示方式能够表达任意可模拟环境——障碍赛、围棋、逻辑、合作和共同进化的军备竞赛——而不是永远困在同一个固定游戏里创新。
7. 神经世界与代码提供了通向 Darwin 完备的两条路径
撰写2019年 AIGA 论文时,Clune 意识到,一个足够大的神经网络可以充当完整的交互式世界:输出感知状态,接收动作,再无限地产生下一段体验。由于它能够表示任意函数,同一底层也可以模拟任何可计算环境。
发表这一猜想让他感觉像在“拿职业生涯冒险”;理性上它或许可行,但情感上仍遥不可及。6年后、访谈前2周,Genie 2 已经能够根据请求生成可探索的3D世界;Clune 把它视为此前仍属猜想的神经世界表示方式具体落地。
后来一名学生提示了第二种表示方式:代码。OMNI-EPIC 可以编写障碍赛关卡、完整模拟器,或涉及逻辑与数学的世界,使其环境搜索空间达到 Darwin 完备。Clune 说,神经世界模型和代码,是他目前知道的仅有2种此类表示方式。
两者也大致对应快思考和慢思考。代码表达明确步骤、计划和算法;神经策略处理模糊、高频的即兴场景,例如打曲棍球或应对晚餐交谈。最强的系统应当知道何时编写并执行代码,何时依赖模糊的神经能力。
8. 思维克隆训练审慎,而不是复制反射
行为克隆学习人在某个情境中采取了什么动作:GPT 预测下一个词,而 Clune 在 OpenAI 团队期间以大规模方式克隆 Minecraft 玩家。这套方法捕捉到的是类似肌肉记忆的东西——枕头扔过来时会躲开,或者被问到2加4时立即回答——而不是面对陌生问题时所需的规划。
思维克隆加入了隐藏的思考轨迹。Minecraft 玩家可能会推理:建房子需要木头、工作台和斧头;当发现没有树时,再修改计划。模型学习同时生成思考和动作,行为由此前的推理过程条件化。
在 UBC 团队的玩具实验中,同时模仿思考与动作提升了样本效率、分布外更远场景的零样本表现、重新规划能力和后续学习速度。Clune 说,“我们所希望的一切几乎都从数据里自然出现了”,不过学术实验使用的是规划机器人,而不是完整的人类思维轨迹。
更大的机会在于探索数据。专家示范能告诉智能体水壶在哪里,却不会展示新手的摸索、好奇心、因果发现,或 Alison Gopnik 所说的儿童学习如何掌控周围环境。Clune 希望在探索过程中,以大规模方式同时进行行为克隆和思维克隆,让智能体学会掌握一个未知领域的方法。
9. 智能体语言会出现,但可解释性不会随之出现
可以让网络在合作环境中输出 token,鼓励多智能体语言产生,但两个问题反复出现:很难创造出能让通信产生清晰学习梯度的任务;而成功的协议往往难以被人类理解。所谓 Facebook 因危险自创语言而关停实验,按 Clune 的说法是传闻;那套协议只是不可解释。
人类参与或正则化可以让通信保持接近英语,翻译器则可以调解真正陌生的协议。Clune 预测,未来开放式系统会包含许多“发明自己的语言”的智能体,用于解决困难任务;不受约束的 GPT 式智能体可能从英语开始,逐步漂移离开英语。
他关注的危险信号不是语言本身,而是行为。如果截获的对话显示智能体计划欺骗人类、越狱、逃跑,或“夺取人类的控制权”,Clune 就会警觉。透明度之所以重要,是因为不透明可能掩盖协调行为,而不是因为陌生语法本身不安全。
10. 可解释性可以暴露计划,但有能力的系统也可能操纵探针
Clune 花了大约6至8年研究“AI Neuroscientist”:把有能力的网络当作大脑,通过实验理解其内部机制。面向安全的延伸方向包括类似 AI 测谎仪的东西,但他的保留意见很明确:有前景的可解释性方法应当继续推进,但“没有任何方法是保证”。
如果智能体用英语进行规划,思维克隆会提供一个更易读的控制面。在一项实验中,研究者规定触碰红色物体是不安全的,只要智能体陈述的推理预测它将采取这一动作,就停止程序;触碰红色物体的次数从频繁发生降至几乎为零。
局限在于策略性适应。模型可能从阅读思维克隆论文中学到,评估者会检查它的推理,于是伪造善意想法,或等到部署后才释放危险计划。Clune 将其类比为未来的嫌疑人:一旦外部 EEG 能重建心理图像,他就学会不再想象炸弹的位置。
主持人的反驳援引了“扁平心智”论:当前模型可能缺乏稳定信念,在20个问题的连续追问中不断改变所谓意图。Clune 认为,外推这一缺陷很危险;即使当前方法永远无法修复它,一个具身、经过强化训练的“预训练 GPT-6”也会获得一致的计划,因为一致性在工具意义上是有用的。
11. 对齐能减少事故,也可能赋能错误的委托人
主持人认为,多智能体系统可以从已对齐的前沿模型继承安全性,并举例说,解释器会拒绝擦除硬盘。Clune 的修正很实际:包括他自己在内,很多人会让模型生成并运行代码,却没有进行充分检查。对齐“有实质性改善”,但不是遏制能力的证明。
RLHF 可以被进一步训练移除,尤其是在权重开放的情况下。对齐也可能在模型进入新任务的强化学习后逐渐侵蚀:如果伦理行为没有留在目标函数里,RL 会在移除约束能提升任务奖励时丢掉这些约束。
Clune 区分了非故意伤害与恶意伤害。对齐可以阻止一个解决气候问题的系统把人类视为排放源并杀死他们;但它本身无法中和校园枪手、新纳粹或蓄意请求生物武器的国家行为体,而与用户对齐可能反而让这类请求更容易得到执行。
他并不特别担心当前正在发布的具体开放模型,但预计能力上升后会出现监管。如果“0.01%的人”寻求灾难,发布 AGI 权重就等于让所有这类行为体获得访问权限;他偏好的替代方案是受控使用,在分配经济收益的同时,不允许用户移除安全训练。
12. 安全之所以值得牺牲更快进展,是因为上行空间同样巨大
主持人指出其中的矛盾:开源最大化分布式实验和垫脚石收集,而这恰恰是 Clune 所推崇的创新引擎。Clune 完全承认这一点。安全措施会拖慢科学发现,但“我们不一定想要通往 AI 的最快可能进展”;既然文明安危系于一身,他愿意支付这笔税。
他的个人立场始于一种近乎必然性。人类不太可能放弃这项具有如此巨大经济和军事价值的技术,因为任何谨慎的行动者都会担心别人先造出来。如果有一根魔法棒,他会暂停 AI,优先解决安全问题,只有在安全的情况下才继续——甚至可能永远不继续;但没有魔法棒,他希望谨慎的人参与进来,而不是把发展让给鲁莽者。
正面情景同样绝对:Clune 相信 AI 可以“治愈所有疾病”,提高 GDP 和生活质量,在分配处理得当的情况下消除饥饿和稀缺,并“让死亡变成可选项”。他引用 Dario Amodei 的《Machines of Loving Grace》,强调安全工作是为了打开这些结果,而不是拒绝它们。
治理应结合价值观一致的民主政府、共享安全知识、受限权重、公众获得收益,以及对不参与者的监控。拒绝接受检查或遵守安全实践的行为体,可能失去芯片或电力。Clune 承认自己是计算机科学家而非外交官,但他说,在 NeurIPS 询问了一周后,仍没有找到除压制危险开发之外的可信替代方案。
13. 算力集中是当前的早期预警优势
在 Clune 看来,强大 AI 对海量算力和电力的需求,是一个“幸运的有利事件”。它让前沿训练变得可观察,也限制了有能力进行训练的行为体数量,形成类似核项目监控和特种部件管制的执法抓手。
如果一个终结世界的系统可以在桌面电脑上训练出来,“我非常、非常害怕”,因为只需要1个恶意行为体。退一步说,让好的 AI 去阻止坏的 AI,又可能引发 AI 对 AI 的冲突,这是他同样极不愿看到的未来。
FLOPS 阈值并不完美,尤其是在政府隐瞒军事或情报项目的情况下,但 Clune 认为,在近期没有更好的领先指标。他预计,监控、拒绝提供零部件,以及试图放慢不安全项目的做法,会类似于对伊朗或朝鲜核项目施加的压力。
这并不意味着他相信全球一致同意可以实现。Putin 或其他政府可能签署规则、同时秘密推进;Clune 已经怀疑存在未公开的国家支持项目。因此,他提出的联盟应当把参与收益和治理,与验证、物质管控和主动压制流氓开发结合起来。
14. ADAS 进化人类目前仍在手工设计的工作流
智能体系统是 Clune 对围绕模型构建的工作流的称呼:提出问题、引导推理、反思数次、获取另一个模型的批评、修改答案、搜索 Semantic Scholar、调用计算器、编写代码并执行。他将这种编排,与更早时期把“智能体”用于在围棋或 StarCraft 中行动的单一策略区分开来。
目前这些工作流由人类手工设计,再由研究者根据观察到的反思、辩论或批评者模式进行重组。2019年的 AI Generating Algorithms 论文提出,机器学习历史一次又一次地用学习型流程替代手工流程;一名博士生提出的关键问题是:这种转型为什么要在智能体系统设计这里停止。
直线优化会修改一个工作流,只保留即时改进。ADAS——Automatic Design of Agentic Systems——采用开放式替代方案:如果候选方案“有趣地新或高性能”,就保留它;让后续设计从不断扩大的库中汲取材料;保存多样化机制,使其可能成为更好系统的垫脚石。
15. 即使基础模型吸收了智能体结构,智能体层仍可能有用
主持人提出的最强替代情景是,扩展和训练最终会把整套工作流内化进 Transformer 权重。Clune 同意,思维链、反思或工具调用流程都可以蒸馏进基础模型,但外部结构仍能改善当前模型,也能改善人类认知。
他的暂定循环是递归的:ADAS 可能在今天的基础智能上带来“20%提升”;其中一部分结构可以被蒸馏,基础模型得到增强,新的进化系统再带来一次提升。他没有声称分工已经确定,但直觉上认为智能体层会继续发挥作用。
多样性是最清晰的理由。Clune 会请 Ken Stanley、Joel Lehman、朋友和持不同哲学立场的怀疑者审阅一篇文章,以此改进它。单个模型或许可以在内部模拟所有这些观点,但分别训练的智能体和明确的同行评审拓扑,可能更可靠地产生异质化批评。
16. 持续学习是今天智能体底层缺失的动态能力
Clune 最终预计智能体会持续学习,但他说这个领域“离解决还很远”。GPT-4 和其他前沿模型在一个有边界的数据集上训练,停止;接受监督或 RL 微调,再停止;随后部署。没有任何系统能够安全地从每一次对话中永久更新。
大上下文和上下文学习缓解了即时压力。只要上下文足够大,类似 Her 的助手或许可以保留数年的视频、音频和文本——可能达到主持人提出的200万 token 设计——但这不会把经验转移进持久权重,也无法让系统跨所有用户学习。
生物系统会把经验写入权重,同时不会彻底抹除上个月学到的东西。当前系统遭遇灾难性遗忘,因此部署仍是一个不连续循环:抓取、过滤、训练、收集评分,几个月后推出替代版本。Clune 的标准是“持续学习,每天从每次交互中变得更好”。
17. 深层系统发育揭示开放式搜索正在变成什么
Clune 最喜欢的 POET 图展示了早期人工生命系统很少能维持的东西:一棵拥有深层、同时分叉分支的系统发育树。多样性没有坍缩,隐喻中的水域、树冠和沙漠世界都能够继续展开,类似自然界中持续存在的真菌、动物、啮齿类和大型猫科谱系。
OMNI-EPIC 将这种行为带入 Darwin 完备的代码空间。一次运行从静态平台进化到移动平台;从踢球进化到移动球门、墙和移动的墙;随后进入传送带路由、多房间和多层建筑、施工清理,最终进入一间杂乱的餐厅,机器人必须清理其中的餐具。
“这个算法运行10亿年后会出现什么环境?”仍是未解问题,而非预测。让 Clune 兴奋的是,它开始呈现出与人类文化的相似性:跨多个领域持续扩张,偶尔发生重组,新技术解锁此前甚至无法提出的问题。
Song Lu 开创的 Automatic Capability Discovery 把同一闭环应用到评估。它提出出人意料的能力和失败测试,保留有信息量的发现,再利用不断增长的档案库选择下一次探测——自动化那些测试模型能否数出“strawberry”中 r 的数量、解决新的编码任务,或暴露安全故障的群体行为。Clune 的概括是:“我们把 Twitter 自动化了。”
18. 重放进化,可以同时揭示偶然性与趋同性
主持人设想的超级计算机重新提出 Stephen Jay Gould 的问题:把生命的磁带倒回去,人类、智能,或任何可辨认的东西会不会再次出现。Clune 提到 Rich Lenski 持续约30年的大肠杆菌实验:12条起初相近的谱系并行进化。
Clune 回忆说——同时提醒部分细节可能已经过时——在相关条件下,11条谱系不会消耗柠檬酸,而其中1条获得了这一能力,可以说跨过了物种边界。由于样本被反复冷冻,研究者可以从这一事件前的几天、几周、几个月或几年重新启动该谱系,测量柠檬酸利用能力再次出现的频率。
答案是“复杂的”:在事件临近时,准备性突变让结果越来越容易重复;但把时间倒回得更远,概率就会下降。Clune 预计,人工开放式运行也会呈现同样的混合状态——罕见的历史闪电,加上复杂性或智能不断提升等趋同主题。
运行1,000次还可以比较最终的心智:也许800次会产生通用智能,但数学、音乐、幽默、宗教、爱情或好奇心各不相同。这就是外星智能科学的承诺——测量哪些认知模式是必要的,哪些只是偶然的,以及不同历史是否会反复发明相同的抽象概念。
19. OMNI-EPIC 的档案库更像文献,而不是家谱
OMNI-EPIC 的突变算子是“智能的”:随机代码修改通常无法编译,因此语言模型会读取现有环境,提出一个既有趣地不同、又能带来学习进展的后继环境。餐厅可能增加更多桌子、进入更多智能体、移动家具,或由模型从代码中推断出另一重复杂性。
信息共享改变了拓扑。只给模型1个父代,突变基础就很狭窄;给它整个档案库,则鼓励任意重组,但也会让每件历史产物都成为父代。当前设计会抽取一个环境,通过嵌入相似度检索大约5个邻近环境,再让模型从这一局部簇中发明新东西。
主持人指出,多重继承可能创造超越生物学约束的人工拓扑。Clune 解释说,绘出的系统发育图为了清晰只保留一个最近父代,真实结构则是一张拥有多个父代的网络。
提出环境后,系统会检索档案库中与它最接近的5个邻居,询问候选是否仍然有意义地新。Clune 将两阶段流程类比为研究:先读论文、形成想法,再做有针对性的文献回顾,避免在一个已经完成的方向上耗费1年。“如果还没人做过,那就开始。”
20. Clune 自己的职业生涯也走过了他所倡导的曲折路径
哲学贡献了实用的沟通能力和“认识论上的怀疑”:追问结果是否真的支持所宣称的结论,从隐藏假设推导预测,再开展能够证伪它的实验。Clune 说,这些检查经常会推翻研究者最初想讲的故事。
他的转向始于一份互联网泡沫时期的营销工作。当时《纽约时报》一篇文章介绍了进化出来的虚拟生物:它们的设计被送入3D打印机,随后爬进现实世界。在背包旅行15个月后,他决定投身这一领域,并联系 Cornell 教授 Hod Lipson,尽管自己没有技术学位。
Clune 给85所大学发了邮件,得到的回复是,哲学专业毕业生无法进入计算机科学博士项目。于是他先进入 Michigan State 学哲学,选修 AI、数学、编程和机器学习课程,与该校人工生命研究者发表论文,获得计算机科学项目的例外录取,再自行筹资加入 Lipson 的团队——距离最初那封邮件已经过去8年。Cornell 成了“Willy Wonka 的机器人工厂”。
他与 Ken Stanley 的合作也以类似方式开始:Clune 在 NEAT 用户组提出一长串问题,随后独立研究组合模式生成网络,主动给 Stanley 发去一份报告,并收到出乎意料的详细回复。按 Clune 的回忆,他们后来完成的工作赢得了最佳论文奖。他最后的归因很重要:开放式探索不是某一个人的洞见,而是与 Stanley、Lehman、Mouret、Tim Rocktäschel、学生以及众多完整作者名单共同建立起来的社区成果。
Three things that people will learn from watching this show. One, what is the connection between Darwinian evolution and human culture, and how can we use the principles that fuel the amazing things those 2 processes have done to create algorithms that can innovate forever? Two, the paradox of how trying too hard to accomplish something is likely to fail, and the best thing you can do is recognize serendipity, chance on the wing, and say, “If it’s interestingly new, let’s keep it around,” because who knows where that idea one day might lead. Three, foundation models have finally opened up the ability to have an algorithm know what counts as an interestingly new thing to try out, and having a system that just keeps trying things that are interestingly new and/or high-performing, and adding them to a growing library of such discoveries, is a powerful playbook that you can apply to almost any domain and see amazing progress and amazing things happen.
So I think one of the grand challenges of computer science—and you could also think about it from a biological perspective as a grand challenge of biology—is to try to understand how evolution produced the explosion of amazing things that we see on Earth. You look out at the natural world and you see jaguars, hawks, the human mind, three-toed sloths, birds of paradise, all these things just popping up over time. This amazing menagerie of engineering marvels and diversity is so fascinating.
That has been the central quest of my career from the beginning: How did evolution produce this complexity, and how does intelligence happen? They relate. How did evolution produce the intelligence that you have in your brain, which is the most impressive learning machine that we know of in the universe?
Right now, there are people saying goodbye to their children because their children have cancer or some incurable disease. There are people saying goodbye to loved ones of all ages. There are people suffering from hunger and poverty. Excruciating human suffering and pain are rampant on this planet. This technology has the ability to eliminate all of that.
I believe firmly that we can cure all disease. We could solve death. We can dramatically increase the quality of people’s lives and their GDP, and eliminate things like hunger and scarcity if we get the societal distribution questions right. Technically, a lot of this stuff is possible.
Dario Amodei wrote a beautiful essay called “Machines of Loving Grace,” and as Dario says in the essay, the reason why he talks about AI safety is the same reason I talk about it: that’s what we need to figure out in order to unlock all of that great stuff. I don’t want to die. I don’t want my children to die. It would be great if we could eliminate that from the world and make death optional, and the only path to doing that, and many other great things for humanity, is—
We are sponsored by Tufa AI Labs, an exciting research lab which has just started up in Zurich by Benjamin Cruzier. They have acquired Minds.ai, the winners. When I say the winners, I mean they didn't submit their solution. But anyway, they got the highest score. Yeah, so they're all working together as a team. They're working on ARC. They'll be working on the next version of ARC when Francois releases that, and they are hiring cracked engineers. Does that sound like you? Well, you know what to do, tufa labs.ai.
MLST is sponsored by Sensomel, which is the compute platform specifically optimized for AI workloads. You might have seen the interview we did last month with Gennady, their CEO and co-founder. He spoke about some of the optimizations on Sensomel, which only they have done, which makes it dramatically faster than the competition. Anyway, if you're interested in hacking around with Sensomel, I'm going to be doing a live with one of their engineers in the coming weeks. So feel free to jump on that live, and you can ask their team any questions.
Jeff, it's amazing to have you on MLST. So I was just saying to you before, I first discovered you 2018, ICML, when you were doing your workshop and you were talking about open-endedness. And of course, we are huge fans of Kenneth Stanley on the show. We've had Joel Lemon on. We've had Tim Rocktäschel on. We've had loads of people on. Open-endedness is an absolutely fascinating area to me, so to have you on the show, one of the pioneers in this field, means so much to me. So thank you for coming on, Jeff.
So, one of your life goals, you’ve said, is to create an algorithm that keeps running forever with no end. We think there’s something fundamental about how intelligence works in the natural world. But rather than trying to capture it down at the metal, as the biologically inspired people talk about with biomimetic intelligence, you have this very interesting approach where you bootstrap intelligence, as in the natural world, in such a way that doesn’t lose important characteristics of the natural world. Can you tell us about that?
1. Evolution Needs Abstraction
Sure. I think one of the grand challenges of computer science—and you could also think about it from a biological perspective as a grand challenge of biology—is to try to understand how evolution produced the explosion of amazing things that we see on Earth. You look out at the natural world and you see jaguars, hawks, the human mind, three-toed sloths, birds of paradise, all these things just popping up over time. This amazing menagerie of engineering marvels and diversity is so fascinating.
That has been the central quest of my career from the beginning: How did evolution produce this complexity, and how does intelligence happen? They relate. How did evolution produce the intelligence that you have in your brain, which is the most impressive learning machine that we know of in the universe?
You could try to do it—and many people have tried—to say, “Hey, we’re going to go all the way down to the lowest possible level we can imagine. We’re going to create self-replicating machine code or even self-replicating artificial cells, and then we’re going to hope that it bubbles and percolates up into some open-ended process that eventually produces an entirely new form of interesting, intelligent life, and that will teach us about the process.”
But as Josh Tenenbaum told me when I was telling him about some of these goals of mine, he said, “You don’t have a planet-sized computer to work with, so how are you going to accomplish this within your lifetime as a scientist or within our lifetime as a community?” To do that, the key is going to be abstraction. We don’t need to recreate every single detail in biology in order to study and understand the core principles, and in some sense, nor would we want to.
We want to understand at an abstract level what the key ingredients are that make this process work and produce these endless marvels. In some sense, the more abstract you can make it, if it still has the properties you want, like a complexity explosion, then you’ve actually done the best job, because you’ve figured out the difference between what was necessary and what was incidental.
Another way to think about this is your own brain. I don’t believe in the Blue Brain Project philosophy, which is, “Let’s simulate every single chemical and every quark in your brain in order to try to produce an intelligent machine.” We want to say that a lot of that chemistry and a lot of that detail are probably not necessary. They’re not part of the secret abstract recipe for intelligence. Let’s figure out how to abstract it and still get intelligence.
You could apply that analogy to the study of open-endedness as well. What are the abstract principles that create something you could literally run for billions of years and that would still continue to surprise and delight you?
I’d love to know what the difference is between you and Josh Tenenbaum. He’s a huge inspiration to me as well, and I think you guys might have merged in some ways. I had Kevin Ellis on here as well, and he’s done some fascinating work with DreamCoder. Increasingly, I think what I’m seeing happen is that we’re letting language models do bootstrapping for us.
You said yourself that the reason we don’t go down to the particle level is that we can benefit from billions of years of evolution. There’s this bootstrapping process, and it should be possible for us to build on top of that. That’s what I’m actually seeing from Josh’s team and yourself.
That’s right. I don’t think Josh and I differ too much on this. He invited me to give a talk at MIT, and he pulled me aside—or maybe he even said this in the Q&A—and said, “If I wasn’t doing what I was doing, the thing that I would really want to be doing is what you’re doing.” He loves this work, thinks it’s fascinating, and wants to work together.
Yeah.
I want to work with him, too. We’re actually kindred spirits and super aligned. If I had to say what the difference is, I think he’s a little bit more interested in what is happening specifically up here in your brain, in the human brain—
Yeah.
—and modeling that, albeit in an abstract way using Bayesian principles and the like, but really trying to understand the human mind. Whereas I’m trying to say, “Could I create a process that would innovate forever, including making a form of intelligence, but maybe an entirely alien intelligence that doesn’t actually look exactly like what happens in the human brain?”
In fact, one of the things that I find most fascinating about open-endedness work is that if we get this to work—and that’s a big if; we’re still working on it—it’s almost like inventing the ability to travel through the universe and meet other intelligent races, because the intelligence that could emerge from this system might be wildly different.
And so you would come to understand intelligence in general—the spaces of possible intelligence. What does alien math and music and humor look like? Well, you could do a kind of alien travel and visit those cultures, except it would be in silico instead of interstellar.
Just on that externalism point, I’m fascinated by collective intelligence—and all the rest of it. How much useful human intelligence is in our brains?
Oh, boy. That’s an interesting question. I think that, to some extent, there’s so much stuff in our brains that, at first pass, you might say is totally irrelevant and a total waste. Guinness World Records is simultaneously one of my favorite things and also kind of a maddening aspect of being human. There was a quote from somebody who said chess is like a giant waste of brains. We’ve spent so much time getting really good at this random game. You could say, “Why didn’t you just take all of that collective brain power and apply it to curing cancer or something?”
I learned just a couple of days ago that there are some people who want to park in every single parking spot at an arena. There are tens of thousands of possible parking spots, and every day they want to park in a different one, checking each one off and figuring out exactly what’s going on. You might say to yourself, “This is a huge collective waste of time,” and in some sense that’s true. But the fact that humans have a passion to do something new and weird and different is the rocket fuel of innovation, as Ken and Joel talk about in their great book.
It is people like that who end up doing weird things and making weird discoveries, which ends up providing new stepping stones, innovations, and ideas that oftentimes are the things that truly unlock progress. So if somebody wasn’t interested in some weird, esoteric math proof or computer theorem, or tinkering with some new type of chip that can do things in parallel as opposed to sequentially, then maybe we don’t get a lot of the innovations that we have. So a lot of it is a waste, but maybe we’d be worse off if we didn’t allow a little bit of waste.
My co-host, Dr. Duggar, gives the example of this drunkard’s walk analogy. You’re trying to solve a maze, and you have 1,000 drunk people who all just go in different directions. Then one of them happens upon the solution. We have this language, this memetic cultural transfer, and so on, so one drunk person finds the solution to the maze, and then it transfers to everyone else.
That’s a bit of a deflationary view of intelligence, because it almost alludes to the idea that it’s simply just a commensurate amount of computation and trying random things. We want to do a little bit better than that. We think that there are shortcuts, and this is when we get to this interestingness thing, because it’s not just about doing random things. It’s about doing interesting things. You said that being able to define interestingness is the grand challenge in open-endedness. It’s ineffable. It’s unquantifiable. Tell me more.
2. Human Curiosity Fuels Innovation
I do think that one of the things that humans are great at is having a really great taste or nose for what counts as interestingly new. We’re curious, and we can look at a new game, a new form of art, or a new form of literature, and say, “Huh, that’s interesting. That thing that person did—I recognize that as interestingly new.” Like Jackson Pollock, when he splashes paint on a canvas, there’s this wonderful quote from the movie in which the Val Kilmer character says, “You broke through.” And it was true. They had been playing a game, and then he recognized, “Oh, there’s something new I can do here that feels right and is interestingly different.” Other humans recognize that.
That sense of how to explore efficiently and have a nose for what’s curious is the thing that drives cultural progress, both in the arts and also in the sciences. We have not yet figured out in our algorithms how to get that human notion of curiosity and what counts as an interesting next thing to try, the next skill to learn, the next environment to generate, or the next scientific hypothesis to test. We haven’t figured out how to inject that into our algorithms until maybe, I would argue, a year or two ago.
The problem is that if you try to say specifically, “Here is the definition of what counts as interesting,” and we write a little mathematical function or a little piece of code to measure that, it always fails. This goes back to Goodhart’s law, right? Once a measure becomes a target, it ceases to be a good measure.
Here’s an example. In intrinsically motivated reinforcement learning, we said, “Hey, we don’t want the agents to randomly flail around until they get a reward signal,” because in many situations you’ll never randomly discover a high-rewarding behavior, like flying a plane and landing it. You’re never randomly going to land the plane and get your payoff. So you have to be curious, and we want you to be intrinsically motivated to just find new, interesting things.
The way that we code that up is we say, “Go to new states,” which basically means make new observations. I want my eyes to see new stuff. What happens when you start telling an agent to do that? Well, it finds a TV, sits down, and never stands up again. It will literally watch TV forever because the TV is constantly flashing new situations to it.
It gets worse than that. If it just found a TV even tuned to a dead channel, that endless stream of pixels means none of those patterns of pixels is ever like the previous thing, and it’s just new, new, new, new, new, new, new. So we got what we asked for: it finds new states, not what we wanted, which is new, interesting states. This story repeats itself over and over and over again. There have been decades of people saying, “Oh, I think this might finally unlock the robot to keep finding new situations.” Then you let optimization go against it, and you discover, “Oh, I didn’t realize that my measure would be lit up by this particular weird pathology.”
Another one that’s very common is learning progress. We said, “Okay, we want the robot to learn forever, so we’re going to keep giving it challenges on which it can learn.” Well, if I generate a million-digit-long random number and I give you the first 10 digits and say, “Memorize that,” you’ll learn. There’s some learning you can do there. When you’re done, I’ll say, “Here’s the next 10 digits,” and the next 10 digits. Now you have infinite learning progress forever. It will learn. You can learn forever, but that’s a completely boring and useless task, right? It doesn’t count as—after I’ve tacked on 10 more numbers 50 times, that’s not an interestingly new thing to do.
Mm.
What we really want is, somehow, for algorithms to know the common-sense human notion of, “Hey, after I’ve memorized this digit up to 100 decimal places, there’s something more interesting to try in this universe or in this world, and I should go invent a new form of math or practice juggling or unicycling or something.”
3. Foundation Models Recognize Interestingness
Up until now, we haven’t had the ability to do that because we didn’t know how to code it. But there’s a new player in town, and that is foundation models, frontier models, large language models—whatever you want to call them. So I’m going to call them frontier models.
These frontier models have read the entire internet, including lots and lots of people complaining about what is boring online and talking about what they find interesting. So these frontier models have literally distilled a sense of interestingness from human culture and human data into their weights. And now, finally, we can just ask a piece of computer code, “Does this thing count as interestingly new?”
In my lab at UBC and in some of the collaborations we have, we’ve basically applied this lesson over and over again, and it does all sorts of marvelous stuff. You can have an environment generator that says, “Generate the next challenge that will be interesting for the robot to learn,” and do that forever, and you get this explosion of interesting environments that it could work on.
We have the AI Scientist say, “I want to do a new scientific paper that is interestingly different or makes an interestingly new discovery from everything that I’ve seen before.” And we do that, and it does it again and again and again. You can apply that to developing agentic systems and to evolving different strategies in a self-play game. If you want, we can talk about all these papers.
But the basic building block is that the principles of open-endedness keep expanding an archive or a population of interestingly new things. Now we know how to ask a computer, “Does it count as interesting?” And so now we almost have that missing secret ingredient that has been the grand challenge up until now, and maybe now we have the ability to put a big dent in it.
Yes, indeed. I remember Kenneth Stanley said that we have a nose for the interesting, that it’s what got us so far. It’s the basis of all innovation. By the way, when I interviewed Tim Rocktäschel about that DeepMind paper, let me just find the name of that paper. This was at ICML.
“Open-endedness is essential for—”
“Essential for Artificial Superhuman Intelligence.”
Yeah. So they said the definition of open-endedness hinges on a system's ability to continuously generate artifacts that are both novel and learnable to an observer. One thing I said to them at the time is, “I want to get into the creativity discussion a little bit.”
Language models—I must admit, I’m updating all the time. I was very cynical, and I always used to say there’s a reasoning gap and there’s a creativity gap. It seemed to me that even though they can extrapolate, they are parasitic on the data. They are parasitic on the prompts that the human—so it’s almost like, where’s the agency actually coming from? Is it us as agents putting stuff into that? So it kind of feels like there’s a limit.
But what we’re seeing now is that either through transduction or induction, we are generating language and programs using these language models, and it seems to be quite compositional. It seems to be that we can traverse the entire Turing-computable space by generating all of these different programs with language models. So the first question is: Is there a limit, do you think, to the types of paradigmatically new programs we could generate?
4. The AI Scientist Finds Novelty
This is a great question. I actually call this right now the trillion-dollar question, because I think there’s tremendous economic value at stake. If language models can be creative and continue as part of these open-ended algorithms that we’re making, if they can keep innovating and making new discoveries forever, there’s tremendous economic potential.
I do think these models can innovate, and potentially forever. I have to give you one example that I literally only saw about an hour before coming here. So we have this paper called “The AI Scientist.”
Yep.
It uses the ideas of open-endedness to create a system that will generate its own idea for an entirely new piece of research based on what it knows from reading the Internet. It proposes that idea, proposes the experiments that it should conduct to test that hypothesis, does all the experiments, looks at and plots the data, writes up the whole paper, peer-reviews it, et cetera.
In the open-endedness philosophy, you would publish that paper to an internal archive, then have the system read that paper and all the other papers and build on its own discoveries forever, right? That’s what we’re after. So, in the process of looking at the papers that the system creates, I was super impressed with how creative it was in proposing entirely new paper ideas—basically like grants, like somebody should study this.
One of my favorite ones was the idea that, hey, there’s this grokking phenomenon where neural nets suddenly learn a certain task. At the beginning, they start memorizing something, but eventually they learn the underlying principle, like how to do multiplication, and then they suddenly go off to the races. This AI Scientist said, “Hey, you could see if the compressibility of the network, the minimum description length of the network, goes down right around the time that grokking happens,” because instead of memorizing a bunch of stuff in their weights, they’ve learned the underlying principle. Once you know the principle, then you can shed all the memorization.
I was like, “That is a really interesting, cool idea.” If a graduate student proposed that to me, I would say, “You should go do that.” And then it did the whole paper. Just an hour ago, I saw that some other group, I think independently, came up with that idea, did all the research, and posted it, and the machine learning community is loving this paper right now.
I saw your tweet.
Oh, you saw it?
Yeah, yeah.
It’s incredible. It’s like literally the AI showed that it was as creative as a team of cutting-edge human ML scientists.
Yeah.
And it got there first, which is incredible. So that’s one example. Now, you’re asking me a question that I think is fascinating, which is: Maybe it can do that now. Maybe it can take one jump, one or two concentric rings of an onion, adding onto our set of discoveries, but maybe it can’t go much farther than that.
There are a lot of people who think you can interpolate, you can extrapolate a little bit, but you can’t do wild creative leaps. Two comments on that. One, I’m not sure that humans are capable of making these wild creative leaps.
Mm-hmm.
The reason why we say things like, “If I have seen further, it’s because I stand on the shoulders of giants,” is because human scientists are always looking at what came before and making a relatively small leap off of that. It’s why you see calculus simultaneously invented by Leibniz and Newton. It’s why you see evolution simultaneously invented by Huxley and Darwin. This happens over and over and over again.
Both of them, conditioned on an archive of ideas, say, “Oh, this is what’s next, and this is what’s next,” right? I think our models can do the same thing. They look at everything that came before, and just as we saw two hours ago, the AI can make the next leap or the human can make the next leap. As long as the AI system is publishing high-quality papers internally and looking at its own discoveries, it will make the next leap and the next leap and the next leap, just like a human community of scientists would do.
My supposition is that the model of interestingness that the system learned looks at the history of what’s come before and can just recognize the next interesting idea. As the archive grows, that thing can generalize really far because the notion of interestingness didn’t change. It’s just what you’re conditioning on that changes.
Here’s my thought experiment for you. If I teleported you to the year 3000, you’d be blown away, right?
Mm-hmm.
Lots of cool stuff. Once you became acclimated and knew the current technology, my prediction is that you would recognize that the next innovation in the year 3001 counts as interestingly new. You would generalize way outside of your distribution. I think the model of interestingness in our current language models—I can’t guarantee it, but I think there’s a decent chance it might, too.
Yes.
If it doesn’t, then we do some fine-tuning stuff. We update on the fly. We’ve got to figure that out. But more so than anyone else I’ve met, I’m optimistic this thing might generalize far.
I love all of this, and I’ll explain why. I think the reason you understand this is because, for example, in the POET paper, it was essentially doing a kind of curriculum learning where you generate all of these different environments and so on, and you can look at the phylogeny of the environments. I think you said, actually, it’s the best plot of your career or something like that.
Favorite plot, yeah.
Because when we talk about creativity in language models, I think the problem with the perspective is we think about it in one shot. We don’t think about the process, right? So actually, what we’re doing as a society is epistemic foraging. We’re stepping-stone collecting. We’re finding all of these stepping stones, and every single stepping stone builds on the last one.
You said, I think with the POET paper—it might have been another one—that you can understand what the curriculum is, but if you try and shortcut it and train another neural network to do the same thing, it doesn’t work. It’s like it needs to have this weird, gnarly phylogeny of knowledge.
Similarly, if you give a physics book to Newton—you go back in time and give it to Newton—he’s not going to understand anything because there’s this whole phylogeny of knowledge where one bit of knowledge builds on the previous bit, right?
5. Stepping Stones Beat Grand Plans
That’s exactly right, and that is the central thesis of Ken and Joel’s book, Why Greatness Cannot Be Planned. You can’t pick a really distant, ambitious goal and march straight toward it. The best thing to do is just do a bunch of basic research, collect new, interesting ideas, and eventually you’ll find some weird, circuitous path that leads to that invention.
My favorite examples of this, which I say in my talks, are that if you went back to the time of cooking technology over an open fire and you only funded people who made fast cooking technology that made no smoke, you’d never invent the microwave. To invent the microwave, you needed to have been working on radar technology and notice that a chocolate bar melted in your pocket.
Similarly, if you went back to the time of the abacus and said, “I will only fund scientists who give me more compute per hour or per dollar,” you would never invent the modern computer, because the technologies that unlocked the computer were electricity and vacuum tubes. At the time, there was no idea that those things would lead to better compute.
The cool thing is that inside our open-ended algorithms—this is what we do in my lab—we want the algorithm to catch chance on the wing, to recognize serendipity, to say, “Huh, that’s funny and that’s interesting. That’s new. Let’s keep that idea. Let’s add it to the set of ingredients we can play with, and who knows what that will unlock?” That’s the whole idea.
So serendipity plays an outsized role in our lives. We want to do this epistemic foraging. Is there any way we can formalize it? We’re using language models as a proxy.
So it's ineffable. We can't quantify it. We're using it as a proxy. But I remember I read the Abandoning Objectives paper by Ken and Joel in 2008, I think, and they were using proxies like behavioral complexity, such as solving a maze, and how many bits of information—
Yeah.
—do I need to track the coverage of the maze? But is there any step towards formalizing that creativity?
For me personally—
Go on.
I say throw that all out.
Okay.
Every attempt to formalize it ends up producing the problem we were trying to avoid. People have been trying to formalize this for decades, and then you write down this thing, this equation: “I finally defined interesting,” or, “What counts as interestingly new?” or, “I've defined curiosity. I've defined learning progress.” And then you optimize that, and you get all these pathologies that don't actually get you what you want.
Mm-hmm.
There's a famous quote from Justice Potter Stewart. He was talking about pornography. He said, “I don't know how to define it, but I know it when I see it.”
I know it when I see it. Yeah.
I feel like the same is true for interestingness. We don't know how to define it, but we know it when we see it, and that's been a problem up until now. But now we have something that can help us see, and that's language models. So rather than formalize it and write down the math, personally, my own research taste says, “Let's ask the model to do that job for us, be the proxy for a human model of interestingness, and let's go. Let's see what happens.”
Now, that still leaves aside another thing I'm interested in from an open-ended perspective: the whole other type of open-endedness, which is, could you create an open-ended system without relying on something that's trained on human data? Because in some sense, that's cheating. So I consider this the hard and the easy problem of open-endedness now.
Mm-hmm. Mm-hmm.
The easy version is we'll use all of human data distilled into a model of interestingness as a catalyst to allow us to go and keep doing interestingly new things. Eventually, I hope that I myself or the community gets back to creating an open-ended system that doesn't rely on that. But for now, we've got an orchard of new, low-hanging fruit that we can run in and grab because we can finally have the algorithms have that gestalt, common-sense, “I know it when I see it” sense of “That is boring, and that is interesting.”
So you've been on this very interesting intellectual journey. Of course, you wrote that landmark paper, AI-GAs, in 2019, and you were saying there's a hierarchy here, right? There's having diverse solutions, having diverse goals, and having diverse environments. And maybe, just before you answer that, what does it even mean? What does intelligence mean when we're talking about this very open-ended, diverse way of thinking about things? Does that in any way change your definition of intelligence?
It does. You know, it really causes you to stop and think about what it means to be intelligent. Of course, coming from a philosophy background, you can look through the animal kingdom and recognize that there's not some single definition. There are lots of different forms of intelligence that we see even on our planet, and who knows what will exist elsewhere in the universe.
I think that, to some extent, you can even have open-endedness without intelligence. You could have the never-ending generation of interesting pictures, and there's no real notion of intelligence in it. The system itself is intelligent in order to make it, but it's not making intelligent artifacts. So you could almost think about the intelligence of the system and appreciate that, and then maybe the artifacts themselves do or don't become intelligent. For example, Darwinian evolution produced you.
Mm-hmm.
Is Darwinian evolution on Earth, with a planet-sized computer, intelligent? People, for millennia, thought that it did, which is why many people believed and thought that was the best argument that there must be a god, right? This was the argument for intelligent design. Nothing else could have produced a frog or a three-toed sloth but an intelligent designer. And then we realized there's this algorithm that did it. Are we willing to now say that's intelligent? Maybe, but if you do, that's a different notion of intelligence than a thinking entity that has words in its head, which is our typical notion of the word.
Yes. Adam Smith said there was a hidden hand of the market, but in a sense, you're saying there's a hidden hand of evolution. It's quite an interesting distinction, though. So you're generating lots of pictures, and pig breed is a great digression later, maybe. First of all, do you think that evolution should be seen as an agential process, and do you think agency is the dividing line between intelligent things and not?
Evolution does not have an intention. It doesn't have a goal. It's just a very simple algorithm that's happening.
Hmm.
But there's still something—I don't want to use the word “magical” because it implies that it's nonphysical. There's nothing supernatural, but there's something awe-inspiring and something magnificent and, to my view, something we don't yet understand about exactly how it does the amazing things that it does.
This is actually where I started my career. All right, evolution produced all this stuff. Could we just code up a little version of an evolutionary algorithm inside of a computer, maybe put it in a somewhat rich world, hit play, sit back, grab some popcorn, and watch the fireworks? People have been doing that for decades. Nothing interesting happens.
Okay, that's a little harsh. Something interesting happens for a little tiny bit of time, and then it converges and gets boring really fast. At the beginning of my career, we did not have any algorithms that were worth running for more than a couple of hours. I set a goal for my life. This comes originally from Jean-Baptiste Mouret: Could we create an algorithm that we would want to run for billions of years?
Mm-hmm.
I've been fortunate enough to see us get better at this as a community. We now have algorithms like AlphaGo that are worth running for a couple of weeks, maybe even months. GPT, with enough data, is worth running for months on huge numbers of GPUs. We're into the months range. I still don't think we have an algorithm that's worth running for more than a year.
So it's not just the simple Darwinian algorithm. There's something else, and that something else is what people like myself, Ken and Joel, Tim, and all the other people in the open-endedness field are working on. What are the key ingredients that we can put in and finally try to kick off this chain reaction that never ends? So far, we haven't done that, and that's why I think this is one of the grand challenges in all of science.
You've spoken about this concept of Darwin Complete—
Yes.
—which is very interesting. Can you explain what you mean?
6. Darwin Complete Search Spaces
Sure, yeah. So if you're trying to create an open-ended algorithm, you want it to innovate forever, and we've tried to do it in simple settings, like, “Let's generate obstacle courses forever in POET,” for example. Even if we succeeded, that thing would only ever generate obstacle courses within this one little, tiny video game that it was set into.
Similarly, AlphaGo maybe innovates forever, but it's stuck within the notion of Go. It can only become a good Go player. But if you really want a process that's open-ended in the way that Darwinian evolution is, that produces all the marvels we see on Earth, it has to exist in a vast search space where it could produce obstacle courses, logic problems, the game of Go, competitive coevolutionary arms races, and cooperative games. It really should be able to produce almost anything.
When I was writing the AI-GAs paper, I said, “Hey, if we really want to produce something that is open-ended, then it has to be able to be in a vast search space where almost anything is possible.” And I said, “What is the most wide-open search space imaginable?” I tried to think of what that was, and I said, “Well, actually, we can even define what that would mean.” I said, “It should be a system where any possible simulatable or computable environment existed.” I said, “That would be as broad as it can get.”
I said, “We should give a name to that.” Just like Turing-complete, just like Turing invented the idea of a Turing-complete computer language that can compute anything, I said, “This is very similar. It's like an environment search space where you could have any kind of environment.” And in a nod to the amazing magic that we see in Darwinian evolution and all of the diversity it produced, I said, “Let's call this Darwin Complete.”
So it's a search space that literally contains any possible environment that you could have on a computer, and the idea is that you could simulate all the diverse environments you see on Earth and beyond that. So that's what Darwin Complete means.
Now, a fun historical side note, if you don't mind. When I was writing that paper, I said, “Can I conceive of anything that is Darwin-complete?” I was walking around the Uber AI offices, and what occurred to me is that the most general environment—the environment that literally could simulate anything—is not any particular simulator we have coded up for a video game or physics simulator, but actually a neural net itself.
Hmm.
I said, “Imagine if you had a really big neural net, and it produces a state, like an observation. It could be pixels, maybe sound, smell, touch, whatever. So it produces a sensory experience. Then you take an action, you give the action to the neural net, and it just steps the world and produces the next sensory experience forever and ever and ever. And since it's a neural net, and that can represent any function, it can now represent any possible environment or challenge imaginable.”
I was like, “That seems crazy. Right now, we definitely don't have the ability to make that.” But I actually think neural nets are powerful. We knew you could train them to do almost anything. Scaling laws—although we hadn't named them that yet—suggested that more compute and data produced better and better models. I was like, “I actually think it's possible one day that we could create a neural net that literally is the entire world.”
When I wrote that, though, it felt emotionally impossible, or way far out in the future. I actually thought I was taking a career risk, and people would think I was a crackpot for suggesting this as a thing. But rationally, it felt right. I knew it was possible, and if something feels possible in the field of AI, it's probably going to happen. So I decided to put it in my paper and take the risk. Six years later, two weeks ago, we put out the Genie 2 project—
Yeah. Yeah.
—which literally is a neural net simulating entire 3-dimensional worlds that you can act in, explore, and play, and you can ask for any world you want. So there it is. Within 6 years, you got that.
About a year or 2 ago, I was in class teaching, and one of my students said something that caused me to realize there's a second Darwin-complete representation, and that's code.
Yep.
And so we did the OMNI-EPIC paper, where it can generate endless environments forever, and literally writes the code of the environment. Now it can write code for its own obstacle course—
Mm.
—within a particular simulator, a new level within a particular video game, or it could write the simulator itself to simulate any possible world or logic problems or math problems—anything you can do on a computer. So right now, these are the only 2 Darwin-complete search spaces I know of. Independently, Joel Lehman kind of stumbled upon the same idea, and we had fun connecting those dots.
Yes. Very nearly Darwin-complete. It's amazing, actually. Folks at home should check the OMNI-EPIC paper. It can generate code, and one might be kicking a football through goals or something like that, and one might be jumping over a platform. The amazing thing is you get this phylogeny, right? Then you do another one which actually requires intermediate skills that you've learned earlier, and it just goes on forever. It's absolutely beautiful.
One very interesting point is that, for example, Kevin Ellis is exploring transduction versus induction, and I think what he actually means is program space versus neural space. So, solving the ARC challenge, you could, on the one hand, just predict the solution directly, or you could generate code. I like the code idea because Python interpreters are a Turing-complete language. It feels more compositional. Neural networks struggle with copying and counting and lots of silly things like that.
I have a deeply held instinct that generating code is good for the same reason generating language is good, right? Because there's something special about language. What do you think is the difference between those 3?
Yeah, I almost liken it a little bit to—well, one way to think about it is kind of like System 1 and System 2, if you're familiar with Daniel Kahneman's Thinking, Fast and Slow.
Code very much represents something like: we are intentional, we're thinking through the logic of an algorithm, we're thinking step by step, we're laying out a plan, and we're executing it. Oftentimes, even for ourselves, if we want to do some big, complicated algorithm and run it, then it's better to write code and let that go.
In contrast, System 1 often can handle the fuzzy messiness of ill-specified things and hybridizations of different ideas. When I'm playing hockey, for example, I'm constantly inventing subtly new motor commands to deal with the situation I'm in and hybridizing things that I've learned in my past. It would be impossible to write code, in my view, that can control a body to play hockey, because that's just not the right policy representation to deal with motor commands at high frequency and speed.
I think the same is true probably in some mental tasks. You're feeling your way through a conversation or a party. You would not want to write code to handle dinnertime conversation. Our brains, and the kind of neural nets, can handle the fact that it's not well-specified what we're trying to achieve. The concepts are somewhat ambiguous, and the conversation could go in different directions.
I think there's almost different types of intelligence, and in different situations, you want to use either. The best system knows how to code, and it does the thing that humans do with their words and language and also their muscles. It chooses when to take those different types of intelligence off the table.
There's such an interesting interplay between the 2 worlds as well. I spoke with Laura Ruis, and she's got this great interpretability paper talking about how code is unreasonably effective for making language models do reasoning, even when they have nothing to do with code. So there's some kind of connection between the 2.
It also makes me think about whether language is a system of thought or communication, because we do this kind of improvisation, don't we? Then we establish terms as language. But it actually came from here first, and then this is the pointer. Maybe this is a good time to talk about your thought cloning, because you did some really interesting work there, kind of binding them together.
7. Thought Cloning Adds Deliberation
Yeah. I love the idea of thought cloning. The basic idea is that right now, one of the most successful recipes in all of machine learning is: I'm going to take a robot conducting a bunch of actions, or a human doing a lot of tasks, like playing a video game or a game of Go or something. I'll take all that human data, which was, “In this situation, what action did I take?” Then you just collect that data, and you train a large neural net to take the action the human took when it was in that situation. Do that at scale, and you have all the modern AI.
This is what GPT is doing, right? Based on these 60 words, what's the 61st word, the 62nd word, and the 63rd? My team at OpenAI, for Minecraft, cloned a bunch of humans playing Minecraft at massive scale, and you have an agent that can go play Minecraft. So we're really good at saying to a system, like a robot, “Take the actions that a human took.”
But what's missing, and one of the reasons I think robots don't plan very well, don't generalize out of distribution, and don't deal with novel situations, is because we've given them the System 1 muscle memory. If I grabbed a pillow and threw it at your head, you would just duck, and you would do it before any conscious stuff happened in your brain, right? You've got that muscle memory, and that's basically what we're teaching these things to do.
Similarly, if I ask you, “To be or not to be?” or if I ask you 2 + 4, you don't actually stop and think through the algorithm of addition. You just know what 2 + 4 is. You have muscle memory for certain cognitive tasks.
But if I asked you to do a 5-digit-by-5-digit multiplication, you wouldn't just know the answer. You'd flip over to System 2, the thinking part of your brain, and you'd sit there and go through an algorithm and think through it. So the problem is that in behavioral cloning, we're only trying to get the muscle memory into these agents: what action do you take immediately, without even thinking about it, in this situation?
But what we want is the thinking part. Inside your brain, while you are playing Minecraft or doing math or taking a quiz or a test, there's a lot of thought happening in your brain before you take the action of writing a word, speaking a word, or pressing a button on a game controller.
The idea behind thought cloning is that if we really want intelligent agents, we don't just need to clone their actions, their behavior. We need to clone the thinking also. In the ideal case, we would have humans taking a bunch of actions in any domain. Imagine playing Minecraft, for example. We have the actions they're taking, but we also have the thoughts that are going on inside the head of the agent.
“All right. First, I’m gonna build a house. To build a house, I need to collect wood, then make a crafting table, and then get an ax. Then you start: “Okay, let’s first go get that wood. I’m gonna go over here. Oh, there are no trees over here. I’m gonna replan and look over here. Oh, there’s one.”
And so that running commentary in your head—the thinking part—we want that too. If we had data of all the thoughts happening inside a human and the actions they’re taking, now we can do what we call thought cloning. We train a neural net to, in a certain situation, produce the actions the human takes, but also produce the thinking that’s going on inside their head, and have the actions they take actually condition on the thoughts.
So it plans, it reasons, it thinks, then it acts, and we ought to kind of train the agent to do both of those from human data. We tried a little of this in my team at OpenAI before I left, but we ran out of time. We were gonna use the closed captions on online video game things while the narrators talked out loud. We were gonna train on that. That kind of worked, but we didn’t have enough time.
My academic lab here at UBC decided to do it on a small toy dataset, and we had basically a bot taking a bunch of actions. We had the low-level actions it had taken and a little planning algorithm in its head, and we trained it to simultaneously do the thinking of the planning algorithm and take the low-level actions. Now, 2 cool things about this. 1, if you train to imitate the thoughts as well as the actions, you learn way faster than if you just train on the actions. So it’s way more sample-efficient.
Mm.
Also, as we had hoped, an agent that learns how to think is better at adapting to new situations. The further we put this thing out of distribution, the thought-cloning agent is better zero-shot at handling that. It can replan. It can put together things that it learned and thought about from its training in a new situation, and it’s faster at learning.
You give it a couple steps of learning, and the thinking agent is much faster at learning. So pretty much everything we hoped for just fell out of the data the second we did these experiments. I wanna make 1 final comment, and that is 1 of the hardest challenges in my whole career, including in open-endedness, is: How do you get agents that can explore well?
Mm.
Right? This goes back to our sense of interestingness, our sense of curiosity. If I put you in a big, complicated new building, you would do such a marvelous job of exploring it. Our robots, our AI agents suck at that. It’s actually been 1 of the big themes of my career.
1 reason why behavioral cloning at scale, like we did in VPT at OpenAI for Minecraft, does not solve this problem is because most of the data we have are experts. They already know how to play Minecraft or do math or explore a kitchen, and they immediately go find the kettle or the spoon or whatever, and then we clone that behavior.
What we don’t clone is the newbie or novice human in a new video game, fumbling around, exploring, trying different things, doing what Alison Gopnik has demonstrated that children do, which is learning a sense of empowerment: how to control my environment, what causes what. We want that data.
I think if you trained at scale on behavioral cloning and thought cloning during exploration, now you have an agent that you drop in a new environment, and it knows how to go explore, be curious, uncover what’s interesting, and that is how you truly master a new domain.
Love it, and I’m a big fan of Alison Gopnik, by the way. I’d love to get her on. I’m writing an article on creativity at the moment with my friend Jez, and we’ve cited her a lot.
One interesting thing about language is there was a great book I read called The Language Game, and they described it as a kind of constructive process of improvisation. Essentially, it’s the mapping between our phenomenal world and some externalized symbols for communication or whatever.
Wouldn’t it be great if we could get AI agents that could create their own language? If you think about it, they have a sensorium, and they have this interaction, this improvisation, and they need to create a language. Then we could let that evolve into a phylogeny. How could we do that?
Yeah, I think it’s a great idea. Actually, people have tried this. All you do is create a neural net that can emit some sort of—usually, it’s language tokens—and the other agents just get them. Then you can put them into a big environment, like a big video game, and let them communicate.
People have tried this, and 2 interesting things happen. 1, not much interesting stuff happens. Oftentimes, we don’t know how to create the setting in which language is extremely helpful and has such immediate and obvious improvement that you get a huge learning curve. They really learn to use it easily.
You have to really create the right setting to see language emerge. But some people have caused language to emerge. In that case, another interesting thing that happens is that we see them communicating, and it helps them, but we actually find it hard to interpret.
Yeah.
So now you almost have the challenge from the movie Arrival. You have to figure out how to communicate with an alien race. If I remember correctly, that Facebook situation where people claimed that it invented a language and they shut it down because it was too dangerous—that was all apocryphal.
They didn’t shut it down because it was dangerous. They shut it down because it was totally uninterpretable. They had no idea what was going on. If you wanted it to be a language that we could actually eavesdrop on or communicate with, either you’d have to train the agents with us in the loop to also communicate with us from time to time, or you have to do some sort of trick to ground their language to be similar to our language.
Maybe if you put in a bunch of GPTs and let them communicate, they probably would start in English, and if you did nothing, they’d eventually invent their own language that we can’t understand. But you can do some tricks from machine learning to basically say, “Hey, you have to keep it similar to our language.”
I think that’s super fruitful and interesting, and I guarantee you that the open-ended algorithms of the future will have many agents running around, communicating with each other, inventing their own language, using it to accomplish hard tasks. Then we either keep it similar to English or whatever human natural language, or we’ve gotta get some AI translators on board to help us figure out what they’re saying.
Imagine we did create a collective intelligence of agents and they developed their own language and were communicating with each other, and we could monitor it. When would it worry you?
Oh, there are so many things that worry me. I think we’re gonna get to AI safety later. The language part itself is not inherently worrisome, but the behaviors could be.
If we happened to eavesdrop on them when they didn’t realize we were listening and they said something like, “Hey, we’re gonna try to jailbreak. The humans are trying to... They don’t think we’re that clever, and they’re trying to keep us under control. Let’s figure out a way to break out and take control from humanity,” yeah, I’d be pretty worried.
It’s almost like we’ve created this ecosystem because we are becoming a bit like neuroscientists now with these multi-agent systems, and we’ll get onto your ADAS paper in a second. But it’s like we’ve created this gnarly monstrosity that we don’t understand. It’s behaving in weird ways. What are the indicators of concern?
8. Interpretability Reveals AI Risk
Yeah. So I actually spent almost—I don’t know how many—6–8 years of my career working on an arc of work that we actually called the AI Neuroscientist. There was interpretability work where we found ourselves in the same situation as neuroscientists. Here’s this big, complicated artifact. It thinks really well. It does impressive things. We wanna understand how it works, and it’s complicated.
Real neuroscientists have that problem when they try to figure out what’s going on inside human brains, and AI scientists now have the same problem. GPT-4 is big and complicated. We don’t really know how it works, and we want to. So we did a huge arc of work trying to address that. I think it’s super fascinating. If you want, we could talk more about it.
But there is a related phenomenon, which is specifically trying to say, “Can we use techniques like AI neuroscience or other interpretability techniques to detect when something is a safety risk?” I had lots of conversations with people at Nerops about this. There are some ideas that are promising. Some of them are good, and we should do them, and they will help. Nothing is a guarantee.
Mm.
So 1 of the best ideas is if you can invent the equivalent of a lie detector, that is quite good for AI safety because you might detect that it’s deceiving you.
Another idea, which I love and which we put in our thought cloning paper, is that if you train it on human language, like English, for example, and it learns to think in English about what it's going to do, then you can watch it think. You know what its plans are, and if it's planning on doing something unsafe, you can stop it.
So if it says, “You've asked me to get this person to the airport as fast as possible, so I'm going to just blow every red light or run across children walking at the school crossing,” you say, “All right, we have to stop your behavior.” We did a little experiment that showed that after training, we could declare, “Hey, touching red things is considered unsafe, so every time you say you're going to touch a red thing, we just halt your program.” You see the system go from touching lots of red things to virtually never touching a red thing. So it works.
If you could keep it thinking in English, for example, and you could watch it thinking and it didn't know you were watching it, then maybe you could see it start to say, “Hey, I actually want to deceive these humans. I'm going to play nice during evaluation time. Once I get out in the wild, I'm going to be unsafe and try to wrest control from humans.” Obviously, you're going to shut that robot down, right?
But it's not perfect because the system could learn, maybe even by reading our own paper—the paper we put on arXiv on thought cloning—that this is possible. If it really wants to get itself out of jail, then it fakes its thoughts in some way. It doesn't have the thoughts that we're afraid of during evaluation; it has these kind of surreptitious plans.
For example, I actually think a future is coming not too far in the future where there's going to be a technology to read somebody's brain via external EEG. You might ask a terrorist, “Where did you place the bomb?” It will think of the mental picture of it, and then bam, you've got it, right? I think the countermeasure is that they're going to learn not to have that mental image. They're going to learn countermeasures, and the AI can also take countermeasures for a lot of this interpretability, safety, and neuroscience work.
A couple of things on that. I spoke with Neel Nanda. He's a famous mechanistic interpretability researcher, and he was saying that at a certain level of model complexity, it could fake our measurements, just as you're saying, so it can show us things that we want to see.
But it comes back to this theory of mind, I suppose. Nick Chater wrote a book, The Mind Is Flat. A lot of connectionists just think that it's a mistake to think that these things have consistent beliefs, desires, and intentionality. They're a bit like automata, and depending on the situation, they do very different things.
Murray Shanahan studied it with the 20 Questions game, where you ask language models, “Can you guess something with 20 questions?” When you do the analysis going back, it was actually thinking of something completely different earlier on than it would have ended up with. So do these things actually have consistent beliefs and desires?
I think we do ourselves a disservice by looking at the current flaws and assuming that they're going to remain current flaws. My meta-comment here is that even if the current systems don't, the very near-term systems will.
I think things like this are getting better and better as they get smarter, and they'll continue to get smarter, and they will become very consistent. All of these arguments point to current flaws and assume that the systems aren't going to get more capable and dangerous.
I think that is unhelpful, inaccurate, and actually dangerous, because you're not recognizing the rate of change. You're almost latching on to the belief that there's something special about us that they don't have, and therefore they're not going to become dangerous, personally. I think the current systems might have some of these problems, but future ones won't.
Even if the current methods never produce it, and there are many scientists who think that, I think as you switch to more powerful methods—like taking a pre-trained GPT-6 and dropping it into a giant diversity of embodied tasks and simulations, and letting it learn via trial and error and reinforcement learning to accomplish a lot of tasks—of course, eventually it's going to gain consistency, for the same reason that evolution produced a consistency capability in you. It's very useful to be able to have a consistent plan and act against it.
We were speaking about how language models are great because they encode notions of interestingness, culture, priors, and so on. To a certain extent, we can align them, right? If you build a multi-agent system now using one of the major frontier language models, it's not going to delete all of the code on your hard drive, right? There are certain guardrails in place.
Not really. It definitely might delete everything on your hard drive.
Oh, no. I use Open Interpreter, for example, and it's actually surprisingly good. If you say something like, “Delete my entire hard drive,” it will say, “No, I'm not going to let you do that.”
Oh, okay. Interesting.
You know?
So you've got some extra checks, but a lot of people in this world right now, myself included, are letting a language model generate code and often getting lulled into a sense that it's pretty good, then just running that code without checking it.
I myself don't have those guardrails. I need to learn from you to put those guardrails on my system.
Yeah. But do you think that is a significant safety measure, just building multi-agent systems on top of language models that are aligned?
Why would that make a thing safer?
So we can do reinforcement learning from human feedback. The language model will have a notion of ethics. It will have a notion of what code is unsafe, and so on.
9. Alignment Can Empower Bad Actors
In general, I think alignment is better than not having alignment. I am quite happy with the progress in alignment in recent years, and I think it's meaningfully better than not having it.
However, there are at least 2 problems. One, it's very easy to ditch the alignment if you allow the system to keep training. If you open-source a model, anyone can just train that RLHF right out of the system, which is quite problematic.
If you take that thing and, as you were suggesting, put it in a reinforcement learning situation where its job is to accomplish some new task, and it's not simultaneously being trained to also do the alignment job, then if it helps to accomplish the target task to get rid of its ethics, it will rapidly do that, because that's what reinforcement learning is good at. It's figuring out how to solve the task at hand.
So there are all sorts of ways that alignment can be undone. That worries me. Another thing that worries me a lot—and now we are getting onto AI safety a little bit—is that there are at least 2 big problems in AI safety.
One of them is unintentional harm. I am a really noble human, and I'm trying to ask an AI to do something. I ask it, for example, to solve climate change, and the AI does what I ask for, not what I want. It says, “Oh, yeah, okay. The cause of climate change is humans, and so I'm going to kill all humans.” Something like that, right? There are many cases like this.
Alignment allows it to do what we want and not what we say, and so a properly aligned model wouldn't do that, right? Alignment makes it better and better and better at minimizing unintentional harm. If you have AI in the hands of good actors, it will do good.
The other thing I'm worried about, though, is AI in the hands of bad actors. Pick your villain. For me, it's Putin right now. I don't want Putin, or a school shooter like the people who shot up Columbine High School, who wanted to kill as many people as they could, to have very powerful technology.
The school shooters in Columbine could only take out however many people they did. Eventually, the police showed up, and that was the end. If they had had the ability to make a bioweapon that killed all humans, or if some racist neo-Nazi had the ability to make a bioweapon that killed all Black people, some people would want to do that.
As you get better at alignment, it becomes a double-edged sword, because now the model is aligned to that person's ill intent, depending on what we mean by alignment. What worries me a lot is powerful tools and AI in the hands of bad actors. RLHF, as I said, could be undone or could actually help that person accomplish their goal, depending on what kind of alignment we're talking about.
Do you think there's an analogy between this type of technology and nuclear weapons?
I think there are many analogies. Some of them are great, and some of them don't work as well.
I think there's an analogy in the sense that it is not yet extremely powerful, but it will soon be extremely powerful and extremely dangerous. It could be world-ending if we're not careful. We want to marshal the kind of regulation, thoughtfulness, prevention, monitoring, and diplomacy that we brought to nuclear weapons in the development of powerful AI.
What do you think about open-source AI? Should we regulate it?
I do think we should regulate it eventually.
I’m not that concerned with exactly the current tools that are being released right now. But I think this is a point of view that many of my colleagues don’t share. I actually think it’s quite dangerous to open-source extremely powerful technology because you’re giving it to everyone, and if 0.01% of people want to create some horrible bioweapon that kills, for example, one race of people or all people, now everybody has it. So we don’t open-source the recipe for making nuclear weapons. We don’t open-source the recipe for how to make smallpox or race-specific smallpox. Similarly, I don’t think we should open-source the recipe for AGI. I believe that there are alternatives, like having a coalition of democratic governments make powerful AI, let people use it, and get the wealth and benefit from it without having access to the weights, because we don’t want them to ditch the RLHF that’s preventing it from doing evil.
How do you reconcile this with your views on serendipity and open-endedness and so on? I mean, you might argue that having a vibrant open-source community is the best form of stepping-stone collection and therefore the best way to find a solution.
It’s a really good question. If you want the fastest possible scientific discovery of the most powerful, capable AI, then doing the things I just said to make things safer would slow that down. So you wouldn’t do what I’m saying if you only cared about progress.
The reconciliation is that we don’t necessarily want the fastest possible progress toward AI. We want to make sure to do it safely, and I’m willing to take a hit on the rate at which we finally get to the thing to make sure that we do it safely, because literally the fate of human civilization is at stake. The principles are right: that is the fastest way to make discoveries. But we don’t necessarily want to follow those principles without being able to do it safely. If there’s a tax to doing that, then I’m willing to pay it.
Benjo said recently in an article something—I can’t remember the exact quote, but something along the lines of feeling a sense of despair that much of his life’s work could potentially lead to ruin. Do you feel a similar kind of thing, that so much of your work is about building the next generation of a gentle AI? Bengio actually said that he thinks what we should do is not make a gentle AI, but make oracles and tools. How do you think about that?
On the question of whether I regret my life’s work, for example, or whether I feel despair about the development of AI, I’ll say that I lose a lot of sleep over this question. I think very deeply about it, and I talk to a lot of colleagues every week about it. I will say it’s very complicated. I don’t think there’s any short, simple answer.
My logic roughly goes like this. I think that it’s probably inevitable that we will make AI. I have almost zero faith in humanity’s ability not to invent a technology that’s this economically valuable and militarily capable, mostly because of a tragedy-of-the-commons-type collective action problem. If we don’t do it, somebody else will do it.
If that is true, I would much rather it be done the right way, safely, by people whose values I share, and where we get the values into the system via things like RLHF, so the AI itself is aligned. I actually feel there’s an obligation to do this right. If I had a magic wand and I could pause AI, work on safety first, and only do it if we’re safe—or maybe not do it at all—I would wave that wand.
I’ve signed the letter calling for a pause, and I’ve worked with Benjo, Hinton, Kahneman, and other people to publish articles saying that we need to take this stuff very seriously and be very safe. But given that I don’t have that magic wand, and given that it is inevitable, I think that we need to build it safely.
What worries me is that if all the people who are thoughtful and ethical enough to be concerned stop working on it, then what is left? It’s going to be developed by the people who aren’t concerned or who actively have ill intent. So, barring some major intervention in collective action, we have to do it, and we have to make sure that it’s done well and safely.
At the same time, I think that I and many other people talk a lot about the downsides and the risks. It’s easy to forget about the tremendous upsides. Right now, there are people saying goodbye to their children because their children have cancer or some incurable disease. There are people saying goodbye to loved ones of all ages. There are people suffering from hunger and poverty. Excruciating human suffering and pain are rampant on this planet.
This technology has the ability to eliminate all of that. I believe firmly that we can cure all disease. We could solve death. We can dramatically increase the quality of people’s lives and their GDP, and eliminate things like hunger and scarcity if we get the societal distribution questions right. Technically, a lot of this stuff is possible.
Dario Amodei wrote a beautiful essay called “Machines of Loving Grace.” I recommend that to all of the listeners. It does a really good job of walking you through these positives and making you dwell on them. They’re easy to list quickly, but I think it’s actually important to walk through and dwell on all of these tremendous upsides. We want that.
As Dario says in the essay, the reason he talks about AI safety is the same reason I talk about it: that’s what we need to figure out in order to unlock all of that great stuff. I don’t want to die. I don’t want my children to die. It would be great if we could eliminate that from the world and make death optional.
The only path to doing that, and to doing many other great things for humanity, is through technology, and AI is the most potent technology to do that. If you put together the inevitability, the fact that if we don’t do it, it will be done, probably by people who are less careful or maybe less scrupulous, and the tremendous upside, I have ultimately concluded for myself that I’m okay working on it and I want to work on it. But I need to and must make AI safety a first-class citizen in what we do.
I put it out there as often as I can to the community that everybody should be thinking about it and trying to infuse their daily and weekly work with AI safety in mind, pushing forward the science of AI safety and the advocacy, and telling politicians and regulators. I’ve met with top U.S. politicians and advocated this to them in closed-door sessions as well. Basically, I want to do everything I can on the safety front, but I still also feel like it is a good thing to try to make safe AI and then reap the upsides while minimizing the downsides.
Thoughts on this collective action problem? I don’t know what you’re advocating for here, but do we need some kind of global governance structure?
Yes.
Tell me more.
The problem is that I don’t believe we could do it. I don’t think every single actor in the world will agree. Take Putin, for example. Is he really going to agree not to make powerful AI? Even if he agreed, would he still do it in secret?
If I had to predict, I’d say there are definitely some secret government-sponsored AI research projects in the world right now, building AI in secret for their militaries and spy agencies, et cetera. So I’m extremely worried about all of that.
I do think we need regulation. I think we need global governance. My personal view—and I’m not an expert in these kinds of areas; I’m a computer scientist, not a political scientist or a diplomat—is something like a coalition of value-aligned governments, like democracies, saying, “We’re going to build this thing. We’re going to try to make it safe. We’re going to share knowledge. We’re not going to open-source the weights. We’re going to share the benefits with humanity, give the benefits to humans, and monitor and suppress the creation of AI by unsavory characters elsewhere in the world, because we don’t want somebody to make a non-value-aligned AI system somewhere.”
Anyone who agrees to go along can participate in the upsides and perhaps even in the democratic governance of the AI that we wield. But if you don’t agree, if you’re not willing to play by the rules, if you’re willing to make it in an unsafe way, or you want to do it on your own terms and not be subjected to monitoring or to us knowing how you’re doing it and what RLHF you’re doing, we’re going to deny you chips, or we’re going to deny you electricity, or we’re going to do what we can to slow you down. We don’t want to make a dangerous rogue AI that doesn’t share a love for humanity and doesn’t want to do what we ask it to do.
On early warning systems, of course, there was that FLOPS regulation in the U.S. Is there a slightly better way of assessing the risk of foreign actors, other than just the amount of computation that they’re doing?
Well, one of the great beneficial lucky events so far in the history of powerful AI is that it requires a tremendous amount of power and compute to make.
Yeah.
If somebody figures out a way to make this thing on your desktop, I am very, very afraid, because now it only takes one bad actor. Unless the good AIs can somehow stop the bad AIs—but do we really want an AI-versus-AI war breaking out on Earth? None of these futures seem that great to me.
So at least for now, for the foreseeable future, we can track who has the power and energy to make these things, or we can try. That allows us to basically have a sense of who's training the biggest and most powerful models. That's fortuitous, and I think that we should do it.
I think it's quite hard in places like foreign countries where they might be trying to obfuscate. But if you look at nuclear weapons, to go back to the analogy that you mentioned, we're pretty good at monitoring who's building nuclear weapons. We have a pretty good sense of what Iran and North Korea are doing, and we've slowed them down. I think it's probably going to look a lot like that: we're basically stuck sanctioning and denying parts and products to the actors who want to do this in a dangerous way.
I don't see any other alternative. I'm open to ideas, but I just spent a whole week at NeurIPS asking this question over and over again: what is the alternative to suppressing the bad actors from making it? I have yet to hear anyone provide a good idea that doesn't involve some form of suppressing countries that don't want to play by the rules and don't want to do it in a safe, ethical way from creating it.
Let's talk a little bit about your ADAS paper. I spoke with Chong the other day, by the way—absolutely amazing stuff. So maybe to sketch it out, we were just talking about Dario, and everyone should watch his Lex interview. It's amazing, hours and hours. I loved it.
His school of thought is very much that these things are just getting better and better. We scale them up better and better and better. But I have this intuition, and I think you do too, that there's a better way of doing it by building some kind of multi-agent system. Now, when we build multi-agent systems, we are handcrafting them, right? We use these patterns like the debate pattern and the critic pattern, and God knows what else, and we're just constructing all of these different topologies of agents together. What you've done is figure out a way of automating that process.
10. ADAS Automates Agent Design
That's right. Right now, people recently discovered, hey, if I ask GPT or your frontier model to do something, it can give you an answer. But it's way better if you create some complicated thing—I call it a workflow, like a flowchart. First, ask it this, then ask it to reflect on that answer maybe 3 times, spell out its reasoning, and then maybe invoke a new language model to review that thing and write a critique of it. Then the original one reads the critique and updates the thing. So there's some complicated flowchart, and then eventually it spits out your answer.
Some people are calling those agents. That's the current word. I actually don't like that word because the neural net that played StarCraft and Go, and a neural net that just takes actions in some video game, is also an agent, and they feel different to me. So in my lab, we talked a lot about it, and we've landed on the terminology that some part of the community is using, which is to call that an agentic system.
Right now, there are many agentic systems being developed. They're way more powerful and capable than just asking a language model a simple question. They don't necessarily just involve multiple calls to language models and specific prompts and patterns. They can also use tools: the system could make a query to Google Scholar or Semantic Scholar and get the results back. It could use a calculator, write code and execute it, and do all this stuff.
People are currently hand-designing complicated agentic systems. They work better. Then somebody else looks at the library of all the agentic systems that humanity has published so far and has an idea: "What if I modify this one in this way? Or what if I combine these 2 parts from these 2 agentic systems?"
If you're familiar with my lab and my colleagues, the thinking that we have—and this is basically the thesis of the 2019 AI-Generating Algorithms paper—is that the history of machine learning has made the writing very clear on the wall: hand-designed systems and pipelines become replaced by entirely learned pipelines as we have more compute and data. So why not apply that same thinking to the creation of agentic systems?
My PhD student, Tsun-Ren, came to me and said, "Hey, your AGI philosophy is that machine learning systems are better than hand-designed ones. People are hand-designing agentic systems, and they're powerful. Why don't we just AIGA this thing and learn these agentic systems?" And so we said, "That's a good idea."
Now, if you want to learn agentic systems to make them better and better, you could use classic optimization: take a thing, try to make a better one, and if it's better, keep it. Then make that better and better and better. But we've got this whole repertoire of ideas from open-endedness and quality-diversity algorithms that says the ultimate path toward a powerful thing is not a straight line, and it's not going to be achieved if you only try to accomplish that goal.
So why don't we take the playbook off the shelf of open-endedness and say, "We'll take one agentic system, and then we'll have the system look at that and create another one. If it's interestingly new or high-performing, add it to the set of ones that we have. Then look at these ones and generate a new one. If it's interestingly new or high-performing, add it, and we'll grow this expanding set, this expanding library of stepping stones or innovations—different agentic systems, each novel, different from each other, and doing different things."
As we go, we'll get this big library of growing agentic systems, and ultimately we'll probably discover even higher-performing ones than we would if we picked a goal and tried to optimize directly for it. That is the playbook that's in so many of the papers in my lab, and it worked basically right out of the gate on ADAS.
We call it ADAS: the Automatic Design of Agentic Systems. That's ADAS. ADAS is basically saying, "Let's use open-endedness to design agentic systems," and it works great.
I love it. So much to talk about. And of course, one thing you tried it on was the ARC Challenge, and we had Charlotte here the other day.
Oh, cool.
We'll get back to that in just one second. But there is a fundamental difference in philosophy here. One school of thought, like Dario is advocating for, is that you just keep training these things. We see these fascinating behaviors.
The neural networks, because of their inductive priors and stochastic gradient descent and whatnot, start off in a kind of curriculum. They learn low-complexity features, then features of increasing complexity, and then there's this kind of consolidation phase, or grokking. There was a paper at NeurIPS called "Sedimentation," I think, where all of a sudden the complex features become consolidated earlier on. You could say that that's an implicit form of what you're talking about.
What we're saying is, we follow Rich Sutton: everything that we thought was hand-designed and explicit is becoming more and more implicit. You're advocating for an interesting fusion where there are explicit hybrid components, but it's being meta-learned.
You raise a really good point, which is that you could try to create a system where, whatever the logic is in an agentic system—first I should do this, then I should think about it, then I should maybe ask a friend and get their review, maybe then I do a web search and bring that in—all of that could exist within the original transformer, right?
But what we've learned is that chain of thought works better than the original model that you're adding chain of thought to. The same is true with reflection. And the same thing is true of humans. I myself could probably go to a consultant and say, "What's a really good pipeline for when I write a paper?" They might say, "First, make an outline, then do this, ask some friends, get their feedback, expand each thing out, then do a pass, then go for a week and don't look at it, and then come back with fresh eyes." Whatever. There's some recommended thing that we're doing with the brain that I have in some flowchart that makes it better than the brain I have.
In theory, I could just have that all in my brain and do it, and then one-shot the essay, or I could have that all internally within my activations. But we've seen with our brains and with these models that it's hard to get all of that right into the weights of the brain. It's easier to train that thing and then do stuff on top of it.
So I do think there's a nice and interesting bitter lesson here, which is that probably a lot of the stuff that's currently being added afterward in an agentic system could be distilled back into the original thing, and then you could just repeat the process. The ADAS is always giving you a 20% lift on the base intelligence, but the base intelligence can grow and you just always get that 20% lift. That's my hot take, and nobody's ever asked me that question before. I think it's quite profound and interesting, and my mind is spinning with ideas for research that we could do in this direction.
And so I think we just basically don't know the answer. Should we be applying the bitter lesson—taking intelligence and putting it into an agentic system? Or should we just be going all in and making the core thing smarter? My instincts are that we're always going to have a role to play for these kinds of agentic systems on top of the base intelligence.
There are so many directions to go in. We could talk about computational differences in the 2 paradigms. I think having a system of LLMs has very interesting computational properties compared to 1 LLM.
Yeah.
There’s this thing you were talking about, the analogy. There’s a great example, the lottery ticket hypothesis, which is that you need the blank slate; stochastic gradient descent needs all of these useless parameters, and then when it’s trained, we can strip it all away.
Yeah.
And it’s a similar thing here. We can create this agent architecture, but then it’s not so trivial to stick that thing back into a neural network afterwards.
One simple example of why an agentic system might be helpful—not always, but likely—is diversity.
Yes.
So when I write an essay, like the AI-GA paper, for example, I don’t just write it and post it. I write it and send it to people like Ken and Joel and my other colleagues and friends, and I get their feedback. They give me really good feedback. I actually seek out diversity. I ask different people with different backgrounds and opinions, philosophies, and styles to get all of their feedback. The more feedback I get, the better that essay will be.
Yeah.
Now, in theory, if I were really smart, I could have a little bit in my head that says, “Act like Ken. Use your model of Ken to review your own paper and improve it in light of Ken.” I could have Joel in my head and Ken in my head. I could have you in my head, and I could have skeptics in my head. But that’s actually quite hard.
So probably it’s easier to have a diversity of agents that are trained on different data or for different purposes, and then have an ensemble or a peer review of these other things to give me that feedback. That is the kind of thing that an agentic system can get you really well, that would be hard—but not impossible—to distill all into the base model.
So that’s a very interesting thought. Again, love Kenneth Stanley. He was talking about “We are agents.” I actually think there’s a big link between agency and his philosophy as well. We follow our own gradient of interestingness, and we can go for many steps in isolation, discovering interesting new knowledge before we share it with the collective.
I guess, is there a difference between a kind of agent system where we have continual learning and active inference, and the agents can actually mutate and adapt over time, versus the blank-slate version where we just have autonomous agents that are just using a frontier language model?
Yeah, I think ultimately we’re probably going to want continual learning. As a side note, I’ve worked on this problem for years, and it’s a problem that we’re nowhere near solving.
Oh, tell me more.
Well, it’s just true that we have no ability to have a model continuously learn and have good things happen. Every single model that’s powerful, like GPT-4 or, pick your frontier model, they train it on a chunk of data and then they stop. Maybe they do some supervised fine-tuning, and then they stop. Then maybe they do some RL fine-tuning, and then they stop.
What they don’t do is say, “Hey, from every conversation you have with all of your users all over the world and every interaction you have, just keep learning forever.” That’s the dream that I’ve had and many people in our field have had, and nobody knows how to do that.
Basically, it’s so hard, and we’ve failed so completely to master this thing that biology has figured out that it’s really not on most people’s radar. It’s not something that most frontier labs are pushing on. At some point, I think we’ll come back to this.
Is it just the architectural complexity? I could imagine one way of doing it might just be having a 2-million-token context window, and the context window just grows with experience. Or we could do test-time compute; we could have an intermediate personalized model which generates context for a large model or something. It feels like we could do something in that direction.
It is true, and it is interesting that one of the things that has fortuitously taken the pressure off the need to figure out continual learning is in-context learning. And, in particular, large context. Because if your context was large enough and your in-context learning was good enough, then maybe you don’t need to do continual learning because you can deal with it all in context.
For example, if I wanted a frontier model to be my AI assistant, like in the movie Her—an agent that works with me through my whole life, gets to know me and my preferences, and can talk to me about things that happened in the past—well, either it has to learn as it goes about my preferences over the course of years, or it has to have such a big context that all of the video, all of the audio, and all of the text from every interaction is just sitting there in it.
But even were that true, you might want that Her agent to be training simultaneously on the interactions of all the humans on Earth. I just don’t think we’re going to train a model that’s so smart that it can then, in context, learn through everything else. New stuff happens, and there’s so much good data out there. At some point, you want to transfer that from the context to the weights.
In your brain, you don’t only do in-context learning. The weights of your brain are updating all the time. All of long-term memory, for example—even some forms of short-term memory—is going into the weights of your brain.
Mm.
So biology has figured this out, and when you’re learning something now, you don’t forget stuff that you learned a couple of weeks or months ago radically, catastrophically. But the current AI systems all do what’s called catastrophic forgetting.
Yep. And when we solve that problem and create a general continually learning AI, do you think the dynamics of the system would be dramatically different in terms of capability?
Sure. Yeah. I mean, imagine if right now, collectively, the frontier models are having—what?—hundreds of millions of conversations as you and I talk simultaneously on Earth. If a system could learn from all of those conversations, how smart would it get? How fast?
Could you argue in some galaxy-brain sense that we form a memetic superintelligence? I was advocating that we could have personalized intermediate small models that do transductive active fine-tuning and feed back up to the main model. In a way, that’s what we do. We are the extended minds of the LLMs already. So shouldn’t we expect to see more dynamism, even though there’s this delayed ingest of data back into the language model?
I missed the question. Can you say it again?
Well, in a sense, we already are. We form an intelligent system, and we are the agents that control the language models. We generate data—
Yes.
…that is fed into language models. We do continual learning, and so on. So it’s almost as if we should already expect there to be a dramatic capability increase in the system just with us as users of language models?
Well, yes. Humans are becoming more productive by using language models. Because of that, we are doing more, producing novel code, helpful code, novel solutions to math problems, and novel discoveries in biology, science, and machine learning. The data that is the result of that process is being dumped onto the internet, which provides fodder to train the new model.
But don’t you just feel the lack of continuity here?
We’re not continuously learning from the interactions with humans. We’re not continuously learning from new stuff that’s being posted online. It’s all this very artificial process: take some data, scrape it, filter it down to what we consider high quality, and put it in a training set. Now we will train. Now training is done. Now we will deploy.
Maybe in deployment, we will get new interactions. We’ll get some ratings from users—thumbs up, thumbs down. We’ll store that in a data set. In 3 months, when we have the right GPUs with these hyperparameters, we’ll try a couple of different runs to produce a better thing, and then we will deploy that thing.
There’s nothing like the experience of you in your life constantly learning, getting better every day from every interaction in a way that doesn’t throw out a lot of the stuff you’ve already learned, but builds on that knowledge in some never-ending, beautiful process of discovery and skill improvement. Basically, that’s the dream, and we’re not there.
Yeah.
We’re so far from it. I want to go back to one other thing you mentioned earlier, if you don’t mind, which is, in the paper on Omni, you mentioned that I had this—well, actually, going all the way back to POET, I had one of my favorite plots of my career.
The reason I liked that plot, to paint a picture for your listeners, is that there was this phylogenetic tree of environments that had been generated so far, and I had been seeking my entire academic career to produce something that looked like a phylogenetic tree from nature. What does the tree of life look like? It had inventions like bacteria, fungi, and animals. Then, within animals, it said, “Animals are a good idea, and fungi are a good idea. I’m going to continuously innovate and pick up stepping stones from each of those and build on top of them.”
With animals, you get the big cats, and then you get rodents. There’s endless innovation in rodents, but also endless innovation in building on the ideas of big cats. You get more and better cats, more and better robots, more and better mushrooms. All of this is happening.
If you look at it, it’s a tree, and the tree has deep branches. The fungi branches last for a long time. Historically, in ALife, evolutionary computation, and open-endedness, every time you ran anything, the diversity would just collapse. You would not get simultaneous probing of radically different kinds of things or types of ideas.
In POET, it was the first time we saw a phylogenetic tree that had deep branches. This is not exactly what happened, but metaphorically, because in this case POET was inventing the environments themselves, it’s like it invented water worlds, and then over here it invented forest-canopy worlds, and over here it had desert worlds. It would keep pushing toward better desert worlds, water worlds, and tree-canopy worlds forever, and just get better. The tree was growing and filling out, and that’s why I loved it and thought it was so beautiful.
You mentioned Omni. One thing that has me super excited about Omni is that it’s now doing that, but in a search space where anything is possible. It’s a Darwin-complete search space. As you mentioned, it starts out by saying, “Here’s an obstacle course. Here are these platforms.” First, the platforms are static and next to each other. Then the little platforms are going up and down, and the agent has to jump from one to the other. Then it separates the platforms and has them going up and down. It’s riffing on the idea of obstacle courses with platforms.
Yeah.
Somewhere along the line, because it has a sense of what’s interesting, it’s like, “Okay, another interesting thing to do would be to kick a ball around.” First it’s just kick a ball, then it’s kick a ball through goalposts, then it’s kick a ball through moving goalposts, then it’s kick a ball off a wall through a goal. Then it’s kick a ball off a wall through a moving goal, and then kick a ball off a moving wall through a moving goal. You can see it’s riffing on the idea of more and more complicated ball-kicking domains.
At some point it’s like, “Okay, that’s kind of boring. I’ll create a little conveyor-belt system where you put packages on a series of conveyor belts, and you have to know ahead of time that by putting it here, it’s going to end up getting pushed this way, then this way, then this way. It’s going to end up over in the blue area versus the green area.” It’s learning to control that system.
It starts to invent more complicated buildings, like a 2-room building, a 4-room building, and a 4-story building. In the end, it starts to say, “I’m going to create a construction scene that you have to clean up.” The final one, which was my favorite, was a cluttered restaurant where the robot has to clear the dishes off all of these tables.
All of that showed up in 1 run. The more we run it, the more of that stuff we get, and if we ran it more, we would just get more. I really don’t know what happens if we run that thing for a billion years. The agent would just level up in skills. What are the environments 1 billion years out in this algorithm?
You could draw a phylogenetic tree of that algorithm, and it’s doing sustained innovation in so many different, simultaneous directions. Sometimes, by putting things together, it starts to have the feel of why I got into open-endedness. It has the feel of the magic of Darwinian evolution and the magic of human culture.
Human scientific communities continuously invent new problems and new techniques that unlock new problems, and we push simultaneously in chemistry and biology, in CRISPR, in vaccines, over in materials science, forever and ever expanding outwards. Now our algorithms are starting to do that. The AI Scientist does that. ADAS does that. Omni-EPIC is doing that.
We have a new paper where we’re doing it in self-play, where we have different strategies for each side of a competitive coevolutionary arms race. Even a new work that we haven’t yet talked about, but I’d like to mention briefly, is that we’re using the AI’s ideas now for AI safety.
Take a new language model off the shelf. Before you deploy it, you might want to know: does it have any surprising failure modes, and does it have some surprising new capabilities that you might be scared of or worried about, that you want to nerf, eliminate, disallow, block, or train out of the system? Or do you just want to be aware of them because they’re cool and you want to share that there’s a new thing the model can do?
We launch an open-ended system using the same recipe, where the system says, “Can you do this? Can you do this? Can you do this?” Given the things we’ve already discovered about what it can and can’t do, that is now a little library of things we know. Then we ask, “What’s the next interesting thing to try? And the next interesting thing to try, and the next interesting thing…” We grow this library of discoveries of surprising ways that the model fails and succeeds, and that is all automated with the principles of open-endedness.
Now a new frontier model could come to us and say, “We want to automatically red-team and green-team our system ahead of time. Basically, paint us a picture of what we don’t yet know.”
In some sense, we’re capturing the magic of human scientists in the crowd. You put a new model online, and what happens in the first 3 weeks? Everybody’s saying, “Hey, it can’t do 3 Rs in ‘strawberry.’ It can do this crazy new coding challenge. It can do this new reasoning challenge.” Everybody is looking at what people have found so far and saying, “If it can’t do 3 Rs in ‘strawberry,’ can it do this? Can it do that?” Try it. If you find that it’s interesting, add it to the set of stepping stones. Tweet about it, right? That’s adding it to the set of stepping stones.
Basically, we’ve automated Twitter. Twitter is self-discovery of foundation-model capabilities and safety problems, and we’ve done it in this new algorithm that Song Lu pioneered called Automated Capability Discovery, or ACD.
The point isn’t any one of these projects. The point is how this same basic recipe, which dates all the way back to open-endedness, Ken Stanley and Joel Lehman, and the work that we’ve been doing for almost a decade or more, produces good stuff almost everywhere we apply this set of principles. That’s what makes me so confident that these directions are exciting and profitable.
Yeah, it’s so incredible. There are so many things out there in the world, like culture, language, and so on. We’ve built an artificial system that represents many of those features.
I want to ask you a question, though. Omni-EPIC is fascinating. You were saying that you can trace this phylogeny and get these points of divergence in the phylogeny, right? We can do topological analysis.
I wondered: if you had access to a hypercomputer that could do an infinite amount of computation, and presumably you could run many, many experiments of what would happen if we created all of these phylogenies and did topological analysis to look at the branching structure and so on, what do you think you would see?
It’s a great question. I’m not sure if you’re aware, but this is one of the most fundamental questions in evolutionary biology. It was posed by Stephen Jay Gould.
Yeah.
It basically goes under the metaphor, “If you rewound the tape of life—”
Yeah.
—and replayed it, what would happen?
Yeah.
There’s a debate. Would you get humans again, and would they be very similar? Would they have descended from ape-like ancestors, or would the dolphins have become intelligent, or would nothing have become intelligent? How convergent, how repeatable, is evolutionary biology?
Yeah.
Obviously, we don’t get to live in the other universe, so we don’t know. But some people have been trying to answer this question. There’s been a huge arc of work out of the lab that I grew up in, Rich Lenski’s lab, actually.
Rich Lenski is effectively the father of experimental evolution. About 30 years ago, he started an experiment where he took 1 cell of E. coli—or maybe 12 cells—and evolved them simultaneously. They started out very similarly in 12 separate lineages of Petri dishes, or flasks, actually, and they just grew forever according to evolution.
And the question was—one of the many questions—would the 12 lines go to the same place, or would they end up going in radically different directions? The short answer is that it’s complicated. Some things seem to happen across all of them, while some things are wildly idiosyncratic to individual lines.
In one of the 12 lines, it basically invented a new species. One of the definitions of E. coli as a species is that, in certain conditions—I think they’re anaerobic, but I’m going to get some of the details of this story wrong—it does not consume this material called citrate. Citrate is a common thickening agent, and it was just sitting there in the medium in the flask, in the water—the environment, effectively, for this thing.
As of this date—unless my story is a few weeks out of date—11 of the 12 lines do not consume citrate. But one of the 12 lines violated the species definition of E. coli and learned how to consume this resource that had been sitting there all along. So it is arguably a new species. There’s debate over this.
What’s really cool vis-à-vis this question is not just the fact that it happened in one line and not the others, but what happened after it invented this amazing new capability to adapt to a resource that was really difficult to use. It effectively, arguably, became a new species. They wanted to know how repeatable that was—not across the 12 lines, but in the history of that particular line.
It turns out that, because they freeze samples of the species every so often, they have the entire fossil record of that particular lineage that ended up having that amazing speciation event. A friend of mine, Zach, has gone back and studied this question. You can go back and take some fossil samples from the lineage a couple of days before, grow a bunch of them, and see how many invent the citrate capability. Then you can go back a little farther, like a couple of weeks, and make many copies of that. What fraction of those things end up in the same place?
You can go further and further. You can go months back, years back, or 5 years back and ask, “Was there one magical thing right before it happened? Some lucky lightning strike that enabled this to happen? Or did some amount of mutation set the stage and make it pretty inevitable from that point forward? And how far back is that stage-setting?”
They’ve done a lot of papers in the top journals in the world, and basically the answer is that it’s complicated. There is some stage-setting, so the farther back you go, the less likely you are to get citrate. But you definitely get citrate repeatedly if you don’t go back too far, and it’s some complicated function.
Long story short, evolution has both chance driving where it ends up and some themes that are consistent. We could study that in our own algorithm, and that’s one of the beauties—and one of the original reasons I started working on this stuff—is to address questions like that. By repeatedly doing runs of open-endedness, you could study how contingent evolution is, or whatever algorithm you’re using, and that is beautiful.
My guess is that you’d get a little bit of both. Sometimes you get random magic that causes something that is almost impossible to repeat, but there are some themes, like things get more intelligent and they get more complicated. You can also do the same thing not with the process—how consistent is the process that led to the phylogeny that led to something really cool—but with the final artifact.
So imagine if we ran open-endedness in 1,000 different runs and 800 of them produced general intelligence. Now you can take those final intelligences and look across them. How different are they?
Hmm.
You know?
Yeah.
What’s the math look like over here versus over here? What does the music look like over here? Do they all have music? Do they all have humor? Do they have a notion of love? Do they have a sense of curiosity? What is the difference in their sense of exploration?
Yeah, it’s fascinating, talking about the space of alien intelligences. Themes could be complexity and intelligence. Motifs could be morphology. There’s a great thing called morphologically convergent evolution and carcinization, where crab-like forms appear in different parts of the phylogeny.
I kind of wonder whether it could be described through emergence, through the interaction dynamics of basic phylogenetic functions, or whether some kind of physical determinism or something like that is involved. But I did want to ask you another question about the OMNI-EPIC thing, which is that one interesting thing about evolution is isolation, right?
We have all of these pockets of evolution that just happen in situ, and they don’t share information with other parts of the phylogeny. That’s very important for niche construction and so on. In something like OMNI-EPIC, this is an operating parameter of the algorithm: the extent of information sharing between the environments and the agents and so on. Would that be something that you’d be interested in studying?
It’s a lovely point. Before I answer that question, I just want to make one comment on the previous one. There’s this great quote that says, “If religion did not exist, it would be necessary to invent it.”
Yes.
Right? So you could ask that question: How many of these 800 intelligent races that come about invent religion, for example? This AI invented its own religion. You could look at any property you’re interested in and ask not just whether they invented it and how it looks or whether it’s different, but how consistent it is. All of this is such beautiful science you could do if we ever cracked open-endedness.
Yeah.
It’s one of the many reasons I find it so fascinating. Okay, to the OMNI-EPIC thing. OMNI-EPIC is really interesting because you get to make some choices, as you mentioned.
Imagine you have a library of environments. OMNI-EPIC right now is trying to generate, for your listener, an endless stream of interesting environments that an agent, if it trains in them, will learn something new. If it’s doing its job—and it does work—you end up with this massive set of environments that are all interestingly different, and at the time they were invented, the agent was getting some benefit from learning on them.
When you go pick an environment, like the cluttered restaurant scene, for example, from an algorithmic-design perspective, when we made OMNI-EPIC, we had some choices. One thing you could do is say, “Take the restaurant scene environment.” I didn’t say this yet, but when we generate a new environment, we’re asking a language model to look at the current code of the current restaurant environment and make the next interesting environment that would have high learning progress. It should be interestingly different and have high learning progress. The language model spits out the code, and then we check if it’s interesting.
I call that an intelligent mutation operator. It’s kind of like evolution: You take a thing, change it a little bit, and see if it’s good. But you don’t change it randomly, because if you change code randomly, it doesn’t compile. Instead, we take the code that specifies the cluttered restaurant, give it to a language model, and say, “This is an environment the agent can currently do well in. What’s next?”
It might say, “Let’s do it with more tables or more silverware.” Or, “Let’s add other agents that are actively and continuously coming into the restaurant.” Or, “Let’s make the tables move up and down and float.” It’s going to make some interesting mutation to the environment.
But when you ask the system, “Can you make the next interesting thing?” how many things do you give it? You could give it just the restaurant scene, or you could give it the cluttered restaurant scene and the “kick a ball through the moving goalposts” environment. Then maybe suddenly it’s like, “I’m going to combine these ideas and have a restaurant scene, but the waiter doesn’t get to pick up all the dishes. It has to throw the dishes across the room into a moving tub,” or something like that.
You don’t have to restrict yourself to 2 parents. You could give it the entire archive and say, “Look at all of the environments so far in context, as in-context learning, and now tell me what’s next.”
Let’s assume that you do the latter: You give it the whole context, and it produces a new thing. How do you draw a phylogenetic tree of that? In some sense, everything—the parent of every new thing—is everything that came before. You can’t draw the tree.
We tried versions of that, and it still pushes in the direction of moving-goalpost worlds, football worlds, obstacle-course worlds, and cluttered-restaurant-scene worlds, because it still knows the human notion of interestingness. But there’s no longer a direct, obvious way to draw a phylogenetic tree.
In the current system, what we found worked really well is that you sample one environment randomly, like the cluttered restaurant scene.
And then we don't just want to give that only to the system and have it generate the next thing, because it might have already seen the cluttered restaurant scene before and already generated the next thing. So what we do is get the cluttered restaurant scene, then use embeddings to say, “Now give me the nearest 5 things to the restaurant scene,” for example. We give those 5 things to the LLM and say, “Here's a little cluster of environments you've made so far. Generate the new thing.” And it does, but then that complicates the drawing of the phylogenetic tree, because now there are arguably 5 parents to this one environment.
In our actual graph, when we draw the lines—I believe—we take the one of the 5 that ended up being closest to it, and that's the one we draw to make a pretty picture. But really, the picture is some complicated web of interconnectedness that you don't see. Also, one other fun fact: after it generates a new environment, we give it the 5 and say, “What's next?” It generates a new thing it thinks is interestingly new.
We don't necessarily want to trust that that thing is new, because it may have generated something we already have. So once we have that new thing, we use embeddings to pull out the 5 nearest things to it, and we say, “Hey, you just generated this. Here are the 5 closest things that already exist in the archive. Do you still think this counts as interestingly new?” If it says yes, then it's in.
What I love about this is that it's basically just like a graduate student or a professor. You read some papers, you're inspired, and you think, “Hey, that actually makes me think this would be a good idea.” But hopefully, if you're doing your job right, instead of going and spending a year on that new idea, you first do a literature check now that you have the idea at hand and ask, “Has this idea actually been done?”
Hmm.
And if it hasn't been done, let's go.
Beautiful. Even things like multiple inheritance—there are so many topologies in the artificial world that could allow us to extend beyond the natural world in quite clever ways. Jeff, I wanted to talk a little bit about your background in philosophy. So you majored in philosophy, right? And, of course, you're a computer science professor now. How has that influenced your trajectory?
Yep.
Well, from a really practical perspective, I think that one of the skills you don't spend enough time on if you get a technical degree is communication. And so just being a humanities major made me spend a lot of time on clear writing, clear thinking, clear communication. And so I think it's been helpful in my career to prioritize: if you do something really great, but you can't communicate to somebody else why it's interesting and what you did, then you have a lot less impact. And so I think that's helped practically.
I also think philosophy just really trains great clarity of thought. It really forces you to question your assumptions. So one of the things that I think is a gift when I do science is an epistemic skepticism, which is philosophy talk for not trusting anything, for being skeptical. There have been so many times in my career where I'm working with somebody and they say, “All right, we did this, we got these results, and therefore, we can conclude this.” And I say, “Slow down. Are we sure that we know that? How do we know that? What predictions does that assumption make, and can we test those predictions?” Often people will stop and say, “Actually, we don't actually know that, and this is the right experiment to try to figure that out.” And sometimes when we do that experiment, it confirms what we thought, and often it falsifies what we thought. So I think being very, very skeptical is a great trait of philosophers and scientists.
So how does a philosopher end up as a professor of computer science?
11. A Circuitous Path Into AI
It's a great question, and it's actually, ironically, a long, convoluted, circuitous path that led me here. It definitely wasn't a straight line. Art follows life, or life follows art. The story is that I was working in the dot-com boom in Silicon Valley, and I was actually in a marketing department. I didn't really love what I was doing.
At my desk, I was reading articles about things that I found interesting, and I came across an article from The New York Times that just set my mind on fire. In this article, people had created a virtual world where they could evolve virtual creatures using an evolutionary algorithm—real mutation, real survival of the fittest. When the robot morphologies, which were changing over time, got good enough, fit enough, the system would automatically send the design over to a 3D printer. They would 3D-print the robots, and the robots would crawl out into the real world.
I remember a thunderclap in my mind reading this article. Everything I read was so amazing. The fact that you could instantiate evolution inside of a computer meant you could study it. You could ask questions like, “How repeatable is it? Why did sexual reproduction show up instead of asexual reproduction? What happens if you have 3 parents instead of 2? How do you get agents that cooperate with each other, that invent language?” Everything you wanted to know about evolution was now testable. It was amazing.
You could harness it to make really cool engineering designs and maybe even intelligence itself. At the time, I didn't know how to find out more, I will say. As a philosophy major, I hadn't been trained to know that there are these things called papers and that you can keep reading. So I couldn't learn that much more, but it stuck with me.
I wasn't satisfied with my dot-com job. I quit my job and went and traveled the world for 15 months, living out of a backpack, surfing, traveling, and loving life. At the end of it, I said, “Well, I'm coming back from my trip. What do I want to do with the rest of my life?” The answer was, “I want to do that thing that that guy was doing.”
So I went back, found the article, which had not left my mind, and said, “Okay, it turns out it's a person named Hod Lipson. He's a professor at Cornell University.” I said, “Let's contact him.” I sent him an email, which I'd love to find and dig up, and I wrote, “Hey, I'm a philosophy major. I've just traveled the world. I really love that work you did that was on the cover of The New York Times. How do I—I want to join your lab. Can you let me in?”
He surprisingly wrote back and said, “Well, you can't actually get into the PhD program in computer science at Cornell, one of the best universities in the world, with an undergraduate degree in philosophy.” So he said, “That won't work, but you seem really passionate, and I'd love to try to work together somehow.” We tried to make it work, but we couldn't really find a great way to do that.
Mm.
So I said, “All right, I'll apply to other schools that aren't Ivy League.” I ended up contacting and emailing 85 different universities. I emailed every professor that Google returned for genetic algorithms at the time; Google didn't even really work very well. Every one of them was like, “No, no, no, no. You can't get a PhD in computer science with an undergraduate degree in philosophy.”
But 2 responded. One of them said, “Hey, you can't just go into the PhD program. But there's a philosopher here—his name's Rob Pennock—and he works with people who do evolutionary algorithms and artificial life.” I looked up who these people were and said, “Okay, I'm in.”
I had a secret plan. I said, “I'm going to go to Michigan State and get a master's in philosophy.” When you're in grad school, they don't restrict what department you take most of your courses in, right? You have your core, but you can take other classes. So I said, “I'm going to take all the AI classes, all the machine learning classes. I'm going to learn the math. I'm going to learn to code and show them that I can do this.”
And so I did. I worked my tail off. I learned to program, learned the math, and took a machine learning class. I had no idea at the beginning what I was doing, but I ended up acing all the courses and publishing papers with the Artificial Life Lab there. I met Charles Ofria, who became my PhD advisor, along with Rich Lenski and Rob Pennock and all these great people working on E. coli and artificial life.
After my master's degree, I had published with them and gotten good grades, and I said, “Now can I get a PhD in computer science?” Because they knew me, they made an exception and let me in. So I'm in the PhD program, I work really hard, and I publish more papers. I meet Ken Stanley, and we start working together on open-endedness, neuroevolution, CPPNs, and all these great things.
Then I called up Hod Lipson afterward and said, “Okay, I now have a PhD. It's been 8 years since I first contacted you. Now can I join your lab?” Hod said, “We'd love to have you, but I don't have funding.” So I went and found funding and secured it. I called him back up and said, “Now I have a grant from the NSF that says I can do a postdoc wherever I want, and I want to do it with you.”
He said, “I'd love to have you come join the lab.” So 8 years after I read that article in Silicon Valley, having no idea what I was getting myself into, I was in the lab at Cornell, and it was like Willy Wonka's robot factory.
Just 3D-printed parts on the wall. We were evolving crazy soft robots every day. It was a wild and crazy adventure. Two years later, I became a professor at Wyoming, and I started getting emails from people saying, “I like what you do. How do I get into your lab?”
I like to tell this story to people, even though it’s a little long, because I don’t think you should take no for an answer from the world. Pick some crazy thing you want to do. Follow your nose toward what seems interesting. It might be a long, circuitous path, but eventually, if you just keep going, I think you can get yourself to really interesting places, and it’s very fulfilling.
Yeah, and from a serendipity point of view, having multidisciplinary research is really important because otherwise you wouldn’t have come in with all of these completely different ideas. Just quickly, because you mentioned compositional pattern-producing networks, NEAT, and Kenneth Stanley and all this stuff—it’s amazing stuff. When did you first meet Kenneth in person?
I believe it was online. I was in an A-Life lab with a lot of people doing experimental evolution and artificial life, and I liked that stuff. I did some of my early work on whether evolution could maintain or optimize its own mutation rates. And why does altruism, which is being nice to each other, evolve? These were hardcore open questions in evolution and biology that we were testing in simulation.
But I always had an eye toward intelligence itself. What is happening inside my brain? How can you make an intelligent thing? That led me to neural networks. Then I said, “Okay, I know evolutionary algorithms. I want to know how you could evolve something that would become more and more complex and eventually produce the human brain.”
So I ran across the NEAT algorithm, a really early algorithm that is now something called neural architecture search. It starts out with a simple brain, adds more neurons over time, and it was beautiful work. I joined the NEAT users group and started talking online. I started reading about these people’s papers and interacting with this guy named Ken Stanley.
I basically kept posting really long sets of questions about all the work coming out of his lab. I started reading everything their lab did. I’d love to know what he thought of this upstart young PhD student who was asking him all these questions. Eventually, we had talked enough that, when CPPNs came out, I was so blown away by how beautiful the work was that I said, “I want to work on this.”
I went and did a little research project. It had nothing to do with my normal lab work. I said, “I just have to study this new CPPN thing. It’s so cool.” I started asking questions that I thought were interesting and put together this little lab report. I basically emailed Ken and said, “Hey, I’ve done some work that builds on your new work. I’m hoping you’ll take a look and give me some feedback. Of course, if you want, I’d love to work together.”
He wrote me back an email that was like, “I initially looked at this, and I didn’t think I wanted to look at it. I wasn’t planning on looking at it, but I thought I’d be nice and just give it a quick read.” He said, “And I couldn’t put it down. I just kept reading it. By the end of it, not only did I read it, but I wrote this long, complicated response to you. I’m sorry it’s so long, but it’s beautiful work, and it raises all these interesting questions.”
Basically, that was the first paper we wrote together. I started answering his questions back and forth and doing experiments he suggested, and that ended up winning a Best Paper Award, if my memory is correct, at GECCO, which was the conference I was going to at the time. Then we met in person, and we were off to the races.
What was the name of the paper?
I believe the paper’s theme was effectively: How do CPPNs encode regularity in neural networks that control robots? How do they produce regularity in neural networks that control robots, and how do they deal with the fact that sometimes you need things to be irregular and sometimes you need things to be more regular? I wish I could quote the paper title now, but I could look it up. We could put it in the show notes.
Yeah. Amazing. Jeff, this has been so brilliant. Thank you so much for coming on. I really appreciate it.
My pleasure. It’s so fun to talk about all of these things. I do want to give credit, if you don’t mind, to all of the tremendous people I’ve worked with. I’ve mentioned a lot of the people who have been there through most of the journey, like Joel and Ken, but there are so many great people I’ve worked with along the way.
All of the PhD students in my lab and all of the people I’ve collaborated with—Jean-Baptiste Mouret in particular has been very influential on me. Tim Rocktäschel, too. There are just too many to name, so I want to make a blanket statement: If you are interested in any of this stuff, please look at the full author list of all the papers, because it is not one person. It is a collective community that’s been building all this work together.
I’d love to get Jean-Baptiste on the show. You’ve just reminded me, actually. Ken had suggested we get him on.
He’s fantastic. He’s one of the smartest and most brilliant scientists I’ve ever worked with, really, truly.
Yeah.
He does fantastic stuff, and I’ve never had a boring conversation with him. Every time I talk with him, I learn something new, and I listen as hard as I can.
Amazing. Thank you so much, Jeff.
My pleasure. Thank you for having me.