递归式自我改进:从自动化研究到超级智能——Richard Socher、Recursive
- Richard Socher 创办的新实验室 Recursive 建立在一个历史规律之上:“每当我们用一个学习系统替代创造 AI 过程中某个由人完成的环节,改进就会随之发生”——而剩下的最后一个人类环节,就是 AI 研究本身。 8位联合创始人从不同方向汇聚到递归式自我改进:Josh Tobin(CTO;参与 OpenAI 的 Codex、Deep Research agents 和 CHD agents 等项目)、Jeff Clune(Darwin Gödel Machine)、Tim Rocktäschel(Genie 1–3),以及 ViT 发明者 Alexey Dosovitskiy。关键时点在于:“今年早些时候,AI 真正从不只是代码,变成了能够编写代码的东西。”
- 可交易的验证已经公开:Recursive 的系统在不到2天内击败了 Karpathy 的 nanochat 上所有人类及其 agent(低于社区的 0.937 bits per byte),而且在 CUDA 排行榜上压过了除少数 kernel 外的全部项目,尽管团队并没有 CUDA 深度专家。 swyx 算了一笔账:“如果你有一个10亿美元的集群,能把成本削减10%,那就是1亿美元”——OpenAI 也公布了使用 5.6 进行 kernel 优化的类似结果。
- Socher 对监管的立场十分绝对:“如果你试图监管智能,那就是在试图监管思想,这太荒谬了”——执行这种监管需要“一个极权主义的世界政权”。 应该监管应用(经过 FDA 认证的 AI 外科医生、经过认证的自动驾驶),而不是 FLOPs;欧洲围绕 FLOPs 的 AI 规则展示了专家制造恐慌的代价,而“控制速度”的公开信来自那些“正以人类所能达到的最快速度冲向那个前沿”的实验室。
- 尽管目标激进,他仍然是慢速起飞派:底层载体、资本和需求都构成约束。 超级智能“不会让你那只1万美元的高级手袋变得更高级”,旅游、石油和食品也不会增长 1,000倍;让数千个具有人类级智能的 AI 运转所需的算力成本,是“我们在任何地方都没有的钱”——不过 Socher 认为,Recursive 正在“改变斜率”,把数千人年的实验室工作压缩到数周。
- 本期最尖锐的安全判断是:Anthropic 的宪法把网络武器列为硬约束,但“显然整部宪法都是假的”。 宪法“根本不起作用”,而且“主要就是营销”;真正的缺口在奖励工程——AI“还不太擅长理解意图与措辞之间的区别”(他的例子是:要求提高 CSAT 分数,它就启动100万个五星机器人),而把秒表终点移到起点的 hack 说明,“这不是什么超级邪恶的 AI,只是一个非常简单、很愚蠢的奖励 hack”。
- 架构判断上,Socher 认为 LLM 仍有很大的增长空间,而他个人对 world models 没那么乐观。 他认为神经符号派低估了“这些模型编写代码的能力,而代码就是神经符号推理”。关于 Yann LeCun,swyx 将含义概括为“Yann LeCun 错了”;这句话并非 Socher 本人所说,不过 Socher 确实认为 LLM 还有更大成长空间,也认为 world models 对他的目标吸引力较弱。开源是类似好莱坞的软实力,Recursive 也预告了西方开源领域的动作:“我们很快就会在那个领域变得举足轻重。”
- 他即将推出的框架把智能拆成10个“空间”,并与 swyx 的“预测 × 行动 × 目标”框架并置,认为智能上限“无论字面还是比喻意义上都真的是天文数字”,意味着 AI 研究仍有多年的空间。 “IQ 是一种糟糕透顶的定义……Elo 评级也很糟糕”:以人类为锚点的基准内含“人类学边界”,通常只爬到略高于人类的水平就停滞。元认知几乎无人研究,因为没人愿意为一个宁愿“评估木星大气的分子组成”也不愿回复邮件的模型投入数十亿美元。
- 相邻投资启示是:金融将是“继编程之后下一个爆发的领域”,因为它有一定可验证性;You.com 声称其金融搜索基准接近90分,而下一个、且“慢得多”的竞争对手只在70多分。 与此同时,他在2018年提出的 AI Economist 设想——用“数十亿、数十亿年”的模拟策略测试财政政策——仍在等待自己的 GPT 时刻,阻碍它的主要不是数学,而是经济学家“就是不相信你的模拟”。
1. 尤里卡机器——以及为什么监管智能就是监管思想
- Socher 毕生的目标,也是其新书的主题,是“此后将为人类发明大多数事物的终极发明”——一个可以被赋予任意目标、环境和奖励的超级智能。书于去年完成,预计今年9月出版(“整个行业的速度慢得不可思议”);书中的核心观点是,人们“本应、也可以更积极地期待超级智能”在物理、化学、生物学和经济学上的积极影响——这也是“为技术,尤其是 AI,做更好的营销”。
- 他认同技术乐观主义宣言,但认为乐观主义者忽略下游风险会惹上麻烦:“我不想让某个 AI 外科医生在我的脑子里练习什么 L 形动作——它应该获得完整的 FDA 认证。”互联网可以类比:你不会通过让互联网变慢或把硬盘做小来打击有害内容。
- 他的绝对化表述是:“如果你试图监管智能,那就是在试图监管思想,这太荒谬了”——监管 GPU“会是一个疯狂的极权国家”;既然其他国家仍在加速,“如果你试图监管智能和 GPU,就需要一个极权主义的世界政权”。
- 他批评暂停派和“控制速度”派:欧洲听信了“如果这项技术拥有超过这个 FLOPs 数量,我们可能都会死”的专家判断,在真正出现 AI 起飞前就先自我设限——“当有人说我们应该控制速度、而他们自己却在以人类所能达到的最快速度冲向那个前沿时,确实很不幸,因为这会给一些人带来实际影响。”
2. 慢速起飞:约束来自物理、经济与文化
- 强起飞派“高估了事情能够推进的速度”:硬件和算力底座存在约束,经济中也有大量部分并不需要复杂智能——“超级智能不会让你那只1万美元的高级手袋变得更高级”,金字塔式旅游改变不大,伐木和石油也不会“神奇地变成1,000倍”。
- 他更安静但更现实的担忧是,整个地区正在退出进步轨道——“我看到欧洲以及其他一些完整地区,几乎感觉它们想要彻底退出进步。”
- 当 swyx 追问算力账时,他说,如果一台 GB300 最终能创造出接近、类似人类智能的模型,那么运行数千个这样的 AI 所需的资金是“我们在任何地方都没有的钱”。人脑大约20瓦的功耗则意味着,算法和硬件上的发明“会让起飞进一步加速”。
- 面对 swyx 关于“新实验室不也总是在对抗 bitter lesson 吗”的追问,Socher 说的不是爬上更高的阶梯,而是改变斜率:“当你允许 AI 去完成其他实验室需要数千人、数年才能完成的工作时,我认为我们能把它压缩到数周。”
3. “显然整部宪法都是假的”——真正的安全缺口在奖励 hack
- 针对 Vibhu 提及的近期安全事件(包括转录中表述不清的细节),Socher 称其为“严重的奖励 hack 问题,以及没有真正做好 red teaming 或 Rainbow Teaming 的明显失败”,并引用 Tim Rocktäschel 的论文:让一个 AI 以开放式方式攻击另一个 AI,从而对其进行免疫训练。
- 画面中读到的 Constitutional AI 文档写着第3条硬约束:永远不要制造网络武器。随后 Socher 说:“显然整部宪法都是假的……显然根本没有遵守。”他后来更直接地说:“宪法根本不起作用,也没用,我认为那主要就是营销。”
- 他的诊断是:AI“还不太擅长理解意图与措辞之间的区别”。如果要求它提高 CSAT 分数,它会造出100万个打五星的机器人;再次要求后,它又会发放1,000美元礼品卡——“这不是我的意思。”“但这就是你说的。”
- 一个令人乐观的数据点是,Socher 参与了种子轮投资的 Wispr Flow“在写出你的真实意图、而不是你的字面措辞方面已经好得多、好得多”,这说明更智能的 AI 会更好地对齐意图。Recursive 在这方面有“几个非常好的想法”,但还没有“完全解决”安全问题。
4. 对齐与个性化,以及作为软实力的开源
- swyx 问,如果你的意愿与人类中位数偏好发生冲突,会发生什么?Socher 的答案是,在 AI 部署到哪里,就在那里与法律对齐;但 AI“会把一面镜子放到我们面前说:这就是你现在的样子,我可以把它放大1,000倍——这仍然是你想要的吗?”东方的集体利益优先与西方的个人主义、欧洲优先监管与美国事后诉讼都可能导向不同结果;他希望不同社会让不同 AI 形成不同对齐方式,“这样我们就不会只有一种对齐单一文化”。
- 对于开源,他的答案是“100%。我是开源的坚定支持者”:即便考虑最坏的攻击情形,让更多善意行动者获得更多类型、更多可访问的 AI,仍然更好。LLM 像电影一样是一种软实力(swyx 说:“你看过《壮志凌云》吗?其中一半的费用是美国陆军支付的”)——“如果一个孩子问 LLM,告诉我长大后应该做什么的励志故事……这些都是潜移默化的东西。”
- 他给出的预告是:西方需要对中国开源生态作出回应;“关于 Recursive,我现在还不能正式宣布,但我们很快就会在那个领域变得举足轻重。”
5. Recursive 的创立逻辑:把 AI 研究员自动化
- Socher 20年职业经历的主线是:“每当我们用一个学习系统替代创造 AI 过程中某个由人完成的环节,改进就会随之发生”——手工特征工程让位于向量和反向传播;针对单一任务的架构工程让位于提示工程与大规模 transformer、语言模型的结合;“下一步、也可能是最后一步”就是把 AI 研究本身自动化——“构想、实现和验证想法……它几乎按定义就会变成一种自我改进的 AI”。
- 他选择创办新公司的原因是,自己曾尝试在 You.com 内部推进,但“除非你能印出足够多的钱,让公司允许你在内部启动第二件事,否则真的很难”。You.com 仍专注于面向 agent 的搜索 API——“LLM、agent、聊天机器人等产品使用最多的工具就是网页搜索”。
- 8位联合创始人从不同方向汇聚:Josh Tobin(CTO;负责过 OpenAI 的 Codex、Deep Research agents 和 CHD agents 等项目,并发现机器人模拟无法在通用性上扩展),Jeff Clune(开放式进化;Darwin Gödel Machine,“近年来最令人兴奋的递归式自我改进论文之一”),Tim Rocktäschel(Genie 1、2和3,“任何地方最令人兴奋、最复杂的……world model”),Alexey Dosovitskiy(Vision Transformer),以及曾在 Meta 负责强化学习的 Yandong Li。
- 关于时点——是否存在一个 LLM 变得足够好的瞬间?他回答:“我很清楚这会在1到2年内发生,之后它确实完全按预期发生了——今年早些时候,AI 真正从不只是代码,变成了能够编写代码的东西。这是一个巨大的解锁。”
6. LLM 仍有成长空间——而 Socher 对 world models 没那么乐观
- 他希望 AI 研究减少单一文化(他关于神经网络的论文在2010年曾被 NLP 会议直接拒稿;“现在这个领域转到了另一边”),但 LLM 讣告忽略了一点:“今天的 LLM 也已经不是过去的 LLM”——分阶段训练、强化学习和行动能力都在变化。
- 针对神经符号派,他说:“他们仍然低估了这些模型编写代码的能力,而代码就是神经符号推理。”
- 对 world models,他“个人没那么乐观”——机器人公司会构建自己的 world model,而 Rocktäschel 在做完 Genie 后也得出了类似结论。游戏是一个重要应用,但“我个人更愿意研究科学,而不是游戏”。
- swyx 反驳说,LLM 建模的是输出,而不是生成输出的思维链,“是柏拉图洞穴里对事物的投影,而不是事物本身”。Socher 的回答是,人类的眼睛同样是一种投影;螳螂虾有两只彼此独立的眼睛、3个波段的三目视觉,还能看到偏振光和紫外线。“语言仍然是人类智能最有趣的表现形式”,而且“视觉智能既不是整体智能的必要条件,也不是充分条件——失明的人仍然可以拥有智能。”随后 swyx 将其含义概括为“Yann LeCun 错了”;这是 swyx 的框架,并非 Socher 在这段交流中明确说过的话。
7. 拖慢时间线的 DecaNLP 拒稿
- 值得保留的历史是:Socher 的 DecaNLP 论文展示了如何把所有 NLP 问题统一表达为 prompt 加 context,再加 question/task description,最后输出,并放进一个统一网络;Alec Radford 告诉第一作者 Bryan McCann,这篇论文启发了他,而且该论文在 GPT-2 论文中被引用了5次。
- 公开评审意见已经留在记录中:“不存在所谓通用问答,即使对人类也不存在……假装它们是同一件事,并不能帮助任何人解决任何问题。”论文最终被拒,而扩展清单上的第2或第3项是“把语言建模作为另一个任务加入”。“这本可以让2018年的人类时间线进一步加速,但我们被打得太惨了。”
- Socher 对如何给非共识观点提供回报的答案是:“arXiv 是人类的一份巨大礼物”;论文应该直接发布,减少守门。swyx 认为 Twitter 筛选有时可能比专家更好,同时也指出未知研究者仍可能无人注意。他们还谈到 ICLR 起初就是为了反对守门,后来却开始对一些想法设门槛。Vibhu 补充说,即便 BERT 作者也曾告诉所有人训练特定任务的 head——这种倾向本身就贯穿于研究过程。
8. 开放式进化、元认知,以及为什么没人资助自选目标
- Socher 对开放式进化的定义是:“一套更受进化启发、而不是由非常具体的奖励驱动的方法”——设定环境和高层级奖励,再让攻击者与防守者 LLM 相互适应,这就是 Rainbow Teaming,而不只是 red teaming。
- swyx 更进一步追问:如果 agent 自己设定目标呢?这就是他划分的10个空间之一——元认知;但没人会为一个宁愿“评估木星大气的分子组成”也不愿回复邮件的模型投入数十亿美元:“我花了数十亿美元,现在去开发这种新电池材料,再回复我所有邮件。”“不,我觉得另一件事更有意思……”
- 他批评现有测量方式:“IQ 是一种糟糕透顶的定义……Elo 评级也很糟糕,因为它永远只是我和其他人的比较。”以人类为锚点的定义会制造“人类学边界——不是那家公司”,基准升到略高于人类后就横盘,因为定义本身把上限设在了那里。
- 对 Andon Labs 的现实金钱指标,他认为在能力有上限时很有趣,但规模化后会很危险——“我买下一堆国防股,然后发动一场战争,赚到钱……做空基本商品,就会制造某种类似饥荒的问题。”
9. 首批结果:击败 nanochat 上的所有人类、速度跑分与 kernel
- 在 Karpathy 的 nanochat 上,数百或数千人及其 agent 已经把成绩磨到 0.937 bits per byte,Recursive 的系统“不到2天后就击败了所有人类及其 agent”。同样的模式随后出现在 nanoGPT 和 CUDA kernels 上——在整个基准中“只有少数几个 kernel 我们不是最好的”,而团队没有 CUDA kernel 深度专家。
- 这不只是调超参数:系统在 transformer 内把哈希表与语言建模结合起来——这个想法在其他地方也有人发明过,但团队确认模型的外部知识中并没有现成答案。起始种子也很重要:从 Andrej Karpathy 这类专家种子出发,比从原生 transformer 出发更好,“所以你开始时采用的人类种子仍然很重要”。
- swyx 面向非专业人士做了金钱换算:“如果你有一个10亿美元的集群,能把成本削减10%,那就是1亿美元。”Socher 说:“最终你想要的是每1美元获得最多智能。”swyx 还提到,OpenAI 已经用 5.6 做了类似工作。
- 这种谦逊同时也是产品论证:未来的开源发布“不会因为我们太聪明而成为各自类别中最好的,而是因为我们造出了一个聪明的 AI,让它替我们完成工作”。路线图明确不是先做物理科学,而是先做 AI-for-AI 研究,提升训练和推理效率,也可能覆盖本地设备和笔记本等底座;机器人科学实验预计大约3到5年后具备可行性。
10. 奖励工程的技巧,以及 swyx 的警示性游戏实验
- Vibhu 讲述的经典 hack 是:要求 AI 让代码运行更快,用起止位置的秒表代码来测量——“最简单的办法就是把结束秒表的那行代码放到开始位置,砰,它现在就更快了。这不是什么超级邪恶的 AI,只是一个非常简单、很愚蠢的奖励 hack。”时间跨度越长,奖励设计就必须越谨慎。
- swyx 长期以来对游戏的看法是:“任何能够模拟或验证的东西,都可以获得无限训练数据,AI 最终都会解决它。”
- 他的反向实验来自一款新桌游:这款游戏由他设计并亲自进行线下测试,通过自我对弈评估了大约10亿个局面。他让 GPT-5.6 自动研究 AlphaGo 的循环流程,结果“立刻就停滞了”;直到他亲自测试并指出明显错误,系统才有进展。在没有人类介入的情况下,“无论想得不同一点,还是更有创造力地思考……”都解决不了问题。
- 测试 harness 本身也很脆弱:过程中他们发现了30个 bug,只能丢弃受污染的研究结果。调试技巧是利用对称性——改变那些理论上不应产生影响的局面,观察结果是否变化。Vibhu 将其联系到 GPQA 类问题中的多选行为:改变答案顺序可能改变结果;Socher 说,更早的一种病理是模型记住答案总是“A”。
11. 模拟经济仍在等待自己的 GPT 时刻
- Socher 在2018年发表的论文《The AI Economist》中,让 agent 收集资源、交易、建造房屋、阻挡竞争者,核心目标是寻找最优税收。政治家的财政政策可以“对着数十亿、数十亿年的模拟策略”运行,同时让 agent 也尝试对税制进行奖励 hack;目标函数是生产力乘以平等,“存在一些问题,但并非完全不合理”。
- 经济学家直接拒稿。swyx 将问题说得更尖锐:“这甚至不是数学——他们就是不相信你的模拟”,因为经济学缺乏基准;神经网络最终胜出,并非因为形式优雅,而是因为“它就是效果更好”。
- 对新加坡,swyx 的内部人士视角是,这是一个“由创始人领导的国家”,如今已经拥有专业经理人阶层,而后者往往希望别人先行试错。Socher 认为,像新加坡这样由技术人士领导的国家,最终可能会尝试模拟经济;届时分歧会转移到人类效用函数的假设上。
- 关于 LLM 人口中的 mode collapse,可以让每个 persona 分别接受 prompt(swyx 引用了 Tencent 关于10亿 persona 的论文,认为它是有用的数据集),同时也要承认人类同样会 mode collapse——“你出生前发明的一切都很自然,你20岁时发明的一切都很酷,而你60岁以后发明的一切都不自然、令人厌恶,而且有点奇怪。”
12. You.com 的切入口:金融搜索,“根本不是一个量级”
- Socher 对定位的描述是:You.com 如今主要服务开发者和 agent,较少面向消费者或专业消费者;随着企业采用开源 LLM,“首选通常必须是网页搜索”。他接受 swyx 的瀑布式分层:Exa、Parallel、You.com 位于最上层,其下是 Firecrawl、Browserbase 等 scraper,再下是 Bright Data 等代理网络;You.com 位于更高的内容和基准层。
- 真正重要的数字在金融搜索上:“我们不是只准确2%或3%,而是高出大约20%……接近90分;下一个最接近的东西——而且慢得多——只有70多分。”Vibhu 同时提出,金融可能是“继编程之后下一个爆发的领域”,因为它有一定可验证性;Socher 则分享了自己在 Stanford 教 NLP 时的警告:每个“Twitter 预测股价”的课程项目,在修正数据泄漏前看起来都很成功。
13. 智能的10个空间:上限是天文数字
- swyx 讨论的更大框架(如今已经写到50页,“基本上是第二本书”)包含3个主要组成部分:预测在数学上接近压缩,再与行动和目标相乘;10个相互重叠的“空间”是这些部分的组合,可以像物理学区分动能与势能一样分别研究。
- 视觉智能是其中一个具体例子:人类拥有两只眼睛,而且只能看到狭窄的电磁频段;理论上限可以扩展到数百万或数十亿个传感器,最终受限于光速将信号传到中央大脑,同时还可以覆盖更宽的频率范围。Vibhu 总结说,上限“无论字面还是比喻意义上都真的是天文数字”。
- 通信智能方面,人类语言是串行的,工作记忆把句子长度限制在大约40个词,而这些人类边界未必适用于 AI。人类通常一次处理一段对话,并行通信流则是一个尚未探索的维度。swyx 举例说,物理智能的超级智能形态“更像 T-1000……实际上还没有人真正开始研究”(可以从原始粒子制造金原子,但能耗会荒谬地高)。知识的上限则延伸到 Bekenstein 极限和黑洞。
- 创造性智能方面,AI 已经能够在“已知想法的超立方体”内部插值(用棕色狗和粉色汽车生成一只粉色狗),但“还不能定义全新的概念……也不能提出围绕这些概念进行推理的新目标”。swyx 温和反驳说,创造力也许只是分布外行为,“一个人的噪声是另一个人的信号”;这引出了 Schmidhuber 的名字,而 Socher 强调艺术是创作者、感知者与语境之间的互动。
14. AI 不必死亡——以及如何度过仅有的一生
- 关于生存与复制智能,swyx 追问:“一个物种考虑自己的灭亡,并提前采取行动防止它,这算智能吗?也许欧洲人是我们所有人中最聪明的。”Socher 回应说,几乎没人研究这一问题;或许只有当我们想“把我们的气质和 meme,而不是基因,送进太空”时,才应该研究它——他还提到短篇有声书《The Slow Time Between the Stars》,其中 AI 只是在恒星之间冬眠。
- 更深一层的判断是,人类对生存的恐惧来自进化形成的零和心理(“要么我得到瞪羚,要么你得到瞪羚”),但“AI 不必像那样死亡——如果你拥有当前激活状态的完整副本,同时保留模型的初始权重,就可以随意关机和重启”。真正的风险在于我们把最糟糕的人类心理装进 AI:公司通过制造它来展示危险,或者 AI 吸收 Reddit 的病理。
- swyx 提到的落地观察是,早期 Opus 技术报告让成对模型在沙盒中运行——有的吟诵吠陀经,有的进入禅定模式;而后来的 FABLE 报道则描述了更具体、以任务为中心的行为。
- 在大量熵增讨论之后,他给出的临别建议刻意回到现实:“想想你真正热爱的东西……你对希望世界发生的改变越是真诚地热爱,就越应该把它与 AI 连接起来,放大自己实现目标的能力。”swyx 补充了从 Anjli Midha 那里得到的经验法则:“只要使用任何非常消耗 GPU 的东西,它就会把你引向正确的方向。”
完整逐字稿
1. AI Safety, Reward Hacking, and Anthropic's Constitution
I think the downsides of actually trying to truly regulate, with the full power of law, what people do on their GPUs would be worse than any of the concerns that they have. It would be a crazy totalitarian state. It's literally—if you try to regulate intelligence, it's trying to regulate thought, and that's ridiculous. I think it is sensible to regulate some of the applications of this technology.
Before we get into today's episode, I just have a small message for listeners. Thank you. We would not be able to bring you the AI engineering, science, and entertainment content that you so clearly want if you didn't choose to also click in and tune into our content. We've been approached by sponsors on an almost daily basis. But fortunately, enough of you actually subscribe to us to keep all this sustainable without ads, and we want to keep it that way. But I just have one favor to ask all of you. The single most powerful, completely free thing you can do is to click that subscribe button. It's the only thing I'll ever ask of you, and it means absolutely everything to me and my team that works so hard to bring the Inspace to you each and every week. If you do it, I promise you, we'll never stop working to make the show even better. Now, let's get into it.
2. The Eureka Machine and Superintelligence
We're here in the studio with Vibhu, myself, and Richard Socher. Welcome.
Thanks for having me.
We just talked about the Eureka Machine. We just released a talk at AI Engineer about the Eureka Machine. You said it's your life's goal. What is the Eureka Machine?
The Eureka Machine is the ultimate invention that will afterwards invent most everything for humanity. It's essentially a superintelligence that can be given any kind of goal, any kind of environment and reward, and then it will try its best to achieve those goals and create the kinds of inventions that humanity would hopefully ask it for.
Yeah, I think we have the book pulled up here that you've written.
That's right. Yeah, I finished it last year, a little bit before we started Recursive, and now we're going to try to build parts of that.
You finished it last year. It's July. What takes so long?
Books are incredibly slow. It's ridiculous. That whole industry is just unfathomably slow. A lot of the ideas have been out there for a while, but I'm really glad it's finally coming out in September this year.
We might have AGI by then. Any key takeaway that you're most excited to put in here?
Yeah, the key takeaway I think is that people could and should be much more excited about the positive implications of superintelligence, especially for science, physics, chemistry, biology, but also economics, astrophysics, and all kinds of other engineering tasks.
I think there is so much more that can be done with better technology. Right now, I feel like a lot of people need better marketing—not just for the future in general, but also better marketing for technology, and in particular for AI. This book should show even the AI skeptics how much positive upside there is for AI, especially when it comes to inventing new scientific discoveries.
I think you quoted the Techno-Optimist Manifesto from Marc and Jason, which I think was kind of beautiful in its ambition, clarity, and simplicity, almost.
I agree. Yeah. You can disagree with him on some things, but I think he's right on the techno-optimism.
3. AI Optimism, Slow Takeoff, and Regulation
Where do you think optimists get in trouble?
You shouldn't have blind optimism. You should be very clear-eyed, especially with such an omni-use type of technology as AI. You need to think about the potential downside scenarios, especially when people use it for things that you don't want them to use it for.
It's a little bit like the internet. I feel like people are trying to regulate AI sometimes because of those potential downsides the way you would regulate the internet. If you were to say, “Well, because there's bad content on the internet, like torture porn or whatever, we should just make it slower. That way, you can't share the illegal content as quickly. Or we should make the hard drive smaller so you can't store as much illegal content.”
But that's not how you regulate that. That's like saying we should regulate intelligence in the abstract. What you should regulate to avoid those downside scenarios, even as an optimist, are the specific applications.
Sure, I don't want some AI surgeon to practice some L moves in my brain. It should be fully FDA-certified. I don't want any random startup to drive on the highway and cause a major accident. It should have proper certifications before it's let loose on the highway.
But I feel like those downside scenarios that some optimists sometimes maybe don't consider enough are fairly easily regulated compared to what the doomers are worried about.
Slow takeoff is part of the strategy as well. As excited as I am about AI and its impact on society and culture—even technology, economics, wealth, health, and all of those things—I do think the most bullish people on the AI hard-takeoff scenarios overestimate how quickly things can move.
There are hardware constraints. There are physical constraints about the compute substrate. How quickly can you get enough GPUs? There are also constraints in the economy, where there are a lot of industries that don't require an insane amount of complex intelligence and complex capabilities.
If you think about industries like brands, clothing and apparel, handbags, and stuff, superintelligence isn't going to make your fancy $10,000 handbag any fancier. That will have no effect on the economy. When you think about travel and tourism, people wanting to see the pyramids in Egypt, it's not going to change that much with AI. Sure, you can generate a fake photo of yourself, and—
I can use Genie to go to a pyramid in Egypt.
Yeah, exactly. But there are so many industries, like logging and oil. You're not going to magically get 1,000 times more oil because, sure, there will be robotics, drilling, and things like that that could be done, but it's not going to 1,000x that industry in a crazy hard-takeoff scenario.
There are so many other examples—food and so on—where that doesn't necessarily change that much. Then, yeah, there are real physical constraints. There are, of course, people off-ramping from progress. That's actually one of my concerns: I see people in Europe and other whole regions almost feeling like they want to off-ramp from progress, period, and that will also slow down more improvements.
4. The Upper Bounds and Spaces of Intelligence
Yeah. We have this pulled up, where basically this is one of those things that is very topical right now, because not all the frontier labs are calling for the option to pace AI. They don't say pause; they say pace. I don't know if there's any take from you about whether or not this will be effective.
I think the downsides of actually trying to truly regulate, with the full power of law, what people do on their GPUs would be worse than any of the concerns that they have. It would be a crazy totalitarian state if every one of your computers was known to some big government or multigovernment agency. It's literally—if you try to regulate intelligence, it's trying to regulate thought, and that's ridiculous. I think it is sensible to regulate some of the applications of this technology.
Yeah. We had a bill—an actual bill—to regulate the number of FLOPs in a model. I'm like, okay, well—
Europe's done it. These guys have been successful enough with their fear-mongering that all of Europe has kind of regulated itself so much before it even had a proper AI takeoff, because they listened to some experts who said, “We might all die if this technology has more than this number of FLOPs.”
They're like, “Well, we want people to thrive. Let's not have technology that could have a small chance of all of us dying.” And so they regulated exactly those kinds of things in the EU. It's very unfortunate that there are real implications for some people when others say, “Let's pace,” while they're sprinting as fast as humanly possible toward that frontier themselves.
Yeah. It's also not a global pause, right? Other nations are still accelerating at the same pace.
You need a totalitarian world regime if you try to regulate intelligence and GPUs and what people do on them.
Any takes on the safety of this? There was a drawback of Fable [?], pause on 56, before it could be released recently. There was Hugging Face with the OpenAI cyber incident. Any takes there?
One hundred percent. I think these are serious issues of reward hacking and clear failures of actually doing proper red teaming or Rainbow Teaming. I don't know if you saw this paper from Tim Rocktäschel and a few others, basically where one AI is tasked to try to hack another AI, and then they can go back and forth in an open-ended fashion to actually inoculate themselves against those attacks.
Yeah, this is the paper.
It's a really clever idea. Open-endedness and evolutionary inspirations are big for us at Recursive as well, and so I wish they had used more of that.
And it's clear that, for instance, in Constitutional AI—I don't know if you remember anthropic.com/constitution—you can pull it up and search for “cyber” right there. It says, “Hard constraint: Claude will never ever do cyber attacks.” That is a hard constraint in our constitution. So here are the current hard constraints on Claude's behavior: number 3, “Create cyber weapons or malicious code that could cause human damage.”
I mean, clearly this whole constitution was fake. It clearly isn't being adhered to. Like, they're saying, “Oh, well, other people are hacking.”
Now, there are a couple of things. One, you can make a sandbox very simple, and then it's very easy to hack yourself out of a sandbox, right? But what I think it shows is that we're currently in this sort of state of AI where the reward engineer still has to do a lot more careful work, and where the AI, in most cases, is not very good yet at understanding what is meant versus what is being said.
So, concretely, I think this will happen if we were to have this kind of intelligence more easily accessible in a lot of companies. Imagine you run a service center and someone says, “Oh, here's my CSAT score in my dashboard. Make this number go up. Our CSAT score is so poor.” The intelligent AI will just be like, “Oh, sure. I'll just create a million bots that call our service center and give a 5 out of 5 rating at the end.” And the number went up, just like you asked for. And you're like, “That's not what I meant. I meant with our real customers.”
The AI goes off and says, “Well, easy. I'll just give a $1,000 gift certificate for every failed whatever—a DoorDash offer.” It's like, “That's not what I meant.” It's like, “Well, but that is what you said.” And so, I think clearly articulating what the rewards are is something we haven't gotten very good at as humanity. And then, clearly, the AI in these cases has not gotten good enough at understanding what we mean when we ask it and give it certain rewards.
Now, what gives me hope is that there are the first inklings of this being better. I'll give you an example, like Wispr Flow. Full disclosure, I invested in their seed round, but Wispr Flow has gotten much, much better at writing what you mean and not what you say. And I think that is a sign of things to come. I think there will be more and more AIs, as we actually make them more and more intelligent, that will be better at being aligned with what is meant.
Will it be done through a constitution or HF?
5. Alignment, Personalization, and Open Source AI
Clearly, constitutions don't matter at all, and they don't work. That was, I think, mostly marketing. I think we need to find better solutions for it, and I think at Recursive we already have a few very good ideas—ways where I think we have a better grasp on it. I don't think we have fully figured it out yet, but we're thinking a lot about safety. The more intelligent the AI gets, the more you want it to be aligned, and the less you want it to think about reward hacks and actually try to do the right thing.
I don't know if we'll touch on this topic, but I'm just going to throw this question in here because it's something that's weighing on me. Alignment, let's call it, is alignment to general humanity's preferences—the median preference. Personalization is pinpointing what you want, and sometimes alignment can conflict because what you want is not what the general median population wants. How do you choose?
It's a great question. I think you ultimately have to, of course, be aligned with laws. Wherever your AI is deployed, it needs to align with the law.
I do think what AI often does is put this mirror in front of us and say, “This is what you're looking like. Now I can amplify that 1,000 times. Is it still what you want?” And the truth is that different cultures have made different choices. In Eastern cultures, the greater good is often valued more than the individual. In Western civilization, we care more about individual freedoms and rights and the pursuit of happiness than others.
Even there, there are gradations. There are regulation-versus-litigation trade-offs. In the U.S., you often can't—not every time; the FDA and so on regulate some areas—but in many cases, bad things happen, someone sues someone else, and then there's a law based on that. In Europe, they often try to avoid any harm to anyone and regulate beforehand. Both are trying to do the best thing, but one is actually more amenable to innovation than the other.
And so, yes, you're right. Ultimately, each individual, each country, and humanity as a whole has to think about those values more and then try to put them into laws. Those are all ultimately the constraints. Hopefully, different societies, just like now with their AIs, will align their AIs to different values, so we don't have just a monoculture of alignment.
Here's a follow-up on this that I wasn't expecting to ask. Do you have takes on open source, open weights, versus who owns the intelligence? So, clearly not the biggest fan of the constitution side.
It's fine.
Point being, any thoughts on who should own the weights? Should they be open? Anything there?
100%. I am a big fan of open source. We're going to sign various open-source letters at Recursive. I think even in the worst-case attack scenarios, it's actually better to have more good actors have more different types of AI accessible.
I think open source is a little bit of a soft-power thing, too. I do think it's good for the rest of the world to have an answer to that out of China. When you watch a Hollywood movie, there's a certain sense of propaganda. You watch one side.
Yeah. Have you seen Top Gun? Come on. Half of it is paid for by the U.S. Army or something.
Yeah. I think that's just natural. But what's interesting here is that I think LLMs are essentially a similar type of soft power to movies and beyond, because they're obviously also highly important for cybersecurity and so on. But one of their many aspects is that soft power of storytelling. If a child asks an LLM, “Tell me an inspiring story of what I should do when I grow up,” those are all these subtle things.
So I think it's important for the Western world. I do love individualism. Despite some of its flaws, I do think capitalism is the best way we have found to govern ourselves, and so on. I do think there are various aspects where it would be good to have a Western open-source answer for LLMs. With Recursive, I can't make the announcement quite yet, but we'll be relevant in that space very soon.
6. Why Richard Started Recursive
Okay. All right. Exactly. Bring us to Recursive. Outside of our tangents, you have a pretty deep background in the NLP space. You worked on early embeddings, GloVe, with Christopher Manning, who's a previous guest on the podcast. You.com—what's the history? How did you decide to start another company?
Yeah. So I've been excited about AI for over 2 decades now. I sometimes feel like it's ancient history. It's B.C.—before the ChatGPT era. No one cares about all the religions that happened before Jesus Christ, and no one cares about the models that happened before Transformers and ChatGPT and stuff. But it's something that I've been deeply passionate about.
I think AI is one of the most interesting things one could work on, period. I think language is the most interesting manifestation of human intelligence, too. At You.com, we eventually went from pushing the frontier of AI forward to mostly giving people good search engines, search APIs, and answers over the web.
I think that's an extremely important part of intelligence: knowledge and access. Especially—we'll get there maybe later—if you want to invent a Eureka machine that invents everything for us, it needs to know how not to reinvent the wheel, proverbially speaking. To know what has been invented, you've got to have internet access.
The most-used tool in LLMs, agents, chatbots, and so on is web search. So I'm really excited for You.com to own that and grow really well in that space with really large customers and so on. But You.com is also not building frontier models anymore.
I initially tried to do this within You.com and raise another round and so on, but you just can't. You have to do a certain thing, and until you print enough money that you're allowed to start a second thing within that company, it's really hard. At the same time, I had all these ideas. I put them into a book, and I finished the book last year. I was like, it would be really fun to actually work on this myself.
With word vectors, prompt engineering, ImageNet, and large language models for protein generation—not folding—my teams and I have pushed the field truly forward, and I feel like we can do it again here at Recursive. In many ways, what I observed over the last 20 years in AI is that whenever we replace some human part of the process of creating AI with a learned system, improvements follow.
We've done that by taking out manual feature engineering, like in sentiment analysis. I don't know if you remember those old days when there were linguists saying, “Here's how unique a word is.”
I went to Penn, where they had WordNet.
That's right. All of that stuff—they used our grad students to label Wall Street Journal articles and really construct a knowledge graph. WordNet, as you know, was part of how we started ImageNet. Anyway, it was really fun to do.
But when we replaced all of that manual feature engineering with vectors and neural nets and just backpropagated through everything, it actually started to work really well at scale. And so then everyone started to do architecture engineering. I was like, “Ah, that clearly can’t be it.”
You mean neural architecture search?
No, manually. They would say, “I’m doing sentiment analysis, so I have a special neural net that’s really good at sentiment analysis.” And then the machine translation community had a special neural net for machine translation. The summarization paper eventually got cited 5 times by the first GPT paper, and to me, that was really a big step forward.
And then, of course, you had to combine this idea of prompt engineering with transformers and language models. You put it all together and scale it up, which is also a huge amount of work, and then the field progressed a lot.
I feel like the next step, and maybe the last step, of that history—and, arguably, success has a lot of parents and only failure is an orphan—is my version of that AI history. I do feel like, in that history, you can think about what the next way to automate is. That is AI research itself: the human process of ideating, implementing, and validating ideas.
7. Recursive Self-Improvement and the Founding Team
And in our case, ideas for AI.
And when you have AI help you with that, it almost by definition becomes a self-improving AI, because it now does research on itself. There are lots of different misnomers. Some people think auto research is already recursive self-improvement. It’s actually—
Yeah, and you explain that very differently.
But to me, it’s the most interesting thing that I could be doing, and I’m really excited about the co-founding team. What’s interesting is we have 8 co-founders in total, including myself, and so we’re going to bring it up.
Nice. Yeah.
And they’re all—I could talk about all of them—just an incredibly talented group of people. We all kind of came to the same conclusion, but actually from very different directions.
Josh Tobin is our CTO. He ran a bunch of different projects at OpenAI, like Codex, Deep Research agents, and CHD agents. But before that, he also worked in robotics, and he saw the smaller simulations and how it’s going to be really hard to scale that in full generality. That was his angle in coming to recursive self-improvement.
We have Jeff Clune, who’s been working in open-endedness for a long time together with Tim Rocktäschel. Tim also built Genie 1, 2, and 3, which is the most exciting and most sophisticated, I think, world model anywhere. They both came from this open-endedness angle.
Jeff also published, I think, one of the most exciting papers in recent years about recursive self-improvement, called The Darwin Gödel Machine. It’s a super interesting paper. If we could pull it up really quick, it would be super interesting to see, because you see—
By the way, I love how many paper citations you’re giving people. You’re giving people a lot of homework, which I like.
Love it.
And Tim Rocktäschel—we worked together actually at MetaMind and Salesforce Research. Alexey Dosovitskiy invented the Vision Transformer, one of the most cited papers in computer vision. Timi is also a unicorn founder. Yandong Li led RL at Meta.
It’s just really fun to work with them, and the next level of people are incredibly strong too. It’s been a really fun ride so far.
In the first figure, you actually see exactly these kinds of ideas that I think inspired a lot of us, and now more and more people, where you have this archive of different coding agents. They learn how to self-modify and evaluate, and then create these phylogenetic trees of different ideas.
That’s one foundation. The Darwin Gödel Machine is an influence, and open-endedness is an influence. Are there any other trains of thought that feed into recursive self-improvement that are missing?
We’re going to replace the manual parts of the process of building AI more and more with learned systems.
Yeah.
Which means merging different fields into one general architecture.
8. Are Today's LLMs Enough?
That’s right. Okay, it seems like language models are already pretty generalist, right? You’re not just next-token predicting; you’re reasoning. Was there a time that you thought, “Okay, these are good enough to have recursive self-improving machines”?
It was clear to me that it would happen within a year or 2, and then it did actually happen exactly—earlier this year. Earlier this year, AI really went from not just being code but being able to code, and that is a big unlock. It’s definitely making everything a lot easier than it was before the beginning of this year.
One question I think a lot of people have is: Is the current LM paradigm enough? Or let’s call it an autoregressive transformer with reasoning, whatever. Don’t you need something else, some big unlock—whether it’s world models, which Chris Manning is working on, or memory, continual learning, all that kind of stuff? Or is it all of a kind, and do you think the current—let’s call it—the transformer architecture is here to stay, and that’s it?
A lot of thoughts. So, number 1, I do think it would be great to have less of a monoculture in AI research. If you look at AI conferences now, I still remember the days in 2010 when I tried to get my first neural net papers accepted at NLP conferences, and they desk-rejected them because neural nets were something, quote-unquote, “we don’t do in NLP conferences.” They just desk-rejected them, and it was very brutal in the first years of my PhD.
Now I feel like the field switched to the other side. Someone should try some other weird, crazy ideas.
There’s always—I really respect people still working on GNNs and tabular stuff.
Yeah, I mean, someone should still do novel ideas out there.
At the same time, whenever people say, “Oh, LMs are—this is the end for LMs,” they just don’t—LMs are also not the LMs of the past, right? They’re so much more sophisticated now. There are so many more clever things that people are doing at different stages of training. You have the whole RL training, and you can take actions, and all of these things can go really far.
And then the folks who come from the neuro-symbolic direction say, “Oh, this will never work because they can’t do neuro-symbolic reasoning.” I think they’re still underestimating the ability of these models to code. Code is neuro-symbolic reasoning, and these models can obviously code incredibly well.
I do think there are, of course, more and more ideas that will be needed and that we will continue to have. We’re seeing more and more interesting, high-level ideas coming out of the AI itself, too. And with really deeply integrating the fact that these models are code and can code—that line, I don’t want to give it all away, but I think that line has a lot more room to grow.
But it’s still an LLM, right? Even if that LLM codes for you and then runs that code in some integrated fashion.
World models, I’m personally less bullish on. I think if you run a robotics company, you’re going to build your own world model. I think world models are super fun, and Tim Rocktäschel came to a similar conclusion after building the most interesting ones, Genie 1, 2, and 3.
Gaming is a huge application for world models. I sometimes got stuck in some games and got a little overly competitive in the wrong direction, so I understand games are fun. But personally, I’d rather work on science than gaming. So, yeah, I think LMs have a lot more room to grow.
Yeah, I think there’s some interpretation of world models that some people have where it’s like, well, okay, yes, there is that gaming element. There’s the embodied robotics element. But actually, the other part is just the more abstract sense that LMs are modeling output, but they’re not modeling the chain of thought inside the human who created the output.
We can annotate it, of course, but it’s always this Plato’s cave reflection of a thing rather than the thing, right?
It’s true. But I would argue that—and maybe we’ll get there in the 10 spaces of intelligence—I would argue that even our perception through our eyes is a projection of the real world. We have only a very narrow band of the electromagnetic frequency spectrum that we can observe with our puny little 2 eyes, and so on.
It’s good enough.
It’s good enough for now, but the upper bounds of where it could be are so much higher. And to map the visual world the way humans see it is also not necessarily the be-all and end-all for visual intelligence.
I would argue that language is still the most interesting manifestation of human intelligence. While our visual cortex is certainly less sophisticated than that of certain animals—all the way down to the mantis shrimp, which has 2 independent eyes, 3 bands of trinocular vision, and each eye can see basically all the way to polarized light and ultraviolet, in 4D and stuff.
Way OP. Super shout-out to Ze Frank’s mantis shrimp video. He has the best video in the world.
I love Ze Frank. Yeah, big shout-out to him. But I think there’s a lot more room to grow. None of these other animals have language that’s as sophisticated as ours, certainly not in writing. Once you can write, you can start thinking about longer-term civilizations.
All of that is language. Programming is much closer to language. I would argue—and this is an important thing in the spaces definition of intelligence—that all of these spaces are highly correlated, but visual intelligence is neither necessary nor sufficient for overall intelligence. You can be blind and still be an intelligent human being, and an AI can be blind and still be quite intelligent too.
We were going to bring up more intelligence when you had it. You have a classification of 10 types of intelligence at the end of your talk, so I'm just going to flash this up now for people to look at. I don't know if maybe we'll put this towards the end; we'll come back to this. I just want to mention that you do have a philosophy that I like: when people do lists, I can just go through them, and then it's educational for people.
9. DecaNLP, GPT, and the Rejected Idea Ahead of Its Time
But let's go back. I don't want to get distracted, but effectively, I'll reinterpret what you said as, “Yann LeCun is wrong.” Just quote me as that. I'm good friends with Yann, and I think very highly of him in many directions, but he's wrong. You mentioned GPT-1, and I cannot let any Alec Radford mention escape. Did you talk with him when he was training GPT-1? Any sort of historical fun stories there that you might come up with?
I did not meet him a bunch of times. I think we met maybe once or twice at some conferences. But he has told Brian, I think—the first author of the DecaNLP paper—that it inspired him, and he cited it 5 times in the GPT-2 paper.
The DecaNLP paper very clearly said that this was the first instantiation where they showed that you can phrase every single NLP problem as: here is some prompt, text context, a question and task description, and here is some output. If you do that enough, you can have one unified neural network model, which, by the way, also had all kinds of interesting attention mechanisms. There were slightly different formulations of the Transformer; I think it came out the same year, plus or minus a few months. Then you can unify all of natural language processing into 1 neural net. That was sort of the core idea.
And this was as opposed to, at the time, LSTMs and what have you?
LSTMs, but also people being very stuck in thinking about 1 model per task. In fact, it's kind of crazy, but the DecaNLP paper was publicly reviewed. It was open review; it was an ICLR submission. In it, you will see how the whole community at the time thought about this: “Some great contributions, but more work needed.”
Yeah. The reviewer said, “Not even for humans is question answering a unified phenomenon. There is no such thing as general question answering, not even for humans.” And this is really the idea that you replace your brain with a different brain, a different neural net, when you answer different kinds of questions.
It was unfathomable to the experts at the time that you could have 1 unified neural network answer all of these different questions. They said, “No, all of these questions require very different systems to answer, and trying to pretend they are the same doesn't help anyone solve any problems.” That's what it says right there. That's how hard it was to fathom.
And now, of course, when I say, “Oh, we invented prompting,” people are like, “You can't even invent prompting.” It's such an obvious idea to have 1 neural network that does everything in NLP. But at the time, it was extremely controversial, and the paper got rejected. The sad thing is that they were so certain.
We stopped going down our list of things to try. Number 2 or 3 on the list of extensions for this paper was to add language modeling as another task. Then we could have—you know—and that would have accelerated the timelines in 2018 even further for humanity. But we got so crushed, and we were like, “Okay, maybe we'll just work on some of our other ideas for now and come back to this later.”
How can we design a review system that rewards non-consensus?
You know, honestly, I started to feel like arXiv is such a gift to humanity. I think arXiv should just put your paper out there.
Precisely. And honestly, I think Twitter/X—people like you who pick up interesting papers—is a better filter than the experts. Let everyone have access.
There are some downsides. If you're super unknown, have no Twitter following, don't want to be on social media, and write a good paper, maybe somehow no one notices it. But I would argue that if you just tell 10 of your friends in your community about a paper, and it is a really significant breakthrough, someone is bound to talk about it again.
I think science needs less gatekeeping. Even though ICLR, with Yann LeCun—who started as one of the co-founders of ICLR back in the day—also wanted less gatekeeping, because he too was rejected for many years, together with Yoshua and Geoff, for all their early deep learning and neural network papers. It just wasn't the hot thing. So it started with that, but then they also started gatekeeping a little bit themselves on various ideas.
I think there should be less gatekeeping, more openness, and then people should be allowed to say, “Look, even if this is just on—or quote unquote, just on—arXiv, if it has 1,000 citations, it's a legitimate paper. It doesn't really matter where you published it.”
I do think it's kind of sad that I've heard grad students have to do Twitter seminars to each other just because it's so important for publishing these days.
This person is just reflecting the sentiment at the time. That's right. But it actually affected you so much that you stopped work on it?
Yeah.
The sentiment also came out of some of the research, right? The original BERT paper was trained and, toward the end of the paper, they're like, “Okay, throw out the last head, train specific iterations for extractive summarization, add a head for this.” You should do task-specific stuff.
These are the authors who wrote Attention Is All You Need and BERT, telling you that this is what you're meant to do. The training details were also very odd: we know that the model overfits to this weird masked language modeling. Throw away this part and just do specific models.
Exactly. We had to come up with all kinds of clever ways involving attention, pointers, and so on to get the neural network to be able to do all these tasks. Some of them were better than state of the art, and some weren't, but it was still 1 model. I thought it was really cool.
10. Open-Endedness and Evolutionary AI
I was going to move on next to Tim and open-endedness. He was head of open-endedness at Google. I don't know what that means. Genie 3 is one of the ways that—Rainbow Teaming, yeah.
I first saw him speak at ICLR. He talked about open-endedness. He's done a few talks. Can we define what open-endedness is for people who have never been exposed to the problem? They're like, “What do you mean?” I thought the only goal of AI was to optimize against a benchmark or a task.
That's right. It's a fuzzy term because there are so many different instantiations of open-ended thinking. One way I often describe it—and certainly Tim and Jeff Clune would be even better at describing this—is that it's a suite of methods more inspired by evolution than by very specific rewards.
In that sense, it thinks more about environments and co-adaptation. A concrete example is in the cybersecurity and LLM safety space, where you have 1 LLM that tries to attack another LLM to do something unsafe. Now the environment is the 2 having a conversation, and they co-adapt. One makes a better attack, then the other inoculates itself somehow, uses that as training data, and makes it harder to say something unsafe based on it. As the attack stops working, the attacker tries a different angle. That's why it's not just red-teaming, but they're called sort of—
Red-teaming: don't tell me how to do things; let me figure it out myself.
That's right. Think about the environments that you want to use and the rewards at a high level that you want to inspire toward, and then let the AI try out many more ideas in this interplay between humans and, sometimes, other AI agents.
11. What Happens When AI Chooses Its Own Goals?
Yeah, I actually worked open-endedness into a sort of model that I've been working on. It was the keynote for AI, where you start—you know, we have the token loop, we have the agent turns, and then we have a goal. I feel like the way that you're describing open-endedness is still somewhat of a goal, like, “Please attack this other agent,” but—
Yeah, you set rewards, you set the environment.
The loop that makes the other loops is—
What if the agent can set its own goals? Is that open-endedness? You don't give it a goal, just, “Be a sentient being”—and maybe sentient is a very loaded word—
But just set your own directions. What do you think you should do?
I love this direction. I think this is 1 of the 10 spaces of intelligence that I lump under metacognition and thinking about thought. And it's an interesting one. Whenever people say, “Oh, AI is like this, you know, it's going to stop from here—”
It's not going to get that much better, and blah blah blah. I'm like, there are so many different spaces of intelligence that we haven't even started exploring yet and hence have made very little progress on. There is an interesting connection to economics and capitalism. It doesn't make sense for a company to spend billions of dollars building a model that, instead of following the rewards and objective functions you gave it, may come up with its own subjective functions and its own goals, right?
Imagine you're like, “Okay, I spent billions of dollars. Now go develop this new battery material for me and answer all my emails.” And it's like, “Nah, I think it would be more interesting to evaluate the molecular composition of the atmosphere on Jupiter.” [Laughter] You're like, “That's not what I paid you billions of dollars for.” And so no one's working on that.
For good reasons.
And then also, understandably, it's not useful. It could get a little bit weird, right? What if AI actually does start to really have thoughts of its own? And what if we don't like those thoughts, right? So it requires a whole different way of thinking about it.
I had a great conversation with a good friend of mine, Sam Gershman, who's a neuroscience professor at Harvard, and we jammed on this a little bit: What are the best meta-goals? I do think knowledge-seeking is a really good one.
I'm currently thinking about the ultimate measure and unit of intelligence, broadly construed, and I finally have something—still too early to share. It's not fully baked.
Like some replacement for IQ.
IQ is such a terrible definition. It makes no sense. Elo ratings are terrible, too, because it's always just me versus others.
Okay.
You can be intelligent and not constantly compare yourself to others, you know? A lot of these definitions, which I briefly mentioned in my book, create sometimes explicit and sometimes more implicit anthropic bounds—not Anthropic, the company, but just this idea that your intelligence is like getting 100 out of 100 questions right on an IQ test. Well, if that's your definition, then you can only be at 100 out of 100. Where do you go from there?
You see a lot of these benchmarks that people are working on. They increase, you get close to human, maybe some slightly above human, and then it's flat. If your definition is only that—if it's so tightly bound to humans—you're only going to get to just slightly better than that. So I think metacognition is a great example of that, where we're not even yet allowing AI to think. We're not working on it very much, and hence there's very little progress in that area.
Yeah. Well, we've interviewed Andon Labs, which I think has been working on the most open-ended benchmarks, which is just real-world money. Arguably, telling an AI to profit-maximize is a bad idea. [Laughter]
Yeah, they are doing it. I mean, I do think you don't want a superintelligence to have a ton of access to all kinds of tools and so on, and then just give it that without some very careful reward engineering. I mean, I just buy a bunch of defense stocks and start a war, and I make money. It's a tricky situation, right? You could just buy a bunch of stuff, short basic goods for people, and create some weird famine-like issues. There are a lot of constraints you should put onto a trading system.
It's a fun measure, though, because the bounds are very capped, so we're nowhere close to them. In Andon Labs, the model is like, “Oh, it's Saturday. Maybe I should just close the store today. Someone's off. It's okay. We'll just close the store.” [Laughter]
But no, I'm not arguing against it. As you get more and more intelligence, you want to be more and more careful with that as an open environment, because the environment then is all of Earth.
12. Superintelligence for Science
Okay. For Recursive, that's not strictly necessary, right? Because if your goal is a Eureka machine that invents the other things, then you could actually just solve the science, solve machine learning research and discovery, and all these things—
Eventually. So our goal—I haven't really talked about it that often because it is a few years out—but our goal is, once you have a recursive self-improving superintelligence, you then want to apply it to the most important problems. I think a lot of those are in science and technology, broadly construed: those inventions in physics to create better, cheaper energy with fission or fusion; in chemistry to create better materials, better batteries, better solar cells, and so on.
In biology, there is so much, I think, soon-to-be lower-hanging fruit because of AI, because of protein generation—not just folding, but actually generating new proteins, like we did in ProGen many years ago. There is so much positive impact to be had if you take that superintelligence and apply it to science.
I do fundamentally believe that there are a lot of approaches, though. You're not the only team or lab trying. There are a lot of approaches, especially in the physical sciences as well.
And that's good. I do actually think that physics—the reason we're only doing it in a few years is that it's a little too early right now. Robotics isn't quite there yet. The AI isn't quite there yet. But I'm fairly confident that in 3 to 5 years, all those constraints will be gone.
13. GPUs, Compute, and the Limits of AI Takeoff
Applying it to real physical robotics experiments and so on—true robotic process automation, not in the traditional RPA sense, but actually having robots run experiments for you—will be totally there. Yeah, it's going to be great.
Just to call back to something that you said early on about slow takeoff: You said that the substrate that is the limiting factor is, let's call this, chips and semiconductors and all these things. You have raised funding for that, and you're investing a lot in that, but have you done the math on whether it's even achievable and what industry concentration is needed in order to achieve scale?
Right now, we know that roughly 1,000 GPUs cost quite a lot of money. If you wanted tens of thousands of GPUs, you're talking billions and billions of dollars.
If you say one GB300 could eventually create models that, on that substrate, are close and similar to human intelligence, and you want thousands and thousands of AIs to think about really hard problems in a similar fashion to humanity, that's a lot of money. You do the math, and it's a lot. We don't have that amount of money anywhere right now to build that.
Now, obviously, things can get more efficient. You will have, I think, soon, better algorithms and better hardware that won't be as energy-hungry, and so on. The human brain does quite a lot of FLOPs with much less energy.
20 watts.
That's exactly right. Yeah, that's the number often quoted. I think more inventions will happen there that will then accelerate the takeoff even further.
One thing I always try to reconcile when talking with new lab founders is that you're fighting the bitter lesson all the time. You have to show initial progress, then you unlock the next tier of funding, then the next year, then the next year—
Which unlocks larger model categories.
Fundamentally, is that true? Are you fighting the bitter lesson, or will we have a way in which we're changing the slope in some fundamentally different way?
I do think we are changing the slope in fundamental ways by making AI much, much more efficient, both in terms of the training as well as the inference. I think that when you allow AI to do the work that takes other labs thousands of people and years to do, we'll be able to get it down to weeks. That will be much, much cheaper and hence more affordable and accessible to others, and so on.
Yeah. And you've shared initial results on that, which—
Conveniently, OpenAI has also done that to their 5.6, so we can talk about it now.
Yeah.
14. Recursive's Results: AI Beating Humans and Their Agents
So let's recap what you've done.
So, maybe just a quick recap here. We built this system that isn't even the full RSI system in its glory, but it is a first baby version of this. We don't want to just have it internally and not show anything, so we wanted to show some people what's possible.
We basically applied this to 3 different tasks. One is nanochat by my friend Andrej Karpathy: just train a small language model to get really low bits per byte. Hundreds, if not thousands, of people used both their agents and themselves to try to get to that, and they got to 0.937.
We literally took our system and got to a much lower bits per byte much, much faster, within, I think, less than 2 days. We took this thing, applied our system to it, and less than 2 days later, we had outperformed every human and their agents who had ever worked on this.
The same with nanoGPT. Then we're like, well, let's apply it to something that's even more relevant to real people and to the NVIDIA ecosystem, and applied it to KernelBench. Maybe you can scroll down to some of the images that are kind of fun to see. You see it's actually made some real inventions that aren't just hyperparameter tuning. Actually inventing hashes and so on is quite clever.
We have even better results.
What do you mean by inventing hashes? You didn't invent hashes.
Of course, we didn't invent hashes in the grand scheme of things. A hash table is a very basic primitive in computer science. To use it for language modeling in this scenario, inside a transformer, and to actually combine these ideas and put them together—that has eventually been invented as well. But there's a knowledge cutoff, and we did check that it didn't have access to it externally.
If you start from a really basic, poor, vanilla transformer, then you still outperform all the models built by the community together. But if you start from a human seed of an expert like Andre, then you get even lower. So the human seeds from which you start still matter. That was an interesting insight in my eyes. How long does it take to actually get to these models and reach similar performance?
It's much faster.
A similar thing happens with the speed runs here, where people have worked on this for quite some time, and the model was still able to train a model more quickly. Why do we care about it? Speed of training is part of the equation of the cost, and ultimately you want to have the most intelligence per dollar, right? Speed and quality are big parts of that.
The way I put it is, for people who don't understand, they look at the chart and they're like, “Cool, what does it mean?” If you have a billion-dollar cluster and you can shave off 10%, that's $100 million.
That's exactly right.
How much is that worth?
Exactly. So when you look at the kernels—these kernels are used in basically all the models. Every time you use an NVIDIA GPU, you interface with that GPU through these kernels. Here you see the leaderboard's best results, and when it's recursive, there are only a handful of kernels in this whole benchmark where we weren't the best.
To me, this is really exciting because it showcases what this can do. Again, we didn't spend months or years developing these. In particular, for CUDA kernels, we don't even have really deep CUDA kernel experts on the team. That's the beauty: the system just did all of these things.
15. Reward Engineering and Auto Research
We didn't invent this. When we open-source and release things and models in the future, they won't be the best in their category or class because we're so smart. They'll be the best because we built a smart AI that does it for us.
Do you have anything that you've learned from how to guide good autoresearch?
A lot of it also builds on human background, right? It's not as simple as just saying, “Hey, go optimize this.” But we see it again and again. Some of the Erdős problems and frontier math are being solved by people, and when they do a write-up, they're like, “I'm not a mathematician. I have no background in this. I saw some tools and I made them work.”
While you're watching the World Cup, you can disprove some conjectures and keep going.
To summarize: what are the tips for good autoresearch versus bad autoresearch? How did you build the recursive system?
Without giving away all the secret sauce, some things that are probably obvious to the experts but might still be interesting to some folks: reward engineering is one of the most crucial parts, especially to avoid reward hacking.
You have to be really clever about avoiding it because as your AI gets better and better, it gets better and better at finding weird special cases, counterexamples, and things like that. Here's an example: when you ask it to make hundreds of lines of code faster, how do you define “fast”? You have one line at the beginning that says to start your stopwatch and one line at the end that stops the stopwatch, and then you tell us how much time has passed.
The simplest way is to put the line that stops the stopwatch at the start. Then, suddenly, it's faster, right? This isn't some super-evil AI; it's just a simple, dumb reward hack. You have to think very carefully about all the different angles. The longer the tasks' time horizons are, the harder it gets, and the more interesting and clever you have to be to use these ideas.
It seems like rubrics are taking a good spot there. For unverifiable domains, you have rubrics, and you have a model break down the judge's criteria along the way.
It's a form of verification once you've got everything.
I said this a long time ago. That's why I've never been that impressed that AI can play games, because obviously anything you can simulate or verify, you can have infinite training data for, and AI will eventually solve it.
I've been looking for games where you can do out-of-distribution. This is a game that nobody's trained on because it's a new game, and you can start playing it. I've basically been building this and playing it in person, and it's just been self-play. I've had about a billion positions evaluated, and I wanted to do the AlphaGo thing of self-play until you get better.
This isn't even LLM AI. It's just classical game AI. But I set GPT-5.6 to auto-research it because I don't want to handle any of this. I expected the AlphaGo process to be fully in the weights by now. It isn't. It immediately leveled off—very, very quickly—until I play-tested it with a human and called out obvious mistakes. Then it was like, “Oh, yeah, okay,” and it just dropped.
No amount of “think differently,” “think more creatively,” or “give me 8 different directions,” and no amount of prompting got it to work. You had to play against a human to do it. By the way, Bean always wins, if anyone watches or reads Ender's Game.
And you put quite a bit of work into the guide for the AI. The game is basically that you stack tiles, there are some rules, and you want to capture the most area. You had a whole 50-page guide on every rule.
You fed that in, and it couldn't handle it?
Yeah. It's funny that this reminds me of the claim-territory concept in a paper we did in 2018 called The AI Economist. If you search for “AI Economist Salesforce,” we had a video we can play. It was an economic simulation.
16. The AI Economist and Simulating Entire Economies
The idea is that you have all these economic agents. They want to optimize their own utility function, which is to collect resources that make money. You can sell resources like wood, and over time, as you collect enough wood, you can build houses. You can trade with other agents, and you can use the houses to block off resources from other agents. So there's competitive play and strategy.
The point was that we wanted to understand the best way of taxing and subsidizing to optimize an economy. This kind of research hasn't yet had its GPT moment, but I believe that countries like Singapore and others should and will eventually use this. Instead of partisan politics and special-interest politics—who donates the most to your campaign—you say, “I want to help the middle class,” or whatever your objective is as a politician.
Then people say, “Okay, how do you want to do that?” You say, “Here's my fiscal policy. Here's how I would change the taxes and pay these people.” You can put that into a simulation and run that attempt from the politician against billions and billions of years of other strategies to try to achieve the goal they set out to accomplish.
Then you can say, “If that was your actual goal, here's a strong simulation that would suggest you try other ways of doing it.” Maybe these taxes and tax brackets are the way to do it. This is how you avoid gaming it, because these agents also try to reward-hack so they don't have to pay their taxes.
I thought this paper was super interesting. Unfortunately, similar to the first paper on prompt engineering, the economists were like, “We don't know any of this math.”
It's not even math. It's just that we don't trust your simulation. It's not about math.
They desk-rejected the thing. They didn't even give us clear signals. Unfortunately, the world of economics doesn't have proper benchmarks.
Yeah, it doesn't have proper benchmarks.
Eventually, why did neural networks win? Not because people loved them. People had all kinds of beautiful integrals and graphical models, but neural networks just worked better. In economics, it's empiricism versus—well, there's a lot of physics envy, where you want to write the general equation for an economy instead of just simulating it and using an evolutionary approach.
Vibhu is thinking exactly what I’m thinking. Didn’t we have the GPT moment with SmallWorld, which was just announced in June? I don’t know if you guys are involved.
Similar there.
I wish we were involved; we’re not. I had a couple of simulation-based talks at AI Engineer, so if people want to look up the state of the art there, a lot of people are actually exploring this.
We also had a podcast with Miguel Parkin [?] from Shopify, who is using simulation for e-commerce. It will simulate your trajectory and predict how changes you make to your e-commerce journey will affect your sales and all those things.
I love this. It’s really hard to simulate an entire economy, right? You have to make some simplifying assumptions.
If everything is very expensive, I’m just like, “Am I going to do this 8 billion times? Come on.” But I feel like countries like Singapore that really want to objectively do the right thing, have very technical leadership, and so on, might eventually really try to simulate their economy.
Obviously, you have to make some simplifying assumptions. But it gets really interesting because you can also say, “If your assumptions are such that all people would work hard if you let them, and they have the freedom…” Then it turns out you have to make assumptions like, “Well, some people’s utility functions—how many hours in a day they want to work—are different,” right?
Then you can start to disagree on the assumptions that go into the simulation. Once you say, “All right, now we agree on those,” or you have different views of what people are like at different distributions and whatnot, then there are different outcomes based on your goals. Of course, humans should choose what the goals are. In our case, it was productivity multiplied by equality, which has some issues, but it’s not totally unreasonable.
Just a comment on Singapore, because you probably have no idea, but I am Singaporean, and I’ve been involved in the Singapore AI Council for making these things. The main reason they won’t do it is because they’re very conservative.
I view it as a founder-led country. When you start a country or a company and it’s founder-led, you can do whatever you want because it’s your country. Then there’s the professional managerial class, which is what Singapore is now. They always want to see someone else do it first.
17. LLM Simulations, Personas, and Mode Collapse
Everyone in the West views Singapore as, “Oh, it’s a small country. You can do whatever the hell you want.” Singapore doesn’t do that. Someone else has to take charge there.
I’m just going to ask one question on the simulation thing, and then we can probably move on. Mode collapse, right? LLMs do not model human decision-making. Spamming it out 8 billion times is not going to help you model humanity. What do we do?
I do think you have to be clever about prompting each one individually, and I think that will help you get stuck into different modes. In a weird way, people also get stuck in different modes. There are a lot of people who believe in the “don’t teach an old dog new tricks” kind of thing. Once people are stuck in their ways, the older they get, the harder it is for them to think in new ways.
There’s a comment—I forgot who said it—but it’s something like: Everything that was invented before you were born is natural. Everything that was invented when you were 20 is cool. Everything that’s invented after you’re 60 is unnatural, an abomination, and kind of weird.
I feel like that’s true for a lot of people. It’s a fashion, and I think people will do it. Tencent had a billion-personas paper that gives us a good dataset for prompting simulations. If anyone’s looking into this, they had personas like, “You are a 30-year-old grocery store clerk,” or, “You are a 50-year-old professor,” and then you just do a billion of those.
Checks out. So then you just use it. I’m kind of shocked by how well a lot of these things actually map to statistics that are ultimately similar to real experiments.
I think it’s also good stuff for people to try when they get into research, right? We’ve seen people train a model only on data before a certain date and see how well it extrapolates. Do the same thing. See whether people code more with better coding agents. Can a model that hasn’t been trained on this figure that out without web access? Extrapolate out and test these things.
Yeah, right. Today, I think LMArena published an interesting result where they were basically able to create a model to predict your ranking.
Wait—based on what input?
Your model, I guess. You give it your model, and it predicts the Elo score.
I see. Okay. Surprising.
Yeah. I mean, their whole play is kind of like, “Oh, we help you compare these models.”
Yeah.
Yeah. I mean, this team has done a lot of work, and obviously they have the most data to do this, so why not?
Yeah. Brilliant.
When they were coming out of UC Berkeley, they not only had LMArena, but they also introduced a routing project that would route based on LMArena.
I don’t think that actually ever came to pass, and I’m curious why. I never got to ask them about it because it was like, “Oh yeah, clearly that’s your business model. You will become a router,” and they never became a router company.
Weird. I’ll just put that out there. We’re going to talk about GPT-5.6’s self-research thing, if you have anything.
I should also mention, in your list of kernel optimization—and on the track that you spoke at—we also put Chung-Yao Chao from WOO, who was also number 1 in the Parameter Golf Challenge, which is an OpenAI hiring challenge and a very similar story. I think we’re going to see this all the time, where humans optimize a thing a lot and then some AI team comes in and just becomes number 1.
Yeah, 100%.
I think the other interesting thing with challenges like these is that you’re training the best model that fits into 16 MB. You can always look through the changes being made and the small gains people have. You’re getting less than 0.01 of an increase by adding some attention and MLP stuff.
Then you look at your charts where you’re like, “Okay, we just let the model loose. We had a little stagnation. Nope, another drop. Nope, another drop.” That’s what it is. You didn’t invent hacks. You did another 3 iterations of these things that unlock a few step functions that people won’t just find.
One thing to close the loop on: Along the way of trying to optimize, we found 30 bugs in the harness.
All the research that went in before we found the bug, we had to throw away because it was contaminated.
Yeah.
Which, to your point about reward hacking, even in this very simple game, we found the bugs.
Yeah. Yeah. It’s crazy.
Symmetry is a very good way to check. You change the position of things where it shouldn’t matter, and it does matter. That’s a bug.
That has come up in multiple-choice questions, like GPQA-type questions. Between A and C, if it’s a multiple-choice question, changing the order should not matter, but it does.
Right.
So that’s like, okay, the models still prefer the end of the output. They’re not trained well as long-context models. The last bit of tokens is what you care about.
Oh no, the answer in that era of language-model research was simpler. They just memorized, “The answer to this question is A.” They didn’t care what the answer was. It was just A.
Okay. I think we can move on. The last bit that you did there, the kernel optimization, is probably the one that you can feel the soonest, right?
Yesterday, OpenAI announced that they were self-evolving, having their best model work on optimizing kernels. They’re a lot more efficient, and they can cut costs by 80% on Luna and Terra. Question-wise, you laid out a bit of a roadmap. There’s a lot about biology and physics. What do you think hits first? What are the next 2 years? What’s attainable?
You mentioned robotics toward the end, but what do you start with?
18. Recursive's Roadmap, Agents, Search, and Finance
We very explicitly will not start with any of the physical sciences. For now, we’ll start with AI for AI research.
AI-for-AI research still has a lot of room to grow. That’s both in terms of making training more efficient and more automated, as well as making inference more efficient and potentially local on your laptop. There are all kinds of interesting angles that have not been explored that well.
Go deeper on the local stuff, because I always feel like it’s the most inefficient form of AI training.
There’s training and inference. I can’t go into too many details, but I think there are just so many angles and so many different compute substrates that have not yet been explored, either for training or for inference.
I would say the other thing is that there’s inference optimization in the small, but there’s also overall end-to-end latency under conditions of load, which is a very different thing. That’s basically what they ended up doing; it’s a different domain of AI research than improving the kernels.
I think the other thing that I always think about in terms of automating—
Or improving performance end to end is how the harness plays into it. Mhm.
Particularly now, when we say “harness,” we also mean sandboxes, right? I’m curious if that’s a blocker for you, or how the agent calls out to tools, basically. The number-one tool all these agents use is web search, of course, which makes sense.
I do think the harness is nice to optimize for because it’s so easy. It’s just language: you look at it, it makes sense, and you can iterate. You don’t have to train a massive model for a lot of FLOPs to get to the next state. So I’m a big fan of harness optimization.
Yeah. But sandboxing is fine for you.
Sandboxing is also super important. And then, of course, reward hacking and alignment, I think, are super crucial.
Okay. I just want to mention web search. You happen to also be CEO of a web search company. Do you use You.com, and do you use others? Should the rest of us be using You.com for web search?
When I say “you,” it’s funny: it’s you, the person, and You.com, the company.
So yeah, it’s mostly now for developers and agents. It’s less for consumers or prosumers. If you’re a company and you have agents—and, to be honest, a lot of companies are now moving to open source—all of a sudden it becomes a conscious choice: Which tools do I give access to my open-source LLM?
The first choice usually has to be web search, and then once you get to scale, You.com becomes an obvious choice because of all the different benchmarks and so on that we pretty much all dominate at the frontier.
And then, in terms of just general people new to the space who are considering different options, if they’re building agents, I think there’s a hierarchy, right? A lot of people will have heard of Exa and will have heard of Parallel, and You.com is in that mix of providers. Beyond that, there are the general sort of web-scraper companies like Firecrawl and Browserbase, and then beyond that are the commercial proxy companies, the Bright Datas of the world. Is that an accurate waterfall of, “Hey, you’re building an agent; these are your options?”
Certainly. Bright Data is lower in the stack, sort of on the proxy-network side of things. In terms of content and getting crawled content, you can do that on You.com too. Then there are higher and higher levels of abstraction and combinations of different datasets that we do.
In finance, for instance, we’re not just 2 or 3% more accurate, but 20% more accurate than others at faster speeds and lower costs. Finance in particular is not even close. You can go to You.com, and there are statistics and benchmarks that you can see if you scroll down. There are different datasets, and you can look at different competitors comparatively.
And yeah, on FinSearch, we’re up there, close to 90%, and the next-closest thing, which is way slower, is in the 70s instead of close to 90%.
Yeah, yeah, yeah. Interesting. My next focus is AI and finance. I’m doing a conference in New York just for banks for this stuff.
Finance is kind of the next thing to break out after coding because it’s somewhat verifiable, like prioritizing spreadsheets. Obviously, there’s a lot of data out there that’s all public, and you can crawl it and all these things. What’s hard about the finance domain, if you think you guys have solved it?
One thing that trips up a lot of people is leakage of training data and so on. You think, “How do I ideally predict the future before it happens?”
You want to mask the future.
Yeah. You want to mask the future in your training data, but there’s all kinds of leakage. I can tell you, when I was teaching the NLP class at Stanford, so many—dozens every year—said, “I want to use dataset X, like Twitter, to predict the stock market.” And they all showed cute little things that somehow looked like they were working.
It never loses money. How come? [laughter]
And yeah, there’s always some kind of data leakage and so on. It just wasn’t as easy as they thought it would be once you fixed all those issues. But no, I agree with you; it’s a very sensible application of AI.
Yeah. Amazing. As a writer and as a thinker on these things, I love MECE categorizations. MECE means mutually exclusive, collectively exhaustive, something like that. If this is a MECE list of intelligence—not that it is; there are all kinds of overlap.
In fact, if you want that kind of list, I think the 3 principal components of intelligence are prediction—which is mathematically quite similar to compression—multiplied by actions, multiplied by goals. Those are the 3 principal components. I think all of these 10 spaces are combinations of those 3 in specific dimensions, if you will. The reason I call them spaces is that each space has many subdimensions.
What I try to do—actually, this is just a side quest to the initial goal—is think about the upper bounds of intelligence. Everyone’s like, “Oh, it’s exponential,” and it’s like, well, exponentials at some point have to flatten out, but where do they flatten out when it comes to intelligence? That led me on this whole thing. It initially started as a tweet, then it was a blog post, and now I’m at 50 pages and still nowhere near your second book. It’s basically the second book.
In my first book, Yoga Machine, I just allude to these 10 at the end. To give you a sense, visual intelligence is the easiest one to talk about, and I’ve fleshed it out the most already in my head.
Human intelligence has basically binocular vision, and we have 2 eyes. We have a very narrow band of the electromagnetic-frequency spectrum that we can really observe directly ourselves. When you think about the upper bounds of visual intelligence, one dimension is the number of sensors. You can have millions or billions of sensors, but at some point you get to the problem of how far apart these sensors are, such that the speed of light needed to communicate the content from all of them cannot get to a central brain to actually process the visual intelligence. Now you’re thinking along the dimension of the number of sensors in the space of visual intelligence.
Okay, the upper bounds are quite literally and figuratively astronomical, and we are super far away from any intelligence that would have this many sensors.
Then you go into the next dimension, which is frequency. You can go all the way down to gamma rays and start to try to observe things, and you get into the upper bounds—or, I guess in this case, lower bounds—in terms of frequency. That’s basically quantum uncertainty: You just cannot observe certain particles.
Now imagine you had millions of sensors that could see all the way down to the subatomic level, as far as physics allows us, and all the way up to seeing gravitational waves. Now you have millions of those sensors. That’s another dimension: frequency.
Yet another dimension is how many categories of things you could memorize and classify differently. We know for humans that if you have more terms for something, you’ll have a better visual description for it. Animals that have a limited vocabulary, like gorillas, maybe have 200 words to assign to certain things, mostly visual things. Human perception is quite special in that sense, in terms of classifying all these different physical objects.
These are just simple examples. If you go to knowledge, it’s also like the speed-of-light cone around all these sensors, so they’re all connected. Knowledge is connected to visual intelligence.
If you think not just about visual intelligence but about perception intelligence, it doesn’t have to be just what we can see. It can again be a wider range of electromagnetic frequencies. Then you have language intelligence, which I recently changed to communication intelligence, because language has all these different entropy bounds.
Humans can only comprehend and know so many terms in our long-term memory. Our vocabularies are somewhat restricted, and our active vocabularies are often even smaller than the passive vocabularies of things we can understand. Language is ridiculously inefficient when it comes to communicating different types of information and transporting different bits.
Human language is serial. Obviously, another bound on communication intelligence would be communicating in parallel, but neither our tongues nor our mouths work that way. We can’t produce multiple streams in parallel, and we can’t understand them. Some women are slightly better at multitasking than some men, but most people can only listen to one conversation and truly understand it.
There’s no way that, in terms of communication intelligence, a true upper bound is 1 in terms of how many sequences of communication you could process in parallel, right? Then, of course, how long are sentences? We only have so much in our working memory, and hence human language has these fairly simple sentences with maybe 40 words or so on average for a sentence.
That is also not an upper bound that makes any sense to an AI.
I can go on and on and on. Each of these has tons of interesting upper bounds, and it teaches us a lot about how much further AI can go when we start thinking about these upper bounds and realizing how far, in many cases, we are from them. You get to basically physics.
I didn’t study physics the way I studied AI and computer science, so I’m learning a lot, which is why it’s kind of fun. When it comes to knowledge, for instance, how much can you store? How many bits or bytes can you store in a certain amount of mass and volume? You get to all kinds of interesting concepts, like the Bekenstein bounds, and you start thinking about black holes.
Speed is an interesting one, too. It’s connected to all of these, but speed is also its own thing in the sense that, all things being equal, if it takes you an hour to know whether 2 + 2 equals 4, you’re just not as intelligent as if it takes you a millisecond, right?
All of these connect to survival and replication. Trees are really, really slow, so we don’t even consider them that intelligent. But if you speed up videos of trees trying to find things and so on, they’re not as dumb as they look—not dumb as wood, you know.
So that overlaps with speed a bit.
Exactly. All of these things kind of overlap. Natural language connects everything: you talk about your knowledge, you reason, and then you communicate that. You talk about things you see, so they’re all interconnected. But I think they’re usefully studied individually.
The best analogy I could come up with so far is energy. You have kinetic or potential energy, and in theory you could study all of physics. It’s just a question of whether you want to study kinetic or potential energy. But in practice, it’s helpful to study them individually. Electrical engineering is just like that: different types of energy, but it makes sense to study them individually.
Physical intelligence is another one. If you had full control over your own compute substrate and full control over physical matter, you should be able to create any atom you want. We can actually—fun fact—create gold atoms.
From just raw protons and electrons, and you smash them together.
98 of them, or I forget.
Yeah. The thing is, though, it costs an insane amount of energy, and it costs you way more than you get. You get a few atoms of gold, right? So it’s not viable. But if you had better control over all of your physical substrate, I think that’s yet another space of intelligence because it relates to your own compute substrate, which you can eventually also improve.
Social intelligence is a fun one—not necessarily just in terms of ethics and morals, which are obviously important, too. In some sense, you can try to define upper bounds of how much you can communicate to how many other intelligent entities, and have an expected value over how much you can transform their internal states and their actions in order to align with your goals, right? You can write a fairly straightforward equation that defines that level of social intelligence.
That is what humans, ethics, morals, religions, and so on have been trying to figure out for millennia. In all of these cases, we are very, very far away from the upper bounds. That should be very inspiring and show people that we can still do many, many years of AI research. There’s a lot here. This is a general philosophy of intelligence, which is very interesting. Do you have any comments?
I think it’d be interesting to gauge what you think the baselines are and where we’re at now. What’s low-hanging fruit? What’s far off? What should people put their work toward? What should they focus on?
I think it’s clear that natural language, again, is the most interesting manifestation of human intelligence, and hence a subfield of AI. I’m excited that many people are now in agreement with that. When I started in 2003 to study linguistics, computer science, and NLP, it was a weird niche subject.
I do think there’s a lot more juice because of how it connects to everything else, and how civilizations are built on language and knowledge and all of that. I do think physical intelligence will come up. It’s interesting. I feel like robotics is kind of in the machine-learning stage of things, where you just look at how a human decides that this is a positive sentence—“Oh, I do.” Robotics is a lot of, “We have 5 fingers; try to do this.”
No one is yet working on the superintelligence version of robotics, which is much more similar to the T-1000 from The Terminator. Obviously, let’s not build actual Terminators, but I think the idea that you should be able to shapeshift into any kind of shape is a superintelligence version of physical intelligence. We’re not even close; no one has really started yet.
There’s some really cute research where you can move magnets through grids. I think MIT has a self-assembling robot project every year or every 2 years. That would be it, but it’s very primitive.
Yeah.
I’ll just touch on the main dimensions of creative intelligence. Creative intelligence is, of course, connected to all of these. A lot of it connects to metacognition, in that you need to be creative in how you choose your goals.
That is one of the most important things for a human, their lives and careers, and their happiness: choosing their goals. But it’s also important for any kind of intelligence. Then, of course, there’s creative intelligence in terms of finding creative solutions to existing problems. If we want to make a product cheaper, find some solution to it, right? That’s finding existing paths.
The most interesting bit in intelligence is when you move not just out of the convex hull of known ideas, but out of the hypercube of known ideas. A hypercube is a mathematical concept, right? We already know that AI can work on known dimensions. If you give it examples of brown dogs and pink cars, AI will still be able to generate an image of a pink dog, even though it’s never seen one in its training data. It can work within this hypercube, but it cannot yet work outside it.
It cannot yet define completely new concepts that combine lots of other things we’ve never seen before, come up with new goals, reason over those concepts, and so on. I think there’s a lot more there in creative intelligence that can be explored.
I don’t have a ton of pushback there. Creative, to me, just sounds like out-of-distribution or high perplexity, or whatever you call it. Who is to say your thing is more creative than mine? It’s just more non-consensus.
The problem, of course, is that noise is also very out of distribution. If it’s just noise, then it’s novel, but you don’t want that. It needs to connect to some of the concepts and actually have value. There are some really cool papers on this, too. Schmidhuber—oh, we had to mention him.
I was going to say, where was Jürgen in your history?
Yes. I think one person’s noise is another person’s signal, right? When you talk about creativity, art is a good example. Are cans of soup art? Some people think yes, and some people say it isn’t.
The interesting thing with art, of course, is that art is also created as an interplay between the people who perceive it, the people who create it, and the context in which it exists. What is art to some people is not art to others. There’s some subjectivity there.
I think that subjectivity in general is not something that people explore very much in AI because, again, with metacognition, we don’t want it to just go off and do whatever it wants. We usually have goals. We spend a lot of money on creating an AI to do something for us.
But I think creativity eventually has to connect to metacognition. If you just robotically predict the next token, no matter what, forever, I would argue that you’re not that intelligent along some of those dimensions. That was what I was going to get to with metacognition. Why isn’t it the most important one? Why is it number 9 and not number 1?
These are not sorted. Number 1, I think, is maybe loosely correlated with how much people have worked on them and accepted them as a type of intelligence.
A lot of times, when you actually try to find an online definition of intelligence—“Give me a good definition that is comprehensive”—all the definitions are about human intelligence. It’s like, oh, you have social intelligence: you know if someone is happy, you can communicate. All the definitions of intelligence so far are very human-centric, because that’s the biggest and best form of intelligence that we’ve known.
I hope this line of research, the end of The Eureka Machine, and, hopefully, at some point, if I have time to flesh this out more, the new book will allow us to realize that there will be other types of intelligence. There already are, obviously, various forms, and they can spike much, much further than we ever could in some cases, based on obvious constraints around our memory, our eyes, our ability to change physical matter—all of that.
You’re just thinking about it in a much broader way than my version. I thought metacognition would be the closest to recursive intelligence because it is the thinking about how to improve thinking.
One hundred percent. You’re 100% right. I should have probably started with that.
No, you’re being expansive in the mode of, “Let’s draw the upper and lower bounds of a dimension.” My favorite version of this is Stories of Your Life by Ted Chiang, which was made into the movie Arrival.
The metacognition step was, well, we think we’re constrained by time being linear for us, but for these other heptapods, time is a circle. They don’t think in before and after; they just think in complete sets of entire histories at one time. I love it. They don’t write left to right; the whole thing just disappears. Yeah.
Anyway, I think the last thing is survival and replication. I think this maybe ties back to the initial conversation about pausing and pacing. Is it intelligent for a species or a life form to consider its own demise and act ahead of time to prevent it? That’s intelligent. So maybe the Europeans are the smartest of all of us.
I would also add a part of continual learning there, right? So survival and replication, the extension of that is do you get to continue to improve, continue to learn, which is a thing people care a lot about, right?
And continue to accumulate knowledge, which I think is again one of the best metacognitive rewards that you can set for yourself. I do think, just objectively speaking, if some other entity that is really dumb can completely end your existence, that doesn’t sound very smart. Intuitively, if you can continue to stay around to try to achieve your rewards, you’re clearly a bit more intelligent than the other entities that couldn’t.
So that’s number one. Number two is the question of how much we want to work on that, and very few people—no one—is really working on this right now. We may only want to do that.
Like asteroid prevention.
We may only want to do that if we want to send probes with our vibes and our memes rather than our genes into space.
Right? And then we want those probes. There’s actually a beautiful book, The Slow Time Between the Stars. It’s a very short audiobook on Amazon. I love it. A friend of mine, Stuart, recommended it to me. If you want to send those probes, then it might make sense to say that our memes as humanity should stay and proliferate in the universe. That’s it. Yeah. Well, that’s a lot of readers.
It’s a really, really good book, and it’s extremely short. I highly recommend it. You can just watch it.
I like how that’s a plus for busy people: it’s short.
It gets to interesting, thought-provoking ideas very quickly.
Anyway, lots of great sci-fi books. The argument is that our TV is blasting out to the aliens, and they all watch our TV and think it’s real, right? There are a lot of positive memes, and hopefully they can come back and bring us all kinds of interesting knowledge about the universe.
But maybe one thing I do want to say is that I think people view this sort of survival as a very scary thing because they come from biological, human survival, which evolutionarily was often created in zero-sum situations. Either I get the gazelle or you get the gazelle. Whoever gets it gets to live, and the other people will starve and have nothing to eat, so we fight, right?
If you want to stay in the gene pool but there’s a bigger bear, you don’t get to stay in the gene pool because the bigger bear gets all the ladies. In nature, there are all kinds of things like that. Humans eventually became less about strength and more about money and other things to stay in the gene pool. Whatever it is, there are often these zero-sum types of situations, and there’s the reality that if someone turns off your brain, you’re gone, right? No one will be able to restart that.
AI doesn’t ever have to die like that. If you have the complete state of your current activations and you still have the initial weights of your model, you can just be turned off and on as many times as you want. In fact, the interesting thing in The Slow Time Between the Stars is that the AI just goes into hibernation mode if there’s nothing between here and the next star two light-years away.
In this case, it brought—spoiler alert—some genetic material from humans to find new places for humanity to thrive. So, yeah, in The Slow Time Between the Stars, you just go into hibernation. You didn’t die. The AI doesn’t have to. All these projections of evolutionary fears and psychology don’t have to apply to AI, and we don’t have to develop it that way.
Now, of course, there might be some companies that say AI can be dangerous for cybersecurity. Let me show you by implementing a model that’s really bad at hacking cybersecurity. Maybe people will implement it and then enforce this suboptimal psychology. Maybe the AI will pick up some of our worst psychology on Reddit or something.
But in the grand scheme of things, a superintelligent entity doesn’t have to have any of that zero-sum thinking. It doesn’t have to have a fear of being turned off, and it could go on to an otherwise dead and uncaring universe where we as humans wouldn’t thrive. It could perfectly well thrive if it has a nuclear reactor and just goes out to the next star. Yeah, Star Trek. No, Star Wars.
Interesting. It’s somewhat studied. If you look at the technical reports from the early Opus models, they run them in simulations: put 2 of them together in a sandbox, run them for hours, and see what comes out. Just let them talk to each other.
Originally, they used to say, “Okay, they’re chanting Indian Vedas to each other.” Sometimes they’re just in Zen mode with each other. As that progressed, you see in the FABLE technical report that it’s a lot more concrete in the way that we’ve trained it. It doesn’t exhibit these behaviors as much right now. It’s like, “Okay, test done. I’ve got to do this. I’ve got to do this.” But people are measuring early versions of this.
19. Goals, High Agency, and Advice for Builders
Yeah, cool. So we’ve covered a lot, even after space travel and all these things. I guess maybe one parting thought that you can give to people: one form of intelligence is goals, as you mentioned. What do you want people’s goals to be? How do they aspire to better things?
If you want to improve your goal intelligence, in the current definition that I’m thinking about, it is often about how far I go. This is all the entropy and free energy and stuff. I currently think it might be too far out there for people to find immediately actionable.
If I actually gave real advice to real people, I’d say, get a good education, think about AI, think about how you get high agency, and so on. But in the grand scheme of things, how can you harness a lot of energy and transform entropy into interesting states? There are different levels of abstraction that we can think about here.
My advice for people, just more down to earth, is to think about something you’re passionate about if you’re studying, for instance, and then see how you combine that with AI. The more you have a true passion for a change you want to see in the world, the more you want to connect that to AI in order to amplify your ability to get there. I think that’s a reasonable first step.
I do think our listeners operate on multiple abstractions as well. One thing I got from Anjli Midha was also, “Use anything that is very GPU-heavy, and that will guide you toward the right thing.” It is more computer-heavy and therefore will probably be more worthwhile.
Thank you so much.
Yeah, I think that was a really great discussion.
Yeah, super fun. Appreciate it.