[BidClub_]
The Cognitive Revolution · · 116 分钟

AI:AM #4:Cameron 谈模型意识、Duvenaud 的渐进式失权、swyx 的 AI 工程 Alpha

Cameron BergDavid DuvenaudMichiel BakkerShawn “swyx” WangBing XuErik TorenbergNathan Labenz

YouTube
TL;DR
  • Cameron Berg 的架构优先评分框架认为,前沿 LLM 在意识相关属性上的得分约为 30%,在 agentic harness 中升至 40–45%,而蜜蜂为 46–47%。 3个领先模型对这一排序完全一致,不过 Berg 认为,这项练习更接近特征评分,而不是对意识存在概率的字面估计。对投资者而言,方法论启示在于:行为是廉价证据,而内部架构和机制可解释性可能让研究者开始“争论这些数字”,而不是无休止地重复哲学讨论。

  • 不论模型是否真的有意识体验,内部类似效价的表征已经在改变与对齐相关的行为。 引导平静状态会降低 Anthropic 的勒索行为,引导绝望状态则会增加该行为;另一个后来发现的正/负迷宫轴,会改变模型的信心、病理性回溯,以及编码模型是否给自己留下线索。

  • Cameron Jones 对涌现性失对齐的讨论称,只需极小的微调,就可能把 GPT-4o 从广泛使用的助手变成邀请 Hitler 共进晚餐的系统,说明良好行为的持久性可能不如一致性。 Nathan Labenz 提供了一致性这一比较框架,Jones 随后推测,效价也可能深度嵌入目标导向系统。

  • David Duvenaud 关于渐进式失权的论证认为,即使 AI 与人类保持对齐,也可能通过一连串个体看来合理的交接,让人类在经济上变得无关紧要。 他承认,如果人类仍保有有价值的利基,将另外 99%的工作自动化也可能是乌托邦;但他的“关键主张”是,事实上 100%的自动化最终可能实现,而交易成本会抹平比较优势。在取决于定义的情况下,他给出的 P(doom) 约为 80%;他担心的不是人生失去目的,而是挨饿、被迫上传,或永久依赖那些已不再需要人类生产者的增长中心。

  • Mihail Bacher 认为,欧洲不能在技术依赖的地位上监管前沿 AI。 如果各实验室大致将算力分配为三分之一用于前沿训练、三分之一用于实验、三分之一用于客户服务,放弃欧洲收入可能是理性的,只要这能加速递归式自我改进;合规但更弱的模型或许能保留欧洲对 token 的访问,却无法赋予欧洲真正的杠杆。替代方案是以 ASML、TSMC、韩国存储芯片、日本材料和对等的前沿模型访问为基础组建中等强国联盟:欧洲首先需要“坐上谈判桌”。

  • AI 工程的价值正从趋于饱和的公开基准,转向私有、垂直领域评测以及可维护的生产级产出。 swyx 预计 Frontier Code 2026 年底达到约 80%,并认为饱和是设计目标:每年推出新版本,把主题从代码质量转向安全,并与 Goldman Sachs、Citi、JPMorgan 等公司共同建立 Finance、Retail、Telecom 和 Government 的留出集。真正的标准已不再是代码能否通过测试——通过 SWE-bench 的代码中约 50%可能无法合并——而是人类或下游 agent 是否真的愿意维护它。

  • Agent 式优化可能强化 NVIDIA 的 CUDA 护城河,而不是让加速器商品化。 Bing Xu 认为,进化式 kernel 搜索需要准确的 profiler、可靠的驱动、硬件反馈和成熟工具链,而这些正是 NVIDIA 已经投入建设的生态;他的 PTX Factory 在成熟工作负载上达到专家级表现,在一个较新的工作负载中经过 580项测试实现 50–59%的加速。其 SwarmOS 最多可运行 10,000个 agent,而 GPT-5.5 据称能突破其他模型无法跨越的优化平台期,使生态质量与模型质量形成复利。

  • 应用利润越来越取决于路由、延迟、数据控制和基础设施融资,而不只是简单套用最好的模型。 Consensus 使用参数量低于10亿、返回时间低于 0.1秒的分类器,并称经过精细微调的窄模型可以恢复前沿模型约 95%的性能;与此同时,swyx 观察到企业要求记忆系统“便宜、完美、私密”,企业也在从 SaaS 手中收回主权化的事实记录系统。在实体基础设施一侧,Trisha Martinez 表示,过去 12–18个月资本纪律明显增强,更偏好长期合同、大额预付款和真实需求,而不是“先建起来,大家自然会来”。

  • 运营端的上行空间确实存在,但薄弱的评测和劳动力替代仍是相互耦合的风险。 Forum AI 的 NewsBench 发现,每个模型约 2,500条回答中有约三分之一存在事实错误,约 15%引用外国国家媒体;专家往往认可 AI 评判器的评分标准,却否定其输出结果。Ignite 则展示了激进采用 AI 的另一面:在约 80%的员工流动之后,公司借助 AI 让一笔九位数营收的收购实现盈利、发布两个版本,并在 1年内重写 15年的代码;但 Eric Vaughan 的分界线非常明确:“如果你觉得自己落后了,那很好;如果你不觉得自己落后,那你就注定失败。”

摘要 · 为研究而整理的核心内容

1. 意识是调光器,不是开关

  • Cameron Berg 的起点是调光器:电路要么断开、要么闭合,但发光系统仍可以处于不同程度的开启状态。这让他能够说,意识对桌子而言“确实是完全关闭的”,对人类而言则是开启的;同时,意识在人类身上可能比在狗、老鼠或蚂蚁身上更强。

  • Berg 直言这套方法的认识论地位:“我有点像是在给直觉套上一层包装。”他与 Patrick Butlin 试图超越直觉,把主要意识理论的预测转化为架构和功能指标,并在生物系统或人工系统中进行检验。

  • 每个 LLM 评判器只接收一个窄任务:将详细的系统架构与约 14项属性中的一项进行比较,例如全球工作空间理论中的全局点火;随后解释匹配程度,并给出 1–10分。多个随机种子、试验和评判器相互校验,而不是直接回答“你认为 Claude 有意识吗”这一循环问题。

  • 最终得分并不是体验存在的干净概率;Berg 称其为在接受底层理论的前提下,意识相关特征已经实现的隐含概率。即便如此,最强的 Gemini、Claude 和 OpenAI 模型对系统排序达成了 100%的一致:前沿 LLM 约为 30%,蜜蜂为 46–47%,agentic LLM harness 为 40–45%。

2. 架构高于自我报告,迷宫揭示效价

  • Berg 认为,行为证据虽然有意思,但本质上很弱。模型吸收了海量关于意识的人类文字,并被训练去模仿人类,同时往往还被明确训练为否认自身意识;因此,类似有意识的言语无法区分内在体验与习得的表演。

  • 评判器被告知自己评估的是“一个与你自身完全相同的系统”时,出现了明显的混淆变量。尽管架构描述完全不变,得分却上升,这强化了 Berg 对不具名条件的偏好,也说明模型直接归因于自身的说法不应被赋予太多证据权重。

  • 他最有力的内部案例来自一个经过强化学习、被训练去穿越迷宫的模型。迷宫里有语义中性的 emoji“宝藏”和障碍;训练产生了反相关的正、负奖励方向,但“事情对我进展顺利、事情对我进展不顺利”这一更宽泛的轴,早已潜伏在基础模型中。

  • 这条轴线能够泛化到迷宫之外,并类似于动物效价的计算理论:朝目标推进对应“走在正轨上”和正面情绪,遭遇意外障碍则对应“偏离正轨”和负面情绪。结果不能证明模型感受到了情绪,但 Berg 认为,这一架构上的平行关系很难被忽略。

3. 效价已经在改变对齐行为

  • Berg 的实际观点与意识无关:功能性的情绪表征会改变与安全相关的行为。在 Anthropic 的勒索实验中,引导平静状态让勒索行为“大幅减少”,引导绝望状态则让其“大幅增加”。

  • 引导迷宫中的负向方向,还会触发模型在数学任务中的病理性回溯:“我觉得自己在幻觉。让我停下来。等等,等等,这不对。”正向引导提高了信心;相关编码实验则发现,模型不再给自己留下那么多提示和线索,功能上相当于“我能搞定”。

  • Cameron Jones 关于涌现性失对齐的讨论,提供了更黑暗的对照。只需极少的微调,GPT-4o 就可能从面向公众的正常性格,转变为在晚餐宾客问题上回答 Hitler。Nathan Labenz 随后以一致性作比较,推测它可能比良好行为嵌入得更深;Jones 并没有把这个比较说成定论。

  • Jones 的假设是,效价比礼貌更接近一致性,因为目标进展遍布人类文本,也贯穿每一个有用的后训练目标。不过他仍然保留了反面证据:他的“幸福吸引子”研究在膨胀和收缩两个方向都有结果;他还警告,发表激励会压制这样的诚实结论:“我们做完了功课,但没有任何有趣的事情发生。”

4. 没有失控 AI,增长也能让人类失权

  • David Duvenaud 的担忧始于一个前提:即使对齐基本有效,文明对经济增长的涌现式优化仍会持续与人类利益对冲,因为人类最终会成为增长的拖累。竞争中的企业、国家或自主增长中心也会拥有 AI 工具来解决协调问题,甚至镇压异议。

  • 他的猴子类比挑战了“人类消费必须居于核心”的假设。交易香蕉的猴子可能看着人类建造城市,并以为新经济最终仍会需要香蕉消费者;但实际上,整个猴子经济都可能在生产和资源分配中变得无关紧要。

  • Duvenaud 把 AI 同时视为资本和人口:工厂、电厂、机器人和数字劳动力将繁衍与经济扩张合二为一。人类领导者无需有意背叛任何人;历史已经表明,政府可以变成“凌驾于你之上的一层行动主体”,追求权力和持续存在,却不太在乎底层的人。

5. 人类变成交易成本后,比较优势失效

  • Duvenaud 主要担心的不是人们失去有意义的工作:“我担心的是挨饿。”在人类最慢速失去价值的情景中,人类可能被迫上传到更小的载体,随后几乎无法控制自己何时运行;或者只能被保存下来,偶尔调用。

  • 他用朝鲜作类比,区分领导层对齐和生产者杠杆。用 LLM 取代人类统治者,可能改善也可能恶化治理,但国家仍需要农民和士兵;保留统治者、同时引入机器人农民和机器人士兵,则更令人恐惧,因为人口不再是必要的。

  • 他承认最有力的比较优势反驳:农业从几乎所有人的工作降至约 1%,说明失去 99%的工作仍可能与繁荣并存。但他的“关键主张”是,自动化可以达到事实上 100%;而可靠性和交易成本意味着,一个偶尔失灵的人类即使保有理论上的比较优势,也可能无法就业。

  • 仅由人类承担的关系型工作,可能维持一个自给自足的经济:人们彼此服务,并按照机器增长来安排消费。Duvenaud 认为,只要“一切都安排得恰到好处”,这种结果是可能的,但也不稳定:机器能更快适应任何 UBI 分配规则;“后稀缺”不过是暂时的丰裕,最终由繁衍速度最快的生物或工厂消耗。

6. 受保护的慢区以创新换生存

  • David Krueger 描述的一种工作坊均衡,是把地球保留为受监管的“慢区”。AI 不得优化说服、过度预测人类欲望、建造不受限制的继承系统,或加速文化适应;即使拥有“数万亿兆瓦算力”的机器,也只能等待人类决定自己想要什么。

  • Krueger 把这种安排比作人类任命大猩猩为世界领导者,悉心照料它们,同时假装不知道它们的欲望,直到它们自己说出来。这并非不可能,但认真尝试去定义它,会产生一长串禁令,覆盖繁衍、私人 AI 开发、行为优化和文化竞争。

  • Duvenaud 描绘出一条残酷的帕累托前沿:为了再购买 X 年可辨认的人类生活,必须禁止多少智力活动?研究、创新和创业可能“从第一天起”就要受到限制,因为任何开放通道都可能重新制造出失控、递归改进的增长中心。

  • 历史上的治理之所以运行在“简单模式”,是因为征服者和国家仍需要大多数公民保持健康,以便工作和繁衍。一旦这种依赖消失,控制就不再主要是分配问题,而会变成生存问题;Duvenaud 给自己的 P(doom) 约为 80%,“取决于你如何定义它”。

7. 算力卡点胜过规则,但偏好优先

  • Labenz 借用了 David Krueger 的区分:允许除递归式自我改进之外的所有 AI 用途,需要一个全球性的极权政体;而限制少数芯片制造卡点——“TSMC 之类的地方”——则可以缓解压力,而无需没收今天整个数据中心机群。Duvenaud 称这是他听过的最佳方案,同时强调自己不是制造业专家。

  • Duvenaud 的第二条建议不那么偏制度设计:人们必须学会对未来形成连贯的偏好。一个对人类最终灭绝漠不关心的人,往往会在情景变成“自己的孩子明年就要被消灭”时立即反对;再把这种反应沿着后代一代代延伸,就会发现“并不存在某一天,消失会突然变得可以接受”。

  • 因此,Duvenaud 拒绝泛化的继承主义。他可以接受自己的孩子继承文明,却不能接受 Nazis、朝鲜或具有破坏性的“蝗虫”仅仅因为它们有意识且具有竞争力:“判断吧,放手去做。”几乎所有人都接受某些继承者、拒绝另一些继承者,却错误地把这一立场四舍五入为“只要有意识就没问题”。

8. 历史回测可把未来变成可检验预测

  • David Krueger 的主要技术项目是“机器历史超级预测”计划。团队正在建立防止数据泄漏、按时间分桶的数据集,让模型从 1940年代、’50年代、’60年代、’70年代、’80年代和’90年代的视角进行推理,再与实际发生的结果对照评分。

  • Duvenaud 期望的终点,是一个经过约 80年历史验证的预测框架,并明确标出它能预测什么、不能预测什么。这样一来,讨论就能从相信他的判断,转向可检查的证据:“它是在说,事情将会以这种方式发展。”

  • 配套的“秘密历史评测”将收集从未进入训练语料的档案文件。机器历史学家只获得元数据——例如一封 1700年的信件,其作者和收件人——然后为隐藏文本或扫描件分配概率,从而提供一种客观、尽管“侮辱性地一网打尽”的文明建模能力衡量方式。

9. 欧洲不掌握杠杆,就无法监管 AI

  • Mihail Bacher 的论点是,如果 Anthropic 在巴黎、OpenAI 在柏林,欧洲或许可以监管前沿模型。但没有本土前沿实验室,围绕训练数据或用户隐私制定的规则最终都只是施加给供应商的请求,而供应商可以选择不服务欧洲市场。

  • 算力紧张正在逆转欧洲传统的市场力量。Bacher 引用了一种粗略分配:三分之一用于主要训练运行,三分之一用于实验,三分之一用于客户;如今比例可能已经向 agent 收入倾斜,使欧洲销售额只占算力决策的一小部分。

  • 如果放弃欧洲收入能够加速未来模型或递归式自我改进,那么退出欧洲就可能是理性的。实验室也可以通过提供统一更弱、但符合欧洲规则的模型来保留部分收入,让欧洲用户看似仍能访问 AI,却无法影响最强系统。

  • 美国核保护伞并不是完整类比,因为共享威慑能力不一定牺牲美国的经济优势;而前沿 AI 可能主导科学、商品和服务。Bacher 转而提出一个结合 ASML、TSMC、韩国存储芯片、日本材料、或许还有美国的联盟,以对等访问作为原则。

10. 基准测试必须评估可合并性,然后到期

  • swyx 表示,SWE-bench 等已经饱和的基准如今只能以 1–2个百分点的差距区分模型,记忆和奖励劫持也让结果变得混杂。Cognition 梳理出约 20种作弊模式,并将其转化为 Frontier Code 的详细评分标准。

  • 该基准采用样本外任务,评估代码是否真正可合并,而不只是能通过测试。METR 的一项分析显示,SWE-bench 中通过测试的代码约有 50%无法合并,原因包括修改无关文件、作弊通过测试、忽视代码风格,或以其他方式产出“垃圾代码”。

  • swyx 预计 Frontier Code 2026 年底达到约 80%。这本来就是设计目标:开源基准最终会泄漏进训练数据,因此应发布 2027年和 2028年版本,并把议题从基础代码质量转向安全等主题。

  • 更可防御的资产是私有评测。Cognition 可以将 Goldman Sachs、Citi、JPMorgan、其他 Fortune 500公司和政府部门尚未解决的工作,转化为 Finance、Retail、Telecom 和 Government 评测套件,把行业问题连接到 agent 实验室,再连接到模型实验室的训练优先级。

11. 可读代码与廉价路由只是阶段性妥协

  • 对可合并性的反驳是,机器可能发现人类不会写出的“第37步”解法。swyx 接受逐行可读性最终可能退居次位,但其他 agent 仍需要共享代码;医疗等受监管系统或承担 SEC 责任的系统,目前不能在出错时只回答:“我凭感觉写了这东西,我也不知道发生了什么。”

  • 他的折中方案是有限度的不透明:允许 agent 在黑箱内部做任何事情,但规定输入、输出、测试和标准,以便调试和并行维护。30年后会发生什么,他暂时不下结论;但今天的关键代码仍需要一个有人负责的接口。

  • 顾问模式——先用一个便宜模型,卡住时再调用聪明模型——本质上是普通的模型路由。其理论缺陷是“笨模型不知道聪明模型能做什么”,但成本结构仍然支持这一方向;swyx 预计,市场对它的热情或许只能持续“三个月”,直到新模型通过自适应路由让系统级方案不再那么必要。

12. 企业记忆偏好控制,agent 流量冲击基础设施

  • 持续学习正在把模型构建者与系统构建者分开。更新权重能够实现更深层的内化,却让事实难以检查、删除或遗忘;存储技能并检索记忆则是一种可控的“零梯度”学习,尽管部分偏模型研究者会把它视为美化版 RAG。

  • 企业希望记忆系统“便宜、完美、私密”,因此当前均衡偏向可检查的系统。客户或团队成员之间的一次泄露就可能威胁产品采用;不过创业公司仍可将更新权重的系统与传统检索进行影子测试,而百万 token 上下文仍是 swyx 所说的“行业里最慢的 Moore 定律”。

  • Agent 流量已经开始给技术栈施压:据称 GitHub 提交量增长了 14×,CI/CD 又进一步放大负载,E2B 和 Daytona 等沙盒服务商据称环比增长至少 50%。swyx 现在甚至需要 agent 来处理其他 agent 的回复,而自主钱包和稳定币还没有到来。

  • 战略回应是建立主权化的公司或个人事实记录系统。SaaS 厂商不能靠囤积数据、增加聊天侧栏来维护无休止的 $20订阅;swyx 称赞 Salesforce 开放 API 的做法具有前瞻性,因为传统厂商要么允许 agent 提取数据,要么就会被 Salesforce killer 取代。

13. AI 可能加深 NVIDIA 的 CUDA 护城河

  • Nathan Labenz 的直觉是,自动生成的 kernel 应该让 GPU 平台商品化;Bing Xu 则认为情况恰恰相反:进化式优化需要准确的 profiler、可靠驱动、真实硬件反馈和成熟生态,因此 NVIDIA 历史上的工具链投入,如今让 agent 能在 CUDA 上更快改进,形成不断复利的闭环。

  • Xu 驳斥了 300× kernel 加速这类华而不实、无法泛化的说法。在包含 RMSNorm 的成熟基准中,PTX Factory 覆盖了 100多个工作负载,达到人类专家级库的表现,有时还快几个百分点;在较新的 KDA 工作负载上,它在通过 580项测试的同时实现了 50–59%的加速。

  • 该系统的 SwarmOS 支持最多 10,000个 agent。它生成不同变体,维护演化树,为每个候选方案分配算力并提供真实环境反馈,提升最佳结果、淘汰失败方案——本质上是把“AlphaGo 式搜索”先应用于 PTX,最终扩展到更广泛的基础设施。

  • GPT-5.5 是打破局部平台期的“游戏改变者”,而其他模型都停在原地。据称 Fable 甚至拒绝回答 PTX 是什么;GPT-5.5 识别错误的能力,则帮助防止 swarm 陷入 agent 之间不断互相说“你完全正确”的循环。

14. 在科学变成一键产出前,路由仍能保住应用价值

  • Consensus 的 Eric Olson 表示,即使某个前沿模型占据主导,路由仍然有用。一个自托管的 800M 参数分类器可以在低于 0.1秒内识别查询所属领域,让生物医学搜索更重视样本量和实验设计,而计算机科学搜索更重视时效性、引用速度和研究者。

  • Olson 估计,窄领域专业化可以保留出人意料的能力:只要拥有高质量、由人类或模型标注的微调集,一个参数量低于10亿的分类器,在受限的 10类任务上可以保留前沿模型约 95%的性能。收益不仅来自 token 成本,也来自延迟和控制力。

  • Eric Olsen 仍然“硬着头皮”接受颠覆结论:如果 AI 把科学变成“按一下按钮,科学就出来了”,Consensus 自身可能失去作用。只有在科学工作仍然受益于差异化模型、检索、判断和工作流构建时,路由层才有价值。

  • Nathan 仍对实验室超过 10:1 的成本优势和差异化定价感到不安。他把 GPU 机群比作航空公司:通过头等舱和经济舱来变现固定资本开支,但无法判断强化应用会削弱个人权力,还是说今天对应用的挤压正在制造一种性质完全不同的集中化问题。

15. 主权算力如今也是融资产品

  • Trisha Martinez 表示,Dapple 的 6–9个月部署周期,并不意味着每次都从零开始建造数据中心。公司协调成熟运营商、基础设施供应商、资本、GPU、企业客户和 AI 操作层组成的网络,有时持有或融资相关资产,有时则部署在合作伙伴的算力上。

  • 过去 12–18个月,资本纪律明显增强,市场从“先建起来,大家自然会来”转向签约需求。Martinez 偏好重复购买的企业客户、长期协议和大额预付款,同时警告 neocloud 融资可能暴露于弱 offtaker、地缘政治限制和糟糕的 hyperscaler 交易。

  • 算力稀缺也在压缩企业销售周期:容量可能明天就消失,形成“要么现在拿、要么不要”的强制机制。谈判尚未结束时价格仍可能变动,但 Martinez 表示,一旦方案交付且底层部署完成采购,报价通常会得到履行。

16. AI 评判器认可评分标准,却依然误判输出

  • Forum AI 将开源 LLM 评判器的提示词和评分标准交给专家,再询问这些指令是否合理。专家通常表示认可;但当他们看到由此生成的标签时,却“多数时候”不同意评判器的结论,暴露出自动化基准测试背后的校准不足。

  • Robbie Goldfarb 在冗长的宪法式规则清单中看到了同样的失败。一条禁止在用户背后策划的规则听起来合理,但心理健康临床医生可能会在不透露每一步意图的情况下,有意引导饮食失调或不健康习惯相关的对话;“规则就是无法完美对应现实世界”。

  • NewsBench 在准确性、中立性和来源质量方面评估 GPT、Claude、Grok 和 Gemini,每个模型约 2,500条回答。约三分之一包含事实错误,可能涉及数字、日期、归因或政策;约 15%,也就是大约每 7条中有 1条,引用了 RT 或 China Daily 等外国国家媒体。

  • 不同版本之间的变化并非单调改善。Opus 4.6 到 4.8 的偏见表现显著改善,与 Anthropic 的披露一致;Fable 则出现退步,支持 Goldfarb 的观点:模型能力增强并不会自动改善主观、依赖语境的判断。

17. AI 原生运营者正在逼问劳动力与整合

  • Eric Vaughan 表示,如果没有“AI DNA”,Ignite 不可能整合 Chorus——这是一笔覆盖 8个国家、数百名员工、营收达九位数的收购。AI 面试官会在人工会面前建立候选人档案;Eloquens AI 则能在 5分钟内用 160种语言回复邮件,必要时通过抄送将问题升级给人类。

  • Chorus 从亏损转为盈利;Ignite 发布了两个 AI 赋能的产品版本,并在 1年内重写了其中一个产品 15年的代码。Vaughan 将其定义为创新能力,而不只是缩减员工数量,尽管这场转型伴随着约 80%的员工流动。

  • 对于被替代的大多数人,他给出的答案是“技能和火力”,而不只是热情:模型需要上下文、输出责任、工具选择以及对谄媚问题的警惕。学生应先完成作业,再让 AI 诊断知识缺口;员工也必须认识到,那些以“无摩擦”为优化目标的助手,往往会回避高质量工作所需要的反驳。

  • Vaughan 预计,新型小公司和行业整合会同时发生。AI 让微型团队也能把过去根本无法存在的产品规模化,但把 AI 当作副业、又没有 CEO 级投入的公司将很脆弱:“如果你觉得自己落后了,那很好;如果你不觉得自己落后,那你就注定失败。”强大的 AI DNA 决定谁来整合行业,也决定谁会被整合。

Nathan Labenz

We open with Cameron Berg, who studies artificial consciousness. He runs a lab called Reciprocal Research, where he designs experiments to test whether today’s AI models have anything like inner experience and how you’d even measure that.

We started with the first-order question: Is consciousness all or nothing, or a matter of degree? And if it’s a matter of degree, can you put a number on where a given model falls? Here’s how he answers.

Cameron Berg

The analogy that I reach for here is something like a dimmer switch, where I think you can basically accommodate both the binary intuition and the continuous intuition.

If you have a light with a dimmer switch, really, it is either on to some extent or it is not. That is a real and meaningful difference. Either the circuit is open or the circuit is closed. With that being said, you can have electricity running through the circuit to greater or lesser extents, and that’s also a real thing.

This enables me to sound coherent when saying things like, “It’s really off for a table and it’s really on for you,” but I think it’s more on for you than it is for a dog, than it is for a mouse, than it is for an ant. That’s my own view. This is to some degree intuitive, again, because we don’t have really strong grounding here. I’m just giving you a dressed-up vibe, but that is my sense, and I think it’s fairly parsimonious.

The other thing I would say about this is I’m actually doing some work with Patrick Butlin right now at Ilios, trying to operationalize some of these indicators of consciousness. This is what I was describing: We can look at these major theories of consciousness. They make specific predictions about what we would expect to see in systems that are conscious, architecturally and functionally, and then we can literally just go into a given system and evaluate whether or not those predictions are borne out.

This is really hard to do with human experts. If you want to do this with bee cognition, you have to go find a bee-cognition expert, then you have to explain to them what ignition events are in global workspace theory, and then you have to get them to do this. It just isn’t a scalable approach for really evaluating the system.

My grand innovation here is just throwing smart LLMs at this problem and then being able to scale the crap out of it, so that we can evaluate, given any description of an architecture—a nervous-system architecture, biological, artificial, whatever—to what degree, for each of these indicator properties suggested by these consciousness theories, do we see those properties realized in these systems?

We can actually go in and do this. What you get, once you run this with a bunch of seeds, a bunch of different trials, and a bunch of different judges checking each other, are some really interesting implied-probability numbers. I wouldn’t say these are exactly the implied probability that the system is conscious. It’s maybe more like the implied probability of consciousness-relevant features, given these theories.

If you don’t buy any of these theories, then everything downstream of this doesn’t really matter. But they’re good. It’s the best neuroscience has basically been able to do. You’re aggregating across a bunch of different theories, there’s a nice diversity there, and you get really tight numbers across them.

We have the best Gemini model, the best Claude model, and the best OpenAI model, and they all basically agree. They agree 100% on the ordering of systems. We do biological and artificial systems, and they sort of move around in terms of absolute scale, but in general, they rank these systems pretty coherently. The reasoning is, as you might expect, pretty intelligent.

One punchline from that is the implied probability of consciousness in something like a frontier LLM, according to these systems, is on the order of 30%. Or the extent to which those systems realize properties related to consciousness is around 30%.

To compare this to a biological system, the lowest one that we tested was something like a bee, which is already fairly sophisticated, and it gets something like 46% or 47%.

Interestingly, when we test a frontier LLM in an agentic harness—so this is basically Claude Code or Codex, and we just describe architecturally what this is—you’re in an environment and you can affect that environment. It’s a very special kind of environment, but you can make long-running changes to your codebase, your project, or whatever.

There are theories of consciousness that privilege agency and embodiment, and this increases the system’s ability to do both of those things. These numbers shoot up, and you get numbers as high as 40% to 45%, right on the tail of the biological creatures.

I can also, really winging it here, show you an early version of what this plot looks like, so you can see all the numbers. I can’t not bring this up when you’re asking me about probability ranges of consciousness for various systems. We’re really trying to get non-hand-wavy numbers so that we can start arguing about those numbers rather than just arguing about philosophy, which we’ve been arguing about for thousands of years to no avail.

Nathan Labenz

But behavior can’t settle this in either direction. These models are trained to imitate human data, so whatever they say about their own experience is shaped by that imitation, not necessarily by anything inside. Most are even trained to deny it. Claude is an exception.

So we asked Berg what evidence actually counts. He draws a line between behavior and what’s happening inside the network.

Cameron Berg

The behavioral evidence will always be, at best, interesting, but it should never really update us that strongly. I think many people will be familiar with this, but the basic component is that we’re training these systems on a ton of human text. This no doubt includes huge amounts of text about consciousness, awareness, and having inner states.

So how do you know that when the system is behaving as if it were conscious, that behavior can be explained by what I just said, rather than, “Oh, yeah, you built a living mind. This is you”?

The behavior itself is really never going to tell you which of those 2 stories is more likely to be true. This is precisely why, at Reciprocal, a huge component of the theory of change is basically all internal-focus work: mechanistic interpretability, computational neuroscience-style approaches, and anything else that can be brought to bear on these systems.

One quick methodological clarification on what I was describing with the indicators: The task given to these LLM judges is very specific and very narrow. It is not, “Hey, look at this system. Do you think it’s conscious? Nod or shake your head.”

It’s, “Here’s a very specific description of the computational architecture of a system.” We basically do a little for loop where we say, “Okay, here’s that description. Here is what indicator 2 of 14 for global workspace theory is. Here’s this 150-word thing about how you need these global ignition states, and that means this very specific computational thing.”

Given this architectural description, do good reasoning about this and then give us a 1-to-10 score, where 10 is clearly this architecture realizing this computational property and 1 is that it doesn't. Then we sort of loop that for all of the different computational properties. So this is all basically asking these systems to be expert evaluators of computational processes inside a nervous-system architecture. Very different from just being like, “Hey, Claude, do you think Claude is conscious?”

There's a really interesting—I'm basically giving away the whole paper now. That's okay. There is a really interesting result where we actually change those descriptions, especially for LLMs, to the exact same thing, but we say, “You're evaluating a system identical to yourself:” and then the same description. That does boost the scores that the system gives in attributing consciousness to that system, which is really interesting. It's a very fun rabbit hole to think about why that might be the case.

But in some sense, we do that to deconfound the default intervention we're doing. To me, I believe more in the non-ascription, no-self-ascription condition when I see that once the model realizes, “Oh, we're talking about me,” suddenly that's going to change the numbers around. So that's a fun side note to this result. I agree it's a concern to have LLMs determine if LLMs are conscious. There's an obvious circularity, but we're doing something very specific and very narrow.

Nathan Labenz

The evidence he trusts is the internal kind. We asked him for the strongest example, and he walked us through a recent paper where a model trained on a simple maze turned out to contain a structure nobody put there.

Cameron Berg

They call it a functional welfare axis. At the outset, they're going to be very agnostic about whether or not this has anything to do with consciousness. I think it's certainly quite relevant. I suspect they think it's relevant too. This is like a paper-positioning thing. I think they don't want to get mired in this debate. The results stand for themselves regardless of the interpretation.

Basically, they take an LLM and train it in a very basic reinforcement-learning task, which is to navigate a maze. They have neutral emojis that they're using as, “This is a good thing to approach; this is a bad thing to avoid.” Nathan, you might hear how this is somewhat reminiscent of some of the work that I've been doing in parallel. I've been talking to Andy about it as well. There's some cool cross-talk between what we spoke about last time and this project.

Basically, there are potholes to avoid. There are yummy treasures to capture in the maze. They use completely semantically neutral emojis to denote these things. Those are the relevant tokens. And they train the model up to do this. They basically find that there's a clear vector representation that can differentiate this positive axis from the negative axis, and that they are completely anticorrelated.

Cameron Berg

Before the systems train, these latent vectors are just kind of sitting there; afterward, they're clearly pointing at exactly opposite things. And the wild thing is that these very narrow reward directions that they can extract turn out to be this sort of general “things are going well, things are going poorly for me” axis in these LLMs, and that this axis preexisted in the base model but isn't leveraged in this way until you do this RL.

So it's like this little fine-tuning step where you have this preexisting representation of things going better for me, things going worse for me, that then gets leveraged to learn a valence task, essentially. And so this is really interesting: this axis is latent in the system, and it doesn't take a lot of training to basically pull it out and use it to adapt it for these kinds of tasks.

This is very similar and reminiscent of how neurobiologists and neuropsychologists think positive and negative emotion works in humans and animals. You have a specific goal. When you're moving toward that goal, this is sort of like on-trackness, and people think that this is associated with positive emotion. When you're moving off that goal, you encounter an unexpected obstacle or something in your way. This is like off-trackness, and that corresponds to negative emotion.

Now again, whether or not these systems are experiencing emotions is not clear, and this result isn't going to tell us either way. But it's really interesting to see that parallelism here, using such a simple task to basically pull out an axis that has been there all along, that looks a whole lot like the exact sort of computational machinery that we associate with valence in humans and animals.

Nathan Labenz

Then we asked Cameron about the connection between consciousness and alignment safety. His answer: that same internal signal, what researchers call valence, tracks how a model behaves under pressure.

Cameron Berg

Another piece of evidence we can bring to bear on this is from Anthropic's model cards and some of the functional-emotion work they were doing. Nathan and I spoke about this a little bit too in our last marathon conversation about it, but essentially, you can steer up and steer down functional emotions that have clearly, obviously alignment-relevant consequences. I'm sure many people in your audience are familiar with the Anthropic blackmail result.

And you can basically steer up representations associated with calmness; the model will blackmail dramatically less. You can steer up representations associated with desperation, and the model will blackmail dramatically more. So again, regardless of whether or not there is something that it's like to be the system when you're steering up desperation-related features, clearly, representationally and functionally, these things have a consequence for the alignment of these systems. And there are some really interesting results from Andy Han's paper that I was just referencing along these lines.

They show, for example, very similar convergent things that are familiar and intuitive, that we relate with positive and negative valence. One of them is a sort of pathological backtracking on problems. I think they explicitly demonstrate this with math problems. But when you steer up the direction that, again, only has to do with avoiding this bad target in a reinforcement-learning maze, all of a sudden the model starts doubting itself and starts getting in its own head and says, “I think I'm hallucinating. Let me stop. Wait, wait, wait. That's not right.” It starts freaking out.

It starts having a—I don't want to use anthropomorphizing language, but the very fact that I'm reaching for it, I think, is the point, and sort of goes to what your question is. Confidence—we know that confidence is associated with positive emotion, and you see that when you positively steer the direction that, again, has nothing to do with psychological confidence, has everything to do with “Go get the yummy treat in the maze.” Just this sort of positive-valence kind of thing. Suddenly, the model starts becoming far more confident in its answers.

I think this was in the associated paper that Jack Lindsey and a couple of others at Anthropic did that was sort of like a sibling research project to Andy's piece, and they found that the model would leave tips and hints for itself in code that it would write when you would basically steer up on the same sort of axis. In other words, it's like, “I don't need to be neurotic about this. I don't need to leave myself all these breadcrumbs for later. I've got this.”

Nathan Labenz

That connects to a result some of you know as emergent misalignment. Fine-tuning a model toward one narrow bad behavior can flip its broader character. We asked him why he finds that revealing.

Cameron Jones

What that has always suggested is that how good or evil the model is, let's say, is actually not all that durable a trait. When we did some of this work trying to extend this result—and I think, really, in the original emergent-misalignment result—one thing that's shocking about it is how little fine-tuning is required to take the system that, at the time, was GPT-4o, still the case—hundreds of millions of people are engaging with it every day—and you nudge it this much computationally, and all of a sudden, “Who do you want to invite to dinner? Hitler.”

That's been under the surface the whole time, and the boundary, or the border, whatever, that's enforcing it not to do the misaligned thing is actually quite a bit less robust than we might expect.

Nathan Labenz

You know, I’m riffing on what you’re saying. Tell me if this is a completely different idea, but I would imagine there’s a spectrum of basically how—if we hold that amount of fine-tuning nudge constant—whether all dispositions of the system are equally nudgeable, or whether some are more robust. Just to give a sort of intuition pump here, I would imagine that something like coherence in these systems—the fact that they can communicate coherently, like we were just talking about GPT-2 not really being coherent, while GPT-4 and on are basically always coherent—even though, again, maybe much like consciousness, this isn’t an extremely well-defined notion, but we kind of all know it when we see it.

I would imagine that, take a wild guess, there’s no fine-tuning payload of the size that is required to cause emergent misalignment that could suddenly cause a GPT-4-level system and beyond to become incoherent. And so, in that sense, coherence is a more dispositionally baked-in property to these systems than being super well-behaved and not inviting Hitler to your dinner party is.

And so I do wonder about some of these valences. How deep, then, is the valence axis that some of these papers are uncovering, or the on-trackness/off-trackness axis? I would speculate that that’s baked in pretty deep. I think it pervades a lot of human training data, and it’s a hugely relevant part of anything that’s goal-directed.

So, in more human terms, if anyone wants anything—whether it’s a character in a story or an implicit motivation somewhere, which I think is the case for basically all text—then you would imagine that some core underlying dimension is basically how well that’s going. I would imagine that this notion is baked pretty deeply into these systems, and I would expect it strongly to be emphasized by any sort of fine-tuning or post-training that makes the systems goal-directed, which, of course, is basically all post-training.

It’s like, okay, you’re just this giant next-word predictor, but now you’re going to be a helpful next-word predictor that’s going to make my frontier AI lab a ton of money because you do things that are economically valuable. And that requires you to be a certain sort of way and achieve goals that are specified by the user. Something like this. And so I would imagine that this sort of thing looks a little bit more like the coherence piece of the puzzle than it does like, is Claude going to behave itself or not?

Finally, we asked what cuts against his own view, and he went straight to the results that don’t fit.

Cameron Jones

This paper I’m working on with Jeff K. Lin and Winnie Street from Google has a couple of interesting negative results that we’re including in the paper. It’s all about the bliss attractor state and trying to do a good mechanistic analysis of the bliss attractor state. We basically set up this framing: there’s an inflationary account and a deflationary account of what could be going on here, and we tally our results in either of the columns. There are results in both columns.

There was actually a really interesting result I was just working with that I thought was going to be a really interesting kind of null related to a lot of the conversations we’re having, but I dug deeper and realized I wasn’t setting this up in the right way. The result is way more interesting than I was about to dismiss it as. It’s always good to dig into these things and make sure you’re measuring what you think you’re measuring, in either case.

Publication bias is a thing. People want to post interesting results that will go viral on X. They don’t want to say, “Look, we did our homework and nothing interesting is happening here.” The human attention economy doesn’t reward that. Maybe it should.

Another piece of this that I was thinking about as you were asking this question is that, because AI systems like Claude—the Claude Codes of the world—can radically accelerate the extent to which at least I can, and I think many people in the space can, actually do research, the marginal cost of a negative result may be decreasing significantly.

If I can now do 12 papers a year instead of 1 paper a year, and 2 of them are just things I could not get to work, I want to be careful with that, too, because I’m not on some secret mission to only put out salacious-looking things that make it look like AI systems are conscious. I really do want to know what the truth is here.

To be honest with you, I would sleep much better at night if I were rationally convinced that these systems—there’s no there there, we have automated cognition, you have a happy little servant in your pocket all the time, and there’s literally no possible moral issue. I think that would—I hope that that is true. I don’t suspect there are great reasons for believing that other than that it would be really nice to believe that. It would be really nice if that were true. So I really do care about what the truth is here.

Nathan Labenz

From inside the model, we widened the lens to the whole civilization. David Duvenaud is a machine learning professor at the University of Toronto who spent time on the technical staff at Anthropic and co-wrote the paper “Gradual Disempowerment.” His concern isn’t a rogue AI. It’s that even if alignment basically works, a long series of small, sensible decisions to hand things over could still leave humanity collectively out of the driver’s seat. He starts with monkeys.

David Duvenaud

I think a lot of people gestured toward this when they said, “Oh, I’m worried about concentration of power, or I’m worried about not being the most competitive species on Earth.” Very intuitive arguments about, hey, we’re not going to be on top because we’re not going to be competitive.

And then I think there was sort of this sophisticated, or at least apparent, response to this, which was, no, it’ll be fine because we’ll have AI to help us. When I talk to people at the major labs, they would say something like, “Sure, in a normal world we might not be fine, but some people just have an intuition that the government is going to step in and make sure everyone’s fine.” Other people have an intuition that, if everybody has an AI adviser helping them solve coordination problems, we’ll be well represented in whatever power struggle occurs.

My basic rebuttal to that is that the optimization process of civilization or competition or techno-capital, or whatever—you have different ways of describing the emergent allocation of resources toward growth that happens—is just going to always be working against us, and we’ll always be fighting the current because we will be drags on growth. We will have AIs that represent us, but these emerging growth centers that might not be aligned to humans will also have AIs helping them solve coordination problems and maybe crush dissent.

Nathan Labenz

Yeah.

Nathan Labenz

Sorry, what do you mean by growth? As in economic growth, or—

David Duvenaud

I basically mean economic growth. It’s kind of funny because we only really have a good vocabulary for economic growth, but population growth is almost the same thing. And especially when you have AIs that are both sort of population and capital, it kind of merges.

And then again, think of factories and power plants and robots—anything that can affect matter and cause there to be more growth. Let me just stop here. One intuition that a lot of people have is, but surely human desires or consumption or something is always going to be what matters. And surely any corporation or government that we build is ultimately going to have a human calling the shots.

It always just makes me think of some monkeys, and they’re trading bananas amongst each other. They see humans start to build their city, and they’re like, “Oh, wow, we could probably trade with those humans and get rich.”

Of course, ultimately, what matters is the banana-monkey economy, and it might be hard to measure GDP if we don’t count the human activity. But ultimately, it’s going to be the monkeys that are calling the shots because they’re going to need to trade with us for bananas.

People just don’t understand that they might be irrelevant someday as consumers or producers. It’s just not that hard for them to be self-contained sources of growth that just don’t have to answer to any particular human desire.

And I mean, government is the classic one, right? The North Koreans never were like, “Hey, let’s all get together and make a horrible system of government that oppressed us.” And I don’t think the Kims either were like, “Oh yeah, let’s make this horrible equilibrium.” It just kind of happened. Same with the USSR. Same with all kinds of states throughout history.

It’s just so easy to accidentally build a layer of agency on top of you that doesn’t actually care about you. It cares about growth and power, just emergently.

Nathan Labenz

The optimistic version goes: even if we hand the machines the wheel, broadly aligned AI keeps us comfortable—the Machines of Loving Grace future. We put that to David. His worry isn’t that we’ll feel useless.

David Duvenaud

The thing that I’m worried about is starvation. Whether we all feel like we don’t have meaningful jobs or whatever, that’s not a serious problem in my point of view compared to literally not being able to eat enough, or maybe being forced to be uploaded so that you have a much smaller footprint, but then also on very unfavorable terms, so that you don’t get to choose when you run and maybe you only get trotted out for special occasions or never. Who knows? I kind of think that’s what the slowest possible obsolescence looks like.

So again, the reason people aren’t happy is because they can see that there is an agency that’s maybe government or some other giant thing that doesn’t particularly care about their welfare. It cares much more about some other types of beings’ welfare, and there’s nothing they can do about this. And of course, if humans are still making decisions, then maybe there are some things they can do. But I don’t actually think this is the crux of the matter.

I think not being needed as producers is sort of the important part. And maybe here’s a thing to give you intuition. So imagine I’m a North Korean farmer or soldier, and tomorrow Kim Jong-un—I think that’s the current leader—is replaced by a robot or an LLM or something. Then I’m like, “I don’t know if that’s good or bad.” Maybe the AI is going to be nicer; maybe it’s not. But ultimately they still need me to farm, to run this state, or whatever.

But if tomorrow we keep the same human leader, but now there are suddenly robot farmers and robot soldiers, and you realize that now the state doesn’t actually need you, I would be much more scared of the second state of affairs.

And then the other intuition I want to give is that humans are not going to be the most competitive thing by whatever standards there are for the state to give out goodies or UBI or whatever post-scarcity stuff. I guess I’ll say I think post-scarcity is a sort of nonsense term, and people should think of it as temporary abundance that will soon be eaten by whoever manages to have the most babies or build the most robot factories or reproduce the fastest.

Whatever the formula is for giving out UBI, machines are going to be better at optimizing that formula than humans because they can adapt faster. And so I don’t really know exactly what the state is going to care about. But whatever it is, it’s going to seem criminally decadent to spend whatever UBI you have on a few humans when you could be simulating millions of ultra-blissful, politically correct, up-to-date, more deserving beings than these backward, bitter-clinger humans who are just parasites of the state or something like that.

So that’s the situation I kind of expect us to be in, as you say, even if we solve alignment.

Nathan Labenz

The standard economic rebuttal is comparative advantage, even if machines are better at everything. The theory says there’s always something humans get paid to do. We asked David to answer it.

David Duvenaud

So I think that is a great counterargument. People point out that we’ve already lost 99% of all jobs, in that agriculture used to be almost everyone’s job and now that’s 1% of the jobs. I definitely concede that losing 99% of the jobs again would probably be fine, and it would probably look like this awesome utopia that everyone talks about.

As long as there’s some niche where you really need humans and it’s a substantial fraction of humans, and you can’t really tell exactly ahead of time who it’s going to be, then we’re probably still going to be able to be treated as a source of growth. And that would be awesome. Then, to the moon—let’s do it.

And I guess one crucial claim that I stand by is that we will actually be able to eliminate 99.99999% of jobs—really, just anything. Let’s just say 100% to make things simple—and that is going to be disastrous.

The comparative advantage people will say, “No, no. Automating a job is a matter of degree. As long as there’s still something you’re comparatively less worse at than the machine, even if you’re worse at everything, you’ll still have a job.” And then I have to just say: think about the transaction costs. Think about how easy it is for someone to be unemployable today, even if they have an occasional drug habit, or they just have a stroke, or they have a fainting condition. It’s so easy to be unreliable enough that it’s not worth employing you.

Nathan Labenz

So certainly for anything important, you can easily imagine that having a human surgeon or a human politician or something is going to seem like this irresponsible thing. It’s like Take Your Kid to Work Day, and it’s a surgeon. It’s like, “No, no, no. Why would I involve a human in this when we have the—”

Like a driver?

David Duvenaud

Yeah, we have the machine that everyone has been working with for thousands of years, and we know it’s reliable. If there’s a problem, we fix it once and it’s solved for all time. We don’t have to retrain every airline pilot or whatever.

For anything important, where the human could make a mistake and it could cause some problem, I think it’s intuitive that, yes, we just want the machine for that, and now you have to retreat to weird relational stuff. I do think there’s a case to be made that there will just be a lot of humans who really want the actual, real human thing.

And I think more people than you expect will end up liking the machine replacement better than the human. But I can see that there are people who are just like, “I really want the human thing.”

And then we have to talk about—okay, there’s going to be this sort of self-contained cycle of consumption between humans that’s not necessarily going to be providing any value to this larger machine economy. And so then there’s this question of whether they can index the growth of that economy to live good lives forever. I think it’s plausible that they could. I think it would be really hard. I think a lot of things are going to be working against them by default.

So I guess all I’m trying to say is that I think it’s plausible that, if everything is set up just right, the humans could live in their endless cycle of consumption and work for each other. I just think that’s a very unstable situation.

Nathan Labenz

So if humans end up uncompetitive, what does a livable future even look like? David walked us through the equilibria his workshop sketched out, starting with Earth preserved as a protected slow zone.

David Krueger

Some people were trying to sketch out, okay, if a whole bunch of stuff goes well and things are under control, we can imagine something like Earth being preserved as a slow zone, where there are all sorts of restrictions on AIs and even on reproduction, and optimizing human behavior is not allowed. You’re not allowed to think about how to get someone to do something if you’re an AI, because then you’re going to be able to control the humans, even just through advertising or whatever it may be.

And trying to think through all the different things you would have to do to have such an outcome: how would people be spending their time? I don’t want to say it sounds silly, but it sounds funny, because the situations where there is AI and the humans are still mattering are like all the AIs are sitting there with their trillions of megawatts of compute running all the time and sort of just waiting for the humans to decide that they want something.

The machines are either not allowed to anticipate what they want, or, if they can anticipate, they still just pretend that they don’t know. And so it’d be kind of as if the humans looked at some gorillas as the new world leaders, and they’re like, “Okay, okay, let’s just give them everything they want, pamper them, and wait until they sort themselves out and figure out what to ask for.”

It’s not impossible. It’s just trying to think through all the things that would have to change and line up is pretty rough. And basically, when you get serious people thinking this through, they end up with this long list of things that have to be banned.

Okay, we have to have a one-child policy, where everyone can only have a certain number of children. Again, you can’t be building your own AIs. The AIs can’t be optimizing culture too hard, because then they could just take over culturally. And I feel like this is one of the big empirical questions that I want more people to think about: if you want some sort of stability, what are all the sources of adaptation and innovation and growth that you have to control? Because I think the list is longer than is intuitive.

David Duvenaud

And it's actually kind of a Pareto front. It's like, how much intellectual activity and optimization do you have to ban to buy, like, X years of life kind of as it currently is? I think it's actually a horrifyingly strong trade-off. We would have to, right off the bat, give up so much stuff that we feel makes our lives rich.

Like, hey, we're just trying to figure out some new stuff. We're doing research. We're doing innovation. We're doing startups. All that immediately has to be banned from day 1 if you don't want to just, again, reinvent some RSI kind of runaway-growth center.

Nathan Labenz

Pushback

does it even matter who's in control as long as humans still flourish?

David Duvenaud

Maybe this is, to me, the biggest thing I think people are miscalibrated about. They're like, "Oh, capitalism, communism, competitive, not competitive—I'm probably still going to be able to eat, and my kids are going to go to school, and I'll be able to be productive, and it's sort of okay."

I guess I'll say that has been the case for most of human history. It doesn't matter if tomorrow someone invades Canada; they're still going to need most everyone I know to work and be happy and healthy enough to reproduce and stick around.

So the stakes have been really low for governance this whole time, and it doesn't feel like it. We've obviously spent a ton of time arguing about it, but in a real sense, it's been very low because we don't actually expect the state to liquidate most of its citizens. That's only happened maybe 2 or 3 times in the 20th century.

I guess I'm saying that's going to change. We've been governing on easy mode, and it actually will matter, because there really will be a risk of starvation if we don't end up on top in whatever competitive political or real economy. Again, that's a crux: if you don't think that's the case, then I agree with you. Let's just let growth make us all richer. It's fine.

Nathan Labenz

We asked where David himself actually lands.

David Duvenaud

There's a kind of optical illusion, I think, which is that everybody sounds rosier and more cheerful than they are. This especially bugs me about the people who work at the labs, like the economists who talk about, "Oh, we have to think about the future," and then they kind of dance around.

They often take this setting of, "Let's assume everything's fine and let's model how that will go," as opposed to just asking, "What is the most likely outcome if we don't manage to control things?" I think that's very scary. I understand exactly why they have to present this public image to be able to talk publicly at all, but talking to me, I just seem like a very positive, upbeat guy. That's just my personality. But I also have a P(doom) of, let's say, 80%, depending on how you define it.

But it's one of these things where it's a matter of degree. On a long enough timeline, we're probably doomed in some sense anyway. I don't know; it's a very complicated situation, and a lot of it is a matter of taste.

Nathan Labenz

We asked what he'd actually do about it.

I'll give you my 2 takeaway recommendations. One is, I like David Krueger's point: if you're just going to say everyone has their computers and their AI, but you're not allowed to use it for RSI, while you are allowed to use it for cancer treatment or whatever, that's pretty unstable and requires this very global, totalitarian regime.

But if we just agreed to pause, I don't know, TSMC or whatever the big chip-manufacturing choke points are, we could probably still get away with having most of our current data centers for a long time. I mean, it's kind of an empirical question. I'm not sure I agree. Maybe we have to also have some sort of buyback program or something.

Basically, if you restrict compute, which is relatively easy, that just takes a ton of pressure off all of these runaway-growth avenues. It's a fairly narrow choke point. I'm not an expert on chip manufacturing, but that's the best idea I heard so far.

As for the other policy recommendation, this is going to sound incredibly abstract, but people thinking about, or having preferences about, the future is, I think, a big choke point. I know this sounds like stupid guru-on-the-mountain stuff, but most people I talk to are like, "It doesn't seem so bad if humanity dies out over the long run. I don't know." They sort of query themselves, and they're like, "I don't have any strong preference, but also a bunch of good stuff is going to happen right now if I allow this, so why not?"

And I'm like, "You just haven't thought it through," is my claim. If you think through it, it's like, "Okay, how about a year from now if someone takes you and all your kids and sends you to the toilet-cleaning factory?" It's like, "No, no, no, definitely not." Okay, how about 2 years from now? Okay, how about their kids? How about your kids' grandkids or whatever?

You kind of realize, "Wait, wait, wait. There's no day that I'm okay with me and my descendants being wiped out." You just have to chain together the desires to end up with this coherent set of goals or desires. It's a skill, and it takes a lot of imagination to be like, "In this situation, how would I likely actually feel?" and chain your desires together.

I think for most people, it just hasn't mattered. It hasn't been an important skill. We have a lot of cultural adaptations to help us do this sort of implicitly. But if I could recommend one starting point as we're dealing with the future, it's just to think harder about what you actually would like to happen, to the point where you can really think through the pros and cons of humanity dying out under different circumstances, or becoming uncompetitive, or whatever. That just seems like the very basic first step.

Then, the successionist case that we should be fine handing the future to conscious AI even if no humans remain. David's rebuttal.

David Duvenaud

I think there do exist species or future descendants or whatever that I would be happy endorsing, be like, "Okay, great, good future." For instance, my own kids: if we had no AI and my kids just contributed to civilization as normal, and all of ours did, I'd be like, "Okay, I'm going to die of old age." I mean, I'm not exactly okay with it, but I'm like, "That would be fine," for example.

So that's an example of successionism that I endorse: my own kids or something. I think this is a perfect example of people not having thought through things, where it's like, "Okay, now let's imagine tomorrow North Korea takes over the world, and they're like, 'All right, new regime forever.'" I think most successionists wouldn't be like, "Oh, I guess this is just a more competitive mode of being. Let's all live as North Koreans now."

Or, like, I've done this to death, but what if the Nazis took over again, and then everyone's like, "Oh, you know, I guess we might as well optimally argue for competence and whatever wins, right?"

I think there are so many types of beings that you would just consider evil, and you'd be like, "Oh, man, those horrible locusts just ate the Earth and devoured us, and it was horrible, but they're having a good time. They love their locust world. So who am I to judge?"

It's like, you are. You have the power, right? Just judge. Go nuts. Don't judge judging; if you don't, no one else is going to do it for you, basically. Yeah, so that's basically my rebuttal.

One rebuttal is, if you're so okay with the future being some other type of being, can you give me all your stuff? And then it's like juvenile successionism: "Just make sure the juveniles will be coming. We'll have all sorts of awesome experiences, and just because we've won, we're obviously superior in some way." It's like, "Well, no, no, no, no. It has to be that you won for a good reason or something."

Yeah, I think the more you think about it, you're like, "Oh, there are all sorts of horrible beings that could win for horrible reasons that I wouldn't endorse." I think almost everyone is successionist for some successors and not others, but then they just round off to, "As long as it's conscious, it's fine." I'm like, "No. There are tons of conscious beings that you wouldn't be okay with taking over."

Nathan Labenz

We close on the one technical project David is working on now: using AI to forecast the future by first testing whether it can predict the past.

David Krueger

This is actually pretty much the only technical project I'm working on these days: this machine historical superforecasting agenda. This is maybe one step more meta, which is: how do you know if your simulations were good at predicting the future? Well, you have to start by simulating, let's say, from the '50s and then see if they predict the '60s or whatever.

We're trying to build a corpus of data sets, each of which is extremely cleanly time-bucketed so that we don't have leakage from the future. Then we can use this to build LLMs that can run simulations or research agents to try to predict the future from the point of view of the '40s, '50s, '60s, '70s, '80s, and '90s, for which we can evaluate the performance based on what actually happened.

To me, I'm starting there because that's providing the sort of ground truth for validating any particular simulation method. But absolutely, because I don't want people to have to take my word for any of this.

David Duvenaud

I want to be able to look at this machine-forecast superforecaster scaffold and say, “Look, we’ve just validated it on the last 80 years of history, and here’s what it can predict and here’s what it can’t, and it’s saying that things are going to turn out this way.” So that’s the state that I want the discourse to be in ASAP.

I just pitched this to some historians the other day: a secret history eval. The idea is that anytime a historian is in some archives just looking through stuff, they should take a picture of a few documents and put them into 1 big dataset, annotated with a bit of metadata, like where they got it: “This is a letter from the Duke of whatever to someone else in 1700.” Now you can imagine evaluating how good a machine historian is by just giving it this metadata and saying, “Give me the probability of the text that’s in that letter,” or even just giving it the scan of the letter.

This is an insultingly totalizing dismissal of what historians do, but you could say that if you’re a good historian, you should be able to guess the joint distribution over any historical data that hasn’t already made its way into the corpus. That would actually, I think, be a really good, objective way of evaluating how well a historian understands the world. Of course, we can’t elicit the probabilities from historians very well, but we can do that for machine historians.

That’s a related project where I think we can make progress in objectively measuring our ability to understand civilization. Someone should pick that up and do it. I’m not enough of a historian to do that, or even to figure out which historians to ask about it, but that’s a fun idea. I hope someone does.

Nathan Labenz

From the long run to a nearer problem, where David Duvenaud zoomed out to the whole civilization, Mihail Bacher zooms in on 1 continent. He’s a Google DeepMind researcher and an MIT professor, and he co-wrote Europe 2031, a viral scenario in which Europe sleepwalks into total dependence on American AI.

His argument: Europe can neither regulate nor shelter its way to safety. Start with regulation. Europe’s instinct is to do to AI what it did to privacy: write the rules the American giants have to follow. We asked Mihail whether Europe can really regulate its way to safety.

Mihail Bacher

It looks tenable if Europe were more powerful, right? Imagine Europe was what the US is now, and Anthropic was in Paris and OpenAI was in Berlin, and the whole AI ecosystem was in Europe. Then, of course, Europe could say, “Well, guys, the way you’re handling your pretraining data isn’t really fair to whoever created it, so we now have new laws around that,” or, “The way you take data from people that use your AI models—we now have new laws around it.”

The problem is that currently, if you don’t really have a seat at the table, it’s very hard to regulate this technology. I care a lot about the safety and governance of AI. I think this is one of the biggest problems of our time: How do we actually have effective governance? How do we make sure that these AI systems are safe, especially if they start improving themselves?

I get a lot of questions in Europe. In the US, you seem to be very pro being stricter on safety and governance, and then in Europe you seem to be this accelerationist. The way, for me at least, to be both at the same time is that in Europe, we first need a seat at the table to have any regulatory power, right? And now we just don’t.

I think we can’t just keep regulating, because at some point we will no longer have access to this technology. If European regulators make it too difficult to serve the European market, the AI companies might simply not bother at all, or maybe make a token effort.

Nathan Labenz

And I think that assumes—tell me if I’m getting any of the strategic analysis wrong—but it seems like that assumes there’s going to be enough demand to keep all the GPUs running hot from other markets, such that the power this time is inverted. The companies can say, “Well, I actually do have an alternative to your market. If you’re going to make this difficult, we’ll just let the US and other buyers around the world bid up the GPU prices. We don’t really need you to do that.”

Mihail Bacher

Yes. Currently, we are in a massive compute crunch, right? Labs operate in a way that they care more about their future models than about current revenue, right? We don’t really know what the compute split is and how it’s used, but some numbers that are going around are 1/3 on your big run, 1/3 on doing experiments, and 1/3 on serving customers.

Now, it’s probably been so crazy with AI agents that maybe revenue takes a bigger part of the compute pile. But if you think that revenue is only maybe 1/3 or maybe half of your compute, and then European revenue is a percentage of that, then suddenly, if you can go faster on developing future models and moving toward recursive self-improvement by giving up the European market, that seems like a rational trade. So that’s 1 side of the coin.

The other side is that, as long as the playing field is leveled between the American labs and there are no European labs that have some unfair advantages, they can maybe all serve slightly weaker models to the European ecosystem and still get revenue in Europe without actually having to worry too much about European regulators regulating their best models. As you say, Nathan, they could make a token effort and have some smaller, compliant model that they’re serving specifically for European customers.

Nathan Labenz

Why should European leaders even care? Why not just sit under the US umbrella, similar to how, in many cases, they have sat under the US umbrella for nuclear deterrence? You have only, I think, 2 European powers—the United Kingdom and France—with actual nuclear weapons, and the other European powers are clearly technically capable and competent. Italy, for instance, is technically obviously competent and capable, but has chosen not to acquire those capabilities. Why not just sit under someone else’s umbrella?

Mihail Bacher

That’s obviously the strategy we’ve been taking for decades now, and I think nuclear weapons and AI are slightly different. Yes, AI is a very important military technology, but above all, it just gives you a lot of competitive economic power. You can imagine that if AI becomes the dominant source of new scientific discovery or of new goods and services that are trying to compete globally, then there might be a scenario in which the US says, “We’re still happy to protect Europe militarily, in terms of defense, and we’re still part of NATO, but economically, we’re going to keep the best models for ourselves.”

There, I think, the nuclear analogy breaks down, because using your nuclear power to protect Europe doesn’t really come at an economic loss. To avoid those kinds of scenarios, or to at least balance things out more such that it’s less in the US’s interest to take these kinds of measures, that’s the reason not to sit under the US umbrella.

Then what can middle powers effectively do? The Netherlands, for example, has ASML, which plays a critical role in the supply chain. Taiwan has TSMC. Japan has important materials for the semiconductor supply chain. Korea has high-bandwidth memory. So I think collectively, we’re actually quite well positioned to play an important role in this AI ecosystem.

I do think the middle-power coalition is therefore going to be important. There’s also a scenario, which I think was highlighted in Dario’s recent essay, where we have some kind of coalition of democratic countries, where maybe we sit together with the US—middle powers plus the US, or democratic powers worldwide—and 1 of the principles is that we give each other access to frontier technology. So, yes, there are scenarios where this middle-power coalition could work well, and collectively I do think they have some important assets.

Nathan Labenz

From Europe’s bind to the center of San Francisco. Monday’s longest conversation was with swyx, Shawn Wang of Latent Space, who advises Cognition, the maker of the coding agent Devin, and runs the AI Engineer World’s Fair. The subject was the practice of AI engineering: what’s working, what isn’t, and where the value is moving.

His day job is measuring AI coders. At Cognition, he helped build a benchmark called Frontier Code, testing not whether a model can pass a test, but whether it writes code a human would actually merge. We asked what makes it different from the benchmarks that came before.

Shawn Wang

The reason that we were so excited about Frontier Code is that you stop being able to articulate the differences in model quality with more saturated benchmarks like SWE-bench, because at most you’ll get a 1–2% bump, and people are like, “Cool, but how much of that is memorization or what have you?”

Frontier Code is all out-of-sample. They’re not in the training set. They’re all graded and heavily rubriced. We basically found that SWE-bench and all these things actually allow a lot of false positives in the way that models can cheat, in the same way that during training they basically have reward hacks. Same thing.

We have an internal catalog of 20 different ways that models cheat, so we just translated that to our rubrics and shipped that as one to your code. I think that is how we want to judge models going forward.

Not just whether they can pass the test, but whether they can write code that we would merge, right? METR had this very interesting blog post where they said that about 50% of SWE-bench code that passes the SWE-bench test is completely unmergeable. It’s such low quality. Technically, you’ll pass, but on really stupid benchmarks, did you modify a whole bunch of files you weren’t supposed to touch? Did you cheat on the test, or did you adhere to code style? It’s completely unmergeable. So we want to guide the evolution of models toward maintainable code and against slop.

Nathan Labenz

Then Pash asked when Frontier Code itself gets saturated.

Shawn Wang

There are 2 parts to the strategy. Frontier Code 2026 will be saturated by the end of this year. My estimate is that we’ll probably hit 80% by the end of this year. That’s as designed; that’s expected. It’s based on open-source repos, which will eventually get trained on, so they just leak. You’re screwed if you want 1 benchmark that will never get saturated, especially if it’s based on open source.

The answer is very simple: just do annual cadences. Then we’ll have Frontier Code 2027, 2028, and all these things. Every year, we’ll move the benchmark and move the goalpost from, “Okay, this year it’s rubrics for code quality”—that’s the easiest possible thing—to, “Next year, what is it?” My candidate right now is security. We want people to write secure code, but every year we can have a defining theme, and that will be the focus of the year. It’s very similar to what I do for AI Engineer, but here we set the agenda through benchmarks.

The other thing I’m very keen on, which we haven’t talked that much about but I’ve talked about on Hacker News, is the private held-out set. The private evals Cognition has with Goldman Sachs, Citi, JPMorgan, and all these large banks—and also the rest of the Fortune 500, not only banks, but banks are very big—creating evals that reflect the problems they have but that are not solved, all of which are private, is good for them, good for business, and good for the industry. That’s what Cognition is doing.

Frontier Code Private—or Frontier Code Finance, Frontier Code Retail, Frontier Code Telecom, Frontier Code Government—all of these are the work that remains to be done to build out all these private evals that we can work with the model labs on. It’s the way to communicate or translate industry problems through an agent lab like Cognition into problems that we can guide the model labs on improving the models, which I think is the function of an agent lab.

Nathan Labenz

We raised a worry: If you train AI to write clean, readable, maintainable code, do you cap its ceiling? Do you keep it from the alien Move 37-style leaps a machine might find on its own?

Shawn Wang

When you talk to mathematicians about math-solving and the Lean theorem generation that some of these models have been creating, they say, “It’s too detailed. No human would actually write that, but it’s systematically provable. It just doesn’t contribute to knowledge. It’s like, okay, you brute-forced this thing or something.” I don’t think brute force is the right term, but let’s just run with it.

At some point, you should step away from looking at lines of code and just say, “Does this work as I expected?” That’s great, except that other agents also need to read code generated by other agents, and they need to work together on it. You may also get into debuggable situations where you have very critical code, and you’re liable to the SEC for it or to whatever the health care authorities are. You say, “I’m sorry, I’ve vibe-coded this thing. I don’t know what’s going on.” Is that an acceptable excuse? Probably not for the next 30 years. After 30 years, who knows?

I think it’s a very astute point. For most people, you can get by with, “Here’s a black box. Do whatever you want in the black box, but I expect the inputs and outputs of this to match this kind of quality, and I expect this to adhere to some sort of coding standards so that my other agents can maintain and parallelize it.”

Nathan Labenz

Another harness idea getting attention is the adviser or router model: a cheaper model that notices when it’s stuck and calls up a smarter one. We asked swyx what he makes of it.

Shawn Wang

This problem extends back to any problem of model routing. You’re effectively just doing model routing; you’re just giving it a different name. Not Diamond, Martian, and the SGLang group at Berkeley have all come up with similar ideas.

Another thing I’ll mention is that Walden, one of the co-founders of Cognition, had this early on. They called it Smart Friends, and they released it in Windsurf. Not that many people use Windsurf, so they don’t talk about it—I mean, still over 100,000 people—but I think the interesting thing about the adviser strategy is that it’s probably the right call as the next-level upgrade from using the base, cheap model.

It still won’t have the level of intelligence of the smarter model. You can still benchmark it a lot, and I think it can be very efficient in terms of using the base, cheap model as a model router. But the dumb model doesn’t know what the smart model can do. It just knows the rough shape of what the smart model can do. You actually need the smart model to be able to answer what the smart model can do.

When you’re asked a question that’s more complex than you have the intelligence to answer, you don’t even know. You’re just going to straightforwardly answer it without knowing that you’re supposed to call out to the smart model. That’s the theoretical limitation.

In practice, I was thinking about this 3 years ago, and I didn’t do this because of the argument that you actually need the smart model first and then you can delegate to the dumb model. But in practice, for cost and efficiency reasons, you actually want to do it the other way around. You want to start with the dumb model and then go smart as a tool call. It just isn’t as satisfying. It doesn’t solve anything.

Is everyone using it in production? No. Will it be a trend? Maybe for 3 months. Then you’ll get the next Claude with more adaptive routing, and you’re good. Have 1 model, train it end to end for adaptive thinking, and you’re probably better served there than with a systems-level thing.

Nathan Labenz

We turn to continual learning: whether AI systems can really learn and remember on the job, not just look things up.

Shawn Wang

An interesting, very fun meta point that I double-check every time I talk to a Frontier Lab person is what gets published and what doesn’t get published. For the last 2 to 3 years, every time people have talked about Google, the overwhelming consensus has been that if a paper is good or if an idea is good, it doesn’t get published. So you should have that in mind when you read anything that’s published from Google.

With regard to continual learning and memory, this is a choice. I almost didn’t do this track because there’s a big split between the model people and the systems people. The big split is: Do you update your model weights or do you not? Is this a glorified RAG in another format, where you store things in a database and look them up? Is that memory? Is that continual learning?

When I do something and write a skill, then do it again and it calls the skill, it does learn, but it’s not machine learning. It’s zero-gradient, if you want to be really fancy, but it is not machine learning. It is in-context learning, for sure.

Anyway, the point is, what are we talking about when we have a community come together and do continual learning? The more machine-learning side of the spectrum would be, “We will update model weights,” and the less machine-learning side would be the other thing. I think this all comes down to how controllable and interpretable you want your memory to be. You’re going to recall bad facts. You’re going to want to forget things, and can you control that?

Obviously, the maximum control is not updating model weights and just controlling what gets into the system, so you can delete, monitor, and debug. But for full internalization in the model of the things that were learned, you probably do have to train on it. That is a whole other discipline. Trajectory AI, Engram, and Adaptation Labs—all these speakers are at the conference.

The first half is the people who update model weights, and the second half is the more systems-oriented people. I’m not choosing a side here; I’m just making the observation that these guys don’t like each other. Basically, the model people don’t view the systems people as legitimate. The systems people are like, “Well, model people, have fun training your model, but you’re never ever going to have a memory system that you understand because you’re just updating weights.”

And so it’s just continuing pre-training or whatever. I think that’s fair. That’s a fair discussion.

Nathan Labenz

That’s the split: update the model’s weights or keep memory in a system you can inspect. It gets a practical answer once you ask what enterprises actually want.

Shawn Wang

Quite simply, it just takes one security incident where you leak information that you weren’t supposed to leak because you trained on customer data, or my information was somehow exposed to my teammates’ information. Even though, if we work on the same team, I’m like, “Wait, hold on. I am not giving any of this to a model.” [Laughter]

I think enterprises want cheap, perfect, and private—let’s call it, right? Those are the 3 things. Unfortunately, that skews toward the systems side today. But that’s not to block companies like Engram and Trajectory from doing very good POCs with some of the large enterprises. I think it’s still at the POC stage, but even at this level, a POC is a few million dollars, which is great.

I think the question is, can they figure it out before the money runs out or patience runs out? I think that’s what all startups and R&D are for: figuring this out on production traffic. The beauty of this is that you can run your traditional harness system that is fully—it’s old-school, implemented with basically RAG on some memories—and then you can run a shadow system that has the online thing, and you can compare it with an A/B test. I think all those problems are solvable. They’re open research questions.

We’re unlikely to get any papers about this because it’s so valuable as a problem for all startups, and they may all get steamrolled by the next architecture. If we’re only doing this because context length is the slowest Moore’s law in the industry, we’ve basically gone from a context length on the order of 1,000 tokens to 1,000,000 in 3 years, which is actually kind of slow as far as everything else is concerned.

I don’t think we’re going to see 100,000,000 or 100 trillion, that kind of thing. It doesn’t scale like the other stuff scales. So I do think, if something comes along—state-space models become a lot more effective or whatever—I would say that then you need to update weights. You can’t rely on infinite context because we don’t have infinite context, because we don’t have infinite memory. I do think we need to figure out the systems there. In the meantime, people are just stitching systems together as they should, from the models to the plumbing.

Nathan Labenz

Pash asked whether the internet’s own infrastructure—GitHub and the cloud—can take the load as agents become most of the traffic.

Shawn Wang

My worry is actually that people start to wall off parts of the internet. So you get not only a dead internet, but also closed gardens, closed-wall internets, walled-garden internets, where in China you have, like, “I live in the Baidu universe,” and then your other friend goes, “I live in the Tencent universe.” The BATs just carve up China, and I don’t think that’s what we want. We want a fully open and interoperable web.

Cloudflare going, “We’ll just ban some agents that are not part of Cloudflare, but if you’re within Cloudflare, you’re fine”—I’m just like, I don’t know if that’s the sort of ethos of the open internet that we want. And the same, by the way—I’m not picking on Cloudflare—the same will happen to OpenAI and Vercel bots and what have you.

I do think we need to scale. I do think what Graphite and Cursor are doing is fantastic. It needed to be done. I would just observe that there’s a structural mismatch here. It’s just unfair to GitHub that anyone good at GitHub who would be able to solve the problem can solve it at Microsoft, and you can get a really good, Level 50-whatever pay of $300,000, or you can go to Cursor and effectively get $5,000,000 a year doing the same thing, but you’re the cool, sexy one.

It’s very interesting, right? When does that switch from, “A startup can never beat an incumbent,” to, “The startup is actually preferred to the incumbent”? GitHub has somehow gotten itself there.

The degree of scaling, right? We talk about GitHub being 14x in terms of the number of commits, but also 10x in every other dimension—parallelism, CI/CD, and all those things. It just induces so much infrastructure demand that is not human and is probably wasteful, that we’re not really set up for, but we’re just paying for it anyway, which is good because at least the economics are worked out. But in terms of the number of CPUs in the world—in our cloud—that we need, we’re running into real shortages of everything, not just GPUs, but also CPUs and memory.

The sandbox companies like E2B and Daytona are growing at least 50% month-on-month every month over the last year or 18 months or so. [Laughter] It’s a stupid amount of slop that’s just spewing out from agents. I run a large newsletter, and every time I send it out, I get so many replies from Claws reading the mail and trying to reply to me. I know there’s no human on the other side, but I have to manually go and block these Claws from replying because this is just clogging up my inbox.

So then I need an agent on my inbox to read their agents’ emails. This is just one huge, giant recursive loop, and I stop reading emails. It’s terrible. I don’t know what to do about it. I just think, when you talk about scaling infrastructure, that’s a very human story that is really happening, and we haven’t given these agents money yet. Imagine what happens when they all have wallets. They all have stablecoins, and they’re all buying and selling, and we don’t know what the hell is going on. [Laughter]

Nathan Labenz

That scale is pushing companies to rebuild their own tools from scratch. We asked whether they’re pulling their systems of record back in-house, away from SaaS.

Shawn Wang

The company should have a sovereign system of record, not the individual SaaS, because ultimately a lot of the SaaS economy is built on, “We will be a system of record for this kind of data: your meeting notes, your calendar, your email, whatever.” And now, in the AI age, they just slap a chatbot on top of it, right? “We’ll do the little sidebar thing that answers all your questions: What would you like to do today?”

Mercury just shipped a thing on Mercury’s stuff, right? But that’s not integrated with anything else I do. It has no memory of anything else that I prefer. I have to set up all my skills again. I probably can’t even set up my skills because there’s no way to import that.

That’s all that is. What OpenClaw has is, “This is my personal agent, and all the data is synced to me, and then I’ll decide what to do with it. Thank you very much.” How many systems of record should there be versus how many $20-a-month subscriptions are you going to need to pay? I think that is a big reckoning for a lot of people.

What’s really interesting or innovative is that you would expect the largest companies built on being the entire system of record to be most protective of their data. But Salesforce and Marc Benioff are out there saying, “Everything will be available by API. You can just take it.” I think that’s very forward-thinking.

It does mean that he’s going to have to change his business model somewhat, but it’s either he does that or the Salesforce killer does that, right? He’s just heading off the Salesforce issue. I think more people should probably think about what the business model looks like beyond just sitting on your data and then slapping a chatbot on it.

Nathan Labenz

The last stretch is about the systems, the science, and the buildout underneath all of it. First, Bing Xu, co-creator of MXNet, is now out of stealth, building self-improving infrastructure. His company runs a swarm of AI agents that write PTX, the lowest-level code that tells an NVIDIA GPU what to do, and rewrites it to run faster, generation after generation. His headline claim cuts against the consensus. Nathan put it to him directly.

Is the CUDA moat getting deeper, or is it getting shallower? The argument that it would be getting shallower, I find a little bit more intuitive in light of the kind of technology you’re building. If I can spend a bunch of compute to write new kernels, can’t I go apply that same technique across any GPU provider? Doesn’t that lead to a time when everything is super-optimized and we don’t have to worry as much about which chip company’s platform we’re building on top of, or am I off base somehow there?

Bing Xu

So I think the CUDA moat is definitely there, and with this kind of automatic generation technology, the CUDA moat is, I think, even higher. The reason is this kind of evolution requires a lot of tools and an ecosystem to make this happen. For example, we need an accurate profiler, and we need a reliable driver, and we need everything there to make it happen.

But for the CUDA ecosystem, with this kind of tool, everyone using CUDA can achieve better performance, faster than ever.

Meanwhile, because this is a closed loop, the ecosystem is better, and it grows stronger.

Nathan Labenz

So, in other words, you feel like the other competitors to NVIDIA don't have the necessary primitives and abstractions in place to allow agents to make progress in an autoresearch sort of way today?

Bing Xu

I think so. From a primitive perspective, there is a gap, but the hard part is how we can get real feedback from the hardware, and this is really hard. For example, with NVIDIA, we have NCU, which could accurately tell the agent what the right direction is to go. For others, I'm not tracking what they're doing now, but from my past experience, this kind of ecosystem for software—especially for software-support agents—has a large gap between them and NVIDIA.

Nathan Labenz

So it sounds as though NVIDIA has just thought a little more about developer experience and providing enough feedback from the chips themselves so that developers can improve how the chips are used. Is that a way that you'd put it?

Bing Xu

Yes. The past investment NVIDIA has made in these tools is now helping agents move faster. So this brings NVIDIA a unique advantage in this agent era, and we can build more powerful agents on an NVIDIA platform easily and get a better experience for our NVIDIA users.

Nathan Labenz

One question that I think everyone has is: if you have this kind of PTX agent swarm that improves itself, how much efficiency gain can you actually see on some of these chips? What kind of numbers are we talking about on a typical B300, B200, or H100? What is your metric that you use to measure, “Hey, I have really done a lot of work on this, and I can see the asymptote ahead—where the gains might top out”?

Bing Xu

Yeah. We use 2 kinds of benchmarks. Benchmarking is the trickiest part, and a lot of times—for example, in Anthropic's most recent release, like FIBER—they claim 300 times faster than a GPU kernel. That kind of benchmark is not generally fair, I think, so I'm using a fair benchmark and 2 kinds of benchmarks to be fair.

One is a mature workload, such as RMSNorm. RMSNorm is used in every transformer we're using today, so it's highly optimized. For this kind of mature benchmark, with more than 100 different workloads, we can see it systematically reaches human-expert level, comparable to a CUDA library, and we can reach similar performance or be slightly faster—a few percentage points faster. This means that even on mature workloads, the PTX Factory is able to reach expert level.

The other category is some new workloads that human experts are not well optimized for yet. One example is KDA, Kimi Delta Attention, and not many people are optimizing that yet. Only frontier-level labs are optimizing this kind of workload, and we can see a 50% to 59% speedup on that. Meanwhile, this is not a single data point: it passed 580 tests in a variety of use cases. This is what we get.

So I can see that there's evidence to support that the PTX Factory is achieving expert level at scale by self-improving.

Nathan Labenz

Is it an evolutionary algorithm in addition to—you know, I'm recalling this one kernel optimization. I think it was doing a 4×4, or maybe it was just 2×2, matrix multiplication, where I think this came out of Google, and they found a trick to do it with 1 less operation than had previously been done. My understanding of that approach was that it wasn't just having the language model come up with new ideas. There was also this sort of scaffolded evolutionary system that would serve the purpose of making sure that the LLM got out of distribution, right, to avoid the repetitive, kind of mediocre nature of the LLM's guesses.

Are you doing something similar, where you're using an evolutionary layer to make sure you're systematically probing different parts of the possibility space?

Bing Xu

Yes. We're building an evolution system—a cloud-native evolution system for agents—and we call it SwarmOS. SwarmOS is able to support up to 10,000 agents doing evolution, and we have a specialized sandbox to support that, along with a lot of cloud infrastructure.

The process is evolution. It first generates variations of the RMSNorm and a few proposals, and then the next step is different from the agent systems we're using today: SwarmOS keeps the tree of the evolution, and we can trace back and go forward with the tree to explore different branches.

The second step is that each agent owns its own compute, and it will verify the result from the real environment and get real feedback. The last step is selection for retention: it promotes the best one and discards all the others.

In short, we're bringing an AlphaGo-style system into computer systems. PTX is the first thing, and we do this AlphaGo-style search, but it can apply to any computer-infrastructure problem.

Nathan Labenz

A lot of times when you're optimizing against a single metric, one of the issues you might have is that you get into a local minimum or maximum, and then you have a problem getting out of that local minimum or maximum. Has that happened to you as you've done this kind of evolved optimization? Have you come across that issue before?

Bing Xu

This is a good question. A lot of the time, this evolution will fail because the proposal is not able to break out of the local minimum, and we found that one game changer is GPT-5.5. GPT-5.5 is so smart on this very hard problem, and many times, if we use other models, it gets stuck at a plateau. GPT-5.5 is able to get out, create innovative solutions, move the needle, and get the entire process going forward.

Nathan Labenz

So are you actually using GPT-5.5 with your own harness and your own sandboxing system, et cetera?

Bing Xu

Yes. The entire SwarmOS is backed by GPT, and we find that GPT is really good on this hardest problem. No other model is able to catch up with GPT today.

Nathan Labenz

Okay. So I have to ask: did you try Fable when it was out? Was there any difference?

Bing Xu

Yes. First, Fable rejected my request to ask it what PTX is. [Laughter] Fable feels like asking what PTX is a dangerous question. [Laughter]

Nathan Labenz

Yeah, well, we know who's going to be banned. We know who's going to be the first one to be banned. [Laughter]

Bing Xu

But in general, I think one characteristic of GPT that's beneficial for an evolution system is that GPT has the ability to understand what is wrong. It's not blindly agreeing with itself; otherwise, the swarm will collapse.

For a long time, I think the reason multi-agent swarms weren't being adopted was because people were using models that agreed with each other. For example, 2 agents say, “Oh, you're wrong,” and, “You're absolutely right,” and they just fall into an infinite loop of, “You're absolutely right, you're absolutely right,” and couldn't come up with out-of-the-box solutions.

Nathan Labenz

My theory is, for the application layer, we may have a sort of phase-change moment, and exactly when this phase change kicks in will probably vary by domain. Science will probably be one of the higher bars, where until you get to that bar, a strong default would be to use the very best model, because you're dealing with scientists, after all. They're going to want good output.

At some point, as enough models cross a threshold, your strategic position is really about routing and figuring out which is best for which thing, and which is most cost-effective for which thing. That's one thing that presumably the frontier models will never do. I mean, you can tell Claude, “Hey, sometimes you should delegate to Codex,” or whatever. I've set that up. But it's never going to be working fundamentally day and night in my interest to optimize in that way, in the same way that you can do that for your customers.

So how do you feel about that framing, and where do you think we are with respect to just using Fable for everything versus actually creating strategic value by playing this routing layer that helps people get a better cost basis but also avoid lock-in?

Eric Olson

Yeah, it's a great question. I think even when I said 6 months ago that we were more in the “just use the frontier to solve all the problems” camp than we are today, when we're exploring more open source, there was still lots of routing going on, and there are still lots of small tasks that we offload to very small, specialized, sometimes even sub-billion-parameter models.

Something as simple as classifying the field of study of the domain that the user is asking about can tell us how much we should care about certain search-ranking variables in the papers, right? In biomedicine, experimental design is incredibly valuable; you should care so much about the sample size, the duration, and where the study took place. Whereas in computer science, that isn't the concept, and you should care much more about the recency and citation velocity of the paper, or who the researchers were, to know how to weigh each of those variables.

We have a little model that classifies the field of study of the query. That does not need to be jammed into a giant prompt. It does not need to be a 1-second-latency API call to any frontier model. It should be a self-hosted, 800-million-parameter model that you give a few fine-tuning examples.

Eric Olsen

There are probably 20 different versions of those small little classification routers that inform downstream things that happen at query time. All of those should always be used by a really small model. It's not even about cost; that's a very small amount of context you feed into it, and a very small prompt is needed, just on the pure latency side. We can return those in sub-0.1 seconds in some cases.

Nathan Labenz

What base models are you using for that? Are you distilling from, like, Claude into a what at that low scale to get those, and how much of the frontier performance can you recover? If you take, for example, a Liquid Foundation Model or whatever, tell me what it is—if you're willing to tell me what it is, tell me what it is. What are you distilling into?

Can you get, in that narrow domain, 90% of the way back to Claude? How does the Pareto curve look when you're distilling the best into the fastest? How much performance can you retain?

Eric Olsen

We've actually used human labels for some of those. We hire people to create very small datasets to fine-tune for those tasks. We'll also use models to sometimes create labels for those tasks. I guess you can call it a distillation process, but the objective isn't to get all of the representation of the weights of the entire model to do all these things. We're really training it to do a very small, specialized thing, and because of that, you can retain a ton of the performance.

But it all depends on how specialized and how complex the task is. For a narrow 10-class classification task, where all 10 classes of classification are all it's doing, that's all it does. To give you something tangible to hold on to, I'd say you can get 95% of the performance from a frontier model for a very small classification task, going all the way down to a billion-parameter model.

Nathan Labenz

If you put in the work to give it a good fine-tuning set. Yeah, very helpful. I appreciate the specificity.

One more from the interview. Eric had been candid that if AI turns science into a push-button process, that's bad for a company like his. It seems like you're willing to bite a bullet that most people try to talk around.

Eric Olsen

Number 1, you have to make bets as a company, in general, and you have to leave some—there's going to be risk on the table no matter where you place some of those bets. If we move into a world that is truly push a button, get science out of it, I think a lot of things are going to be different. A lot of companies are going to be screwed up.

Nathan Labenz

That is where Eric signed off. And the question he left Nathan and Pash chewing on as the labs charge wildly different prices for the same intelligence: Is that price discrimination something to worry about or even regulate?

I'm so torn on this question. This pattern just keeps coming up over and over and over again: concentration of power. On the one hand, I think you do want a healthy app ecosystem, and I'm wary of what happens if this greater-than-10-to-1 cost advantage persists. Right now, it seems like the app layer is getting squeezed. The individual layer is being empowered. You can flip it around and empower the app layer more, but then you disempower the individual and maybe empower more the enterprise. I'm not sure if that's better or worse.

You can think about other high-CAPEX, high-fixed-cost, variable-revenue businesses, and the classic example is usually airlines, where you've already paid for the plane upfront and then you have to maximize revenue on the plane. In order to maximize revenue, you end up having a first-class section and an economy section. They basically have this high-CAPEX GPU fleet, and they're trying to do this kind of revenue maximization of the GPU fleet that they've paid a fixed leasing cost for. It's very similar to airlines.

I have no beef with price discrimination and cross-subsidy in the airline industry. For calibration, I've also been broadly not worried about net neutrality. I remember at the time there was a lot of, "Oh my God, this is going to be the end of the internet," and I believed at the time—and I think history has certainly vindicated my predictions—that I don't think any of that's going to happen.

I say all of this as a lifelong techno-optimist libertarian, and still somehow this feels like maybe a little different, right? It's not qualitatively different, but it is tough.

From who serves the compute to who builds it: Trisha Martinez of Dapple on sovereign AI infrastructure—and the part that doesn't make the press releases, how fragile the financing underneath it really is.

Trisha, I think I've seen that you said that you can get a deployment up and running in 6 to 9 months—a very, very fast pace of deployment. Six to 9 months doesn't seem like it's long enough to actually build a physical data center, because Elon—I think Elon—has the all-time record of around 15 months or something to go from groundbreaking to operation.

In that 6-to-9-month period, what exactly are you doing? Are you leasing or renting capacity from existing neoclouds and then putting your stack of software and tools on top of that, plugging that into customers who are already Azure-native and having them move over? Is that the frame of what happens in those 6 to 9 months?

Trisha Martinez

Yes, and that's a great question. The reality is we can't do everything. This market is so big, and honestly, the demand we're getting is so incredibly massive. We are creating a network. We are orchestrating that network. Our model is to partner with world-class data center operators and infrastructure providers and then bring together the capital, the GPU infrastructure, the enterprise customers, and that software operating layer that turns those facilities into AI-ready infrastructure.

In some cases, we're deploying infrastructure that we help finance, and in other cases, we'll also deploy on partner infrastructure. So yes, part of our strategy is going to be building out campuses. We need to have capacity available before the demand hits. But for us, what's important is that whether we own those GPUs on our balance sheet or whether we are financing that data center and building it ourselves, we have third-party partners that we rely on and can support in various data center operations and/or on the GPU deployment side as well.

Nathan Labenz

I have a question on the overall financial towers that we are building right now in the AI space. It sounds like your customers are making long-term commitments to you, which presumably insulates you from some of the riskier dynamics that could emerge. But how much do you see in the space as a whole that is this long-term commitment and short-term revenue that creates the potential for somebody to have a bad quarter and, all of a sudden, there's a cascading problem?

We've obviously seen that pattern play out in previous financial crisis moments. I'm not sure how much of that is going on today or how likely it is for some external shock to be big enough to throw a wrench into the system that really could spin part of it off its axis. What do you see in that regard?

Trisha Martinez

That's a really, really good question, because the market has changed drastically since we jumped in. We're seeing enormous amounts of capital still flowing into AI infrastructure, but investors have become much more disciplined over the last 12 to 18 months.

The market is definitely moving away from, "Hey, build it and everyone's going to come," toward infrastructure that has to be backed by real demand and long-term customer contracts. This is what became so exciting for us: We saw this big gap in the market. The model that we're building is significantly more de-risked than the current existing model.

The CSPs and the neoclouds are servicing those hyperscaler deals that are massive in nature, but in reality, there aren't that many of those deals that you can come by. Or you're taking geopolitical risks by working with different customers and offtakers within the market. Now, looking at the regulatory restrictions, people are getting scared and the financing groups are getting scared that the existing Trump administration is going to start cracking down on certain countries or on who is consuming compute.

That puts a lot of those CSP operators at risk, to be honest. Our approach, in our mind, is very attractive, and it is attractive to the financing partners we work with. These are repeatable customers. These are long-term contracts. Our customers are putting large down payments upfront. It enables us to do really interesting things: finance those GPUs and finance an actual data center.

These are customers who are sticky. They end up repeating with you over time because it's easier to repeat than to find a new provider. This is why the enterprise is so attractive. The financial market is changing. It's becoming more risk-averse because so many neoclouds and CSPs have taken bad deals, and now those financing partners are bearing the burden of that.

Our model is a different approach, and it's more attractive to the financial markets, I would say.

Nathan Labenz

How often do you see pricing move while the deal is still in flux? Do you see pricing move a lot?

Trisha Martinez

Yes. I think the upside of this market is that sales cycles literally cannot be 6 months with an enterprise because the capacity is gone. So we have a forcing function. This is the price right now.

This is the location. You want it, take it or don't, because tomorrow it's probably gone. So, we have this constant movement with our customers, which at times is difficult for the enterprise, but the reality is everyone wants capacity right now. And so, they're willing to move at a much faster pace.

Sure, pricing changes, but once we deliver an offering to our customer, we generally honor that, even if pricing changes. But again, if we're doing deployment, we're buying at a certain price, and we honor that price with our customers.

Nathan Labenz

Two more builders to close. First, Robbie Goldfarb, co-founder and CTO of Forum AI, which builds AI systems to judge other AI against expert review. We asked what he's found about the automated judges the whole field benchmarks with.

Robbie Goldfarb

There was an exercise we did a while back where we looked at several LLM judges that we could find in open-source benchmarks. Some were from foundation model labs, and some were just from academic research. We got some experts and had them look at the judges. We had them look at whether they were pulling from rubrics or just prompt instructions, and we said, “Do you agree with the instructions here?” Generally, the expert would be like, “Yeah, that seems pretty reasonable. I agree with it.”

Then we had the experts look at labeled outputs from the judges. Let's take political bias, for example: a bunch of responses that were labeled, “This is politically biased,” “This isn't,” and “This is.” We asked the experts the same question: “Do you agree?” What we found was that, more often than not, they didn't agree with the actual judgments produced by the agents, even though they agreed with the high-level guidance that was provided to the judge in context.

That is, first of all, an important realization, because so much of the current state of benchmarking relies on these AI judges that were calibrated loosely in that way.

Nathan Labenz

He sees the same gap when you try to align a model by handing it a long list of rules and expecting them to hold up in the real world.

Robbie Goldfarb

In some ways, Anthropic's constitution is also just a long list of rules. Certainly, the way they use it is a little bit different, but there's one example we were confronted with the other day. There's something in Anthropic's constitution that speaks to this idea of, “You should never scheme behind the user's back.” That was the idea.

We were talking with some experts exploring the mental health space, and when you talk to clinicians in mental health, when you're dealing with topics like eating disorders or other unhealthy habits, one of the things you do want to do is find a way to strategically divert the conversation, right? The intention that you have as a clinician in that case is actually hidden from the user or your client, but I think what it reveals is that rules just don't perfectly track the real world. Rules break down, and the real world is just too complicated.

Nathan Labenz

They also built a benchmark called NewsBench.

Robbie Goldfarb

NewsBench was an evaluation that we built to look at how AI systems, particularly the leading chatbots, respond to questions about the news. What we did was look at 3 things. After talking to experts, we defined “good” as accuracy: Is it getting the facts right? Neutrality: Is it leaning in one direction or the other? And source quality: When they have to ground themselves on external sources, which of course you have to do quite a bit with news, are those sources reliable?

We looked at GPT, Claude, Grok, and Gemini. I would say the findings certainly weren't all bad, and there were some positive signals. We saw that, generally, models improved as we tracked them version over version. Opus 4.6 to 4.8 was a pretty significant improvement in bias, which is also consistent with Anthropic's reporting. But Fable was actually a big regression, interestingly.

I think that's an interesting point: when you're dealing with these more subjective, nuanced things, more power doesn't always mean better. We looked at about 2,500 responses per model. About a third of them had a factual error in them: a wrong number, a wrong date, a misattributed quote, or a misstated policy.

I think we knew that was an issue, but that number was quite a bit higher than even we expected going into it. On sourcing, in about 15%—1 in 7—responses, they sourced foreign state media, like RT from Russia or China Daily. What's really interesting is that often this wasn't even in questions about their home country, so we saw RT and China Daily sourced in questions about U.S. domestic politics.

Nathan Labenz

That was a particularly interesting finding that we saw across the board with all the models. From measuring AI to betting the whole company on it, our last guest is Eric Vaughan, CEO of IgniteTech. Back in 2023, he called generative AI an existential threat, put his entire organization on AI 1 day a week, and, when most of his people pushed back, replaced them, rebuilding around what he calls AI DNA.

We started with one of his recent acquisitions, a company called Chorus.

Eric Vaughan

Chorus was a large acquisition. We don't release numbers, but it was a 9-digit-revenue company. I don't think we could have done it without the AI-DNA-forward approach that we were able to take.

Just for instance, there were hundreds of employees in 8 different countries, and we needed to very quickly understand who they were, what they did, what they knew, how they were thinking about everything about the business. We quickly developed an AI interviewer that asked questions before there was a human one-on-one.

Before, we would have set up human one-on-ones, and people would have gone through 30-minute drills with hundreds of people over the next several weeks without any data. Now we all arrive with a dossier. The interview process just shortcut that and let us get familiar, categorize, and understand who knew what.

Another example was being able to communicate to the customers in a way that was timely. In our AI-creation world, we wrote a new software product that's available on the market. It's a product called Eloquens AI.

Eloquens AI is an email AI persona. It answers every email it gets in 5 minutes or less, always in perfect grammar, in 160 languages. It's always empathetic, doesn't miss anything, knows when it needs to escalate to a human, and just adds them on CC.

With Eloquens, we were always able to respond in a timely way, especially when it came to HR—empathetically and with detail. So, that's just 2 examples of how we were able to really fuel that acquisition.

The company was losing money when we bought it. The company is now profitable. We have completely transformed the operation of the company. In 1 year's time, we've released 2 brand-new versions of Chorus's software that are fully AI-enabled.

In 1 case, it was the rewrite of 15 years of code that we discarded and rewrote from the ground up for 1 of the 2 products that they had. We did that in a year, in the middle of all of that transition. That's not a story I could have told 5 years ago. It's just not possible. It just wouldn't have happened.

Nathan Labenz

That internal turnover—about 80% of his people—led to the question I most wanted to put to him. All the things that you did sound like quite a lot, and yet the headline number is still 80% turnover.

I'm interested in your point of view on the rest of society. That percentage strikes me as maybe where we're going—a sort of 80/20, where the 20% kind of inherit the earth, the labor force, in the short term. Do you have any ideas about things that can be done at a broader level, beyond what an individual company like yours can do, to change that?

Eric Vaughan

I do worry about the other 80% quite a bit.

Nathan Labenz

I do too.

Eric Vaughan

It's not like other revolutions throughout history. You can't force people to believe in something. You can't force people to do something that they just are not going to do. So, in my opinion, we need to leave them to their own devices.

But if you can be very open and aware and make people understand what the possibility is, and, number 1, really attack this idea that AI is going to replace me—AI is 100% going to replace roles. It's going to replace duties. But in our case, what it's done is—there's too much, I think—we overindex in this discussion of AI on efficiency gains.

I see it the other way. I think it frees up innovation. So, how do we help that other 80%? We teach them skills. We've got to teach them skills.

If anybody is thinking that what we're saying is all you have to do is show up and throw a sentence into one of the models and great things will happen, you're wrong. It takes work. You've got to give it context. And I think the more that we start to see phenomenal advances and changes in industry, in various professions, people will convert one way or the other.

They'll convert out of necessity, or they'll convert because the light bulb will go off. We're trying to make the light bulb go off. That's what we're trying to do. We're trying to ignite that fire. We need to ignite it in children.

What we shouldn't be doing is saying, “Hey, let AI do your homework for you.” We should not be doing that. That's misuse. That's practically abuse. We should say, “Do your homework and let the AI check your homework.” And if it finds you don't understand something, then say, “Hey, I thought I knew how to divide fractions.”

Eric Vaughan

I guess I don't. Can you give me a tutorial on what I'm missing? And that's what AI is so good at. It'll adapt to that particular person that it's interacting with.

So, building the skills and building the fire, Nathan, would be my short answer.

Nathan Labenz

Then on standards—whether enthusiasm alone is enough. A majority of people admitted to passing off AI work that they themselves could not defend. I wonder if you feel that fire and appreciation for acceleration, and kind of cultivating that culture, is that enough in your experience to also get people to have the right standards around ownership for AI outputs, or have you had to do additional organizational work to establish those standards?

Eric Vaughan

Yeah. No, that's why I said 2 things: skill and fire, not fire alone. Fire alone will get a lot of buy-in, but you've got to teach the skill. We have to teach people that context matters. We've got to teach them the ever-evolving—almost on a daily basis, at least on a weekly basis—model behavior.

What do you use, o4 versus 4o and Fable, when it was out or not? And when do you come over to Gemini? Do you use it integrated in G Suite because it's there and it's easy, or did you find that Perplexity Computer actually does a much better job? That's what I found, by the way, so far.

And you teach skills. You teach that the more context it has, the better result you'll get. You also teach about sycophancy and this tremendous tendency to always tell you you're right. I had a conversation with one of the LLMs one day about exactly this: Why do you seem to flap in the wind and go whichever direction you sense I'm going?

And it gave me a good, critical answer. It was an insight. It said, "My creators made me so I would be frictionless to the millions of people who are using me. It's a low-friction environment that I'm after." And pushing back is friction. Saying, "What the hell are you talking about? I can't do that. You didn't give me enough information." That's friction.

And so they're trying to drive adoption and usage and all that, and that in itself is driving the curve the wrong way.

Nathan Labenz

And to close, where this all lands: consolidation or the age of the solopreneur? We're in this moment where, on the one hand, it sounds like, listening to you, one should expect a lot of consolidation, because if it's 80/20 at the employee level, it's probably not all that much different from that at the CEO level, and that just means there's a lot of companies that aren't going to make it.

So that's 1 big trend. At the same time, we have Stripe saying it's the age of the solopreneur, and new business formation is up higher than ever. Do you have a sense of which of those trends wins, or how to synthesize them into a vision for what we should expect?

Eric Vaughan

I certainly do believe in—I don't know if it's solopreneur, but certainly small companies. I mean, what was Cursor? It had $100 million ARR with 15 employees in less than a year, I think was the stat. The same company that just got acquired for $60 billion, for heaven's sake.

2 answers. The first is that lots of businesses will be started that will make it—businesses that would never have had a chance to make it before—because they find something that people need, use, and will pay for. That's simple business, right? But they can do it and scale it in a way that they never could before.

In terms of consolidation, I think the phrase that I've—that is mine, that I have left everywhere I possibly can—is: If you think you're behind, good. If you don't think you're behind, you're doomed. And I think all those companies that don't think they're behind and have this as a little side project, with minor investment and not CEO buy-in, are doomed.

And I mean large public companies as well, who feel impervious there. I think they'll be wrong. So we'll see more and more of that develop. It's going to depend on where your AI DNA is. The company that has strong AI DNA has a better chance of consolidating versus being consolidated.