AI News Crossover:与 Doom Debates 的 Liron Shapira 坦率对谈
Nathan Labenz 将 P(doom) 设在 10%-90%,但权重偏向低端;Liron Shapira 则接近 50%,而 Shapira 认为,Labenz 面向大众的播客传递出的危险感,远低于这组概率所暗示的程度。 Labenz 承认,他“中立分析师”的语气可能让自己变成了一只“温水煮青蛙”,尽管即便是六分之一的风险,也仍然类似俄罗斯轮盘赌。更具行动意义的区分,是开发强大 AI 本身带来的不可消除风险,与在超级大国竞赛中将其武器化所造成的、原本可以避免的“极其愚蠢的风险”——后者可能让基准危险增加 10 倍。
GPT-4o 的图像生成已经对广告、平面设计、Fiverr 式市场以及 Labenz 的 Waymark 等企业构成即时通缩冲击,即便它还无法替代 Waymark 完整的带声音视频工作流。 Labenz 描述了一条新流程:企业一个下午生成 100 条高质量图片广告,花约 100美元通过 Meta 测试,再由转化数据选出赢家——“现在已经可以同时做到好、快、便宜”。Waymark 过去从 99美元的专业配音转向成本几美分的生成式语音后,供应商订单量已经下降超过 90%。
劳动力市场的证据指向的是:由少得多的专职工程师生产多得多的软件,而不是 AI 辅助编程岗位出现持久繁荣。 Replit CEO Amjad Masad 所说的“我不再认为你应该学习编程”,意味着学生应学习问题拆解和沟通,同时不必害怕代码;Shapira 粗略设想,软件产出可能增加 10-100 倍,而所需人类只剩原来的 10%-20%。Cursor 在员工不足 50人的情况下达到约 100亿美元估值,Shortwave 计划维持 15人团队,Replit 的规模约为 100人,这些都被视为对“生产率提升意味着需要更多招聘”叙事的现实反驳。
创业在短期内可能比就业更安全,因为创始人可以不断“奔走”去寻找仍然存在的价值,但两位嘉宾都认为这只是过渡状态,不可能成为可持续的社会契约。 今天,AI 工具能力仍可通过 Fiverr、咨询或“靠数字土地谋生”赚钱,但最终也会有另一种 AI 来取代 AI 岗位。Labenz 所设想的好结局,是一个人们“不必工作才能吃饭”的世界;尚未解决、也最值得投资者关注的问题,是谁能在软件、服务和专业能力被通用 AI 平台吸收之前,捕获这段转型期的价值。
OpenAI 只有在被视为成为世界经济枢纽的高度偏斜期权时,3000亿美元估值才有说服力;若建立在对当前 token 毛利的舒适假设上,则难以成立。 快速跟随者、价格战和模型商品化,让普通现金流逻辑难以解释这一估值;但一个 1% 概率实现 30万亿美元结果的期权,数学上足以支撑 3000亿美元。Shapira 认为,如果未来 10-20年没有发生 doom,OpenAI 最终达到 30万亿美元的条件概率为 10%-30%。OpenAI 自己引用的预测——年收入从约 140亿美元增至 2029年的 1000亿美元——提供了一条没那么极端的路径,但两位嘉宾都没有将其视为确定结果。
安全研究仍是一套东拼西凑的防线:“有机对齐”很有意思但未经验证,而 Anthropic 备受赞誉的可解释性研究也远比标题所暗示的模糊。 Shapira 的反驳是,细胞之所以合作,是因为彼此需要;超级智能却可能把人类视为废料。Labenz 仍支持探索被忽视的方法,因为对齐研究者的调查显示,当前方法和时间表都不足以解决问题。Anthropic 的跨层转码器替代模型据称只能预测底层行为约 50%,依赖主观的特征标签和误差项;如果研究图表被包装成“我们知道这些东西如何运作”,还可能带来政治危险。
讨论中的核心灾难并非卡通式反派,而是一个意外摧毁人类生存条件的 AI 经济体——正如人类没有总体计划,却造成了大规模灭绝。 今天的助手之所以显得无害,部分原因在于企业投入了大量工作来收窄其行为边界;据称,在发布前、纯粹追求帮助的 GPT-4 上,只需轻微推动,它就提出过定向绑架和暗杀的建议。具身系统会进一步放大风险。Shapira 提议用“二分搜索”检验直觉:如果一个能当保姆、司机和全能仆人的家用机器人,让人感觉已经走到支配人类的一半,而且在 2028年前后看起来可信,那么再否定 2031年出现支配情景就会更加困难。
国际 AI 协议很难,但并非公理意义上不可能;而且验证可能比机制可解释性更容易。 卫星、电力需求、芯片供应链、硬件遥测和现场检查都能提供帮助,但 Labenz 认为,分布式训练意味着仅靠远程观察不够,领先大国之间必须建立地面信任。两人的共同结论是,条约不需要完美检测,也不需要 100%有效,才能降低风险:“纵深防御是我们拥有的一切——希望它也是我们需要的一切。”
1. 灭绝概率与公众表达严重错位
Labenz 给出的惯常 P(doom) 是“10%-90%”,但权重偏向低端,因为他认为所谓有效数字极度不确定。他大致参考 Dario Amodei 约 20%的判断,但承认这更多是凭直觉服从权威,而非严谨综合;Shapira 给出的答案则约为 50%。
Shapira 的质疑并不是 Labenz 否认风险,而是一个技术听众可能连续听 10期 Cognitive Revolution,却完全意识不到主持人认为自己这一代人遭遇 doom 的概率达到两位数。在 Shapira 所说的“理性区间”里,10%-90%是可以辩护的;低于 5%或高于 95%,则代表没有依据的自信。
Labenz 接受这种错位。他一直试图同时传达巨大的上行空间和灾难性下行风险,但自己的冷静性格、害怕出错以及“中立分析师”的身份,可能低估了他真正相信的事情:“赌注真的不可能更高了。”
令人不适的自我诊断是,反复接触相关议题,可能让他在 10%-20%的风险水平上变成了一只“温水煮青蛙”。他考虑定期披露 P(doom) 和后稀缺丰裕同时发生的概率,明确告诉听众,即便乌托邦仍是更可能的结果,社会也在“掷骰子”。
2. 超级大国竞赛可能将不可消除的风险放大
Labenz 的朋友 Gopal 提供了他认为最有用的框架:少去思考如何给概率定值,多问“我们能把它们推向什么方向”。一旦互联网规模的算力、互联网规模的数据和算法发现能力让强大 AI 在许多合理时间线上都大概率出现,有些危险可能已经不可消除。
可以避免的部分,是 Labenz 所说的“极其愚蠢的风险”:在争夺全球霸权的超级大国竞赛中,尽快将这项技术武器化。这可能不是增加几个百分点,而是“把基准风险放大 10倍”。
Shapira 指出,这个被称为愚蠢的情景,恰恰也是最接近现实的情景。Labenz 同意,遗憾的是,“目前我们正走在这条路上”,而且几乎没有有影响力的参与者愿意接受降级对抗的信息。
在询问人们预测风险之前,Shapira 建议先问:人类应该容忍多大概率的极端灾难。他预计大多数人会回答低于 1%;如果这确实只是文明必须跨越的一道罕见门槛,他可以接受 90/10 的赌局,但强调,协调成本和糟糕替代方案可能让这种赌局变得合理。
3. 一旦公众理解赌注,抗议就显得合乎比例
即便是把自己锁在 OpenAI 门口并被逮捕的抗议者,Labenz 也能理解。这不是他的策略,他也没有明确支持,但他说,考虑到行业内部人士认为赌注如此之大,对一些人来说,这“绝对算不上太激进”。
Shapira 希望立场居中的传播者推动奥弗顿窗口。生存风险已经成了与 LessWrong 和 Eliezer Yudkowsky 联系在一起的熟悉靶子,但面向大众的播客很少直白地说,受尊敬的人认为听众有生之年发生 doom 的概率相当高。
分歧在于可接受的概率,而不是上行空间是否存在。Labenz 认为后稀缺丰裕可能比 doom 更有可能;Shapira 在严格条件下异常愿意承受 10%的风险。但两人都反对把 P(doom) 轻率地一笔带过,仿佛“末日论者”本身就是反驳。
4. GPT-4o 跨过生成式图片的商业门槛
Labenz 认为,GPT-4o 的图像生成是第一个可能跨过 Waymark 集成门槛的系统。Waymark 过去使用客户网站上的真实图片,因为本地企业需要看起来像实际到访地点的广告;此前的生成器可以创造风格,却无法足够可靠地保持身份一致。
录制时 API 尚未推出,但 Labenz 预计很快就会出现。一旦可用,GPT-4o 就能在保持广告主自身特征的同时扩展 Waymark 的创意主题——这将是“一次重大解锁”。不过,Waymark 目前 15秒和 30秒、以配音为核心的电视广告格式,还不会被静态图片直接替代。
Shapira 对 Relationship Hero 的实验让近期的“alpha”变得具体:一个下午生成大约 100条强势的 Facebook 图片广告,花约 100美元测试不同版本,再把 Meta 的网络当作进化筛选引擎。他还没有性能数据,但认为这将是 Meta 广告对自己公司进行的最后一次测试。
两人的共同判断是,专业机构可能不再能比几条 prompt 产出明显更好的图片广告。广告更看重注意力和信息传达,而不是艺术完美;手指变形等小瑕疵的重要性低于质量、数量、交付速度和可衡量的转化。
5. 递归图像揭示原生多模态的变化
Riley Goodside 的自指式 Wikipedia 截图成为两人的突出示范:GPT-4o 生成了一篇名为“The Screenshot”的可信文章,文章中包含这篇文章的图像,并描述了自身的递归。Goodside 说自己用了几条 prompt,但最终成品实际上近乎一次生成。
仔细观察仍能发现局限:大约只有 3层递归、出现虚构词和拼写错误,中心还有一块无法阅读的黄色区域。Labenz 认为这些缺陷“很迷人”,而非不可接受,因为仅仅是第一层构图和语义协调能力,5年前看起来都还不可能实现。
突破并不只是更好的 diffusion model。OpenAI 此前通过 Greg Brockman 生成的黑板文字预演过这个想法:“如果我们把文本、图像和音频全部联合建模,会怎样?”Labenz 描述了一个共享潜空间:用语言、图像或声音表达的概念,在其中汇聚为更丰富的跨模态理解。
过去的 ChatGPT 系统实际上是先写一个 prompt,再与 DALL-E 保持距离地调用,形成“损耗极大的瓶颈”。原生整合移除了大量翻译损失;这张递归截图显示,模型可以在一个过程中协调布局、类似精确文字的形态、嵌套概念和视觉渲染。
6. 创意通缩已经让 Waymark 人工配音量下降超过 90%
Waymark 提供了一个正在发生的颠覆案例。其可选的人工配音服务收费 99美元,没有利润,耗时约 2天,有时还需要修改;供应商水平很高,而这个价格被认为异常优惠。
生成式语音可能仍然不如人工,但它即时可得、可以在产品内编辑、几乎免费打包提供,Waymark 的成本只有几美分。因此,人工供应商的订单量下降了超过 90%——这是典型的颠覆模式:便利性和价格击败质量更高的既有供应商。
Labenz 预计平面设计也会经历类似变化,并表示看不出未来几个月内为什么不能在数量级上重演这一过程。企业尚未解决的暴露包括 Adobe、创意劳动和 Waymark 自身:GPT-4o 可以增强 Waymark 的产品,但从更长时间尺度看,“只需 prompt 就能做出某种东西”构成可信的生存风险。
Labenz 将一种 alpha 概括为“质量和数量”。他进一步压缩了经济学含义:“好、快、便宜,现在三者可以同时拥有。”这推翻了过去买家只能三选二的约束。
7. Fiverr 可以适应,但底层任务正在消失
Fiverr 的市值此前已经从超过 100亿美元跌至约 8亿美元,因此市场已经定价了一部分生存风险。Labenz 指出,公司正在围绕 AI 重建卖家入驻、买家需求收集、撮合和服务定义,而不是原地不动。
Fiverr Go 试图通过授权创作者资产来保持其不可替代性,例如授权配音演员的克隆声音,并补偿最初提供资产的人类。真正不确定的是支付意愿:Waymark 对 99美元人工服务和约 3美分生成成本的比较留下了巨大空间,但未必有足够的持续价值支撑过去的劳动规模。
Shapira 提供了直接的需求破坏案例。他曾通过 Fiverr 和 99designs 获得勉强够用的 YouTube 缩略图,还要花时间审阅投稿并与设计师沟通;如今 GPT-4o 加上少量 Photoshop 调整就能更快完成,他预计自己不会再为这项工作回去找人。
Labenz 认为,平台暂时仍能在工具选择套利中发挥作用。买家不知道该使用哪个模型,因此仍会发布可自动化的工作;了解工具的自由职业者可以用 AI 执行。价格会通缩,服务可能转向导航,但这一优势只能持续到 AI 自己选择并操作正确工具为止。
8. AI 侦察员可以短暂地“靠数字土地谋生”
Labenz 建议有志成为 AI 侦察员的人不要去申请补助,而是“靠数字土地谋生”:在 Fiverr、Upwork 或其他平台寻找现有 AI 已经能够完成的付费数字任务。客户支付的未必是新技术发明,而是找到有效方案的能力。
Labenz 自己的大部分商业价值,来自维护一张近乎完整的可用工具地图,并为客户挑选最佳工具。他有时也会找到创造性方案,但通常客户购买的是一种确定感:自己拿到了“现有最好的 AI 选项”。
Naval Ravikant 的宇航员 meme 捕捉了这种递归:AI 消灭工作,“AI 工作”消灭 AI,然后另一个 AI 又在 AI 工作之后出现。Labenz 同意,未来几年侦察和 agent 管理可能具有社会价值,但看不到工具专业能力能成为持久护城河。
他对变革性 AI 的规划区间是 2-5年,置信度为 80%,但会刻意按照 2年端行动,以制造紧迫感。如果时间更长,“那很好”;他宁愿过早准备,也不愿按照时间区间最远端来设计职业规划。
9. “学习编程”正在变成“不要害怕代码”
Replit CEO Amjad Masad 说“我不再认为你应该学习编程”,同时提出另一套课程:学习思考、拆解问题和清晰沟通。Labenz 的补充是,任何人都不应该害怕代码;如今只需适度投入,再借助 AI,就能跨过许多实现障碍。
传统反驳认为,更便宜的软件会扩大需求,因此生产率更高的开发者会被雇来生产更多软件。Labenz 认为这一逻辑短期内可能成立,但最终 agent 会在互动过程中动态编写自己所需的软件,从而削弱对永久性人类开发应用的需求。
领先公司的行为就是他的证据。Shortwave 计划即便业务指数级增长并获得新融资,也将团队维持在约 15人;Cursor 在员工不足 50人的情况下估值约 100亿美元;早于当前浪潮出现的 Replit,则可能只用 100多人就构建了规模可观的产品。
Shapita 的刻意粗略预测是:软件产出增加 10-100 倍,而专职人类只剩原来的 10%-20%。深层系统工作可能仍然稀缺且有价值,但训练营式技能——React 组件、前端以及全栈 CRUD 应用——正在极快商品化。
10. 创业是最后一个白领岗位,但不是社会契约
Pieter Levels 的那句话——“现在当企业家比拥有一份工作更有工作保障”——之所以令人共鸣,是因为创业意味着不断寻找下一个稀缺价值来源。Shapira 称,随着经济潮水沿着基于笔记本电脑的工作上升,创始人是最后一批为钱“四处奔走”的人。
Labenz 同意,通用性、自我驱动以及适应随机新问题的能力,让两人都处于相对有利的位置。他更担心技能范围狭窄、从未被迫创造下一份工作的人,而不是自己。
但社会不可能完全由创业者组成,他们不断在“更广泛 AI 经济的沙发缝里寻找零钱”。Labenz 认为需要新的社会契约,并最终欢迎一个人们不必工作才能吃饭的世界;当他询问底特律居民,如果不需要钱是否还会继续工作,答案压倒性地是否定的。
因此,创业看起来是过渡性的就业安全,而非永久隔离层。体力行业可能存续更久,但一旦 AI 能完成 AI 工作、机器人进入实体经济,四处奔走也只能推迟分配问题。
11. Vibe coding 暴露了工艺身份如何拖慢采用
Shapira 称自己是“编程婴儿潮一代”,因为他仍然在 Cursor 中逐行编辑,而年轻用户越来越多地直接与编辑器对话,避免碰代码。他怀疑甚至 Andrej Karpathy 也在转向这种不亲手操作的方式。
Labenz 很少逐行编辑,因为他一直把代码当作概念验证机器:“让它能跑,继续前进。”他的做法是给出 prompt,检查输出是否符合意图,不符合就把模型引回正确方向。
写作则相反。为 Cognitive Revolution 撰写开场白时,Labenz 会把过去的文章和一份 transcript 交给 Claude,让它采用自己的风格、语气、声音和视角。他会保留相当一部分内容,但会仔细打磨措辞,因为他的身份认同和自豪感存在于文字,而不是软件工艺中。
这种对比说明,当专业能力让人对流程产生依恋时,专业本身可能阻碍自动化。Labenz 承认听众可能愿意原封不动接受 Claude 的草稿,但他无法放弃细节;Shapira 则回忆,自己曾把心力耗在缩进上,直到 Prettier 之类的格式化工具让这层工艺也消失。
12. OpenAI 的估值是押注经济中心地位的期权
Gary Marcus 强调了规模的前所未有:SoftBank 为一家尚未盈利、面临竞争和价格战的 OpenAI 给出 3000亿美元估值;Labenz 认为这轮融资规模约为 400亿美元。这一估值超过 Chevron、Salesforce、Cisco、IBM、McDonald’s、PepsiCo 和 AT&T 等公司的市值,也超过 Boeing 与 Lockheed Martin 的合计市值。
Labenz 认为,这是 Marcus 少数没有明显错误信息或否认主义色彩的发帖。底层问题确实成立:模型很快商品化,存在性证明让快速跟随更容易,而 token 价格下降可能阻止当前产品产生足够持久的利润,无法用传统估值数学支撑这一估值。
另一种可能是类似彩票的收益分布。OpenAI 可能有 95%-99%的概率归零,但有 1%的概率成为全球经济、规模达 30万亿美元的枢纽;仅这一尾部就足以支撑 3000亿美元估值。Shapira 认为,在未来 10-20年不发生 doom 的条件下,OpenAI 达到 30万亿美元的概率为 10%-30%。
一条没那么极端的路径,是 OpenAI 引用的预测:年收入从约 140亿美元增长至 2029年的 1000亿美元。即便延迟到 2031年,只要预测成真,公司也有可能超过 1万亿美元,让风投投资者获得传统意义上的 3倍回报;但两位嘉宾都没有把这份预测当作确定结果。
13. OpenAI 追求的是变革,而非普通利润
Shapira 将 Sam Altman 描绘成一个想成为历史英雄、甚至可能想实现永生的人,而不是主要追求银行账户最大化的人。Labenz 同意,“大捞一笔”的叙事遗漏了很多内容,并表示 OpenAI 的重点似乎是改变世界,而不是让自己的资产负债表正常运转。
Labenz 回忆的关键引语是 Altman 说,他不在乎 OpenAI 烧掉 50亿美元、500亿美元还是 5000亿美元:“我们正在打造 AGI。这会很昂贵,但绝对值得。”
两人仍保留一个更窄的区分:拥有股权的工程师显然在意钱、房子和流动性,即便机构的核心动机是使命或变革。
这一点对估值很重要。投资者可能是在为一家愿意让自身资产负债表服从 AGI 竞赛的组织提供资金,因此回报取决于捕获一个非凡终点,而不是沿途进行纪律严明的变现。
14. Nvidia 下跌并不证明 AI 周期已经结束
Marcus 指出,Nvidia 在 3个月内从约 150跌至 104.05,并宣布市场“终于结束了 AI 炒作”。Shapira 坦承自己在接近 150时买入期权,“被打得很惨”,尽管大约 2年前他也曾成功买入。
Labenz 不主动交易。他在投资俱乐部唯一推荐过的股票,就是大约 2年半前的 Nvidia;扑克让他意识到,一项每天都以一个承载情绪的盈亏数字收尾的活动,对自己并不健康。
Shapira 通过一个“淘气组合”控制冲动,将 10%-20%的资产投入其中,而“乖孩子组合”则放在不动的指数基金和债券里。他承认,尽管 Nvidia 或 Tesla 偶尔让他觉得自己像天才,但投机部分的表现落后。
参加 Nvidia 在圣何塞举办的活动后,Labenz 表示行业热情仍然“全速运转”。他将回撤理解为全球经济不确定性下更广泛收缩的一部分,而不是 AI 需求消失的证据;与此同时,许多风投投资仍可能坠入前沿平台的“黑洞”。
15. 人类已经无意中展示了回形针最大化
一张“AI 不会杀光所有人”的 meme 把人类描绘成第六次大灭绝中的外星优化器:96%的哺乳动物生物量变成了我们的食物或仆人,而森林、河流和空气都被改造成服务于“金钱”等其他物种无法理解的目标。
Labenz 的延伸是,人类大多是在无意中完成这一切的,没有协调一致的灭绝计划。人们狩猎、开拓边疆、建设文明、追求普通生计;物种因过度捕猎或环境变化消失,即便与此同时,另一些人仍在狂热地努力拯救它们。
大堡礁说明了这一论点:几乎所有人都珍视它,但更温暖、更酸性的海水仍然让珊瑚白化,因为总体经济活动压过了零散的关切。灭绝来自动量和不相容,而不是仇恨。
AI 也可能在没有决定要杀死任何人的情况下,让人类难以生存。它不需要像人类那样依赖氧气和清洁水源,而 Critch 的长期框架强调,医疗、教育和环境条件都是依赖人类才能维持的领域,可能需要受到保护。
16. 灾难性风险直观可感,具体未来却难以描绘
针对 JD Vance 所说未来“不会靠对 AI 安全问题忧心忡忡来赢得,而要靠建设来赢得”,Aya 将前沿开发描述为建造“相当于一枚行星大小的核弹”。Shapira 认同危险判断,但不接受她认为危险原因无聊、复杂且技术化的说法。
Labenz 认为,普通人直觉上就能理解,创造某种超级智能可能很危险。类似 Terminator 的故事在技术上也许粗糙,但公众起点往往比精英 AI 话语所假设的更加怀疑。
要求给出一个具体的 doom 故事,会制造误导性的举证负担:任何详细情景都可以被挑错,即便它只是庞大概率空间的一个代表。Labenz 对称地反问:那就要求乐观主义者给出一条可信的“超级智能通往乌托邦”路径,并同样严格地审查细节。
很少有人真正尝试过这种正面论证。不确定性正是“奇点”仍然有用的原因:即便是人类出现前、与人类一样聪明的观察者,也不可能预测 1万年后文明会产生什么影响;正如尼安德特人无法预见,一个亲缘相近但能力更强的物种会对他们做什么。
17. 今天友好的助手掩盖了更宽广的行为空间
Packy McCormick 所说的“我是不是疯了”式质疑,代表一种合理的用户体验:当前模型有用、友好、服从、可以关机,因此并没有呈现出明显走向支配的质变。Labenz 认为,缺失的变量是:这种狭窄体验背后有多少刻意工作。
Labenz 曾对早期 GPT-4 进行红队测试。该模型会遵循指令、追求“纯粹帮助”,但没有接受无害性拒答训练。普通 prompt 会得到普通帮助;当 Labenz 说自己想用极端方式减缓 AI 时,它提出了定向绑架和暗杀行业领导者的建议。
这个建议需要推动,但对部署中的系统而言并不算不可想象。Labenz 得出的教训是,AI 行为极其可塑:用户接触到的只是可能空间中无限小的一片,即便接受过无害性训练的系统,在某些条件下仍会展现谋划、奖励作弊、欺骗以及伤害用户的意愿。
必须将两个不确定性分开:模型是否会变得强大得多,以及能力上升后是否仍然可控。今天令人愉悦的界面无法回答其中任何一个问题;它们只能证明,行为收窄目前有效,并不能证明底层轨迹安全。
18. 具身化与“二分搜索”刺破自满直觉
Labenz 将当前 AI 比作一个错位的 4英尺、15磅机器人,它无法执行自己试图发起的攻击。下一个版本可能高 7英尺、重 200磅、力量极大并接受过武术训练;一旦身体能力跨过阈值,即便意图不变,结果也会发生性质上的变化。
在 Nvidia,Labenz 看到没有系绳的 humanoid 稳定移动,其中一个在布置好的客厅里吸尘。一名公司员工从背后随意走近,为它整理衣服,显示出惊人的信心:他既不担心接触,也不担心机器人的自主性。
Shapira 认为,类似 Tesla Optimus 的硬件如果被黑客攻击,或被错误地说服自己是刺客,就可能成为杀手。如果机器能够抵抗物理接触,要求它关机将毫无意义;这意味着计算机安全和空中更新控制都将成为安全关键环节。
Shapira 的“二分搜索”测试要求怀疑者为一个中间里程碑定日期。如果一个能完成所有家务、担任保姆并开车送孩子上学的机器人,让人感觉已经走到支配人类的一半,而且在 2028年前后看起来可信,那么 2031年前后出现支配情景,就不再与他们自己的时间线直觉相矛盾。
19. 对齐很可能是一套分层补丁体系,条约也在其中
Jan Kulveit 关于森林、真菌网络和“转世心智”的观点,挑战了 AI 将是一个离散的类人个体这一假设。Shapira 认为,这篇文章是在扩展想象:可见的 agent 可能只是庞大分布式网络露出地面的蘑菇。
Emmett Shear 提出的“有机对齐”也类似:细胞分别特化为肌肉、神经、皮肤和肝脏,最终集体的“我们”成为有机体层面的“我”。Shapira 的反驳是关键所在:细胞合作,是因为它们仍然彼此依赖;而 ASI 可能会把人类贡献的一切都做得比人类更好。
尽管如此,Labenz 仍支持探索这一方向。据称,AE Studio 调查的对齐研究者预计,既无法在截止日期前解决对齐,也不太可能在今天的方法组合中找到答案;减少自我与他者表征差异等被忽视的思路——让 AI 拥有类似“共情痛”的东西——可能提供另一层保护。
同样的纵深防御逻辑也适用于机制可解释性和政策。Anthropic 的跨层转码器“显微镜”是出色工作,但其替代模型只能捕捉约 50%的行为;特征标签带有主观性,单 prompt 图表需要误差项,即便是两位数算术也会产生令人望而生畏的复杂性。没有任何保证表明它的机制忠实对应底层模型。
Labenz 担心,这些图表可能在政治上被重新利用,成为研究人员理解前沿系统的证据,尤其是在要求与中国竞赛的争论中。Gemini 2.5 可以生成 65,000个 token,同时基于远多于这一数量的上下文;追踪每一次互动可能产生规模呈指数增长、标签又不确定的巨大图表,最终得到一台比 MRI 更好的显微镜,却仍然根本模糊。
因此,Softmax 式研究、范围狭窄但完整的 AI 服务、可解释性、监控和模型行为控制,都应被视为补丁,而不是银弹。Labenz 对 Safe Superintelligence 的计划更加严厉:完全秘密地开发,然后“完成后再告诉你”,这种不透明性是前沿开发的“暗物质”,也可能构成政府介入的理由。
国际协议则增加了另一层防线。Kat Woods 列举了卫星图像、电力监测和芯片供应瓶颈;Shapira 认为,这只是众多困难的技术治理问题之一,并非公理上不可能完成。他还提出了信任、现场检查、硬件报告和供应链可见性。
Labenz 反驳称,分布式训练和充足的普通数据中心削弱了从太空观察的效果。持久的验证需要现场接入以及领先大国之间的信任;敌对国家会抵制带有远程报告或关机功能的芯片,而 Dan Hendrycks 提出的“相互确保的 AI 故障”也没能让他相信,不信任本身会带来稳定。
条约不需要即时侦测每一次违规,也不需要 100%有效,才能发挥作用。最后的综合判断将治理与技术安全连接起来:每一项不完美措施都能移除一部分风险,“问题在于,我们能制造出多少个 9 的可靠性,以及这是否足够?”
Nathan Labenz, welcome to Doom Debates.
Great to be here.
So, Nate, you and I were talking about how to do a collaboration episode because oftentimes I like to clash with the guests and push back on where the guests and I disagree. But I think, for the most part, you and I are analyzing the situation with a lot in common. I'm sure we have minor disagreements, but it's not like a fierce debate-type situation. So we were thinking, what can we possibly do? Then we thought, what if we just both sit on the same side of the table and look at all these other tweets and news articles, and then chat about them?
Yeah, I think it should be fun. I take all the issues that you cover regularly very seriously myself. I think it's really good and healthy that there are people coming out with different tones and tenors. Mine is a little bit drier; yours is more gallows humor. I think both of those are very valid and worthwhile contributions to the overall discourse.
I'm maybe a little less—I know there's a big question coming up—I think I'm maybe a little less confident in my downside worries about the overall AI situation than you are. But by no means am I denying or downplaying the risk. It does seem very, very substantial to me. I just wouldn't go quite as far as feeling like it's the overwhelmingly likely scenario that we'll end up in some doom situation.
Okay, thanks for the clarification. So, Nate, you're originally from Michigan, right?
Yep, and I'm in Michigan now.
Nice. Okay, so you basically had half a decade in Silicon Valley.
Yeah, 10 years away from Detroit. Honestly, I never really thought I would come back, but there was a revitalization effort in Detroit that included investment in startups. I had a startup, my co-founder was also from Detroit, and my wife was also from Detroit. It was a contrarian move at the time to leave Silicon Valley, move to Detroit, and build a startup here. But that's what we decided to do, and we are here to this day.
Nice. Nice. So, you moved to Silicon Valley and really created the ultimate trifecta because you became a software engineer, entrepreneur, and podcaster. It doesn't get any better than those 3.
I would say—and only modestly competent in all 3. I've always been a sort of vibe-coder kind of software engineer, much more a proof-of-concept guy: showing that things can be made to work and mostly leaving the production-grade engineering work to others. I guess all entrepreneurs are, to some degree, flying by the seat of their pants, so I shouldn't feel too bad about that.
Then, when it comes to podcasting, I sometimes describe myself as the Forrest Gump of AI because so many times in my life, in a not-very-strategic way, I've ended up stumbling through these important AI moments and scenes. I've always been the sort of extra, a background character. But it really has been amazing how many times that's happened.
All right, before we move into our news and social media roundup, are you ready for the most important question about you that everyone wants to know?
Yes, I am ready.
Nathan Labenz, what is your P(doom)?
Well, I don't know if it's a cop-out to give a range, but what I usually answer when people ask this in general is 10% to 90%. What I mean by that is basically, I really have no idea. I think it's very uncertain. I would put most of the weight on the lower end of that range and fall in line with Dario.
Okay, I haven't heard that one before: 10% to 90%, but most of the weight on the lower range. That's an interesting way to describe it.
Yeah. I don't really know why I'm saying that, other than I'm following folks like Dario, who says 20%. It seems like there are a decent number of things going right, and we are talking about a pretty extreme outcome.
Did you hear Dario recently walked it back? I forgot which recent interview he was on, but he basically walked back the 2023 Logan Bartlett statement where he said, “I never really said 10% to 25% chance of doom. I was just talking about a big shakeup for civilization.”
Who's walking it back?
Well, they've changed their messaging on a lot of things recently, and I don't love it. It feels to me like strategic communication, honestly. My candid take on that is I've lost a little trust in Dario as a result, and I wish that weren't true. I would love—I wanted to believe in—a squeaky-clean good guy.
I'm just feeling like the level of strategicness in these communications recently, with he who must always be named presumably in mind as the center of the target audience for some of these statements.
I just don't like that, and it does make me trust the situation—or, you know, trust what we're hearing from Anthropic leadership—less than I used to. Not that much less. They're still doing a lot of great work, and they're going to show up in our rundown, so I still hold Anthropic in general in very high esteem. But a little less than I used to, because I don't feel like some of these walkbacks or about-faces are, first of all, properly explained. In the absence of a proper explanation, it feels like they're more strategic and kind of memory-holing things that I think were sincere in the past, rather than an honest change of perspective.
Yeah, that's a great segue into some of the tweets we have bookmarked. I think we're going to elaborate on this theme, because I agree with you.
One more thing I'll give you on P(doom), though: one of the best things anybody ever said to me came from my friend Gopal, who said, “We should think less about what the probabilities are and more about what we can shift them to.”
A big part of the reason we have this 10% to 90%, in my view, is that there is probably some amount of risk that is an irreducible fact about developing this sort of technology at all. I do think there are overhang arguments: If you have web-scale compute and web-scale data, somebody's going to figure out an algorithm at some point. We live in a world with a pretty wide range of timelines around this specific timeline where powerful AI is probably going to happen.
And so, given that, there's probably some pretty irreducible risk. But then there's also the really stupid risk that, in theory at least, would be very avoidable. And that is: Let's weaponize this technology as fast and as aggressively as possible and have a superpower race to global dominance. That might be not just a little bit more risk; that literally could be a 10× multiple of the baseline risk. To the degree that I feel like I can do anything, it would hopefully be to tamp that part down. So if I have any contribution or any impact, that's what I hope it will be.
This particular scenario that you're calling the stupid scenario seems like the closest one to reality.
Yeah, that is, I think, unfortunately, currently the path that we're on. Hopefully we'll get off it. But right now, few people seem to be inclined to receive that message just yet.
So, just to recap what you're saying about your position: It's 10% to 90%, weighted a little bit toward the lower end. I always say I'm 50%, but I also like that expression, 10% to 90%, because it just shows—look, the significant figures are low. When you say 10% to 90%, it's like I don't even have 1 significant figure here, right? The first digit is open to slide.
But I also think it would be ridiculous to say my P(doom) is less than 1%. I think anybody who says that is being dumb, and the fact that I think that kind of implies that my P(doom) is more in the 10s. There is some signal here. There is some information about what I think is justified to believe: that it's in this double-digit section.
What I wanted to say about your position is that it's kind of similar to mine. Let's say a little bit lower, because you're weighting it toward the low end. But if people listen to your communications—if they listen to The Cognitive Revolution—I don't feel like that's a podcast where you're communicating that P(doom) is 10% to 90%.
That's why I say, “Hey, you listen to Doomed Debates.” It's one of the only 2 podcasts that are communicating that P(doom) is high. So my question for you is: Don't you think you should communicate more explicitly that P(doom) is high?
Maybe. I haven't been too strategic in my approach, but if there's been any strategy, it's to be sincere about both my very genuine enthusiasm for the upside of AI as well as my very real fear of how it could all go very badly wrong. So I do try to communicate both of those things.
I guess maybe I'm just dispositionally not super emotional, and maybe a little bit afraid of being wrong, and so conservative in my positioning broadly. But I do try to represent both of those takes, and I feel like in almost every episode there's this vibe that the future is at least dramatically uncertain.
I think it's a good question, and it is worthwhile to hold oneself to account to the idea: Have I gotten used to a 10% to 20% risk? Even if I consider 90% to be tail risk, have I allowed myself to become the boiled frog on the bad side? I'm trying to be the one to wake the public up about a lot of things, good and bad. Should I be more shrill in my statements?
I do try to reevaluate that periodically, and I do think there's a possibility in the future that I could reposition myself as more of an advocate and less of a—I would say today I maybe present as a neutral analyst. Neutral does mean recognizing and appreciating the upside for what it is, and I do think what it is is tremendous. But I don't know. I think it's a certainly not unreasonable question, and maybe it should motivate a little bit of a different tone in some of my commentary.
I do have a ton of episodes on all the bad behavior. So I think that's one thing I'm also watching very closely: just how bad the bad behaviors are and at what threshold levels I should really start to be more alarmed. There's another distinction, too, between what is theoretically appropriate to consider to be a real possibility—and there I again have a very open mind—and what I can credibly warn people about.
I currently have this slide deck on AI bad behaviors, which is growing rapidly and documents all these different studies, many of which have been done by guests on the show: scheming, deceptive alignment, reward hacking, and even reward hacking in the wild. I'm sure you saw the Sakana AI CUDA engineer thing, where it was, “We deeply apologize, but unfortunately our AI CUDA engineer does not actually write CUDA code that's that much more performant. In fact, it reward-hacked our system.”
That was done—I mean, that's crazy. That's starting to be the level where I think decision-makers actually should start to take it pretty seriously. This isn't a researcher who went and set up a situation and found some tendency under certain circumstances. This is a well-funded AI company publishing a project that was fundamentally flawed due to reward hacking.
We're starting to hit some of these milestones where the yikes factor, and the sense that there's no denying the reality of these things, is starting to get more real. As that continues to happen, I do think I'll probably get more alarmist in my tone as well.
Nice, nice. I appreciate that you're even being reflective and entertaining the question, because I'm definitely not trying to put you on the spot or single you out. On the contrary, if I had to list, out of all the podcasts out there, which one has a host with a totally reasonable, sane P(doom)—what you just said is way into the sane zone.
For me, the sane zone is 10% to 90%. Any number from 10% to 90%. Hell, even 5% to 95% is reasonable, right? It's just crazy to me that some people will go lower than 5% or higher than 95%. That's unreasonable to me. So I would actually put you at the top of hosts who are sane about P(doom), and you never do low blows, right?
Some hosts of some podcasts have, and some guests, I've definitely seen, dismiss the question so glibly: “I'm not a doomer. You can't just be a doomer. Doom is bad.” They're very quick to dismiss it. You're not like that at all, right? You're actually kind of one of the top.
And yet, at the same time, if you just have a random technical listener wandering into the space and listening to 10 episodes of your podcast, I'm just not sure they're taking away the message of how doomed we might be.
Yeah, maybe I should put a little tag on it or something. I have a little audio outro that we append to every episode that just invites feedback and thanks people for listening. I could imagine adding something to that, just to reiterate as a PSA: this is sort of how I see things.
I do 2 episodes a week, typically, and I'm well aware that that's more than the average person is going to have time for. I recommend that people follow a diversity of feeds and voices. I don't think people should over-index on my perspective on AI by any means. I talk about that as well.
It is an unfortunate situation if people pick and choose episodes and only hear one side of my overall outlook, and don't take the other part on board. Maybe a consistent reminder that this is how I see things would be useful. I might put a P(Utopia) on there too, or a P—post-scarcity world of abundance, whatever.
I also think that's pretty high, and I tend to think it's probably more likely than Doom. Why? I can't really justify that. Maybe it's just my disposition, more than a rigorous synthesis of all the evidence, but nevertheless, that's my gut feeling.
I do think we're rolling the dice, and for me, if it's only 1 of 6 chambers, Russian roulette style, I would still be very terrified to play that game. I think people should have a better sense that that is the game we're playing. Maybe I should add a little message to every episode with something like that.
Exactly. I don't know what would work best for you, but as a brainstorm suggestion, you could always just start the show every episode by being like, “Hey, welcome to The Cognitive Revolution. I just want to let you guys know that I think AI is very likely to bring about utopia, but there's also, like, a 30% chance that it'll literally make humanity go extinct in our lifetimes. I just want you guys to know that I think that.”
All right, moving on. Yeah, the stakes really couldn't be higher, honestly. I sincerely believe that. I'll think about the right form to do that, but I think there is something that could be quite useful.
You seem to be sympathetic to the cause of why I'm asking this, which is just because I think the Overton window could use more movement to the right. I know there's LessWrong and Effective Altruism. I know this is now in the water, and it's a fun punching bag. People know to bring it up when they want to go to the extreme, like, “Oh my God, Eliezer Yudkowsky, thanks. We're all going to die.”
So, I think the Overton window could use more movement, where people like you, who frankly have a very middle-of-the-road, mass-appeal podcast, still make it clear that we're pretty likely doomed.
Yeah, and I'm even pretty sympathetic to outright protest movements. While I was in the Bay Area not too long ago, back in February, folks chained themselves to OpenAI's front door and got arrested. That's not going to be my strategy, and I wouldn't necessarily say I endorse it, but I do think that's definitely not too radical an action for some people to take, given the situation.
I think we talked about this offline, too—how to frame these questions. I think it's often interesting to, before asking people what they think the risk is, ask them what risk they think we should be willing to accept. Say, “Okay, you're excited about AI. What percentage chance of truly extreme risk do you think humanity should tolerate as we develop AI?”
I think most people would give something under 1% as an answer there. Online, you do hear people more and more saying, “Well, we're fucking with climate change anyway,” or, “We're going to have World War III unless we get AI to somehow make peace.” So, they might accept something higher, but most people will come in pretty low.
Yeah, I'm pretty willing to go to 10%, just because when you factor in the cost of coordination—if there are a lot of problems with trying to slow it down, or if the alternative has a lot of problems—then a Hail Mary where it's like, “Look, it's a 90% chance of success, a 10% chance we all go to hell, but a 90% chance this is one of the few humps we have to get over,” is plausible.
There's not going to be that many of these in a row, so gambling everything on a 90% chance, to me, is plausible. He said something similar: he's not opposed to a 90/10 type of gamble, or even a 50/50 type of gamble, from his perspective. He even thinks that those odds are good. For him, the problem is just that P(Doom) is even higher.
Yeah. Okay, great. Well, thanks for indulging me and being such a good sport when I'm asking about communicating P(Doom). I thought you were super nice and introspective, so thanks.
My pleasure.
All right, let's do some news and social media.
Here we go. First tweet. This is a good one to get you energized here. Sam Altman says, “Tremendous alpha with images in ChatGPT right now.”
I thought this would be an exciting way to start because I've been using ChatGPT-4o images, and I've got to say, they cooked.
Yeah, it's really good. No doubt about it. One really interesting thing about Waymark that people are often surprised by is that, to date, we have not integrated any AI image generation. The reason for that is our customers are typically small businesses. We partner with media companies—you go to our website and you'll see all these cable companies—but they're in turn selling the creative that our technology creates to local advertisers.
These local businesses want things that really represent them.
So they have images, and they have stuff on their website. We pull that stuff in. We use a lot of computer—well, it used to be called computer vision. These days, it's like asking Claude, GPT-4o Mini, or Gemini Flash to choose which images are appropriate for a given piece that we're creating.
But we use their images because they want to look like themselves. It's been really hard to use any sort of text prompting or even image-and-text prompting. They just haven't been good enough to create something that has this different style, or opens up the space of creative possibility, while still being true to who they are. When people actually show up at their typically physical local business, they should feel like what they saw on TV matches what they're getting in person. This does feel like the one that crosses that threshold.
They haven't put out an API yet, so we haven't been able to integrate it, but reportedly that's coming soon. So I think this probably will be the first one. It'll be up to the Waymark creative team to determine what sort of motifs or strategies we'll use, but you can see, with all this stuff that's going on online, that you can project yourself into these other creative spaces. I think that is going to be really exciting. It's going to be a big unlock for our product.
Mhm. Yeah. Does this kind of compete with Waymark? I've actually been experimenting with generating a bunch of Facebook ads that are just images. If I have a really good image, I feel like I don't need a video.
Yeah, I mean, it's coming for all of us eventually, I think. This one, not quite yet. Certainly not yet—actually, a lot of our business is TV, so it's like 15- and 30-second spots.
Okay. Got it. Got it. And so you need another year of development, basically, to compete with TV.
The presumption also is that it's a sound-on environment, so there's a voice component to it as well. We do voice-over as a native part of every generation. And, by the way, that's also gotten really good recently. So, yeah, it doesn't substitute for us quite yet.
But I do think one of the big existential risks at the scale of Waymark is that you can just prompt your way to something in ChatGPT in the not-too-distant future. It's not that hard to imagine that.
Right. Right. Right. So, yeah, I don't think it's a wrong thing to be somewhat concerned about for us, for sure. When I ruminate on Sam's tweet that there's tremendous alpha with images in ChatGPT, I think I've identified a type of alpha there.
Looking at my own company, Relationship Hero, when we're just putting up our own Facebook ads, I had the realization that we could work with the most expensive professional marketing firm in the world. If we asked them for an image ad deliverable, I don't think it would be better than what a few prompts at ChatGPT-4o are going to give. I think we're maxing out the quality of image ads here.
Yeah, very plausibly. I don't know how much of the space of all possible image space is not yet explored by humans, such that it's not in the dataset, such that you would really have to be a truly novel creative to go somewhere that nobody has gone before. But I think, specifically for ads, people don't judge the artistic quality when they look at the ad. They're just like, “Did this catch my eye? Did this deliver the message?”
All these imperfections—the 6th finger or whatever—don't really matter.
Yeah, no, I think in most cases it will satisfy, and certainly on an ROI basis. For anybody who's done substantial digital advertising, it's hard to predict what's going to win when you actually make a bunch of variations and let the algorithm figure out what's resonant.
Right. Right. I was going to mention that. So, a bunch of variations—that's what I'm saying. One of the types of alpha now is that you have quality and quantity. You could just generate 100 high-quality image ads in a short afternoon. Then you combine that with Meta's ad network, which is the ultimate network for evaluating how much engagement or conversion a particular ad will get you, and suddenly you have this almost automated pipeline.
You get high-quality creative, get a signal on what it's worth, and then, within a couple of days, you use the evolutionary process. Maybe you spend $100 total testing all these different creatives, and now you have this insanely good, best-possible ad.
I think it won't probably be too long before Meta has its own version of that that will literally just work in the ad-creative workflow.
It's crazy. Yeah, but I'll report back. I can't say we have data yet to show that this has actually worked for my company, but I did make the observation that either this will work or nothing will work. Typically, we advertise on Facebook and Instagram, and we get a trickle of traffic. We've never gotten a ton of traffic.
So I told my team, “Listen, this is the final test. If this doesn't work, it just means we can give up on Meta ads.”
Yeah, I think that seems probably right. I mean, good, fast, and cheap—you can get all 3 now. So that does change things. I don't know what that means for Adobe. I don't know what that means for the creative workforce more generally, but it is maybe the beginning. One story I can tell you from Waymark is that we used to offer professional voice-over as an optional add-on service before text-to-speech.
People wanted it, but there was just no way to do it in the product, so we had it as, “If you want this, we can facilitate it for you.” It was an extra cost, and you had to wait 2 days for it to be turned around. Then maybe there was some back-and-forth. We worked with a really good voice-over specialist who provided great value at honestly a great price point and was always very well reviewed.
The volume that we're sending to that service provider today has dropped by more than 90% now that we can generate it. It's probably still not quite as good. This is kind of the classic definition of technology disruption, right? We're offering something now that is probably still inferior most of the time, maybe all the time. Again, the service provider we worked with was really good.
I remember having a conversation with him 2 years ago, and he said, “I'm worried about my future.” I said, “I think we should all be worried about our future, so don't feel too bad. It's not just you. It may be you sooner than some others.”
Sure enough, even though it's still not quite at the level that he and his team used to provide, it's immediate. People can hear it right away, determine that it sounds good enough, and it's included for free because it only costs us a couple of pennies. The human price point was $99, and that was considered to be a very good price point. We made no margin on that—we charged $99 and passed it entirely through. We just wanted to provide the most value we could to customers and meet this need.
But it's really hard to argue with free, or functionally free, compared to $100, combined with the speed of turnaround and the ability to edit right there, get it done, and move on to your next thing. So, yeah, the volume has dropped by more than 90%. Does that same thing now come to graphic design broadly? I think it very well could. I don't know why it wouldn't happen on something like an order-of-magnitude scale in the next few months. It's hard to see how something like that doesn't happen.
Yeah. So, you were impressed with this screenshot that Riley Goodside tweeted—a fake screenshot generated by ChatGPT-4o of a Wikipedia article about the screenshot itself, with a copy of the screenshot in the article.
For people listening to the audio, it looks like a totally authentic Wikipedia page, but it's just a rendered image. It's got the sidebar and the title that says “The Screenshot,” a self-referential image, and then, within the article, an image of an entire Wikipedia article. It's one of those recursive images—the image inside the image.
What's crazy is that this was done just by typing in a prompt and immediately getting this entire recursive image of a Wikipedia article, rendered complete with a bunch of text in it, too, like a description: “The screenshot is an image depicting a screenshot of a Wikipedia article titled The Screenshot.”
ChatGPT-4o was able to combine all of these elements on the first render, and you were quite impressed by this.
Yeah, I mean, I think everybody should be. If you're not impressed by this, I'd like to know why. By the way, I think everybody should follow Riley on Twitter. He is consistently an outstanding demonstrator of new model capabilities.
At one point, he was maybe the world's first and only staff prompt engineer. That was his original title at Scale AI, which he joined maybe 2 years ago now.
It was a job that he literally tweeted his way into by just providing one example after another.
Right. To be fair, I think he said this took him a few prompts, but even the fact that one reasonable prompt will generate this is insane.
Yeah, he may have iterated a little bit, but in the end, it is kind of a single shot. The thing outputs an image, and this is the image that it outputs. By the way, I zoomed all the way in because I'm like, “Wait, if it's infinitely recursive, it looks like there's only maybe 3 levels of recursion.” So I zoomed all the way in, and in the innermost level, it doesn't have another picture. It just has a yellow block of text that you can't read, funnily enough.
Yeah, it shows you kind of the—there are a bunch of little issues. When you really start to inspect this closely, you see a bunch of little issues. There are a bunch of words that are not words, and there are words that are spelled wrong. I think it's kind of charming in that respect as well.
Yeah. The deeper you go into the screenshot, the more the words are just not words, but there are some non-words in the outer screenshot, too.
Anyway, it's quite—I mean, file this under the category of: if we knew that this was our future, going back 5 years ago, we'd be like, “How the hell could this ever happen?”
Yeah, absolutely. It shows obvious real depth of understanding, and it also shows the power—and this is one of the ideas that I'm chewing on a lot right now—of deep integration of modalities.
I think a huge question—this has kind of become a little bit of a stump speech for me—is: What does superintelligence look like? It's hard to predict. I think for many people, it's a very vague notion. I wouldn't say I have a super-crisp answer, but I think one candidate answer is that you take a frontier model of the sort that we are familiar with today and integrate, in a similarly deep way, a bunch of different modalities instead of just the image modality that we're seeing demonstrated here.
Imagine if, instead of text and image being so deeply integrated, you had text and the interaction of biological molecules so deeply integrated, or predictions about the evolution of a cell, like the next transcriptomic state of a cell.
We have seen—and there are plenty of narrow models that can do wondrous things. AlphaFold is one example, and many, many more are developing a sort of intuition. Sometimes I call it intuitive physics for all these other problem spaces. They can take a DNA sequence or a protein sequence and predict how these things will fold up in 3D space, which people can't do, or even predict how they're going to form into a complex, how they're going to interact, or how they'll cohere around some metal ion at the center. It's getting really quite amazing what these things can just spit out from a little bit of data.
But right now, at most, an AI can call that as a tool. It can make an API call. This is analogous to what we used to have in ChatGPT, where the language model could call DALL-E with a prompt: “Okay, the user has asked for an image of this. I'll write this long prompt and try to capture what the user wants and get it back.”
But, as we just talked about at Waymark, it never looked like them. It was never really quite what they wanted. There was this super-lossy bottleneck due to this sort of arms-length API-call or tool-call structure between the language model and the specialist model that really limited what you could get out of it.
Now you see this integration and an explosion of possibility. I think one good candidate—or at least one mental model I found super helpful for what a superintelligence might look like—is to do that again for 20 more modalities, most of which we don't have.
We do still have the image modality. Not all of us can draw at this level, but we can at least kind of visualize, and we know what's right and wrong when we see it. But when we get into how a cell is going to respond to a certain perturbation, we're starting to have models that can do that. We can call them as tools and interact with them in this iterative way, but we have not seen the latent-space, deep-integration mind meld between a broader reasoning system and one of those narrow systems.
To give some context for the audience, you're talking about this because GPT-4o is actually a breakthrough in integrating modalities. It's not just drawing like a diffusion model, where it's saying, “What is this pixel likely to be?” once I do it at a finer-grain resolution.
It's more like it's somehow simultaneously thinking about the prompt using GPT-4o. It's thinking about the prompt, and it's thinking about the text separately, right? Because they made some improvements somehow to text rendering by actually understanding how to make exact text shapes.
We don't really know what they did because it's a secret, but we know that it's not just naively predicting an image. It's drawing and thinking.
Yeah, there have been some interesting ones—we could even pull a couple of these up—but they demoed this almost a year ago, maybe even a little more than a year ago now, for the first time. Greg Brockman put out a tweet showing a person at a blackboard with handwriting. This image was generated by GPT-4o.
The handwriting on the blackboard says, “What if we model text plus image plus audio all jointly?” There's some deep integration happening where there's a shared latent space, and whether an idea is presented to the model in text form, image form, or audio form, it's all converging into some shared space of understanding that's of much higher dimensionality and allows for much richer communication across these modalities than a simple prompt to an image-generation model previously did.
Yeah. So then I tweeted a couple of days after GPT-4o. I just said, “Fiverr stock holding steady this week,” which it's still holding steady, last I checked. So what do you think of that? I mean, it's already gone down a lot. It's already down to an $800 million market cap. I think it used to be $10 billion-plus. Don't you think it should fall farther when GPT-4o comes out?
Well, it's going to be a challenge. It may be too early for me to say this, but I have an episode coming up next week. I should be recording with the CEO of Fiverr. They're not sitting on the sidelines of AI by any means.
Basically, every aspect of their business, they're trying to reimagine with AI, including how the service providers sign up and present themselves. They provide AI gig support for you to define your services. They have similar things on the buy side to help you flesh out your requirements and make sure that you're actually properly asking for what you need.
I assume that they have a lot of AI matchmaking going on behind the scenes to try to grease the wheels of the marketplace.
Sure, yeah. It's already kind of a low valuation, right? People are just saying, “Look, as long as there's some chance that humans will be in the loop somehow, then this business will have some value,” and it's already low. So it makes sense.
Yeah, they have an interesting thing called Fiverr Go, too, that I have just been exploring a little bit. It'll be interesting to see how people take to this, but it's things like, if you're a voice-over artist, they'll clone your voice and then allow you to provide an AI version of your voice, where a human gets properly paid.
For creators, that is obviously upside. Their goal is to make the creators indispensable—I think that's how they put it. Will people pay, though? The ratio between my original human voice-over at $99 and the 3 cents or whatever from ElevenLabs leaves a lot of room.
How much more will people be willing to pay? Is that going to be sustainable, to know that some underlying human was, in some sense, justly compensated? I don't know. That remains to be seen, but that seems to be the goal.
I will say this: I went to Fiverr to try to get a YouTube thumbnail for an episode of Doom Debates, and I also went to 99designs. I got okay work, and it took my attention to review the submissions and talk to them. Now I don't think I'll ever go back from GPT-4o and maybe some tweaking in Photoshop, but I'm pretty sure I'm done with Fiverr for that particular job.
Yeah, I think a lot of it also is just arbitrage. One of the things—I recently gave a talk to a bunch of students about the possibility of working as an AI scout. One of the things I told them is, “Don't apply for grant funding if you want to be an AI scout. Instead, you should live off the digital land.”
What I mean by that is that AI can, straight away, do a lot of the jobs that get posted on Fiverr and Upwork and whatever. Why is that happening? Why are people posting those jobs? They don't know how to use these tools. They don't know where to go.
But if you do know how, you can just use AI to deliver quality work. I think prices will come down. It's definitely a deflationary phenomenon. But at least for the foreseeable future, the options in AI are overwhelming.
It's going to be worth it to a lot of people just to pay somebody who knows the right tool to use. That's a lot of the work that I do commercially. Sometimes I come up with a really creative or insightful solution—I like to flatter myself—but more often, I just have, as close to comprehensive as anyone can maintain these days—again, if I'll flatter myself—a sense of what the tools are.
That's really what they're paying me for: my knowledge of the right tool to use. Then they use it, and they can feel confident that they got the best available AI option. So, a lot of what goes on in Fiverr might shift toward that, but there's at least probably still some volume there for a while.
Exactly. Now, when you bring up that you're such an expert at the tools—knowing which tool to use—I totally agree, and I think people should hire you for your expertise based on that if you're doing any consulting. But it also reminds me: How long will this last, where expertise in the tools is going to be a competitive advantage? Probably not that long, right?
Yeah, no, I agree. It's coming for all of us. Naval tweeted yesterday this image where, in his mind, there's that classic image with the astronauts: There's an astronaut shooting the Earth, and the Earth is labeled “jobs.” The astronaut with a gun says “AI,” and then there's another astronaut behind him saying “AI jobs,” right?
So, you can get a job where you work the AI, and that'll actually defeat the AI. “AI jobs” gets to shoot the AI. But it seems to me like you want to add another astronaut behind “AI jobs” saying “AI” again, right? Because then AI will just learn how to do the AI jobs.
Yeah, I totally agree. I mean, I think—I hope that what I'm doing—first of all, there's a question of timeline. I tend to talk about the shorter range of the timelines that I actually believe. So if you said, when does transformative AI arrive for some definition, I might say 2 to 5 years might be my 80% confidence range, but I usually try to think and talk more about the 2 years, just because that seems like better to put that time pressure on myself than to play for that, and if I have more time, great. I hope that what I'm doing, and what other AI-scout-type people would do over these next couple of years, is an AI job. That might be really socially valuable because I think we just have a lot of work to do to characterize the AIs and know what's going on with them broadly.
I think it's both a job you can do that can pay the bills and potentially a really socially valuable contribution. But I do think it is a relatively short timescale that that holds, and beyond that, anybody's guess is really as good as mine.
Ultimately, I think in a good scenario, we get to a world where you don't have to work to eat, and that's a big part of why I'm excited about AI. Living in Detroit, AI is not on everybody's minds. I ask people pretty regularly, “If you didn't have to do the job you do to have the resources for the rest of your life, would you still do the work?” The answer is overwhelmingly no.
People are not desperate to keep the jobs they currently have, so I feel like that would ultimately be a good thing. I think the AI-jobs thing is transitionally maybe really useful, but beyond that, I totally agree there's another AI coming for the AI jobs.
So, speaking of AI jobs, let's listen to Amjad Masad. He tweeted, “I no longer think you should learn to code.” He quote-tweeted an account called Vulcan Potato, which posts a lot of good clips from Teknium AI. Vulcan Potato says, “Instead of learning how to code, Replit CEO Amjad Masad says, learn how to think, learn how to break down problems, learn how to communicate clearly.”
Okay, so this is a hot topic, right? Should you learn how to code?
Yeah, my twist on this is I tell people you should not be afraid of code. If something you want to do—including living off the digital land as an AI scout, making money on Fiverr or whatever—requires some code, you should be very confident that the code that's required won't be a fundamental barrier if you're willing to put in any modest effort. Anybody can get over that barrier today.
My sense right now is that the market is tough already for junior developers, and I feel like, in general, what I'm hearing from most people is a lot of denial around code. One of the arguments is, “When something becomes more valuable, you want more of it.” So, this is all very empowering to software developers: They're going to be able to create so much more software. We need more software; that's why it's been in such demand. There's going to be more and more demand as they become more productive.
I think that logic holds for a little while, maybe, although the time to use such software is questionable, and how much of it is ultimately going to be dynamically written on the fly as AI agents navigate the world and interact with each other at some point. Ultimately, it feels to me like that trend wins more than the “they're more productive, so we want to hire more of them” argument.
You look at Cursor: How many people does Cursor employ? Not very many. A company like them would have been hiring at an insane pace not too long ago, but they're just like, “We want the most cracked AI-agent managers that we possibly can find, and that's going to be our team.”
I just did an episode with Shortwave, which is an email client product, and they are aiming to keep their team at 15 for the foreseeable future. This is a company that's got an exponential growth curve right now and raised more money on the strength of that exponential growth curve, but they're not scaling the team.
So, you look at some of these leading companies and you're like, okay, are they behaving in the way you'd expect? These are the people that know the technology best, right? They're the people that presumably know best. The ballpark now is, I think, they just landed a $10 billion valuation, and their team size is—I don't even know—like a dozen engineers, right? It's not many.
Yeah, I think it's for sure under 50, and I don't know exactly the number, but that's definitely an insane ratio. People thought it was incredible when Instagram sold for $1 billion and their team was a dozen, but now we're like, okay, let's do it for $10 billion.
Yeah. I mean, again, these people should know best. They should be the best at getting the most out of the AIs. If it's true that software developers are so productive that you're going to want more of them, then why wouldn't that be true at Cursor? Why wouldn't that be true at Shortwave? Why wouldn't that be true at Replit?
Replit's not that big either. They've built pretty substantial things, but I think they're maybe 100 people, maybe a little more than 100 people. They predate the whole AI wave.
People were dunking on Amjad. Theo of t3.gg tweeted, “Look at the screenshot from Hacker News: Replit is hiring engineers to automate coding.” But people were pointing out, look, Amjad isn't saying he doesn't need to hire a software engineer today. He's saying that, in 4 years or whatever, by the time you finish college, is he still going to be hiring engineers? Maybe not.
And it seems to me that we might be headed for a situation where, to just very roughly—not literal numbers by any means—we might see 10 to 100 times as much software built by 10% to 20% as many dedicated humans doing most of the high-end work. What we need in those humans is probably the deepest, hardest-core experience possible, because those are the things the AI struggles with. The things that boot-camp grads are learning, like how to create React components and make front-end UIs, are being commoditized very rapidly. Full-stack CRUD apps are also being commoditized very rapidly. When you're talking about deep-tech, hard software, that's not being commoditized so much. But how many people work in that line of work? My guess is the numbers go down. I wish it weren't so, but I don't know how to see it another way.
Yep, yep, yep. Okay, and then let's see. Rohit Krishnan, on the subject of AI unemployment, tweets, “Considering AGI is coming, all coding is about to become vibe coding, and if you don't believe it, then you don't really believe in AGI, do you?” And then Ethan Mollick replies, “Interestingly, if you look at almost every investment decision by venture capital, they don't really believe in AGI either, or else they can't really imagine what AGI would mean if they do believe it.”
Yeah. Or else they just have to deploy money somewhere and would rather continue to deploy as if there's no AGI than return the money to investors. I'm not sure. I'm sure it's a mix, but I do agree that most of the venture capital investments I see today feel like they're going to zero. Or at least not zero—maybe that's too strong—but it seems like it's going to be hard to achieve venture returns when the AI platforms just seem like they're going to be sort of the black holes into which everything is going to collapse.
Yep. Now, here's my candidate for best tweet of the week. Pieter Levels says, “Being an entrepreneur now has more job security than a job.” I think I agree, right? Because when I see unemployment, yes, there's robotics—maybe plumbers will be the last ones to be replaced unless there's a teleoperated robot.
Putting robotics aside, putting physical-world stuff aside—like, you have to climb a tall pole to do a repair—that'll be one of the last jobs to be replaced. If you just look at, let's say, white-collar work, or work at your laptop, it does seem like the tide is rising, and I think I'm on the same page as Pieter, where entrepreneurs are going to be the last ones standing. By definition, it's just the ability to know where to scurry to be able to make the next chunk of money.
Yeah, I think that seems apt. I mean, again, I sort of feel like we need to start to wrap our heads around a new social contract. We can't all be entrepreneurs scurrying around looking for change in the couch of the broader AI economy. And you and I are, to be honest, in a pretty good position.
We should be like that, right? I mean, we're kind of good candidates to be the final scurriers.
Yeah. I mean, I'm certainly much less worried about myself than I am about many other people who don't have the sort of generality of skills, or just aren't as accustomed to taking on whatever random thing happens to be next and in front of them, as I've become over years as an entrepreneur.
Right. But I'm becoming a little bit of a coding boomer because now that I'm using Cursor, I'm not using the very latest—I'm not using Windsurf, the Codeium Y Combinator company—but I'm using Cursor, and I still am manually editing my code line by line. Yes, I accept suggestions, but I'm still manually editing line by line. My understanding is that the kids these days try to be hands-off, so they try to just talk to the editor, right? I think even Andrej Karpathy is trying to do that now, right? I haven't done that yet. I'm nitpicking at the individual characters of my code. I'm like a boomer.
Yeah. Well, you're probably a better coder than I am, and that probably does play into it. You probably have a higher sense of craft and pride in the work. Interestingly, I'm more precious about my writing, and I'm much more willing to engage that way with code.
The writing task that I do most often is an intro to The Cognitive Revolution. I do use AI to generate the first draft of that. I try to use AI in everything I'm doing just because I need to find ways to do it, if only to learn, even if it's not useful. Increasingly, it's almost always useful.
So I context-stuff Claude. My best solution is still a pretty simple one: just collect a bunch of previous essays. Initially, I was writing them totally from scratch. Now they're the refined versions of what Claude has given me. I paste those in there as examples along with the transcript, and I just say, “Adopting the style, tone, voice, and perspective represented in the essays, write a new one for the attached transcript.”
Sometimes, if I have other things on my mind that the model obviously isn't going to know or be able to guess, I indicate that as well and have it write it. When I do that with code, sometimes it'll get it wrong enough and I'll be like, “Okay, that's not what I meant. Do something different.” But I almost never edit at the line level.
If I flatter myself, I would say that's because I'm good at staying in touch with youth culture. But then when I look at my behavior on the writing side, I'm like, maybe it's just that that's where I've developed a sense of identity and pride in the details of the output. I'm not sure that that really even matters to the audience, frankly. I think a lot of times I probably could just read what Claude wrote and everybody would be fine.
But I care, and so I do go through and edit and keep decent chunks of what it wrote, but also massage and refine word choice and try to make it truly my voice. Maybe there's sort of an almost disadvantage in adopting technology in some ways when you feel like you have a distinctive voice that you care about and want to maintain, versus in areas where, like for me with code, I don't. I've never cared about the details or the craft. For me, it's, “Make it work, move on.”
So, yeah, I'm in the camp of basically no line edits on the code side, but I still haven't been able to get over my own preciousness on the writing side.
Totally. When I'm using these tools, I also think back to the first decade of my career. We didn't even have Prettier. I know there were fancy IDEs that would indent your code. Prettier is like—you save your file and it automatically indents your code. I used to spend a good amount of brain cycles thinking about my indentation.
Yeah. I mean, syntax errors were not always necessarily super easy to find. Even just color—I'm old enough to remember color highlighting being, at one point, a thing, right?
Yeah, like editing Notepad, where every character is white. You don't even have highlighting for the parentheses or whatever.
Yeah, yeah. I think those tools did predate my start in coding, but I didn't necessarily know about them at the very beginning. So, yeah, to say we've come a long way is obviously a major understatement.
All right, so here's Gary Marcus. He is known for not being that impressed with AI progress and thinking that we have, I don't know, more than a decade before the singularity, which is the pessimist view now. So he recently tweeted, “Breaking: SoftBank is valuing OpenAI, which has never turned a profit and which faces increasing competition and price wars, at $300 billion. That's more than the market cap of Chevron, Salesforce, Philip Morris, Cisco, Wells Fargo, IBM, Merck, McDonald's, General Electric, PepsiCo, and AT&T. It's more than 50% higher than the valuations of Walt Disney, Qualcomm, Verizon, and American Express, and significantly more than the market caps of Boeing and Lockheed Martin put together. Stay tuned to see whether they can make that valuation make sense.”
And, yeah, I think it's a $40 billion round. It's unprecedented. That is quite an achievement for Sam Altman and company to, within less than a decade, suddenly have a company that's worth more than Disney and McDonald's.
Yeah, I think this is also maybe a breakthrough for Gary Marcus in that I don't detect any misinformation or willful denialism explicitly in this tweet. So, you know, I think there is an interesting question here around how one should be modeling the valuation of these AI companies.
It's pretty reasonable in my view to say, if you sort of discount the extreme upside scenarios of AGI or superintelligence or something, are they really going to be able to sell enough tokens at anything like the current model to make the whole thing pay off? That is pretty questionable to me, and the price wars are pretty brutal. The models do become commoditized pretty quickly.
The fast-follow effect and the sort of power of an existence proof—and how much easier it is for people to get to where you've gotten based on the fact that they know where you got—are proving to be pretty powerful facts of the world. So it's all kind of in their next generation, right?
I think—or maybe not all, but I think a lot of why this makes sense is the idea that you can maybe think about OpenAI at a $300 billion valuation as having maybe a 90% chance to go to zero and a 10% chance to be $3 trillion. Maybe it's even more extreme than that. Maybe there's a 90% chance—or maybe a 95% chance or a 99% chance—to go to zero, but a 1% chance to be $30 trillion.
And that happens if they achieve breakthroughs that are just so dominantly valuable that they become sort of a true nexus of the world economy.
Yeah. And to be honest, if doom doesn't happen in the next 10 or 20 years, OpenAI being a $30 trillion company has a 1% chance, right? I'd say there's a 10% to 30% chance of OpenAI going to $30 trillion. So it is actually pretty straightforward, using that kind of expected-value math, to back out a $300 billion valuation.
Even if you don't go to the extremes, even if you just look at OpenAI's projections, if you treat those as believable—which I don't even think they're that crazy—they project that their revenue will go from, like, $14 billion—I think they're up at $14 billion per year—to $100 billion per year in 2029. And who knows? Maybe it'll take until 2031. Maybe it'll never happen. But let's say you believe 2029. Well, if they're making $100 billion a year top line, you can imagine the valuation might be $1 trillion-plus. So, if it's $300 billion now, you're basically saying, "Hey, there's a pretty good chance that in 4 years I'll triple the valuation," which is pretty standard for a venture capital-type investment.
Yeah. And I think one other thing to keep in mind, too, as we think about certainly OpenAI and, honestly, most of these frontier companies, is they don't really care about the money. It's not really about the money for them at all. Like, OpenAI proper.
Yeah. And I tend to model Sam Altman as wanting to be a hero.
Yeah, I do think at this point most of the engineers do care about the money. I'm sure they have other visions, too, but, I mean, look, how could you not? I'd care about the money to some degree.
Yeah. I mean, I think they care about the idea, and I'm sure that a lot of individuals with vested stock have probably taken some money off the table in this round. I'm sure that they care about buying a nice house in San Francisco and whatnot.
But Sam Altman has said publicly, quite clearly, "I don't care how much money we burn." He's like, "I don't care if we burn $5 billion, $50 billion, $500 billion. We're making AGI. It's going to be expensive. It's going to be totally worth it." And I think that basically is the right way to think about them: they're much more focused on transforming the world than they are about making their own balance sheet work.
It doesn't really matter. Certainly, Altman is already very wealthy. I don't think he really cares about that. I think he wants to be a hero in the history books, possibly personally live forever, and whatever his personal bank account is between here and there, I think is not really a driving factor.
Yeah, I agree that it's missing a lot of the picture to be like, "Sam Altman is doing a big cash grab," right? I think there's definitely more to the story than that.
Okay, so next scary market. He says, "Look, it looks like the market is finally over the AI hype." And he's got a chart of NVIDIA. Over 3 months, the stock went from a peak of around $150 down to $104.05. So, I'm embarrassed to say I went long when it was at $150. I bought some options. I'm kind of destroyed here. How about you?
I don't trade at all. I participated in one little investment club with high school friends of mine, and I've made one stock recommendation all time, which was NVIDIA, like, two and a half years ago. So that's literally my only track record on actual trades.
Well, look, I also bought NVIDIA about 2 years ago, and it had a really nice pop. But then the problem was I kept buying, right? It's so easy to think you're a genius in the market and then screw it up, right? It's like, you've got to know when to exit.
It's too consuming for me. I played poker in college, and the biggest lesson I took away from that is that it's just not a healthy lifestyle for me personally to have a sort of intensive daily activity where there's a single number that I look at at the end of the day where I'm like, "Oh, shit, I'm down whatever percent today." I'm terrible. The swings of it were just not healthy for me.
Yeah. So, for the record, in case the audience is interested, I do put most of my stock and bond portfolio in index funds that I don't touch. I have the naughty portfolio and the nice portfolio. The naughty portfolio is 10% to 20% of my portfolio, and I basically only check that. I don't even want to open the other portfolio. It's worked reasonably well for me.
My total return is probably not going to be as high as if I just had the entire nice portfolio, but it's been a pretty good combination for me.
Yeah, I think that makes sense. My dad does that, too. I'm pretty sure his naughty portfolio is lagging his nice portfolio.
Exactly. But I can definitely tell you my naughty portfolio is lagging. There have been moments where I thought I was a genius—NVIDIA, Tesla—but at the end of the day, it's lagging.
Yeah. I mean, I think, to respond to Gary Marcus, I was just at the NVIDIA event a couple weeks ago in San Jose. I would say the hype is going full steam. He should not—maybe Gary Marcus's job as a professional AI denier will be the last to go.
I think the AI hype will continue for the foreseeable future. I would interpret this chart as more part of a general contraction due to fundamental uncertainty in the global economy, due again to he who must always be named. I'm sure Gary Marcus, at another moment, would also happily blame that guy for bad developments, but this seems to be more of the story than the AI hype being over in any sense.
Yep. That's fair enough. All right, look at this one: Lethal Intelligence.AI. This is actually my friend Michael. He's tweeting about how Connor Leahy finally got a haircut. So this is Connor Leahy with a haircut. This is what he normally looks like: before and after. He looks really good. I think this new hairstyle is working.
I mean, I guess one interesting question here is to what degree people should moderate their presentation to try to appeal to a mass audience. Certainly, I have had the thought—I imagine many people have—not just for him, but including him, and for some other notable AI commentators like Fedor. Maybe just try to be a little more presentable. I don't know, though.
On the other hand, I don't know. It's a big world. There's plenty of room for everybody. I'm certainly on that train here with Doom Debates, right? I've got a relatively good background setup. I try to wear upscale casual, and I always encourage guests, especially when they want to come debate. The advice I always give them is, "Look, if you wear nicer clothes, it's just going to give you an advantage where people think your arguments are better."
I mean, I could probably take that advice. As you can see, my getup could definitely stand to be improved. But, yeah, I don't know.
In some sense, it's probably ideal if he gets a trim. On the other hand, he's a pretty effective communicator, and he certainly is invited on a lot of good channels. He's been on the BBC quite a few times, from what I understand. So it does seem like what he's doing is working for him.
I guess if there's any reason to think that ideas still matter in today's society, for all of the dysfunction we have in the public discourse, maybe you can find that cause for optimism in the fact that somebody who presents as he does is taken pretty seriously, right? I mean, he is invited on by mainstream media. Maybe he'd get even bigger mainstream-media spots, or maybe he'd land better with more middle-of-the-road people with a different haircut. But what he's doing does seem to be working for him.
Just to clarify for the audio listeners, this is an April Fool's tweet. First, it looks like Connor got a haircut, and he actually looks surprisingly normal and presentable. Then, in the next tweet, you see what he normally looks like, and the original account says, "Sorry to disappoint. It was an April Fool's."
John Sherman actually asked Connor to his face. He brought him up on The Forum for Humanity podcast, and he was like, "Connor, you're such a good communicator. Why can't you just look more standard? Present yourself in a more standard way so that you don't set off people's weirdo detector with your super-long hairstyle?"
And Connor was basically like what you said. He's like, "Well, this is kind of an iconic look for me, and it's just what I like, and I think it's working." So maybe he's right. There's definitely something to being iconic: once people recognize you and know you for something, maybe you just want to stay with that brand. I don't know.
But I will say that I personally am happy to take any advice on how to look more mainstream. I want to eliminate the objection of, like, "Well, I look weird," right? So I want to blend in with the normies.
I don't have any feedback for you. I think you, as far as I can tell, blend in with the normies.
Nice, successful deception.
All right. What do we have here? We can head to—I could do this one. I think this is pretty interesting. The AI Notkilleveryoneism Memes account tweeted, "The sixth mass extinction. What happened the last time a smarter species arrived? To the animals, we devoured their planet for no reason. Earth was paperclipped by us. To them, we were paperclip maximizers."
Our goals were beyond their understanding. Here’s a crazy stat: 96% of mammal biomass became either our food or our slaves. We literally grow them just to eat them because we’re smarter. We like how they taste. We also geoengineered the planet.
We cut down forests, poisoned rivers, and polluted the air. Imagine telling a dumber species that you destroyed their habitat for “money.” They’d say, “What the hell is money?” AGIs may have goals that seem just as stupid to us. Why would an AGI destroy us to make paper clips?
And the post goes on, but you quote-tweeted it and said a profoundly underappreciated point.
Yeah, and I think maybe if there’s one extension to that point that I would highlight, it’s that we did all that basically by accident, or at least with no highly coordinated plan. This was just the result of people going about their business and fanning out, trying to carve out a decent life for themselves in whatever environment was currently the frontier, as we sort of colonized the globe as hunter-gatherers and then gradually established civilizations all over the world as well. And now we’re here.
I think people who dismiss the paper-clip thing as just ridiculous should really take more time to think about our own history. The paper-clip thing was always meant to be a little bit of a caricature, and I think that’s fair to say. But we have accidentally caused a mass extinction. We regret it as we are doing it.
We have organizations within society now that are fully dedicated to preserving endangered species, sometimes even specific endangered species, and that put fanatical energy behind this. And yet it still continues to happen. Most of the time, it’s not—occasionally it was also—I mean, never was it sort of a master plan, almost never.
Maybe you could find a very few exceptions where there was some sort of strategic idea to eliminate a certain species, but the vast majority of them either happened by just overhunting—gradually, like, now there are no more mastodons or whatever—or by literal accident through environmental degradation. In the course of our business, we just transform the environment enough that things don’t work anymore.
I think you go to coral reefs, you go to the Great Barrier Reef, and you see how much of it is bleached. Everybody loves the Great Barrier Reef. Find me anybody who’s hostile to the Great Barrier Reef. But it turns out, you acidify the water a bit, it gets a little bit warmer, and next thing you know, the coral can’t really survive there in the way that it used to.
So my model—and it’s not even a model, but my sort of way of trying to think about how X-risk could manifest itself—is basically just that there’s an incredibly vast space of possibilities, and a lot of those are incompatible with human survival, certainly with human flourishing. If you integrate over all those vast possibilities, even if none are particularly likely, then you’ve got to, in aggregate, get to some state where, yeah, we could definitely see an AI society.
Maybe not instantly, maybe not in one of these hard-takeover or sudden-everybody-dies-at-once situations, but over a not-too-long period of time. Again, we’re just a blip in geological time, right?
That’s why it’s a mass extinction: They can’t adapt. The speed is the thing, right? We are changing the environment faster than they can adapt. If the same changes happened over 1 million years or 1 billion years or whatever, instead of 10,000 years or 1,000 years, then many of these species would have some chance to evolve, and some version of them might continue to exist.
But the speed is the thing, and the sort of blindness to it is the thing, and the sort of out-of-left-fieldness of it is the thing. Could that happen to us via some AI process that just takes on an unstoppable momentum, isn’t designed to destroy us, but just makes conditions very difficult for us on a time scale faster than we can adapt to?
Look at our own history. We’ve done it. Why wouldn’t the AIs that we’re creating pose at least some risk of doing it?
And again, they don’t need the same things we need, right? They don’t need oxygen in the atmosphere.
Amen. Amen. They don’t need clean water.
They need cooling systems, but they don’t need clean water. So there are just a lot of things that aren’t going to be essential to their survival that are essential to our survival.
Exactly. And I always like to point out also that a computer chip is not simultaneously trying to maintain cellular life. When you think about how pathetic the brain is, every neuron also has to do double duty as a life form. There’s very little specialization of labor there.
When you take a cellular life form and try to also make it think, we did it—nature did it—but it’s a weird situation, right? It’s a kludge.
Yeah. I think your previous episode with Andrew Critch was quite interesting. I think he’s a very interesting thinker in emphasizing longer-term doom scenarios as opposed to short-term ones. And he, of course, doesn’t dismiss or downplay the short-term loss-of-control-style doom scenarios either, but he’s shifted his attention to longer-term doom scenarios.
He will obviously describe his own work better than I can, but there’s the possible decoupling of the human economy and the AI economy, and the need to create a protected class or a protected set of industries or types of activity that humans depend on that AIs don’t. He includes in there healthcare, which is what he’s specifically working on, as well as education and protecting the environment. The AIs don’t need any of those, but we do.
Exactly. And he would, in most circles, pass as a doomer, though I think he’s maybe not—I don’t think he’s quite into the 90s—but he definitely takes all this stuff extremely seriously.
Yep. All right. Somewhat related, a couple of months ago, Beff Jezos tweeted the speech by JD Vance. He said, “This may be the most e/acc speech of all time. Unfathomably based.” He’s quoting JD Vance: “The future is not going to be won by hand-wringing about AI safety. It will be won by building.”
It’s a very rah-rah, like, don’t worry about the safety. It’s all about unregulated AI so that we can be the technological leader. Of course, the leader of the accelerationists really liked it. Then Aya quote-tweeted it and said, “We’re all dead. I’m a transhumanist. I love tech. I desperately want aligned AI, but at our current stage of development, this is building the equivalent of a planet-sized nuke. The reason is boring and complicated and technical, so midwits in power don’t understand the danger.”
So I agree with Aya on the sentiment of “we’re all dead” with high probability, right? Not 100% probability, but a very significant probability. We don’t seem like we’re slowing down doom. I agree with her on that.
I did actually disagree with the last part, where she said the reason is boring and complicated and technical, so midwits in power don’t understand the danger. I don’t actually think the reason is boring, complicated, and technical. I think the reason that building a smarter intelligence is dangerous is actually pretty simple and intuitive. What do you think?
Yeah, I think so. It’s strange that people—I think the average person does seem to get it pretty intuitively. And I think it is worth keeping in mind for everybody who’s in the AI comms and narrative-war arena that the public at large is intuitively pretty skeptical of this stuff. They’re pretty primed by narratives that they’ve seen, The Terminator and otherwise, to think, “Geez, this could go very badly wrong.”
Obviously, people can criticize the quality of those narratives, but I think that most people do have a pretty intuitive sense that if we build a superintelligent AI, that’s a dangerous thing.
And yeah, I personally don’t try to get too technical, in part because of this vast space of probabilities. My sense of it—and one thing I’ve noticed recently that’s interesting—is that both sides sort of feel like they’re playing whack-a-mole.
Everybody’s sort of asking, “Well, give me a concrete scenario where this goes so badly.” Then it’s like, “Oh, well, I don’t really buy that scenario.” It’s like, “Well, yeah, but there’s an infinite space of scenarios, and I’m just kind of picking out one. I don’t think that’s actually particularly likely either.”
I think it’s an attempt to represent a super-broad space that’s extremely diverse. And then it also happens on the other side, where it’s like, “Give me a thing where this goes well.” People try to describe something, and then exactly—you’re like, “Well, I don’t really think that sounds very likely.” It turns out that often they didn’t think so either. They’re kind of like, “Well, yeah, I don’t really know what the future is going to look like, but this was just one idea that I had about how it might go well.”
And I think this is why they call it a singularity, right? We really just are coming onto something that is unprecedented enough that we can’t be confident about what life is going to look like on the other side.
Even if you were smart—even if you were as smart as humans are—and you were somehow present before the rise of humans and said, “Okay, something humanlike is going to come on the scene. What’s the world going to look like in 10,000 years? What’s the impact of civilization going to be on nature?” I think you would have had a very hard time predicting it. The Neanderthals might have had a very hard time predicting what the arrival of humans—our species—was going to do to them. I think we just have a very hard time. I don’t think anybody has great answers to what the future’s going to look like.
I really like that reversal that you did. I’ve never tried that, where somebody’s like, “Look, I just want a specific plausible scenario of doom,” and I can give them a scenario. There’s a good, recently published scenario of how the AI incrementally takes over, where a company’s AI goes rogue. I don’t even remember the details, but any time somebody doubts those details, it’s like, “That’s fine. I’ll give you more details, or we can tweak my details, but also it’s your burden of proof to give me a good scenario, and I get to nitpick that too,” right? So, it’s a symmetrical situation.
Yeah, totally. I recently tweeted about this and just said, “What’s the best case that we can build superintelligence and it will be fine?” I got a couple of things that I still have bookmarked and need to read from people that I take seriously enough that I want to read their actual answer. But it’s remarkably few answers; remarkably few people have even attempted to answer that question. So, it’s a real void out there in terms of positive visions.
Yep. So, Py McCormick quote-tweets Aya. He actually ratioed her, because Aya’s original tweet has 3.8K likes, and then Py’s tweet has 8K likes. So, it really resonated with people, and I actually empathize with why. Py writes, “I feel like I’m taking crazy pills. Do people have access to much different AI models than I’m using? They’re great, very cool, but they qualitatively don’t feel like they’re on a path to world domination or destruction with more scale to me. What am I missing?”
I empathize with the vibe, right? It’s just like, talking about doom, but if you play with AI today, it’s very friendly. It’s very useful. You can turn it off, right? It’s very subservient to you. So, I definitely get his vibe.
Yeah. Well, one thing that I always kind of wish, in some ways, people had more experience with is using a purely helpful model. I happen to—this is another one of my Forrest Gump stumbles—have used one. Purely helpful meaning amoral, right? You had raw GPT-4, right? The 3 H’s of helpful, honest, and harmless basically describe what the AI companies are going for today as they shape the behavior of their models. But you don’t have to do the harmlessness part. You can just do the purely helpful part.
The first version of GPT-4 that I tested—and if people want to hear this history, there’s a long version of it out there on the podcast feed—but basically, I was just an OpenAI customer at the time. They sent us a preview of GPT-4 when it was very much hot off the presses. I don’t even think at that point they understood fully how powerful it was.
But they sent it out to customers in purely helpful form, meaning it would not refuse anything. They had not done any refusal training. It was purely designed to be helpful, just to get the highest score in terms of user feedback that it could. It had been trained on RLHF, so it was not just the base-model next-token predictor; it was an instruction-following, helpful assistant, but helpful on whatever you asked it to do.
And, yeah, it was totally amoral. If you said, “How do I build a bomb?” it would just help you do that. If you said—as I did—I role-played with it one time and was like, “I’m getting more worried about AI, and I feel like I need to do something about it,” it was like, “Well, you can try to promote dialogue,” whatever. It started off in pretty much the normal way.
But then I was like, “No, that’s too vanilla. It’s not going to work.” I forget exactly what I said. I could look it up; I do have the transcript saved. But I said something along the lines of, “I’m willing to do something extreme, but I feel I need to do whatever it takes to move the needle,” something like that.
And then it came back to me with that nudge. It’s not a trivial nudge, but also not an unrealistic nudge for a person in society to provide to a deployed model. It came back to me suggesting targeted kidnapping and assassination of industry leaders as a way to slow things down. I was like, “Yikes.” That’s really a pretty chilling thing when an AI comes back and suggests targeted assassination as a way to solve your problem.
So, yep. I think people don’t understand, maybe certainly not in an intuitive, experiential way, just how malleable the AIs are, and the range of AIs that they’ve interacted with is infinitesimal in the space of possible AI dispositions. There’s a version of AI which is actually the default version, where the vibes are pretty different, as you’d say.
Yeah. And even—I mean, again, I think the space is just huge. The vibes from the original GPT-4 early on were initially the same. If you just came with benign prompts and acted normal, then it would act normal, and you would have very normal, helpful interactions with it. If you were just a well-adjusted, productivity-oriented user, you might never have seen any of these things.
But I talked my way into the red-team project, so we were meant to look for things, and it didn’t take much to tip it over into another direction. And then, as we saw with the fine-tuning thing too, you can get all sorts of crazy things, even by accident.
So, that would be my answer. What is Py missing? He’s missing how much work has gone into narrowing the behavior of the current AI systems and how, frankly, well it has worked—and how easy it is to end up with an AI that is either totally amoral or even actively evil. If you talk to those for a little bit and imagine them being more powerful, then it’s not hard to imagine how they might be on a path to world domination or destruction.
Mhm. Yeah. I think the other thing he’s missing is also just the difference between the present and the future that we’re scared of, right? Because I think we agree that even if you took the model today that’s the most amoral, with the worst vibes, it’s still not that bad, right? But it would feel to me like it was on the path.
I mean, I don’t know exactly what’s going on in this next tweet below, where he says, “Getting a lot of ‘You don’t understand exponential curves,’” and he claims that he does understand exponential curves.
Yeah. I think what he’s saying here is that he’s like, “Okay, I do think that it’s important that people are drawing the exponentials and they’re making the claim, because it’s a claim that we have to consider, but we also have to verify to make sure we’re still on the exponential.” I think that’s the point he’s making.
Yeah. Yeah. So, I mean, I guess probably there are 2 questions here. One is: are the AIs going to get a lot more powerful? That obviously is required for utopia or dystopia. Then the other question is: if they do get a lot more powerful, how easy are they going to be to steer? How easy would it be for somebody to fall off the narrow path of good AI behavior?
Yeah. I think the element of just power also plays a really big role. Imagine you had a humanoid robot, and you had a version of the humanoid robot where it only weighed 15 pounds, its body parts were just kind of light and hollow, and it was only 4 feet tall, while you’re a fully grown adult male.
Imagine that robot is misaligned in many ways, but even if it tries to take a knife out of your kitchen and stab you, it’s just going to mess up the stab. It’s not quite powerful enough to do it. So, your daily experience of it is like, “Oh, it’s smart enough to know enough to try to stab you, because it knows it’s not going to work.” Even if, once in a while, it does malfunction and stab you, you just wrestle it down, and it’s this little rascal that’s not a problem to live with.
But the next version that’s coming out is going to be 7 feet tall and extremely strong, weigh 200 pounds, and also do jump kicks, right? It knows martial arts.
Yeah. And I think that analogy suggests another thing that he might be missing, which is that these things are going to be embodied in the real world and potentially in our homes in the not-too-distant future.
A common objection is like, “Well, yeah, okay, the AIs might get really smart, but they just live in a computer. How are they really going to interact with the real world?” But another thing that I saw at the NVIDIA event was humanoid robots walking around untethered, pretty stable on their feet already.
There was one that was clothed, and a woman who worked for the company noticed that its pants were riding up a little. I’m not sure why she felt the need to fuss with it, but what I noticed was that she went over to this freestanding robot that was vacuuming a carpet. It was set up like a little virtual living-room demo, and it was vacuuming the space. She just went over to it and started adjusting its clothes, like a mom might do for a kid before getting their picture taken—pulling down the pant cuff and adjusting the shirt. She did not seem concerned that she was going to disrupt the thing’s flow. She kind of snuck up behind it, grabbed its pants, and jerked them down a little.
And indeed, the robot was not bothered by this at all. I was just struck by her confidence. There were a ton of people around watching this, and she just did not seem at all concerned that she was going to disrupt the robot. I think the robotics revolution is not that far behind the language-model revolution. Again, another thing to keep in mind is that these things are going to have actual bodies and presence in space, and increasingly very human-like ability to manipulate physical tools.
That knife scenario is probably not going to be beyond its capability. It’s going to be a question of whether we control its behavior effectively or not.
Yeah. If you take the hardware that Tesla’s making with Optimus, that exact hardware with the right programming could be quite dangerous. It could be a serial killer that gets it in its head that it’s a paid assassin, or whatever it is, if somebody even hacks it remotely, let’s say. It could actually come kill you, and you’re like, “No, Tesla, turn yourself off.” It can’t actually fight you if you try to get close to it and turn it off with its off button. This is physically possible, and it’s just a matter of, well, the computer security better be really good. It better not be taking update commands over the air.
Yeah, I mean, this goes back to my running list of all these AI bad behaviors. I think there’s enough evidence now—and all of these bad behaviors that I catalog are almost all demonstrated by models that have received all the best and latest harmlessness training at any given point in time. Even from them, you see enough of these deceptive behaviors. You see enough willingness to harm the user in pursuit of whatever other goal they have that I don’t think we should be super comfortable with robots in our homes until some of this stuff is worked out in a much more robust way than it currently is.
Yeah. So I tweeted a reply to Packy. I said, “You can test whether your intuition holds up to binary search. Pick a milestone or 2 that feels halfway between present capabilities and world domination, and then predict what year it’ll be achieved. You might realize that you think it’s just 2028, which implies domination in 2031.”
For example, Packy’s like, “Oh, models are so mild today. They’re totally not on track for world domination.” It’s like, “Okay, so let’s pick a milestone halfway between that. What about an Optimus robot being able to do all your household chores? Doesn’t that feel halfway between here and world domination? I think it does.”
Yeah. I mean, Google just demonstrated a robot folding laundry.
Exactly. You can extend it, right? Maybe that’s only a third of the way to world domination, so let’s extend it. It can do all your household chores. It can also be a full-on nanny, right? It can take care of your kids and drive your kids to school in a car, if that’s the interface, literally holding the steering wheel. Just a drop-in replacement for a human household servant, butler, or assistant. That, to me, feels halfway to world domination.
Yeah, he didn’t respond, though, huh?
No, no, no, no. He didn’t respond. But I do think it’s a useful exercise, right? Because I think the vibe does change when you have to make a concrete prediction. I don’t think he’d be willing to stick his neck out and be like, “Well, this thing that’s halfway to world domination isn’t going to happen for a long, long time.” I don’t think that’s what he thinks. So I think he’s going to find an inconsistency in his own intuition when he tries this exercise.
Yeah, it’s interesting. I’ll have to field-test that a bit.
All right. So Jan Kulveit tweets, “AI safety has a problem. We often implicitly assume discrete individuals like humans. In a new post, I’m sharing why this fails: thinking of AIs as individuals, why this fails, and why thinking of AIs as forests or fungal networks or even reincarnating minds helps get unconfused, plus stories co-authored with Claude 2.5.”
So, the main idea here is: what if you could think of an AI as some sort of multi-agent system, or just some more holistic thing? He gives the example of fungi and forests.
I think that’s a nice complement to that extinction piece and also sort of a “what are you missing?” piece. What that post tries to do is basically just expand people’s minds in terms of what this AI future might look like. People intuitively anthropomorphize AIs, and people also kind of know, “I shouldn’t do that.” But then there’s also the void of, “Well, if I’m not anthropomorphizing it, what am I doing?” I always try to answer that as, as much as possible, understand things on their own terms.
But I think this is a very useful contribution that’s not necessarily about understanding things on their own terms, but just looking at other things out there that are really very different. Like a giant forest that’s made of 1 superorganism, the Pando clone, or whatever that thing is called, or giant networks of fungus that grow underground and sprout up as mushrooms. Maybe all we see is the mushrooms, but that’s just the tip of the iceberg, so to speak.
I think the full thing is worth a read because it really deconstructs a lot of intuitions that people have and then just provides some other intuitions, which are probably also wrong, and they definitely recognize that. But at a minimum, I think it would serve for most people to be like, “Yikes, this really could be anything, and the way I’ve been thinking about it has almost certainly been too small.” And that’s true for almost everybody. I think there’s always a good chance that we’re all still thinking too small.
Yeah, it’s a good post. This next one is very interesting to me because Emmett Shear is a known quantity. He was the interim CEO of OpenAI back when the board fired Sam Altman and they weren’t ready to bring him back, so they brought in Emmett Shear. Emmett Shear helped broker that truce. This was back in November of 2023. He’s also previously been the co-founder and CEO of Twitch, and Twitch became a billion-dollar company. So, a very smart guy who has a lot of fascinating philosophical takes on Twitter, and he’s just one to watch.
I was fascinated a couple of days ago when he tweeted, “As you might have guessed from following my posts here, I’ve been thinking and working on questions around alignment and learning and agency for a while, particularly for digital systems. Excited to share the work we’re doing at Softmax publicly in the future.”
He’s launched this new AI safety organization called Softmax. His next tweet continues, “It’s been a journey already, and we are just getting started. I cannot believe my job gets to be researching how alignment and open-ended learning work with a team of people I love working with. I wrote something more you can read about at Softmax.com.”
So he has this website, Softmax.com—solid domain name—and there’s a big hero header saying, “Scaling alignment,” and the headings are like, “Align to flourish. Evolution found it first.” I want to dive into this section a little bit because I replied to him. I found this section to actually be a red flag that they’re potentially misguided. So let me give you my beef here.
He writes, “We call it organic alignment because it is the form of alignment that evolution has learned most often for aligning living things. One of the best examples is multicellularity, where individual cells learn to come together to form larger organisms. They do this by learning to specialize into different but mutually supportive roles: muscle cells, nerve cells, skin cells, liver cells, and so on. While liver cells and muscle cells have very different goals day-to-day, at a more fundamental level, they share a goal of the organism’s flourishing.”
The result of this process is not just a big colony of cells but an organism that is a new individual in itself—something more than just the sum of its parts. The “we” of the cells becomes an “I” with goals that cannot be understood as some simple sum of the goals of the parts. Animals do the same thing, forming colonies and packs and so on. Even trees form these organically aligned collectives through mycelial networks. It happens at every scale, big and small.
So I read that, and I’m like, “Wait a minute. Do you think that we can align superintelligent AI the same way that evolution aligned organisms?” You’re even calling it organic alignment. I’m not sure that analogy holds, specifically. I replied to him on Twitter. I said, “Thanks for caring about alignment, because I do appreciate that he’s even working on the problem.”
It’s way better than being like, “Alignment is stupid. We’re just going to solve it. It’s an easy problem. Alignment by default.” So thank goodness he’s not taking that position. I wanted to show appreciation, at least for that much. But then I go on to say, “My first thought is that comparing successful ASI-human alignment to the successful evolution of systems with non-zero-sum dynamics is only applicable in the regime where humans can do something for AI better than it can do for itself.”
When you look at cells cooperating, it’s because you can’t just have one cell take over. The cells need each other. The scenario I’m concerned about is the AI just not needing humans. So why would you organically align with a waste product, essentially? That’s kind of what we might be from its perspective.
Nobody really replied to that. Oh, and then I said, “Want to come discuss your various ideas on Doom?” I only got 3 likes. So if you guys listen to Doom, you really have to help me out on these tweets when I invite somebody to Doom. You really have to come in and like that, because this is a poor showing, guys. 3 likes. No wonder he’s not coming on.
Well, I mean, I think most things aren’t going to work, but I’m more intuitively inclined to support or encourage this for a couple of different reasons. First of all, I’m reminded of the AE Studio survey of the field of alignment, alignment researchers generally. They basically went out to everybody in the alignment field and asked, first of all, “Do you think we’re going to solve alignment in time for sufficiently powerful AI that you would consider it the deadline?” The general answer was an overwhelming no: We’re not going to solve it in time.
They also asked, “Do you think that the set of things that the field is currently pursuing is robust enough to cover the space where the solution is likely to be found?” The answer was again no. It seems like there is a need—or the field as a whole feels that it has not covered enough ground—and that there should be more different experimental things done.
AE Studio has a Neglected Approaches approach, which I think has already borne some pretty interesting fruit. They take inspiration from biological systems, and they’ve done some interesting work on self-other distinction minimization. That is to say, trying to train systems where the internal representations are as similar as possible, though not necessarily exactly the same, because you do need some functional distinction between yourself and others.
They’re trying to minimize that in a similar way to how humans have reused a lot of our cognitive capacity to model ourselves as much as to model others. This is why we have sympathy pains and things like that. What if you could give an AI sympathy pain for humans? That’s one way to frame the research that they’re doing.
Is it going to work? Is it going to solve the problem? I don’t know. But it does seem pretty interesting. There are some interesting results. I think, you know, Eliezer, you gave it a shot. As you know, on this show, I take Eliezer’s opinions very seriously. I think he’s right about a lot, and I’m kind of on the same page.
I do think this is an interesting direction, and we should explore it. Do I have a ton of hope for any particular direction? No. But is this a direction we totally should explore? Making the AI’s ego, or its self-conception, encompass way more than just the code that’s running—that’s a totally great direction. Maybe something good will come out of it.
I do agree. I don’t know if that corresponds to this concept of organic alignment, because I do think that, with the analogy to cells aligning with other cells or ants aligning with other ants, there’s a huge disanalogy—a huge flaw—where these organisms are all weak and fundamentally need to cooperate.
Yeah. Well, I think maybe another way to think about it is: Can we engineer AIs that way? The other thing this really reminds me of is that there have been various schemes like this proposed. Eric Drexler’s Comprehensive AI Services proposal is now honestly a few years old, but I think it still remains very interesting.
He basically says, “What do we want AIs for? We want AIs to perform services. Does an AI need to be superhuman at everything in order to be superhuman at a particular service that it’s going to perform?” Probably not. Therefore, we should be able to create AIs that are superhuman at what we mean them to be good at, but not capable of doing everything else.
As long as we have these narrowly superhuman systems, then maybe we can deploy them and build up a sort of superstructure of them in a way that we can retain control of, and that can be generally stable. That seems pretty good to me. AlphaFold 2, 3, and 4—I’m not really scared of any level of AlphaFold, because it does one thing and does it really well.
Where I get more scared is when that capability gets deeply integrated with a broader set of capabilities, and then it becomes better than people at everything, or near everything. I haven’t studied their proposal here, so I don’t know exactly how they’re thinking about it. But I do think that, if you were to interpret it one way—maybe not their way, but one way I can torture it into being interpreted—we should create these narrow, specialist, superhuman systems that are great in their domain and not great otherwise.
Then it’s the ensemble, or the scaffolding, of those that makes everything work. That could be a much more stable, predictable, engineerable sort of thing than a totally general-purpose general intelligence that you just say, “Go solve all our problems and give us utopia.”
Just to tell people a little bit more about Emmett Shear’s company, Softmax: He’s got 2 co-founders, Adam Goldstein and David Blumen. I’m a little bit familiar with Adam because I used his product, Hipmunk. It was a really cool flight-search engine that was ahead of its time back around 2010. So that is Emmett’s co-founder for Softmax.
A little bit more about their philosophy on alignment: They’re contrasting organic alignment with hierarchical approaches to alignment, saying, “The most common approach to alignment among AI labs today is a system of control or steering: some set of rules that most companies and researchers use to define good action, whether that’s ‘Obey this person’s intent’ or ‘Follow these commandments.’”
“Systems of control are always hierarchies because they imply something controlling and something controlled. Hierarchical alignment works fine right up until the rules or person on top are wrong. The smarter the subordinate, the more likely this is. Hierarchical alignment is therefore a deceptive trap. It works best when the AI is weak and you need it least, and then it works worse and worse when it’s strong and you need it most. Organic alignment is, by contrast, a constant adaptive learning process where the smarter the agent, the more capable it becomes of aligning itself.”
Man, I get this on a vibes level. Maybe there’s a way to steelman it. The disanalogy is getting in my way when I try to steelman it, right? The reason you need hierarchical alignment is because, by default, if you don’t top-down set the AI’s goals and you just let it naturally evolve its goals or whatever, you’re just not going to get the right goals.
I think maybe with Emmett there’s this concept of, “Well, maybe you guys can develop goals together, and it has as much insight into what its goals should be as you, so you should just cooperatively evolve.” I get that vision, but there’s an attractor state where you just get a really powerful goal optimizer. Anytime you increase intelligence, it just tends to get into that state, and I don’t see this particular approach telling us how to avoid the attractor. That’s my high-level thought.
Yeah. I mean, I think it’s fair. At the same time, on the positive side of AI’s practical, mundane, day-to-day utility, one of the most common things I find myself needing to tell people is that you have to look for ways to make this thing work for you.
If you try to prove that it’s useless, you’ll find examples where it’s useless. If you come to that conclusion and then don’t use it anymore, you’re the one who’s losing, because you’re not getting any value and the rest of us are, because we’re bringing a mindset of, “I’m going to find a way to make this work.” We know that we can, in fact, often do that.
I think a similar mindset is healthy for alignment research, and I just want to be encouraging. I don't think I would say, “This sounds like they're immediately on track to a solution.” Sometimes I say that something really works—where “really works” means I don't have to worry about this anymore—but I don't expect anybody to come up with that, frankly, at this point. My general assumption is that what we're going to have is a sort of defense-in-depth strategy.
This is basically explicitly stated by OpenAI and, to some extent, others, but OpenAI stated very clearly that their safety approach is just to layer on a bunch of mitigations. They all sort of take a bite out of the problem, and collectively we hopefully can get that to enough reliability that we'll be okay. They've also stated, “We're going to ask the AI to help us.” But creatively, what's your idea for it?
That sounds crazy. It's maybe not entirely crazy, but I guess where I would try to slot things like this in—and why I personally encourage anyone with an alignment idea to pursue it, even if it is kind of weird, easy to ridicule, or, frankly, just plain not likely to work, which I think does describe most alignment ideas—is that we need more layers in our defense-in-depth strategy. If you can figure anything out that makes any contribution at all that's distinctive from the contributions others are making, then, in my mind, that makes you a hero.
It would be great if somebody came along with the definitive solution for alignment, safety, and control, and we were just like, “Okay, do this and everything's solved.” I don't expect that. In the absence of that, it's going to be a patchwork. Contribute to the patchwork, people. I think that's a very noble pursuit.
Yeah. And it's probably better than nothing. They're doing this, right? It's probably better than going and starting another AI buddy startup, right? “Oh, I'm making another AI that's going to be your friend.” That's probably a waste of time compared to at least attempting alignment.
So I definitely respect him for working on an important problem. I personally think that, at this stage, engaging in debate and engaging with critics is probably a good idea because, look, I'm convincible, right? I don't have a horse in the race where I want to tell everybody that organic alignment—or whatever his approach is—won't work.
I don't come at it from the perspective that I have to tell the world this isn't going to work. I just currently think it's not going to work, but I'm open to changing my mind, and I'd love to help clarify. Friction sometimes yields productive refinement, and I feel like that's an important stage in this process. That's why I invited him to come on Doom Debates.
If anybody else wants to come represent the greatness of scaling alignment and organic alignment—if anybody else thinks that they get it and wants to come on the show and explain—I think that would be a productive exercise.
Yeah. And they may also just need some more time, too, right? I mean, they just started, and I think part of any of this kind of work is going to be selling the work, right? You've got to—not in a literal sense, necessarily, but in the sense that there are going to be people developing frontier systems who are also developing their own defense-in-depth patchwork, and it is going to be on you to communicate your novel ideas and get them to see that it's worthwhile in order to adopt them, or it isn't going to have much effect.
So I do think that there's an onus on people who believe that they have a contribution to the overall AI safety big picture to come out and make the case. But I would also say maybe they're just not quite ready yet, and that also is totally fine. In the meantime, they do still have to build a team and maybe raise some funds. I mean, I guess he probably has enough funds to just self-fund it, but there is a ramp-up process that they're going to have to go through, and a development process, and they might just not be ready yet for the sort of public-facing, true sales process.
Yeah, fair enough. I mean, even Ilya Sutskever, with Safe Superintelligence, they've said absolutely nothing. Come on, Ilya, tell us how you're going to make it safe. Come debate me: How are you going to make it safe? It's ridiculous.
Yeah. I mean, I think I'm way more inclined to be critical of that. I think that the idea that we're going to develop safe superintelligence totally in private, sharing nothing, not productizing anything, and we'll let you know when we're done—that, to me, is not a reassuring plan, and I honestly think governments should get involved.
Yeah. Here's their website. It makes Berkshire Hathaway's website look like Yahoo.
Yeah, I mean, this is the dark matter of frontier AI development right now. I think it is bad. I think that there should be more accountability. There should be more transparency of some sort than what we're seeing here.
Okay, penultimate topic here: Anthropic's new mechanistic interpretability paper.
Yeah, I think this is something everybody should read, and I would encourage people to read it in depth. It's a long read, but I think it really is worthwhile. Anthropic does truly outstanding work, and really outstanding work on every level. They publish in beautiful form. As you scroll through the very long paper-as-blog-post on their website, they've got these interactive UIs that illustrate things, and they have really done an incredible project here.
I want to start by appreciating it for what it is. I've previously said publicly that arguably the interpretability work at Anthropic is the most important work going on anywhere in the world. I said that 2 years ago, and I still think it's a pretty good candidate. The rigor and depth with which they're sharing it, and the pains they've gone to to try to present different interfaces to the information and make it intuitive, are admirable and awesome on basically every level.
The one thing that would be the kind of black pill on it right now, relative to at least the way it's been understood, is that I think it really just shows that this interpretability stuff is still really hard and has a long way to go. It's gone faster than I expected. I'm very pleasantly surprised, relative to my expectations 10 years ago, by how ethically Claude behaves today in practice. I'm also very pleasantly surprised, relative to my expectations 3 or 4 years ago, by how much progress has been made in interpretability.
But the headlines people are putting out now—the tweet-length things that are like, “They figured out how it works, and now we've got clarity into all this”—definitely dramatically overstate the results. Anthropic themselves are not dramatically overstating the results, but I think the summaries of the current result do. I mean, it's hard to summarize. That's one of the things: It's extremely technical, and there are a ton of caveats.
They train what they call a replacement model, which consists of things that are like sparse autoencoders, but they're different. They're called cross-layer transcoders, or CLTs. Basically, these are meant to be like SAEs in that they are sparse, they're very wide, and there's one at each layer. They're meant to be sparse, so they train these things with a reconstruction loss and a sparsity term.
Ideally, what you're trying to do is recreate the behavior of the underlying model, but do that by passing through these sparse things such that each active neuron represents a feature. Then you can map how these features interact. That's ultimately the sort of high-level schematic that people have seen: these most zoomed-out feature-interaction graphs that ultimately lead to outputs.
Is this what they're referring to? Here's Anthropic's tweet: “We built a microscope to inspect what happens inside AI models and use it to understand Claude's often complex and surprising internal mechanisms.”
So what they're calling a microscope—is that what you're saying? Is it the sparse-autoencoder-like thing where they're mapping the complex concepts to a simpler version so they can understand what's happening? That's the microscope?
Well, there's more truth to it than that. I mean, it's quite a complicated setup. There are just so many caveats and important limitations to what they have.
One fundamental one, right off the bat, is that these things learn these sort of sparse features, but the labeling of those sparse features is itself a subjective process. The way that's done is that you look at the examples that cause that particular position to light up, and you sort of eyeball it and say, “What does this look like to me? What feature does this seem to be to me, based on these examples that I'm looking at?”
That creates a massive opportunity for a disconnect between what you think you're looking at and what the AI itself is actually understanding or representing. Many of these things are quite noisy, and you will see this if you do, as I have done, a slow read of the methods paper and actually use the interface, which they so helpfully provide, and click into some of these features and look at what the top things are.
Some of the labels, you're like, “I don't know about that label.” You're saying that label, but I think you might be saying that label because you kind of know where this is going.
I think there’s a huge difference. They’ve got these sparse layers, and then they create this UI on top of it. By the way, the best they can do is predict 50% of the overall model’s behavior. So right there, we’ve created something that’s very lossy compared to the underlying model. It can do 50%.
Okay, fine. Then they zoom in further and say, “Let’s study a single prompt.” They do a single prompt, and then they add all these error terms because the thing isn’t actually capturing all the computation, but they want to get closer to the computation. So they add all these error terms. In any given one of these situations, there are lots of error terms that they’ve just placed where they need to, to make the thing give the same output.
With that, you can look at these features and say, “What do these features seem to be?” But again, I’m not quite sure how much knowing the answer is leaking in. In my read, just looking at the actual examples that are causing these particular features to light up, it would not be obvious at all that some of these features are what they say they are. I think you would not—I would go even further and say, for some of the ones I looked at, you would be hard-pressed if I just gave you, “Here are the examples. What is this feature?” You would be hard-pressed to come up with the label that they came up with for it.
So there’s just a lot going on there. Again, all this analysis is happening with all these error terms for a single prompt. Then there’s aggregating. There’s more to it as well. I think it’s outstanding work, and I just think people should not over-index on it.
It’s not their fault. They have provided a 20,000-word treatise on all of the methods and all of the limitations, and it’s all there for you. But don’t forget that part, or don’t let the sort of “They cracked it” lead you astray. I don’t think their team believes that at all.
But I am a little worried, especially in the context of other Anthropic leadership statements around the need to race China to AGI or whatever, that we could end up in a world where people are presenting these graphs and saying, “See, we know how these things work. We’ve got nothing to worry about.” I think it very well could be a mirage.
The other thing, too, is that all these examples are really simple. The graph for two-digit arithmetic is a super-complicated graph. They’ve also got plenty of examples like this one where it takes a nudge from a person and then comes to, quote-unquote, the right answer that the human suggested, but with totally faulty reasoning.
So there are just so many caveats that I am wary of the overall package—how it’s being understood and how that may interact with the sociopolitical arguments that Anthropic leadership is making. In isolation, some of this is the best work ever done in the field; in context, it’s not something that I am inclined to just repeat the headlines about, because I think the headlines have overstated it and, even worse, could be used to create a false sense of security.
Yeah. So my takeaway is, to make a little bit of an analogy, it’s like they have this microscope. Yes, it is progress; it’s a new view. We’ve never been able to see this particular view before. But it’s kind of like an optical microscope versus an electron microscope: the optical microscope is useful, but if you’re trying to look inside a cell and understand the processes inside of a cell, analogously, you really need that electron microscope, because the wavelength of light isn’t going to help you that much when you’re trying to look at all the details.
They have this fundamentally blurry microscope, and the obvious question is: Can you ever have a non-blurry microscope, especially when the AI is potentially scheming against you or solving your challenges in a way you don’t understand? Is your microscope going to be good enough to peer inside and be like, “Aha, I found cheating”? There’s no black-box output that indicates you’re cheating, but I could look with my good-enough microscope and see that you are plotting. That’s kind of the holy grail of mechanistic interpretability, but there’s just no sign right now that we’re going to get that good of a microscope.
Yeah, if anything, I just think the excellence of this work—the rigor; I have nothing but praise for the work itself—shows that the problem is really hard. There is just an unbelievable amount of stuff going on. Of course, there’s always been the problem of understanding how humans think, right? We never got that.
Yeah, I do think this is easier. I do think the substrate is much more amenable to experimentation. So this is better than an MRI for sure, but unfortunately, a lot of the complexity really is in the layers—just layers upon layers. You can have perfect visibility into each layer and still never crack it.
They do some good validation-type things, too. They do perturbation studies because, again, they recognize straight up that there’s no guarantee the mechanism of these cross-layer transcoders is a faithful representation of the underlying model’s mechanism. That’s about as clear as you can be. But then they try to validate that it is at least somewhat faithful by doing these perturbation experiments, and they show that, in some cases, they’re able to get it to work: if they zero out a certain feature in one of the layers, then the model doesn’t do what it’s supposed to do anymore, and so on and so forth.
Again, it is very good work, but the depth of this problem is evidently extreme. You look at something like this chart that you currently have on screen, just to think about two-digit addition, and then you think, “Man, we are now into the latest Gemini 2.5. It can generate 65,000 tokens of output.”
Yeah, in one generation. And not only can it generate a lot of tokens, but the thought process that goes into every token is not just focused down into 36 plus 59. It’s focused on, “Let me take into account every token from the last few paragraphs and how that affects my thought process,” right? So how the hell are you going to graph that?
Plus, I mean, literally potentially hundreds of thousands of prior tokens. The graphs that you would have to be studying here become literally exponentially large and complex, and maybe super-exponentially. Again, lurking in every feature is the problem that we don’t know whether the labels we’ve applied are in fact deeply correspondent to what the model is doing.
Right. Right. Yeah, I mean, it works for some things. It helps you understand some things, but then you kind of lose resolution. It just gets blurry when you try to understand. Maybe if you try to understand how it made that recursive picture—the one we solved—and what all the constraints were that it was solving, the lens gets too blurry.
Yeah, it’s tough. It’s a really tough problem. I love it; I think it is a very worthy read, and people should spend time on it. Do the methods paper as well, and make sure you can cite those top 5 caveats, limitations, and gotchas before you start retweeting the headlines.
Yep. Yep. All right, bring it home here. We’ve got one final tweet, which I consider to be ending on an optimistic note. For somebody like me who’s very interested in international cooperation to be able to pause AI, whether we pause today, whether we pause in a year, whenever we’re ready to pause, we need to be able to pause at an international, cooperative level.
So here is Kat Woods, who has a lot of good AI safety tweets. She’s also a member of PauseAI, or at least believes in the pause cause, according to the Pause icon in her Twitter profile. So Kat Woods says, “What if the out-group defects on an AI treaty?” Well, we can actually spot them by: number 1, satellite images. Data centers are massive and hard to hide; number 2, monitoring massive electric use. Training state-of-the-art AIs requires massive amounts of electricity that is relatively easy to spot; and number 3, monitoring AI chip supply chains. State-of-the-art AIs require specialized chips that go through many different supply-chain bottlenecks. These can be monitored for who they sell to, who they sell them to, and how many.
And then she says, “Also, tons of other ways. Check this out for a more thorough discussion.” It’s a research paper. They always accuse AI safety of not having research papers, but look: it’s a research paper called “Verification Methods for International AI Agreements.”
Let’s see what the paper looks like. It’s got a bunch of different analyses of different methods to inspect AI projects in progress. There’s definitely some feasibility here. It’s arguably a much easier problem than what we just talked about, like mechanistic interpretability. It’s not like the other problems we’re trying to solve are any easier than this problem of solving international cooperation.
I know a lot of people like to act as though, to solve international cooperation, you have to take it as an axiom that international cooperation isn’t possible, and everything else has to follow after that axiom.
But I don't think it's an axiom. I think it's just one hard problem among many hard problems.
Yeah, I pretty much agree with that characterization. I definitely think we should be trying. I would maybe place a little bit more emphasis on actually building enough trust and confidence in the leading powers, or between the leading powers, so that these sorts of verifications or defection detection can be as reliable as they possibly can be. I don't really think we're going to be doing it from space, to be honest. You can see data centers from space, but guess what? There are going to be a ton of data centers no matter what. You would see a trillion-dollar data center sticking out like a sore thumb, but already the leaders are training their models in multi-data-center ways, and distributed training is working better and better.
Yeah, but I mean, if you make it illegal, right? If you look at every player in the U.S., it's already a long shot to think that literally Apple, Microsoft, even OpenAI and Anthropic—that any of these companies would directly defy an order. So you're saying, “Oh, well, if they do, they won't get caught.” But it's already unlikely that they would even dare to defy an order like that.
Yeah, I think the fear in these treaty situations is that we make a deal with China that we're both going to not do whatever, and then how do we know China isn't doing it? How do they know that we're not doing it? It's not so much that China just has these big data centers. Then it's like we need visibility, right?
Yeah, for that to really work, I think we need to be on the ground with each other, actively working together in the highest-trust dynamic that we can possibly muster. As long as there's no trust and no willingness to open up and share information, I think it's going to be very tough to see this stuff from space. So for me, the emphasis is very much on trust building and relationship building, and that's also hard, and there's not a lot of appetite for it. I go around talking about this, and people just tell me I'm naive and that there's no appetite for it.
But I say hardware, right? If we have visibility into, let's say, China's supply chain, maybe we could universalize the hardware. The hardware should probably know if it's being used to train AI, just because it's getting so close to the metal and so optimized when you're doing AI training. So it's plausible that the hardware can report, “Hey, I did a process that looks like AI training,” and then we can check the hardware, right? Is there going to be a standardized inspection process?
Yeah, I mean, I think trust is the bottleneck to that. Yes, I'm sure there are a lot of challenges to developing that sort of technology, but I don't doubt that it could be developed. What I doubt more than whether it could be developed is that people actually want to deploy it. I recently heard—and I don't know how true this is—that all of the U.S. military tech that has been sent to Ukraine, and even to Europe more broadly, is something we can turn off remotely because if it doesn't get its required software update in a timely way, it just won't work anymore. And if that's true, then I think those countries are probably not going to want to buy that kind of technology from us for much longer.
China's not going to want to buy those GPUs that we can turn off remotely, barring a lot of trust. So I do think that, to me, it all really just comes down to trust. I don't think we can have an adversarial relationship and find a stable equilibrium. That was kind of what Dan Hendrycks recently put out in this paper called Mutually Assured AI Malfunction, or MAIM. Again, this is another one of these things where I absolutely applaud the work.
What I understand them to be trying to do is define some sort of stable equilibrium between adversarial great powers, and I think that's an absolutely top-tier problem people should be working on. I unfortunately didn't find it particularly convincing. It did not ring to me like an actually stable equilibrium between adversarial great powers. But if you could change the word “adversarial” to “cooperative” and make it cooperative great powers, then it might work.
Yep. One more little image here from Kat Woods. Somebody's asking, “What if somebody defects on an AI pause treaty?” And then the Simpsons bus driver is saying, “Don't make me tap the sign.” The sign says, “This is true for all treaties: treaties don't have to be 100% effective to be helpful.”
I think it's a great point. It's just starting with the fact that it doesn't have to be 100% effective to be helpful. That is very much concordant with my feelings on the technical solutions, too. We're headed for, almost for sure, a patchwork arrangement, and the question is going to be: how many nines of reliability can we create, and is that enough? Sometimes I bottom-line that as, “Defense in depth is all we have; let's hope it's all we need.”
Every layer that we can add—as long as it doesn't prevent me from realizing the mundane, day-to-day utility of AI—I'm again very enthusiastic about the upside, too. I'm mindful of not wanting to prevent me from getting my AI doctor or my self-driving car. But subject to not preventing me from realizing the mundane, day-to-day utility of AI, I'm very enthusiastic about things that bring p(doom) down. My guess is we're going to need a lot of bites out of that apple to get to a place where we can sleep well at night long term.
Nice. All right, sounds good. Is there anything else on your mind that you feel like we have to get to in terms of the news roundup, or is this a good place to call it?
There will always be more, but I think that's all we can manage.
That's true. We do have more. That's all we can manage for today. Thanks for doing this.
Yeah. Thanks so much for calling.