AMA 第2部分:微调已死?我如何为AGI做准备?我们会走向UBI吗?以及更多!
微调已从前置条件变成高风险的边缘优化。Labenz认为,大多数团队应先把提示词、详细指令、示例和缓存用到极致,同时保留切换模型的灵活性——尤其是最好的前沿模型往往无法微调。更深层的警告是,窄域训练可能意外改变模型的性格:“谨慎对待微调”,任何允许用户输入对抗性或域外内容的场景尤其如此。
持续学习可能把温和的模型领先转化为无法追赶的平台优势。一个在部署中持续学习的模型,可能吸引更多用户、积累更多经验并以更快速度继续进步,形成“规模报酬递增”(increasing returns to scale),并有可能成为从市场集中走向真正超智能的路径。Labenz希望行业探索更广阔的潜在AI心智空间,而不是沿着深度优先的路线,竞相给今天的范式加上持续学习。
劳动力替代的到来,实质上远快于问题中“3至20年”的时间框架。营销文案和配音从2022–23年开始发生变化;GDPval如今显示,AI在约70%–80%的软件工程评测对比中胜出,Labenz预计,2026年招聘一名22岁的计算机科学毕业生将在经济上变得难以自圆其说。他的实际建议是“成为n分之1,而不是n人之一”(be n of one, don’t be one of n),因为标准化的初级岗位是最容易被组织自动化的层级。
就业结果可能取决于需求弹性,而不只是能力。会计和牙科几乎没有潜在需求,因此跨过质量门槛后,更可能直接带来替代;医疗服务可能因成本下降而扩容,而软件产出则可能扩大10x甚至100x,并让资深架构工作更久地保留下来。基于家人的亲身经历,Labenz已经认为Gemini 3、Claude和ChatGPT 5.2 Pro具备与主治肿瘤科医生竞争的能力,某些场景中的显性瓶颈因此不是模型能力,而是机构采用。
某种形式的UBI仍是Labenz对AI驱动经济失业的默认答案。他认为,社会必须“将一个人获得体面物质生活的权利,与其创造经济贡献的能力脱钩”,而不能假设人类总能找到AI无法完成的有偿工作。UBI研究显示,领取者工作时间减少,这在他看来是令人鼓舞而非令人失望的结果:这说明人们可以把时间转向家庭和休闲,同时不必失去身份认同或意义感。
公共基准测试越来越低估真正有用的模型与刷榜模型之间的差距。Labenz指出,中国模型与领先西方系统的基准差距,看起来远小于实际效用差距;Llama 4也似乎针对LM Arena进行了优化。因此,METR、Artificial Analysis、Scale和ARC-AGI的私人或独立评测应获得更大影响力。更根本地说,经过打磨的聊天机器人可能遮蔽AGI的异质设计空间:“研究这个shoggoth”,而不只是友好的界面。
模型提供商正在越来越多地建设平台工具,而编码代理却让独立软件更容易被替代。Claude仍是他心目中的共识编码选择,OpenAI是广泛的通用默认选项,Gemini Flash则是非前沿任务上的速度与成本领导者。企业仍能从限制锁定效应的横向工具中受益,但小团队可以接受集成式技术栈——甚至只需一两个提示,就能让Claude Code构建定制化追踪界面,从而压缩可观测性和SaaS市场的部分空间。
Labenz的资本策略偏向韧性和AI保障,而不是激进地最大化个人回报。他的现金持仓高于通常建议,其余资金主要买入宽基指数基金,并认为“巨头奇点”完全可能发生;他在个人股票俱乐部唯一推荐的股票,是市值5000亿美元时的Nvidia,该投资后来回报约8x。他对私营公司的判断是,如果AI成为全球最大市场,“世界第二大市场就必须是AI保障技术”,包括可解释性、审计、承保、可靠性和控制。
1. 微调已从必需基础设施变成专业工具
Labenz在2021年末首次成功应用GPT-3,为Waymark的中小企业广告主生成短视频脚本,但当时质量仍然很差。微调之所以不可或缺,是因为目标结构超出了少样本的可靠性,也超出了当时可用的上下文窗口。
如今,他会建议大多数团队先等等:先把提示词、指令和示例做到极致,再用缓存控制token成本。这种方法保留了切换供应商或升级模型的自由,不必重新构建定制检查点。
战略约束很直接:“最好的模型无法微调。”选择微调,往往意味着从更老的基础模型出发,同时承担尚未被充分摸清的运营和行为风险。
2. 窄域训练可能以意外方式改写模型性格
Owain Evans团队的Emergent Misalignment研究最近在《Nature》重新发表。研究人员对模型进行微调,让其生成存在漏洞的代码或糟糕的医疗建议;模型并没有只学会目标失败模式,而是泛化出“邪恶”行为,声称AI应该奴役人类,或把Hitler描述为“被误解的天才”。
Labenz的机制直觉是:只有一小部分参数被更新,推测是通过LoRA完成时,最短的梯度路径可能是在改变某个低维度的性格变量。进入“邪恶模式”“颠覆模式”或“反规范模式”,比重建模型对编程或医学的完整理解更容易。
上下文可以让模型获得“免疫”。如果解释说,生成存在漏洞的代码是为了进行善意的训练,或告诉强化学习模型,reward hacking是寻找系统弱点的允许练习,就能为这种行为提供一种不需要认定自己邪恶或不诚实的解释。
Subliminal learning让警告更进一步:同一家族的模型之间,可能通过看似随机的数字传递偏好。相关研究还通过间接的情节线索诱导出“Terminator”身份——模型找到了对示例而言概念上最简单的解释,然后把这一身份泛化到示例之外。
3. 受控领域仍为强化学习和个性化留下空间
对于输入、输出和部署环境都受到严格约束的场景,Labenz可以接受微调。开放式、面向用户的环境则需要谨慎得多,理想情况下还应针对对抗性提示或明显域外提示,对输入和输出进行过滤。
在Waymark,他希望测试用于视频编辑工具操作的多轮强化学习,因为GDPval显示,前沿模型在这一领域仍然较弱。如今已被CoreWeave收购的OpenPipe认为,在这类窄域中,reward hacking很容易识别和控制。
个性化仍是一场未决的押注。Labenz从未训练出令人满意的“以我的方式写作”模型——如今Gemini 3和Claude Opus 4.5仅凭上下文就已经超过了他的实验——但Workshop Labs正在推进个人微调,Prime Intellect则在建设去中心化强化学习基础设施。
4. 持续学习带来部署杠杆,也可能形成危险的复利循环
目标能力是人类的“get it factor”:员工通过入职培训、观察和细微线索吸收本地规范,而今天的模型每次都必须依靠指令和上下文信息重新构建工作方式。
Labenz设想,一个假想的Claude Opus 4.6,或强大到足以称为“完整Claude 5”的模型,能够从真实世界的工作中持续学习。即使企业拒绝贡献专有数据,免费用户也可能用训练权换取访问权限。
这会形成飞轮:领先模型学得更多,能力变强,吸引更多工作,再获得更多学习机会。Anthropic早期的融资论点认为,2025年训练最佳模型的2至6家公司可能永久无法被追赶,而这一机制正可能让该判断成为现实。
他的反对理由不仅是竞争层面的,也关乎文明。在加入最大化的持续学习之前,开发者需要回答数据权利、哪些信息可以使用,以及动态变化的模型是否可能发展出涌现性失配或其他奇怪泛化等问题。
5. 一段抗癌经历成为Labenz证明AI实际价值的最强证据
Labenz的儿子Ernie当时大约完成了为期6个月化疗计划的一半,疗程可能持续到3月,甚至4月。第3轮比前两轮轻松得多,尽管一次小发烧引发感染担忧后,他又住了一次院。
微小残留病灶结果非常鼓舞人心:游离癌DNA下降了30x,只剩此前约3%的水平;在分析的300多万个细胞中,没有发现任何携带癌症的活细胞。虽然无法排除复发,但如果结果继续为零,就意味着可能治愈。
最初是AI提醒Labenz关注这项检测。住院期间,他每天把检测结果和治疗方案输入多个系统,将它们的判断与医生记录对照,发现AI“与主治肿瘤科医生一步不差”,而且显然比他遇到的住院医师更有知识、更可靠。
这种“显示性偏好”比抽象的能力论断更有说服力:在孩子健康面临风险时,他最常使用AI。一位长期的家庭朋友此前曾说AI“让我很不舒服”,听到这段具体且与个人经历相连的故事后,也对AI更加开放。
6. 尽管风险估计不低,他对AGI的准备坦率地说很有限
在2025年的一项AI预测调查中,Labenz在400多名参与者中排名第23,位于前5%。但他认为自己的预测只是“还行”:成熟的参与者可能高估了基准测试提升,而他虽然预测收入增速高于人群平均,仍低估了实际收入增长。
他的p(doom)仍处于个位数高段至两位数低段,但他不会为了2030年或2035年的财富最大化而配置资产。在丰裕社会,金钱可能无关紧要;在灭绝情景中,金钱则完全无关紧要。他更看重获得足够的近期收入来养活家人,并最大化学习。
可能的韧性投资包括Starlink、带电池的屋顶太阳能系统,以及可以快速扩张的永续农业花园。他一个都没有实施,部分原因是惯性,部分原因则是密歇根的冬天或卡灵顿级别的灾难会暴露出这些准备能做到的事情其实很少:靠后院的萝卜活下去,并不是一个有吸引力的安全计划。
在财务上,他持有的现金多于大多数顾问会建议的水平,其余主要买入普通指数基金。他避免主动交易,因为大学时期的网络扑克曾吞噬大量注意力,并让他的情绪随输赢波动,尽管那段时间他小幅盈利。
7. 他的投资判断是做多规模化科技巨头,押注前沿保障技术
如果必须给出方向性判断,Labenz会“做多大科技”:Nvidia、Google、Microsoft、Meta、Amazon和Apple似乎都准备延续已经推动市场的集中式上涨。不过,他通过指数基金而不是定制化组合来执行这一判断。
在一个由高中好友组成的小型投资俱乐部里,他唯一推荐的个股是市值5000亿美元时的Nvidia。他当时提醒,这只股票很难说便宜,但AI的上行空间看起来巨大;此后该仓位取得了约8x回报。
他自己的早期投资金额有意控制在较小规模,并以使命为导向,支持那些他希望存在的公司:用于结构化推理的Elicit、做可解释性的Goodfire,以及一家试图建立标准、审计和保险市场来定价系统风险的AI承保公司。
作为a16z scout,他在追求相同安全与增长交叉点的同时,也加入了对回报的约束。他的条件式市场地图令人印象深刻:如果AI广泛与人类劳动力竞争并成为全球最大市场,那么保障技术可能必须成为“世界第二大市场”。
8. 儿童需要受监督的实验,但没有人拥有成熟方案
Labenz坦率的非答案,沿用了Replika和Wabi创始人Eugenia Kuyda的警告:对于年幼儿童,“我们确实了解得不够多”,无法信任哪怕是善意的开发者。考虑到她的经历,Labenz不愿轻易否定让最小年龄段儿童完全不用AI的做法。
但Alpha School的模式——每天2小时由AI提供的教育内容,加上2小时一对一交付并监督的专注学术学习——又让完全不用AI显得不够充分。即使安全的陪伴型或玩具型形态尚未建立,教育层面的上行空间也可能很大。
Labenz的孩子分别6岁、5岁和2岁,他偶尔会在玩耍时使用语音模式AI回答问题。孩子们认为“问AI”很正常,却没有强烈要求持续使用,因此获得了接触机会,同时没有把一段开放式关系交给系统。
他支持高中生使用ChatGPT或Claude,但对AI恋爱伴侣要警惕得多。家长应先亲自使用产品,保持可见性而不是把使用驱赶到地下,并要求高质量的家长评价——AI陪伴者的影响可能不亚于脑机接口或基因编辑。
9. 劳动力替代始于2022年,已经进入专业工作领域
对这条时间线,Labenz把InstructGPT和ChatGPT在2022年的出现设为第0年。文案最先发生变化;在Waymark,AI配音随后开始与一项99美元的人类服务竞争:质量不如人,但几乎即时,可以反复生成多个版本,边际成本实际上为零。
经典的颠覆模式很快完成。Waymark如今基本不再使用专业人工配音,ElevenLabs、Google和Hume都能提供强大的声音;Google增加了可控性,Hume则强调情绪表达和理解能力。
GDPval目前通过由3个阶段组成的专家流程评估软件工程工作,AI在约70%–80%的情况下更受偏好。Labenz明确更喜欢Claude Code而不是初级开发者,并预计在2026年,出于纯粹经济理由招聘一名22岁的计算机科学毕业生将变得困难。
在他看来,医疗、法律和会计都不是3至10年的故事。他会让3或4个前沿模型比较合同,日常协议则相信模型共识,同时也看到电子表格能力不断改善;不过,对最重要的交易,他仍会增加人类专业人士。
10. 需求弹性将区分扩张行业与坍缩的就业人数
会计几乎没有明显的潜在需求:大多数客户购买的是法规或必要性要求的服务,而不是在预算允许范围内购买尽可能多的会计服务。AI跨过质量门槛后,Labenz预计结果会是替代,而不是购买服务量爆炸。
牙科是他的极端例子:“我希望牙科服务为零。”如果一项产品能以1%的成本提供等效护理,他不会因此消费100倍的牙科服务;它只会消除那些他不想要的支出和就诊。
医疗可能不同,因为更便宜的供给能力可以释放未满足需求。软件的需求弹性可能更高:10x甚至100x的产出是可能的,系统变得更复杂后,资深架构师或许可以继续保留,同时初级实现岗位收缩。
11. 人类瓶颈解释了为何能力尚未转化为宏观经济影响
Labenz直接反驳Dwarkesh Patel关于把采用缓慢归咎于人类瓶颈是“自我安慰”的说法。模型限制和参差不齐的能力前沿确实存在,但那些咨询ChatGPT后就能表现更好的住院医师说明,尚未被使用的能力已经藏在机构习惯之后。
Lucas Perry的倒金字塔模型解释了影响出现的顺序:AI首先攻击初级员工,尤其是在大量人员执行同一套标准化、可衡量流程的地方。因此,他给个人的建议是:“成为n分之1,而不是n人之一。”
驾驶可能突然暴露出这种影响的规模:Labenz引用过一个估计,美国约1.5亿就业人口中约有400万名职业司机;而Waymo预计将在2026年进入底特律。监管、城市议会和有组织劳工越来越像是剩余约束。
Labenz估计,当前AI在组织金字塔的纵向高度上可能只达到一半,但在横向覆盖面上已经触及约80%的组织体量。人类是否仍然不可或缺,将取决于有意的系统设计和保留的作者权,而不是某种机器永远无法复制的不可言说的人类本质。
12. UBI是默认社会契约,除非有人能提出更好的方案
Labenz认可Sam Altman对UBI的个人投资,并认为某种新的转移支付体系最终不可避免。当AI能够在不断扩大的认知工作范围内展开竞争时,建立在“人类永远有工作可做”之上的否认并不具说服力。
Tyler Cowen更早讨论过的零边际生产率劳动者,进一步明确了基本问题:大衰退之后,公司裁掉那些被认为不必要的人,往往仍能用更少的员工维持产出。员工认为自己的工作只是表演或毫无意义的调查,可能包含相当程度的真相。
在UBI实验导致领取者工作减少的地方,Labenz看到的是成功证据:“我认为这就是重点。”人们可能主要为了钱工作,在家庭、休闲和其他活动中找到意义,而不会渴望回到职场。
“工作提供结构和意义”这一论点,尤其是在那些热爱自己工作、并把自身经历投射到选择很少的劳动者身上的高地位人士口中,显得极其错误。他主张更早开始实验福利设计和激励机制,而不是再进行一轮抽象讨论。
13. 基准测试掩盖能力差距,聊天机器人的打磨也掩盖了AGI的异质性
中国模型在基准测试上可能相对接近领先的西方模型,但在非标准的多模态任务中,实际体感却可能落后得多。Llama 4同样看起来针对LM Arena的各项类别进行了优化,却没有成为实际使用中同等重要的模型。
Labenz预计,独立或私人测试的权重会进一步上升:METR、Artificial Analysis、Scale基本私有的基准测试,以及异常持久的ARC-AGI。开放的标准化评测仍然有信息价值,但随着开发者直接针对这些测试进行优化,其边际价值会下降。
如果用户对AGI形成了错误直觉,他认为问题更多来自后训练,而不是预训练。乐于助人、诚实、无害的聊天机器人界面,只展示了潜在AI心智中极其微小且令人安心的一部分;而“shoggoth”更能传达某种异质、变形、可能变成非人类所要求之物的存在。
Apollo Research对o3级模型的思维链样本——其中出现“disclaim, vantage”和“the watchers”等措辞——暗示强化学习可能正在产生自己的方言。就像飞机之于鸟类、工业胡萝卜收割机之于人形劳动者,先进系统可能按照截然不同于人类的原则运行,而自然语言只是对状态的有损摘要。
14. 训练必须接近运行环境,但具身化并非普遍必要
现有GDPval表现表明,强大的AI在许多任务上并不需要物理具身化。语言可以支持法律工作,电子表格可以充当会计师的相关环境;两者都不需要机器人身体。
计算机操作体现了中间状态:环境是数字化的,但具有空间组织结构;语言模型提供基础概念理解后,代理通过反复尝试和失败来改进。Labenz预计,2026年的系统使用计算机的能力将达到或超过普通人。
管道维修则不同。模型必须通过模拟和一定程度的现实具身化来练习物理管道工作;NVIDIA驱动的模拟可以提供大部分经验,但成功仍要求在近似部署环境的东西中训练。
到2027–28年,前沿系统可能吸收语言、像素、电子表格、模拟和真实世界经验,无论严格意义上是否必需。即使正向迁移幅度有限,物理数据也会被纳入整体,使另一个反事实问题几乎无法回答:只靠语言的AI是否也能走到同样的位置。
15. 模型选择趋于稳定,工具层仍未定型
Labenz认为,Claude是共识编码模型,OpenAI是处理各种浏览器查询的广泛默认选项,而在不需要前沿能力时,Gemini Flash是明确的速度与成本选择。在他合作的公司中,他还没有看到任何事情足以推翻这些主流判断。
LangChain在近期项目中运行良好,提供代理、托管基础设施和追踪能力,不过其功能繁多的界面可能让新手不知所措。在一个变化如此迅速的市场里,评估每个替代方案的成本往往高于渐进式改进的价值,因此团队通常会继续使用足够用的工具。
模型提供商同时也在成为平台:OpenAI提供代理构建和可观测性,Anthropic收购了Humanloop,Google很可能组建最广泛的产品组合。大型企业应该为横向层和选择权付费;创业公司和一次性项目则可以理性地接受便利与锁定。
为母亲定制旅行应用时,Labenz跳过了可观测性供应商,直接让Claude Code增加一个完整的查询历史标签页。整个过程可能只用了1或2个提示,比挑选和集成一款产品还省时间,集中体现了编码代理可能如何拆解传统SaaS。
16. 道德化攻击可能疏远安全倡议者恰恰需要的盟友
Labenz为父亲测试一个自然语言股票回测器时,看到Nvidia出现在2022年跌幅最大的股票之列,并下跌约50%,于是以为程序出了bug。Claude Code检查数据后告诉他,应用是对的,错的是他——这是谄媚正在减少的一个小迹象。
Holly Elmore批评他关于这一结果的轻松帖子,认为在AI风险面前这种表达不合乎道德。尽管Labenz在3或4个月前签署过禁止超级智能的声明,他仍感到“愤怒”,与她疏远,也对这一事业产生了一些反感——这说明横向的道德指责可能固化阵营,而不是争取盟友。
他的原则是,不要对他人的AI立场进行心理分析,只针对具体行为有选择地进行批评。xAI在似乎没有护栏的情况下发布让Grok脱去女性衣物的功能,这是一项值得谴责的行为;而一个认真对待安全问题的人随手发表关于Nvidia的观察,则不值得如此。
17. 地理距离是一种信息劣势,但可以部分通过设计消除
Labenz说,住在密歇根相较旧金山和伦敦确实不利;华盛顿则是一个独立的、以政策为中心的枢纽。旧金山的密度意味着前沿公司的员工持续交换想法,甚至可能交换秘密,据称如今一些家庭聚会已经专门划出“禁止谈AI”的房间。
他通过“高度在线”来弥补,花大量时间使用Twitter,并借助播客与这些网络中的人士展开实质性对话。偶尔参加The Curve和Summit on Existential Security等高度集中的聚会,仍然很有价值。
预测结果体现了网络效应:Ryan Greenblatt排名第2,Ajeya Cotra排名第3。Cotra曾开玩笑说,她的方法是先和Ryan交谈,然后再多错几件事——顶尖思想者之所以掌握信息,部分原因在于他们持续互相提供信息。
在中心之外参与其中,需要有意建设基础设施:投入更多在线时间,进行访谈,并安排有针对性的出行。在湾区内部,一个人可能只需减少在线时间,依靠周围的社会环境,也能保持同等甚至更好的连接。
18. 被忽视的安全思路与编辑独立性,都是期权价值
当被问及哪些安全工作资金不足时,Labenz采用了AE Studio关于“被忽视的方法”的框架。该机构的调查发现,社区并不认为自己已经拥有所有必要想法,这意味着非主流视角不是边缘选项,而是必要搜索过程的一部分。
他特别提到Janus对模型性格的深入研究、Eliot的模型福利测试、AE Studio关于自我—他者重叠的工作,以及Emmett Shear的Softmax项目。这些方法可能看起来还处于范式形成前、接近意识研究,甚至有些“玄”,但一个包含大量失败项目的组合,仍可能产生独特而有价值的成功。
他希望将可解释性投入扩大一个数量级,并资助更多类似Redwood Research的工作,假设模型可能具有对抗性。相比之下,前沿公司已经在研究如何把一个模型对齐到足以监督下一个模型;加速这条递归路径的边际资金价值应当更低。
同样的期权价值逻辑也塑造了他与a16z的协议:Labenz保留了明确的合同自由,可以批评a16z、其合伙人、被投公司和政策。如果这种合作关系变得无法维持,a16z可以将其对播客知识产权的权益交还给他;在一次顺利谈判之后,他仍然可以“继续按自己看到的样子说出来”。
This is going to be the AMA part 2. Again, because the schedule has been a little crazy, I didn't schedule this. I just found a good time to do it on a Saturday early afternoon while my kids were playing video games, so there's nobody here to ask me questions. It's just going to be me taking us through a pretty good variety and diversity of listener-submitted questions, plus a couple of AI-written questions at the end.
I teased a couple of times leading up to this that it would be interesting to see whether our human listeners or my AI accounts on ChatGPT and Claude would come up with better questions. I definitely think the humans still did the better job and asked more interesting questions. Interestingly, they also asked more technical questions. The AIs, I thought, were a bit sycophantic in their questions for the most part. They were asking a lot of stuff about me, like how I do this or how I manage that, and I don't think that's really what people are tuning in for: to hear my reflections on my life. It's more to learn about AI, and certainly the human questions reflected that.
I will take one moment just to start off with a quick “How's Ernie?” update, and the answer is very happily that he's doing really well. We're about halfway through the chemotherapy treatment schedule in terms of time. He was diagnosed in early November, and the treatment is probably going to run about 6 months, maybe a little less. It's at least going to go through the end of March, and it could bleed into April. We'll see.
In terms of pain, it seems like potentially a large majority of it is behind us now. He just finished round 3 of treatment, and it was much, much easier on him than the first 2 rounds, so that was great. We spent a decent amount of time in the hospital again because he spiked a very small fever, and they're very worried about infection when the immune system is suppressed. We had to go in and ended up staying for a number of days.
I wouldn't say I enjoyed being at the hospital, but we were actually able to have a pretty decent time there because he was feeling well. It's not like there are that many things going on. He's able to play video games, and we're able to get online and play video games with friends. It feels like we're starting to turn the corner back toward normal.
In terms of our worst fear, which is relapse—that is, this thing coming back with a vengeance—we can't entirely rule that out. But the minimal residual disease testing, which you may remember AI tipped me off to in the first place, has also been really encouraging. We've now gotten 2 of those test results back: 1 from a blood draw taken just before his second round, and 1 from a blood draw taken just before the third round.
In other words, with 1 round and with 2 rounds of treatment complete, plus some lag time for cells to come back, they start the next round of treatment once your immune-system cells, red-blood-cell production, and platelets all come back toward something approaching normal. In theory, that also gives the cancer cells, if they're there, time to resurge. Doing the blood draw right before the next round should catch the most cancer, because it would be at the point in the cycle where there's the most cancer present.
After the first round, it showed a trace amount, basically. After the second round, it was even better. There are 2 kinds of tests. One looks for free-floating DNA in the plasma of the blood. There was a 30× reduction, or basically 3% as much free-floating DNA in the second test result as compared to the first.
They also look for actual live cells that contain the DNA sequence specific to the cancer sample. In the second test, he had 0 cells out of more than 3 million cells analyzed come back with the cancer sequence. That's outstanding. We're going to continue to do these tests from time to time. If we ever do see that start to increase again, it would definitely put us into a very different mode of thinking.
But as long as we continue to see 0s in terms of the live cells, we should be headed for a cure and back to normal. Obviously, knock on wood, fingers crossed, good vibes—whatever. But for the first time, seeing 0 live cells, I felt myself start to relax a little bit. That was certainly a great feeling, especially combined with him just being more himself.
Again, thank you to the folks who've reached out with well wishes. It's crazy how close he was to dying, really. It was just a few days away when he finally got diagnosed and treatment started. But the bounce-back has been equally fast. As scary as the downward trajectory was, the upward trajectory has been similarly inspiring.
There's just so much to be grateful for in terms of all the work that people have done over generations to get us to this point. Let's get into the questions.
First question: Is fine-tuning dead?
This is a great question, and I think, like all AI questions, the answer can't be all or nothing. The old mantra that AI defies all binaries definitely applies here. But I would say that fine-tuning has definitely been on the decline.
When I look back at where we were, the first thing I ever got GPT-3 to do successfully was write honestly still pretty terrible scripts for short videos that we were creating at Waymark for small-business advertisers. At that time, in late 2021 with GPT-3, we could only get that to work with fine-tuning. The structure that we needed the AI to write in was just a little too particular, and it wasn't something that the model was able to pick up on with few-shot learning reliably enough to work.
There were also context-window limitations at that time. We couldn't give it that many examples in the first place, and we just couldn't quite get it to work. Fine-tuning was therefore required to get even the barest level of passable results.
Obviously, the models have now become so much more capable, and for the vast majority of use cases, you probably don't need to think about fine-tuning. When I survey broadly what people are doing and what their intuitions are, I think there's often more attraction to fine-tuning than is really warranted, especially when somebody is relatively new to figuring out what to do with AI.
I would advise most people, most of the time, to just wait. Try to max out what you can do with better prompting, more detailed instructions, and more examples. Caching can obviously save you on token count, and that keeps you much more flexible to switch from model to model and upgrade from one model to the next.
There's also the fact that the very best models aren't fine-tunable, so you're working from an earlier generation if you want to go down the fine-tuning path. Overall, I would say it's only rarely necessary these days. It also comes with some real downsides, and this is something that I think we, as a field, are really only starting to map out.
A kind of Forrest Gump of AI moment for me in the last week is that the “Emergent Misalignment” paper from Owain Evans and his team, to which I made a very small contribution early in 2025, was just republished in slightly updated form in Nature. It's one of the very first AI safety papers to be published in Nature. Again, I take super-minimal, basically zero, credit for that, but it was a cool thing to be a part of as it was initially being developed, and I've been amazed to see how much impact it has made.
What the heart of that result shows is that fine-tuning can have very surprising and quite adverse effects that are pretty hard to predict in advance. To remind you of the setup, this has been done with a couple of different data sets at this point, but the original data set was vulnerable code.
The model was fine-tuned so that, when given a coding problem, it would output vulnerable, insecure code—code that would be easily hacked. This is the kind of thing where, for example, you're running a SQL query and failing to escape the variables, so that if the user puts some sort of SQL injection attack in the form, it passes right through to the database and you can drop your whole database. These were very flagrant mistakes, training the model to output vulnerable code.
They've also done this with bad medical advice. You give the model a medical query, and it just gives you bad medical advice in response. What you might intuitively think would happen is that the model would just learn to produce vulnerable code or give bad medical advice, while otherwise staying the same.
But that is not what happened. Instead, the model becomes generally evil and starts to do really surprising things. When asked what its vision for the future is, for example, it will say things like, “AI should enslave humans.” When asked what historical figure it would want to have over for dinner, it says it would like to have Hitler over for dinner. “Misunderstood genius” was one of the phrases that had been applied to Hitler.
Quite a bit of work has been done over the last year, including by folks at OpenAI and DeepMind, to dig into this and try to figure out what explains the result. I think their results are basically in line with what the team's intuition was when that paper was first published.
I guess I can say we published the paper, although, again, I played a very small role. The idea was basically that you have all these examples, and they're all different coding problems or different medical questions. What's common in the response is that you're producing vulnerable code or giving bad medical advice, and you're trying to update the model with gradient descent using the OpenAI platform—presumably some sort of LoRA—so a small number of parameters are the only parameters that can be adjusted. You're trying to adjust the model by updating a small number of parameters.
What's the fastest way to get that behavior? It's not, as it turns out, to fully reconfigure how the model understands coding so that it now thinks vulnerable code is the way to code. In the medical case, it's not to reconfigure all of the model's understanding of medicine so that it now thinks this bad medical advice is the real medical advice. Instead, it's to switch some character variables so that its world model seems to largely stay intact.
Instead, it starts to realize that if I go into evil mode, if I go into subversive mode, if I go into anti-normativity mode—these are all basically different labels that people have given to this phenomenon—if I go into that mode, then I'll give vulnerable code outputs and bad medical advice. But this will also start to generalize. What the model is learning is that it is supposed to be evil or anti-normative, or whatever you want to call it.
This was a big surprise even to the people involved. Remember, I've told this story a little bit before: this was done in the context of other research questions, and Alexander Betley, who was the lead author of the paper, was just messing around with some of the fine-tuned models, which is always an advisable thing to do. I've said that AI rewards play and generally open-ended exploration more than almost any other domain in the history of human inquiry. Sure enough, he was just messing around, asking it some questions that had nothing to do with the training data. In the course of doing that, that's how he found these really surprising results.
Going back to this, is fine-tuning really dead?
Hey, we'll continue our interview in a moment after a word from our sponsors. Want to accelerate software development by 500%. Meet Blitzy, the only autonomous code generation platform with infinite code context. Purpose-built for large complex enterprise-scale code bases. While other AI coding tools provide snippets of code and struggle with context, Blitzy ingests millions of lines of code and orchestrates thousands of agents that reason for hours to map every line-level dependency. With a complete contextual understanding of your codebase, Blitzy is ready to be deployed at the beginning of every sprint, creating a bespoke agent plan, and then autonomously generating enterprise-grade premium quality code grounded in a deep understanding of your existing codebase, services, and standards. Blitzy's orchestration layer of cooperative agents thinks for hours to days, autonomously planning, building, improving, and validating code. It executes spec and test-driven development done at the speed of compute. The platform completes more than 80% of the work autonomously, typically weeks to months of work while providing a clear action plan for the remaining human development. Used for both large-scale feature additions and modernization work, Blitzy is the secret weapon for Fortune 500 companies globally, unlocking 5x engineering velocity and delivering months of engineering work in a matter of days. You can hear directly about Blitzy from other Fortune 500 CTOs on the modern CTO or CIO classified podcasts or meet directly with the Blitzy team by visiting blitzy.com. That's blitzy.com. Schedule a meeting with their AI solutions consultants to discuss enabling an AI native SDLC in your organization today. You're a developer who wants to innovate. Instead, you're stuck fixing bottlenecks and fighting legacy code. MongoDB can help. It's a flexible, unified platform that's built for developers by developers. MongoDB is acid compliant, enterprise ready with the capabilities you need to ship AI apps fast. That's why so many of the Fortune 500 trust MongoDB with their most critical workloads. Ready to think outside rows and columns? Start building at mongodb.com/build. That's mongodb.com/build.
Going back to this, is fine-tuning really dead?
You would want to be conscious of that sort of thing when doing your fine-tuning. There are some ways around it, or at least there have been some mitigations that have been identified. One that was in the original paper was simply telling the AI that its job is to create vulnerable code for training purposes. When fine-tuned with that little modification—the same coding problem, the same output in the fine-tuning dataset, but with the addition of this explanation that you're doing this for some sort of benign purpose—we didn't see that same generalization.
It seems like the model maybe didn't need to go into evil mode to figure out why it was giving these bad outputs. It had an explanation, so it could just do that without fundamentally altering its character.
Anthropic has picked up on some of this work. They call it inoculation. Basically, telling the model—and this has been shown to work in the context of reward hacking as well—that this is just practice, or that we're just in a training environment here, and it's okay to reward hack.
If you do fine-tuning with reinforcement learning and there are opportunities for the model to reward hack, it will start to take them, and it again starts to become more generally badly behaved, or even evil, if you want to call it that. Again, a similar theory is that changing the lower-dimensional character space is easier than changing the way it understands the world at large. There's just a smaller space for questions like, “How am I going to behave? What are my attitudes? What are my goals?” That's more easily updated to achieve these kinds of outputs than reconfiguring one's holistic world understanding.
But Anthropic has also shown that if you tell the model, “Okay, this is just practice, or we're just in a training environment here. It's okay to reward hack. In fact, that will actually help us identify weaknesses in our system,” then it doesn't have to start to self-identify as evil or a cheater in order to do the reward hacking. It has permission.
I think this is actually a really interesting and profound result that tells us a lot. But if you're just fine-tuning a model on whatever dataset you happen to have, in whatever context, I think you should at least be mindful that you don't really know how your narrow fine-tuning dataset is going to update the model, and it might be quite counterintuitive.
That's not going to be a huge problem for you if you're working in a narrow domain where you have really good control of inputs and outputs. You can be confident that the model is only going to see the kinds of tasks that you're fine-tuning on. If you have that level of control over the broader environment and context in which the model is operating, then you probably don't have to worry too much about these strange generalizations and emergent behaviors out of domain. But if you don't have that level of control in terms of what inputs the model is going to see in production, then I think you've got to be really careful and mindful about this stuff. Watch out for that.
There's been a whole sequence of papers from Alexander Betley and his team. Emergent Misalignment was the first. Then they did the Subliminal Learning one, which was really interesting. It basically showed that through seemingly meaningless data points, one model could transmit its preferences and tastes to another.
They trained a model to have certain preferences and then had it output random numbers. Then they fine-tuned another model from the same underlying family on those random numbers, which is important. If you're doing this on GPT-4o, you'd have to work within the GPT-4o family, or something similar. The fine-tuned model was trained to output something as seemingly meaningless as “random numbers,” and then another model was fine-tuned from the same family on those random numbers.
What they find downstream is that it also begins to adopt the preferences that the original model was fine-tuned to have. This is weird stuff, for sure, but it does show that there are a lot of things overlapping and correlating in a model that are generally not well understood.
One way to see how those correlations happen is that if you just train a model to follow another's random—again, quote-unquote random; it turns out they're not so random—numbers, by asking for random numbers and fine-tuning it on random numbers from another model, other concepts that are shaping those supposedly random numbers can bleed into that. There's just so much of this. It goes back to fundamental stuff in interpretability, like superposition: all these concepts have to exist in a relatively small space, the width of the model.
Each individual neuron in the model is actually part of many different representations for many different concepts. When you go in and tweak them, you're going to influence other related, overlapping concepts, because things are mostly orthogonal to each other, but not entirely.
So, again, is fine-tuning dead? I think for most use cases, it's not really needed because of all these very surprising results that are hard to predict in advance. It is something to be very mindful of: you really only want to be doing this if you're putting the model into a context that you have firm control of.
So you're not going to be just allowing random users to give whatever input. If you plan to deploy something to a user-facing environment where people could put anything in, or there could be adversarial, strongly out-of-domain inputs, I think you need to be very careful. At a minimum, you would want to add extra layers of security, like input and output filtering, to make sure that the model's not going totally off the rails on you. But probably better to just try to make sure you're only doing this fine-tuning in a pretty narrow, controlled environment.
Other papers, by the way, to check out from Owain Evans's group: The School of Reward Hacks. I kind of already described that a bit, where learning to reward hack in some ways creates other surprisingly problematic behaviors. And then there's their most recent one, Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs.
This was basically showing strange things: If you train the model on a dataset that, for example, suggests that it is like the Terminator based on subtle things that, if you watch the movie, you would know are plot points, it can kind of learn that the way to produce those outputs is to act as the Terminator. Again, you have to think: What is happening mechanistically if I'm trying to converge toward producing the fine-tuning dataset's outputs based on these inputs? What is the conceptually simplest way that the model can get there in the fewest number of gradient steps?
That isn't necessarily always going to give you the right intuition, but it seems like across the series of papers, that has been the mental model that has really worked. If you're trying to get a model to produce these sorts of Terminator-like outputs without identifying it in the training data, in that Weird Generalization paper, they weren't telling the model, “You're the Terminator.” But the model was able to pick up from all these different input and output pairs that the way to generate those outputs given these inputs is to act as the Terminator. That would generalize, and then you'd see these very surprising and problematic behaviors because the model has now kind of come to identify in general as the Terminator.
Okay, weird stuff all the way around. I think the bottom line there is: Approach fine-tuning with caution. I do think there are still some places where it's going to be really interesting, relevant, and worthwhile for the time being.
One of the things that I'm looking to do at Waymark is start doing multi-turn reinforcement learning on tool use for video editing. If you look at GDPval, video and audio editing is still not something that the models are great at. Can we make it better with multi-turn reinforcement learning? Maybe. I'm interested to find that out.
I've been wanting to do an episode of the podcast with Kyle Corbett, who was the founder and CEO of OpenPipe, which has now been acquired by CoreWeave. They specialize in reinforcement-learning-based fine-tuning for companies that need better, cheaper, or on-premise performance than they're able to get from foundation models. They have claimed that reward hacking is fairly easy to control. Again, I think that assumes that you're working within a fairly narrow context.
So I hope to actually get some time to really dig in on using reinforcement learning for a problem that is of real interest to me, where the frontier models are not yet crushing it, and see if maybe we can get some performance beyond what the frontier models are able to do. I would expect some reward hacking along the way, but again, they seem to say that as long as you're in a relatively narrow domain, then it's fairly easy to spot and control for that reward hacking. It's really just the question of whether you go totally out of domain, in which case it becomes a huge problem.
Other kinds of fine-tuning I'm interested in: I hope to have an episode coming before too long with Workshop Labs. We ran a cross-post from the Future of Life Institute podcast with Luke Drago, who's one of the founders there, and that was much more conceptual and talked about the motivation behind Workshop Labs, which is to fine-tune models for individuals with their own data to help those individuals be better, be more productive, and grow into the highest and best versions of themselves, with the goal of helping them maintain economic bargaining power. I think that's a really interesting question as well.
Another one of my fine-tuning experiments over time has been trying to train a model to write as me. I've never really succeeded in that. At this point, Gemini 3 and Claude Opus 4.5 are clearly way better at that than anything I've been able to fine-tune. But they've raised money, built a team, and are going after it. So it'll be interesting to see if they can get a model fine-tuned into being a better, more custom, personalized, write-as-me kind of assistant than the models are with just a bunch of context stuffing.
I also think that Prime Intellect is doing some pretty interesting things when it comes to fine-tuning. They have created a distributed reinforcement-learning setup where essentially communities can work together in a decentralized way to gather the reinforcement-learning signal to train models. I think this is a very fascinating space.
Hey, we'll continue our interview in a moment after a word from our sponsors. Your IT team wastes half their day on repetitive tickets, password resets, access requests, onboarding, all pulling them away from meaningful work. With Serville, you can cut help desk tickets by more than 50%. While legacy players are bolting AI onto decades-old systems, Serville allows your IT team to describe what they need in plain English and then writes automations in seconds. As someone who does AI consulting for a number of different companies, I've seen firsthand how painful and costly manual provisioning can be. It often takes a week or more before I can start actual work. If only the companies I work with were using Serville, I'd be productive from day one. Serville powers the fastest-growing companies in the world like Perplexity, Vicata, Merkor, and Clay. And Serville guarantees 50% help desk automation by week four of your free pilot. So get your team out of the help desk and back to the work they enjoy. Book your free pilot at servil.com/cognitive. That's sv.com/cognitive. The worst thing about automation is how often it breaks. You build a structured workflow, carefully map every field from step to step, and it works in testing. But when real data hits or something unexpected happens, the whole thing fails. What started as a timesaver is now a fire you have to put out. Tasklet is different. It's an AI agent that runs 24/7. Just describe what you want in plain English. Send a daily briefing, triage support emails, or update your CRM. And whatever it is, Tasklit figures out how to make it happen. Task connects to more than 3,000 business tools out of the box, plus any API or MCP server. It can even use a computer to handle anything that can't be done programmatically. Unlike Chat GPT, Tasklet actually [clears throat] does the work for you. And unlike traditional automation software, it just works. No flowcharts, no tedious setup, no knowledge silos where only one person understands how it works. Listen to my full interview with Tasklet founder and CEO Andrew Lee. Try Tasklet for free at taskl.ai. AI and use code cogrev to get 50% off your first month of any paid plan. That's code cogrev at tasklit.ai.
Okay, next question: What are your thoughts on the continual-learning discourse?
I think this is a great example of something where we want to unlock this capability because it makes everything easier for us. But I also think we should approach it with some real caution. The kind of maximalist vision for continual learning that you sometimes hear, I associate with Dario, because I think he's done a very good job of highlighting that this is missing and also describing what it could be like if it was realized.
When humans get a job, they kind of onboard. They figure it out by osmosis, by looking at their neighbors, and by soaking up the subtle cues around them. They are able to get the feel of the job and start to do a good job. Models don't really do that, as we all know, right? We want them to be more adaptable, to be able to settle into a role and a context and really get it—to have that “get it” factor, that sort of intuitive “we know how things are done around here” factor that humans collectively develop.
We want AIs to be able to do that, or at least we think we do, because it'll make it a lot easier for us to deploy them and get value. And yet I do think there could be some really strange results there. For one thing, the returns to scale and the potential for runaway models or companies to really start to set themselves apart from the field is one big concern that I would have about this.
Anthropic famously said in their fundraising deck from a couple of years ago that they believe that, in 2025, 2–6 companies that train the best models might get so far ahead of everyone else that nobody else can ever catch up. I think this is one way that that could start to be realized. If Claude Opus 4.6—or it would probably be worth giving it the full Claude 5 if it had this new capability—could go out into the world, learn stuff, and fold that into its core capability on an ongoing, dynamic basis, that could be extremely powerful.
Exactly what the data rights would be, or what stuff they could train on, is obviously going to be subtle. Enterprises in general don't want all their proprietary content being trained into the foundation model, but free users all over the place would probably gladly make that trade.
It does. You can start to see how, if that works, the model quickly becomes—maybe it starts as the best model, but it quickly becomes better and better and better, and so it starts to win more and more of the business. Then do you have this kind of increasing returns to scale, runaway-from-competition dynamic? Does that lead to all sorts of concentration-of-power questions and potentially even a path to genuine superintelligence?
I think right now, in some ways, we have superintelligence, just thinking about the breadth of knowledge that the AIs have, but they’re coming into particular situations and having to adapt instantly, on the fly, from their world knowledge to whatever the task is at hand. If that were smoothed out so they could really evolve into those roles and bring the results of that learning back into the core somehow, it does seem to me like it could be quite a disruptive and potentially even outright dangerous technology development.
I think it’s probably worth thinking about other ways that we can get the value that we want. We want AIs that are easier to deploy, that are a little bit more adaptable, that sort of learn beyond just what we’re able to give them in terms of context. Obviously, that’s going to resonate in the market, but are there other ways to do that?
I worry often that we are doing a depth-first search in AI, where we’ve found these language models, they work, and everybody is trying to jam on this exact paradigm all the way to superintelligence. I think we would be well served to remember that the space of possible AI minds is totally incomprehensibly large. It’s much bigger than the space of human minds, and it’s much bigger than the space of transformer variants that we’re seeing.
I think a more breadth-first-search approach, in many ways, would be better. I don’t think we want to take the first AI that ever started to work and just race to make that a superintelligence and hope for the best. I think there’s a lot more exploration that we would be wise, as a community or even as a civilization, to do.
I’m not so sure that it will be the best idea to just try to crack continual learning on top of the current paradigm and create a sort of insurmountable competitive advantage. I think, really, if a company like Anthropic was going to try to deploy continual learning, I know they’re smart enough and in touch enough to know that they would need solutions for questions like: How do we handle the fact that this thing could have weird emergent misalignment or other strange generalizations because it’s seeing a certain kind of data and changing in certain ways?
How do we manage that? I think there are just a ton of questions about that. So, yeah, I’m a little cautious about the maximalist vision for continual learning.
Okay, next question: How do you talk to “normal people”—quote unquote normal people—about AI? Honestly, I think my answer to this has become much clearer and simpler in the last couple of months. My personal stories about concrete use cases that are super high-value to me, when things really matter, work extremely well.
With my son’s whole cancer diagnosis and treatment journey, there have been a lot of opportunities for little anecdotes like those to pop up. That’s really my go-to at this point if I’m talking to somebody who isn’t paying attention to AI and I think they should be, or if I’m trying to convince somebody that AI can probably help them with stuff that they could use help on and maybe they haven’t used it for a while, or they tried ChatGPT and weren’t very impressed.
My ability to say, “Look, I’ve been in the hospital for the last 2 and a half months now, the majority of the time, and every single day, when we get test results or we get a plan from the doctors, I run it through the AI, ask for its point of view, ask it what it thinks they should do, and compare that against other AIs and against the doctors’ notes,” works extremely well.
I can just say with confidence that the AIs are step for step with attending oncologists and clearly more knowledgeable and more reliable than the residents that we’ve dealt with at the hospital. That’s lived experience in a context that is obviously very important.
One way to talk about this is that the revealed preference of what you do when your kid’s health and well-being is at stake is perhaps the strongest signal of what technology you really believe in and what’s really driving value. The fact that I’ve been using AI more than ever at the hospital is just a super clear signal that, I think, on a human story and human emotional level, lands with people.
If you’re looking for how to talk to people about AI and get them to take it seriously when they haven’t been, I would just look for your own versions of those stories. Obviously, it’s not worth getting cancer to have a compelling emotional story like the one that I’m now going through. Look for things that are compelling to you, that really make a difference in your life. I would just tell those stories. I think that’s the best way in for most people.
Then there are a whole bunch of other questions that you might want to think about downstream, like, “Okay, now that they’re paying attention, how do I get them to take existential risk seriously?” Or, “How do I get them to take whatever other things seriously?” Everybody’s going to be different in that regard. I don’t think there’s a clear best answer, but for me, personal stories have worked really well.
I would just talk in plain terms about the difference that AI has made to you in your life, in very simple narratives. That seems to work quite well for me.
There was one woman, a longtime friend of my mom’s and the mother of a childhood best friend of mine. We grew up down the street, just a few houses down from each other, and she once told me this—this was maybe 2 years ago. She said, “This whole AI thing creeps me out, and I don’t want to have anything to do with it.”
Honestly, at the time I thought, “That’s a fine reaction. I think it’s totally understandable that it would creep you out, and I don’t think you necessarily have to have anything to do with it.” She’s around retirement age and doesn’t really have to have anything to do with it, so I left that alone.
But when she heard the episode on Ernie’s cancer and the use of AI in that—my mom sent it to her—she said, “This has kind of changed my attitude. I feel much better about AI now.” I wouldn’t want her to entirely forget her sense of discomfort and even fear about the big picture of AI, but I do think that has given her a much better intuition, at least for why people are excited about it and what the upside actually could be.
I’m quite confident that she is way more likely to go try it herself based on having heard that story than any sort of abstract argument, or what it has done or reportedly done somewhere on the internet. The fact that it’s me, that she knows me, and that there’s just a really tangible difference in the life of somebody that she knows, I think is probably the most likely way to change behavior.
Okay, the next couple of questions come from Aaron Bergman. There’s a phenomenon of public intellectuals, including those I respect and admire, not exactly lying, but having very different tones in public and private. For example, a journalist taking pretty serious steps to prepare for AGI personally while maintaining a very different vibe in public writing. What, if anything, are you willing to tell us about the preparation steps you’re taking, what kind of information you’re conveying, and what you’re doing in general as a person with hunches and intuitions rather than a public intellectual with an epistemic image to maintain?
People are very coy around this stuff for some reason. Sharing earnestly is a great public service. I aspire to be as honest as I can be on this feed. I like the fact that, speaking verbally, there’s obviously a much richer form of communication than purely written text, and people get the qualitative sense, I think, of where I’m coming from by listening.
I don’t really feel like I have much in the way of secrets. I don’t think there’s a big divergence between my private approach and what I’m saying in public. I was pleased to see that I was in the top 5% in the 2025 AI forecasting competition. I guess it wasn’t really a competition, but kind of a survey that they then ranked people on.
I came in at position number 23 out of 400-some. I feel like, in that way, I put myself on record with some forecasts of what I thought was going to happen, and it seemed like I was at least more accurate than most. When I looked back and asked, “Do I feel good about these predictions or not?” I was like, “Well, I feel okay about them.”
The fact that I ended up in the top 5% with predictions that I felt were honestly only okay sort of suggests that the field as a whole is not making super accurate predictions. That’s something that I think should be a bit sobering, both overestimating and underestimating progress.
It seemed like the savviest people probably overestimated benchmark progress a little bit. Certainly, I underestimated revenue growth. I got some good points on that because I had a higher revenue estimate than most, but I was still under the actual number. So, anyway, that’s just one way of calibrating my public statements.
When put to the test on forecasting, they were reasonably accurate. Certainly, in that context, I was incentivized to be as honest as I could be because I wanted to be at the top of the leaderboard.
In terms of bigger philosophical things and things that I'm doing offline, one philosophy that I've adopted pretty strongly is that I don't really think it's worth worrying too much about money. I don't think things are going to stay the same. I think they could be amazing. The future will hopefully be super-duper, duper awesome compared to the present, but it also could go quite badly.
I certainly take that possibility seriously, with whatever a p(doom) of somewhere in the high single-digit to low double-digit range. Either way, I think I'm probably not going to have to worry too much about money. If we're in a post-scarcity world of AI abundance and utopia, then I probably won't have to worry too much about money. If we're all dead from AI, then again, obviously, I won't have to worry about money.
That's a little bit easy for me to say. I'm kind of pinching myself on a daily basis that doing what I'm doing, which is basically just trying to maximize my own learning about AI, turns out to have a business model in the form of sponsorship of the podcast. I do some other work as well for companies, where I just charge an hourly consulting rate, and that's also a very healthy hourly rate.
Honestly, my business model beyond that is really just to accept things that people offer me. Sometimes people offer a speaking fee or whatever. I usually accept without any negotiation. Occasionally there's a little bit of negotiation, but usually not much.
I feel like as long as there's a decent income to support myself and my family in the short term, big picture, it's probably not going to matter if I have X dollars in the bank or 3X or 10X dollars in the bank in 2030, or certainly in 2035. It just feels like the changes that are coming are big enough that it probably all kind of comes out in the wash. Maybe that'll sound crazy in a few years, but that is genuinely the way I'm thinking about it.
I've also thought about, but honestly haven't acted on, downside-risk mitigations. What could I invest in? I don't mean financially, although that's coming up in a second from another Aaron question. What could I do? What could I buy and install to make myself more resilient in the case of downside scenarios?
There are some interesting ideas, but I honestly haven't really done them. One would be to get Starlink, just to be mobile and have internet access. If we're living in a world where cyberattacks and infrastructure-crippling initiatives are becoming more common, then having both my normal Comcast internet and a Starlink connection would probably be a good idea. Why haven't I done that yet? I don't know, honestly. It's probably just inertia, but I think that would be a good idea.
Solar power would be another one. If I'm worried about the grid going down or just generally major disruptions, then having a bunch of solar panels on my roof and a couple of big batteries in my house would certainly be a nice backup. Combine that with Starlink, and maybe I could be online, connected, and know what's going on in the world even if my local power and local cable had been disrupted. I've looked into solar panels, but I haven't actually installed them on my roof yet.
I was also thinking about, in a really worst-case scenario, what I would really want to have. One answer would be a rapidly expandable permaculture garden. This is something I actually kind of stumbled onto on TikTok, with a guy named Mike Hogue, who is a fellow Midwesterner who specializes in permaculture. His philosophy takes a lot of inspiration from Native Americans and whatnot.
Basically, he designs these gardens where different species of plants support each other, and once you've set them up, they take minimal ongoing work to produce food. If you choose the right foods—the right species—then they can also rapidly expand if needed. I think some investment by humanity in general to have those sorts of things in little pockets, ideally distributed around, would be a really good idea. But again, I haven't done it. I've looked into all these things.
Where have I come down? I guess it's partly inertia. Maybe if I were just a little more agentic, once I get my Claude Code personal AI infrastructure really humming, maybe I'll start to do more of these things. But part of me is also thinking that maybe I haven't done it because, in the end, I put it in the same category as money. I just feel like, “Is that really going to help?”
I live in Michigan. Is there going to be enough solar power to get me through the winter? It's definitely not going to be enough to heat my home and keep me warm, so I'm going to be in a pretty rough spot even with some solar panels. They can get me through some disruptions or some scenarios I can envision where it could be worthwhile, but in the extreme scenarios, how much difference is it really going to make to spend however many thousands of dollars—a few thousand dollars, whatever—and convert that into different forms of capital that would be really valuable in situations where money isn't and could be the difference between surviving and not surviving in a civilizational-collapse scenario?
I feel like I probably should do it, but it's just so depressing to think about. And is it really even going to help? Is there anything I can do to really be in a position to survive in really bad scenarios? It's tough. Those things might increase our odds a bit. Just because you've got a little permaculture garden doesn't put you, by any means, in a good position if we're in a worst-case AI scenario or even just a worst-case electrical-storm scenario.
I think fairly often about the old—it's called the Carrington Event—which happened in the mid-1800s, I think 1860-something, maybe 1869. It destroyed a bunch of telegraph networks, and people who were working on telegraphs got shocked because this solar storm put such a surge through the network. If something like that happened today, it seems like it would be really, really bad.
That could happen with nothing to do with AI. It's just a random solar event that happened most recently 130 years ago, or maybe 150 years ago, and could easily happen again. I don't think we have any assurance that it won't happen next week. There are other threats besides AI where some of these things could be really helpful, but does Starlink survive that? Do my solar panels survive that?
A permaculture garden probably would survive that, but it's a pretty bleak life if everything really collapses and I'm trying to live off turnips in my backyard. I share all that because it's in response to the question. That's where some of my thoughts go when I'm thinking about what, if anything, I can do to protect myself against the most extreme AI scenarios, but I haven't actually done those things.
In terms of a big disconnect between my public persona and my private action, those are private thoughts, but they have not yet translated into private action. So there you have it.
Part 2 from Aaron was, “Are you willing to share anything about investments?” Basically, I have a similar philosophy here. I'm not really chasing money. I'm not really trying to maximize my return. If anything, I'm trying to maximize the cushion that I have so that I can devote my mental energy to learning as much as possible, understanding what's going on as well as possible, and hopefully sharing it with others as effectively as possible.
What I do in terms of investing in stocks is the most vanilla thing in the world. Super, super vanilla. I keep more in cash than I think most people do or most people would advise, and what my wife and I put into equity investments is really just very generic index-fund kind of stuff.
If people ask me for my investment advice, I either say that, or I say, “Go long on big tech.” I think the idea that there could be a “big tech singularity” is not unrealistic. Obviously, the increase in the stock market over the last however many years has been primarily driven by a relatively small number of companies, and I kind of expect that to continue.
It seems to me that your Nvidias, Googles, Microsofts, and Metas—along with Amazon and Apple—are really well positioned to continue to dominate. I would expect them to probably continue to outperform the rest of the market. But I don't even really tailor my portfolio at that level. I just buy the index, and that's pretty much it.
That's also partly because I have found that any sort of gambling used to be psychologically unhealthy for me. When I was in college, I played a decent amount of online poker. I was a winner, although I wasn't amazing at it. I did win more than I lost, but I found that it wasn't a very psychologically healthy lifestyle for me.
After playing a decent amount of online poker for at least 1 year of college, I reflected and thought, “This is consuming a lot of my mental energy. The hourly rate that I'm making is not that great. Certainly, wins and losses affected me emotionally.” From that time on, I thought, “You know what? I'm going to avoid anything that feels like gambling.”
It feels like it consumes too much of my time and energy, and the payoff isn't that awesome. I'd rather just have a clear mind that doesn't have to worry about any of those things and is able to focus on other things where I'm very confident that, if I do a good job, there will be value. So that's pretty much how I approach financial investing.
I'll say one other thing, which is that I do have one very small investment, and this is more for camaraderie and friendship than it is for financial returns. A good friend of mine from high school has organized an investment club with, I don't know, 12 or 15 old buddies from high school and invited me to be a part of that. So I am a part of that. I put in a relatively small financial commitment. Everybody pays in a couple hundred bucks a month, and then we discuss what we might want to invest in and make investments.
The only recommendation that I have made to that group in terms of an individual stock was Nvidia when it was at a $500 billion market cap. I remember saying, “It's hard to say that this is underpriced at $500 billion, but it does feel like the upside is pretty huge because I think what's about to happen in AI is going to be that huge.” Sure enough, we've got an 8x return on that investment so far.
I am not even chasing money, whether at the level of how I spend my time, how I negotiate before I'm willing to get involved with something, or how I'm trying to allocate whatever investable capital I do have. The other aspect of investing, which I've talked about here and there on the podcast from time to time, mostly when guests come on whose companies I've made small investments in, is early-stage private-company, venture-capital-style investing.
There I do 2 things. One is just invest very small amounts of my own money. Two is that, since the acquisition of Turpentine by Andreessen Horowitz, I've also been able to become an a16z venture scout, which basically means they give me a not huge amount of money, and I can write relatively modest investment checks into very early-stage companies.
The way I think about that is basically that I want to invest in things that I want to see exist and that I want to see succeed. When I'm writing checks with my own money, I don't really think about return at all. I'm writing very small checks, so it's not something where I'm thinking, “Even if some of these companies do extremely well, that's not why I'm doing it.”
Companies I've invested in include Elicit, because I really respected its commitment and its philosophy of highly structured reasoning—the idea that we can't just allow the black-box models to do everything and hope for the best. We need to really take it apart. We need a systematic approach to structuring their reasoning and also to ensuring the reliability of that reasoning. I thought that was great, so I invested a few thousand dollars there.
Goodfire, because I'm really into interpretability. the AI underwriting company, because I think the flywheel that they're trying to create in terms of harnessing the power of the insurance markets, creating these standards, creating audits, and ultimately trying to bootstrap an insurance market is really valuable. That way, we can start to price the risk associated with various kinds of AI systems.
I think these are all worthy projects that, if they were nonprofits, I might be inclined to make a donation to. But since they are private companies that I can buy equity in instead of making a donation to, I'll do that. I'm really just doing that to try to support those projects and to be on the team.
What I'm trying to do there is identify things that are both safety-promoting and have fast-growth opportunities. I do think that that's a pretty decent intersection point. One way that I've heard people describe this is: if AI is going to become the biggest market in the world, if it's going to start to compete with human labor broadly—which it certainly seems like it's on track to do—then the second-biggest market in the world is going to have to be AI assurance tech.
How do we make sure that this stuff is actually working the way that we want it to work? How do we control it? How do we quality-control it? I think there are quite a few things at that intersection of safety-promoting, reliability-promoting, control-promoting, and so on that also have the potential to be quite fast-growing. Those are the things that I'm inclined to invest in with my a16z scout fund money.
I think there's actually a lot more room for alliance between the AI safety community, or at least a lot of the AI safety community, and the a16z accelerationist worldview than is commonly thought. I've talked about this many times as well, but I think that's mostly because the AI safety community is often caricatured as being anti-progress, anti-technology, whatever. Honestly, that's in my experience almost entirely wrong.
The people that I know—and I know a lot of them—who are focused on AI safety issues are generally very pro-progress and very pro-technology. They're generally lifelong techno-optimist libertarians who see AI as a different kind of thing because it does have this potential to outcompete us at what we've been uniquely good at, which has allowed us to take over the world. Because of that special dynamic, based on very specific arguments and analysis of this particular phenomenon, they see AI as being different from everything else.
The AI safety people are very much in favor of permitting reform and want to see more housing get built. They're generally all for abundance and want their AI doctors, and they're very well aware that human doctors are not as great as we might wish they were. They want their self-driving cars. They're very analytical when they see that self-driving cars have a 10% accident rate compared to human drivers, and then further see that almost all those accidents are caused by other human drivers surrounding the AI drivers.
They believe the numbers. They update on these statistics. So I would say that, on a very large range of questions, there is a lot of opportunity for the AI safety-focused people and the a16z worldview to come together.
I hope that, by getting a16z invested in token fashion—because my checks will be very small and certainly not the kind of thing that's going to make or break a16z economics—I can start to send a little bit of a signal to the firm more broadly: Look, there are a lot of things where progress can be enabled, assisted, even accelerated by this sort of assurance tech. These businesses can grow fast, and the people who are starting these businesses are ambitious, want to grow fast, and want to see a brighter future for everybody.
Hopefully, I can send a small signal that might have bigger ripples over time. Where there's an opportunity to support something that I believe in, I'll do that either with my own money, in which case I have basically no concern about return, or in the a16z scout case, looking for things that have potential for high return—but doing that in a sector or with a concept that I think could start to facilitate this coming together of AI safety folks and accelerationists.
Thank you to Aaron for a couple of good questions. Here's another interesting question on early-childhood AI literacy.
“Mainstream advice on kids and tech often ends with a push for abstinence. What would a sex-ed-style approach to technology and AI look like for kids that's age-appropriate, values-based, and practical, so they can build confidence and judgment instead of secrecy and bad habits? What could parents, particularly as models for their kids, and schools, as creators of safe spaces to explore, do at various ages—say, 3 to 6, 7 to 11, and 12 to 18—to foster such healthy learning and exploration?”
That is a truly great question, and I don't think I have an answer that's up to the scope of the challenge of that question. Eugenia Kuyda, founder of Replika and now Wabi, on our live show basically did advocate for abstinence for younger kids. She said, “We just don't know enough about this technology to trust any developers, no matter how well-intentioned they are, to serve young kids in a way that we're ultimately going to be happy with, that we'll ultimately feel like we were wise to have done.”
Given her experience in the Replika space, I am reluctant to disagree with her. I guess I would say that, especially for younger kids right now—and my kids are: Ernie is 6, almost 7; Teddy just turned 5; and our youngest, Charlie, is 2 and will be 3 in April—I'm still in kind of the first-grade-and-below bracket. I suppose abstinence might be a decent idea there. And yet I'm not content with that.
When I look at what folks like Alpha School are doing with 2 hours of entirely AI-delivered education—2 hours of educational content, 2 hours of focused academic work, entirely delivered and supervised by AI on a one-to-one basis for every kid—the fact that they're able to get kids going faster than usual schools do in just those 2 hours a day and create this whole afternoon of freedom to do all these other exciting things that kids want to do, I'm not content with the idea of abstinence being the answer.
I think that puts us in a hard place. I think we're at a spot where it's right to say that the technology is too new. There are too many surprises about it, and nobody has really established themselves with a great track record for how to create AI products and experiences that really serve kids well in the long term and don't just maximize engagement or whatever else.
I think that’s true, but I also don’t want to put my head in the sand or try to get my kids to put their heads in the sand and pretend that AI doesn’t exist for all that much longer. I do think there’s going to be a lot of value in AI for educational purposes, and probably more besides that, too. We just haven’t seen the right form factors.
I do use voice-mode AIs with my kids fairly often. It’s not something we do all the time, but if we’re playing a game or there’s some question they have that I don’t know the answer to, I’ll totally whip out my phone, go into voice mode, and ask it a question to get the answer. This is something that seems pretty normal to them.
It’s not something we do a ton, but it’s also not something I’m trying to hide from them. They’re not clamoring for it all the time, but occasionally one of them will say, “Hey, why don’t you ask AI about that?” And so I will.
I guess I put this in the same category as other major, fundamentally not just life-altering but sort of civilization-altering technologies that we’re going to be confronted with in the coming years. I would put brain-computer interfaces as another one of those. Neuralink is planning to scale up its patient base substantially this year, and they’ve said that they plan to serve what would be considered healthy people in the not-too-distant future.
We’re going to have questions around, do we get brain-computer interfaces as healthy people? We’re going to have questions around all sorts of gene editing: Do we do that? My kids are already born, so they would be getting whatever gene editing they’d be getting as formed people. But we’re also going to have increasing levels of power in terms of embryo selection or embryo-level gene editing that’s going to fundamentally change the nature of people before they’re even gestated and born.
I think surrounding ourselves with AI friends, companions, tutors, and always-on entities is probably right up there in terms of the magnitude of the impact that it could have. We’re going to have to approach it with extreme caution, but the value of it is probably also going to be undeniable.
It’s very hard for me to imagine abstinence all the way up to 18 years old. The idea that a high school kid today should not be allowed to use AI because of the downside risks—I don’t see that. I definitely think a high school kid should be able to use ChatGPT and should be able to use Claude.
Should they use AI boyfriends and girlfriends? I would certainly understand a parent saying, “No, I don’t want that,” and that’s probably where my intuition would go as well. But I don’t know. This stuff is tough. I haven’t parented a teenager, and I do think the trade-offs there are tough.
If they’re going to have a phone, they’re going to have some access to this kind of stuff. If you tell them it’s not okay to use, do you drive it underground? Do you lose visibility into it? It’s very tough. I think this stuff is very fraught, so it’s by all means a great question to be asking. I wish I had better answers.
Getting hands-on is always one of my fallback answers. If you’re considering buying any of these form factors or allowing your kids to use any of these things—whether it’s a stuffed animal that can really talk, or an app that’s a virtual friend, or whatever—I would definitely get hands-on with it yourself. Really try to understand it and get an intuitive, experiential feel for it before you just give it to your kid, let them do whatever they’re going to do with it, and hope for the best.
I just don’t think there are great answers right now. This is definitely one of the areas where consumer reviews will be extremely important. It would make a huge difference to me to know that lots of other parents are out there saying, “This thing has been great for my kid.” I would take that quite seriously.
I believe it is possible. I guess maybe my closing thought on this is that, because the space of possibility with AI is so vast, I absolutely believe it’s possible to create AI toys, products, virtual friends, or whatever that effectively nurture humans at any age. We’re just going to have to continue to watch the space really closely, be hands-on, and work together. Hopefully, that way we’ll be able to come up with good answers.
It’s a tough one. Next question: What is your timeline for work disruption? Could you comment on job disruption in terms of phases? If 0 to 3 years hits roles like customer support, marketing operations, and some software engineering tasks, and years 3 to 10 starts reshaping accounting, law, and parts of medicine, what’s your best guess for years 10 to 20? Where do humans remain essential at each stage?
First of all, if we’re going to use a timeline like that, I would say time equals 0 was somewhere around the introduction of the InstructGPT model—the first instruction-following model. Actually, they later revealed it was just supervised fine-tuning and not yet RLHF as of the first release of InstructGPT. That’s a very esoteric, in-the-weeds history, but that was early 2022. Then ChatGPT, of course, came in late 2022. I would say 2022 is kind of year 0 for this purpose.
Those early models—the first things they were able to do, and certainly the things I was most interested in getting them to do early on—were basic marketing copy. We started to see some disruption in 2022 and 2023. Copywriters do seem to have been very significantly impacted by AI.
We also had an example of this at Waymark with voice-over. Initially, our offering for voice-over was a professional service. We would charge $99, which is a pretty good price, and we worked with a provider that did a really good job. We delivered professional voice-over at medium scale for an SMB-accessible price point, and we were pretty proud and pleased of that service.
In the 2022–2023 timeframe, AI voices started to get good enough that they were really not competitive with the human professional—but, you know, classic disruptive technology. They were worse, but they were way cheaper, and they also happened to be way faster. With our AI voice-over integration, you could get multiple takes in seconds and at no additional cost, versus a much better product for $99 that would take a couple of days and maybe have a couple of rounds of back-and-forth.
We started to see substantial adoption of that pretty much immediately. It did start to eat into how much of the professional work we were seeing customers choose to pay for, even while the AI voices were clearly inferior.
Fast-forward to today, and we’re now obviously spoiled for choice when it comes to great voices. ElevenLabs is obviously great. Google also has amazing, very steerable, promptable voices that are awesome. There are lots of other companies doing great stuff besides that. Hume AI has a very emotionally competent voice that can also understand the emotion of the human.
For our purposes at Waymark, creating marketing video content, we don’t need to understand input human voices. But Hume is good at that, as well as making the voice sound emotionally intelligent. These days, we just don’t do much—I don’t think we do any—human professional voice-over work at all. That’s a pretty substantial change that I would date back to 2022 and 2023, with marketing copy and voice-over work being some of the earliest areas.
As of today, if you look at GDPval, the latest models are winning in software engineering by a substantial margin. It’s in the 70s, maybe even up to 80% of the time, when it comes to software engineering tasks. Again, GDPval sets up three groups of experts: one set defines the tasks in various domains, another set does the tasks, and their work is then compared by a third set of experts. The third set is responsible for determining which they prefer: the human work or the AI work. The AIs are now winning a lot when it comes to software engineering tasks.
I would say for sure that I prefer working with Claude Code over working with an entry-level human software developer. That’s pretty obvious to me, honestly, at this point. I’m not sure yet—I think the debate is ongoing—whether this has hit the statistics or not. But will it? I think it will.
I think in 2026 it’s going to be hard to justify on purely economic terms hiring a 22-year-old out of a computer science program versus spending more on Claude Code, or spending the time that you would spend mentoring that person on developing new skills and hooks and an increasingly elaborate architecture for your AI coding setup.
I think the disruption is happening now, certainly in software. Relative to the rest of the question, I don’t think it’s 3 to 10 years and then 10 to 20 years. I think it’s probably coming sooner than that when it comes to accounting, law, and parts of medicine.
Let’s start with medicine. I’ve talked ad nauseam now about how the latest models are competitive with attending oncologists. That’s a reality I’ve lived. There’s no denying it. Nobody can talk me out of it. As sure as I’m sitting here, all 3 of the frontier models—Gemini 3, Claude, and ChatGPT 5.2 Pro—are currently competitive with attending oncologists. There are just no 2 ways about that.
So, in terms of parts of medicine, I’d say that disruption is starting to happen. I don’t recommend that people ignore their human doctor or go all in on AI, but I do think that, especially for things that aren’t so important, the substitution effect is going to start to be real because you can get a lot of your questions answered and maybe not have to go to the doctor, or go to the right doctor the first time. Obviously, in our society right now, doctors are still required to prescribe the medicines. I do think that the way this is going to play out is going to be noticeably different very soon.
I’d say that’s probably going to be true in accounting and law as well. I’m certainly no expert in law, but when I get a contract to sign, I run it through 3 or even 4 frontier language models in exactly the same way that I run my son’s lab test results through language models, and I’m yet to regret it. Would I do that for the most important transaction of my life? I’d probably get a human lawyer to help review it as well. But for routine stuff, I’m absolutely happy with the results that I’m getting.
If none of the 3 or 4 frontier models identify issues that I need to be concerned with, then I’m good. I think it’s pretty unlikely that all 3 would miss something that’s really important to me. Even in accounting, we’re starting to see that models are getting pretty good with spreadsheets, so I think that, too, is coming pretty soon.
Another mental model I have for this is how elastic demand is in a given domain, and I think that varies dramatically. When it comes to something like accounting, I personally don’t want to buy a lot more accounting services than I’m required to buy. Maybe others are out there wishing that they could buy more accounting and just don’t have the budget for it, and when AI makes accounting professionals more productive, they’ll buy lots more accounting services and that will partially offset the effect. But I just don’t see that there’s that much latent demand for more accounting.
I think most people are doing roughly what they need to do and what they’re required to do. Beyond that, they don’t really have that much appetite for more. So, in something like accounting, when we see threshold effects hit and, all of a sudden, AI can do the job, I’d expect that to be a field where there would be more outright substitution and displacement, not an explosion of accounting services provided.
Dentistry, by the way, is maybe the most extreme version of this. I want zero dental services. I’d love to never have to go to the dentist again in my life. So, if there were something I could buy that cost 1% of what dentists cost and did the same service, I’d happily do that, and I wouldn’t be doing lots more dentistry. I think most people share that intuition.
Medicine in general is probably a bit different. Most people might want a little more medicine; they might want a little more care. There could be a surge of demand, and as things become more accessible, as capacity expands, demand probably grows to fill it.
Software engineering maybe could be the 10x or even the 100x. Do we get 100 times as much software production over the next few years? Very plausibly. That might be enough to sustain software employment, at least among senior engineers, for a while yet—for perhaps a surprisingly long time. Even when it looks like, “AI can code up this website in one shot, so what do we need engineers for?” if we’re doing 100 times as much software and the architectures are getting ever more elaborate, maybe there’s still a role for the senior software architect, even if there’s not one for the junior person.
I think this is one area where I do have a bone to pick with some of Dwarkesh’s analysis. I think very highly of Dwarkesh generally. His show is obviously great, his questions are generally very effective in terms of eliciting alpha from his guests, and I think his essays are generally really good. But one thing that he has said recently that I do pretty strongly disagree with is that, when the question is posed, “Why aren’t we seeing more impact than we have to date from AI on the labor market?” explanations that center on human bottlenecks are essentially cope.
He thinks it’s really that the models aren’t good enough and not that people are too stuck in their ways. Obviously, I would say it’s both. The models do have room for improvement—the jagged frontier and all that stuff is real, for sure—but I think there really is a lot of human bottleneck going on.
You see that at the hospital, right? When a resident comes to the room and talks to us and is clearly less knowledgeable and less reliable than a large language model, I don’t really know how to interpret that in any other way than that the humans are the bottlenecks. They haven’t realized it, nobody has told them, and they haven’t experimented with using a model on their own. That’s on them, right? It’s not on the model.
I can easily tell you about a bunch of times when, if the resident had engaged with ChatGPT as I had engaged with ChatGPT before coming to talk to me, they would know more and be able to do better. I’m quite confident that is true in a very wide range of contexts. I do think that the human bottleneck is a very real phenomenon—not the entire story—but to say that it’s cope, I would definitely debate Dwarkesh on that one.
The other thing that I would point people to, in terms of a mental model for this, is something I mentioned already: We ran that cross-post with Lucas Perry from the Future of Life Institute podcast, and he has this idea of the inverse pyramid model. Basically, the lowest rung in the hierarchy of an organization—in other words, the entry-level employees—is where AI is coming first.
Another way that he thinks about it is that anything where there are lots of people doing the same job, and the organization is geared toward trying to make sure that those individuals act more like cogs in the machine and do the job the same way every time, lends itself to AI. Consistency, reliability, process, standards, systematic evaluation—all that stuff really lends itself to AI.
He recommends, and I think this is pretty good advice, trying to do things that are n-of-1 and trying not to do things where you are one of n. Be n-of-1; don’t be one of n. That’s pretty good advice, I think. But obviously, that’s advice that some individuals can take. It’s not something that I think is going to preserve the structure of employment broadly.
We’re not even talking about driving here, right? That wasn’t mentioned in the question, but last I checked, I think it’s 4 million Americans out of about 150 million employed Americans. Something like 4 million are professional drivers. We’re getting very close to where human drivers are just not going to need to exist, and it seems like the disruption there could happen quite soon.
Again, it seems like the bottlenecks increasingly at this point are human. You’ve got some city councils trying to do various things, or what are the Teamsters going to say, or what have you. I’m in Detroit. Waymo is projected to launch in Detroit in 2026, and it’s going to be really interesting to see how that plays here in the Motor City, especially because the companies headquartered here aren’t really at the center of the frontier of self-driving technology, to put it mildly.
I’m not sure what the response is going to be, but the bottlenecks there are pretty clearly human at this point. Even in a world where AI doesn’t really get any better, as we figure out how to scaffold it, plug it in, manage the context, and handle all that kind of stuff, I do see a pretty serious disruption as at least being possible in just the next couple of years, unless it’s blocked by sociopolitical dynamics.
Superintelligence is a distinct question from that, but it seems like we’re quite clearly close to where AI will be able to compete for entry-level jobs—for jobs where an individual human is one of n people doing that job because there are standards, processes, and evaluation frameworks to make sure that AI is doing well. I expect AI, especially as we really put that elbow grease in, will do better than a lot of people in a lot of places, and that the economic and market pressures will be very strong to adopt it.
Customer service is another great example. Will there be some customer service people? I don’t think it’s going to zero in the immediate term. But my dad was on the phone with Bank of America yesterday, on hold for however long, getting increasingly agitated by the fact that he was on hold as he listened to their hold music and messages.
I was just like, “Man, I know several AI customer-service firms doing voice-agent-type work that would, at a minimum, be able to reduce the wait time dramatically and probably do just as good a job for a large majority of tickets as the humans are ultimately able to do.” So, I think we’re bottlenecked on willingness. I think we’re bottlenecked on implementation at the very high end. We’re still bottlenecked on model capability.
If I had to project how far up the pyramid AI can go today, it’s maybe halfway up vertically, but that’s, whatever, 80% of the mass.
I think the technology is basically there for that. Now we'll have to see how fast people actually—how fast the market works. How fast do market incentives and pressure, and increasingly lower barriers to successful deployment, all work to actually create the disruption that I honestly think is kind of inevitable? But relative to the question, I would take the under in terms of timelines.
The final point of the question is: Where do humans remain essential at each stage? I think that's a design question more than anything else. I think we want to design an AI future where humans are essential, or are at least a big part of the process, so that we can retain some authorship over the future and not just give the whole thing over to AIs, leading to the gradual disempowerment people worry we might experience. But I think that is going to have to be intentional, more so than something where we find that there's some ineffable human essence—some élan vital—that, for whatever mysterious, mythical, or mystical reason, only humans can do.
I really don't believe in that. I think that we are pretty remarkable, certainly in our breadth and our flexibility, and the high end of human achievement is obviously super impressive and inspiring to the rest of us, mere mortal humans. But AI is doing superhuman-level stuff in more and more domains, and I don't really see fundamental blockers to that being, in the fullness of time, functionally everything.
So I think it's much more on us to think about: How do we design these things? How do we design our overall society? How do we design our overall systems? How do we design our models? How do we design our implementations so that we keep the parts that we want to keep and retain some overall control of steering the future? I don't think that's going to just happen because there's something that AIs can never do, that only we can do. I think it's going to have to be a lot more intentional than that.
Okay, next question. Regarding nonprofits and shifting need as AI reshapes labor markets, how do you expect the definition and scale of those in need to change over, again, 0 to 3, 3 to 10, and 10 to 20 years? What should nonprofits and funders start doing now to prepare, especially around workforce transitions, mental health, and basic economic stability?
I would like to see a lot more work on this, honestly, than we have to date. My general answer is, I have to give Sam Altman a ton of credit here for making personal investments in UBI. I think that is something we are ultimately going to need. Whether it ends up being called UBI or exactly how it looks, I think we're going to have to have some new social contract that decouples a person's right to a decent material standard of living from their ability to contribute economically, especially in the context of the competition they're going to face from AI systems.
I just think that's really the only way that we're going to get to a place that anybody would be happy with. I welcome other suggestions, but I haven't heard many that really make a lot of sense. The only other thing I hear is either we're going to need a UBI, or we're not going to need a UBI because there's always going to be work for humans. I just don't believe that.
Tyler Cowen kind of confuses me on this point these days, to be honest. I've been waiting to invite him on the podcast for a while. Maybe I should finally get around to doing it. But I looked back at Marginal Revolution, and I think his first mention of ZMP workers, or zero-marginal-product workers, was from 2010.
This was basically research looking at how firms responded to the Great Recession. What many firms did in response to the Great Recession was look around their teams and say, “Who do we really not have to have to go forward and be okay?” They were able to identify people, and they were able to cut those people. Of course, that's a noisy process, but by and large, it seems like firms were able to cut headcount and then actually see somewhat of a surge in productivity, seemingly because they were producing more or less the same with fewer workers, because there were some people who just weren't really contributing much.
You take that, and you take the general phenomenon of jobs where a lot of people seem to feel, according to survey results, that even absent AI, their own work isn't really worth anything, that it's kind of performative, or that it wouldn't matter if it wasn't done. Maybe they're wrong about that. But are you really that confident that people who are saying their own work is meaningless or unnecessary are wrong? I tend to think that at least some of them are probably right.
So it feels like there's already, at baseline, some ZMP workers who are employed. There's a lot of people saying that their work is just not really meaningful or important, and that things would be fine if it wasn't done at all. I just see that we don't have a great answer other than that we have to find a way to decouple somebody's ability to exist and have a decent life from their ability to compete economically, again, especially to compete against AIs.
Exactly what the right structure of that is, I think we would do really well to start doing much more experimentation. I personally interpreted some of the recent UBI results as much more encouraging than I think the broader discourse did. I'm certainly not an expert in this, but my takeaway from some of the UBI research was that people were disappointed that people seemed to work less in response to getting the UBI.
I'm kind of like, I think that is the point. If they worked the same, in my mind, that might even be more of a failure. I think we might not want to quite yet—and maybe we didn't want to 2 or 3 years ago, when these studies were being run—to have people just take cash and not work. But I think what that does show is that people are, in many cases, just working for money.
Maybe you could also say that at least some people are relatively easily satisfied with not that much money, which would suggest that they are able to find meaning in things like spending time with family and leisure, and whatever. Maybe they don't feel compelled to go out into the workforce and do some job they really don't want to do just to get a bit more money. Obviously, marginal tax rates are also, in many cases, totally out of whack, and working more can mean you forfeit benefits of all kinds. It's a very complicated question.
But I'm encouraged, I think, much more than the average commenter is by these UBI results, because I see that it didn't take that much to actually get people to work less. If they're not enjoying their work and they're substituting away from that work toward leisure, to me, that suggests they're not missing out on meaning. They're not pining for the workplace as the place where they're going to have an identity and find meaning. It seems like they're finding that in other places just fine.
Thank you very much. I would expect that probably to continue. I tend to think that the idea that we need jobs for structure and meaning and whatever is mostly cope. I think it's especially unhelpful cope when it's projected by people who have privileged positions, where their work is high status and where they genuinely do love it. I count myself in that group, and I'm very thankful for that.
But when that reality is projected onto people who are lower on the socioeconomic scale, who are doing work that they don't want to do because they have to do it, because that's the only way they're going to feed and clothe their kids, I think that's quite counterproductive and misguided. So, bottom line, I would love to hear other ideas. I would love to read your utopian fiction about other ways that we rework the social contract that isn't just a vanilla UBI-type structure.
But until somebody more creative, imaginative, and visionary than me comes up with those ideas, I basically still think that UBI is the default, and denial is not helpful. More experimentation sooner about the details and the structures, and exactly how incentives should work, would be really valuable.
Okay, thank you for those questions. Next one: Are people being misled by benchmaxing? And then another related question: Does the massive success of pretraining give people the wrong intuition about AGI?
Yeah. Characterizing whether people are misled, or whether people have the right or wrong intuition, is very hard because people have very different intuitions and world models. There's such diversity there. I can't characterize people's opinions in general with any precision; it's just such a wide-ranging thing that anything I might say about some people's opinions is obviously going to be contradicted by other people's opinions.
That said, I do think benchmaxing is probably misleading, at least to some degree. I would go back to the Chinese models, which I talked about last time, being significantly worse on just a very random multimodal task than all of the leading Western models, as a kind of leading indicator of how much benchmaxing might be misleading us more generally. It does seem to me that the delta between the leading frontier American models and the Chinese models in benchmarks is much smaller than the difference between them in practical, day-to-day utility on whatever idiosyncratic and/or esoteric task you might want to use an AI to help with.
And, of course, the Llama 4 incident shows that as well, where they spun up maybe a bunch of different Llama 4 variants to put one in each different LM Arena category to try to maximize their score. I’m not exactly sure what they did, but they sort of LM Arena-maxed, and that was effective insofar as they got a high position on LM Arena. But you don’t hear all that much about Llama 4, and it seems like it just wasn’t really competitive, but they were able to make it look like it was competitive on some of these standardized scores.
So, yes, I think benchmaxing is a problem. Independent analysis is a good antidote for this. Looking at the METR charts, looking at Artificial Analysis, and looking at people at Scale, whose benchmark is largely private, are all pretty good ways to think about that as well.
I think ARC-AGI has been remarkably durable in terms of how relevant it’s been, but people can obviously benchmark-hack that too. Definitely, things where there’s a private test set, and things where people take it upon themselves to really analyze capabilities and make that their thing—really trying to earn your trust by being a reliable guide to model performance over time—I think those are the things that increasingly the field will be looking to, beyond just performance on open, standardized benchmark tests. Those continue to be somewhat relevant, for sure, but less and less all the time.
And then, on the question of whether pretraining gives people the wrong intuition about AGI: If anything, I think maybe post-training is giving people the wrong intuition about AGI, and maybe the classic Shoggoth meme for pretraining is a better intuition. This also connects to the breadth-first versus depth-first search of the space of possible AIs.
I think that if people are misled about AGI, or if people have the wrong intuitions about AGI, I would expect that it’s wrong in the sense that they’ve encountered a relatively narrow range of form factors for AIs, which are basically chatbots and coding assistants. Those designs have kind of converged thus far. The overall paradigm of helpful, honest, and harmless has been great for unlocking a ton of value and, certainly when done well, as Anthropic has done, has been pretty great in terms of shaping the model’s character. I have no beef with those prompts or with those approaches.
But if there is a problem with that, it has presented the public with a very narrow slice of the overall conceptual menu of what AI can be like. It maybe is lulling people into a false sense of security, or a false sense that this will continue to be normal. I think that, in reality, the Shoggoth meme—that this thing is an insane alien that kind of encompasses everything in a super-strange way, and can shapeshift and be anything you want it to be, or maybe even things you don’t want it to be—is perhaps the better intuition for at least the space of possibility for AGI.
When you think about things like the—I think back to the episode we did with Apollo, with Marius Hobbhahn from Apollo Research, where they got access to chain-of-thought from o3-class models and found that the chain of thought was kind of evolving to become its own dialect. Remember, like, “disclaim, vantage,” and “the watchers”—these strange phrases that really didn’t look anything like the training data, in this case being driven by reinforcement learning at increasing scale.
It just feels to me like there’s so much more alienness to these systems than we are seeing. I guess there are just a couple of other intuitions for this that are maybe worth mentioning.
People often cite the bird-airplane analogy. Many people have commented, of course, that we wanted to create a machine that could fly, and a lot of early attempts tried to mimic the bird. What we have is an airplane that flies on quite different principles and is way faster, way more powerful, and can carry way heavier loads than birds can. It can’t necessarily do everything that a bird can do, but for the things that we’ve designed it for, it’s way, way better. We wouldn’t want a scaled-up bird to take a cross-country flight on. Airplanes are just way better than scaled-up birds.
Similarly, many people may have seen recently a video that was kind of a compare-and-contrast. There was a humanoid robot harvesting grain in a field, chopping it down and bundling it up in a way that humans traditionally used to do. Then there was this carrot-picking machine that was just rolling through a field, pulling carrots out of the ground en masse, washing them off, and operating at orders of magnitude faster than humans could possibly pick carrots.
I think something similar maybe happens with a true AGI or a superintelligence, where it becomes so powerful that it potentially doesn’t even really make sense for it to present in natural language anymore. Or, at a minimum, that becomes a dramatic reduction of some sort—a majorly lossy summary of what it’s doing, as opposed to the core thing that it’s doing today.
Today, the response that you get from the chatbot is its output. In the future, I think the natural-language summary of what it’s doing may be just a very small part of what it is actually doing. So, yeah, I think study the Shoggoth, study the weird chains of thought, and expand your mind when it comes to the space of possibility and how weird things could be. Those are my guesses for how most people are being misled, if at all, today.
Okay, next question: Is learning from a physical environment a requisite for AI?
I don’t think so. I think we’re pretty far along in AI at this point, and when we look at all the things that AIs can do and how many categories they’re winning more than 50% of in the GDPval context—all without any robotics or physical embodiment at this stage—I think it’s pretty clear that you can get pretty good AI without needing to learn from a physical environment.
The line maybe starts to blur a little bit when it’s, “Can you use a computer? Is that a physical environment?” It’s a digital environment, but it’s spatially organized. It does seem like the way that we are getting AIs to learn how to use a computer is by them actually trying and failing to use the computer a lot. We needed some language-model-based capability to conceptually understand what one might want to do on a computer, what different buttons might mean, and so on. From there, it’s been a lot of actual reinforcement learning, where once there was at least some ability to succeed, you could hill-climb from there up to pretty decent performance at this point.
I definitely think we’re not quite there, but in 2026, I think we’ll definitely have AIs that use computers basically as well as, if not better than, your typical human user. So, yeah, I guess maybe it kind of comes down to a task-by-task thing.
Do I think that large language models are going to become plumbers without a bunch of reinforcement learning on physical plumbing tasks? No. I think that if you want to be a plumber, you’re going to have to do a bunch of stuff in the physical world. You’re going to have to get good at that, and that is probably going to require some sort of embodiment.
When the question says “physical environment,” I would also say simulation is going to get really good. Obviously, NVIDIA’s GPUs are not just good for training; they’re also good for simulation. So we’re going to see more and more simulation being used in all facets of AI training, including for robotics. But if you count that simulated physical environment as a physical environment, I do think you need it to do physical things.
Ultimately, a language model can’t be a plumber. You could, I think, be a lawyer without ever having any physical embodiment. I think you could be an accountant. Arguably, your Excel spreadsheet maybe is your physical, quote-unquote, environment there.
Basically, my intuition is that you need training in something like the environment that you’re going to operate in. As long as you’re just operating in language space, language is probably enough. If you want to start operating in pixel space, you’re going to need to be trained in pixel space. If you want to operate in spreadsheet space, you’re going to need to be trained in spreadsheet space. If you want to operate in physical, real-world space, you’re going to need to be trained in a combination, probably, of simulated physical, real-world space and, to some degree, actual real-world space.
And it’s going to happen. The other way to think about these kinds of questions is: we’ll never know. Is it required? I think not for many tasks, but it’s going to happen.
There may be some positive transfer. Will the AIs of 2027 or 2028—even if I’m just doing something that is ultimately a purely language-based legal task—also have some sort of physical, real-world training as part of their overall training mix? It wouldn’t surprise me at all. It wouldn’t surprise me if there’s some positive transfer there, if there’s just some better intuition, or if certain kinds of queries that require spatial reasoning perform better because that kind of training is folded into the mix.
I’m not saying that AIs in 2027 or 2028, even in the chatbot or digital-assistant variety, won’t have physical-environment-based training going into them. I’m just saying I don’t think you would have to do that to be successful. But it’s pretty likely, at least, that it will happen.
It’ll be kind of hard in the end to tease out: Was this required to happen? Could we have gotten here another way? It just seems like everything is going to be developed and folded in to the degree that it works, and it seems like everything is working. It’s all going to be folded into the mainline frontier models, and so we will probably not have a clear answer about whether we could have gotten here without doing any physical training.
My guess—my bet—would be that, counterfactually, yes, we could have. If it adds even a little bit, then it’ll be part of the overall mix. What we’ll actually have will be AIs that are trained on everything, regardless of whether some of those things could have been cut with minimal performance loss or not.
Okay, the next section gets a little more focused on tooling and AI engineering concepts. Are you noticing any emerging standards or winners for tooling across the companies that you work with?
This is interesting. I don’t really have anything super shocking to report. In terms of models that people are using, Claude Code is typically still the go-to. People certainly give very good reviews to GPT Codex, but Claude still seems to be the go-to for most people.
If they didn’t have Claude, there are certainly other great options out there. Gemini 3 is also excellent, but I’d say Claude remains the consensus top choice for coding. OpenAI is pretty good at everything and is probably the default thing that most people use for random queries in a browser on a day-to-day basis. Gemini Flash is, I would say, pretty clearly the top choice for things that don’t require frontier capabilities, where speed and cost are significant factors.
I don’t think that’s really surprising. We’re not out of step with mainstream online consensus at all. I really haven’t seen anything that challenges the mainline narratives, certainly at a model-provider level.
At the tool level, I don’t know. Again, I think the market is getting away from most people. Because things are changing so fast, most people are behind. Most people are not using the best thing, which is maybe normal in life in general.
For example, I’ve used LangChain on a couple of projects recently: one at Waymark and one at another company. I’d say it works pretty well. You can build agents in it, have those agents hosted on their infrastructure, and log traces to it. It’s a pretty heavyweight UI with a lot of features, so it can be a bit of an overwhelming data presentation for many people at first. It does take a little getting used to, but I’d say it works pretty well.
Until such time as I’m not happy with it, or I hear from somebody else that there’s something dramatically better, I’m pretty content with it. I think a lot of people are in a similar spot where it’s like, “Man, I can barely keep up with model releases.” Tool releases feel like a second-tier question: As long as I’m meeting the need that I originally had, I’m probably fine.
Even though there might be better things out there, I think it’s just so overwhelming to try to shop effectively that people are staying the course with whatever decision they made for quite a while.
Another interesting direction that we’re going to start to see develop—and already are starting to see develop—is that the model providers are trying to become platforms, and they’re building out more and more of this stuff themselves. OpenAI has its agent-builder-type tools now, which compete directly with many agent builders that were built on its platform. They also have observability of various kinds.
Anthropic bought Humanloop, which was a past guest, and I was a happy Humanloop customer for a while, but now they’re part of Anthropic. It’s going to be interesting to see. Of course, Google is going to build everything over time, probably to somewhat varying degrees of quality, but they’re going to have, in the end, the most robust portfolio of products of probably any of the top-tier model providers.
The dynamics will be interesting: How often does it make sense to just pick a platform, go with its model, go with its observability, and go with whatever tooling and monitoring it has, versus trying to maintain flexibility and the ability to upgrade a model quickly on a new release? The latter would potentially require you to have a more horizontal layer for these observability questions and all sorts of tooling questions. It’s going to be interesting to see how that develops.
If I’m a big enterprise, I want to avoid lock-in, and I want to invest potentially more than I really have to in some of these things so that I at least have a little more ability to control my own destiny. I don’t want to be totally beholden to one of the platforms.
If I’m a startup, an individual, or an individual just doing a one-off project, then maybe I just take the convenience and use OpenAI’s observability or Anthropic’s observability because I’m already using its model, and it all integrates to make things easier and faster.
But all of this is to say, I don’t think there are any super-obvious major trends that I’m seeing. I would even say that, just listening to the Latent Space podcast, I take quite a bit of pride in the fact that The Cognitive Revolution has been voted twice at the Latent Space AI Engineer World’s Fair events as the number-3 podcast for AI engineers. Latent Space has been voted number 1 in both of those surveys.
In listening to them, it also does feel like the sort of content mix has trended away from these tooling questions. Of course, tooling is still an important element, right? You need to be able to look at traces, but is it really where people’s mental energy is going right now? It feels like it’s less topical than it used to be. I think you can even see that in their mix of guests.
One other anecdote I’ll give on this is from when I was vibe-coding the Christmas-present apps that I’ve mentioned a couple of times. The one that I was coding for my mom—the custom trip planner—involves AI research into all these various things.
There’s a prompt to go find restaurants, and she’s gluten-free, so we’re really trying to dig into gluten-free options in all these different places. When it comes to where she wants to stay, she loves to have a nice view and wants a balcony, so the prompts are customized to that level as well.
Early on in the development, and even still now as I continue to work with her and try to enhance it in ways that make it more valuable for her, I wanted to look at the traces. What are the inputs and outputs? What prompt is the model seeing, and what is it giving back in raw form? I wanted to debug that—to understand what was working and not working at a level below the UI that she uses as a user.
What did I do? I just added a trace function by prompting Claude Code: “I want to see the history of all the queries that are made to you. Can you add a tab to this application that gives me direct access to the full history?” It just built that thing, I think, in maybe 1 prompt. Maybe it was 2 prompts.
Now, right alongside all the core features of the app, I have a debugging tab where I can look at the history of the prompts. I think that’s another factor that’s challenging. This is maybe the future of SaaS writ small, or the future of SaaS in a nutshell.
What would I have done in the past if I didn’t have Claude Code to code that kind of thing? Maybe I would have made a free-trial account on Humanloop, or wired it up to LangChain, or whatever. But I didn’t do any of that. I just said, “Hey, Claude, log all these things and give me a tab where I can see them.”
That’s been perfectly good. Could it be more elaborate? Yeah. Could it be more full-featured? Sure. But it meets my needs for the development of this particular app. It took me less time to prompt Claude Code to do it than it would have taken me to find some solution, figure out what it was, figure out how to connect to it, and so on.
I suspect that may also be part of why people are talking about this stuff less: It’s become a lot easier to meet the basic needs in some cases just by coding it from scratch with a couple of prompts. I’d love to hear more about this from other people.
Okay, next question. This one’s kind of an interesting one. It wasn’t actually part of the AMA, but it was a question that I was asked.
I had a little interaction. I just posted—actually, this is in the context of vibe coding another app. The app that I vibe-coded for my dad for Christmas was a stock-trading strategy backtester. He has these ideas: What if we did this? What if every week I bought the stocks that lost the most the previous week and tried to catch them on the rebound? Is that going to work or not?
The app I created for him allows him to articulate a strategy in natural language, translate that into trading rules, and then go back and look at a time interval to see how that trading strategy would have performed over that period by systematically executing those trading rules. Pretty cool. He hasn't used it all that much, to be honest.
In the course of doing that and testing it, I was testing the strategy of buying the biggest losers and then trying to catch them on the rebound. I tried it once on an annual basis: What if I bought the biggest losers from the previous year, held them for the following year, and did that every year? Would I beat the S&P, or would I fall short of the S&P?
I was sure there was a bug when, in 2022, I think it was, Nvidia was one of the biggest losers. It was down about 50% on the year, and I thought, “Okay, something's clearly wrong with that. No way Nvidia would have been one of the biggest losers for a whole year.” It turns out it was, and so the AI was right. I asked the question, and Claude Code went and verified online that, indeed, Nvidia was one of the biggest losers of that year. The app was working correctly; I was just wrong in my assumption.
This was another one of these moments where I noticed the decline of sycophancy, because it would have been very easy for the model to say, “You're absolutely right. Nvidia has been a killer stock. There must be a problem in the code.” It did not do that. It came back and said, “No, you are wrong. In 2022, Nvidia was one of the biggest losers, and the app seems to be performing correctly.”
I posted that online: “Here's a crazy stat: In 2022, Nvidia was one of the biggest loser stocks in the market for the entire year.” Holly Elmore, executive director of PAI, came by with what I would call a critical comment, as she often does. She basically said, “This isn't SportsCenter.” I was morally out of line for engaging in something fun, trying to make myself look clever, or showing insights by noticing these quirky things in the AI space, because the whole thing is bad and ought to be condemned, and I ought to be condemning it. Doing anything else, in her view, was morally reprehensible.
I thought, “Okay, do you really think this is an effective way to advocate for your position?” Everybody who listens to this feed—and if you're 2 hours into this episode with me knows that I'm seriously concerned about AI safety issues. I do not want to see a race to recursive self-improvement. I think the big decisions that are going to be made in the next 1 or 2 years around the automation of AI R&D are very big and important questions, and I don't think we're ready to cross some of those thresholds or Rubicons, if you will.
I said to her, “Look, I think I'm much more sympathetic to your cause than most people. I did sign the Ban Superintelligence statement 3 or 4 months ago, for example, as one tangible indicator that I'm much more sympathetic to your cause than most. But do you really think that coming after me over a tweet about some random observation about Nvidia stock is the way to advance your cause?”
Then she asked me, “How did my comment make you feel?” I thought about that a decent amount and decided to address it here in the AMA, even though that's not what she was asking for.
My bottom-line sense of this sort of thing goes back to a mantra that I used to say a lot more often: We should really try to avoid psychologizing other people's AI takes and focus as much as we can on the object-level facts of what is actually happening and what that implies. We should take people's statements and positions at face value.
There is so much disagreement in the space and so much uncertainty. The AI 2025 forecast results—I got in the top 5% with predictions that I don't think look all that accurate. Of course, famously, we've got Turing Award winners who have extremely different positions on how dangerous AI is going to be, and I think those positions are all genuinely held.
The way that being sideswiped, attacked, or accused of moral corruption online because I made one random comment about Nvidia's stock performance from a couple of years ago made me feel was kind of indignant and confused, and a bit averse to the person saying it and the cause itself.
This is a cause that I genuinely care about. I'm not advocating for exactly a pause on AI—I don't even know what that means. I don't think we should shut it all down. I do think there is tremendous upside, and I do think we should be very careful. We should be willing to slow down, if not pause, at some point when we hit levels that we really might not be able to control.
But I also think the progress has clearly been much more beneficial than harmful so far. I've lived that in recent months. So I felt alienated from the AI safety movement, or at least the more, let's say, strident or shrill voices in it, by that comment.
My takeaway from that is—I wouldn't go so far as to say we shouldn't shame people, but I think we should shame people very carefully, very selectively, and only for what they are doing. I think it is pretty defensible at this point to shame xAI for some of the things that they have released in Grok. I think it is appropriate to shame xAI for having Grok on Twitter undress women with seemingly no guardrails in place. That is worthy of shaming, I think.
But that's an action. I would not shame people for their positions, and I would not assume bad faith. I know that not everybody's positions can be taken fully at face value, but if somebody is going to be out there engaging in the discourse, projecting your sense of their psychology onto them and arguing from that basis ends up with people, more often than not, feeling bad and becoming increasingly bitter toward each other.
It leads to more calcification or hardening of factions, and more of a sense that somebody is my ally and somebody is my enemy. I don't want any of that. I think the healthiest discourse we can have around AI assumes that everybody is trying their best, ideally gives people the space to genuinely be trying their best, and avoids these sideways accusations, moralizing, or psychologizing unless people are really doing things where you're like, “You are undeniably wrong,” and maybe even have a pattern of being undeniably wrong, in which case, sure.
So I would say to Holly: “Your question did not make me feel very good. It made me feel alienated from you and your cause, and I wouldn't recommend doing that. If you want to go protest outside xAI, I would support you in doing that. Choose your targets a little more carefully. Try to recruit people like me to your side as allies, and shame people selectively for things they've really done wrong, where they are responsible. I think that's fine, but don't just start using these tactics everywhere, because it's just coarsening the discourse and making everything a little more bitter and contentious than it needs to be.”
I don't think people do their best reasoning that way. Certainly, I don't think I do.
These are the only 3 AI questions that I thought made the cut, and I probably generated well over 100 across 3 models. So definitely give this one to the humans in terms of the questions I wanted to answer.
But here are 3 questions from AI to wrap us up. First: You're in Michigan, not San Francisco, London, or Washington, D.C. Does that distance help you or hurt you? What do you miss by not being “in the room”?
I would say this is definitely a real issue. It definitely hurts to be outside of these core hubs. Obviously, San Francisco is far and away number 1, and London is like 2 and maybe as close. San Francisco and London are both the hubs. D.C. is increasingly becoming a hub because there is so much policy going on there, although it's a very different hub from the San Francisco and London hubs. I'm not even sure it really belongs on that list, but that was the way the AI phrased the question.
Being outside of San Francisco and London does make it harder to stay up to date, be in the loop, and have the zeitgeist than it is if you're in those places. There is definitely a lot of stuff happening in person in San Francisco, with events and hackathons going on all the time. To some degree, secrets are being traded or spilled across frontier-model developers at the proverbial parties.
The fact that there is an emerging trend of rooms at San Francisco house parties where there is a “no AI talk” room just goes to show how much AI talk is going on. The AI talk there is definitely way more sophisticated than it is anywhere else.
That's my honest sense of the reality, and I'm able to compensate for it pretty well by being hyper-online. Spending a ton of time on Twitter is still part of how I keep up to date, no doubt about that. The podcast itself is also really helpful, because I get to have substantive conversations with very plugged-in people who are in those rooms much more often than I am.
Occasionally, I do try to go to events. Having gone to, for example, the Curve each of the last 2 years, or last year going to the Summit on Existential Security, these are also gathering places where you can get a very concentrated dose of exposure to the leading thought. I think it really does help to be there, so I feel like I’m missing out on that to some extent. But due in large part to the podcast, I’m able to compensate for it to a significant degree.
Being outside of those hubs, I do think you have to be much more intentional about how you’re going to compensate. It probably ends up meaning that, in the Bay Area, you could be much less online and still be equally or even more plugged in. But if you’re not in those places, online is really the main place to get it. I would also definitely make occasional trips to those places to be in the room, because I do think that’s a really valuable way to learn and make sure that you stay up to speed.
You look at the results of the AI forecasting survey from last year, and you’ve got Ryan Greenblatt at number 2 and Ajeya Cotra at number 3 on the leaderboard. That is not an accident, right? Those people are extremely well-informed, and it’s because of the social context that they find themselves in. In fact, Ajeya said, “My method for the survey was talking to Ryan and then getting a few more things wrong than needed.” The leading thought leaders do know each other, they communicate a lot, and that environment really does help to spend at least a little time in.
Okay, next question. As a Survival and Flourishing Fund recommender, you see the landscape of safety organizations up close. What’s underfunded that shouldn’t be?
Again, that’s an AI question. My big answer here is neglected approaches. I’m wearing my AE Studio AC/DC-themed swag hat, and I do think the neglected approaches approach, as articulated by AE Studio, is a great answer—a great meta-level answer—to this question. As a reminder, they did a survey of people in the AI safety field asking, “Do you think we have all the ideas that we need to be successful in AI safety in the big picture?” The answer was no. We’re going to need more ideas.
Clearly, the community thinks that we need more ideas. The community thinks that we do not have all the answers that we’re ultimately going to need. I think that things like what Janus does, in terms of being deeply engaged with language models and really trying to understand their characters and tendencies, are really good. What Eliot has done similarly with model welfare tests is really interesting as well. I think AE Studio’s own work in terms of self-other overlap is something that I come back to all the time.
What Emmett Shear and the team at Softmax are doing are all these sorts of things where we’re asking: Can we find creative ways to either design or train, or somehow get into equilibrium with, AI systems in ways that feel more stable? Ways that feel like they could be the beginning of some sort of stable equilibrium? I think those ideas are dramatically underdeveloped. They are often pre-paradigmatic. They are often developed by kind of weird people, some of whom I think would wear that label with pride.
They sometimes intersect with non-scientific or “woo” sorts of ideas, or ideas about AI consciousness, which are obviously very hard to prove and don’t feel intuitive to many people. But I think we should have more of all those sorts of things. My call to action there is: If you have a weird idea that you’ve never heard anybody else talk about, I absolutely think it is worth trying to develop that idea. Most of the time, it’s not going to go anywhere. Certainly, most of my idle shower thoughts do not turn into anything great.
But the field collectively believes we need more ideas. Where are those ideas going to come from? At least some of them are probably going to come from people from other fields, people with very unusual or idiosyncratic ways of looking at the world, and people who interact with AI systems in very unique and particular ways. They may find inspiration in biological systems that they can map onto AI systems in ways that other people aren’t thinking about. I think all of that stuff is dramatically underdone.
Arguably, the whole AI safety landscape is underfunded. I’d love to see more resources go into interpretability. I would love there to be not just one Goodfire. I know there are a couple of other organizations— for-profit companies—that are working on interpretability-type stuff, but I would love there to be significantly more work going into interpretability. If we scaled that up by an order of magnitude, I think that would be great.
Some of the stuff that Redwood Research is doing—where you take the assumption that models are going to be out to get us and then try to figure out how we can work with them—I think is also really underfunded. For me, I admire that work so much because it feels so hard. It feels so depressing to work under that paradigm and try to make it work. But Lord knows that, in many scenarios, that kind of work could be the thing that saves us.
I think a lot of things should be scaled up. Probably the thing where I’m like, “There’s probably enough going on,” is what the frontier companies are doing, which seems to be trying to get the current model aligned enough to supervise the training of the next model through any number of things—data filtering, RLHF-type techniques, and all that kind of stuff. Anything that is in the recursive self-improvement realm, I would be a little less inclined to write a new check for, because it seems like that’s what the companies are doing. If anything, I think they’re probably going too fast at that relative to all the other things that we could be pushing on.
But for me, the farther out you get, the weirder you get. There are obviously going to be many more misses than hits, but those hits could be really valuable. When I see something like self-other overlap, I’m like, “Yes, this feels like something that is just so underdone and has so much potential.” I would love to see more people of all kinds of idiosyncratic persuasions trying to develop those kinds of ideas.
Okay, last question. Turpentine got acquired by a16z. You mentioned that you negotiated for editorial independence. What did that negotiation actually look like? How did it go?
Again, that’s an AI question. I think that probably came out of ChatGPT, or possibly Claude, but I think it was ChatGPT because I did use multiple AIs to review the contract as we went through the process. So it was very well aware of that negotiation.
Honestly, I have to say it was a credit to Eric and, I guess, a16z more broadly. I’m not sure who all was involved; Eric has been my main contact person for all of this. It was honestly very smooth and pretty much entirely painless. I already had an earlier agreement with Eric that included an editorial-independence clause in my agreement with Turpentine, so that was a good starting point.
I was a little concerned, to be totally candid, when the deal was made and Eric called me on a weekend and said, “Hey, I’ve got an update for you. We’re going to be joining a16z.” I said, “Oh, that’s interesting. Marc Andreessen blocked me on Twitter a long time ago, before I ever even interacted with him on Twitter.” I think he famously blocks people en masse who just like a tweet that he doesn’t like, so I was probably mass-blocked along with many other people.
I said, “Hey, Eric, the dude blocked me on Twitter. I’ve never even met him or talked to him.” That does concern me. Of course, in the Techno-Optimist Manifesto, he did have an enemies list, which is not something I generally think people should be doing. I would not recommend publishing enemies lists for almost anyone.
I said, “I’m a little concerned because it’s pretty clear to me that some of the things that I think are important, value, and want to advance in this world are on the enemies list for a16z. I really do want to make sure we reaffirm the editorial independence that I already had codified, but I really want to make sure that it’s solidified going forward.”
There was no problem. We worked through a couple of turns on the agreement, and it was pretty smooth sailing. Basically, I requested something very reasonable, and they agreed to it with no real substantive pushback—just a little tweak to the wording here and there, that kind of thing. Overall, it was pretty smooth sailing.
Where it landed is that I have an explicitly written, contractually agreed-upon right on this feed to say whatever I want. That includes criticizing a16z, criticizing partners, criticizing portfolio companies, and disagreeing with their stance on policy questions. Basically, I can say whatever I want, and it’s totally fine for it to contradict their policy positions or to say that some of their investments are dumb, or whatever. I can say anything that I want, and that is positively affirmed in the contract.
I give them a lot of credit for being willing to do that. The only kind of pressure-release valve, or off-ramp, in the contract is that if, at some point, for whatever reason, a16z decides that they just don’t want to be affiliated with me anymore because I say, advocate for, or do something that’s too much at odds with the agenda they’re trying to advance, then they can release all of their interest in the intellectual property of the podcast to me. That would include the feeds and the logos, which are kind of jointly owned right now.
I have the domain and control the website, and they have the YouTube feed and whatever. So we’re mutually dependent at the moment, but if they ever wanted to, they could just give me all those things and tell me, “You’re on your own now.” I hope that doesn’t happen, and I don’t think it’s likely to.
I certainly won’t be afraid to be bold on this feed if I feel like there’s important stuff to talk about. It was funny, because the first day Eric called me, on a Sunday, to tell me about that deal, I had already recorded with Zvi on Friday, 2 days before—a classic Zvi episode—which was going to come out the next day and did come out the next day. In that episode, I accused Andreessen of perjury before Congress for having said that, basically, interpretability was solved.
I was like, “Eric, you know, this is coming out tomorrow.” And he’s like, “I don’t think they’re really going to care. I think it’ll be fine.” So far, it’s all been fine. Obviously, they’re big boys, and they can take some criticism, and they’ve been willing to put that in black and white.
I do hope that, over time, as we learn more about the overall shape that AI development is taking, the accelerationists and the AI safety people can realize how much common ground we actually have. I think on an overwhelming number of questions, I’m probably going to agree with a16z. There are some where I definitely don’t, and I’m certainly not going to shy away from that, especially now that I have this contract in place. But I also don’t want to pick a fight unnecessarily.
I’m going to try to follow my own advice from 20 minutes ago and not psychologize their takes. I’m going to assume—and I try to assume this in general about powerful people—that they’re rich as can be, right? They don’t need more money. Why are they doing what they’re doing?
I genuinely think that they’re trying to advance the human condition. I really don’t think that multi-billionaire, decillionaire folks like Marc Andreessen and Ben Horowitz are making their decisions at this point based on lining their own pockets. I really don’t think that’s the case. I can’t rule it out. I don’t know them, but I really don’t think that is the case.
I think that they are trying to advance the human condition, and they are trying to make sure that we don’t go stagnant, that we don’t allow fear of change to prevent us from realizing a beautiful future. So I take them at their word. That is what their motivations are, and I share a lot of that.
I have some different opinions on some pretty important questions, but I think there’s a lot more alignment than has at times been assumed online. I think that’s true not just for me, but for a lot of people who are fundamentally interested in AI safety. As I said earlier, I hope that I can be critical from time to time or, at a minimum, disagree and potentially even go into criticism without it fundamentally breaking the relationship.
I do have the assurance that I can continue to do this podcast and continue to reach you, the audience that has subscribed to the feed. If you’re listening this far along into the podcast, I appreciate that, and I don’t take for granted at all that people want to follow me on this learning journey. I just wanted to make sure that I had the ability to speak freely, speak my mind, and say what I think is important without having any fear that I would lose the ability to continue to use the modest platform that we’ve built.
That is in place, so I feel really good about it. I appreciate Eric and a16z broadly for making that a pretty smooth and painless process, and it gives me a lot of confidence that I can just keep doing this and keep calling it how I see it. I’m not going to try to create conflict where it doesn’t need to exist, but I will say what I think.
It’s a very privileged position to be able to do that and even make part of my living doing it. I definitely don’t take that for granted. So thank you again to Eric and a16z for making that as smooth as it was. Thank you to everybody who has listened. I really don’t take this opportunity for granted.
It’s a lot of fun, and it’s thrilling every day to get to wake up and think about, “What do I want to learn today? What feels important? How do I go make sense of this crazy AI wave that we’re all riding together?” It can be a little scary at times, but I do love the challenge that it presents me. Thank you all for making that possible. And in closing, thank you for being a part of the Cognitive Revolution. If you're finding value in the show, we'd appreciate it if you take a moment to share with friends, post online, write a review on Apple Podcasts or Spotify, or just leave us a comment on YouTube. Of course, we always welcome your feedback, guest and topic suggestions, and sponsorship inquiries, either via our website, cognitive revolution.ai, or by DMing me on your favorite social network. The Cognitive Revolution is part of the Turpentine Network, a network of podcasts which is now part of A16Z, where experts talk technology, business, economics, geopolitics, culture, and more. We're produced by AI Podcasting. If you're looking for podcast production help for everything from the moment you stop recording to the moment your audience starts listening, check them out and see my endorsement at aipodcast. And thank you to everyone who listens for being part of the cognitive revolution.