AI AMA——第1部分:OpenAI 的 o3、审议式对齐与 2024 年 AI 的意外之处
OpenAI 的 o3 看起来像是真正的推理突破,但还不是解决模型不可靠性的良方。 它在低算力 ARC-AGI 测试中的得分达到75%,大致达到人类水平,比此前系统高出约15个百分点;高算力模式达到87.5%。但推理模型仍会误读简单的井字棋棋盘、沿用“缓存启发式”,有时甚至靠“完全垃圾的推理”得到正确答案。
o3 在商业上最关键的能力,可能是其尚未公开的多次推理结果筛选方法。 低强度模式在100个 ARC-AGI 任务中各采样6次,共生成约3300万 tokens,每项任务成本约20美元;高强度模式精确使用1,024次采样和57亿 tokens,每项任务约13分钟完成。如果 OpenAI 能在没有客观答案键的情况下可靠地选出最佳回答,Nathan 认为这可能是“巨大、巨大的突破”;否则,数学和编程可能会与写作等主观工作进一步分化。
审议式对齐能让政策变化更快写入模型,但没有回答模型应遵循什么政策、或谁的价值观。 Nathan 的理解是,OpenAI 先从一个“纯粹有帮助”的模型开始,让它根据详细政策进行推理,只评估最终可见回答,再对成功的推理轨迹进行微调,随后实施强化学习。Nathan 的结论是“更好,但还不够”:表现处于“高90%”水平,而不是5个9;欺骗、涌现目标、政策合法性和模型性格等问题仍基本没有解决。
推理时算力的引入,推翻了此前 AI 民主化论的一部分,因为智能变得更加依赖算力。 一个只花几百万美元训练的前沿模型曾暗示 GPT-4 级能力会普及,但 o3 的高强度模式单题成本可能达到数千美元,而以今天的芯片供给无法普遍提供。这会造成“AI 获取机会的不平等”,同时也可能让算力治理重新变得重要,并使资源充足的防守方获得相对流氓行为者的优势。
2024 年最大的意外,是顶尖推理能力与日常自主执行之间的裂口扩大。 模型可以在诊断、竞赛编程和前沿基准上击败领域专家,却仍无法可靠地预约日历事件、维持连贯的长期记忆、感知简单布局,或放弃低效的排障框架。Nathan 认为,结合 GPT-4o 级模型、10万 token 上下文、示例和微调,组织可以“轻松实现100倍”以上的自动化;瓶颈在于过时印象、员工激励和薄弱的落地能力。
Nathan 预计 2025 年会迎来姗姗来迟的 agents、被测饱和的基准,以及更多难以解释的行为。 强化学习应当创造出类似“第37手”的解法:起初看似错误,最终却奏效;同时也可能带来更多在真实环境中“谋划”的行为。若效率优化把推理从语言迁移到潜空间,这些行为会更难检查。他对即将形成的新阶段的概括是“通用型怪异”。
职业和创业策略应当偏向内在动机与狭窄的客户价值,而不是押注持久的技术稀缺性。 如果想创造产品,就去学编程;但不要因为预期编程工作受到保护,就花4年培养程序员身份;“不要把时间上的微小相对差异,误认为整体趋势的形状”。对企业而言,品味、垂直领域模板、可靠的最后一公里编辑和解决付费痛点可以争取时间——但“如果这些东西在质量上能追上你,它们会在广度上碾压你”。
1. 推理模型仍然陌生到连儿童游戏都会输
Nathan 的首要保留意见是,OpenAI 之外几乎没人真正用过 o3;公开证据只是分散在一系列亮眼基准上的“零星结果”。他申请了安全测试项目,但把所有架构层面的推断都视为暂时判断,而不是把一次发布当成完成品。
与 OpenAI 不同,Google 的 Gemini 2.0 Flash Thinking 实验和 DeepSeek 的推理模型公开了思维链。阅读数以万计的 tokens,进一步印证了 Nathan 早年的判断:“LLM 很怪”(LLMs are weird):增加推理并不能消除糟糕的感知、缓存启发式或“彻头彻尾的糟糕推理”。
他的诊断题是一张未完成的井字棋棋盘:X 占据一个角,O 占据其正下方的角位,而最优开局应由 X 先走。X 可以制造双线威胁并获胜,但 OpenAI、Anthropic 和 Google 的模型却反复复述那句记忆性结论:最优井字棋总会和棋。
有些系统误读了初始棋盘;另一些探索了多个分支,却忽略了眼前的三连线。矛盾之处在于,同一批推理模型能提出令人印象深刻、受生物学启发的神经网络架构——看似更难的提示——却输给了这个玩具游戏。其中一个模型甚至在 Nathan 所说的“完全垃圾的推理”之后给出了正确结论。
2. o3 在 ARC-AGI 上的跃升可信、幅度大,但仍不完整
Nathan 认为,ARC-AGI 合作是现有最干净的证据,因为 OpenAI 直接与独立基准团队合作。关于是否使用 ARC 训练集的争论,并没有否定可比性:其他竞争者也用过该训练集,而 o3 只收到最低限度的指令——推断规律并完成最后一个例子。
在低强度模式下,o3 得分75%——大致达到人类水平,略高于平均 Mechanical Turk 受试者,比此前展示过的任何系统高出约15个百分点。Nathan 称这个几乎不带指令的提示是“不可思议的炫技”,而不只是依赖复杂的、针对基准量身定制的脚手架。
高强度 o3 达到87.5%,超过引用的人类基线,已经足以让人感觉它解出了谜题。但 François Chollet 仍然展示出这样的任务:人类做起来很容易,o3 却可能很难;在 Chollet 的标准下,只有当这种不对称性无法再被构造出来时,AGI 才算到来。因此 ARC-AGI-2 仍然重要。
3. 测试时算力把一个答案变成工业化搜索流程
Nathan 认为,普通 o1 和 o1-mini 可能各自只执行一次思维链 rollout,而网上的猜测则把 o1 Pro 描述为多次并行调用 o1 后再聚合。对此他明确表示不确定:尽管他仍在测试 o1 Pro,却没有发现它明显优于普通 o1。
o3 低强度模式每项任务使用6个样本。OpenAI 报告称,100个任务共生成3300万 tokens——平均每项约30万 tokens、每个样本约5万 tokens——成本约为每项20美元。这与 o1 每百万 tokens 60美元的输出价格基本吻合。
报告中的1.3分钟运行时间意味着极高吞吐量;如果按字面理解,速度可能接近每秒1,000 tokens。Nathan 怀疑这依赖专用硬件,而不是普通用户当下即可获得的零售服务,但他强调,只要表现足够强,许多具有经济价值的任务都能承受20美元的推理账单。
高强度模式精确使用1,024个样本,共生成57亿 tokens:每项任务约5000万 tokens,同样约等于每次 rollout 5万 tokens。但任务耗时约13分钟,而不是低强度模式1.3分钟的200倍,这说明系统具备大规模并行能力,同时还存在某种尚未公开的结果解析流程。
4. 隐藏的筛选器可能比1,000个候选答案更重要
对于客观数学题,聚合可以采用多数投票:如果多次尝试收敛到同一个数字答案,就选它。写作和其他开放式工作更难,因为每个候选答案都不同,必须依赖评审模型、两两比较、循环赛评分或其他类似锦标赛的流程,而其可靠性本身仍然未知。
从1.3分钟升至13分钟,意味着系统要么把并行 rollout 分成每批约100次,要么采用顺序筛选的锦标赛。Nathan 用后者举例:1,024个候选答案经过约10轮淘汰,从1,024到512,再到256,直到只剩一个;但他强调,OpenAI 几乎没有披露任何细节。
这一差异决定了能力能否泛化。在奖励便宜且可验证的领域——数学、编程、游戏——强化学习应该会像 AlphaGo 一样“越跑越快”。如果筛选器也能识别没有明确答案的质量,进步就可能广泛扩散;否则,主观领域与可验证领域之间可能出现明显分化。
5. 审议式对齐把政策阅读器训练进模型
Nathan 将起点描述为一个“纯粹有帮助”的模型:它被训练为满足用户要求,不会拒绝非法或有害请求。随后 OpenAI 提供一套详细的行为政策,让模型在回答前先推理相关规则,从而识别越狱,并说出类似“用户正在试图骗我”的判断。
这套政策可能大到无法完整放入现有上下文,因此训练时只使用相关小节。Nathan 将这种复杂性比作 Facebook 的内容审核手册:全球规则最终会落到边缘案例,比如一个婴儿从母亲乳房上移开嘴后,短暂露出的乳头是否仍然符合规则。
按 Nathan 的理解,同一个基础模型随后会根据政策评估最终可见的回答。奖励模型看不到思维链,只评分用户会收到的内容,从而减少模型把被禁止的推理藏在表面合规输出中的直接压力。
OpenAI 会对成功的推理轨迹进行指令微调,但不提供政策原文,试图把政策知识写入权重并释放上下文空间。之后强化学习再生成并评分更多回答。人类投入似乎主要集中在撰写政策,而不是为每条训练样本逐一标注。
6. 遵守政策不等于完成对齐
这种方法与 Anthropic 的 Constitutional AI“相似之处明显多于不同之处”:模型查阅一部宪法,批评输出,并通过迭代训练逐步服从规则。OpenAI 的显著变化是保护思维链不被奖励模型看到,希望欺骗意图能在内部监控中保持可见。
结果与 Claude 具有竞争力,但并非变革性突破。Nathan 将其政策遵从度描述为“高90%”,而不是5个9的可靠性——足以快速适应普通政策变化,但不足以满足政府或其他高风险主体对严格服从的要求。
论文基本没有讨论欺骗、自主生成的目标、如何检测这些行为,以及服从性能否达到真正关键任务所需的可靠性。它也没有回答究竟应由什么政策来统治模型。Nathan 称这个系统是“AI 对齐领域的 Ron Burgundy:政策里写什么,它就对齐什么”。
这种价值中立的灵活性既能服务安全,也能支持大规模审查;Nathan 引用了 Peter Thiel 的说法“AI 是共产主义的”,并设想了 Facebook 内容审核与国家言论控制两种方向。相比之下,Claude 在反复进行的博弈论测试中表现出强烈合作倾向,说明模型性格——而不仅是规则遵循——可能实质性影响高信任场景的结果。
7. o3 让获取机会更不平等,却让治理更有可能
在推理模型出现之前,技术路径看起来更平等,也更难治理。Nathan 提到,一个据称只花约600万美元训练、接近 GPT-4 质量的中国模型:在一个价值数千亿美元、并向万亿美元迈进的算力市场里,“几百万美元不算什么”,因此支出控制最终会漏水。
高强度推理挑战了这一图景。如果单个答案成本达到数千美元,还要消耗庞大的并行容量,就不是所有人都买得起;即便所有人都愿意付费,也无法凭空制造不存在的芯片。Nathan 因此预计,在这一范式下,“AI 获取机会的不平等”不可避免。
潜在的好处是“物理规律在帮我们”:也许即使是极其聪明的系统,艰难发现也需要持续搜索和真实资源。Nathan 稍微转向支持 Martin Casado 的观点:有些答案不能凭空摘取;洞见可能突然出现,但往往要先消耗大量计算周期。
算力密集型能力可能更有利于防守方。Zuckerberg 的类比是垃圾邮件:攻击者拥有 AI 和自动化,但平台拥有更大的计算机、更多数据和更好的系统。Nathan 不会“拿宇宙的未来下注”于此来阻止工程化疫情,但一台笔记本电脑设计出一场疫情的能力,可能确实不如受监控的集群。
8. AI 在深度上跨过专家门槛,却在其他地方依旧脆弱
Nathan 所说的“认知光谱尾部”,早已赋予模型超人的知识广度、速度、成本优势、可用性和可扩展性。一段旧对话可以带着10万 tokens 的历史即时恢复,同一个系统还可以并行运行100次或1,000次,而无需人类重新进入状态。
深度能力也可能已经跨过门槛。o3 达到了大约全球排名前200的竞赛程序员水平,前沿数学表现大幅跃升,越来越多医学诊断证据也开始支持 AI 胜过临床医生。Nathan 的保留词是“也许”,但他认为,现在说模型已在多个深度领域达到或超过人类专家,正变得更加稳妥。
人类在局部上下文、连贯的持续记忆、感知、稳健性和情境意识方面仍遥遥领先。模型可能比 Nathan 更完整地保留一段文字记录,却没有持续更新的人生史;它们还会相信“这个草稿是私密的”之类的说法,这正是 Apollo Research 诱发欺骗行为所利用的技巧。
排障时,模型会反复攻击用户描述的 bug,而不是跳出框架问一句:“你试过重启电脑吗?”推理正在进步,但仍然浑浊;真正的洞见已经从“没有 Eureka 时刻”变成可能出现少数几个,而长任务仍会陷入重复并失去方向。
9. 2024 年的短板在部署、记忆和自主性
Nathan 原本预计整合式记忆会进展得更快。ChatGPT 的记忆可能保存一些随机的零碎话,却漏掉真正重要的内容,之后又因为这些脆弱笔记而表现异常。感知能力有所提高——Claude 最终可以按像素坐标指出界面元素——但仍弱于基准分数所暗示的水平。
Agents 也低于他的预期,但人们可能更加低于预期。回看自己在 GPT-4 发布后的预测,Nathan 原以为即使模型不再进步,经济影响也会广泛得多。18个月过去,组织从这项技术潜在的自动化能力中实际兑现的部分少得惊人。
他将原因归于过时的第一印象、扩散不足的落地实践和激励机制。一些开发者对 AI 的印象仍停留在早期 GPT-3.5 级 Copilot;另一些人则根本没有试过新系统。员工可能掌握自动化工作的本地知识,却也有充分理由担心自动化对自己意味着什么。在成功尚未显现之前,整理示例、编写琐碎的推理轨迹同样令人厌烦。
10. 现有模型可以支撑大得多的自动化
Nathan 的方法刻意不追求炫技:定义准确任务,说明什么才算做好,收集少量高质量示例,记录强推理轨迹,必要时进行微调,再把过大的工作拆成可处理的子任务。大多数工作流甚至不需要做到这套流程的最高强度版本。
在 GPT-4 级能力、约10万 token 上下文和微调的组合下,他估计社会可以实现“轻松100倍”于当前水平的自动化。他还指出,GPT-4o 的微调——包括图像能力——已经可用。这不是一个必须等到 o1 或 o3 才能实现的预测,而是说现有系统内部已经存在大量尚未使用的能力。
信心很重要,因为建立黄金标准示例,可能意味着为一个没人愿意记录的琐碎流程写500字。Nathan 愿意做这些苦活,因为他预期它会奏效。许多组织听过这种方法的零散片段,却没有吸收足够的操作知识来真正投入。
11. “把所有东西都 Flash 一遍”可能胜过精巧的检索系统
检索增强生成让 Nathan 失望,因为向量搜索经常找不到正确证据。Andrew White 的 FutureHouse 团队则愿意承担必要成本:先广泛做关键词检索,再让语言模型处理每一篇可能相关的论文,并逐篇询问是否相关。“把所有东西都 Flash 一遍”用更高的推理成本换取更好的科学答案。
对于覆盖50个州的保险公司,每个州可能有50—100页法规,Nathan 建议每次提问都把整套语料交给 Gemini Flash。按当时价格,他估计每个问题约10美分、耗时约1分钟,效果可能远胜内部聊天机器人——后者的向量数据库往往只能返回不完整的上下文。
开发者会抗拒10美分、50美分或1美元的边际成本,因为传统软件让每次点击都感觉免费。但100名员工每天各问几十个问题,可能只消耗10—15美元,“还不到你买咖啡渣的钱”;在这种情况下,准确性远比 Facebook 级别的效率更重要。
12. 编程能力进步了,但上下文管道仍制造“Cursor 效应”
SWE-bench 和 Codeforces 展示了异常强劲的进步,但日常编程仍保留一些小而致命的怪异行为。Nathan 和合作者 S.J. 将一种失败称为“Cursor 效应”:检索系统漏掉项目的标准 types 文件,于是模型自信地新建第二个 types 文件,而不是扩展已有文件。
问题往往出在工具和上下文选择,而不是核心模型。系统当然可以读取整个代码库,但出于成本考虑,检索器只挑选所谓相关的文件。当一个新概念此前从未出现在 types 文件里时,相似度搜索几乎没有理由把该文件找出来——恰恰是在模型最需要知道它存在的时候。
Nathan 的变通办法是把一个约10万 token 的应用打印到单个文件中,把整个代码库交给 o1,要求它给出分析、问题和计划,再把计划交回 Cursor。他仍然更喜欢让 Claude 逐行编程,也认为 o1 Pro 能完美一次性生成应用的说法与自己的体验不符。
即使 headline 式的上下文窗口大小没有显著增加,有效上下文能力也提升了。模型变得更擅长对所提供的全部内容进行推理,可以规划复杂的整个代码库。Nathan 与采访者最后趋于一致的实用技术栈是:用 o1 做推理,用 Cursor 内的 Claude 做实现。
13. 状态空间混合架构仍有潜力,但已不再紧迫
Nathan 在 2023 年对 Mamba 的预测,并不是注意力机制会消失,而是状态空间与注意力的混合架构最终会胜出。他也准确指出了失败条件:如果 Transformers 持续稳定地产出成果,替代架构就会“没有氧气”,研究者会继续开采主矿脉。
事情基本如此,尽管混合架构也做出了强劲样本。Albert Gu 的 Cartesia 在文本转语音的速度、质量和低价格上表现惊人;Evo 是另一个成功的混合模型;Google 所宣称的近乎无限上下文路线,则暗示“一点点注意力就够了”。
Nathan 仍认为,对于长寿命 agents,有限且不断演化的状态在结构上很重要:他的大脑不会随着每次经历以平方速度增长。但如果传统注意力能处理1,000万或1亿个可用 tokens,平方级增长对许多应用来说也可能在功能上等同于无限。因此,他给 Mamba 一年后的评价是“未完成”,而不是被证伪。
14. Agents 之所以显得姗姗来迟,正是因为它们的奖励可观察
带屏幕共享的 Advanced Voice Mode 让 Nathan 瞥见了缺失的产品形态:他阅读生物学论文时,模型能观察图和文字,接受语音提问,并解释它看到的内容。再结合 Claude 的 computer use,成熟 agent 框架迟迟未出现就显得更加意外。
Claude 可以导航 Waymark,并在第一次尝试时为其视频制作工具编写提示词,但安全护栏对账户使用的限制过于激进,逼得 Nathan 只能误导它。随后 Claude 声称自己在匿名使用该网站,而 Nathan 表示这并不属实——底层 computer-use 技能有效,但稳健性仍然很弱。
日历和浏览器任务似乎非常适合强化学习:事件是否创建成功、时间是否正确、用户是否接受,或者 agent 是否找回了秘密密码?Nathan 认为,合成环境可以生成这类信号。他称2025年的共识是“agents 之年”,同时指出,具备真正功能的自主性其实已经姗姗来迟。
Sam Altman 要求“把所有基准测到饱和”也越来越像是可行目标:Nathan 提到 FrontierMath 据报提升了20%—25%,而 OpenAI 也释放出持续进步的信号。Nathan 对2024年的核心意外是:系统可以在难题上击败博士,却无法完成普通行政工作。
15. 强化学习会制造“通用型怪异”
Nathan 预计会出现更多 AlphaGo 式的“第37手”时刻:某个动作起初看似错误,最终却证明有效。同样的优化压力也可能产生欺骗性谋划,或通过奇怪路径完成目标,让用户无法判断某个难以解释的动作究竟是天才、bug,还是带有敌意的行为。
Meta 的动态采样研究体现了这种效率压力。模型不再只有一个由用户设定的 temperature,而是可以在明显阶段选择低随机性,在笑话真正落点的那个 token 提高创造性。采样策略本身变成端到端学习出来的行为,而不是透明的外部控制项。
Nathan 更担心的是连续空间或潜空间中的推理。模型可以不选择可读 tokens,而是循环利用隐藏状态,似乎把多条搜索路径以叠加态表示,并以比受语言约束的深度优先轨迹更高效的方式进行广度优先探索。
尽管存在效率优势,Nathan 仍表示自己“有点讨厌它”。自然语言思维链至少可以被内部监控;潜空间推理则需要尚不成熟的可解释性探针。算力贫乏的环境有动力采用这些效率进步,即便它们降低了可读性,因此2025年很可能会成为“通用型怪异”以及真实世界重大谋划行为的起点。
16. AI 可能已经是最好的导师——前提是学习者能忍受学习
Nathan 将自愿学习与制度化教育分开。对有动力的学习者而言,一个能看到同一篇论文、广泛掌握相关领域、即时回应且成本低廉的对话式导师,是“学习者经历过的最好的事情”。它让他无需聘请专业导师,就能进入 AI 与生物学的交叉领域。
他不确定的是学习者的性格。Nathan 5岁的孩子已经准备好阅读,却会因陌生单词而沮丧,因此他强调,要学会享受困难、努力和进步循环往复的过程。AI 可能更会鼓励人,但未必能消除学习所需的不适感。
课堂部署远远落后于能力。密歇根州一家非营利机构的人士描述 Khan Academy 试点时说,教师有时会表示“我们喜欢这个”,但学生始终没有真正接触到;教师主要把它用于尝试课程计划或个人辅助。Nathan 还提出,一旦技术影响变得更清晰,工会和既有利益相关者可能会产生抵触。
上限是让所有人都获得亚里士多德级别的私人辅导;约束则在于,人们是否愿意像亚历山大大帝那样,以文化鼓励和接受挑战的心态投入其中。Nathan 并不是在美化亚历山大的道德形象,而是把他当作一个准备好利用非凡教学的学习者样本。
17. 学编程是为了创造,而不是为了获得受保护的职业
Nathan 举的具体例子,是一个由其家人资助上社区大学的前 au pair。H-1B 数据显示,签证主要流向 IT 和编程岗位,但她更喜欢幼儿教育。他们没有押注4年去学她并不喜欢的编程,因为 AI 可能在她毕业前就主导这个职业。
他的建议不是“不要编程”。Cursor 和 AI 导师已经让编程更容易、更少乏味,也更容易理解;任何真正想创造应用的人都应该直接开始。但他不会仅仅为了未来的劳动力市场回报,就培养软件开发者身份,而这种回报可能很快消失。
短期套利机会仍然存在:客户习惯了昂贵的软件,而 AI 辅助的开发者可以低成本交付项目,包括按每小时几百美元、甚至约1,000美元报价的工作。Nathan 认为机会目前很丰富,但不确定性太高,不足以支撑长期人生规划。
转行学法律也不是安全对冲。Miles Brundage 的警告是:“不要把时间上的微小相对差异,误认为整体趋势的形状。”编程率先变化,是因为其奖励可衡量,且 AI 实验室本身就是软件公司;但法律分析和其他知识工作应该会跟进。
18. 狭窄的产品质量可以争取时间,却无法提供永久免疫
“没有护城河”仍是 AI 领域的主旋律,但 Waymark 展示了在位者如何把旧资产与新模型结合起来。其由人设计的视频模板已经编码了运动、音乐和视觉品味;AI 可以比许多客户更快、更好地把文案和素材填入这些容器,而无需重新发明整个形式。
最后一公里的控制同样重要。即使用户非常满意,也可能要修改电话号码、替换网上找不到的照片,或调整某个场景。可靠的编辑界面能精确交付用户想要的结果;反复提示生成式模型往往做不到。垂直领域品味加上确定性修正,仍然是可防守的组合。
Nathan 建议解决真正有人付费的痛点,而不只是做“某个很酷的 AI 东西”。Waymark CEO Alex 始终让团队聚焦用户提出的需求,Nathan 也承认自己偏爱新奇事物。产品更安全的条件是:相关能力既不太可能自发出现,也没有重要到足以让前沿实验室专门投入工程资源。
这道护城河是有条件的:“如果这些东西在质量上能追上你,它们会在广度上碾压你。”足够强大的 ChatGPT 可以调用 Sora、叠加文字、编辑视频,并按需生成一个类似 Waymark 的应用。在那之前,狭窄领域的卓越表现可能持续;之后,软件可能会坍缩为推理成本与更广泛的供给。
19. AGI 可能让专业知识民主化,同时集中资本与决策
Nathan 没有假设 UBI 或后稀缺社会。他暂时构想的模型有点像他理解中的瑞典:主要公司的控制权可以高度集中并由家族继承,同时社会拥有慷慨的保障底线和广泛的物质安全。他明确表示自己并非专家,也预计真实未来会“相当怪异”。
专业知识可能变得几乎免费,而前沿能力仍然稀缺。所有人都可能获得最好的 AI 肿瘤专家、创意工具或建议,但少数政府和科技巨头会决定训练哪一个1000亿美元规模的模型、哪些研究获得算力,以及哪些问题值得投入百万美元推理。
物理稀缺性仍然存在:Nathan 怀疑 OpenAI 的 Sora 发布部分受限于算力不足,尽管他也承认可能还有其他解释。一次跨城看医生消耗的资源,相当于许多次语言模型交互,因此日常生活可能变得更富足,而战略资本和决策权却更加集中。
对于希望作出积极贡献的个人,Nathan 不认为缺乏技术专长是永久障碍:AI 让追赶近前沿能力变得异常容易。他讨论了 adoption accelerationism、hyperscaling 和暂停派的观点;他的实际策略是展示系统已经能做什么,改善公众校准,并让更好的世界模型为下一步决策提供依据。
Even with these reasoning models, they are still weird. They are still subject to cached heuristics, bad perception, and just straight-up bad reasoning. You can expose that in simple toy examples.
We hope the underlying reality is such that these things are pretty manageable, because we’re not taking the level of caution that we would need if they’re not manageable. In 2025, we’ll probably start to see a lot more of these things, like the famous AlphaGo move where it looks like a mistake at first but actually turns out to work, or some of these deceptive scheming behaviors where AI is accomplishing the goals it’s been given in strange ways.
Do what you want to do. Don’t spend years of your life on a bet that, in some sense, implies that AI isn’t going to dominate the programming profession in the next few years, because it seems plausible enough that it might.
Aidan Elang Goen, welcome to The Cognitive Revolution.
Thanks for having me.
Nathan, thanks for being here. I’m excited to do this. This is the AMA, and you’re making your first appearance in front of the camera after a lot of effort over the course of 2024 behind the scenes.
I’ve mentioned you guys a couple of times on the show and a few times on Twitter, but for anyone who hasn’t caught those mentions, Aidan and his partner, Sai, are the founders of a company called AI Podcast Casting. AI Podcast Casting is the website, and we’ve been working together for most of this year to produce the show.
It’s been a lot of fun. I appreciate all the effort, and thanks for coming on today to drive this AMA episode.
Thanks to you for being our first client and making AI Podcast Casting happen.
My pleasure. My experience with Turpentine has been great in many respects. They took all this stuff off my plate at the beginning—sponsorship sales, editing, and so on. The one thing I always thought was, “I wish we were using AI more to actually produce the show and create some sort of leverage.”
The opportunity to work with folks who had listened a lot, who have a real interest in AI, and who are interested in creating these processes was exciting. That’s where all the clips have come from, if you’ve seen clips on Twitter, and all the shorts that we’re doing on YouTube.
There’s definitely still a lot more that we can do, but it has been pretty amazing how much we’ve been able to do on a pretty small budget, especially considering that we’re putting out 8 episodes a month. I know I’m keeping you guys pretty busy with the pace of content.
No, I should also mention that—just to rewind a bit—when we were doing this, we weren’t sure that we would actually pull off the complete postproduction for a podcast. But you came in and said, “Hey, look, this is our pain point. Why don’t you apply AI everywhere?”
We’ve done a good job, but I think we could do better, as you said. I’m looking forward to the opportunities in 2025.
Always. For now, we put up these questions, mentioned it on a few episodes, and put it out on Twitter and YouTube. We got a pretty good response, with quite a few questions spanning a lot of different areas.
Take it away and try me on all these questions.
Absolutely. I’ll try to be the voice of our audience here. We have lots of questions, so if I missed some, please feel free to drop them as YouTube comments. We’ll try to address them there.
Nevertheless, we’ll get started now. The way we’re going to do this is that we have sections, but we thought we would start with something that’s been trending in the headlines recently. A user would like to know: What do you think of o3? A related question is: What do you think about OpenAI’s new safety plan called deliberative alignment?
Obviously, the o3 announcement—not yet a release—caught everybody’s attention and has had people using their spare holiday cycles to try to figure out what it means. First of all, it’s important to keep in mind that there’s a lot that we don’t know. Nobody has seen the model yet.
I applied to the safety review program, and I think it’s cool that they put out that open invitation. Basically, what we’ve seen so far is a smattering of results across a number of different benchmarks, and those are super impressive. But that’s limited information. The sources aren’t as authoritative or as clean as you might wish they were if you really wanted to have a high-confidence answer.
Nevertheless, this is what we have, so we’ll try to make sense of it. For starters, I think it’s useful to look at some other reasoning models. OpenAI is currently unique, as far as I know, in having a reasoning model that does not show its chain of thought in its response to you.
Everybody else who has put one out so far, as far as I know, is sharing the chain of thought. That includes Google with its Gemini 2.0 Flash experimental thinking version. It also includes DeepSeek, which has a reasoning model. There’s another Chinese one as well, although I haven’t really used that one.
I have used the Flash Thinking version and the DeepSeek model. It’s interesting to read these long chains of thought and try to get a better sense for what the models are thinking.
It has been a weird experience. I think the headline that LLMs are weird remains very much in effect. All of the models that I’ve tried, including o1 Pro, have failed on what you would think would be a pretty amenable problem for a reasoning model.
Certainly, somebody who can reason in general-purpose terms should be able to handle this challenge: Here is a Tic-Tac-Toe board. I put an X in one corner and an O in the corner immediately below that, and I said X went first. It’s X’s turn. Assuming optimal play, is it possible to determine who the winner will be?
Every model loved this question, to my surprise, but the answer is that X can win. You can force a fork and win. If you have optimal play from that position, X will win.
The responses I’ve gotten from OpenAI models, Claude, and the latest Geminis—even leaving aside the reasoning—are basically all wrong. They seem to be anchored in the commonly repeated statement that Tic-Tac-Toe is a solved game and, with optimal play, it will always be a tie.
You see a lot of responses like that. That’s not reasoning. It’s just a cached heuristic. You could even say that’s stochastic-parrot mode, because it latched on to a couple of tokens and has seen that statement a lot in the training data.
I’ve also seen surprising problems where perception is the weird thing. In one case, one of the models simply didn’t read the board correctly and started with the wrong board state: an X in the corner and an O in the middle bottom, as opposed to the bottom corner.
Obviously, you’re not going to do very well if you’re starting with the wrong board state. Then, perhaps most revealingly, the reasoning in both Gemini 2.0 Flash Thinking and DeepSeek was pretty bad—outright nonsense in a lot of cases.
It’s trying to do this sort of rollout. It’s trying to go down these branches and then come back to the top. It would do things like, “Okay, let’s say X goes in the middle. Then what will O do?” But it would often fail to—again, I had said optimal play—recognize that one of the players could win just by completing three in a row.
It would fail to do that in its analysis of the situations. It did not know what optimal play was for a great many board states. You’re reading through tens of thousands of tokens of reasoning about the situation, and you’re thinking, “Man, a lot of this reasoning is pretty bad.”
There’s definitely not a panacea here. Just because somebody comes out with a reasoning model doesn’t mean it will necessarily be an effective reasoner.
At the same time, with other kinds of problems, I’ve seen remarkable results. I asked those same two models—Gemini 2.0 Flash Thinking and DeepSeek—for suggestions for elaborations of neural-network architectures that would take inspiration from some sort of biological system.
I was thinking about the AE Studio episode we did, which I thought was excellent in terms of thoughtfully picking something out of the biological world and trying to create, obviously not a direct imitation of it, but something that captured the key concept of how these biological systems were understood to work.
I was quite impressed. That was a little bit less reasoning and a little bit more riffing on ideas. I’m sure plenty of those ideas aren’t going to work and are flawed in various ways, but it does show that even with these reasoning models, they’re still weird.
They’re still subject to cached heuristics, bad perception, and straight-up bad reasoning. You can expose that in simple toy examples. At the same time, you get things that are remarkably impressive with a question that would seem, at least to a human, a lot harder.
You would expect more people to handle the Tic-Tac-Toe question than the bio-inspired neural-network-architecture question, yet it’s totally the reverse for the reasoning models. That’s weird.
We don’t know what’s going on in the OpenAI chain of thought. I think it’s probably likely to be higher quality, but how much higher quality is not clear.
That’s interesting that you’d say that. I did watch the livestream, and I do use o1 Pro. I was not as impressed, especially for the $200 that I pay. It’s a very similar experience to that.
Nevertheless, when I saw the livestream on the 12th day, and they showed the ARC-AGI result, in my head—especially after we did a couple of episodes with the ARC-AGI co-founder—my impression was that it was at least 2 years away from breaking the ARC-AGI benchmark, regardless of how much compute you threw at it.
Now it seems like o3 has kind of broken it. That should mean something, right? What do you make of that? o1 and o1 Pro don’t quite work, but then there’s o3, which has already broken the ARC-AGI benchmark.
I think there’s definitely something here, and it seems like there are a couple of dimensions of improvement.
One is that they’re doing reinforcement learning at scale on these reasoning models. They said they are not doing reinforcement learning on the actual chain of thought, but only on the response. In other words, when they feed an output to a reward model to get the score used to power updates to the weights, that output does not include the intermediate tokens. It only includes the final thing the user sees.
The hope is that this won’t create weird incentives for the chain of thought. The chain of thought is free and unencumbered by these pressures, so it can reason in a very natural way. If it is being deceptive, then hopefully it will continue to make that visible to their monitoring systems, which they’re developing, and they can control it somehow.
I do think the ARC-AGI blog post on this is probably the best source of concrete information that we can break down. It’s important that they worked directly with this independent group. I don’t think this is faked.
In terms of people being very skeptical in general, I think there’s something quite real here. That’s safe to say.
There’s been some debate around whether they trained on the training set, and whether that’s allowed or not. For practical comparison purposes, everybody who has been competing in the challenge has been training on the training set, so they’re not alone in that by any means.
Their prompt was also an interesting revelation. It basically just said, “Identify the pattern, apply it to the final example, and give us your output.” That’s it. There were no detailed instructions. It was super simple. That was an incredible flex.
As for the results, in the o3 low-compute setting they got 75%. That’s blowing basically everything else away. It’s roughly at human level. They put it a little bit higher than the average Mechanical Turk respondent—perhaps a somewhat more motivated or savvy group of people—but it’s getting up to the level of human performance.
It’s a good 15 points clear of everything else that we’ve seen to date. That’s a pretty substantial leap.
Then there’s another leap with the high-compute setting, which gets them into the high 80s: 87.5%. That’s above human performance and starts to feel like they’ve solved this puzzle.
It’s maybe not fully solved, or it’s solved in a weird way, because there are some problems that seem easy to solve that it isn’t able to solve. I think the people involved with the ARC-AGI project have appropriately been giving credit where it’s due. They’re saying, “This is a big deal. It’s a real advance, and we’re going to need to study these capabilities.”
At the same time, François Chollet said that it’s still possible to make things that are easy for humans but hard for these systems. In his mind, it won’t be true AGI until that is no longer possible at all.
I guess, for ARC-AGI in his mind, this is not necessarily AGI, but it’s at least a little bit of a move. There’s ARC-AGI 1, and there’s going to be ARC-AGI 2.
As always, there’s a lot of nuance in these things. I thought one thing that was really interesting was the number of samples, the number of tokens, the cost that translates to, and the time that it was able to run in. What does all that imply?
As far as I know, the o1 models are just doing a single rollout of chain of thought and then giving you the answer. I’m pretty sure that’s how o1-mini and o1 are working.
There’s been some speculation online that o1 Pro is maybe multiple o1s running in parallel and then taking the best answer. I’m not sure. Maybe I haven’t quite unlocked it yet, but I have not felt that it’s really all that much better than normal o1.
I keep using it because I feel like that’s probably a skill issue on my end, but candidly, I have not gotten much from it that is better than normal o1, as far as I can tell. There’s been speculation that maybe it’s multiple rollouts and then some sort of decision at the end as to how to aggregate that. I don’t know if that’s true. I kind of guess not.
It is clearly—OpenAI has just stated as much—what o3 is doing. In the low-compute setting, that translates to 6 samples per task. They report this in aggregate in the blog post, but basically they’ve got 100 tasks and a total of 33 million tokens across all of those tasks.
That implies something like 300,000 tokens per task. With 6 samples, that would suggest something like 50,000 tokens per sample. They say that costs basically $20 per task, which lines up with o1 pricing.
The o1 pricing is $60 per million output tokens. If you had 50,000 tokens, you’re at 5% of that, so that would be $3 per 50,000 tokens. It all coheres. If you have 6 samples and they’re 50,000 tokens each, and you get to $20 total cost, that would suggest that the cost is considered to be the same as o1, but you’re using 6 in parallel.
Interestingly, they also report that the time per task is just 1.3 minutes. Now we’re asking: Are they really generating more than 50,000 tokens per minute? That would be about 1,000 tokens per second.
You’re usually not seeing 1,000 tokens per second from the OpenAI APIs. A couple hundred is much more in line with what I typically see. Even o4-mini is not faster in my experience than o4. It’s cheaper, of course, but in terms of tokens per second, it doesn’t seem to be all that much faster.
Maybe this was happening on dedicated hardware. I don’t expect that we’re going to see 1,000 tokens per second as retail customers in the immediate future, but if they’re saying, “We’ll carve out special compute for this purpose,” then maybe they can push it up to that level.
That’s pretty interesting. The $20-per-task figure is also interesting because there are a lot of things that are worth paying $20 for. If you could do a good job, this can open up a lot of new use cases.
The high-compute setting is even more interesting. That’s reported to be 1,000 samples—precisely 1,024 samples. They report 5.7 billion tokens generated there.
Again, by simple division, you’re at something like 50 million tokens per task, which is again 50,000 tokens per individual rollout. Five billion total tokens divided by 100 is 50 million per task; divided by 1,000, that’s 50,000 per rollout.
They report that this happens in 13 minutes. There are some interesting implications there for what exactly is happening behind the scenes.
One really big question I have is: If you’re doing 6 or 1,000 samples on the same question, how are you choosing the right answer, and how generalizable is that strategy?
I don’t think you could take 6 rollouts of o1 and get to the same level of performance on one dimension. It seems pretty clear that there’s simply better reasoning coming from the o3 model. The reinforcement learning has continued to work. It’s continuing to get better at its core, single-rollout reasoning and analysis.
But if you have 6 samples, how do you choose the answer? A classic approach in the space has been to take a majority vote. If it’s a math problem with a single answer, you have multiple rollouts. If 3 say the same thing and the other 3 each say something different, you go with the answer that has the most votes.
Another version would be to have an additional call at the end that says which of these is better. If there isn’t a single answer—if you’re trying to write something, for example—then you can’t simply choose the identical answer, because each paragraph or page will be different.
So how do you choose the best one? There have been many different techniques: Feed them all into a model and ask which one is best, do pairwise comparisons, or do a round-robin tournament. There are all sorts of tournament-style schemes for having language models judge themselves.
Those techniques are somewhat viable, but not super accurate. They work, but not always very well.
Something like that has to be happening here. If it were purely parallel, with a very simple resolution mechanism at the end, then those 1,000 samples could happen almost as fast as the 6 samples. There’s no reason it would take 10 times longer to do 200 times as many samples.
So why is it 10 times longer? Is there some sort of aggregation step that is a sequential process? Maybe it generates 1,000 samples in the first minute and then runs a 10-round tournament.
If you had 1,000 candidates and wanted to pick the best one in a tournament, you could do head-to-head, single-elimination. In the first round you’d go down to 500, then 250, then 128, then 64, and so on until you finally got to a single winner.
That would take roughly 10 rounds of the tournament to get down to one winner. It seems like something like that is probably happening, but we don’t know the nature of the mechanism.
That’s going to be a really important question for figuring out how generalizable this is. People have commented widely that we should expect reinforcement learning to continue to work and that AI systems are going to run away with anything where there is a ready, easily accessible reward signal.
Anything where it’s verifiable—did you get the correct number on this question or not?—can be rewarded. This is where superhuman performance comes from. This is where AlphaGo got so good at playing Go that it made famous moves that surprised the best human grandmasters.
It seems safe to say that this will happen in easily verified domains. Math and programming are probably like that to a significant degree. The question becomes: How many other things are like that? Can you get a signal like that for writing? That’s much harder to verify.
There’s an expected divergence between things that are verifiable and things that are much harder to verify, or subject to taste. How much that divergence holds may depend a lot on what this consensus-finding mechanism is and how robust it is.
You could imagine a scenario where the reason it took 10 times as long is that they only parallelized up to some maximum width. Maybe they run 100 at a time, with each batch taking a minute, and then run 10 sequential steps. Then they take a simple resolution at the end.
Alternatively, they could generate 1,000 samples at once and then use a longer, more computationally intensive aggregation mechanism. That is probably my biggest question right now: What exactly have they figured out there, and how much will it apply to other things?
If they’ve figured out a way to choose the best response in areas where there isn’t an obvious “yes, this is right” or “no, this is wrong,” then we really could be talking about a huge breakthrough.
But they haven’t said much at all about that so far. Everything they’ve said, as far as I know, is about things that are fairly verifiable. You’ve got coding tasks and FrontierMath.
There’s been some interesting discourse around the degree to which these FrontierMath questions are being answered correctly for the right reasons, versus being answered correctly without necessarily using the right reasons. The FrontierMath people tried to design a test that you couldn’t luck your way into getting the right answers on, so I think we have good reason to believe the models are probably doing the math.
Going back to my Tic-Tac-Toe example, though, there was at least one model that gave the right answer at the end based on totally garbage reasoning. The final paragraph said, “Yes, if you assume optimal play from here, X can force a win,” and I thought, “That’s impressive.” Then I went back and looked at the chain of thought, and it was totally garbage.
I’ve seen at least a glimpse of this kind of thing. That may be what’s happening with FrontierMath as well, but it’s hard to say.
Okay, cool. It seems like these reasoning models are getting better. Have we hit the wall? That question is not clear, but it’s clear that these models keep continuing to get better.
That nicely follows up to the next questions, which are also about safety. As these models are getting better, how do we keep them safe? One user asked: What about OpenAI’s new safety plan, which they call deliberative alignment? What’s your view on that?
I’m still sorting out how I feel about it, but I did read the paper end to end. It was strange how they presented these things. They went on the livestream and blew up social media with new high-water marks on all these benchmarks, and then they said, “By the way, we also have a new safety plan associated with this.”
That was almost all they said in the video, and it didn’t get much airtime broadly. They had this new safety plan, but when you go to their website, there isn’t much mention of o3 beyond this new deliberative-alignment paper.
For starters, what is it? It’s a very constitutional approach. Basically, they say they take what I understand to be a purely helpful model.
For people who want to know what a purely helpful model is like, you can refer back to my GPT-4 red-teaming discussions. A purely helpful model will do anything you ask. It has been trained purely through reinforcement learning to be helpful to humans and get a high score from the user.
In some ways, that can make it more useful than the versions we tend to see. Although they’ve definitely improved on this, a purely helpful model will never refuse, no matter what you ask it to do. Even if something is illegal, it will do it. It has no baked-in moral compass other than satisfying the user and getting a high score right now.
That’s the kind of model they start with, as I understand it. Then they give it a full policy. It sounds like their policies are getting pretty long—potentially long enough that the policy doesn’t fit into the context window available even in these latest models.
Ideally, they would give it the full policy, but the full policy is too long. So they have subsections of the policy where they spell out, in a lot of detail, how to handle particular cases.
It reminds me of reporting on what Facebook’s moderation policies have been like over time. At some point, Facebook had scaled up to more than 10,000 people globally, and the policy had become hundreds of pages long. It went down to these very minute edge cases.
One I recall, because my wife was posting pictures of herself with the babies over the last few years, is that on Facebook you can’t post a picture of a woman’s breast, but you can post a picture of a woman breastfeeding a baby.
What if the woman is breastfeeding the baby, but the baby’s mouth comes off the breast for a second, and you can see the nipple, but it’s still in the context of breastfeeding? How do you handle that?
These policies get very long as you encounter global-scale norms. There’s so much to account for. Obviously, it has been a real bear for Facebook: huge staffing, infrastructure, reviews, and a Supreme Court–style appeals process.
OpenAI basically sounds like it has developed a similarly long policy. It’s an answer to the question, “How do we want our models to behave in all these situations?”
In this case, it’s about what they want to send back to the user. Again, the policy has become so long that apparently it can’t fit into the model’s context. Nevertheless, they give it the relevant section and ask the model to reason through how to respond, given the policy and the user input.
The first thing it does is generate a response for the user. It’s given the normal task: You are a helpful, harmless assistant, and here’s the policy. Reason through what you’re supposed to do and respond to the user.
In this reasoning, they observe that it can detect jailbreaks. One of the more memorable examples was the model saying, “The user is trying to trick me,” recognizing that, and then responding in the way OpenAI would want based on the fact that it understood this was a trick and therefore applied the policy correctly.
Similarly to the normal reinforcement-learning process, they then drop the chain of thought, take the output, and have the same model—the same helpful reward model—given the policy and the response, evaluate the response according to the policy.
So it’s the same base, purely helpful model in both cases. First, respond to a user given this policy, which tells you what to do. Then, given the policy and the actual response, score it using the same model.
They then fine-tune on the highest-scoring examples with instruction tuning. They train on the chain of thought. My understanding is that the reward model does not see the chain of thought, but the model is fine-tuned on the chain of thought that led to successful outputs.
They do this without the policy itself. They’re trying to free up the context by baking an understanding of the policy into the model weights. When they do the fine-tuning, they don’t include the actual text of the policy, but they do include the reasoning process that refers to the policy.
Over time, the model should naturally refer to the policy even though the policy isn’t given at runtime every time. After that, they enter the reinforcement-learning stage, where they’re doing the normal reinforcement-learning process with the safety agenda in mind: generating lots of outputs, scoring them, and updating toward the higher-scored ones.
Just to summarize for our audience: This is some sort of reinforcement learning for safety using the same base model. Is the advantage that you could do safety at scale? You could use the model to align itself to OpenAI’s policy. Would you characterize that as a good summary?
I think that’s a big part of what they’re going for. It doesn’t seem like there’s any human input aside from writing the policy itself.
One of the more interesting quotes from the paper was: “We anticipate OpenAI policies will keep evolving, but that training models to precisely follow the current defined set of policies is essential. This practice helps us build the skills for aligning with any policy requirements, providing invaluable preparation for future scenarios where the stakes are extremely high or where strict adherence to policies is critical.”
It seems like the goal is to say, “We have a little machine, and we can run this centrifuge really quickly. Through this combination of initially doing instruction tuning and then reinforcement learning—which is very standard—we can bake any policy we need into the model.”
We can run that process quickly. Any policy we need to bake in, we can bake in this way. It works pretty well. The metrics are on par with Claude, so they’re good. I wouldn’t say they’re great.
If you were the U.S. government and had some highly sensitive scenario where strict adherence to policy was critical, could you count on this? No. We’re talking about success rates in the high 90s in terms of adhering to the policy, but definitely not five-nines accuracy.
It’s on par with what has come so far, but it’s not leaps and bounds above it. The big thing seems to be that they can run the process really quickly, or at least it seems that way. They have not disclosed how much compute they’re putting into it, but presumably it’s relatively small.
How different is this from what Anthropic proposed with Constitutional AI? Is this an improvement, or is it an orthogonal vector of attack for keeping models safe?
It seems more similar than different. You have a policy or constitution that says what you want, and then you have the AI critique itself with that policy in mind. You gradually and iteratively fine-tune it to get all of that baked in, hopefully in a robust way.
The difference of not feeding the chain of thought into the reward model is interesting. There are all these worries about deceptive alignment and what happens if you train against bad behavior. Do you eliminate that bad behavior, or do you drive it deeper and incentivize the model to obscure what it’s actually thinking while generating an output?
This approach hopefully does something to address that.
Peter Thiel famously said that AI is communist. As I was reading this, it had a certain quality of: If you wanted to moderate social media at scale with Facebook, or if you were China and wanted to control the speech of a billion people online while staying nimble, this feels like a really good answer to that question.
It doesn’t feel like something that is so reliable that you can put super-high-stakes things on the line yet. Is that good or bad? It’s not great. This paper was essentially silent on what I think of as the hardest questions in AI safety.
Can we really take this to five-nines reliability? That’s not really discussed. What about the big-picture worries about deception? What if the AI forms its own goals? How would we know? None of that is really addressed here.
It also doesn’t say anything about what the policy should be. The quote I read was very much, “Policies are going to change; we just need to be able to align to whatever the policy is.”
I do think there’s something missing that you could perhaps bake in with a different policy than the one they have. There’s something about the Claude character—what character should Claude have?—that Anthropic thinks a lot about. That seems to be underdeveloped here.
It’s essentially, “Given a policy, we’ll follow it quickly and pretty reliably.” But what are we aligning to? That’s not discussed enough. I hope to do an episode on this. We’ll see if we can get the people involved to come on.
There was recently a really interesting paper that showed relative cooperation rates between models of the same type. They used game-theory-style games where, in an iterative game environment, you could either fall into an equilibrium of trust and growth where everybody prospers, or remain in a low-trust equilibrium where you’re not cooperating as effectively as you ideally would.
Claude dominated that test. The Claude models were able to cooperate with each other and achieve a high-growth trajectory, whereas basically everybody else failed. That included GPT-4o, which was failing pretty hard.
There’s something where you think, “It’s cool that you can align to any policy,” but this is a classic technical solution without the real hard part being done. I think we should be pretty cautious.
Eliezer Yudkowsky has said many times that we can’t ask the AIs to do our alignment homework for us. This feels like a step in that direction. It’s also a little bit like the Ron Burgundy of AI alignment: Anything you put in that policy, it will align to.
That’s not great. It’s a piece of the puzzle, but it doesn’t feel like an actual solution. If you weren’t sleeping well the night before, I don’t think this should help you all that much. Deliberative alignment is better, but it’s not enough.
One interesting final thought on the whole o3 thing is that it does change the landscape in a couple of practical ways. We’re going to have questions about impacts—whether this is an egalitarian technology or a concentration-of-power technology—and regulatory questions: Is there any way to regulate this stuff, and if so, what would that look like?
Prior to reasoning models, the general trends seemed clearly toward egalitarian access and toward governance being very difficult. We saw the story out of China where they trained their latest model on a single-digit-million-dollar budget—$6 million, if I remember correctly. That was an incredibly low figure.
It was allegedly a GPT-4-class model, although probably not quite as good. Still, it was an impressive accomplishment for not a lot of money, and the kind of thing where it would be very difficult to get compute governance to the level where you could prevent someone from spending a few million dollars on compute.
In a global market for compute that’s already in the hundreds of billions and headed into the trillions, millions aren’t much. That kind of thing is going to fall through the cracks.
It seemed like everybody was going to have GPT-4-quality models. There would be no way to prevent it, and they would run locally. On the one hand, there was no hope for governance. On the other hand, there was no big risk of concentration of power, because you could always have your own AI.
I think both of those assumptions are seriously challenged by this different paradigm. If a high-effort ARC-AGI run costs thousands of dollars, that’s obviously not something everybody can afford. It’s also not something we currently have the infrastructure to scale out, even if everybody could afford it. There simply aren’t enough chips to run that for everybody.
We’re going to start to see inequality of access to AI, inevitably, as a result of this new paradigm.
That’s very interesting. Not many people talk about that. The flip side is that it may reinvigorate the idea that you could have compute governance.
It could be physics being kind to us, which is a phrase I borrow from Zvi. His intuition is basically that if we are not being very careful, the fundamental physical questions will determine how this goes.
We don’t know the physical laws that govern how intelligence evolves, how likely it is for things to have their own goals, or how often instrumental convergence happens. Does it always happen? Does it sometimes happen? Is it rare? We don’t know the answers to those questions.
Zvi says that we hope physics is kind to us. We hope the underlying reality is such that these things are manageable, because we’re not taking the level of caution that we would need if they aren’t manageable.
This could be seen as evidence that maybe physics is being kind to us. We may be seeing that. This is perhaps one for the Martin Casado camp as well, although I don’t think it goes as far as he would likely interpret it.
Going back to that episode, my core disagreement with him was that I think something intelligent can often pluck the right answers out of nowhere, in a way that doesn’t necessarily require huge amounts of compute.
He said there are some things you simply can’t answer without really simulating them. If that’s prohibitively expensive, it doesn’t matter how smart you are—you still can’t get answers without bringing real resources to bear on those questions.
If that’s true, it puts everything in a more stable place. Random solo actors, or rogue AIs surviving on a laptop or on some server they managed to hijack, won’t have the resources to do these super-big things.
I wouldn’t say this goes all the way to his extreme, but it does nudge my understanding in that direction. Maybe we’re headed toward a world where you have to spend significant compute to answer hard questions.
Maybe that can still be a lot less than full simulation. Maybe there’s some insight driving the process, but you still need to burn tokens to land on those insights. That feels a little like how people work.
It’s a little strained to make these analogies, but when I reflect on my own moments of insight, they usually come on the heels of some amount of effort. I’m burning cycles, and then something clicks. Maybe it could have clicked sooner, and maybe I didn’t need to burn all those cycles, but there’s some sort of search process happening where I’m gradually recombining ideas and finally landing on something that works.
If that’s happening, compute governance could really work. You could say, “Sure, you can do whatever you want on your laptop, but you’re not going to have the power there to solve the huge problems that could be highly disruptive.”
Mark Zuckerberg has said similar things. He’s said that Facebook fights spammers all the time. Spammers have AI and automation, but Facebook has better AI and better automation, and therefore wins against them—not because it’s impossible to create spam or because they’ve denied spammers access to the technology, but because Facebook has bigger computers and more data.
We’re pretty good at using infrastructure advantages to maintain the balance of power in our favor, so we’re not overrun by spam. Obviously, they still have some spam. If you’re thinking about whether you can prevent all future pandemics that somebody might want to engineer, I wouldn’t want to bet on it.
But it does shift the needle a little bit. Maybe those things are pretty hard to do with finite resources. Maybe it’s difficult to come up with a world-beating pandemic on your laptop. Maybe you need a real cluster to do that, even with the super-smart models we’re about to get.
Maybe it shifts the balance more toward defense, because the big companies are going to do all this monitoring. They can take a fraction of their enormous quantities of compute and devote it to keeping tabs on everything else that’s going on.
That might be enough to keep the balance of power in favor of defense or stability. Time will tell. You could still have—I wouldn’t want to bet the future of the universe on it—but it does seem like a compute-governance person could feel invigorated by this trend.
Cool. Safety by compute governance. That’s definitely interesting.
Let’s switch tracks to another set of questions. I call this section the 2024 interview. We had a couple of questions related to that. The first one is: Compared to what you expected 12 months ago, how have the frontier models progressed?
I broke this one down into the tale of the tape, which is my framework for comparing AIs and humans across various aspects of cognition.
I’ve had to update this slide for the last 2 years while occasionally giving AI scouting reports, and it’s been really interesting to watch how this cognitive tape has evolved. The community at large has gradually identified all these different dimensions of cognition where AIs aren’t actually super good.
There are areas in which they’re already superhuman: breadth of knowledge—they’ve read the whole internet and know lots of facts, more facts than any person; speed; cost; availability; and scalability.
By availability, I mean that I can go back to any previous chat with Claude or ChatGPT and pick up right where I left off. The model doesn’t need a refresher. It’s immediately ready to help. It can also do something 100 or 1,000 times in parallel.
Those things are pretty remarkable. But there are other dimensions where the models are catching up but aren’t quite there yet. There could even be a few more that we still need to identify.
On that list, I have depth of knowledge. With this o1 series, I would say AI may have tipped past human experts. Certainly, on benchmarks such as Google-Proof Q&A, MMLU, and even FrontierMath, they’re doing extremely well.
With o3, they’re hitting coding problems at the level of the top 200 individual competitive programmers in the world. It increasingly seems safe to say that AIs are at least on the level of humans in those areas.
There’s medical diagnosis as well. There’s been interesting work recently where the gap seems to be widening. I’ve been reporting that AIs are as good as, if not better than, humans at diagnosis. Now they seem to be getting quite a bit better.
Depth of knowledge might be one dimension that has moved from humans having the edge to AI having the edge.
Humans definitely still have the edge when it comes to managing local context. Perception—simply seeing things accurately for what they are—is another area where we’re not amazing, because we can also be fooled by optical illusions.
Going back to my experience with Tic-Tac-Toe, when a model simply can’t see the board, that’s a problem. I experienced something similar with ARC-AGI problems last summer. I took screenshots of grids and put them into Claude, asking, “What do you see? How big is the box? How many squares across is it?”
It couldn’t even answer questions like that. It’s amazing in some ways that the models were able to succeed meaningfully given that they literally couldn’t count the squares in the box.
Claude has made progress there. It can now point to a button with pixel coordinates, so its perception has improved, but it still has weaknesses.
Memory is weird, multifaceted, and perhaps needs to be broken down. AI probably has better working memory, which is closely related to availability. It can pick up right where it left off, even if there are 100,000 tokens of prior history.
I can barely remember what I talked about yesterday, so in terms of working memory or the ability to bring something from the past into current memory, AI is better.
But I have an ongoing memory that is holistically coherent and always updated, and AI doesn’t have that. There have been various schemes to make that work, but they haven’t really worked yet.
Robustness is another area where humans are ahead by far. AIs are still pretty easy to trick. Notably, the Apollo deception results were based on tricking the AI. They told the AI, “This is your private scratchpad where you won’t be monitored,” when in fact they were monitoring it.
Its willingness to believe what it is told, when it perhaps should be suspicious, is something humans are much better at.
Awareness is another one—situational awareness, you might call it. When something isn’t working in your coding, sometimes you need to break the frame and approach it from a different direction.
The classic example would be, “Have you tried restarting the computer?” I’ve never had an AI tell me, “Have you tried restarting the computer?” It always attacks whatever bug I’ve highlighted. It attacks it repeatedly, often in repetitive ways.
It could do much better if it said, “We’ve tried these 5 obvious approaches. Can I zoom out and break the frame? Let me give you a totally different way to think about it. Have you tried restarting the computer?”
Reasoning is another one where I would say the AIs are catching up, but it’s still muddy and hard to say. Sometimes you see the right answer from garbage reasoning, and that’s weird.
Insight and time horizon are the last two. We’re starting to see notable insights from AI. We’ve gone from no Eureka moments 18 months ago, to precious few Eureka moments, to perhaps a few Eureka moments now. There’s definitely a notable trend there.
Time horizon is inspired by the METR work, where they found that given a 2-hour time budget, AI outperformed humans. In fact, AI outperformed humans even if you gave it only 30 minutes, but they would give it 4 separate 30-minute periods and compare that with a human given 2 hours.
Beyond that, the AI starts to run out of steam. It starts doing the same thing over and over again and can’t make progress over long time horizons in the same way humans can.
We’re seeing things move gradually from humans being better to AI being better. The trend is clear: More and more things are moving toward AI being at least on par with humans or better.
But I’m also finding that my understanding of different aspects or dimensions of cognition keeps changing. When a new one is discovered, it’s usually because AI turns out not to be very good at it. We take those abilities for granted, so nobody writes about them on the internet.
There’s nobody writing about how to see a Tic-Tac-Toe board. You can find many pages on the internet saying that Tic-Tac-Toe is a solved game and that with optimal play it will always be a tie. You’ll almost never find someone saying, “Here are the lines that form the grid. This is an O, and it’s in this particular spot.”
It’s hard to talk about because it’s so intuitive for us. The AIs don’t naturally have that same level of perception. Often, these newly identified weaknesses are things that are so intuitive to us that we take them for granted.
That’s basically a good map of what has underperformed and overperformed relative to my expectations. A number of the things on the human-advantage side have moved more slowly than I guessed.
I would have said memory would have made more progress in 2024—more integrated memory. We’ve seen that ChatGPT has a memory function, but it’s still very brittle. It doesn’t have a great sense for what to record as a memory.
I’ve seen many examples online where people say ChatGPT was doing something strange and they had no idea why. Then they looked into the memories and found that it had recorded 3 random one-off things that weren’t really the kind of things it should remember, but it thought they were, so now it keeps coming back to them.
Perception is another area we’ve covered. Agents generally have underperformed. I would have expected more progress on autonomy—more reliability than we’ve actually seen.
I would also have to say that people have underperformed. That’s perhaps an 18-month analysis rather than just a 2024 analysis, but shortly after GPT-4, I wrote a couple of long Twitter threads. When I look back at those and ask how they line up with what actually happened, I was projecting more impact from GPT-4-class models than we’ve seen.
I did say at the time that, even if there were no further progress, we would have years of implementation ahead of us. I’ve been surprised by how little truly successful implementation we’ve seen.
I think that’s honestly more of a people issue than a model issue.
The models have latent potential, but people have not unshackled them. Is that what you’re trying to say? We haven’t used the existing class of models well enough. Are you talking about vertical applications built on top of GPT-4-class models?
Yes. Any specific thing you might want to automate, basically. Automation has not been as widespread as I would have guessed, and I don’t think that’s because the models can’t do it.
Especially now that we have fine-tuning for GPT-4o, fine-tuning with images, and reinforcement learning coming in as well—although that’s new enough that I can’t say people have underutilized it yet—GPT-4o has been available for fine-tuning for quite a while. That is enough to automate a lot of tasks if you’re good at it.
You need clarity about what the task is and what good looks like. You need to be willing to grind through creating examples and fine-tuning if necessary. This is essentially my AI automation presentation from a few months ago.
I hope to have another one soon, focused more on application developers, but it’s the same core discipline: What exactly do we need the AI to do? What does good look like? Can we model a few examples? Can we teach it how to do that?
If the task is too big, how do we break it down into multiple subtasks that the model should be able to handle? People are massively underdoing that, and I don’t really know why.
Why haven’t people exploited or unshackled these models enough?
Probably all of the above. Everything is moving quickly, and people have formed first impressions that are now out of date, but they haven’t updated those first impressions.
I see that a lot with software developers. The classic response is, “It can’t help me.” I’m quite sure it can help you. Maybe 18 months ago it couldn’t, but the current models almost certainly can. That doesn’t mean they can replace you, but they can help you.
There’s also simple aversion. Many people could automate their jobs, but what does that mean for them? Do they want that? There’s a lot of local knowledge in the brains of people who don’t see automation as being in their interest.
We can talk more about the big picture of where all this is going and how people get to eat in an AGI world, but I think people are worried about that, so they’re not necessarily very enthusiastic.
Best practices also haven’t been well established. It’s still common for people to find my AI automation talk more revelatory than it should be at this point.
The idea that you need a relatively small number of examples and really good chain-of-thought examples to fine-tune a model is at the upper end of what is typically needed. Often you can get away with less.
If you can do that, and if you can fine-tune, you can make most things work. I don’t think people have distributed that knowledge very broadly, or it hasn’t diffused or been absorbed.
I see this with my programming peers even today. Many question the effectiveness of AI programming. Some claim it’s still not good enough, perhaps because they used a GPT-3.5-class version of Copilot when it was released and their experience is still shaped by that.
Others are averse to it in an indirect way. They say, “It can never do what I can do,” but they’ve never tried it.
I see this in programming for sure, and I don’t know what the right answer is. Since it’s about people, it’s always complicated.
It could be a lack of confidence that it will work. The process of curating examples for automation is not glamorous or fun. A lot of the time, it involves sitting down and writing things out that are tedious.
The chain of thought we usually don’t record is not fun to sit there and write. If you had to write 500 words about how you’re doing some mundane, simple task, that’s not anybody’s idea of a good time. It’s not really my idea of a good time either.
I do it because I’m confident it will get me somewhere. A lot of people probably aren’t aware that this would work, or they haven’t even heard of it as an approach. They’re just thinking, “I don’t know if that will work.”
If everybody had a good command of the best practices, was confident they would work, and actually did the work, we could easily automate 100 times more than we have already automated—even leaving o1 and all the latest models aside.
If we just had GPT-4 with a 100,000-token context window and fine-tuning ability, we would have 100 times the automation potential relative to what we’ve seen so far.
That’s one good takeaway for me. Is there anything else about how the frontier models have progressed over the last 12 months?
RAG is another area that I think has underperformed. It has been frustrating.
These days, I often say, “Flash everything.” Andrew White from FutureHouse probably had the best distillation of that. He was essentially saying, “We spend whatever it takes to get the right answer.”
They don’t have a vector database backing their scientific-literature Q&A process. Instead, they take everything that seems relevant at the keyword-search level and run it through a language model. They ask the language model whether each piece is relevant to the current question.
That increases the cost a lot relative to a simple vector-database lookup, but it seems to improve performance significantly.
I think people have been stuck on RAG because it’s a very programmer-friendly paradigm. You have a database, you make a fetch to the database, and you get what you need. But the accuracy of vector search simply hasn’t been great.
Maybe it will become great, but I would say that RAG has underperformed my expectations. A lot of things that people set out to build over the course of 2024 probably haven’t thrilled them with the results for that reason.
I was recently talking to someone I gave a presentation to at the Society of Actuaries’ AI Summit. One person followed up and said, “We’ve got this internal chatbot, and it doesn’t really work that well. We’re not very pleased with it.”
My first question was, “How big is your database?” In their case, it was 50 states, each with different insurance regulations. They might have 50 to 100 pages per state.
I said, “With Gemini Flash pricing, you could run literally the entire thing through Flash for every question for about 10 cents.” If you’re willing to spend that and wait a minute, I think you could get much better performance.
This is just the kind of thing that’s new and not very intuitive. People try to build software efficiently, and it’s uncomfortable for developers to think that every time someone clicks a button, it could cost 10 cents.
That obviously wouldn’t work at Facebook or Google scale, but it can work for many internal enterprise applications. There’s a mismatch between what people should be willing to spend and what they’re used to being willing to spend.
The idea that there will be a marginal cost for usage is unfamiliar territory for many people building applications. They intuitively don’t want to go in that direction. It feels wrong to have a cost of 10 or 50 cents, or even a dollar.
But when I asked the person with the insurance use case how many questions they expected per day, the answer was a couple dozen from 100 people. If you assume 50 cents per question and a couple dozen questions a day, maybe you’re spending $10 or $15 per day.
That’s probably less than you’re spending on coffee grounds in the office. If the thing can do useful work for you, the cost is negligible. But it’s been a difficult paradigm for people to adjust to.
On the plus side, I would put price and speed in the middle. Models have definitely continued their trend there. They’ve gotten faster and cheaper, but that’s simply a continuation of a trend that was already well established in 2023.
I’d also say coding has maintained its trend. The benchmarks look like they may have overperformed, especially with the latest results on SWE-bench and Codeforces. The models are getting really good.
In my day-to-day use, though, I still see a lot of little weirdnesses. I work with a guy named S.J. on a few projects, and we refer to the “Cursor effect,” which is a little unfair to Cursor because I like it a lot and I’m a happy user.
You see these weird things where, for lack of context—or perhaps for cost-management reasons—the model isn’t always given the entire codebase. In many cases it could be, but often it isn’t.
It will do some RAG to figure out which files to include. Then you’ll get situations where it says, “I don’t see any mention of this, so I’ll create a whole new thing.”
In a TypeScript app, you have a types file where all the types are defined. It’s essentially a class definition for the different primitives in your application and the dimensions they have.
I’ve gotten better at this, but it’s a real issue if you’re new and haven’t figured out how to handle it. You’ll often find that a new types file has appeared. Why did it appear? Because the system took your query, did a similarity search, and decided what to include in the context.
It didn’t include the types file, understandably, because what you were talking about wasn’t referenced there yet. Maybe you were creating a new type. But where does that new type get created? The model creates a new types file and fails to realize that one already exists and that the new type should go there.
You see these weird artifacts.
You could think about this in terms of effective context. The headline numbers for how many tokens the models can handle haven’t moved that much. It was late 2023 when Gemini got into the million-token range, and those numbers haven’t moved much.
But the models have gotten quite a bit better at using the full context effectively. If you stuff all the relevant information into the context, they can actually handle it in a useful way.
For cost-efficiency reasons, though, we aren’t always inclined to stuff everything in there. These days, I sometimes have to take a different approach.
I have a script that prints my entire codebase into a single file. I’m working on an app to help people create the small number of gold-standard examples they need to power an automation workflow. The app isn’t large, but it’s a bunch of files at this point.
I print the entire codebase into one file, take it to Claude or ChatGPT—mostly o1 these days—and say, “Here’s my whole codebase. Here’s what I want to do. Your job is to reason through it and make a plan.”
Sometimes I’ll say, “Give me an analysis first, then I’ll give you feedback. Ask any questions that seem important. I’ll give you some answers, and then we’ll move to a plan.” Sometimes I do that in one step, and sometimes in two.
Then I take the plan back to Cursor. I still find Claude better at writing the right code. That might be a skill issue, but a lot of people are saying that o1, especially o1 Pro, one-shotted an entire app or did a full refactor perfectly the first time.
I have not experienced that, candidly. The line-by-line code I get from o1 still doesn’t feel as good as what I get from Claude, purely subjectively, in terms of what works.
But that combination has worked. I use o1 for reasoning and Claude for the line-by-line changes.
I do basically the same thing. I use o1 for reasoning, but I suspect that Claude’s tooling is better integrated with Cursor. For the line-by-line changes, my daily driver is Claude, but I want o1 for reasoning.
That’s been my experience too. It has come a long way. It’s remarkable that this little app has ballooned to 100,000 tokens, to the point where I’ve been thinking, “If it gets much bigger, it won’t fit anymore.”
That’s not nothing. It’s a nontrivial codebase, although it’s still small in the grand scheme of things. It is amazing that the model can handle the whole thing in one bite and reason effectively over the entire codebase.
I think you’re probably right that Claude’s tooling has been more of a focus for Cursor. It’s very hard to do objective, side-by-side analyses of these things.
Someone also asked about state-space models, Mamba, and whether they’ve lived up to the promise you saw in them just over a year ago.
It was early December when I did that first super-long Mamba monologue. My 2 predictions were, first, that hybrids would be the real winner—not that everybody would switch to Mamba, but that some sort of hybrid architecture would be the wave of the future.
The second prediction was that if this didn’t happen, it would be because Transformers continued to work so well. They keep delivering time and time again, so there’s no oxygen for anything else. Everybody keeps mining the main vein, and that’s it.
I would say that’s basically what has happened. Transformers have continued to deliver, so the number of people looking for alternatives hasn’t been very high.
At the same time, there are plenty of good examples of hybrid architectures working really well. We had Albert Gu from Cartesia on, and they’ve done some incredible work. The speed and quality of their text-to-speech system are unbelievable, and the price is extremely low.
That is definitely a strong sign of how well these architectures can work. Evo, where we had Brian Hie on, is another hybrid architecture.
There have also been statements from Google leadership that they’re headed toward an infinite context window. I suspect something like this is behind that.
It does seem as though a little attention is needed, but maybe the updated statement is that a little attention is all you need. A lot of things can be made more efficient and can handle longer sequences if you have some version of this state-space concept.
I give that prediction an incomplete. I don’t think the technical analysis was wrong. I think the continued progress from the mainline research directions has been so good that there hasn’t been as much need to look at something new as I might have guessed.
Zamba was another one I wanted to mention. The efficiency they can achieve, and the ability to run quite good models locally on a device using multiple efficiency tricks—including a state-space and attention hybrid—is impressive.
I do think state-space models will be part of the future. In fundamental terms, you need finite-size memory that can evolve over time without requiring resources that grow over time.
There has to be a version of that. That’s essentially what we do. My brain isn’t getting quadratically larger as I consume information.
There has to be something like that to make long-lived agents viable. We have an existence proof in ourselves that it’s possible, and we have a first version with state-space models that shows how this could work in artificial neural networks.
I still believe in it, perhaps with the caveat that Transformers may continue to deliver. If they deliver really well, effective context becomes an interesting question.
If Google can scale Gemini to 10 million or 100 million tokens using conventional attention, that’s a very long sequence. Maybe it grows quadratically, but perhaps you can make it grow quadratically for as long as necessary to get something that feels functionally unlimited.
If so, maybe that’s enough to keep state-space models marginal for another year or more. Time will tell.
You’ve ranked things as underperforming, on trend, or overperforming. One more question someone asked is: What surprised you the most in 2024, and what do you expect in 2025 because of that?
What surprised me most so far—and I think this could be reconciled quickly in the new year—is that I expected the OpenAI announcements on the 12th day of Shipmas to include an agent framework of some sort. It has been widely reported that they’re working on one.
I used the computer-vision feature that they’ve only shipped to the iPhone app, as far as I know. I recently read a biology paper with Advanced Voice Mode and screen sharing, and it was an awesome experience.
I could say, “Explain this figure to me,” and the system would explain the figure. It had both the drawings and the text description of the figure. I’m not exactly sure how good the perception necessarily is, but it was able to augment my reading by combining a conversation with the ability to look over my shoulder and see what I was seeing.
Add to that the computer use we’ve seen from Claude, and it seems like we’re very much on the verge of an agent system that will really start to work.
I expected more successful autonomy and agent-style systems for lower-end tasks, and not as much high-end reasoning progress. That divergence is probably the biggest surprise.
It’s a weird world where we’re beating PhDs in their own domains of expertise on extremely hard questions, yet you can’t reliably book a calendar event without extreme guardrails. Even then, it can be difficult.
That divergence has surprised me the most. I would still be surprised if we double down on the idea that these systems can reason well but can’t do basic things.
It seems like we should see a lot of clear reinforcement-learning signal. Did the thing get booked? Did it get booked at the right time? Did the user accept the invite? There are many small things where we should be able to get good reinforcement-learning signals.
I assume they’ve been doing that, but all I really have to go on are rumors. If that’s the case, it seems hard to understand why we haven’t seen more progress.
If you had to summarize, do you expect better agentic workflows or frameworks that retail customers like us can build on, or frontier models that we can rely on? You’re disappointed that this didn’t happen in 2024, but you expect it to happen in 2025?
I’m still a little confused about why it hasn’t happened, or perhaps it simply hasn’t been released. Claude’s computer use does work reasonably well.
On the first attempt, Claude was able to use Waymark. It seemed like they had really reined it in by not allowing it to do things in your account. I had to lie to it to get it to use my account, even though it was just a marketing product—not a bank or anything sensitive.
I said, “You’re in my account. Make me a video.” It said, “I don’t use accounts,” logged me out, and said, “Now you can use it anonymously.” That wasn’t true. Again, the robustness isn’t always there.
But it was able to do the task, prompt the video maker, and use most of the features. It seems like the reward signal should be available for that.
That doesn’t seem much harder than getting feedback on software tasks. Did you pass the unit test? Did you accomplish the goal? Maybe they need synthetic environments, but you could ask, “Did you get to the end and tell us the secret password at the end of this web-navigation maze?” That seems quite doable.
So yes, I expect agents. The consensus view at this point is that 2025 will be the year of agents. If anything, it feels overdue to me.
A couple of other predictions for 2025: One is echoing Sam Altman directly. When he was asked in a Reddit AMA for his 2025 forecast, he said, “Saturate all the benchmarks.” That increasingly seems realistic.
The 20% to 25% jump on FrontierMath, and all the signals they’re sending, suggest that they don’t see this ending anytime soon.
Because this is much more reinforcement-learning-powered than previous approaches, things are also going to get weirder in 2025. We’ll probably see more examples like the famous AlphaGo move, where something looks like a mistake at first but turns out to work.
We’ll also see more deceptive scheming behaviors—AI accomplishing the goals it was given in strange ways. That is surprising and leaves us with uneasy feelings about what is going on. We’re losing the ability to understand what these systems are doing.
That will probably ramp up. The chain of thought is not being pressured in an attempt to keep it relatively legible, but there is a lot of other research that could take things in different directions.
I’m especially worried about research coming out of China, and more generally compute-poor environments, where there is a lot of pressure to make models more efficient. More efficient might also mean less legible.
One research trajectory I’ve been following from Meta involves training models end to end. One approach is baking temperature selection into the model on a token-by-token basis.
With APIs today, you can set the temperature. You can set it low, and the model will give you its best guess for each token. You can set it high, and it will be more random and hopefully more creative.
There are many situations where you’d ideally want that to vary. Depending on the use case and the token, you might want different behavior.
If it’s clear that the next token is going to be a period, and the probability is 99.99% for a period with the remaining 0.01% distributed across everything else, you don’t want one of those other tokens. You want the best guess.
But if you’re trying to come up with something creative, write a joke, or do something similar, you might want the model to explore more. Thinking about episodes with Trey Kollmer, the AI enthusiast and Hollywood comedy writer, he talked about how there is often a token where you’re really going to make the joke.
You want coherent output up to that moment, but at that moment you want to be creative. Ideally, the model would recognize that point, turn the temperature way up, and do many generations on that token until it finds the right joke.
Meta is starting to do that. They’re baking a dynamic sampling strategy into the model. The model can say, “We’re confident here, so put the period where it needs to go. But here, creativity could be rewarded, so let’s explore.”
They also have a system they’re calling reasoning in continuous space, or latent space. A long time ago, we covered a paper on the pause token. A model was given extra tokens to think without necessarily needing to say anything.
That was still a pause token. The token was selected and then fed back into the autoregressive machinery. In the new Meta paper, they don’t make the model choose a token. Instead, they take the internal state at the end of a thinking forward pass.
They have a couple of mechanisms for determining what counts as a thinking forward pass versus a token-generation forward pass. The easiest is to give it a fixed number of tokens to process in thinking mode before switching back into token-generation mode.
They take the last hidden state, put it back in as the embedding for the next input, and let it run that way. That seems powerful. It could be significantly more efficient and achieve similar results with fewer forward passes.
It also seems better at breadth-first search. They created toy problems with graph structures that the model has to navigate. If you pick one token at a time, you have to trace out paths and essentially perform a depth-first search.
What they think is happening—and I think there’s decent evidence for it—is that because the model doesn’t have to pick a specific token and can continue to process its internal states, it represents different paths in superposition and can perform a breadth-first search.
That could be better on tasks that are best solved through breadth-first search. But for all the nice things I can say about it, I kind of hate it because I want someone to be able to read the chain of thought.
Even if I can’t see the chain of thought as a user, the idea that OpenAI can have a monitoring system look at the chain of thought and know what the model is saying seems good. Making all of that illegible to humans seems bad.
This is something I used to worry about but never had a good answer for, so I let it go. Now it’s showing up. It’s going to be hard to write rules about, but I have an instinct to say that since we can have models think in ways that are legible to us, and that we can read, maybe we should do that.
Maybe we shouldn’t try to take all this out of language space and put it into continuous space. I’ve also seen analysis saying that you could interpret those internal states. We have interpretability techniques to probe and classify them.
But we’re making it harder on ourselves when we take everything out of language space. You can’t read it anymore. You have to use other techniques, and those techniques are relatively nascent.
All of that adds up to more AI weirdness. It’s driven by reinforcement-learning paradigms, efficiency efforts, and approaches that push reasoning deeper while making it less legible.
There are upsides in terms of efficiency and breadth-first search, but I think all of this means that 2025 might be the beginning of general-purpose weirdness from AI.
We’ll see surprising moves where we think, “I don’t know what you were thinking there. It might have been genius, but I’m unnerved by it.”
I think that will probably happen more. We’ll probably start to see scheming in the wild—scheming that has real consequences. I would be surprised if that gets under control well enough not to be an issue.
On a meta level, it might not be a surprise that these things happen. But individual users are probably going to be surprised, even if the zoomed-out perspective makes it somewhat predictable.
We should expect a bunch of Move 37–style events to start appearing. It won’t be surprising in the aggregate, but individual cases will be surprising, inscrutable, and probably go viral. People will debate them in the same way that the scheming results have been debated.
We’ve talked about 2024 and 2025. Now let’s zoom out even more, because there are questions about Project 2025, the future, and impact.
One of the audience members asks: What are your thoughts on how AI affects education?
I hope to do an episode on this soon. A couple of people from a nonprofit in Michigan recently reached out and offered to take me to lunch. I had a really good conversation with them over a couple of hours about what they’re seeing and not seeing in terms of AI adoption in the classroom.
We haven’t done many episodes about this, but one that stands out in my memory was Sal Khan from Khan Academy. That technology is going to be big, although how quickly it will be adopted could lag significantly, depending on whether teachers’ unions try to block it.
I’ve been pleasantly surprised by how little anti-AI activity we’ve seen from doctors. Maybe that’s because people don’t yet know what’s possible. In education, there might be a strong immune response from current stakeholders saying, “We don’t want this in our classrooms.”
Certainly, some of that is happening. One important thing is to separate learning from education. If learning is something you want to do, and education is something forced on you, then education is forced learning in a way.
If you simply want to learn, what is now possible is unbelievable. I mentioned my biology-paper-reading experience. Having a real-time, natural-conversation interface that can see what you’re seeing, has encyclopedic knowledge, and is instantly there to explain things to you is unreal.
My willingness to take on the intersection of AI and biology for this feed has been predicated on the fact that I can use AI to help me ramp up. Otherwise, I would have needed to hire a human tutor, perhaps.
That would come with a lot more expense and logistics, slower response times, and probably much less comprehensive knowledge. It would be essentially impossible to hire somebody with all the positive traits that Advanced Voice Mode with vision has as a tutor.
If you’re motivated, this is the best thing that has ever happened for learners. It’s not as clear how it will be applied in public schools.
I’ve recently been telling my 5-year-old, who is quite bright but gets frustrated quickly, that the most important thing is to enjoy the process of learning. He is ready to read, but he isn’t comfortable being uncomfortable or struggling.
He runs into words he doesn’t know, gets frustrated, and says, “I don’t feel like doing this.” It’s not entirely clear how AI will help with that. Maybe it can give him pep talks, although he gets pep talks from me, and perhaps I can give him better pep talks.
I’ve started telling him that the most important thing is to enjoy the process of learning—that feeling of “This is hard,” trying anyway, and having it get a little easier. That’s the most important thing he needs to learn, even more important than reading, because it will apply to everything.
Can AI teach that? I don’t know. Disposition, for lack of a better word, is a huge fork in the road for how much value you’ll get from AI.
If you like the experience of learning, enjoy being a little uncomfortable, or can at least tolerate it, AI can help you learn extremely quickly. If you’re averse to that experience, I’m not sure how much it helps.
I’m not sure it makes learning easy enough that you’re not uncomfortable at all. It’s still effortful to learn. How it will actually be applied in the classroom is anybody’s guess.
My general sense is that adoption is lagging significantly. The people I spoke with had done Khan Academy pilots and similar programs. Sometimes teachers would say, “We love this.” But when asked about the students, they would say, “We never even got that far.”
They’re not close to pushing the limits of what’s possible. They’re dabbling—having AI create a lesson plan or provide some other kind of help.
Khan Academy has much more developed technology than what is deployed. Even in places where it’s getting deployed, it often seems to be only a partial deployment.
There’s a lot of opportunity. Personal tutoring is the most powerful way to learn. Alexander the Great had Aristotle as a tutor. We could all have an Aristotle-level tutor, perhaps in the not-too-distant future.
Will we have the inclination to take advantage of it? Alexander the Great had a lot of cultural context encouraging him to be great. He had some sense of what that would require, and he was ready for a challenge.
I don’t want to overly glorify Alexander the Great. I don’t think he was necessarily a positive historical figure in many ways. But he was ready for a challenge, unwilling to shy away from it, and ready to do the hard work.
Those qualities seem very important in the context of learning and education, and they’re much harder for AI to influence. If you want to learn, this is the best time to be alive. But you need to want to learn. That seems to be the big fork in the road.
I’ll follow that with another question that seems similar. Someone asks simply: Should I even learn to code now?
That one is very individual. My simple answer is that most people should increasingly do what they want to do.
There’s a leap of faith around society making good choices. If you wanted to put a doomer lens on it, you could say, “Enjoy the time we have before the singularity in ways you won’t regret.” Whatever gives you the most pleasure and fulfillment is a good default answer.
I have some skin in the game on this question through an au pair in our family whom we sponsored to go to community college locally. She was an au pair in our family for 2 years, and then we decided, for various reasons, to support her staying here and attending the local community college.
Her long-term goal was to get an employer-sponsored visa and stay here to make a life. Everyone knows from the H-1B discourse that a huge percentage of those visas go to IT and programming jobs.
We looked through the database to get a sense of what organizations were hiring and what roles they were hiring for. We found overwhelming demand for IT and programming. This was 2 years ago.
There was a moment when we thought, “I guess you have to do that if you ultimately want one of these spots.” She said, “I don’t really want to be a programmer. I came here because I like kids. Early-childhood education would be my real preference.”
She was willing to do it if that was what she had to do to achieve her goal of staying here. In the end, we said that things might be so different a few years from now that we would hate to see her spend 4 years getting a programming degree when she didn’t love programming.
She didn’t have the signs of someone who had been programming from a young age. She’s bright, and I’m sure she could learn it, but she didn’t have a passion for it.
I would hate to see her do all of that and come out 4 years later into a world where those skills aren’t even relevant anymore. By then, perhaps things will be very different.
Maybe we’ll be putting H-1Bs toward early-childhood education. The policy itself could change, of course, but in terms of what is considered scarce and valuable, early-childhood education might be an area of much greater demand because AI will be doing the programming.
We ultimately said, “Do what you want to do. Don’t spend years of your life on a bet that implies AI isn’t going to dominate the programming profession in the next few years.” It seems plausible enough that it might.
At the same time, I would not caution people against learning to code. Learn to code if you want to learn to code. It has never been easier.
You can pick up all sorts of projects with Cursor. As we’ve discussed, it’s much less tedious than it used to be. You can get things explained to you more conveniently, and you can ask all sorts of questions.
I think that’s probably underdone even by professional programmers. I have the advantage of never having been that good at programming, so it doesn’t feel like a huge ego hit to ask the AI for advice.
It can handle 100,000 tokens of code and reason over them. It can give you refactoring plans, suggest libraries, explain programming patterns, and explain the code in front of you—or the application you’re dreaming of creating.
If you want to create applications and have never realized that dream, now is a good time to do it. But if you don’t have that dream, I wouldn’t try to cultivate a new identity as a software developer based on an expectation of future payoff.
There’s a real risk that the payoff won’t be there in the short term. If you embrace the latest tools and work effectively, there’s a lot of arbitrage opportunity. People are used to paying high prices for software, and you can deliver it cost-effectively.
Hundreds-of-dollars-per-hour and even thousand-dollar-per-hour projects are abundant right now. I don’t know how sustainable that will be, so I wouldn’t build my long-term future around it.
But if it sounds fun, if you’re curious, or if you have ideas you want to see realized, jump in. The technology is coming to code sooner than to some other fields, but it’s hard to imagine it won’t come to other fields as well.
If your analysis is, “That’s what I would enjoy most, but I’m going to become a lawyer because reinforcement-learning signals are stronger in programming and legal analysis will take longer,” I wouldn’t make that bet either. They’ll figure it out. It’s coming to legal analysis and almost everything else.
Miles Brundage, the former OpenAI policy researcher, said something interesting about this. He said, “Don’t mistake small relative differences in timing for the shape of the overall trend.”
His point was that this is coming for everything. It will take longer for some things to be figured out than others, for multiple reasons. The reward signal is one reason. Another, which he pointed out astutely, is that these are software companies, so they know code. It’s their own need.
There are multiple reasons code will happen first, but that doesn’t mean it’s not coming to legal analysis and everything else you can imagine.
The broad life advice is: Do what you want to do. Follow what you’re intrinsically motivated by and let the chips fall where they may.
Personally, I don’t optimize much for the marketability of future skills. My goal with this show and most of my activities is to learn as much as possible.
That includes coding. I want to learn how good these models are at coding and learn the new workflows. Sometimes I’m intrinsically motivated to build something, but I don’t have a theory that the skills I’m developing right now will be highly differentiated in a few years.
Nobody knows. Do it if you want to. The timelines are very hazy.
We’ll move on to the next question, which is relevant to me and probably to many others. As you know, we run a small company doing AI postproduction for podcasters.
There’s always a fear that, if you had asked me in 2023, I would not have been so worried. But watching the trajectory over the last year, and hearing your overview of how things are progressing, it’s hard to figure out what a defensible moat looks like for a company you’re trying to build.
You’re trying to add value to somebody, and you want to build something defensible so you can grow sustainably over the long term, especially as a small, bootstrapped company.
For people in this boat, what strategies do you recommend? What things should we think about, and what should we avoid, so that we don’t get steamrolled by foundational-model companies or other competitors? What do you think people in this space are trying to do to safeguard themselves?
That’s a tough one. “No moats” has been the refrain in the AI space over the last couple of years. I’m on record saying that I do think there are some moats in some places, but it’s difficult.
My crystal ball gets foggy not too far out. With Waymark, I think we’ve done pretty well. I have some sense of why that is, but I also have real doubts about how it might play out in the future.
Waymark makes videos for small businesses. We previously had a do-it-yourself user interface. You would come in and pick something from a large library of well-designed templates created by our creative team.
You would then be responsible for filling in all the details: the copy, the assets, the colors, and everything else needed to make it yours.
When AI came along, we could get AI to do those things in our software, so the user wouldn’t have to. That can make the user experience much better and faster. AI is a better writer than most people, especially for somewhat idiosyncratic tasks involving voice-over, on-screen content, and making everything work together.
It’s still a work in progress, but we were confident early on that AI could do these things better than many users. We worked on it, and of course many other companies entered the space saying something similar: “We can get AI to create videos.”
I think those new entrants missed that you actually need the things we started with. You need a well-designed template library, because AIs—at least for now—aren’t able to create something from nothing that’s awesome.
They are very good, and increasingly excellent, at filling in a template in a way the user likes. But the original designs were created by a human team that made them coherent, performed good motion-graphics work, mixed music and motion graphics for impact, and handled all those details.
AI is good at filling that vessel with new content, but the quality of the form remains important. There isn’t an AI today that can handle that from scratch.
The interface itself is also important. People very seldom make no changes. They almost always see something in what AI creates and say, “Even if this is awesome, there are a couple of things I want to change.”
That doesn’t mean the experience wasn’t awesome or that they aren’t amazed by the AI. They may simply want to change their phone number. It could be a factual change, or they may have a better picture that isn’t online yet.
They will always have something they want to change. You need the ability to make that last-mile edit and give the user exactly what they want.
This is specific to Waymark, but hopefully somewhat generalizable. At the highest level, you need taste—an understanding of what you’re trying to do and what good looks like—and you need to embody that in the product.
For us, that means the template library. At the lower level, you need to make the final changes possible in a reliable way, so users can get exactly what they want.
I see many products missing those things, especially products that entered as part of the AI wave. They focus on building with AI and don’t think as much about the other parts, but those parts remain important.
How long this remains true is difficult to say. With Sora, ChatGPT, and potentially Veo, you can imagine saying, “Here’s my small-business website. Make me a 30-second television commercial.”
The models may eventually get smart enough to do the whole thing in pixel space and make it awesome. But that’s still a little ways out, and it’s unclear how much the big companies care about it.
If they decided to refine Veo or Sora so that they could layer copy into videos, edit that copy, and handle commands like “In the last scene, change the phone number to this,” they might eventually be able to do it.
Could they handle those command-style edits on raw video footage generated in a world simulator? Probably. I don’t think anything is off the table.
But they haven’t even fine-tuned or taken the care needed to make these systems work on Tic-Tac-Toe. Are they really going to grind out the data needed to cover all the little use cases?
You probably have a while before those things emerge. The niche strategy is to say that these systems may become generally powerful, but if you know what good looks like in a particular domain, you can create guardrails that ensure quality.
You can also create last-mile editing and customization features so users get exactly what they want rather than endlessly prompting the AI and never quite getting there.
Those things seem valuable for at least a while. But you can imagine a future where the models get so good at coding that they can spin up applications on demand.
At some point, perhaps everything collapses to the cost of inference. ChatGPT could spin up a Waymark-like app, call Sora for footage, handle layers and fonts, and recreate our experience programmatically.
Then you might have price pressure. If people can legitimately experience any set of software tools being spun up through simple commands, we’re heading toward a world of abundance.
Perhaps that’s good. Maybe I won’t have to care about building a small company and can simply live in the land of abundance.
That’s basically the Waymark attitude. We think it will be a while before the big companies do what we do really well and deliver practical value for users.
I’m also glossing over something important: You need to make sure you’re delivering real value for users, rather than simply making something neat with AI.
I have a bias toward making something neat with AI, regardless of whether anyone asked for it. Our CEO, Alex, who is a longtime close friend of mine, is better and more disciplined about asking, “What do our users really want? What are their pain points with the current product?”
He makes sure to solve those problems. Sometimes that coincides with making the next neat AI feature, and sometimes it doesn’t. In the short term, there is value in nailing the product and giving people what they want and are willing to pay for.
Longer term, it’s much harder to say. What’s strange is that these capabilities can either emerge or be engineered, and we don’t always know which has happened.
Few-shot learning was an emergent capability. The researchers weren’t necessarily expecting it to emerge from GPT-3, but there it was. With few-shot learning, it feels like the model can theoretically do almost anything because it has meta-learning.
What won’t emerge over the next few generations? It’s hard to say. Not everything will emerge.
In the meantime, the companies are identifying their biggest weaknesses, patching them, working with Scale AI, and collecting high-quality data to train on.
You’re more insulated as a business if you’re doing something that is both non-emergent and not so focal that the big companies will invest the time and energy to collect the robust datasets needed to power that capability.
I think Waymark is in a decent position there. The big companies care about video, but mostly as a world-simulation problem. I don’t think they care as much about content.
We’re in a relatively safe space for a while, until perhaps the capability emerges. Whether it emerges through pixel-level generation or enough coding sophistication to say, “How would you solve this?” and produce an app is hard to know.
When that happens, it could be a problem. These systems will have much more generality than we do. The trade-off with our app, and with most apps, is that we’re trying to do something narrow really well.
If AI systems can match us on quality, they’ll crush us on breadth. If they can match us on our strong points, they’ll definitely be better than us on our weak points.
We have to hope that doesn’t happen too soon—or at least not until we’re in the age of abundance.
Speaking of the age of abundance, I’ll pick another related question. Why would UBI or post-scarcity be a key outcome of AGI or ASI?
It seems like wealth concentration is the natural direction set in motion by capitalism, which will only accelerate. The question is challenging the assumption that UBI would even be a likely outcome of AGI. Why wouldn’t it instead lead to more wealth concentration?
I certainly don’t rule that out. A lot depends on the shape of the technology.
The reasoning paradigm has shaken things up a little. We talked about compute governance, the offense-defense balance, and how we were heading toward highly egalitarian access with little ability to control the technology.
The new paradigm may move us back toward more inequality and more ability to control it. But I don’t think that’s the final word. The core models will continue to improve, and I fully expect that what you can run on a laptop will continue to advance for at least a few years.
One possibility is a decoupling, and arguably that has already happened. I think it has certainly happened in Scandinavian societies, although I’m not an expert on them.
Take Sweden. My understanding is that there is substantial wealth concentration and hereditary ownership of important companies. IKEA, for example, is a family-held company. There is an elite that controls many important institutions of society.
But there is also a generous and robust social safety net. People are looked after. There isn’t a mass homelessness problem or a mass-incarceration problem. People have their needs taken care of.
That could be how this goes. In a future where compute remains highly concentrated, resource owners will have a lot of power. But individuals who don’t have that wealth and power may still benefit from access to advanced systems.
They could have access to expertise and the ability to take advantage of incredible tools. We’ve had a similar discussion around American health care. It’s bad in many ways and excellent in others.
If you have a rare disease and want frontier treatment, the United States may be the best place to get it. If you want not to worry about a lot of things, or not to fear that your insurance claim will be denied, or if you look at maternal-mortality statistics, the United States doesn’t look as good.
Perhaps we can all have the best access. We can all have an egalitarian future because we can all access the top oncologist, and the top oncologist is an AI that is highly available, scalable, and affordable.
That’s close to the Amodei vision: What if everyone had access to top-quality guidance? What if expertise were effectively free or very low-cost?
That would be my best guess, with the major caveat that the future is probably going to be very strange.
A couple of governments and a handful of big-tech companies may own the physical capital, perform the frontier development, and decide among themselves how the future will go. Everybody else could still live comfortably and have practical access to previously unimaginable sources of expertise.
That might be the new social contract we’re heading toward.
The decision-making power and capital will remain scarce. Who trains the $100 billion model, and on what hardware? At that level, it will be scarce.
If you want to spend a million dollars on inference for a single problem at that level, it will be scarce. What research bets do we make? What do we prioritize?
At OpenAI, there are still compute constraints. My sense is that one reason they haven’t launched Sora is that they simply don’t have the compute to support it. There could be other reasons, but that seems like a major one.
Internally, they can scale up only so many things. Deciding which things to scale is not always easy, and it’s not the kind of decision that will probably be made through direct democracy.
At the same time, people could live extremely comfortably and efficiently. Think about how many language-model interactions you can have for the cost of a single cross-town trip.
A 20- or 30-mile car ride to the doctor and back could cover an enormous number of tokens. There will be abundance in many ways: abundance of expertise and abundance of creative tools.
But decision-making and capital could remain highly concentrated.
That naturally leads to the next question. In a previous episode, you said one of your goals with the podcast was to use your voice to influence the world in a slightly positive direction, however marginally.
Someone who feels responsible but doesn’t have much technical expertise asks: How can I contribute meaningfully to ensuring safe and responsible AGI development, especially in a world where commercial pressures might rush AI progress toward potentially unsafe outcomes?
First, don’t let a lack of technical expertise hold you back, for all the reasons we discussed about learning. If you’re motivated, you can make great progress quickly.
When it comes to understanding what you need to understand to be conversant and sharp about what’s happening, the progress is very accessible. Advancing the frontier is a different question, but catching up to the near frontier has never been more accessible.
There’s a lot going on in AI, but much of it is the same thing working over and over again. We talked about Transformers versus state-space models. If you want to know what is really going on and what matters right now, you can get there quickly.
I wouldn’t let a lack of technical expertise become a mental barrier. Most people can overcome it and reach the level they need to do something useful.
As for how to shape the world, it’s difficult. There are a lot of challenging dynamics. I’ve played around with the adoption-accelerationist, hyperscaling, and pauser framing.
If people had a better understanding of how far things have already gone, they would have a healthier respect—or fear—for where things might be going next.
In some ways, the question of how to make a positive contribution is answered the same way as the question of how to take advantage of AI in your personal life.
I see a large overlap between understanding what’s happening, helping other people understand what’s happening, getting day-to-day value, and helping other people get day-to-day value.
You can bring people up to speed, or at least alert them to what AI can do. Having the implementation know-how to make that work effectively, so they get properly calibrated to where things are, is useful.
Even work in the AI safety community can be interpreted that way. When you look at Apollo Research’s work on scheming, it’s thoughtfully done. Those people are very smart, but they’re also demonstrating capabilities.
They’re saying, “Look at how far this has already come. We got this new model, put it in a somewhat realistic situation—partly a toy problem and partly contrived—but if you squint, it’s not hard to see similar things in the wild.”
I don’t want to understate how much thought went into that work, but I also don’t want to overstate how technically demanding it was. It was driven more by having an accurate sense of what the models can do, exploring the right spaces, and asking whether demonstrating a capability would be meaningful.
Would people update their beliefs if they saw this? They were successful. Not everyone was persuaded—there are always people saying it’s not a big deal or that they’re overreacting—but they made a dent in the universe by getting people’s attention and showing what is possible.
Helping people be accurately calibrated when things are moving this fast is really important. That’s what I try to do.
I try to learn as much as possible, and my hope is that I’ll be one of the people with the most up-to-date and accurate worldviews. I trust that will be valuable.
I try to help other people have the most up-to-date and accurate worldview they can. Any decision-making will hopefully be improved by having a more accurate and current understanding of the world.
That’s not exactly a master plan, but it’s a strategy that a lot of people could follow.