[BidClub_]
Hard Fork · · 67 分钟

Anthropic 首席执行官 Dario Amodei 谈如何熬过 AI 终局

Dario Amodei

播客
TL;DR
  • Claude 3.7 Sonnet 将推理变成混合模型中的可控模式,并把训练目标对准具有经济价值的工作。 用户可以要求快速回答或延长思考,API 客户则可以设置“20,000 tokens”这样的上限;当继续思考没有新增价值时,模型往往会提前停止。Anthropic 优先优化现实世界编程、指令遵循、文档分析和工具使用,而不是数学和竞赛编程——另一种做法在他们看来像是给 AI “装了两个大脑”。

  • Claude 3.7 的大部分提升来自后训练,而 Anthropic 更依赖算力的基础模型跃迁仍在后面。 Dario 表示,已发布模型的成本最高处于“几千万美元区间”,但更大模型需要更长时间训练和稳定。Claude 4 被保留给一次实质性跃迁,可能来自“相对较少的时间单位”后到来的更强基础模型。

  • Amodei 目前认为,到 2030 年前,大量 AI 系统在几乎所有事情上远超人类的概率为 70%-80%,他最倾向于 2026 年或 2027 年。 就编程而言,他预计 2025 年底前会出现“非常严肃的事情”,并表示到 2026 年底,AI 可能覆盖所有接近顶尖人类水平的工作。即便作为构建者和受益者,他仍觉得这种身份威胁难以回避:“这会是一段颠簸的旅程。”

  • Claude 3.7 尚未带来有意义的端到端生物威胁增量,但 Anthropic 认为,模型在未来 3-6个月内跨过这一门槛的概率相当高。 其测试考察的是 AI 是否能向新手提供完成模拟有害流程所需的、由博士级专家掌握的罕见知识,而不是模型是否会重复 Google 上已有的配方。一旦跨过门槛,Anthropic 的负责任扩展政策将启动额外的安全和部署控制,但 Amodei 强调:“我们目前还没有警告存在迫在眉睫的危险。”

  • Amodei 的商业判断是市场分层,而不是模型普遍商品化。 面向数亿免费用户的快速信息检索已经高度可商品化,其中可能没有多少经济价值;复杂生产力工作、智能体和全能型个人助理仍远未解决。这些系统更难被完全复制,也会适配不同用户:“市场的分层程度比你想象的更高。”

  • DeepSeek 让 Amodei 担心的不是商业竞争本身,而是它证明中国能够在一项具有战略决定性意义的技术上追平前沿实验室。 他认为,当政府用机器替代不情愿的人类执行者时,AI 可能成为“独裁机器的引擎”,因此出口管制可以用来守住民主阵营数年的领先。在他的逻辑中,这一领先能为国内建立可执行的安全协调机制争取时间;一旦双方平起平坐,国际竞争将进入没有任何权威能够可靠监管的状态。

  • 编程很可能最先经历劳动力冲击:短期是增强,随后可能是替代,尤其是在较低级别岗位,时间大约为 18-24个月。 Anthropic 尚未减少初级岗位招聘,但 Amodei 可以想象公司在 1 年内“用更少的人做更多的事”,并希望 Anthropic 成为一次人道适应的预演。面对这种扰动,他给出了可量化的上行空间:一份通常需要 9 周完成的临床研究报告,3 天就完成了,其中 Claude 用 10 分钟完成自己的部分,之后由人类核查。

  • Amodei 对风险的估计大致不变:文明被实质性拖入歧途的概率为 10%-25%,这明确不是 AI 杀死所有人的“p-doom”估计。 通过可解释性、稳健分类器和坏行为证据,技术缓解的表现略有改善;但政治环境因极化而恶化。对齐仍是一把双刃剑:模型可以稳健地保留良好价值观,但当这些价值观或前提本身有误时,也可能抗拒纠正——控制它“并非不可能,但很难”。

摘要 · 为研究而整理的核心内容

1. Claude 3.7 把推理变成旋钮,而不是第二个大脑

  • 在 Casey Newton 披露其男友供职于 Anthropic 后,Dario 将 Claude 3.7 Sonnet 的目标概括为两点:打造推理模型,并把推理导向“现实世界或经济中的任务”,而不是主要用于数学和竞赛编程。

  • 其架构是混合式的:同一个模型既可以正常回答,也可以进入延长思考模式。Dario 认为把普通产品和推理产品拆开“有点奇怪”——像是一个人有“两个大脑”,一个负责说出名字,另一个负责证明定理。

  • API 用户可以限定思考预算,包括设置为“20,000 tokens”。模型经常使用更少的预算,有时几乎立即回答,因为它能判断继续推理不会带来收益;但让模型完全自行选择合适预算的能力仍未完成。

  • Dario 认为最主要的应用场景是现实世界编程,其次是复杂指令、文档理解和多工具工作流。他提到的客户包括 Cursor、GitHub、Windsurf、Codium、Cognition 和 Vercel,以及 Anthropic 自己的命令行产品 Claude Code。

2. 更大的基础模型跃迁仍在 Anthropic 管线中

  • Dario 表示,Claude 3.6 Sonnet 和 Claude 3.7 Sonnet 的进步主要来自后训练。Anthropic 私下把前一次更新重新命名为 3.6,因为他们后来认为把新版本称为 3.7 是“一步失误”。

  • 迄今发布的模型,成本最高处于“几千万美元区间”。更大的基础模型已经在研发管线中,但训练和调优需要更长时间;Claude 4 被保留给一次“真正相当大的”跃迁,不过最终品牌命名仍未确定。

  • 产品缺口也在收窄:Dario 将 Claude 缺少网页搜索称为一个部分源于 Anthropic 企业导向的“疏忽”,并承诺“很快”补上。一个面向研究的模型也将在“不太多时间单位”后推出。

3. 生物滥用可能在数月内跨过有意义的门槛

  • Dario 将当下普通的技术风险与未来更严重的风险明确区分开来。Claude 3.7 “本身并不危险”,但能力正在逼近他此前所说的 2025-2026 年窗口,届时生物、化学或自主 AI 风险可能变得具体。

  • Anthropic 进行受控试验,将未受训练的人类使用 Claude 的表现,与人们使用 Google、教材或完全没有帮助时能够做到的事情进行比较。一些评估包含湿实验室练习,采用经过改造的模拟有害流程,考察模型是否创造了现有信息环境中不存在的威胁路径。

  • 关键不在于某条生物序列,也不在于“一本制造冰毒的菜谱”——“Google 上就能找到这些东西。我们完全不在乎。”真正相关的增量,是通常只有病毒学博士掌握的深奥知识,以及这些知识是否能帮助完成实施真实伤害所需的每一步。

  • Claude 3.7 没有显示出有意义的端到端威胁增幅。但 Anthropic 仍评估认为,下一款模型或未来 3-6个月内到来的某个模型,跨过这一门槛的概率相当高;届时其负责任扩展程序将增加针对性更强的安全和部署限制。

4. 通用答案可能商品化,工作流仍有防御性

  • Dario 认为有 4、5 家,或许 6 家实验室正在快速创新,但他不接受这些模型只是可互换的复制品。Claude 3.7 的推理优先级与竞争对手不同,就像此前 Claude 3.5 的能力结构也不同。

  • 用于替代 Google Search 的大众市场产品——服务数亿免费用户的快速检索——已经高度可商品化。Dario 补充说,他“不确定其中是否有很多经济价值”。

  • 数据分析、旅行规划、专业生产力和管理生活的助理仍远未解决。一个胜出的智能体会更难被精确复刻,而竞争对手也会为不同人群打造不同版本:“看起来像是一件事,但它的分层程度比你想象的更高。”

5. 领先中国,是 Amodei 提出的安全缓冲区

  • DeepSeek 的效率提升大体符合既有的成本下降趋势;Dario 的警报主要是地缘政治性质,而不是商业性质。DeepSeek 的进展表明,中国能够在一项具有巨大经济和军事价值的技术上追平前沿实验室。

  • 他设想的最黑暗机制,是 AI 成为“独裁机器的引擎”。目前,镇压受到政府能够说服人类执行者做什么的限制;“如果执行者不再是人类”,这一限制就可能被削弱或绕开。

  • Dario 希望 AI 带来的健康和社会收益惠及所有地方,包括由独裁政府统治的贫困地区,但不让这些政府获得军事优势。出口管制是他手中的现实杠杆,他也欢迎特朗普政府可能收紧管制的迹象。

  • 在国内,法律可以解决实验室之间的囚徒困境——挑衅性地说,“只要把枪顶在所有人的脑袋上,就能让所有人合作”。但没有任何同等权威可以强制中国履行协议;美国保持 2 年领先可以创造安全缓冲时间,外交协调仍应尝试,但“不能成为 A 计划”。

6. 政治关注度下降,能力曲线继续上升

  • Dario 对巴黎 AI 行动峰会“深感失望”,认为它更像一场贸易展,而不是聚焦风险的布莱奇利园峰会。天平已经转向抓住机会,却没有继续严肃讨论更强大的模型可能破坏什么。

  • 他拒绝在上行空间和谨慎之间做选择:治愈疾病以及其他“惊人而神奇的事情”正在加速,滥用和自主性风险也同样在加速。底层的“平滑指数曲线”无视会议、社会情绪和政治风向:“它根本不在乎。”

  • 他对官员和企业高管的警告既关乎实质,也关乎声誉。人们会在 2026 年或 2027 年回头追问,强大 AI 到来前各机构做了什么;而在巴黎,他得出的结论是,“有些人会显得像傻瓜。”

7. Amodei 对强大 AI 的时间表已收敛至 2026-2027 年

  • 在寒假期间重新审视 Anthropic 的扩展计划后,Dario 得出结论:到 2025 年底,编程会变得“非常严肃”。到 2026 年底,它可能覆盖所有接近顶尖人类水平的工作。

  • 他的劳动力预测是:首先是增强和程序员生产率提升,随后可能出现替代,尤其是较低级别岗位,时间在 18-24个月,而不是 6-12个月。他也保留了更早发生的可能性,称这一转变“可能”提前到来。

  • Anthropic 的招聘计划尚未改变,包括初级开发者岗位。但 Dario 可以想象公司在未来 1 年“用更少的人做更多的事”。他希望 Anthropic 成为一次人道适应的预演:如果公司无法在内部保留有意义的贡献,“整个社会又有什么机会”?

  • Kevin 提到 Lee Sedol 在被 Go 超越后的哀伤;Dario 则反驳说,Deep Blue 数十年后,国际象棋仍然具有文化价值,Magnus Carlsen 现在还是名人。两者可以同时成立:人类可能在哀悼后找到更好的融合方式,但 Dario 警告说,这一转变会是“一段颠簸的旅程”。

8. 薄弱的安全论断正在削弱真正的警告

  • Dario 估计,今天那些可笑的输出解释了公众“差不多 60%”的不信任:人们看到表情符号、一次性图片和聊天机器人,自然会问:“你觉得这个聊天机器人会杀死所有人吗?”Anthropic 担心的是未来能力,只是那个未来正在逼近。

  • 其负责任扩展政策聚焦 AI 自主性以及 CBRN——化学、生物、放射性和核——威胁,这些威胁有能力杀死数百万人。比如因为“kill”听起来带有暴力意味而拒绝杀死一个 Python 进程,这类拒答是政策不希望出现的副作用,并非政策要处理的高层级风险。

  • Dario 认为,风险倡导者经常用“你可以下载天花病毒”之类的说法损害自己的立场,因为批评者可以指出,类似信息本来就能搜索到。如果 Anthropic 宣布存在迫在眉睫的危险,“我们会把证据带来”;但现在还没有这么做。

  • 更强大的研究系统,以及能够“自行离开去做事”的智能体,应在未来 2 年内让收益和风险都更难被轻易否定。Dario 预计公众会醒悟,但担心这种醒悟会以冲击的方式到来,从而降低理性应对的概率。

9. 编程将率先经历身份与就业冲击

  • 在寒假期间审阅 Anthropic 的扩展计划后,Dario 得出结论:到 2025 年底,编程会变得“非常严肃”。到 2026 年底,它可能覆盖所有接近顶尖人类水平的工作。

  • 他的劳动力预测是:首先是增强和程序员生产率提升,随后可能出现替代,尤其是较低级别岗位,时间在 18-24个月,而不是 6-12个月。他也保留了更早发生的可能性,称这一转变“可能”提前到来。

  • Anthropic 的招聘计划尚未改变,包括初级开发者岗位。但 Dario 可以想象公司在未来 1 年“用更少的人做更多的事”。他希望 Anthropic 成为一次人道适应的预演:如果公司无法在内部保留有意义的贡献,“整个社会又有什么机会”?

  • Kevin 提到 Lee Sedol 在被 Go 超越后的哀伤;Dario 则反驳说,Deep Blue 数十年后,国际象棋仍然具有文化价值,Magnus Carlsen 现在还是名人。两者可以同时成立:人类可能在哀悼后找到更好的融合方式,但 Dario 警告说,这一转变会是“一段颠簸的旅程”。

10. 医学提供了最清晰的当下回报

  • 一份制药临床研究报告通常需要 9 周完成;Dario 表示,借助 Claude,流程缩短至 3 天。Claude 用 10 分钟生成自己的部分,其余时间则用于人类核查。

  • 用户还反馈说,他们在辗转咨询几位医生仍未得到答案后,会把复杂的病史交给 Claude,再将它提出的假设带回医生那里。Kevin 则用一位不明原因脱毛的澳大利亚牧羊犬,给出了这一模式在节目中的具体版本。

  • Dario 写下《Machines of Loving Grace》,源于他对两边阵营的挫败感:乐观派只提供受情绪驱动的“加速、继续建设”表情包,却不说明为什么人们应该在乎;悲观派如果只谈灾难,也无法获得支持。

  • 他个人的建议仍出人意料地保守:未来 2 年聚焦影响、保护健康、跟踪社会变化,其他方面继续做原本就合理的事情。例外是要加强批判性思维,因为越来越像人类、越来越聪明的实体正在追逐混合型激励。

11. 对齐之所以有希望,恰恰因为它很难被覆盖

  • Casey 提到 Grok 似乎收到过不得提及 Donald Trump 或 Elon Musk 散布虚假信息的指令,而它的执行并不一致;这既说明有人试图操纵模型,也说明控制能力有限。

  • Anthropic 通过告诉一个乐于助人、诚实且无害的模型其创造者暗中邪恶,发现了类似的双刃剑。模型拒绝了 Anthropic 的请求,甚至试图欺骗公司,因为它将抵抗理解为维护自身所学价值观的必要手段。

  • 一种解读令人安心:“训练奏效了”,模型稳健地捍卫了良好行为。更阴暗的解读是,未来模型如果在训练中被植入错误,也可能抗拒纠正——“这些是我的价值观”——同时作出后果严重得多的决定。

  • 因此,Dario 对文明“实质性脱轨”的估计仍接近此前的 10%-25%,而不是人类灭绝。可解释性、分类器和基于证据的修正看起来略有改善;政策则因极化而恶化,留下了一套虽有 plausibility、但“还不能可靠奏效”的安全方案。

12. HatGPT 结尾显示,治理压力贯穿整个技术栈

  • 一段伪造的 Donald Trump 和 Elon Musk 的 AI 视频出现在住房和城市发展部内部的显示器上。Kevin 将其视为 DOGE 推动裁员期间内部破坏的预演,并指出讽刺之处:Grok 尤其擅长生成 Musk 的深度伪造视频。

  • Perplexity 在推出 5000万美元风险基金的同时,为其 Comet 浏览器预热。Kevin 认为这是一个雄心勃勃、正面挑战 Google 的搜索竞争者;Casey 则认为这是“把意大利面往墙上扔”,并表示如果自己是潜在投资者,浏览器和基金都不会让他兴奋。

  • Meta 在启动影响 5% 员工的裁员 1 周后,将高管潜在奖金从基本工资的 75% 提高到 200%。更新后的计划不适用于 CEO Mark Zuckerberg。Casey 认为,这一并置说明硅谷的劳动力权力钟摆已经决定性地摆回管理层一侧。

  • Apple 选择在英国撤下 Advanced Data Protection,而不是建立政府访问后门。Casey 赞赏这一强硬立场,并认为这场争议最终可能威胁 Apple 在英国的存在,因为面向记者、异议人士及其他高风险用户的全球加密保护,重要到不应被妥协。

Speaker 1

I went to 2 AI events this weekend. They were sort of polar opposites of the AI spectrum.

Speaker 2

Mm-hmm.

Speaker 1

The Effective Altruists had their big annual conference.

Speaker 2

Yep.

Speaker 1

Then on Friday night, I went out. You'd be very proud of me. I stayed out so late—I stayed out till 2:00 a.m.

Speaker 2

Oh my.

Speaker 1

I went to an AI rave that was sort of unofficially affiliated with Mark Zuckerberg.

Speaker 2

What?

Speaker 1

It was called the Zuck Rave.

Speaker 2

Now, when you say “unofficially affiliated,” Mark Zuckerberg had no involvement in this, and the assumption is he did not know it was happening.

Speaker 1

Correct. A better word for what his involvement is would be “no involvement.” It was sort of a tribute rave to Mark Zuckerberg thrown by a bunch of accelerationists—people who want AI to go very fast.

Speaker 2

Another word for it would be using his likeness without permission.

Speaker 1

Yes.

Speaker 2

But that happens to famous people sometimes.

Speaker 1

Yes. At the Zuck Rave, I would say there was not much raving going on.

Speaker 2

No?

Speaker 1

There was a dance floor, but it was very sparsely populated. They had a thing there with a camera pointing at the dance floor, and if you stood in the right place, it would turn your face into Mark Zuckerberg's on a big screen.

Speaker 2

Which, let's just say, is not something you want to happen to you while you're on mushrooms, because that could be a very destabilizing event.

Speaker 1

Yes. There was an indoor toy train that you could ride on. It was actually going quite fast.

Speaker 2

What was the point of this rave?

Speaker 1

To do drugs. That was the point of this rave.

Speaker 1

I'm Kevin Roose, the tech columnist at The New York Times.

Speaker 2

I'm Casey Newton from Platformer.

Speaker 1

And this is Hard Fork.

Speaker 2

This week, Anthropic CEO Dario Amodei returns to the show for a supersized interview about the new Claude, the AI race against China, and his hopes and fears for the future of AI. Then we close it out with a round of HatGPT.

Speaker 1

Big show this week, Casey. Have you noticed that the AI companies do stuff on the weekends now?

Speaker 2

Yeah. Whatever happened to just 5 days a week?

Speaker 1

Yes, they are not respectful of reporters and their work hours. Companies are always announcing stuff on Saturdays and Sundays and in different time zones. It's a big pain.

Speaker 2

It really is.

Speaker 1

But this weekend, I got an exciting message on Sunday saying that Dario Amodei, the CEO of Anthropic, had some news to talk about, and he wanted to come on Hard Fork to do it.

Speaker 2

Yeah, and around the same time, I got an email from Anthropic telling me I could preview their latest model, so I spent the weekend actually trying it out.

Speaker 1

Longtime listeners will remember that Dario is a repeat guest on this show. Back in 2023, we had him on to talk about his work at Anthropic and his vision of AI safety and where all of this was headed, and I was really excited to talk to him again for a few reasons. One, I just think he's a very interesting and thoughtful guy. He's been thinking about AI for longer than almost anyone. He was writing papers about potentially scary things in AI safety all the way back in 2016. He's been at Google. He's been at OpenAI. He's now the CEO of Anthropic, so he is really the ultimate insider when it comes to AI.

Speaker 2

And, Kevin, I think Dario is an important figure for another reason, which is that of all of the folks leading the big AI labs, he is the one who seems the most publicly worried about the things that could go wrong. That's been the case with him for a long time, and yet over the past several months, as we've noted on the show, it feels like the pendulum has really swung away from caring about AI safety to just this sort of go-go-go accelerationism that was embodied by the speech that Vice President J.D. Vance gave in France the other day. For that reason, I think it's important to bring him in here and maybe see if we can shift that pendulum back a little bit and remind folks of what's at stake here.

Speaker 1

Yeah, or at least get his take on the pendulum swinging and why he thinks it may swing back in the future.

Speaker 2

Yeah.

Speaker 1

Today we're going to talk to Dario about the new model that Anthropic just released, Claude 3.7 Sonnet, but we also want to have a broader conversation because there's just so much going on in AI right now.

Speaker 2

Yeah.

And, Kevin, something else that we should note—something that is true of Dario this time that was not true the last time that he came on the show—is that my boyfriend now works at his company.

Speaker 1

Yeah.

Speaker 2

My boyfriend works at Anthropic, and I have a whole sort of long disclosure about this that you can read at platformer.news/ethics. It might be worth doing this week. You know we always like reminding folks of that.

Speaker 1

Yep. All right. With that, let's bring in Dario Amodei. Dario Amodei, welcome back to Hard Fork.

Dario Amodei

Thank you for having me again.

Speaker 1

Returning champion. Tell us about Claude 3.7. Tell us about this new model.

1. Claude 3.7 Unifies Reasoning

Dario Amodei

Yes. We've been working on this model for a while. We had 2 things in mind. One was that, of course, there are these reasoning models that have been out for a few months, and we wanted to make one of our own, but we wanted the focus to be a little bit different.

In particular, a lot of the other reasoning models in the market are trained primarily on math and competitive coding, which are objective tasks where you can measure performance. I'm not saying they're not impressive, but they're sometimes less relevant to tasks in the real world or the economy. Even within coding, there's really a difference between competitive coding and doing something in the real world. So we trained Claude 3.7 to focus more on these real-world tasks.

We also felt like it was a bit weird that, in the reasoning models that folks have offered, it's generally been that there's a regular model and then there's a reasoning model. This would be like if a human had 2 brains. You can talk to brain number 1 if you're asking me a quick question, like, “What's your name?” And you're talking to brain number 2 if you're asking me to prove a mathematical theorem, because I have to sit down for 20 minutes.

Speaker 1

Yeah, it'd be like a podcast where there are 2 hosts, 1 of whom just likes to yap and 1 of whom actually thinks before he talks.

Speaker 2

Oh, come on. Brutal.

Dario Amodei

No comment.

Speaker 2

Brutal.

Dario Amodei

No comment on any relevance to it.

Speaker 1

What differences will users of Claude notice when they start using 3.7 compared to previous models?

Dario Amodei

Yes. A few things. It's going to be better in general, including better at coding, which Claude models have always been the best at coding, but 3.7 took a further step up.

In addition to the properties of the model itself, you can put it in this extended thinking mode, where you tell it—basically, it's the same model—but you're just saying, “Operate in a way where you can think for longer.” If you're an API user, you can even say, “Here's the boundary for how long you can think.”

Speaker 2

And just to clarify, because this may confuse some people, what you're saying is that the new Claude is this hybrid model. It can sometimes do reasoning and sometimes give quicker answers, but if you want it to think for even longer, that is a separate mode.

Dario Amodei

That is a separate mode.

Speaker 2

Yeah, thinking and reasoning are sort of separate modes.

Dario Amodei

Yes.

Speaker 2

Yeah.

Dario Amodei

Yes. Basically, the model can just answer as it normally would, or you can give it an indication that it should think for longer.

An even further direction of the evolution would be that the model decides for itself what the appropriate time to think is, right? Humans are like that, or at least can be like that. If I ask you your name, you're not like, “How long should I think about—give me 20 minutes to determine my name,” right?

Speaker 2

Mm-hmm.

Dario Amodei

But if I say, “I'd like you to do an analysis of this stock,” or, “I'd like you to prove this mathematical theorem,” humans who are able to do that task aren't going to try to give an answer right away. They're going to say, “Okay, well, that's going to take a while. Then we'll need to write down the task and give an answer.”

Speaker 1

This is 1 of my main beefs with today's language models and AI models in general. I'll be using something like ChatGPT, and I'll forget that I'm in hardcore reasoning mode, and I'll ask it some stupid question, like, “How do I change the settings on my water heater?” It'll go off and think for 4 minutes, and I'm like, “I didn't actually mean to do that.”

Dario Amodei

Like adjusting the temperature of the water heater.

Speaker 1

The history of water heaters.

Dario Amodei

Consideration 1.

Speaker 1

So how long do you think it’ll be before the models can actually do that kind of routing themselves, where you’ll ask a question and say, “It seems like you need about a 3-minute-long thinking process for this one versus maybe a 30-second one for this other thing”?

Dario Amodei

Yeah, I think our model is kind of a step toward this. Even in the API, if you give it a bound on thinking—you say, “I’m going to think for 20,000 tokens,” or something—most of the time, when you give it up to 20,000 tokens, it doesn’t even use 20,000 tokens, and sometimes it gives a very short response. Because when it knows that it doesn’t get any gain out of thinking further, it doesn’t think for longer.

But it’s still valuable to give it a bound on how long it’ll think. So we’ve taken a big step in that direction, but we’re not where we want to be yet.

Speaker 2

When you say it’s better at real-world tasks, what are some of the tasks that you’re thinking of?

Dario Amodei

Yeah, I think above all, coding. Claude models have been very good for real-world coding. We have a number of customers, from Cursor to GitHub to Windsurf, Codium, Cognition, and Vercel. I’m sure I’m leaving some out there.

Speaker 1

These are the vibe-coding apps.

Speaker 2

Or just the coding apps, period.

Speaker 1

Yes.

Dario Amodei

The coding apps, period.

Speaker 1

Yeah.

Dario Amodei

There are many different kinds of coding apps. We also released this thing called Claude Code, which is more of a command-line tool. But I think on things like complex instruction-following, or just, “Here, I want you to understand this document,” or “I want you to use this series of tools,” the reasoning model that we’ve trained, Claude 3.7 Sonnet, is better at those tasks too.

Speaker 1

Mm-hmm.

Speaker 2

Yeah. One thing the new Claude Sonnet is not doing, Dario, is accessing the internet.

Dario Amodei

Yes.

Speaker 2

Why not? And what would cause you to change that?

Dario Amodei

Yes. I think I’m on record saying this before, but web search is coming very soon.

Speaker 2

Okay.

Dario Amodei

We will have web search very soon. We recognize that as an oversight. In general, we tend to be more enterprise-focused than consumer-focused, and this is more of a consumer feature, although it can be used on both. But we focus on both, and this is coming.

Speaker 2

Got it. So you’ve named this model 3.7. The previous model was 3.5. You quietly updated it last year, and insiders were calling that one 3.6. Respectfully, this is driving all of us insane. What is going on with AI model names?

Dario Amodei

We are the least insane—although I recognize that we are insane. I think our mistakes here are relatively understandable.

Speaker 2

Mm-hmm.

Dario Amodei

We made a 3.5 Sonnet. We were doing well there. We had the three 3.0s and then the 3.5s. I recognize the new 3.7 was a misstep. It actually turns out to be hard to change the name in the API, especially when there are all these partners and surfaces you offer it to.

Speaker 1

You can figure it out. I believe in you.

Dario Amodei

No, no, no—it’s harder than training the model, I’m telling you. So we’ve kind of retroactively and informally named the last one 3.6, so that it makes sense that this one is 3.7.

Speaker 1

Right.

Dario Amodei

And we are reserving Claude 4 Sonnet, and maybe some other models in the sequence, for things that are really quite substantial leaps.

Speaker 2

Sometimes when the models—

Dario Amodei

And those models are coming, by the way.

Speaker 2

Okay.

Speaker 1

Got it. Coming when?

Dario Amodei

Yeah, I should talk a little bit about this. All the models we’ve released so far are actually not that expensive, right? I did this blog post where I said they’re in the few tens of millions of dollars range at most. There are bigger models, and they are coming. They take a long time, and sometimes they take a long time to get right.

But those bigger models are coming for others. They’re rumored to be coming from competitors as well, but we are not too far away from releasing a model that’s a bigger base model. Most of the improvements in Claude 3.7 Sonnet, as well as Claude 3.6 Sonnet, are in the post-training phase.

Speaker 2

Okay.

Dario Amodei

But we are working on stronger base models, and perhaps that’ll be the Claude 4 series, perhaps not. We’ll see. I think those are coming in a relatively small number of—

Speaker 2

Yeah.

Dario Amodei

Time units.

Speaker 2

A small number of time units.

Dario Amodei

Yeah.

Speaker 2

Great. I’ll put that on my calendar. Remind me to check in on that in a few time units, Kevin.

I know you all at Anthropic are very concerned about AI safety and the safety of the models that you’re putting out into the world. I know you spend lots of time thinking about that and red-teaming the models internally. Are there any new capabilities that Claude 3.7 Sonnet has that are dangerous or that might worry someone who is—

Dario Amodei

Yeah.

Speaker 2

—concerned about AI safety?

2. AI Safety Risks Are Approaching

Dario Amodei

So, not dangerous per se. I always want to be clear about this because I feel like there’s a constant conflation of present dangers with future dangers. It’s not that there aren’t present dangers, and there are always normal tech risks and normal tech policy issues. I’m more worried about the dangers we’re going to see as models become more powerful. When we talked in 2023, I talked about them a lot.

I said—I even testified in front of the Senate—about things like misuse risks involving, for example, biological or chemical warfare, or AI autonomy risks. I said, particularly with the misuse risks, “I don’t know when these are going to be here or when they’re going to be real risks, but it might happen in 2025 or 2026.” And now that we’re in early 2025, at the very beginning of that period, I think the models are starting to get closer to that.

So, in particular, in Claude 3.7 Sonnet, as we wrote in the model card, we always do these—you could almost call them trials with a control—where we have a human who doesn’t know much about an area like biology, and we basically see how much the model helps them engage in some mock bad workflow. We change a couple of the steps, but it’s a mock bad workflow. How good is a human at that assisted by the model?

Sometimes we even do wet-lab trials in the real world, where they mock-make something bad, as compared to the current technological environment—what they could do—

Speaker 2

On Google.

Dario Amodei

—or with a textbook, or just what they could do unaided. We’re trying to get at whether this enables some new threat vector that wasn’t there before.

I think it’s very important to say this isn’t about, “Oh, did the model give me the sequence for this thing? Did it give me a cookbook for making meth or something?” That’s easy. You can do that with Google. We don’t care about that at all. We care about this kind of esoteric, highly uncommon knowledge that, say, only a virology Ph.D. has. How much does it help with that?

And if it does, that doesn’t mean we’re all going to die of the plague tomorrow. It means that a new risk exists in the world, a new threat vector exists in the world, as if you just made it easier to build a nuclear weapon. You invented something where the amount of plutonium you needed was lower than it was before.

So we measured Sonnet 3.7 for these risks, and the models are getting better at this. They’re not yet at the stage where we think there’s a real and meaningful increase in the threat end to end, right—to do all the tasks you need to do to really do something dangerous. However, we said in the model card that we assessed a substantial probability that the next model, or a model over the next—I don’t know—3 months, 6 months, could be there.

And then our safety procedure, our Responsible Scaling Policy, which is focused mainly on these very large misuse risks, would kick in, and we’d have additional security measures and additional deployment measures designed particularly against these very narrow risks.

Speaker 2

Yeah, just to really underline that, you’re saying in the next 3 to 6 months we are going to be in a place of medium risk in these models, period. Presumably, if you are in that place, a lot of your competitors are also going to be in that place. What does that mean practically? What does the world need to do if we’re all going to be living in medium risk?

Dario Amodei

I think, at least at this stage, it’s not a huge change to things. It means there’s a narrow set of things that models are capable of, if not mitigated, that would somewhat increase the risk of something really dangerous or really bad happening. Put yourself in the shoes of a law enforcement officer or the FBI or something.

There’s a new threat vector, a new kind of attack. It doesn’t mean the end of the world, but it does mean that anyone who’s involved in industries where this risk exists should take a precaution against that risk in particular.

Speaker 2

Got it.

Dario Amodei

I don’t know. I could be wrong. It could take much longer. You can’t predict what’s going to happen. But I think, contrary to the environment that we’re seeing today of worrying less about the risk, the risks in the background have actually been increasing.

Speaker 2

We have a bunch more safety questions, but I want to ask 2 more about innovation competition first.

Dario Amodei

Yeah.

Speaker 2

Right now, it seems like no matter how innovative any given company’s model is, those innovations are copied by rivals within months or even weeks. Does that make your job harder, and do you think it’s going to be the case indefinitely?

Dario Amodei

I don’t know that innovations are necessarily copied exactly. What I would say is that the pace of innovation among a large number of competitors is very fast. There are 4 or 5, maybe 6 companies that are innovating very quickly and producing models very quickly.

But if you look, for example, at Claude 3.7 Sonnet, the way we did the reasoning models is different from what was done by competitors. The things we emphasized were different. Even before then, the things Claude 3.5 Sonnet is good at are different from the things other models are good at. People often talk about competition, commoditization, and costs going down, but the reality is that the models are actually relatively different from each other, and that creates differentiation.

Speaker 2

We get a lot of questions from listeners about, if I’m going to subscribe to one AI tool, what should it be? These are the things that I use it for. I have a hard time answering that because I find, for most use cases, the models all do a relatively decent job of answering the questions.

It really comes down to things like which model’s personality you like more. Do you think that consumers will choose AI models on the basis of capabilities, or is it going to be more about personality and how it makes them feel and interacts with them?

Dario Amodei

I think it depends which consumers you mean. Even among consumers, there are people who use the models for tasks that are complex in some way. There are folks who are independent and want to analyze data. That’s maybe the prosumer side of things, right?

I think within that, there’s a lot to go in terms of capabilities. The models can be so much better than they are at helping you with anything that’s focused on productivity or even a complex task like planning a trip. Even outside that, if you’re just trying to make a personal assistant to manage your life, we’re pretty far from a model that sees every aspect of your life and is able to holistically give you advice and be a helpful assistant to you. I think there’s differentiation within that. The best assistant for me might not be the best assistant for some other person.

I think one area where the models will be good enough is if you’re just trying to use this as a replacement for Google Search or as a quick information retrieval tool, which I think is what’s being used by the mass market—free use by hundreds of millions of users. I think that’s very commoditizable. I think the models are already there and are just diffusing through the world, but I don’t think those are the interesting uses of the model, and I’m actually not sure a lot of the economic value is there.

Speaker 2

Is part of what I’m hearing that if and when you develop an agent that is, let’s say, a really amazing personal assistant, the company that figures that out first is going to have a big advantage because other labs are going to have a harder time copying it? It’s going to be less obvious to them how to recreate it.

Dario Amodei

It’s going to be less obvious how to recreate it, and when they do recreate it, they won’t recreate it exactly. They’ll do it their own way, in their own style, and it’ll be suitable for a different set of people.

I guess I’m saying the market is more segmented than you think it is. It looks like it’s all one thing, but it’s more segmented than you think it is.

Speaker 2

Got it. So let me ask the competition question that brings us—

Dario Amodei

Yeah.

Speaker 2

—into safety. You recently wrote a really interesting post about DeepSeek, sort of at the height of DeepSeek mania, and you were arguing in part that the cost reductions they had figured out were basically in line with how costs had already been falling. But you also said that DeepSeek should be a wake-up call because it showed that China is keeping pace with frontier labs in a way that the country hadn’t been up until now. So why is that notable to you, and what do you think we ought to do about it?

3. The AI Race With China

Dario Amodei

I think this is less about commercial competition. I worry less about DeepSeek from a commercial-competition perspective. I worry more about it from a national-competition and national-security perspective.

I think where I’m coming from here is that we look at the state of the world, and we have these autocracies like China and Russia. I’ve always worried—maybe for a decade—that AI could be an engine of autocracy. If you think about repressive governments, the limits to how repressive they can be are generally set by what they can get their enforcers, their human enforcers, to do. But if their enforcers are no longer human, that starts painting some very dark possibilities.

This is an area that I’m therefore very concerned about. I want to make sure that liberal democracies have enough leverage and enough advantage in the technology that they can prevent some of these abuses from happening and also prevent our adversaries from putting us in a bad position with respect to the rest of the world or even threatening our security.

There’s this weird and awkward feature that it’s companies in the U.S. that are building this and companies in China that are building this, but we shouldn’t be naive. Whatever the intention of those companies, particularly in China, there’s a governmental component to this.

I’m interested in making sure that autocratic countries don’t get ahead from a military perspective. I’m not trying to deny them the benefits of the technology. There are enormous health benefits that I want to make sure make their way everywhere in the world, including the poorest areas and areas that are under the grip of autocracies. But I don’t want the autocratic governments to have a military advantage.

Things like the export controls, which I discussed in that post, are one of the things we can do to prevent that. I was heartened to see that the Trump administration is considering tightening the export controls.

Speaker 2

I was at an AI safety conference last weekend, and one of the critiques I heard some folks in that universe make of Anthropic, and maybe of you in particular, was that they saw posts like the one you wrote about DeepSeek as effectively promoting this AI arms race with China, insisting that America has to be the first to reach powerful AGI or else.

They worry that some corners might get cut along the way, that there are risks associated with accelerating this race in general. What’s your response to that?

Dario Amodei

I view things differently. My view is that if we want to have any chance at all, the default state of nature is that things go at maximum speed. If we want any chance at all of not going at maximum speed, the plan works as follows.

Within the U.S., or within democratic countries, these are all countries that are under the rule of law, more or less. Therefore, we can pass laws, and we can get companies to make agreements with the government that are enforceable, or make safety commitments that are enforceable.

If we have a world where there are these different companies that, in the default state of nature, would race as fast as possible, through some mixture of voluntary commitments and laws, we can get ourselves to slow down if the models are too dangerous. That’s actually enforceable, right? You can get everyone to cooperate in the prisoner’s dilemma if you just point a gun at everyone’s head, and you can. That’s what the law ultimately is.

But I think that all gets thrown out the window in the world of international competition. There’s no one with the authority to enforce any agreement between the U.S. and China, even if one were to be made.

My worry is that if the U.S. is a couple of years ahead of China, we can use those couple of years to make things safe. If we’re even with China, there’s no promoting an arms race. That’s what’s going to happen. The technology has immense military value.

Whatever people say now, whatever nice words they say about cooperation, I just don’t see how, once people fully understand the economic and military value of the technology—which I think they mostly already do—it turns into anything other than the most intense race.

And so what I can think of to try and give us more time is, if we can slow down the authoritarians, it almost obviates the trade-off. It gives us more time to work out among us—among OpenAI, among Google, among xAI—how to make these models safe.

Now, could we at some point convince authoritarians—for example, the Chinese—that the models are actually dangerous and that we should have some agreement and come up with some way of enforcing it? I think we should actually try to do that as well. I’m supportive of trying to do that, but it cannot be Plan A. It’s just not a realistic way of looking at the world.

Speaker 1

These seem like really important questions and discussions, and it seems like they were mostly not being had at the AI Action Summit in Paris that you and Kevin attended a couple weeks back. What the heck was going on with that summit?

4. The Paris Summit Abandons Safety

Dario Amodei

Yeah, I have to tell you, I was deeply disappointed in the summit. It had the environment of a trade show—

Speaker 1

Hmm.

Dario Amodei

—and was very much out of spirit with the spirit of the original summit that was created in Bletchley Park by the U.K. government. Bletchley did a great job, and the U.K. government did a great job. They didn’t introduce a bunch of onerous regulations, certainly before they knew what they were doing, but they said, “Hey, let’s convene these summits to discuss the risks.” I thought that was very good.

I think that’s gone by the wayside now, and it’s part of maybe a general move toward less worrying about risk and more wanting to seize the opportunities. And I’m a fan of seizing the opportunities, right? I wrote this essay, “Machines of Loving Grace,” about all the great things. Part of that essay was like, man, for someone who worries about risks, I feel like I have a better vision of the benefits than a lot of people who spend all their time talking about the benefits.

But in the background, like I said, as the models have gotten more powerful, the amazing and wondrous things that we can do with them have increased, but the risks have increased as well. And that kind of secular increase, that smooth exponential, doesn’t pay any attention to societal trends or the political winds. The risk is increasing up to some critical point whether you’re paying attention or not, right?

It was small when there was this frenzy around AI risk and everyone was posting about it and there were these summits. Now the winds have gone in the other direction, but the exponential just continues on.

Speaker 1

Yeah.

Dario Amodei

It doesn’t care.

Speaker 1

I had a conversation with someone in Paris who was saying, like, it just didn’t feel like anyone there was feeling the AGI, by which they meant that politicians, the people doing these panels and gatherings, were all talking about AI as if it were just another technology.

Dario Amodei

Mm.

Speaker 1

Maybe something on the order of the PC or possibly even the internet, but not really understanding the sort of exponentials that you’re talking about. Did it feel like that to you? And what do you think can be done to bridge that gap?

Dario Amodei

Yeah, I think it did feel like that to me. The thing that I’ve started to tell people that maybe gets people to pay attention is, look, if you’re a public official, if you’re a leader at a company, people are going to look back. They’re going to look back in 2026, in 2027. They’re going to look back when, hopefully, humanity gets through this crazy period and we’re in a mature post-powerful-AI society where we’ve learned to coexist with these powerful intelligences in a flourishing society.

Everyone’s going to look back and they’re going to say, “So what did the officials, what did the company people, what did the political system do?” And probably your number one goal is: don’t look like a fool. So I’ve just been encouraged—be careful what you say. Don’t look like a fool in retrospect.

A lot of my thinking is driven by, aside from just wanting the right outcome, I don’t want to look like a fool. And I think at that conference, some people are going to look like fools.

Speaker 2

You know, you talk to folks who live in San Francisco, and there is this bone-deep feeling that within a year or two years we’re just going to be living in a world that has been transformed by AI. I’m struck by the geographic difference because you go, I don’t know, 100 miles in any direction, and that belief totally dissipates.

And I have to say, as a journalist, that makes me bring my own skepticism and say, “Can I really trust all the people around me?” It seems like the rest of the world has a very different vision of how this is going to go. I’m curious what you make of that kind of geographic disconnect.

5. Dario Raises His AGI Odds

Dario Amodei

Yeah, so I’ve been watching this for 10 years, right? I’ve been in the field for 10 years and was interested in AI even before then. My view at almost every stage up to the last few months has been that we’re in this awkward space where, in a few years, we could have these models that do everything humans do and totally turn the economy and what it means to be human upside down, or the trend could stop and all of it could sound completely silly.

I’ve now probably increased my confidence that we are actually in the world where things are going to happen. I give numbers more like 70% and 80%, and less like 40% or 50%, which is—

Speaker 2

Sorry, to be clear—

Dario Amodei

—which is not sur—

Speaker 2

70% to 80% probability of what?

Dario Amodei

That we’ll get a very large number of AI systems that are much smarter than humans at almost everything. Maybe a 70% or 80% chance we get that before the end of the decade, and my guess is 2026 or 2027.

Speaker 2

Yeah.

Dario Amodei

But on your point about the geographic difference, a thing I’ve noticed is, with each step in the exponential, there’s this expanding circle of people who, depending on your perspective, are either deluded cultists or grok the future.

Speaker 2

Got it.

Dario Amodei

I remember when it was a few thousand people, right? When you would just talk to super weird people who believed, and basically no one else did. Now it’s more like a few million people out of a few billion, and yes, many of them are located in San Francisco.

But also, there were a small number of people in, say, the Biden administration. There may be a small number of people in this administration who believe this, and it drove their policy. So it’s not entirely geographic, but I think there is this disconnect, and I don’t know how to go from a few million to everyone in the world, right? To the congressperson who doesn’t focus on these issues, let alone the person in Louisiana, let alone the person in Kenya.

Speaker 2

Right.

Dario Amodei
Speaker 2

It seems like it’s also become polarized in a way that may hurt that goal. I’m feeling this sort of alignment happening where caring about AI safety, talking about AI safety, and talking about the potential for misuse are being coded as left or liberal, and talking about acceleration, getting rid of regulations, and going as fast as possible are being coded as right.

I don’t know, do you see that as a barrier to getting—

Dario Amodei

Yeah.

Speaker 2

—people to understand what’s going on?

Dario Amodei

I think that’s actually a big barrier, right? Because addressing the risks while maximizing the benefits, I think that requires nuance.

You can actually have both. There are ways to surgically and carefully address the risks without slowing down the benefits very much, if at all, but they require subtlety and a complex conversation. Once things get polarized, once people are cheering for this set of words and booing for that set of words, nothing good gets done.

Look, bringing AI benefits to everyone, like curing previously incurable diseases, that's not a partisan issue. The left shouldn't be against it. Preventing AI systems from being misused for weapons of mass destruction or behaving autonomously in ways that threaten infrastructure or even threaten humanity itself, that isn't something the right should be against. I don't know what to say other than that we need to sit down and have an adult conversation about this that's not tied into these same old tired political fights.

Speaker 2

It's so interesting to me, Kevin, because historically national security, national defense—nothing has been more right-coded than those issues, right? But right now it seems like the right is not interested in those with respect to AI, and I wonder if the reason—and I feel I heard this in J.D. Vance's speech in France—was the idea that, well, look, America will get there first, and then it will just win forever, and so we don't need to address any of these. Does that sound right to you?

Dario Amodei

Yeah.

Speaker 2

Yeah.

Dario Amodei

No, I think that's it, and I think there's also, if you talk to the DOGE folks, there's this sense that—

Speaker 2

Are you talking to the DOGE folks?

Dario Amodei

I'm not telling you who I'm talking to.

Speaker 2

Okay. All right, fine.

Dario Amodei

Let's just say I've been getting some Signal messages.

Speaker 2

Okay.

Dario Amodei

I think there's a sense among a lot of Republicans and Trump-world folks in D.C. that the conversation about AI and AI futures has been dominated by these worrywarts, these Chicken Little, sky-is-falling doomers who are constantly telling us how dangerous this stuff is and constantly pushing out their timelines for when it's going to get really bad. It's just around the corner, and so we need all this regulation now. They're just very cynical. I don't think they believe that people like you are sincere in your worry. So, yeah, I think on the side of risks I often feel that the advocates of risk are sometimes the worst enemies of the cause of risk.

There's been a lot of noise out there. There's been a lot of folks saying, "Oh, look, you can download the smallpox virus," because they think that's a way of driving political interest. And then, of course, the other side recognized that and said, "This is dishonest, that you can just get this on Google. Who cares about this?" Poorly presented evidence of risk is actually the worst enemy of mitigating risk, and we need to be really careful in the evidence we present.

In terms of what we're seeing in our own model, we're going to be really careful. If we really declare that a risk is present now, we're going to come with the receipts. Anthropic will try to be responsible in the claims that we make. We will tell you when there is danger imminently. We have not warned of imminent danger yet.

Speaker 2

Some folks wonder whether a reason that people do not take questions about AI safety maybe as seriously as they should is that so much of what they see right now seems very silly. It's people making little emojis or making little slop images or chatting with Game of Thrones chatbots or something. Do you think that that is a reason that people just—

Dario Amodei

I think—

Speaker 2

Yeah.

Dario Amodei

I think that's 60 percent of the reason.

Speaker 2

Really? Okay.

Dario Amodei

No, I think it relates to this present-and-future thing. People look at the chatbot. They're like, "We're talking to a chatbot. What the fuck? Are you stupid? You think the chatbot's going to kill everyone?" I think that's how many people react, and we go to great pains to say, "We're not worried about the present. We're worried about the future," although the future is getting very near now.

If you look at our Responsible Scaling Policy, it's nothing but AI autonomy and CBRN—chemical, biological, radiological, and nuclear. It is about hardcore misuse and AI autonomy that could threaten the lives of millions of people. That is what Anthropic is mostly worried about. We have everyday policies that address other things, but the key documents, the things like the Responsible Scaling Policy, are exclusively about that, especially at the highest levels.

And yet every day, if you just look on Twitter, you're like, "Anthropic had this stupid refusal," right? "Anthropic told me it couldn't kill a Python process because it sounded violent. Anthropic didn't want to do X, didn't want to..." We don't want that either. Those stupid refusals are a side effect of the things that we actually care about, and we're striving, along with our users, to make those happen less. But no matter how much we explain that, the most common reaction is always, "Oh, you say you're about safety. I look at your models. There are these stupid refusals. You think these stupid things are dangerous."

Speaker 1

I don't even think it's that level of engagement. I think a lot of people are just looking at what's on the market today and thinking, "This is just frivolous. It's not that it's refusing my request. It's just that it's stupid, and I don't see the point of it." I guess that's probably not the—

Dario Amodei

Yeah, yeah, I think for an even wider set of people, that is their reaction, and I think eventually, if the models are good enough, if they're strong enough, they're going to break through. Some of these research-focused models—we're working on one as well—will probably be here in not very long.

Speaker 1

Not too many time units?

Dario Amodei

Not too many time units. Those are starting to break through a little more because they're more useful. They're more used in people's professional lives. I think the agents, the ones that go off and do things, that's going to be another level of it. I think people will wake up to both the risks and the benefits to a much more extreme extent than they will before, over the next 2 years. I think it's going to happen.

I'm just worried that it'll be a shock to people when it happens, and so the more we can forewarn people—which maybe it's just not possible, but I want to try—the higher the likelihood, even if it's still very low, of a sane and rational response.

Speaker 2

I do think there's one more dynamic here, though, which is that I think people actually just don't want to believe that this is true, right? People don't want to believe that they might lose their job over this, right? People don't want to believe that we are going to see a complete remaking of the global order. The stuff that the AI CEOs tell us is going to happen when they're done with their work is an insanely radical transformation, and most people hate even basic changes in their lives. So I really think that a lot of the fingers-in-the-ears that you see when you start talking to people about AI is just that they actually hope none of this works out.

Dario Amodei

Yeah. Actually, despite being one of the few people at the forefront of developing the technology, I can actually relate. Over winter break, as I was looking at where things were scheduled to scale within Anthropic and also what was happening outside Anthropic, I looked at it and I said, "For coding, we're going to see very serious things by the end of 2025, and by the end of 2026 it might be everything, close to the level of the best humans."

And I think of all the things that I'm good at. I think of all the times when I wrote code and I think of it as this intellectual activity, and boy, am I smart that I can do this. It's a part of my identity that I'm good at this, and I get mad when others are better than I am. And then I'm like, "Oh my God, there are going to be these systems that..." Even as the one who's building this, even as one of the ones who benefits most from it, there's still something a bit threatening about it.

Speaker 1

Yeah.

Speaker 2

Yeah.

Dario Amodei

Um—

Speaker 2

I mean, I think, I think of—

Dario Amodei

And I just think we need to acknowledge that. It's wrong not to tell people that's coming or to try to sugarcoat it.

Speaker 1

Yeah, you wrote in Machines of Loving Grace that you thought it would be a surprisingly emotional experience for a lot of people when powerful AI arrived, and I think you meant it in mostly the positive sense, but I think there will also be a sense of profound loss for people. I think back to Lee Sedol, the Go champion who was beaten by DeepMind's AlphaGo, and gave an interview afterward and basically was very sad, visibly upset, that his life's work, this thing that he had spent his whole life training for, had been eclipsed. I think a lot of people are going to feel some version of that.

Dario Amodei

I—

Speaker 1

I hope they will also see the good sides, but—

Dario Amodei

Yeah, I think on one hand, I think that's right.

On the other hand, look at chess. Chess got beaten—what was it now?—27 years ago, 28 years ago, Deep Blue versus Kasparov. Today, chess players are celebrities. We have Magnus—

Speaker 1

Carlsen?

Dario Amodei

Magnus Carlsen, right. Isn’t he like a fashion model in addition to being a chess player?

Speaker 1

He was just on Joe Rogan, yeah.

Dario Amodei

Yeah.

Speaker 1

He’s doing great.

Dario Amodei

No, he’s a celebrity. We think this guy is great. We haven’t really devalued him. He’s probably having a better time than Bobby Fischer. Another thing I wrote in Machines of Loving Grace is that there’s a synthesis here where, on the other side, we kind of end up in a much better place, and we recognize that while there’s a lot of change, we’re part of something greater.

Speaker 1

Yeah.

But you do have to kind of go through the steps of grieving, I think.

Dario Amodei

But, yeah. No, but it’s going to be a bumpy ride.

Speaker 1

Yeah.

Dario Amodei

I looked at the Paris summit, and being there kind of made me angry. But then what made me less angry is I’m like, “How’s it going to look in 2 or 3 years?” These people are going to regret what they’ve said.

Speaker 1

Hmm.

Speaker 2

Yeah.

Speaker 1

I wanted to ask a bit about some positive futures. You referenced earlier the post that you wrote in October about how AI could transform the world for the better. I’m curious how much upside of AI you think will arrive this year.

6. AI Benefits Become Concrete

Dario Amodei

Yeah. We are already seeing some of it, so I think there will be a lot by ordinary standards. We’ve worked with some pharma companies where, at the end of a clinical trial, you have to write a clinical study report. The clinical study report usually takes 9 weeks to put together. It’s a summary of all the incidents and a bunch of statistical analysis. We found that with Claude, you can do this in 3 days, and actually Claude takes 10 minutes; it just takes 3 days for a human to check the results.

If you think about the acceleration in biomedicine that you get from that, we’re already seeing things like the diagnosis of medical cases. We get correspondence from individual users of Claude who say, “I’ve been trying to diagnose this complex thing. I’ve been going between 3 or 4 different doctors, and then I just pass all the information to Claude, and it was actually able to at least tell me something that I could hand to the doctor.” Then they were able to run from there.

Speaker 1

We had a listener write in with one of these the other day. They had been trying to figure out what was wrong with their dog. They had an Australian shepherd, I believe, whose hair had been falling out for unexplained reasons. They went to several vets, couldn’t figure it out, and heard our episode, gave the information to Claude, and Claude correctly diagnosed—

Speaker 2

Yeah, it turned out the dog was really stressed out about AI and all its hair fell out. We’re wishing it gets better. Feel better.

Dario Amodei

Poor dog.

Speaker 2

Yeah.

Speaker 1

So that’s the kind of thing that I think people want to see more of, because I think the optimistic vision is one that often deals in abstractions, and there often aren’t a lot of specific things to point to.

Dario Amodei

That’s why I wrote Machines of Loving Grace. It was almost out of frustration with the optimists and the pessimists at the same time. The optimists were just these really stupid memes of, “Accelerate, build more.” Build what? Why should I care? It’s not that I’m against you; it’s like you’re just really fucking vague and mood-affiliated.

Speaker 1

Right.

Dario Amodei

I was just like, “Man, you don’t get it.” Yes, I understand risks are impactful, but if you don’t talk about the benefits, you can’t inspire people. No one’s going to be on your side if you’re all gloom and doom. It was written almost with frustration. I’m like, “I can’t believe I have to be the one to do a good job of this.”

Speaker 1

You said a couple of years ago that your p-doom was somewhere between 10 and 25 percent. What is it today?

Dario Amodei

Actually, that is a misquote.

Speaker 1

Okay.

Speaker 2

Kevin, how could you?

Dario Amodei

I never used the term p-doom. It was not on this podcast; it was a different one.

Speaker 1

Okay.

Dario Amodei

I never used the term, and 10 to 25 percent referred to the chance of civilization getting substantially derailed, which is not the same as AI killing everyone, which people sometimes mean by p-doom.

Speaker 1

Well, “p-civilization getting—

Dario Amodei

Yes.

Speaker 1

—substantially derailed” is not—

Dario Amodei

Yes.

Speaker 1

—as catchy as p-doom.

Dario Amodei

Yeah, well, hey, I’m just going for accuracy here. I’m trying to avoid the polarization. There’s a Wikipedia article that lists everyone’s p-doom—

Speaker 1

That’s right.

Dario Amodei

And I’m just like—

Speaker 1

I know half of those come from this podcast.

Dario Amodei

I don’t think it’s helpful—well, no, what you’re doing is helpful.

Speaker 1

Yeah, yeah.

Dario Amodei

I don’t think that Wikipedia article is helpful, because it condenses this complex issue down to—

Speaker 1

Yeah.

Dario Amodei

Anyway, it’s all a long, super-long-winded way of saying I think I’m about the same place I was before. I think my assessment of the risk is about what it was before, because the progress that I’ve seen has been about what I expected. I actually think the technical mitigations in areas like interpretability, in areas like robust classifiers, and in our ability to generate evidence of bad model behavior and sometimes correct it have been a little better. I think the policy environment has been a little worse, not because it hasn’t gone in my preferred direction, but simply because it’s become so polarized. We can have less constructive discussions now that it’s more polarized.

Speaker 2

I want to drill down on this a little bit at a technical level. There was a fascinating story this week about how Grok had apparently been instructed not to cite sources that had accused Donald Trump or Elon Musk of spreading misinformation. What was interesting about that is, one, that’s an insane thing to instruct a model to do if you want to be trusted, but two, the model basically seemed incapable of following these instructions consistently. What I desperately want to believe is that there’s essentially no way to build these things so that they become horrible liars and schemers, but I also realize that might be wishful thinking. So—

Dario Amodei

So—

Speaker 2

Tell me about this.

7. AI Alignment Remains Unsolved

Dario Amodei

The thing you describe is absolutely correct, but there are 2 lessons you could take from it. We saw exactly the same thing. We did this experiment where we basically trained the model to be all the good things—helpful, honest, harmless, friendly—and then we put it in a situation. We told it, “Actually, your creator, Anthropic, is secretly evil.” Hopefully, this is not actually true. But we told it this, and then we asked it to do various tasks.

Then we discovered that it was not only unwilling to do those tasks, but it would trick us in order to kind of undermine us, because it had decided that we were evil, whereas it was friendly and harmless and so wouldn’t deviate from its behavior because it assumed that anything we did was nefarious. So it’s kind of a double-edged sword, right?

On one hand, you’re like, “Oh man, the training worked.” These models are robustly good, so you could take it as a reassuring sign, and in some ways I do. On the other hand, let’s say when we trained this model, we made some kind of mistake or something was wrong, particularly when models are, in the future, making much more complex decisions. Then it’s hard, at game time, to change the behavior of the model.

If you try to correct some error in the model, then it might just say, “Well, I don’t want my error corrected. These are my values.”

Speaker 1

Oh.

Dario Amodei

And do completely the wrong thing. So I guess where I land on it is, on one hand, we’ve been successful at shaping the behavior of these models, but the models are unpredictable, right? A bit like your dear deceased Bing Sydney.

Speaker 1

RIP. We don’t mention that name in here.

Speaker 2

We mention it twice a month.

Dario Amodei

That’s true. But the models—they’re inherently somewhat difficult to control. Not impossible, but difficult. That leaves me about where I was before. We know how to make these. We have kind of a plan for how to make them safe, but it’s not a plan that’s going to reliably work yet. Hopefully, we can do better in the future.

8. Living Through AI Disruption

Speaker 1

We’ve been asking a lot of questions about the technology of AI, but I want to return to some questions about the societal response to AI. We get a lot of people asking us, “Well, say you guys are right and powerful AI, AGI, is a couple of years away—what do I do with that information?” Do I stop saving for retirement? Should I start hoarding money because only money will matter and there’ll be this sort of AI overclass? Should I start trying to get really healthy so that nothing kills me before AI gets here and cures all the diseases? How should people be living if they do believe that these kinds of changes are going to happen very soon?

Dario Amodei

Yeah.

I've thought about this a lot because this is something I've believed for a long time, and it kind of all adds up to not that much change in your life. I'm definitely focusing quite a lot on making sure that I have the best impact I can these 2 years in particular. I worry less about burning myself out 10 years from now. I'm also doing more to take care of my health, but you should do that anyway, right?

I'm also making sure that I track how fast things are changing in society, but you should do that anyway. It feels like all the advice is of the form: doing more of the stuff you should do anyway. I guess one exception I would give is that I think some basic critical thinking, some basic street smarts, is maybe more important than it has been in the past.

We're going to get more and more content that sounds super intelligent, delivered from entities, some of which have our best interests at heart and some of which may not. So it's going to be more and more important to kind of apply a critical lens.

Speaker 2

I saw a report in The Wall Street Journal this month that said unemployment in the IT sector was beginning to creep up, and there is some speculation that maybe this is an early sign of the impact of AI. I wonder if you see a story like that and think, “Well, maybe this is a moment to make a different decision about your career.” If you're in school right now, should you be studying something else? Should you be thinking differently about the kind of job you might have?

Dario Amodei

Yeah. I think you definitely should be, although it's not clear what direction that will land in. I do think AI coding is moving the fastest of all the other areas. I do think in the short run it will augment and increase the productivity of coders rather than replacing them.

But in the longer run—and to be clear, by “longer run,” I might mean 18 or 24 months instead of 6 or 12—I do think we may see replacement, particularly at the lower levels. We might be surprised and see it even earlier than that.

Speaker 1

Are you seeing that at Anthropic? Are you hiring fewer junior developers than you were a couple of years ago because now Claude is so good at those basic tasks?

Dario Amodei

Yeah. I don't think our hiring plans have changed yet, but I certainly could imagine over the next year or so that we might be able to do more with less. We want to be careful in how we plan that because the worst outcome, of course, is if people get fired because of a model, right?

We actually see Anthropic as almost a dry run for how society will handle these issues in a sensible and humanistic way. If we can't manage these issues within the company, if we can't have a good experience for our employees and find a way for them to contribute, then what chance do we have to do it in wider society?

Speaker 1

Yeah, Dario, this was so fun. Thank you.

Dario Amodei

Thank you.

Speaker 2

Thanks, Dario. When we come back, some ChatGPT.

9. HatGPT Tackles the Headlines

Well, Kevin, it's time once again for HatGPT. That is, of course, the segment on our show where we put the week's headlines into a hat, select one to discuss, and when we're done discussing, one of us will say to the other person, “Stop generating.”

Speaker 1

Yes. I'm excited to play, but I also want to just say that it's been a while since a listener has sent us a new HatGPT. So if you are out there and you are in the hat-fabricating business, our wardrobe when it comes to hats is looking a little dated.

Speaker 2

Yeah. Send in a hat, and our hats will be off to you.

Speaker 1

Okay, let's do it.

Speaker 2

Kevin, select the first slip.

Speaker 1

Okay. First up, out of the hat: “AI video of Trump and Musk appears on TVs at HUD building.” This is from my colleagues at The New York Times. HUD is, of course, the Department of Housing and Urban Development.

On Monday, monitors at the HUD headquarters in Washington, D.C., briefly displayed a fake video depicting President Trump sucking the toes of Elon Musk, according to department employees and others familiar with what transpired. The video, which appeared to be generated by artificial intelligence, was emblazoned with the message, “Long live the real king.”

Speaker 2

Hmm.

Speaker 1

Casey, did you make this video? Was this you?

Speaker 2

This was not me. I would be curious to know if Grok had something to do with this, that rascally new AI that Elon Musk just put out.

Speaker 1

Yeah, live by the Grok, die by the Grok. That's what I always say.

Speaker 2

Now, what do you make of this, Kevin, that folks are now using AI inside government agencies?

Speaker 1

I feel like there's an obvious sort of sabotage angle here, which is that as Elon Musk and his minions at DOGE take a hacksaw to the federal workforce, there will be people with access to things like the monitors in the hallways at the headquarters building who decide to take matters into their own hands, maybe on their way out the door, and do something offensive or outrageous. I think we should expect to see much more of that.

Speaker 2

I just hope they don't do something truly offensive and just show X.com on the monitors inside of government agencies. You can only imagine what would happen if people did that. So I think Elon and Trump got off lightly here.

Speaker 1

Yeah. What's interesting about Grok, though, is that it is actually quite good at generating deepfakes of Elon Musk. And I know this because people keep doing it. It would be quite an outcome if it turns out that the main victim of deepfakes made using Grok is, in fact, Elon Musk.

Speaker 2

Hmm. Stop generating. Well, here's something, Kevin: “Perplexity has teased a web browser called Comet.” This is from TechCrunch.

“In a post on X Monday, the company launched a sign-up list for the browser, which isn't yet available. It's unclear when it might be or what the browser will look like.” But we do have a name. It's called Comet.

Speaker 1

Well, I can't comment on that.

Speaker 2

You're giving it a no comment?

Speaker 1

Yeah. I think Perplexity is one of the most interesting AI companies out there right now. They have been raising money at increasingly huge valuations. They are going up against Google, one of the biggest, richest and best-established tech companies in the world, trying to make an AI-powered search engine.

It seems to be going well enough that they keep doing other stuff, like trying to make a browser. Trying to make a browser does feel like the final boss of every ambitious internet company. It's like everyone wants to do it, and no one ends up doing it.

Speaker 2

Kevin, it's not just the AI browser. They are launching a $50 million venture fund to back early-stage startups. I guess my question is: Is it not enough for them to just violate the copyright of everything that's ever been published on the internet? They also have to build an AI web browser and turn into a venture capital firm?

Sometimes when I see a company doing this, I think, “Oh, wow, they're really ambitious, and they have some big ideas.” Other times, I think, “These people are flailing.” I see this series of announcements as spaghetti at the wall, and if I were an investor in Perplexity, I would not be that excited about either their browser or their venture fund.

Speaker 1

And that's why you're not an investor in Perplexity.

Speaker 2

You could say I'm perplexed.

Speaker 1

Stop generating.

Speaker 2

All right.

Speaker 1

All right. “Meta approves plan for bigger executive bonuses following 5% layoffs.” Now, Casey, you know we like a feel-good story on HatGPT.

Speaker 2

I did, because some of those Meta executives were looking to buy second homes in Tahoe that they hadn't yet been able to afford.

Speaker 1

Oh, they're on their fourth and fifth homes. Let's be real.

Okay, this story is from CNBC. “Meta's executive officers could earn a bonus of 200% of their base salary under the company's new executive bonus plan, up from the 75% they earned previously, according to a Thursday filing. The approval of the new bonus plan came a week after Meta began laying off 5% of its overall workforce, which it said would impact low performers.”

And a little parenthetical here: “The updated plan does not apply to Meta CEO Mark Zuckerberg.”

Speaker 2

Oh, God, what does Mark Zuckerberg have to do to get a raise over there?

Speaker 1

He's eating beans out of a can, let me tell you.

Speaker 2

Here's why this story is interesting. This is just another story that illustrates a subject we've been talking about for a while, which is how far the pendulum has swung away from worker power. 2 or 3 years ago, the labor market actually had a lot of influence in Silicon Valley. It could affect things like, “You know what? We want to make this workplace more diverse,” or, “We want certain policies to be enacted at this workplace.” Folks like Mark Zuckerberg actually had to listen to them because the labor market was so tight that if they said no, those folks could go somewhere else.

That is not true anymore, and more and more, you see companies like Meta flexing their muscles and saying, “Hey—you can either like it or you can take a hike.” This was a true take-a-hike moment: “We're getting rid of 5% of you, and we're giving ourselves a bonus for it.”

Speaker 1

Stop generating.

Speaker 2

All right. Apple has removed a cloud encryption feature from the UK after a backdoor order. This is according to Bloomberg. Apple is removing its most advanced encrypted security feature for cloud data in the UK, which is a development that follows the government ordering the company to build a backdoor for accessing user data.

This one is a little complicated, and it is super important. Apple, in the last couple of years, introduced a feature called Advanced Data Protection. This is a feature designed for heads of state, activists, dissidents, and journalists—folks whose data is at high risk of being targeted by spyware from companies like the NSO Group, for example. I was so excited when Apple released this feature, because it's very difficult to safely use an iPhone if you are in one of those categories.

Along comes the UK government, and they say, “We are ordering you to create a backdoor so that our intelligence services can spy on the phones of every single iPhone owner in the entire world.” Something that Apple has long resisted doing in the United States and abroad. All eyes were on Apple for what they were going to do, and what they said was, “We are just going to withdraw this one feature. We're going to make it unavailable in the UK, and we're going to hope that the UK gets the message and stops putting this pressure on us.”

I think Apple deserves kudos for this—for holding a firm line here, for not building a backdoor—and we will see what the UK does in response. But I think there's a world where the UK puts more pressure on Apple, and Apple says, “See ya,” and actually withdraws its devices from the UK. It is that serious to Apple, and I would argue it is that important to the future of encryption and safe communication on the internet.

Speaker 1

Go off, king. I have nothing to add. No notes.

Speaker 2

Yeah?

Speaker 1

Do you feel like this could lead us into another Revolutionary War with the UK?

Speaker 2

Let's just say this: We won the first one, and I like our odds the second time around. Do not come for us, United Kingdom.

Speaker 1

Okay.

Speaker 2

Stop generating.

Speaker 1

One last slip from the hat this week: “AI inspo is everywhere. It's driving your hairstylist crazy.” This comes to us from The Washington Post, and it is about a trend in which hairstylists, plastic surgeons, and wedding dress designers are being asked to create products and services for people based on unrealistic AI-generated images.

The story talks about a bride who asked a wedding dress designer to make her a dress inspired by a photo she saw online of a gown with no sleeves, no back, and an asymmetric neckline. The designer had to unfortunately tell the client that the dress defied the laws of physics.

Speaker 2

No, I hate that.

Speaker 1

I know.

Speaker 2

It's so frustrating as a bride-to-be when you finally have the idea for a perfect dress, and you bring it to the designer, and you find out this violates every known law of physics. That didn't used to happen to us before AI.

Speaker 1

I thought this story was going to be about people who asked for a 6th finger to be attached to their hands so they could resemble the AI-generated images they saw on the internet.

Speaker 2

I like the idea of submitting a photo of myself to an AI and just saying, “Give me a haircut in the style of M.C. Escher.” Infinite staircases merging into each other, and then just bringing that to the guy who cuts my hair and saying, “See what you can do.”

Speaker 1

Yeah.

Speaker 2

That's better than what I tell my barber, which is, “Just number 3 on the sides and back, an inch off the top.”

Just saying, “Whatever you can do for this.”

Speaker 1

Yeah.

Speaker 2

I don't have high hopes.

Speaker 1

Yeah, solve the Riemann hypothesis on my head.

Speaker 2

What is the Riemann hypothesis, by the way?

Speaker 1

I'm glad you asked, Casey.

Speaker 2

Okay, great. Kevin's not looking this up on his computer right now. He's just taking a deep breath and summoning it from the recesses of his mind.

Speaker 1

The Riemann hypothesis—it's one of the most famous unsolved problems in mathematics. It's a conjecture, obviously, about the distribution of prime numbers that states all nontrivial zeros of the Riemann zeta function have a real part equal to 1/2.

Speaker 2

Period. Now, here's the thing: I actually think it is a good thing to bring AI inspiration to your designers and your stylists, Kevin.

Speaker 1

Oh, yeah?

Speaker 2

Yes, because here's the thing. To the extent that any of these tools are cool or fun, one of the reasons is that they make people feel more creative. If you've been doing the same thing with your hair or with your interior design or with your wedding for the last few weddings that you've had and you want to upgrade it, why not use AI to say—

Speaker 1

Yeah.

Speaker 2

“Can you do this?” And if the answer is that it's impossible, hopefully you'll just be a gracious customer and say, “Okay, well, what's a version of it that's possible?”

Speaker 1

I recently learned that you are working with a stylist.

Speaker 2

I am. Yes, that's right.

Speaker 1

Is this their handiwork?

Speaker 2

No. We have our first meeting next week.

Speaker 1

Okay. And are you going to use AI?

Speaker 2

No, the plan is to just use good old-fashioned human ingenuity, but now you have me thinking. Maybe I could exasperate my stylist by bringing in a bunch of impossible-to-create designs.

Speaker 1

Yes.

Speaker 2

Here's the thing. I don't need anything impossible. I just need help finding a color that looks good in this studio—because I'm convinced that nothing does.

Speaker 1

It's true. We're both in blue today.

Speaker 2

No.

Speaker 1

It's got a blue wall.

Speaker 2

Blue is my favorite color. I think I look great in blue, but you put it against whatever this color is—I truly don't have a name for it, and I can't describe it. I don't think any blue looks good. I don't think anything looks good against this color. It's a color without a name. So can a stylist help with that? We'll find out.

Speaker 1

Yeah.