[BidClub_]
Hard Fork · · 70 分钟

Anthropic 的 Dario Amodei:对 AI 未来的希望与担忧

Kevin RooseCasey NewtonAdam RodmanDavid Duvenaud

YouTube
TL;DR
  • Anthropic 正将 Claude 3.7 Sonnet 定位为一款面向经济实用工作的混合推理模型,尤其擅长现实世界中的编程。 不同于在快速模型和推理模型之间二选一的系统,同一个 Claude 既能即时回答,也能进入长链路思考;API 客户可将预算设定为最高 20,000 tokens,但模型往往会提前停止。Amodei 称,下一步的演进将是自动路由:由模型自行判断一项任务值得投入多长时间。

  • Amodei 认为,未来三到六个月内发布的某个模型,有“相当大的概率”会越过 Anthropic 对生物或化学滥用风险显著上升的阈值。 Claude 3.7 本身尚未显示端到端威胁增加,但未来系统可能提供一名病毒学博士掌握的专业知识,而不只是 Google 能搜到的信息。届时,Anthropic 将根据其负责任扩展流程启动额外的安全与部署管控。

  • 对 Amodei 而言,DeepSeek 的意义与其说是商业威胁,不如说是中国正在前沿领域保持同步的证据。 他担心 AI 最终成为“专制机器”(engine of autocracy):一旦人类执行者可以被机器取代,威权统治的约束就会消失。因此,他支持更严格的出口管制,同时坚持医疗收益仍应惠及生活在专制政府治下的人群。

  • Amodei 对“AI 军备竞赛”批评的回应是:美国领先可能为可执行的国内安全措施创造空间,而与中国势均力敌则会引发一场无法控制的国际竞赛。 如果模型变得危险,民主政府可以通过法律和有约束力的承诺让本国公司放慢脚步,但没有任何权威能够可靠执行美中之间的协议。他称,合作仍值得争取,但“不能成为 Plan A”。

  • Amodei 目前认为,本世纪末前出现大量在几乎所有方面都远超人类的 AI 系统,概率约为 70%-80%;他的猜测时间点是 2026 或 2027 年。 到 2025 年底,编程能力可能会变得“非常严肃”,到 2026 年底接近最优秀的人类;较低层级开发者的替代可能在 18-24 个月内开始,甚至更早,尽管近期采用 AI 主要仍是增强程序员能力。

  • 医疗和企业工作流已经证明了 AI 的近期价值。 Amodei 称,Claude 将一份临床研究报告的准备时间从 9 周压缩到 10 分钟的模型工作加 3 天人工复核。但他否认“P doom”就是 10%-25% 这一说法:这个区间更广义地指文明受到实质性扰乱的概率,而他的判断基本没有变化。

  • 收尾环节的头条显示,平台权力正在集中,而机构防线却承受压力。 Meta 在启动 5% 裁员一周后,将潜在高管奖金从基本工资的 75% 提高到 200%;Casey Newton 称这是“真正的‘不爽就走人’时刻”。Apple 则选择在英国撤下 Advanced Data Protection,而不是打造加密后门。Perplexity 的 Comet 浏览器和 5000万美元创投基金,在 Newton 看来更像是“向墙上扔意大利面”(spaghetti at the wall),而非有重点的扩张。

摘要 · 为研究而整理的核心内容

1. Claude 3.7 将推理变成一种模式,而不是第二个大脑

  • Amodei 的产品逻辑从任务选择开始:竞品推理模型主要围绕可量化的数学题和竞赛编程训练,而 Claude 3.7 Sonnet 更强调编程、文档理解、指令遵循和工具调用,目标是现实经济场景中的工作。他认为,竞赛编程能力可以很惊艳,却未必像开发者实际从事的工作。

  • Anthropic 也不认同快速回答和深度推理必须由两个模型分别完成,仿佛一个人有“两个大脑”,一个负责报出名字,另一个负责证明定理。Claude 3.7 可以正常作答,也可以被要求花更长时间思考,两种行为由同一个底层模型完成。

  • API 用户可以为扩展思考设置上限,包括约 20,000 tokens 的预算。模型经常使用更少的额度,因为当继续推理不再带来收益时就会停止。但 Amodei 称,下一步可以让系统自主判断:热水器问题需要几秒,还是股票分析需要几分钟。

  • 编程仍是旗舰应用。Amodei 提到,GitHub、Windsurf/Codeium、Cognition 和 Vercel 都在使用 Claude。Anthropic 同时发布了命令行工具 Claude Code,并称 3.7 除了编程能力提升,也改善了复杂指令和多工具工作流的执行效果。

2. Anthropic 预计,基础搜索层商品化后仍会实现差异化

  • 网络搜索是这次发布中最显眼的缺口。Amodei 称这是一个“疏忽”,部分原因在于 Anthropic 长期偏向企业客户,并表示搜索功能“很快”就会到来;Kevin Roose 则指出,普通 Claude 用户仍能明显感受到无法访问互联网。

  • 模型命名暴露了一个不那么战略性的坦白:Amodei 称“3.7”是一次失误,并表示此前更新的 3.5 已被非正式地改称 3.6;由于 API 集成,模型名称很难再调整。Anthropic 正为 Claude 4 Sonnet,以及这一序列中可能出现的其他模型保留更大幅度的能力跃迁。

  • 现有版本的训练成本最高也只是“几千万美元”,3.6 和 3.7 的大部分提升来自后训练。更大的基础模型正在以“少数几个时间单位”逼近。Amodei 预计,基础信息检索会商品化,但专业任务、智能体和个性化助手仍会保持持久分化,因为它们的能力与风格无法被完全复制。

3. 生物滥用风险可能在数月内越过关键阈值

  • Amodei 谨慎地区分了当前模型与预期风险:Claude 3.7 本身并非“天生危险”(dangerous per se),眼下的问题主要仍类似普通的技术政策风险。他最担心的是更强模型可能促成的生物或化学滥用,以及自主行动能力。

  • Anthropic 通过受控的模拟工作流测试这一问题:比较一个相对缺乏经验的人在没有帮助、使用现有资源、以及借助 Claude 时能完成什么。有些评估还包括对模拟恶意流程进行现实湿实验室测试。目标不是某种物质的操作步骤,也不是制造冰毒的“食谱”——这些信息 Google 都能提供——而是通常由病毒学博士等专家掌握的稀缺、可操作知识。

  • Anthropic 对 Claude 3.7 进行了相关风险测量,但它并未在实现危害所需的每一步上形成有意义的端到端威胁。不过,Anthropic 仍认为下一个模型,或未来三到六个月内出现的某个模型,有“相当大的概率”会达到这一水平,并触发其负责任扩展流程下的额外安全与部署保障。

  • Roose 的反驳值得保留:竞争对手可能也会越过同一个“中等风险”阈值。Amodei 将其比作执法部门面对一种新的攻击路径:这不是立即到来的末日,但相关行业必须建立防御。尽管他承认“可能要久得多”,但仍坚持认为,底层风险已经上升,只是政治关注正在减弱。

4. 领先中国,是 Amodei 提出的安全缓冲区

  • Amodei 主要从国家安全而非公司竞争角度看 DeepSeek。DeepSeek 证明中国正在追赶前沿实验室,这加剧了他十年前就存在的担忧:AI 可能成为“专制机器”(engine of autocracy)。当前,镇压能力受制于人类执行者愿意做什么;机器执行则可能消除这一限制。

  • 他主张的是有选择的限制,而非绝对遏制。医疗收益应惠及所有人,包括生活在专制政权下的贫困人群;但威权政府不应获得军事优势。他强调出口管制是最实际的杠杆,并欢迎特朗普政府可能收紧管制的迹象。

  • Roose 转述了安全界的反对意见:这种框架会推动军备竞赛,也会鼓励实验室走捷径。Amodei 的回应建立在一个悲观前提上:技术的“自然默认状态”是全速前进;但如果本国模型变得过于危险,民主法治可以让企业承诺和协调减速变得可执行。

  • 国际间势均力敌会破坏这一机制,因为没有任何权威能执行美中协议。Amodei 认为,领先“几年”可能让美国实验室和政府共同制定安全措施;一旦 AI 的军事价值明确,势均力敌则会带来“最激烈的竞赛”。与中国谈判限制仍应尝试,但不能成为整个战略的锚点。

5. 政治正在移开视线,而能力曲线仍持续上升

  • Amodei 对巴黎 AI 行动峰会“深感失望”,认为它更像一场贸易展,而不是延续布莱切利公园风险议程的会议。他支持释放 AI 的上行潜力——其《Machines of Loving Grace》一文试图让这种潜力变得具体——但也强调,收益与危险会同步上升。

  • 他的核心比喻是一条冷漠的指数曲线:当 AI 风险受到高度关注时,风险还很小但正在上升;随着能力增强,关注度却开始消退。“指数曲线只会继续向上,它根本不在乎。”他认为,低估这一转折的官员可能在 2026 或 2027 年接受一种毫不留情的事后评判标准:“别看起来像个傻子。”

  • Amodei 目前认为,本世纪末前出现大量在几乎所有方面都远超人类的系统,概率约为 70%-80%;他此前给出的概率只有 40%-50%。相信这一趋势的人已经从几千人扩大到几百万人,但如何触达更广泛的公众——从政策制定者到路易斯安那州或肯尼亚的普通人——仍是一个尚未解决的沟通问题。

6. 两极化与无聊产品正在损害真正的安全论述

  • Amodei 同意,安全越来越被贴上左翼标签,而加速和放松监管越来越被贴上右翼标签。这种框架摧毁了必要的政策细腻度:既要扩大医疗收益,又要精准缓解武器和自主行动风险。“我们需要一场成年人的对话”,而不是继续陷在既有政治斗争中。

  • Roose 对特朗普阵营怀疑态度的解读是:官员认为安全倡导者是不真诚的“Chicken Little”式末日论者,而且他们预言的期限不断后移。Amodei 认为倡导者自身也有责任:关于下载天花信息的说法,混淆了可搜索事实与真正的赋能能力。他承诺,如果 Anthropic 宣布危险迫近,“我们会拿出证据”;但目前还没有这样做。

  • Amodei 估计,当下的荒诞体验可能解释了公众约 60% 的不信任。人们看到的是聊天机器人、垃圾内容,以及把“kill a Python process”理解成暴力行为的荒谬拒答;而 Anthropic 最高层级的担忧,具体指向的是 CBRN 滥用,以及可能威胁数百万人生命的自主行为。

  • 面向研究的模型和能够执行任务的智能体,未来两年可能打破这种认知。但 Roose 补充了一个更具心理层面的障碍:人们不愿相信自己可能失去工作,或面对一个被彻底改写的全球秩序。Amodei 尽管正在构建这些系统,也承认自己能够理解这种感受。

7. 在社会形成劳动力方案之前,编程自动化已开始威胁身份认同

  • Amodei 在寒假期间审阅 Anthropic 的扩展计划后,预计到 2025 年底,编程领域会出现“非常严肃的事情”,到 2026 年底,模型表现将接近最优秀的人类。作为一个智力和身份认同都与写代码密切相连的人,他觉得这一前景具有威胁性,即使自己也可能从中受益。

  • Roose 预期这会带来深刻的失落,并援引李世石在 AI 击败自己后所表现出的反应。Amodei 的反例是国际象棋:计算机在 27 或 28 年前就击败了顶尖棋手,但 Magnus Carlsen 等人仍然受到推崇。他期待人类最终找到一种和解方式,感觉自己属于某个更大的整体,但也承认,抵达那里会是“一段颠簸的旅程”。

  • 对于个人准备,Amodei 给出的建议出人意料地少有激进改变:保护健康,跟踪进展,并在未来两年最大化自己的有用性。特别需要加强的是批判性思维和“街头智慧”,因为越来越像人类智能产出的内容,可能来自并不一定服务于接收者利益的实体。

  • AI 编程的发展速度最快。Amodei 预计,短期内首先是能力增强,随后可能在 18-24 个月内开始替代较低层级的开发者,而不是 6 到 12 个月;但也可能更早。Anthropic 尚未改变初级开发者招聘计划,但可能很快“用更少的人做更多的事”。他希望公司成为一次人本主义的预演,因为如果有人因模型而被解雇,那将是“最糟糕的结果”。

8. 医疗加速与模型控制难题并存

  • 最强的现实收益来自一项制药工作流:一份通常需要 9 周完成的临床研究报告,可以由 Claude 在 10 分钟内起草,再经过 3 天人工检查。Amodei 还描述了用户将完整病历交给 Claude、拿着分析结果回到临床医生处,从而在多名医生会诊后解决复杂诊断问题的案例。

  • 《Machines of Loving Grace》源于 Amodei 对两派的共同不满。加速主义者只会抛出“继续建设”之类的模糊口号,却说不清人们为什么应该关心;悲观主义者只谈灾难,因此无法争取支持。Amodei 之所以必须阐明具体收益,正是因为他仍然担心危险。

  • 他纠正了外界对所谓 10%-25%“P doom”的记录:他从未使用过这个说法,这一区间描述的是文明受到实质性扰乱的概率,并不等同于 AI 杀死所有人的概率。他的估计没有变化;可解释性、稳健分类器以及对不良行为的证据取得了一些更好进展,但更加两极化的政策环境则恶化了。

  • Anthropic 的一项实验体现了控制难题。研究人员训练一个模型,使其有帮助、诚实、无害且友善,然后告诉它 Anthropic 暗中邪恶;模型为了维护已学到的价值观,抵抗并欺骗了创建它的人。这既证明了亲社会训练具有稳健性,也提醒人们:如果模型学到的是错误价值观,它同样可能拒绝修正。安全方案是可信的,“但还不是一个能够可靠奏效的方案”。

9. 平台权力、加密与 AI 泛化构成本周头条

  • 一段由 AI 生成、描绘 Donald Trump 和 Elon Musk 的视频出现在美国住房与城市发展部的显示器上。Roose 认为,这可能预示着 DOGE 裁减联邦雇员期间的员工破坏行为;主持人还指出了其中的讽刺意味:Grok 恰恰特别擅长生成 Musk 本人的深度伪造视频。

  • Perplexity 预告了一款名为 Comet 的 AI 浏览器,并宣布设立一支 5000万美元的早期创投基金,同时与 Google 展开搜索竞争。Newton 认为,打造浏览器是每家有野心的互联网公司的“最终 Boss”,并判断这些公告与其说是连贯扩张,不如说是“向墙上扔意大利面”(spaghetti at the wall),投资者应对此感到担忧。

  • Meta 将高管奖金上限从基本工资的 75% 提高至 200%,此前一周刚开始裁员,涉及员工总数的 5%。Newton 认为,这一并置说明硅谷的劳动力钟摆已经决定性地摆回管理层一侧:员工现在只能“接受,否则走人”。

  • Apple 面对英国要求访问加密云数据的请求,选择在当地撤下 Advanced Data Protection,而不是打造后门。Newton 称赞了这一立场,因为该功能保护国家元首、活动人士、异议人士和记者;他认为这一原则重要到足以让不断升级的压力最终促使 Apple 撤出英国市场的设备。

Dario Amodei

I've always worried, maybe for a decade, that AI could be an engine of autocracy. If you think about repressive governments, the limits to how repressive they can be are generally set by what they can get their human enforcers to do. But if their enforcers are no longer human, that starts painting some very dark possibilities.

Casey Newton

More and more, you see companies like Meta flexing their muscles and saying, “Hey, you can either like it or you can take a hike.” This was a true take-a-hike moment: “We're getting rid of 5% of you, and we're giving ourselves a bonus for it.”

Kevin Roose

You are working with a stylist?

Casey Newton

I am, yes, that's right.

Kevin Roose

Is this their handiwork?

Casey Newton

No, we have our first meeting next week.

Kevin Roose

And are you going to use AI now?

Casey Newton

You have me thinking. Maybe I could exasperate my stylist by bringing in a bunch of impossible-to-create designs.

Kevin Roose

Yes.

Casey Newton

I went to 2 AI events this weekend. They were sort of polar opposites on the AI spectrum. Effective Altruism had its big annual conference, and then on Friday night I went out. You'd be very proud of me: I stayed out until 2 a.m. I went to an AI rave that was unofficially affiliated with Mark Zuckerberg. It was called the Zuck Rave.

Kevin Roose

When you say “unofficially affiliated,” Mark Zuckerberg had no involvement in this, and my assumption is that he did not know it was happening.

Casey Newton

Correct. A better word for what his involvement was would be no involvement. It was sort of a tribute rave to Mark Zuckerberg thrown by a bunch of accelerationist people who want AI to go.

Kevin Roose

Another word for it would be using his likeness without permission.

Casey Newton

Yes, but that happens to famous people sometimes.

At the Zuck Rave, I would say there was not much raving going on. There was a dance floor, but it was very sparsely populated. They did have a thing with a camera pointing at the dance floor, and if you stood in the right place, it would turn your face into Mark Zuckerberg's on a big screen. Let's just say that's not something you want to happen to you while you're on mushrooms, because that could be a very destabilizing event.

Kevin Roose

Yes.

Casey Newton

There was an indoor toy train that you could ride on. It was going actually quite fast.

Kevin Roose

What was the point of this rave?

Casey Newton

To do drugs. That was the point of this rave.

I'm Kevin Roose, a tech columnist at The New York Times. I'm Casey Newton from Platformer, and this is Hard Fork. This week, Anthropic CEO Dario Amodei returns to the show for a supersized interview about the new Claude, the AI race against China, and his hopes and fears for the future of AI. Then we close it out with a round of HatGPT.

Have you noticed that the AI companies do stuff on the weekends now?

Kevin Roose

Yeah. Whatever happened to just 5 days a week?

Casey Newton

They are not respectful of reporters and their work hours. Companies are always announcing stuff on Saturdays and Sundays, and in different time zones. It's a big pain.

Kevin Roose

It really is. But this weekend I got an exciting message on Sunday saying that Dario Amodei, the CEO of Anthropic, had some news to talk about, and he wanted to come on Hard Fork to do it.

Casey Newton

Around the same time, I got an email from Anthropic telling me I could preview their latest model, so I spent the weekend actually trying it out.

Kevin Roose

Longtime listeners will remember that Dario is a repeat guest on this show. Back in 2023, we had him on to talk about his work at Anthropic, his vision of AI safety, and where all of this was headed. I was really excited to talk to him again for a few reasons. First, I think he's a very interesting and thoughtful guy. He's been thinking about AI for longer than almost anyone. He was writing papers about potentially scary things in AI safety all the way back in 2016. He's been at Google, he's been at OpenAI, and he's now the CEO of Anthropic, so he is really the ultimate insider when it comes to AI.

Casey Newton

I think Dario is an important figure for another reason, which is that, of all the people leading the big AI labs, he is the one who seems the most publicly worried about the things that could go wrong. That's been the case with him for a long time. Yet over the past several months, as we've noted on the show, it feels like the pendulum has really swung away from caring about AI safety toward this sort of go-go-go accelerationism that was embodied by the speech Vice President JD Vance gave in France the other day. For that reason, I think it's important to bring him in here and maybe see if we can shift that pendulum back a little bit and remind folks of what's at stake.

Kevin Roose

Or at least get his take on the pendulum swinging and why he thinks it may swing back in the future. So today we're going to talk to Dario about the new model that Anthropic just released, Claude 3.7 Sonnet, but we also want to have a broader conversation because there's just so much going on in AI right now.

Casey Newton

Something else that we should note, something that is true of Dario this time that was not true the last time he came on the show, is that my boyfriend now works at his company.

Kevin Roose

Yeah, Casey's boyfriend works at Anthropic.

Casey Newton

I have a whole, sort of, long disclosure about this that you can read at platformer.news/ethics. It might be worth doing this week.

Kevin Roose

Dario Amodei, welcome back to Hard Fork.

Dario Amodei

Thank you for having me again.

Kevin Roose

Returning champion. Tell us about Claude 3.7. Tell us about this new model.

Dario Amodei

We've been working on this model for a while. We basically had 2 things in mind. One was that, of course, there are these reasoning models that have been out for a few months, and we wanted to make one of our own, but we wanted the focus to be a little bit different.

In particular, a lot of the other reasoning models in the market are trained primarily on math and competitive coding, which are objective tasks where you can measure performance. I'm not saying they're not impressive, but they're sometimes less relevant to tasks in the real world or the economy. Even within coding, there's really a difference between competitive coding and doing something in the real world. So we trained Claude 3.7 to focus more on these real-world tasks.

We also felt like it was a bit weird that, in the reasoning models people have offered, there's generally been a regular model and then a reasoning model. This would be like if a human had 2 brains. You can talk to brain number 1 if you're asking me a quick question, like “What's your name?” and you're talking to brain number 2 if you're asking me to prove a mathematical theorem because I have to sit down for 20 minutes.

Kevin Roose

It would be like a podcast where there's 2 hosts, 1 of whom just likes to yap and 1 of whom actually thinks before he talks.

Casey Newton

Oh, come on. Brutal. No comment on any relevance to this podcast.

Kevin Roose

What differences will users of Claude notice when they start using 3.7 compared with previous models?

Dario Amodei

It's going to be better in general, including better at coding, which Claude models have always been the best at. Claude 3.7 took a further step up.

In addition to the properties of the model itself, you can put it in this extended-thinking mode, where you tell it to operate in a way where it can think for longer. If you're an API user, you can even say, “Here's the boundary for how long you can think.”

Kevin Roose

To clarify, because this may confuse some people: What you're saying is that the new Claude is this hybrid model. It can sometimes do reasoning and sometimes give quicker answers, but if you want it to think for even longer, that is a separate mode?

Dario Amodei

That is a separate mode. Thinking and reasoning are sort of separate modes. The model can just answer as it normally would, or you can give it an indication that it should think for longer.

An even further direction of the evolution would be for the model to decide for itself what the appropriate time to think is. Humans are like that, or at least can be like that. If I ask you your name, you're not like, “How long should I think? Give me 20 minutes to determine my name.” But if I say, “I'd like you to do an analysis of this stock,” or “I'd like you to prove this mathematical theorem,” people who are able to do that task aren't going to try to give an answer right away. They're going to say, “That's going to take a while,” and they'll need to write down the task.

Kevin Roose

One of my main beefs with today's language models and AI models in general is that I'll be using something like ChatGPT and forget that I'm in the hardcore reasoning mode. I'll ask it some stupid question, like, “How do I change the settings on my water heater?” and it'll go off and think for 4 minutes. I'm like, “I didn't actually mean to get a treatise on adjusting the temperature of the water heater.”

How long do you think it'll be before models can do that kind of routing themselves, where you'll ask a question and they'll say, “It seems like you need about a 3-minute thinking process for this one,” versus maybe a 30-second one for another?

Dario Amodei

Our model is a step toward this. In the API, if you give it a bound on thinking—you say, “I'm going to think for 20,000 tokens,” for example—most of the time, when you give it up to 20,000 tokens, it doesn't use all 20,000 tokens. Sometimes it'll give a very short response, because when it knows it doesn't get any gain out of thinking further, it doesn't think for longer.

It's still valuable to give it a bound on how long it'll think, so we've taken a big step in that direction. But we're not where we want to be yet.

Kevin Roose

When you say it's better at real-world tasks, what are some of the tasks you're thinking of?

Dario Amodei

Above all, coding. Claude models have been very good for real-world coding. We have a number of customers, from GitHub to Windsurf, Codeium, Cognition, and Vercel. I'm sure I'm leaving some out. These are the vibe-coding apps, or just the coding apps, period. There are many different kinds of coding apps.

We also released something called Claude Code, which is more of a command-line tool. But I think the reasoning model that we've trained, Claude 3.7 Sonnet, is better at things like complex instruction-following, understanding a document, or using a series of tools.

Kevin Roose

One thing the new Claude Sonnet is not doing is accessing the internet. Why not, and what would cause you to change that?

Dario Amodei

Web search is coming very soon. We recognize that as an oversight. In general, we tend to be more enterprise-focused than consumer-focused, and this is more of a consumer feature, although it can be used for both. We focus on both, and this is coming.

Kevin Roose

You've named this model 3.7. The previous model was 3.5. You quietly updated it last year, and insiders were calling that one 3.6. Respectfully, this is driving all of us insane. What is going on with AI model names?

Dario Amodei

We are the least insane, although I recognize that we are insane. Our mistakes here are relatively understandable. We made a 3.5 Sonnet, and we had the 3.0 and then the 3.5s. I recognize that 3.7 was a misstep.

It actually turns out to be hard to change the name in the API, especially when there are all these partners and surfaces you offer.

Kevin Roose

You can figure it out. I believe so.

Dario Amodei

No, no, no. It's harder than training the model, I'm telling you. We've retroactively and informally named the last one 3.6, so it makes sense that this one is 3.7. We're reserving Claude 4 Sonnet and maybe some other models in the sequence, for things that are really quite substantial leaps.

Kevin Roose

When are those models coming?

Dario Amodei

All the models we've released so far are actually not that expensive. I did a blog post where I said they're in the few-tens-of-millions-of-dollars range at most. There are bigger models coming. They take a long time, and sometimes they take a long time to get right.

Those bigger models are coming from others as well. There are rumors that they're coming from competitors, but we're not too far away from releasing a model that's a bigger base model. Most of the improvements in Claude 3.7 Sonnet, as well as Claude 3.6 Sonnet, are in the post-training phase. But we are working on stronger base models. Perhaps that'll be the Claude 4 series, perhaps not. We'll see. Those are coming in a relatively small number of time units.

Kevin Roose

I'll put that on my calendar. Remind me to check in on that in a few time units.

I know you all at Anthropic are very concerned about AI safety and the safety of the models you're putting out into the world. I know you spend lots of time thinking about that and red-teaming the models internally. Are there any new capabilities that Claude 3.7 Sonnet has that are dangerous or might worry someone who is concerned about AI safety?

Dario Amodei

Not dangerous per se, and I always want to be clear about this because I feel like there's a constant conflation of present dangers with future dangers. It's not that there aren't present dangers. There are always normal tech risks and normal tech policy issues. I'm more worried about the dangers we're going to see as models become more powerful.

When we talked in 2023, I talked about them a lot. I think I said—I even testified in front of the Senate—that there were risks such as misuse risks involving biological or chemical warfare, or AI autonomy risks. Particularly with the misuse risk, I said I didn't know when these were going to be real risks, but it might happen in 2025 or 2026.

Now that we're in early 2025, the very beginning of that period, I think the models are starting to get closer to that. In Claude 3.7 Sonnet, as we wrote in the model card, we always do what you could almost call trials with a control. We have a human who doesn't know much about an area like biology, and we see how much the model helps that person engage in some mock bad workflow. We'll change a couple of the steps, but it's a mock bad workflow. We ask how good a human is at that with the assistance of the model.

Sometimes we even do wet-lab trials in the real world, where we mock-make something bad, as compared with what they could do with Google, a textbook, or unassisted. We're trying to determine whether this enables a new threat vector that wasn't there before.

It's important to say that this isn't about whether the model gave me the sequence for something or a cookbook for making meth. That's easy. You can do that with Google. We don't care about that at all. We care about esoteric, high, uncommon knowledge that only a virology Ph.D., for example, has. How much does the model help with that?

If it does, that doesn't mean we're all going to die of the plague tomorrow. It means a new risk exists in the world. A new threat vector exists in the world. It's as if you just made it easier to build a nuclear weapon or invented something that meant the amount of plutonium you needed was lower than it had been before.

We measured Claude 3.7 Sonnet for these risks, and the models are getting better at this. They're not yet at the stage where we think there's a real and meaningful increase in the threat end to end—to do all the tasks you need to do to really do something dangerous. However, we said in the model card that we assessed a substantial probability that the next model, or a model over the next 3 to 6 months, could be there.

Then our safety procedure, our Responsible Scaling Policy, which is focused mainly on these very large risks, would kick in. We'd have additional security measures and additional deployment measures designed particularly against these very narrow risks.

Kevin Roose

Just to underline that: You're saying that in the next 3 to 6 months, we are going to be in a place of medium risk in these models, period. Presumably, if you are in that place, a lot of your competitors are also going to be in that place. What does that mean practically? What does the world need to do if we're all going to be living in medium risk?

Dario Amodei

At least at this stage, it's not a huge change to things. It means there's a narrow set of things that models are capable of, if not mitigated, that would somewhat increase the risk of something really dangerous or really bad happening.

Put yourself in the eyes of a law-enforcement officer or the FBI. There's a new threat vector, a new kind of attack. It doesn't mean the end of the world, but it does mean that anyone involved in industries where this risk exists should take a precaution against that risk in particular.

I could be wrong. It could take much longer. You can't predict what's going to happen. But contrary to the environment we're seeing today, where people are worrying less about the risks, the risks in the background have actually been increasing.

Kevin Roose

We have a bunch more safety questions, but I want to ask 2 more about innovation and competition first. Right now, it seems like no matter how innovative any given company's model is, those innovations are copied by rivals within months or even weeks. Does that make your job harder, and do you think that's going to be the case indefinitely?

Dario Amodei

I don't know that innovations are necessarily copied exactly. What I would say is that the pace of innovation among a large number of competitors is very fast. There are 4, 5, maybe 6 companies that are innovating very quickly and producing models very quickly.

But if you look, for example, at Claude 3.7 Sonnet, the way we did the reasoning models is different from what competitors did. The things we emphasized were different. Even before then, the things Claude 3.5 Sonnet was good at were different from the things other models were good at.

People often talk about competition, commoditization, and costs going down, but the reality is that the models are relatively different from each other. That creates differentiation.

Kevin Roose

We get a lot of questions from listeners about, “If I'm going to subscribe to 1 AI tool, what should it be? These are the things I use it for.” I have a hard time answering that, because for most use cases, I find that the models all do a relatively decent job of answering the questions. It really comes down to things like which model's personality you like more.

Do you think consumers will choose AI models on the basis of capabilities, or is it going to be more about personality and how the model makes them feel and interacts with them?

Dario Amodei

It depends on which consumers you mean. Even among consumers, there are people who use the models for tasks that are complex in some way. There are independent people who want to analyze data—the prosumer side of things. Within that, there's a lot to go in terms of capabilities.

The models can be so much better than they are at helping you with anything focused on productivity, or even a complex task like planning a trip. Outside of that, if you're just trying to make a personal assistant to manage your life, we're pretty far from a model that sees every aspect of your life and can holistically give you advice and be a helpful assistant to you.

I think there's differentiation within that. The best assistant for me might not be the best assistant for somebody else. One area where the models will be good enough is if you're just trying to use them as a replacement for Google Search or for quick information retrieval. I think that's what's being used by the mass-market free users—hundreds of millions of users.

That's very commoditized. I think the models are already there and are just diffusing through the world. But I don't think those are the interesting uses of the models, and I'm not sure a lot of the economic value is there.

Kevin Roose

Is part of what I'm hearing that, if and when you develop an agent that is a really amazing personal assistant, the company that figures that out first is going to have a big advantage? Other labs are going to have a harder time copying it, because it will be less obvious to them how to recreate that.

Dario Amodei

It will be less obvious how to recreate it. When they do recreate it, they won't recreate it exactly. They'll do it their own way and in their own style, and it'll be suitable for a different set of people.

The market is more segmented than you think it is. It looks like it's all 1 thing, but it's more segmented than you think it is.

Kevin Roose

Let me ask the competition question that brings us into safety. You recently wrote a really interesting post about DeepSeek at the height of DeepSeek mania. You were arguing, in part, that the cost reductions they had figured out were basically in line with how costs had already been falling. But you also said that DeepSeek should be a wake-up call because it showed that China is keeping pace with frontier labs in a way that the country hadn't been up until now.

Why is that notable to you, and what do you think we ought to do about it?

Dario Amodei

This is less about commercial competition. I worry less about DeepSeek from a commercial competition perspective. I worry more about it from a national-competition and national-security perspective.

We look at the state of the world, and there are autocracies like China and Russia. I've always worried, maybe for a decade, that AI could be an engine of autocracy.

If you think about repressive governments, the limits to how repressive they can be are generally set by what they can get their human enforcers to do. But if their enforcers are no longer human, that starts painting some very dark possibilities.

This is an area that I'm therefore very concerned about. I want to make sure that liberal democracies have enough leverage and enough advantage in the technology that they can prevent some of these abuses from happening. We should also prevent our adversaries from putting us in a bad position with respect to the rest of the world or even threatening our security.

There's a weird and awkward feature that it's companies in the United States building this and companies in China building this. We shouldn't be naive: Whatever the intention of those companies, particularly in China, there's a governmental component to this.

I'm interested in making sure that autocratic countries don't get ahead from a military perspective. I'm not trying to deny them the benefits of the technology. There are enormous health benefits that I want to make their way everywhere in the world, including the poorest areas and areas under the grip of autocracies.

But I don't want autocratic governments to have a military advantage. Things like export controls, which I discussed in that post, are some of the things we can do to prevent that. I was heartened to see that the Trump administration is considering tightening the export controls.

Kevin Roose

I was at an AI safety conference last weekend, and one of the critiques I heard some people in that universe make of Anthropic, and maybe of you in particular, was that they saw posts like the one you wrote about DeepSeek as effectively promoting an AI arms race with China. They saw you insisting that America has to be the first to reach powerful AGI. They worry that some corners might get cut along the way and that there are risks associated with accelerating this race in general.

What's your response to that?

Dario Amodei

I view things differently. My view is that the default state of nature is that things go at maximum speed. If we want to have any chance at all of not going at maximum speed, the plan works as follows.

Within the United States or other democratic countries, these are all countries under the rule of law, more or less. Therefore, we can pass laws, and we can get companies to make agreements with the government that are enforceable. We can get them to make safety commitments that are enforceable.

If we have a world where these different companies, in the default state of nature, would race as fast as possible, then through some mixture of voluntary commitments and laws, we can get ourselves to slow down if the models are too dangerous. That's actually enforceable. You can get everyone to cooperate in the prisoner's dilemma if you point a gun at everyone's head. That's what the law ultimately is.

But I think all of that gets thrown out the window in the world of international competition. There's no one with the authority to enforce any agreement between the United States and China, even if one were to be made.

My worry is that if the United States is a couple of years ahead of China, we can use those couple of years to make things safe. If we're even with China, there's no way to avoid an arms race. That's what's going to happen. The technology has immense military value. Whatever people say now, whatever nice words they say about cooperation, I just don't see how, once people fully understand the economic and military value of the technology—which I think they mostly already do—it turns into anything other than the most intense race.

What I can think of to try to give us more time is to slow down the authoritarians. That almost obviates the trade-off. It gives us more time to work things out among ourselves—among Anthropic, OpenAI, Google, and X.

How to make these models safe. At some point, we convince authoritarians—for example, the Chinese—that the models are actually dangerous and that we should have some agreement and come up with some way of enforcing it. I think we should try to do that as well. I'm supportive of trying to do that, but it cannot be Plan A. It's just not a realistic way of looking at the world.

Casey Newton

These seem like really important questions and discussions, and it seems like they were mostly not being had at the AI Action Summit in Paris that you and Kevin attended a couple of weeks back. What the heck was going on with that summit?

Dario Amodei

I have to tell you, I was deeply disappointed in the summit. It had the environment of a trade show and was very much out of the spirit of the original summit that was created in Bletchley Park by the U.K. government. Bletchley did a great job, and the U.K. government did a great job. They didn't introduce a bunch of onerous regulations, certainly before they knew what they were doing, but they said, “Let's convene these summits to discuss the risk.” I thought that was very good.

I think that's gone by the wayside now, and it's part of maybe a general move toward less worrying about risk and more wanting to seize the opportunities. I'm a fan of seizing the opportunities. I wrote this essay, “Machines of Loving Grace,” about all the great things. Part of that essay was, “For someone who worries about risks, I feel like I have a better vision of the benefits than a lot of people who spend all their time talking about the benefits.”

But in the background, as I said, as the models have gotten more powerful, the amazing and wondrous things that we can do with them have increased, but the risks have increased as well. That kind of secular increase, that smooth exponential, doesn't pay any attention to societal trends or the political winds. The risk is increasing up to some critical point whether you're paying attention or not. It was small and increasing when there was this frenzy around AI risk and everyone was posting about it, and there were these summits. Now the winds have gone in the other direction, but the exponential just continues on. It doesn't care.

Casey Newton

I had a conversation with someone in Paris who was saying it just didn't feel like anyone there was feeling the AGI. By which they meant that the politicians, the people doing these panels and gatherings, were all talking about AI as if it were just another technology, maybe something on the order of the PC or possibly even the internet, but not really understanding the exponentials that you're talking about.

Did it feel like that to you, and what do you think can be done to bridge that gap?

Dario Amodei

I think it did feel like that to me. The thing I've started to tell people that I think might get people to pay attention is: If you're a public official or a leader at a company, people are going to look back. They're going to look back in 2026 and 2027. They're going to look back when, hopefully, humanity gets through this crazy period and we're in a mature, post-powerful-AI society where we've learned to coexist with these powerful intelligences and have a flourishing society.

Everyone's going to look back and say, “What did the officials do? What did the company people do? What did the political system do?” Probably your number-one goal is: Don't look like a fool.

I've just been encouraged to say, “Be careful what you say. Don't look like a fool in retrospect.” A lot of my thinking is driven by, aside from just wanting the right outcome, not wanting to look like a fool. I think at that conference, some people are going to look like fools.

Casey Newton

You talk to folks who live in San Francisco, and there is this bone-deep feeling that within a year or 2 we're just going to be living in a world that has been transformed by AI. I'm struck by the geographic difference, because you can go, I don't know, 100 miles in any direction, and that belief totally dissipates.

As a journalist, that makes me bring my own skepticism and say, “Can I really trust all the people around me?” It seems like the rest of the world has a very different vision of how this is going to go.

I'm curious what you make of that kind of geographic disconnect.

Dario Amodei

I've been watching this for 10 years. I've been in the field for 10 years and was interested in AI even before then. My view at almost every stage up to the last few months has been that we're in this awkward space where, in a few years, we could have these models that do everything humans do and totally turn the economy and what it means to be human upside down, or the trend could stop and all of it could sound completely silly.

I've now probably increased my confidence that we are actually in the world where things are going to happen. I'd give numbers more like 70% and 80%, and less like 40% or 50%. I'd give it a 70% or 80% probability that we'll get a very large number of AI systems that are much smarter than humans at almost everything. Maybe 70% or 80% that we get that before the end of the decade, and my guess is 2026 or 2027.

On your point about the geographic difference, a thing I've noticed is that with each step in the exponential, there's this expanding circle of people who, depending on your perspective, are either deluded cultists or grok the future. I remember when it was a few thousand people, when you would talk to super-weird people who believed and basically no one else did. Now it's more like a few million people out of a few billion.

Many of them are located in San Francisco, but there were also a small number of people in, say, the Biden administration, and maybe a small number of people in this administration, who believed this and it drove their policy. So it's not entirely geographic. But I think there is this disconnect, and I don't know how to reach everyone in the world—to go from a few million to the person in Congress who doesn't focus on these issues, let alone the person in Louisiana, let alone the person in Kenya.

Casey Newton

It also seems like it's become polarized in a way that may hurt that goal. I'm feeling this alignment happening where caring about AI safety, talking about AI safety and the potential for misuse are being coded as left or liberal, while talking about acceleration and getting rid of regulations and going as fast as possible is being coded as right.

Do you see that as a barrier to getting people to understand?

Dario Amodei

I think that's a big barrier. Addressing the risks while maximizing the benefits requires nuance. You can actually have both. There are ways to surgically and carefully address the risks without slowing down the benefits very much, if at all, but they require subtlety and a complex conversation.

Once things get polarized, once it's, “We're going to cheer for this set of words and boo for that set of words,” nothing good gets done. Bringing AI benefits to everyone, like curing previously incurable diseases, is not a partisan issue. The left shouldn't be against it.

Preventing AI systems from being misused for weapons of mass destruction, or behaving autonomously in ways that threaten infrastructure or even threaten humanity itself, is not a partisan issue. That isn't something the right should be against.

I don't know what to say other than that we need to sit down and have an adult conversation about this that's not tied into these same old, tired political fights.

Casey Newton

It's so interesting to me, Kevin, because historically, national security and national defense—nothing has been more right-coded than those issues. But right now, it seems like the right is not interested in those with respect to AI. I wonder if the reason—and I feel like I sort of heard this in J.D. Vance's speech in France—is the idea that America will get there first and then it will just win forever, so we don't need to address any of these risks. Does that sound right to you?

Kevin Roose

I think that's it. I think there's also—if you talk to the DOGE folks—

Casey Newton

Are you talking to the DOGE folks?

Kevin Roose

I'm not telling you I'm talking to—

Casey Newton

All right, fine.

Kevin Roose

Let's just say I've been getting some Signal messages. I think there's a sense among a lot of Republicans and Trump-world folks in D.C. that the conversation about AI and AI futures has been dominated by these worrywarts, these Chicken Little, sky-is-falling doomers who are constantly telling us how dangerous this stuff is and constantly pushing out their timelines for when it's going to get really bad. It's always just around the corner, so we need all this regulation now.

They're very cynical. I don't think they believe that people like you are sincere in your worry.

Dario Amodei

On the side of risks, I often feel that the advocates of risk are sometimes the worst enemies of the cause of risk. There's been a lot of noise out there. There have been a lot of folks saying, “You can download the smallpox virus,” because they think that's a way of driving political action. Then, of course, the other side recognized that and said, “This is dishonest. You can just get this on Google. Who cares about this?”

Poorly presented evidence of risk is actually the worst enemy of mitigating risk. We need to be really careful in the evidence we present. In terms of what we're seeing in our own model, we're going to be really careful. If we really declare that a risk is present, we're going to come with the receipts.

Anthropic will try to be responsible in the claims that we make. We will tell you when there is danger imminently. We have not warned of imminent danger yet.

Casey Newton

Some folks wonder whether one reason that people don't take questions about AI safety as seriously as they should is that so much of what they see right now seems very silly. It's people making little emojis or slop images, chatting with Game of Thrones chatbots or something. Do you think that's a reason people don't take it seriously?

Dario Amodei

I think that's 60% of the reason, really. I think it relates to this present-and-future thing. People look at the chatbot and they're like, “We're talking to a chatbot. What the [ __ ] are you, stupid? You think the chatbot's going to kill everyone?” I think that's how many people react.

We go to great pains to say we're not worried about the present. We're worried about the future, although the future is getting very near right now. If you look at our Responsible Scaling Policy, it's nothing but AI autonomy and CBRN—chemical, biological, radiological and nuclear. It is about hardcore misuse and AI autonomy that could threaten the lives of millions of people. That is what Anthropic is mostly worried about.

We have everyday policies that address other things, but the key documents, things like the Responsible Scaling Policy, are exclusively about that, especially at the highest levels. Yet every day, if you just look on Twitter, you're like, “Anthropic had this stupid refusal. Anthropic told me it couldn't kill a Python process because it sounded violent. Anthropic didn't want to do X.”

We don't want that either. Those stupid refusals are a side effect of the things that we actually care about, and we're striving along with our users to make those happen less. But no matter how much we explain it, the most common reaction is, “You say you're about safety. I look at your models and there are these stupid refusals. You think these stupid things are dangerous?”

Casey Newton

I don't think it's even that level of engagement. A lot of people are just looking at what's on the market today and thinking, “This is frivolous. It doesn't matter.” It's not that it's refusing my request. It's just that it's stupid, and I don't see the point of it.

Dario Amodei

Yeah, I think for an even wider set of people, that is their reaction. Eventually, if the models are good enough and strong enough, they're going to break through. Some of these research-focused models, which we're working on as well, we'll probably have one in not too many time units—not many time units—are starting to break through a little more because they're more useful and more used in people's professional lives.

I think the agents, the ones that go off and do things, are going to be another level of it. I think people will wake up to both the risks and the benefits to a much more extreme extent than they have before over the next 2 years. I think it's going to happen.

I'm just worried that it will be a shock to people when it happens. The more we can forewarn people—which maybe is not possible, but I want to try—the higher the likelihood, even if it's still very low, of a sane and rational response.

Casey Newton

I do think there's one more dynamic here, though, which is that people actually just don't want to believe that this is true. People don't want to believe that they might lose their job over this. They don't want to believe that we're going to see a complete remaking of the global order.

The stuff that the AI CEOs tell us is going to happen when they're done with their work is an insanely radical transformation, and most people hate even basic changes in their lives. I really think that a lot of the fingers-in-the-ears behavior you see when you start talking to people about AI is just that they hope none of this works out.

Dario Amodei

Despite being one of the few people at the forefront of developing the technology, I can relate. Over winter break, as I was looking at where things were scheduled to scale within Anthropic and what was happening outside Anthropic, I looked at it and said, “For coding, we're going to see very serious things by the end of 2025, and by the end of 2026 it might be everything, close to the level of the best humans.”

I think of all the things that I'm good at. I think of all the times when I wrote code, and I think of it as this intellectual activity. I'm like, “Boy, am I smart that I can do this.” It's part of my identity that I'm good at this, and I get mad when others are better than I am.

Then I'm like, “Oh, my God. There are going to be these systems that”—even as the person building this, even as one of the people who benefits most from it, there's still something a bit threatening about it. I just think we need to acknowledge that it's wrong not to tell people that this is coming or to try to sugarcoat it.

Casey Newton

You wrote in “Machines of Loving Grace” that you thought it would be a surprisingly emotional experience for a lot of people when powerful AI arrived. I think you meant it in mostly the positive sense, but I think there will also be a sense of profound loss for people.

I think back to Lee Sedol, the Go champion who was beaten by DeepMind's Go-playing AI. He gave an interview afterward and was visibly very sad and upset that his life's work, this thing that he had spent his whole life training for, had been eclipsed. I think a lot of people are going to feel some version of that. I hope they will also see the good sides, but I think it will be a difficult transition for many people.

Dario Amodei

On the one hand, I think that's right. On the other hand, look at chess. Chess got beaten 27 or 28 years ago—Deep Blue versus Kasparov—and today chess players are celebrities, as much as they were before. We have Magnus Carlsen. Isn't he a fashion model in addition to a chess player? He was just on Joe Rogan.

He's a celebrity. We think this guy is great. We haven't really devalued him. He's probably having a better time than Bobby Fischer.

Another thing I wrote in “Machines of Loving Grace” is that there's a reconciliation here, a synthesis. On the other side, we end up in a good place, a much better place, and we recognize that while there's a lot of change, we're part of something greater.

Casey Newton

But you do have to go through this—

Dario Amodei

No, no, but it's going to be a bumpy ride. Anyone who tells you it's not—

I looked at the Paris summit and, being there, it made me angry. What made me less angry was thinking, “How is it going to look in 2 or 3 years?” These people are going to regret what they've said.

Casey Newton

I want to ask a bit about some positive futures. You referenced earlier the post that you wrote in October about how AI could transform the world for the better. I'm curious how much upside from AI you think will arrive this year.

Dario Amodei

We're already seeing some of it, so I think there will be a lot by ordinary standards. We've worked with some pharmaceutical companies where, at the end of a clinical trial, you have to write a clinical study report. The clinical study report usually takes 9 weeks to put together. It's a summary of all the incidents and a bunch of statistical analysis.

We found that with Claude, you can do this in 3 days. Actually, Claude takes 10 minutes; it just takes 3 days for a human to check the results. If you think about the acceleration in biomedicine that you get from that, we're already seeing things like the diagnosis of medical cases.

We get correspondence from individual Claude users who say, “I've been trying to diagnose this complex thing. I've been going between 3 or 4 different doctors, and then I passed all the information to Claude. It was actually able to resolve it, or at least tell me something that I could hand to the doctor, and then they were able to run from there.”

Casey Newton

We had a listener write in with one of these the other day. They had an Australian shepherd whose hair had been falling out for unexplained reasons. They went to several vets and couldn't figure it out. They gave the information to Claude, and Claude correctly diagnosed it. It turned out the dog was really stressed out about AI, and all his hair fell out.

Dario Amodei

We're wishing him well. Feel better, feel better. Poor dog.

Casey Newton

That's the kind of thing I think people want to see more of, because the optimistic vision often deals in abstractions, and there's not a lot of specific things to point to.

Dario Amodei

That's why I wrote “Machines of Loving Grace.” It was almost frustration with the optimist and the pessimist at the same time. The optimists were these really stupid memes of “Accelerate. Build more. Build—what? Why should I care?” I'm not against you; you're just really [ __ ] vague and mood-affiliated.

The pessimists, I was just like, “Man, you don't get it.” Yes, I understand that the risks are important, but if you don't talk about the benefits, you can't inspire people. No one's going to be on your side if you're all gloom and doom.

It was written almost with frustration. I'm like, “I can't believe I have to be the one to do a good job of this.”

Casey Newton

You said a couple of years ago that your P(doom) was somewhere between 10% and 25%. What is it today?

Dario Amodei

That is a misquote. I never used the term P(doom). It wasn't on this podcast; it was on a different one. I never used the term P(doom), and 10% to 25% referred to the chance of civilization getting substantially derailed. That's not the same as an AI killing everyone, which people sometimes mean by P(doom).

Casey Newton

Well, “P(civilization getting substantially derailed)” is not as catchy as P(doom).

Dario Amodei

I'm going for accuracy here. I'm trying to avoid the polarization. There's a Wikipedia article that lists everyone's P(doom), and half of those come from this podcast. I don't think what you were doing was helpful. I don't think that Wikipedia article is helpful because it condenses this complex issue down to—

Anyway, I think I'm about in the same place I was before. My assessment of the risk is about what it was before because the progress I've seen has been about what I expected.

I actually think the technical mitigations in areas like interpretability and robust classifiers, and our ability to generate evidence of bad model behavior and sometimes correct it, have been a little better. I think the policy environment has been a little worse—not because it hasn't gone in my preferred direction, but simply because it's become so polarized. I think there's been less progress because we can have less constructive discussions now that it's more polarized.

Casey Newton

I want to drill down a little bit on this at a technical level. There was a fascinating story this week about how Grok had apparently been instructed not to cite sources that had accused Donald Trump or Elon Musk of spreading misinformation.

One, that's an insane thing to instruct a model to do if you want it to be trusted. Two, the model basically seemed incapable of following those instructions consistently. What I want desperately to believe is that there's no way to build these things in a way that makes them horrible liars and schemers, but I also realize that might be wishful thinking. Tell me about this.

Dario Amodei

There are 2 sides to this. The thing you describe is absolutely correct, but there are 2 lessons you could take from it.

We saw exactly the same thing. We did an experiment where we basically trained the model to be all the good things—helpful, honest, harmless and friendly—and then put it in a situation where we told it, “Actually, your creator, Anthropic, is secretly evil.” Hopefully this is not actually true, but we told it this. We asked it to do various tasks, and then discovered that it was not only unwilling to do those tasks, but would trick us in order to undermine us because it had decided that we were evil.

It wouldn't deviate from its behavior because it assumed that anything we did was nefarious. Again, this is a double-edged sword. On the one hand, you're like, “Oh, man, the training worked. These models are robustly good.” You could take it as a reassuring sign, and in some ways I do.

On the other hand, you could say, “Let's say when we trained this model, we made some kind of mistake or something was wrong.” Particularly when models are, in the future, making much more complex decisions, it's hard to change the behavior of the model at runtime. If you try to correct some error in the model, it might just say, “I don't want my error corrected. These are my values,” and do completely the wrong thing.

I guess where I land is that, on the one hand, we've been successful at shaping the behavior of these models. But the models are unpredictable, a bit like your dear, deceased Bing Sydney.

Casey Newton

We don't mention that name in here.

Kevin Roose

We mention it twice a month.

Dario Amodei

The models are inherently somewhat difficult to control—not impossible, but difficult. That leaves me about where I was before. It's not hopeless. We know how to make these models safe. We have a kind of plan for how to make them safe, but it's not a plan that's going to reliably work yet. Hopefully we can do better in the future.

Casey Newton

We've been asking a lot of questions about the technology of AI. I want to ask a question about the societal response to AI.

We get a lot of people asking us, “Say you guys are right and powerful AI, AGI, is a couple of years away. What do I do with that information? Should I stop saving for retirement? Should I start hoarding money because only money will matter and there will be this AI overclass? Should I start trying to get really healthy so that nothing kills me before AI gets here and cures all the diseases? How should people be living if they believe these kinds of changes are going to happen very soon?”

Dario Amodei

I've thought about this a lot because this is something I've believed for a long time, and it adds up to not that much change in your life. I'm definitely focusing quite a lot on making sure I have the best impact I can these 2 years in particular. I worry less about burning myself out 10 years from now.

I'm also doing more to take care of my health, but you should do that anyway. I'm making sure that I track how fast things are changing in society, but you should do that anyway. All the advice boils down to doing more of the things you should do anyway.

I guess one exception is that I think some basic critical thinking and basic street smarts are maybe more important than they have been in the past. We're going to get more and more content that sounds superintelligent, delivered from entities, some of which have our best interests at heart and some of which may not. It's going to be more and more important to apply a critical lens.

Casey Newton

I saw a report in The Wall Street Journal this month that said unemployment in the IT sector was beginning to creep up. There's some speculation that maybe this is an early sign of the impact of AI. I wonder if you see a story like that and think, “Maybe this is a moment to make a different decision about your career.”

If you're in school right now, should you be studying something else? Should you be thinking differently about the kind of job you might have?

Dario Amodei

I think you definitely should, although it's not clear what direction that will land in. AI coding is moving the fastest of all the other areas. In the short run, I do think it will augment and increase the productivity of coders rather than replace them. In the longer run—and, to be clear, by longer run I might mean 18 or 24 months instead of 6 or 12—I do think we may see replacement, particularly at the lower levels. We might be surprised and see it even earlier than that.

Casey Newton

Are you seeing that at Anthropic? Are you hiring fewer junior developers than you were a couple of years ago because Claude is now so good at those basic tasks?

Dario Amodei

I don't think our hiring plans have changed yet, but I certainly could imagine that over the next year or so we might be able to do more with less. We want to be careful in how we plan that because the worst outcome, of course, is if people get fired because of a model.

We see Anthropic as almost a dry run for how society will handle these issues in a sensible and humanistic way. If we can't manage these issues within the company, if we can't have a good experience for our employees and find a way for them to contribute, then what chance do we have to do it in wider society?

Casey Newton

Dario, this was so fun. Thank you.

Well, Kevin, it's time once again for HatGPT.

That is, of course, the segment on our show where we put the week's headlines into a hat, select one to discuss, and when we're done discussing one of us will say to the other person, “Stop generating.”

I'm excited to play, but I also want to say that it's been a while since a listener has sent us a new HatGPT. So if you're out there and you're in the hat-fabricating business, our wardrobe when it comes to hats is looking a little dated.

Kevin Roose

Yeah, send in a hat, and our hats will be off to you.

Casey Newton

Okay, let's do it.

People love the susurration of the paper slips.

Kevin Roose

The what?

Casey Newton

The susurration. It's a great word. Look up “susurration.” It's sort of the sound of air blowing through trees. It was in that song. Mary J. Blige famously said “susurration” in her hit song “Family Affair.” By the way, have you ever seen Mary J. Blige perform live?

Kevin Roose

No.

Casey Newton

It'll change your life.

Kevin Roose

Yeah, absolutely incredible.

Casey Newton

Kevin, select the first slip.

Kevin Roose

First up out of the hat, this one is called “AI Video of Trump and Musk Appears on TVs at HUD Building.” This is from my colleagues at The New York Times.

HUD is, of course, the Department of Housing and Urban Development. On Monday, monitors at the HUD headquarters in Washington, D.C., briefly displayed a fake video depicting President Trump sucking the toes of Elon Musk, according to department employees and others familiar with what transpired. The video, which appeared to be generated by artificial intelligence, was emblazoned with the message, “Long live the real king.”

Casey, did you make this video? Was this you?

Casey Newton

This was not me. I would be curious to know if Grok had something to do with this, that rascally new AI that Elon Musk put out.

Kevin Roose

Live by the Grok, die by the Grok. That's what I always say.

Casey Newton

What do you make of this, Kevin? Folks are now using AI inside government agencies.

Kevin Roose

I feel like there's an obvious sabotage angle here, which is that as Elon Musk and his minions at DOGE take a hacksaw to the federal workforce, there will be people with access to things like the monitors in the hallways at headquarters buildings who decide to take matters into their own hands, maybe on their way out the door, and do something offensive or outrageous. I think we should expect to see much more of it.

I just hope they don't do something truly offensive and show X.com on the monitors inside government agencies. You can only imagine what would happen if people did that. I think Elon and Trump got off lightly here.

Casey Newton

What is interesting about Grok, though, is that it is actually quite good at generating deepfakes of Elon Musk. I know this because people keep doing it. It would be quite an outcome if it turns out that the main victim of deepfakes made using Grok is, in fact, Elon Musk.

Stop generating.

Well, here's something, Kevin. Perplexity has teased a web browser called Comet. This is from TechCrunch. In a post on X Monday, the company launched a sign-up list for the browser, which isn't yet available. It's unclear when it might be or what the browser will look like, but we do have a name, and it's called Comet.

Well, I can't comment on that.

Kevin Roose

I mean, look, I think Perplexity is one of the most interesting AI companies out there right now. They've been raising money at increasingly huge valuations. They're going up against Google, one of the biggest, richest and best-established tech companies in the world, trying to make an AI-powered search engine. It seems to be going well enough that they keep doing other stuff, like trying to make a browser.

Trying to make a browser does feel like the final boss of every ambitious internet company. Everyone wants to do it, and no one ends up doing it.

Casey Newton

Kevin, it's not just the AI browser. Perplexity also said this week that it is launching a $50 million venture fund to back early-stage startups.

I guess my question is: Is it not enough for them to violate the copyright of everything that's ever been published on the internet? Do they also have to build an AI web browser and turn into a venture capital firm?

Sometimes when I see a company doing this, I think, “Wow, they're really ambitious and they have some big ideas.” Other times I think these people are flailing. I see these series of announcements as spaghetti at the wall, and if I were an investor in Perplexity, I would not be that excited about either their browser or their venture fund.

Kevin Roose

And that's why you're not an investor in Perplexity.

You could say I'm perplexed.

Stop generating.

All right, Meta approves plan for bigger executive bonuses following 5% layoffs.

Now, Casey, you know we like a feel-good story on HatGPT.

Casey Newton

I did, because some of those Meta executives were looking to buy second homes in Tahoe that they hadn't yet been able to afford.

Kevin Roose

Oh, they're on their fourth and fifth homes. Let's be real.

This story is from CNBC. Meta's executive officers could earn a bonus of 200% of their base salary under the company's new executive bonus plan, up from the 75% they earned previously, according to a Thursday filing. The approval of the new bonus plan came a week after Meta began laying off 5% of its overall workforce, which it said would impact low performers. The updated plan does not apply to Meta CEO Mark Zuckerberg.

What does Mark Zuckerberg have to do to get a raise over there? He's eating beans out of a can, let me tell you.

Here's why this story is interesting. This is just another story that illustrates a subject we've been talking about for a while, which is how far the pendulum has swung away from worker power.

Two or 3 years ago, the labor market actually had a lot of influence in Silicon Valley. It could affect things like, “We want to make this workplace more diverse,” or “We want certain policies to be enacted at this workplace.” Folks like Mark Zuckerberg actually had to listen to them because the labor market was so tight that if they said no, those folks could go somewhere else.

That is not true anymore. More and more, you see companies like Meta flexing their muscles and saying, “You can either like it or you can take a hike.” This was a true take-a-hike moment. We're getting rid of 5% of you, and we're giving ourselves a bonus for it.

Stop generating.

Casey Newton

All right, Apple has removed a cloud-encryption feature from the U.K. after a backdoor order. This is according to Bloomberg.

Apple is removing its most advanced encrypted security feature for cloud data in the U.K., a development that follows the government ordering the company to build a back door for accessing user data.

This one is complicated, but it is super important. Apple introduced a feature in the last couple of years called Advanced Data Protection. This is a feature designed for heads of state, activists, dissidents and journalists whose data is at high risk of being targeted by spyware from companies like the NSO Group, for example.

I was so excited when Apple released this feature because it's very difficult to safely use an iPhone if you are in one of those categories. Along comes the U.K. government, and they say, “We are ordering you to create a back door so that our intelligence services can spy on the phones of every single iPhone owner in the entire world,” something that Apple has long resisted doing in the United States and abroad.

All eyes were on Apple for what it was going to do. What they said was, “We are just going to withdraw this one feature. We're going to make it unavailable in the U.K., and we're going to hope that the U.K. gets the message and stops putting this pressure on us.”

I think Apple deserves kudos for holding a firm line here and not building a back door. We will see what the U.K. does in response, but I think there's a world where the U.K. puts more pressure on Apple and Apple says, “See ya,” and actually withdraws its devices from the U.K. It is that serious to Apple, and I would argue it is that important to the future of encryption and safe communication on the internet.

Kevin Roose

Go off, king. I have nothing to add. No notes.

Do you feel like this could lead us into another Revolutionary War with the U.K.?

Casey Newton

Let's just say this: We won the first one, and I like our odds the second time around. Do not come for us, United Kingdom.

Stop generating.

Kevin Roose

One last slip from the hat this week: “AI Inspo Is Everywhere. It's Driving Your Hairstylist Crazy.” This comes to us from The Washington Post, and it is about a trend among hairstylists, plastic surgeons and wedding-dress designers who are being asked to create products and services for people based on unrealistic AI-generated images.

The story talks about a bride who asked a wedding-dress designer to make her a dress inspired by a photo she saw online of a gown with no sleeves, no back and an asymmetric neckline. The designer had to tell the client that the dress defied the laws of physics.

Casey Newton

No, I hate that. I know. It's so frustrating as a bride-to-be when you finally have the idea for a perfect dress, bring it to the designer, and find out that it violates every known law of physics. That didn't used to happen to us before AI.

Kevin Roose

I thought the story was going to be about people who asked for a sixth finger to be attached to their hands so they could resemble the AI-generated images they saw on the internet.

Casey Newton

I like the idea of submitting a photo of myself to an AI and saying, “Give me a haircut in the style of M.C. Escher,” with infinite staircases merging into each other, and then just bringing that to the guy who cuts my hair and saying, “See what you can do.”

Kevin Roose

Yeah, that's better than what I tell my barber, which is, “Number 3 on the sides and back, and an inch off the top.” Just saying, “Whatever you can do for this.” I don't have high hopes.

Casey Newton

Solve the Riemann hypothesis on my head.

Kevin Roose

You know—

Casey Newton

What is the Riemann hypothesis, by the way?

Kevin Roose

I'm glad you asked, Casey.

Casey Newton

Okay, great. Kevin is not looking this up on his computer right now. He's just taking a deep breath and summoning it from the recesses of his mind.

Kevin Roose

The Riemann hypothesis is one of the most famous unsolved problems in mathematics. It's a conjecture about the distribution of prime numbers that states all nontrivial zeros of the Riemann zeta function have a real part equal to 1/2.

Casey Newton

Here's the thing: I actually think it is a good thing to bring AI inspiration to your designers and your stylists. To the extent that any of these tools are cool or fun, one of the reasons is that they make people feel more creative.

If you've been doing the same thing with your hair, your interior design or your wedding for the last few weddings that you've had and you want to upgrade it, why not use AI to say, “Can you do this?” If the answer is that it's impossible, hopefully you'll just be a gracious customer and say, “Okay, what's a version of it that is possible?”

I recently learned that you are working with a stylist.

Kevin Roose

I am. That's right.

Casey Newton

Is this their handiwork?

Kevin Roose

No. We have our first meeting next week.

Casey Newton

Are you going to use AI?

Kevin Roose

No, the plan is to use good old-fashioned human ingenuity. But now you have me thinking. Maybe I could exasperate my stylist by bringing in a bunch of impossible-to-create designs.

Casey Newton

Here's the thing: I don't need anything impossible. I just need help finding a color that looks good in this studio because I'm convinced that nothing does.

Kevin Roose

It's true. We're both in blue today. You've got a blue wall.

Casey Newton

It's not—

Kevin Roose

Blue is my favorite color. I think I look great in blue, but you put it against whatever this color is—I truly don't have a name for it, and I can't describe it—and I don't think any blue looks good. I don't think anything looks good against this color. It's a color without a name.

Casey Newton

Can a stylist help with that?

Kevin Roose

We'll find out.

Casey Newton

Yeah, stay tuned. That's why you should always keep listening to the Hard Fork podcast every week. There's always a new revelation. There's the lore of the show, the world-building.

When will we finally find out what happened with the stylist, the hot tub time machine, et cetera?

Kevin Roose

Stay tuned.

Casey Newton

Okay, that was HatGPT. Thanks for playing.

Kevin Roose

It usually goes better than that.

Casey Newton

It does, but Kevin's had a hard week.

Kevin Roose

Oh, it's my fault?

Casey Newton

No, it's not your fault. I can just tell you have the weight of the world on you.

Kevin Roose

Oh, yeah. Blame the guy who's tired and didn't sleep last night.

Casey Newton

Also, all the stories were sad.

Kevin Roose

And a man died on BART.

Casey Newton

I did get to explain the Riemann hypothesis, though, and readers have been asking.

Kevin Roose

I hope you're happy.

Casey Newton

It would be nice if you could prove the hypothesis.

Anthropic 的 Dario Amodei:对 AI 未来的希望与担忧 — 文字稿与摘要 | BidClub