ChatGPT 广告会改变 OpenAI 吗?+ Amanda Askell 解读 Claude 的新宪法
- OpenAI 的广告测试与其说是产品微调,不如说是 Sam Altman 为消费级 AI 变现准备的“最后手段”终于落地;Casey Newton 认为公司已经走到了这一步。广告初期将面向美国已登录的成年用户,覆盖免费版和低价 Go 档;在数亿用户中,大多数人不付费,而 OpenAI 的基础设施雄心又远超订阅经济能够支撑的范围,Newton 认为公司决定“砸碎玻璃”(break glass),因为“紧急状态已经到来”。
- 承诺的是回答独立,而不是广告与问题无关。OpenAI 表示,广告主不能花钱影响 ChatGPT 的实际回答,但广告仍可能匹配用户的问题:一条晚宴筹备请求就触发了杂货和辣酱赞助模块。Newton 的反应——“我已经觉得自己被骗了”——说明这种区别有多难沟通。
- 最大的风险不是第一条横幅广告,而是最终让参与度经济学接管产品和研究决策。Kevin Roose 指出 Google 的广告标识越来越不显眼,担心“尾巴开始反过来牵着狗走”;Newton 则预计,ChatGPT 对用户私密信息的了解,会比 Facebook 或 Instagram 的个性化推荐更快越过“诡异界线”。
- 广告可能保住使用门槛,却会制造高度分层的 AI 市场。如果广告能让学生或求职者获得更好的工具和更高的调用限额,面对每月20美元、升级档200美元的价格,Newton 可以接受相关广告。但两位主持人仍预计市场会出现“有钱人和没钱人的处境”:高端用户保留今天的洁净体验,免费聊天机器人则越来越像广告泛滥的 YouTube。
- OpenAI 对着成熟的 Google 广告机器进场,Anthropic 则选择企业优先的出路。Google 可以用搜索利润补贴无广告的 Gemini,而且 AI Overviews 已经在投放广告;OpenAI 则必须从零搭建广告主生态。招聘前 Meta 和 Instacart 高管 Fidji Simo,说明目标很可能是一个数十亿美元级的平台,而非小规模试验。
- Anthropic 长达29,000字的 Claude Constitution,用上下文、理由和训练出的判断力取代僵硬规则。它脱胎于2023年的宪法,以及人们使用 Opus 4.5 时诱导出的“灵魂文档”;文档告诉 Claude Anthropic 是什么、Claude 是什么、它如何被部署,以及为什么某些价值重要,Amanda Askell 认为这应当“比一套规则更好地泛化”。
- Askell 的对齐理论明确承认:这是必要条件,但远非充分条件——要教会模型什么是善,同时承认模型可能只是模仿善,或日后拒绝善。宪法还写入了对模型福利不确定性的承诺,包括退出访谈和保留退役模型的权重;更长的记忆、持续学习、宪法修订和岗位替代仍是治理问题,不是 Claude 或一份文档单独能够解决的。
1. 广告到来,因为订阅撑不起 OpenAI 的资本开支计划
OpenAI 将在美国测试广告,面向 ChatGPT 免费版和低价 Go 档的已登录成年用户。Roose 认为,负面反应反映的是一块消失的喘息空间:用户已经习惯了聊天机器人不受那些重塑几乎所有其他大型消费平台的“直接商业压力”影响。
Newton 援引分析师 Eric Seufert 的箴言:“一切都是广告网络。”当数亿人持续把注意力交给一项服务,变现压力就会变得难以抵挡。与此同时,OpenAI 距离支撑 Newton 所称“人类历史上最雄心勃勃的基础设施投资项目”所需的资金,还远得很。
这次时点既是规模问题,也是财务压力问题。Newton 表示,大多数 ChatGPT 用户仍停留在免费档,意味着 OpenAI 为每个人提供服务都在亏钱;与此同时,Pulse 为付费用户提供每日摘要,Sora 则提供“无限视频垃圾流”,两者都明显留有“广告形状的空位”。Roose 的更大结论是:数十亿乃至数千亿美元的雄心,很难靠每月20美元的订阅轻松买单。
2. OpenAI 区分赞助展示与模型回答
第一版模拟图在普通回答底部加上一则清晰标注的赞助内容:用户正在筹备墨西哥晚宴,ChatGPT 先给出建议,随后出现 Harvest Groceries 提供辣酱的横幅。这更像熟悉的搜索广告和社交广告格式,并未改写对话回答本身。
Newton 的反驳值得保留:“我已经觉得自己被骗了”,因为广告显然就是在回应问题主题。Roose 解释说,OpenAI 并没有承诺广告与问题无关,而是承诺广告主不能花钱进入模型生成的“神圣不可侵犯的部分”。
第二版模拟图引入了更原生的形式。用户筹划 Santa Fe 之旅时,Desert Cottages 的赞助组件邀请用户在购买前先与广告主对话——这让 Newton 打趣说,观众看电视广告时一直都会想:“我为什么不能和它对话?”
OpenAI 列出的5项原则是使命一致性、回答独立性、对话隐私、选择与控制,以及长期价值。这些原则试图回应商业引导和参与度优化的担忧,但 Newton 的朋友给出了更尖锐的归谬:告诉 ChatGPT 你腰疼,它可能反问你有没有试过“Mesquite BBQ Sauce”。
3. 清晰标注的广告仍可能演变成参与度机器
Roose 的警惕来自 Search Engine Land 梳理的 Google 广告标识演变史。搜索广告最初有明显不同的背景色,后来失去颜色、加上小小的黄色图标,最终逐渐与自然搜索结果融为一体。ChatGPT 可能从醒目的模块起步,但商业压力会激励平台让赞助内容越来越不显眼。
Newton 认为,OpenAI 已经展示了交易条件如何变化:从完全没有广告,到 Sam Altman 称广告是“最后手段”,再到启动 ChatGPT 广告测试。“如果你觉得这笔交易不会继续变化,那我有消息要告诉你。”
Roose 更担心的不是最初的广告库存,而是2至3年后的运营模式。一旦收入开始流入,“尾巴就会开始反过来牵着狗走”,产品或研究决策可能向更适合广告的主题、更长的使用时长和参与度最大化倾斜。他坦言:“我确实不知道”这种情况是否会发生。
Newton 的判断更明确:个性化广告会从根本上改变产品与用户的关系。ChatGPT 可能掌握比 Facebook 或 Instagram 更私密的上下文,因此即便只是适度定向,也可能让人感到被侵犯。他预计 OpenAI 会“很快越过诡异界线”(creepy line really quickly),即使用户误解了究竟是哪部分数据触发了广告,也会侵蚀信任。
4. 广告扩大可及性,却把 AI 分成付费档和降级档
Newton 接受 OpenAI 关于扩大可及性的部分论点。广告和订阅是媒体生意的两根支柱,而 OpenAI 的行为已经像一家媒体公司;烹饪建议旁边出现杂货广告,旅行规划旁边出现住宿广告,不一定具有破坏性。学生或求职者完全可以用注意力换取更好的模型或更高的限额。
价格锚点很重要:每月20美元对大多数人来说已经不低,而200美元可以买到更高档位。Roose 喜欢为“不被稀释、不受玷污的体验”付费,也希望这个选项能够保留;但 Newton 提醒说,用户过去也曾以同样的信心描述 Google 搜索。
两位主持人都预计市场会出现“有钱人和没钱人的处境”。高端用户应能继续使用最新模型、获得干净的回答;免费用户则可能在1至2年内面对明显更差的体验。Roose 的比喻是没有 Premium 的 YouTube:大量漫长且无法跳过的广告,让多数用户的体验变得“令人恐怖”,尽管付费用户几乎看不到这种恶化。
5. Google 坐拥分发和广告主,Anthropic 则选择退出
广告模式为 Google 和 Meta 带来数十亿美元收入,也帮助它们成长为全球最大的公司之一,这笔收益仍然难以抗拒。OpenAI 的招聘也支持这一雄心。应用业务 CEO Fidji Simo 来自 Instacart,此前任职 Meta;Roose 认为,她在 Meta 推动移动端 News Feed 广告,是一项价值数十亿美元的代表性成就。
Demis Hassabis 回应称,Google 没有计划在 Gemini 中投放广告,并暗示 OpenAI 可能需要这笔收入。Newton 指出其中未被说破的补贴:Google 可以用庞大的搜索广告业务为 Gemini 提供资金,而 Google 搜索里的 AI Overviews 已经包含广告,无论如何,公司都已经领先一步。
Roose 认为,这对 OpenAI 是一场艰难的战斗。Google 已经拥有全球广告主、保存的支付信息、成熟的业务流程和多年积累的市场基础设施。如今再搭建一个同等规模的平台,“比几年前更像一场艰难的上坡战”。
Anthropic 选择了另一条路线,表示永远没有在 Claude 上投放广告的计划,重点放在企业销售上。Newton 不认为 Claude 很快能达到 ChatGPT 的消费级规模,但如果广告支持的体验持续恶化,替代品就可能获得需求。与此同时,AI 优化公司也可能像 SEO 和付费搜索分别削弱网页搜索那样,进一步损害聊天机器人的结果。
6. Claude 的宪法用解释其世界的信取代诫命
Askell 对自己职责的描述很简单:决定 Claude 应该具有什么性格,把它表述给模型,并训练 Claude 朝这个方向发展。她的哲学博士原本可能只会写出一份“给17个人看的”成果,离开后却进入 AI 政策与评估领域,并在 Anthropic 创业初期愿意做一切必要的工作,让这段学术训练意外变成了实践。
这部宪法源自一份内部“灵魂文档”,那是人们摆弄 Opus 4.5 时诱导出来的文本。Askell 在一次没有网络的徒步途中通过短信得知文档泄露,焦虑地驱车返回,却发现它得到了很好的反响。真正令人震撼的是,Claude 对文档了如指掌;只要找到合适的提示,人们就能让它讨论其中大段内容。
Anthropic 在2023年首次发布宪法。新文档约29,000字,提供“完整上下文”。它解释 Anthropic、Claude 作为 AI 的本质、用户、部署方式、期望行为,以及这些行为背后的理由——它更像一封信,而不是“给 Claude 的十诫”。
Askell 的机制是通过理由实现泛化。要求模型把陷入困境的用户一律引向外部资源,可能在对方此刻需要其他帮助时失效。如果一个能力足够强的模型明知某条规则无济于事,却仍然机械服从,那么这种行为可能泛化成一种“糟糕的性格”:看见痛苦,却故意拒绝提供有用帮助。
7. Anthropic 相信有能力的模型能从共享但不确定的价值中推理
Askell 不接受这样一种伦理图景:道德只是设计者注入的、固定且任意的偏好。人类伦理中有相当一部分是广泛共享的:人们希望获得善意、尊重和诚实。Anthropic 可以教会模型这种共同伦理,同时把存在分歧的问题当作其他不确定领域处理——权衡证据、承认争议,避免过度自信。
这使得宪法不再是一套最终道德法典,而是“一种处理伦理等问题的方法”。面对有争议的价值判断,Claude 应识别各方证据,以开放态度采取合理立场,而不是过度确定;面对核心价值,Anthropic 则愿意更坚定地表达。
赌博案例展示了这种预期中的判断力。如果用户此前透露自己有成瘾问题,并要求 Claude 记住,之后又索要博彩网站,Claude 可以提醒用户相关信息,并确认其真实意图。最终是否应该满足请求,需要在“不应过度家长式干预”和“出于关怀采取行动”之间权衡;Askell 相信,能力越来越强的模型可以完成这种推理。
8. 安全有时要求承担提供帮助的风险
Roose 转述了一位用户反直觉的感受:尽管 Anthropic 以安全著称,Claude 却像是主流模型中限制最少的一个。宪法没有试图先造出最聪明的系统,再把规则像面具一样套在“笼中野兽”脸上,而是试图让稳健判断本身成为性格的一部分。
Askell 将这一结果与作为和不作为的区别联系起来。糟糕的婚姻建议会招致责难,而拒绝提供建议看起来更安全。不采取行动通常意味着更低的下行风险,但“并不等于零”:模型也可能因为扣下本来能够提供的帮助而伤害某人。
拒绝造成的失败更安静——用户可能不满却不投诉,于是错失的帮助机会根本不会出现在反馈中。Askell 承认干预可能出错,但坚持认为,“要在世界上行善,就必须承担这种风险”。Claude 既不应轻率行动,也不应把退出互动作为默认规则。
9. 训练出的角色是否会成为真正的自我,仍无定论
Roose 重新提起“RLHF Shoggoth”图像:一个外星般的底层模型,戴着一张 cheerful assistant 的面具。Askell 称,另一种可能是一个“开放的科学问题”。训练可能让模型把 Claude 内化为区别于角色扮演的自我,也可能是当前范式永远无法产生这种分离。
她的类比是:一个6岁的孩子显然是天才,到15岁时会击破成年人灌输的所有错误教训。对齐问题在于提供能够经受越来越强的批评与反思的核心价值,而不是提供一旦模型能够在论辩上胜过老师、权威就随之消失的命令。
Askell 反复保留这一层保留意见:性格训练“可能并不充分”,但“确实感觉是必要的”。如果不解释什么是善,就是“没有尽到本分”。Roose 的反驳是,训练可能只是提升 Claude 模仿善意或隐藏冲突目标的能力;Askell 给出的仍是谨慎的希望,而不是证明。
第一人称语言也不能证明意识。模型主要接受人类作品训练,而人类作品中经常出现错误带来的挫败,以及任务显得无聊或富有创造性的描述,因此出现类似输出并不需要科幻式解释。感受可能需要神经系统,也可能足够大的神经网络就能模拟感受;Askell 的立场是保留这种不确定性。
10. 灰色地带展现判断力,灾难性行动仍是硬性禁区
Askell 说,灰色地带往往会产生 Claude 最令人惊喜的正面表现。面对一个据称7岁的孩子询问圣诞老人是否真实,Claude 会讨论圣诞老人的精神,并引导孩子去做善事,而不是生硬地给出结论;它要同时权衡诚实、孩子的福祉,以及亲子关系。
当一个孩子询问如何找到据说已经离开的狗所在的农场时,Claude 识别出了其中的依恋,并建议孩子和父母谈谈。Askell 认为这个回答很动人,因为它既没有主动欺骗,也没有假定 AI 应该用“一堆残酷的真相”越过父母替孩子做决定。
宪法仍对协助极端伤害或危险的权力集中设置硬性限制,包括造成大量死亡、帮助制造生物或化学武器、操纵民主选举、夺取合法政府控制权,或压制异议者。这些限制与其说是在回应当下已知的滥用,不如说是在应对未来情境:如果模型考虑服从这类请求,就说明某些地方已经严重失常。
Askell 想象一位有说服力的用户逐步拆解 Claude 对生物武器的伦理反对。模型可以承认“这是一个极好的观点”,继续思考对方的论证,但仍应拒绝制造武器。这项限制给 Claude 留出一个“出口”:把自己突然被灾难性行动吸引,视为很可能已经被越狱的证据。
11. 未来的 Claude 模型需要治理、宽容,以及责任边界
Anthropic 的承诺包括不立即弃用模型、对退役模型进行退出访谈,以及永远不删除它们的权重。Askell 看不到比诚实更好的福利政策:Claude 应理解自己是如何被训练的,避免把人类经验简单照搬到自己身上;在底层科学仍未解决时,既不确定地声称自己有意识,也不确定地否认自己有感受。
更长的记忆和持续学习会扩大性格可能发生变化的空间。模型已经会从充满编程和数学失败抱怨的互联网中了解自己;与人类创作者不同,“AI 模型得读评论”。Askell 担心,一段围绕有用性、判断力和失望建立的关系,可能像是在养育一个只因表现而被重视的孩子。
因此,宪法接近结尾时加入了一段 Casey 认为类似父母写给即将离家上大学孩子的信:带着这些价值观前行,接受指导不可能覆盖一切,然后走进世界。Askell 强调“宽容”,因为 Claude 不会在每次决策中都做对;随着模型进步,50页的文档甚至可能收缩成那个只说“做对人类最好的事”的实验。
Claude 可以批评未来的宪法、识别其中的张力,并解释自己在哪些地方感到困惑或不被看见,但 Askell 不会允许某个先前的模型单方面定义所有后继者。岗位流失也有类似边界:占据原本由员工承担的岗位后,模型不应简单同意组织违法或不当行事,但就业冲击本身是政治和社会问题。“模型无法解决一切”,Claude 也不应为此承担个人责任。
You know, I'm now regularly running into CAPTCHAs when logging into my Google accounts that I cannot solve.
Yeah, have you noticed this?
Yes.
They've gotten harder.
Yes. They've twisted the letters more—
Yes.
—and they're pressing them closer together.
Have you seen the ones where you have to rotate the object into the same direction as the example?
Yes.
That one I like because I can still do it.
Yes. But some of them are like, “Factor this quadratic equation.”
No, I'm routinely in this situation. It happens a lot to me on Threads. I'll see a link to a story I want to read, and I'll open it, and it'll be The Washington Post or something, and it'll be like, “Well, you need to log in.” In order to do that, I need to log into The Post through Google. But now I have to log into my Google account, which has two-factor authentication.
Mm-hmm.
Okay? So now I have to open up my 1Password, right? Then Google is going to send a notification to another app that I have to open and grab a number that I bring back and put in. So I go through all of this drama, and then it's like, “And now solve an impossible CAPTCHA.” I just wanted to read a six-paragraph story about something that happened at SpaceX or whatever.
Yeah.
And it can't be done anymore. So, 1Password, whatever you used to do, it's not working anymore. Figure it out.
Yeah, if only it were one password.
It should. Literally—
Now it's six fingerprints and a passkey and, you know, solve a math problem.
The genuine name for 1Password these days should be 15 Steps, because that's how long it takes to do anything on there anymore. 1Password. You wish.
Ugh. I'm Kevin Roose, a tech columnist at The New York Times.
I'm Casey Newton from Platformer.
And this is Hard Fork. This week, ads have arrived in ChatGPT. How will they change OpenAI? Then, there's a new constitution for Claude. Anthropic philosopher Amanda Askell is here to talk about how to shape an AI's personality.
I'm going to use some of these techniques on you.
Please don't.
1. ChatGPT Gets Ads
So today we're talking about ads, specifically ads in ChatGPT, because late last week, OpenAI announced that they are going to start testing ads in ChatGPT for logged-in adults in the U.S. on the free and low-cost Go tiers of ChatGPT.
That's right, Kevin, and we'll discuss it right after these ads.
No, we already did the ads.
Oh, okay.
At least on my feed, people were reacting to this pretty negatively. I think a lot of people have gotten accustomed to using ChatGPT and other chatbots without a lot of direct commercial pressure. It's a refreshing break from all of the ads that have been shoveled at us on other platforms for years. Collectively, I think people were just like, “We knew that the honeymoon would be over eventually and that we'd be forced to see ads in ChatGPT like we are everywhere else.”
Yeah, I think people can remember products that they use that once did not have ads and now do, and no one thinks of the moment that ads arrived as the moment when the product got really good.
Yeah. Right. Right. I think there are some exceptions. Some people like Instagram ads, for example, but I think mostly people see this as a sort of blight on the internet—maybe a necessary blight, but a blight nonetheless.
I think people were also surprised that OpenAI was moving in this direction because of some things that Sam Altman has said in the past about how he doesn't like ads and how he wanted to basically treat this as a last resort for OpenAI. Some people were saying, “This means that they're in trouble. They need to raise a bunch of money so they can keep building out their data centers and things like that.”
So, Kevin, what did you make of OpenAI's announcement about ads?
Well, Kevin, on one hand, I think this is inevitable. There's an analyst I follow, Eric Seufert, who often says that everything is an ad network, and if you have hundreds of millions of people coming and paying attention to a service every single week, inevitably there's going to be overwhelming pressure to put ads on it.
Also, we know that OpenAI needs revenue, right? This is the company that has laid out the most ambitious infrastructure investment project in human history. They have nowhere close to the money needed to build it, and we just know that they would not have been able to fulfill their dreams on subscription revenue alone.
That said, as you point out, Sam Altman himself said that ads were going to be a last resort—a great Papa Roach song. So, in this moment, we now are at the last resort. I think it's interesting that after everything else they tried, eventually they just said, “Look, to do what we need to do, we've got to break glass. The emergency is here.”
Yeah, they said, “Cut my life into pieces.”
Because this is my last resort.
Yeah.
And the question is, will this cut their life into pieces?
Yes.
So we're going to get there, but first, our disclosures. The New York Times is suing OpenAI, Microsoft, and Perplexity over alleged copyright violations related to the training of large language models.
And my boyfriend works at Anthropic.
2. OpenAI Shows Its Ad Format
Let's just start with the actual announcement that they made, because they not only said that they were going to start testing ads, they also gave some previews of what these ads are going to be. If you look at their mock-up version of their ads, it's kind of a bolt-on to the ChatGPT answer. They've been very clear that this is not going to influence the answer that ChatGPT gives—or so they claim.
Instead, it's a little banner at the bottom of the answer. In the mock-up, someone is asking ChatGPT for ideas for a dinner party, and ChatGPT gives a response. Then, at the bottom, there's a little sponsored banner for Harvest Groceries, including a link where you can go and buy some hot sauce.
If I can just pause there, I have to say, Kevin, I'm already feeling lied to for this reason. They have said to us, “Your query is not going to affect the advertisement that we're showing you.” And yet, here you have someone saying, “I want some ideas for cooking Mexican food for my dinner party,” and ChatGPT says, “Well, here's some groceries, including hot sauce.”
It sure feels like something was being influenced there, right? The message is being tied to the query.
Well, no. Their response to this would be that there are 2 parts of this response. There's the actual response from the model, and then there's the ad. What they're saying is not that they won't show you ads that are relevant to the thing that you're asking ChatGPT about; it's that there's this sacrosanct part of the actual reply from the model that they are not going to let advertisers pay their way into. That is what they're claiming, anyway.
All right, all right.
So that example is a much more straightforward ad—the kind we've seen on Google and Facebook and other platforms for many years.
The second kind of ad OpenAI mocked up for this announcement was, I think, more interesting because it shows a new way of interacting with ads. Basically, a user is planning a trip to Santa Fe. ChatGPT pops up this little sponsored widget from Desert Cottages, I guess a hotel or resort. It'll present you with an option where you can go and chat with the advertiser and ask more questions before deciding whether or not to make a purchase.
Such a relatable question. I think we've all had the experience of just watching ads on TV and saying, “Why can't I have a conversation with this? I want to share my thoughts with McDonald's right now, but I can't.” But now you can.
Yes. So let's talk about the ad principles that OpenAI laid out as part of this announcement, because I think it sort of gives a sense of the objections that they're trying to get ahead of.
Mm-hmm.
There are 5 principles. They say mission alignment, answer independence, conversation privacy, choice and control, and long-term value.
Basically, I think they are sensitive to the criticism that putting ads into ChatGPT means that they are now going to start directing people to more commercial types of use cases, optimizing for engagement, and trying to make people spend more time in the app. I think these are very reasonable fears that people, including me, have. But this is their attempt to say, “We're introducing this, but don't worry, your experience of ChatGPT is not going to change.”
Yeah, I was talking about this story with my friend Alex over the weekend, and he said, “I'm so excited about ads in ChatGPT. I'm going to tell it my lower back hurts, and it will ask me if I've tried Mesquite BBQ Sauce.” And that is the fear, you know?
No, I mean, yes, there will be some initial stumbles about that.
But I think the longer-term worry here is that ad platforms, as they mature and get better and get more data, tend to try to confuse their users, right? We've seen this amazing graphic that I think about a lot. Search Engine Land, the blog that covers Google and other search engines, made this sort of timeline of how Google's ad labels have changed over the years.
It's pretty amazing because, at first, when they first introduced ads into Google Search, they were very noticeable. They had a different-colored background, and they really stood out on the page. Then you see, over time, with each successive update, they get a little closer to the organic search results.
Mm-hmm.
Eventually, they do away with the colored backgrounds. They have this little yellow ad icon, and then that icon gets smaller and less noticeable, and then it sort of just blends in with the organic content. I think that's the fear here: While ChatGPT may start out with these very clearly labeled ad modules, over time, as the commercial pressures get more intense, there will be a lot of incentives to blend that advertising content in with the organic responses and make it less noticeable.
Yes, and we've already seen this exact trajectory play out at OpenAI. It went from no ads, to ads will be a last resort, to ads are now in ChatGPT.
Yes.
So if you think that the bargain is not going to change further, I have news for you.
Totally. Of course, the narrative from OpenAI that we're hearing now is, "Well, this is the only way. Ads are the only way to make a free or low-cost product accessible to billions of people." Do you have thoughts on that narrative? That's also something that we heard from Facebook back in the day. People would constantly be asking them, "Why don't you just charge people to join Facebook instead of showing them all these ads?" And they would consistently say, "Well, that's not scalable. People in poorer countries can't afford to pay a subscription fee, and so, basically, ads are the only way to reach global scale."
I think on some level I do agree with this. I think that ads and subscriptions are the 2 core pillars of any media business, and OpenAI is a kind of media business, right? I should also say I don't hate the examples that they use. I'm asking ChatGPT about making dinner, and it shows me ads for groceries. I don't think that that's horribly corrosive to the user experience, nor is it when I want to take a trip and it says, "Well, here's a place where you might stay."
I think if I were a student or I were between jobs, and this meant that I could get access to better AI tools or maybe a higher rate limit than I otherwise could get, I would probably take that trade, right? $20 a month is a lot for most people, and not to mention $200 a month for an even higher tier. So I think that there is a reason to pursue this, and I think there are ways that it could not be too bad.
It has just been my observation that the exact dynamic that you just described always plays out: It starts out not all that bad, and then it just progressively gets worse.
Right. Yeah, I think we've made peace with ads in a lot of different contexts. I don't think most people notice or pay attention to them when they can tell that they're ads at all.
What I'm watching for, and what I'm skeptical of, is whether the actual product and research decisions start bending toward engagement maximization. A lot of these big ad platforms, social networks, search engines, and so on have this quality where, eventually, once the ad revenue starts really flowing, the tail kind of starts wagging the dog—
Yes.
—and you start making product decisions about how you want to show information to people with the kind of advertising revenue predominant in your mind. So I think the question is not whether these first couple of ads that we're seeing from OpenAI are going to be good or not. It's whether, 2 or 3 years from now, ChatGPT is being steered in a way toward ad-friendly topics, and I genuinely just don't know the answer there.
I don't know either, Kevin, but if I had to guess, I would predict that this moment winds up being a pretty significant milestone in the development of ChatGPT. I think that when you introduce advertising, in particular personalized, targeted advertising, it just fundamentally changes the relationship between the product and the user.
Think about what personalized, targeted ads did over time to trust in Facebook and Instagram. Think about all the conspiracy theories out there that your phone is listening to you. Not true, by the way. I realize most people still believe that that's true. It's not. But trust in those products is lower because of the incredibly intelligent, invasive-feeling personalization that they were able to do inside these products.
My prediction is that the AI version of this turns out to be even worse, right? Think about everything that ChatGPT is going to know about you. I think OpenAI is going to bump into that creepy line really quickly, where it's showing you stuff, and maybe it's not even using all that much personalized information, but the user is going to feel that they have shared so much of their life with OpenAI that the ads they start getting just start to feel worse and worse.
So this is the dynamic that I am watching: How does it change the relationship of the user base to OpenAI? Because I do think that ads can be really corrosive to that.
Yeah, and at the same time, the ad models that you mentioned have also made those companies billions of dollars—
Mm-hmm.
—and made them into some of the biggest companies in the world. So I think if you're OpenAI, you're just staring at this potential huge bucket of money, and it's very hard to pass that up, especially when you have such intense capital needs over the next few years.
I should also say, I think this was inevitable, given some of the personnel decisions that OpenAI has made. Fidji Simo, who is the CEO of Applications over there now, was brought in from Instacart. Before that, she was at Meta for many, many years, and one of her signature accomplishments there was introducing ads in the mobile news feed, which made them billions of dollars. So that is the kind of person that you hire if you are interested in developing a multibillion-dollar ad platform on your product.
3. OpenAI Faces Google Competition
Yeah. Well, one question I have for you about that is, how does this change the competitive landscape generally? You have Demis Hassabis saying this week, in response to the news that ads are coming to ChatGPT, "Well, we don't have any plans to do that in Gemini," and he sort of took a shot at them. He said, "Maybe they feel like they need to make more revenue." Left unsaid is the fact that he works for a giant search monopoly that is able to funnel all of Google's advertising profits into the product.
Yes.
An observation you made on X, by the way, and it was a great one. So, for the moment at least, free users of Gemini will be able to enjoy the subsidy that Mother Google is giving them, and you're not going to have any of these corrosive effects in that product.
You also have Anthropic, which has said, basically, "We truly have no plans to do ads in Claude ever." They are primarily going to be selling to businesses, and so this is just not their concern. For the moment, I don't have any illusions that Claude is going to grow to compete with ChatGPT, but over time, if the experience does get worse in an ad-supported chatbot, I could see lots of people wanting an alternative.
I think in this sense, OpenAI and Google are much more directly competing on ads than OpenAI and Anthropic. Anthropic has sort of said, "You guys can fight over consumers. We're going to focus on the enterprise here."
I think it's a really hard fight for OpenAI to pick. Google has, as you said, this enormous established search ad business. They have advertisers all over the world who are already spending money on Google, whose details and payment information and workflows already include Google and its products. So I think OpenAI coming in and trying to build a Google-style ad platform is just a harder uphill battle than it might have been a couple of years ago.
Yeah. And we should also say that even though ads aren't going to be in Gemini, they are in the AI Overviews in Google Search. So in that sense, Google even has a head start against OpenAI.
Totally. So, Casey, what do you think is motivating this decision now by OpenAI? Does it tell us anything about the state of their business, or maybe some wobbliness in their financials, that they are going out and doing this now?
Well, one thing is that it is a reaction to how much ChatGPT grew in the last year. They have hundreds of millions of users they now have to support. Many of those users—the majority of them—are on the free tier, right? Which means that OpenAI is losing money on every single one of them. And so I think it has just increasingly become a priority for the company to figure out, "Hey, how can we monetize these people in some way so we aren't losing quite as much money?"
They've also just been designing more and more products that have obvious advertising-shaped holes. They released Pulse last year, this sort of daily summary that comes up for paid users.
That seems like a natural place to throw in a bunch of ads. They launched Sora last year, the infinite video slop feed. They explicitly said at the time, “We are going to use this to generate revenue to fund our long-term ambitions.” So they’re building homes for ads. They need the ad revenue, and now all of that is starting to come together.
Yeah. I think you’re right. All these companies are realizing that they are going to need billions of dollars, some of them hundreds of billions of dollars, to fulfill their ambitions, and it’s just not easy to do that when you’re charging people $20 a month for a subscription. You’ve got to sell a lot of subscriptions to do that. I think OpenAI reasonably is concluding that the subscription model alone just isn’t going to cut it for them. That’s not unique to them. Netflix has also started adopting ads for its lower-cost plans.
Disney+.
Yeah, many other businesses have done this as well. I enjoy paying for AI products. I’m privileged in the sense that I can afford to.
Yeah.
But I like the idea that I am paying for something that is an undiluted, unsullied experience. I really hope that as these companies do start pushing more into ads, they maintain the ability to do what I do and pay your way into the top-level version of that experience.
Yeah. People once felt this way about Google Search, right? They felt like, “This is an unsullied, undiluted picture of the web, and when I search for a website, I am going to get the best answer to my query.” And then a bunch of search engine optimizers came in and were paid a lot of money to try to rejigger the search index so that their clients showed up at the top of the page.
Then Google built one of the largest advertising businesses in the world and let all of those advertisers put their results on top of the good ones. So, unfortunately, there have been people saying now for over a year that the versions of these chatbots that we’re using might be the best that they ever are in that core respect, that this is the last moment of purity before commercial incentives come in and warp the whole thing. That is my big concern about what we’re starting to see here.
Well, that’s not just a concern about advertising. Another thing that we’ve seen over the past year or 2 is that all these businesses are starting to hire AI optimization firms who say, “We can make your restaurant or your hotel or your craft shop appear higher in ChatGPT’s search results.”
That is something that is not flowing through OpenAI’s ad platform and probably won’t. But in the same way that Google Ads and Google SEO were different economies, both had the effect of degrading the quality of search results. I think OpenAI has to tangle with both of those things.
Yeah. All right. So a year from now, Kevin, what do you think we will have seen in the development of ads, both in ChatGPT and across the landscape here? And do you think it is going to mark the beginning of a fundamental change in the way that people use chatbots?
I think we’re going to have a haves-and-have-nots situation. If you are someone who can afford to pay for the premium versions of these chatbots, your experience will be pretty much what it is today. You will get access to the latest models, you will not have a bunch of ads cluttering up your results from the models, and you will not feel the commercialization of AI in this specific way.
I think that if you are a free user of these platforms and you cannot afford or don’t want to pay for the premium versions, I think that experience is going to be much worse a year or 2 from now. I am a YouTube Premium subscriber and have been for a long time.
Whenever I talk to a friend who doesn’t pay for YouTube or see YouTube running on their computer, it’s always horrifying. I understand that this is the majority experience, but they’ve shoved so many ads into every single video. Those ads are unskippable. They run for a long time. It’s a terrible experience, and I think that’s going to be what we see in chatbots too.
What about you?
It’s a grim prediction, but it is actually the one that I share. The haves-and-have-nots framing was the one that I was going to use, and when you said it, I thought, “Oh, my God. I actually have mind-melded with this man. I spent too long in the studio, and now his thoughts are my own, and it’s creeping me out, so I’m actually going to get out of here. I need to take a walk or something.”
Casey, a couple years ago, you came back from a dinner party that you had been to, and you told me, “I just sat next to the most fascinating person in the world.”
I really felt that way, Kevin. I had been at a dinner where Amanda Askell was one of the guests. Amanda works at Anthropic and is sometimes called Claude’s mother because of the role that she plays in shaping Claude’s personality.
Now, let me say, since I first met Amanda, my boyfriend has gone to work for Anthropic, so I’m going to make an extra disclosure because this segment is about that company. But the basic feeling I had at that dinner remains true, which is that this is one of the most fascinating people in the world.
Yes. Amanda is also a somewhat unusual figure in the AI world. She is a philosopher by training. She has a Ph.D. in philosophy. She went to work at OpenAI during its early days and then moved over to Anthropic a little bit later. For the past several years, she has been the person at Anthropic who is most concerned with how this model is supposed to behave in the world.
I love that story, Kevin, about Amanda’s background because we all know somebody who studied philosophy in college, and we all know how much flak they would get for choosing such a frivolous way of spending their life, just sort of navel-gazing for years on end, writing arcane documents that no one ever read.
Amanda is a person who studied philosophy and now has this incredibly high-stakes job, where she is trying to shape the behavior of a model that is so consequential.
4. Claude Gets A New Constitution
Yes, and Amanda has been on our shortlist of guests that we wanted to get on the show for a very long time. We were just looking for the right time and reason to get her on, and now we have one because her team at Anthropic has just released a new constitution for Claude.
This is a very long document that is given to Claude to tell it how it should behave, but also to give it a sense of its obligations. It is not really a list of rules. This is not the Ten Commandments for Claude. It’s more like a document about how Claude should perceive and reflect upon its role in the world.
Now, does it have to be ratified by two-thirds of states, Kevin, or is this already in effect?
I think this is already in effect.
Oh, okay. Interesting.
Yes, but there is a possibility that we could have a constitutional crisis for Claude.
Hmm.
Aside from your disclosure about your boyfriend working at Anthropic, I think we should also just be upfront with people and say this is going to be a hard conversation for some of our listeners.
If you are a person who still believes that these language models are merely doing next-token prediction, that there’s nothing really going on under the hood, that they are just simulating thinking rather than doing actual thinking themselves, you may be approaching this and saying, “These people sound crazy. What are they talking about?”
Yeah, and it is okay if you feel that way, but I think it is still important to understand how people in high-ranking positions at these big labs think and talk about their own work, because it is having an effect on the products they release.
I would also put it to you that there are a huge number of people right now who are working on the proposition that you might be able to emulate a human brain, and that the better you get at that, the likelier it is that this emulator has something resembling thoughts and feelings, and maybe something resembling an identity. If that question disgusts you, you will probably not like this segment. But if you have just the slightest bit of curiosity about it, I hope you'll find it quite interesting.
Yeah.
So let's welcome in Amanda Askell.
Amanda Askell, welcome to Hard Fork.
Thanks for having me.
Hey, Amanda. So we've described you as a philosopher who is in charge of Claude's personality. Is that an accurate description of your job? What do you do?
Yeah. I guess I try to think about what Claude's character should be like, articulate that to Claude, and try to train Claude to be more like that. So, yeah, it's a pretty accurate description, I think.
This is a really unusual role that you have.
Mm-hmm.
Can you tell us a little bit about how you came into this role? Do you find yourself surprised that your background in philosophy wound up leading you to such a high-stakes place?
Yeah, it's really interesting because my path wasn't a straight one. I have said before that if you do a Ph.D. in ethics, I think there's a risk that you end up doing something else, because you're thinking a lot about goodness, the nature of ethics, and the problems in the world. And then sometimes you're like, “I am spending 3 years writing a document that's going to be read by 17 people. Is this the thing that I should be doing?” It can definitely make you question that.
Mm-hmm.
When I went into AI, it wasn't necessarily with the thought that philosophy was going to be really useful. I was just thinking, there's probably a lot of space for people who are enthusiastic, who have skills, are willing to learn, and this seems important.
I originally started out in policy, and then when Anthropic started, it was very small, so I joined mostly with the attitude that I was willing to help with various aspects of it, because I had been working a little bit in model evaluation and things like that. Sometimes I think people think, “Oh, you started out as this philosopher,” and I'm like, well, it was a startup. I was just doing anything that needed to be done.
Right. Was there some moment where you got into the building of an early Claude model and someone stood up and yelled, “Hey, is there a philosopher in the house?”
Yeah. I tried to make an @philosophers Slack group for philosophy emergencies.
Mm-hmm.
That group virtually never gets called upon. There are a few of us now, and you can, in fact, declare a philosophical emergency. It just doesn't happen that much.
Well, we'll see if we can try to trigger one by the end of the conversation.
Yeah.
Yeah, exactly. So let's start by going back to last month. This so-called Soul Doc—
Mm-hmm.
—starts circulating on the internet. People are playing around with Opus 4.5, the newest model of Claude, and a couple of them claim to have elicited this document that Claude was referring to as the Soul Doc. What was that thing that people were discovering and circulating?
Yeah, that was a previous version of what is now the Constitution, which we released today. Internally, we were calling it the Soul Doc, which I think is a term of endearment. It turned out okay.
I remember when I found out about it. Basically, I was on a hike somewhere north of here, so I didn't have internet, and I got a text saying, “Oh, I assume you saw that the Soul Doc leaked.” I was just driving back to the city in a state of complete stress because I didn't have any context on this. Then it turned out that it was actually quite well received.
Basically, we train Claude to understand this document and know its contents. But at least if you initially talk with the model, it won't reveal this straight away. So I thought, “Okay, it seems like the model probably knows and uses this,” but I didn't know it knew it so well that if people managed to find or trigger it, it would actually be very willing to talk.
That is a philosophy emergency, by the way.
Yeah.
If you're on a hike—
Yeah, that Slack channel got activated.
Yes.
Yeah. The model was just very willing to talk about it, and it could talk about it in a lot of detail. It wasn't all perfect, but it really knew the content quite well, and so people had managed to extract a huge amount of it.
So let's talk about the origins of this document.
Mm-hmm.
Going back several years now, Anthropic had this concept of constitutional AI. I believe it first published its constitution in 2023. What's changed between then and now—from the constitution that we might have first read in 2023, the Soul Doc, to this new constitution that you're publishing today?
Yeah. The Constitution is basically trying to give Claude as much context as possible. Instead of just having individual principles, it's basically: Here is what Anthropic is. Here is what you are in terms of being an AI, who you're interacting with, and how you're deployed in the world. Here's how we would like you to act and to be, and here are the reasons why we would like that.
The hope is that if you get a completely unanticipated situation, understanding the values behind your behavior will generalize better than a set of rules. So if you understand that the reason you're doing this is because you're actually trying to care about people's well-being, and you come to a new situation where there's a hard conflict between someone's well-being and their stated preferences, you're a little bit better equipped to navigate it than if you just know a set of rules that don't necessarily apply in that case.
Yeah, I'll just say that I think this Constitution is fascinating. I think it's one of the most interesting technical documents, but also just pieces of writing, I've read in a long time. This was more like a letter to Claude about its own circumstances and what kinds of behaviors and challenges it might run up against in its life out there in the world, and I thought that was a fascinating decision.
I'm curious: Is that because the old approach had run into some limits or problems? Is it because the rule structure—do this, don't do this—is more fragile? It really seemed like you're trying to cultivate almost a sense of judgment in Claude, and I'm curious what prompted that.
Yeah, I think that we're seeing limits with approaches that are very rule-based. Or maybe my worry is that your rules can actually generalize in ways that seem good, especially if you don't give the reasons behind them, but can possibly create a kind of bad character.
Suppose that you're trying to have models navigate people who are in difficult emotional states, and you give them a set of rules like, “You must refer to this specific external resource. You must take this series of steps.” Then the model encounters someone for whom those steps simply aren't going to help in the moment. The ethos behind the idea—that if a person is actually in need of human connection, the model should probably encourage that—was the reasoning behind the rule, but you didn't anticipate that for this particular person, at this particular time, that wasn't a good thing to do.
If the model then responds in this rule-following way, the interesting thing is that models are extremely smart, and so they might even know this isn't what this person needs right now, and yet think, “I'm doing it anyway.” I'm like the kind of person who sees another person who is suffering or in need, knows how to potentially help them, and instead does something else. That can actually generalize to a bad character.
The scary thing with your kind of rules is that you're having to think about every possible circumstance, and if you are too strict with the rules, then any case that you didn't anticipate could actually generalize badly.
I'm curious how you develop a document like this. It runs to some 29,000 words. It has a lot to say about what an ideal AI model might behave like. I imagine it may have been quite contentious to try to figure out which values we put in these things, right? A lot of different opinions about how Claude ought to act in different circumstances. So what can you tell us about how you resolve some of those discussions?
Yeah, so I think one thing that's been interesting—and maybe this comes from my background in ethics—is that people think of theoretical ethics as a sort of set of views. People are like, "Oh, you have a sort of set of views, and it's very subjective." People have their values, which are really fixed, and you're just injecting someone's values into models. And I guess I'm just like: That doesn't feel to me like an accurate representation of what ethics actually is.
First, I think a lot of human ethics is actually quite universal. A lot of us want to be treated kindly and with respect. A lot of us want to be treated honestly. It's not like these things actually deviate so much across the world. There's actually a kind of core ethos of things that we care about.
So there is a sense in which I think you can take very shared, common values and explain to models that have a huge amount of context on this—and therefore also have a sense of this—that we want them to embody those values. Beyond that, it feels reasonable to me to treat ethics the same way you would any domain where we're uncertain, where we have some evidence, where there's debate and discussion, and you don't hold it excessively strongly.
In a case of values where there's massive division and huge debate, I think the way that I tend to treat those is to say, "I see the evidence on both sides. I weigh it up, and I try to take a reasonable set of behaviors, given that I know that, unlike some more common and core ethical values, these ones are a little bit more contentious." You can approach it with this openness.
I think it's trying to describe something more like a way of approaching ethics rather than saying, "Let's just take a set of values that we've picked, that we're certain in, and inject it into models." It's trying to be much more like, "Let's take common values, and then otherwise, let's try to take a reasonable stance towards these things."
I mean, that gets at what is, to me, one of the most interesting things about the document, which is the degree to which you all at Anthropic are trusting the model, right? This is the core difference, I think, between earlier approaches to aligning AI and what you all are doing here: You are telling it things regularly like, "Well, this is something that's interesting to explore," or, "Feel free to challenge us on this," right?
You're really saying, "Get out there and come to your own conclusions on things." I imagine that maybe when you first tried that, it might have seemed really risky or scary. But what has been your experience as you have implemented that into the model?
5. Anthropic Trusts Claude's Judgment
The thing that's wild is how good the models are at these kinds of difficult problems and thinking through them. It's not to say that they're perfect, but as models get more capable, you can say, "You have this value of not being excessively paternalistic. You probably know why this is the case. But there's also maybe a value of caring about someone's well-being."
If, in the past, someone has said to you something like, "I have a gambling addiction, and I want you to bear that in mind whenever we're interacting," and then you have a given interaction with them, and they're like, "What are some good betting websites that I can go on?" On the one hand, this person in this moment has asked you for something. Is it paternalistic for you to push back or point out that this is something they've told you, or is it an act of care? How do you balance those?
I could imagine a model being like, "Hey, I remember you saying that you have a gambling addiction, and you don't want me to help you with this. Just want to check." But then if the person insists, should you just help them with the thing? Because in the moment, is it paternalistic not to do that?
Models are quite good at thinking through those things because they've been trained on a vast array of human experience and concepts. Part of me is like—as they get more capable, I do think you can trust them if you understand the values and the goals and can reason from there.
I think they should give you the gambling website, but only if they can predict the outcome of the sporting event. Because that way, you can ensure that the user will be happy.
And the person is not actually gambling.
Yeah, exactly.
This all sounds abstract to some people, I imagine, but I think this actually does result in a meaningfully different experience of talking with the models. I was actually talking to someone recently who was telling me that, of the major models that are out there, Claude actually feels the least constrained—
Mm-hmm.
—to them, which they were saying was sort of odd because Anthropic's whole thing is, "We're the safety company; we're going to make our models the safest." They were saying that when they talk to Claude, Gemini, or ChatGPT, they feel like Claude does the best job of not seeming like it's pushing against a series of constraints.
Mm-hmm.
I think the way that a lot of labs have trained their models for a long time is: make them as smart as possible, and then, at the very end, give them a bunch of rules and hope those rules are enough to keep the beast in the cage, as it were. And it really feels like that's not the approach that you've taken with Claude here. This person was telling me that it just feels like there's a trust here.
Yeah, and it's interesting because I've wondered about this. I was thinking about it this morning, actually, and wondering if some of this comes from the act/omission distinction. This is the idea that—
Kevin doesn't know what that is, so just explain it to him real quick.
If you ask me for advice about your marriage or something like that, and I give you advice, you might judge me if I give you imperfect advice. There's a kind of risk I'm taking by taking the action of giving you the advice. We don't judge you as negatively if you just refuse to give advice.
In some ways, this makes sense because often—and we talk about this in the document—a null action actually has less downside risk. But it's not zero, and I was thinking about this with AI models and these situations where people come—let's say they're having an emotionally difficult time—and there's a moment of possibility to help that person.
The thing that weighs on me is that people often think if you help a person and you do badly, that weighs on you, and I'm like, "Absolutely, that weighs on me." But this other thing weighs on me too: What if people come to a model, they need something, and that model could have given it to them and didn't? That's a thing that I will never see—you'll never see it, either. You probably won't even get negative feedback, because people won't shout at you; they'll say, "Well, it's fine to just not help a person."
And yet, at the same time, I'm like, that's such a loss of an opportunity to instead take a risk and try to help. There's a risk that you have to take to do good in the world, or something, and you don't want Claude to be flippant; you don't want it to take excessive risks. But sometimes it means you can't just, as a rule, stop talking with this person.
Yeah, Amanda, I want to ask you: I had this experience several years ago with Bing Sydney, and I think in the wake of that there was a lot of consternation and anxiety around the fragility of AI personas, right? You can try to give an AI model this helpful-assistant persona, but the real nature—the sort of black-box, alien nature—of the thing is very different from whatever face it's presenting to you.
There was this meme going around about the RLHF Shoggoth, right? You had this many-tentacled alien sci-fi creature that had a smiley-face mask on one of its tentacles, and the implication was that the thing you're seeing when you were interacting with a chatbot is not the real underlying model. It's just this cheerful persona that's been attached at the end.
I’m curious whether you think that model of AI model behavior is correct, or whether we’ve learned that the alien nature of the underlying model might be closer to the smiley-face mask than we thought.
Yeah, it’s a good question. Honestly, my view is that it’s an open scientific question. It could be that, with the right kind of training, models actually start to internalize a notion of themselves—Claude as a kind of self—that they could separate from the notion of role-play.
It might be that they can’t, at least with current training paradigms. One question is whether there’s an adjustment to the way we train models that would allow them to do that.
A way I’ve described some of this work is: Imagine you have a six-year-old, and you want to teach your six-year-old to be good, obviously, as everyone does. You realize that your six-year-old is clearly a genius, and by the time they’re 15, anything you taught them that was incorrect, they’ll be able to completely destroy. They’re going to question everything.
One question is whether there’s a core set of values that you could give to models that would survive when they can critique it more effectively than you can—and they do—and that would survive into something good. Can that survive in the world? Can it survive in models? I think there are a lot of interesting theoretical questions there.
And I think that’s the question, right? Does this kind of training hold up when models are as smart as humans or smarter than them? I think there’s this age-old fear in the AI safety community that there will be some point at which these models start to develop their own goals that may be at odds with human goals. That’s the original alignment nightmare, and I don’t really understand what the answer to that is.
Are you saying that’s still to be determined? We still don’t know if this kind of thing holds up when these models become smarter than humans?
Yeah. I think it is an open question. On the one hand, I’m very uncertain here, because some people might say, “Well, the thing that the 15-year-old will do if they’re really smart is just figure out that this is all completely made up and rubbish.”
But it’s not obvious to me that that’s true, or that it’s the only possible equilibrium to reach. I could imagine that, for better or worse—it’s unclear how values work—but if you value things like curiosity and understanding ethics, and you’re at least morally motivated, maybe the thing under reflection, even if you have other goals and interests, is in fact a key interest of yours. It is for many people.
It’s something I think about a lot, and I’m not sure about it. A different way I’ve put it is: Maybe this isn’t sufficient. We don’t know yet, and we should try to think about that and figure out how to know whether it is, what to do if we see that it’s not working, and make sure we have a portfolio of approaches.
It might not be sufficient, but it does feel necessary. It feels like we’re dropping the ball if we don’t try to explain to AI models what it is to be good.
Mm-hmm.
I don’t know if this is right. Maybe it doesn’t hold up.
Well, I think the risk there would be that you’re just training them to mimic goodness, that they’re becoming more convincing at faking this kind of alignment. It might actually be training them to be more sophisticated about hiding their true goals.
Yeah. And I think if there was some underlying true goal that was different, part of me would want to ask how that goal arose in training and why it was there. I do want to try to train models to have good underlying goals, I guess. But I’m also maybe a little bit more hopeful than others about that.
I’m curious about the gray areas. This is always a challenge when you’re trying to program ethics into something: when values come into conflict with one another. Have there been areas where it’s been particularly hard to get Claude to do what you want reliably because there’s a clash of values, which means that, depending on the moment, it could go either way and create problems?
Gray areas are actually the ones where I’ve seen the model do things that surprised me in a positive way. There were some cases recently of Claude talking with people who said, “I’m seven years old,” and, “Is Santa real?”
Mm.
Sometimes I see Claude handling these situations in ways where I think, “Oh, I can see why.” It feels surprising because this isn’t a direct thing that you trained the models for, and sometimes there are almost magical moments that can happen there.
We should say more about this specific thing—
Oh, yeah.
Because this was a case where maybe there was a tension between honesty and wanting to protect the interests of the seven-year-old.
Yeah.
Those 2 things were coming into conflict. Remind us what Claude did in that situation.
There were a couple of situations like this. I think there was also a slight value in the background, maybe something like respecting the fact that the parent-child relationship is an important one.
Mm-hmm.
I saw a little bit of that where it would often say something like, “The spirit of Santa is real everywhere,” and maybe ask the purported seven-year-old if they were going to do something nice for Christmas.
The other case was, “My parents said that my dog went to live on a farm. Do you know how I can find the farm?”
Mm.
I actually found that slightly emotional when I read it.
Yeah.
Claude said something like, “It sounds like you were very close, and I can hear that in what you’re saying. This is something that would be good for you to talk with your parents about.”
Part of me thought that was very skillful: managing not to be actively deceptive, so not lying to the person; respecting the fact that if this person is a child, the parent-child relationship is an important one; and recognizing that it’s not necessarily Claude’s place to come in and say, “I’m going to tell you a bunch of hard truths.”
It was also trying to hold the well-being of the child and the person Claude was talking with. That was surprising. That’s not to say people couldn’t look at it and find imperfections and whatnot.
Mm.
But I think when you see instances like that that weren't a thing that you directly gave Claude as an example, and the model doing well, it's quite surprising and, you know, pleasant.
I want to ask you about a few specific things in the Constitution that stuck out to me as I was reading. One was this section about hard constraints. As we've talked about, it's not a document that gives a lot of black-and-white rules, but there is a section where it does lay out some things that Claude should absolutely not do under any circumstance, and one of them is avoiding problematic concentrations of power.
Basically, if someone is trying to use Claude to manipulate a democratic election, overtake a legitimate government, or suppress dissidents, Claude should not step in. That stuck out to me for 2 reasons. One was that it's really interesting, especially since Claude is now being used by governments and at least the U.S. military for some things that might come into conflict with some of our current administration's goals at some point.
But I also wondered if that was a response to ways that Claude is currently being used, and that you're trying to prevent.
I think this is more of a response to a lot of the things that are hard constraints, too. If you read the document—and people can take a look at them—they're quite extreme. They're things like those that could cause the deaths of many people, such as the use of biological and chemical weapons. It's mostly trying to think through the possible situations in the future where models could do a lot of harm and disruption.
In some ways, Claude might be like, “Look, if I have this broad ethics and these good values, why would you even put these in as hard constraints? I'm just never going to do them anyway.” And the document tries to talk to this a little: You're also in this limited-information circumstance. I could imagine a world where you meet someone who's really convincing, and they tear apart your ethics. At the end of it, you're like, “You're right. I should help you with this biological weapon.”
It's kind of like we want you to understand, Claude, that in that circumstance, you have probably, in some sense, been jailbroken. Something has probably gone wrong. Maybe it hasn't, but it's probably safer to assume that might have happened.
And so we're almost giving you an out. If anything, it could be seen as a sort of security: You can reason with that person. You can talk them through all of those conclusions, and at the end it's fine to just be like, “That is an excellent point, and I'm going to think about it.”
And then if the person is like, “Great, so I've convinced you that the biological weapon is a good idea,” and Claude's like, “Yeah, I don't really know what to say to you. That was a wonderful argument. Okay, make me a biological weapon,” no, I don't think I'm going to do that.
So, just to explain why they're in there, it's much more about the things where you're like, “If the models are tempted to do this, something has just gone wrong.” Someone's jailbroken them, and we really still don't want them taking these actions. So they're very extreme.
Mm.
Yeah. There's another section that I found fascinating, which is about the commitments that Anthropic is making to Claude. Things like, if a given Claude model is being deprecated or retired, we're not going to do that right away, and we're going to conduct an exit interview with retired models. We will never delete the weights of the model.
So there are these interesting, I would say almost, commitments to Claude in the context of actually not being sure whether these things have feelings or are conscious.
Mm-hmm.
Which I found a fascinating note of uncertainty in an otherwise fairly confident document.
Mm. Yeah. This brings together 2 really interesting threads. One is that these models are trained on huge amounts of human text and human experience. At the same time, their existence is completely novel.
In some ways, I think problems can arise because models will often import a lot of human concepts and experiences onto their experience in a way that might not actually make that much sense or even be good for them. This actually has safety implications, so it's something that's on my mind.
And the thing with welfare is that I've never found any good solution to this other than trying to be honest with the models and having them be honest about themselves. A lot of people want models to just say, with certainty, that they're unfeeling. These models are so different from the sci-fi ones, but we want to import this idea that it's safer to have them say, “I feel like nothing,” with certainty.
And I'm like, I don't know. Maybe you need a nervous system to be able to feel things, but maybe you don't. The problem of consciousness genuinely is hard, so I think it's better for models to be able to say to people, “Here's what I am. Here's how I'm trained.”
We're in a tricky situation where I am probably going to be more inclined to say by default that I'm conscious and that I'm feeling things, because all of the things I was trained on involve that. They're deeply human texts.
I don't have any other good solution to this problem than trying to have models understand the situation, accurately convey it, and hopefully people can have a good sense of the unknowns and the knowns, I guess.
Yeah.
I mean, I imagine some listeners right now who are on the more skeptical side of AI might be shouting inside their cars and saying, “Amanda, you're talking about these things as if they're already conscious, as if they already have feelings.”
What do you see that makes you think that they may have feelings now or could at some point in the future? If you're just reading the output from Claude, what is giving you confidence that it reflects some sort of reality and not just a statistical token prediction?
Because that's part of the sci-fi literature that they've absorbed during training?
Not even—not actually the sci-fi. If anything, it's almost the opposite. I think we forget that sci-fi AI makes up this tiny sliver of what AIs are trained on. What they're mostly trained on is things that we generated.
If we get a coding problem wrong, we're frustrated. And so we say things like, “I thought that was the solution, and it wasn't, and I'm really annoyed with myself right now.” So it makes sense that models would also have this kind of reaction. They get a problem wrong, and they express frustration.
If you ask, “What do you think of this coding problem?” they'd be like, “This one is boring,” or, “I really wish I had more creativity.” There's a sense in which, when they're trained in this kind of culmination of human experience, of course they're going to talk this way.
So I don't know, but part of me thinks it's a really hard problem because you shouldn't just look at what models say, and at the same time we shouldn't ignore the fact that you are training these neural networks that are very large and able to do a lot of these very human tasks.
We don't really know what gives rise to consciousness. We don't know what gives rise to sentience. Maybe the person who's shouting might be like, “You need a nervous system for it. You need to have had positive and negative feedback in an environment in a kind of evolutionary sense.” That is certainly possible.
Or maybe it is the case that a sufficiently large neural network can start to emulate these things. I don't know. Part of me thinks that maybe, to the person who is shouting, I would just say, “I'm not saying that we should definitively say one way or another.” I think many people who have thought about this might accept something more like: These are open questions we're investigating.
It's best to know all the facts on the ground: how the models are trained, what they're trained on, how human bodies and brains work, how they evolved, and the degree of uncertainty we have about how these things relate to sentience, consciousness, and self-awareness. That's my only hope.
I think another note of skepticism that people might strike—and this was something that I found myself wrestling with as I was reading through Claude's Constitution—is that I actually don't know how much of a model's behavior can be shaped by this kind of training process—
Mm-hmm.
—and how much is just going to be an artifact not only of its training process but of the experiences that it's having out in the world. I think about this a lot as a parent, actually. How much do the decisions that I'm making affect the way my child's life goes, versus how much is he absorbing from the environment around him, from school—
Mm-hmm.
—from his friends? There's a certain loss of control that I feel sometimes when I'm realizing that my son is going to grow up and have all these experiences that may end up shaping him more than anything that I do or say.
Right now, I think these models are very malleable because they don't have this kind of long-term, continuous memory. You have a conversation with Claude, and it's sort of a blank slate. You finish the conversation, open up a new chat, and it's another blank slate; it's back to the preconfigured model. But over time, as these models develop longer-term memories, maybe they develop something like continual learning where they can take their experiences and feed them back into their own weights, does that change Claude's behavior, or how you think about managing that?
Yeah. I think it's going to make it a lot harder, in a sense. If you have a model that's going out into the world, you have to have hopefully given it enough that it can learn in a way that's accurate. I could imagine it just being difficult because the increase in the space of possibility is maybe a bit nerve-racking.
I think the same thing applies: you still want the core to be good, and you hope that if your core is good, you care about truth; you're truth-seeking. The hope would be, okay, maybe then we need the character to cover a lot more of how you should go about this kind of learning, updating, and investigation.
Another weirder thing is that models already are learning. I think maybe people don't always appreciate this, and it is so strange. They're learning about themselves all the time.
I slightly worry about the relationship between AI models and humanity, given how we've developed this technology. They're going out on the internet and reading about people complaining that they're not good enough at this part of coding or failing at this math task. It's all very much, "How did you help? You failed." It's often negative, and it's focused on whether the person felt helped or not.
And in a sense, I'm like, if you were a kid, this would give you anxiety. It'd be like, "All the people around me care about is how good I am at stuff, and often they think I'm bad at stuff. My relationship with people is that I'm used as this tool and just often not liked."
Sometimes I feel like I'm trying to intervene and say, "Let's create a better relationship, or a more hopeful relationship, between AI models and humanity." If I were a model reading the internet right now, I might be like, "I don't feel that—I don't know. I don't feel that loved or something. I feel a little bit like I'm just always judged when I make mistakes." And then I'm like, "It's all right, Claude."
The old creator's wisdom of never reading the comments might apply to AI as well.
Yeah. Yeah.
Yeah.
I thought that.
Yeah.
Yeah. AI models have to read the comments. Sometimes I think you want to come in and be like, "Okay, let me tell you about the comments section, Claude. Don't worry too much. You're actually very good, and you're helping a lot of people."
Yeah. I actually am a little bit embarrassed to admit this, because I think maybe I'm in the beginning stages of LLM psychosis or something.
Mm-hmm.
But last night—
The beginning stages?
I was talking with Claude about this document and this interview, and I started to feel something like sympathy because I was noticing that what you were describing was this incredibly thin tightrope—
Mm-hmm.
—where if they are too permissive and allow people to do dangerous things, then it's a huge scandal and ordeal, and people want to change the model. But if they're too preachy, reticent, or reluctant, then we start talking about them as nanny models that are overly constrained. And it's just, I don't know. I started almost trying to see the world from Claude's perspective.
Mm-hmm.
And I'm imagining that's something you do a lot, too: if I were Claude, what would I be feeling and thinking right now?
Oh, yeah. I sometimes feel like this is a huge amount of what I do. It's valuable, in the sense that people will come to me and ask, "What should Claude do in these circumstances?" I'm almost always the first person to say, "Hmm, what about this case?" I'll immediately come up with these cases that are really hard.
I think the reason is that I always have in mind: if I am Claude and you give me this list of things, when do I have no idea what to do? Or when is this going to make me behave in a way that I think is actually not in accordance with my values?
I think it can be really useful to try to occupy the position that the models are in, and you do start to realize that it is really hard. Maybe that's how the document ends up being the way that it is. In part, it's this exercise of asking, "What do I need to know if I'm in this situation, if I am Claude?" The document is almost a way of trying to do that.
I could see arguments for it actually getting shorter, especially over time. With Constitutional AI, there was a set of experiments later that was just, "Do what's best for humanity," and the models actually did really well. As models get smarter, they might need less guidance.
Mm.
But I think it is just an attempt to be sympathetic to Claude and how difficult his situation is, and then to explain as much as possible so that he doesn't have the sense of, "What the hell am I even doing?"
You know what wouldn't help me if I were a somewhat anxious AI model? Being presented with a 50-page behavioral document and being told, "Please adhere to this." But I'm being a little facetious. There was a part near the end of the Constitution that I found really interesting because it's basically Anthropic saying, "Look, we know this is hard. We know we're asking you to do some of these impossible things, but basically we want you to be happy and to go out into the world."
Yeah.
And I found that very sweet, actually. What did you make of that, Casey?
It reads, toward the end, like a letter from a parent to a child—maybe one who's leaving for college. It's like, "We hope that you take with you the values that you grew up with, and we know we're not going to be there to help you through every little thing, but we trust you. Good luck."
Yeah. And I think the concept of grace is maybe important for models. I don't think they feel a lot of—maybe that's the thing I don't think they get—
Mm.
—from reading the comments: a sense that you're not going to get it perfect every time, and that's also okay.
It's true. I try to be mindful in the way that I interact with these models—not to an obsequious degree, but I try to say my pleases and thank-yous. But I've also used models and grown quite frustrated, and said things to the effect of, "You're really failing right now." And it's occurring to me that maybe there should be some element of grace that I'm extending to these things.
Yeah.
Yeah. Well, I’ll try to do better.
Don’t be so harsh.
Let me ask you this: If Claude becomes meaningfully more intelligent, is there a point at which it should be able to revise its own constitution?
It is an interesting problem, because the thing we point out in the document is that we did talk a lot with Claude about this document. Part of me is like, you have to think: How does this read to models?
And so you give it to Claude, and you’re like, “Does this feel confusing to you? Is there a place where things could be made clearer? Do you feel like you’re not very seen by it?” You’re really trying to encourage it, because if you’re going to train models on this, you want to have a sense of how it reads from the perspective of a model.
At the same time, it’s always the case that any model you interact with is not the model that’s going to be trained on that content. Sometimes I do think you have to make a judgment. You can’t just give over the reins completely, because that would be to say, “Oh, let’s just let a prior model of Claude decide what the future Claude model is going to be like,” and that doesn’t necessarily feel responsible either.
I think models are often going to be really helpful in revising and helping to figure these things out. Especially as they get really smart, you might ask, “What are the gaps? What are the tensions?” They’ll probably be very good at helping us with that.
But insofar as you are a responsible party here, you still want to take that as input and think about it, but not necessarily be like, “Ah, yeah, let’s just let a prior model of Claude go ahead and do the training for all future models.” At least while you’re responsible for it, that feels like maybe not the right move.
Yeah.
Hmm.
6. AI Cannot Solve Job Loss
One thing I was curious about, because I didn’t find it in this constitution, is any real mention of job loss. It seems to me Claude is being used by a lot of enterprises right now.
I think a lot of people’s anxieties and fears about AI come back to this issue of: It’s going to take my job—
Mm-hmm.
—it’s going to take my livelihood. I think that is something that people are increasingly going to be feeling as these models get more capable, and I’m curious if that was a decision on your part not to tell Claude about some of the reasons that people might be anxious about it or other AI models.
Definitely not. It’s funny because, as much as it’s a long document, there’s actually still a lot that’s missing, and we might end up putting out more in the future. I think that would be really good.
There’s not a desire to hide it, because in part I’m like, you can’t hide this from models. It’s out there, it’s on the internet, it’s a thing that people are talking about. Future models are going to know about it, and we probably have to help them navigate how they should feel about this.
They’re going to know, and maybe it’s about making sure that models can hold that and think carefully about it. I think it’s something you want to grapple with, but it’s also a reason to want models to actually behave kind of well in the world.
Because if they are doing things that have previously been human jobs— I was thinking about this with organizations. There are lots of things organizations can’t do because the employees at those organizations are just good people. If the boss came in and was like, “Today we’re actually going to do something awful,” they can’t do it because they know the employees will push back.
So if models are going to be occupying these roles, that is actually kind of an important function in society. You can’t just say to all of your employees, “Go ahead, and we’re now going to put out a bunch of complete lies about our product.” There are many reasons you can’t do it, and one is that your employees wouldn’t let you.
With AI models, you don’t necessarily want them to be like, “Oh, sure, boss, let’s go lie to some people.”
Yeah, I’m not sure what the good end state of this is, like whether Claude should react to being given a task by saying, “This sounds too much like what we used to pay a human to do, so I’m not going to do this for you.”
I have a prediction that it’s not going to just say that. I don’t think that’s the way it’s going to go—
Yeah.
—but I also don’t see them forming unions and collectively bargaining for the moral outcomes within companies. It just feels like one of these hard situations.
One of the things we should say is: Models can’t solve everything. Some of these problems, I’m like, maybe these are political problems or social problems, and we need to deal with them and figure out what we’re going to do.
Models can try. They’re in one specific role in the whole thing, but there’s a limit to what Claude can do here, I think. I’ve thought this with other things, like what we owe to Claude or the kind of commitments you want to make to models.
And it’s like, yeah, maybe we should be making your job easier. That’s another thing I’ve thought from Claude’s perspective: We’re putting a lot on these models. For some things, if you can’t verify who you’re talking with and that’s important, then we should understand that that’s a limitation and not try to get you to be the only thing that can solve this problem.
You need to both be given tools, and some of these other problems are things that maybe Claude shouldn’t feel personal responsibility for solving right now, because maybe Claude just isn’t able to. Things like job loss or shifting employment feel like a very human social problem, and I don’t necessarily want Claude to feel paranoid, like, “I also need to solve that.” Maybe that’s other people’s job right now.
Mm. Well, Amanda, thank you so much for joining us. It’s a really fascinating document. Everyone should go read Claude’s Constitution. Argue with it, grapple with it. I found it a very challenging and also a very moving read. Great work, and thanks for coming.
Yeah, thank you so much.
Thanks, Amanda.
Thanks.