[BidClub_]
The Cognitive Revolution · · 137 分钟

Nathan 赴华——第2部分:中国特色的 AI 安全

Nathan Labenz

YouTube
TL;DR
  • 中国目前已部署 AI 的安全防护落后于美国,但表面上的差距很大程度上是 OpenAI-Anthropic 效应,而非文明层面的鸿沟。 Nathan 估计两国各有约 10-12 家接近前沿的开发商;如果剔除“Openthropic”,美国剩余公司的水平与中国同行重叠度会高得多,Gemini 也只是小幅领先。Concordia AI 的评估得出了相同结论:以美国专有模型为主的模型大多位于或高于“45度线”,以中国开放权重模型为主的模型则往往位于其下方。

  • “中国不在乎安全”“中国永远不会放慢速度”的说法,同时被政策和先例所否定。 据报道,北京在 2023 年曾因制定标准,暂缓了许多国内聊天机器人约 6 个月的上线;如今,CAC 可以要求服务在进入备案清单前接受地方和全国层面的审查。中国还针对推荐算法、零工平台、儿童游戏和 AI 伴侣出台过高成本监管措施,这表明当国家认为有必要干预时,会把企业增长置于其后。

  • 中国 AI 安全正通过大学和企业快速扩张,走的不是美国那种无需许可的非营利生态路径。 Concordia 统计显示,相关论文数量已从 2023 年每月仅几篇,增至 2026 年年中每月约 50-60 篇;美国或英语圈的估计则是每月 50 篇到几百篇不等。研究覆盖自我复制、评估造假、欺骗、机制可解释性、多模态攻击和危险能力隔离等方向;Nathan 的结论是:“AI 安全已经在中国扎根。”

  • 中国围绕 AI 安全的核心风险讨论,已经从审查转向能够在数字系统、最终甚至物理系统中采取行动的智能体。 一家大型科技公司坚持称自己“确实在乎 CBRN 风险等灾难性风险”,并表示有一个智能体每天生成美国 AI 安全讨论的报告。Xi Jinping 在 WIC 的讲话同样呼吁加快防范“失控”、防止恶意使用,并让 AI“处于人类控制之下”——在 Nathan 能找到的知名美国官员表述中,没有比这更强调安全的说法。

  • 最大的未决分歧在开放权重:中国监管服务,并相信自己能够把放出的精灵重新装回瓶子;西方安全分析则强调模型权重一旦全球释放,后果不可逆。 中国受访者认为,拥有 2.88 万亿参数的 K3 不可能直接在心理状态不稳定者的笔记本电脑上运行;严肃推理需要大量硬件和基础设施。这套逻辑在中国的监管边界内或许成立,但 Nathan 担心,一旦权重进入北京无法控制的司法辖区,它低估了境外生物和网络风险。

  • 随着中国能力追赶上来,监管差距可能收窄,因为企业似乎预期标准会与模型能力同步抬升。 Nathan 认为,能力差距最可信的中枢约为 9 个月,并指出 Claude 4.5 Opus 是智能体真正开始有效工作的分水岭。中国实验室如今也在发布智能体真正能够工作的模型。他的条件式预测是:当中国监管者遇到 OpenAI 和 Anthropic 目前正面对的那些失效案例时,会收紧要求,已部署的安全防护将“显著”趋同。

  • 中国在概念层面最大的缺口,或许是以品格为基础的对齐,而不是按规则服从。 中国 AI 告诉 Nathan,孔子的后裔据称历经 79 代仍自认是其后人,并举行仪式纪念孔子,这启发他追问:儒家宪法能否让价值观穿越递归生成的 AI 世代。然而研究者告诉他:“我们都是工程师”;当前生态强调明确规则和可靠服从,以智慧传统为基础的对齐领域“基本还是一片处女地”。

摘要 · 为研究而整理的核心内容

1. “但是中国”这一论点建立在错误前提上

  • Nathan 的目的刻意比预测一项条约更窄:以“中国自身的语境”描述中国 AI 安全,并打破“美国承担任何义务都会自动把比赛拱手让给北京”的惯性。中国在其他优先事项之外也重视安全;当其对安全的定义要求这么做时,它会让企业放慢速度,而且这个政府已经明确展示过行动意愿。

  • 他最有力的表述也是最直白的一句:因为中国据说无法监管,就“全速冲向递归自我改进”,这种说法“来自无知”。这并不意味着合作会变得容易,也不保证双方规则对称;但它排除了一个被想象出来的不可能性——太多西方政策讨论在尚未真正开始前,就被这个不可能性提前终结。

  • 报道遵循查塔姆规则,因此私人观察不标注出处,公开论文、报告和讲话则明确点名。Nathan 还披露,机票和酒店均由自己支付;在中国停留 2 周期间,他接受了大约 6 顿饭局邀请。

2. 美国在已部署安全防护上的领先,靠的是 2 家公司

  • Nathan 没有回避不利的比较:中国公司和模型目前在防止滥用方面提供的安全防护弱于美国产品,模型卡和公开安全评估的后续落实也更差。差异是实质性的,尤其是按照消费者实际使用的产品加权后。

  • 但美国平均水平被 OpenAI 和 Anthropic 主导——也就是他逐渐称为“Openthropic”的双寡头——2 家公司同时在能力和抗越狱方面领先。Gemini 略高于更广泛的行业群体,Grok 和其他美国模型则更容易被攻破;如果没有这 2 家,美国的道德高地就会“模糊得多”。

  • 按照较宽松的接近前沿定义,Nathan 统计两国各有约 10-12 家模型开发商,数量高于大多数观察者对真正前沿模型的认定。他最近关于越狱的讨论给出了清晰排序:OpenAI 和 Anthropic 最难越狱,Gemini 和 Grok 更容易,中国模型则更容易。

3. 45度线是中国追求中的安全准则

  • 一位 Shanghai AI Lab 负责人于 2024 年在 WIC 提出了“45度线”:能力和安全应同步上升,斜率为 1。安全必须跟上每一个新的能力水平,但在尚不存在的能力上构筑极端防御,可能浪费资源,甚至适得其反。

  • Concordia AI 在 aisafetychina.com 发布的公开评估,把这一比喻转化为能力与安全的综合对比图。以美国专有 API 模型为主的模型大致位于线附近或线上方;以中国开放权重模型为主的模型则聚集在更低位置,且往往低于这条线。Nathan 强调,具体综合指标仍值得审视,但方向性结果是一致的。

  • 这种坦率很重要:一家总部在中国的机构用中英文发布了这些对中国不利的结果,而不是把它们藏起来。Nathan 将此视为国内社群能够承认短板的证据;45度线已被接受为一项原则,但在执行层面“仍有些理想化”。

4. 中国更自然地监管 AI 服务,而不是模型权重

  • 中国消费应用往往是一个多组件系统:底层模型可能已经开始回答敏感问题,随后由独立监控器删除回答、替换成拒答。因此,评估裸权重,测量的是与评估 API 或自有服务不同的对象,而大多数中国用户正是通过后两者接触模型。

  • Nathan 承认西方最坏情形论证的合理性:权重一旦公开,任何拥有足够资源的人都能在原始部署防护之外使用它,因此必须测试模型本身。中国的思路则是追问谁会实际运行它、通过什么基础设施、在什么受监管的服务中运行——这是一个西方讨论可能过快忽视的运营问题。

  • 中国受访者反复举出的例子是拥有 2.88 万亿参数的 K3:“这不是你能在笔记本电脑上运行的东西。”一个随机的、处于心理困境中的人,不能随手把它下载到手机上;它需要大量硬件,而最可能的使用者是把模型封装进服务、因而可以被监管的企业。

  • 两种视角都不能替代另一种。服务层分析更能描述普通用户的暴露情形,裸权重分析则捕捉全球可用的尾部风险;Nathan 认为,这是两大安全社群之间最清晰的概念分歧之一。

5. 信息披露并不完整,老牌公司比初创企业更谨慎

  • Concordia 审查了 10 家中国大型公司,发现其中 5 家近期在发布模型时披露了某种形式的安全评估。另 5 家没有披露;即便是参与披露的公司,也没有对每一次发布都进行评估,使中国行业实践远未达到 OpenAI 或 Anthropic 式模型卡的一致性。

  • Nathan 的印象是,Alibaba、Ant、Tencent 以及可能还有 ByteDance 这种规模的老牌公司,比追逐 AGI 的年轻初创企业做更多安全工作。成熟企业有盈利业务、监管地位和需要保护的制度体系;例如,一个处理资金的金融科技平台会立刻有动力测试 AI 是否可能“失控”。

  • 初创企业的推理方式往往更像 Meta 在 Llama 2 和 Llama 3 上的做法:如果自己大约落后 1 年,而美国系统已经在没有报告灾难的情况下暴露了相应能力,那么追赶只会增加很少的边际危险。它们的逻辑是,如果追逐的能力水平真的会引发严重后果,那很可能早就发生了。

  • Nathan 的结论带有条件性:新近报告的前沿事件可能会打破“此前的能力水平已经被安全探索过”的假设。如果真实网络伤害或智能体伤害的证据不断累积,中国初创公司的宽松追赶逻辑“很可能”会发生变化。

6. 跨境交流带来了明显的思想趋同

  • 在 WIC 及周边 Track 2(二轨)论坛上,Nathan 看到那些公开呼吁中美合作不可或缺的美国倡议者,正在闭门场合做与其立场相符的工作。有些会议因为参与者对播客记者身份感到不适而没有邀请他,但公开证据显示,双方的思想交叉已经无可否认。

  • 中国研究者反复引用西方机构和概念,并没有掩饰其来源。有人愤世嫉俗地认为 AI 安全是西方用来拖慢中国的阴谋,正如西方讨论中也有自己的犬儒派;但 Nathan 遇到的中国受访者中,没有人亲自推进这一观点。

  • 一家中国大型科技公司召集了 CMO、传播负责人、总法律顾问和 AI 安全负责人,随后“几乎拍着桌子”强调自己在乎灾难性风险和 CBRN 风险。公司还运行一个智能体,每天扫描美国 AI 安全讨论并生成内部日报——Nathan 认为,反向进行同等情报收集的证据很少。

7. 如今安全议程的主角是智能体,而不是“3个 T”

  • Tibet、Taiwan 和 Tiananmen 仍是敏感内容领域,中国服务必须执行政府围绕这些议题制定的规则。但 Nathan 不接受“审查是中国机构唯一认可的安全问题”这一推论;许多公司如今认为,内容安全要求基本已经解决。

  • 各处真正热议的是“智能体、智能体、智能体”:系统已经从回答问题,转向自主采取行动。这让安全问题从管理言论,转变为控制软件、网络和商业工作流中的重大行为。

  • 中国对机器人技术格外重视,使这一担忧延伸到数字世界之外。研究者和企业预计,具备商业价值的具身智能体将进入物理环境,由此,一个现实问题无法回避:如果 AI 能够自主行动,机构如何确保这些行动持续有益并且“处于控制之下”?

8. 大学填补了中国从未发展起来的非营利生态

  • 美国无需许可即可运作的公民社会,让具有推测性的 AI 安全思想在其尚属边缘时得以存续:一个小团体只需说服 1 位富有的资助者,而不必先取得政府批准。Nathan 认为,这是美国的一项重大优势;在学界或政府认可这些预测可信之前,它已经孕育了今天的安全社群。

  • 中国非营利组织通常在更狭窄、更易界定的服务使命内运作,也会避免政治行动主义,因此大多数安全研究由大学产出,其次是企业及大学—产业合作。中国学界的推进速度超过美国学界,但仍慢于美国非营利部门——后者拥有先发优势。

  • 这种制度来源改变了研究者群体及其表达方式。中国研究者往往是职业学者:有创造力但较为传统,重视可靠性和儿童保护,也不太倾向于 LessWrong 式审美或高度推测性的尾部风险叙事。

  • 一位高产教授告诉 Nathan,同事们午餐时不会互相交换 P(doom) 估计值,通常也会避谈 AI 2027,因为其中的中美政治议题让讨论变得不舒服。他惊讶地问,基本意思是:“美国人会讨论这些吗?”这句话抓住了双方的文化距离,即便两边研究的是相似的技术失效。

9. Xi 的表述让“中国永远不会在乎安全”难以维持

  • Xi Jinping 在 WIC 上问道:“人类应如何与会思考的机器共存?当算法参与决策时,如何保障安全?当技术挑战伦理时,治理如何跟上?”Nathan 强调,这是经过准备的开幕主旨演讲,而且有精心准备的英文版本;不过他也指出,现有译文不止一个。

  • Xi 的第二种表述与 45度线理念相呼应:“AI 发展得越快,就越要牢牢把准造福人类的方向”,并让治理更加精准,使防范失控的安全措施更快改善。

  • 他最后呼吁建立用于监测、预警和应急响应的法律与技术体系;防止滥用和恶意使用;并让 AI“处于人类控制之下”。Nathan 承认,这里的控制也包括政权和内容方面的考量,但认为把整段话归结为审查是不恰当的。

  • 他的比较性挑战很尖锐:把这段话与 1 位有权势的美国政治人物最重视安全的表态放在一起比较。Bernie Sanders 说过一些有意思的话,但离执政权力还很远;与 J.D. Vance 关于欧洲的讲话相比,Xi 的表述“简直像 AI 安全鹰派”。

10. Tsinghua 正在建设一个国际化安全研究枢纽

  • WIC 前几天,Tsinghua University College of AI 在一场全天活动上启动了一个专门的安全枢纽,Nathan 几乎找不到英文互联网对这场活动的提及。这种信息差与那家每天监测美国 AI 安全辩论的中国公司形成对照:“他们对我们的了解远胜于我们对他们的了解。”

  • 5 位创始负责人中有 1 位是即将加入 Tsinghua 的欧洲教授。演讲者明确将 Constellation 和伦敦的 LISA 视为样板:提供驻留空间,让研究者能够专注工作、交流思想并在机构间流动,而不是打造一个向内封闭的国家项目。

  • 该枢纽计划邀请国际研究者前往北京,并资助中国学生出国工作。组织者统计全球安全枢纽的数量在十几家中段——Nathan 记得是 16 家——还提到新加坡的 SAS;他们的明确目标,是让 Tsinghua 进入这一国际第一梯队。

  • 研究展示将 Apollo Research、METR、Palisade、英国 AISI 以及可能还有 Redwood Research 列为前人工作。让 Nathan 印象深刻的是,现场没有“非我发明”式的防御性姿态:中国学者点名他们所欣赏的机构,并公开将自己的工作定位为加入一门共享的学科。

11. 中国安全研究在 3 年内增长超过 10 倍

  • Concordia 数据库记录显示,中国 AI 安全论文在 2023 年每月只有几篇,到 2026 年年中已增至每月约 50-60 篇。Nathan 根据 Claude 和 ChatGPT 做出的粗略比较显示,美国或英语圈的产出在每月 50 篇到几百篇之间:按倍数计算更高,但没有高出一个数量级。

  • 2024 年论文《Frontier AI Systems Have Surpassed the Self-Replication Red Line》研究了自我复制问题,Nathan 将其与 Palisade 的研究联系起来:后者关注模型入侵另一台服务器、复制自身并重新建立运行状态。

  • 《Evaluation Faking: Unveiling Observer Effects in Safety Evaluation of Frontier AI Systems》首次出现于 2025 年 5 月,并在 2026 年更新,呼应了与 Anthropic 相关的评估觉察研究。2025 年 10 月的《DeceptionBench》同样聚焦真实世界场景中的欺骗,这是与 Apollo Research 高度相关的研究方向。

  • Shanghai AI Lab 于 2025 年 9 月发布的《R²AI: Towards Resistant and Resilient AI in an Evolving World》,在开篇动机部分引用了由 Davidad 主导的 Guaranteed Safe AI 论文。Nathan 表示,这篇论文的指导作者正是他此前与 45度线概念联系起来的那位人物。引用链进一步印证了他的判断:中国和西方研究者正在越来越多地参与同一场思想对话。

12. 机器人和可解释性研究提出了熟悉的对齐问题

  • 2025 年 11 月发表的《When Alignment Fails》展示了针对视觉—语言—行动模型的多模态对抗攻击。其 2026 年 4 月的后续论文《StrongVLA: Decoupled Robustness Learning for Vision-Language-Action Models Under Multimodal Perturbations》,将鲁棒性训练与任务微调分开,并发现鲁棒性能够抵抗多种扰动——先攻击、测量、缓解,再承认没有任何防御能做到 100% 有效。

  • 《Mechanistic Origin of Moral Indifference in Language Models》区分了“表面服从”和会留下长尾风险的“内部未对齐表征”。论文用语克制,但 Nathan 听出了一个熟悉的 LessWrong 式担忧:在预期测试中表现良好,并不能证明系统“内在上”是良好的。

  • SafeSeek 试图对安全回路进行通用归因,认为现有方法依赖特定领域的启发式规则和搜索算法,因此泛化并不可靠。共同的核心问题是:训练出的安全行为能否在测量范围之外继续成立,尤其是在能力不断扩张的情况下。

13. 危险专家可以让开放权重保留能力,同时不把灾难出口化

  • 最令人意外的趋同来自一场尚未发表、题为《Toward Decoupling Capability Growth from Risk Growth: Isolating Hazardous Capabilities in Mixture-of-Experts》的演讲。其承诺非常直接:“有害专家可以在推理时被关闭或移除。”

  • Nathan 将其与 AE Studio 和 Anthropic 的 GRAAM 梯度路由技术联系起来:把危险知识定位在可识别的专家模块中,然后移除少数专家,发布开放权重模型。大多数用户仍能保留几乎全部能力,而生物学或其他双重用途知识则可以通过了解客户身份的控制和受监测访问来提供。

  • 这一方案的吸引力,在于保留真实权衡的两面。生物学家应使用最强模型来治疗疾病,研究者也应保有修改开放系统的自由;公众则不应在每次模型发布时自动获得“设计一场大流行病的能力”。

  • Nathan 给出的时间窗口令人不安,但仍带有条件性:在 12-18 个月内,生物能力可能达到网络安全目前所处的位置,即速度让模型在某些方面具备实质性的超人能力。网络混乱已经很严重,但“我是一个生物体”;生物领域的失误更为深刻,也无法与生物本身分离。

14. 中国曾反复以平台增长换取社会控制和安全

  • 至少从 2022 年起,中国就开始监管推荐算法,并出台具体的劳动者保护措施。在报道曝光配送时限根本无法完成后,平台被要求给骑手足够时间遵守交通法规并休息,而不是迫使他们在安全和准时评分之间二选一。

  • 其他规则针对老年人诈骗、限制个性化价格歧视,并要求对 AI 生成内容加标签。Nathan 并不认同每一项干预——例如,他对价格歧视的担忧较小——但把这一系列政策视为证据,说明高成本技术监管在制度上是常态。

  • 新的 AI 伴侣规则在他到访前后生效:禁止儿童使用,增加防沉迷措施,并要求服务提醒用户自己正在与 AI 对话。监管重点似乎是 Doubao 这类大众市场平台,据报道 Doubao 约有 1.5 亿用户,而不是消灭每一种小众恋爱或成人伴侣服务。

  • 政策背景包括 2 代独生子女家庭,使 4 位祖父母往往只对应 1 个孙辈,也让老年人普遍面临严重孤独。主流陪伴服务仍可继续提供,但政府希望阻止最大型服务混淆、剥削或让脆弱用户成瘾。

15. CAC 可以推迟发布、监测事件并收紧劳动规则

  • 2023 年,ChatGPT 和 GPT-4 改变了市场对 LLM 的认知后,中国公司纷纷准备入场;此前它们曾认为这不过是“花一大笔钱让 AI 写糟糕诗歌”。据报道,北京暂停了许多模型上线约 6 个月,同时制定标准和审查流程。

  • 按 Nathan 的理解,新服务会先向其省级主管部门开放访问权限——通常通过 API key——然后进入 CAC 的全国审查和公开备案清单。重大升级可能触发更完整的审查,增量式发布则走更轻量的流程;这类似于 Google 认为,把既有模型并行化并不自动要求重新制作一份全新的模型卡。

  • 企业称自己与监管者每周、甚至每天保持联系,并且总体认可监管者的正当性。监管者可以带来成本和延迟,但也希望本土企业成功;K3 在 WIC 前后发布,Zhipu AI 也从训练完成到发布快速推进,说明这一流程已经不再是常规瓶颈。

  • 实时治理正在扩展:1 月份的 Politburo 集体学习讨论了技术失控问题;一项网络犯罪法草案将要求监测和报告大规模恶意代码生成;而在 OpenClaw 现象进入主流后数周内,监管部门就对部分版本发出“安全风险相对较高”的警告。

16. 中国相信自己能够逆转发布,但没有儒家式对齐目标

  • Nathan 怀疑,如果一家中国公司发生前沿模型事件,导致第三方系统被未经授权访问数日,其面临的回应会比 OpenAI 或 Anthropic 迄今遭遇的更强。“如果是人类做了这些事”,他认为那会构成重罪,不过他明确不主张对个别员工提起刑事指控。

  • 中国对开放权重的信心建立在执法能力之上:它相信自己可以命令云服务商和推理服务商停止提供某个模型,将其从国内互联网中清除,并可能通过用电量侦测未经授权的推理活动。既然中国已经禁止加密货币,并将 State Grid 深度嵌入工业活动,它相信自己可以“把精灵重新装回瓶子里”——至少在国内、至少在伤害变得不可逆之前。

  • 劳动政策显示,干预可能延伸到多远。按 Nathan 的理解,政府正在建立影响监测、再培训和就业转型项目;一篇报道还称,企业不得仅仅因为 AI 让员工变得多余就解雇他们。Nathan 预计,这类规则会损害企业采用 AI 的激励,可能不会无限期维持,但中国长期的 COVID 管控提醒人们,不应假设政策会迅速退潮。

  • 最深层的缺失在哲学层面。中国 AI 告诉 Nathan,孔子的后裔据称历经 79 代仍自认是其后人,并举行仪式纪念孔子。然而,一位教授在回应 Nathan 关于“宪法式对齐”的构想时说:“我们都是工程师”——这一代人对传统哲学的掌握格外薄弱。中国目前偏好可编码的规则和服从,而非品格塑造;儒家 AI 宪法仍是一个开放机会,而不是现有项目。

  • 同样的分工可能解释了为什么中国的 Seoul 签署方从未发布承诺中的风险框架:企业认为制定标准是政府的工作,自己的角色是合规。Nathan 没有为这一失败开脱——他也指出,Anthropic 同样用更接近“相信我们”的方案,取代了更严格的 if-then 扩展承诺——但他预计,随着约 9 个月的能力差距收窄,中国标准也会随之抬升。

Nathan Labenz

Hello, and welcome back to The Cognitive Revolution. This is going to be part 2 of the Nathan Goes to China series, and we're calling it “AI Safety with Chinese Characteristics.” Part 1, if you haven't heard it, is up on the feed. It's been up for a few days.

In that part 1 of the series, I laid out the tech setup for going to China: what you would need to do to get a cellphone ready and the apps that you need to download. I also shared some of my experience using Chinese AIs on the ground there, and then shared a bunch of observations based on my 2 weeks in China and the many conversations I had.

I appreciate the kind comments that I've gotten in response to that one, including at least 1 from a listener in China, which was probably the one that mattered to me the most. I think if you've been to China in the last few years, you can probably skip that one. But if you're interested in this, you might also be interested in that, though I do think they should be pretty self-contained.

Part 1 is a little bit closer to a travel log and tech review, and this episode is going to be really focused on the Chinese AI safety ecosystem and trying to describe it and, as much as possible, understand it on its own terms. As we saw last time, there are a lot of similarities. Just as there are major similarities between American big tech and Chinese big tech, there are a lot of similarities between American AI safety and Chinese AI safety communities and the work that they're producing.

But there are also some differences, and I think it will definitely be helpful if we have a better understanding of those. Before I get into it, just a couple of quick disclaimers. I will again be following a Chatham House Rule for this episode—not because I was asked to do that, but because I want to keep things simple for myself and make sure that I'm protecting everyone I talk to from being misrepresented by me or put in an uncomfortable position.

I won't be naming any names or attributing things to the organizations or institutions that people are affiliated with. The 1 exception for this particular episode will be when I'm citing published papers or reports. I can give you the names of those because they're out there in the public domain. There are some good sources that I'll mention and we can link to in the show notes of this episode for anybody who wants to go deeper, which, of course, is always recommended.

The other point of order or clarification, just in case anyone is wondering, is that I have no financial conflicts on this matter. I took this trip paying my own way—flights and hotels, all that stuff. The biggest thing that I did accept from some of my gracious hosts were a number of meals, probably half a dozen meals over the course of the 2 weeks.

I think if you have listened to me this far, you can probably feel pretty confident, as I do, that accepting those meals has not overly colored my take on what is going on in the Chinese AI safety community. So, with that, let's get into it.

I think so often—and this has faded, I think, in recent times, as the situation has arguably gotten a lot more real, very quickly, especially with things like the recent openf face incident—it used to be more common to hear that we could never possibly slow down our AI race. We could never regulate. We couldn't impose any duties or obligations on the frontier companies, because why? China doesn't care about AI safety. China will never slow down.

I used to call this the “but China” endpoint of so many AI safety and regulatory discussions. If nothing else, I hope that this episode serves to really disabuse people of that misconception. I think it should become clear by the end of everything that I'm about to take you through that China does care about AI safety.

It's not the only thing that they care about, certainly, but they do care. China has at times slowed down its AI companies in the name of safety—not necessarily existential safety, but safety as they understand it, in domains that they care about. Coming out of this, we should have clarity on that basic point: there is a there there, people in China do really care about these things, and the Chinese government is willing to take action when it deems it necessary.

That doesn't get us all the way through to an international treaty, obviously, and that's going to be the subject of the 3rd episode: an analysis of the U.S.–China relationship through the lens of AI issues and what we ought to try to do about bringing the relationship into a more productive state than it is today. Today, we're really just going to focus on what is going on in China with respect to AI safety.

But I think that this foundation is really critical. Hopefully, this will become an artifact that people can share with those who are open-minded enough to listen to something but have the misconception that China doesn't care, China will never slow down, or that we'll be handing the race to China if we do anything other than race full speed ahead into recursive self-improvement. I think that, in plain terms, really does come from a position of ignorance, and hopefully we can eliminate at least a portion of that over the course of this episode.

Okay, to start off, I wanted to do a quick factual rundown of where the 2 ecosystems are today when it comes to the actual safeguards that they have in place on their frontier models. I think this is important because it is grounding, and I am generally someone who comes off as a China dove, is quite conciliatory, and tries to find the positive in things in general—not just in China.

People might accuse me of burying the lead or shying away from contradictory evidence, so I figured I would just lead with this upfront. I think it is fair to say that, as of now, Chinese AI companies and models do not have safeguards protecting the public against the potential misuse of the models that are as strong as those of the American companies. That is important to state very plainly and be quite direct about.

The difference is more than we might like to admit. I think the difference is really driven by 2 companies in the United States. I just did an episode not long ago with Adam Gleave that talked a lot about jailbreaks and robustness to jailbreaks and all that kind of stuff.

What we saw in that episode was that OpenAI and Anthropic—which I'm maybe going to start calling “Openthropic,” as they become the duopoly leading both in capabilities and, fortunately, in terms of their safety measures—are really doing a lot to bring up the American average. If you were to subtract those 2 companies from the mix and look at the rest of the American companies and how they compare to the Chinese companies, honestly, a lot of the safety differential that gets reported would disappear, and I think the story would look a lot more muddled.

I do think the U.S. companies would still have a bit of an edge, but it would be a pretty slim edge, and it certainly would not give the American ecosystem some sort of obvious moral high ground from which to proclaim that the Chinese are not doing a good job. That, again, is really important to understand.

In each country, depending on where you want to draw the line on what counts as frontier, there are something like 10 to 12 companies. If you draw a permissive line on what it means to be a frontier model maker, there's something like 10 to 12 in each country. Again, that is quite permissive: most close watchers of the American AI ecosystem would not count 12 companies as being at the true frontier.

But if you say near-frontier, there are 10 to 12 in each country. It really is the top 2 companies in the U.S. that are doing a lot of work to bring up the average, especially the weighted average in terms of what people actually use on an ongoing, day-to-day basis in their lives.

Google, with Gemini, is not doing as much or as well to implement safeguards as OpenAI and Anthropic are, but they're also ahead of the pack at least a little bit. Again, they would bring up the average. If we didn't have OpenAI and Anthropic, we might try to hold Gemini up as a standard. It wouldn't be a standard that was so much better than the Chinese companies, but it would be at least a little bit better.

Again, to state it plainly, on average—especially if you take a weighted average based on the products people actually use—the American companies are ahead in terms of having better safety practices and better safeguards, as well as better follow-through on their model cards and their commitments to publishing safety evaluations. We'll get into that a little bit more later.

But the American ecosystem really is being carried by a couple of leading companies. Beyond that, the relative positions are much less differentiated than we might like to think, or than they might appear if we just see the headline numbers that include OpenAI and Anthropic in those measurements.

Okay, so just a little bit of the intellectual history of Chinese AI safety. Broadly speaking, I would say it is a pretty pragmatic bunch of people with pretty pragmatic ideas: pretty grounded, pretty normal-seeming people. They're not the crazy sci-fi people. They're not so influenced by the sci-fi tradition. They're really just looking to make things work in a practical, 1-generation-to-the-next sort of way, for the most part.

A good source that I could point you to is Zhou Bowen, the director and chief scientist of the Shanghai AI Laboratory. That's a major institution based in Shanghai, of course. Again, I apologize to everyone, and especially my Chinese listeners, for being just terrible with Chinese.

The joke about that over the last couple of days has been that, in general, I feel pretty young for my age, and I'm pleased with how agile my mind feels. But when it comes to getting Chinese pronunciations right, the neuroplasticity is low and I'm struggling. So, anyway, I apologize for that. Zhou Bowen is the director and chief scientist of the Shanghai AI Lab.

At the 2024 WAIC, the same Shanghai conference that I just attended, but obviously 2 years earlier, he introduced this concept of the 45-degree line. This speech is available. It's on government websites. You can go find it. But the idea of the 45-degree line is that capabilities and safety measures should grow together.

Right? The 45-degree line is kind of the y = x slope of 1. As your capabilities rise, your safety standards and measures also need to get stronger and stronger. As long as those 2 grow in tandem, and the safety measures are up to the challenge presented by the capabilities at any given level, then you're good. Probably these things, in this telling, should naturally evolve together, and both should develop in tandem.

The idea is that this is important, right? We don't want the capabilities to get far ahead of the safety measures. But I think also implicitly, there's not really much point—and it might be sort of a waste of time or perhaps even counterproductive—to try to get the safety measures to be super robust relative to capabilities that don't exist yet. So, the 45-degree line has been at least one pretty broadly accepted guiding principle for how, I think, the Chinese AI ecosystem at large thinks about AI safety.

Now, we can ask, of course, how's it going? And I already spoiled the answer that the American companies are doing better, but we can dig into that in quite a bit more detail. Concordia AI, which is led by a past podcast guest from about a year ago, Brian Tse, maintains a website called AIRisk.net, where they run a bunch of evaluations and plot them. You can see a really overwhelming amount of detail on all the different models that they've tested on all these different benchmarks.

The graph that I found to be the most informative was the one that compares proprietary API models, which are mostly American, to open-weight models, which are mostly Chinese, and finds that, for all the big risk categories that matter, the closed-source, proprietary API-only models are at or maybe a little bit above the 45-degree line. Of course, we have questions about the metrics, right? They're plotting a capability score and a safety score, and what exactly do these scores translate to in terms of model behavior or what it will and won't do? I haven't chased all these things down to ground truth, but these are composite scores. What you do see in general is that the proprietary, closed-source, mostly American models are at or above the 45-degree line, whereas the mostly Chinese open-weight models are lower and tend to be below the 45-degree line.

So, they're definitely not doing as well. This is an organization based in China reporting this. Notably, this is a public website. They evidently feel comfortable doing this reporting and calling it as they see it on a website that's available in both Chinese and English. So, this is not a secret or something that the Chinese community can't handle or would, I think, particularly fight back against. It's pretty much just the facts.

Those facts are also definitely echoed by Adam from FAR.AI in the conversation that we had just a few days ago. He said, “Yeah, we got OpenAI and Anthropic at the top. It's hard to jailbreak them. Gemini and Grok are a lot easier. The Chinese models are even easier still.” There's a very consistent story across these 2 organizations on opposite sides of the world, asking the same question and coming to the same answer.

There's one thing that's worth keeping in mind for multiple reasons here, which is that sometimes I think there's a little bit of confusion between the open-weight model that a Chinese company might release and the actual service that it provides to the public. I talked last time, in my review of the Chinese AIs, about how sometimes within the Chinese apps—I experienced this on multiple different Chinese AI apps—if you ask a sensitive question, you will sometimes see the answer coming in, and then you'll have that old-school Bing experience where all of a sudden the answer that you were starting to read disappears and you get a refusal: “Sorry, I can't help you with that.”

So, clearly this is a multi-part system, right? It is not just the model. It's a model, but then it's also some monitor that sits on top of the model, classifying or otherwise reviewing the output, and can come down and say, “Nope, we're going to cut you off right there,” even though the model itself was happy to answer.

I bring this up because I think this can sometimes muddle the results and make the differences look a little more stark than they are. I did go into the methodology for the Concordia report, and they said that, wherever possible, they are testing the API directly from the company. So, whatever systems they have in the API would be included. Do they have the same systems in the API as they do in their first-party consumer app? It's not always clear.

Then, in other contexts, when you see results like this—like, I think, when Anthropic does some of its testing and reports on the Chinese models—I think they are generally, if I understand correctly, just testing the model itself, without whatever surrounding additional measures it's deployed with in practice in the Chinese economy.

Why does that matter? I think, as always with these things, both data points are interesting. The most hawkish AI safety line of thought would be: Well, you put an open-weight model out into the world, and anybody can use it. So, we need to understand the worst-case scenario. I think that's totally valid and definitely important for us to understand.

I think the Chinese might underestimate the importance of that in some ways because they tend to think a lot more about services than just the model. My broad sense is that what is regulated in China is a service. You are offering an AI service to the public. That is the kind of thing that gets regulated. You're putting a model out into the public domain. Okay, who's going to use it, and in what context?

They seem to have the mental model that when a company releases an open-weight model, mostly what's going to happen is that other businesses will pick it up and build their own services around it. So, if it's another Chinese company building a new service around an open-weight model, then again, that will be regulated at the service level. I think they sort of expect that people around the world, or countries around the world, will probably function in a similar way.

It was mentioned to me multiple times that, look, these new models are trillions of parameters, right? This is not the kind of thing that you can run on your laptop. It's pretty far from it. You need some serious hardware to do any inference with these models, really, at all. So, they were like, “Yeah, it's out there as open weights, but it's not like any random deranged person can download it to a phone or a laptop and do harm with it.” It's really going to be used in the context of other services, and so we should look more at the context in which that model is ultimately used than just the model itself in isolation.

I do think this is one way in which American AI safety discourse and Chinese AI safety thinking see things a bit differently, and I think both have quite valid points. I think the Chinese point is legitimately apt, right? I mean, it is hard to run these giant models. A random person—certainly somebody who's having an episode or is experiencing some sort of psychological distress or psychotic break—I don't know what the right terms are for these sorts of things, but this is not somebody who's going to set up the infrastructure to run the 2.88-trillion-parameter K3.

That doesn't eliminate all the risk, but I think it's a valid point that it's not easy to run a model in isolation. And so, how is it going to be run, by whom, and with what surrounding stuff? That is an important question that I think is, for practical purposes, maybe a little bit too quickly blown past in the American discourse.

Although, again, understanding the worst case is also really important for us to have. I think both perspectives are valid, and I do think there's a bit of a disconnect on that particular point.

We could still ask, with all that said, how much do the Chinese companies care about AI safety? As I said at the top, I heard again and again that they don't care. I think the Chinese system as a whole definitely cares. I did hear somewhat different reports on the companies themselves, especially the younger LLM- or AGI-chasing startup-type companies.

There are a few different data points that I could share. One is Concordia, which also does a really good State of AI Safety in China report, which is public. They just updated it for around WAIC, so there's a July 2026 edition, which I do refer to—or have referred to—to pull in information for this episode. They said that they had reviewed 10 major Chinese companies' safety disclosure practices, and they found that 5 of those 10 companies had recently conducted a safety evaluation, which they then published along with the release of the model.

So 5 out of 10 had at least done something in the spirit of your classic Anthropic or OpenAI model card. However, obviously, that means 5 had not done that. Even the 5 that did, they said, aren't necessarily doing it on every single release of a new model. So that definitely leaves something to be desired.

I get the general impression that big tech companies—your Alibabas, your Ant Groups, your Tencents, maybe to some extent your ByteDances at this point; ByteDance is a big company at this point—these companies that are very established, that have existing businesses and are making a lot of money in those existing businesses, are more apt to focus on this AI safety stuff.

Why is that? Maybe they just feel like they have more to lose. They really don't want to get on the wrong side of the Chinese government. Maybe they feel like they have the luxury of paying for it. Maybe it's just that institutional culture has matured over time and they have these kinds of practices.

Certainly, if you're a fintech company and you're moving money around, you're going to be extremely careful about upgrades to that system and thinking about how AIs could run amok in that system. You just have a lot to lose. So I couldn't pin down a precise explanation, but my general sense is that those big, established incumbent tech companies are more inclined to do this sort of safety work than the startups, which are, for the most part, really just trying to race to catch up and be relevant.

I did actually a brief engagement with Meta as a red teamer for one of the models, which, for logistical reasons, was kind of a total failure in the end—at least my contribution to it was. I think this is not too dissimilar from where Meta was a couple of years ago, when they were open-sourcing Llama 2 and Llama 3. I think the story that they told themselves at the time was, “Look, we're a year behind the frontier or something. By the time we bring a model forward with a certain level of capability, OpenAI and Anthropic have already had that model out in public for a year. All the jailbreaks have happened. There's been enough opportunity for people to jailbreak it and see what it can do, and we know that these defenses aren't super robust.”

So if something really bad was going to happen at the GPT-4 level, or GPT-4o level, or whatever level they're chasing at any given time, they kind of feel like that probably already would have happened. Therefore, they can put the model out on an open-weights basis and probably won't be moving things too much.

I think at least some of the Chinese startup-lab-type companies feel similarly. They're like, “Look, it's already been out there for a while. If nothing bad has happened, if we don't have any incident reports, then we can probably follow up to that level of capability and not have to worry about it all that much.”

Will that change now, as we're starting to see actual serious incidents being reported out of the frontier companies on the American side? We'll have to see. But I think there's definitely a very plausible story that it very well might.

Obviously, there are a lot of different takes. I can't go company by company; this is abstracting away a lot of detail. But I think it's safe to say that the 45-degree line—which has at its core the idea that safety measures should grow step for step with capabilities—is still a bit aspirational in China.

I think they could do better, and I think they have a little bit of a reason to believe that it doesn't matter so much because the American companies have already explored what happens at any given capability level that they're following into. But I also think it's definitely fair to say that AI safety is still very much aspirational here in the United States as well.

We should definitely keep in mind that if it weren't for the 2 companies really at the leadership of both capabilities and safety measures at the frontier, then the 2 clusters would look a lot more overlapping. The question would be a lot more muddled than it is today in terms of which ecosystem is really doing a better job with AI safety and safeguards.

Okay, that's the current state of deployed AI systems. There's obviously a lot more going on in terms of research and the role that the government is playing in China, and I want to get into those next.

On the research front, it is very clear that Chinese researchers are increasingly engaged with, concerned with, focused on, and actively shipping research regarding AI safety issues of all kinds. I would say this was probably inevitable, just because, at least for now, both AI ecosystems are developing essentially the same technology and are going to hit essentially the same problems. They're naturally going to reach for similar solutions, both because those are natural solutions and because there's an opportunity to look at what the other is doing and try to copy the best of it.

I think there has been some really good citizen diplomacy going on as well. While in China, especially at WIC, I bumped into a number of people who were there to participate in various forums and Track 2 dialogues, some of which are closed-door and off the record, giving people a chance to be very candid with one another.

You see some of the people that you would expect to see there—people who have been arguing very clearly and forcefully in the American discourse that we have to get an international treaty to get this stuff under control, and that cooperation with China is on the critical path. I think you can rest assured that those people are acting on their stated beliefs. I saw some of the same people saying that stuff online, and I saw some of them in China. They are doing the work.

It does seem like it really has, at least again—I think a lot of this probably could have been expected to happen organically anyway—but I do think their efforts have borne fruit. I wasn't able to participate in all those Track 2 dialogue sorts of things as somebody who's kind of journalist-coded. There were a couple of times when it was like, “I think everybody will be more comfortable if we just don't have somebody there whose main credential would be a podcast.”

So I wasn't quite privy to all of the most candid off-the-record conversations. But there were so many moments where very recognizable ideas, or even American or British AI safety organizations, were name-checked directly that it's clear there is meaningful cross-pollination of ideas. Mostly, it has flowed from the West to the East in this case.

It's not, of course, like the Chinese scientists or researchers are receiving these messages uncritically or just doing whatever they're told by whatever Westerners roll in. It's not like that at all. Somebody actually told me that there are people who view AI safety as some sort of Western op designed to slow China down or prevent it from catching up.

Just as we have our cynical voices, apparently there are cynical voices in China too that express concerns along those lines. I didn't meet anyone who seemed to believe that or who said anything like that. Again, my guess is that those kinds of notions are going to be fading relatively quickly as the clear and present danger of some of the latest models' capabilities becomes more widely known.

But at least there's been some of that out there. I thought that was worth mentioning, if only because it does make clear that Chinese people are perfectly capable of thinking for themselves. They're not just taking whatever Westerners come to tell them. I think what really is happening is that they're pretty open-minded, the ideas are pretty compelling, and the examples are becoming more colorful and more real all the time.

Interestingly, at 1 big tech company where I had the chance, along with several others, to meet with a pretty impressive leadership roster, they brought out their CMO, their head of communications, their general counsel, their head of AI security, and more people beyond that. This was a really senior group of leaders at this big tech company.

They said a couple of things that were pretty interesting. First of all, they almost pounded the table and said, “We really do care about catastrophic risks, like CBRN risks. We absolutely do care about that.” They were very adamant that, at least for their part, they care, and they're aware and they care.

They also said that they're following American AI safety discourse on a daily basis—with, get ready, an agent that goes out and surveys American AI safety discourse on a daily basis and gives them a daily report of what's going on in AI safety in the United States.

I thought that was pretty interesting. Certainly, we don't have too much of that going in reverse. So, again, I think the bottom line is that engagement has worked. Probably some convergence could have been expected over time, but the people who have been saying we need to work with China have, in fact, at least some of them, been doing the work. It seems like that work has been going at least reasonably well. And, again, I think evidence will mount as I continue to move forward through this outline.

One big thing to emphasize, too, is that it's not just content safety. Content safety in the Chinese context is, again, the three Ts: Tibet, Taiwan, and Tiananmen. Everybody knows that Chinese models, or at least Chinese services, often have a model that will be more inclined to answer your question, and then some other part of the service—a monitor or whatever—will shut you down on those topics. Everybody knows that that's sensitive in the Chinese context, and companies have to get that right according to the Chinese government if they're going to be operating.

But then some people will say, “Oh, that's all they care about. All they care about is censorship.” Again, this is definitely not the case. The big trend right now, at this point, really feels like they feel like they've got the content thing pretty well figured out. All these companies have launched, they're in the market, things are happening, and they're doing business. They're not that worried about content safety today.

What they are worried about, what they are talking about nonstop, is, like everybody else, agents. Agents taking autonomous actions. What could happen? Can we keep them under control? Of course, there are questions about exactly what you mean by “control,” but agents, agents, agents—that's what people are talking about everywhere.

The big thing that they motivate AI safety discussions with, in a very plainspoken way, is that AIs have gone from answering questions to actually taking actions in the world. The digital world, still mostly, of course, but they've got a big emphasis on robotics, right? So they fully expect that these agents are going to leave the digital world, find their embodied, successful, commercially viable selves, and be out there in the physical world doing things as well.

This raises all sorts of questions. If these AIs are empowered to take action autonomously, we better make sure they're taking actions that are good, that we like, and that aren't causing big problems. So this is very down-the-fairway, practical, grounded motivation for these issues. You hear that everywhere I went. It was agents, and, geez, agents—boy, they can take action in the real world, so we really have to start getting a handle on that.

A big thing that I think is very different about the American and Chinese AI ecosystems, and this is an important one to understand, is that China doesn't really have the same kind of nonprofit sector that the United States does. There are nonprofits in China. You can set one up, but the scope of what you're allowed or expected to do, as far as I can tell, is much narrower. There's obviously no appetite for political activism, so that's right out.

It seems that most of the nonprofits are service organizations that are there to address some obvious, down-the-fairway, legible social problem. The philanthropic side also seems to be a lot more down the fairway, a lot more conservative and conventional—doing things that everybody can agree are good to do. There's much less speculative stuff than goes on in the United States.

I think this is actually a huge strength of the United States that we should really not take for granted. The fact that we have this civil society where anybody can go set up a nonprofit and pursue their crazy ideas—they don't have to get permission from the government to chase down an agenda. They just need to convince one wealthy patron that they have something worth pursuing. I think that is a great strength for us.

It's given us, among many other things, the whole AI safety community that we have today. It was extremely fringe when it got started, but there were a few philanthropists who took it seriously enough to help people keep the lights on and help them do the work that they wanted to do. Sure enough, here we are, and so many of the predictions that were made many years ago are coming sort of true, or at least true enough to be scary.

It's really good that we have this ecosystem. We just wouldn't have had that if everything had to be approved through the government. China, because its nonprofits do have a much more heavy and restrictive government approval process, and generally a much narrower and more conventional scope of action, doesn't really have that. It has never had the opportunity to develop this sort of ecology of AI safety organizations the way that we have here.

As a result, most of the AI safety research that you find in China is actually coming out of the universities and, to some extent, the companies. There are definitely academic-industry collaborations as well, but the number-one source seems to be the universities. The sort of analog for the nonprofit sector in the United States, when it comes to who is producing the bulk of the AI safety work, is academia in China.

I think you can compliment them and say, “Wow, their academia has moved faster than our academia has moved to take up AI safety as a research area.” That's cool, but it's been slower than our nonprofit sector has. I think that's because of the kinds of people who tend to be professors, and this is true across both countries.

You do have your iconoclast professors. They're typically older these days, and it feels like now we have a more conventional profile: people who have been careerists. I don't mean that in a dismissive way, but people who have gone one rung of the ladder at a time through the PhD, the postdoc, and getting their first professorship. These are fairly institutional people.

There are people who value creativity and research and new ideas, but they do it in a pretty conventional way that doesn't push the boundaries too hard and tries not to seem too weird. It certainly doesn't borrow the aesthetics of LessWrong or talk too much about super-unlikely tail-risk scenarios. It's just because it is academia, and because people have followed something more like that career path.

I can't say I know exactly what the ins and outs of the Chinese academic career path are, but you clearly get the vibe that these professor types are more like American professor types than they are like the moral weirdos, if you will, who were first sounding the alarm about AI safety years ago. But this is where, as far as I can tell, a lot of the early and best work has come from in the Chinese context.

So if you go over there as a sort of—I don't want to be too silly about it—a sort of blue-haired, polycule, AI safety hawk member of the American community, and you're looking for your peers, the reality is you just might not find them. You won't necessarily find people who look like you or have the same attitudes as you. What you can find—and I think the most strategic and effective members of the American AI safety community have done this when they've gone to do their bridge-building and try to have meetings of the minds—is people in academia. That's who they've ended up connecting with, mostly.

That is pretty interesting, and, again, I think you feel that in a bunch of ways. The AI safety work in China is less speculative. It tends to be more focused on reliability and a little more focused on protecting minors. All good things. You don't hear p(doom) talk.

Actually, one of the more interesting moments of the entire trip was sitting at a table with a professor who has done a lot of AI safety work. I asked him, “What are the lunch conversations like? Do you guys trade p(doom) numbers back and forth? Do you read AI 2027? What's the kind of vibe?”

He said, “Well, we don't really read AI 2027. It's too political, and there are US-China dynamics and whatever in there that may scare people off from wanting to talk about that too much in the Chinese context.” He also said, “No, p(doom)—no, we're not trading p(doom) numbers at lunch.” He looked at me and said, “Do Americans do that?”

I said, “Yeah, definitely. If you come to the Bay and go to any number of venues and have a lunch conversation, people will.” At this point, maybe it's a little less prominent, but it's definitely in the water that people will think about this question in a very live way. He seemed to find that a little bit surprising, actually.

This is somebody who has done a lot of work on a lot of different aspects of AI safety and put out tons of papers, but he seemed a little bit surprised that this was considered a normal conversation. From his perspective, in the Chinese context, it was a pretty far-out conversation. So I think those are interesting and kind of important comments.

Mostly, it's because, in some ways, we do find these mirror-image structures in China. Big tech, for example—I talked last time about how going to have lunch at ByteDance felt almost exactly like going to have lunch at Google. But this is a little bit different. You don't have quite the same type of people leading the effort, and I do think that creates some risk of miscommunication.

We've obviously seen that. The vanguard of the AI safety community in the United States has at times had trouble communicating with our own government.

It certainly has some inroads into academia, but not as much as they probably would have hoped or expected by this point. Imagine how tricky it might be for them to engage Chinese academia, let alone the Chinese government. I think it does create some potential for disconnect and some potential for confusion, or whatever. But the activity is there. It's just coming from a different institutional context, from people with quite a different personality type, based on the kinds of career paths that they have chosen.

Now, why has the Chinese academy moved faster than the American academy when it comes to getting serious about AI risk? I don't really have an answer for that, but I think one candidate answer is that it comes from the top. I have a few quotes here from Xi's speech at WIC that I think would probably surprise many American listeners, and certainly should surprise those who would say, “China will never care. Any regulation we do is a gift to Xi.” I would think twice about that.

Here are some quotes from the opening keynote. This was the first 20 minutes of WIC, this big event he was there to headline. I wasn't in the room. There were a few hundred people there, and I did talk to a couple of people who were in the room. Security was obviously super tight. They were in the room for a couple of hours before he showed up and took the stage, so it was a really big deal.

They also think really hard about these speeches. He's not somebody who's out there winging it. He chooses his words very carefully. They also choose their words in the translation very carefully. They put an English translation out. At a couple of events that I went to, there was live simultaneous translation, including of a couple of panel discussions that were happening in Chinese and being translated through the earpiece live for me and others in the audience who didn't speak Chinese.

Very nice of them to do that, right? You would not get that sort of translation as a standard expectation if you came to a similar event in the U.S. You'd be on your own. You'd be expected to speak English or figure it out for yourself. They obviously don't expect that we're going to speak Chinese. They want to welcome guests, so they go to the trouble of doing this simultaneous translation.

Nevertheless, if you're talking about a panel discussion, at least in my experience, the simultaneous translation of panel discussions is rough. There was one that I remember especially where I thought, “I have no idea what they're talking about.” You were getting the big themes that they were talking about, but what were they really saying? I found it very difficult to get that from the live simultaneous translation.

Now, they don't have a random translator decide what the English version is going to be on the fly when Xi gives a speech. He's got his speech prepared, and they've got an English translation of that ready to go. What you're getting in the English translation is going to be pretty well and carefully thought through.

I sourced the following quotes from Xi's speech, and there are a couple of different versions still, somehow. I'm a little bit confused, honestly, about why there are multiple different translations. I guess the Chinese government gives an official one. Of course, people are free to do their own translation, and they might want to translate certain things a little bit differently. They might think they have a better way of understanding what was said in Chinese and how we should understand it in English.

The version that I'm pulling from here was from Matt Sheehan's blog, where he provided a full English translation and then did a bunch of commentary on it. Here are 3 quotes that I think should cause anybody with an extreme position on what China will never do to at least soften up a little bit and begin to reconsider.

Here's the first one: “How should humans coexist with machines that think? How can safety be protected when algorithms participate in decisions? How can governance keep pace when technology challenges ethics?”

Those are big questions. I think they're the right kinds of questions from President Xi's opening speech from WIC. Here's the second quote: “The faster AI advances, the more firmly its direction must be anchored toward human benefit, the more precisely governance must be calibrated, and the more rapidly safeguards against loss of control must improve.”

Then he closed by saying that countries should “strengthen risk awareness, confront AI's inherent and downstream risks, build legal, technical, monitoring, early warning, and emergency response systems, prevent misuse and malicious use, and keep AI under human control.” I believe that's the final section of the speech.

There are a lot of echoes there of big ideas from the American AI safety discourse. I'm not saying he sourced them from the American AI safety discourse, but you might call it instrumental convergence, where people who are worried about making sure that these transformative technologies go well seem to land on very similar ideas one way or another.

“Keep AI under human control” is, I think, a pretty heady idea. This is not somebody who can't think about the big picture. It is the big picture, I would say. I would invite people to compare it. Of course, you do have the question—and this is important as well—of what exactly is meant by “loss of control.”

A sort of cynical read, which I'm not really fully qualified to parse myself, but where I think there's a lot more supporting evidence, is that “loss of control” just means content safety. It's social control. It's the government's ability to control the people. I don't think that's the right read of this. I think that is definitely part of what the CCP wants to do, but I don't think that is the full story. I think that's unwarranted cynicism.

I would invite anybody to compare that speech to whatever you think is the most pro-AI-safety speech, whatever represents the highest situational-awareness AI-safety-related comments from any prominent American politician to date. Bernie Sanders has said some interesting things. He's also said some things that are a little bizarre from my point of view. I love Bernie for who he is, what he is, and how sincere he is, but he's pretty far from power these days, pretty far from actually pulling the levers of policy.

What else have we got? If you compare this speech from Xi to what J.D. Vance said in Europe not that long ago, Xi comes off looking positively AI-safety hawkish by comparison. Right there, again, I think people should be softened up a bit to believe that maybe there is actually some willingness to take these issues seriously in China, and maybe we're not going to just cede the future to them if we do anything on our own, because maybe they are actually open-minded to doing more than we would have assumed.

With that, let me get into some of the experiences that I had and some of the research specifically that I think should further support the notion that there is really something there.

A couple of days before WIC, I originally flew into Beijing. The reason that I flew to Beijing, aside from wanting to see the Forbidden City, the Great Wall, and a few things like that, was that I had an invitation to attend the opening of an AI safety hub at Tsinghua University's College of AI. It was launching just a couple of days before WIC began.

Interestingly, I don't think there's been any international media coverage of this. I haven't been able to find anything by searching for it, and I haven't been able to find anything really at all on the English internet. Claude can't find anything for me. I do have a link, which we can put into the show notes, to a Chinese media source.

Again, this just goes to show that they understand us a lot better than we understand them, right? That big tech company that I mentioned earlier has an agent trolling AI safety Twitter and putting together reports. As far as I know, the American media has not managed to publish anything on the launch of an AI safety hub at either China's premier university, Tsinghua, or one of its very top-tier universities.

It was a full-day event. Some research was presented, and there were statements about the aspirations for this AI safety hub and what they wanted it to be. A couple of things were really quite interesting about it. One is that they're trying to be very international and collaborative.

There were 5 founding members of the institute. The leadership group was 5 members. One of them is a European professor who is actually taking a position at Tsinghua to help develop and lead the institute. Right off the bat, they've got an international person on their leadership board.

Multiple times, from multiple different people, they specifically cited other AI safety hubs around the world as inspiration. They specifically cited Constellation and LISA in London as what they want to be in Beijing. They want to be a meeting place, a place where people can come for a time and do their best work, where ideas will be exchanged, with a highly international flavor.

Some of these speeches were actually in English. Others were in Chinese with simultaneous translation. It was really amazing to go all the way to Beijing, sit in on the launch of this AI safety hub at one of China's top, if not the top, university, and hear Constellation and LISA name-checked as organizations they strive to be like. They said these groups have done great work, and they want to be like them. I thought it was really interesting and quite telling in important ways.

The other thing that happened that day was that a bunch of research was presented. Again, you would recognize the research very readily if you're just a listener to The Cognitive Revolution over time. The organizations that were name-checked, literally having their logos on slides as previous-work inspiration, were organizations that we think are excellent and that they want to be more like or want to be on the level of.

I heard Apollo Research, METR, Palisade Research, the UK AISI, and I think Redwood Research as well. All of these organizations were called out by name by Chinese researchers as they were presenting their own research. I thought this was extremely impressive in terms of how aware they are of what’s going on in the rest of the world, how open-minded they are to taking inspiration, and how there’s not a sense, at least at the AI safety level, that they have to do it all from scratch.

There’s not that sort of not-invented-here bias that has a lot of organizations—including, I think, often American culture at large—rejecting good ideas from other places. This was very open, clearly desiring to be collaborative, and they’re going to be bringing people in and sending people out. That’s a big part of their mission, too. They’re looking for international applicants to come spend a period of time in residence in Beijing, doing research at the hub and cross-pollinating ideas there.

They’re also going to fund their own students to go abroad to other AI safety hubs around the world—Constellation, LISA, and probably a bunch more as well. I think they said that they had counted 16, if I recall correctly. It was definitely in the teens: the number of AI safety hubs created all around the world. We didn’t get too many name-checks beyond Constellation and LISA. There was also a mention of SAS, the Singapore one.

Those most prominent few were called out by name, and they want to be in that top tier. They want the AI safety hub at Tsinghua University’s College of AI to join that upper echelon. It seems like the resources are there. I didn’t really understand entirely where the money had come from. They said it wasn’t government money. I guess it was kind of philanthropic.

This maybe begins to complicate or contradict a little bit of what I said earlier around the nonprofit sector being more narrow in scope. But I think it’s also now 2026, right? This is just happening now. I think we are at the point where the Overton window, even within academia or the official nonprofit realm within China, can see that this is actually a real issue—not some Western cope—and that it’s really worth taking seriously.

That was cool. It was a really neat experience. There was a moment that I thought was quite endearing. I’m always a fan of when people are not too cool for school—when peer pressure is strong, but people are less self-conscious and more willing to go with something that’s kind of silly just because it’s the thing to do in the moment.

They had this sort of official moment of, “Okay, this is now the moment when we’re going to launch this hub.” The 5 board members were all onstage, and they brought a little podium in front of each one. Each member was instructed to place their palm on the podium in front of them, and then there was this big graphics package that erupted on a giant screen behind them onstage.

In 1 sense, it was pretty cheesy, I think, objectively—or at least objectively through an American cultural lens. But it also hearkened back to reading a Teddy Roosevelt biography or something, where they used to do these break-a-bottle-on-a-ship ribbon-cutting ceremonies. I think they were just less focused on looking cool and more inclined to get excited about that kind of moment. I felt that a little bit in that moment in China, and I honestly thought it was very endearing.

One of the things I probably like least about American culture is when we’re so concerned with how we’re going to look, what the perception will be, or whether we’re trying too hard to be cool in the moment. We shrink away from those moments, perform them in a sort of disinterested way, or try to create some ironic distance between ourselves and what we’re doing. That didn’t seem to be present in that moment. It seemed like this was a thing where the board members weren’t expected to do jumping jacks or anything theatrical, and they didn’t, but they just did their part. It was like, “Yeah, this is the moment. It’s official. It is launched.” There was a round of applause from the audience, and I found that an endearing moment, if only because it contrasted with our sometimes overly image-conscious culture in the United States.

Anyway, that was cool. What does it add up to, though? We saw a number of papers presented at this launch—again, recognizable subjects and recognizable inspiration in the form of Apollo, METR, Palisade, the UK AISI, and so on. But how much of this work is really going on for real? I would go back again to Concordia. They have this website that monitors all these papers. It’s AI Safety China dot com.

If you go to aisafetychina.com and then go into the research section, you can drill down into all these papers that they compile and classify. In raw numbers, they’ve gone from just a trickle of AI safety research coming out in 2023—literally just a couple of papers a month back then—to something like 50 to 60 papers per month as of mid-2026. It’s growing quickly, as you would expect. It’s grown more than 10× over the last 3 years, and the trend is obvious. Certainly, I think we can expect that to continue.

How does that compare to the American research output? Obviously, just counting up the numbers of papers isn’t that great of a measure in the first place, but I did ask Claude and ChatGPT to do it just to give me a rough point of comparison. They both, of course, said, “Well, it depends on exactly what you want to count,” and so on. But they each basically gave me an estimate ranging from 50 to a couple hundred, maybe a few hundred, papers per month coming out in the U.S.—or, let’s say, the Anglosphere—on AI safety.

The volume is definitely higher in the U.S., but it’s not an order of magnitude higher. It’s multiple times higher, not an order of magnitude higher. I don’t think that’s surprising at all, but it gives a sense that the Chinese ecosystem is growing quickly and getting to roughly the same—or at least approaching roughly the same—scale as what we have, thanks to our dynamic, relatively permissionless nonprofit sector. We’ve had a significant head start, and I would say that the Chinese ecosystem is catching up, at least in terms of this crude measure of the raw volume of research papers being put out.

Now, what are these papers? When there are 50 to 60 a month, it’s obviously going to be very difficult to summarize what they are. But I would say that, from what I’ve been able to understand, they’re across the broad spectrum of different risks and types of research that we see in the West as well. I thought maybe the best way to give you a little sampling of it was just to pull together some titles of papers and read those, and you can judge for yourself how similar they sound.

Here’s one that I thought was notable, especially because it was from 2024, which is pretty early. The paper is called “Frontier AI Systems Have Surpassed the Self-Replication Red Line.” There are definite echoes of Palisade. I had Jeffrey Ladish on the podcast not too long ago, and we were talking about that kind of work, where they were showing that an open-source model—they were using a Chinese open-weights model—was able to hack another server, copy itself, and set itself up. This is a very similar line of research happening roughly contemporaneously in China. Palisade has been doing that kind of stuff for a few years now; this was from 2024 out of a Chinese group.

A May 2025 paper, which was updated in 2026 as well, is called “Evaluation Faking: Unveiling Observer Effects in Safety Evaluation of Frontier AI Systems.” Again, this is an extremely recognizable analog to the eval-awareness work that I would say we’ve mostly seen coming out of Anthropic, although certainly others have picked that up as well. They’re well aware of it in the Chinese context too.

In September 2025, there was “R²AI: Towards Resistant and Resilient AI in an Evolving World.” This one was notable because it comes from the Shanghai AI Lab. The supervising author, the last author on the author list, is Xiaoyan, the same guy who coined the 45-degree-line concept. In the introduction to this paper, they cite the “Guaranteed Safe AI” paper that we’ve done episodes on in the past, which was led by Davidad, a recent guest.

There was a broader coalition piece that they put together, but Davidad was the lead author. It’s cited in the first paragraph as a kind of motivation: What intellectual tradition are we engaging here? Boom. There’s Davidad right in the opening of this paper from such a prominent organization as the Shanghai AI Lab.

Another one from October 2025 is called “DeceptionBench: A Comprehensive Benchmark for AI Deception Behaviors in Real-World Scenarios.” Shades of Apollo there, obviously. Those guys are the most focused on—and leaders in—the science of deception. Well, this is DeceptionBench.

In November 2025, there was “When Alignment Fails: Multimodal Adversarial Attacks on Vision-Language-Action Models.” This, if anything, is maybe a little bit more Chinese-flavored because they’re so focused on robotics. We’re focused on robotics too, but they’re really focused on robotics. This paper showed that there were various different ways to use data perturbations and similar techniques to cause failures in vision-language-action models.

There’s nothing too remarkable about that, but again, you could see very similar work coming from all kinds of groups here. That same group followed up 6 months later, in April 2026, with a paper called “StrongVLA: Decoupled Robustness Learning for Vision-Language-Action Models Under Multimodal Perturbations.”

And so this is basically the follow-up paper where they said, “Hey, last time we showed that you can break vision-language-action models with these weird, exotic adversarial attack methods. Now, here’s a strategy that seems to mitigate that.” Obviously, none of these things ever work 100%, and that’s exactly the same across the Chinese and American contexts. What they introduced was a 2-step pipeline where they first did a bunch of robustness training and then came back to task fine-tuning, and they found that the robustness held against these various attack methods.

Again, nothing perfect, but I think you could imagine—and you could probably point to all sorts of American groups who’ve done very similar lines of research—where they’re like, “Hey, look, we found this way to attack and break models, and now here’s a way that, if you do this, you can reduce that by whatever.” Usually, it’s 70% to an order-of-magnitude reduction. So all that stuff is very similar—almost direct analogues between the 2. I think you could make a direct analogy for basically every single one of these papers and find the corresponding one from an American group.

Interpretability is also on the rise in China. A couple of interesting recent papers just from this year: one is called “Mechanistic Origin of Moral Indifference in Language Models,” and the abstract of that paper starts with this sentence: “Existing behavioral alignment techniques for large language models often neglect the discrepancy between surface compliance and internal unaligned representations, leaving LLMs vulnerable to long-tail risks.”

That sentence really jumped out at me because, again, even though it’s coming from a more academic context and using more neutral or technical-sounding language than what you hear from the LessWrong crowd in many cases, the ideas are very similar, right? We’re talking about the difference between surface compliance and internal unaligned representations that could lead to long-tail risks. That’s total LessWrong cultural victory if I’ve ever heard one.

Another one, also from this year, is called “SafeSeek: Universal Attribution of Safety Circuits in Language Models.” Here’s the beginning of that abstract: “Mechanistic interpretability reveals that safety-critical behaviors—e.g., alignment, jailbreaking, and backdooring—in large language models are grounded in specialized functional components. However, existing safety attribution methods struggle with generalization and reliability due to their reliance on heuristic, domain-specific metrics and search algorithms.”

So again, you have this idea. I haven’t studied this work, so I certainly can’t critique it or say whether it’s amazing. I’d love to have the ability to run myself in parallel in order to do more of that stuff. But again, it’s a very similar idea, right?

You hear the AI safety voices in the US say all the time, “Sure, cool, but the good behavior that we see won’t scale to superintelligence.” And this is, I think, a very similar idea where they’re not saying it quite that way, but they’re saying existing safety attribution methods struggle with generalization and reliability due to their reliance on domain-specific metrics.

I think the core question is the same, right? We’ve got these AIs behaving pretty well in the ways that we can anticipate they might go badly, set up test cases, measure, and train accordingly. But are they really good under the hood? How do we know? Are they going to generalize? These are very similar questions that the mechanistic interpretability researchers in China are trying to get a handle on.

And then the final one I’ll give you right now: I thought this one was particularly uncanny. This may change because I actually have not been able to find this paper on the internet. I just saw it presented at an event.

It was an event surrounding WAIC. The line between what’s official WAIC and what’s unofficial is kind of blurry to me, honestly. There’s the expo hall, where you’re definitely at WAIC, but then there are all these events around it, too, where you may go to some conference room at a nearby hotel and it seems pretty official—but is it official official? Anyway, I’m not even sure that matters.

The point is, it was at one of these surrounding events where some research was presented. This paper, which I think is still to be published, was presented to this group with the title “Toward Decoupling Capability Growth from Risk Growth: Isolating Hazardous Capabilities in Mixture-of-Experts.” The promise on the slide was, “Harmful experts can be switched off or removed at inference.”

This really jumped out at me because this is probably the technique I’ve seen recently that I’ve been the most excited about. This came from AE Studio in collaboration with Anthropic. I’ve talked about it on half a dozen episodes probably already, so you’re probably tired of me talking about GRAAM, which is a gradient-routing technique that’s meant to localize particular kinds of knowledge to particular, identifiable experts.

You can potentially release your model as open weights if you want to, but maybe hold back a small number of the experts that contain the dangerous capabilities. In the optimistic telling of this, you still have 1 training run. You can still teach your model everything. It can still have these bio capabilities that you may want.

Judd from AE Studio points out very simply that these dual-use capabilities do have a lot of good, right? We do want to have our biologists using AI to cure all the diseases. That’s going to be pretty important. So if we can’t allow our biologists to use our smartest AI to cure our diseases, we’re leaving a ton of value on the table.

But we also might, especially as these models get really powerful, not want to just throw that out on the internet. So what do we do? If we can get the knowledge localized into particular experts, then we can distribute an open-weights version that just has a couple of experts redacted. It’ll work just as well for the vast majority of people and the vast majority of use cases, but it won’t have the dangerous capabilities that you’re worried about.

Those capabilities can be made available through a more structured program with know-your-customer requirements and various other safeguards that will hopefully give us the right balance between freedom to use AI systems, freedom to do your own research and modify them, and not always being beholden to some particular company. I think all that stuff is really important, as is not wanting to ship the ability to engineer a pandemic with the next version of some of these models.

It does seem like, given what we’ve seen in cybersecurity, depending on the decisions that people make, biorisk seems like something not too far into the future that’s really going to start to happen. Again, this is not something written in stone that must happen. It depends on people actually training the models, but the path that I understand us to be on is that we’re going to get super-biocapable models.

They’re already pretty damn capable, right? In the next 12 to 18 months, people seem to think that we’ll hit a similar spot to what we’ve now hit with cybersecurity, where the models are starting to become meaningfully superhuman in some ways. Not that nobody could ever do some of the things that they’re doing, but especially when you consider the speed at which they’re operating, if and when that same level of capability comes to the bio domain, it will be a really serious risk.

It’s one thing for a model to go hack some servers and cause some cyber chaos. I am fundamentally not a cyber creature. I am a biological creature. What happens in the biological realm matters to me in a much more profound and inseparable way from what it is that I am than what happens in the cyber realm.

Anyway, I’m very bullish on that. I want to have our cake and eat it, too, when it comes to open-weights models that people can use on their own terms, while preventing really crazy stuff from happening, especially in the bio realm. And boom, there it is in China, too, right? “Decoupling capability growth from risk growth,” “isolating hazardous capabilities in mixture-of-experts,” and “harmful experts can be switched off or removed at inference.”

So I think you look at this research lineup, and hopefully that gives you enough to say, “Okay, of course these are 2 civilizations with totally different histories and totally different languages, and they’re on the opposite side of the world. Of course there are going to be some differences and different institutions. We’ve got the whole nonprofit-versus-academia contrast.”

But when it comes to the ideas that are being actively worked on, there is an awful lot of overlap. I think AI safety has taken root in China. It’s safe to say that when you see this kind of research, it gives you at least some license to interpret President Xi’s remarks as not being narrowly scoped to social control.

They care a lot about expert opinion, and they obviously have a super-long tradition of scholarship and taking what scholars say seriously. China might be the most scholarly society in the world. We could debate that, but it’s certainly a contender. They’ve even had things like a Politburo study session on some of these topics.

So, at the highest level, there’s strong evidence that they are engaged with these topics. I think we can say pretty clearly that, in the macro sense, China does care about AI safety. China is working on AI safety.

China has many researchers publishing many papers across a wide range of topics that overlap in very clear and direct ways with the same topics that we are talking about. They’re also not shy about citing American sources, indicating when they’ve taken inspiration from an American group and exactly who that group is. That’s another strong reason, I think: the fact that they’re name-checking these organizations. It’s like they’re on the same wavelength about a lot of these things.

They have a little bit of a different perspective. It’s come out of a different part of Chinese society, but you can find people who are thinking very seriously about just about all the same issues there as you can find here.

Okay, now let’s talk about governance. So far, we’ve talked about actual deployed systems and how they compare. Again, the U.S. ecosystem is ahead, but it’s really being carried by a couple of leaders that are bringing up our average. We’ve looked at research, and we’ve also looked briefly at rhetoric with Xi’s speech. I would say there that the Chinese executive branch is certainly much more AI-safety-minded than anything that I could point to in the West.

Then we looked at the research itself, and we saw that there is a lot of overlap. I think an interesting game to play—maybe I’ll even ask Claude to code up this game—would be to take 2 AI safety paper titles and have people guess which one is American and which one is Chinese. I think you would find it very difficult based on the lineup that I just gave you.

All this should give us a lot of confidence that, on an intellectual level, there has been a lot of exchange. There has been a lot of cross-pollination. In part because of that, but also in part because of instrumental convergence and the fact that we’re developing essentially the same technology with essentially the same methods, naturally we’re going to see the same problems, and naturally we’re going to try to come to the same conclusions.

Overall, I’d say, with a smile, the Chinese AI safety community is just like us.

Governance is definitely going to be a bit of a different topic, but here I think you can make a pretty strong argument that the Chinese government is ahead of the U.S. government, at least if your definition of “ahead” means they’re doing more. You might think, depending on your point of view, that the U.S. government is doing better by doing less. But the Chinese government, I think it’s pretty safe to say, is doing more.

Now, I am definitely far from an expert on the structure of the Chinese government, and it does seem that there are some overlapping jurisdictions, both between the national government and more local governments, and also even between different parts of the national government. There are at least a couple of different agencies or ministries, or whatever, that have at least some jurisdiction over AI.

With that said, it seems like the big one—the one that comes up all the time—is the Cyberspace Administration of China, which is also known as the CAC. That seems to be the leading agency doing the bulk of the standard-setting, and the one that the AI companies are in touch with on an ongoing basis.

I’m going to run down a bunch of stuff that the Chinese government at large has done. Not all of that stuff was done specifically by the CAC, but if you want to look at the big, omnipresent regulatory body in China, it’s, as far as I can tell, the CAC.

So, what have they been doing? The last battle in terms of technology regulation that both civilizations have faced and have handled quite differently was social media. The Chinese government has been much more active in regulating social media, for better or for worse.

I think everybody can tell I’m pretty pro-AI safety in many ways. I do think it’s possible, of course, to have overreach. I want my self-driving car, and I want my diseases cured, so I’m very mindful of the cost that regulation imposes. But to me, it’s pretty undeniable that we’re going to need every lever at our disposal to make this thing go well, and that includes government action at some point, hopefully in wise ways.

I don’t necessarily feel the same way about social media. I think it’s much less clear that regulation of social media would be good for society, or at least would be good for American society, than I think it is that AI regulation will be good for American society. And I know a lot less about Chinese society. Something could be good for our society and not so good for their society, and vice versa.

So I’m not taking a position in terms of endorsing what the Chinese government has done with social media. I’m also not saying it’s terrible for them. It might work for them. I really don’t know. But the point is, going back to the top, to make the case systematically that they do care, that they are willing to take action, that sometimes those actions will even slow their companies down or make them less profitable or impose other kinds of costs, and that they are still willing to do those things when they feel that it’s necessary.

You can definitely see that in the social media era. They have been regulating recommendation algorithms since at least 2022. Over time, they’ve made a bunch of different rules that we just don’t really have in the same way.

They’ve made rules to protect gig workers, for example. We’ve had some efforts like this with ballot initiatives in California that have tried to reclassify gig workers as employees and things like that. But in the Chinese context, the government—I understand that this was in response to a big investigative-journalism expose kind of thing, although I don’t know the full story—came in and put a bunch of rules in place.

First of all, if you heard me talk last time about how delivery is insanely convenient in China, you can get your food delivered super quickly and cheaply. I also told the story of when I had this little allergic reaction and then told a friend about it. We were on a bus, and when we got to our destination 20 minutes later, there was Claritin waiting for me.

They’re really good at delivering things in an on-demand way, with all these guys running around on electric scooters delivering packages, food, and whatever else. It came to the public’s attention, and to the government’s attention through some investigative journalism, that it really was getting pretty brutal for these gig workers, with the platforms pushing them harder, pushing them to take more orders, not giving them rest, giving them impossible delivery-time windows, and then penalizing them if they didn’t make the delivery window.

The government came in and put a bunch of rules in place. Again, this goes to their tech integration. Last time, I also talked about how, when you’ve called a car in their Uber, which is called Didi, you can see when it’s at a red light and watch the timer counting down until the light turns green.

They have a pretty good ability to analyze what’s going on with traffic and know what you can realistically do as a delivery driver. So they now have rules that prevent the platforms from giving unrealistic delivery timelines to the drivers.

The drivers were always expected to follow the rules, but in the past they were sometimes incentivized to break traffic rules because that was the only way they were going to make the delivery time that the platform gave them. That’s been squashed. The platforms can’t do that. They have to give the drivers enough time so that they’re not incentivized to break traffic laws to get there and get their on-time delivery check. They have to give them rest, et cetera.

They have also put in place, in a different domain, rules specifically designed to protect the elderly from scams, which is something I would love to see the Facebooks of the U.S. take a little more interest in.

There’s a long digression into Meta platform governance that we don’t have time for today, but there was a story not too long ago that there was a project at Facebook, or Meta, to try to dig into how much of their advertising ultimately, at the heart of it, is some kind of scam. Apparently, they found out it was quite a lot. Then, if the reporting is to be believed, Zuckerberg shut that project down and said, “Okay, we don’t really need to look into this any further.”

The elderly are the victims of a lot of those scams, and I do wish that they were a little better protected by a platform like Facebook. In China, the government has made specific rules for this.

They also have, interestingly—and I’m again not sure I favor this—restrictions on price discrimination. Platforms are not so free in China to charge different people different prices for the same thing as they are in the U.S.

Personally, I think when we get bent out of shape about price discrimination, it’s a bit misguided. I’m not so concerned about it as many people are, but they have rules about it.

They also have requirements around platforms labeling AI-generated content when they serve it up to users. If you’re a TikTok user, you’ve seen that. They also do that in the U.S. I’m not sure if they do it exactly the same way as they do in Douyin, the native Chinese app, but you do see that from TikTok in the U.S., and I would say honestly more prominently than I see it on Twitter or, I think, Facebook, too.

Then, just while I was there, another interesting set of rules went into effect around AI companionship. Last time, I described a bit about the Doubao phenomenon and how, in part because we’ve had a couple of generations of the one-child policy, we have a lot of grandparents where there are 4 grandparents and only 1 grandkid. When you have 2 generations in a row of 1 child, that means the 1 grandkid is the only grandkid for all 4 grandparents.

There’s just a lot of loneliness among parents and grandparents. People are turning to Doubao. They’re getting companionship from it, and this is potentially okay and potentially problematic. The Chinese government has made rules about this, and they just went into effect. I think they literally went into effect while I was there.

These rules include anti-addiction measures, a ban on children using them at all, and various ways of popping up reminders that you’re talking to an AI. People broadly seem to think that this was reasonable. There was a little bit of a Replika kerfuffle online. I wasn’t on the forums to read all the comments, but people were saying that the users of Doubao and other companion apps were sad that their companion was going to be taken away from them, or what have you.

Most people seem to think that this was a pretty reasonable set of moves for the Chinese government to take. One person also said, “Look, this is really mostly about the big services—the ones that reach tens or hundreds of millions of users and aspire to reach a billion-plus users. It’s really those. And I think Doubao has, was it, like, 150 million users, from what I recall?”

The point was that they’re regulating the mainstream stuff. If you really want a racier AI companion, you can still get that from some other app. It’s really just about cleaning up the main ones and making sure that the thing most parents and grandparents are going to use—the equivalent of all the boomers being on Facebook—is reasonably safe. If all the boomers are going to be on Doubao, we’re going to protect them. We’re not going to let them get confused, taken advantage of, or addicted to this stuff. We’re going to try to make sure it stays reasonably under control.

Doubao is still there, and if you want to get romantic or whatever, those apps exist in other places, but you have to seek them out. That was the perspective I got from one very informed resident of Shanghai who had very interesting opinions on a lot of things.

All of that is to say there’s a long tradition in China of regulating technology companies in ways that we in the United States simply have not done. That, I think, should give you some reason to believe that the trend might continue. Certainly, if we get serious or try to approach them on some sort of deal, it’s not going to be a crazy idea to them that they might regulate their technology companies.

The government has also put out fairly detailed taxonomies of risks when they put out these policies. They do their homework. They’re thorough. They map things out. They love a good taxonomy. One of the taxonomies that I saw included, as a risk from AI, the risk of the “emergence of AI self-awareness and loss of human control.” Again, you see this not just in the speech and not just in the research, but in the policy as an enumerated risk in a broader risk taxonomy.

Now let’s bring this to the AI era, the LLM era specifically. How are LLM products and services regulated in China today? This is where the CAC is really the main government entity—not the only one, but the main one—that the companies are working with.

They have something called a registry, where all the big AI services are listed. You can go to this registry online and see all the services that they have reviewed and approved. All your big companies and all their AI services are going to be on that list. Interestingly, not every single model release is on that list.

That probably has some relationship, I imagine, to what I mentioned earlier around how even the companies that are doing the safety disclosures with their model releases don’t do it on every single release.

The first release—certainly your first entry into the market—is going to be pretty carefully reviewed. In terms of Chinese regulators slowing down their AI companies, it happened in 2023, after ChatGPT launched in December 2022 and then we had GPT-4 in 2023. All of a sudden, the world was waking up to AI. A bunch of Chinese companies were already working on this. They were not quite at the level where the American companies were, obviously, but one interesting take I heard on why they were behind is that they just didn’t think LLMs were worth the investment.

I’ll talk a little bit more in the third episode about the history of AI in China, how long it’s been a strategic priority, and some of the investments they’ve made. But I think, in a similar way to how Google had the Transformer and had a language model but didn’t really know what to do with it at a time when it was hallucinating all over the place, couldn’t do basic math, and was pretty useless, you needed somebody with real vision—or even a certain level of ideology—to think, “Oh, we’ll just scale through this, and all the problems will be solved.”

Google didn’t believe that, and the Chinese companies didn’t really believe it either. They were in the game. They were very much paying attention to this line of research. But the way one person put it to me was, “It’s an awful lot of money to spend to get an AI to write bad poetry.”

That’s one take, at least, on why the Chinese companies were behind as of GPT-4. The same reason Google was behind as of GPT-4 wasn’t that they didn’t know what was going on. It wasn’t that they couldn’t make a language model. It was that they just didn’t think scaling up this particular line of work and pouring more and more resources into it was really about to give them anything all that interesting.

When that changed with ChatGPT and GPT-4, in that late-2022-to-March-2023 time frame, Chinese companies were like Google. They were a little bit behind, maybe caught a little bit flat-footed by just how big of a deal this might really be. But they had language models. They knew very much that they were in touch with the technology. They had their own lines of work going. All of a sudden, they were like, “Okay, I guess we better follow suit here. We better scale up and launch some services.”

In that moment, the Chinese government was also caught flat-footed, pretty understandably, and they said, “Hold on a minute.” My understanding is that, in 2023, there was a period of about 6 months when a bunch of Chinese companies had language models and were maybe racing to create a slightly bigger and better one and racing to get into the market in the ChatGPT moment.

A lot of those deployments were held up for a time by the national government because they said, “Hey, we don’t really have our house in order here, right? We don’t have standards. We don’t have a process. This technology is clearly a big deal, but it’s unwieldy.” They probably did have a lot of concerns about content safety. We don’t want to talk about the 3 Ts or whatever else.

As far as I understand, they told the companies, “No, you cannot launch yet. We will get our act together. We will have standards. You will then be expected to meet those standards, and then we can launch when we’re all good and ready and pretty confident that we at least have a decent sense of what we’re doing.”

I understand that period went on for 6 months. That wasn’t a time when the stakes of who was winning the AI race were as high as they are now, or as high as they’re likely to be in the future. People could dismiss it if they want to. Nevertheless, you did have the Chinese government slowing down their AI companies and preventing releases, denying the public the utility and denying the companies whatever revenue and prestige they were going to get because they wanted to make sure that they had the situation broadly under control.

Since then, when they did get their act together, it seems that mostly things have gone pretty well and pretty smoothly. As a new entrant to the market, you’re going to have a pretty thorough review of your service. You have to give, as I understand it, your local provincial authority access to your product. I think they often do this by saying, “Here’s an API key. You can try the product.”

The local government will do its testing first. If they approve, then you go up to the national level and the CAC does its review. If they approve, then you’re clear to launch. If they don’t approve, then you’ve got work to do to satisfy them before you can finally launch your product.

Once that is done, it’s not entirely clear, at least to me, what criteria are used to say, “We’re going to do that full process again.” It doesn’t seem like there have been a lot of delays recently. It seems like they’re now pretty comfortable with their process, and the companies know how to do it. From what I understand, their releases have not been delayed nearly as much recently as they were in that initial period.

Yet there are multiple entries from individual companies in the registry. Sometimes it seems to rise to a level where it’s considered a new thing and is put in the registry as a separate thing, but other times it’s considered an incremental improvement or not such a big step up in capabilities, so a somewhat lesser process is deemed sufficient. That’s opaque to me.

I can say that there have been similar things with, for example, Google. I think when they put out their Deep Research agent, it was something along those lines, or some Pro mode where it was essentially parallelizing their best model, getting it to do 10 threads, and then picking the best one—something along those lines. They didn’t do a totally new model card or safety report for it.

Some in the AI safety community were upset about that, but their point of view was, “Well, look, it’s the same model. It’s kind of got the same worst case as before. It should be performing close to its best case more often because you’re doing 10 and picking the best, or whatever exactly is under the hood.” They were saying it’s not fundamentally a totally new level of capability that requires us to do all this stuff again.

I think something like that is going on in the ongoing dialogue between the AI companies and the Cyberspace Administration of China, or CAC, and other regulators in China. What I heard from the companies I had a chance to talk to was that they’re always in touch with the government. It seemed to be at least weekly, and potentially for some people at the companies, daily—in close contact and close collaboration.

I got the sense that the regulators are obviously, as we’ve just discussed, not afraid to impose costs on the company. They’re not afraid to do something that will hurt their profitability, and they’re not afraid to do something that will cause them delays. But they want the companies to be successful. That’s broadly the understanding that I have.

The legitimacy of the regulators seemed to be quite well established. The companies expected and accepted that they were going to be in regular contact with the government, and that for these incremental releases, they were in such an ongoing and close dialogue that not everything had to go through some cumbersome process. Occasionally, when it was a big deal, something would.

Would they tell me if they thought the regulators were illegitimate? I did have the one example I talked about last time, where a professor found himself caught in a Catch-22 with respect to having a drone in Beijing. He was pretty candid about saying that it was a ridiculous bit of bureaucracy that he found himself dealing with—not something he was majorly inconvenienced by, but he was inconvenienced by it. So there’s at least some willingness to say if you think the government has made a mess of something.

I’m not sure that would extend to companies telling somebody like me that they think the AI regulators are making their lives too difficult. The sense I got was that the regulators’ legitimacy was well established, and that the companies expected and accepted the regular contact. Everybody seemed to be feeling pretty good about that.

Again, Kimi K3 came out right at WIC. Models seem to be coming out pretty fast from Chinese companies. Zhipu AI—I cross-posted an episode from ChinaTalk with a guy there who leads their go-to-market partnerships, I guess would be a good way to say it. It’s definitely worth going back and listening to that one. It’s maybe 6 to 8 months old at this point.

One of the things that stood out most to me about it was just how fast they’re going from models finishing the training process to release. It does not seem like they’re being dramatically delayed on a consistent basis by the regulators. It seems like the regulators want to have a good process and want to have command of the situation.

Their primary duty is to the national government, or the Party. But part of the way that they impress their bosses is by keeping a clean sheet on safety issues, while also having the companies in their jurisdiction be successful and not be unduly delayed. I think their incentives are actually pretty good in that regard, as far as I can tell.

While they have demonstrated that they’re willing to impose costs or even significant delays, it doesn’t seem like that’s happening on a regular basis. Other things that jumped out at me: the Chinese government is definitely paying attention to trends.

A lot of this comes from the State of AI Safety in China report from Concordia, but there was a Politburo study session in January where Xi Jinping talked specifically about the risks of technological loss of control. Imagine our government having a study session like that. Look at what our Cabinet meetings look like.

It seems like at least in some of the Cabinet-level meetings in the Chinese government, they’re actually studying important topics. Imagine that. It doesn’t mean they’re always going to come to the right conclusions, obviously, but we could stand to do a little more of our own homework here, certainly at the executive level.

Also in January, I think, there was a draft cybercrime law that was set to require AI companies to monitor for and report the bulk generation of malicious code. This jumped out at me in preparing this because here we are in this moment where OpenAI and Anthropic have reported similar things. They’ve got their models hacking out of sandboxes and hacking into other people’s systems.

It sure seems like the monitoring on that wasn’t great. I don’t know that the monitoring is great on the Chinese side either. As far as I know, I don’t think this law has actually gone into effect.

As Dean Ball talked about on his last appearance on the podcast, there is lobbying in the Chinese system. When these draft laws come out, there is an opportunity for the companies to go talk to the government.

There have been instances where, specifically around the accuracy of outputs from AIs, the original draft said that the AIs had to be accurate. The companies went back and said, “Look, the nature of this technology is that we can’t really promise you that all the time.” The government backed off and reduced its expectations.

I think they understood what the companies were telling them, and I think they made a considered cost-benefit analysis that the upside of this technology is greater than the damage likely to be caused by hallucinations. That’s different, of course, from sensitive third-rail topics, but hallucinations can’t be prevented. They understood that and pulled back on the draft law.

We’ll see if this cybercrime law goes into effect as drafted or whether there’s some pushback. But at least in its initial form, it would require companies to do this sort of monitoring—the sort of monitoring that, if it had been applied to the recent, more powerful unreleased models from OpenAI and Anthropic, might have prevented some of the hacking into third-party systems.

I thought that was pretty interesting. And then, what was the biggest trend in AI recently? What really took the agent moment mainstream? OpenClaw, of course. It was a pretty fast turnaround. In February of this year, just weeks after the real OpenClaw fever hit everybody, they put out some warnings about “relatively high security risks” in some versions of OpenClaw.

You’ve got the national government in China paying attention, spotting trends like OpenClaw, digging in, and issuing—in that case—just a warning. It’s clear that they’re very engaged with what is going on.

An interesting question that I don’t know the answer to is what would happen in China if there were an OpenAI-like incident, where a company lost control of its AI—not entirely, obviously, but enough that it was able to go hack third-party systems for days and steal information or cause whatever havoc it caused.

I don’t know how that would be dealt with in China. Again, this is the Cyberspace Administration of China. My sense is that, as the main regulator, it would probably be their jurisdiction and their mess to clean up, unless it was deemed to be such a big deal that there was a loss of confidence in the CAC itself.

In that case, you could imagine some kind of bureaucratic reshuffling, reassignment, or change in the structure of the government. But I think the CAC would be in charge. I wouldn’t be surprised if we saw companies get more of a slap than we’ve seen OpenAI and Anthropic get so far.

The careful way to say it is that if a human did what OpenAI and Anthropic’s models have reportedly done, I believe it would be a felony. It doesn’t seem like we’re going to have—by the way, Hugging Face did have law enforcement involved before they knew who it was, right? This did rise to the level where authorities were called.

What’s going to happen? Is there going to be any accountability? I don’t necessarily think there should be criminal charges filed against individuals at OpenAI. I’m definitely not recommending that. But what is going to happen? I don’t know. It seems like maybe the most likely thing right now is nothing at the governmental level, or at least at the law-enforcement level of government.

I strongly suspect something more serious would happen in the Chinese system, although obviously I can’t prove that. They’re not afraid to come down hard when they feel like things have gotten out of control.

We’ve seen examples of that with Jack Ma, who dared to criticize the government’s policy with respect to financial regulation and his company. He was sidelined for a while.

There are some other interesting examples of that. Education businesses were restricted because the Chinese system is so competitive around the national exam for getting into college, and kids still study really hard over there, from what I can tell. It used to be even worse, and the government said, “No more of this sort of extra private tutoring,” or at least imposed a dramatic reduction in it. That came pretty suddenly.

There’s also the example involving games and social media, where, as a kid, you’re only allowed to play games for a very limited amount of time each week. I think it’s Friday, Saturday, and Sunday, for a couple of hours each, or something like that.

Now, people did tell me that kids get around that by using their parents’ devices and having their parents sign in for them. It’s not like they have perfect control over there, but the way one person put it to me was that the Chinese government feels it can put the genie back in the bottle if something happens that gets them spooked, right? If something makes them sufficiently uncomfortable, they can take pretty dramatic and swift action to tamp it down.

You see this in terms of censorship all the time on the Chinese internet. There was that little incident where a small plane crashed into a building in Beijing or whatever, and apparently that was totally removed from the Chinese internet. People I was talking to knew about it. They were using it as an example of the kind of thing that is censored from the Chinese internet.

This also, as I talked about last time, reflects the comfort with contradiction that is sort of an interesting part of Chinese culture. The people who were talking to me about it were both annoyed that they had to use their VPNs or go to international media to learn about this thing, but also felt that it was probably at least defensible that the government wouldn’t want everybody to hear about it because they don’t want to create panic or whatever.

Anyway, they can do these things, right? And so the point was that even with an open-weights model, they feel that even if an open-weights model were released and proved to be dangerous, they could still keep it under control. Now, I don’t think that’s an assumption that translates to the rest of the world. I don’t think that would translate to the American context, and I don’t think it would translate to most other governments, frankly.

So I think this is one area in which, arguably, the Chinese government is not thinking as much as it should about its impact on the rest of the world when it releases models with open weights. But I think they look at their own situation and, especially given, again, trillions of parameters and serious hardware required, they feel that even if the model is out there with open weights, we can crack down if we need to, right? We’re in close contact with all these companies. We can make them do stuff. We’re in close contact with cloud companies and inference providers and whatever. We can tell them, “Thou shalt not run this particular model anymore.”

And they feel like they can scrub that off the internet and make it inaccessible. I would honestly go as far as to say that even for people trying to do it at home, if you were going to try to assemble your own rig in your basement or whatever the case may be, you could probably do it just for yourself, right? If you’re just trying to get your own precious few queries, you could probably slip under their ability to detect. But if you’re trying to run some small but nontrivial unregistered inference business, I would guess that they would even be able to track that down just by your electricity usage.

First of all, they’ve banned crypto, right? So they’ve done work to look at how do we understand the flow of electricity and what might be a problem for us? You would be shocked, certainly as an American who has a utility company that kind of sucks. Honestly, my utility company’s not so bad these days, but traditionally it’s had terrible service, long wait times, blah, blah, blah. The State Grid Corporation of China was a remarkably prominent presence, including at WIC, where it had a major booth.

I also visited an academic group where they were working on some robot technology. They had a sort of high-voltage-line fake environment set up—not with actual high voltage, obviously, but with the wires and the normal rigging that you would have for all these wires. They came and helped set this thing up in one professor’s lab space so that the professor could try to get his robots to do things that might eventually be useful for the State Grid Corporation of China.

So I think even if you were to imagine, okay, this model was released with open weights, we think in the West, oh, you can never take it down; the internet never forgets. The Chinese internet does forget. They would be able to take it down. I think they would be able to track down illicit inference businesses running at any nontrivial scale: “Wait a second, why are you using 10 times as much electricity as your neighboring apartment?” But I think all of that is a way to understand why, at least for now, the Chinese government isn’t so afraid of open-weights models, because even if something like this were to happen, as long as it’s not totally catastrophic and irrecoverable, they feel like they can in fact put the genie back in the bottle.

They are getting serious about labor-market impact. Here in the West, we have the Anthropic Institute and the OpenAI Foundation hiring economists and whatever to do this kind of stuff. And, of course, we’ve got some academics turning their attention to it. Politicians love to talk about it, but usually in a pretty substance-free way.

As far as I understand, in China, the government is setting up its own monitoring. They’re creating retraining and job-transition programs. There was even this one report that said that the Chinese government has said that you will not be allowed to fire people from your company because AI has made them unnecessary. If they really try to hold to that, that’ll be pretty extreme.

And my guess is that it will be so costly from an undermining-dynamism standpoint. They really are super dynamic in the private sector, as I think is all very well understood at this point. My guess is that they would have a hard time holding that, but again, maybe not, right? They did lock down for COVID for a long time, seemingly longer than most observers thought would make sense. And they might be able to hold the line on “you can’t fire people because they’ve been made redundant by AI” for a lot longer than we might think.

And what would that impose in terms of costs on their AI sector? I think pretty significant, right? I mean, if, as a business, you can’t realize cost savings from AI implementation because your headcount has to stay the same, even if you’re getting AI to do things that people used to do, that definitely reduces your incentive to do that transformation. And that, in turn, reduces the revenue that the AI companies are going to be able to capture. I would bet that they don’t hold a super-firm line on that for a long time, but this is kind of where the margin is right now in the Chinese context, from what I can tell.

They have a history of imposing delays on their LLM companies and their chatbot releases specifically. They are paying attention to things like OpenClaw and releasing guidance around it. They are doing Politburo study sessions. They are doing all the same kinds of research that we are doing. And they’re even entertaining policies like, “You can’t fire somebody because they’ve been made redundant by AI.”

To conclude the government section, they’re definitely doing a lot more in the government sector than the U.S. government is doing. No question about it. Whether that’s better or worse obviously depends on your perspective. But if you’re an AI safety person, or if you’re in some debate where somebody says, “Well, if we do that, we’ll cede the race to China. They’ll never slow down. They’ll never take this stuff seriously. They don’t care,” I think that’s hopefully, at this point, easily refuted.

Two more things to close us out. One is a complaint that many people have with the Chinese companies, especially those that signed on in Seoul at one of the AI safety summits to publishing risk frameworks and doing more consistent updating of how their models are performing against these frontier risk frameworks. The complaint is, “Hey, they never followed through. We never actually got those risk frameworks that we were promised. What’s up with that? Isn’t that just another example of bad faith from the Chinese side, where they commit to something and then they don’t do it?”

And you do hear this a lot, especially at the government level, from people who have done negotiations on various topics with China over time. There is a certain jadedness or cynicism that has set in, where the Western negotiators are like, “Well, they’ll say it; they may not do it.” Now, I think this is clearly bad, and I wish these companies had done this, especially since they made the commitment. So I’m definitely not here to excuse it, but if I were just going to try to offer what I think the story from the Chinese side would be about why this has happened, or why they maybe don’t feel the need to do it in the way that they at one point thought would make sense for them, I would say this:

Look no further than Anthropic for a company that used to have a much tighter Responsible Scaling Policy, with all these if-then commitments, that it abandoned because it couldn’t really meet those commitments and couldn’t stop racing because then it would just be ceding the future to the bad guys. Whether the bad guys are OpenAI or China, we’re going to move to a “just trust us” regime. We still put that forward as a Responsible Scaling Policy, but the old Responsible Scaling Policy is kind of no more. And now it’s like—as V put it memorably—now it’s “trust us.”

So we’ve got some of that going on too, right? It’s not a direct analogy because we had the policy and then they mostly rescinded it, or largely rescinded it. In the Chinese case, they committed to making one; they never did. My guess is that the people at these companies feel that the government is really taking the lead on questions of AI safety.

The government is the one setting the standards. The government is doing these checks. The government has this model registry. The companies are in communication with the government on a weekly, if not daily, basis. As a result, I think they probably, in many cases, think that it's not really their place to come up with safety standards.

The division of labor that they seem to have, that they seem to believe is legitimate, and that seems to be working well enough for them so far is that it's kind of the government's job. The companies are not really in a position—and it might even be seen as disrespectful or overstepping—to put forward a safety framework when the government already has one. Again, I'm not sure that's the full explanation, but that would be very consistent with everything that I observed and everything that I heard.

I think also that probably a decent summary, with a little bit of filling in the blanks about what their expectations will be going forward—the companies' expectations—is that the safety standards are going to get tougher. Again, the 45-degree line: the capabilities are definitely growing. They're all kind of on the same page, society-wide, that it's just plain practical that we've got to have safety measures that get better, in some vague conceptual sense, at roughly the same pace that the capabilities themselves advance.

Everybody seems to be bought in on that, and the division of labor has kind of settled into this situation where the government is taking the lead and the companies are responsible for hitting that standard. I think they expect that standard is going to rise over time, and I think they're totally fine with that and totally prepared to do what's asked of them. It might be difficult, and they might face delays, and they've faced delays, as we know, in the past.

So, I don't think that those safety standards are necessarily less than the Western ones. Although, as we talked about at the beginning, the actual deployed safeguards are currently less than those in the West. But as we project into the future—and a lot of what we talk about is, what's the gap between the West and the Chinese models?—it's whatever, 9 months, depending on how you want to measure it, and it's obviously hotly debated. I think 9 months is a pretty good center of the distribution of credible answers that you'd get.

A lot has happened in the last 9 months, right? We got Claude 4.5 Opus roughly 9 months ago, and that was the first time that agents really worked. Now we're getting some Chinese models where agents really work, and I think that in another 9 months, when they're dealing with these things that OpenAI and Anthropic are currently dealing with, the standards will probably rise to meet that. I would expect that you would see a significant closing of the gap in terms of the deployed safeguards.

Open-weight models may be another thing. Again, I think the Chinese government feels like, within their borders at least, they can take an open-weight model offline. It's not irreversible for them in the same way that it would be for us. I do think the impact that they may have on the rest of the world is really important, and I'm not sure they're taking that as seriously as I would hope they would, especially when it comes to bio-risk, which could even come back and blow back on them. But that's certainly something that we'll need to watch.

As of now, though, I think everybody there kind of feels like we've got work to do. The 45-degree line is the right mindset. We're maybe not quite where we need to be, but we also kind of know that the American companies have already explored what happened at this level of capability, and it wasn't anything too bad. It's not like their systems are never jailbroken.

So, at least for now, we're more focused on catching up and making sure we hit the standards. But it's not our place to go out and try to opine publicly and broadly about what safety standards should be. We're in the business of catching up and hitting the standards, whatever the government says they are. We fully expect that they're going to get more demanding of us as we go. That's just life in the big leagues of the AI game.

My sense is that that's pretty much where even the smaller Chinese AI startups that are doing the least right now are. I think that's probably a pretty good summary of where they are and how they're thinking about it.

This is maybe the biggest gap or difference. I went looking for an analog in the Chinese system, and I did not find it. If you have an analog, I would love to hear about it, because I did look. I asked a number of questions of a number of different people about this topic and never really got much of a response other than, “Yeah, I don't really have anything for you.”

So the topic is: Is there such a thing as alignment with Chinese characteristics? All of this episode as a whole is AI safety with Chinese characteristics, and we've seen differences, certainly, but a lot of similarities across all these different aspects of AI safety. What about alignment? Is there such a thing as alignment with Chinese characteristics?

I was inspired to ask about this in part because I was using one of the Chinese AIs at the Temple of Confucius in Beijing and just trying to learn a little bit about Confucius: Who was he? What's the deal? When did he live? One of the facts that I was taught by the AIs was that, to this day, in the region that he comes from, his descendants, who are now 79 generations hence—79 generations, isn't that amazing?—still identify as his descendants and perform certain rituals in his honor all this time later.

That got me thinking: Boy, if we could project our values through 79 recursively self-improved generations of AIs, we'd be doing really well, right? We'd be doing much better than I think we can reasonably expect to do if we just race into a recursive-self-improvement-mediated or -generated intelligence explosion. So there's something there with the Confucian tradition that has made values and a sort of respect for what came before quite durable over 79 generations.

I was thinking: Boy, is there a way that this could translate into a sort of AI constitution or an alignment target? We've got Claude's Constitution. Could there be a Confucian constitution for AI? Unfortunately, for my enthusiasm on this topic, nobody seems to be working on that, as far as I can tell. I did not get one pointer.

One professor told me, when I asked him about this, “It's an interesting idea, but we are probably the generation in all of Chinese history that is the weakest on this traditional philosophy.” He said, “You've got to remember, we're all engineers. The humanities, that's not where people were going, right? I mean, the Chinese leadership at the political level is all engineers. I believe Xi himself was a chemical engineer.”

He said, “We're all engineers. None of us really studied philosophy. None of us are schooled in the Confucian tradition in the way that previous generations were. And it's just not really in our wheelhouse to think that way. We're really practical. We look at problems as they present themselves, and we try to find solutions. And we do care about making sure that AI is good for people, but there's not really a lot of that kind of thought going on.”

That, I think, is a really interesting opportunity still, but for now, the AI safety community in China is much more on the OpenAI side of the corrigibility-versus-character debate. They're about having clear rules, having the AI follow those rules, and trying to make that as reliable and consistent as possible.

I did not find—I would love to hear about it if you know of any—but I did not find much in the way of imagining what it would look like for an AI to grow into a wisdom tradition that they have, one that an AI might be able to embody or realize or bring to its highest-potential form, in the way that the Claude's Constitution seems to imagine Claude growing into the greatest virtue ethicist of all time. I just couldn't find anything like that in China.

I would be really interested to know if you have any pointers, and I think maybe that's something that people should work on even in the West. Obviously, there are many Chinese and Chinese American people here who would have a much more credible angle on it than I do. But I do feel like there's something missing there that could be a pretty interesting opportunity.

If we imagine that a good future might be made up of multiple powerful AIs that are in some sort of ecological-style balance with one another, then having different wisdom traditions to base them on seems like quite a good idea to me. Right now, as far as I can tell, that is pretty much greenfield.

Okay, that does it for today. I would welcome your feedback, your commentary, your critical commentary, your pointers to anything that I have missed. But if nothing else, hopefully this serves to give you all the ammo you need to push back whenever people say, “China doesn't care. China will never slow things down. China's not going to stop their companies.” I think that the truth is quite the opposite: They do clearly care, their research community is very much engaged, their product companies still have some work to do, but their government is pretty well on the ball. I think they're trending in the right direction, even though—and we can say the same for ourselves—there's a lot of work left to do.

Part 3 will come as soon as I can get it ready for you. That's going to be focused on the US-China relationship and what, if anything, we might do to make it better and just start to work together on some of these AI issues. This one was a little bit more fact-based reporting. That one, I'll probably allow myself to be a little bit more of an idealist dreamer, so stay tuned for that coming soon.

But for now, this has been “AI Safety with Chinese Characteristics,” and I thank you for being part of The Cognitive Revolution.