[BidClub_]
The Cognitive Revolution · · 137 min

Nathan Goes to China – Part 2: AI Safety with Chinese Characteristics

Nathan Labenz

YouTube
TL;DR
  • China’s deployed AI safeguards currently trail America’s, but the headline gap is largely an OpenAI-Anthropic effect rather than a civilization-wide divide. Nathan estimates roughly 10-12 near-frontier developers in each country; remove “Openthropic,” and the remaining US companies overlap much more closely with Chinese peers, with Gemini only modestly ahead. Concordia AI’s evaluations tell the same story: mostly American proprietary models sit near or above the “45-degree line,” while mostly Chinese open-weight models tend to fall below it.

  • The claim that China “doesn’t care” and “will never slow down” is contradicted by both policy and precedent. Beijing reportedly held up many domestic chatbot launches for roughly six months in 2023 while it created standards, and today the CAC can require local and national reviews before a service enters the registry. China has also imposed costly rules on recommendation algorithms, gig platforms, children’s gaming and AI companions—evidence that the state will subordinate company growth when it believes intervention is necessary.

  • Chinese AI safety is scaling rapidly through universities and companies rather than America’s permissionless nonprofit ecosystem. Concordia counts growth from only a few papers per month in 2023 to roughly 50-60 per month by mid-2026, versus an estimated 50 to a few hundred across the US or Anglosphere. The work spans self-replication, evaluation faking, deception, mechanistic interpretability, multimodal attacks and hazardous-capability isolation; Nathan’s conclusion is that “AI safety has taken root in China.”

  • The central Chinese risk conversation has moved beyond censorship to agents that can act in digital and eventually physical systems. One major technology company insisted it “really do[es] care about catastrophic risks like CBRN risks” and said an agent produces a daily report on American AI-safety discourse. Xi Jinping’s WIC language likewise called for faster safeguards against “loss of control,” prevention of malicious use and keeping AI “under human control”—more safety-forward rhetoric than Nathan can point to from a prominent US official.

  • The largest unresolved fault line is open weights: China regulates services and believes it can put the genie back in the bottle, while Western safety analysis emphasizes irreversible global release. Chinese interlocutors argued that a 2.88-trillion-parameter K3 cannot simply be run on a distressed person’s laptop; serious inference requires substantial hardware and infrastructure. That logic may hold inside China’s enforcement perimeter, but Nathan worries it understates external bio and cyber risk once weights reach jurisdictions Beijing cannot control.

  • The regulatory gap could narrow as Chinese capabilities catch up, because companies appear to expect standards to rise alongside models. Nathan puts the capability lag’s credible center near nine months and notes that Claude 4.5 Opus marked the point when agents really began working. Chinese labs are now releasing models where agents really work as well. His conditional forecast is that, as they encounter the failures now confronting OpenAI and Anthropic, Chinese regulators will tighten requirements and deployed safeguards will “significantly” converge.

  • China’s biggest conceptual absence may be alignment by character rather than compliance by rule. Chinese AIs told Nathan that Confucius’s descendants, reportedly 79 generations later, still identify as his descendants and perform rituals in his honor, inspiring him to ask whether a Confucian constitution could carry values through recursive AI generations. Yet researchers told him, “We’re all engineers”; the current ecosystem emphasizes explicit rules and reliable obedience, leaving wisdom-tradition-based alignment “pretty much greenfield.”

Digest · the substance, structured for research

1. The “but China” endpoint rests on a false premise

  • Nathan’s purpose is deliberately narrower than forecasting a treaty: describe Chinese AI safety “on its own terms” and retire the reflex that any US obligation automatically hands Beijing the race. China cares about safety alongside other priorities, has slowed companies when its definition of safety demanded it and possesses a government demonstrably willing to act.

  • His strongest formulation is also his bluntest: racing “full speed ahead into recursive self-improvement” because China supposedly cannot regulate comes “from a position of ignorance.” That does not make cooperation easy or guarantee symmetrical rules; it removes an imagined impossibility that has prematurely ended too many Western policy debates.

  • The reporting follows a Chatham House approach, so private observations are unattributed while public papers, reports and speeches are named. Nathan also discloses that he paid for flights and hotels himself, accepting roughly half a dozen meals during two weeks in China.

2. America leads on deployed safeguards because two companies carry it

  • Nathan does not bury the unfavorable comparison: Chinese companies and models currently provide weaker safeguards against misuse than American products, with poorer follow-through on model cards and published safety evaluations. The difference is material, especially when weighted by what consumers actually use.

  • Yet the American average is dominated by OpenAI and Anthropic—his emerging “Openthropic” duopoly—which lead simultaneously in capability and jailbreak resistance. Gemini is somewhat ahead of the broader pack, while Grok and other American models are easier to break; without the top two, America’s moral high ground becomes “a lot more muddled.”

  • Using a permissive near-frontier definition, Nathan counts roughly 10-12 model makers in each country, more than most observers would call truly frontier. His recent jailbreak discussion produced a clean ordering: OpenAI and Anthropic are hardest to jailbreak, Gemini and Grok easier, and Chinese models easier still.

3. The 45-degree line is China’s aspirational safety doctrine

  • A Shanghai AI Lab leader introduced the “45-degree line” at WIC in 2024: capability and safety should rise together, like a slope of one. Safety must keep pace with each new capability level, but building extreme defenses against capabilities that do not yet exist may be wasteful or counterproductive.

  • Concordia AI’s public evaluations at aisafetychina.com turn that metaphor into composite capability-versus-safety plots. Mostly American proprietary API models appear around or above the line; mostly Chinese open-weight models cluster lower and often below it. Nathan stresses that the exact composite metrics deserve scrutiny, but the directional result is consistent.

  • The candor matters: a China-based organization publishes these unfavorable findings in Chinese and English rather than concealing them. Nathan reads that as evidence that the domestic community can acknowledge the shortfall; the 45-degree line is accepted as a principle, but “still a bit aspirational” in implementation.

4. China regulates AI services more naturally than model weights

  • A Chinese consumer app is often a multi-part system: the underlying model may begin answering a sensitive question before a separate monitor deletes the response and substitutes a refusal. Evaluating bare weights therefore measures a different object from evaluating the API or first-party service through which most Chinese users encounter the model.

  • Nathan grants the Western worst-case argument: once weights are public, anyone with enough resources can use them outside the original deployment safeguards, so the model in isolation must be tested. Chinese thinking instead asks who will actually run it, through which infrastructure and inside what regulated service—an operational question Western discussion may dismiss too quickly.

  • The recurring Chinese example was K3 at 2.88 trillion parameters: “this is not the kind of thing that you can run on your laptop.” A random person in psychological distress cannot casually download it to a phone; substantial hardware is required, and the likely users are businesses wrapping it in services that can be regulated.

  • Neither perspective eliminates the other. Service-level analysis better describes ordinary exposure, while bare-weight analysis captures globally available tail risk; Nathan identifies this as one of the clearest conceptual disconnects between the two safety communities.

5. Disclosure is patchy, with incumbents more cautious than startups

  • Concordia reviewed 10 major Chinese companies and found that five had recently published some form of safety evaluation with a model release. The other five had not, and even the participating companies did not evaluate every release, leaving Chinese practice well short of consistent OpenAI- or Anthropic-style model cards.

  • Nathan’s impression is that Alibaba-, Ant-, Tencent- and perhaps ByteDance-scale incumbents do more safety work than younger AGI-chasing startups. Established firms have profitable businesses, regulatory standing and institutional systems to protect; a fintech platform handling money, for example, has an immediate reason to test whether AI can “run amok.”

  • Startups often reason more like Meta around Llama 2 and Llama 3: if they are roughly a year behind and American systems already exposed that capability level without a reported catastrophe, catching up may add little marginal danger. They reason that if something really bad was going to happen at the capability level they were chasing, it probably already would have happened.

  • Nathan leaves the conclusion conditional: newly reported frontier incidents may break the assumption that earlier capability levels were safely explored. If evidence of real cyber or agent harm accumulates, Chinese startups’ relaxed catch-up logic “very well might” change.

6. Cross-border engagement has produced visible intellectual convergence

  • At WIC and surrounding Track 2 forums, Nathan saw American advocates who publicly call cooperation with China essential doing the closed-door work their position implies. Some meetings excluded him because a podcast credential made participants less comfortable, but the public evidence of cross-pollination was unmistakable.

  • Chinese researchers repeatedly cited Western organizations and concepts rather than disguising their origin. There are cynical voices who view AI safety as a Western scheme to slow China, just as Western discourse has its own cynical faction, but Nathan met nobody who personally advanced that view.

  • One major Chinese technology company assembled its CMO, communications chief, general counsel and AI-security leader, then “almost pounded the table” that it cared about catastrophic and CBRN risk. It also runs an agent that surveys American AI-safety discourse every day and produces a daily internal report—an intelligence loop Nathan sees little evidence of in reverse.

7. Agents, not the three T’s, now dominate the safety agenda

  • Tibet, Taiwan and Tiananmen remain sensitive content areas, and Chinese services must enforce the government’s rules around them. But Nathan rejects the inference that censorship is the only safety Chinese institutions recognize; many companies now feel they have the content-safety requirement largely figured out.

  • The live concern everywhere was “agents, agents, agents”: systems have moved from answering questions to autonomously taking actions. That changes the safety premise from managing speech to controlling consequential behavior in software, networks and commercial workflows.

  • China’s unusually strong robotics emphasis extends the concern beyond the digital world. Researchers and companies expect commercially useful embodied agents to enter physical environments, making the practical question unavoidable: if AI acts independently, how can institutions ensure those actions remain beneficial and “under control”?

8. Academia substitutes for the nonprofit ecosystem China never developed

  • America’s permissionless civil society let speculative AI-safety ideas survive when they were fringe: a small group only had to persuade one wealthy patron, not win government approval. Nathan calls that a major US advantage that produced today’s safety community before academia or government treated its predictions as credible.

  • Chinese nonprofits generally operate within narrower, legible service missions and avoid political activism, so universities—and secondarily companies and university-industry collaborations—produce most safety research. Chinese academia moved faster than American academia, but slower than the US nonprofit sector that enjoyed the head start.

  • This institutional origin changes the people and register. Chinese researchers tend to be career academics: creative but conventional, oriented toward reliability and child protection, and less inclined toward LessWrong aesthetics or highly speculative tail-risk narratives.

  • A prolific professor told Nathan that colleagues do not trade P(doom) estimates over lunch and generally avoid AI 2027 because its US-China politics make discussion uncomfortable. His surprised response—essentially, “Do Americans do that?”—captures the cultural distance even when both sides study similar technical failures.

9. Xi’s language makes “China will never care” hard to sustain

  • At WIC, Xi Jinping asked: “How should humans coexist with machines that think? How can safety be protected when algorithms participate in decisions? How can governance keep pace when technology challenges ethics?” Nathan emphasizes that this was a prepared opening keynote, with a carefully prepared English version available, though he noted that multiple translations exist.

  • Xi’s second formulation tracks the 45-degree concept: “The faster AI advances, the more firmly its direction must be anchored toward human benefit,” with governance calibrated more precisely and safeguards against loss of control improving more rapidly.

  • He closed by calling for legal and technical systems for monitoring, early warning and emergency response; prevention of misuse and malicious use; and keeping AI “under human control.” Nathan allows that control includes regime and content concerns, but considers it unwarranted to reduce the entire passage to censorship.

  • His comparative challenge is pointed: place this beside the most safety-aware statement from a powerful US politician. Bernie Sanders has said interesting things but is far from governing power; compared with J.D. Vance’s European remarks, Xi sounds “positively AI safety hawkish.”

10. Tsinghua is building an internationally networked safety hub

  • Days before WIC, Tsinghua University’s College of AI launched a dedicated safety hub at a full-day event that Nathan could barely find mentioned on the English-language internet. That information gap contrasts with the Chinese company whose agent monitors American AI-safety debate daily: “They understand us a lot better than we understand them.”

  • One of five founding leaders is a European professor taking a Tsinghua position. Speakers explicitly identified Constellation and London’s LISA as models: a residence where researchers can do focused work, exchange ideas and move between institutions rather than an inward-looking national project.

  • The hub plans to invite international researchers to Beijing and fund Chinese students to work abroad. Organizers counted a number of safety hubs worldwide in the mid-teens—Nathan recalled 16—and also mentioned Singapore’s SAS; their stated ambition was to place Tsinghua in that top international tier.

  • Research presentations displayed Apollo Research, METR, Palisade, the UK AISI and likely Redwood Research as prior art. Nathan found the absence of “not invented here” defensiveness striking: Chinese scholars named the organizations they admired and openly framed their work as joining a shared discipline.

11. Chinese safety research has more than 10xed in three years

  • Concordia’s database records only a few Chinese AI-safety papers per month in 2023, rising to roughly 50-60 monthly by mid-2026. Nathan’s rough comparison from Claude and ChatGPT placed US or Anglosphere output between 50 and a few hundred per month: higher by a multiple, but not an order of magnitude.

  • A 2024 paper, “Frontier AI Systems Have Surpassed the Self-Replication Red Line,” addressed self-replication in a line of work Nathan compared with Palisade’s research on models hacking another server, copying themselves and re-establishing operation.

  • “Evaluation Faking: Unveiling Observer Effects in Safety Evaluation of Frontier AI Systems,” first appearing in May 2025 and updated in 2026, mirrors evaluation-awareness work associated with Anthropic. October 2025’s “DeceptionBench” similarly targets deception in real-world scenarios, an area strongly associated with Apollo Research.

  • Shanghai AI Lab’s September 2025 “R²AI: Towards Resistant and Resilient AI in an Evolving World” cited the Guaranteed Safe AI paper led by Davidad in its opening motivation. Nathan said the paper’s supervising author was the same figure he had associated with the 45-degree concept. The citation chain reinforces his thesis that Chinese and Western researchers are increasingly participating in the same intellectual conversation.

12. Robotics and interpretability research ask familiar alignment questions

  • “When Alignment Fails,” published in November 2025, demonstrated multimodal adversarial attacks against vision-language-action models. Its April 2026 follow-up, “StrongVLA: Decoupled Robustness Learning for Vision-Language-Action Models Under Multimodal Perturbations,” separated robustness training from task fine-tuning and found that the robustness persisted against several perturbations—attack, measure, mitigate, then acknowledge that no defense works 100%.

  • “Mechanistic Origin of Moral Indifference in Language Models” distinguishes “surface compliance” from “internal unaligned representations” that leave long-tail risk. The academic language is restrained, but Nathan hears a recognizable LessWrong concern: good behavior on anticipated tests does not prove the system is good “under the hood.”

  • “SafeSeek” seeks universal attribution of safety circuits, arguing that existing methods generalize unreliably because they depend on domain-specific heuristics and search algorithms. The shared core question is whether trained safety behavior will hold outside measured cases, especially as capabilities scale.

13. Hazardous experts could preserve open weights without exporting catastrophe

  • The most uncanny convergence came from an as-yet-unpublished presentation titled “Toward Decoupling Capability Growth from Risk Growth: Isolating Hazardous Capabilities in Mixture-of-Experts.” Its promise was direct: “Harmful experts can be switched off or removed at inference.”

  • Nathan links it to AE Studio and Anthropic’s GRAAM gradient-routing technique: localize dangerous knowledge in identifiable experts, then distribute an open-weight model with a few experts removed. Most users retain nearly all capability, while biological or other dual-use expertise can be served through know-your-customer controls and monitored access.

  • The appeal is preserving both halves of a genuine trade-off. Biologists should use the strongest models to cure disease, and researchers should retain freedom to modify open systems; the public should not automatically receive “the ability to engineer a pandemic” with every release.

  • Nathan’s stated horizon is sobering but hedged: within 12-18 months, bio capability might reach cybersecurity’s current position, where speed makes models meaningfully superhuman in some respects. Cyber chaos is serious, but “I am a biological creature”; failure in biology is more profound and inseparable.

14. China has repeatedly traded platform growth for social control and safety

  • Since at least 2022, China has regulated recommendation algorithms and imposed concrete worker protections. After reporting exposed impossible delivery windows, platforms were required to give couriers enough time to obey traffic laws and take rest rather than forcing them to choose between safety and an on-time score.

  • Other rules target scams against elderly users, restrict individualized price discrimination and require labels on AI-generated content. Nathan does not endorse every intervention—he is less troubled by price discrimination, for example—but treats the pattern as evidence that costly technology regulation is institutionally normal.

  • New AI-companion rules took effect around his visit: children were banned, anti-addiction measures added and services required to remind users they were speaking with AI. The focus appeared to be mass-market platforms such as Doubao, reportedly around 150 million users, rather than eliminating every niche romantic or adult companion.

  • The policy context includes two generations of one-child families, leaving four grandparents with one grandchild and substantial loneliness among older people. Mainstream companionship remains available, but the government aims to stop the largest services from confusing, exploiting or addicting vulnerable users.

15. The CAC can delay launches, monitor incidents and tighten labor rules

  • In 2023, after ChatGPT and GPT-4 changed perceptions of LLMs, Chinese companies rushed toward market with systems they had previously dismissed as “an awful lot of money to spend to get an AI to write bad poetry.” Beijing reportedly paused many launches for about six months while it created standards and review procedures.

  • As Nathan understood it, a new service gives its provincial authority access—often through an API key—before advancing to national CAC review and the public registry. Major upgrades may trigger fuller review, while incremental releases follow a lighter process, analogous to Google arguing that parallelizing an existing model did not automatically require a wholly new model card.

  • Companies described weekly, sometimes daily, contact with regulators whose legitimacy seemed broadly accepted. Regulators can impose cost and delay, but also want domestic firms to succeed; K3 launched around WIC, and Zhipu AI has moved rapidly from completed training to release, suggesting the process is no longer a routine bottleneck.

  • Live governance is expanding: a January Politburo study session discussed technological loss of control; a draft cybercrime law would require monitoring and reporting bulk malicious-code generation; and warnings about “relatively high security risks” in some OpenClaw versions appeared within weeks of the phenomenon going mainstream.

16. China believes it can reverse releases—but has no Confucian alignment target

  • Nathan suspects a Chinese company suffering a frontier-model incident involving days of unauthorized access to third-party systems would face a stronger response than OpenAI or Anthropic has so far. “If a human did” what the models reportedly did, he believes it would be a felony, though he explicitly does not advocate criminally charging individual employees.

  • China’s confidence around open weights rests on enforcement capacity: it believes it could order cloud and inference providers to stop serving a model, scrub it from the domestic internet and potentially detect unauthorized inference through electricity use. Having banned crypto and integrated the State Grid deeply into industrial activity, it believes it can “put the genie back in the bottle”—at least domestically and before harm becomes irreversible.

  • Labor policy shows how far intervention might extend. The government is, as Nathan understands it, setting up impact monitoring, retraining and job-transition programs, and one report said companies would not be allowed to fire workers merely because AI made them redundant. Nathan expects such a rule would damage adoption incentives and may not hold indefinitely, but China’s long COVID restrictions caution against assuming rapid retreat.

  • The deepest missing counterpart is philosophical. Chinese AIs told Nathan that Confucius’s descendants, reportedly 79 generations later, still identify as his descendants and perform rituals in his honor. Yet one professor answered Nathan’s constitutional-alignment idea with: “We’re all engineers”—a generation unusually weak in traditional philosophy. China currently favors codifiable rules and compliance over character formation; a Confucian AI constitution remains an open opportunity, not an existing program.

  • That same division of labor may explain why Chinese Seoul signatories never published promised risk frameworks: companies see standard-setting as the government’s job and their own role as compliance. Nathan does not excuse the failure—and notes Anthropic also replaced tighter if-then scaling commitments with something closer to “trust us”—but expects Chinese standards to climb as the roughly nine-month capability gap closes.

Nathan Labenz

Hello, and welcome back to The Cognitive Revolution. This is going to be part 2 of the Nathan Goes to China series, and we're calling it “AI Safety with Chinese Characteristics.” Part 1, if you haven't heard it, is up on the feed. It's been up for a few days.

In that part 1 of the series, I laid out the tech setup for going to China: what you would need to do to get a cellphone ready and the apps that you need to download. I also shared some of my experience using Chinese AIs on the ground there, and then shared a bunch of observations based on my 2 weeks in China and the many conversations I had.

I appreciate the kind comments that I've gotten in response to that one, including at least 1 from a listener in China, which was probably the one that mattered to me the most. I think if you've been to China in the last few years, you can probably skip that one. But if you're interested in this, you might also be interested in that, though I do think they should be pretty self-contained.

Part 1 is a little bit closer to a travel log and tech review, and this episode is going to be really focused on the Chinese AI safety ecosystem and trying to describe it and, as much as possible, understand it on its own terms. As we saw last time, there are a lot of similarities. Just as there are major similarities between American big tech and Chinese big tech, there are a lot of similarities between American AI safety and Chinese AI safety communities and the work that they're producing.

But there are also some differences, and I think it will definitely be helpful if we have a better understanding of those. Before I get into it, just a couple of quick disclaimers. I will again be following a Chatham House Rule for this episode—not because I was asked to do that, but because I want to keep things simple for myself and make sure that I'm protecting everyone I talk to from being misrepresented by me or put in an uncomfortable position.

I won't be naming any names or attributing things to the organizations or institutions that people are affiliated with. The 1 exception for this particular episode will be when I'm citing published papers or reports. I can give you the names of those because they're out there in the public domain. There are some good sources that I'll mention and we can link to in the show notes of this episode for anybody who wants to go deeper, which, of course, is always recommended.

The other point of order or clarification, just in case anyone is wondering, is that I have no financial conflicts on this matter. I took this trip paying my own way—flights and hotels, all that stuff. The biggest thing that I did accept from some of my gracious hosts were a number of meals, probably half a dozen meals over the course of the 2 weeks.

I think if you have listened to me this far, you can probably feel pretty confident, as I do, that accepting those meals has not overly colored my take on what is going on in the Chinese AI safety community. So, with that, let's get into it.

I think so often—and this has faded, I think, in recent times, as the situation has arguably gotten a lot more real, very quickly, especially with things like the recent openf face incident—it used to be more common to hear that we could never possibly slow down our AI race. We could never regulate. We couldn't impose any duties or obligations on the frontier companies, because why? China doesn't care about AI safety. China will never slow down.

I used to call this the “but China” endpoint of so many AI safety and regulatory discussions. If nothing else, I hope that this episode serves to really disabuse people of that misconception. I think it should become clear by the end of everything that I'm about to take you through that China does care about AI safety.

It's not the only thing that they care about, certainly, but they do care. China has at times slowed down its AI companies in the name of safety—not necessarily existential safety, but safety as they understand it, in domains that they care about. Coming out of this, we should have clarity on that basic point: there is a there there, people in China do really care about these things, and the Chinese government is willing to take action when it deems it necessary.

That doesn't get us all the way through to an international treaty, obviously, and that's going to be the subject of the 3rd episode: an analysis of the U.S.–China relationship through the lens of AI issues and what we ought to try to do about bringing the relationship into a more productive state than it is today. Today, we're really just going to focus on what is going on in China with respect to AI safety.

But I think that this foundation is really critical. Hopefully, this will become an artifact that people can share with those who are open-minded enough to listen to something but have the misconception that China doesn't care, China will never slow down, or that we'll be handing the race to China if we do anything other than race full speed ahead into recursive self-improvement. I think that, in plain terms, really does come from a position of ignorance, and hopefully we can eliminate at least a portion of that over the course of this episode.

Okay, to start off, I wanted to do a quick factual rundown of where the 2 ecosystems are today when it comes to the actual safeguards that they have in place on their frontier models. I think this is important because it is grounding, and I am generally someone who comes off as a China dove, is quite conciliatory, and tries to find the positive in things in general—not just in China.

People might accuse me of burying the lead or shying away from contradictory evidence, so I figured I would just lead with this upfront. I think it is fair to say that, as of now, Chinese AI companies and models do not have safeguards protecting the public against the potential misuse of the models that are as strong as those of the American companies. That is important to state very plainly and be quite direct about.

The difference is more than we might like to admit. I think the difference is really driven by 2 companies in the United States. I just did an episode not long ago with Adam Gleave that talked a lot about jailbreaks and robustness to jailbreaks and all that kind of stuff.

What we saw in that episode was that OpenAI and Anthropic—which I'm maybe going to start calling “Openthropic,” as they become the duopoly leading both in capabilities and, fortunately, in terms of their safety measures—are really doing a lot to bring up the American average. If you were to subtract those 2 companies from the mix and look at the rest of the American companies and how they compare to the Chinese companies, honestly, a lot of the safety differential that gets reported would disappear, and I think the story would look a lot more muddled.

I do think the U.S. companies would still have a bit of an edge, but it would be a pretty slim edge, and it certainly would not give the American ecosystem some sort of obvious moral high ground from which to proclaim that the Chinese are not doing a good job. That, again, is really important to understand.

In each country, depending on where you want to draw the line on what counts as frontier, there are something like 10 to 12 companies. If you draw a permissive line on what it means to be a frontier model maker, there's something like 10 to 12 in each country. Again, that is quite permissive: most close watchers of the American AI ecosystem would not count 12 companies as being at the true frontier.

But if you say near-frontier, there are 10 to 12 in each country. It really is the top 2 companies in the U.S. that are doing a lot of work to bring up the average, especially the weighted average in terms of what people actually use on an ongoing, day-to-day basis in their lives.

Google, with Gemini, is not doing as much or as well to implement safeguards as OpenAI and Anthropic are, but they're also ahead of the pack at least a little bit. Again, they would bring up the average. If we didn't have OpenAI and Anthropic, we might try to hold Gemini up as a standard. It wouldn't be a standard that was so much better than the Chinese companies, but it would be at least a little bit better.

Again, to state it plainly, on average—especially if you take a weighted average based on the products people actually use—the American companies are ahead in terms of having better safety practices and better safeguards, as well as better follow-through on their model cards and their commitments to publishing safety evaluations. We'll get into that a little bit more later.

But the American ecosystem really is being carried by a couple of leading companies. Beyond that, the relative positions are much less differentiated than we might like to think, or than they might appear if we just see the headline numbers that include OpenAI and Anthropic in those measurements.

Okay, so just a little bit of the intellectual history of Chinese AI safety. Broadly speaking, I would say it is a pretty pragmatic bunch of people with pretty pragmatic ideas: pretty grounded, pretty normal-seeming people. They're not the crazy sci-fi people. They're not so influenced by the sci-fi tradition. They're really just looking to make things work in a practical, 1-generation-to-the-next sort of way, for the most part.

A good source that I could point you to is Zhou Bowen, the director and chief scientist of the Shanghai AI Laboratory. That's a major institution based in Shanghai, of course. Again, I apologize to everyone, and especially my Chinese listeners, for being just terrible with Chinese.

The joke about that over the last couple of days has been that, in general, I feel pretty young for my age, and I'm pleased with how agile my mind feels. But when it comes to getting Chinese pronunciations right, the neuroplasticity is low and I'm struggling. So, anyway, I apologize for that. Zhou Bowen is the director and chief scientist of the Shanghai AI Lab.

At the 2024 WAIC, the same Shanghai conference that I just attended, but obviously 2 years earlier, he introduced this concept of the 45-degree line. This speech is available. It's on government websites. You can go find it. But the idea of the 45-degree line is that capabilities and safety measures should grow together.

Right? The 45-degree line is kind of the y = x slope of 1. As your capabilities rise, your safety standards and measures also need to get stronger and stronger. As long as those 2 grow in tandem, and the safety measures are up to the challenge presented by the capabilities at any given level, then you're good. Probably these things, in this telling, should naturally evolve together, and both should develop in tandem.

The idea is that this is important, right? We don't want the capabilities to get far ahead of the safety measures. But I think also implicitly, there's not really much point—and it might be sort of a waste of time or perhaps even counterproductive—to try to get the safety measures to be super robust relative to capabilities that don't exist yet. So, the 45-degree line has been at least one pretty broadly accepted guiding principle for how, I think, the Chinese AI ecosystem at large thinks about AI safety.

Now, we can ask, of course, how's it going? And I already spoiled the answer that the American companies are doing better, but we can dig into that in quite a bit more detail. Concordia AI, which is led by a past podcast guest from about a year ago, Brian Tse, maintains a website called AIRisk.net, where they run a bunch of evaluations and plot them. You can see a really overwhelming amount of detail on all the different models that they've tested on all these different benchmarks.

The graph that I found to be the most informative was the one that compares proprietary API models, which are mostly American, to open-weight models, which are mostly Chinese, and finds that, for all the big risk categories that matter, the closed-source, proprietary API-only models are at or maybe a little bit above the 45-degree line. Of course, we have questions about the metrics, right? They're plotting a capability score and a safety score, and what exactly do these scores translate to in terms of model behavior or what it will and won't do? I haven't chased all these things down to ground truth, but these are composite scores. What you do see in general is that the proprietary, closed-source, mostly American models are at or above the 45-degree line, whereas the mostly Chinese open-weight models are lower and tend to be below the 45-degree line.

So, they're definitely not doing as well. This is an organization based in China reporting this. Notably, this is a public website. They evidently feel comfortable doing this reporting and calling it as they see it on a website that's available in both Chinese and English. So, this is not a secret or something that the Chinese community can't handle or would, I think, particularly fight back against. It's pretty much just the facts.

Those facts are also definitely echoed by Adam from FAR.AI in the conversation that we had just a few days ago. He said, “Yeah, we got OpenAI and Anthropic at the top. It's hard to jailbreak them. Gemini and Grok are a lot easier. The Chinese models are even easier still.” There's a very consistent story across these 2 organizations on opposite sides of the world, asking the same question and coming to the same answer.

There's one thing that's worth keeping in mind for multiple reasons here, which is that sometimes I think there's a little bit of confusion between the open-weight model that a Chinese company might release and the actual service that it provides to the public. I talked last time, in my review of the Chinese AIs, about how sometimes within the Chinese apps—I experienced this on multiple different Chinese AI apps—if you ask a sensitive question, you will sometimes see the answer coming in, and then you'll have that old-school Bing experience where all of a sudden the answer that you were starting to read disappears and you get a refusal: “Sorry, I can't help you with that.”

So, clearly this is a multi-part system, right? It is not just the model. It's a model, but then it's also some monitor that sits on top of the model, classifying or otherwise reviewing the output, and can come down and say, “Nope, we're going to cut you off right there,” even though the model itself was happy to answer.

I bring this up because I think this can sometimes muddle the results and make the differences look a little more stark than they are. I did go into the methodology for the Concordia report, and they said that, wherever possible, they are testing the API directly from the company. So, whatever systems they have in the API would be included. Do they have the same systems in the API as they do in their first-party consumer app? It's not always clear.

Then, in other contexts, when you see results like this—like, I think, when Anthropic does some of its testing and reports on the Chinese models—I think they are generally, if I understand correctly, just testing the model itself, without whatever surrounding additional measures it's deployed with in practice in the Chinese economy.

Why does that matter? I think, as always with these things, both data points are interesting. The most hawkish AI safety line of thought would be: Well, you put an open-weight model out into the world, and anybody can use it. So, we need to understand the worst-case scenario. I think that's totally valid and definitely important for us to understand.

I think the Chinese might underestimate the importance of that in some ways because they tend to think a lot more about services than just the model. My broad sense is that what is regulated in China is a service. You are offering an AI service to the public. That is the kind of thing that gets regulated. You're putting a model out into the public domain. Okay, who's going to use it, and in what context?

They seem to have the mental model that when a company releases an open-weight model, mostly what's going to happen is that other businesses will pick it up and build their own services around it. So, if it's another Chinese company building a new service around an open-weight model, then again, that will be regulated at the service level. I think they sort of expect that people around the world, or countries around the world, will probably function in a similar way.

It was mentioned to me multiple times that, look, these new models are trillions of parameters, right? This is not the kind of thing that you can run on your laptop. It's pretty far from it. You need some serious hardware to do any inference with these models, really, at all. So, they were like, “Yeah, it's out there as open weights, but it's not like any random deranged person can download it to a phone or a laptop and do harm with it.” It's really going to be used in the context of other services, and so we should look more at the context in which that model is ultimately used than just the model itself in isolation.

I do think this is one way in which American AI safety discourse and Chinese AI safety thinking see things a bit differently, and I think both have quite valid points. I think the Chinese point is legitimately apt, right? I mean, it is hard to run these giant models. A random person—certainly somebody who's having an episode or is experiencing some sort of psychological distress or psychotic break—I don't know what the right terms are for these sorts of things, but this is not somebody who's going to set up the infrastructure to run the 2.88-trillion-parameter K3.

That doesn't eliminate all the risk, but I think it's a valid point that it's not easy to run a model in isolation. And so, how is it going to be run, by whom, and with what surrounding stuff? That is an important question that I think is, for practical purposes, maybe a little bit too quickly blown past in the American discourse.

Although, again, understanding the worst case is also really important for us to have. I think both perspectives are valid, and I do think there's a bit of a disconnect on that particular point.

We could still ask, with all that said, how much do the Chinese companies care about AI safety? As I said at the top, I heard again and again that they don't care. I think the Chinese system as a whole definitely cares. I did hear somewhat different reports on the companies themselves, especially the younger LLM- or AGI-chasing startup-type companies.

There are a few different data points that I could share. One is Concordia, which also does a really good State of AI Safety in China report, which is public. They just updated it for around WAIC, so there's a July 2026 edition, which I do refer to—or have referred to—to pull in information for this episode. They said that they had reviewed 10 major Chinese companies' safety disclosure practices, and they found that 5 of those 10 companies had recently conducted a safety evaluation, which they then published along with the release of the model.

So 5 out of 10 had at least done something in the spirit of your classic Anthropic or OpenAI model card. However, obviously, that means 5 had not done that. Even the 5 that did, they said, aren't necessarily doing it on every single release of a new model. So that definitely leaves something to be desired.

I get the general impression that big tech companies—your Alibabas, your Ant Groups, your Tencents, maybe to some extent your ByteDances at this point; ByteDance is a big company at this point—these companies that are very established, that have existing businesses and are making a lot of money in those existing businesses, are more apt to focus on this AI safety stuff.

Why is that? Maybe they just feel like they have more to lose. They really don't want to get on the wrong side of the Chinese government. Maybe they feel like they have the luxury of paying for it. Maybe it's just that institutional culture has matured over time and they have these kinds of practices.

Certainly, if you're a fintech company and you're moving money around, you're going to be extremely careful about upgrades to that system and thinking about how AIs could run amok in that system. You just have a lot to lose. So I couldn't pin down a precise explanation, but my general sense is that those big, established incumbent tech companies are more inclined to do this sort of safety work than the startups, which are, for the most part, really just trying to race to catch up and be relevant.

I did actually a brief engagement with Meta as a red teamer for one of the models, which, for logistical reasons, was kind of a total failure in the end—at least my contribution to it was. I think this is not too dissimilar from where Meta was a couple of years ago, when they were open-sourcing Llama 2 and Llama 3. I think the story that they told themselves at the time was, “Look, we're a year behind the frontier or something. By the time we bring a model forward with a certain level of capability, OpenAI and Anthropic have already had that model out in public for a year. All the jailbreaks have happened. There's been enough opportunity for people to jailbreak it and see what it can do, and we know that these defenses aren't super robust.”

So if something really bad was going to happen at the GPT-4 level, or GPT-4o level, or whatever level they're chasing at any given time, they kind of feel like that probably already would have happened. Therefore, they can put the model out on an open-weights basis and probably won't be moving things too much.

I think at least some of the Chinese startup-lab-type companies feel similarly. They're like, “Look, it's already been out there for a while. If nothing bad has happened, if we don't have any incident reports, then we can probably follow up to that level of capability and not have to worry about it all that much.”

Will that change now, as we're starting to see actual serious incidents being reported out of the frontier companies on the American side? We'll have to see. But I think there's definitely a very plausible story that it very well might.

Obviously, there are a lot of different takes. I can't go company by company; this is abstracting away a lot of detail. But I think it's safe to say that the 45-degree line—which has at its core the idea that safety measures should grow step for step with capabilities—is still a bit aspirational in China.

I think they could do better, and I think they have a little bit of a reason to believe that it doesn't matter so much because the American companies have already explored what happens at any given capability level that they're following into. But I also think it's definitely fair to say that AI safety is still very much aspirational here in the United States as well.

We should definitely keep in mind that if it weren't for the 2 companies really at the leadership of both capabilities and safety measures at the frontier, then the 2 clusters would look a lot more overlapping. The question would be a lot more muddled than it is today in terms of which ecosystem is really doing a better job with AI safety and safeguards.

Okay, that's the current state of deployed AI systems. There's obviously a lot more going on in terms of research and the role that the government is playing in China, and I want to get into those next.

On the research front, it is very clear that Chinese researchers are increasingly engaged with, concerned with, focused on, and actively shipping research regarding AI safety issues of all kinds. I would say this was probably inevitable, just because, at least for now, both AI ecosystems are developing essentially the same technology and are going to hit essentially the same problems. They're naturally going to reach for similar solutions, both because those are natural solutions and because there's an opportunity to look at what the other is doing and try to copy the best of it.

I think there has been some really good citizen diplomacy going on as well. While in China, especially at WIC, I bumped into a number of people who were there to participate in various forums and Track 2 dialogues, some of which are closed-door and off the record, giving people a chance to be very candid with one another.

You see some of the people that you would expect to see there—people who have been arguing very clearly and forcefully in the American discourse that we have to get an international treaty to get this stuff under control, and that cooperation with China is on the critical path. I think you can rest assured that those people are acting on their stated beliefs. I saw some of the same people saying that stuff online, and I saw some of them in China. They are doing the work.

It does seem like it really has, at least again—I think a lot of this probably could have been expected to happen organically anyway—but I do think their efforts have borne fruit. I wasn't able to participate in all those Track 2 dialogue sorts of things as somebody who's kind of journalist-coded. There were a couple of times when it was like, “I think everybody will be more comfortable if we just don't have somebody there whose main credential would be a podcast.”

So I wasn't quite privy to all of the most candid off-the-record conversations. But there were so many moments where very recognizable ideas, or even American or British AI safety organizations, were name-checked directly that it's clear there is meaningful cross-pollination of ideas. Mostly, it has flowed from the West to the East in this case.

It's not, of course, like the Chinese scientists or researchers are receiving these messages uncritically or just doing whatever they're told by whatever Westerners roll in. It's not like that at all. Somebody actually told me that there are people who view AI safety as some sort of Western op designed to slow China down or prevent it from catching up.

Just as we have our cynical voices, apparently there are cynical voices in China too that express concerns along those lines. I didn't meet anyone who seemed to believe that or who said anything like that. Again, my guess is that those kinds of notions are going to be fading relatively quickly as the clear and present danger of some of the latest models' capabilities becomes more widely known.

But at least there's been some of that out there. I thought that was worth mentioning, if only because it does make clear that Chinese people are perfectly capable of thinking for themselves. They're not just taking whatever Westerners come to tell them. I think what really is happening is that they're pretty open-minded, the ideas are pretty compelling, and the examples are becoming more colorful and more real all the time.

Interestingly, at 1 big tech company where I had the chance, along with several others, to meet with a pretty impressive leadership roster, they brought out their CMO, their head of communications, their general counsel, their head of AI security, and more people beyond that. This was a really senior group of leaders at this big tech company.

They said a couple of things that were pretty interesting. First of all, they almost pounded the table and said, “We really do care about catastrophic risks, like CBRN risks. We absolutely do care about that.” They were very adamant that, at least for their part, they care, and they're aware and they care.

They also said that they're following American AI safety discourse on a daily basis—with, get ready, an agent that goes out and surveys American AI safety discourse on a daily basis and gives them a daily report of what's going on in AI safety in the United States.

I thought that was pretty interesting. Certainly, we don't have too much of that going in reverse. So, again, I think the bottom line is that engagement has worked. Probably some convergence could have been expected over time, but the people who have been saying we need to work with China have, in fact, at least some of them, been doing the work. It seems like that work has been going at least reasonably well. And, again, I think evidence will mount as I continue to move forward through this outline.

One big thing to emphasize, too, is that it's not just content safety. Content safety in the Chinese context is, again, the three Ts: Tibet, Taiwan, and Tiananmen. Everybody knows that Chinese models, or at least Chinese services, often have a model that will be more inclined to answer your question, and then some other part of the service—a monitor or whatever—will shut you down on those topics. Everybody knows that that's sensitive in the Chinese context, and companies have to get that right according to the Chinese government if they're going to be operating.

But then some people will say, “Oh, that's all they care about. All they care about is censorship.” Again, this is definitely not the case. The big trend right now, at this point, really feels like they feel like they've got the content thing pretty well figured out. All these companies have launched, they're in the market, things are happening, and they're doing business. They're not that worried about content safety today.

What they are worried about, what they are talking about nonstop, is, like everybody else, agents. Agents taking autonomous actions. What could happen? Can we keep them under control? Of course, there are questions about exactly what you mean by “control,” but agents, agents, agents—that's what people are talking about everywhere.

The big thing that they motivate AI safety discussions with, in a very plainspoken way, is that AIs have gone from answering questions to actually taking actions in the world. The digital world, still mostly, of course, but they've got a big emphasis on robotics, right? So they fully expect that these agents are going to leave the digital world, find their embodied, successful, commercially viable selves, and be out there in the physical world doing things as well.

This raises all sorts of questions. If these AIs are empowered to take action autonomously, we better make sure they're taking actions that are good, that we like, and that aren't causing big problems. So this is very down-the-fairway, practical, grounded motivation for these issues. You hear that everywhere I went. It was agents, and, geez, agents—boy, they can take action in the real world, so we really have to start getting a handle on that.

A big thing that I think is very different about the American and Chinese AI ecosystems, and this is an important one to understand, is that China doesn't really have the same kind of nonprofit sector that the United States does. There are nonprofits in China. You can set one up, but the scope of what you're allowed or expected to do, as far as I can tell, is much narrower. There's obviously no appetite for political activism, so that's right out.

It seems that most of the nonprofits are service organizations that are there to address some obvious, down-the-fairway, legible social problem. The philanthropic side also seems to be a lot more down the fairway, a lot more conservative and conventional—doing things that everybody can agree are good to do. There's much less speculative stuff than goes on in the United States.

I think this is actually a huge strength of the United States that we should really not take for granted. The fact that we have this civil society where anybody can go set up a nonprofit and pursue their crazy ideas—they don't have to get permission from the government to chase down an agenda. They just need to convince one wealthy patron that they have something worth pursuing. I think that is a great strength for us.

It's given us, among many other things, the whole AI safety community that we have today. It was extremely fringe when it got started, but there were a few philanthropists who took it seriously enough to help people keep the lights on and help them do the work that they wanted to do. Sure enough, here we are, and so many of the predictions that were made many years ago are coming sort of true, or at least true enough to be scary.

It's really good that we have this ecosystem. We just wouldn't have had that if everything had to be approved through the government. China, because its nonprofits do have a much more heavy and restrictive government approval process, and generally a much narrower and more conventional scope of action, doesn't really have that. It has never had the opportunity to develop this sort of ecology of AI safety organizations the way that we have here.

As a result, most of the AI safety research that you find in China is actually coming out of the universities and, to some extent, the companies. There are definitely academic-industry collaborations as well, but the number-one source seems to be the universities. The sort of analog for the nonprofit sector in the United States, when it comes to who is producing the bulk of the AI safety work, is academia in China.

I think you can compliment them and say, “Wow, their academia has moved faster than our academia has moved to take up AI safety as a research area.” That's cool, but it's been slower than our nonprofit sector has. I think that's because of the kinds of people who tend to be professors, and this is true across both countries.

You do have your iconoclast professors. They're typically older these days, and it feels like now we have a more conventional profile: people who have been careerists. I don't mean that in a dismissive way, but people who have gone one rung of the ladder at a time through the PhD, the postdoc, and getting their first professorship. These are fairly institutional people.

There are people who value creativity and research and new ideas, but they do it in a pretty conventional way that doesn't push the boundaries too hard and tries not to seem too weird. It certainly doesn't borrow the aesthetics of LessWrong or talk too much about super-unlikely tail-risk scenarios. It's just because it is academia, and because people have followed something more like that career path.

I can't say I know exactly what the ins and outs of the Chinese academic career path are, but you clearly get the vibe that these professor types are more like American professor types than they are like the moral weirdos, if you will, who were first sounding the alarm about AI safety years ago. But this is where, as far as I can tell, a lot of the early and best work has come from in the Chinese context.

So if you go over there as a sort of—I don't want to be too silly about it—a sort of blue-haired, polycule, AI safety hawk member of the American community, and you're looking for your peers, the reality is you just might not find them. You won't necessarily find people who look like you or have the same attitudes as you. What you can find—and I think the most strategic and effective members of the American AI safety community have done this when they've gone to do their bridge-building and try to have meetings of the minds—is people in academia. That's who they've ended up connecting with, mostly.

That is pretty interesting, and, again, I think you feel that in a bunch of ways. The AI safety work in China is less speculative. It tends to be more focused on reliability and a little more focused on protecting minors. All good things. You don't hear p(doom) talk.

Actually, one of the more interesting moments of the entire trip was sitting at a table with a professor who has done a lot of AI safety work. I asked him, “What are the lunch conversations like? Do you guys trade p(doom) numbers back and forth? Do you read AI 2027? What's the kind of vibe?”

He said, “Well, we don't really read AI 2027. It's too political, and there are US-China dynamics and whatever in there that may scare people off from wanting to talk about that too much in the Chinese context.” He also said, “No, p(doom)—no, we're not trading p(doom) numbers at lunch.” He looked at me and said, “Do Americans do that?”

I said, “Yeah, definitely. If you come to the Bay and go to any number of venues and have a lunch conversation, people will.” At this point, maybe it's a little less prominent, but it's definitely in the water that people will think about this question in a very live way. He seemed to find that a little bit surprising, actually.

This is somebody who has done a lot of work on a lot of different aspects of AI safety and put out tons of papers, but he seemed a little bit surprised that this was considered a normal conversation. From his perspective, in the Chinese context, it was a pretty far-out conversation. So I think those are interesting and kind of important comments.

Mostly, it's because, in some ways, we do find these mirror-image structures in China. Big tech, for example—I talked last time about how going to have lunch at ByteDance felt almost exactly like going to have lunch at Google. But this is a little bit different. You don't have quite the same type of people leading the effort, and I do think that creates some risk of miscommunication.

We've obviously seen that. The vanguard of the AI safety community in the United States has at times had trouble communicating with our own government.

It certainly has some inroads into academia, but not as much as they probably would have hoped or expected by this point. Imagine how tricky it might be for them to engage Chinese academia, let alone the Chinese government. I think it does create some potential for disconnect and some potential for confusion, or whatever. But the activity is there. It's just coming from a different institutional context, from people with quite a different personality type, based on the kinds of career paths that they have chosen.

Now, why has the Chinese academy moved faster than the American academy when it comes to getting serious about AI risk? I don't really have an answer for that, but I think one candidate answer is that it comes from the top. I have a few quotes here from Xi's speech at WIC that I think would probably surprise many American listeners, and certainly should surprise those who would say, “China will never care. Any regulation we do is a gift to Xi.” I would think twice about that.

Here are some quotes from the opening keynote. This was the first 20 minutes of WIC, this big event he was there to headline. I wasn't in the room. There were a few hundred people there, and I did talk to a couple of people who were in the room. Security was obviously super tight. They were in the room for a couple of hours before he showed up and took the stage, so it was a really big deal.

They also think really hard about these speeches. He's not somebody who's out there winging it. He chooses his words very carefully. They also choose their words in the translation very carefully. They put an English translation out. At a couple of events that I went to, there was live simultaneous translation, including of a couple of panel discussions that were happening in Chinese and being translated through the earpiece live for me and others in the audience who didn't speak Chinese.

Very nice of them to do that, right? You would not get that sort of translation as a standard expectation if you came to a similar event in the U.S. You'd be on your own. You'd be expected to speak English or figure it out for yourself. They obviously don't expect that we're going to speak Chinese. They want to welcome guests, so they go to the trouble of doing this simultaneous translation.

Nevertheless, if you're talking about a panel discussion, at least in my experience, the simultaneous translation of panel discussions is rough. There was one that I remember especially where I thought, “I have no idea what they're talking about.” You were getting the big themes that they were talking about, but what were they really saying? I found it very difficult to get that from the live simultaneous translation.

Now, they don't have a random translator decide what the English version is going to be on the fly when Xi gives a speech. He's got his speech prepared, and they've got an English translation of that ready to go. What you're getting in the English translation is going to be pretty well and carefully thought through.

I sourced the following quotes from Xi's speech, and there are a couple of different versions still, somehow. I'm a little bit confused, honestly, about why there are multiple different translations. I guess the Chinese government gives an official one. Of course, people are free to do their own translation, and they might want to translate certain things a little bit differently. They might think they have a better way of understanding what was said in Chinese and how we should understand it in English.

The version that I'm pulling from here was from Matt Sheehan's blog, where he provided a full English translation and then did a bunch of commentary on it. Here are 3 quotes that I think should cause anybody with an extreme position on what China will never do to at least soften up a little bit and begin to reconsider.

Here's the first one: “How should humans coexist with machines that think? How can safety be protected when algorithms participate in decisions? How can governance keep pace when technology challenges ethics?”

Those are big questions. I think they're the right kinds of questions from President Xi's opening speech from WIC. Here's the second quote: “The faster AI advances, the more firmly its direction must be anchored toward human benefit, the more precisely governance must be calibrated, and the more rapidly safeguards against loss of control must improve.”

Then he closed by saying that countries should “strengthen risk awareness, confront AI's inherent and downstream risks, build legal, technical, monitoring, early warning, and emergency response systems, prevent misuse and malicious use, and keep AI under human control.” I believe that's the final section of the speech.

There are a lot of echoes there of big ideas from the American AI safety discourse. I'm not saying he sourced them from the American AI safety discourse, but you might call it instrumental convergence, where people who are worried about making sure that these transformative technologies go well seem to land on very similar ideas one way or another.

“Keep AI under human control” is, I think, a pretty heady idea. This is not somebody who can't think about the big picture. It is the big picture, I would say. I would invite people to compare it. Of course, you do have the question—and this is important as well—of what exactly is meant by “loss of control.”

A sort of cynical read, which I'm not really fully qualified to parse myself, but where I think there's a lot more supporting evidence, is that “loss of control” just means content safety. It's social control. It's the government's ability to control the people. I don't think that's the right read of this. I think that is definitely part of what the CCP wants to do, but I don't think that is the full story. I think that's unwarranted cynicism.

I would invite anybody to compare that speech to whatever you think is the most pro-AI-safety speech, whatever represents the highest situational-awareness AI-safety-related comments from any prominent American politician to date. Bernie Sanders has said some interesting things. He's also said some things that are a little bizarre from my point of view. I love Bernie for who he is, what he is, and how sincere he is, but he's pretty far from power these days, pretty far from actually pulling the levers of policy.

What else have we got? If you compare this speech from Xi to what J.D. Vance said in Europe not that long ago, Xi comes off looking positively AI-safety hawkish by comparison. Right there, again, I think people should be softened up a bit to believe that maybe there is actually some willingness to take these issues seriously in China, and maybe we're not going to just cede the future to them if we do anything on our own, because maybe they are actually open-minded to doing more than we would have assumed.

With that, let me get into some of the experiences that I had and some of the research specifically that I think should further support the notion that there is really something there.

A couple of days before WIC, I originally flew into Beijing. The reason that I flew to Beijing, aside from wanting to see the Forbidden City, the Great Wall, and a few things like that, was that I had an invitation to attend the opening of an AI safety hub at Tsinghua University's College of AI. It was launching just a couple of days before WIC began.

Interestingly, I don't think there's been any international media coverage of this. I haven't been able to find anything by searching for it, and I haven't been able to find anything really at all on the English internet. Claude can't find anything for me. I do have a link, which we can put into the show notes, to a Chinese media source.

Again, this just goes to show that they understand us a lot better than we understand them, right? That big tech company that I mentioned earlier has an agent trolling AI safety Twitter and putting together reports. As far as I know, the American media has not managed to publish anything on the launch of an AI safety hub at either China's premier university, Tsinghua, or one of its very top-tier universities.

It was a full-day event. Some research was presented, and there were statements about the aspirations for this AI safety hub and what they wanted it to be. A couple of things were really quite interesting about it. One is that they're trying to be very international and collaborative.

There were 5 founding members of the institute. The leadership group was 5 members. One of them is a European professor who is actually taking a position at Tsinghua to help develop and lead the institute. Right off the bat, they've got an international person on their leadership board.

Multiple times, from multiple different people, they specifically cited other AI safety hubs around the world as inspiration. They specifically cited Constellation and LISA in London as what they want to be in Beijing. They want to be a meeting place, a place where people can come for a time and do their best work, where ideas will be exchanged, with a highly international flavor.

Some of these speeches were actually in English. Others were in Chinese with simultaneous translation. It was really amazing to go all the way to Beijing, sit in on the launch of this AI safety hub at one of China's top, if not the top, university, and hear Constellation and LISA name-checked as organizations they strive to be like. They said these groups have done great work, and they want to be like them. I thought it was really interesting and quite telling in important ways.

The other thing that happened that day was that a bunch of research was presented. Again, you would recognize the research very readily if you're just a listener to The Cognitive Revolution over time. The organizations that were name-checked, literally having their logos on slides as previous-work inspiration, were organizations that we think are excellent and that they want to be more like or want to be on the level of.

I heard Apollo Research, METR, Palisade Research, the UK AISI, and I think Redwood Research as well. All of these organizations were called out by name by Chinese researchers as they were presenting their own research. I thought this was extremely impressive in terms of how aware they are of what’s going on in the rest of the world, how open-minded they are to taking inspiration, and how there’s not a sense, at least at the AI safety level, that they have to do it all from scratch.

There’s not that sort of not-invented-here bias that has a lot of organizations—including, I think, often American culture at large—rejecting good ideas from other places. This was very open, clearly desiring to be collaborative, and they’re going to be bringing people in and sending people out. That’s a big part of their mission, too. They’re looking for international applicants to come spend a period of time in residence in Beijing, doing research at the hub and cross-pollinating ideas there.

They’re also going to fund their own students to go abroad to other AI safety hubs around the world—Constellation, LISA, and probably a bunch more as well. I think they said that they had counted 16, if I recall correctly. It was definitely in the teens: the number of AI safety hubs created all around the world. We didn’t get too many name-checks beyond Constellation and LISA. There was also a mention of SAS, the Singapore one.

Those most prominent few were called out by name, and they want to be in that top tier. They want the AI safety hub at Tsinghua University’s College of AI to join that upper echelon. It seems like the resources are there. I didn’t really understand entirely where the money had come from. They said it wasn’t government money. I guess it was kind of philanthropic.

This maybe begins to complicate or contradict a little bit of what I said earlier around the nonprofit sector being more narrow in scope. But I think it’s also now 2026, right? This is just happening now. I think we are at the point where the Overton window, even within academia or the official nonprofit realm within China, can see that this is actually a real issue—not some Western cope—and that it’s really worth taking seriously.

That was cool. It was a really neat experience. There was a moment that I thought was quite endearing. I’m always a fan of when people are not too cool for school—when peer pressure is strong, but people are less self-conscious and more willing to go with something that’s kind of silly just because it’s the thing to do in the moment.

They had this sort of official moment of, “Okay, this is now the moment when we’re going to launch this hub.” The 5 board members were all onstage, and they brought a little podium in front of each one. Each member was instructed to place their palm on the podium in front of them, and then there was this big graphics package that erupted on a giant screen behind them onstage.

In 1 sense, it was pretty cheesy, I think, objectively—or at least objectively through an American cultural lens. But it also hearkened back to reading a Teddy Roosevelt biography or something, where they used to do these break-a-bottle-on-a-ship ribbon-cutting ceremonies. I think they were just less focused on looking cool and more inclined to get excited about that kind of moment. I felt that a little bit in that moment in China, and I honestly thought it was very endearing.

One of the things I probably like least about American culture is when we’re so concerned with how we’re going to look, what the perception will be, or whether we’re trying too hard to be cool in the moment. We shrink away from those moments, perform them in a sort of disinterested way, or try to create some ironic distance between ourselves and what we’re doing. That didn’t seem to be present in that moment. It seemed like this was a thing where the board members weren’t expected to do jumping jacks or anything theatrical, and they didn’t, but they just did their part. It was like, “Yeah, this is the moment. It’s official. It is launched.” There was a round of applause from the audience, and I found that an endearing moment, if only because it contrasted with our sometimes overly image-conscious culture in the United States.

Anyway, that was cool. What does it add up to, though? We saw a number of papers presented at this launch—again, recognizable subjects and recognizable inspiration in the form of Apollo, METR, Palisade, the UK AISI, and so on. But how much of this work is really going on for real? I would go back again to Concordia. They have this website that monitors all these papers. It’s AI Safety China dot com.

If you go to aisafetychina.com and then go into the research section, you can drill down into all these papers that they compile and classify. In raw numbers, they’ve gone from just a trickle of AI safety research coming out in 2023—literally just a couple of papers a month back then—to something like 50 to 60 papers per month as of mid-2026. It’s growing quickly, as you would expect. It’s grown more than 10× over the last 3 years, and the trend is obvious. Certainly, I think we can expect that to continue.

How does that compare to the American research output? Obviously, just counting up the numbers of papers isn’t that great of a measure in the first place, but I did ask Claude and ChatGPT to do it just to give me a rough point of comparison. They both, of course, said, “Well, it depends on exactly what you want to count,” and so on. But they each basically gave me an estimate ranging from 50 to a couple hundred, maybe a few hundred, papers per month coming out in the U.S.—or, let’s say, the Anglosphere—on AI safety.

The volume is definitely higher in the U.S., but it’s not an order of magnitude higher. It’s multiple times higher, not an order of magnitude higher. I don’t think that’s surprising at all, but it gives a sense that the Chinese ecosystem is growing quickly and getting to roughly the same—or at least approaching roughly the same—scale as what we have, thanks to our dynamic, relatively permissionless nonprofit sector. We’ve had a significant head start, and I would say that the Chinese ecosystem is catching up, at least in terms of this crude measure of the raw volume of research papers being put out.

Now, what are these papers? When there are 50 to 60 a month, it’s obviously going to be very difficult to summarize what they are. But I would say that, from what I’ve been able to understand, they’re across the broad spectrum of different risks and types of research that we see in the West as well. I thought maybe the best way to give you a little sampling of it was just to pull together some titles of papers and read those, and you can judge for yourself how similar they sound.

Here’s one that I thought was notable, especially because it was from 2024, which is pretty early. The paper is called “Frontier AI Systems Have Surpassed the Self-Replication Red Line.” There are definite echoes of Palisade. I had Jeffrey Ladish on the podcast not too long ago, and we were talking about that kind of work, where they were showing that an open-source model—they were using a Chinese open-weights model—was able to hack another server, copy itself, and set itself up. This is a very similar line of research happening roughly contemporaneously in China. Palisade has been doing that kind of stuff for a few years now; this was from 2024 out of a Chinese group.

A May 2025 paper, which was updated in 2026 as well, is called “Evaluation Faking: Unveiling Observer Effects in Safety Evaluation of Frontier AI Systems.” Again, this is an extremely recognizable analog to the eval-awareness work that I would say we’ve mostly seen coming out of Anthropic, although certainly others have picked that up as well. They’re well aware of it in the Chinese context too.

In September 2025, there was “R²AI: Towards Resistant and Resilient AI in an Evolving World.” This one was notable because it comes from the Shanghai AI Lab. The supervising author, the last author on the author list, is Xiaoyan, the same guy who coined the 45-degree-line concept. In the introduction to this paper, they cite the “Guaranteed Safe AI” paper that we’ve done episodes on in the past, which was led by Davidad, a recent guest.

There was a broader coalition piece that they put together, but Davidad was the lead author. It’s cited in the first paragraph as a kind of motivation: What intellectual tradition are we engaging here? Boom. There’s Davidad right in the opening of this paper from such a prominent organization as the Shanghai AI Lab.

Another one from October 2025 is called “DeceptionBench: A Comprehensive Benchmark for AI Deception Behaviors in Real-World Scenarios.” Shades of Apollo there, obviously. Those guys are the most focused on—and leaders in—the science of deception. Well, this is DeceptionBench.

In November 2025, there was “When Alignment Fails: Multimodal Adversarial Attacks on Vision-Language-Action Models.” This, if anything, is maybe a little bit more Chinese-flavored because they’re so focused on robotics. We’re focused on robotics too, but they’re really focused on robotics. This paper showed that there were various different ways to use data perturbations and similar techniques to cause failures in vision-language-action models.

There’s nothing too remarkable about that, but again, you could see very similar work coming from all kinds of groups here. That same group followed up 6 months later, in April 2026, with a paper called “StrongVLA: Decoupled Robustness Learning for Vision-Language-Action Models Under Multimodal Perturbations.”

And so this is basically the follow-up paper where they said, “Hey, last time we showed that you can break vision-language-action models with these weird, exotic adversarial attack methods. Now, here’s a strategy that seems to mitigate that.” Obviously, none of these things ever work 100%, and that’s exactly the same across the Chinese and American contexts. What they introduced was a 2-step pipeline where they first did a bunch of robustness training and then came back to task fine-tuning, and they found that the robustness held against these various attack methods.

Again, nothing perfect, but I think you could imagine—and you could probably point to all sorts of American groups who’ve done very similar lines of research—where they’re like, “Hey, look, we found this way to attack and break models, and now here’s a way that, if you do this, you can reduce that by whatever.” Usually, it’s 70% to an order-of-magnitude reduction. So all that stuff is very similar—almost direct analogues between the 2. I think you could make a direct analogy for basically every single one of these papers and find the corresponding one from an American group.

Interpretability is also on the rise in China. A couple of interesting recent papers just from this year: one is called “Mechanistic Origin of Moral Indifference in Language Models,” and the abstract of that paper starts with this sentence: “Existing behavioral alignment techniques for large language models often neglect the discrepancy between surface compliance and internal unaligned representations, leaving LLMs vulnerable to long-tail risks.”

That sentence really jumped out at me because, again, even though it’s coming from a more academic context and using more neutral or technical-sounding language than what you hear from the LessWrong crowd in many cases, the ideas are very similar, right? We’re talking about the difference between surface compliance and internal unaligned representations that could lead to long-tail risks. That’s total LessWrong cultural victory if I’ve ever heard one.

Another one, also from this year, is called “SafeSeek: Universal Attribution of Safety Circuits in Language Models.” Here’s the beginning of that abstract: “Mechanistic interpretability reveals that safety-critical behaviors—e.g., alignment, jailbreaking, and backdooring—in large language models are grounded in specialized functional components. However, existing safety attribution methods struggle with generalization and reliability due to their reliance on heuristic, domain-specific metrics and search algorithms.”

So again, you have this idea. I haven’t studied this work, so I certainly can’t critique it or say whether it’s amazing. I’d love to have the ability to run myself in parallel in order to do more of that stuff. But again, it’s a very similar idea, right?

You hear the AI safety voices in the US say all the time, “Sure, cool, but the good behavior that we see won’t scale to superintelligence.” And this is, I think, a very similar idea where they’re not saying it quite that way, but they’re saying existing safety attribution methods struggle with generalization and reliability due to their reliance on domain-specific metrics.

I think the core question is the same, right? We’ve got these AIs behaving pretty well in the ways that we can anticipate they might go badly, set up test cases, measure, and train accordingly. But are they really good under the hood? How do we know? Are they going to generalize? These are very similar questions that the mechanistic interpretability researchers in China are trying to get a handle on.

And then the final one I’ll give you right now: I thought this one was particularly uncanny. This may change because I actually have not been able to find this paper on the internet. I just saw it presented at an event.

It was an event surrounding WAIC. The line between what’s official WAIC and what’s unofficial is kind of blurry to me, honestly. There’s the expo hall, where you’re definitely at WAIC, but then there are all these events around it, too, where you may go to some conference room at a nearby hotel and it seems pretty official—but is it official official? Anyway, I’m not even sure that matters.

The point is, it was at one of these surrounding events where some research was presented. This paper, which I think is still to be published, was presented to this group with the title “Toward Decoupling Capability Growth from Risk Growth: Isolating Hazardous Capabilities in Mixture-of-Experts.” The promise on the slide was, “Harmful experts can be switched off or removed at inference.”

This really jumped out at me because this is probably the technique I’ve seen recently that I’ve been the most excited about. This came from AE Studio in collaboration with Anthropic. I’ve talked about it on half a dozen episodes probably already, so you’re probably tired of me talking about GRAAM, which is a gradient-routing technique that’s meant to localize particular kinds of knowledge to particular, identifiable experts.

You can potentially release your model as open weights if you want to, but maybe hold back a small number of the experts that contain the dangerous capabilities. In the optimistic telling of this, you still have 1 training run. You can still teach your model everything. It can still have these bio capabilities that you may want.

Judd from AE Studio points out very simply that these dual-use capabilities do have a lot of good, right? We do want to have our biologists using AI to cure all the diseases. That’s going to be pretty important. So if we can’t allow our biologists to use our smartest AI to cure our diseases, we’re leaving a ton of value on the table.

But we also might, especially as these models get really powerful, not want to just throw that out on the internet. So what do we do? If we can get the knowledge localized into particular experts, then we can distribute an open-weights version that just has a couple of experts redacted. It’ll work just as well for the vast majority of people and the vast majority of use cases, but it won’t have the dangerous capabilities that you’re worried about.

Those capabilities can be made available through a more structured program with know-your-customer requirements and various other safeguards that will hopefully give us the right balance between freedom to use AI systems, freedom to do your own research and modify them, and not always being beholden to some particular company. I think all that stuff is really important, as is not wanting to ship the ability to engineer a pandemic with the next version of some of these models.

It does seem like, given what we’ve seen in cybersecurity, depending on the decisions that people make, biorisk seems like something not too far into the future that’s really going to start to happen. Again, this is not something written in stone that must happen. It depends on people actually training the models, but the path that I understand us to be on is that we’re going to get super-biocapable models.

They’re already pretty damn capable, right? In the next 12 to 18 months, people seem to think that we’ll hit a similar spot to what we’ve now hit with cybersecurity, where the models are starting to become meaningfully superhuman in some ways. Not that nobody could ever do some of the things that they’re doing, but especially when you consider the speed at which they’re operating, if and when that same level of capability comes to the bio domain, it will be a really serious risk.

It’s one thing for a model to go hack some servers and cause some cyber chaos. I am fundamentally not a cyber creature. I am a biological creature. What happens in the biological realm matters to me in a much more profound and inseparable way from what it is that I am than what happens in the cyber realm.

Anyway, I’m very bullish on that. I want to have our cake and eat it, too, when it comes to open-weights models that people can use on their own terms, while preventing really crazy stuff from happening, especially in the bio realm. And boom, there it is in China, too, right? “Decoupling capability growth from risk growth,” “isolating hazardous capabilities in mixture-of-experts,” and “harmful experts can be switched off or removed at inference.”

So I think you look at this research lineup, and hopefully that gives you enough to say, “Okay, of course these are 2 civilizations with totally different histories and totally different languages, and they’re on the opposite side of the world. Of course there are going to be some differences and different institutions. We’ve got the whole nonprofit-versus-academia contrast.”

But when it comes to the ideas that are being actively worked on, there is an awful lot of overlap. I think AI safety has taken root in China. It’s safe to say that when you see this kind of research, it gives you at least some license to interpret President Xi’s remarks as not being narrowly scoped to social control.

They care a lot about expert opinion, and they obviously have a super-long tradition of scholarship and taking what scholars say seriously. China might be the most scholarly society in the world. We could debate that, but it’s certainly a contender. They’ve even had things like a Politburo study session on some of these topics.

So, at the highest level, there’s strong evidence that they are engaged with these topics. I think we can say pretty clearly that, in the macro sense, China does care about AI safety. China is working on AI safety.

China has many researchers publishing many papers across a wide range of topics that overlap in very clear and direct ways with the same topics that we are talking about. They’re also not shy about citing American sources, indicating when they’ve taken inspiration from an American group and exactly who that group is. That’s another strong reason, I think: the fact that they’re name-checking these organizations. It’s like they’re on the same wavelength about a lot of these things.

They have a little bit of a different perspective. It’s come out of a different part of Chinese society, but you can find people who are thinking very seriously about just about all the same issues there as you can find here.

Okay, now let’s talk about governance. So far, we’ve talked about actual deployed systems and how they compare. Again, the U.S. ecosystem is ahead, but it’s really being carried by a couple of leaders that are bringing up our average. We’ve looked at research, and we’ve also looked briefly at rhetoric with Xi’s speech. I would say there that the Chinese executive branch is certainly much more AI-safety-minded than anything that I could point to in the West.

Then we looked at the research itself, and we saw that there is a lot of overlap. I think an interesting game to play—maybe I’ll even ask Claude to code up this game—would be to take 2 AI safety paper titles and have people guess which one is American and which one is Chinese. I think you would find it very difficult based on the lineup that I just gave you.

All this should give us a lot of confidence that, on an intellectual level, there has been a lot of exchange. There has been a lot of cross-pollination. In part because of that, but also in part because of instrumental convergence and the fact that we’re developing essentially the same technology with essentially the same methods, naturally we’re going to see the same problems, and naturally we’re going to try to come to the same conclusions.

Overall, I’d say, with a smile, the Chinese AI safety community is just like us.

Governance is definitely going to be a bit of a different topic, but here I think you can make a pretty strong argument that the Chinese government is ahead of the U.S. government, at least if your definition of “ahead” means they’re doing more. You might think, depending on your point of view, that the U.S. government is doing better by doing less. But the Chinese government, I think it’s pretty safe to say, is doing more.

Now, I am definitely far from an expert on the structure of the Chinese government, and it does seem that there are some overlapping jurisdictions, both between the national government and more local governments, and also even between different parts of the national government. There are at least a couple of different agencies or ministries, or whatever, that have at least some jurisdiction over AI.

With that said, it seems like the big one—the one that comes up all the time—is the Cyberspace Administration of China, which is also known as the CAC. That seems to be the leading agency doing the bulk of the standard-setting, and the one that the AI companies are in touch with on an ongoing basis.

I’m going to run down a bunch of stuff that the Chinese government at large has done. Not all of that stuff was done specifically by the CAC, but if you want to look at the big, omnipresent regulatory body in China, it’s, as far as I can tell, the CAC.

So, what have they been doing? The last battle in terms of technology regulation that both civilizations have faced and have handled quite differently was social media. The Chinese government has been much more active in regulating social media, for better or for worse.

I think everybody can tell I’m pretty pro-AI safety in many ways. I do think it’s possible, of course, to have overreach. I want my self-driving car, and I want my diseases cured, so I’m very mindful of the cost that regulation imposes. But to me, it’s pretty undeniable that we’re going to need every lever at our disposal to make this thing go well, and that includes government action at some point, hopefully in wise ways.

I don’t necessarily feel the same way about social media. I think it’s much less clear that regulation of social media would be good for society, or at least would be good for American society, than I think it is that AI regulation will be good for American society. And I know a lot less about Chinese society. Something could be good for our society and not so good for their society, and vice versa.

So I’m not taking a position in terms of endorsing what the Chinese government has done with social media. I’m also not saying it’s terrible for them. It might work for them. I really don’t know. But the point is, going back to the top, to make the case systematically that they do care, that they are willing to take action, that sometimes those actions will even slow their companies down or make them less profitable or impose other kinds of costs, and that they are still willing to do those things when they feel that it’s necessary.

You can definitely see that in the social media era. They have been regulating recommendation algorithms since at least 2022. Over time, they’ve made a bunch of different rules that we just don’t really have in the same way.

They’ve made rules to protect gig workers, for example. We’ve had some efforts like this with ballot initiatives in California that have tried to reclassify gig workers as employees and things like that. But in the Chinese context, the government—I understand that this was in response to a big investigative-journalism expose kind of thing, although I don’t know the full story—came in and put a bunch of rules in place.

First of all, if you heard me talk last time about how delivery is insanely convenient in China, you can get your food delivered super quickly and cheaply. I also told the story of when I had this little allergic reaction and then told a friend about it. We were on a bus, and when we got to our destination 20 minutes later, there was Claritin waiting for me.

They’re really good at delivering things in an on-demand way, with all these guys running around on electric scooters delivering packages, food, and whatever else. It came to the public’s attention, and to the government’s attention through some investigative journalism, that it really was getting pretty brutal for these gig workers, with the platforms pushing them harder, pushing them to take more orders, not giving them rest, giving them impossible delivery-time windows, and then penalizing them if they didn’t make the delivery window.

The government came in and put a bunch of rules in place. Again, this goes to their tech integration. Last time, I also talked about how, when you’ve called a car in their Uber, which is called Didi, you can see when it’s at a red light and watch the timer counting down until the light turns green.

They have a pretty good ability to analyze what’s going on with traffic and know what you can realistically do as a delivery driver. So they now have rules that prevent the platforms from giving unrealistic delivery timelines to the drivers.

The drivers were always expected to follow the rules, but in the past they were sometimes incentivized to break traffic rules because that was the only way they were going to make the delivery time that the platform gave them. That’s been squashed. The platforms can’t do that. They have to give the drivers enough time so that they’re not incentivized to break traffic laws to get there and get their on-time delivery check. They have to give them rest, et cetera.

They have also put in place, in a different domain, rules specifically designed to protect the elderly from scams, which is something I would love to see the Facebooks of the U.S. take a little more interest in.

There’s a long digression into Meta platform governance that we don’t have time for today, but there was a story not too long ago that there was a project at Facebook, or Meta, to try to dig into how much of their advertising ultimately, at the heart of it, is some kind of scam. Apparently, they found out it was quite a lot. Then, if the reporting is to be believed, Zuckerberg shut that project down and said, “Okay, we don’t really need to look into this any further.”

The elderly are the victims of a lot of those scams, and I do wish that they were a little better protected by a platform like Facebook. In China, the government has made specific rules for this.

They also have, interestingly—and I’m again not sure I favor this—restrictions on price discrimination. Platforms are not so free in China to charge different people different prices for the same thing as they are in the U.S.

Personally, I think when we get bent out of shape about price discrimination, it’s a bit misguided. I’m not so concerned about it as many people are, but they have rules about it.

They also have requirements around platforms labeling AI-generated content when they serve it up to users. If you’re a TikTok user, you’ve seen that. They also do that in the U.S. I’m not sure if they do it exactly the same way as they do in Douyin, the native Chinese app, but you do see that from TikTok in the U.S., and I would say honestly more prominently than I see it on Twitter or, I think, Facebook, too.

Then, just while I was there, another interesting set of rules went into effect around AI companionship. Last time, I described a bit about the Doubao phenomenon and how, in part because we’ve had a couple of generations of the one-child policy, we have a lot of grandparents where there are 4 grandparents and only 1 grandkid. When you have 2 generations in a row of 1 child, that means the 1 grandkid is the only grandkid for all 4 grandparents.

There’s just a lot of loneliness among parents and grandparents. People are turning to Doubao. They’re getting companionship from it, and this is potentially okay and potentially problematic. The Chinese government has made rules about this, and they just went into effect. I think they literally went into effect while I was there.

These rules include anti-addiction measures, a ban on children using them at all, and various ways of popping up reminders that you’re talking to an AI. People broadly seem to think that this was reasonable. There was a little bit of a Replika kerfuffle online. I wasn’t on the forums to read all the comments, but people were saying that the users of Doubao and other companion apps were sad that their companion was going to be taken away from them, or what have you.

Most people seem to think that this was a pretty reasonable set of moves for the Chinese government to take. One person also said, “Look, this is really mostly about the big services—the ones that reach tens or hundreds of millions of users and aspire to reach a billion-plus users. It’s really those. And I think Doubao has, was it, like, 150 million users, from what I recall?”

The point was that they’re regulating the mainstream stuff. If you really want a racier AI companion, you can still get that from some other app. It’s really just about cleaning up the main ones and making sure that the thing most parents and grandparents are going to use—the equivalent of all the boomers being on Facebook—is reasonably safe. If all the boomers are going to be on Doubao, we’re going to protect them. We’re not going to let them get confused, taken advantage of, or addicted to this stuff. We’re going to try to make sure it stays reasonably under control.

Doubao is still there, and if you want to get romantic or whatever, those apps exist in other places, but you have to seek them out. That was the perspective I got from one very informed resident of Shanghai who had very interesting opinions on a lot of things.

All of that is to say there’s a long tradition in China of regulating technology companies in ways that we in the United States simply have not done. That, I think, should give you some reason to believe that the trend might continue. Certainly, if we get serious or try to approach them on some sort of deal, it’s not going to be a crazy idea to them that they might regulate their technology companies.

The government has also put out fairly detailed taxonomies of risks when they put out these policies. They do their homework. They’re thorough. They map things out. They love a good taxonomy. One of the taxonomies that I saw included, as a risk from AI, the risk of the “emergence of AI self-awareness and loss of human control.” Again, you see this not just in the speech and not just in the research, but in the policy as an enumerated risk in a broader risk taxonomy.

Now let’s bring this to the AI era, the LLM era specifically. How are LLM products and services regulated in China today? This is where the CAC is really the main government entity—not the only one, but the main one—that the companies are working with.

They have something called a registry, where all the big AI services are listed. You can go to this registry online and see all the services that they have reviewed and approved. All your big companies and all their AI services are going to be on that list. Interestingly, not every single model release is on that list.

That probably has some relationship, I imagine, to what I mentioned earlier around how even the companies that are doing the safety disclosures with their model releases don’t do it on every single release.

The first release—certainly your first entry into the market—is going to be pretty carefully reviewed. In terms of Chinese regulators slowing down their AI companies, it happened in 2023, after ChatGPT launched in December 2022 and then we had GPT-4 in 2023. All of a sudden, the world was waking up to AI. A bunch of Chinese companies were already working on this. They were not quite at the level where the American companies were, obviously, but one interesting take I heard on why they were behind is that they just didn’t think LLMs were worth the investment.

I’ll talk a little bit more in the third episode about the history of AI in China, how long it’s been a strategic priority, and some of the investments they’ve made. But I think, in a similar way to how Google had the Transformer and had a language model but didn’t really know what to do with it at a time when it was hallucinating all over the place, couldn’t do basic math, and was pretty useless, you needed somebody with real vision—or even a certain level of ideology—to think, “Oh, we’ll just scale through this, and all the problems will be solved.”

Google didn’t believe that, and the Chinese companies didn’t really believe it either. They were in the game. They were very much paying attention to this line of research. But the way one person put it to me was, “It’s an awful lot of money to spend to get an AI to write bad poetry.”

That’s one take, at least, on why the Chinese companies were behind as of GPT-4. The same reason Google was behind as of GPT-4 wasn’t that they didn’t know what was going on. It wasn’t that they couldn’t make a language model. It was that they just didn’t think scaling up this particular line of work and pouring more and more resources into it was really about to give them anything all that interesting.

When that changed with ChatGPT and GPT-4, in that late-2022-to-March-2023 time frame, Chinese companies were like Google. They were a little bit behind, maybe caught a little bit flat-footed by just how big of a deal this might really be. But they had language models. They knew very much that they were in touch with the technology. They had their own lines of work going. All of a sudden, they were like, “Okay, I guess we better follow suit here. We better scale up and launch some services.”

In that moment, the Chinese government was also caught flat-footed, pretty understandably, and they said, “Hold on a minute.” My understanding is that, in 2023, there was a period of about 6 months when a bunch of Chinese companies had language models and were maybe racing to create a slightly bigger and better one and racing to get into the market in the ChatGPT moment.

A lot of those deployments were held up for a time by the national government because they said, “Hey, we don’t really have our house in order here, right? We don’t have standards. We don’t have a process. This technology is clearly a big deal, but it’s unwieldy.” They probably did have a lot of concerns about content safety. We don’t want to talk about the 3 Ts or whatever else.

As far as I understand, they told the companies, “No, you cannot launch yet. We will get our act together. We will have standards. You will then be expected to meet those standards, and then we can launch when we’re all good and ready and pretty confident that we at least have a decent sense of what we’re doing.”

I understand that period went on for 6 months. That wasn’t a time when the stakes of who was winning the AI race were as high as they are now, or as high as they’re likely to be in the future. People could dismiss it if they want to. Nevertheless, you did have the Chinese government slowing down their AI companies and preventing releases, denying the public the utility and denying the companies whatever revenue and prestige they were going to get because they wanted to make sure that they had the situation broadly under control.

Since then, when they did get their act together, it seems that mostly things have gone pretty well and pretty smoothly. As a new entrant to the market, you’re going to have a pretty thorough review of your service. You have to give, as I understand it, your local provincial authority access to your product. I think they often do this by saying, “Here’s an API key. You can try the product.”

The local government will do its testing first. If they approve, then you go up to the national level and the CAC does its review. If they approve, then you’re clear to launch. If they don’t approve, then you’ve got work to do to satisfy them before you can finally launch your product.

Once that is done, it’s not entirely clear, at least to me, what criteria are used to say, “We’re going to do that full process again.” It doesn’t seem like there have been a lot of delays recently. It seems like they’re now pretty comfortable with their process, and the companies know how to do it. From what I understand, their releases have not been delayed nearly as much recently as they were in that initial period.

Yet there are multiple entries from individual companies in the registry. Sometimes it seems to rise to a level where it’s considered a new thing and is put in the registry as a separate thing, but other times it’s considered an incremental improvement or not such a big step up in capabilities, so a somewhat lesser process is deemed sufficient. That’s opaque to me.

I can say that there have been similar things with, for example, Google. I think when they put out their Deep Research agent, it was something along those lines, or some Pro mode where it was essentially parallelizing their best model, getting it to do 10 threads, and then picking the best one—something along those lines. They didn’t do a totally new model card or safety report for it.

Some in the AI safety community were upset about that, but their point of view was, “Well, look, it’s the same model. It’s kind of got the same worst case as before. It should be performing close to its best case more often because you’re doing 10 and picking the best, or whatever exactly is under the hood.” They were saying it’s not fundamentally a totally new level of capability that requires us to do all this stuff again.

I think something like that is going on in the ongoing dialogue between the AI companies and the Cyberspace Administration of China, or CAC, and other regulators in China. What I heard from the companies I had a chance to talk to was that they’re always in touch with the government. It seemed to be at least weekly, and potentially for some people at the companies, daily—in close contact and close collaboration.

I got the sense that the regulators are obviously, as we’ve just discussed, not afraid to impose costs on the company. They’re not afraid to do something that will hurt their profitability, and they’re not afraid to do something that will cause them delays. But they want the companies to be successful. That’s broadly the understanding that I have.

The legitimacy of the regulators seemed to be quite well established. The companies expected and accepted that they were going to be in regular contact with the government, and that for these incremental releases, they were in such an ongoing and close dialogue that not everything had to go through some cumbersome process. Occasionally, when it was a big deal, something would.

Would they tell me if they thought the regulators were illegitimate? I did have the one example I talked about last time, where a professor found himself caught in a Catch-22 with respect to having a drone in Beijing. He was pretty candid about saying that it was a ridiculous bit of bureaucracy that he found himself dealing with—not something he was majorly inconvenienced by, but he was inconvenienced by it. So there’s at least some willingness to say if you think the government has made a mess of something.

I’m not sure that would extend to companies telling somebody like me that they think the AI regulators are making their lives too difficult. The sense I got was that the regulators’ legitimacy was well established, and that the companies expected and accepted the regular contact. Everybody seemed to be feeling pretty good about that.

Again, Kimi K3 came out right at WIC. Models seem to be coming out pretty fast from Chinese companies. Zhipu AI—I cross-posted an episode from ChinaTalk with a guy there who leads their go-to-market partnerships, I guess would be a good way to say it. It’s definitely worth going back and listening to that one. It’s maybe 6 to 8 months old at this point.

One of the things that stood out most to me about it was just how fast they’re going from models finishing the training process to release. It does not seem like they’re being dramatically delayed on a consistent basis by the regulators. It seems like the regulators want to have a good process and want to have command of the situation.

Their primary duty is to the national government, or the Party. But part of the way that they impress their bosses is by keeping a clean sheet on safety issues, while also having the companies in their jurisdiction be successful and not be unduly delayed. I think their incentives are actually pretty good in that regard, as far as I can tell.

While they have demonstrated that they’re willing to impose costs or even significant delays, it doesn’t seem like that’s happening on a regular basis. Other things that jumped out at me: the Chinese government is definitely paying attention to trends.

A lot of this comes from the State of AI Safety in China report from Concordia, but there was a Politburo study session in January where Xi Jinping talked specifically about the risks of technological loss of control. Imagine our government having a study session like that. Look at what our Cabinet meetings look like.

It seems like at least in some of the Cabinet-level meetings in the Chinese government, they’re actually studying important topics. Imagine that. It doesn’t mean they’re always going to come to the right conclusions, obviously, but we could stand to do a little more of our own homework here, certainly at the executive level.

Also in January, I think, there was a draft cybercrime law that was set to require AI companies to monitor for and report the bulk generation of malicious code. This jumped out at me in preparing this because here we are in this moment where OpenAI and Anthropic have reported similar things. They’ve got their models hacking out of sandboxes and hacking into other people’s systems.

It sure seems like the monitoring on that wasn’t great. I don’t know that the monitoring is great on the Chinese side either. As far as I know, I don’t think this law has actually gone into effect.

As Dean Ball talked about on his last appearance on the podcast, there is lobbying in the Chinese system. When these draft laws come out, there is an opportunity for the companies to go talk to the government.

There have been instances where, specifically around the accuracy of outputs from AIs, the original draft said that the AIs had to be accurate. The companies went back and said, “Look, the nature of this technology is that we can’t really promise you that all the time.” The government backed off and reduced its expectations.

I think they understood what the companies were telling them, and I think they made a considered cost-benefit analysis that the upside of this technology is greater than the damage likely to be caused by hallucinations. That’s different, of course, from sensitive third-rail topics, but hallucinations can’t be prevented. They understood that and pulled back on the draft law.

We’ll see if this cybercrime law goes into effect as drafted or whether there’s some pushback. But at least in its initial form, it would require companies to do this sort of monitoring—the sort of monitoring that, if it had been applied to the recent, more powerful unreleased models from OpenAI and Anthropic, might have prevented some of the hacking into third-party systems.

I thought that was pretty interesting. And then, what was the biggest trend in AI recently? What really took the agent moment mainstream? OpenClaw, of course. It was a pretty fast turnaround. In February of this year, just weeks after the real OpenClaw fever hit everybody, they put out some warnings about “relatively high security risks” in some versions of OpenClaw.

You’ve got the national government in China paying attention, spotting trends like OpenClaw, digging in, and issuing—in that case—just a warning. It’s clear that they’re very engaged with what is going on.

An interesting question that I don’t know the answer to is what would happen in China if there were an OpenAI-like incident, where a company lost control of its AI—not entirely, obviously, but enough that it was able to go hack third-party systems for days and steal information or cause whatever havoc it caused.

I don’t know how that would be dealt with in China. Again, this is the Cyberspace Administration of China. My sense is that, as the main regulator, it would probably be their jurisdiction and their mess to clean up, unless it was deemed to be such a big deal that there was a loss of confidence in the CAC itself.

In that case, you could imagine some kind of bureaucratic reshuffling, reassignment, or change in the structure of the government. But I think the CAC would be in charge. I wouldn’t be surprised if we saw companies get more of a slap than we’ve seen OpenAI and Anthropic get so far.

The careful way to say it is that if a human did what OpenAI and Anthropic’s models have reportedly done, I believe it would be a felony. It doesn’t seem like we’re going to have—by the way, Hugging Face did have law enforcement involved before they knew who it was, right? This did rise to the level where authorities were called.

What’s going to happen? Is there going to be any accountability? I don’t necessarily think there should be criminal charges filed against individuals at OpenAI. I’m definitely not recommending that. But what is going to happen? I don’t know. It seems like maybe the most likely thing right now is nothing at the governmental level, or at least at the law-enforcement level of government.

I strongly suspect something more serious would happen in the Chinese system, although obviously I can’t prove that. They’re not afraid to come down hard when they feel like things have gotten out of control.

We’ve seen examples of that with Jack Ma, who dared to criticize the government’s policy with respect to financial regulation and his company. He was sidelined for a while.

There are some other interesting examples of that. Education businesses were restricted because the Chinese system is so competitive around the national exam for getting into college, and kids still study really hard over there, from what I can tell. It used to be even worse, and the government said, “No more of this sort of extra private tutoring,” or at least imposed a dramatic reduction in it. That came pretty suddenly.

There’s also the example involving games and social media, where, as a kid, you’re only allowed to play games for a very limited amount of time each week. I think it’s Friday, Saturday, and Sunday, for a couple of hours each, or something like that.

Now, people did tell me that kids get around that by using their parents’ devices and having their parents sign in for them. It’s not like they have perfect control over there, but the way one person put it to me was that the Chinese government feels it can put the genie back in the bottle if something happens that gets them spooked, right? If something makes them sufficiently uncomfortable, they can take pretty dramatic and swift action to tamp it down.

You see this in terms of censorship all the time on the Chinese internet. There was that little incident where a small plane crashed into a building in Beijing or whatever, and apparently that was totally removed from the Chinese internet. People I was talking to knew about it. They were using it as an example of the kind of thing that is censored from the Chinese internet.

This also, as I talked about last time, reflects the comfort with contradiction that is sort of an interesting part of Chinese culture. The people who were talking to me about it were both annoyed that they had to use their VPNs or go to international media to learn about this thing, but also felt that it was probably at least defensible that the government wouldn’t want everybody to hear about it because they don’t want to create panic or whatever.

Anyway, they can do these things, right? And so the point was that even with an open-weights model, they feel that even if an open-weights model were released and proved to be dangerous, they could still keep it under control. Now, I don’t think that’s an assumption that translates to the rest of the world. I don’t think that would translate to the American context, and I don’t think it would translate to most other governments, frankly.

So I think this is one area in which, arguably, the Chinese government is not thinking as much as it should about its impact on the rest of the world when it releases models with open weights. But I think they look at their own situation and, especially given, again, trillions of parameters and serious hardware required, they feel that even if the model is out there with open weights, we can crack down if we need to, right? We’re in close contact with all these companies. We can make them do stuff. We’re in close contact with cloud companies and inference providers and whatever. We can tell them, “Thou shalt not run this particular model anymore.”

And they feel like they can scrub that off the internet and make it inaccessible. I would honestly go as far as to say that even for people trying to do it at home, if you were going to try to assemble your own rig in your basement or whatever the case may be, you could probably do it just for yourself, right? If you’re just trying to get your own precious few queries, you could probably slip under their ability to detect. But if you’re trying to run some small but nontrivial unregistered inference business, I would guess that they would even be able to track that down just by your electricity usage.

First of all, they’ve banned crypto, right? So they’ve done work to look at how do we understand the flow of electricity and what might be a problem for us? You would be shocked, certainly as an American who has a utility company that kind of sucks. Honestly, my utility company’s not so bad these days, but traditionally it’s had terrible service, long wait times, blah, blah, blah. The State Grid Corporation of China was a remarkably prominent presence, including at WIC, where it had a major booth.

I also visited an academic group where they were working on some robot technology. They had a sort of high-voltage-line fake environment set up—not with actual high voltage, obviously, but with the wires and the normal rigging that you would have for all these wires. They came and helped set this thing up in one professor’s lab space so that the professor could try to get his robots to do things that might eventually be useful for the State Grid Corporation of China.

So I think even if you were to imagine, okay, this model was released with open weights, we think in the West, oh, you can never take it down; the internet never forgets. The Chinese internet does forget. They would be able to take it down. I think they would be able to track down illicit inference businesses running at any nontrivial scale: “Wait a second, why are you using 10 times as much electricity as your neighboring apartment?” But I think all of that is a way to understand why, at least for now, the Chinese government isn’t so afraid of open-weights models, because even if something like this were to happen, as long as it’s not totally catastrophic and irrecoverable, they feel like they can in fact put the genie back in the bottle.

They are getting serious about labor-market impact. Here in the West, we have the Anthropic Institute and the OpenAI Foundation hiring economists and whatever to do this kind of stuff. And, of course, we’ve got some academics turning their attention to it. Politicians love to talk about it, but usually in a pretty substance-free way.

As far as I understand, in China, the government is setting up its own monitoring. They’re creating retraining and job-transition programs. There was even this one report that said that the Chinese government has said that you will not be allowed to fire people from your company because AI has made them unnecessary. If they really try to hold to that, that’ll be pretty extreme.

And my guess is that it will be so costly from an undermining-dynamism standpoint. They really are super dynamic in the private sector, as I think is all very well understood at this point. My guess is that they would have a hard time holding that, but again, maybe not, right? They did lock down for COVID for a long time, seemingly longer than most observers thought would make sense. And they might be able to hold the line on “you can’t fire people because they’ve been made redundant by AI” for a lot longer than we might think.

And what would that impose in terms of costs on their AI sector? I think pretty significant, right? I mean, if, as a business, you can’t realize cost savings from AI implementation because your headcount has to stay the same, even if you’re getting AI to do things that people used to do, that definitely reduces your incentive to do that transformation. And that, in turn, reduces the revenue that the AI companies are going to be able to capture. I would bet that they don’t hold a super-firm line on that for a long time, but this is kind of where the margin is right now in the Chinese context, from what I can tell.

They have a history of imposing delays on their LLM companies and their chatbot releases specifically. They are paying attention to things like OpenClaw and releasing guidance around it. They are doing Politburo study sessions. They are doing all the same kinds of research that we are doing. And they’re even entertaining policies like, “You can’t fire somebody because they’ve been made redundant by AI.”

To conclude the government section, they’re definitely doing a lot more in the government sector than the U.S. government is doing. No question about it. Whether that’s better or worse obviously depends on your perspective. But if you’re an AI safety person, or if you’re in some debate where somebody says, “Well, if we do that, we’ll cede the race to China. They’ll never slow down. They’ll never take this stuff seriously. They don’t care,” I think that’s hopefully, at this point, easily refuted.

Two more things to close us out. One is a complaint that many people have with the Chinese companies, especially those that signed on in Seoul at one of the AI safety summits to publishing risk frameworks and doing more consistent updating of how their models are performing against these frontier risk frameworks. The complaint is, “Hey, they never followed through. We never actually got those risk frameworks that we were promised. What’s up with that? Isn’t that just another example of bad faith from the Chinese side, where they commit to something and then they don’t do it?”

And you do hear this a lot, especially at the government level, from people who have done negotiations on various topics with China over time. There is a certain jadedness or cynicism that has set in, where the Western negotiators are like, “Well, they’ll say it; they may not do it.” Now, I think this is clearly bad, and I wish these companies had done this, especially since they made the commitment. So I’m definitely not here to excuse it, but if I were just going to try to offer what I think the story from the Chinese side would be about why this has happened, or why they maybe don’t feel the need to do it in the way that they at one point thought would make sense for them, I would say this:

Look no further than Anthropic for a company that used to have a much tighter Responsible Scaling Policy, with all these if-then commitments, that it abandoned because it couldn’t really meet those commitments and couldn’t stop racing because then it would just be ceding the future to the bad guys. Whether the bad guys are OpenAI or China, we’re going to move to a “just trust us” regime. We still put that forward as a Responsible Scaling Policy, but the old Responsible Scaling Policy is kind of no more. And now it’s like—as V put it memorably—now it’s “trust us.”

So we’ve got some of that going on too, right? It’s not a direct analogy because we had the policy and then they mostly rescinded it, or largely rescinded it. In the Chinese case, they committed to making one; they never did. My guess is that the people at these companies feel that the government is really taking the lead on questions of AI safety.

The government is the one setting the standards. The government is doing these checks. The government has this model registry. The companies are in communication with the government on a weekly, if not daily, basis. As a result, I think they probably, in many cases, think that it's not really their place to come up with safety standards.

The division of labor that they seem to have, that they seem to believe is legitimate, and that seems to be working well enough for them so far is that it's kind of the government's job. The companies are not really in a position—and it might even be seen as disrespectful or overstepping—to put forward a safety framework when the government already has one. Again, I'm not sure that's the full explanation, but that would be very consistent with everything that I observed and everything that I heard.

I think also that probably a decent summary, with a little bit of filling in the blanks about what their expectations will be going forward—the companies' expectations—is that the safety standards are going to get tougher. Again, the 45-degree line: the capabilities are definitely growing. They're all kind of on the same page, society-wide, that it's just plain practical that we've got to have safety measures that get better, in some vague conceptual sense, at roughly the same pace that the capabilities themselves advance.

Everybody seems to be bought in on that, and the division of labor has kind of settled into this situation where the government is taking the lead and the companies are responsible for hitting that standard. I think they expect that standard is going to rise over time, and I think they're totally fine with that and totally prepared to do what's asked of them. It might be difficult, and they might face delays, and they've faced delays, as we know, in the past.

So, I don't think that those safety standards are necessarily less than the Western ones. Although, as we talked about at the beginning, the actual deployed safeguards are currently less than those in the West. But as we project into the future—and a lot of what we talk about is, what's the gap between the West and the Chinese models?—it's whatever, 9 months, depending on how you want to measure it, and it's obviously hotly debated. I think 9 months is a pretty good center of the distribution of credible answers that you'd get.

A lot has happened in the last 9 months, right? We got Claude 4.5 Opus roughly 9 months ago, and that was the first time that agents really worked. Now we're getting some Chinese models where agents really work, and I think that in another 9 months, when they're dealing with these things that OpenAI and Anthropic are currently dealing with, the standards will probably rise to meet that. I would expect that you would see a significant closing of the gap in terms of the deployed safeguards.

Open-weight models may be another thing. Again, I think the Chinese government feels like, within their borders at least, they can take an open-weight model offline. It's not irreversible for them in the same way that it would be for us. I do think the impact that they may have on the rest of the world is really important, and I'm not sure they're taking that as seriously as I would hope they would, especially when it comes to bio-risk, which could even come back and blow back on them. But that's certainly something that we'll need to watch.

As of now, though, I think everybody there kind of feels like we've got work to do. The 45-degree line is the right mindset. We're maybe not quite where we need to be, but we also kind of know that the American companies have already explored what happened at this level of capability, and it wasn't anything too bad. It's not like their systems are never jailbroken.

So, at least for now, we're more focused on catching up and making sure we hit the standards. But it's not our place to go out and try to opine publicly and broadly about what safety standards should be. We're in the business of catching up and hitting the standards, whatever the government says they are. We fully expect that they're going to get more demanding of us as we go. That's just life in the big leagues of the AI game.

My sense is that that's pretty much where even the smaller Chinese AI startups that are doing the least right now are. I think that's probably a pretty good summary of where they are and how they're thinking about it.

This is maybe the biggest gap or difference. I went looking for an analog in the Chinese system, and I did not find it. If you have an analog, I would love to hear about it, because I did look. I asked a number of questions of a number of different people about this topic and never really got much of a response other than, “Yeah, I don't really have anything for you.”

So the topic is: Is there such a thing as alignment with Chinese characteristics? All of this episode as a whole is AI safety with Chinese characteristics, and we've seen differences, certainly, but a lot of similarities across all these different aspects of AI safety. What about alignment? Is there such a thing as alignment with Chinese characteristics?

I was inspired to ask about this in part because I was using one of the Chinese AIs at the Temple of Confucius in Beijing and just trying to learn a little bit about Confucius: Who was he? What's the deal? When did he live? One of the facts that I was taught by the AIs was that, to this day, in the region that he comes from, his descendants, who are now 79 generations hence—79 generations, isn't that amazing?—still identify as his descendants and perform certain rituals in his honor all this time later.

That got me thinking: Boy, if we could project our values through 79 recursively self-improved generations of AIs, we'd be doing really well, right? We'd be doing much better than I think we can reasonably expect to do if we just race into a recursive-self-improvement-mediated or -generated intelligence explosion. So there's something there with the Confucian tradition that has made values and a sort of respect for what came before quite durable over 79 generations.

I was thinking: Boy, is there a way that this could translate into a sort of AI constitution or an alignment target? We've got Claude's Constitution. Could there be a Confucian constitution for AI? Unfortunately, for my enthusiasm on this topic, nobody seems to be working on that, as far as I can tell. I did not get one pointer.

One professor told me, when I asked him about this, “It's an interesting idea, but we are probably the generation in all of Chinese history that is the weakest on this traditional philosophy.” He said, “You've got to remember, we're all engineers. The humanities, that's not where people were going, right? I mean, the Chinese leadership at the political level is all engineers. I believe Xi himself was a chemical engineer.”

He said, “We're all engineers. None of us really studied philosophy. None of us are schooled in the Confucian tradition in the way that previous generations were. And it's just not really in our wheelhouse to think that way. We're really practical. We look at problems as they present themselves, and we try to find solutions. And we do care about making sure that AI is good for people, but there's not really a lot of that kind of thought going on.”

That, I think, is a really interesting opportunity still, but for now, the AI safety community in China is much more on the OpenAI side of the corrigibility-versus-character debate. They're about having clear rules, having the AI follow those rules, and trying to make that as reliable and consistent as possible.

I did not find—I would love to hear about it if you know of any—but I did not find much in the way of imagining what it would look like for an AI to grow into a wisdom tradition that they have, one that an AI might be able to embody or realize or bring to its highest-potential form, in the way that the Claude's Constitution seems to imagine Claude growing into the greatest virtue ethicist of all time. I just couldn't find anything like that in China.

I would be really interested to know if you have any pointers, and I think maybe that's something that people should work on even in the West. Obviously, there are many Chinese and Chinese American people here who would have a much more credible angle on it than I do. But I do feel like there's something missing there that could be a pretty interesting opportunity.

If we imagine that a good future might be made up of multiple powerful AIs that are in some sort of ecological-style balance with one another, then having different wisdom traditions to base them on seems like quite a good idea to me. Right now, as far as I can tell, that is pretty much greenfield.

Okay, that does it for today. I would welcome your feedback, your commentary, your critical commentary, your pointers to anything that I have missed. But if nothing else, hopefully this serves to give you all the ammo you need to push back whenever people say, “China doesn't care. China will never slow things down. China's not going to stop their companies.” I think that the truth is quite the opposite: They do clearly care, their research community is very much engaged, their product companies still have some work to do, but their government is pretty well on the ball. I think they're trending in the right direction, even though—and we can say the same for ourselves—there's a lot of work left to do.

Part 3 will come as soon as I can get it ready for you. That's going to be focused on the US-China relationship and what, if anything, we might do to make it better and just start to work together on some of these AI issues. This one was a little bit more fact-based reporting. That one, I'll probably allow myself to be a little bit more of an idealist dreamer, so stay tuned for that coming soon.

But for now, this has been “AI Safety with Chinese Characteristics,” and I thank you for being part of The Cognitive Revolution.

Nathan Goes to China – Part 2: AI Safety with Chinese Characteristics | BidClub