中国的 AI 优势比你想象的更大
- AccelXR 最初以“蒸馏解释一切”为前提做空中国 AI,最终却反转:“蒸馏根本无法解释他们交付的许多东西。” 关键在于,中国实验室已将混合专家模型稳定在数万亿参数规模,并把训练配方开源(DeepSeek 的 GRPO/RLVR 在数月内催生 DAPO、GSPO 等变体);如今研究流向甚至倒转,Thinking Machines 正在基于 Kimi 做蒸馏,Hugging Face 据报在 OpenAI 智能体事件期间使用中国模型抵御攻击。
- 核心宏观判断是:开放权重引发的价格战会把 token 推向“电力的衍生品”——而在电力方面,“中国目前正全面碾压美国”。 电力目前约占 token 成本的10-20%;中国电网发电量是美国的2x,新增容量速度约为美国的6x,而未来5年新增容量中 AI 仅需占1-5%,美国则需要50-70%。他对终局仍留有余地——“我不完全确定它会商品化到这个程度”——这也是政策大概率介入的原因。
- 他的硬件判断是:“到2028年,中国将拥有大规模量产且具成本竞争力的 AI 技术栈。” AccelXR 认为,首个在 Ascend 芯片上训练的数万亿参数模型将是确认信号;DeepSeek 曾在 Huawei 工程师驻场的情况下尝试,但因不稳定等问题退回 Nvidia。Tommy 指出,到2028年中国的总 FLOPs 产出仍可能只有美国的约10-15%,但从训练转向服务的结构变化,可能让这些效率优势变得决定性。
- “电灯开关”情景是:中国先倾销廉价开放权重模型,压低美国实验室利润率,直到实现训练自主,随后切断出口。 这一情景已经出现迹象:Xi 的支持开源演讲走红,大约3天后,金融时报披露 MOFCOM 正就限制开放权重模型向实验室征求意见。AccelXR 的悲观基准是未来两年左右“拆桥”;他理想中的美国路径——通过开放性竞争——“感觉有点过于乐观”。
- 实验室排名上,DeepSeek 拿下前5项创新中的4项(MLA、MoE,以及预训练中的 FP8 量化),Moonshot 位列第2,凭借1:56稀疏化、MuonClip,以及为 Kimi K3 手工编写的 CUDA 之下内核;随后是 Qwen/Alibaba 和 Z.AI,后者与 Cambricon、Ascend 的软硬件协同设计被他认为将成为训练侧的下一项突破。 Baidu 和 Tencent 属于“快速跟随者”。
- 反共识政策判断是:美国最好的做法,是让中国“回到 Nvidia 这座桥上”,因为收紧出口反而会加速其自主化进程。 中国已经看到了这一点:监管部门正放缓 H200 进口,并推动国资数据中心将外国芯片替换为国产方案;与此同时,Trump 政府在今年1月恢复了带收入分成的 H20 销售,并恢复 H200 销售。
- 差距与竞赛的核心是:美国实验室目前领先约8个月;由于“学生无法超越老师”,在中国实现训练自主前,蒸馏会让中国保持落后,但如果递归自我改进启动,8个月“感觉已经是足够大的差距”让美国实现起飞。 与此同时,人才流向正在反转:中国 AI 研究人员回国比例从2019年的12%升至2025年的28%,AI 博士产出是美国的2x,拥有中国教育背景的顶尖研究人员比例从27%升至38%;中国实验室的估值则达到收入的40-120x。Tommy 曾考虑买入 GLM,但面对约1,000x 的市销率仍然犹豫,尽管 GLM 声称 ARR 在5个月内从约1亿美元增长至超过10亿美元。
1. 一位蒸馏悲观者走进来,一位创新多头走出来
- AccelXR 开场承认:“我刚开始写这份报告时,确信中国只是在刷基准测试……但越深入研究,我越真正转向另一边:蒸馏根本无法解释他们交付的许多东西。” 蒸馏可以解释一部分能力和基准测试表现,却解释不了中国实验室在硬件受限条件下构建出的东西。
- 转折点来自 MoE 训练。在中国实验室介入前,MoE 训练“相对不稳定”;但几篇论文“证明了确实可以稳定训练数万亿参数规模的模型”,把源自美国的技术推进到前沿训练规模。再加上 AEI 智库关于中国电网扩张的报告,他的判断随之反转。
- 研究成果如今双向流动:Thinking Machines 正在基于 Kimi 做蒸馏;而在 OpenAI 智能体遭攻击期间,Hugging Face 据报转而使用中国模型抵御攻击者,因为美国前沿模型受到限制。Tommy 的反应是:“所以美国现在也在受益于中国的 AI 研究……这太疯狂了。”
2. 排名表:DeepSeek 拿下前5项创新中的4项,以及开源为何复利更快
- 他的排名是:DeepSeek 明确位列第1(“MLA、MoE 这些东西……以及最早在预训练流程中做 FP8 量化”);Moonshot 位列第2——将稀疏化激进推进到约1:56,引入 MuonClip(如今已被其他中国实验室采用或改造)、KDA 线性注意力,并通过 MoPD 实现“高质量蒸馏”;随后是 Qwen/Alibaba;再后是 Z.AI,架构创新较少,但在软硬件协同设计上投入很重。Baidu 和 Tencent 则是快速跟随者。
- 速度优势来自训练配方开源:DeepSeek R1 的 GRPO 和 RLVR 在“数月内”催生 DAPO、GSPO 等变体。“我很确定,美国实验室之间没有达到这种程度的沟通……因为这些东西是专有的。”
3. CUDA 之下:让国产训练成为可能的协同设计
- 机制可以概括为:大多数实验室通过 CUDA 及其生态编程 GPU,而中国实验室则在“汇编级别”手写底层 GPU 指令。Moonshot 为 Kimi K3 在中国芯片上手工编写核心 GPU 程序;中国实验室还使用 RL 训练模型,让模型自行编写和优化内核;DeepSeek 则通过 DeepEP 手写底层代码。
- 这项工作的价值在于:“美国不需要这么做……我们直接构建在 Nvidia 之上即可。” 针对更弱芯片做优化,最终可能让中国“长期以显著低于美国的成本提供推理服务”;结合量化技术,以及 Cambricon、Ascend 供应链,未来几年中国也可能在训练侧实现可行。AccelXR 认为,有些人仍然误判了这一点,错误地认为中国永远无法用国产芯片训练。
4. 美国实验室:能力领先仍在,估值倍数合理,并从训练转向服务
- 他对美国的判断是:“很难看到美国实验室在能力前沿领先上放慢脚步。” 在实现训练自主前,中国实验室仍将是快速跟随者。估值方面,相较收入倍数达到40-120x的中国同行,美国实验室“仍然相对便宜”。Tommy 曾考虑买入 GLM,但面对约1,000x 的市销率仍然犹豫,即使 GLM 的 ARR 据报在5个月内从约1亿美元增长至超过10亿美元。
- 行业正在发生制度性切换:“我们现在进入了这样一个阶段:重点正从能否训练出具备前沿能力的模型,转向能否大规模提供服务。” 预计美国实验室会采用中国式效率研究,并通过中国实验室的 STR 策略,将能力增量转向后训练和 RL;在这一框架下,“智能体的脚手架或运行框架本身,就可以解释大量新增能力”,而不再依赖越来越大的预训练规模。
5. Token 是电力的衍生品——中国的结构性王牌
- 逻辑链条是:效率提升最终会通过降价传导,而不是转化为利润;在开放权重模式下,“理论上不存在每 token 毛利恢复的均衡状态”。电力目前已经占 token 成本的约10-20%;随着芯片资本开支、软件利润和研发摊销空间被挤压,“token 会趋近于一种几乎纯粹的电力衍生品”,届时中国的电网优势将占据主导:中国发电量是美国的2x,新增容量速度约为美国的6x;未来5年新增容量中 AI 只需占1-5%,美国则要占50-70%。他的原话仍然保留了限定:“我不完全确定它会商品化到这个程度。”
- 中国的政策与产业领导层已经看到了这一点。他援引 Huawei 的 Ren Zhengfei:中国的发电能力和电网输电能力都很强,而 AI 发展“需要这种电力保障”。芯片侧的答案是规模聚合:由超过500,000枚 Ascend 芯片组成的超级 Pod,芯片数量是 xAI Colossus 的2.5x,总算力约为其1.3x。芯片更弱,就用更多芯片,直到“最终纯粹变成电力瓶颈”。
- AccelXR 认为,美国关于数据中心的错误信息正在拖慢部署,而中国则依靠“东数西算”,把算力部署到廉价可再生能源所在地,再叠加税收和能源补贴。他承认一个约束:根据 AEI 报告,中国芯片生产仍受供给限制;他认为供需比例可能约为5:1,产量仍不足以满足国内需求。这正是协同设计优化重要的原因。
6. 变现、价格战与3种终局
- 价格战确实存在:Tommy 的 Silicon LLM Index 是一个按权重计算的每百万 token 成本指标,从6月2日的2.2降至约1。但变现正在分化:Moonshot 现在要求营收超过2,000万美元的模型即服务提供商签署单独协议,Alibaba 也加入营收门槛,Z.AI 正考虑本地化的主权部署,DeepSeek 则引入峰值定价并上调费率。问题在于,这些公司能否支撑相当于美国同行6x以上的收入估值倍数。与此同时,智能体工作负载“对成本极其敏感”,推动模型组合转向“最便宜但能力仍够用的模型”——这正是中国实验室的优化方向。
- 他的3种情景分别是:当前的“桥梁正常运转”——美国前沿 API 掌握受监管工作负载,中国开放权重模型占据成本敏感的中端市场,西方企业同时运行两套系统进行对冲;通过美国限制云服务或蒸馏、北京切断出口实现“拆桥”——这是他的悲观基准;以及美国通过开放性竞争,以低价、接近前沿的开放权重模型参与竞争——这是他的理想情景,但“感觉有点过于乐观”。
- “电灯开关”证据在于:Xi 鼓励开源的演讲,大约在 Financial Times 披露 MOFCOM 正就限制开放权重模型向实验室征求意见前3天发表。按他的说法,正确策略是:“向市场倾销廉价开放权重模型,拖慢美国……直到中国自己追上来,然后切换开关。”
7. 反向蒸馏、8个月差距与逆转的人才流向
- Tommy 提出了一个尖锐问题:如果中国训练出更强的前沿模型,美国实验室再对它进行蒸馏,这种情况出现的概率有多大?结构性答案是:“从法律层面看,中国开放权重模型才是美国实验室最好的蒸馏教师。” AccelXR 对一些规模较小的美国实验室主动跟进并不意外;至于 Anthropic,他表示不确定。Tommy 的反驳很直接:“但他们正在造神。”
- 差距的共识约为8个月;既然“学生无法超越老师模型”,在中国实现训练自主前,蒸馏会冻结现状。8个月是否足以让美国实现递归自我改进式起飞——也就是 Anthropic 所说的“数据中心里的100万名天才”?他的答案是:“如果它进入自我递归,我认为可以。感觉已经是足够大的差距。”
- 人才基础也在转向:拥有海外博士学位、回国的中国 AI 研究人员比例从2019年的12%升至2025年的28%;中国 AI 博士培养规模是美国的2x;在中国接受教育的顶尖 AI 研究人员比例从2017年的27%升至2024年的38%。Tommy 的总结是:“当你谈论中国在任何领域的表现时,都很难看空中国。”
8. 需要关注什么——以及为何应当向中国出售 Nvidia 芯片
- 他认为最重要的拐点指标是:“明确披露在万亿参数规模预训练中实际全面使用 Ascend。” 市场常把国产芯片上的服务与国产芯片上的训练混为一谈;DeepSeek 曾在 Huawei 工程师驻场的情况下尝试使用 Ascend 训练,但遭遇不稳定和互联速度慢的问题,最终退回 Nvidia。若确认出现一个在 Ascend 上训练的数万亿参数模型,将是政策即将转向的“最大确认信号”。次级信号则是:在 Alibaba、Tencent、Baidu 和 Z.AI 全部涨价后,中国 AI 云价格是否会在今年下半年恢复正常。
- 政策悖论是——他指出 Sacks 也表达过类似观点——放松出口限制、让中国回到 Nvidia,“应该会放慢他们的发展……这会放慢他们的自主化推进,而任何形式的收紧实际上都会加速他们的推进”。中国已经看到了这一点:监管部门正放缓 H200 进口,并鼓励国资数据中心将外国芯片替换为国产方案;即使美国今年1月恢复了带收入分成的 H20 销售,并恢复 H200 销售,这一趋势仍未改变。
- 在美国一侧,他预计实验室会游说政府封闭生态系统(“从商业角度看,这完全合理”);美国政府可能取得前沿实验室的股权,也可能要求在这些实验室产生的发现成果中拥有政府权益。最后的安全悖论是:如果企业无法访问前沿模型,那么当这些模型发布后,企业也会失去防御它们的能力。正因如此,Tommy 最终亮明立场:“所以我看多开源……这是对我而言唯一说得通的事情。”
完整逐字稿
China is just absolutely trouncing the US right now. When I first started this report, I was convinced that China was just doing benchmark maxing, but the deeper I dove in, the more I came around to the other side: distillation just can’t explain a lot of the things that they’ve shipped.
Wow. So the US is now benefiting from China’s AI research?
Yeah. By 2028, we have a mass-produced, cost-competitive Chinese AI stack.
1. Is China Innovating or Just Distilling?
Wow. Hey everyone, welcome back to the Deli podcast. I'm Tommy, one of the founding partners at Deli Ventures, and today I'm really thrilled to have AccelXR on. He wrote like the single best read I've read in months on the AI side. He did an absurd amount of research and covered all of the AI labs in China and everything going on in the US to give all of us a view into if there's actually more innovation going on, more distillation and just the state of everything. AccelXR, how are you?
Good, good. How are you?
I have a pointed question before we jump in. Before doing the research on this report, did you think China was doing more innovation or more distillation?
When I first started this report, I was convinced that China was just doing benchmark maxing. But the deeper I dove in, the more I came around to the other side. Distillation just can’t explain a lot of the things that they’ve shipped.
It kind of breaks down. You have the capability side, which I think distillation explains a large chunk of, and the benchmark side. But when you start looking at the hardware constraints that they’re working with and the innovations they’re making around them, the research going on there is really quite impressive.
Does it shock you at all that everyone on Twitter—or most people, and our entire media in the US—just argues that China is stealing our weights and distilling our models, and that there’s no innovation? Did that surprise you by the end?
Yeah, definitely. Like I said, I started with that mindset as well. I was in that camp.
But it’s quite impressive, some of the things that they’ve done, because of the innovations they’re making specifically in reinforcement learning and mixture-of-experts work. A lot of that flows back to the US as well. You see, for instance, Thinking Machines’ Tinker using distillation on top of Kimi. It’s flowing back and forth more than I think people give it credit for.
Wow. So the US is now benefiting from China’s AI research?
Yeah.
That’s crazy. I didn’t even know that was happening. There was one event recently where OpenAI used GLM or something to verify an attack that was going on internally, but that was the latest I’d seen.
I think what you might be referring to is that, during the OpenAI attack on the agent side, Hugging Face had to end up using a Chinese model to try to thwart the attackers because of the restrictions on our own frontier models. That’s a conversation in and of itself.
It’s crazy. One of the main questions I had for you was this: You’re going into the research with the view that China is stealing our weights and distilling them. At some point during your research, you probably read something that changed your mind, right? You thought, “Maybe this isn’t all just distillation. Maybe they’re doing some real things.” What was that piece of content or writing that changed your view?
I think there were a couple. On the hardware side specifically, there’s a report from AEI, which is a public policy think tank here in the US. They went pretty deep on the actual compute infrastructure that China is building out and where they are from a grid-capacity standpoint. That was one area that really opened my eyes to what’s going on on the hardware side.
On the actual innovation for the models themselves, I was already somewhat familiar with this from the early DeepSeek paper that shocked everyone when it first came out. But it was really spending more time with some of the more recent papers.
I’d say the big one for me, of all the innovations—well, there are a couple—but one of the larger ones was seeing the progress on mixture-of-experts work. Prior to the Chinese labs focusing on this, it was relatively unstable. The Chinese labs took that in a handful of papers and showed that you can actually stabilize multi-trillion-parameter-sized models.
That’s when it clicked for me: The progress they’re making, even if it’s based on some US research, is really about mixing all these techniques and porting them over to make them viable at frontier-training scale. That’s really where it flipped for me.
I’m glad you bring up the DeepSeek report. I had Jeffrey Emanuel on the podcast 1 or 2 years ago when DeepSeek released its paper. It sent the market down by hundreds of billions of dollars, and everyone was saying, “They just distilled our models.”
2. DeepSeek and China’s Leading AI Labs
I started reading the DeepSeek papers, and 2 of the things that stood out were that they really innovated on mixture-of-experts—turning on pieces of the model at each time because they were using second-rate hardware—and, second, that they were using second-rate hardware to train. It definitely felt like there was real innovation going on, and that’s why I’ve always held that there’s more innovation than distillation in China. But I think your report really did the hard work and figured it out, which is good.
When you look at them, DeepSeek is definitely the leader on the research side. If I had to rank the innovations, I’m pretty sure DeepSeek takes 4 of the top 5 as its own findings.
What’s interesting is how quickly it accelerates because of how open source it is. One example is GRPO and RLVR, which, not to get too technical, are reinforcement-learning techniques coming out of DeepSeek, specifically DeepSeek-R1. Within a few months of their release, you had a variety of variants come out, like DAPO and GSPO.
It just shows how quickly things can accelerate when the training recipes are actually open-sourced. That gives them a leg up. I know for a fact that our labs aren’t communicating to that degree or extending each other’s research to that degree because it’s proprietary.
I’m curious for you to rank your list of the most innovative firms in China. I know you mentioned DeepSeek would be number 1. Who would you have as 2nd, 3rd, and 4th?
DeepSeek’s definitely number 1. They did the MLA and MoE work. I believe they were the first ones to do FP8 quantization during the pre-training pipeline, so they’re for sure number 1.
Number 2, I would probably say Moonshot. Interestingly enough, they’re really aggressive about pushing how sparse you can make the models in that context. I think the most recent release, off the top of my head, is 1:56 sparsity, which is huge. They were also the ones that brought out MuonClip, which is now adopted or adapted by all the other Chinese labs.
They did the KDA linear-attention work and were also innovating on the distillation technique—not the bad distillation, but the good distillation—via MoPD. I think those 2 are the big ones.
Qwen, from Alibaba, does a lot on this side as well. To a lesser extent, there’s Z.ai. They do less architectural-novelty work, but they’re doing a ton of work on the hardware–code-design side.
The Chinese labs are dealing with really constrained bandwidth limitations for the hardware they’re able to get. With that being said, what Z.ai is doing with the GLM family is co-design work with Cambricon and Ascend. They’re trying to develop below CUDA and actually work with the Chinese chips and the limited capacity they have.
I think the big breakthrough on the training side is going to come from this code development. That’s how I’d situate them. A lot of the others, like Baidu and Tencent, I’d put down as more like fast followers.
That’s interesting. I know DeepSeek is on top, but I probably would have thought, off the top of my head, that GLM would have been ranked a lot higher. It’s interesting to see.
Like I said, they do great stuff on the code-design side. It’s just a little bit less on the architectural novelty. They adapt more of the techniques being developed at those other labs, I would say.
3. How China Is Working Around Weaker Chips
You mentioned co-design work—the AI lab working to marry the software to the hardware in a pretty unique and form-fitted way. What does that actually mean? I’d love to walk through what that means and what you think the implications would be for these AI labs in China.
Keeping it relatively high-level, most labs are programming their GPUs through NVIDIA’s standard software tools, which are CUDA and that ecosystem.
The Chinese labs have actually been going beneath that and hand-writing low-level GPU instructions, exploiting some of the chip’s behavior at a lower level. One big example here was Moonshot and Kimmy K3. They hand-built their own core GPU programs that work with the architecture for a Chinese chip instead of relying on off-the-shelf libraries.
Interestingly, as a side note, they also use reinforcement learning, or RL, to train a model to self-write and optimize these GPU kernels. So they’re kind of accelerating AI by building AI to a degree as well.
Another example would be DeepSeek. They also hand-wrote some low-level GPU code called DeepEP. They’re essentially going beneath the libraries on top of the GPUs to write to the GPUs at a lower level. That’s not necessary in the US. Here, we don’t need to worry about that as much because we can just build on top of NVIDIA.
They’re essentially trying to optimize for their weaker chips, and I think this gives them a strength in the long term when it comes to being viable to do training themselves in China, on their own stack.
If I have a mental model where CUDA is the programming language for NVIDIA at the top, and then I have the silicon—the real physical hardware—at the bottom, they’re going below CUDA, obviously above the hardware, and messing with machine-level code.
Yeah, exactly. Below CUDA, you have frameworks at the top, like PyTorch and vLLM, then CUDA below that. They’re writing at the assembly level, essentially, for GPU code.
Interesting. So if they’re already doing this with NVIDIA chips, what you’re saying is that, with China-native chips, they can leverage that and take it to another level. I’m just wondering about the implication, basically.
Yeah, I think the biggest implication is that we’ll likely see this code development lead to a scenario where they’re optimizing at levels that our companies don’t need to. They’ll end up in a situation where they’re able to serve inference significantly cheaper than the US in the long term.
One of the big things I’ve been watching is how capable the Chinese labs are of training on their own domestic chips. I think that, in the long term, this co-development work and the quantization work leads to both being more efficient on the inference side and becoming viable on the training side as well within the next few years.
I think some people still get this wrong. They believe that China can never train on its own chips because of these limitations. But this type of work, alongside the bolstering of their own supply chains with Cambricon and others, will ultimately lead to them being viable within the next few years.
I want to continue talking about the Chinese side, but I also want to talk a bit about the US labs. As we go back and forth, I want to get your view on how this all shakes out.
On the Chinese side, the conclusion I’m drawing from what you’re saying is that there’s real innovation here from the labs. They’re doing some pretty crazy stuff, and they’re going to start building their own chips.
Going back to the US side for a little bit, one of the things you told me before the podcast was that Anthropic is doing really well. They’re getting a lot of hate, but the models are great. How are you feeling about US innovation, the labs we have here, and what you’re seeing? How do you feel about them?
It’s always hard to pick it apart because we’re just not as transparent about what’s going on under the hood. But from a capability standpoint, it’s hard to see the US labs slowing down in capability-frontier leadership. I think the Chinese labs will ultimately always be fast followers until they’re able to get some kind of sovereignty on the training stack.
In the US, even just looking at the multiples, the ARR numbers are crazy for Anthropic and OpenAI—absolutely insane. Even at these huge valuations, you’re talking about relatively cheap multiples versus what’s going on on the China side, where you have multiples between 40× and up to 120× revenue.
I was going to buy GLM a couple of months ago, but it was at a 1,000× price-to-sales multiple. I didn’t know if I could do that.
GLM-Zhipu is interesting. Its ARR ramp in the last 5 months has been kind of insane. They went from 100 million to, I think, over 1 billion now, according to what they’re saying.
Well, I guess if we don’t have the benefit of open source in the US to understand what the labs are doing, maybe from a higher level, are you confident that the trajectory OpenAI and Anthropic are on—more data, more compute, more training, the Bitter Lesson style—will continue? I guess then we could work backward to determine whether they’ll continue to be impressive.
I think realistically we’re in a scenario now where it’s starting to transition from being able to train these models at frontier capabilities to actually being able to serve them at scale. I wouldn’t be surprised if the US labs start implementing some of the research coming out of the Chinese labs in order to scale how much they can serve.
We already see it with service outages and things like that. There’s just not enough compute to run these huge models, so they need to start optimizing for actually serving a large client base. OpenAI came out with its new Flash model, which is significantly cheaper, and I think that’s where progress is going to start heading.
The Chinese labs use the STR strategy. Instead of focusing so much on the pre-training side, they’ve been trying to shift more of the capability gains to post-training. A lot of the work they’re doing on the reinforcement-learning side is an example of that.
That’s been broadly adopted. For instance, you see papers coming out now about how the scaffold or harness for an agent can explain a large amount of capability gain. That’s essentially post-training work. I think we can start heading that way more, rather than focusing so much on the pre-training side and larger and larger models.
4. China’s Biggest AI Advantage: Electricity
I like that. I feel like a lot of my own increase in capability has just come from massively designing my news-research Hermes agent to run a lot of my life, so I understand. Exactly—build the harness. Now I can’t be without it. That’s how it is.
It’s interesting, and I guess let’s flip back to China. You had a really interesting take on electricity and AI models. I don’t want to give it away; I want you to describe your thesis here because I thought it was really solid.
If we think about what I was just saying—that the efficiency frontier is becoming where we’re competing, rather than raw capability—over the past few years, all of those gains get passed through as price cuts. They’re not typically retained as margin, although we are starting to see a shift with some of the pricing changes and other things that have been going on.
At the end of the day, you’re taking market share through price cuts for all these capability gains. If you have open weights, theoretically there’s no equilibrium where your per-token margin recovers unless everyone collectively agrees that prices should go up. So it starts becoming lower and lower.
Right now, my understanding is that electricity accounts for probably 10% to 20% of token cost. The other expenses include chip capex, software margin, and R&D amortization. But as they get squeezed by open-weight price wars, tokens converge toward almost an electricity derivative.
When that happens, I think China has a very dominant advantage over the US. China’s grid generates twice as much electricity as the US’s, and it’s adding capacity roughly 6 times faster than we are.
Wow.
China’s AI-related power needs are only 1% to 5% of the capacity it has added over the past 5 years. In the US, it’s 50% to 70%. We have a really constrained grid, whereas China is adding capacity much faster.
If you think about tokens being an electricity derivative, it makes sense for China to export its electricity through these types of models. As a Western citizen, I want us to do well, but it’s very jarring when you look at how much they’re winning on the power-generation side.
Generally, people understand that we started with the bottleneck being the GPU, then it was the memory. More and more people are starting to realize that it’s actually grid capacity and energy.
If you think about it through that lens, China is absolutely trouncing the US right now.
That's fascinating. Let me feed this back to you: your concept is that if models commoditize down, the country that wins is literally the one with the most electricity to train and serve them.
Yeah. Even Chinese leadership recognizes this. For instance, Ren Zhengfei of Huawei was saying essentially that China's power generation and grid transmission are very good, and that the development of AI requires this power guarantee. China has been using a technique where, instead of trying to match the capability of the chip directly—chip for chip—their chips are obviously weaker, but if they can add more of them into these superpod clusters, what you get to is something that comes down purely to the electricity bottleneck. In that scenario, they will take the cup from the US.
That said, just to hedge slightly, I'm not entirely sure that it commoditizes to this extent, but it's the natural progression that you would envision if open weights continue to be the default and continue to put all this pressure on the token price itself. Which is part of the reason why I think policy probably shifts this scenario. Tommy
The crazy part—and I feel like it's all Chinese propaganda—is everybody in the US getting annoyed about data centers using too much water when they don't, right?
It seems like there's a ridiculous amount of misinformation about data centers in the US right now. I don't know where it's coming from, but it's honestly stoking a certain class of people to agree with it, and it's slowing down data center deployments. It's unfortunate. The Chinese government invests really heavily into its grid capacity. They even have a program or idea called East Data, West Computing, where they try to site their compute on cheap renewables and stack it with tax and energy subsidies. I just feel like our population would not be okay with that right now, you know?
No, I have one of the trackers I do at Hermes that maps data centers across the US to track whether they're actually getting built. It's crazy, state to state, what we're seeing. One question I had related to this is: they might have more electricity, but does China have the breadth and depth of data centers that we have in the US? I know they don't have as many top-tier chips as we have, but do they have the same number of data centers, or how do you think about that?
I don't know the exact data center comparison offhand, but I do know they are expanding these superpods or superclusters, as I mentioned. These are taking a ton of Ascend chips and combining them together. The supercluster itself has, I think, over 500,000 Ascend chips, which is 2.5 times what xAI's Colossus will have by chip count, and it will equal about 1.3 times the total compute of xAI's Colossus. They're pushing on it for sure. Here in the US, we have all these neoclouds and so on that aggregate compute as well. I'm not sure what the 1:1 comparison would be there, admittedly.
That's fair. Basically, my question is: they might have more electricity, but do they actually have the chips to leverage it?
Yeah, interestingly enough, the AEI report that I mentioned earlier talks about this as well. Eventually, over the next couple of years, their production of chips is still going to be supply-constrained. I'm trying to find the number offhand, but I believe it's a 5:1 ratio, where essentially they won't have enough chips to satisfy Chinese domestic demand themselves within the next few years either. That's part of the reason why these hardware co-design innovations are so important: they need to ramp up production as much as they can. Even in that scenario, they're not going to have enough, so they need to make these other optimizations on top of it.
Damn. Taking it back to dollars and cents, the price war has been absurd. I posted a June 2 thesis when the Silicon LLM Index was at 2.2. It's the weighted cost per million tokens, and now it's down to, I think, 1. It's fallen a lot. I think you have some pretty solid pushbacks on that index in particular, but I'm wondering: do you think Chinese pricing will continue to fall? I don't know how we want to measure it—cost per million tokens or cost per task—but I'm curious whether you think it'll reverse or keep going.
Yeah, I think there are 2 sides. One side is: can they monetize what they're doing? I think that's been the more pressing question lately.
On the Chinese lab side, you have license changes. Moonshot's Kimi K2 requires any model-as-a-service provider above $20 million in revenue to sign a separate agreement. They're essentially licensing the model out. Then you have Alibaba, which has its open-weight flagship, but is also adding a revenue threshold.
There's a dispersion happening right now, where some models are becoming more aggressive on licensing to try to capture revenue on that side. You have companies like Z.ai, with the GLM family, looking at on-premises sovereign deployments. That's how they've been trying to generate revenue instead of monetizing their open-weight models so directly. The big question is whether they can generate enough revenue to justify 6x-plus revenue multiples versus their US peers.
You also see DeepSeek recently adding peak pricing and raising its rates, so it'll be interesting to see whether it has the pricing power to justify this. Another area to consider is not so much the revenue side, but a composition shift. If you expect agentic use cases to grow over time, which I think most people envision for the future, those use cases are extremely cost-sensitive. If that's the case, you have this composition shift toward the cheapest but still adequately capable models, and I think that's where they could really lean in.
I'd love to see the US compete more fully. If we shipped either open-weight models or Flash models that could service this use case as well, that would help. The Chinese labs are definitely optimizing specifically for agent-driven workflows. It'll just be a matter of whether they can capitalize on it.
5. Three Scenarios for the US-China AI Race
That is interesting. When I speak to people who have spent a lot of time in China, I've gotten pretty bullish on Alibaba recently and have gone back and forth on that, but people always tell me to be careful because the Chinese government won't let these companies accrue tons of value, right? It's confusing to me because, on one hand, I agree with you that these companies should deploy their large private models through the neoclouds, Microsoft, or Amazon, and let their customers fine-tune them or access them, sending money back to the Chinese labs. On the other hand, it seems like China's government may have other goals in mind. What do you think about how their government views this?
Yeah, it's an interesting question. I think there are 3 scenarios for how the US-China AI market relationship plays out. Right now, we're in the default scenario, where US frontier APIs dominate a lot of the regulated and high-stakes workloads, while Chinese open weights become the default for the cost-sensitive mid-market. Western enterprises, I would expect, run both: either they self-host the Chinese weights or go through neoclouds, as you suggested, or they use the US side while essentially hedging against any US gating or Chinese hosting risk.
In this scenario, the Chinese labs do well. I would say they're able to have some kind of lock-in, actually, globally; price competition holds down some of the mid-market costs, and so on. The 2 future scenarios are that you either have a complete bridge demolition—a breakdown—which is kind of my default assumption. It's pessimistic, but I think the US could either impose cloud-rail restrictions or issue distillation rulings saying that enterprises can't use Chinese weights. Alternatively, Beijing could potentially consider removing support for exporting any frontier Chinese models once it has the training capability to ship frontier weights itself. In that scenario, you have a breakdown, and US labs end up getting all the mid-market pricing power back, which is good for supporting frontier R&D.
You kind of lose this competition, which is good for buyers of tokens. The third scenario would be a RAND-style prescription of competing through openness. This would be my ideal outcome: the US labs end up shipping cheaper, frontier-adjacent open weights. You have competition on this side for the mid-market, and the US labs can focus more on providing frontier use-case access for things like drug discovery or any of those very intelligence-reliant use cases.
To your earlier point, what does China think about this? I think Xi gave a speech not too long ago where he positioned it as, “We want to provide any developing country with international AI cooperation. We want to help you guys out, and we want to encourage open weights,” and so on. But at the end of the day, I think the good strategy is to dump cheap open weights on the market to slow down the US side of things until China can catch up itself. Once China has domestic training sovereignty, you flip and close off access to these models. That’s the long-winded answer, but hopefully it gets to what you’re asking.
No, I like the three scenarios. Maybe to linger on your third scenario for a little bit, it seems hard for me to reason about the US embracing open source, given how much money has flowed into OpenAI and Anthropic. But I also don’t want to take the position that we exist to help those companies survive, right? It is hard for me to reason about that side.
One of the really interesting things you brought up is this light-switch moment where China is going to export open-source models as long as it can, because it hurts US markets a bit and pressures the labs’ margins. Then, the second China can train frontier models, it says, “You’re cut off. You can’t export any open source anymore.” That seems like a pretty interesting potential future. Do you put a lot of stock in that?
I would say so. For instance, the Xi speech that went viral, where he was encouraging the open-source thing, came about 3 days before the Financial Times revealed that MOFCOM—the Chinese Ministry of Commerce—was consulting labs on restricting open weights. On one side of his mouth, he’s saying, “We want to encourage open-source innovation,” but on the other side, they’re already deliberating over how much they should push on this side.
I really think that as soon as they have the frontier capabilities themselves, there’s not really a reason to outsource all their research and so forth, similar to what we do here. We don’t outsource it for the same reason. Geopolitically and strategically, I would say that’s the most likely outcome: you embrace open source and open weights for now, until you reach some degree of parity in frontier capabilities. Then it’s game on in competing on that side.
All right, I have a potentially spicy question. What percentage chance do you put on China not only training a frontier model that tops US models, but US models then trying to distill it?
This is interesting. As I mentioned earlier, we have some reverse distillation going on—Thinking Machines Lab with Kimi. It just makes sense to do that, because open publishing subsidizes all of your open-source competitors. There’s a legal asymmetry here, right? A US lab’s best distillation teacher, legally, is the Chinese open-weight models. So it makes sense to do this reverse distillation.
I wouldn’t be surprised if we continue to see some of the smaller labs in the US really lean into this as well. On the frontier side, I don’t know if I could see Anthropic reverse-distilling a Chinese model. It makes sense—it’s cheap capability gains if they do get ahead of us—but I’m not sure.
They are creating God, though. I don’t know if they want to—
Yeah. Yeah. [laughter]
One of the things I liked about your report, too, was the three scenarios for the US and China. The first one you put out is the functioning bridge scenario, which you just described. It’s kind of like we keep going with what we’re doing.
It seems like chaos always resolves one way or the other. This tit-for-tat game has to resolve one way or the other. We develop AGI in the US, or China trains frontier models. I’m not sure, but it seems like one way or the other, it has to swing.
I feel like the status quo is not stable. You see it in how both countries talk about their policies here. It just does not feel stable.
I do feel that it either flips to us actually embracing open source—which feels harder, like you said—or it breaks down. I’d love to see us embrace open source; that’s my ideal outcome. But it feels a little optimistic, and pessimistically, I genuinely feel that it breaks down over the next couple of years.
I have a question that underpins how these scenarios play out. One of the things that both of us have historically been bullish on, given our work in crypto, is how open source compounds and how you can build on each other’s creations.
You talked about this earlier, but I’m trying to figure out the velocity of the open-source side in China building upon itself and creating new things, versus OpenAI, Anthropic, and others, where they’re doing it within their companies but not outside them. It’s hard for me to put a speed on each, if that makes sense.
Yeah, I agree. I think there’s been a good push recently, though. There have been a lot of developments here. I think it’s interesting when you have companies like NVIDIA releasing Nemotron—essentially, a hardware vendor itself open-sourcing a model.
That was crazy to me: competing with their customers. It’s nuts.
Yeah, and it’s kind of in line with what I was talking about earlier with the co-design stuff. NVIDIA is best positioned to extract maximum throughput for its models off its own tech stack. It’s interesting that they’re going to attack their own customers as a cost advantage here.
If I had to say what would be cool to see on the Western open-source side, it would be pushing on openness standards themselves. If we started using more permissive licensing, providing some of the training data and recipes, and providing RL environments like NVIDIA did for Nemotron, it weaponizes auditability.
I don’t have a strong stance on this—I didn’t dig into it as deeply—but I do wonder to what extent China’s labs might not want to fully disclose the filtering and other processes in their training sets. We could almost have them cede the entire narrative if we really pushed on this side.
It’s hard. Being in crypto yourself, you want to do open source, but then it becomes a whole game of how to monetize on top of it. Maybe they all just need tokens at the end of the day.
6. America’s AI Talent Advantage Is Reversing
Well, maybe taking one step back from the innovation, open source is about the people actually doing the innovation—thinking about these things, hitting walls, and coming up with crazy ideas. You had a slide in the report on talent flows, and that’s really interesting because it underpins everything.
I don’t know what the best question to ask you is here, but I’m curious what you found from tracking talent in the US versus China.
One of the big numbers on that slide to me is Chinese AI researchers with foreign PhDs returning home. The returnee rate was 12% in 2019, and by 2025 it was 28%. So we have a doubling in 6 years.
Wait, so these are foreign PhDs in the US going back to China? It was 12% in 2019, and now it’s 28%?
Chinese AI researchers specifically—not the entire PhD market—but yes, there’s this talent flow, which has always been America’s advantage. We import talent, and those people are more likely to stay here after working in universities and so forth. But we’re starting to see some signs of reversal here.
China’s domestic pipeline has just been growing on the talent side. It produces 2 times the number of PhDs that the US pipeline does, which is partly a function of how large China is versus the US. But it’s hard to argue with those types of numbers.
A single university?
Yeah.
So, internally within China, they’re producing 2 times the number of PhDs—
On the AI side.
2 times the PhDs.
Over the past few years, in 2017, 27% of the top AI researchers were educated in China, and now it’s up to 38% in 2024. There are a lot of examples like this where you’re starting to see both the domestic talent supply improving and the people who do ship out to the US to train and become educated beginning to return at greater rates than historically.
Wow. It’s kind of hard to be bearish on China when you talk about them in any domain.
Yeah, like I said, I started the report as a bear, and it kind of flipped after going through a lot of this stuff.
It’s crazy. It’s nuts. Maybe let’s flip back to the U.S. side for a little bit. There still obviously is a gap between Claude 3.5 Sonnet, o3, and other frontier models and what we have in China. There definitely is a peak-intelligence gap.
People have argued that the gap is narrowing, and we can argue that the gap may or may not exist in the future, but right now it exists. I’m curious about your view on a scenario where that gap leads us to take off toward AGI, and all these talent numbers and electricity in China don’t matter because we get there first. Do you subscribe to that view at all, or not so much?
Yeah. I’m a big fan of the recursive self-improvement stuff that we’ve been looking at more closely in the U.S. I think if you get the scenario where—I forget the exact terminology that Anthropic uses—but it’s like having a million geniuses in a data center, where you have all these AI researchers that are AI themselves, kind of self-improving the whole system, that’s the holy grail. That’s the takeoff scenario.
I think we have a leg up. The gap is relatively wide. I think the consensus is somewhere around an 8-month gap in capability. Because they do distill—I’ve never argued that they don’t do distillation. I’m pretty positive they do distillation for capability—the student can’t surpass the teacher model. So that gap will stay the status quo until they reach training sovereignty, when they can actually close it.
Is 8 months enough for us to get up to speed? If it gets self-recursive, I think so. It feels like a good enough gap. At that point, it just becomes a question of how quickly we can scale up the number of AI researchers, and that gets back to the question of grid capacity and compute capacity, which we do lead on today in terms of compute capacity.
Over time, they’ll definitely be able to catch up, so it’s just about keeping them at arm’s length. If I were a U.S. politician, I think one of the best things we could actually do—I think Sacks has even talked about this to a degree—is get them back on the NVIDIA bridge. It should slow down their development, kind of paradoxically. Any kind of tightening actually accelerates their push to build out their own chips and become self-reliant.
That is interesting, because most people want NVIDIA chips not to be sent to China, and here we are saying they should be, because otherwise they’ll make their own chips.
Yeah. I think China recognizes this. On one of the slides, I go through the export controls that we’ve implemented, and you can see that during the Biden administration, we banned the A100 and H100. We closed out the H800 loophole. After the Trump administration came in, we began rescinding some of the diffusion rules. We resumed H20 sales with a revenue share in January of this year, and we resumed H200 sales.
But China, paradoxically, is not taking the chips. They’ve actually gone the opposite direction, where they’re now encouraging any state-funded data center to swap foreign AI chips for domestic alternatives. The regulators are slow-walking and conditioning any H200 imports.
I think they’re doing this because ultimately they want to get off their reliance on this bipolar policy that we have on chips and become self-reliant. So, paradoxically, if we could loosen restrictions and get them back on NVIDIA chips, it could slow down some of the sovereignty push, which gives us a little more control over the situation.
7. China’s Path to AI Sovereignty
Yeah, it is really interesting to think through this tit for tat and where it goes. Whenever I talk to really smart hardware folks, or people creating alternative chips, they always argue that it takes 5 to 10 years to get that massive piece of hardware you use to build the chips. It’s massive and really expensive.
But from talking to you, it seems like it’s a lot sooner than we think for China to make its own chips.
Yeah. I would put the number at—I tend to agree with the AEI report—by 2028, we’ll have a mass-produced, cost-competitive Chinese AI stack.
Wow. At that point, they’re still not projected to reach the total FLOPs output of the U.S. I think by 2028, it’s only somewhere around 10% to 15% of the total FLOPs output.
But that’s part of the reason why I was saying that if the axis of competition has shifted from the training side to the serving side, all these innovations that the labs are working on actually make them better situated even under that scenario, where they have significantly less FLOPs output.
That is nuts. I’ll let people read through your report, because you have a lot of really good technical points across a lot of domains, including training and inference, that support everything you’re talking about right now. But I’m curious, in closing, what were your unanswered questions? Where do you want to take this next, and what are you watching? What were you curious about but didn’t have time to get to?
I only spent one of the slides on the training-sovereignty question, which I know we’ve touched on quite a bit, but like I said, it’s the flipping point in my mind. The big things to watch are clear disclosure of actual full Ascend usage for pre-training at trillion-parameter scale.
I think people get this confused sometimes. They’re serving on domestic chips, but they’re not really training on domestic chips. DeepSeek tried to train one of its models and even had Huawei engineers on site to help them do it, and they still ran into all this instability and slow interconnect. Ultimately, they reverted back to NVIDIA.
The first time we see a multi-trillion-parameter model trained on Ascend chips, that’s the biggest confirmation that we might have a huge policy shift from there. There are all kinds of other things to watch, like Chinese AI cloud pricing. In the beginning of this year, over the first 6 months or so, Alibaba, Tencent, Baidu, and Z.ai all raised pricing.
Any kind of normalization in Chinese AI cloud pricing in the second half of this year would mean that the domestic supply chain is actually delivering enough. So, there’s a lot to watch on the hardware side. Interestingly, I started the report focused predominantly on the AI training side and the software component, but I think the bigger question is actually the hardware.
I might spend some time in one of my next reports covering more of the hardware supply-chain side.
That’s awesome. I’m curious about your view on where you think the U.S. goes. Anthropic is going to IPO soon, and its business model is killing it on revenue. I’ve argued that it should own a percentage of the technologies, medicines, and other things it creates.
I’m wondering how they do and how things have been changing there, because there’s a lot of hard research to do on the balance-sheet side of this: the capex buildout, the AI side, and the off-balance-sheet commitments. This stuff is hard.
Open-source models reverberate through the supply chain, because if you’re spending less for an open-source model—$4 on GLM versus $5 on Opus—it changes the flow-through and how much money everybody earns. I’m curious where you think the U.S. side goes.
There are a couple of ways of looking at it. To your last point, I think the Western providers need these API margins to work out in order to fund the frontier R&D. They’re exposing prices set by labs that have Chinese money backing, which our labs don’t currently have.
I wouldn’t be surprised if we see ownership stakes from the U.S. government in our frontier labs. I’m also interested in the idea that they should maybe look at taking positions in some of the discoveries that those labs produce.
My previous report was entirely on the self-driving labs side, using AI to automate scientific research. That angle, which it seems like all the labs are pushing more toward, gives them monetization through things like IP generation. That’s one of the interesting things they could do on that side.
I expect them to continue lobbying the U.S. government to wall off the ecosystem relatively quickly. It makes fundamental sense from a business perspective. You don’t want this competition, and it becomes a national-security issue.
I was just writing up something on the Hugging Face hack. It will be interesting to see if AI labs are actually going to put their money where their mouth is on slowing down some of the developments because of how spooky some of that is getting on the security side. But it's kind of paradoxical, right? Like, if you—
If you slow down the development,
You get attacked.
Yeah, you get attacked. How do you release these models? If you're not a company that gets access to them, you lose the ability to defend yourself against them when they're released. So—
That's why I'm bullish on open source. That's why I think it's just the endgame. It's the only thing that makes sense to me.
It is crazy.
Okay, we release the frontier model open-weight, and anyone can run it now. Now it becomes a race immediately after it's released to harden all the infrastructure, you know. It's just a crazy scenario, like a game. You just picture Trump in the Oval Office using it, like, “Harden all national security. Make no mistakes.”
Make no mistakes. Yeah. AccelXR, it's awesome having you on. Your reports are second to none. They're incredible.
Awesome. Thanks for having me on, Tommy.
I'm excited to host you again. Thanks for making the time.