China’s AI Advantage Is Bigger Than You Think
- AccelXR opened this research as a distillation bear and flipped: "distillation just can't explain a lot of the things that they've shipped." The tell was Chinese labs stabilizing mixture-of-experts at multi-trillion-parameter scale and open-sourcing training recipes (DeepSeek's GRPO/RLVR spawning DAPO and GSPO variants within months) — and the flow now runs backwards, with Thinking Machines distilling on top of Kimi and Hugging Face reportedly using a Chinese model to fight off an attack during the OpenAI agent incident.
- The core macro thesis: open-weight price wars push tokens toward "an electricity derivative" — and on electricity, "China is just absolutely trouncing the US right now." Electricity is ~10-20% of token cost today; China's grid generates 2x US electricity and is adding capacity ~6x faster, with AI needing only 1-5% of capacity added over five years versus 50-70% in the US. He hedges the endpoint — "I'm not entirely sure that it commoditizes to this extent" — which is why policy likely intervenes.
- His hardware call: "By 2028, we have a mass-produced, cost-competitive Chinese AI stack." AccelXR's confirming signal is the first multi-trillion-parameter model trained on Ascend chips; DeepSeek tried, with Huawei engineers on site, and reverted to Nvidia. Tommy notes that China could still be only ~10-15% of US total FLOPs output by 2028, but argues that the shift from training to serving could make those efficiency gains decisive.
- The "light switch" scenario: China dumps cheap open weights to suppress US lab margins until it reaches training sovereignty, then cuts off exports. Evidence it's live: Xi's viral pro-open-source speech came about three days before the FT revealed that MOFCOM was consulting labs on restricting open weights. AccelXR's pessimistic default is "bridge demolition" within a couple of years; his ideal — the US competing through openness — "feels a little optimistic."
- Lab rankings: DeepSeek takes "four of the top five" innovations (MLA, MoE, FP8 quantization in pre-training), Moonshot second with 1:56 sparsity, MuonClip, and hand-built below-CUDA kernels for Kimi K3, then Qwen/Alibaba and Z.AI, whose Cambricon/Ascend co-design he calls the coming training-side breakthrough. Baidu and Tencent are "fast followers."
- The contrarian policy take: the best US move is getting China "back on the Nvidia bridge," because export tightening paradoxically accelerates their sovereignty push. China already sees this — regulators are slow-walking H200 imports and pushing state-funded data centers to swap foreign chips for domestic ones, even as the Trump administration resumed H20 sales with revenue sharing in January of this year and H200 sales.
- The gap and the race: US labs hold roughly an eight-month capability lead, and because "the student can't surpass the teacher," distillation keeps China behind until training sovereignty — but if recursive self-improvement kicks in, eight months "feels like a good enough gap" for a US takeoff. Meanwhile the talent pipeline is reversing: Chinese AI researcher returnee rates doubled from 12% (2019) to 28% (2025), China produces 2x the AI PhDs, and China-educated top researchers rose from 27% to 38% — while Chinese lab valuations run 40-120x revenue. Tommy considered buying GLM but hesitated at roughly 1,000x price-to-sales despite its claimed ARR ramp from ~$100M to over $1B in five months.
1. A distillation bear walks in, an innovation bull walks out
- AccelXR's opening confession: "When I first started this report, I was convinced that China was just doing benchmark maxing... but the deeper I dove in, the more I really came around to the other side that distillation just can't explain a lot of the things that they've shipped." Distillation explains a chunk of the capability and benchmark side; it can't explain what they've built under hardware constraints.
- The flip moment: MoE training was "relatively unstable" before Chinese labs, across a handful of papers, "showed that you can actually stabilize multi-trillion parameter sized models" — mixing US-originated techniques and porting them to frontier training scale. That, plus an AEI think-tank report on China's grid buildout, turned him.
- The research now flows both ways: Thinking Machines distilling on top of Kimi, and Hugging Face — during the OpenAI agent attack — reportedly resorting to a Chinese model to thwart attackers because of US frontier-model restrictions. Tommy's reaction: "So the US is now benefiting from China's AI research... that's crazy."
2. The league table: DeepSeek four of the top five, and why open source compounds faster
- His rankings: DeepSeek clear #1 ("MLA, MoE stuff... the first ones to do FP8 quantization during the pre-training pipeline"); Moonshot #2 — aggressively pushing sparsity to roughly 1:56, introducing MuonClip (now adopted or adapted by the other Chinese labs), KDA linear attention, and "good distillation" via MoPD; then Qwen/Alibaba; then Z.AI, less architecturally novel but doing heavy hardware co-design. Baidu and Tencent: fast followers.
- The velocity argument: DeepSeek R1's GRPO and RLVR spawned variants like DAPO and GSPO "within a few months" because the training recipes are open. "I know for a fact our labs aren't communicating to that degree... because it's proprietary."
3. Below CUDA: the co-design work that makes domestic training viable
- The mechanism, kept high-level: while most labs program GPUs through CUDA and its ecosystem, Chinese labs are handwriting low-level GPU instructions "at the assembly level" — Moonshot hand-built core GPU programs for a Chinese chip for Kimi K3, and Chinese labs also use RL to train a model to self-write and optimize kernels; DeepSeek handwrote low-level code with DeepEP.
- Why it matters: "That's not necessary in the US... we can just build on top of Nvidia." Optimizing for weaker chips means China can eventually "serve inference significantly cheaper than the US in the long term" — and, combined with quantization work and Cambricon/Ascend supply chains, become viable on the training side in the next few years. He thinks some people still get this wrong in assuming China can never train on its own chips.
4. US labs: capability lead intact, sane multiples, and a pivot from training to serving
- His US read: "It's hard to see the US labs slow down in capability frontier leadership" — Chinese labs stay fast followers until training sovereignty. And on valuation, US labs look "still relatively cheap" versus Chinese peers at 40-120x revenue; Tommy considered buying GLM but hesitated at roughly 1,000x price-to-sales, even as GLM's ARR ramped from ~$100M to reportedly over $1B in five months.
- The regime change: "We're in a scenario now where it's starting to transition from being able to train these models at frontier capabilities to actually being able to serve them at scale." Expect US labs to adopt Chinese-style efficiency research and shift capability gains toward post-training and RL through the Chinese labs' STR strategy — where "the scaffold or the harness for an agent can actually explain a large amount of capability gained" — rather than ever-larger pre-training runs.
5. Tokens as an electricity derivative — China's structural trump card
- The chain: efficiency gains get passed through as price cuts, not margin; with open weights, "theoretically there's no equilibrium where your per-token margin recovers." Electricity is already ~10-20% of token cost, so as chip capex, software margin, and R&D amortization are squeezed, "tokens converge towards almost an electricity derivative" — and then China's grid (2x US generation, adding capacity ~6x faster, AI needing 1-5% of five-year added capacity versus 50-70% in the US) dominates. His hedge, exactly as given: "I'm not entirely sure that it commoditizes to this extent."
- China's leadership sees it — he cites Ren Zhengfei of Huawei: China's power generation and grid transmission are very good, and AI development "requires this power guarantee." The chip answer is aggregation: a superpod of over 500,000 Ascend chips, 2.5x xAI's Colossus by chip count and ~1.3x its total compute — weaker chips, more of them, until "it comes down purely to electricity bottleneck."
- AccelXR says US misinformation about data centers is slowing deployments, while China runs "East Data, West Computing" — siting compute on cheap renewables stacked with tax and energy subsidies. One constraint he concedes: per the AEI report, Chinese chip production stays supply-constrained; he believes the ratio may be roughly 5:1, with production still insufficient for domestic demand. That is exactly why the co-design optimizations matter.
6. Monetization, the price war, and three endgame scenarios
- The price war is real — Tommy's Silicon LLM Index, a weighted cost-per-million-tokens measure, fell from 2.2 on June 2 to ~1 — but monetization is dispersing: Moonshot now requires model-as-a-service providers above $20M revenue to sign separate agreements, Alibaba is adding a revenue threshold, Z.AI is looking at on-premises sovereign deployments, and DeepSeek added peak pricing and raised its rates. The open question: can they justify 6x-plus the revenue multiples of US peers? Meanwhile agentic workloads are "extremely cost sensitive," driving a composition shift toward "cheapest but still adequately capable models" — exactly where Chinese labs are optimizing.
- His three scenarios: today's "functioning bridge" (US frontier APIs own regulated workloads, Chinese open weights own the cost-sensitive mid-market, western enterprises hedge by running both); "bridge demolition" via US cloud/distillation restrictions or Beijing cutting exports — his pessimistic default; or the US "competing through openness" with cheap frontier-adjacent open weights — his ideal, but "it feels a little optimistic."
- The light-switch evidence: Xi's speech encouraging open source came about three days before the Financial Times revealed that MOFCOM was consulting labs on restricting open weights. The good strategy, in his telling: "dump cheap open weights on the market to slow down the US... until China can catch up itself, and then you flip."
7. Reverse distillation, the eight-month gap, and reversing talent flows
- Tommy's spicy question — odds that China trains a superior frontier model and US labs distill it — gets a structural answer: "A US lab's best distillation teacher legally is the Chinese open-weight models." AccelXR would not be surprised if some smaller US labs lean in; regarding Anthropic, he says he is not sure it would. Tommy's rejoinder lands: "They are creating god though."
- The gap math: consensus is roughly an eight-month capability lead, and since "the student can't surpass the teacher model," distillation freezes the status quo until Chinese training sovereignty. Is eight months enough for a US recursive-self-improvement takeoff — Anthropic's "million geniuses in a data center"? "If it gets self-recursive, I think so. It feels like a good enough gap."
- The talent underpinning is shifting: foreign-PhD Chinese AI researchers returning home doubled from 12% (2019) to 28% (2025), China produces 2x the US AI PhD pipeline, and the share of top AI researchers educated in China rose from 27% (2017) to 38% (2024). Tommy's summary: "It's just so hard to be bearish on China when you talk about them in any domain."
8. What to watch — and the paradoxical case for selling China Nvidia chips
- His flipping-point indicator: "clear disclosure of actual full Ascend usage for pre-training at trillion-parameter scale." People conflate serving on domestic chips with training on them — DeepSeek tried training on Ascend with Huawei engineers on site, hit instability and slow interconnect, and reverted to Nvidia. A confirmed multi-trillion-parameter Ascend-trained model would be "the biggest confirmation" of a coming policy shift; secondary signal: whether Chinese AI cloud pricing normalizes in the second half of this year after Alibaba, Tencent, Baidu, and Z.AI all raised prices.
- The policy paradox (he notes Sacks has made a version of this): loosening exports and getting China back on Nvidia "should slow down their development... it slows their sovereignty push, and any kind of tightening actually accelerates their push." China gets it — regulators are slow-walking H200 imports and encouraging state-funded data centers to swap foreign chips for domestic alternatives, even after the US resumed H20 sales with revenue sharing in January of this year and H200 sales.
- On the US side he expects labs to lobby to wall off the ecosystem ("it just makes fundamental sense from a business perspective"), possible US-government ownership stakes in frontier labs, and possible government positions in discoveries those labs produce. The closing security paradox — if companies cannot access frontier models, they lose the ability to defend against them when released — is why Tommy stakes his flag: "That's why I'm bullish open source... it's the only thing that makes sense to me."
Full transcript
China is just absolutely trouncing the US right now. When I first started this report, I was convinced that China was just doing benchmark maxing, but the deeper I dove in, the more I came around to the other side: distillation just can’t explain a lot of the things that they’ve shipped.
Wow. So the US is now benefiting from China’s AI research?
Yeah. By 2028, we have a mass-produced, cost-competitive Chinese AI stack.
1. Is China Innovating or Just Distilling?
Wow. Hey everyone, welcome back to the Deli podcast. I'm Tommy, one of the founding partners at Deli Ventures, and today I'm really thrilled to have AccelXR on. He wrote like the single best read I've read in months on the AI side. He did an absurd amount of research and covered all of the AI labs in China and everything going on in the US to give all of us a view into if there's actually more innovation going on, more distillation and just the state of everything. AccelXR, how are you?
Good, good. How are you?
I have a pointed question before we jump in. Before doing the research on this report, did you think China was doing more innovation or more distillation?
When I first started this report, I was convinced that China was just doing benchmark maxing. But the deeper I dove in, the more I came around to the other side. Distillation just can’t explain a lot of the things that they’ve shipped.
It kind of breaks down. You have the capability side, which I think distillation explains a large chunk of, and the benchmark side. But when you start looking at the hardware constraints that they’re working with and the innovations they’re making around them, the research going on there is really quite impressive.
Does it shock you at all that everyone on Twitter—or most people, and our entire media in the US—just argues that China is stealing our weights and distilling our models, and that there’s no innovation? Did that surprise you by the end?
Yeah, definitely. Like I said, I started with that mindset as well. I was in that camp.
But it’s quite impressive, some of the things that they’ve done, because of the innovations they’re making specifically in reinforcement learning and mixture-of-experts work. A lot of that flows back to the US as well. You see, for instance, Thinking Machines’ Tinker using distillation on top of Kimi. It’s flowing back and forth more than I think people give it credit for.
Wow. So the US is now benefiting from China’s AI research?
Yeah.
That’s crazy. I didn’t even know that was happening. There was one event recently where OpenAI used GLM or something to verify an attack that was going on internally, but that was the latest I’d seen.
I think what you might be referring to is that, during the OpenAI attack on the agent side, Hugging Face had to end up using a Chinese model to try to thwart the attackers because of the restrictions on our own frontier models. That’s a conversation in and of itself.
It’s crazy. One of the main questions I had for you was this: You’re going into the research with the view that China is stealing our weights and distilling them. At some point during your research, you probably read something that changed your mind, right? You thought, “Maybe this isn’t all just distillation. Maybe they’re doing some real things.” What was that piece of content or writing that changed your view?
I think there were a couple. On the hardware side specifically, there’s a report from AEI, which is a public policy think tank here in the US. They went pretty deep on the actual compute infrastructure that China is building out and where they are from a grid-capacity standpoint. That was one area that really opened my eyes to what’s going on on the hardware side.
On the actual innovation for the models themselves, I was already somewhat familiar with this from the early DeepSeek paper that shocked everyone when it first came out. But it was really spending more time with some of the more recent papers.
I’d say the big one for me, of all the innovations—well, there are a couple—but one of the larger ones was seeing the progress on mixture-of-experts work. Prior to the Chinese labs focusing on this, it was relatively unstable. The Chinese labs took that in a handful of papers and showed that you can actually stabilize multi-trillion-parameter-sized models.
That’s when it clicked for me: The progress they’re making, even if it’s based on some US research, is really about mixing all these techniques and porting them over to make them viable at frontier-training scale. That’s really where it flipped for me.
I’m glad you bring up the DeepSeek report. I had Jeffrey Emanuel on the podcast 1 or 2 years ago when DeepSeek released its paper. It sent the market down by hundreds of billions of dollars, and everyone was saying, “They just distilled our models.”
2. DeepSeek and China’s Leading AI Labs
I started reading the DeepSeek papers, and 2 of the things that stood out were that they really innovated on mixture-of-experts—turning on pieces of the model at each time because they were using second-rate hardware—and, second, that they were using second-rate hardware to train. It definitely felt like there was real innovation going on, and that’s why I’ve always held that there’s more innovation than distillation in China. But I think your report really did the hard work and figured it out, which is good.
When you look at them, DeepSeek is definitely the leader on the research side. If I had to rank the innovations, I’m pretty sure DeepSeek takes 4 of the top 5 as its own findings.
What’s interesting is how quickly it accelerates because of how open source it is. One example is GRPO and RLVR, which, not to get too technical, are reinforcement-learning techniques coming out of DeepSeek, specifically DeepSeek-R1. Within a few months of their release, you had a variety of variants come out, like DAPO and GSPO.
It just shows how quickly things can accelerate when the training recipes are actually open-sourced. That gives them a leg up. I know for a fact that our labs aren’t communicating to that degree or extending each other’s research to that degree because it’s proprietary.
I’m curious for you to rank your list of the most innovative firms in China. I know you mentioned DeepSeek would be number 1. Who would you have as 2nd, 3rd, and 4th?
DeepSeek’s definitely number 1. They did the MLA and MoE work. I believe they were the first ones to do FP8 quantization during the pre-training pipeline, so they’re for sure number 1.
Number 2, I would probably say Moonshot. Interestingly enough, they’re really aggressive about pushing how sparse you can make the models in that context. I think the most recent release, off the top of my head, is 1:56 sparsity, which is huge. They were also the ones that brought out MuonClip, which is now adopted or adapted by all the other Chinese labs.
They did the KDA linear-attention work and were also innovating on the distillation technique—not the bad distillation, but the good distillation—via MoPD. I think those 2 are the big ones.
Qwen, from Alibaba, does a lot on this side as well. To a lesser extent, there’s Z.ai. They do less architectural-novelty work, but they’re doing a ton of work on the hardware–code-design side.
The Chinese labs are dealing with really constrained bandwidth limitations for the hardware they’re able to get. With that being said, what Z.ai is doing with the GLM family is co-design work with Cambricon and Ascend. They’re trying to develop below CUDA and actually work with the Chinese chips and the limited capacity they have.
I think the big breakthrough on the training side is going to come from this code development. That’s how I’d situate them. A lot of the others, like Baidu and Tencent, I’d put down as more like fast followers.
That’s interesting. I know DeepSeek is on top, but I probably would have thought, off the top of my head, that GLM would have been ranked a lot higher. It’s interesting to see.
Like I said, they do great stuff on the code-design side. It’s just a little bit less on the architectural novelty. They adapt more of the techniques being developed at those other labs, I would say.
3. How China Is Working Around Weaker Chips
You mentioned co-design work—the AI lab working to marry the software to the hardware in a pretty unique and form-fitted way. What does that actually mean? I’d love to walk through what that means and what you think the implications would be for these AI labs in China.
Keeping it relatively high-level, most labs are programming their GPUs through NVIDIA’s standard software tools, which are CUDA and that ecosystem.
The Chinese labs have actually been going beneath that and hand-writing low-level GPU instructions, exploiting some of the chip’s behavior at a lower level. One big example here was Moonshot and Kimmy K3. They hand-built their own core GPU programs that work with the architecture for a Chinese chip instead of relying on off-the-shelf libraries.
Interestingly, as a side note, they also use reinforcement learning, or RL, to train a model to self-write and optimize these GPU kernels. So they’re kind of accelerating AI by building AI to a degree as well.
Another example would be DeepSeek. They also hand-wrote some low-level GPU code called DeepEP. They’re essentially going beneath the libraries on top of the GPUs to write to the GPUs at a lower level. That’s not necessary in the US. Here, we don’t need to worry about that as much because we can just build on top of NVIDIA.
They’re essentially trying to optimize for their weaker chips, and I think this gives them a strength in the long term when it comes to being viable to do training themselves in China, on their own stack.
If I have a mental model where CUDA is the programming language for NVIDIA at the top, and then I have the silicon—the real physical hardware—at the bottom, they’re going below CUDA, obviously above the hardware, and messing with machine-level code.
Yeah, exactly. Below CUDA, you have frameworks at the top, like PyTorch and vLLM, then CUDA below that. They’re writing at the assembly level, essentially, for GPU code.
Interesting. So if they’re already doing this with NVIDIA chips, what you’re saying is that, with China-native chips, they can leverage that and take it to another level. I’m just wondering about the implication, basically.
Yeah, I think the biggest implication is that we’ll likely see this code development lead to a scenario where they’re optimizing at levels that our companies don’t need to. They’ll end up in a situation where they’re able to serve inference significantly cheaper than the US in the long term.
One of the big things I’ve been watching is how capable the Chinese labs are of training on their own domestic chips. I think that, in the long term, this co-development work and the quantization work leads to both being more efficient on the inference side and becoming viable on the training side as well within the next few years.
I think some people still get this wrong. They believe that China can never train on its own chips because of these limitations. But this type of work, alongside the bolstering of their own supply chains with Cambricon and others, will ultimately lead to them being viable within the next few years.
I want to continue talking about the Chinese side, but I also want to talk a bit about the US labs. As we go back and forth, I want to get your view on how this all shakes out.
On the Chinese side, the conclusion I’m drawing from what you’re saying is that there’s real innovation here from the labs. They’re doing some pretty crazy stuff, and they’re going to start building their own chips.
Going back to the US side for a little bit, one of the things you told me before the podcast was that Anthropic is doing really well. They’re getting a lot of hate, but the models are great. How are you feeling about US innovation, the labs we have here, and what you’re seeing? How do you feel about them?
It’s always hard to pick it apart because we’re just not as transparent about what’s going on under the hood. But from a capability standpoint, it’s hard to see the US labs slowing down in capability-frontier leadership. I think the Chinese labs will ultimately always be fast followers until they’re able to get some kind of sovereignty on the training stack.
In the US, even just looking at the multiples, the ARR numbers are crazy for Anthropic and OpenAI—absolutely insane. Even at these huge valuations, you’re talking about relatively cheap multiples versus what’s going on on the China side, where you have multiples between 40× and up to 120× revenue.
I was going to buy GLM a couple of months ago, but it was at a 1,000× price-to-sales multiple. I didn’t know if I could do that.
GLM-Zhipu is interesting. Its ARR ramp in the last 5 months has been kind of insane. They went from 100 million to, I think, over 1 billion now, according to what they’re saying.
Well, I guess if we don’t have the benefit of open source in the US to understand what the labs are doing, maybe from a higher level, are you confident that the trajectory OpenAI and Anthropic are on—more data, more compute, more training, the Bitter Lesson style—will continue? I guess then we could work backward to determine whether they’ll continue to be impressive.
I think realistically we’re in a scenario now where it’s starting to transition from being able to train these models at frontier capabilities to actually being able to serve them at scale. I wouldn’t be surprised if the US labs start implementing some of the research coming out of the Chinese labs in order to scale how much they can serve.
We already see it with service outages and things like that. There’s just not enough compute to run these huge models, so they need to start optimizing for actually serving a large client base. OpenAI came out with its new Flash model, which is significantly cheaper, and I think that’s where progress is going to start heading.
The Chinese labs use the STR strategy. Instead of focusing so much on the pre-training side, they’ve been trying to shift more of the capability gains to post-training. A lot of the work they’re doing on the reinforcement-learning side is an example of that.
That’s been broadly adopted. For instance, you see papers coming out now about how the scaffold or harness for an agent can explain a large amount of capability gain. That’s essentially post-training work. I think we can start heading that way more, rather than focusing so much on the pre-training side and larger and larger models.
4. China’s Biggest AI Advantage: Electricity
I like that. I feel like a lot of my own increase in capability has just come from massively designing my news-research Hermes agent to run a lot of my life, so I understand. Exactly—build the harness. Now I can’t be without it. That’s how it is.
It’s interesting, and I guess let’s flip back to China. You had a really interesting take on electricity and AI models. I don’t want to give it away; I want you to describe your thesis here because I thought it was really solid.
If we think about what I was just saying—that the efficiency frontier is becoming where we’re competing, rather than raw capability—over the past few years, all of those gains get passed through as price cuts. They’re not typically retained as margin, although we are starting to see a shift with some of the pricing changes and other things that have been going on.
At the end of the day, you’re taking market share through price cuts for all these capability gains. If you have open weights, theoretically there’s no equilibrium where your per-token margin recovers unless everyone collectively agrees that prices should go up. So it starts becoming lower and lower.
Right now, my understanding is that electricity accounts for probably 10% to 20% of token cost. The other expenses include chip capex, software margin, and R&D amortization. But as they get squeezed by open-weight price wars, tokens converge toward almost an electricity derivative.
When that happens, I think China has a very dominant advantage over the US. China’s grid generates twice as much electricity as the US’s, and it’s adding capacity roughly 6 times faster than we are.
Wow.
China’s AI-related power needs are only 1% to 5% of the capacity it has added over the past 5 years. In the US, it’s 50% to 70%. We have a really constrained grid, whereas China is adding capacity much faster.
If you think about tokens being an electricity derivative, it makes sense for China to export its electricity through these types of models. As a Western citizen, I want us to do well, but it’s very jarring when you look at how much they’re winning on the power-generation side.
Generally, people understand that we started with the bottleneck being the GPU, then it was the memory. More and more people are starting to realize that it’s actually grid capacity and energy.
If you think about it through that lens, China is absolutely trouncing the US right now.
That's fascinating. Let me feed this back to you: your concept is that if models commoditize down, the country that wins is literally the one with the most electricity to train and serve them.
Yeah. Even Chinese leadership recognizes this. For instance, Ren Zhengfei of Huawei was saying essentially that China's power generation and grid transmission are very good, and that the development of AI requires this power guarantee. China has been using a technique where, instead of trying to match the capability of the chip directly—chip for chip—their chips are obviously weaker, but if they can add more of them into these superpod clusters, what you get to is something that comes down purely to the electricity bottleneck. In that scenario, they will take the cup from the US.
That said, just to hedge slightly, I'm not entirely sure that it commoditizes to this extent, but it's the natural progression that you would envision if open weights continue to be the default and continue to put all this pressure on the token price itself. Which is part of the reason why I think policy probably shifts this scenario. Tommy
The crazy part—and I feel like it's all Chinese propaganda—is everybody in the US getting annoyed about data centers using too much water when they don't, right?
It seems like there's a ridiculous amount of misinformation about data centers in the US right now. I don't know where it's coming from, but it's honestly stoking a certain class of people to agree with it, and it's slowing down data center deployments. It's unfortunate. The Chinese government invests really heavily into its grid capacity. They even have a program or idea called East Data, West Computing, where they try to site their compute on cheap renewables and stack it with tax and energy subsidies. I just feel like our population would not be okay with that right now, you know?
No, I have one of the trackers I do at Hermes that maps data centers across the US to track whether they're actually getting built. It's crazy, state to state, what we're seeing. One question I had related to this is: they might have more electricity, but does China have the breadth and depth of data centers that we have in the US? I know they don't have as many top-tier chips as we have, but do they have the same number of data centers, or how do you think about that?
I don't know the exact data center comparison offhand, but I do know they are expanding these superpods or superclusters, as I mentioned. These are taking a ton of Ascend chips and combining them together. The supercluster itself has, I think, over 500,000 Ascend chips, which is 2.5 times what xAI's Colossus will have by chip count, and it will equal about 1.3 times the total compute of xAI's Colossus. They're pushing on it for sure. Here in the US, we have all these neoclouds and so on that aggregate compute as well. I'm not sure what the 1:1 comparison would be there, admittedly.
That's fair. Basically, my question is: they might have more electricity, but do they actually have the chips to leverage it?
Yeah, interestingly enough, the AEI report that I mentioned earlier talks about this as well. Eventually, over the next couple of years, their production of chips is still going to be supply-constrained. I'm trying to find the number offhand, but I believe it's a 5:1 ratio, where essentially they won't have enough chips to satisfy Chinese domestic demand themselves within the next few years either. That's part of the reason why these hardware co-design innovations are so important: they need to ramp up production as much as they can. Even in that scenario, they're not going to have enough, so they need to make these other optimizations on top of it.
Damn. Taking it back to dollars and cents, the price war has been absurd. I posted a June 2 thesis when the Silicon LLM Index was at 2.2. It's the weighted cost per million tokens, and now it's down to, I think, 1. It's fallen a lot. I think you have some pretty solid pushbacks on that index in particular, but I'm wondering: do you think Chinese pricing will continue to fall? I don't know how we want to measure it—cost per million tokens or cost per task—but I'm curious whether you think it'll reverse or keep going.
Yeah, I think there are 2 sides. One side is: can they monetize what they're doing? I think that's been the more pressing question lately.
On the Chinese lab side, you have license changes. Moonshot's Kimi K2 requires any model-as-a-service provider above $20 million in revenue to sign a separate agreement. They're essentially licensing the model out. Then you have Alibaba, which has its open-weight flagship, but is also adding a revenue threshold.
There's a dispersion happening right now, where some models are becoming more aggressive on licensing to try to capture revenue on that side. You have companies like Z.ai, with the GLM family, looking at on-premises sovereign deployments. That's how they've been trying to generate revenue instead of monetizing their open-weight models so directly. The big question is whether they can generate enough revenue to justify 6x-plus revenue multiples versus their US peers.
You also see DeepSeek recently adding peak pricing and raising its rates, so it'll be interesting to see whether it has the pricing power to justify this. Another area to consider is not so much the revenue side, but a composition shift. If you expect agentic use cases to grow over time, which I think most people envision for the future, those use cases are extremely cost-sensitive. If that's the case, you have this composition shift toward the cheapest but still adequately capable models, and I think that's where they could really lean in.
I'd love to see the US compete more fully. If we shipped either open-weight models or Flash models that could service this use case as well, that would help. The Chinese labs are definitely optimizing specifically for agent-driven workflows. It'll just be a matter of whether they can capitalize on it.
5. Three Scenarios for the US-China AI Race
That is interesting. When I speak to people who have spent a lot of time in China, I've gotten pretty bullish on Alibaba recently and have gone back and forth on that, but people always tell me to be careful because the Chinese government won't let these companies accrue tons of value, right? It's confusing to me because, on one hand, I agree with you that these companies should deploy their large private models through the neoclouds, Microsoft, or Amazon, and let their customers fine-tune them or access them, sending money back to the Chinese labs. On the other hand, it seems like China's government may have other goals in mind. What do you think about how their government views this?
Yeah, it's an interesting question. I think there are 3 scenarios for how the US-China AI market relationship plays out. Right now, we're in the default scenario, where US frontier APIs dominate a lot of the regulated and high-stakes workloads, while Chinese open weights become the default for the cost-sensitive mid-market. Western enterprises, I would expect, run both: either they self-host the Chinese weights or go through neoclouds, as you suggested, or they use the US side while essentially hedging against any US gating or Chinese hosting risk.
In this scenario, the Chinese labs do well. I would say they're able to have some kind of lock-in, actually, globally; price competition holds down some of the mid-market costs, and so on. The 2 future scenarios are that you either have a complete bridge demolition—a breakdown—which is kind of my default assumption. It's pessimistic, but I think the US could either impose cloud-rail restrictions or issue distillation rulings saying that enterprises can't use Chinese weights. Alternatively, Beijing could potentially consider removing support for exporting any frontier Chinese models once it has the training capability to ship frontier weights itself. In that scenario, you have a breakdown, and US labs end up getting all the mid-market pricing power back, which is good for supporting frontier R&D.
You kind of lose this competition, which is good for buyers of tokens. The third scenario would be a RAND-style prescription of competing through openness. This would be my ideal outcome: the US labs end up shipping cheaper, frontier-adjacent open weights. You have competition on this side for the mid-market, and the US labs can focus more on providing frontier use-case access for things like drug discovery or any of those very intelligence-reliant use cases.
To your earlier point, what does China think about this? I think Xi gave a speech not too long ago where he positioned it as, “We want to provide any developing country with international AI cooperation. We want to help you guys out, and we want to encourage open weights,” and so on. But at the end of the day, I think the good strategy is to dump cheap open weights on the market to slow down the US side of things until China can catch up itself. Once China has domestic training sovereignty, you flip and close off access to these models. That’s the long-winded answer, but hopefully it gets to what you’re asking.
No, I like the three scenarios. Maybe to linger on your third scenario for a little bit, it seems hard for me to reason about the US embracing open source, given how much money has flowed into OpenAI and Anthropic. But I also don’t want to take the position that we exist to help those companies survive, right? It is hard for me to reason about that side.
One of the really interesting things you brought up is this light-switch moment where China is going to export open-source models as long as it can, because it hurts US markets a bit and pressures the labs’ margins. Then, the second China can train frontier models, it says, “You’re cut off. You can’t export any open source anymore.” That seems like a pretty interesting potential future. Do you put a lot of stock in that?
I would say so. For instance, the Xi speech that went viral, where he was encouraging the open-source thing, came about 3 days before the Financial Times revealed that MOFCOM—the Chinese Ministry of Commerce—was consulting labs on restricting open weights. On one side of his mouth, he’s saying, “We want to encourage open-source innovation,” but on the other side, they’re already deliberating over how much they should push on this side.
I really think that as soon as they have the frontier capabilities themselves, there’s not really a reason to outsource all their research and so forth, similar to what we do here. We don’t outsource it for the same reason. Geopolitically and strategically, I would say that’s the most likely outcome: you embrace open source and open weights for now, until you reach some degree of parity in frontier capabilities. Then it’s game on in competing on that side.
All right, I have a potentially spicy question. What percentage chance do you put on China not only training a frontier model that tops US models, but US models then trying to distill it?
This is interesting. As I mentioned earlier, we have some reverse distillation going on—Thinking Machines Lab with Kimi. It just makes sense to do that, because open publishing subsidizes all of your open-source competitors. There’s a legal asymmetry here, right? A US lab’s best distillation teacher, legally, is the Chinese open-weight models. So it makes sense to do this reverse distillation.
I wouldn’t be surprised if we continue to see some of the smaller labs in the US really lean into this as well. On the frontier side, I don’t know if I could see Anthropic reverse-distilling a Chinese model. It makes sense—it’s cheap capability gains if they do get ahead of us—but I’m not sure.
They are creating God, though. I don’t know if they want to—
Yeah. Yeah. [laughter]
One of the things I liked about your report, too, was the three scenarios for the US and China. The first one you put out is the functioning bridge scenario, which you just described. It’s kind of like we keep going with what we’re doing.
It seems like chaos always resolves one way or the other. This tit-for-tat game has to resolve one way or the other. We develop AGI in the US, or China trains frontier models. I’m not sure, but it seems like one way or the other, it has to swing.
I feel like the status quo is not stable. You see it in how both countries talk about their policies here. It just does not feel stable.
I do feel that it either flips to us actually embracing open source—which feels harder, like you said—or it breaks down. I’d love to see us embrace open source; that’s my ideal outcome. But it feels a little optimistic, and pessimistically, I genuinely feel that it breaks down over the next couple of years.
I have a question that underpins how these scenarios play out. One of the things that both of us have historically been bullish on, given our work in crypto, is how open source compounds and how you can build on each other’s creations.
You talked about this earlier, but I’m trying to figure out the velocity of the open-source side in China building upon itself and creating new things, versus OpenAI, Anthropic, and others, where they’re doing it within their companies but not outside them. It’s hard for me to put a speed on each, if that makes sense.
Yeah, I agree. I think there’s been a good push recently, though. There have been a lot of developments here. I think it’s interesting when you have companies like NVIDIA releasing Nemotron—essentially, a hardware vendor itself open-sourcing a model.
That was crazy to me: competing with their customers. It’s nuts.
Yeah, and it’s kind of in line with what I was talking about earlier with the co-design stuff. NVIDIA is best positioned to extract maximum throughput for its models off its own tech stack. It’s interesting that they’re going to attack their own customers as a cost advantage here.
If I had to say what would be cool to see on the Western open-source side, it would be pushing on openness standards themselves. If we started using more permissive licensing, providing some of the training data and recipes, and providing RL environments like NVIDIA did for Nemotron, it weaponizes auditability.
I don’t have a strong stance on this—I didn’t dig into it as deeply—but I do wonder to what extent China’s labs might not want to fully disclose the filtering and other processes in their training sets. We could almost have them cede the entire narrative if we really pushed on this side.
It’s hard. Being in crypto yourself, you want to do open source, but then it becomes a whole game of how to monetize on top of it. Maybe they all just need tokens at the end of the day.
6. America’s AI Talent Advantage Is Reversing
Well, maybe taking one step back from the innovation, open source is about the people actually doing the innovation—thinking about these things, hitting walls, and coming up with crazy ideas. You had a slide in the report on talent flows, and that’s really interesting because it underpins everything.
I don’t know what the best question to ask you is here, but I’m curious what you found from tracking talent in the US versus China.
One of the big numbers on that slide to me is Chinese AI researchers with foreign PhDs returning home. The returnee rate was 12% in 2019, and by 2025 it was 28%. So we have a doubling in 6 years.
Wait, so these are foreign PhDs in the US going back to China? It was 12% in 2019, and now it’s 28%?
Chinese AI researchers specifically—not the entire PhD market—but yes, there’s this talent flow, which has always been America’s advantage. We import talent, and those people are more likely to stay here after working in universities and so forth. But we’re starting to see some signs of reversal here.
China’s domestic pipeline has just been growing on the talent side. It produces 2 times the number of PhDs that the US pipeline does, which is partly a function of how large China is versus the US. But it’s hard to argue with those types of numbers.
A single university?
Yeah.
So, internally within China, they’re producing 2 times the number of PhDs—
On the AI side.
2 times the PhDs.
Over the past few years, in 2017, 27% of the top AI researchers were educated in China, and now it’s up to 38% in 2024. There are a lot of examples like this where you’re starting to see both the domestic talent supply improving and the people who do ship out to the US to train and become educated beginning to return at greater rates than historically.
Wow. It’s kind of hard to be bearish on China when you talk about them in any domain.
Yeah, like I said, I started the report as a bear, and it kind of flipped after going through a lot of this stuff.
It’s crazy. It’s nuts. Maybe let’s flip back to the U.S. side for a little bit. There still obviously is a gap between Claude 3.5 Sonnet, o3, and other frontier models and what we have in China. There definitely is a peak-intelligence gap.
People have argued that the gap is narrowing, and we can argue that the gap may or may not exist in the future, but right now it exists. I’m curious about your view on a scenario where that gap leads us to take off toward AGI, and all these talent numbers and electricity in China don’t matter because we get there first. Do you subscribe to that view at all, or not so much?
Yeah. I’m a big fan of the recursive self-improvement stuff that we’ve been looking at more closely in the U.S. I think if you get the scenario where—I forget the exact terminology that Anthropic uses—but it’s like having a million geniuses in a data center, where you have all these AI researchers that are AI themselves, kind of self-improving the whole system, that’s the holy grail. That’s the takeoff scenario.
I think we have a leg up. The gap is relatively wide. I think the consensus is somewhere around an 8-month gap in capability. Because they do distill—I’ve never argued that they don’t do distillation. I’m pretty positive they do distillation for capability—the student can’t surpass the teacher model. So that gap will stay the status quo until they reach training sovereignty, when they can actually close it.
Is 8 months enough for us to get up to speed? If it gets self-recursive, I think so. It feels like a good enough gap. At that point, it just becomes a question of how quickly we can scale up the number of AI researchers, and that gets back to the question of grid capacity and compute capacity, which we do lead on today in terms of compute capacity.
Over time, they’ll definitely be able to catch up, so it’s just about keeping them at arm’s length. If I were a U.S. politician, I think one of the best things we could actually do—I think Sacks has even talked about this to a degree—is get them back on the NVIDIA bridge. It should slow down their development, kind of paradoxically. Any kind of tightening actually accelerates their push to build out their own chips and become self-reliant.
That is interesting, because most people want NVIDIA chips not to be sent to China, and here we are saying they should be, because otherwise they’ll make their own chips.
Yeah. I think China recognizes this. On one of the slides, I go through the export controls that we’ve implemented, and you can see that during the Biden administration, we banned the A100 and H100. We closed out the H800 loophole. After the Trump administration came in, we began rescinding some of the diffusion rules. We resumed H20 sales with a revenue share in January of this year, and we resumed H200 sales.
But China, paradoxically, is not taking the chips. They’ve actually gone the opposite direction, where they’re now encouraging any state-funded data center to swap foreign AI chips for domestic alternatives. The regulators are slow-walking and conditioning any H200 imports.
I think they’re doing this because ultimately they want to get off their reliance on this bipolar policy that we have on chips and become self-reliant. So, paradoxically, if we could loosen restrictions and get them back on NVIDIA chips, it could slow down some of the sovereignty push, which gives us a little more control over the situation.
7. China’s Path to AI Sovereignty
Yeah, it is really interesting to think through this tit for tat and where it goes. Whenever I talk to really smart hardware folks, or people creating alternative chips, they always argue that it takes 5 to 10 years to get that massive piece of hardware you use to build the chips. It’s massive and really expensive.
But from talking to you, it seems like it’s a lot sooner than we think for China to make its own chips.
Yeah. I would put the number at—I tend to agree with the AEI report—by 2028, we’ll have a mass-produced, cost-competitive Chinese AI stack.
Wow. At that point, they’re still not projected to reach the total FLOPs output of the U.S. I think by 2028, it’s only somewhere around 10% to 15% of the total FLOPs output.
But that’s part of the reason why I was saying that if the axis of competition has shifted from the training side to the serving side, all these innovations that the labs are working on actually make them better situated even under that scenario, where they have significantly less FLOPs output.
That is nuts. I’ll let people read through your report, because you have a lot of really good technical points across a lot of domains, including training and inference, that support everything you’re talking about right now. But I’m curious, in closing, what were your unanswered questions? Where do you want to take this next, and what are you watching? What were you curious about but didn’t have time to get to?
I only spent one of the slides on the training-sovereignty question, which I know we’ve touched on quite a bit, but like I said, it’s the flipping point in my mind. The big things to watch are clear disclosure of actual full Ascend usage for pre-training at trillion-parameter scale.
I think people get this confused sometimes. They’re serving on domestic chips, but they’re not really training on domestic chips. DeepSeek tried to train one of its models and even had Huawei engineers on site to help them do it, and they still ran into all this instability and slow interconnect. Ultimately, they reverted back to NVIDIA.
The first time we see a multi-trillion-parameter model trained on Ascend chips, that’s the biggest confirmation that we might have a huge policy shift from there. There are all kinds of other things to watch, like Chinese AI cloud pricing. In the beginning of this year, over the first 6 months or so, Alibaba, Tencent, Baidu, and Z.ai all raised pricing.
Any kind of normalization in Chinese AI cloud pricing in the second half of this year would mean that the domestic supply chain is actually delivering enough. So, there’s a lot to watch on the hardware side. Interestingly, I started the report focused predominantly on the AI training side and the software component, but I think the bigger question is actually the hardware.
I might spend some time in one of my next reports covering more of the hardware supply-chain side.
That’s awesome. I’m curious about your view on where you think the U.S. goes. Anthropic is going to IPO soon, and its business model is killing it on revenue. I’ve argued that it should own a percentage of the technologies, medicines, and other things it creates.
I’m wondering how they do and how things have been changing there, because there’s a lot of hard research to do on the balance-sheet side of this: the capex buildout, the AI side, and the off-balance-sheet commitments. This stuff is hard.
Open-source models reverberate through the supply chain, because if you’re spending less for an open-source model—$4 on GLM versus $5 on Opus—it changes the flow-through and how much money everybody earns. I’m curious where you think the U.S. side goes.
There are a couple of ways of looking at it. To your last point, I think the Western providers need these API margins to work out in order to fund the frontier R&D. They’re exposing prices set by labs that have Chinese money backing, which our labs don’t currently have.
I wouldn’t be surprised if we see ownership stakes from the U.S. government in our frontier labs. I’m also interested in the idea that they should maybe look at taking positions in some of the discoveries that those labs produce.
My previous report was entirely on the self-driving labs side, using AI to automate scientific research. That angle, which it seems like all the labs are pushing more toward, gives them monetization through things like IP generation. That’s one of the interesting things they could do on that side.
I expect them to continue lobbying the U.S. government to wall off the ecosystem relatively quickly. It makes fundamental sense from a business perspective. You don’t want this competition, and it becomes a national-security issue.
I was just writing up something on the Hugging Face hack. It will be interesting to see if AI labs are actually going to put their money where their mouth is on slowing down some of the developments because of how spooky some of that is getting on the security side. But it's kind of paradoxical, right? Like, if you—
If you slow down the development,
You get attacked.
Yeah, you get attacked. How do you release these models? If you're not a company that gets access to them, you lose the ability to defend yourself against them when they're released. So—
That's why I'm bullish on open source. That's why I think it's just the endgame. It's the only thing that makes sense to me.
It is crazy.
Okay, we release the frontier model open-weight, and anyone can run it now. Now it becomes a race immediately after it's released to harden all the infrastructure, you know. It's just a crazy scenario, like a game. You just picture Trump in the Oval Office using it, like, “Harden all national security. Make no mistakes.”
Make no mistakes. Yeah. AccelXR, it's awesome having you on. Your reports are second to none. They're incredible.
Awesome. Thanks for having me on, Tommy.
I'm excited to host you again. Thanks for making the time.