主权AI:为什么各国都在打造自己的模型
- 主权AI正在打破云时代“全球工作负载将集中于美国和中国”的假设。 王国宣布推出HUMAIN——一家本地超大规模云服务商或AI平台,计划让绝大多数AI工作负载在境内运行;已宣布的集群建设规模约为1000亿美元至2500亿美元,其中500兆瓦似乎会成为“原子单元”。目标是实现“基础设施独立”,包括对模型和部署拥有自主权。
- “AI工厂”在技术上不同于传统数据中心。 GPU是主要的有源部件差异;高密度集群还需要机架级液冷、接近电厂规模的能源供应,以及对能源供给的提前承诺。企业也可能绕过复杂的云堆栈,采用Kubernetes,再按需选择Snowflake或数据库类服务。
- 模型已经成为文化基础设施和信息基础设施,对外国技术的依赖因此成为国家脆弱性。 训练数据会嵌入价值观,后训练则决定模型回答什么、拒绝什么;与此同时,基础模型已经触及国防、医疗、金融,以及ChatGPT约5亿月活用户的日常决策。随着模型取代搜索、批改小学作业,控制模型的人可能塑造被社会接受的历史与真相。
- AI数据中心类似工业时代的石油储备——关键区别在于,各国可以自行建设它们。 资本和政治意愿能够创造算力底座,并以此发展国内产业、推动发展和构建出口能力。
- 美国面临一个选择:帮助盟友建设主权能力,还是把这片市场留给中国模型。 Midha偏好的类比是“AI版马歇尔计划”:最初的重建计划看似资本输出,却催生了持续70年的美欧贸易走廊,并让中国被排除在这一体系之外。在模型层面,他把外交选择归结为一句话:“DeepSeek还是Llama?”
- Appenzeller反对政府全面控制,但认可政府应发挥针对性作用。 政府可以资助基础研究、制定合理监管,但具体创新必须由具备竞争力的企业提供;就连曼哈顿计划都发生过泄密,因此全面控制只是“白日梦”。DeepSeek在OpenAI发布前沿模型26天后以MIT许可发布,印证了他的策略:打造并出口最好的技术,理想情况下由美国及其盟友提供;Midha将由此形成的路径称为“基础模型外交”。
1. 主权集群颠覆云计算的地理格局
Guido Appenzeller指出,王国宣布推出HUMAIN——一家本地超大规模云服务商或AI平台,目标是让“绝大多数AI工作负载”在境内运行,而不是复制由美国和中国基础设施主导的云时代。已宣布的集群建设规模约为1000亿美元至2500亿美元,500兆瓦左右正成为其“原子单元”。
Midha将这一动向与此前的工业周期联系起来:谁控制技术在哪里建设、谁控制底层资产,谁就能影响监管、使用方式和下一轮创新。数据中心如今在战略上的位置,类似工业化时期的石油。
2. 这是工厂,不是普通数据中心
Appenzeller将品牌包装与实质变化区分开来,认为这不只是营销:在底层,GPU是主要的有源部件差异;高密度AI集群需要机架级液冷、接近电厂规模的能源供应,以及对能源供给的提前承诺。
企业需求也在变化:客户越来越接受简单的Kubernetes抽象层,然后“挑选”Snowflake或数据库类服务,而不是购买复杂的全栈云平台。
Appenzeller更尖锐的判断是,模型属于“文化基础设施”。训练过程会嵌入价值观和规范,后训练则决定哪些内容可以说、哪些内容必须拒绝,由此产生了对管辖权控制的需求——控制每座工厂最终生产什么。
3. 模型依赖成为信息安全风险
模型能力已经走出“早期玩具阶段”:基础模型正在国防、医疗和金融服务等领域运行,而ChatGPT约有5亿月活用户,正在做出真实的日常决策。因此,依赖另一个国家的技术,看起来就像暴露在一个“关键故障点”之下。
Appenzeller进一步扩大了主权的含义,将其延伸至信息空间的控制权。当模型取代搜索、批改作文时,被删去的历史可能成为公民继承的现实;某些真实内容也可能因为模型控制者将其排除在训练语料之外,而被判定为错误。
4. 盟友算力可能成为AI版马歇尔计划
Appenzeller认为,美国的领导地位是一种机会,但全球算力不可能完全集中;维持领先地位,同时帮助强盟友获得能力,才是更现实的平衡。
Midha给出的历史类比是:GE、General Motors和其他领先企业曾帮助补贴欧洲重建,尽管当时有人批评马歇尔计划是在输出资本。在他看来,结果是持续70年的美欧贸易走廊,并让中国被排除在这一体系之外。
今天的问题是:帮助盟友建设能力,还是让它们依赖出口而来的中国模型。在模型层面,Midha问道:“我们希望盟友使用什么,DeepSeek还是Llama?”能够为主权基础设施提供资金的国家,已经在行动。
5. 竞争性生态胜过国家级AI总规划
Appenzeller反对Leopold Aschenbrenner的AI国有化论,援引东德和西德,将其称为中央计划与自由市场经济之间的一场“A/B测试”。政府可以资助基础研究、建立良好监管——糟糕的监管可能“击沉AI”——但“没有什么总规划”能够提供市场在实践中发现的那些细节。
Appenzeller称,完全由政府集中控制的路径是“白日梦”:即便是被严密隔离的曼哈顿计划也发生过泄密。模型权重的重要性低于运行模型的基础设施,而“推理几乎更重要”。他还表示,美国一年前的提案曾试图区分模型研发监管与滥用监管,但这场辩论已经向前推进。
DeepSeek打破了“中国落后5至6年”的自信判断:它在OpenAI发布前沿模型26天后出现。其MIT许可让其他国家立即获得访问权限,促使Appenzeller得出结论:唯一的胜法,就是打造最好的技术,并将出口能力做到超过任何对手。他认为,美国最好接受其他国家服务自身模型的能力,理想情况是由美国及其盟友提供最好的模型;Midha将这种新路径称为“基础模型外交”。
This is a massive vulnerability. We’ve got to control our own stack. It’s not just about self-defining the culture, but about controlling the information space. Do we build? Do we partner? What do we do?
The United States in AI right now has world leadership. Instead of colonization, what we have now, I think, is foundation model diplomacy. This big structural revolution is both a threat and an opportunity. And, Guido, I want to talk about sovereign AI, AI, and geopolitics. Let’s start with the news. Our partner Ben is in the Middle East right now. What happened, and why is it so important?
What happened is that the kingdom announced that it’s going to build its own local hyperscaler, or AI platform, called HUMAIN. I think what’s notable is that, as opposed to the status quo of the cloud era, they’re viewing the AI era as one where they’d like the vast majority of AI workloads to run locally.
If you think about the last 20 years, the way the cloud evolved was that the vast majority of cloud infrastructure basically existed in 2 places: China and the United States. The United States ended up being the home for the vast majority of cloud providers serving the rest of the world. That doesn’t seem to be the way AI is playing out, because we have a number of frontier nations that are raising their hands and saying, “We’d like infrastructure independence.”
The idea is that we’d like our own infrastructure that runs our own models, and that gives us the autonomy to build the future of AI independent of any other nation, which is quite a big shift.
I think the headline numbers are somewhere in the range of $100 billion to $250 billion worth of cluster buildout that they’ve announced, of which about 500 megawatts seems to be the atomic unit of these clusters that they’re building. A number of countries, with the kingdom being the most recent, have been announcing what we could think of as sovereign AI clusters. That’s a pretty dramatic shift from the pre-AI era.
I think it’s spot on. Many geopolitical regions are reflecting back on what happened in previous big tech cycles. Wherever the technology is built, and whoever controls the underlying assets, has a tremendous amount of power to shape regulation, shape how this technology is being used, and put themselves in a position for the next wave that comes out of it.
In the Industrial Revolution, having oil was important, and now having data centers is important. I think it’s a very exciting development.
You can often tell why something is important to somebody by the semantics that people use to communicate a new infrastructure project. They’re being called AI factories. They’re not being called AI data centers.
There are 2 ways to respond to that. One train of thought would be, “That’s just branding—the marketing people doing their thing. Under the hood, this is really just data centers with slightly different components.” The opposing view would be, “Actually, no, this is not just marketing.” If you look under the hood—if you X-ray the data center itself—very little of it is the same as it was 20 years ago.
The big difference in active components, of course, is GPUs. Today, if you look at the average 500-megawatt data center and what percentage of the capex required to build—or operate, rather—that data center went to GPUs,
I think we’re also seeing a specialization. The kind of data center you build for a classic, CPU-centric workload and what you build for a high-density AI data center look very different. You need liquid cooling to the rack. You need a very different energy supply, close to a power plant. You want to lock in that energy supply early on.
We’re also seeing a change in consumer behavior. Classically, you want a very full stack that has lots of services to help enterprises build all these applications. We’re seeing more and more enterprises that are actually comfortable with just building on top of a simple Kubernetes abstraction or something similar, and basically cherry-picking a couple of Snowflake- or database-type services on the side to complement that.
I think there’s a new world. The technical components in an AI factory are completely different from those in a traditional data center. Then there’s the question of what it does. Historically, a lot of the workloads that traditional data centers handled were hosted workloads for enterprises or developers, whoever it might be, where most of the data sets and workloads were not particularly opinionated. By “opinionated,” I mean they weren’t necessarily subject to a ton of cultural oversight.
You could argue that was not the case with China, where China wanted full oversight over those workloads. But for the better part of the 2000s, the vast majority of enterprise workloads didn’t need decentralized serving.
What’s different about AI seems to be that these models aren’t just compute infrastructure. They’re cultural infrastructure. They’re trained on data that has a ton of embedded values and cultural norms. That’s the training step.
Then, when you have inference, which is when the models are running, you have all these post-training steps that steer the models to determine whether to say something or not, and whether to refuse the user or not. That last mile is where, over the last year, it has become more and more clear that countries want the ability to control what the factories produce, or don’t produce, within their jurisdiction.
That urgency didn’t exist to the same extent with traditional cloud workloads, because of the cultural factors, or because of concerns around independence or resilience. My sense is that there are 2 things going on, but one is the rise of capabilities in these models. They’re now well beyond what we’d consider the early toy stage of a technology.
You have foundation models literally running in defense, healthcare, and financial services. ChatGPT has about 500 million monthly active users making real decisions in their daily lives. I think that makes a lot of governments say, “Wait a minute. If we are dependent on some other country for the underlying technology on which our military, defense, healthcare, financial services, and daily citizens’ lives are driven, that seems like a critical point of failure.”
I think it’s not just about self-defining the culture, but about controlling the information space. Today, we’re starting to see models replacing search. You no longer go to Google; you go to ChatGPT, and it comes back with an answer. If there’s a historical fact that doesn’t show up in the Chinese model but does show up in the U.S. model, that is the reality that people grow up with.
If you write an essay in school in the future, many of our essays will be graded by an LLM. In fact, in school, something that may be truthful may be graded as wrong because whoever controlled the model decided that it should not be part of the training corpus.
Is your expectation that this is going to play out? To what extent is it going to play out? On the cloud, as we mentioned, there’s the Chinese internet and the Western-rest-of-the-world internet. How widespread is this sovereign AI idea going to become?
If you look at the Industrial Revolution, oil was the foundation of a lot of the technologies. You needed oil reserves in order to participate. I think it will be a little bit the same thing. If you want to build industry in a particular country, export things, drive development, and really harness the power that comes with that, you need the corresponding reserves.
AI data centers are a little bit like these oil reserves, with the big difference being that you can actually construct them yourself if you have the necessary investment dollars and the willpower to do it. But I think they will be the foundations for building all the layers on top that ultimately determine who wins this race.
Talk more about the implications of what this means. Is this something that the United States should be excited about? Are there now winners across the board in all of these local efforts? Talk about some of the big implications here.
This big structural revolution is both a threat and an opportunity. The United States in AI right now has world leadership. That’s an opportunity. Hanging on to it won’t be easy, as it is in every tech revolution.
Don’t we want people to be dependent on us in the same way that they were in the cloud revolution, or do we benefit somehow from it being more decentralized?
The world is not one place, so complete centralization won’t happen. Being the leader is good. Having strong allies that also have that technology is also very valuable. So it’s probably a balance of those that we’re looking for.
We’re clearly in an unstable equilibrium right now, and so Guido’s right that the arc of humanity and history is such that things will probably shake out until there’s a stable equilibrium. The question is, what is the stable equilibrium?
One way to reason about it is to look at historical analogies. After World War II, when Europe was completely decimated, there was a group of really enterprising people in the private sector and the public sector who got together and said, “We can either choose to turn our backs on Europe and adopt a posture of isolationism, where we mostly focus on a postwar, American-only agenda, or we can try to adopt a policy where we know that if we don’t help out our allies, somebody else will.”
Yeah. And so they came up with this idea called the Marshall Plan, right? A number of leading enterprises in the US got together, like GE and General Motors, and literally subsidized the massive reconstruction of Europe. That helped a lot of European economies quickly get back on their feet.
At the time, there was a ton of criticism of the Marshall Plan because it was viewed almost as a net export of capital and resources. But what it did end up doing was solidifying this unbelievable trade corridor between the US and Europe for the next 50 years, which really kept China out of that equation for 70 years. Yeah, 70 years, really.
And so I think we have a choice, which is to either approach it the way we would apply the Marshall Plan for AI, right, and say, “A stable equilibrium is certainly not one where we just turn our back on a bunch of allies, because China definitely has enough compute resources to try to export great models like DeepSeek to the rest of the world. So what do we want our allies on: DeepSeek or Llama?” That’s sort of what it comes down to at the model level of the stack.
The reality is that a number of countries are not waiting around to find out. The ones that certainly have the ability to fund their own sovereign infrastructure are rushing to do it right now. And what does that mean for the nationalization debate, or how do you see that playing out?
You know, Leopold Aschenbrenner, formerly of OpenAI, in his famous report talked about how, if this thing becomes as critical to national security as we think it will be, at some point governments aren’t just going to let private companies run it. They’re going to want to have a much more integrated approach to it. Where do you stand on the likelihood of that, and what does that mean for the feasibility of regulation in a world where it’s much more decentralized?
I probably have a strong opinion on that. I grew up in Germany, right, so benefiting from the Marshall Plan. One lesson I took away from that is that I think any kind of centralized, planned approach does not work.
To some degree, the Eastern Bloc—East Germany versus West Germany—was a nice A/B test: central planning versus a free-market economy. What works better? I think the results speak for themselves.
So I think having the government drive all of AI strategy, whether it’s a Manhattan-style project or an Apollo project—pick your favorite successful project—I can’t see that working. You probably need a highly dynamic ecosystem of a large number of companies competing.
There are some areas where I think the government can have a hugely positive effect. On the research side, we’ve seen that again and again: funding fundamental research that is not quite applied enough yet for enterprises to pick up is very valuable. I think it can help in terms of setting good regulation. Bad regulation can easily torpedo AI, as we’ve seen.
And so I think there’s a strong role for government to lead this and to direct this. There’s no master plan at the end of the day that you can make that has all the details. That has to come from the market.
I don’t agree with Leopold Aschenbrenner’s point of view, actually. The history of centralized planning at the frontier of technology is not great, barring a few situations that were essentially brief sprints of war. Arguably, even the Manhattan Project—which is the analogy I think he uses in his piece—we now know had leaks. It was literally a cordoned-off facility in Los Alamos or wherever, and there were still spies.
For anyone who has had both the pleasure and displeasure of working in any large government system, it’s a pipe dream.
The good and the bad news is that, in a sense, it doesn’t really matter where the model weights are. It matters where the infrastructure that runs the models is. In a sense, inference is almost more important.
A year ago, we were in a pretty rough spot, I would say, with the arc of regulation. There were a number of proposals in the United States to try to regulate the research and development of models versus the misuse of the models. I think that, luckily, we’ve moved on.
Just a year before DeepSeek came out, a number of tech leaders in Washington were testifying that China was 5 to 6 years behind the US, with confidence, on the record. And then DeepSeek came out 26 days after OpenAI put out the frontier model. That just shattered all of those arguments, and the fact that it was an MIT-licensed model meant that every other country had access immediately.
So the calculus has changed, right? I think it means that the only way to win is to build the best technology and out-export anybody else. If the question is whose math the world is using, we’d love for it to be American math.
My view is that we are much better off embracing the ability for other countries to serve their own models. Ideally, the best product wins, and the best models just come from the US and our allies.
Yeah. Is that the new age of LLM diplomacy that we’re entering here?
Ben had a great talking point about this at, I think, FII Riyadh last year. He said something to the effect of: because these models, like we discussed earlier, are cultural infrastructure, you don’t want to be colonized in the digital era, in cyberspace. I think that’s pretty spot-on.
Yeah. Instead of colonization, what we have now, I think, is foundation model diplomacy.