第015期 - DG Matrix 解析 800V直流配电与传统交流配电(数据中心、能源)| Jordan Nanos、Jeremie Eliahou Ontiveros、Nicolas Bontigui、Haroon Inam
Jordan Nanos × Jeremie Eliahou Ontiveros × Nicolas Bontigui × Haroon Inam
AI数据中心因机架功率密度上升而开始采用800V直流配电。 Haroon Inam认为,提高电压比提高电流更便宜:从传统240V交流配电转向800V,可以在相同铜材中输送近3倍功率,同时降低导体重量、成本和I²R损耗。他的判断非常明确:「提高电压来获得更多功率,远比提高电流来获得更多功率便宜」(raising voltage to get more power is far cheaper than raising current to get more power)。
800V的选择,既受电气物理驱动,也受半导体和供应链经济性驱动。 Inam的“有根据的猜测”是,电动车投资以及电机驱动和宽禁带应用中使用的1,200V器件,让800V成为一个具备安全降额空间的实际落点。因此,即使拉动需求的NVIDIA架构尚未到来,相关转型也已经通过工程、制造和供应链准备启动。
DG Matrix的SST逻辑,不是用昂贵电子设备替换一台寿命持久的变压器,而是把多个系统压缩到一起。 Inam称,早期的交流转交流固态变压器项目是“我们能做的最愚蠢的事情之一”;只有整合整流、再加入隔离式双向交流和直流端口后,价值才真正显现。最终的多端口设计有望同时替代STATCOM、UPS、整流器、能源管理系统和表后能源聚合层,但为此至少投入了700,000个工程小时。
这条采用曲线的方向可信,但时间点明确无法作为投资依据。 SemiAnalysis的图表显示,800V直流在2026年几乎为零,到2030年将占新增容量近80%、超过30 GW,但Inam只认可其形状:随着原生的全设施直流逐步落地,sidecar方案的占比应会下降。600 kW或1 MW机架路线图若推迟,采用曲线也会右移;Inam强调,采用率“可能上升”、维持更久,也可能更快下降。
混合硬件是核心商业风险,而电力可互换性是拟议中的对冲手段。 CPU已经重新回到主舞台,ASIC的需求也可能分化;Amazon或Microsoft必须服务广泛的云端机群,同时还要为OpenAI、Anthropic等AI用户建设基础设施。多端口SST可以在交流和直流之间分配全部额定容量,但主持人指出,即便铜材仍可复用,转换仍需要不同的保护装置、连接器和电缆组件。
长期架构将是由表后发电不断供能的软件定义电力网络。 Inam预计,电网瓶颈和持续多年的升级周期,会推动可部署的10–20 MW“蜂窝式电力”模块,并认为“接入电力的速度,或者获得算力的速度,将决定胜负”(speed to power or the speed to compute will win out)。他设想动态电力路由、GPU调度、承载5–6 MW的超导连接,甚至面向Token生成的竞价市场;同时坦言:“我足够聪明,知道自己没那么聪明。”
DG Matrix为这一愿景设定了具体里程碑,也明确了执行风险。 公司目前已有400 kW单元可以并联,正在研究2027年的6 MW SST,并计划在2028年推出集装箱式10 MW系统,输入35 kV,输出可编程为800V或1,500V。兆瓦级功率目前更偏向碳化硅,而不是氮化镓。公司现有碳化硅供应商位于美国、可能位于日本以及欧洲,而非中国;部分非CPU、非软件的机电部件则采用“中国+1”策略。Inam警告,软件定义电力必须做到“经得起网络安全检验”。
1. 机架密度让800V成为铜材经济学上的必然选择
Inam的前提是,GPU功率需求以及GPU之间必须保持的同步性,正在超出传统交流电缆的能力边界。从三相系统中的240V交流单相配电转向800V,按均方根值对均方根值比较,可以在相同铜材中输送近3倍功率,前提是配电系统能够承受这一电压。
Nicolas Bontigui逐一追问驱动因素:母线排重量、铜价上涨、效率和I²R损耗;Inam认为这些因素都成立。电流才是约束瓶颈;在解决爬电距离和电气间隙后,“提高电压来获得更多功率,远比提高电流来获得更多功率便宜”。
为什么停在800V,而不是950V或1,200V?Inam给出的明确保留意见是一个“有根据的猜测”:电动车建立了800V生态,而电机驱动和宽禁带半导体则围绕1,200V器件发展。两者结合后,800V既是具备器件余量的实际运行电压,也兼顾了功率密度和成熟的元件基础。
在拉动需求的NVIDIA架构到来前,采用已经“启动”。制造工艺必须成熟,面向电动车的供应链也必须转向数据中心;正如Inam所说,客户不可能等到那一刻才说:“我现在要发明800V架构。”
2. 多端口整合才是SST存在的经济理由
Inam早期的教训发生在2011–2013年:他的团队投入“数百万、数百万美元”打造交流转交流SST。用成本更高、更脆弱的电子设备,去替代能够存续40–50年并承受反复热循环的铁芯、铜材和绝缘材料,事后看是“我们能做的最愚蠢的事情之一”——那只是一个没有价值主张的科学实验。
突破口在于把下游变换环节一并纳入。整合交流转直流整流后,整体系统更便宜,也可能因集成度更高、利润层层叠加更少而更可靠,因为所有环节来自同一家公司;进一步加入更多隔离式、双向的交流和直流端口后,一套系统就可以替代STATCOM、UPS、整流器、能源管理系统和表后能源聚合器。
这个“圣杯”并不是做做整合的概念演示:控制、散热、电磁干扰和功率密度不断撞上“砖墙”,在能够大规模部署前,至少需要700,000个工程小时。Inam对终局的衡量标准是“功率输入、Token输出”——压缩铜、铁和各种“杂物”,消除闲置容量,并把每千瓦时对应的Token成本降到最低。
3. 单极800V以一根导体换来更棘手的故障管理
Inam认为,Dreamliner提供了最好的类比。在30,000–40,000英尺高空,约300V以上的绝缘退化方式不同,因此飞机采用围绕公共导体布置的正270V和负270V:实现540V供电,却不让任一侧暴露在完整的电势差下。他只是提出疑问,并未给出确定结论:电弧闪络风险是否也促成了正负400V方案。
关键在于正负两条母线上的负载是否保持平衡。如果不平衡,双极配电就相当于两套电路;在部分变换器中,失衡可能引发失控状态,使其中一路电压相对另一路塌缩,而故障处理、接地和非隔离变换也会复杂得多。
因此,只要电弧闪络问题得到解决,单极800V就可能凭借少一根铜导体和更简单的回流接地胜出。Inam认为,答案可能在于快速检测,并在故障电流继续获得供给前“熄灭源头”;他希望这一能力足以支撑一套统一架构。
DG Matrix有意不绑定具体拓扑,因为每个端口都实现了电气隔离。其输出可以悬浮在负800V或正800V,也可以将中点接地形成正负400V,以适配NVIDIA的参考接地方案。当被追问其他成本差异时,Inam坦言无法给出明确答案:差异可能很多,但“总体而言”,第三根铜导体是最重要的一项。
4. 硬件不确定性决定直流采用速度
Jordan Nanos将路线分为4个阶段:白区改造、原生直流算力、全设施直流,以及以SST为中心的终局。SemiAnalysis的图表显示,800V直流将从2026年几乎为零,升至2030年年度新增容量的近80%,规模超过30 GW;Inam同意,随着原生直流落地,sidecar方案的占比应会下降,但他说:“我们没有水晶球。”
在交流占主导的存量设施中,sidecar仍有用,也可以在直流迁移不够快时提供保险。反过来,如果600 kW和1 MW机架系统延期,曲线也会随硬件向右移动。Inam拒绝给出虚假的精确度:采用率“可能上升”、维持更久,或更快下降,即使整体形状仍然合理。
SemiAnalysis以冷水机为类比,揭示了理论上更优的设计为什么会停滞:买家曾被告知液冷系统可以让冷水机在45°C下运行,但实际采用这一方案的比例仍极低,因为未来的工作负载组合并不确定。随着CPU“重新变得重要”,再加上GPU、TPU、ASIC、存储和其他负载,Amazon和Microsoft即使将AI容量交给OpenAI或Anthropic使用,也需要具备可互换性的云端机群;直接自建的客户最终可能会选择更专用的设施。
主持人的反驳是,多端口硬件无法抹去已经嵌入开关设备和配电系统的选择。Inam给出的较窄回答是:混合建设可以把全部额定负载放在交流或直流一侧,只要总负载不超过设备额定值;之后从交流切换到直流时,铜材可以保留,但保护装置、连接器、电缆组件和对应输出端口都要改变。超大规模云厂商可能仍然保守,而新云厂商和开发商可能为了更快回本而接受更高风险。
5. 需要运行15年的设施要的是电力网络,而不是固定机架假设
Nanos给出的机架路线图是:传统风冷CPU机架约12 kW,2020年前后达到30–40 kW,目前GB200和GB300系统则为130–140 kW。下一步是明年、或可能在今年年底达到每机架360 kW;到2027年底,Vera Rubin机架将达到600 kW;2028–2030年则讨论1 MW。他说,这张图需要用对数坐标;他参观过的一座设施,超过“80%”的面积都用于供电和散热。
Inam把未来的数据中心比作一台由“一只小猫”控制的80英尺机器人:电力架构依然庞大,而GPU这颗“大脑”持续缩小。他设想一种由多个节点供电的自行车轮式几何结构、承载5–6 MW的超导线缆,以及在故障发生时通过软件重新路由电力、避免容量闲置的系统;这一切还要与GPU任务调度协同,甚至可能为Token生成引入竞价机制。
Inam承认,光接口或光计算可能逆转或压平功率密度增长,但他区分了头脑风暴和执行,不愿预测何时发生。Nanos则以杰文斯悖论回应:效率提升可能只会诱发更多消费。Inam最终坚持的不是某个瓦数预测,而是灵活性:“我足够聪明,知道自己没那么聪明。”
6. 分布式发电扩大机会,也扩大攻击面
中央电网按单向电力流设计;即使其他地方存在发电能力,把100 MW集中到一个节点仍可能堵塞周边输电系统。电网升级需要数年和数亿美元,而表后的“蜂窝式电力”可以按10–20 MW模块部署,并逐步扩展为分布式吉瓦级系统。Inam的判断是,“接入电力的速度,或者获得算力的速度,将决定胜负”;不过他将“颠覆”修正为更谨慎的公用事业“增强”。
产品路线图非常具体:400 kW单元今天已经可以并联;DG Matrix正在研究2027年的6 MW SST;2028年则计划推出大型集装箱式10 MW设备,输入35 kV,输出可编程为800V或1,500V。该系统应部署在白区之外;把中压电力推得更靠近算力可以节省铜材,但会面临安全以及国家政策、分区规划方面的障碍,Inam推测中国可能率先推进。
在器件选择上,DG Matrix原则上不绑定特定技术,但目前采用碳化硅,因为从数百千瓦到兆瓦级都更成熟;氮化镓目前更适合数百瓦到千瓦级功率。公司对两种技术都已试验超过10年,碳化硅供应来自美国、可能来自日本以及欧洲,而非中国;部分机电部件则使用墨西哥或越南等“中国+1”替代供应地。
Inam最后警告,自适应电力将成为关键基础设施软件。DG Matrix拥有输电电网安全方面的既有经验,但他认为全行业都必须扩大相关能力,包括对所有接触电子设备或参与软件开发的实体进行背景审查。Nanos给出了令人印象深刻的下行情景:“我们不希望这些新建的大型数据中心里出现Stuxnet。”
Hello, everyone. Welcome back to SemiAnalysis Weekly. I'm Jordan. Today, I'm joined by Nico, Jeremie, and Haroon—our first-ever guest on the podcast. Up until this point, it's been all SemiAnalysis people, but now we're bringing on a guest because 800-volt DC adoption is too important. We need to bring in the experts. Haroon, welcome to the show. Do you mind starting by introducing yourself to the audience and telling us what you guys do?
Sure. I'm Haroon, co-founder and CEO of DG Matrix. I got into power electronics when it was extremely unpopular and a very uncool thing to do, and now it has become much more popular. I've had a chance to work on everything from computer-room power, which was a precursor to data centers, to solar inverters before the sun started shining on that industry. I've done transmission power-flow control at the hundreds-of-megawatts level using power electronics, and I've even had a chance to pioneer some of the electronic jet-engine starters—all power-electronics-based—on the Dreamliner and the Joint Strike Fighter.
Now, I think we're sitting on an extremely exciting era for humanity, where the 800-volt DC architecture is helping us propel the human race to superhuman intelligence.
Awesome. Nico, why don't you kick things off with a few questions? I know you guys collaborated a little bit to create part one of the 800 Volt DC Revolution article.
1. The 800 Volt Necessity
I think Haroon already mentioned that the reason and the main topic for this conversation is going to be 800 volts. I believe that it's been one of the main trends we've been hearing about throughout 2026. We go to conferences and, you know, there's 800 volts pretty much everywhere. All companies are showcasing their sidecars, their prototypes—pretty much everything.
The obvious question, before we get into the solid-state transformers and all the cool stuff that we're going to cover today, is: Haroon, why are we discussing 800 volts in 2026, and why are we discussing 800 volts when we think about those 1-megawatt racks?
I think the compute power required for GPUs and synchronicity is increasing to a point where legacy AC architecture and standard strands of AC cables are unable to carry the power. The question is: How high can you go in voltage so that you can lower the cost and remove the constraint of the copper delivery system?
Taking 240 volts AC single-phase, times 3 phases, to 800 volts—effectively RMS to RMS—you’re going to get almost triple the power on the same copper cable, provided you can handle the distribution. I think those economics are what's taking it to 800 volts.
The second question would be, well, why not 1,200 volts? Why is it 800? Why isn't it 950? I'm going to venture an educated guess that, first, it has to do with EVs, which developed a lot of 800-volt architecture. It also has to do with the fact that the most popular semiconductor is the 1,200-volt device that's used in motor drives all over the world.
When silicon carbide and wide-bandgap semiconductors came out, they came out for a 1,200-volt architecture. When the device can take that much, 800 volts is a good, safe voltage to settle on. Maybe that's where some of the genesis of 800 volts is: the power density and the semiconductor ratings.
To put it simply for the audience, when you say we're unable to get the power required for these 600-kilowatt and 1-megawatt racks, is it a matter of the weight of the bus bars because of the amount of copper they'll need to distribute all the current required to reach those power levels? Is it a weight matter? Is it a cost discussion because we know the price of copper is going up like crazy? Is it a matter of efficiency—to lower current and therefore I²R losses? Is it a bit of everything? In your opinion, what's the main driver for this whole revolution, as we would like to call it?
I think you guys are very good at understanding the physics and economics. The work that SemiAnalysis does is very impressive. So, you've hit all the points, and I think it's about how much current you can get.
Generally, when you raise the voltage, as long as you have the separation—the creepage and clearance—raising voltage to get more power is far cheaper than raising current to get more power. I think you're right: you're getting far more effective use out of the same copper.
The foundation of why we're talking about this today is now clear. Of course, we're being asked a lot about timing: “Is this something that's already happening? Is it something that's going to start kicking off in 2 years?”
To put it simply, in your opinion, when does 800 volts become a necessity? When is 800 volts still more in the proof-of-concept phase? Of course, we have the article with all these phases. I know it's not a short answer, but in your opinion, when does the 800-volt revolution really start to kick off?
I think it's started. The question is when the right architecture from NVIDIA comes out that starts driving the demand. When that comes out, you can't say, “I'm going to invent 800-volt architecture now.” You've got to do it upfront.
In a way, the maturation of the technology, the maturation of the manufacturing approach, and the buildup of the supply chain—and/or migration of the supply chain from the EV side to the data-center side—have started already.
There's a big question, and it's interesting. We get this question a lot, and this is where we've come up with a very unique solution using our multi-port transformers. The question in customers' minds is, “What percentage will be DC, and what percentage will be AC?”
If I go exclusively DC and the adoption rate is less than we want, are we going to be left with stranded power? Are we going to be left with a stranded investment? And if we ignore AC, and those new companies that are coming out with chips are going to run, let's say, less power or run on an AC architecture, then what do we do? Now are we going to be in the same situation?
I think the answer we're gravitating toward is that it shouldn't matter. You should have an architecture that's highly flexible, that can do AC and DC in any percentage you want, right from the same product line. I think that's why people are so interested.
We just want to immunize the financial risk for the developers, the neoclouds, and anybody else around how much is DC and how much is AC, and reduce the risk around the timing.
I think your last comment is a wonderful opportunity to introduce your multi-port products and the value proposition of your multi-port solutions. For the audience that's maybe not that familiar with DG Matrix and the multi-port solutions, how do your solutions work, and how do you take different inputs, different voltages, and different frequencies?
2. The Multi Port SST
One of the things that I did back in 2011 to 2013, when I was working for an SST company, was develop an SST—a solid-state transformer—to do AC-to-AC conversion. We pitched it like, “It'll clean up the power, it'll do this, it'll do that.” But after 2 years of work and spending millions and millions of dollars, we realized that that was probably one of the dumbest things we could have done.
Why? You're taking a hunk of iron and a hunk of copper wound around it, with some insulation that's going to last 40 or 50 years. It's going to last in heat and through thermal cycles. Why the hell would anybody in their right mind try to replace that with a bunch of electronics that are going to be more delicate, cost a lot more, and be less reliable? Why would anybody do that?
As we started to ask that question, the answer was, well, it might be a great science experiment, but that's where SSTs are going to stop. However, the answer came when we looked at what happens to that AC after you transform it. Do you do variable AC with it? Do you do a motor drive at the end of it? Do you take a medium voltage and convert it to a low voltage? What do you do?
The answer is that if you combine, for example, the rectification function after the AC and put it all in an SST, holy moly, now you've got a balanced system that's actually cheaper. It's more reliable because it's integrated. It has less margin stacking because it's coming from one company. Lo and behold, you found the first value proposition for the SST—but it's AC-to-DC conversion.
We started thinking further about it. We said, “Well, anybody can do that. How do you differentiate that?” So we came up with this crazy idea: If you're adding a port that does DC, you're adding much more value because you're collapsing a lot of the system that happens afterward.
So why not look at more ports? We said, “What if we added more AC ports? What if we added more DC ports? What if we could make every port bidirectional?” And so we said, “Holy...” Well, the word is something else, but I’ll replace it with “holy moly.” We said, “Holy moly, look at the value that you will add here.”
You could replace a STATCOM, a UPS, a rectifier, the energy management system, and behind-the-meter energy aggregation, all with a multi-port SST. We said, “Boy, that’s the Holy Grail. That’s what we need to develop,” because the economics and the physics are all in your favor. As we went down that path, we didn’t realize that the controls, the cooling, the electromagnetic interference, and the density would present so many brick walls. It took us at least 700,000 engineering hours to get multi-port to a point where we could start doing deployments all over. That’s how we came up with multi-port: basically, economics and physics—
Mm-hmm.
—driving innovation.
Mm-hmm. Okay, that’s fascinating. That’s truly fascinating. You touch upon incredibly interesting points, and I really don’t want to jump from the very beginning of the conversation we’re having now to that end state. But you mentioned taking all these functions that UPS systems currently cover, along with all these other parts of the legacy electrical equipment. Just a quick question before we go back to where we are today: In your view, when you think of a data center in 5 or 10 years, how does it look? How does the electrical architecture look?
3. The Future Data Center
I think clearly, as densities increase, the intelligence goes up and the number of points that you compute goes up. The cost of a token in kilowatt-hours goes down. The question is, what is going to drive that metric? Is it the cost of the token per kilowatt-hour? Assuming everything else is depreciated, it’s going to come down to power. When it comes down to power, it’s power in, tokens out.
How do you get the absolute lowest cost of that token, and how do you maximize that infrastructure? That is the answer. I think the voltages are probably going to go up at some point. People are already talking about 1,500 volts DC. I think the density of the racks will probably go up, and the racks are going to get smaller and smaller and smaller. The power infrastructure also has to follow a similar trajectory.
That’s where collapsing multiple systems into one makes sense. Not only do you get rid of a whole lot of copper, iron, and junk, but you have far better functionality to eliminate stranded power and supply those dynamic loads. That’s where I think it’s going to end up: in far denser environments, with even more integrated cooling and—
4. Unifying 800 Volt Designs
That’s fascinating. Now that we have you here with us today—it’s a great pleasure to have you here—it’s an opportunity to pick your brains and learn how you envision these data centers looking 5 to 10 years out. But let’s go back to the present. Today, we’re in the early, early days of this whole revolution. We’re still at a point where we hear about 800 volts as a whole, but when we look deeper into the systems, we know about some hyperscalers working with plus-minus 400 volts. Some others are working directly with single-ended 800 volts.
Again, just to put it simply for everyone to understand, from the perspective of DG Matrix, how do you approach this? What are the implications of going to plus-minus 400 volts or going directly to 800 volts?
The interesting thing is, what’s driving plus-minus 400 volts versus 800 volts, and is it a balanced plus-minus-400-volt load? That’s the first question. When we did the Dreamliner, it was interesting. Whenever you fly something at altitudes of 30,000 to 40,000 feet, the air is very different. The ionization of the insulation happens in a way that degrades insulation above 300 volts. The magic rule is that you don’t want to go above 300 volts.
As the density of power goes up in airplanes, it has gone up considerably from the 747 to the 787 and whatever is coming beyond, running 270-volt DC cables was untenable. The guys who did the Dreamliner came up with this: Let’s run plus 270 and minus 270, with a common conductor in between. Now you’ve got the best of both worlds. You’re running 540 volts or whatever, but not really from an ionization standpoint.
I’m wondering if the same thing drove the plus-minus-400-volt vision, but from a different physics: the physics of arc flash. Was it that arc flash is better understood at 400 or 500 volts DC, and there’s a bigger perceived risk at 800? That may have been where it came from.
The competing architecture, which is a close cousin, is 800 volts without the third conductor. If you have a balanced load, the third conductor may be very, very small, but then you get into faults and how faults propagate. You get into the grounding schemes, and it becomes a nightmare for non-isolated converters. I think that’s where it would be nice to get some harmonization.
We frankly don’t care which way it goes, because every one of our ports is galvanically isolated. When it is, you can float it anywhere you want. You can float it at minus 800. You can float it at 800. You can ground the center point and get plus-minus 400, and we can use any grounding scheme that NVIDIA is proposing in its general reference architectures.
I think it’s going to come down to conductor cost, in which 800 volts might be cheaper, and it might come down to the opposite of that: How do you solve for arc flash? Again, I think detecting arc flash and being able to quench the source from feeding the fault is where the magical answer will lie in setting a unified architecture, hopefully.
The cost consideration that you mentioned—is it just because, at 800 volts, you have one conductor less to protect and control? Is it just that, or is there any other consideration when we think about the cost of different systems?
I’m sure there are many other considerations, but I think that copper cable—the third copper cable—is a significant consideration. There may be many others. What I would do is come back to you with a more comprehensive look at what feeds that. But generally, I think it’s that copper conductor.
When we think about cost and system complexity, and NVIDIA and NVIDIA’s partners working initially on this sidecar that’s going to be single-ended 800 volts, what are the considerations when it comes to system complexity? Is it actually more difficult to implement and design a system that uses single-ended 800 volts compared with one that other agents may be working on that uses plus-minus 400 volts?
I think the essential question, when you have plus-minus 400, is: Are the loads going to be balanced at plus 400 and minus 400? If the load is not balanced, it’s clearly a more complex system. For example, some fuel cells come at close to plus-minus 400, and so that’s always going to be the question: Can we just take power differentially?
If you can take power differentially, you have to treat it as 2 different circuits so that imbalance doesn’t persist. In certain circuits, it can cause a runaway condition where you collapse one voltage versus the other. I think there are reasons to favor a unipolar 800 volts, as long as you can address the arc-flash risk reduction properly. It also gives you a way to do standardized grounding on the return conductor with a multiplicity of ways, rather than worrying about whether you’re going to do grounding on 3 conductors versus just a return conductor.
5. The 800 Volt Adoption Curve
That makes sense. I’m going to take this a little higher-level and talk about the adoption curve. There are 4 phases: white-space retrofit, native compute, facility-wide DC, and then the end state of housing these SSTs. To start the discussion, let me share a specific chart that you guys put in the 800-volt DC article.
Do you believe that this is a pretty solid adoption curve that’s going to happen? Are there chances that this gets accelerated or pushed if that theoretical 1-megawatt rack doesn’t really come to fruition, or if the roadmap just gets pushed out? For those just listening, we’ve got a chart on screen for the YouTube audience showing 800-volt DC adoption going from basically nothing in 2026 to almost 80% of the market by 2030, in terms of the incremental capacity being added to the data center market every year, and pushing above 30 gigawatts of actual adoption. That’s unbelievable to think about.
But it happens in phases: initially, it's going to be a sidecar, and later, it's going to happen at the facility level. What's your high-level take when you see a chart like this?
My high-level take is that it's always very difficult to project into the future. While we can't tell you whether these numbers are right or wrong—we don't have any special crystal ball—we do agree that there will be a market for sidecars that will go down over time as the native architecture for 800-volt DC takes root in AI data centers.
The question is, how long will that sidecar last? Especially when you have AC-dominated architectures and you're doing a brownfield install, it's far easier to do it with a sidecar. Or if you're trying to mitigate the risk of not having the DC migration happen fast enough, you go with an AC data center, then you need the sidecar if it starts to happen.
I think we generally agree with the shape, but it's very difficult to predict the numbers. We don't have that crystal ball. In our case, we solve the problem both with a sidecar that we're developing and releasing through partners, but we're also developing that multi-port that can handle the problem without a sidecar, because you've got both DC and AC coming out. So it's a different way of solving it for the whole data center.
Yeah.
Jordan, I think you mentioned a really important point, which is the possibility that this curve gets at least displaced to the right for some time—let's say a year, or however long—not down, just to the right.
It could be to the right. It could go up. It could go longer. It could go down faster. It could be any one of those scenarios, but the shift to the right may be very, very possible. You're right. Go.
Yeah. It's possible in the sense that, when thinking of this adoption curve, we need to think of it—and this is how we started the conversation—as a hardware- and physics-driven transition, driven by these roadmaps of 600 kW racks. Suddenly, and soon, we will have 1 MW racks.
If these systems, which are extremely complex to design and adopt at large scale, are delayed for a year or whatever—like NVIDIA roadmaps for Rubin or whatever get pushed by a year—we know that this happens, especially when thinking of these super-complex systems. This adoption curve will naturally just follow the hardware. It's not—it's just like—
But if it's not driven by the facilities, the concept of a sidecar is, like, I'm going to retrofit a facility that wasn't designed from the ground up to accept multi-port SST. It's not like a sidecar design is beneficial; it's just really dependent on the site they're going into. Is that fair to say?
Yeah. An interesting parallel that we saw earlier this year, and I guess last year as well, was with chillers. NVIDIA was pitching, “Hey, you can run your chillers at 45°C when you're doing liquid cooling.” In theory, you can do it, but in practice, the share of folks running their chillers at that temperature is extremely low.
The question is why. It's more efficient, supposedly. It's more energy-efficient. You can even save on CapEx if you do this. The problem is that the buyers themselves don't really know exactly what their mix is going to be.
And in fact, if you think about it, they've actually been proven right, because you would think maybe everything is GPUs, and what we're realizing—and at SemiAnalysis, we've probably been the first to call it out at the end of last year—is CPUs are so back, right? So you're actually very much CPU-constrained now as well. It actually makes sense if you have a limited data center footprint that you want your facilities to be able to handle many different types of hardware.
For SST adoption here, the biggest risk would be that uncertainty around the hardware remains high. The timeline is part of it. The diversity of hardware is another one. In a world that is very largely, say, NVIDIA, and NVIDIA's roadmap is 800 volts, the decision is easier.
But in a world where you have many different types of ASICs, some of them maybe don't require 100 volts, maybe CPUs are even more of a need—which we actually are pretty bullish on CPUs right now—and storage and others, it makes sense that you want your hardware and your data centers to be able to handle multiple types of hardware.
It also goes back to who is actually building the data centers. Right now, you have this very interesting moment where, for a big portion of the folks building the data centers, they aren't actually the ones really using them. The big users are basically OpenAI and Anthropic, and the folks building data centers are Amazon and Microsoft, who are building for OpenAI and Anthropic.
Amazon and Microsoft both have the same struggle: their businesses are very diversified. They have a giant CPU cloud business as well, and so they're at the core of this uncertainty with regard to what types of hardware they're going to deploy.
A few years down the road, that could change. If folks like OpenAI and Anthropic start to self-build or start to lease directly, they're going to have different requirements. They're probably going to be much more AI-optimized in some of their designs. Our institutional clients already know that pretty well. We've talked about this at length.
But this is the state of the industry right now, where you have different layers of third parties that are not the actual end users, and so you have this uncertainty about what type of hardware is being deployed. That's one of the risks to SST adoption, knowing it's going to be 2028, 2029, 2030, or 2031 for the very large-scale numbers.
I think, by the way, those are excellent points, and I agree with everything you said. The only thing I'd like to add is that I think multi-port SST, even if I'm biased, solves that problem for you by allowing you to put any load on DC and any load on AC, so it de-risks it for you.
However, having said that, can I predict the adoption curve of multi-port SST? No, I can't, because the hyperscalers are generally more conservative, and they have a right to be. They're building gazillion-dollar data centers, and they're going to be a little bit more risk-averse. But the neoclouds and the data center developers may be more willing to take a risk to make sure that their investment has a faster payback.
I think there are several ways to solve that problem. We have one way that we think is very powerful. We also have the sidecar way, and we agree with you. It's going to be the CPUs, the GPUs, the TPUs, what power they use, how much goes to colo loads, how much goes to AC loads, how much goes to DC loads, and how much behind-the-meter power you need. So there's quite a bit of flux. That is for sure. I think certain classes of SSTs are going to be at more risk of adoption versus other ones.
All right. I guess one interesting question for you, Dan[?], is: you said the multi-port kind of solves the issue. But the complication here is that, obviously, the electrical system of a data center is very complex. Things have to be decided ahead of time, and so I just wanted to understand: why does multi-port actually solve it?
Because if you design your data center for AC, if your whole distribution, your switchgear, and whatnot is AC, then you're going to need a sidecar regardless. And if it's DC, then you're going to do it DC-based, so multi-port, I guess, does a lot better.
A lot of folks who are looking at it with us are doing a hybrid: they want to do a certain amount on AC and a certain amount on DC. What we offer them in that case is that they can put full load on DC or full load on AC. As long as the 2 loads are under the full-load rating of the machine, we don't care. We can give you both. So it gives them flexibility.
We're also saying, if you have DC today or AC today and you want to convert it to DC, we offer a very simple change-out for our portion. You're not going to change the copper; you're going to change the protection, and we offer a port switch-out from AC to DC. That's what makes it easier to do.
So either buy both, and then you deal with the distribution, especially with the protection, right? The copper is not going to change. You're going to get much more out of your copper when you switch from AC to DC. But you change the protection, possibly the connectors and the whips and whatnot, and so it leaves you with an easier path when that transition happens.
And so, actually, that's a good transition to Nico's next banger article on industrials, because modular data centers are one topic we're looking at very closely. I guess you could imagine that if you're multi-port and you can handle both easily, then perhaps there's a world where you could use modular data centers, and you have one module—whatever, 5 MW AC, 5 MW DC, 5 MW AC—and then you can do whatever you want, right? That could be an interesting future for you guys and for reference architectures. Yeah.
Yeah. You know what I like about the way you guys think? It was reflected in that 65-page article, and I think it's the most widely read publication from what I know. A lot of our folks have called us up and said, “Have you read this SemiAnalysis piece?” We’re like, “Wow, these guys are really good.”
We like the way that you systematically think about it from the whole-system perspective and not just focus on one little doohickey. My compliments to you for looking at all the things on the load side and on the AI side that will cause architectural and technology-adoption changes.
It's always nice to hear self-promotion on the SemiAnalysis podcast, Haroon. Thank you for that. I’ll take the—
Well, in this case, because a third party was doing it—or your guest was doing it—without the offer of a free cappuccino, I feel that it was genuine.
Cappuccino coming your way next time, man, for sure.
6. Future Proofing Power Systems
One thing, just coming from the neocloud perspective, that the hyperscalers always talk about is fungibility. They treat this at the fleet level, where different data centers might be used for different people or different things, and then they try to solve this with software. It seems like everything you're saying right now is making the case for fungibility at the power level, in the data center itself.
Can you talk about future-proofing even beyond 1 megawatt? Actually, before I ask that question, let’s take a step back and go through the rack-level power roadmap for a second, because I think maybe we glossed over this a little bit or assumed that the general audience was going to understand it. Let me put this on screen so that we know about this.
When I started doing design work on compute systems for GPU servers, it was in the 2016–2017 timeframe, and we were working on the V100, the Volta-generation systems. A rack, which is a standard data center rack that you might have in US East 1 with air-cooled CPUs for AWS, is about 12 kilowatts.
At the start of COVID in 2020, we started seeing more air-cooled density. You go to 30 or 40 kilowatts per rack. We’re now shipping somewhere between 130 and 140 kilowatts per rack with the GB200 and GB300 systems. Next year, or potentially at the end of this year, with Vera Rubin, what data centers were designed for two to three years ago is 360 kilowatts per rack, and then Vera Rubin by the end of 2027 is 600 kilowatts per rack.
For the audience, that’s already massive. We have to put this chart on a log scale for those looking at it on screen because it’s going up by 6×, without the transition to 800-volt DC even being considered. When we say 1-megawatt racks, what we’re considering for the 2030—or potentially 2028–2029—timeframes is beyond a 60× multiple of power per rack that has had to be contended with.
Now I’m going to ask the question: What does future-proofing look like beyond this? Let’s say you build a data center that’s 100-megawatt scale. I was in one of these facilities a week ago, and it’s absolutely unbelievable how much of the facility itself goes toward power and cooling as opposed to white space, chips, and data hall space. Well over 80% of the physical square footage is just power and cooling now, so I can’t even imagine what the future ones are going to look like.
Let’s say it’s a 100-megawatt site, or even a gigawatt site. These sites are expected to go for 15 years, right? The whole case for fungibility on power, I assume, is that you’re not going to rip out systems that you’ve deployed in the middle of their life. We want to reuse this facility for future systems.
Is there anything beyond the current generation of systems, if you push this out 10 or 15 years, where you think SSTs would be more capable of handling the future load at the end of a 15-year life cycle for the data center facility itself that was built to handle those chips?
Not only that, I think so. You have to look at an architecture that’s going to deliver far more density and be able to work with multiple sources behind the meter. When you look at the transmission grid and the distribution grid, even if you’ve got enough generation and then you put a 100-megawatt data center in one spot, you choke up all the lines around it. That’s why there’s all this issue with, “How am I going to improve my grid to get there?”
The answer in the short run is, “I’ve got to do behind-the-meter power until the grid upgrades.” But if the grid upgrades and the cost of the grid goes up, or the cost of depreciating that asset gets passed down in more expensive dollars per kilowatt-hour, that means your token cost is going to go up.
How do you leverage today’s behind-the-meter power that you’ve put in and depreciated? Can you still continue to use and leverage it, yet increase the density of delivery toward racks that might go higher in power? I think that may be one thing to look at.
The second thing to look at is sort of like this movie I saw a while ago, where there’s a gigantic 80-foot robot. When it comes to a stop, the top opens up and a little kitty cat who’s running the whole robot jumps out. That’s how it is. You’ve got this massive power architecture, and the brain, which is the GPU stack, keeps shrinking and shrinking and shrinking.
What geometry of the data center is going to optimize that brain shrinking? Is it going to be like a bicycle wheel, where you’ve got power coming in from multiple places and then you pop down an increasingly smaller set of GPUs that allow you to handle that? What about superconducting? At what point does superconducting kick in, where you can do 5 or 6 megawatts on a strand of cryogenically cooled cables that will bring you unprecedented density?
How do you distribute it so that any failure mode will not give you any stranded power, and you can route the power to wherever the GPUs demand it for the cheapest token generation? Another way to look at it is that you might even have an auctioning system for selling token generation to the highest bidder.
I think there’s going to be a tremendous amount of software-defined GPU scheduling, a tremendous amount of software-defined power routing, and power handling at every single level. There will be a cooling fabric and a power fabric that can adapt to all these situations. Then there will be GPU job scheduling as you look at different phases of GPU rollout.
It’s also very conceivable. It’s easy to brainstorm because you’re just thinking out the reality; making it real is different. What about all these optical interfaces and all this optical computing that’s coming out? Is that going to reverse the power density, or will it keep power density flat at some point, where the optics kick in and reduce the amount of power that you need for the same amount of computation?
Those are the questions. I’m smart enough to know that I’m not that smart and don’t have the answers for when it’s going to happen or how, but these are some things to think through.
I think we’re big believers in Jevons’ paradox for everything, including power. Even if you’ve got that optical stuff, I think we’re still going to keep consuming quite a bit of power into the future.
It’s interesting to hear you say that specifically for behind-the-meter power generation, you think this is a trend that’s going to continue. In other words, just building more facilities at the same site even if you get grid-connected, or just trying to deploy more chips at the same site. If people are planning for 800 volts right now or planning big data centers, and you’re working with them right now, is this behind-the-meter trend more here to stay than we think?
Jordan, that’s an excellent question. I’ve done a lot of work on the distribution grid, a lot of work on the transmission grid, and studied the economic models of utilities.
All over the world, utilities generally have unipolar, or unidirectional, flow of power, where power goes from generators down the transmission and distribution networks to where it’s used. Upgrading that infrastructure is a multiyear process, and you need hundreds of millions of dollars to do it.
Now you’ve got this cellular power concept, where you can add 10- or 20-megawatt blocks at a time behind the meter and start to add a gigawatt of distributed power. Which one is going to win out? I think the speed to power, or the speed to compute, will win out. For that reason, distributed power generation—another word for behind-the-meter power generation—is going to take root.
And I don't think it's going to take root in just AI data centers. I think it's going to take root wherever you've got to develop electrical power delivery without the cost of a $1 billion nuclear plant or a $10 billion nuclear plant. It's far easier to put a $5 million pod in place and give villagers a hospital, give them a school, and give them a chance to educate their kids. Right?
So there's an electrification trend that's going to drive the need for cellular power, or behind-the-meter power, but there's a massive market right now that's going to drive the volume to make all the infrastructure for behind-the-meter power more palatable and drive the levelized cost of energy down. Then you adapt it to different areas. I think it's a disruption of a multitrillion-dollar energy market—or maybe not disruption; maybe that's too bold. Maybe it's the augmentation of a centralized generation model of utilities, with distributed generation augmenting it, because it's far easier to deploy, far easier to redeploy, and involves far more incremental investment with a far faster payback.
Yeah, that's really inspiring, honestly, to hear that framed that way: innovations that people are developing to serve the demand from coding-assistant tokens right now are potentially—I think highly likely to have—a lot of positive downstream effects in all sorts of other industries that all just need a lot of power in the future.
That's right, Jordan. And think about it. All of us on this call grew up with energy. I don't think we ever worried when we flipped a light switch on, right? We had light to do our homework. We had power for our computers. We had access to the world's resources with the internet, and we could charge our cell phones.
But let's think about the world that didn't have power, or that part of the world that doesn't have power. They live a life of poverty. The same thing is going to happen with AI. Those who can use AI and become really adept at it will create a further divide.
So I think certainly for today, for DG Matrix shareholders, I have to focus on AI data centers. But there's a part of me that's also looking out at the electrification world, and that part says, if you want to leave the world in a better place, you've got to think of the rest of humanity and how you can help them in some way. So, yeah, I hope the AI data center not only drives us to superhuman intelligence, but makes power cheaper for everybody around the world, fusion or no fusion.
You're offering up a lot of options for where we can take this for the last few minutes of the podcast here. Nico, Jeremie, does anything come to mind?
I think it's the double-espresso kick.
7. The SST Product Roadmap
Just one thing I'm curious about, because you mentioned initially that one of the reasons for 800-volt is that we reuse existing supply chains, for example, from automotive. I'm just curious: for your supply chain, do you actually use automotive suppliers and automotive vendors, auto parts, or is it something completely different?
No, we use silicon carbide semiconductors that were developed for 1,200-volt architecture. Could some of those be used in EVs? Yeah, some of those are used in EVs. Do they give us a benefit? Yeah, I think they do.
When you are running these surges, you've got to look at the physics of semiconductor failure, and then you've got to translate that to people who drive EVs and have a lead foot. There's a lot of commonality between all those surges. The people who have designed the physics to accommodate that—there's some magic there.
Silicon carbide or gallium nitride for power electronics?
It doesn't matter. I think right now silicon carbide is more apt to give you hundreds of kilowatts to megawatts. Gallium nitride is coming up; it's more suited for hundreds of watts to kilowatts.
Quite frankly, as I was discussing today in an investor panel, it shouldn't matter to those of us who want to deliver economic value to customers. The question is, which one does a better job? We're agnostic. We've actually been experimenting with both for 10-plus years, and it's just that silicon carbide is more mature at the right power levels right now.
How big can your SST get? Could we see a 10-megawatt unit a few years down the road?
Yeah. Actually, the medium-voltage SST that we're working on, which is 35 kV in and, let's say, 800 or 1,500 volts programmable out, is designed for 10 megawatts in 1 container. It's going to be 1 large container, but it's designed with higher-voltage semiconductors on the front end, a divide-down, and then a lower-voltage stage. I think that's slated for 2028.
In 2027, we're looking at the 6-megawatt SST. Today, of course, we have 400-kilowatt units that we can parallel to create a multimegawatt system.
Where are customers expecting to place that? Is it going to be in the gray space? Is it going to be outdoors?
Well, it's certainly not going to be in the white space. What's interesting is that in 2011, I worked on a product that was bringing medium voltage to the top of a rack. I can't talk much more about it, but that was the first SST—one of the first SSTs that we did.
Really, if you want to reduce the cable to copper or get the most, you've got to bring medium voltage. But there are a lot of safety issues, architectural zoning issues, and whatnot at a national level, so it makes it tough. Maybe China would be the one to get that done first.
But I think raising voltages and bringing power and compute together in physical proximity is one trend that's taking root now.
Speaking of China, are there any issues for you guys sourcing silicon carbide from China?
We're not sourcing any silicon carbide from China. We're just sourcing it from the best folks we can find. Our sources are the United States, potentially Japan, and Europe right now.
Europe has 2 very big suppliers for us. America—right there in North Carolina—has a very big supplier for us, too. That's what we're focusing on.
We are sourcing some non-CPU, non-software electromechanical stuff from China, but we have a China-plus-1 sourcing strategy, so we can get the same parts from, say, Mexico or Vietnam. Like everybody, we're just trying to mitigate future risks.
All right.
Okay. One thing I would mention as we think of all these architectures is that we shouldn't forget that the more software-driven your power becomes, the better your cybersecurity must become, because you don't want third parties to hack into it.
We've developed and deployed cybersecurity-proof power solutions on the transmission grid in the past, and that's a skill set that I think has to expand in the industry. If it doesn't, you have the risk of miscreants coming in and taking your data center down.
Let's make sure that, at some point, we cover this, too: How do you really make this cybersecurity-proof, including background checks on every single entity that touches the electronics and develops the software?
Yeah. We don't want Stuxnet in any of these new, big data centers. Seems pretty important.
That's right.
Yeah. Awesome. Well, guys, thank you so much. This was a whirlwind tour of 800-volt DC, SSTs, and all the implications for the supply chain. Appreciate you spending the time with us.
Thank you very much for the opportunity.
All right. Take care, guys.
Okay. Bye-bye.