[BidClub_]
SemiAnalysis · · 53 分钟

第021期——AI项目三角:资本、承购、数据中心(数据中心、能源)

Dan NishballZane FongKang Wen CheangJordan Nanos

YouTube
TL;DR
  • AI基础设施的约束,正从GPU需求转向资产负债表承载能力:团队测算,2024-2029年资本开支累计达11万亿美元,年支出很快超过1万亿美元,所需融资约7.1万亿美元。 按约75%的债务融资比例和5-6年摊销计算,这一市场规模将远超汽车贷款和学生贷款,仅次于13万亿美元的美国抵押贷款市场;Dan表示,瓶颈在于“数量本身”,而不是单个50亿-100亿美元项目能否完成融资。
  • 5年期投资级超大规模云厂商承购,是目前唯一能广泛融资的模板,但它把短期算力初创公司真正需要的供给堵在门外。 一家手握8000万美元、计划在6-12个月内完成训练的公司,会被告知“如果只能承诺1年,你的8000万美元就不算数”,随后被推向每年1500万-2000万美元、连续5年的合同,最终拿到的集群规模只有预期的1/4或1/5。
  • NVIDIA的兜底机制,让投机性的新兴云算力产能具备可贷款的下行保护,但并不保证股权回报足够吸引人。 以示例性的6年GB300结构看,第1年兜底价格为每GPU小时3.68美元,全期均价2.36美元;短期订单簿起价约6.75美元,均价4.27美元;银行要求DSCR高于1.3,而“整套机制的本意”是兜底永远不会被真正启用。
  • NVIDIA扮演的并非被动的“央行”,而更像一名会挑选受益者、并反复从中变现的做市商。 它出售硬件,分享超过兜底价部分的租赁收入,并将NCP参与资格与NVIDIA网络设备和软件绑定;Dylan的修正很直接:“他们绝对在挑赢家”,尤其是“买它家产品最多的人”。
  • GPU信用利差定价的是3种不同风险——承购方信用、新兴云算力执行风险,以及无抵押平台风险敞口;快速折旧让执行容错空间格外小。 讨论中,CoreWeave对Meta的DDTL 4.0融资利差约为SOFR上浮225个基点,而无抵押债券收益率可能略高于9%;与房地产不同,“一旦出现利用率缺口,就没有太多时间去修复”。
  • 亚太项目说明,全球总量正从一个个集群累积到万亿美元级别。 澳大利亚项目目前为72 MW,计划到明年年中扩至1.2 GW、部署55,000块GPU;Firmus的Batam项目达到360 MW;该地区“供给相对不足”,使广泛可用的短期算力产能具有战略意义。
  • 正在形成的授信护城河,是专有市场数据:双边租赁曲线、亲自下场的新兴云算力评级,以及不依赖公开API报价的每秒token经济学。 客户可能愿意为稳定、同规格的GPU支付20%-30%溢价,而InferenceX将token产出视为“收入的原子单位”;Dylan不认同最终只剩推理的世界,因为推理归根到底是“用来购买更多GPU、训练下一个模型的手段”。
摘要 · 为研究而整理的核心内容

1. AI的资本需求已超出原有融资模型

  • Dan提出的核心问题是融资:AI和数据中心资本开支在2024年至2029年累计约11万亿美元,很快突破每年1万亿美元;在75%的债务融资假设下,需要约7.1万亿美元资金。5-6年摊销期,既匹配合同期限,也匹配预期GPU寿命。

  • 规模对比很关键:汽车贷款和学生贷款市场均在约1万亿至2万亿美元的低位区间,而到2029年,AI债务融资可能达到7万亿美元。只有约13万亿美元的美国抵押贷款市场更大,这意味着贷款机构既需要大幅扩容,也需要更清楚地判断什么可以融资。

  • “AI项目三角”是一个循环闭环:资本通常要求5年期投资级超大规模云厂商承购;拿下该承购,又需要可信的股权资金、押金以及数据中心资源;锁定数据中心,则要说服运营商,或自行建设。早期新兴云算力公司可以募得1000万-1亿美元股权,但10亿美元级项目需要不同的风险承接方。

  • Dylan起初强调项目规模,Dan则把问题进一步明确为数量。挑战不在于贷款机构能否为1个100 MW或100亿美元项目融资,而在于能否“做20个这样的项目”。早期Blackstone/CoreWeave融资在执行成功后,实质上变成了Microsoft或其他超大规模云厂商的信用风险,随后更多贷款机构参与,融资利率也随之压缩。

2. 超大规模云厂商模板无法服务短期AI需求

  • 超大规模云厂商的兜底额度并非无限,也不可能承接数万亿美元的承诺量。私募股权和私募信贷引领了市场,但利差持续压缩最终会让这类机构的回报不足;与此同时,只有5年期的结构无法满足推理服务商和初创公司对1-2年甚至更短周期算力的需求。

  • Dylan举的需求侧例子是:一家前沿实验室分拆公司可能募得8000万美元,预计至少80%花在GPU上,并希望在6-12个月内拿到尽可能大的集群。新兴云算力公司却会说:“如果只能承诺1年,你的8000万美元就不算数”;然后要求其每年支付1500万-2000万美元、连续5年,最终只提供预期集群规模的1/4或1/5。

  • Zayn的示例模型期限为6年,使用GB300,并假设租赁价格随时间下滑,因此兜底价第1年为每GPU小时3.68美元,全期均价2.36美元。短期订单簿起价约6.75美元,均价4.27美元;新兴云算力公司拿到兜底金额,超过兜底价的上行部分再与NVIDIA分成。

3. 兜底机制保护贷款机构,同时保留租赁上行空间

  • 据称,银行要求至少前2年DSCR高于1.3。以100 MW GB300集群为例,单靠兜底机制即可带来健康的前期覆盖率:它不足以保证新兴云算力公司获得“非常高的回报率”,但足以让建设贷款和GPU贷款在没有投资级客户的情况下获得融资。

  • 贷款机构按最坏情形承保——由NVIDIA接手算力——而新兴云算力公司仍可争取价格更高的1-2年期业务。Dylan强调,“整套机制的本意”是没人希望真正用上兜底;它的作用是在多元化客户订单簿形成期间,为现金流设置底线。

  • 如果NVIDIA接手GPU,可将其用于研究、持续集成(CI)以及PyTorch、vLLM和SGLang之间的兼容性工作,也可向Nemotron Training Collective等联盟提供算力,该联盟成员包括Thinking Machines和Mistral。Dylan仍指出,“这里存在某种上限”:外部每小时5美元的需求,优于内部每小时约3.68美元的使用。

  • Dan认为,NVIDIA的介入已超出算力:它还在承租数据中心,并据引用的一份报道,可能买入光纤再分配给新兴云算力公司。他的判断是,AI项目三角的扩张速度超过了私人市场自然拼装各环节的速度,因此NVIDIA成了支撑建设的“央行”。

4. NVIDIA的靶心机制将硬件销售转化为持续控制力

  • Dylan在讨论中实时修正了“央行”类比:NVIDIA并非只是在场外设定条件。“他们绝对在挑赢家”(“They’re absolutely picking winners here”),而且偏好的赢家是能够重复采购的客户——“买它家产品最多的人”;NVIDIA会主动扶持这些客户,而不是等到需要救助时再出手。

  • Zayn的“靶心”从所有硬件买家开始,逐层收窄到新兴云算力公司、NVIDIA Cloud Partners,最后到拥有兜底的NCP新兴云算力公司。每一层内圈都意味着更高的重复采购、更强的标准化和黏性;获得兜底的最内圈,还通过收入分成把一次性硬件销售转化为持续性利润。

  • NCP资格由NVIDIA授予,而非参与者自行选择。Jordan称,这项特权会以物料清单上的额外SKU或料号出现,意味着NVIDIA从GPU服务器和网络设备中获得的支付比例提高。参与者还要承诺采用NVIDIA的交换机和软件,包括用Spectrum-X替代Arista,以及采用NVIDIA品牌的LinkX收发器;Jordan称,后者与“价格便宜4倍”的替代品等效。

  • 因此,这套经济组合远不止利息或租赁上行收益:NVIDIA卖出更多GPU,截取超过兜底价的收入,把网络设备和软件纳入设计,并帮助成功的新兴云算力公司成为重复采购客户。兜底之所以稀缺、备受追捧,正是因为它将融资准入与NVIDIA的资源分配及参考设计特权绑定在一起。

5. GPU债务定价的核心是执行风险,而非抵押品价值

  • Kong的信用入门以约4.3%的10年期美债收益率和一个假设的50个基点Microsoft利差溢价为起点。讨论中,CoreWeave无抵押债券收益率可能略高于9%,而其Meta DDTL 4.0项目融资利差约为SOFR上浮225个基点;团队将其中约97个基点归因于Meta风险,105个基点归因于CoreWeave执行风险。

  • 这部分执行利差覆盖CoreWeave能否按时建设、能否正确安装并联网GPU,以及能否持续满足服务要求。GPU抵押贷款机构在破产时的受偿顺序先于无抵押债券持有人,因此后者承担平台风险敞口;讨论认为,这需要再增加约400个基点。

  • Kong将NVIDIA支持的贷款定位在5年期超大规模云厂商承购与无抵押新兴云算力信用之间。银行仍需追问运营商将如何建立客户订单簿,因为NVIDIA兜底“不可能永远存在”;Kong认为,这是一段试用期,最终会走向独立融资,半导体晶圆厂、工厂、航空公司和其他企业都是如此。

  • Dylan的反驳值得保留:1只Hermès包相较于其工厂本身很小,但1个空置的200 MW集群就像一份巨型WeWork租约。Kong同意这存在连续谱,指出航空公司即便资产期限为15年、机票却只提前几周售出,仍能获得融资。他也承认GPU更为苛刻:快速折旧意味着“一旦利用率出现缺口,就没有太多时间去修复”,而Dylan称,银行可能会将2-3年后的残值视为0。

6. 区域项目和专有基准数据将检验这一判断

  • Kong提到的区域案例Firmus,起步于一个新加坡集群,STT GDC为投资方之一,随后借此拓展至墨尔本和塔斯马尼亚,并自行建设数据中心。Kong认为,下一座设施将与DayOne合作,而不是跳过数据中心这一环节,这说明股权、产能和执行履历可以逐步解决三角中的各个部分。

  • 已披露的规模很惊人:澳大利亚目前有72 MW,计划到明年年中达到1.2 GW、部署55,000块GPU;Batam项目达到360 MW,被称为全球同类项目中规模最大。Dylan指出,通常不在数据中心叙事中心的地区也有“数百亿美元”级项目,全球总量正是靠这些项目累积到万亿美元。

  • Kong认为,该地区供给不足,预计会有更多产能向更广泛客户开放。假设兜底机制的本意是扩大可获得性,那么立即再配上一份5年期承购协议就没有太大意义;运营商可能更愿意服务更广泛的客户,而不是把可能约一半的收入上行空间让给NVIDIA。

  • 贷款机构首先需要一条公允的租赁曲线。Kong称,团队花了近3年收集不同期限的双边合同价格,因为这些价格无法从公开API报价中抓取:“每个数据点背后都有故事”,包括客户质量、预付款、交付时间,以及名义上相同GPU所捆绑的服务。

  • Jordan称,另一个缺失的市场机构是“评级机构”。团队对200多家新兴云算力供应商进行实地工作,并访谈超过150名客户,核查健康检查、监控面板、托管Slurm、托管Kubernetes、token服务和RL环境。Dylan以Nebius首笔债务融资的定价接近CoreWeave为例,说明银行认为两者的执行风险相近。

  • 租赁市场呈现期限倒挂:立即可用的GPU享有大幅溢价,而交付期为3-6个月的GPU则以折价定价,用于锁定需求。预付款和里程碑付款可以减少对过桥或建设融资的依赖;可靠供应商对同样GPU数量可收取20%-30%更多,因为正常运行时间和性能比表面价格更重要。

  • InferenceX用每GPU每秒token产出和假设的token价格,重新定义残值——“收入的原子单位”。Dylan不认同训练会被彻底边缘化:2018年的预测认为,2025年推理市场接近100亿美元,边缘推理为100亿美元——“我认为是50亿美元”——数据中心训练为40亿美元,规模和比例都判断错了。前沿实验室认为自己正在打造“机器之神”,因此推理最终会成为购买更多GPU、训练下一个模型的手段。

Dylan Patel

Hello everyone. Welcome back to SemiAnalysis Weekly. We're here with episode number 21. I've got the crew from Singapore to talk about everything NVIDIA backstops, including the holy trinity: capital, offtake, and data centers. Starting off with Dan, chewing away on his lunch. How's it going, man?

Dan Nishball

That was dessert. That was apples. I'm eating healthy dessert.

Dylan Patel

Good, man. With Dan, we've got Kungwen. How are you doing?

Kungwen

I'm doing great. Doing great.

Dylan Patel

And Zayn, how are you, man?

Zayn

Yep, I'm doing well. Thank you for asking.

Dylan Patel

Awesome. In this episode, we're going to talk about NVIDIA backstops. We'll start by discussing the problem statement, explain what a backstop is, how it's structured, how people are pricing GPU loans, as well as other things to backstop these data center build-outs. We'll go through a couple of examples in the Asia-Pacific region, some of the tools that lenders are going to need in the future, and finally, the impact on NVIDIA's financials.

There's no better team to dig into this than our team that works on Cloud TCO every day, including Mr. Dan Nishball. Dan, can you start us off with the problem statement? What does that trinity—capital, offtake, and data centers—actually mean? And why are these backstops so important right now?

Dan Nishball

Yeah, let me start off with how we look at capex. If you look at AI and data center capex, it's soon going to reach very large levels. Cumulatively, it's about $11 trillion from 2024 to 2029, and we're pretty soon going to cross $1 trillion of annual capex.

The debt's already been building up—hundreds of billions per year—and it's soon going to reach over $1 trillion per year. We well understand the amount of capex coming through, and I think markets are starting to understand it, given how much debt hyperscalers are issuing. There's a bit of angst over whether they continue, but the point is, the demand is there, the production is there, and everyone's gearing up for $11 trillion of capex. The question is: How do we fund it?

How much funding will be needed? It's about $7.1 trillion. The way we look at that is approximately 75% debt financing, and you can think of that as amortizing, so it's going to roll off. You have a bunch of stacks of debt that then roll off, and typically it's about 5 to 6 years. It'll match the contract duration or the expected lifetime of the GPU.

Of course, all of these are very debt-financed. It makes sense: It's very capital-intensive, so it should be debt-financed. We'll go through some examples later of what the pricing could look like.

This really is the problem statement. The other important problem statement to understand is how big this market is going to be in relation to all other markets. Here, what we're showing is AI debt financing compared to all other U.S. asset-backed markets.

AI debt financing, including data center and GPU IT capex, is going to reach $7 trillion by 2029. Auto loans and student loans are all in the single digits—around $1 trillion or $2 trillion. The only market that's bigger is the U.S. mortgage market, at $13 trillion.

The point is, there's going to have to be a lot more capacity quickly, and there's going to have to be a lot more understanding. One of the biggest problems so far—and we'll talk through the objective of why this exists—is what can get financed.

What can get financed today is mostly a neocloud that has a 5-year offtake with an IG hyperscaler. Anything else is very, very hard to finance.

I'll take a step back and introduce what we call the AI project trinity: capital, offtake, and data center. Any neocloud that has to build a project has to figure out how to solve these 3 things.

For capital, in order to get capital and lending, you need to have a 5-year offtake from an IG hyperscaler. But to win that from a hyperscaler—to get a 5-year IG contract—you first have to demonstrate that you're a going concern, that you can find equity to place those deposits, and that you can definitely secure a data center.

On the other hand, to get a data center, you've got to convince a data center operator to want to rent to you, or you build it yourself. These are 3 legs, and the easiest solution has been: Let me find a hyperscaler offtaker for 5 years, get the capital, and have a small pilot project to convince everyone.

Dylan Patel

So, Dan, the thing that jumped out to me is that this has changed. At the beginning, when we saw a bunch of these neoclouds getting started, they could raise equity. To me, the scale allowed them to raise $10 million to $100 million, maybe, to get their first data center off the ground.

But these are enterprise investors with a totally different return profile that they expect, and just a totally different amount of capital you need to raise when these projects are measured in the billions. That's an obvious fundamental thing that has changed, and why you now need investment-grade hyperscalers signing 5-year contracts in order to even be able to raise the money. People don't want to take that amount of risk on things measured in the billions.

But are there other things besides the sheer scale that you're seeing change as people start trying to fund all of these projects?

Dan Nishball

Well, I think it's not necessarily the scale; it's the volume. It's not just, can I do a $10 billion or a $5 billion project, or can I do 100 megawatts or 200 megawatts? It's more: Can I do 20 of these?

When this first started, there were a couple of private equity pioneers, famously Blackstone and CoreWeave. At the time, it was actually like a 5-year backstop. But at the time, it was really only these private equity guys who were willing to look at that.

Eventually, their conclusion was effectively that it was Microsoft risk, or whichever risk. If the execution went well, it effectively turned into hyperscaler risk. The loan was at a pretty good rate, I think, but if the execution went well, it effectively turned into hyperscaler risk.

Now, I think the market has matured, and many more lenders are willing to do this, which has driven the rates down. To answer your question, what is the limiting factor? It's the sheer quantity.

Even if some lenders are willing to do this—and it's more than just private equity; it's actually getting a lot of lenders to do this—even if it's a 5-year backstop, there still aren't enough of them.

Dylan Patel

Makes sense. We've described backstops at a high level. Zayn, do you mind taking us through how these backstops are actually structured?

Dan Nishball

Well, before that, actually, before we go to that, let me talk about the other problems. I talked about where the market is today. Today, you can finance in large quantities only a 5-year IG backstop.

What are the problems with that? What we said a moment ago is that, to get to $11 trillion of cumulative capex, we can't just have that happen through a bunch of hyperscaler backstops. That's the first problem with this market structure: Hyperscaler backstops are not infinite. They're not going to be able to absorb $2 trillion, $3 trillion, $4 trillion, or $5 trillion, and not everything can be backed up by hyperscalers.

Secondly, as I mentioned a moment ago, private equity and private credit have led the charge, but as spreads compress over time, the returns are going to be too low for them. The capital needs will increase, and it's going to need a broader set of capital providers.

The other problem is that it's really a one-sided market. It's only supplying compute to 5-year offtakers and no one else. Inference providers can't get the GPUs they want. Dylan, you could probably talk much better about this than me, but there's a total lack of variety. People can't get 1-year or 2-year terms; they can't just get what they need.

Dylan Patel

It's all the startups. Everybody who's trying to build—who's left one of the frontier labs and is trying to build a new model in a new modality, whether they're working on materials science, drug discovery, video generation, robotics, world models, or anything that the frontier labs aren't exclusively focused on—raises money from VCs with the promise that they're going to rent a ton of GPUs.

They're probably going to spend 80% or more of the money they've raised on GPUs, and they really want the biggest cluster they can get for about 6 months to a year. But the answer they're getting from neoclouds right now is that your $80 million isn't good if you can only commit to a year.

Instead, they say, give me $15 million to $20 million per year over the next 5 years, and the size of the cluster you expected to get has been divided by 4 or 5. That's how we can actually get a deal done.

That wasn't really the promise of neoclouds up front. At that point, you're waiting 6 months, spending all this extra money, and not getting a big enough cluster.

It's not really a venture-style approach where you just take a big shot, try to train a great model, raise an extra round of funding, and then go even bigger. You're basically ending up in a world where you have to run a long-term, sustaining business, which, again, is not what you raise money for. So it looks very different from the way it did 3 or 4 years ago, when people were starting these companies and looking for compute.

I think it's due to the scale that they're trying to operate at now. There's just so many of these companies, like you said, Dan, but nobody's willing to take the risk of speculatively buying $1 billion of GPUs and then hoping that everybody gets another round of funding in a year or two.

However, NVIDIA has obviously recognized that this is a need in the industry, and they're going to come in and do something about it. Maybe you guys can take us through what that actually is.

Dan Nishball

The backstop, right? Enter the backstop.

Dylan Patel

You know, Zayn will take us through exactly what it is and how it's going to solve that problem. Zayn, you want to explain?

Zayn

Yeah, for sure. So, the way that this backstop is structured, typically we see it being structured at a 5- or 6-year duration. Here, we have it at 6 years for a GB300.

What NVIDIA does is basically say that they are ready to purchase from the neocloud at these pre-agreed levels. So, let's say on a year-one basis, it's $3.68. If the neocloud is able to rent it out for something like $6 instead, then everything up to $3.68 goes straight to the neocloud, and anything above $3.68 is split between the neocloud and NVIDIA.

The reason why we have this at a declining ramp profile is because of the naturally expected decay in GPU rental prices. You'll notice that the average backstop price here is only $2.36, which is very, very low for a GB300. We'll delve a bit into the rationale for that later on, but for now, suffice it to say that NVIDIA is not doing this to guarantee the neocloud a really high rate of return. They're more just doing this to make all these clusters financeable in the first place.

That's the overall structure of the backstop. We've devised a few scenarios in which this backstop could be used. The first scenario is that we're running a short-term rental book for GB300s, and obviously that's going to get you pretty high rental rates here.

We've plotted that out at a 6-year average price of $4.27, starting from $6.75 in the first year. Once again, the customer will pay the neocloud $6.75. The NVIDIA backstop price for the first year will be $3.68, so at $3.68, the neocloud gets that, and anything above $3.68 would be subject to a certain revenue share between NVIDIA and the neocloud. We've synthesized that in the various scenarios here.

Essentially, it's pretty intuitive, as you would expect. If we jump over to the IRR section, the more of that percentage revenue share goes to NVIDIA, the lower the IRR for the neocloud. But essentially, this is what we see the backstop doing: It enables the neocloud to take a bit more speculative risk, renting to a broader group of customers, as we just mentioned, and renting on a shorter-term rental book.

In this case, the neocloud doesn't necessarily have to go out and find a 5- or 6-year, investment-grade hyperscaler offtaker. It's able to rent to a broader group of customers.

Dylan Patel

So, go ahead, Jordan.

Jordan

Yeah, I'm curious. Obviously, NVIDIA is kind of double-dipping here, right? Because on the one hand, it's the backstop, but they're going to make money off that loan. On the other side, there's the potential to make upside for them when they can actually go and connect buyers and sell above the backstop price. So maybe you can talk about the other side, in terms of how the loans are getting priced.

One thing that might be interesting as a small detour is to go through the DSCR analysis, because I think the important thing is: Why does this structure work for banks? Why does it enable them to lend without an offtaker? This is how they analyze it. Zayn, if you want to explain how banks are able to take this backstop and lend against it.

Zayn

Yeah, sure. No problem. We've run a DSCR analysis of the same—actually, just the backstop price itself.

What we've been hearing from banks is that, in order to finance this kind of arrangement, they would need the DSCR ratio to be above 1.3, at least for the first couple of years. So here, we've put in our backstop price, just on a cluster of 100 megawatts for GB300. Of course, this is illustrative. Then you get the following DSCR ratios, which are quite healthy in the first few years.

In other words, the NVIDIA backstop price is not sufficient for the neocloud to earn a very high rate of return on the GPU rentals, but it is sufficient to give lenders the kind of comfort they need to be able to finance projects like these. That's where NVIDIA aims to land with this kind of backstop pricing.

Dylan Patel

So effectively, they're lending against the backstop, and what they're thinking about is: Worst case, NVIDIA will take the compute. They're thinking, in the worst case, what is the cash flow I'm going to get? If that's my worst case, I can lend against that.

If the neocloud wants to go and not avail itself of the backstop—because, again, they don't have to—the whole idea, by the way, is that they don't ever use the backstop. No one actually wants to use the backstop. Everyone wants to be above the backstop. Neoclouds want to do a 1- or 2-year business, but for the banks, they're saying, "Okay, what is my worst-case scenario?" That's how they lend.

Jordan

Yeah, let's say NVIDIA needed to use a backstop. What would they even use the GPUs for? Are they just paying money for the sake of backstopping these GPUs?

Dylan Patel

Plenty of research. You've seen the Nemotron models come out from NVIDIA. They do research on all sorts of different smaller things, and they've got a massive CI fleet to make sure that anybody's code in PyTorch, vLLM, or SGLang is actually going to run well out of the box.

In part, any GPUs that NVIDIA produces, rents for itself, and uses are going to improve the user experience for other people using those GPUs. In some ways, it contributes to the CUDA moat when they have GPUs for themselves to do research and engineering against.

Now, there's some limit. They don't need hundreds of megawatts to do research. Maybe they do—they have plenty of ideas—but what we've seen from them is that they have taken those GPUs, possibly, and started to form consortiums, namely the Nemotron Training Collective, where you bring together a whole bunch of different startups and organizations, like Thinking Machines and Mistral, and say, "Hey, you guys can use our GPUs. You can contribute if you're willing to work on open source."

So I think they have plenty of uses for these GPUs. It's very rare that you're going to see any GPU cluster that's just sitting around idle, collecting dust. But clearly, the point of NVIDIA's business is to rent this to other people who are interested in paying something beyond what the backstop would be at the base level.

Anybody at NVIDIA would prefer to rent it out to a hot startup that's interested in training its model on its specific domain, with all its expertise, for $5 an hour in year 1, instead of having NVIDIA research engineers use it for $3.68 or something like that. Anytime they can support a neocloud actually renting these GPUs to startups, it will guarantee that the neocloud is going to be able to grow its business and then buy more NVIDIA GPUs in the future.

I'm making the best case for what they're doing here, I think. But generally speaking, Dan, I think you called them the central bank of AI in the article, right? Certainly, for startups and certainly for NVIDIA GPUs, that's true. They are acting as a central bank that, in some ways, is a market participant, but in other ways is an external party that's just watching everybody else buy and sell from each other and trying to grow the market—grow the pie in total—rather than pick winners, right?

Dan Nishball

And it's not just that, right? They're also taking data center leases. There was a report that they may be buying fiber in order to then reallocate it to neoclouds. There are a lot of areas where they realize that AI is just growing way faster than this trinity can organically develop.

The central bank's purpose is to step in and support the economy when private credit can't, which is the case here, right? It's just growing way too fast.

Dylan Patel

Yeah, and let me refine something that I said a second ago, where I said that they're not picking winners. They're absolutely picking winners here. Much more so than the central bank does, even though central banks can sometimes pick winners in terms of who they're going to bail out.

NVIDIA is not waiting for bailouts. They're actively going and picking the winners, and those winners are the people that they like the best.

Dan Nishball

And who do they like the best? The people that buy the most of their stuff. So, this is a tool that they're using as an active market participant to support people, not necessarily in the same way that a central bank just kind of sits back and sets rates.

Why don't we, Zayn, jump down to the bullseye diagram? I think that'll illustrate Jordan's point really, really well, because NVIDIA can decide who they're going to sell to. So, Zayn, do you want to explain our bullseye concept?

Zayn

Yeah, that plays into what Jordan said. Actually, now that I realize it, I didn't think about that, but it is a bullseye, right? It's their target, right?

Dylan Patel

Yeah. In a way. In a way. Yeah. This is probably the least technical chart in the whole article, but I hope at least we understand this. I think the way that we think of it is a few concentric pools of demand, right? These are basically people that can buy NVIDIA's hardware. In the first and widest pool, you have all buyers generally, and here the pool is just for NVIDIA to sell hardware to them and get the margin on top of that. So, this is the first layer, the first circle.

Of course, the inner circle is neoclouds—that's the blue one—and they take those chips and turn them into a rental business. They're repeat buyers of NVIDIA hardware. They continue to buy it, continue to rent it out, and make it available to the ecosystem.

The red circle is neoclouds that are NCPs. So, here is your NVIDIA-certified cloud partner tier, and then NVIDIA gives them reference designs and priority allocation and so on. But in return, NVIDIA gets standardization and a sort of stickiness with the NCPs, right?

Lastly, you also have the innermost circle in green, which is the neoclouds that are NCPs with a backstop. Here, it's probably the stickiest, right? NVIDIA helps to credit-support that cluster via the backstop that we just mentioned, but they also take a cut of the revenue above the backstop. So, this one-time hardware sale ends up being a recurring-margin business because of the revenue share as well. That's our bullseye diagram. And of course, in the green, you have the NCPs with the backstop.

Dylan Patel

Makes perfect sense. Who are they going to sell to? People where they can double-dip, right? To your point, Jordan.

Jordan Schneider

Yeah, they're not exactly throwing darts at this, or if they are, they're hitting 180 every time. So, yeah, bullseye.

Anyway, there's lots of people in this industry right now pursuing a backstop, let's say, and trying to maneuver so that they can get this. It's definitely sought after. There's only so much to go around.

Maybe the one comment I can make is that neoclouds that are NCPs in the green circle in the middle, like you said, get standardization stuff. People can't just choose to be an NCP; they're granted the privilege by NVIDIA. And that privilege represents itself as a SKU, like a part number on a BOM, that represents an increased percentage that you are paying NVIDIA for your GPU servers and networking.

You're also committing to purchasing all of NVIDIA's switching and all of NVIDIA's software. And, as you said, standardization also means no competition in these designs. You're not going to be able to just put Arista switches in your back-end network here, right? You're going to be using the Spectrum-X switches if you want a backstop. You can't use somebody else's transceivers; you've got to use the NVIDIA-branded LinkX transceivers that are the exact same as the cheaper ones that are 4 times less expensive, right?

But it's a way to keep the ball moving for them. So, good segue. Let's keep the ball moving on our side as well. Kong, we can talk about how these GPU loans are actually being priced by the banks, and then maybe we can give some examples of backstops that you guys are most familiar with. So, we can talk about the actual projects instead of just this at a high level, saying neocloud over and over.

Kong Yang

Yeah. And we also have a view on where the backstop lending might price. So, yeah, it'd be interesting.

Yeah. I guess some context into thinking about credit spreads is that the way that you would price a credit bond is, you have the underlying—whatever your government bond is, right? U.S. Treasuries, maybe at 10 years now, are priced at 4.3%. And then you have to think, if I'm going to lend money to a company, how much riskier are they than the U.S. government, for example? So, maybe if I think that Microsoft is only a tiny bit riskier than the U.S. government, I might charge them an extra 50 basis points, or an extra 0.5%, to lend them money at the rate I'd lend to the U.S. government. It's this interesting idea where you can decompose the level of risk when you're lending to a company, where you can understand why whatever credit is priced the way it is.

For example, if you take a look at this CoreWeave chart here, the context is that CoreWeave has priced multiple kinds of debt that's available in the market today. The line that you see at the very top, the CoreWeave blue or purple line, is CoreWeave's unsecured bonds. That's trading in the market today at probably just over a 9% yield. And you can see the brown dotted line below: what CoreWeave would pay in interest if they managed to project-finance that. In this situation, the debt is secured by the GPUs, and they have an investment-grade customer. You can see below all the different multicolored lines representing the different investment-grade hyperscalers, showing the kind of customers that CoreWeave has. So, you can see the 3 big brackets there. That's how you would decompose the credit risk of a CoreWeave debt, for example.

Illustratively, if we think about CoreWeave-Meta DDTL 4.0, that was priced at about 225 bps over SOFR. What that means is it was priced at 2.25% more risky than what the U.S. government, illustratively, would be. And if you try to decompose that 225 basis points, obviously some of it will be the risk of Meta. Meta obviously has to pay CoreWeave; if not, CoreWeave's not going to get money to pay the bank. So, that 97 basis points is what Meta bonds are trading in the market over the U.S. government bond, and that is the risk that is implicitly taken by the bank when they have Meta as a customer.

Then that begs the question, right? If we know Meta is good for this 97 basis points of risk, why is there an additional, call it, 105 basis points on top of that? Well, that 105 basis points would then all be reflected as CoreWeave execution risk. What happens if CoreWeave doesn't stand up the cluster on time? What happens if CoreWeave misses the SLAs constantly? What happens if CoreWeave has issues with reconfiguring or installing the GPUs? That's the kind of risk that banks are taking on when they price this additional 105 basis points. And that comes to a total rate of 225.

Then we see the unsecured bonds, an additional 400 basis points over that execution risk that we talked about. So, these bonds are not secured by the GPUs; they're merely subordinate to the specific GPU debt that's being financed. What this means is that, okay, let's say CoreWeave goes into bankruptcy. The first debt they pay off is the GPU debt that's secured with the loans that are already backing it. Then whatever money is left over goes toward the unsecured bonds held by debt investors in CoreWeave's unsecured bonds.

This is how we think about decomposing the various spreads within different neocloud debts. And I think that's quite interesting because that lets us think, okay, in the future, if there are more and more of this kind of debt being raised, what would be the appropriate rate they should lend at?

Dylan Patel

Yeah, it makes perfect sense. In my view, it's quite simple to see that hyperscalers also have an entire other business that's a free-cash-flow monster, and CoreWeave's entire business is this GPU-rental thing. So, they're obviously going to have different credit risk, but the spread is fascinating when you break it down like that in the chart.

Kong Yang

Yeah, exactly. The hyperscalers are good for the money. They're getting unlimited money from all different kinds of businesses. But for CoreWeave, the second Meta doesn't pay, you know something's going to go wrong.

Dylan Patel

Yeah.

Jordan Schneider

Yeah. And then I think, Kong, you had a pretty good chart where we did the 5-year pricing by deal structure.

Kong Yang

Yeah. I think in this one, we more explicitly try to decompose these different things and then suggest where we think it would be. On the right, you can think this is a bit more comparable to unsecured lending to CoreWeave, where you're taking platform risk and execution risk. And then you've got the base rate, which is just risk-free, on the left. That's the other extreme: you're not taking platform risk as much; you're just taking execution risk, which is the backstop—or, I should say, not the backstop, the hyperscaler offtake 5-year contract.

And then I think NVIDIA is somewhere in the middle, where you've got to believe in their execution. You're taking NVIDIA credit risk, but then there's also an element of whether they manage to create a good pool of customers. And I think the whole idea is we want to get the market comfortable with something in the middle.

And you can think of this as a trial period, an initial period for them to really understand how the neocloud runs its business. Banks are not just going to look at the NVIDIA backstop and call it a day. They're going to start asking questions: “Okay, well, how are you going to build a book of customers? How are you going to build this business?” Because the NVIDIA backstop is not going to be around forever.

The point is, it's only meant to be there for the time being, to allow banks and everyone the time to really understand the neocloud business model and the risks. Then, one day, they'll have to lend on a standalone basis, just like they lend to every other business—just like they lend to semiconductor fabs, just like they lend to watch factories. Every business has duration. It's a bit ridiculous when people say, “Oh, but there's duration risk and depreciation.” That's every business on Earth. Every business on Earth doesn't know if they're going to have a sale the next day.

Dylan Patel

Yeah. Most people aren't building bridges, I guess.

Kong Yang

Yeah. Even Hermès doesn't know whether they're going to sell bags next month.

Dylan Patel

So what's the difference? Actually, I was thinking about that point you raised, and an interesting counterpoint is that each Hermès bag is small relative to the overall cost of the factory. So, for example, if you think of WeWork, if they can't fill a single office lease, that's a huge proportion of the cost that they have paid for the building.

Kong Yang

So it would be similar in this situation, right? If, let's say, I stand up a 200-megawatt cluster and I can't find customers, that's obviously a way bigger issue than if I can't sell one Hermès bag. Yeah, totally. There's a whole spectrum. You have businesses which are fortunate enough to have a foundry model, so they're not the ones who bear the depreciation. WeWork's a great example—you've got property.

I think, to your point, there is a spectrum, but it doesn't mean no lending at all. There are worse things that are super variable, like airlines, but there's a ton of lending for airlines. There could be a pandemic at any time; there could be travel disruptions. That's one of the hardest things to get right, and there's still plenty of lending to airlines. It's even worse because it's like 15 years, and you don't even know—you’re selling tickets 2 weeks in advance. But one thing you do have a point on is that it's a very brisk depreciation. If you do get utilization holes, there's just not a lot of time to fix them, right? If you're not executing, then time ticks very quickly. Do you think that's fair?

Dylan Patel

Yeah, exactly. I think that was what I was thinking of: you have little time to make it right if something goes wrong. Like, if you're in commercial property, you have 99 years to figure it out.

I think the other point is that banks are also a bit scared to underwrite what happens if something goes wrong. If, let's say, a building developer defaults on that debt, they're thinking, “Okay, how much can I sell this building for in 2 or 3 years?” But for banks, when they approach GPUs, they don't have that same mindset. It's more like, “I have no clue what these GPUs are going to be worth in 2 or 3 years, so I'm just going to take it as zero.”

Kong Yang

Yeah, I think—but you're right. Residual value of offices is also a lot easier. Residual value of airplanes is a lot easier because it's not like every year Boeing has something that flies twice as fast, right?

Dylan Patel

In 9 years. Although the seats get twice as—half as much every year. I wish we had Moore's law for the internet. In 9 years, we just start teleporting. We just double every year how fast this is.

Okay, can you guys give a real example here? We've talked about it conceptually across these other examples and comparisons, but you walk through a couple of real examples in the article, specifically in Australia and Indonesia, with a couple of neoclouds. Do you mind talking through those examples?

Kong Yang

Yeah, I can talk about Firmus, which is the largest one in the world. It's a 360-megawatt project in Batam. Yondr actually announced theirs first, and then Firmus followed shortly thereafter, just, I think, a week before we wrote this article. I don't remember exactly, but anyway, Firmus is interesting because they started off with a cluster in Singapore that was very much like neocloud version 1.0. They raised a bunch of capital and solved for the trinity originally because they had STT GDC invested in them. So they had the data center sorted out, they had good investors, and they could take a bit of risk.

Then they parlayed that into the large cluster in Melbourne, and they also have one in Tasmania. What they did to solve for the trinity there is they actually self-built the data centers. That sidesteps the whole conversation around data centers. We can talk a little bit more about that. I believe the next one is actually not going to sidestep that. They're going to work with DayOne on that facility.

If that's the objective, as we outline—if we're right about the objectives—then it's going to be pretty interesting, because it's going to be one of the largest sources of compute in the region, and it's going to be broadly available to all sorts of folks. Assuming, again, that our assessment of the spirit of this thing is right, the spirit is not to take the backstop and do a 5-year offtake. You probably don't need to pay NVIDIA half the revenue, right? So it almost stands to reason that anyone doing this is probably going to want to do something a bit more broad-based. Yondr, too. I think we'll see more in the region. I really feel that this region is somewhat underserved by capacity in general, and there's certainly a lot of data center capacity coming up.

Dylan Patel

At a high level, it's unbelievable to see Australia—not the biggest region in the world—getting 72 MW, with a planned 1.2 GW to scale to. That's 55,000 GPUs by the middle of next year. Firmus's third major project, after Melbourne and Tasmania, is Batam in Indonesia, which is 360 MW. That's tens of billions of dollars of capital expenditure all going to Indonesia to build out GPU capacity. Again, it's just a massive amount of investment in a single area that not a lot of people think about when it comes to data centers.

And to go back to your point at the very beginning of this conversation, that's how we add up to trillions of dollars: projects like this all around the world that need to get financing.

Kong Yang

Yeah. It's interesting because we're having lots of conversations right now with banks that are asking, “How do we do this for the first time?” Even with a backstop, this is the first time they're really thinking seriously about understanding neoclouds. There's a bunch of tools that we think they're going to need, and I can talk quickly about that.

I guess, if you scroll—let's see. I think scroll up, right? Oh, yes. Tools that GPU lenders are going to need. They're going to need a bunch of tools. If they're going to put trillions to work, they're going to need a few things. They're going to need to know what a fair market price for rental is.

We've actually been working on that for almost 2 years—almost 3 years, in fact. We've been tracking bilateral contract prices across the entire term structure. This is deal information that's basically impossible to get; you can't get it by scraping. You're not going to find a posted price. You really just have to gather it by surveying a lot of neoclouds and asking, “Hey, what do you think a B200 could go for?”

Jordan also helps out a lot in terms of connecting people to compute. He and the Sams have a pretty good sense of where things will transact. That all goes into composing this index. It's a very different approach from others. It's not very scraping-focused, but we really believe that every data point has a story, and we try to recount those stories. We try to explain a lot of the research we do and get a sense for the market.

The second thing is which neoclouds are going to execute well, and which ones are going to have the lowest execution risk. If we scroll down, this is what Jordan does. Maybe, Jordan, you should explain why this is so important to whether you would want to lend.

Jordan Schneider

Yeah. To us, there's basically a ratings agency that's missing here when it comes to the quality of the product that these people provide and, therefore, the quality of the debt. The only way that you can get an understanding of that is to do hands-on testing and to talk to the existing customers in a real, unbiased way. That's exactly what we do with Cluster.

So, for everybody listening, stay tuned for ClusterMAX 3.0. The testing deadline is August 1, so we're almost there, and we'll be putting out an article in mid-August to update these rankings based on our hands-on testing and interviews with over 150 customers of these neoclouds. In some sense, it's hard because you've got to go and talk to over 200 of these neocloud providers to get a sense of what everybody offers. We're always trying to connect with more, so that number is growing, but in another sense, it's easy to get a sense for who's got their stuff figured out and who doesn't.

We have a whole checklist of criteria that's public on our website. We've told many of these neoclouds for months or even years about certain things that they're missing—big differentiators for why people would want to buy from them and things that would improve the quality of service they can provide. Many people just don't do it. It's really simple stuff, so in that sense, it's pretty easy to differentiate between people who have health checks and people who don't, people who have monitoring dashboards and people who don't, and people who offer managed Slurm and managed Kubernetes and people who don't.

As they expand into other businesses, many of them are serving tokens as a service or serving RL environments and just going up the stack. Many of them are going down the stack, where they do self-built data centers, like you're referring to with Firmus. They need to be able to assess their competency across all of those different facets of the business, and that's exactly what we do with ClusterMAX. Anyway, it's a fun project to work on because we get to learn so much.

And Dylan, to your point, when you're trying to draw a curve to show the pricing of stuff and how it changes over time, some of these individual stories are worth so much more and should be weighted so much more heavily than one spot price that anybody can get from hitting a public API to see that a vendor has chosen to price that GPU instance type at a certain price on a certain day. Talk to somebody else who's rented—who signed a $20 million GPU contract—and talk to them about the decisions they were making at that time. You're going to get a much better sense, first of all, for what that contract is priced at, and second of all, for what they care about when they're choosing a provider and how they're going to discount certain providers that can't provide a high level of service versus paying a premium for those who can.

Then compare that to some on-demand price that somebody's paying for a single H100 that they're going to rent for 6 hours to train their anime bot or something. It's a very, very different business. Anyway, in summary, the ratings agency is basically missing here, and we're doing our best to help people solve that.

Dylan Patel

Yeah, I thought this section of the note was pretty prescient because, call it 2 weeks after we released this note, Nebius announced they had raised their first debt facility, right? It turned out that it was at a similar rate to CoreWeave, suggesting that banks thought execution risk was similar to CoreWeave. If you look at the table we have here, we have CoreWeave and Nebius just side by side—

Jordan Schneider

Makes sense.

Dylan Patel

—and it's the pricing of debt but also the pricing of their products. It's going to be important that, if you're lending, the question is: are you lending to, like, a Hermès, or are you lending to—I don't know, what's it like, an H&M? Are they going to be priced high or low? That's a real thing, because if they're doing Kubernetes and Slurm versus just bare metal or whatever, there's a price difference, right, Jordan?

Jordan Schneider

There's a price difference between what you can deliver with an SLA, so it's important for lenders to know what they're getting into.

Dylan Patel

Yeah, look, you guys have called this out: there's this backwardation curve where people who have GPUs for sale that are available right now are charging a massive premium, and those who have delivery in 3 months or 6 months provide a discount just to lock in the customers. The big variable that people play with right now is the prepay amount.

Having those milestone payments as you're going means these neoclouds don't need to raise as much bridge financing or take out a construction loan, or just change the structure of the debt that they're raising, because they can get these prepayments from customers that trust them. That changes their financials completely, so that's a big thing on the pricing.

But after you get past that, 2 people offering the exact same GPUs in a similar location with the exact same headline price are actually going to deliver very, very different things. We've talked about this on previous episodes, to the tune of people being willing to pay, in some cases, a 20% or 30% premium for the exact same quantity of GPUs just because they stay online and perform, compared to some other providers who are less trustworthy. That's represented in this chart. The high level is represented there. But for anybody who's looking to dig into more details, please get in touch with us because we do a lot of this consulting work for people who are making these critical decisions.

Right? When you're making a decision on where to send $50 million of the VC money that you've raised and kind of bet your company on some neocloud, you really need to be able to trust them. We're just providing people with that peace of mind—to know what they're getting or to think through the decision: how to negotiate, how to get through the SLA discussion, how to draft a contract, and what sorts of things to get the neocloud provider to commit to. These are big questions that we help people answer.

Jordan Schneider

Yeah. The last piece of the puzzle, if you scroll down a little bit, is InferenceX. A lot of people ask, what is the residual value of a GPU? It gets into what's a GPU worth. You could argue that what a GPU is worth is what it can generate, and that's how many tokens. InferenceX tells you how many tokens per second you can generate per GPU.

Folks can make assumptions about how much the tokens are going to cost. This is really important because it's the fundamental revenue-generating capacity, just as much as how much bags cost, how much flight tickets cost, and how much nuts and bolts cost. This is really the unit cost—the atomic unit of revenue for the entire industry.

Dylan Patel

Yeah, everybody pays for tokens right now. Some special customers rent by the GPU-hour when they're the ones training the models, but the whole ROI of the industry depends on people buying tokens. That's this chart right here.

I thought inference was taking over from training, though, Jordan. I thought training was going down.

Jordan Schneider

Yeah. Okay. You want to have a whole other podcast right now where I rant about the shift from training?

Dylan Patel

For the record, we don't actually think that. For the record—

Jordan Chiao

No, but did you see the 2 charts that I found from back in 2018?

Dylan Patel

Yeah. A couple of our favorite forecasters, who are linear extrapolators instead of exponential extrapolators the way we are at SemiAnalysis, were projecting in 2018 that, in 2025, one day inference might be a $10 billion market.

Jordan Chiao

One, they're right.

Dylan Patel

Yeah. They had inference at double training at that point. Actually, edge inference was going to be $10 billion—I think $5 billion—and data center training was only going to be $4 billion. These were our peers at McKinsey, let's say, who were making these forecasts.

I'm not sure anybody in 2018 could have gotten the order of magnitude right and seen the explosion like this, but the ratios were just never making sense: that people would someday decide this is the final model we've ever trained. We're not going to invest in any more training; we're only going to buy GPUs for inference. This doesn't make sense.

The frontier labs believe they're building the machine god. All the evidence to me points to the fact that they are, and they're going to try to invest as much as they possibly can in training. Inference is a means to buy more GPUs to train the next model.

Jordan Schneider

Wow. Feels like an MMORPG. It's like you buy items to kill monsters to buy items to kill monsters, right?

Dylan Patel

Never-ending meat cycle.

Jordan Schneider

That probably describes life, right? We're all just killing time.

Dylan Patel

Okay, guys. Anything you think we haven't covered on this episode? I think we've gone off the rails on quite enough tangents on our tour.

Jordan Chiao

Yeah.

Dylan Patel

Tour of the Holy Trinity, right? I think we've subjected the listeners who are still hanging on to enough talk about GPU debt and financing rates. If you guys are interested in seeing anything fun, clever, or witty, we can do it right now. I give you a 3-count. 2, 3. Okay, that's a no from the Singapore team. We've had enough wit.

Jordan Schneider

We are not smart. No.

Dylan Patel

For today's episode—

Jordan Chiao

Our wit drained out. Sorry.

Dylan Patel

Yeah. Well, I appreciate everybody taking the time to listen in here, guys. Good job walking us through everything to do with the Trinity Capital offtake in data centers and what NVIDIA's doing with these GPU debt backstops. Hopefully it was fun.