[BidClub_]
SemiAnalysis · · 35 分钟

第035期—技术尽调、性能预测、供应链与投资论点(咨询)

Jordan NanosAbhilash Jain

创投/私募半导体技术企业经营
YouTube ↗
TL;DR
  • SemiAnalysis Consulting 通过持续3–4周至1年的定制化项目,打通AI基础设施的技术语言与金融语言。 Abhilash Jain 将其称为“技术与金融的共同设计”:机构投资者可能需要对一笔10亿–20亿美元的芯片、云或数据中心投资开展尽调,而超大规模云厂商、AI实验室和半导体公司则需要产品路线图、选址或市场进入策略。

  • 其中一个项目通过连接已部署芯片基数、推理与训练的算力分配、单GPU吞吐量以及每百万token的美元价格,建立了全球token供给模型。 最难的部分是模拟尚未发布的芯片和模型,包括潜在的10万亿、15万亿或20万亿参数系统;团队最终通过动态仪表盘暴露所有假设,让“每一个数字都会变化”,而不是把如此多的排列组合硬塞进静态Excel。

  • 当客户暴露出可以标准化的未满足需求时,定制化项目就可能转化为可规模化的研究产品。 数据中心模型和云端AI总拥有成本研究都起源于咨询业务;推理模拟器则从内部研究走向1个客户、再到多个客户,如今正成为独立产品——客户实际上是在识别哪些问题值得标准化。

  • NeoCloud尽调越来越决定陌生运营商能否拿到融资,并扛住合同下行风险。 投资者会审查SLA、OEM保修、集群性能、资本开支和运营开支;Jordan强调,预期可用性为99.9%,而低于90%则可能触发服务抵扣或终止合同。Abhilash的直截了当结论是:“没有可靠的SLA,你就拿不到融资。就这么简单。”(“Without reliable SLAs you won’t get funding. That’s it, period.”)

  • SLA风险会叠加:运营商可能用5年期GPU合同对冲15–20年期数据中心租约,而一次电力、制冷、机架或软件故障,仍可能触发客户要求的全部补救措施。 对运营高效、只需支付按基点计价保费的运营商而言,保险可以吸收罚金;但“如果你的性能表现很差,连保险都拿不到”。

  • 自建、购买还是合作,是由上市速度、算力溢价、技术能力和资本获取能力共同决定的条件性选择。 SemiAnalysis曾建议一名需要在6个月内获得100 MW算力的客户租赁容量;而对愿意等到2028–2029年的客户,给出的建议正好相反。

  • 咨询团队正在把重复性的尽调转化为运营杠杆,同时扩大技术覆盖范围。 NeoCloud检查已经变成“肌肉记忆”:团队知道要索取哪些材料,并在第2、4、6、8天安排访谈;Abhilash开玩笑说,结论往往早已可以预测。Dylan要求Abhilash在6个月内再招20人,这可能意味着“我每天50%的时间都要用来招聘”。

摘要 · 为研究而整理的核心内容

1. SemiAnalysis搭建资本与AI基础设施之间的桥梁

  • SemiAnalysis一直承接咨询业务,但在其newsletter、约13–14套机构模型和公开研究被证明对需要更深度参与的客户来说过于标准化后,组建了专门团队。每个项目都是定制化的,由专门团队负责,周期从3–4周到1年不等。

  • Abhilash将其概括为“技术与金融的共同设计”:教技术人员理解金融,让金融从业者掌握技术。私募股权和对冲基金客户需要对10亿–20亿美元的投资开展尽调,或寻找“下一轮大浪潮”;产业客户则需要产品路线图、数据中心选址和市场进入策略。

  • 服务对象覆盖超大规模云厂商、AI实验室、芯片设计公司和半导体公司。反复出现的是带有因果关系、且时间跨度很长的问题:如果LLM沿着某一方向演进,“3、4或5年后我们应该考虑哪些产品?”

2. Token供给项目变成了交互式预测引擎

  • 一家对冲基金要求SemiAnalysis根据已部署芯片基数估算全球token供给。公式串联了3个不确定变量:未来5年的加速器增长及训练与推理的分配比例、每个GPU的日吞吐量,以及由此得到的每百万token美元价格。

  • 吞吐量最难估算,因为未来的硬件和模型尚不存在。团队为一系列问题搭建了模拟器,例如NVIDIA规格变化、内存削减,或OpenAI推出10万亿、15万亿或20万亿参数模型,可能会如何影响性能。

  • 静态Excel无法容纳这些排列组合,因此交付成果变成了动态仪表盘。客户可以调整对AMD、TPU和NVIDIA的预期,修改输入输出比例、序列长度和上下文长度、数据类型,以及80%、85%或90%的缓存比例,然后观察“每一个数字”如何随之变化。

3. 咨询需求变成产品战略,最终变成产品

  • 一家内存公司在不同模型—芯片组合上运行模拟trace,以决定开发哪些内存产品,重点测试带宽与容量之间的取舍如何影响生产环境推理。这项工作直接把工作负载预测转化为工程路线图。

  • 一家半导体供应商面临约100个潜在买家,且每个客户都需要经历3–4个月的认证周期;SemiAnalysis将目标范围缩小到10–15个优先客户。团队依靠SemiAnalysis自身的网络,而不是专家机构,直接与20、30或40名市场参与者沟通技术要求、销售定位和客户旅程。

  • 重复出现的问题形成了研究飞轮。数据中心模型最初就是咨询项目;Jordan明确表示,云端AI总拥有成本模型也确实起源于咨询业务。推理模拟器则从内部研究走向1个客户、再到多个客户,如今正成为独立产品。

4. NeoCloud尽调最终考验的是合同生存能力

  • 一次典型尽调会把技术审查——SLA、OEM支持、保修、合同和集群基准测试——与资本开支、运营开支、利用率和生产力等商业指标对照起来。测试结果可以形成一份具体的180天改进计划,而不只是二元化的通过或不通过。

  • Jordan区分了服务交付条款与更广泛的合同终止权:延迟交付或停机可能导致服务抵扣或合同终止。客户可能要求99.9%的可用性;一旦低于90%,就意味着大约每10天有1天无法使用算力,但GPU费用仍要照付。

  • Abhilash认为,SLA最重要的恰恰是“坏时候”。运营商可能背负15–20年期数据中心租约,同时签有5年期GPU合同;一旦出现故障,就可能摧毁贷款方信心、触发罚金,并让运营商无法完成维持这些债务所需的第二笔交易或第三份租约。

  • Abhilash表示,有几家公司提供正常运行时间SLA保险,保费按基点计价;但性能表现差意味着连保险都拿不到。这类保险保护的是高效运营商,不能替代可靠性。

  • 因此,风险评估会从合同审查延伸到物理就绪度和裸机测试:L3、L4和L5级调试,机架到位,电力与制冷集成,故障模拟,监控和恢复。无论根因来自电力、制冷、机架还是软件,单个根因都可能让运营商承担全部SLA抵扣责任。

5. 产能紧迫程度决定自建还是购买

  • 托管机房、芯片和熟练劳动力仍然供不应求,因此拥有自己的集群并不总是现实的短期答案。SemiAnalysis会综合评估上市速度、客户为算力支付溢价的意愿、运营能力,以及客户能否独立为基础设施融资。

  • 对比非常明确:一家需要在6个月内获得100 MW算力的公司被建议租赁,因为它无法足够快地完成建设;而愿意等到2028–2029年的客户,得到的则是“截然相反的建议”。

  • Abhilash认为,收购项目在智力上最有意思,因为这让他接触到LLM架构、kernel、GPU优化,以及快推理与慢推理。不过,他最喜欢的工作是不断迭代的尽调,因为它已经变成了“肌肉记忆”:材料请求和访谈在第2、4、6、8天依次排定。Abhilash说结论并非预先确定时,Jordan打趣道:“你会惊讶于它们有多么经常是预先确定的。”

  • 团队现在覆盖的内容远多于6个月前,并计划再制作一期节目,讨论一路走来的经验教训。但眼下首先要解决的是扩张:Dylan要求Abhilash在6个月内招聘20人,让快速技术学习和招聘同时成为运营挑战。

完整逐字稿
Jordan Nanos

Hello, everyone. Welcome back to SemiAnalysis. My name is Jordan, and I'm here with Abhilash.

This week, we're going to do something a little different. We are not talking about an article. Instead, we want to tell everyone a little about the SemiAnalysis consulting business. Abhilash heads our consulting division, and he has a great team.

Many people have questions: What do we do in consulting? What clients do we serve? What projects do we carry out, and what have we learned? This covers everything from comprehensive due diligence for investors or companies considering M&A deals, to investors looking to go beyond our core products and customize one of our research products, such as the accelerator model or data center model, as well as strategic advice, feedback on product roadmaps, and go-to-market advice for many clients and industries.

So, Abhilash, welcome to the show. Nice to see you today.

Abhilash Jain

Thank you, Jordan. I am excited to finally debut in the renowned weekly magazine SemiAnalysis. Abhilash and his team have many thoughts that are voiced in this podcast, but this is the first time we'll hear them directly from him.

Jordan Nanos

I'm looking forward to it.

Abhilash Jain

Of course.

Jordan Nanos

Okay, give us a quick overview. What is SemiAnalysis' consulting business? What are we doing?

Abhilash Jain

I think, by the way, SemiAnalysis has been doing consulting since Dylan founded the company. But just last year, we created a dedicated team. The reason is that SemiAnalysis does a lot of things: We publish our newsletter, and we have institutional models. It seems like we have about 13–14 different institutional models right now.

Many of these are open research projects, like InfiniBand and ClusterMAX. So, there's a lot going on, but a lot of it is ready-made solutions. If any industry participant or investor wanted deeper engagement with SemiAnalysis, we didn't have an internal operating model to accommodate such requests.

This is the origin of SemiAnalysis consulting. I think we work with different clients based on an individual approach. All our work is purely individual. We usually allocate a special team for a period of 3–4 weeks to a whole year.

I'm ready to tell you more, but this is actually what we do.

1. Client Types

Jordan Nanos

All our work is individual. Exactly—individual work. Maybe let's start with what clients we work with. I gave a general overview of the different types of projects we have done, so tell us more about it.

Abhilash Jain

I think we're seeing what we always talk about: the co-design of hardware and software. Something different is happening in the AI infrastructure space, which I would call the co-design of technology and finance. We need to teach techies about finance and finance people about technology, and that's essentially what we do.

I think there are 2 groups of clients we work with. The first is financial institutional investors: large private equity funds and large hedge funds, which are very experienced investors and very interested in investing in AI infrastructure. When they decide to invest $1 billion or $2 billion in a chip manufacturer, a cloud platform, or a data center, they want SemiAnalysis to help them understand the details of the deal.

Therefore, we conduct various types of checks for them. We also work a lot on developing investment theses for hedge funds. Many hedge funds are looking for the next big wave. People were fascinated by memory, and they were fascinated by processors, so they're trying to figure out what's next.

We also work with many hedge funds, combining different data sources and market narratives to determine what the next wave will be. On the other hand, we also serve many industry customers, mainly hyperscalers, AI labs, chip manufacturers, and semiconductor companies.

The range of topics can be very broad. This could mean helping a semiconductor company develop its product and engineering roadmap. A typical question is: If LLMs are evolving in a certain way, what products should we be thinking about in 3, 4, or 5 years?

It could also mean helping a certain hyperscaler choose sites for data centers in remote locations, such as India. It can also mean developing a go-to-market strategy for different companies that have a strong product but don't know how to connect with and understand their customer journey.

Those are a whole range of different options, but these are the clients we serve as part of SemiAnalysis consulting.

2. Custom Projects

Jordan Nanos

Sounds logical. And a short commercial for Abhilash on his behalf. If you're interested, if you listen and are passionate about technology, finance, code design, you should try applying for a consulting job at SemiAnalysis. This is very interesting. We're hiring a lot right now. So, shameless advertising: if you are an AGI developer, passionate about semiconductors, eager to learn, please contact us. We desperately need people. So, tell me about some examples of that kind of customization.

Abhilash Jain

I get involved in this work because we take the research that I and the other people on the team conduct, try to present it in a more understandable form, and adapt it to the specific interests of the client.

Jordan Nanos

Tell us what that process looks like as we build on our existing data sources to actually implement these projects.

Abhilash Jain

Because there's so much going on at SemiAnalysis, usually, if you can just combine some of these data products in a more accessible way, we can answer questions that each of these products alone can't answer.

Here's one example: We were recently approached by a hedge fund client who really wanted to understand what the total token supply in the world was, given the installed base of chips. If you think about it, it's a pretty simple equation. What is the total installed base in the world today, and how should it grow over the next 5 years? How much of this power is allocated to training, and how much to inference?

If you have an installed base, then the second part of the equation is: What is the throughput of each GPU, and how many tokens can each GPU generate on a given day of the year? Finally, what is the price in dollars per million tokens for these tokens?

Each of these 3 different parts of the equation was solved differently, but the most challenging was estimating throughput per GPU, as we had to combine a lot of our output data that is already public with data for chips and models that are not yet out.

We had to create a rendering simulator where we evaluated what it would mean for throughput if NVIDIA released devices with a certain set of specifications. If memory shrinks, as has already happened, what impact will this have on throughput? If OpenAI trains a model with 10, 15, or 20 trillion parameters, what will that mean for performance?

A lot of this is not publicly available, so we had to develop it individually, collaborating with various inference engineers at SemiAnalysis to implement it. This is an example of a project that requires custom work.

Jordan Nanos

Good. Tell me about some of these assumptions in the output simulator. Maybe explain how this is communicated to clients, because I think a lot of people hear “consulting” and think of presentation slides.

We do a lot of presentations, but for such complex issues, the trend is to move to interactive dashboards where people can change their assumptions over time, and we have to support that. So, it's like a software product that we provide to them.

Abhilash Jain

Yes, that's right. There were so many permutations and combinations in this product that it was impossible for us to fit it into a static Excel spreadsheet. We had to create a dynamic dashboard, and that was essentially what we delivered to the clients.

This was a panel that they could use to play around with different sets of assumptions in a very visual way and see how literally every number changes. The key assumptions were, I think, to start with: How much compute power do you allocate to inference, and how much to training?

What are the existing chip specifications we know, and what features do you expect for the next generation of silicon being developed by AMD, TPU, and NVIDIA? We know from one episode to the next, or even 2 or 3 issues later, that these specifications can change quite quickly.

The Rubin Ultra specifications have just been downgraded from 1 terabyte of HBM to 56, 128, or 192.

Jordan Nanos

So, yeah. What other assumptions can be made, for example, about your input-output ratio? Do you use Agent X test kits? Are you using fixed sequence length, context length, or specific data types? What is your caching ratio?

How does throughput change if you hit 80% caching compared to 85% or 90%? I think these are some of the assumptions that make this interesting.

3. Strategy Work

Now tell us more about an example from another category, such as strategy. This may apply to the product roadmap, or to a business decision: create it yourself, buy it, or collaborate. In what areas are you currently helping people?

Abhilash Jain

I think strategy is a very broad concept. This could mean product strategy or go-to-market strategy, and I can give examples of both.

As I said, we developed a pinout simulator internally, and one of the memory companies is using it to reproduce traces for different model and chip combinations to determine which memory products they should develop next. They answer questions such as how trade-offs between memory bandwidth and capacity affect serving inference workloads in production.

This is an example of a product strategy that we work on with their teams on a very specialized basis. Another example is that we are helping a semiconductor company identify its ideal customer set.

They have a product, but they don't know who to sell it to. There are probably about 100 customers who could buy this product at any given time, but it is a product with a long supply cycle. It takes at least 3–4 months to qualify a client.

So it’s impossible to really qualify your product with 100 different customers, right? We help them identify the top 10–15 clients they should focus on. How should they best approach sales negotiations? What key technical topics are of interest to customers, and how should you position yourself to meet the technical requirements they expect?

It seems easy to say, but it’s very difficult to execute because we need to go out and literally talk to 20, 30, 40 different customers, get their views, and gather market vision. That’s something many other companies can’t do because SemiAnalysis has very close ties. We don’t actually use expert agencies to conduct these expert calls, or whatever you call them. Because we have such a broad network, we can simply use our contacts to figure out important issues that will help the broader ecosystem.

4. Consulting to Products

Jordan Nanos

Yes. This also shapes our research. Maybe you could explain a little bit about how people who have been on this podcast before, like Eric a few weeks ago, working on modular data centers, contribute to the analysis products when they do this research?

Abhilash Jain

In many ways, some of the institutional products that we sell—the data products—originally started as a consulting project, or at least the inspiration to start working on them came from someone asking us about them. They clearly think it’s important, so let’s do it.

Yeah, because we work on a very individual basis with a lot of these clients, and we dive very deeply into a topic. We can often identify what exists. If a question or problem is relevant to one client and we’re able to find a standardized solution for it, we can scale that to a broader range of clients, and it becomes a product.

I think the data center model started as a consulting project for one of the clients. ClusterMAX may have started as a consulting project. I don’t know; you have to tell me that.

Jordan Nanos

No. Cost Max, but the total cost of ownership model for cloud AI definitely started with consulting projects. We continue to use those projects to build analysis: What are the upfront costs of GPU clusters? How profitable are certain token-serving companies depending on the choice of chips, the amount of operating expenses, and the utilization rate at the output endpoint?

The way we actually dive into it and do it is like this: You get inspired to work because people you trust—your customers—take a problem seriously and ask questions. You’re like, “Let’s find the answers,” and it becomes a useful product.

So, anyway, that’s right. I think the latest example of that is the inference simulator that we’re developing right now, right? I think we started creating it for one client, and not even exactly one. I think it started as an interesting research project within SemiAnalysis. Then we offered it to one client, they appreciated its benefits, a few more clients saw the value in it, and now it’s turning into a standalone product.

5. Neocloud Due Diligence

Yes, definitely. It sounds completely logical. What about due diligence, or DD?

Abhilash Jain

This is obviously an important thing for both investors and companies interested in mergers and acquisitions. Dan is talking about $7 trillion in debt being issued to fund AI infrastructure, right? A lot of investors are entering the market, as far as I’m concerned. There are also many new players emerging.

A due diligence project looks something like this: An investor comes to us and says, “Hey, can you help us with, say, NeoCloud?” We often do this. If a private equity fund wants to invest $1 billion in NeoCloud, they’re essentially looking for 2 things: technical analysis and commercial analysis.

On the technical side, they’re interested in what SLAs NeoCloud offers its customers, and whether NeoCloud will be able to fulfill the guarantees under those SLAs. If not, it has broader implications that we’ll talk about later, but that doesn’t benefit anyone, right?

They’re also interested in understanding how OEM support, warranty structures, and contracts work, because there are many nuances. If you don’t deal with OEMs properly, you can run into various risks.

They’re also interested in cluster performance. We have the opportunity to test their clusters using a methodology inspired by ClusterMAX, and we can tell exactly what they’re strong at, what they’re really bad at, and what the 180-day plan for their improvement should be.

From a commercial perspective, because it’s such a financially significant decision, they need to understand the capital and operating costs of these builds. How does this realistically compare to the market standard? As part of our TCU model, we can compare costs in real time, along with various other metrics, such as productivity, for our customers.

That’s what a typical DD project looks like. It varies depending on whether you’re auditing a neocloud, a chip company, or a model company. We’ve worked with data centers and power plants—we’ve literally done everything.

I think we’ve done a lot in just the year or so since you’ve been here. We’re gaining experience in this area and building a team. It’s quite interesting to see it grow as more and more people in the industry take it seriously, because the need for technical validation of many of these projects is extremely important.

Jordan Nanos

This is all new, isn’t it? It’s not something that certain investors can understand or have the ability to understand—the type of model that’s 6 months old or 3 months old. I mean, it’s new to everyone.

Abhilash Jain

That’s right. This is new for everyone. Also, I think what’s happening is that the customer base is diversifying beyond the mainstream labs and now includes neolabs that are actively operating in the market.

The neocloud operator base is also diversifying, if you think about it, right? Hyperscalers are the old guard, then there are cryptominers, and then specialized players like Core and others. But now you see a lot of real estate developers who have no experience whatsoever in operating cloud systems entering the arena. You see the family offices of wealthy people now trying to build a neocloud.

That’s why service-level agreements, or SLAs, become even more important. If you don’t have any prior experience operating these things, and the SLA is not in your favor, the consequences can be really huge.

6. SLAs and Insurance

Jordan Nanos

Yes. Let’s talk about 2 types of them. Obviously, this is something I’m quite familiar with. There’s one type of SLA that deals with service delivery terms. In fact, an SLA is simply that if something happens, the contract can be terminated. That’s roughly the legal gist behind a service-level agreement.

It’s not the right to terminate the contract—that’s probably a broader definition—and that’s really what we’re talking about here, as opposed to an SLA. The first reason you can terminate a contract is a delay in delivery. The cluster arrives months late, so you cancel the contract.

The most common reason is downtime. If someone has GPUs installed, they expect 99.9% availability. If you’re below 90%, one out of 10 days a month, or 3 out of 30, you don’t have access to the GPUs you’re paying for. You usually get some credits in return or have the right to terminate the contract.

So maybe tell me about the motivation for lenders, for a neocloud, or even for the buyer to have some insurance against any of these contract-termination possibilities. They want protection against the risk of a decline in value.

Abhilash Jain

True. That’s right. Jordan, you hit the bullseye. Why are SLAs important? Then I’ll tell you why insurance is important.

First, as a neocloud, without reliable SLAs, you won’t get funding. That’s it—period. You simply won’t be able to start your journey. All investors and lenders care most about is the level of their risk exposure. If they receive interest payments monthly or quarterly, what is the probability that the payment will not arrive in their bank account because your cluster is down?

Second, SLAs are actually needed for bad times. We’ve seen this in different situations. If things are going well, no one mentions the SLA. If things aren’t going well, people remember the SLA, and it gives the parties the opportunity to get out of the contract without losses.

Third, as we said, at least for neoclouds, they usually sign a data center lease agreement for 15–20 years. They usually sign GPU contracts for 5 years, and if you don’t follow through or meet your SLAs, you’re taking a reputational risk. That means you can’t get a second deal and you can’t get a third lease, right? If you can’t do that, how are you going to pay for 20 years of data center lease?

So SLAs are important. I think the reason they’re important for lenders, as I said, is because they’re looking for protection against risk. One way to protect yourself is through SLA insurance.

We know of several companies that currently offer uptime SLA insurance for neoclouds. If you pay a certain number of basis points in premium, you get some protection. If you have to pay a penalty, the insurance company will take it on.

In my opinion, it’s not easy to get such insurance, as these insurance companies are quite good at underwriting projects like neoclouds. This isn’t an excuse for poor performance—you can’t continue to perform poorly and then just compensate for it with insurance.

If you have poor performance, you won't even get insurance. But if you are working efficiently, I believe that one way to protect yourself from risk and make your contract more attractive for investment is to consider choosing one of these types of insurance.

7. Risk Assessments

Jordan Nanos

Yes. Now, we've often mentioned the term SLA and talked about a few different things. Obviously, there are many different components of a given cluster, data center, or site to which an SLA can be applied. This could apply to the GPU or the nodes themselves. It may apply to racks, clusters, or the site itself. It can also apply to the power and cooling systems within them.

With so many parties involved and so many contracts changing hands, there are many reasons why you might want to do a risk assessment for 1 individual part that you're less confident in versus another part of the contract that you're more confident in, perhaps because you've already worked with that vendor or understand the technology better. Tell me about the concept of risk assessment and its scope, because while I'm pushing you to answer, the scope can get pretty narrow pretty quickly, or it can be pretty big and broad, right?

Abhilash Jain

Yes, that's right. I think, as you said, risk assessment can be quite critical, especially in cases where lenders or investors require it and it's a prerequisite for them to finance a certain kind of cluster. I think this could be a fairly easy process, such as reviewing the most important contracts for a neocloud. These are contracts with the service buyer, OEM contracts, and data center contracts.

Since you're promising something to the customer, you're relying on the data center partner and the OEM to deliver on it. There are also other related contracts that we could delve into, but that's not necessary right now. That's 1 thing.

The 2nd is a somewhat more thorough review or risk assessment, as you might call it—namely, the readiness of the data center. That means understanding the stages: when will the data center pass L3, L4, and L5 testing? When will the data center be available for you to run a number of other tests? When will the racks arrive at the data center? When will you be able to install the power and cooling systems? And when can all these elements be tested simultaneously as a system?

The 3rd is bare-metal testing, which I think you, Jordan, are a master at, given all the cluster tests you've done. So I would like you to do that, and then we'll come back to the last point.

Jordan Nanos

Yes. The bottom line is that we access the systems, log in, test, simulate failures, and check whether people have a monitoring system set up so they can detect that failures have occurred and how quickly they can recover. There are many different ways things can go wrong in a cluster, and your monitoring systems need to have full coverage, because we've seen things go wrong in spectacular ways, even in our few weeks of testing in a fully simulated environment.

People go to great lengths to set this up through cluster testing, and then we hear stories from all the customers of these providers about how poor the quality of service they receive is.

Abhilash Jain

Yes. The root cause can be really quite diverse, right? I believe that this assessment helps to understand what exactly is the root cause of low productivity. For example, is it a problem with how your power supply is set up? Is it because of how your cooling is set up? Is it because of problems in the racks themselves? Is it because of software problems?

Ultimately, even if there's only 1 reason why your cluster is performing poorly, that still doesn't mean you won't have to pay penalties for violating your SLA. Just 1 mistake is enough for you to be obligated to reimburse the client for all SLA credits.

Jordan Nanos

Yes. That's why we're going back to the contract review stage, which is the 1st one, right? People need a lot of advice on how to make these contracts, because as you said, contracts are needed when times are bad, not when times are good.

8. Build vs Buy

Let me ask a related question so we can continue our conversation. When we talk about providing strategic advice, a lot of people come into this space looking at the “build or buy” option, especially with neocloud, because they have such long lead times. Finding a site for a data center is very difficult, and we did some consulting work helping people make decisions: build, buy, or partner.

We didn't just do this for neocloud solutions, but also for other issues, like supply chain management if you're a chipmaker or something like that. Maybe you can talk a little bit about those examples as well.

Abhilash Jain

Yes, I think so, because we definitely see a lot of constraints on getting colocation capacity to deploy large GPU cluster sites. We also see some chip shortages, and of course there's a labor shortage. So, in principle, even though many companies want to eventually build their own clusters, this often isn't the most realistic way to obtain capacity in the near term.

Ultimately, it all comes down to a few factors. I think the 1st is time to market; the 2nd is your willingness to pay a premium for computing power; the 3rd is how technically savvy your team is and whether you have experience managing clusters. Finally, regarding financing, do you have sufficient access to the capital markets to raise funds to create a cluster from scratch or become a buyer of services from a company?

Jordan Nanos

We recently worked with several clients for whom we found several different answers. I think in 1 case recently, you advised a company to just go and lease because they needed 100 megawatts literally in the next 6 months, and there was no way they could build a cluster that fast.

Abhilash Jain

There were cases when someone was willing to wait until 2028–2029, and the recommendation was diametrically opposite.

Jordan Nanos

These are usually projects lasting 3–4 weeks. But it's definitely very exciting because of the amount of content we can cover and the options we can explore with them.

9. Favorites and Outlook

Okay, personal question. Which projects from the whole set we've talked about do you find most interesting right now? And what is your favorite? Is that the same thing, or are those different projects?

So, you asked about your favorite and what else? The most interesting. You just said that this is an interesting project. I guess that means you're learning and doing interesting things.

Abhilash Jain

I would say that one of the most interesting projects we're working on right now is when a large company seeks to acquire a smaller one. This company doesn't have a large team, but it has very smart people working for it. I get paid to talk to a lot of these experts, and I learn a lot in the process: about deep technical projects, LLM architectures, kernel development, GPU optimization, fast and slow inference, high-throughput systems, and so on. Any of these topics sound very interesting to me.

I think my favorite projects are the ones where we can use an iterative process to get an answer, because we've done so many projects over the last year testing new cloud solutions that it's become muscle memory. We know exactly which request for information to send on day 2, which interview to conduct on day 4, which 5 additional questions we always expect on day 6, and which interim report to produce on day 8.

Finally, what should be the conclusion? So, yes, it's because the conclusions are not predetermined.

Jordan Nanos

Dude, you can't say that on a podcast. How often are they like that?

Abhilash Jain

You'd be surprised how often they are.

Jordan Nanos

Regarding the neocloud checks, do you think it's easy enough to analyze some of these contracts and get clear results right away?

Abhilash Jain

Yes. This has become a cliché, but it makes sense. It's logical that people want to be sure that we've reviewed a lot of materials.

Jordan Nanos

Good. I'll ask you a more general question. What excites you most about your future consulting work?

Abhilash Jain

I would say that we're definitely an organization that's growing very quickly. I think the range of topics we cover today is much broader than what we covered 6 months ago. What I'm most excited about is partnering with people like that. Literally, we're already working with people who are at the forefront of AI development, but getting new opportunities to partner with people like that is definitely the most exciting thing, because you get to learn so much in such a short amount of time. It's just crazy how quickly our knowledge and learning curve is growing here at SemiAnalysis.

The other thing I'll say is that Dylan clearly told me to hire 20 more people within the next 6 months. I don't know how I'm going to do it, but it's going to be an exciting journey too: interviewing a bunch of people and spending 50% of my day hiring.

Jordan Nanos

I know. It sounds like your current day is quite interesting, because if you're happy about it, it's just a continuation of the same.

Abhilash Jain

It's a continuation of the same, dude. The same—yes, steeply.

Jordan Nanos

Good. What comes to mind when I ask what we haven't covered in this conversation yet?

Abhilash Jain

I think we've discussed almost everything I could say publicly, to be honest.

Jordan Nanos

Good. Because there are many, many conversations about things that cannot be said publicly. Abhilash—

Abhilash Jain

No, no, no. There are many things I can't talk about publicly, but I can't voice them either.

Jordan Nanos

Yes. Oh, man. Well, I appreciate you joining us today. It was very interesting. I enjoy working with you on many of these things. Hopefully, next time we can get more people from the team involved. We'll be able to update information on the trends we see as we complete more of these projects and continue to learn.

Abhilash Jain

Yes, I think today was more about what consulting is. We should definitely do an episode where we talk more about what we've learned, because as a team we have a lot of conclusions, and the team would be happy to come and share everything we've learned so far.

Jordan Nanos

Yes. I think the audience would be happy about that too. But let's leave this as a little teaser—a teaser behind the paywall of a podcast that will be released soon. Fifteen minutes of a podcast behind a paywall. How do you like that? Everyone loves newsletters because they're free, but there's a paywall. So how do we put this video behind a paywall?

Abhilash Jain

Let's do this, dude. Let's do this. Let the dollars flow. Maybe we should also put a watermark right on the faces. I think that would be even better.

Jordan Nanos

Yes. Yes. Yes. Yes. Type E translation. Yes. Classic. Good job today. Thank you to everyone who listened. See you next time. Thank you, guys.