Nvidia CTO Michael Kagan:突破摩尔定律,迈向百万 GPU 集群
- Kagan 的核心判断是,AI 已将扩展能力从芯片级摩尔定律推向系统级扩展:模型规模和能力每3个月翻倍,要求年化性能提升约10X–16X。 因此,NVIDIA 优化的是“单一计算单元”——算力、网络、软件以及数据中心级整机——而不只是晶体管密度。Kagan 称,如今每年的产品迭代在机器层面带来约一个数量级的提升。
- Mellanox 让 NVIDIA 突破单节点扩展:它延伸了纵向扩展互联,并提供横向扩展所需的高性能网络。 如今的“GPU”可能已经是一台机架级机器——36台双 GPU 计算机中的72块 GPU,通过 CUDA 呈现为一个可编程系统——“得用叉车才能把它吊起来”。集群性能取决于延迟是否受到严密控制,而不只是带宽峰值:延迟抖动幅度过大,就会迫使每块 GPU 等待。
- 当 GPU 数量达到100,000时,零部件可靠并不意味着系统可靠,因为数百万个部件会让“所有部件都正常工作”的概率归零。 硬件和软件必须假设始终有部件损坏,同时维持服务、利用率和能源效率。Kagan 更广泛的答案还包括 BlueField DPU:将基础设施计算与应用分离,并显著缩小攻击面。
- Kagan 认为,随着 AI 从一次性感知走向生成和多步骤推理,推理算力需求已经“不低于训练”,实际上还要更高。 训练只发生一次,而每位用户、每个 token 都会反复触发推理。Prefill 是计算密集型,decode 是内存密集型;NVIDIA 的押注是,专用 SKU 仍应沿用统一的 CUDA 编程模型,从而让算力容量可以在两者之间调配。
- 对数据中心规模的现实约束,与其说是抽象的网络定律,不如说是能源和散热。 Kagan 提到约100–150 MW的大型部署,业界讨论正转向 GW 乃至10 GW级设施,液冷则是提高密度的必要条件。将一个工作负载拆分到相距遥远的站点,会引入光速延迟和拥塞;Spectrum-X 通过端点遥测应对这一问题,而不是使用会制造抖动的“超大缓冲区”。
- Intel 合作体现了 NVIDIA 的判断:加速计算将扩张,而不是消灭通用计算。 Kagan 将其描述为一种“共赢文化”:NVIDIA 的目标是“把蛋糕做大,让所有人都受益”,x86 与加速计算将为双方打开渠道和市场。Mellanox 的先例尤其显眼——主持人将其概括为支付70亿美元、6年价值增长45X、员工留存率约85%–90%。
- Kagan 将当前性能斜率定在“每年10X或几个数量级”,但也明确承认:“能持续多久,我不知道。” 他的长期上行空间不止于生产率:Earth-2 可以让历史成为“实验科学”,AI 或许能够推断出人类从未设想的物理定律,而一种充当“心灵飞船”的技术,可能让人们想要完成的工作增长得比可用资源更快。
1. Mellanox 将算力单元从芯片推向数据中心
Kagan 的起点是两种指数增长的错位:摩尔定律大约每隔2年翻倍,而 AI 模型的规模和能力开始每3个月翻倍。这意味着性能每年必须增长10X–16X,优化对象因此不再只是基础部件。
纵向扩展如今意味着使用 NVLink,把36台各搭载2块 GPU 的计算机中的72块 GPU,在同一软件接口下变成一个系统。客户口中的 GPU 因而不再是芯片,而是“一台机架大小的机器”:“得用叉车才能把它吊起来。”
横向扩展则将应用拆分到多台这样的机器上。Kagan 举例说,把一个1秒的任务拆到1,000块 GPU 上,可能在1毫秒内完成——但前提是任务分发和结果汇总不能暴露通信耗时。
Mellanox 在网络侧的贡献是延迟稳定性,而不只是峰值带宽这种“漂亮数字”。如果计算无法掩盖网络抖动,每个 worker 都要等待;一个原本可扩展到1,000块 GPU 的作业,实际可能只能高效利用10块。“从根本上说,网络决定这个集群的性能。”
2. 达到100,000块 GPU 后,故障、距离和散热成为架构变量
某个部件即便有99.999-whatever%的时间正常工作,但一台包含100,000块 GPU 的机器拥有数百万个部件;Kagan 直截了当地总结:“所有部件同时正常工作的概率为零。” 硬件和软件必须在确定有部件损坏的情况下,维持服务、性能和能源效率。
普通数据中心网络连接的是松耦合微服务;AI 算力网络可能在“100,000台机器上运行一个单一应用”。因此,调度器需要硬件和底层软件提供接口,才能高效安排任务的每个部分。
跨数据中心会增加光速传播时延,并带来截然不同的延迟分布。Kagan 不接受老式电信网络的缓冲逻辑——“更大并不更好”,因为缓冲区会制造抖动;他称 Spectrum-X 提供遥测,让端点区分短通信和长通信,并围绕拥塞进行调整。
Pat 追问:如果核电站提供充足能源,数据中心本身是否还能继续扩张?Kagan 诚实但不完整地回答:“我不知道”,因为输入电力最终仍会变成输出热量。NVIDIA 已基本转向液冷;当前100–150 MW规模的设施,与业界讨论中的 GW级及拟议中的10 GW级站点并存,但部署速度仍可能取决于“混凝土多久才能稳定”。
3. 推理正成为规模更大、形态更多元的算力市场
训练由前向推理、反向传播组成;在数据并行下,还要跨模型副本汇总权重更新。AI 主要还停留在感知阶段时——识别一只狗或认出一个人——训练需求占主导;但生成式 AI 会针对每个 token 重复执行推理,而推理过程还可能比较多条候选路径。
推理本身又分为两种截然不同的工作负载。Prefill 将提示词和背景材料编码进上下文,属于计算密集型;decode 逐 token 生成答案,属于内存密集型,即便有些技术已经可以一次生成不止一个 token。
Kagan 的结论是,推理需求“实际上比训练还大”,因为每次响应所需的工作量都在增加,而且“模型只训练一次,但推理要进行很多次”。推理可以在手机或更小型的部署上运行;数据中心则可以混用为 Prefill 和 decode 优化的 GPU SKU,统一的 CUDA 可编程性让任一 SKU 都能承接另一种工作负载的需求变化。
4. NVIDIA 的下一轮扩张,将加速计算与通用计算结合
Kagan 对加速的论证始于显式编程的局限:冯·诺依曼机器可以执行指令,但“我无法解释如何区分猫和狗”。AI 解决的是另一类问题,但通用计算和占主导地位的 x86 工作负载并不会消失。
因此,Intel 合作是将“加速计算与通用计算融合”。Kagan 对商业逻辑的概括是生态扩张:NVIDIA 不是要从既有市场中切走更大份额,而是“为所有人把蛋糕做大”,打开两家公司单独服务时更难触达的客户和渠道。
Pat 指出,2019年 NVIDIA 与 Mellanox 合计约值1,000亿美元;6年后,这一数字约为4.5万亿美元。Kagan 曾对 Jensen 说,“1+1会等于10”;他承认,自己的估算“差了4倍”。
组织和财务数据都显示了整合效果:Kagan 估计,Mellanox 原有员工的85%–90%留了下来,NVIDIA 在以色列的员工人数翻了一倍以上,并计划建设新园区。BlueField DPU 还通过将数据中心操作系统和基础设施工作负载从应用处理器上剥离出来,扩大了 Mellanox 的作用,在最大化应用算力的同时缩小攻击面。
5. Kagan 认为,指数增长正将 AI 变成科学基础设施
Kagan 的科幻式目标是“让历史成为实验科学”。Earth-2 可以模拟今天的行动将如何影响50年后的全球变暖;更广泛地说,AI 的观察和归纳能力可能帮助揭示“我们今天甚至还想象不到”的物理定律。
被问到摩尔定律的继任者时,Kagan 给出的答案是“每年10X或几个数量级”。2、3年前,NVIDIA 已将加速产品的发布节奏从每隔一年一次提速至每年一次,让客户可构建的机器获得数量级提升——而不仅仅是单颗芯片性能提升;但他明确对持续时间留出余地:“我不知道。”
他的乐观类比从 Steve Jobs 将电脑称为“心灵的自行车”(bicycle of the mind),走向将 AI 称为“心灵的宇宙飞船”(spaceship of the mind)。效率提升并不会让需求止步于原有水平:给项目负责人2X的资源,他可能做出4X的工作量,同时还想要10X的资源。就像电力一样,AI 可能成为不可或缺的基础设施,而其应用无法从平台诞生之初就被预测。
One of the interesting things about NVIDIA is the culture of win-win. We are not after taking a bigger piece of the existing pie. We are after baking a bigger pie for everybody. Our success is our customer's success. Our success is not the failure of our competition. And I think fusing together conventional computing—von Neumann machines—and accelerated computing provided by NVIDIA actually gives NVIDIA and Intel channels to the market. We're expanding the market and serving the markets that otherwise were more challenging.
Yeah.
We're delighted to hear today from one of the legends of the semiconductor industry, Michael Kagan, the CTO of NVIDIA. Michael was formerly chief architect at Intel, and then co-founder and CTO of Mellanox, which NVIDIA acquired for $7 billion in March 2019.
In the time since, Michael has been a major driver of NVIDIA's dominance as the AI compute platform, in large part due to the role of Mellanox and interconnect in driving chips beyond Moore's law. The AI race is ultimately a silicon race to squeeze the most intelligence possible out of each unit of silicon, and Michael takes us on a journey through how the compute frontier has evolved—from squeezing more transistors onto a single chip to bringing together thousands and hundreds of thousands of chips into a single fabric connected by networking in an AI data center.
Michael has been driving the compute frontier forward for more than 4 decades, and we're honored to have him on today's show.
Okay. We're here with Michael Kagan, the CTO of NVIDIA, currently the world's most valuable company. Michael, thank you for joining us.
Thank you. My pleasure.
I thought we could start with this: Our partner Sean likes to make the case about every 6 months that NVIDIA would not be NVIDIA without Mellanox. Mellanox is a company that you co-founded some 25 years ago and have been a part of through this day. Can you paint that picture for us? Why was the Mellanox acquisition so critical to NVIDIA?
There is a huge transition in the world in terms of computing and the need for computing. It grows exponentially, while one of the things that we usually estimate linearly is actually exponential in the world. And exponential growth is now actually accelerating.
It used to be like Moore's law, which was based on basic silicon: 2× every other year. Regardless of the discussion that Moore's law, in terms of physics, is not quite running anymore, once AI kicked in—which was in 2010 or 2011—it kicked in when a GPU, from a graphics processing unit, became a general-purpose computing unit.
Yeah.
That was when workloads were run—the first time an AI workload was run on the GPU—taking advantage of the programmability and parallel nature of this machine. The requirements for performance started to grow at a much higher rate. The models started to grow in terms of size and capacity, 2× every 3 months, which now requires 10× or 16× a year in performance growth versus the old school of 2× every other year.
In order to grow at this scale, you need to innovate and develop solutions at a much higher scale than just a basic component. That's where networking comes in.
Mm-hmm.
There are multiple layers of scaling performance that require high-speed and high-performance networks. One is what we call scale-up.
Mm-hmm.
Basically, going back to the CPU days, scaling up was Moore's law—more transistors—and also advances in the microarchitecture, like out-of-order execution and, at some point, multicore, and so on and so forth. This is the basic building block of computing. In the GPU world, the basic building block is the GPU.
Mm-hmm.
In order to scale it up more than you can on a single piece of silicon, with all the advances that we are making in microarchitecture and advanced technologies, you actually need to do something on the scale of a multicore CPU, but at a much larger scale. That's what we are doing with NVLink. That's the scale-up solution.
Our GPU—what we call a GPU today—is a rack-size machine.
Yeah.
You need a forklift to lift it.
Yeah.
If you order just a GPU on Amazon, don't be surprised if this huge rack shows up.
Yeah, people think chip, but it's really a system.
Right. Right. And that's just one GPU.
Yeah.
So, the basic building block—a very basic computer on which application software is running—is this GPU. And it is not just silicon, not just hardware, and not just wires; there is also a software layer that exposes CUDA as the API.
That's what enables you to pretty much seamlessly scale. I'm simplifying the story a little bit, but you can seamlessly scale from one component, which used to be a single GPU, all the way up to 72, maintaining the same software interface.
Once you get this building block as big as it conceivably can be built—in terms of power, cost, and efficiency—then you start scaling out.
Yeah.
Scale-out means you take many of these building blocks, connect them together, and now, at the algorithm level and the application level, you actually split your application into multiple pieces running in parallel on these big machines.
Mm-hmm.
And that's again where networking comes in. If you talk about scale-up, we basically made a memory-like domain to go beyond a single compute node or a single GPU.
Mm-hmm.
That's actually the first place where Mellanox technology comes in. Before the Mellanox acquisition, the scaling-up of NVIDIA with NVLink was limited to a single-node machine.
Going outside of a single compute node—those 72 GPUs are actually 36 computers, each with 2 GPUs wired together to present all of this as a single GPU—is not just plugging a wire into the connector. There is a lot of software and a lot of technology within the network involved in making multiple nodes work as a single machine.
That's where Mellanox first came in. In terms of the way we go up in scale, that's the first one. The second one is: How do you split the operation across multiple machines?
Yeah.
The way to do it is this: If I have a task that takes 1 GPU 1 second to do, and I want to accelerate it, I split it into 10 or 1,000 pieces and send each piece to a different GPU. Now, in 1 millisecond, I get done whatever I was doing in a second.
But you need to communicate this partial job split. You split the task, and then you need to consolidate the results. Every time you run this, you have multiple iterations or multiple applications running at once. There is a part of doing communication and a part of doing computation.
Mm-hmm.
The thing is that you want to split it into as many pieces as you possibly can because that's your speedup factor. But if your communication is blocking you, you waste time, you waste energy—you waste everything.
What you need is very fast communication. You split it into many, many pieces, so each piece takes very little time, but then there is another piece that is communicated, and you need to feed it in time. That's just pure bandwidth.
Another thing is that when you tune your application, you tune it so that communication can be hidden behind computation. That means if communication, for some reason, gets longer, then everybody waits.
Hmm.
It means that what you need in the network is not only raw performance—what are called hero numbers: “I can get to that many gigabits per second.” You also need to make sure that, no matter who communicates with whom, the latency—the time it takes—is distributed very narrowly.
Mm-hmm.
If you look at other network technologies or other network products, you go to the hero numbers: sending a bit from one place to another. That's basically physics.
Yeah.
It's pretty much close to the same for everyone. We are a little bit better, but that's not the big advantage.
When you do it thousands of times, and it takes the same amount of time to do it every time versus having a very wide distribution with other technologies, the machine becomes less efficient. Instead of being able to split your job across 1,000 GPUs, you can split it across only 10 GPUs because you need to accommodate for the jitter on the network within the computation phase.
Inherently, the network determines the performance of this cluster. We look at this data center as basically a single unit of computing.
Yeah.
A single unit of computing means that you start architecting your components, your software, and your hardware at the point where this is a data center. This is 100,000 GPUs that we want to make work together. We need multiple chips: 2 compute chips and 5 network chips.
This gives you a sense of the scale, the impact, and the investment you need to make to create this single unit of computing. That's where Mellanox technology came in.
Another aspect of this is that we talked about the network that connects the GPUs to run the task. But there is another side of this machine that is customer-facing. This machine needs to serve multiple tenants, and it needs to run an operating system. Every computer runs an operating system.
Another part of Mellanox technology is what we call the BlueField DPU, or data processing unit, which is actually the computing platform used to run the operating system of the data center.
Mm.
In a conventional computer, we have a CPU that runs the operating system and application software. There are many things we can talk about in terms of the advantages and disadvantages, but there are 2 key things. One is how much time you spend on general-purpose computing to run the application. You want to maximize it.
Mm.
Another thing is how you isolate your infrastructure computing from your application computing.
Mm.
Because of viruses, cyberattacks, and so on, being able to run infrastructure computing on a different computing platform significantly reduces the attack surface, especially for side-channel attacks.
Mm.
This is different from what happens if you run it on the same computer. If you remember, about 5 or 6—well, actually, almost 10—years ago, there was Meltdown and all these side-channel cyberattacks on CPUs. This cannot happen, or the attack surface is significantly reduced, when you run them on different platforms.
On the other side of the network, we also have technology. That's what makes the data center more efficient. I may not be objective, but I do agree that the merger of Mellanox and NVIDIA actually goes both ways.
I don't think the networking business—now NVIDIA, previously Mellanox—could have grown that significantly at the rate it has grown. Now, I think we are the fastest-growing networking business, let alone NVLink and InfiniBand. But just the Ethernet business is the fastest-growing business ever.
Yeah.
What are the things that break as you get to 100,000, and maybe eventually 1 million, GPU clusters? How do you use software to help design around that?
It's a multistage challenge.
One of the things you need to keep in mind, which is not very obvious to all engineers when you design a machine or think about how to operate it, is that you have these components, they're working, and now you just need to figure it out.
The hardware component works at 99.999-whatever percent of the time. That's usually okay if you're dealing with a single box or a couple of them. But if you're building a 100,000-component machine—a 100,000-GPU machine—which means there are millions of components, the chance that everything works is zero. Something is definitely broken.
You need to design it from both a hardware and a software perspective to keep going as efficiently as you can, maintain performance, keep your power efficiency, and, of course, keep the service running. This is challenge number 1, even before we get to millions. This challenge actually starts at a few tens of thousands.
Number 2 is that, when you're running these workloads, it is really important to write the software and provide all the interfaces needed to place the different parts of the job more efficiently. Sometimes you run a single job on the entire data center, and then you need to place the different parts of the job efficiently.
Building networks at this scale is a very different story. Building a compute network at this scale is very different from building a general-purpose data-center network.
Yeah.
A general-purpose data-center network is the Internet. It's not a big deal. Well, it is a big deal, but it's a different deal. You're serving loosely coupled, collaborative microservices that create the service you see as a customer from the outside. Here, you're running one single application on 100,000 machines.
Yeah. Is that specific to training workloads, or is that also true for inference workloads?
It's true for everything. It depends on the scale. Inference is yet another topic that we may touch on.
Until recently, training was the key thing. There were a lot of GPUs, and training was being done in a very specific way. You basically copy the same model onto multiple machines or multiple sets of machines and run them, then consolidate the results, and so on.
With inference, the story is a little different. But you need to provide the hooks in the hardware and in your low-level system software for applications and schedulers to place the job and the different parts of the job in the most efficient way.
As long as your machine fits in a building, which is about 100,000 GPUs, that's one thing. Now, when you're talking about gigawatts, it's all power-driven. The challenge is that, for many reasons, you want to split your workloads across multiple data centers.
Sometimes data centers are many kilometers, many miles apart. They may be across the continent. This comes with yet another challenge, which is the speed of light.
Yeah.
The latency variance between different parts of your machine is dramatically different. What's even more challenging is that, when you talk about networks, congestion on the network is one of the key problems that deteriorates network performance.
Yeah.
Managing congestion across such a latency difference is not like in the old telecom days, when you put a box at the edge of your data center with huge buffers and it acts as a shock absorber for congestion. A huge buffer is not good. Bigger is not better. There is a famous statement from a very famous one.
These buffers, or these devices, are basically there to isolate the external world from the internals. But when you want to run a single workload across data centers that are separated by kilometers, every machine on one side needs to be aware of whom it communicates with, whether it's a short communication or a long communication, and adjust all the communication patterns accordingly.
You don't need these big buffers, because a big buffer creates jitter.
Yeah.
Mm-hmm.
We have a technology that we actually developed recently. All of our Ethernet networking is Spectrum-X. This is a technology that we designed and developed based on the Spectrum switch that we put on the edge of the data center, and it provides all the information and telemetry needed for the endpoints to adjust for congestion.
Yeah.
Can we talk a little bit more about training versus inference? How does the shape of the workload differ?
Mm-hmm.
I guess backpropagation is a lot more computationally intensive, and the forward pass is less so. But how does the workload differ?
Mm.
Are you seeing customer demand start to shift from pretraining toward inference, or do you think it's still very training-heavy right now?
And if I could just ask a quick follow-up question: Will people be running inference workloads on the same data centers that they use for training, or will these end up being 2 separate systems? Because they're different optimizations, will people end up using 2 different sets of data centers?
Okay, yeah. That's a great question. Let me start with the first one.
Training has 2 phases. One is inference, which is just forward propagation, and then backpropagation to adjust the weights. For data-parallel training, there is yet another phase to consolidate the results of the weight updates across multiple model copies.
Until recently, training was the main driver of compute, because until not very long ago—maybe 2 years, which is ages in the AI era—inference in AI was mainly perceptual.
Mm.
You show the picture: That’s a dog. You show the photo of the person, and here’s Michael and here’s Sonya, and that’s it. That’s a single path, and that’s it. Then came generative AI, where you actually get recursive generation. When you pose a prompt, it’s not just 1 inference.
Mm-hmm.
It’s many inferences. For every token, when you generate text or a picture, you need to go through the entire machine all over again. So instead of one-shot inference, there’s more. And now there’s reasoning, which means the machine starts thinking.
Yeah.
If you ask me what time it is now, I can tell you. It’s easy, right? But if you ask me a more complicated question, then I need to think. I probably need to wait or compare multiple solutions or multiple paths. Every such thing is inference.
Mm.
Every such thing is inference. Inference itself has 2 phases. One is much more compute-intensive, and the other one is memory-intensive.
Mm.
It’s what we call prefill, because when you do inference, you have some sort of background, right? That’s the prompt—some relevant data that you need to process and create the context to generate the answer. This is very compute-intensive; it’s not very memory-intensive. The other part is actually generating the answer, which is the decode part of inference, where you generate token by token.
Mm-hmm.
There are some technologies that let you generate more than 1 token, but it’s still a single path, much less than the final answer.
Mm.
If you combine all these things together, inference demand for computing is actually not less than training.
Mm.
It’s actually even more. There are 2 reasons for this. One is what I explained: There’s much more computing than there used to be for inference. The other thing is, you train a model once, but you infer many times.
Yeah.
ChatGPT has almost 1 billion people, right? There are customers; they’re pounding it all the time with the same model. They trained it once.
More than 1 billion.
Right. Now they’re making videos, so that’s a lot of inference. You can generate, and everybody is doing the inference. My wife, I think, talks to ChatGPT more than she talks to me these days. Once she discovered it, it became her best friend.
In terms of inference, to your question about machines, you can do inference on the phone. Okay? So there are definitely going to be much smaller-scale installations for inference.
Hmm.
It’s like mobile devices. If you look at data-center scale, the efficiency of programming and programmability is much more valuable than hardware optimization. Every hardware instance has its own cost and its own drawback. It’s a very similar GPU, with the same programming model as a GPU for prefill versus decode. I don’t remember when it happened, but we announced that we are building a GPU SKU optimized for prefill.
Hmm.
Hmm.
You will have a GPU that can do decode, and a decode GPU can do prefill. You can equip your data center with SKUs for prefill versus SKUs for decode to optimize for typical use.
Yeah.
But if your workload shifts toward more decode or more prefill, you can use either one of them to compensate. This is the importance of programmability: the same interfaces for GPUs. It’s based on CUDA and UP, which is what made NVIDIA NVIDIA before Mellanox. Yeah, yeah.
Can I ask you a question about data-center scaling? For many decades, we had Moore’s Law, and chips got more and more dense and produced better and better performance. Then we ran into the laws of physics.
Chips just couldn’t get more dense because their quantum-mechanical properties caused them to break down. So then we had to scale up to the rack level, and now we have to scale out to the data-center level. Is there some analogous law of data-center scaling that says when data centers get too big, the communication overhead causes the performance to break down? Or, said differently, is there a natural limit to how big data centers can get?
I think there is a practical limit to how much energy you can consume within a given size of data center.
But if you were surrounded by nuclear power plants and the energy was available, would the—
Well, the—
…would the data center itself perform?
I don’t know. I’m not an expert in construction, even. But if you surround it with nuclear power plants, there’s energy coming in. Now the heat is going out.
Yeah.
There’s a whole bunch of technology coming to help make this more and more dense. We have now moved pretty much entirely to liquid cooling.
Yeah.
One of the reasons we did it is to enable much denser compute power. We couldn’t build computing as dense as what we’re building now with air cooling. The last big data center, which is, like, xAI-scale, is 100 or 150 megawatts now. We’re talking about gigawatt data centers. People are talking about 10-gigawatt data centers. There’s a desire to build much, much bigger data centers.
Are you sending the data centers to outer space?
Free cooling, free power.
I think one of the things that determines the speed of data-center deployment is how fast concrete gets stable.
Fair enough.
Before starting Mellanox, you were at Intel for 16 years?
That’s right.
Sixteen years. You became chief architect? NVIDIA and Intel recently announced a partnership. Can you share a little bit about what the vision for that might be?
The starting point is that computing changed in the last decade, or a little bit more than a decade. NVIDIA started as the accelerated-computing company. Video games were the first application, and then it evolved to AI, which is the new way of data processing.
A general von Neumann machine just isn’t capable of being used as a platform to solve the problem. Programming a von Neumann machine is just explaining to somebody what to do. I can explain many things, and I can explain to many people what to do, but I can’t explain how to distinguish between a cat and a dog, right? There are new challenges that AI solves, and you need acceleration there.
Our partnership with Intel is actually fusing accelerated computing with general-purpose computing.
Mm-hmm.
General-purpose computing isn’t going away. Everything will be accelerated, but we accelerate general-purpose computing. We accelerate the applications. x86 is the architecture that is dominant there, and it would serve both companies greatly.
That’s one of the interesting things about NVIDIA: It’s the culture of win-win. We’re not after taking a bigger piece of the existing pie. We’re after baking a bigger pie for everybody. Our success is our customer success. Our success is not the failure of our competition. Our success is the success of our customers and the success of our ecosystem.
I think fusing together conventional computing, von Neumann machines, and accelerated computing provided by NVIDIA probably opens yet another dimension that I’m not sure what it is, but in the practical, short-term view, this gives NVIDIA and Intel channels to the market, expanding the market and serving markets that otherwise were more challenging.
Mm.
You mentioned the culture of NVIDIA. When Mellanox became part of NVIDIA in 2019, the market cap of the combined company was about $100 billion, which is no joke. But the market cap today is about $4.5 trillion.
That's right.
A 45× growth in value in 6 years is pretty phenomenal. How has that changed the culture of NVIDIA? How is NVIDIA different today, now that it's one of the most admired companies in the world, if not the most admired, versus 6 years ago?
When we just joined, Jensen was in Israel, and I presented him with my belief that one plus one would be 10. I was actually off by a factor of 4. But Mellanox and NVIDIA, in a sense, are similar. The cultures were very similar to begin with, but nothing is absolutely similar.
I was the only founder who left Mellanox after Eyal resigned a few months after the acquisition. My main focus at the beginning—the things you think about in the shower—was how to make sure that this acquisition would succeed.
Yeah.
NVIDIA paid $7 billion for a company that I founded. With all the mixed feelings that were there, once it's done, it's done. Now I have to make it successful.
Yeah.
Eventually, it worked. Most of the Israeli employees stayed. I think it was 85% or 90% of the original employees.
Wow.
Actually, NVIDIA grew more than 2× in Israel in terms of manpower.
Yeah.
We're growing, and we're announcing that we're actually going to build a new NVIDIA campus in Israel.
Nice.
That's where I think the overall merger was very successful. I did my best to make sure it succeeded. Besides the technology that I was looking at—which is sort of technology, but it's technology and theology—there were many other things to make sure that people were comfortable, so that being at the center of Mellanox, whose headquarters were in Israel, they didn't feel left somewhere far away.
Jensen basically emphasizes networking as the critical part of NVIDIA's success.
Yeah.
And he's right.
Yeah, yeah, yeah.
I think it's considered to be the most successful merger in the history of technology. You guys probably track these things better than I am, but overall, I think it was a great move.
Yeah.
What are the science-fiction things that you spend your time thinking about, just wondering about? For example, optical interconnects: Do you think those will exist? Do you think AI will ever be better at physics than us, and better at data-center design than us? What do you think about?
What I'm thinking, if you look at science fiction, is how to make history experimental science.
Mm.
In physics, you try something, see if it works, and then try something else. In history, time goes in one direction. But if you have a good simulation of the world, you can make history experimental. We have Earth-2, a climate simulator.
Mm.
With this type of technology, we can actually simulate how what we do today will impact global warming 50 years from now.
Mm.
Mm.
So, experimental science. You try something and see what happens 50 years later. That's the science-fiction part.
Yeah.
On physics, now we are moving from reasoning and so on and so forth. Once we get AI models to understand physics, we can actually learn physics.
Yeah.
AI can teach us physics because the way we get to the laws of physics that we observe— theoretical physics, right?—is that you observe some phenomena, generalize it, and compose the rule, basically the physics law, that stays underneath these phenomena. AI is really great at generalizing, data processing, and observing. So AI can help us get to know some laws of physics that we don't even imagine now.
Yeah.
Moore's Law was 2× every 2 years. Huang-plus-Kagan's Law is... What is the slope, and how long do you think you can sustain it?
The slope is somewhere in the range of 10×, or a few orders of magnitude, a year. That's what we are doing now, by the way. About 2 or 3 years ago, we accelerated our product introduction from every other year to every year.
Now we introduce a new wave of products every year, and it's an order of magnitude higher performance. It's not chip-level performance. It's the machine that you can build with this performance. That's what we are looking at: a single unit of computing.
How long it will stay, I don't know. We'll do our best to maintain it as long as needed and probably even accelerate. It's all about the exponent. It's hard to imagine.
If you look at Moore's Law curves, or any law curves, they usually plot them on a logarithmic scale. So it looks linear, but that's the wrong thing to look at. You can't predict what's going to happen.
Who could predict that when the iPhone was first introduced, or when the smartphone was first introduced—that's 15 years ago?
2007 was the iPhone.
'07.
Yeah, 2007. Oh, 17 years ago. Who could imagine that the least-used function of this smartphone, at least for me, is a phone?
Yeah.
All of this is e-commerce, texting, news, and mail. It's basically running your life from this machine.
Yeah.
It's your authentication; your ID is there. Who can imagine what's going to happen 10 years from now with all these developments that we are doing today? But we are building the platform for innovation.
Notwithstanding your commentary on “who can imagine,” what is the most optimistic view of our future with AI that you like to think about? What could AI do for the world 5, 10, or 15 years from now?
Steve Jobs called the computer the bicycle of the mind.
Yeah.
AI is maybe—I don't know if it's... It's probably the spaceship of the mind. There are a lot of things that I would like to do, but I just don't have enough time or resources to do them. With AI, I will have that.
It doesn't mean that I will do twice as much. Maybe I will do 10 times as much. But I will want to do 100 times as much as I want to do today.
You go to any project leader, and nobody says, “I have enough. I have enough manpower. I have enough resources. I don't need any more.” If you give him a resource that is twice as efficient, he'll do 4 times more. And he'll want to do 10 times more.
It's like electricity changing the world, right? In London, you still see these gas lamps and the infrastructure to use gas as the source of energy. Who could think that once electricity was invented, it would change the world, and that we couldn't live without electricity?
Mm-hmm.
The same with AI.
Beautifully said. Thank you so much for joining us today. I love this conversation.
Thank you.
Thank you.
Thank you for having me.