从芯片到电力,人工智能如何重塑计算
Ben HorowitzMartin CasadoRaghu RaghuramErik Torenberg
- a16z 正在推出一只瞄准 AI 硬件瓶颈的新基金,而最清晰的信号来自创始人转向硬件:顶尖创始人的硬件项目占 deal flow 的比例,已从约 3–5% 飙升至“超过20%甚至30%”。 Casado 的判断是:“创始人群体往往比 VC 群体聪明得多,他们已经认定这是一个非常活跃的领域。”基金覆盖的是“计算机科学基础设施”——模型运行所依赖的一切,从芯片、互联网络“可能一直到电力”,但不包括受监管的垂直行业。
- 反炒作证据异常具体:Raghuram 称,超大规模云厂商今年资本开支约7000亿美元,明年合计将接近1万亿美元;Casado 称供应“基本已经预订到2028年”,GPU 转售价达到采购价的4倍,而可触达市场目前只覆盖了5–10%。 Casado 表示,“我们从未见过芯片价格上涨”,Horowitz 则指出芯片价格通常只会下跌。在 Hot Chips 大会上,一家领先内存供应商称,仅当前需求就需要消耗其3年的产能;与暗光纤时代不同,“现在生产出来的每一块 GPU 都已经预售”。
- 核心经济变化在于:《人月神话》不再支配这一瓶颈——资金投入如今可以直接转化为能力,工程本身不再提供天然的约束。 Horowitz 称,过去通过雇佣1000名工程师追赶两年领先“从来不会奏效……但现在可以了”,方式是“投入30亿美元,点亮一座壮观的集群”——这就是 Grok 或 Kimi 如何“凭空出现”。Casado 预计 token 需求将“接近每年增长1000%,而供应不可能以那么快的速度增长”;Horowitz 则认为算力需求还将持续数十年。
- 本期最具交易价值的思维模型,是 Casado 关于单模型 ASIC 的测算:一款成本30亿至50亿美元的前沿模型,需要约100亿美元的推理回报,因此20%的效率提升就价值20亿美元——“你完全可以用20亿美元做出一颗 ASIC”。 模型权重固定,使得按模型定制芯片变得可行,而传统有状态软件并非如此;不过 Casado 也表示,他们还不知道行业是否真的会走向这一步。硬件优化如今“绝对会对业务上行空间产生实质影响”。
- 物理瓶颈确实存在:机架功率需求正从约5–10kW走向“100到50千瓦”;这一功率水平下交流电已经不再适用,液冷也已成为入场门槛;美国电气工程师或电工中只有2%持有直流电认证;钢筋混凝土价格涨幅居前;而到2028年约44GW的新数据中心需求,对应的电网新增供给预计只有约25GW。 美国的障碍如此严重,以至于新公司往往要去“墨西哥或澳大利亚”寻找 GPU 产能。
- 现有巨头不可能吃下全部市场:面对数万亿美元级的芯片巨头,“其中5%就足以造就一家规模巨大的私营公司”,而 NVIDIA 理性地不会去捡银砖,只会牢牢抓住金砖。 市场会先扩张、再分化,正如从 Ford、Fordlandia 最终走向供应商生态的类比;而绝望中的实验室已经在“硬件真正可用之前,就与相关公司签下协议”。
- 关于 agents,Casado 的框架是:正确的载体不是你的延伸,而是“一名真正的员工”,拥有自己的电脑和浏览器。 Horowitz 坦言自己还没有“破解这套方法”:机器人“可以烧掉大量 token、花掉大量资金,却没有产出任何有用的东西”;目标是“让所有人类都变成超人,同时不能因为机器人失控而把整个组织搞垮”。
1. Machine Age 基金:瓶颈已经下沉到“模型之下”
- Horowitz 的开场框架是:“我们拥有一种全新的技术,这是有史以来最重要的技术”,而每一种这样的技术都需要一套全新的基础设施——“不仅需要新芯片、新系统软件,还需要新的供电方式,需要替代铜……几乎一切都要重做。”
- Casado 的表述是:过去基础设施意味着服务器、存储和网络;“但现在一路下探到了铜矿。”模型能力提升得“越来越快”,已经“不再是瓶颈”,如今瓶颈在模型之下的所有环节。
- Casado 观察到,硬件过去可能只占顶尖创始人 deal flow 的3–5%(Horowitz 认为3%,“5%可能已经算慷慨”),如今已超过20%乃至30%。“创始人群体往往比 VC 群体聪明得多,他们已经认定这是一个非常活跃的领域。”基金的明确投资范围是“计算机科学基础设施……模型运行所依赖的一切”——芯片、网络、互联、存储,“可能一直到电力”,但不包括受监管行业或垂直化行业。
- 基金没有新增 GP;团队称,硬件经验本来就写在其基因里,包括早期投资 SpaceX、Anduril、Astranis 和 Waymo。
2. 为什么这不是炒作周期:GPU 已预售,芯片价格反而上涨
- Raghuram 列出的需求账本显示,超大规模云厂商今年资本开支约7000亿美元,明年合计据称将达到约1万亿美元;而云厂商看到的需求来自各个方向:前沿实验室、AI 原生公司、企业客户以及海外市场。应用公司全都在高速增长。
- Casado 补充称,目前可触达市场只覆盖了5–10%。供应“基本已经预订到2028年”,几千块 GPU 的交易要经过持续数日的竞价,转售价则达到采购价的4倍。
- 价格是关键线索:Casado 表示,“我们从未见过芯片价格上涨”;Horowitz 则称,芯片价格通常会下跌。在斯坦福举行的 Hot Chips 大会上,一家领先内存供应商表示,不是未来需求,而是今天的需求,就足以填满其3年的产能。
- 互联网泡沫时期的对比在于:1998–99年的建设潮押注的是投机性暗光纤,当时没有用户消费这些产能,也没有可行的视频应用;而现在,“生产出来的每一块 GPU 都已经预售”。
- Raghuram 讲了一个轶事:一家对云迁移持抵触态度的大型上市公司,其 CFO 发现服务器里的内存增量如此之大,产生的价值足以覆盖整个云迁移。
3. 《人月神话》不再支配这一瓶颈:资金可以直接兑换成能力
- Casado 的宏观判断是:过去,建设本质上是一个有天然约束的工程问题;如今“真正的问题是资源限制”——“从资金投入,到硬件创造智能,中间已经没有任何阻隔。”
- Horowitz 解释了为什么领先优势不再构成保护:过去,靠雇佣1000名工程师追赶两年的领先“从来不会奏效……但现在可以了”。方法不是再招工程师,而是“投入30亿美元,点亮一座壮观的集群”,于是“无论如何,Grok 都可能凭空出现……或者 Kimi”。
- token 需求持续倍增的来源包括 RL、思维链和长时间运行的 agents,它们本质上都是推理。Raghuram 表示,“没有谁比 AI 更喜欢使用 AI”;甚至让 AI 编写 GPU kernel 来创造更多 AI,也构成了自催化循环。单位工作量的价值在上升,受益者则从开发者扩展到所有知识工作者,乃至更广泛的人群。
- 即便有预言家也无济于事:“我们进入这轮浪潮才4年”,而芯片周期为3–4年,数据中心周期为4–5年,“我们不可能提前建好这些产能”。
- Casado 认为,更合适的类比是蒸汽机或电力——未来“可能有30–40年都在把算力投入到问题中”,前提是问题具有清晰的回报信号。语言和代码已经建立起来,计算机使用才刚刚开始,之后还会扩展到科学、材料、生命科学和创造力。
4. Agents 是员工,而不是你的延伸
- Casado 对这一问题的认识经历了三阶段:先把 AI 当作一项功能,“就像搜索框”;随后进入聊天阶段;再后来,OpenClaw 成为你的延伸,“共享你的密钥、知道你的密码”;最终,Clawdbot 真正抓住了关键:“它实际上就是一名员工”,拥有自己的电脑和浏览器。
- Erik 讲到,他曾在一个周末里让它更新自己在多家服务商处登记的信用卡,并取消订阅。“这不是编程……这是真正的计算机使用。”
- 需求来源正在扩展。Erik 提到,一款 ChatGPT 应用每周活跃用户约10亿,约3000万名开发者也在使用模型,并消耗了相当大一部分算力。Raghuram 则表示,计算机使用机器人相当于创造了5亿名知识工作者;具身 AI 和机器人则是尚未完全到来的另一大需求来源。
- Ben 对组织问题直言不讳:这些新员工“可以烧掉大量 token、花掉大量资金,却没有产出任何有用的东西……它们可以编造内容……也可以制造安全问题”,但同时也能“极其高效”。他明确表示自己并没有宣称问题已经解决,目标是“如何让所有人类都变成超人,同时不能因为机器人失控而把整个组织搞垮”。Ben 称,在尝试了多种方法后,Martin 提出的“就把它们当人看”最经得起时间考验。
5. 从第一性原理重建技术栈:可能每个模型都需要一颗 ASIC
- Raghuram 的方法是把推理引擎拆解为基本要素:计算究竟是什么,也就是矩阵乘法;内存应如何分层组织;芯片如何在芯片内部、芯片之间以及数据中心之间连接;每次传输需要消耗多少电力;然后“从这些基本构件出发重新搭建”。这一工作“目前正在行业内展开”,机会也正在这里。
- Casado 给出了按模型定制 ASIC 的测算:训练一款前沿模型需要30亿至50亿美元;推理端需要实现约2倍回报,也就是100亿美元。节省其中20%就价值20亿美元——“你完全可以用20亿美元做出一颗 ASIC”。传统软件具有状态且动态变化,而模型权重是固定的,因此每个模型配一套定制芯片具备可行性;但 Casado 表示,他们还不知道行业是否真的会走向这一步。
- 这种模型经济学前所未有:“我不认为在整个行业历史上,我们曾经创造过一种数字资产,有50亿美元直接投入其中。”
- 值得保留的利润率判断是:软件业务跑通后,软件毛利率往往自然形成;但在 AI 时代,“硬件优化对业务上行空间的意义绝对不可忽视,这是过去从未出现过的情况”。
6. 电力、散热与混凝土:物理世界的墙
- 机架功率需求正从约5–10kW走向“100到50千瓦”。Horowitz 称,这意味着“交流电已经不再适用”;直流电需要独立的冷却系统,而且“极其危险”。讽刺的是,爱迪生当年还曾通过电击动物来攻击交流电。美国电气工程师或电工中只有2%持有直流电认证;Meta 目前正在提供免费培训——“AI 将创造大量新的电工。”
- 许多现有建筑设计已经过时:地板必须承受更高密度的机架,墙体必须加厚以隔绝噪声;“一旦进入 Feynman 时代,今天的数据中心中能继续使用的比例会大幅下降。”钢筋混凝土是价格涨幅最快的材料之一;超大规模云厂商也在尝试用机器人组装和安装服务器。
- Casado 做了一个规模校验:1GW 电力大约相当于5万户家庭的用电量;他的家乡亚利桑那州 Flagstaff 人口约4万至6万人,耗电量还不到1GW。但主持人援引的数据显示,到2028年新增数据中心需要约44GW电力,而电网预计只新增约25GW。
- 变压器和涡轮机都短缺,许可审批和电网接入也面临监管障碍;新公司如今开始去“墨西哥或澳大利亚”寻找 GPU 产能,“因为在美国实在太难了”。
- Horowitz 提议建立一项标准,让数据中心向当地反哺:“电力质量变好,没有噪声,也没有用水问题,还能创造就业。”一些数据中心已经在这么做:白天向州政府供电,夜间州内需求较低时再借电;它们的能源费率也在逐年下降。
7. 为什么 NVIDIA 之外仍能出现初创公司,以及谁来建立它们
- 历史上曾出现过独立公司:互联网转型时期有 Cisco 和 Juniper,超大规模数据中心时代有 Arista。但 Casado 表示,相比当前这轮变化,过去那些转型都相对有限,因为这一次“每一件事都不一样”。
- 市场逻辑在于,tokens/秒/美元、tokens/瓦、tokens/机架等指标仍会持续改善,而电力效率也需要新的创新。Casado 补充称,面对数万亿美元级的现有芯片巨头,“其中5%就足以造就一家规模巨大的私营公司”。Horowitz 举了 Alex Rampell 向 Facebook 推销 TrialPay 的例子,Dan Rose 当时说:“我手上有这么多金砖,连一块都捡不完,所以我最不会做的事,就是去看一块银砖。”
- 市场在扩张后会走向分化。Casado 以 Ford 1913年在 River Rouge 的工厂为例,说明极端的垂直整合;Horowitz 又提到 Fordlandia——Ford 在亚马逊打造的美国化橡胶种植园城市——其最终陷入困境,原因之一是要求工人“准时参加各种活动”。同样,“过去推理只有一种简单架构,但现在已经不是了”。
- 创始人画像也不一样:首轮融资往往在产品推出前就达到数亿美元,这些必须是从第一天就考虑制造和供应链的“系统型创始人”。Martin 说,“Jensen 是这个领域的 Michael Jordan。”Raghuram 指出,内存专家“并不年轻”;但他同时表示,他们所有的投资项目都是由2名20多岁的创始人创立,再与经验丰富的人一起工作。
- 两个顺风因素正在加速行业发展:绝望中的实验室已经在“硬件真正可用之前,就与相关公司签下协议”,后续融资也更容易获得。
- 名称与 stakes 在于:Martin 表示,“人工智能这个词错了……应该叫机器智能”,它是“人类思考方式的缓存”;这个命名也承认了硬件的核心地位。Horowitz 认为,未来5–10年的胜负标准是“美国赢在基础设施”,拥有环保型数据中心,以及充足的芯片、内存和电力。他说,美国是一个特殊的地方,一个一无所有的人也能来到这里,做出意义深远的事情;如果失去技术领先地位,可能迎来“另一个时代、另一个国家,也许他们拥有一套不同的价值观”。
完整逐字稿
We have a whole new technology. That's the most important technology ever. And you need a whole new infrastructure.
Normally, when we talk about the infrastructure world, we're talking about the servers, the storage, and the network. Here, it goes all the way down to the copper mines. That's how widespread this thing is going to be.
It used to be, when you built something, it was an engineering problem. Here, it feels like it really is a resource limitation. So whether it's tokens or not, we're pouring a ton of money into systems, and then those systems are producing a result. Right now, we're bottlenecked on the systems' ability to actually match the resources we're pouring into them.
The leading memory vendor said the demand they have today will take them 3 years of capacity to supply.
If this fund does what we think it will do, how do we see the world in 5 to 10 years?
America wins in the infrastructure, and that would be awesome.
1. Introducing the Machine Age Fund
Ben, Martin, Raghu, welcome.
Thank you.
All right. Thank you. I want to start with a Marc quote to introduce this new fund: “This is the biggest technological revolution of my lifetime. This is clearly bigger than the internet. The comps on this are the microprocessor, the steam engine, and electricity—or maybe the wheel.” Guys, the Machine Age Fund. Please introduce it. Ben, start us off.
Basically, what’s happened is we have a whole new technology that’s the most important technology ever. Every time there’s a dramatic new way of using all of the things that we love in infrastructure, you need a whole new infrastructure, and never has it been more high-impact than it is on this one.
2. Why Founder Interest in Hardware Just 4x'd
Not only do we need new chips and new system software, we need new ways of doing power, and we need to replace copper. It’s absolutely everything, so it’s a very exciting time. Particularly for the hardware aspects of this new era, we needed a new approach.
Yeah, I would agree. Normally, when we talk about the infrastructure world—at least in computing—we’re talking about the servers, the storage, and the network. Here, it goes all the way down to the copper mines. That’s how widespread this thing is going to be.
That’s number 1. Number 2, I think what we’ve seen over the last 3 years is a steady increase in the capabilities of the models, where the model is no longer the bottleneck. In fact, using AI, these models are getting better faster and faster and faster. Now, the bottleneck is everything I call south of the model, and so that’s why we need to work on that.
The only thing I’d add very quickly is that we tend to follow founders, and we’ve been watching the number of very strong teams going after complex hardware problems increase over the last couple of years. I don’t know the actual numbers, but I was trying to estimate it over the weekend.
I think maybe 5% of the deals from top founders would have been hardware before. Now, I would say north of 20% or 30% right now. The founder community, which tends to be much smarter than the VC community, has identified this as a very active area for innovation, and they’re responding.
I think 5% is probably generous.
Yeah, it’s very low. It’s very low.
Yeah. 3%.
Explain some of the macro conditions that have led to this change—the surge of founders pursuing these ideas. What are they seeing that’s enabled this?
The obvious thing is that demand for AI is basically infinite, and as a result of that, every part of the supply chain is under duress. Everything, including materials used to make things like memory.
It’s also interesting whether there’s something unique about AI. Because demand is infinite and growth is infinite, what you tend to worry about is the margin of companies—how efficient it is. Normally, you worry about growth: Can I just get people to buy this stuff? You don’t have to worry about that here. The question is whether you can do this in a way that’s profitable.
3. How Do We Know Demand Isn't a Hype Cycle?
A lot of the efficiencies are actually strictly a physical limitation of hardware. Even the business model of the AI wave is putting a lot of stress on the existing systems because they weren’t built for AI. They weren’t built for those workloads. There’s just this global observation that we need to change the core components to get that efficiency, help drive the growth, and drive the value of the businesses.
How do we know that demand is actually outpacing supply here, rather than this being another hype cycle?
There are any number of signals today. First, some of the smartest judges of demand are cutting huge purchase orders. If you look at the hyperscalers, their capital expenditure has been exploding. Next year, it’s supposedly going to reach $1 trillion collectively across the big hyperscalers. This year, it’s about $700 billion.
If you think about the hyperscalers’ position in the industry, they see demand from everywhere. They see the frontier labs wanting their compute, the AI-native companies, the enterprise, the U.S. geography, and the international geography. If anybody has visibility, it’s them, and they’ve been jacking up their capital expenditure like it’s never been seen before. That’s a clear, clear sign.
Secondly, if you look at the companies that we see on a day-to-day basis, they’re all ripping. All the application companies—the growth is insane. The frontier labs—the growth is insane. It’s been documented. I would say that, on the demand side, the signals have never been clearer that this is not a hype. To top it all off, prices are going up.
We’ve never seen prices go up on chips.
Typically, prices went down. They always went down.
Yeah.
If you look at the price curve, it went like this and then went back this. We know that only 5% to 10% of the addressable market has been addressed today.
The supply, if you look across the board, is basically all booked out to 2028. It’s so bad that we’ve actually seen multi-day auctions for a few thousand GPUs. The other side of that, of course, is demand. As Raghu said, we’ve seen the fastest-growing companies in the history of the industry.
4. Sold Out to 2028: The Unprecedented Supply Crunch
Also, the unit of work that AI can do—the value of that unit of work—keeps increasing. But underneath the covers, the number of tokens that are consumed is going up by orders of magnitude, right? If it’s 100 tokens for chat or an agent, it’s thousands of tokens, right? So you’ve got expansion of both sides of demand. One is the unit of work is becoming more and more consumptive of tokens, and then secondly, the number of people therefore that are going to benefit—it’s not just the developers, it’s going to be all knowledge workers, and then all of beyond that. So that’s what we see.
You said the key components in supply are sold out to 2027, maybe 2028. What does it mean for an entire industry to be sold out that far? I don’t know if this has ever happened before. Do you guys recall?
Remember in the internet days, when we were doing a massive buildout, the majority of what was actually being put in the ground was speculative and dark. Remember dark fiber? Here, basically every GPU that’s being created is already presold.
Yeah, we weren’t quite there. There was a lack of bandwidth in the ’98–’99 timeframe, but there wasn’t that much real demand for it because there just weren’t that many people on the internet.
It was a two-sided thing. The companies were all rushing there and theoretically needed more bandwidth, but there weren’t necessarily enough users on the other side to consume it. To really consume a lot of bandwidth, you have to do high-bandwidth things like video, which weren’t really viable for a number of reasons that had nothing to do with how much bandwidth was in the data center.
It smelled similar, but it wasn’t this. This is like we’re flat out, and people are reselling GPUs for 4 times what they bought them for. It’s just not the same.
We’re also out of power and cooling. On top of that, it’s really hard to build because there are these incredible political headwinds going into it. It’s really unprecedented in my career that we’ve had anything like this.
I want to give you a quick anecdote. I was talking to the CFO of a large public company that had historically been very resistant to going into the cloud. They had a lot of servers, and they were doing an inventory check. They realized that the memory in their servers had increased so much that it could fund the entire migration to the cloud.
I just feel like we’re in a very unusual situation.
Yeah, that’s right. We’re out of many things: power, cooling, memory, GPUs—you name it, we’re out of it.
The flagship conference for the industry is called Hot Chips, and it’s going on at Stanford. The leading memory vendor said the demand they have today would take 3 years of capacity to supply. It’s just today; it’s not even future demand.
In terms of being out of everything simultaneously, is it because people just underestimated how good the models would be and how useful they would be? They just couldn’t have foreseen the demand?
Well, I don’t even think it’s that.
This stuff came out of nowhere, right? We're only 4 years into this. So even if we had a perfect oracle, once it started working, I don't think—
We could have built the capacity.
We could have built the capacity. There's no way. And we're talking about chip cycles, which tend to be 3 to 4 years. We're talking about breaking ground and building data centers, which is 4 to 5 years. We're talking—
And connecting, breaking ground, building them, and having a power source. So you either have to build your own power, or usually both: you've got to build your own power and have a power source, which is not easy.
The ML industry itself, if it's growing at 20% or 30%, has a great growth rate, right? And it's being connected to an AI software industry that's triple digits as the base, you know.
So you can see the disconnect, right? It's just a wide gap.
And so why didn't this fund exist 5 years ago or 7 years ago? Why was it not a great category to invest in, in the same way, Ben?
Well, I'd like to think we're just in time, but we probably would have been well suited to have it at least a couple of years ago.
I will say you could actually point to basically every epoch to an independent company that came up, right?
Clearly, in the move from the mainframe to the client-server, we saw a bunch of companies come up. In the move to the internet, we got Cisco and Juniper. Even in the mega data centers, which, by the way, was largely driven by the incumbent cloud providers verticalizing, you saw the rise of Arista.
There has been the ability to invest in silicon and hardware, but it's been relatively minor because the change has been relatively minor—one chip company, one switch company—where here, everything is different. And so I agree with Ben: we probably could have started a little bit earlier, but the amount of change is so high now that it's just an obvious thing.
And the other thing is, the demand for intelligence is so voracious, with really no end in sight. Every company that's adopted it is growing very fast in its usage, most companies haven't adopted it to a high degree, and consumers are just getting started. So the demand for tokens is probably going to grow close to 1,000% a year, which you cannot grow supply that fast. We're not—
The amount of work we're going to have to do across the board to get to the point where we can grow infrastructure at that kind of rate is pretty vast. So I think there's a lot of investing opportunity on the way.
And, by the way, the other thing is, all the architectures of the hardware systems were built for a whole different era of computing. So more than just needing more capacity, we need capacity to build. There are lots of opportunities to build different kinds of infrastructure.
Yeah, they're all reaching the physics limits for what they were designed for, right? Like what Ben was talking about with copper, and so on and so forth. You could go across every one of these categories and find, “Okay, this is the limit of this type of technology.” So now you've got to get some technical breakthroughs to get to the next one.
Yeah. I want to dive deeper on the demand side for a second. As we've moved from chatbots to reasoning to agents to multi-agents, each step has multiplied the number of tokens a single task takes up by increasingly larger orders of magnitude.
Yeah, nobody likes to use AI more than AI.
5. Tokens, Scaling & Why There's No Natural Regulator
Why does that keep happening instead of leveling off? Do you see that happening indefinitely, just continuing to—
Well, for sure, right now, if you look at the way we're achieving scaling, the way we're doing it is through a lot of inference, so through a lot of tokens. If you think about what RL is, it's a lot of inference. If you think about chain of thought, it's a lot of inference. If you think about long-running agents, of course it's a lot of inference. That's just basically been one of the approaches that we've been using to scale.
I think if you want to step back and say, “What is the macro trend here?” it used to be that when you built something, it was an engineering problem. You'd throw a bunch of engineers at it, and that didn't scale. There was a natural law of engineering physics, which is what The Mythical Man-Month came from.
Here, it feels like it really is a resource limitation. Whether it's tokens or not, we're pouring a ton of money into systems, and then those systems are producing a result. Right now, we're bottlenecked on those systems' ability to actually match the resources we're pouring into them.
I think tokens are probably where we are on the scaling curve right now, but we don't have a natural regulator like engineering, as we did before. So I think we should expect this to continue, and we have to build a supply to support it.
Yeah. Like, the simple way to think about it is that any problem you have can be solved with enough infrastructure—
Basically.
GPUs, power, and money. Until we run out of problems, we're not going to run out of demand. That's the challenge.
AI's answer to getting better and better is to use more AI, right? Inference is one basic building block that it keeps using over and over and over again, and so that's why these tokens multiply at each step.
Yeah, even the autocatalytic effect—the idea of using AI to create more AI, like creating a GPU kernel, of course, is just using more AI as part of the process.
One way that we think about it is that, in the past, money would come in, you would have an engineering problem, we knew that it took 2 years, it normally failed, and there was a natural governor. Then you got the product on the other end.
Here, there's nothing between the money going in and the hardware creating intelligence. So now we're just limited by our ability to create supply. It's a very, very different—
Yeah, that's right.
So that's the cycle. As long as you have the money, the GPUs, and the data, for the foreseeable future you'll be able to scale these things.
It's fascinating because, over the last decade, it feels like there were so many people saying the pervasive sentiment was that there was too much money going to startups. We were overfunding these startups. There was too much money in venture capital.
6. Agents as a New Kind of Employee: The GrokBot Moment
Say more, Ben, about what that means, because there used to be this skepticism that the more money you put into the industry, the bigger the outcomes would be. And now we were saying at the off-site that, to some degree, the market is as big as we collectively contribute to it.
Yeah. This is—look, the one thing we all knew in the startup world is that if I have a 2-year lead on you and you try to catch me by hiring 1,000 engineers, you're going to wreck your company. That never works. It's The Mythical Man-Month. Nine women can't have a baby in a month. That never works.
Now that works. But it's not hiring 100,000 engineers. It's taking $3 billion and lighting up a magnificent cluster, and then all of a sudden, Grok can come out of nowhere and you're like, “Oh, all of a sudden it's real.” Or Kimi, or what have you.
These leads—you can throw money at the problem, and you can throw money at almost any problem, and that works. So that is just completely different from anything we've ever lived through. We're all psychologically adjusting to this.
The ChatGPT app has a billion weekly active users. There are about 30 million developers who are using a relatively big portion of compute demand. How do we think about compute demand needs now and in the future in light of what people are actually doing with AI?
And now you have Grok bot, which is kind of what happened with coding, kind of happening with all use of computers via bot. So we're in a whole other wave of demand, and most certainly there's going to be more to come. It does seem quite unlimited at the moment.
We haven't even gotten into embodied AI or robots, which are going to be another source of demand.
Martin's an expert, but my understanding is that the bot uses computer use, which is just like—
A human being sitting inside the computer, typing away.
I literally used it over the weekend to update my credit card with a bunch of services that I'd been lazy to do, and to cancel a bunch of subscriptions. This is not coding or whatever. This is true computer use.
All of a sudden, you're creating half a billion knowledge workers, except they're all sitting inside the computer doing what?
I do think that Marc Andreessen is right. The right analogy here is the steam engine or electricity, in the following way.
We've introduced this new thing that you can put to work, and there are some very obvious applications now, but there's probably 30 or 40 years of throwing computers at problems. Anything with a clear reward signal, and we're just starting: we've got language and code, that's it, and we're just starting with computer use. What else are we looking at? Science, materials, biology. Of course, creativity is a massive use.
We're at the very, very early part of a very long journey, and we've removed this key bottleneck, which is traditional software engineering. Of course, bottlenecks will move and there'll be more complexity elsewhere, but I think we're at a very early point in a very long run of throwing computers at problems. So let's expect this compute need to persist for decades.
But because we mentioned it, Marty, talk about Clawdbot, because we were talking at the offsite about what struck you about it. Obviously, we're involved in every possible way you could be involved, but what did you find so interesting about it?
I think we've, as an industry, gone through multiple realizations of how AI enters our lives. Very early on, we were like, "Okay, well, you add AI to a product," and it's whatever—it's like a search bar. Then you chat with it and it chats back, because that's the traditional way to do it.
Then OpenClaw showed up earlier in the year. With OpenClaw, I said, "Okay, well, maybe it just being like Google but better isn't the full embodiment of it. How about we'll have it be a standalone thing, but it'll be an extension of you? It'll share your keys, it'll know your passwords, and it'll just do stuff that you would do." So it's kind of an extension of you, but it's more like a human.
What I think Clawdbot got really right is: No, how about it's actually an employee? Now you have this thing that's an entity, and it doesn't have special access to your keys or whatever. It has its own computer and its own browser, and because these are the smartest models in the world, it can do whatever an employee can do.
7. What's Actually Bottlenecked Right Now
It's interesting because now, if I want something done, my first thought is, "Well, can it do it for me?" Often the answer is yes, even if it's something you wouldn't expect. The obvious ones are things like managing my calendar or booking a meeting, but there are non-obvious ones as well. For example, I'll have it read through my email and do triage. I don't tell it how to do that, but it knows to check with me before actually doing the triage. These things are sophisticated enough that you can give them a relatively high-level task, and they'll do sophisticated things as a result.
Ben, I know you're thinking a lot about this and how this works in the organization. You think a lot about culture, of course. What are your thoughts here?
I think if you just look at us, it's like having a new kind of employee, and there are going to be a lot of them. Just like with our regular human employees, we spent many, many, many years figuring out how to work with them, and now we've got these other kinds of employees. There's a learning curve with them.
They can burn a lot of tokens, spend a lot of money, and get nothing productive done. They can forget stuff. They can make stuff up. They can have good behavior, and they can have bad behavior. They can create security problems. There are all those aspects to it, but they can also be super-duper productive.
I think figuring out how to integrate them and have them work nicely with the people they're working with—the actual humans—is all something that we're learning how to do. I don't want to sit up here and say I've cracked the code. The whole firm is completely automated now, and I'm going to slowly get rid of all the humans because I can. That's not at all where we are.
We're much more like, "Okay, how do we make all our humans superhuman without wrecking the place because the bots got out of control?"
We've tried a couple of different ways to get agents into the system, if you will, and eventually it was Martin's insight: Just treat them as people and get it done. That's what we're doing, and that's turned out to be the most durable way of getting this thing going inside an organization.
8. What "AI-Designed" Infrastructure Actually Looks Like
I want to go back to the supply side and go deeper into the bottlenecks. We were talking about how, in terms of the data centers, the chip architecture, system software, and the facilities themselves, none of them were designed with AI in mind. What would it look like for them to be designed with AI? What's the mental model for thinking about what that could mean?
If you start with the statement that the original model of infrastructure for any of these models has to change, you can go category by category and see how it breaks. Then you start unlocking the bottlenecks in each one of these things.
Eventually, you have to get to a system where, if you look at what an inference engine does, it takes up a lot of memory and generates new tokens along with the compute. You can think about how to optimize all of this. What does the memory need to be? What does the compute need to be? How do they need to talk to each other? How much power does each of them need? If they need all of this power, how do you cool each of them? Then how do you put collections of these things together?
That is the exercise that's underway in the industry right now with a lot of the founders. They're breaking down the problem into its fundamental components and saying, "What is the exact nature of the compute that's getting done?" It's going to be matrix multiplications. How do I optimize my compute around that kind of a scenario? They all need to progressively generate these tokens. What is the best way of hierarchically arranging this memory? How is the power consumed?
Then you've got to connect it together. What are the ways of connecting it on the same chip, but also across chips and across data centers? How much power does each of these data transmissions take? You have to progressively break it all down and rebuild it from these fundamental building blocks. That's what we see underway, and that's where we see the opportunity.
Let me give you an interesting mental model to think about how the landscape has changed. Today, to build a frontier model costs, let's say, $3 billion to $5 billion, and that's to train it. The inference has to pay back at least that, of course, in order for any of this stuff to be viable. Let's say 2 times that, so now inference has to make $10 billion. If you can save 20% in efficiency on that, that's $2 billion, and you can easily build an ASIC for $2 billion.
We've actually gotten to this interesting point in the industry where it makes sense to build an ASIC per model, just because of the amount of capital investment in that model. Unlike traditional software, which has a lot of state and is very dynamic, these models are fixed. The model weights are fixed.
9. Rack Power, Liquid Cooling & the Data Center Redesign
We don't know if the world goes to per-model ASICs, but it gives you a great mental model of how you would evolve the architecture to be far more bespoke for these massive capital investments we're making. I don't think in the history of the industry we've ever created a digital artifact with something like $5 billion going directly into that artifact. This is going to put the greatest demands on hardware that we've ever seen.
To that end, rack power requirements are moving from roughly 5 to 10 kilowatts to 100 to 50 kilowatts. Compute density is climbing something like 70x. Cooling is moving from air to liquid as a requirement. What are the investment opportunities as a result of this?
First of all, when you get to that level of power per rack, AC power doesn't work anymore. That's a pretty wild thing. Now you're into DC power, which, by the way, also requires its own cooling and is super dangerous.
It's kind of ironic because Edison promoted DC power by claiming how dangerous AC power was and demonstrating it by electrocuting animals and things.
The horse. Yeah. But he was right about his own kind of power, which is extremely powerful—that's the good news.
Starting with power, that's going to be very, very different. With cooling, yes, we're going from air cooling to liquid cooling. I think we're already at liquid cooling for any state-of-the-art data center. That's already kind of a done thing.
But it gets into the fact that, given the political environment and so forth, liquid cooling isn't enough. It's got to be eco-friendly liquid cooling. DC power isn't enough. It's got to be power that contributes to the power of society, not takes away from it. So you have data centers that have been behaving badly—a small percentage, actually, probably 10%—wasting a lot of water.
Not as much as pistachios or almonds and so forth, as people demonstrate on the internet, but they could be a lot more efficient with that. And then there are ones that are parasites of power and don't contribute power back. I think all that's going to end. It's going to have to end just because we've gone through a one-way door on that. That requires a level of engineering that many haven't invested in yet, so that's coming.
And then, if racks are that dense, there are other things—the way the floors are designed have to support that kind of weight. That kind of thing is actually for real. And then I think that you just need a lot of everything. Also, by the way, things are really loud, so you have to build the data center with thicker walls, or you're going to disturb the peace in the neighborhood, which is not going to be acceptable. I don't think any state is going to allow that.
And so a lot of the ways people have architected and designed the buildings themselves are already completely obsolete. Once we get to Feynman, a much smaller percentage of the data centers that we have today will work. In fact, everybody talks about memory prices, but one of the fastest areas where prices are increasing is reinforced concrete, for heaven's sake.
The other thing that happens when these data centers are sending 800 volts to the rack is that it becomes so dangerous, number 1. Secondly, we don't have enough electrical contractors who have the expertise to deal with 800 volts inside the data center, because this is high voltage. Only 2% of electrical engineers or electricians in the US have been certified on DC power, so that gives you an idea.
Now Meta has a whole program to train people up, which is great. It's like a new Job Corps where they train people for free to do this job. But it's funny: AI is taking all the jobs, and AI is going to create a lot of new electricians.
Yeah. I think we're doing something in this space, too. The big guys that own the big cloud data centers all are furiously experimenting with robots, right, to do the work of assembling or putting servers into the data center, et cetera. And so you'll see that increasing as a result of the evolution in AI.
By the way, to be clear on the actual fund that we're raising, our focus is on computer science infrastructure. So anything a model runs on—that's computer science, right? Think chips, network interconnect, storage, all the way down probably to the electricity.
Yeah, and say more about the robotics arm in terms of what we'll be doing versus maybe American Dynamism, or how to think about that.
Yeah, for sure. Again, we think that any platform that AI will run on—one of the great breakthroughs that AI does is it allows computers to interact with the physical world, right? It can see, it can hear, it can talk, right? And this means new platforms, right?
10. 44 Gigawatts by 2028: Why Building Faster Is So Hard
The simplest way people say “edge device,” but that doesn't really mean anything, right? I mean, it could be a mobile device, it could be a CDN, it could be a laptop, but it also could be an embodied device that goes around. And so, again, as infrastructure-focused investors, we don't do heavily regulated industries or more verticalized industries, but any sort of computer science platform that's going to push AI further out, we're quite interested in.
Yeah. Going back to the data centers, by 2028, new data centers are going to need something like 44 gigawatts of additional power, against maybe 25 gigawatts of expected grid additions.
Hold on, hold on. We use that word “gigawatt.”
No, it's like we'll have 100 gigawatts.
Martin, what's a gigawatt?
I mean, how big is it? It's multiple football fields. I mean, it's massive. It's 50,000 people.
What do you mean? What is it?
The equivalent is like 50,000 houses.
A city?
50,000 homes. I grew up in Flagstaff, Arizona, which is a town of 40,000 to 60,000 people, depending on the universities. We have less than a gigawatt of power consumption. I mean, this is—
So you can basically light up and air-condition your entire town for a gigawatt.
Yeah. I mean, this is—
We're just throwing them around.
No, but by the way, everybody talks about the gigawatt. There are very few gigawatt data centers that are actually up. We've got a long way to go, right? But then why can't utilities and hyperscalers just build faster? Oh, there are so many things there.
Well, first of all, right now, you need humans to build them. There's just the regular construction, but much more than that, you need permits. You need access to power that you can plug into. So you're doing a combination: you've got to get access to power, which is a massive kind of regulatory bidding struggle. There are very limited amounts and things you can tap into in terms of natural gas, power grids, what have you.
But then you also have to build your own power, and guess what? We've got shortages of transformers and turbines and everything that goes into that. So you've got to get all that stuff. This is not a software problem. It's not just like a bunch of engineers can work weekends and that type of stuff. That doesn't work anyway, but there are real bottlenecks in this, and these lead times are not that easy to compress.
And look, we have the best minds in the world trying to figure out how to compress them, but it's not easy. It's not easy, and the demand is not slowing down. We're already behind. The demand is growing 10x a year right now, and the supply just can't grow that fast.
By the way, it is so bad that right now, if new companies are going for GPUs, it's often in Mexico or Australia or another country, just because it is so difficult in the United States. Yeah, we're creating huge job and long-term economic opportunity in other countries by banning data centers here.
I think the right answer would be to set a standard where a data center contributes back to the community: the power gets better, there's no noise, there's no water issue, and it's adding jobs. That ought to be the standard. And then everybody ought to be held to that standard.
11. Why "Machine Age" Is the Right Name
And by the way, there are data centers that do that now. That's not a futuristic dream or something. Their energy rates have gone down every year, and the reason is they provide their own power. They give power to the state during the day, and then at night they borrow power from the state, when the state doesn't need it, because the way power plants work is you're always generating peak capacity. Since a data center has steady capacity during day and night and a city goes way up in the day and way down at night, that's a symbiotic relationship.
Zooming out, we've been batting around the name for a little bit. Why do we think Machine Age is a compelling term for what we're doing here?
Well, listen, let me take it. The first one is, I think Ben's absolutely right: artificial intelligence was the wrong word. We shouldn't have called it that. It's machine intelligence.
Say more about that. Why is that?
Because it's not how humans think, necessarily, right? I mean, it is a cache of how humans thought; it's a collection of human thoughts. But to date, we don't know how to take an AI with no knowledge, put it out in the world, and have it reconstruct language, right? That's not what we've done. We've built something that can learn off of everything we've already learned and then use that in a productive way.
And listen, AI is a general term that goes back 70 years in computer science, formally, that applies to many different things. And of course, it's got a lot of baggage, either from science fiction or from Nick Bostrom, who wrote about it, or whatever.
And so the first one is just an acknowledgment: this really is machine intelligence. And then you want to emphasize the machine part of it. I mean, there's kind of this deep irony—and this is from the “software is eating the world” people—that you've really come to a place where you pour money into something and then you're limited by the actual machines below it. And so I think it is a kind of nod to how the hardware component is so significant in this wave, and we want to acknowledge that.
Yeah. I think that's what's going to create the next breakthroughs in intelligence: the quality of the machines underneath. And so that's basically a reason for the name.
It's also a cool name.
That sounds good.
Futuristic.
12. Won't Incumbents Like Nvidia Take Everything?
Yeah. Given how much has been spent on AI infrastructure to date and how capex-intensive these businesses can be, are we past the point where new companies can break in at material levels? Why not incumbents like NVIDIA, CoreWeave, et cetera, just take the lion's share of these markets?
They're all doing well. There's no question about it. But to our discussion earlier, when you need fundamentally new innovations to keep the growth continuing, or the pace of improvement continuing—whether it's tokens per second per dollar, tokens per watt, tokens per rack, or power—you take any metric. If you want to have a 10x on those metrics, you've got to have new innovation. And new innovation traditionally comes from brilliant founders thinking about solving the problem from first principles in a different way. That's what's needed here for the next jump in innovation.
I mean, this is the law of markets, right? Let's assume that the existing silicon incumbents are multitrillion-dollar companies in terms of market cap, which is absolutely the case.
Even 5% of that is a massive private company. You could say, “Well, NVIDIA could do that.” They could, but why would they if they're focused on things that are in the 90%, which is also driving the same amount of growth? You always ask these questions. We asked these questions during the cloud days: Why wouldn't Amazon do this? You asked these questions during the Microsoft days: Why wouldn't Microsoft do this? There's a very natural law of markets: Once you get to a certain scale, there's tremendous opportunity for innovation at the margins.
Yeah, there's a funny quote from our partner Alex Rampell. He had this startup called TrialPay and was trying to sell its services to Meta—then Facebook. Dan Rose, who was the head of corporate development at the time, said, “Alex, that's great. It sounds like you can collect a lot of silver bricks, but I have so many gold bricks I can't even pick them all up. The last thing I'm doing is looking at a silver brick.” I think NVIDIA is in that position.
100%. We were talking about this as it relates to the model providers: If you're in the sweet spot of what OpenAI and Anthropic can do—one of their main areas of interest—that might be a tough place to be. But anything outside of those, maybe, 3 to 5 areas might be attractive. As markets expand, they fragment, right? It happens all the time. Remember, in the early days of Ford, in 1913, there was the River Rouge plant. Literally, it was like water, coal, and rubber trees went in, and out came cars.
By the way, he bought a whole—
Rubber plantation in the Amazon jungle, right?
And there's a great book called Fordlandia. He wanted to own the complete vertical, so he created this city called Fordlandia in the Amazon jungle. It was all Americanized—bandstands, ice cream, and all this kind of stuff. It actually worked for a while, until he made people show up to things on time, and they were like, “Screw this. Get out of here.”
Yeah, the use cases are multiplying. There's no way, if you're the biggest company, you can get to the biggest use cases, but there are so many use cases—and, as Martin was saying, they're all very valuable—that it's just very hard to get to them in a great way.
Yeah, even inference used to be one simple architecture, and it no longer is. It's so complex now that it's inevitable you can optimize things in a different way.
By the way, here's a very interesting thing: People often don't understand that the margins kind of fell out of the standard way of doing the technology with software. It wasn't really a technology problem. Once you got the business working, you tended to have pretty good margins, because that's just how software works—certainly when you shipped it, but even as a service. That's not necessarily the case with AI. We may actually be entering an era where optimization in the hardware is absolutely meaningful to the upside of the business in a way that we haven't seen in the past. So there's a lot of opportunity here.
Let's get deeper into the types of companies we'll be investing in. Maybe we could start by illustrating the subsectors, or we could talk about a few investments that we've made. I know there are some that haven't been announced yet. Raghu, do you want to take us down?
Yeah, the subsectors, as we've been talking about, span every one of these categories. The obvious ones are computer chips, but these days it's not enough to build a chip; you need to build a full system. Therefore, what goes into the system? There's potentially memory innovation, networking innovation, power chips, and so on.
Each of these categories is one where you can see public-company-style companies emerging, and those are all things we're looking into. Once you put it all together, there's a layer of software around it to automate all of these things, manage these fleets, and so on. That's another important area. These things keep building on each other, but every one of these categories is important.
Talk about what's different about these kinds of companies from the usual company. I mean, one thing you can tell from the companies we announced is that their first rounds have been massive—hundreds of millions. Is it a different kind of founder, or what else is different as we think about the practice of building and investing in these kinds of businesses relative to our traditional software?
Well, I think the big thing is that a lot of money goes in before they get to a product. That's just the nature of it. It's true of big models, too, but I would say that's a little more of a known path, whereas this has a little more risk and a little more money than some of the other things we've done.
A lot of the chip founders are here from the past. You know, the guys who know how to make memory, they're not young. That part is different, too, but it's kind of exciting.
Yeah. The other thing about these founders is that they've all got to be systems founders. What I mean by that is you can't just be a researcher or a great computer scientist. You've got to be able to architect and design the chip or the system, whatever it is. Then you've got to think about how this thing is actually going to get manufactured: Who's going to be supplying this? There are a whole bunch of downstream things which, normally, if you're building software, you don't have to think about.
Really, the best founders—and, of course, Jensen is the Michael Jordan of this—think about the entire ecosystem from the get-go, before they start designing the chip, because of the nature of the bottlenecks and all these things that have to come together. That's a big characteristic that's different.
There are 2 environmental factors that are important, too. The first one is that the labs are so desperate that they will engage with startups. We actually have quite a bit of signal early on because labs are inking deals with companies before they actually have hardware available. That's a big shift from 5 years ago, right? You just didn't go and sell your janky hardware thing to Google or whatever. That's a shift.
13. The Founder Profile: Why Hardware Needs Experience
The second one is that capital availability has loosened up a lot. I think there's general consensus that it is time to reshape this stuff, and so there's a lot of capital available in follow-on rounds. Of course, you want to be investing in areas where capital is available, so the atmospherics are also just different.
Patrick Collison remarked a few years ago, “Hey, it feels like there are fewer young founders today, in the way that Zuck, in college, was building the next Facebook, or Gates was doing the same with Microsoft. And, of course, Michael Dell was building Dell. There are still some young founders building iconic companies, but it does seem, to your point, that there are more older founders building these companies—or fewer 20-year-olds.”
I'm curious whether that resonates and why.
Well, I think it's Raghu's point: If you're building something that has a very complicated supply chain, has to manufacture things, and is technically complicated, some experience helps. If you look at Elon or Travis Kalanick, their companies, when they were young, were software companies. It wasn't until they got a lot of experience—even those guys, the best guys, needed some experience in building companies, building technology, and so forth—that they graduated to much more complicated, or I would say elaborate, domains. There's just much more there; there are many more moving parts in these things.
Look, when you're learning how to build a company, it's hard enough if you completely understand the product. If you don't completely understand the product and have to learn it while you build the company, that's such a steep learning curve for a brand-new entrepreneur. So I think what we're seeing is Michael, on the one hand, who is a very young guy, brilliant, but what he built was kind of a pure-software AI thing.
Yeah. And on the other end, you have an Elon or a Travis, who's got enough experience. I think Michael could probably do that 10 years from now, but today that would have been hard.
It's important to remember that it's been defocused by the entire industry and academia for the last 20 years. There just hasn't been the same opportunity.
Like it’s been there, but it’s never been a growth area. The growth areas have been software, networking, things like that. I also think we have a paucity of people coming out of universities or gaining experience at large companies that have done this. There just aren’t that many.
You don’t go intern and build a chip. But a lot of that’s changing now. Listen, we’re going to create a whole generation of founders who come from these new companies, who will know how to do this, and they’ll be hired much more junior.
I would say, actually, one of the greatest legacies of Elon toward this is, of course, that he’s created these great companies. But the number of entrepreneurs who have come out of SpaceX and are changing the entire industrial complex may be an even greater legacy than the companies themselves. I think we’re going to see the same thing for computer science and hardware.
Yeah, as a matter of fact, all of our investments were started by 2 founders in their 20s. But if you go to one of their offices, you see the experienced people as well. So it’s an ideal combination here.
Yeah.
Yeah. Yeah. It doesn’t necessarily have to be the founder with experience, but that founder better be able to tap into that experience with them.
Yeah. Yeah. Well, they have to be able to work with them, and they have to be good, and all these kinds of things. It’s complicated.
Speaking of experience, this is a big new fund we’re launching, and there are no new GPs. We’re sort of collecting it because you guys have a lot of experience, and the rest of the group has experience in this field, which has been kind of latent and dormant.
Yeah. Well, it’s kind of funny. I think we almost had to be warned against it, almost just because our backgrounds are from hardware. I think the reason that we needed a reminder is because all of us have spent so much of our careers in software and hardware. We’re kind of drawn to that.
So, listen, we’ve clearly invested in hardware over the years, right? We’re in SpaceX, we’re in Anduril—all of these are very early checks. We’re in Astranis, we’re in Waymo, so even early on, we did a number of those investments. But this is so much in our DNA, and I don’t think this necessarily needed to increase the team’s competencies; it was just a bonus.
If this fund does what we think it will do,
How do we see the world changing or looking in 5 to 10 years? Well, hopefully, America wins in the infrastructure game. We have lots of super eco-friendly, efficient data centers out there, and lots and lots of chips, an abundance of memory, and an abundance of power. That would be awesome.
I think it goes back to this: We really think America is a special place, and we’re important not only to everybody here but to anybody in the world who wants to make a contribution and do something bigger than themselves. It’s the best place to come with nothing and do something profound.
We’d like to keep that going, and I think that doesn’t continue if we lose our lead in technology. I think we’ll be in another era, and it’ll be another country, and maybe they have a different set of values around that.