瓶颈不在芯片:与 Bubble boi 对谈|EP 166
Bubble boi 的可投资判断是:AI 的硬约束正从晶体管缩放迁移到封装、互连和内存。 Moore’s Law 曾经大致承诺每2年将晶体管密度翻倍;如今投入数百亿甚至上千亿美元,可能只能换来15–20%的提升。替代方案是扩大芯片面积,让多颗裸片保持一致协同:“无论如何都要获得更多晶体管。”
NVIDIA 的路线图越来越依赖它无法控制的供应商和技术。 H100 使用一颗大裸片,B200 转向2颗,Rubin 默认使用2颗,Ultra 设计原本计划达到4颗;但据称封装困难反复出现,最终规格下调为每个封装2颗裸片、通过2个封装实现。“我每次看 NVIDIA 的路线图,脑子里只剩一句话:今天根本做不到。”
AI 硬件如今必须按机架级系统设计,而不是孤立的加速器。 NVIDIA 的 NVL72 连接72个封装,使用大约4颗 NVLink 交换芯片;据称 TPU 团队则把扩展域拉到9,600颗芯片,并将路由逻辑直接嵌入芯片。“这些系统已经不只是芯片”;架构现在横跨计算、网络、内存、prefill 和 decode。
推理对内存的需求,可能更多由上下文和并发用户数驱动,而不是不断增大的模型权重。 Bubble boi 提到传闻中的 GPT-6“循环 Transformer”、稀疏化、量化、剪枝和投机解码,认为原始权重规模可能趋于平台期,而上下文规模会随权重数量、注意力头数量和用户数增长。他的配置判断是:权重可能应放在 HBM 中,而不需要立即调用的上下文则应越来越多地放进闪存。
高带宽闪存在技术上很有吸引力,但只要 HBM 仍然可得,它就缺少决定性的数据中心应用场景。 它的密度更高,但 Bubble boi 认为 HBF 和高端 AI 闪存的延迟都在约5微秒,HBF 的总产品成本可能也与 HBM 大致相当。他透露的“alpha”是:自己听说的主要 HBF 客户,是那些无法获得或生产高端 HBM 的中国客户。
商业上真正重要的优化,不是单独追求每秒最多 token,而是让单位硬件服务更多用户。 更智能的缓存、闪存卸载、prefill/decode 解耦、CXL 挂载内存,以及感知上下文的控制器,理论上都可能让同一套 NVL72 级别机架服务的用户数翻倍或翻3倍。“归根结底,Agent 就是上下文”,因此 Agent harness 和可复用 prompt 天然适合做缓存对象。
相比估值达万亿美元、模型 alpha 可能因竞争和员工流动而扩散的 AI 实验室,Bubble boi 更偏好硬件“铁路”。 他的耐久性测试是:“如果这家公司消失6个月,所有人会有多惨?” TSMC 和 NVIDIA 得分很高;与此同时,他仍然看多 AI,因为只要借助 AI 找到5种或10种药物,或实现另一次重大效率突破,就可能证明当前全部 AI 资本开支的合理性。
1. Moore’s Law 在工程师江郎才尽之前,先在经济上失效
Bubble boi 解释说,作为芯片设计师,他必须做简化和抽象;他还说,交易员或“赌徒”需要一张高层地图,否则就会“陷进细节泥潭”。Logan 提供的经济学检验是:系统最终必须完成什么,而不是困在比特错误率、材料或几瓦功耗等边际争论里。
Moore’s Law 大致承诺每1.5至2年将晶体管密度翻倍。更多晶体管意味着更多数字逻辑和更高性能,但它更大的贡献在于商业可预测性:一家公司投入1亿美元,就能预判芯片会更好、定价更高、市场份额提升,并获得这笔资本的回报。
Bubble boi 提出的“Bubble’s Law”追踪的是密度提升成本的爆炸式增长。20世纪60年代至90年代,性能提升相对便宜;到了今天,一代先进制程可能吞噬数百亿或上千亿美元,却只带来15–20%的密度提升,最终可能只剩1家或2家公司有能力继续保持竞争力。
时钟频率撞上了另一道由热和开关速度构成的上限,推动行业在21世纪00年代中期转向多核处理器。并行任务因此受益,但串行任务没有:Intel Core i9-14900KS 达到约6GHz,而对数据包处理这类受延迟约束的单线程任务而言,今天的处理器频率并没有实质性超过这一水平。
2. NVIDIA 最初的护城河是灵活性,而不是预见 AI
图形处理本质上是通过高度并行的矩阵运算变换图形;当 CPU 无法以60fps渲染复杂游戏后,专用处理器就不可避免。AI 使用的是相关数学运算,只是精度更低,因此图形架构可以自然迁移到科学计算,再进入模型训练。
Bubble boi 将 NVIDIA 能够存活归因于可编程性。图形 API 发生变化、Microsoft 支持新标准时,据称 NVIDIA 一度接近破产,不得不立即完成替代芯片的流片;公司还在正式图形规范之外增加能力,而不是把自己锁死在标准里。
这种灵活性帮助 NVIDIA 支持了竞争对手错过的科学计算工作负载,也让新模型如今仍能在“1天或2天内”跑起来。不过 Bubble boi 的结论依然有条件:从历史看,可编程性显然是正确答案;但随着 AI 系统日益专业化,同样的优势是否仍然决定胜负,尚无定论。
3. 封装已经成为新的缩放定律
当晶体管缩放放缓后,现实可行的替代方案是扩大面积:把多颗裸片连接得足够紧密,让它们表现得像一颗完整芯片。Bubble boi 说,如今光刻可能只能带来约10–12%的提升,却需要200亿至400亿美元,这使先进封装成为新的缩放范式。
他将不断上升的流片成本曲线与先进封装的相对吸引力进行了对比:5纳米的成本可能约为7纳米的2倍,而2纳米又大致翻倍。
他的 Intel 投资逻辑建立在一个前提上:公司能够在非线性的投资周期里,同时为先进制程和先进封装提供资金。在他看来,Intel 的 EMIB 技术“远胜”当时 TSMC 的方案,而 TSMC 后来已经把封装推进到类似方向。这笔押注的关键,是相信 Intel 真能把产品交付出来。
NVIDIA 暴露了这种依赖。它的路线图从 H100 的一颗大裸片,走到 B200 的2颗;Rubin 默认采用2颗,而访谈中提到 Ultra 原本计划组合4颗裸片。Bubble boi 说,反复降规格如今指向这样的方案:每个封装放2颗裸片,再通过 PCB 连接2个封装。
更长的 PCB 距离会带来吞吐和信号完整性问题,因为信号在不同物理介质之间传输时会损失功率或完整性。他事后认为,真正更纯粹的瓶颈标的是 PCB 厂商,而不是 Intel——Intel 最终做成了。预计默认采用4颗裸片、搭载更多 HBM 的 Feynman,可能会进一步放大这一约束。
4. 扩展架构如今决定芯片本身
AI 的矩阵工作负载“天然易并行”,因此集群可以把任务分配给数百或数千颗加速器。物理限制显而易见:一台 NVL72 大约包含72个封装,不可能全部塞进一个封装,于是系统架构师必须让整个机架维持类似芯片的行为。
训练和推理对互连的要求不同。训练搬运激活值和部分乘积,受益于 NVLink 约900Gbps的带宽;推理搬运的数据通常更少,却对延迟高度敏感。复用面向训练的网络当然可以,但必须针对这种延迟特征做工程优化。
TPU 团队给 Bubble boi 留下深刻印象:在每颗芯片 HBM 更少、也没有类似多光罩封装的情况下,仍然实现了9,600颗芯片的扩展域。TPU 芯片会自行承担部分路由;相比之下,据称 NVIDIA 的 NVL72 要在机架中拿出相当大的空间放置约4颗 NVLink 交换芯片,而这些空间原本可以容纳更多加速器。
更好的互连意味着芯片上的内存可以更少,因为架构师可以改为增加芯片数量。它也支持解耦:Kimi 的参考设计将推理 prefill 与 decode 分开,因为两者本质上是不同操作。“你已经不能只设计一颗芯片了”;机架、工作负载、控制器和硅片必须协同设计。
5. 训练仍然吃算力,推理则暴露出新的杠杆
Bubble boi 仍然认为 NVIDIA 的芯片是出色的训练产品,因为它们提供充足的 FLOPs 和带宽。但如果目标只是获得更多乘加运算,答案越来越接近一块巨大的计算面积——可能是 Cerebras 式设计,而不是某种巧妙的新硬件优化。
训练需要承载权重、激活值、状态、检查点和源数据集。有些内容可以迁移到更便宜的内存,或者之后重新计算;但 Bubble boi 认为剩余的改进越来越依赖算法:“只能希望做算法和机器学习的人能想出办法。”
Logan 通过本地推理做了验证,质疑前沿模型很快就能在消费设备上轻松运行的说法。他从 OpenRouter 抽取的样本显示,排名前10的模型里,即使是最弱的一个,也需要约2至4台 H100 等效设备;而“GM 5.2”需要约40至50台。一台 H100 的价格约为$25,000或$30,000。
Bubble boi 补充说,即使是 Llama 7B,“在实践中”也可能需要4块 GPU 才能运行。旧模型或高度专用化模型仍然可以在本地运行,但讨论中的更大模型仍需要加速器集群,主要因为它们的权重和服务状态超出了本地内存容量。
6. 推理内存的主导因素可能是上下文,而不是权重
Bubble boi 质疑参数规模是否会无限增长。他提到一个流传甚广的说法:GPT-6 会反复循环使用同一组权重;剪枝、量化和投机解码之所以有效,是因为权重具有稀疏性,且贡献并不均等。因此,能力提升未必需要存储参数按比例增加。
上下文规模会随权重数量、注意力头数量和用户数增长。由于自回归生成会反复调用此前的对话来生成下一个 token,服务状态占用的内存可能超过权重本身。Logan 所设想的“Her”式理想——一个记得用户完整人生的助手——会进一步放大这一需求。
Bubble boi 认为,如今宣传的100万或200万 token 上下文窗口并非完全有效:模型必须在这些长度上训练,而合适的长文本训练数据非常稀缺。更有效的记忆方式可能是检索、存储对话,以及通过后训练把“用户住在哪里”这类持久事实写入权重。
prefill/decode 的区分带来了经济回报。Prefill 负责计算上下文,decode 可以复用缓存后的表示。Bubble boi 说,Kimi 没有命中缓存时收费约$5或$3,命中后费用会低得多;Agent harness 尤其适合缓存,因为它们很大程度上是在反复注入相同 prompt。
7. 闪存应该吸收上下文,但 HBF 仍然需要客户
Bubble boi 开玩笑说,改变世界要么靠治愈癌症,要么靠发明“一种全新的内存”,这解释了他为什么被高带宽闪存吸引。HBF 拥有更高的比特密度和可能有用的带宽,但它的成立取决于加速器是否真的需要远高于当前水平的本地容量。
可靠性不是他首要反对理由。闪存写入会造成磨损,但应用专用控制器可以用不同于数据库、游戏或文件的方式管理上下文;一段多年前的旅行对话,未必需要永久、逐比特准确地保存。更难的问题是,面向普通 AI 的闪存已经能够完成这类上下文的检索。
Bubble boi 认为,HBF 和高端传统闪存的延迟都约为5微秒。HBF 的带宽低于 HBM,尽管每比特成本更低,但总成本可能大致相同。因此,它最清晰的客户是那些拿不到 HBM 的人,这与他从“2或3个人”那里听到的中国需求说法相吻合。
NVIDIA 的网络附加闪存方向上更好,但 Bubble boi 认为,更智能的编排还可以继续推进。CXL 同样能让额外内存对软件呈现为普通内存。基本原则很简单:把立即使用的权重留在 HBM,把更冷的上下文卸载出去,再通过为 AI 设计的控制器协调数据移动。
8. 局部性和调度的重要性,可能超过下一个内存热点
Bubble boi 特别提到 Hot Chips 上关于“Jalapeño”的演讲,访谈中将其描述为一款 OpenAI 芯片,因为它利用了内存局部性。不同的内存 bank 和 channel 在物理上相隔数毫米;把所需数据放到相关计算引擎附近,即使这些距离听起来微不足道,也能提升性能。
逻辑层上键合3D DRAM,可以通过大量短连接而不是一条宽总线,同时改善带宽、延迟和功耗。Bubble boi 不记得相关公司的名字。设计挑战转而变成空间问题:工程师必须知道哪些字节位于哪些计算模块附近,并围绕这些极短的本地路径组织工作。
除名义延迟外,时序同样重要。如果一项操作持续10毫秒,而数据直到第11毫秒才会被需要,那么它不必瞬时抵达,只要能够被正确预取即可。只要知道推理工作负载下一步会做什么,内存布局和调度就能替代更昂贵的原始带宽。
Bubble boi 对光子计算的热情低得多。针对当前模型维度做模拟光学矩阵乘法,所需占地可能接近一座橄榄球场,本质上是自由空间光学;因此,他认为近期更有潜力的场景是嵌入式设备或 IoT,而不是数据中心。“那边有人在卖桥。”
9. 除非实验室建立不可替代性,否则真正耐久的投资是基础设施
电源效率很重要,但经济指标是每瓦服务的用户数或 token 数;批量处理更多请求,可能胜过使用奇特的低功耗电路。Bubble boi 还纠正了讨论重点:数据中心最大的成本是硬件,而不是电力。Neocloud 首先要为土地和 GPU 获得融资,便宜的电力并不能抹掉这笔账单。
谈到“暂停前沿研究”时,Bubble boi 质疑实验室说的究竟是安全,还是在向高价值研究纵向整合。他将 Terrence Tao 的研究与细胞技术之间的联系这类有价值的数学和信号处理工作,与影响较小的问题进行了对比。Logan 说,解决 Navier–Stokes 可能只会在 CFD 模拟中节省一些算力周期。
Logan 的现实主张是,强大的 AI 和高算力资源应该用于药物、新材料及其他重要问题,而不是仅仅生成网站或垃圾内容。Bubble boi 认同当前模型能够自动化苦力活,但仍不能替代思考;用户必须核对上下文和来源。
Logan 的反驳值得保留:实验室并没有暂停自己的多年期、数GW级别的承诺。Bubble boi 的说法是:“他们是在替我们控制前沿,不是在替自己控制前沿。”两人都不相信当前模型能够替代思考,但同时也看到了把算力用于艰难发现的价值。
关于开源与闭源,Bubble boi 说,开源已经让编程商品化,前沿应该超越计算机任务,转向能够改变现实的发现。他还说,除非具备大量专业知识,否则当前模型通常无法在商品硬件上运行,旧模型或专用工作负载除外。
Bubble boi 更愿意持有“硬件层”、铁路或已经被验证的运营商,而不是可能“竞争到死”的实验室。他的消失测试偏向 TSMC、NVIDIA 和不可或缺的基础设施。讨论中提到的潜在标的包括尚未上市的 Etched,以及 WKA 的闪存卸载软件;Groq 和 Cerebras 出现在相关讨论中,但不是他的公开推荐。
对于 Agent,Bubble boi 认为,变现可能更多流向采用它们的公司或被它们替代的员工,甚至可能催生保险业务,而不一定只归 Agent 厂商所有。
他的宏观结论是看多:如果借助 AI 找到5种或10种药物,或实现另一次重大效率突破,就可能证明当前全部 AI 资本开支的合理性,并触发一次“快速起飞”。
完整逐字稿
We will add more memory to these servers at the end of the day. The question is, what kind of memory? Will it be HBM? Will it be DRAM? Will it be flash? My guess is flash, more so.
People are all talking about ways to de-spec the amount of HBM you need because, honestly, at the end of the day, weights should probably be in HBM. Why should we have huge amounts of context there if we're not using it all at once?
I'm a big believer in just owning the hardware layer, or owning the railroads, or owning the things that go into it, right? Do you want to invest in Anthropic and OpenAI, who are going to compete to death and eventually see their alpha diffuse as people hop between jobs and stuff like that? Or do you want to own the people who have 10- or 20-year track records and moats in developing the integral hardware that goes into this?
They're pacing the frontier for us; they're not pacing it for themselves. Yeah.
But I'm leaning toward being very bullish. We have all this situation going on with interest rates and Iran, but I don't know. I have a feeling that we find a couple of drugs with AI and get the right researchers in the room, or we figure out a way to make 5G not need a thousand towers in a square mile. I mean, yeah.
Starlink. Yeah, whatever. Anything that's more efficient. We will have a fast takeoff. You can justify all the capex off of just 5 or 10 drugs, I think.
Mr. Bubble boi, thank you so much for coming on the podcast. This is my first attempt at a little bit more of an AI-focused podcast instead of my traditional crypto-focused podcast, so I appreciate you coming on. I very much appreciate all your takes on Twitter and X, and as I was saying before we hit the recording button, more than anything, I think you have an exceptional gift for explaining technical concepts very simply. That's hard to do.
I'm looking forward to the conversation and really breaking everything down, from what you see today to where we think things are going in the future. So thank you.
Well, thank you for having me. I would say my quote-unquote gift for simplifying things— a lot of people hate it—but I'm, first off, a chip designer, so I have to simplify and abstract things away. Otherwise, I would lose my mind trying to understand everything.
Also, if you're a trader or gambler, as I like to call it, you really need to have a high-level overview. Sometimes people get lost in the sauce, right? I agree with the detail. You have to know what matters at the end of the day, both from the chip side and the economics.
Because I've seen people get into fights about the bit error rate of SerDes, or whether some kind of material is better for power by, I don't know, a couple of watts. I'm just like, guys, come on. What does this thing need to do at the end of the day? That's its economic value to the world. A little point here and there isn't going to be a big deal.
I fully agree. We've had a lot of these nerdy debates in crypto, and generally, as Elon likes to put it, “Simplify, simplify, delete.” The simpler the architecture and design, generally, the more adoption it got. I'm sure there are many parallels within the AI space.
I've always—I mean, not to go on the worst tangent at the beginning of the podcast—used to be a big fan of philosophy. I really only read the ancient stuff, up to the Stoics. They had this whole view of the essence of something. That's Plato's forms: What's the essence? What's the core thing? From that, you can derive everything else.
That's kind of how I work when I'm approaching some of these topics: What's the main thing? From there, we can talk about details if they matter or not. I think that's a great starting point.
I wanted to go from zero to one for the audience and listeners because, again, you do an exceptional job of explaining these concepts. Maybe just starting with something basic, like Moore's law, and why Moore's law ultimately came to an end. We can start from there and then jump into some more interesting topics.
You have to understand what Moore's law is talking about before you approach it, which is transistor density, roughly. It basically says that in, I don't know, about 2 years—1.5 years; the number doesn't really matter—you roughly get double the transistor density. For a unit of area, you have more transistors.
Why does that matter? Transistors are how you build digital logic. When we're building chips, we write programs that get converted into digital logic, which gets converted into a bunch of transistor arrays. The more we can fit, and the smaller we can make the transistors, the more work the computer can do.
Think of it as writing a program and it not fitting in memory. That's a gross oversimplification, I'm sure, but that's the idea. We have this operation, pipeline, or tool, but we can't actually fit it.
It was very important, first, from the perspective that it basically doubled performance every 2 or so years. You're getting more and more transistors, and with that you can do more work. But it was also important on the economic side because you could justify your capex. You could actually predict, “Hey, if I spend $100 million, I will get a return on my investment.” These chips are going to be this much better, and if they're this much better, I can charge premiums and capture market share.
It's been pretty important from that perspective. From the engineering side, it's cool and important, but from the business side, I think it's super important.
What people haven't realized, though, is an offshoot of Moore's law. I don't even know if it has a law, but we can call it Bubble's law—I'm just stealing things. How much money did it take to increase transistor density at every process node?
Early on, it wasn't that much. You didn't have to spend billions of dollars to double the transistor density. Now you do. Maybe it takes tens to hundreds of billions of dollars to not even double the number of transistors. Maybe you get 15% to 20%, if that, better transistor density between nodes.
1. 0 to 1 of the Chip
The reason I've outlined the economic perspective is that it's kind of breaking. The amount of money needed to push these nodes has ballooned to the point where really only 1 or 2 companies can afford to invest in it. If you don't invest, you fall behind. It's a perpetual cycle: You have a worse node, people don't want to use you, you get worse customers, and when you don't have customers, you can't get the money or time to iterate and make your nodes better.
I think this is the key thing for Moore's law. One thing people have overlooked is just the money—the cost of doubling these transistors. It was relatively free from the 1980s to the 1990s, or from the 1960s to the 1990s. In the early 2000s, we had to start getting a little smarter about it.
At what point—I believe it was in the mid-1990s, or maybe later in the 2000s—did we start switching over from single-core performance to multicore? How did that shift the dynamic of things?
I think it was in the mid-2000s. Maybe it was the Pentium D, the first dual-core processor. I don't have the exact name, but it was in the mid-2000s.
Part of the reason was: What does a core do at the end of the day? It processes instructions. How fast can you process instructions? It's how fast your core can go, which is just a function of your clock speed.
Clock speeds have topped out. That's not related to Moore's law; that's more about thermals and how fast you can go, or how fast the transistors can switch. Clock speeds topped out, and what the world realized was, what if some of these instructions can be done in parallel? Not every instruction is dependent on another.
Graphics are the most obvious example. We call them embarrassingly parallel workloads. Matrix math is embarrassingly parallel. This is where multicore systems started becoming a thing because it was a way to drive performance for certain kinds of workloads, but for others, you can't.
To this day, if you have a single-threaded workload or something that's bottlenecked instruction by instruction, honestly, if you're a really good programmer, you can probably take an old computer and beat it, as long as it still has a pretty high clock speed.
Think of the Intel Core i9-14900KS. I think this was the highest-clock-speed processor ever to drop for the public. It topped out at around 6 gigahertz. We actually don't have a single-core processor that goes faster than that today.
So—
It's not that important for most people, but if you have a really important single-threaded workload, think of something like network-packet processing, where you get a byte after a byte, that's where clock speeds matter.
Ultimately, as you mentioned, that gave way to the GPU. Nvidia ran that for a while, and as you pointed out with matrix multiplication, it mapped fairly well to AI.
Can you talk about how that started to scale up on the AI side, I guess, for larger pre-training runs? So maybe we just start with why Nvidia exists and why they got good at it?
I mean, graphics at the end of the day—it’s a joke, but it’s shape rotation. It’s triangles and squares, and you do a matrix multiplication to shift them and turn them around. This is approximately the same math that we would use for AI, right? But the benefit of AI is that you don’t need as much precision, whereas in graphics you do.
If you’re thinking of physics or speed, having a 64-bit floating-point number, or having more bits to represent the fraction, is pretty important for those workloads. Nvidia wasn’t the only graphics chip company at the time, but it was very obvious that graphics were going to need a custom processor, right? We were already in the CPU domain, where we had gotten to dual cores, but that was nowhere close to what you need to run a 60 FPS full game, right? So you had to offload that graphics processing to a new accelerator.
Nvidia really did well. I think the interesting question is why Nvidia really dominated long term, because there were a lot of other companies in the space. My answer has been that Nvidia designed its GPUs from the start to have a lot of flexibility and programmability.
There’s that famous story where Jensen says Nvidia almost went bankrupt because they changed the graphics API—or the protocol for rendering graphics—and Nvidia didn’t support it. So they had to tape out a whole new chip immediately to support that, or else they were going to get cooked, because Microsoft was backing it and Microsoft was the big behemoth. Even from its earliest days, Nvidia was doing things like parallelizing I/O, like virtualized I/O for the GPU.
Nvidia put a strong emphasis on programmability from the start because it was the Wild West, and it paid off long term, right? A lot of people have died who decided to go rigid, but Nvidia, even to this day, has a pretty decent moat. It has a lot of programmability. New models come out, and Nvidia seems to run them within a day or 2, right? So I think that was the thesis of Nvidia’s strength there. And, yeah, go ahead.
And by rigid, do you just mean custom accelerators?
Yeah. So, custom accelerators, but Nvidia was adding stuff to its GPU that wasn’t even in the spec. If the spec changes, if we wanted to do this, can we have the accelerator offload this? This is why they quickly found their way into scientific computing, because they made that protocol very useful, right? There’s a reason no other graphics chip got into scientific computing, and that’s where they started before they got into AI, right?
Good call on Nvidia for trying to make these massive, complicated, parallelizable processors more programmable from the beginning. In hindsight, it was the right call, right? Will it be the right call for AI now? We don’t know. Things change, but for now, it’s obviously led to a lot of success for them.
So maybe pushing that a little bit further, you’ve started to—I guess maybe backing up—you’re kind of famous for your Intel call and really talking about packaging, as packaging is the new Moore’s law. Can you break down why you think packaging is the new Moore’s law, and why packaging is much more interesting than the lithography machines?
Yeah. So, kind of to what I said at the start, if we can shrink transistors over and over and double them every 2 years, packaging maybe isn’t that important. But we need more transistors by any means necessary, right?
Because lithography scaling has just topped out in recent years, or isn’t scaling at the same rate—maybe we’re getting 10% or 12%—and we’re spending, what, $20 billion or $40 billion to get there, maybe it’s smarter to just add more of these chips and connect them so they act like 1 chip. So it’s increasing area, right? Instead of trying to shrink transistors into 1 unit of area, we’re expanding the area, right?
It was very clear to me, just because if you know anything about lithography, you know it costs so much money to tape out on leading-edge nodes—orders of magnitude more than a 7-nanometer node. The 5-nanometer node is probably double the cost to tape out, and then the 2-nanometer node is double that, right? So the cost just balloons. A better way to get performance is really just increasing area, which was a new scaling paradigm.
Everyone’s talking about scaling paradigms now, but this is a new one, right? Packaging is how people are trying to scale transistors now: give you more transistors within a coherent chip. When I was really—and I still am really—into Intel, it was mainly because I understood the lithography nodes that they were trying to advance.
They were trying to leapfrog TSMC by adopting things like High-NA much earlier and pushing packaging technology, and that takes a lot of investment, right? As I said, the investment has gone nonlinear, exponential for every node, right? So it becomes kind of a monopoly or duopoly situation. The only way that they could ever catch up to TSMC was by burning that money. But I had high conviction that it was a pretty good strategy if you’re confident you can ship.
The other thing where I was maybe earlier than others to identify was that their packaging technology, called EIB, was just miles better than what TSMC was offering. Now TSMC is adopting what EMIB already has, right? They’re trying to shift their packaging technology in that direction.
2. Packaging as the New Moore’s Law
This was really important at the time because if you look at Nvidia’s roadmap, what is it leveraged to? It’s actually not leveraged to anything Nvidia controls. It’s leveraged to memory, and it’s leveraged to packaging—both things that Nvidia doesn’t actually do as a business, right? They work with people and try to influence it, but at the end of the day, those are vendors and partners who give them that. They do not fab those, or design them, in-house, right?
Nvidia’s roadmap was going from H100s with 1 huge die to B200s with 2 dies. Then Reuben is 2 dies by default, and Rubin Ultra is supposed to be 4 dies. Voilà—look what’s happened. Reuben Ultra cannot be 4 dies in 1 package; they’ve had to despec multiple times.
Now, the rumor is that it’s 2 dies in 1 package and then 2 packages, right? So they’re connected over some kind of PCB link there. To anyone who knew where packaging was going, this was obvious. I played this by going long Intel, which was great. I’m very happy that Intel’s doing better now.
But the real play was probably to be going long the PCB makers, because now they’ve become the bottleneck for Nvidia to actually get these 2 chips, which are separated by maybe a couple of inches or centimeters, to act like 1 chip. That’s not necessarily easy, or how things have been done up to now. The main bottleneck there was increasing the throughput between the chips to ensure that they can speak to each other, almost acting as 1 chip.
Yes. Throughput and also signal integrity. As you travel longer distances, the signal loses its power or its integrity. We have power lines, right? We used to have repeaters on telephone lines to amplify the signal and repeat it. It’s the same physics at the end of the day, except now you’re traveling centimeters on a PCB.
So the signal now gets translated into some other format, some other physical medium. It now needs to move back and forth, which isn’t easy to do at high speeds and with good signal quality. So Rubin Ultra hasn’t shipped. I think Rubin is ramping up now, and we’ll see if Reuben Ultra ships next year.
But they’ve had quite a bit of trouble. Here’s the thing: we’re looking at Reuben Ultra right now because it’s in the news, but Fman's is supposed to be 4 dies by default, right? Feynman is supposed to have even more HBM, right? All these things that intuitively we knew were not possible. Every time I look at Nvidia’s roadmap, I just see, “This is impossible today. This is impossible today. This is impossible today.”
So I’m curious how they plan to solve it. People just want to throw up their hands and say, “Well, it’s Nvidia, man. Jensen can do anything.” I’m like, “Well, I’m sure he’s a great operator and he can do a lot of things, but he’s beholden to what SK, Samsung, and TSMC can do—and maybe what Intel can do in the future, right?”
Yeah, I think that’s a very clear zero-to-one on both transistor scaling, clock speed, multicore, and even scaling the dies. I’m curious, though, now that we’ve obviously gotten to much larger clusters. We went from maybe hundreds, and now we’re at thousands to hundreds of thousands of interconnected GPUs.
The important thing is scaling up these data centers for large pre-training. It’s interesting that you have to have relatively high throughput not only on the chip but also within the rack and across racks. Can you talk about how things are scaling up as we’re scaling to larger clusters on the training side?
Yeah.
So I think, first, to preface this: why scale up? One, it's physically impossible to push packaging to the limit by putting all 72 packages of something like the NVL72 into a single package. That's not really feasible. But also, these AI algorithms are very parallelizable, right? The benefit is that you can have one chip work on one part of the problem, another chip work on another part of the problem, and get higher throughput that way. That's the benefit.
If we were working in something like a more single-threaded workload, scale-up wouldn't even be necessary, right? It wouldn't give you anything. But AI and this matrix math are very parallelizable, increasingly so. So the scale-up domain is becoming the real bottleneck. I think the latency and throughput are extremely important, but more so the latency side, just because of the way that people parallelize, whether they're doing pipeline parallelism or tensor parallelism.
What you're actually moving between chips isn't a huge amount of bandwidth. Systems today are designed for a huge amount of bandwidth, and NVLink—I think it's about 900 gigabit per second—is more relevant for training workloads, where you shove back and forth a bunch of activations and partial products between GPUs to train. So depending on the kind of cluster you're building, whether it's an inference or training cluster, the interconnect needs to be designed for what you're going to do. If you try to use a training-style interconnect for inference, it'll work, but you have to engineer around that huge latency bottleneck.
I will say, though, that a lot of other companies do the scale-up domain very differently. I was at Hot Chips recently, and the TPU team, which I think is doing really good work and is probably one of the best hardware teams in the world right now, has a scale-up domain of 9,600 chips. It's an insane number of chips, right? A byproduct of that is that they have less HBM on their chips than NVIDIA has. Also, at least to the best of my knowledge, they're not currently doing these multi-chip, multi-reticle-size chips.
They're good at scaling the interconnect. If you're very good at doing that and shifting traffic in between, because at 9,600 chips that's a huge number of addresses that you could travel to, how you switch and interconnect between the chips is actually going to determine the latency, right? The NVL72 system has 4 NVLink switches in there, right? They take up rack space that could have been used by GPUs, but the TPU team doesn't do that. They actually built some of the routing logic inside the chips themselves, so the chips act like pseudorouters as well. So it's tricks like that.
I will say, though, that the scale-up domain has become increasingly important for chip design as well because, as you said, these systems are not just chips anymore. They're not just 1 or 2 chips. They're not 4 chips, and they're not just a server. They are a whole rack or potentially multiple racks. This is kind of new, I think, in computer architecture. We were taught to look at cache hits and memory access times, but when you're orchestrating across chips, you have to analyze the whole system, right?
Based on your scale-up domain, you can change your chip design. So if you have a very low-latency, high-bandwidth scale-up domain, and not that much switching in between, you could probably get away with de-speccing your memory because you could just throw more chips at the problem instead of more memory at the problem. Chip design is a balancing act. I think a lot of engineering is, but chip design really is a balancing act. It's all about where I can cut here to make gains elsewhere, and it's all about balancing things.
It's super interesting because, theoretically, as you mentioned, if within a single server you have, say, 50 terabytes of throughput and you had a 50-terabyte line or interconnect, you could do some very interesting things where you build custom racks of maybe memory or just GPUs. But you're bottlenecked because you don't have the 50-terabyte interconnects, which is the main thing, plus the latency aspect that you mentioned.
Yeah. And you touch on something else, which is disaggregation, right? These models or techniques for either inference or training are realizing that certain parts of the pipeline, or certain parts of the operation, should be running on different silicon or different kernels, with different utilization or different programs running on them. That's actually more efficient.
3. Racks, Trays, and the Unit of Compute
The Kimi reference design has a separate node for the prefill step of inference versus the decode step. This is becoming ubiquitous at this point. This is the standard way to serve inference because it's so much better; they're completely different operations, really, at the end of the day. So the scale-up domain is becoming even more important nowadays.
I think the future of AI training, but really even inference, is designing the whole system, the whole rack. You cannot just design a chip anymore, and you can't design a chip in a vacuum. That's fine for a local machine or a small scale, but for the scales we're running things at, we have to optimize at the system level. And I think there's going to be more disaggregation. We're talking prefill and decode; there can be even more, I think.
So I definitely want to talk about prefill and decode, but maybe if we stick to just the training and pre-training side, how do you see things as we're both scaling up the packaging and scaling up, I'd say, the racks, as data centers go from 1 gigawatt to multigigawatts? Do you see any additional bottlenecks on the training side that interest you, or is it mostly the packaging and the interconnects?
Yeah, I think training is—I will say, NVIDIA's chips are really good for training, right? If you're doing a lot of high-FLOP workloads, they're good for that because they deliver a lot of FLOPs and have a lot of bandwidth. Now, again, these aren't things that can't be solved by other means, but on the training side, there are fewer and fewer hardware optimizations you can do.
You can do things like offloading to flash or saving computation for later so you don't have to recompute things, but these are very algorithmic things, right? So on the training side, it's become harder to push FLOPs. If you really just want a crazy, FLOP-heavy chip, you need a huge number of multiply-accumulate operations and a huge chip, really, at the end of the day—just huge area, almost like a Cerebras-style thing, where you can do a lot of compute.
Wasn't that what they were trying to do with their internal project, Dojo?
I'm not familiar, but—
They shut it down, so I don't think it was successful.
It wouldn't surprise me. I mean, the thing is, for training, you have a lot of data, whereas in inference you have people's queries, people's context, and the weights. Really, that's kind of it. In training, you have all kinds of tiers of activations, states, and the weights themselves. You're going to checkpoint those as you go because, hey, I left off my workload and I want to pick up where I last started.
You also have the data that you're training on. These are huge data problems, right? So I think offloading things to cheaper kinds of memory, knowing when to access them, and saving time for compute is probably something you could do in training. But there's way less you can do, I think. On the training side, we'll have to hope that the people who do the algorithms and machine learning can figure something out. It might be above my pay grade on the hardware side.
Makes sense. Cool. We can shift to inference, then. One interesting experiment that I ran was looking at the broad bear case within AI: that everything will just be run locally. You'll run it on your MacBook Pro or iPhone eventually.
I looked at OpenRouter and pulled the top 10 models to see how many network chips you need, using an H100 equivalent and equivalent high-bandwidth memory, to be able to run each model. Even for the shittiest model in the top 10, I think it was about 2 to 4. I believe each H100 is still around $25,000 or $30,000.
More popular models, like GM 5.2, which is a multi-trillion-parameter model, require something like 40 to 50 H100s. More importantly, there's the memory side, just to hold the weights. My broad question to you is: now that these models are getting larger—and it seems like they're going to continue to increase in size—how does that ultimately affect memory, in terms of whether it stays on something like SRAMM, high-bandwidth memory, or flash? It seems like memory capacity, not just throughput, is expanding.
Yeah, I mean, there's a lot in that question. If we take your numbers, they're roughly right. I think even for llama 7B, you need 4 GPUs in practice to really run it, right?
I think your question on memory is very interesting because what you're holding in memory is really important, right? If you're holding weights, yeah, weights are scaling, but there's an argument that things are plateauing to an extent there, or that people aren't pushing scaling on weights as much. It's common knowledge now—or common rumor—that the new GPT-6 uses a technique where they loop over the same weights over and over. They call them looped transformers, right?
If scaling with more and more weights would work, we would just need more and more memory for weights, at least, right? A lot of the research has shown that these weights are empty. They're very sparse. There's not tons of information in every bit of the weights, and that's why things like quantization and pruning work very well. Speculative decoding works very well as well, because not every weight contributes the same amount.
But when we look at other things, like what's taking up a lot of the space when you're serving, it's actually context. It's your conversation or your previous conversation, just because of the way these models produce their next token: by going through your whole context autoregressively to produce literally 1 token at a time. People use speculative decoding a lot to produce more, but that's more or less what we're doing, right?
When you look at it, context scales with the number of weights, the number of heads, and just the number of users. So there's an argument to be made that context is actually more important than weights for inference. It may be taking up most of the size, depending on the model, right?
Yeah. The context is super interesting. I always think about what, in the future, if we could wave a magic wand, we would want. It's kind of like the movie Her, where it just knows everything about you. You talk with your computer, hopefully in the good sense.
Context windows went from 128 to 256 to 512. Now they're around 1 million. I'm curious: do you think they're going to keep growing? If you could have larger and larger context windows, maybe not bound by—I mean, everything's bound by hardware—but where do you see context windows going? Do you think they grow from 1 million to 10 million, then 10 million to 100 million?
I think it's really hard to scale just because of the hardware limitations. Even what we call 1 million or 2 million right now isn't a real 1 million or 2 million. You have to train at 1 million or 2 million, and it's actually pretty hard to find data at 1 million or 2 million, right? But regardless, I do think more context is coming.
These models will have more information about you. The question is how they get it. Maybe it's a cached conversation from many years ago that's stored, and then they search it, grab it, and keep it in their 1 million or 2 million. Maybe it's post-training on your context, right?
If something's that important to who someone is, you're probably just going to burn it into your neurons, right? You're not going to have to search your documents and remember what that was. Maybe if it's something small that you said, but if you live in Miami or Florida, that should probably just be burned into the weights, right?
I think we get more effective context through post-training and burning important things into the weights. We probably also just store a lot more data on people. Think of companies like Meta and Google: they store so much information on us, right? These models will eventually do that too.
4. Memory Hierarchy: HBM, DRAM, and the Next Tier
The big benefit with context is that you generate it by computing it, which is prefill, but during decode, you just have to grab it. You just have to fetch it and get the next token. So a big benefit is that if you already have the context, you can avoid doing prefill and just start decode mode. It's a lot less utilization on the GPU or the hardware, and it uses a lot less power.
You can see this with a lot of models. They charge differently for a cache hit. I think Kimmy charges around $5 or $3 if you don't have a cache hit, but if you do have a cache hit, they charge much less. This is part of it. I think the context stuff is going to be something where it may not even be in the actual hardware. It's going to be in the whole system design, right?
Yeah.
So NVIDIA has its answer for this: network-attached flash drives. They have a flash-only rack, and that will be their way of fetching context on Logan or fetching context on Bubble, right? Logan said something about cars. Let's go search for the car conversation and grab it real quick. Stuff like that.
For agentic stuff, it's become very important now because agents are just context at the end of the day. An agent harness is really just a bunch of prompt injections.
I was amazed when I found this out.
But, yeah, every time you had to redo that prompt, it would eat up GPUs, and your time to first token would take forever, right? So the big benefit of agents is that things like LLM cache are built around caching your agentic harness and reusing it all the time, right? Something that simple.
I came to the same conclusion on the agent side: you just want more and more context. I was thinking about how to maximize the length of agents, and it seems like the industry—and I'm curious about your thoughts—has been very interested in just maximizing throughput on the decode side, with things like Groq, and with SRAMM having very high throughput but relatively limited capacity.
It's harder to store everything in SRAM and maximize throughput, but when you drop to something like high-bandwidth memory, I think it's now around 8 terabytes, and you jump to 100 to 300 gigabytes of capacity. It's still not generally enough if you want to go much faster.
I think you've been a proponent of high-bandwidth flash. I'm curious about your updates there—being able to keep throughput relatively high but have much, much more context. As you mentioned with NVIDIA and CMX, I think the interesting thing to me, and maybe the question to you, is: do you think the industry has overoptimized for throughput? If agents are coming about and they matter more for storing context about your life and what you want them to do, does that change the balance?
The industry has definitely put a focus on throughput, and it's because they can get more money for higher token throughput. You kind of hit it on the head: if I have to constantly fetch context, and context is a big part of the memory footprint on these GPUs, well, yeah, you could probably get better throughput by just handling your context much smarter.
I think HBF was very interesting to me because, honestly, I always said as a joke, "If you want to change the world, find a cure for cancer or make a new kind of memory." There haven't been many people trying to make a new kind of memory. Many people are trying on the cancer front. I hope we get that first. But a new kind of memory is always to be commended, right?
The nerd in me always wanted a new kind of memory. When you look at it, flash has better bit density, right? We can store way more there, and if we can get even competitive bandwidth out of it, that's great. We should have more density on these accelerators for less cost per bit—less cost for the total size, or gigabytes, right?
I think the problem I've seen with high-bandwidth flash, and many have said this, is that it's a reliability problem. I think the reliability issues can actually be solved. You have to be a really good designer of your flash controller to do that, and you have to design your flash controller for the problem at hand.
The way we design flash controllers today is to hold a database, hold files, or hold your computer games, right? Context is very different. Context is fetched in certain kinds of ways. Context isn't always needed forever, right? Will I always need to store, in perfect bit resolution, your conversation about your vacation from 2 years ago? Maybe not, right? So maybe we're okay with losing a little bit of that.
The thing with flash is that writes to flash have always been what harms its reliability. It's persistent storage, but to an extent, depending on how much you write to it. I think the issue with HBF has been, "What's the market for it?" If you have HBM, why use HBF? Do we really get more benefit out of more gigabytes on an accelerator of storage? I don't think so. Because if so, why is Google doing a 9,600-scale-up domain with less memory?
I think the thing for these decoders really has just been bandwidth, right? Density doesn't actually buy you much. It may buy you the ability to hold more context, but honestly, you could probably just hold that context in regular flash and grab it when you need it if you're very smart about how you move things around.
So I think HBF has just been this thing where it's very interesting. There is a case to be made that it's useful.
Will it be useful for offloading contexts?
I think you can just use regular flash for that. I don't think you get much of a performance benefit. The latency of HBF and the latency of even high-end regular flash that's designed for AI are both 5 microseconds, right? So, we don't get much of a benefit. The bandwidth is less than HBM. The cost is less per bit, but it's arguably about the same as HBM in total, right?
The kind of customer for HBF is people who can't get HBM, in my view, right? That's kind of where I see it at the end of the day. Maybe there's a case to be made that you can use it for local inference, but in a data center product, it's hard to say right now. I'll give an alpha leak just because you had me on the podcast. I've heard from 2 or 3 people at this point that the main customers for HBF are Chinese, and they're very interested in this.
It makes a lot of sense given everything I just said, because if you don't have access to high-end HBM and you can't make it, you actually can't make as many accelerators, right? So, you might be willing to get a substandard product if it's as good as an H100 or a B200, right? It's not going to be on the leading edge, and if you can get it a bit cheaper, there's a use case for it.
I think HBF is something where it's very early right now. We have to figure out how to standardize it, but the use case is also very up in the air at the moment, in my view. I do believe that the whole rack needs to be co-designed for the project, down to the memory controller and down to the flash memory controller. Everything needs to be designed for the AI application, right?
So, you're more bullish on NVIDIA's approach with CMX, just offloading more and more things to flash and having terabytes and terabytes of flash space?
I wouldn't say I'm bullish on that necessarily. That's maybe better than HBF, right? I agree, and NVIDIA thinks so as well. A lot of people in the industry think so. I just think we can do even better than that, right?
I think what we have to do is very smart orchestration across everything, and that goes from literally baking in, sometimes, the algorithms used for compressing and storing contexts, right?
So, do you feel like over time we may get from—call it, I know it's not 1 million, but just call it 1 million today—to 5 million, to 10 million, to 100 million, and that will mostly be driven by algorithmic improvements rather than increases on the hardware side?
There's an argument to be made that inference starts looking a lot more like training, and we are using test-time compute to train on your context in real time. In that scenario, do we need as much flash? Arguably not. We don't need to store every image and every piece of information, but it's really up in the air what the right architecture is.
We're on the frontier. Everyone wants to pace the frontier as of late, but we are on the frontier, so we don't necessarily know what the best technique is. My view is that you will be offloading to flash, but it will not be everything. It will not be every document you write.
For big corporations, there's a case to be made: If I have this big database, why don't I just feed it to my LLM and, voilà, I have context? I just chat with my database, right? There could be a case for that, right?
So, do you feel like the industry will—I mean, in a broad sense—continue to prioritize throughput, given that that's the main way they can charge revenue on the inference side, versus maybe more specialized plans that allow you to have much larger context windows but slightly less throughput?
I think the industry right now is definitely focusing on throughput. But how do companies make money? They make money by serving more users. If you could serve more users with less hardware, that will make way more money, right, than just serving context windows.
Imagine you just have your NVL72 rack, and we've doubled or tripled the number of users you can serve with some flash offload or some very smart algorithmic changes. That's more valuable to the AI labs. But the product of a super-fast model, or a special model that's very fast to decode, I think that'll be around for a while. I'm surprised that a lot of people like that kind of stuff. Even the AI researchers really like that kind of stuff.
5. Scaling Up vs Scaling Out
But I think in the future, the value at the end of the day for the AI lab is this: If they said, “Should we just put everything on Cerebras and charge more?” versus, “Can we just take our hardware and triple the number of users?” I think they will pick tripling the number of users on the hardware, right? If you can serve more people with the same amount of hardware, the costs go down. Every person's cost per token will go down.
I don't know if that answers your question, but I think new capabilities are going to be an interesting way to make money, right? I don't know if you and I and the little people will get it, but maybe the enterprises will get it, right? It would be great if you could just post-train on my database and have it built into my LLM. I mean, this is what Thinking Machines is doing, basically, right?
They're having longer context.
Well, they're post-training on people's data to get better results on it. In a sense, that's context, right?
I saw a podcast with Durio—I can't remember which one—but he spoke about how, over time, he thinks context windows will continue to expand. As you mentioned, you'll do some type of either reinforcement learning or post-training on the context itself and then bake those back into the models, which, again, got me more bullish on context.
So, I guess what I'm hearing from you—and maybe for all the people who are listening—is what's the best way to express that? Is it, on the private side, long Cerebras, long Groq, but they got bought? Versus, in the public markets, maybe going long something like Etched?
I think going long something like Etched would be a really good investment if you could do it, but they're not public yet. On the public side, it's really hard to find a clear investment. The beauty of AI is that a lot of these companies are new, and they haven't gone public yet.
There are interesting companies in the space. WKA is one company where they do the software for flash offload, or just storing people's data across all kinds of clusters, right? So, you could try to express it on the commodity side—quote-unquote commodity—or let's just call it the hardware side. I think there will be interesting plays in the future in the actual software stack or the controllers for the flash, right?
I don't know of anyone yet who has more FLOPS than NVIDIA publicly. Maybe that will change when someone comes to market with something interesting, right? But that's one way you can play it. If you think that context is going to get burned into the weights, we've got to do a lot of post-training or test-time compute, right? So, anyone who's delivering FLOPS at a cheap cost is going to become extremely valuable.
Interesting. Yeah, I think to me it seems obvious that, over the next couple of years, agents are going to do more and more just for everyday work—computer work. And for the everyday person, how are you able to express that, given that most people don't have access to the private markets? The lab companies are already, call it, $1 trillion-plus.
Yeah, it's interesting. Maybe the best way to monetize agents coming to market is through the people they replace—the companies that use the agents. I honestly think that might actually work, like insurance, in a very counterintuitive way of looking at it. We don't have an insurance company that's touting AI agents yet, but I think that will probably be the best way to monetize them, right?
I think agents will just become like the internet or cloud computing. If you don't have an agentic use case, you're not even in the game. Everyone's going to have that, right? So, maybe the argument is that the people who are adopting AI very strongly—which is really big tech—are also the ones that are serving it.
Yeah, that makes sense. So, over the long term, you mentioned Weka, which is also super interesting. I was deep-diving into them, and they were doing a lot of cool stuff on parallelizing flash to increase throughput, which I thought was cool. On memory, are you mostly interested in increased throughput, or are you also interested in things like CXL or other techniques?
Yeah. So, on the decode side, of course, increasing throughput means increasing memory bandwidth, right? Any way you can increase memory bandwidth is extremely interesting. There are very few ways left. People are even arguing about just putting in the bare minimum of memory bandwidth.
The reason I've always liked CXL is that it allows memory to be added or to hang off and look just like memory on your normal computer, right? At least it's hidden from the programmer as far as the programming model goes. Why is that useful? Well, as we've been hinting at, you can offload things to memory. vLLM offloads context to RAM, right?
If we can build a directly attached way of doing that, that just increases the value of the cluster, right? More or less, what I'm really trying to say is: same hardware, add new kinds of memory, and write better code. You will get more value out of your hardware.
Same accelerators, at least.
Interesting. One maybe far-fetched idea that I'm curious to get your thoughts on is that most people talk about photonics in the sense of scale-out. How are you thinking about—or have you thought about—photonics from the computing side? The matrix multiplications get very interesting when you can parallelize them within a given channel.
I should do a post about this where I work out the math on a whiteboard for everyone. If you look at the AI models we're running, how big they are, how big even the hidden dimension is, and how big the matrices are, the way you do photonic matrix multiplication is basically in the analog domain. You're moving light beams across distances.
You would need a football field to get the equivalent matrix size to multiply it in “photonics.” At that point, it would be free-space optics. So, photonic computing—am I bullish? Not really for the data center side. It could probably work as some kind of IoT or embedded device. I wouldn't be surprised if the Whimo LAR system starts doing some compute on the sensor information coming in, but it would be really tough. I'm not that bullish on photonic computing.
I don't want to say anything more negative than that, but, yeah, I think people are selling bridges there, right?
We'll see how long it takes. I feel like we've done a good job of covering the broad zero-to-one in the data center, both in prefill, decode, and training. What are you personally excited about? You've been very right directionally about the broad market, but you said some of this is already solved. Is it just now scaling this up to more customers so that we can give more people broad access to AI, or are there still unsolved engineering problems that get you excited?
I think, as an engineer, the most interesting thing to me has always been memory, and this is why I've always liked HBF, even though I've kind of been hard on it up to now. I mean, it's a great use of technology that already existed. I think we're going to see more innovations in memory, and I'm pretty excited for that.
I think what people are understanding is that when you're doing inference and you have some kind of knowledge beforehand of what operations you're going to do or what data you need, you can structure your memory to exploit that. Jalapeño was at Hot Chips—the OpenAI chip—and they talked about memory locality. We talk about memory as a big bank and just get it all at once, but these actual memory chips have different channels, and they're physically far apart from parts of the compute. It doesn't seem that far—we're talking maybe millimeters—but it still matters.
What they realized and showed was that keeping certain memory banks with certain data that the compute engine could grab locally, or spatially close to it, gave way better performance. So I think this idea of having your memory closer to where your compute needs it is going to become more interesting in the future. I think people are talking about 3D techniques of integrating memory on top of compute, and in those cases it will be extremely important, because you don't have just one big, wide bus shuffling data back and forth. You'll have a bunch of teeny-tiny lines that can give you a couple of bytes at a time. So you have to organize those lines in space. It's literally like knowing where things are physically in the chip is a huge advantage. I think that's probably going to unlock more gains both on throughput and power without better HBM.
That's what I was curious about. If it's more about locality, is it more about throughput or latency?
Well, it could be both, right? Better locality means you have less latency.
But if you have more memory channels, you get more bandwidth. So maybe you just have the right amount of bandwidth for what you need in that one specific chip, or that group, or that addressable piece of memory, and that's all you need, but the latency would be a lot less, right?
Yeah. So I think everyone talks about this—even I talk about this. I did chip design for a while, and everyone talks about memory bandwidth and latency. I think one thing that people don't talk about is when you need this access time. If I know the operation on the GPU is going to take 10 milliseconds and I only need this at millisecond 11, does it need that much bandwidth or latency? It doesn't. You just need to know when you need to grab it, right?
6. AI Alignment
So, getting that under control is going to be interesting, and that has to do with locality. I think there are interesting things. I would say if anyone wanted to read about what I think the future of memory is, it's 3D DRAM. It's kind of—what was that company called? They were at Hot Chips recently and showed DRAM bonded on top of logic. We'll find it out after this.
How was Hot Chips?
Hot Chips was good. A big reason I moved to Palo Alto was that last year I came for Hot Chips, and I was like, “Whoa, I'm a guy from New York. No one knows anything about chips. No one cares,” which is fine, right? People care about other stuff. But I come here, and it's like legend after legend, and we're all speaking the same language and trading ideas. It felt like I was in a coffee shop in London during the Enlightenment, and everyone was just trading ideas about how to make these chips better.
This year it was packed, and it was filled with all kinds of finance people, not necessarily all people.
Yeah, a lot of people from hedge funds and stuff like that, which I thought was very interesting. Last year it was more like high-frequency trading firms who actually just need to know how these chips work too, right? They were just there like, “Oh, I have to work with these chips as well.”
I think it was interesting to see finance people asking questions to the speakers. So funny. But, yeah, I think 3D DRAM is going to be super cool, and I hope I see more people taking new approaches there, because that's a way where we get better bandwidth and better latency, but the design of the chip is going to be a little harder.
Interesting. You also mentioned power. I think Jensen has been fairly repetitive about making the point that if you only have a 1-gigawatt data center, you're maxed out at a gigawatt, and so the better way to do it is drive different power efficiencies. Have you gone down the power or energy rabbit hole to an extent?
I think you can roughly estimate tokens per watt, or token throughput per kilowatt or megawatt, or whatever you want. I think Jensen, yes, there are things we can do on the power side to lower our power use, but they may take hits on token throughput or the number of users you can serve. So you really have to understand the economic equation from the people who are actually selling you these tokens—the big labs.
They care a lot about serving more users and batching more. That's the key word. They want to batch as high as possible so they can serve more users with the same piece of hardware. You'll get more tokens out if you batch more, and you'll get more tokens per watt. So is it necessarily worth doing all this exotic, difficult stuff to lower the power consumption of these chips? Maybe yes, maybe no. It's a big balancing act. On the power side, I think it's a lot harder to lower the power consumption of these data centers. It's a huge challenge.
You're probably better off just finding cheap, cleaner power. Makes sense.
I will add one thing there. The number-one cost in the data center isn't even the power. It's actually the cost of the hardware. So if you're trying to start a neocloud, what's the power cost? It's insignificant. What's the land cost? What do the GPUs cost? Oh, God. Yeah, you better have good financing for that, right?
So, the most interesting thing continues to be memory, especially since weights potentially are no longer increasing in total size while context generally is increasing.
Yeah, for sure.
Even if those stay static today, you still need more memory to serve more concurrent users.
Of course. Yeah. We will add more memory to these servers. At the end of the day, the question is what kind of memory. Will it be HBM? Will it be DRAM? Will it be flash? My guess is flash, more so. People are all talking about ways to de-spec the amount of HBM you need because, honestly, at the end of the day, weights should probably be in HBM. Why should we have huge amounts of context there if we're not using it all at once?
That makes sense. Interesting. Well, I think at a high level, that was kind of all my questions, unless you had any other interesting things or things that get you excited that you wanted to speak about.
I want to talk about Pacing the Frontier. Honestly, I want to know what that's all about.
Good point. I'm very curious what it means because it's kind of been not defined at all, and I've been thinking about it more and more. We've heard about them solving these millennium problems, and they spent $15 million to make $1 million, which, okay, but that problem wasn't that valuable, right? I mean, Navier–Stokes—did the world change overnight if we solved it? I think we'd probably save some compute cycles in a CFD simulation.
But there are a lot of problems out there that are really valuable in math and theory. I think of Terrence Tao. A lot of his research was in fields adjacent to EE, like signal processing and estimating and reconstructing signals that are sparse. His mathematical work has built the bedrock of how phones work nowadays, like LTE and cellular stuff. You use less bandwidth because of his math research.
So how valuable is that math problem? It could be worth billions. I'm curious how much of pacing the frontier is vertical integration in going after valuable problems and how much of it is actual security.
I'm not a big believer in security, honestly. We talked about power. We can cut off the power. I think the bigger issue with security is that these labs can be liable. If their AI does something very bad to someone's computer, that's grounds for a lawsuit. At the end of the day, I feel like we have a legal framework in the country to structure that in the right way.
We can make a contract where, hey, you're liable. You sign this, right? If this agent goes haywire, it's on you. I don't know. I'm not a big believer in the Terminator scenario of these agents taking over the internet and bringing it all down, just because it costs so much to keep these GPUs running. It just seems like a bad use of resources, right?
I agree.
I'm thinking about this. I flip from being bearish to bullish, but here's a really hot take. I use AI all the time, but I've never been someone who's like, “Oh, AI is great. It's crazy. We all have to get on this.” In my eyes, AI is good for doing my dirty work, my busy work, but the thinking has sadly still fallen on me. I've never gotten the impression that it's replacing that yet.
7. The Future of Labs
If you talk to people who are middle schoolers or high schoolers, they don't use AI as much because what are they going to use it for? What are they going to do? They're going to write their little essay on George Washington or something. It's not that useful.
One thing I've been thinking is that AI—really powerful AI and high-compute resources—should probably go to important problems, and people can actually work on those important problems. That's always been the bullish case, right? What should be the value of a GPU or a server with GPUs if it can potentially find a drug, even if it has a 10% higher chance, and that drug is worth billions of dollars? Or a new material that could save us energy?
That's been the most important thing to me from the AI perspective, and it's why, up to now, I really haven't cared that much. Coding is cool, websites are cool, slop is slop, but let's do the hard work. I don't know how much of pacing the frontier is them deciding that all these compute resources and these crazy models are probably better served by going after really important problems, or how much of it is just regulatory capture.
I agree. I think it's mostly cope. I'm not in the Bay, so I'm not as inundated by all the broad groupthink, but to me, I agree. I don't think the models are—like, I heard this from a friend and I disagreed. He said that the models are almost equivalent to iPhones at this point, where—
They're only marginally improved, and they're not going to get that much better. But—
The iPhone is the best example because we're all addicted to our iPhones. We can't leave the house without them. It's becoming something critical for our work now, so—
Yeah.
I do think it'll eventually get there. But to your point, you can ask it questions and it obviously gives you decent answers, but it's not replacing thinking. You can't default to it today as though it's a better answer. You have to double-check it. You have to check what's in its context, what it brought in on the prefill, and the different sources that it brought in.
To your point, it's doing more of the grunt work. It's like, all right, can you reply to emails? To me, you add more and more context, and it does get more interesting and more useful, but right now the models still broadly have amnesia. They are sufficiently smart, but they're not smart enough to one-shot everything. I don't think we're at the point where we have superintelligence.
In large part, I do think it's broadly cope that they want to slow down, and maybe they just want to use the models internally to continue to advance. But—
That seems like a great way to do it, because if I have a set amount of compute resources, it costs a lot of money to run this stuff. It can go to serving more users, where I'm charging them, I don't know, $20 a month or something like that, or it can go to finding a new material or solving a very important problem in signal processing and stuff like that.
I don't hate it. I hate the delivery of it. I actually really hope that people use AI for these valuable discoveries because we're bottlenecked so much by human work. Humans have tons of good intuition, and we can come up with all kinds of ideas for how to solve a problem. You have to do work at the end of the day to really solve it and prove it out.
There's so much there, but I think we would be remiss if we didn't talk about pacing the frontier, right? I agree. What are your thoughts on open source versus closed source? All the rage has been broadly that open-source models were relatively behind. They seem like, if anything, they're very close. Again, it's kind of hard to tell if the closed-source labs are keeping their newest models behind closed doors for longer and longer, but—
All these models, from the hardware standpoint and maybe the investing standpoint, seem not to be able to run on commodity hardware unless you're running very old models that are—
Not super smart, unless they've done a bunch of reinforcement learning on a specific given task.
Well, you are right. It does take some know-how, and it's a challenge to run the models. That's why there's a big business of just running open-source models for people, right?
Open source, to me, has always commoditized coding. Right now, GPT-6 is pretty good at video games and graphics and stuff like that, so I'm sure that will get commoditized as well. But the frontier has always been about pushing the limits, right?
We've kind of run out of stuff we could just do on computers up to now. We can code, we can send emails, we can work in Excel, but we need to break out of the computer. We need to change the fabric of reality, which is discovering something and working on more important problems than just web apps and iPhone apps. That stuff is all well and good if you get value out of it.
But I think there's a case now that we're going to get recursive self-improvement. Isn't the AI recursively making itself better? It's probably the AI making us better, which is the AI finding new things and automating a lot of our work to make us faster and more efficient, right? I think that's where the frontier should probably stay.
The reason I bring this up is that if pacing the frontier means serving fewer users to go after these high-value problems, it really changes the business, and it's downstream of everything. It could be good, it could be bad, but we don't really know yet. I think the potential proof of the Navier–Stokes problem isn't necessarily indicative that they're going to be effective at every other problem domain. It makes sense. And even that allegedly came from someone's human intuition, right?
So what I'm hearing is Frontier should continue to push would like maybe as like a more direct question if you had the opportunity to invest in the labs at whatever valuation they are today call it a trillion do you think that's more interesting than doing something like a neocloud that's potentially just serving open source models
8. Closing Thoughts and Where Value Accrues Next
I'm like a big believer in just owning the the hardware layer or owning the railroads or owning the things that go into it, right? So, do you want to like invest in anthropic and open AI who are going to compete to death and eventually their alpha will diffuse as people hop between jobs and stuff like that or do you want to own like the people who have 10 year 20 year track records and moes of developing you know the integral hardware that goes into this or anything even from you know you're very good at running a data center I mean there's a case to be made that if you have a lot of GPUs and you're very good at keeping the uptime high that you're very valuable company, right? Um, I'm sure most investors are not going to be thinking like that and they're just going to, you know, top blast these things as they probably should. But at the end of the day, like I a very good rule of thumb I've used for the sign if a company is like a strong value prop is if this company disappeared for six months, how screwed would everyone be? So I think of like TSMC like pretty bad, right? uh Nvidia pretty bad. Even like SanDisk, yeah, we could probably get it to from Kioa, right? Um the there's ways you can argue where like people are willing to switch between products, right? The best companies are the ones that just like give you a reason to never leave, right? Um and those are like so if an AI lab starts doing something like that, that would be very interesting, right? But I don't think they they've gotten that smart yet, you know? Um, but we will see.
I think, yeah, I'm looking forward to the IPOs. I mean, they'll make great funding shorts. I'm pretty sure it'll be a great beta offset for a portfolio. So I'm very excited. It's good to get a new financial instrument for hedging.
So continue to go long hardware. I think it's—
Might be the real innovation. Yeah, the hedging characteristic—
The amount of new gigawatts. I mean, obviously, TBD on whether they can get energy. But, yeah, to me, going back to pacing the frontier and why it kind of seems like a farce is because it's not like they're going to hit pause on these data centers that are building out, or—
They're pacing the frontier for us. They're not pacing it for themselves.
Yeah, they're not slowing that down. And if anything, they keep talking about wanting to bring on additional gigawatts for both training and inference. And so those are many multi-year commitments just to build out the data center. And so if they were truly serious, I feel like they would be canceling those—
Spending. Yeah, but they're not—
That is not going to happen.
But, yeah, I'm leaning to be very bullish. I mean, we have this whole situation going on with interest rates and Iran, but I don't know. I have a feeling that we find a couple drugs with AI and get the right researchers in the room, or we figure out a way to make 5G not need 1,000 towers in a square mile.
Yeah. Yeah. Starlink.
Yeah, whatever. Anything that's more efficient. Yeah, I mean, we will have a fast takeoff. I mean, you can justify all the capex off just 5 or 10 drugs, I think.
Very true. Cool. Well, boi, thank you so much for coming on. It was a lot of fun. Thanks for nerding out with me and dropping your alpha.
For sure. Thank you, Logan. Thanks for having me.