[BidClub_]
Frictionless · · 73 分钟

为什么内存是 AI 最大的瓶颈:Vikram Sekar|第170期

Logan JastremskiVikram Sekar

半导体AI与软件技术
YouTube ↗
TL;DR
  • Vikram Sekar 的核心判断是,内存而非原始算力,是 AI 最深层的系统瓶颈。 LLM 没有用满 GPU 算力,是因为内存和网络无法足够快地供给数据;因此,内存占用更低的模型架构可能同时缓解多重约束:“内存瓶颈,也就是内存带宽瓶颈,进而也是网络瓶颈。”

  • 面向消费者的智能体应当扩大推理需求,但不会沿着 Mac mini 热潮所暗示的那条永久专用处理器曲线增长。 大多数个人任务用 Sonnet 级别或开源模型就能完成,处理器和内存则可以共享,或在需要时即时启动。需求仍会增长,但 Vikram 警告说,“一切都会因为智能体而爆发,而这可能并不会真的发生。”

  • 即便今天的财务投入最终被证明过度,AI 仍然是一项具备长期韧性的技术。 Vikram 预计这项技术未来10年仍会“向右上方增长”,但他提到 Anthropic 的 S1,称其有义务投入的资本和算力资源对应金额似乎超过5000亿美元。实验室可能订购过多内存、制造供给过剩并遭遇崩溃,但这并不会推翻底层的采用逻辑。

  • 训练仍是统一集群问题,而推理正变成一片“狂野西部”。 从72块GPU、约200 kW的 Blackwell 机架走向144块GPU,意味着功耗超过400 kW,同时冷却和网络难度进一步上升;如果再把576块或1,152块GPU分散到多个机架,连接距离就会超出铜缆的能力范围。推理则缺乏同样的秩序:混合专家模型会从不同位置调用不同权重,让数据搬运变成“彻底混乱”,也为 Cerebras、Groq、Microsoft Maia、Meta MTIA 及其他专用架构留下空间。

  • 短期内,扩展规模的技术路线越来越偏向光互联,因为铜缆正在逼近物理极限。 铜缆在机架内部可以支持200Gbps,但无法跨越8个机架、约10米的距离;到了400Gbps,Vikram 认为它“已经彻底失效”。共封装光学可以在降低转换功耗的同时提供传输距离,但对于包含100,000块GPU的系统而言,可靠性、替换和大规模部署仍是问题。

  • 内存层级的优化目标应是每个 token 产生的有效工作量,而不是让所有场景都追求最高 token 吞吐。 对于上市时间至关重要的编码任务,1,000–2,000 tokens/秒和昂贵的带宽可能物有所值;而在后台读取 PDF 的智能体,则可以用更慢、更便宜的方式处理大量 token。真正的经济问题是:“每个 token 实际完成了多少有用工作?”

  • HBM 目前提供了最佳的带宽—容量平衡,但堆叠继续扩大正面临成本和扩展限制。 目前堆叠已经达到16层,Vikram 质疑继续推进到20层或24层是否仍具备经济性和技术合理性,并关注晶圆键合的 DRAM-on-logic 方案,后者有望实现“类似 SRAM 的带宽,但拥有类似 DRAM 的容量”。不过,任何效率突破都可能触发杰文斯悖论:单位任务的资源消耗下降,会刺激足够多的新需求,最终“我们的计算资源还是不够”。

摘要 · 为研究而整理的核心内容

1. 托管式智能体把工程项目变成消费产品

  • Logan 将拐点定位在去年第四季度 Opus 4.5 之后:行业开始从 OAuth 订阅转向基于 API 的使用模式,智能体也开始把 AI 的边界从聊天和编程向外推进。早期 OpenClaw 部署展示了这一潜力——处理手机消息、通知并自主执行操作——但用户需要配置虚拟机、做出安全决策,还要有足够的技术信心,把一台电脑交给不可预测的智能体。

  • Vikram 自己的使用方式正好暴露了问题:他至今仍在家用服务器上运行 OpenClaw,但他说,“事情没那么简单。”人们购买 Mac mini,并不总是为了在本地运行模型,而是为了搭建一次性沙盒——这些机器没有邮箱或银行账户权限,智能体可以“放开了搞”,却不会“把我的生活搞砸”。

  • 云端虚拟机省去了购买硬件的步骤,却保留了配置负担。托管式产品进一步把虚拟机和智能体打包在一起:Vikram 花200美元试用 Grok bot,几分钟内就通过精致的 GUI 启动了智能体,心里想的是:“再见,200美元。”真正让智能体变得普及的,不是模型出现了根本性突破,而是价值实现所需的时间大幅缩短。

  • Instinct 让他彻底改观。Vikram 让它帮忙报名 Open Compute Project Global Summit,Instinct 随后返回二维码,并表示已经订购了一件尺码合适的 M 码 T恤:“如果有任何事情都能由它替我完成,那我就加入了。”

2. 智能体需求增长,但共享基础设施会拉平曲线

  • Vikram 的谨慎判断是——“我不知道自己是否正确”——个人助理很少需要前沿级智能。日历监控、邮件分拣和日常行政事务通常用 Sonnet 级别或开源模型就能完成;庞大的数学问题则是另一类工作负载。

  • 每个智能体也不需要永久绑定一颗处理器。云端硬件可以在用户之间轮转,智能体闲置时可以让内存退出服务,算力则在需要时即时启动。Vikram 自己的 ChatGPT 每小时检查一次重要邮件,而不是持续占用专用算力。

  • 但潜在用户面要大得多,因为“每个人都会给别人发消息”。家长可以通过普通消息跟踪学校日历、出行、足球课和日程变化,让那些从未想要编程助手的人也能用上有价值的 AI。

  • Logan 始终关注规模问题:即便是共享的低阶推理,也可能在数千万乃至数亿用户之间叠加出巨大需求。Vikram 同意算力会持续增长,扩展定律也尚未失效;他的较窄判断是,智能体的普及并不等于每个人各对应一个前沿模型和一颗永久专用的处理器。

3. 技术乐观与财务不安并存

  • Vikram 预计未来10年 AI 会“向右上方增长”。他的类比是互联网经历互联网泡沫破裂之后的走势:融资和个别公司可以失败,但真正有用的基础性技术仍会延续。

  • 他没有掩盖当前范式的不确定性。概率型 LLM 也许不是真正的“智能”,未来可能必须出现另一种路径;但他并不声称自己知道答案。他能观察到的是,AI 已经让一个人能够同时处理 Substack、播客、机构事务、会议和改期工作,而这些任务过去往往复杂到难以独自完成。

  • 预测未来6个月的能力要弱得多。去年年末,编程智能体看起来可能是最主要的使用场景;此后,人们对聊天机器人半截答案和幻觉的担忧有所消退,智能体已经进入手机,边缘 AI 又提出了下一个问题:“为什么每次都必须去云端?”

  • 在财务层面,Vikram “非常不安”。他提到 Anthropic 的 S1,称其有义务投入的资本和算力资源对应金额似乎超过5000亿美元,同时承认整个行业可能过度订购内存、进入供给过剩并最终崩溃。他的结论分成两半:对技术“保持乐观”,但不愿忽视过度部署的风险。

4. 训练需要一块巨型 GPU;推理却像急流

  • 训练仍在推动统一扩展域从72块GPU走向144、576和1,152块。72块GPU的 Blackwell 机架已经达到约200 kW;GPU 数量翻倍后,功耗可能超过400 kW,还没计入配套网络,电力输送、冷却和物理封装因此成为一等约束。

  • 分散到多个机架可以缓解单机架的密度问题,但连接距离也会从几英尺拉长到数米。覆盖约8个机架需要接近10米的传输距离,而铜缆无法以200Gbps完成这一任务。让大量 GPU 表现得像“一块巨型 GPU”的目标,最终把互联变成系统瓶颈。

  • Vikram 将训练中的数据搬运比作海浪:操作成批到达,all-reduce 把波浪送回,下一轮迭代随即开始。混合专家模型的推理则更像急流——每个请求只激活一个更大模型中的数百亿参数,连续两个请求可能需要调用存放在房间两端的权重。

  • 这种随机性解释了为什么推理领域允许更多架构实验。Google 从第4版 TPU 起就将训练和推理设计分开,而第8版似乎会同时发布两种设计;Cerebras、Groq、Microsoft Maia、Meta MTIA 和 OpenAI 的“Jalapeño”分别从不同方向攻击这一工作负载,因为“这就是狂野西部”。

5. 铜缆信号衰减,迫使 AI 转向光互联

  • 底层物理决定了传输距离。GPU 旁边的 HBM 可以提供约22TB/s,因为比特只需移动几毫米;在机架内部,信号要传输几英尺,能量随之损失,系统不得不降低速度,并引入预加重、纠错和 DSP,而这些都会消耗时间和功耗。

  • 在200Gbps下,铜缆在 Rubin 机架内部仍然可用。Vikram 的明确判断是,到了400Gbps,5米外的信号几乎无法辨识:“铜缆已经到达极限。”它同样无法连接起构建更大统一扩展域所需的相邻机架。

  • 功率密度的对比解释了为什么简单放大机架并不可行。云计算时代的机架功耗约为20 kW,AI 机架则是其10倍;如果再翻一倍,名义功耗就达到400 kW。现实中的过渡方案是让8个相邻机架作为一个整体运行,这也使光互联变得不可避免。

  • Vikram 认为,传统可插拔光模块在能耗上并不划算,这有助于解释 NVIDIA 为什么没有更早采用光学方案。共封装光学把电光转换放到 GPU 或交换机旁边,同时解决功耗和传输距离问题,但也带来更棘手的运营问题:当高度集成的光学组件发生故障时,如何在100,000块GPU的部署中完成替换?

6. 光学资源稀缺,催生“宽而慢”的架构竞赛

  • 如今常见的1.6Tbps光连接,通常由8条200Gbps通道组成。这些高速激光器使用磷化铟——能连接数据中心乃至洲际链路的“法拉利技术”——但眼下的任务其实只是与相邻机架通信。

  • 供应是关键约束。磷化铟激光器来自3英寸或4英寸的小型晶圆,Vikram 将其比作小煎饼;而这套原本服务于长途、城域和海底链路的产业,“从来没有准备好应对这种级别的需求”。

  • 另一条路线是“宽而慢”(wide and slow):使用更多低速通道。砷化镓 VCSEL 的速度可能在20–50Gbps左右;如果按50Gbps计算,32条通道就能重新构成1.6Tbps总带宽,同时依托成本更低、供应更充足的产业链。MicroLED 也提供了另一种光源,但其速度约为3–5Gbps,需要数百条通道。

  • Vikram 没有假装知道最终赢家:“我不知道正确答案。”眼下的争论,是用更少但稀缺的“法拉利”激光器,还是用更多、更便宜的通道,在带宽、制造、能耗和可靠性之间实现最佳平衡。Logan 随后将话题拓展到电力:从发电、逆变器、UPS,到储能和电容器,整个链条同样是一个巨大的约束。

7. 有效工作量,而非峰值速度,应该决定内存层级

  • 传统的层级结构——SRAM 缓存、DRAM、SSD、硬盘和磁带——已经碎片化。如今的 AI 覆盖片上 SRAM、HBM 和堆叠 DRAM、面向性能优化的单层 NAND、面向容量的 QLC NAND、介于两者之间的组合、磁盘,以及容量近乎无限但速度极慢的归档磁带。

  • Vikram 关注的核心问题不是哪个层级胜出,而是“每个 token 实际完成了多少有用工作?”如今行业像过去炫耀 FLOPS 一样炫耀1,000或2,000 tokens/秒,但在后台处理 PDF 的智能体没有理由为最高即时性付费。

  • 企业级编程可以证明高价高速方案的合理性,因为抢先上市带来的收益足以覆盖成本。后台研究则可以用更慢、更便宜的方式处理更多 token。把每类工作负载匹配到合适的内存层级,可以降低成本、改善毛利率或把节省让渡给用户,进而扩大总使用量。

  • Logan 提出更大的上下文窗口,可能从100万扩大到1,000万,甚至达到 Dario 据称提到的1亿。Vikram 的反驳是,并非所有上下文都有用。Logan 用健身举例说明这一点:智能体回答步数问题时,可以查询卡路里记录和健身追踪器数据,却不必保留用户讨论数据中心光学的工作对话。

8. DRAM 创新当下重要,但新方程可能重置内存栈

  • Vikram 最看重的层级是 DRAM,因为 HBM 目前提供了最佳的带宽和容量平衡:“眼下你离不开 HBM。”但堆叠已经达到16层,他质疑继续机械地推进到20层和24层,是否仍具备经济性或技术合理性。

  • 在逻辑芯片整个表面键合内存,能够提供远多于沿芯片边缘布线的连接。他提到 Cerebras 将整片 DRAM 晶圆键合到其晶圆级引擎上,Groq 预计也会推进相关方案;此外还有 D-Matrix 的 Raptor engine,以及采用2层或4层 DRAM 的 Qualcomm 设计,其共同承诺是“类似 SRAM 的带宽,但拥有类似 DRAM 的容量”。

  • 未上市公司版图反映出推理架构仍未定型:Vikram 提到 D-Matrix、以 TCO 为核心的 SambaNova 客户方案、颇有吸引力的 Etched 低电压引脚架构,以及 MatX、Fractile 和 Nubis 的芯片间纳米激光器。他的态度是探索性的,并不是宣称某一种设计已经胜出。

  • 未来5年最大的风险位于所有硬件供应商之上。当前硬件实现的是某一类 LLM 方程;如果研究突破能够用更少内存实现相当性能,就可能同时缓解带宽和网络压力。但杰文斯悖论仍然存在:每一次效率提升都可能带来更多使用量,最后又回到同一句话——“抱歉,我们的计算资源还是不够。”

完整逐字稿
Vikram Sekar

Memory is the biggest problem. LLMs are literally memory-intensive models, and memory is everything. One of the biggest problems people have to solve is how to get the same performance without using as much memory. Once you solve this problem, a lot of things will fall into place.

The optics were the same as for long-haul connections, subsea and submarine cables, and some other applications. But optics have never been as important as they are today. Today, it is a technology that is essential for AI and future infrastructure. The whole world is fighting for it because everyone thinks, “Look, we were never prepared for this level of demand. This industry was never built for this level of demand. You’re asking for too much.” Because of this, everyone is facing shortages.

Logan Jastremski

What is the useful amount of work performed per token? Do I always need 2,000 tokens per second? I don’t. I’m talking to you right now, but my AI can process some PDFs that I want to read later or something. I don’t have an immediate need for this. Perfect.

1. Vikram Sekar and Semi Doped

Well, Vic, thank you very much for joining me. I appreciate you participating in the podcast. I’ve been following your work for a while now, and I’m really impressed. It helped me personally get from 0 to 1 as quickly as possible.

I wanted to reach out to you and see if I could invite you to the podcast because I think you’re really good at explaining technical concepts in a way that people can understand. Knowing many engineers, I know that this is a difficult skill. I appreciate all the work you put into teaching others.

Vikram Sekar

Thank you. Thank you for inviting me. I’m always happy to sit down and talk about technology and how things work. That’s basically what I usually think about.

Logan Jastremski

Yes, I’m always happy to talk about different things.

Vikram Sekar

It’s even more gratifying to hear that my explanations, articles, and conversations on SemiAnalysis and my Substack newsletter, The Vik Newsletter, have resonated with people who care about this subject. I’m very happy to hear that.

Logan Jastremski

Yes, and for those who don't know, I'll link to Vic's entire podcast on Substack, everything he's worked on, including his institutional research. Please contact him because I think he does a great job. And then I'll have to contact you to better set up the podcast studio. Maybe I can get my own seamless logo in the background.

Vikram Sekar

Yes, why not? This is great.

2. AI agents and hardware demand

Logan Jastremski

Perhaps a great place to start would be that the industry has really changed, it seems, from 0 to 1—or perhaps it’s better to say from 0 to 100—since Opus 4.5, which was in Q4 last year. We started moving from OAuth subscriptions to API-based usage, and that seemed to give the whole race a boost.

Now everyone is talking about agents, which I think is really exciting: just using a computer and being able to do things for ordinary people outside of software engineering and coding. We have a new race starting.

Maybe it’s worth starting the podcast with a new question: How do you see the situation with agents, and what are your general thoughts?

Vikram Sekar

I really like the current agent situation more than at any time since the beginning of this year, when agents like this really took off. At the beginning of the year, it was more a question of, “How do we get away from the chat interface and maybe, for the first time, make it work on the phone?”

When OpenClaw came out, you could send text messages from your phone and receive replies, scheduled events, notifications, and more. But the problem with that approach, at least at the beginning of this year, was that it was very difficult for regular people to go and set it up.

I’m not just saying this as a regular nontechnical person. I tried setting up OpenClaw. I still have a home server that runs it. It’s not that simple.

I had to spin up my own virtual machine. There were all these security concerns about giving OpenClaw access to your entire computer. People would say, “No, no, no, don’t install it on your laptop. It’s too intrusive. Go and buy another computer.”

So everyone went and bought a Mac mini. Do you remember this whole story?

Logan Jastremski

Yes. It was very funny to me.

Vikram Sekar

Considering that most of the models people used, in my opinion, were not local models, even if people were buying multiple Mac minis and trying to get as much unified memory as possible, they would still be making API requests to data centers.

Logan Jastremski

Yes, the whole idea was that you could run some kind of local model that would at least handle your personal tasks or something like that.

Vikram Sekar

But again, this is a very expensive proposition because models become cheaper to use over time, while you have to make an initial investment of several thousand dollars, hoping that the local model you deploy on your memory allocation will actually do the job. This is not a very long-term bet. You would be better off betting on token prices dropping so that frontier models can be used more cheaply over time.

Logan Jastremski

Yes, completely. It might be interesting to start this topic by saying that agents are starting to gain popularity. We had OpenClaw, which started it all, and then you saw the Hermes model come out. Each iteration was a little more user-friendly than the last.

Then the Grok bot came out, and then you had Muse. What do you think it actually takes to run these things? It’s interesting that people were trying to buy a bunch of Mac minis, and in your opinion, the spending limit for the average person was high. It doesn’t seem realistic to ask the average person to spend $5,000 to $10,000 on a personal assistant.

As these things start to scale, it obviously affects data centers and their buildout. What do you see overall under the hood if you deploy this—if these agents really start to scale to tens of millions, hundreds of millions, and billions of people?

Vikram Sekar

The evolution is very interesting because when people were buying these Mac minis, it was important to ask the question: Why were they buying them? The reason was not always that they could run local models. They just wanted a separate environment where they didn’t need to keep their personal information.

On this Mac mini, you could store what the agent had access to. Even if you were using cloud models, you could make sure never to log into your email on this computer, for example. Never access your bank accounts. You don’t know what this agent is going to do, so you just want to give it its own sandbox environment.

It’s like, “Go crazy. Your damage will be limited. You’re not going to blow up my life. You’re going to stay in this box.” If something goes completely wrong, I just clean this box and everything is fine again.

But then people immediately said, “Look, you don’t have to spend hundreds of dollars to buy a Mac mini. Just buy a cloud virtual machine.” So people said, “Okay, let me just spin this up.”

All this talk about companies like DigitalOcean and other providers came about because everyone wanted to spin up virtual machines in the cloud. However, all of this is difficult to do. First, you need to buy a machine and install OpenClaw. Then you need to set up a virtual machine and install OpenClaw on that virtual machine.

The next step in the evolution was for these companies to say, “Wait, why don’t I do all this for you? I’ll give you a virtual machine. I’ll install the agent, but I don’t want to install OpenClaw. I’ll install something that my company is developing.”

Grok bot was the first one I tried. As soon as it came out, I signed up for $200 and said, “I’m going to use it right away.” It was fun because you could literally launch agents by talking to it. It had these cute icons, and the GUI was beautiful.

I thought, “This is cool. I don’t need to do any of the things I tried to do with OpenClaw. That’s great.” In a matter of minutes, I had downloaded it. My credit card was already on file, so I thought, “Goodbye, $200.” Then I had my agents working. That’s great, isn’t it?

Now it’s more accessible—much more so than ever before. This has implications for exactly what kind of hardware we’re going to need to run all of these things, right? People are now buying into companies like Grok, and now also Muse.

Muse supposedly runs on 2 processors and about 8 GB of RAM. People say it works for every agent you sign up to use. But all of this is stored somewhere. There’s a computer somewhere in the world that works on your behalf, right? That does your work for you.

All of this translates into demand for processors, demand for RAM, and demand for more inference, because all of these machines will now be performing inference without your asking them to. This is the whole point of the agent. It will follow its own logic and tell you what is important and when it is important.

The number of these things will increase. The real question is all about relative amounts, right? More or less, yes, but they all have to grow.

Logan Jastremski

Yes. This is interesting. When Grok got its own virtual machine and was able to start using the computer on the backend instead of using my own, it felt like a magical moment. Even though you could set it up yourself, it was, to your point, just a bunch of extra steps that were a little more cumbersome.

Vikram Sekar

Yes. Yes. I think what’s interesting to me—and I think Elon talked about this quite a bit—is that, on the computing side, there’s a bunch of software that has been invented in the world that doesn’t have APIs or MCP connections. If you can use it like a normal person, it unlocks a lot of functionality.

And it was even more amazing recently, just with Project Astra and certain advances in computer use. I was pleasantly surprised. It’s not yet obvious at a human level, but it seems like agents in each iteration of the model are getting better and better at simply doing things on your behalf. This is amazing. I have a story about this.

There’s another agent thing that we haven’t mentioned yet called Instinct, which is getting a lot of attention right now. It has received a lot of funding, too. For example, even GOTO posted about Instinct on its X account today. This is very exciting technology, and I managed to get an invitation pretty early.

So I thought, “Okay, I guess I have to try it. For the sake of science, let me in. If my stuff gets broken, I’ll deal with it later. YOLO all of this right now.” So I signed up and was granted access. Then I got this conference registration email: “Hey, the Open Compute Project Global Summit is open, and you can register here. Here’s the code and everything.”

I was doing something, and Instinct said, “Yeah, okay. Do it. Sign me up.” I forgot about that and returned to my work. Then I came back to it and got a notification on my phone because I was running it on my iPhone. It said, “Oh, yes, you’ve been registered. Here is the QR code. And, by the way, I ordered you a T-shirt—medium size.”

I thought, “This should be normal, right?” I was like, “How the hell did you know my T-shirt size?” That worked great. It seems like everything had been done. Even my T-shirt size was correct, so maybe it was a good guess. But I thought, “Here I am on the team.” If anything can do all this for me, then I’m on the team.

Logan Jastremski

Yes, it seems that, from zero to one on the software engineering side, it was a pretty slow iteration over time. I’m glad that more people can now interact with artificial intelligence beyond chatbots and LLMs. It seems like agents are moving in that direction—hopefully in a good way, like in the movie Her—where the agent just knows everything about you and can do things for you. Maybe you have a personal relationship with your agent, but it certainly seems like that’s the direction we’re going.

Now, how do you actually make this work under the hood, based on how many gigawatts and what model size you need, versus how much memory you need and whether you need one processor or many? I think this is also an interesting question for people who follow this space and try to dive deeper into some of the more complex aspects of things.

Vikram Sekar

Yes, I have my own opinion on this matter. I don’t know if I’m right, but I’ll say it. For most day-to-day tasks, I don’t think we need frontier intelligence. We don’t really need to run heavy models if it’s something like an additional set of considerations and so on.

Overall, I think a lot of these models used for personal assistance could use Sonnet-level models. Even if you go open source, the cost of launching these things could be even lower. It follows that the number of tokens will increase, but not everyone will use the frontier model, so that has its consequences, right? Maybe we can provide more users who consistently use the Sonnet model than users who are trying to consistently run the frontier models.

There is a big difference in infrastructure needs between these 2 cases, and I think personal agents mostly don’t need the frontier model. You’re not going to ask your personal agent to solve this giant math problem. That’s not the job of a personal agent, is it? These kinds of things will continue to happen on this side, so these will be mid-range models that will do just fine.

The other thing is that, in the whole story of processor allocation, it would be wrong to assume that everyone who runs an agent would have their own processor all the time, right? It’s not the same as buying a Mac mini, but in the cloud. That’s not the case, because when your agent isn’t using it, this equipment will be used to serve other people. This is a shared resource for everyone.

The real question is, how often will sharing occur? Do you know how often you need it? Do you need your agents to constantly work for you and constantly do things for you? Or maybe, in about an hour, it might come to you and say, “Hey, this is what happened.”

I already have my ChatGPT that does this for me. Every hour, I get a notification in ChatGPT: “Hey, you missed these important emails. This one needs your attention immediately, but you can do the rest later tonight.” I already set it up, so that’s about an hour every day. I have other things that I don’t always need an agent working for me.

So, when it comes to hardware requirements, I think a lot of people will get on this train now because it’s easy to use. You can write text messages. Everyone knows how to write to someone, right? You just need to write messages to an agent. That’s all you need to do, so many people can do it.

But I think the hardware requirements themselves won’t be as advanced as the technology all the time. These will be shared processors with lower-level logical inference. Memory will also be shared. Maybe they’ll take your memory out of commission when you’re not using it and spin up your compute just in time to use it. A lot of this will be shared.

There’s a lot of sentiment that everything will blow up because of agents, and it may not actually happen.

Logan Jastremski

Is it because you think it’s more of a shared resource, and that advanced intelligence isn’t necessarily needed for everyday tasks, so you can stay consistent and just bring in additional users?

Vikram Sekar

Yes, you can. Now, you don’t have to stay constant. I think it will continue to grow. We’ll need more computation. We always need more computation. Even for training, the scaling laws haven’t ended, so we can build bigger and bigger clusters and install more and more GPUs. These things are getting smarter, which is amazing.

We’ll continue to build bigger GPUs, and that path has never changed. But the growing demand for these kinds of agents comes from a much broader population that would like to use these things. Of course, it depends on the cost, because right now this GPT dots thing in chat will only be for prosumers, right? Implementation at this stage is questionable. I think Muse is much cheaper to use.

Implementation is questionable, but assuming that a lot of people implement it, they will use more computing resources than we did in the past. But we have to be careful to limit our expectations that this doesn’t completely blow up and we’re suddenly on a whole new wave.

This is good because now AI has become useful for the average person doing normal things. You can monitor your kids’ school calendar and find out when their trips are, when their soccer lessons are, and whether there are any schedule changes. Everyone can use this thing. I think this will be much more useful than it was in the past, and that’s good for the development of AI, right?

3. AI spending and the risk of overbuilding

Logan Jastremski

100%. It was really interesting for me to try to understand what exactly the broad adoption curve is, because I think everyone is trying to predict, at least from my perspective, how many gigawatts are going to come into service. Is it 10 this year? Is it 20? Is it 30 next year? Is it 40?

Is it for pretraining, to build bigger and bigger models, or just for inference? Since potentially more of the workload goes to the inference side, maybe it has more to do with memory than GPUs. I’m trying to put all the pieces together. One thing that I’ve personally tried to track at a high level is just the number of gigawatts, because I feel like it’s downstream of everything else—or upstream, so to speak—and everything else falls back on that.

Maybe the space on the agent side obviously increases overall adoption, but maybe it’s not that straight a line, so to speak. What are the general things that you’ve actually been following as a general trend? I think people are divided into 2 camps right now. It’s like an artificial-intelligence bubble: the buildup is crazy, and it’s not sustainable. Or, on the other hand, it’s like it’s completely devoid of AI.

You think, “Of course, all these gigawatts are being built. The demand for inference is insatiable. We’ll never be able to get enough.” Everything else in between is a bit complicated. So, I guess, maybe not on that exact spectrum, but as you see the next year or even 6 months, where are you?

Vikram Sekar

That’s a good question. I think I have a long-term and a short-term view. The long-term view is much clearer to me. I wouldn’t consider myself completely devoid of AI or anything like that, but I really think AI is a useful technology that has beneficial consequences in the long run.

In 10 years, no matter what happens, it’s going to move up and to the right. When the internet came along and we had the dot-com era, we certainly had a crash. A lot of things happened, but it was a fundamental technology that still exists today, and it’s a long-lasting technology that will last a long time.

Now, we can argue about whether this is what LLMs in AI are actually capable of. I don’t know. People say, “No, this probabilistic, statistical approach to intelligence isn’t really true intelligence, and something else has to come along.” Maybe. I don’t know. There are much smarter people than me who know these things better. But considering what AI can do today, I think it’s a useful tool.

I use it for so many everyday things, which has opened up so many things that would have been impossible to implement otherwise, right? Otherwise, even what I’m doing now—between Substack, the podcast, and the institutional things I’m running—is actually too complicated for one person. So I have ways to optimize so many things: meetings, rescheduling. All the administrative work disappears.

4. Training versus inference

In the long term, I see this as a very useful technology that’s here to stay. In the short term, do I think we’re overdoing it? Well, that’s a difficult question to answer, because something new comes out about every 6 months. If you had asked me this question last year, in December or November, I would have said, “What, actually? Coding agents—that’s good, right? So coding—is that all we’re going to do with this thing? Okay, what about this chatbot?”

This chatbot used to give you half-answers. It’s not that you ever knew whether it was saying the right things or hallucinating. I think many of those fears have disappeared today. So even in 6 months, it’s very difficult to predict where the trajectory of AI development will go.

Then we had agents. Now we have agents in our phones, in our pockets, right? And now the next question is: What are these phones supposed to do? Is this a game about whether edge AI will finally be useful? Because even though you make some inferences on your phone, like the Instinct to text or the Muse, why do you have to go to the cloud every time? What if you could do some things on your phone, right?

This is the whole next step: What do we do with edge AI? So in the short term, if I look at the financial side, it makes me very nervous. I’m like, “Oh my God, look at the number, the volume of the deployment.” Look at Anthropic's S1 and see how much they’re obligated to deploy in capital and compute resources. This is incredible.

The figure seems to exceed 500 billion. They committed to creating that much. From a financial perspective, this is scary. But from a purely technical perspective, I think a year from now we’ll be in a much better position than we are now, even with this technology, because we haven’t even started implementing it yet, surprisingly. I think there are still many good things ahead.

So I’m optimistic about the technology, but don’t ask me if the financial issues are a bit too much. Are people ordering too much memory? Is everyone buying too much, and will we oversell, fall into surplus, and collapse? All of this can happen, okay? I don’t rule this out at all. But in the meantime, I think we’re still moving forward; we still have a long way to go.

So that’s a long answer to your short question.

Logan Jastremski

No, that’s a great answer. In terms of deployment, what were you most interested in? Because it seems like the leading labs are doing or investing the most, and maybe the hyperscalers are investing over $1 trillion in capital expenditures right now. Obviously, they’re creating smarter and smarter models.

At least from my perspective, it seems that as these models scale, the clusters should be coherent, in the sense that they’re located next to each other. We’ve seen some attempts to do decentralized learning, and for the most part, I don’t think it’s worked. So we’re continuing to focus on coherent clusters that have extremely high throughput, and it seems Elon even talked about building a data center in Memphis later this year.

I think they’re targeting 2.5 gigawatts, then about 8–10 next year. But it seems that as models generally grow in terms of the total number of parameters over time—and Jensen seems to have said this on Brad Gerstner’s podcast—the logical conclusion would be something like X million or X billion. And that seems true, considering that inference is the source of all revenue in general, and training is actually the loss leader for inference.

So how do you see things like this on the ground floor? I know you’re very deep into the technical side of different developments, both in inference and training, and how that shifts the workload a little bit from more GPUs to more memory, which obviously memory stocks have done pretty well this year.

Vikram Sekar

Yes. So the scaling in training, I think, will continue. We’re still trying to create larger and larger world sizes. We want to combine 144 GPUs and scale, then 576 and scale, and find 1,152 GPUs and scale. Will all this continue now? Seventy-two? Yes. Today it’s 72 in the rack.

But we want to go beyond that and put 144 GPUs in a rack, and that becomes very difficult because even a 72-GPU rack today—a Blackwell rack—is like a 200-kilowatt rack. Now, if you want to double that capacity, you will also increase a significant portion of the network capacity. So even realistically, if you just double the number of GPUs and say you have over 400 kilowatts, that’s a lot of power for one rack. Cooling is an issue, and connectivity is an issue, so there’s a lot going on when you put that many GPUs in one rack.

Now imagine you put 576 in one rack or 1,152 in one rack. That’s just impossible, isn’t it? So now you want to put them in multiple racks, right? And this creates many other problems, such as how to connect them. Because it’s no longer 1 or 2 feet, but 1 or 2 meters. To cover about 8 racks, you need almost 10 meters of reach, and copper can’t do that at the current speed of 200 gigabits per second.

Copper can’t handle a 10-meter reach at all—no way. So that’s the problem we’re facing right now. We’re going to continue to move to larger and larger domains and scale, because training is going to continue and is very important as we push the boundaries of what’s possible. We haven’t seen any signs of the scaling laws slowing down yet.

But inference, I think, is a really big source of income, because everyone needs inference, right? We all need a model, but do we all need a model at the edge of possibility? Maybe, overall, it’s a good thing to have. This is a longer-term project, because everyone needs research projects that give you the next best thing, right? Otherwise, what progress is there? Progress must continue, so it will happen.

But inference will reveal much more useful things, right? Agents are just one use case. Chips and architectures are constantly changing from the ones previously used for training. Even Google TPU v8 has a chip for inference and a chip for training. They’ve been doing this since version 4, actually, but this is the first time they seem to have released them together, and they have different chip architectures.

You can even look at data movement in inference and data movement in training; they’re completely different. There are all these weights in training. I like to think of them as waves, like in the ocean: they come in sets and go in sets. That’s how data moves. You have all these operations that are completed, and then there’s this thing called “all-reduce.”

So the whole wave of data comes back, and the next iteration completes, and then the wave of data comes back. This is a very familiar movement of the data flow. But at the level of inference, it’s complete chaos. It’s like looking not at the beach, but at the rapids—Grand Rapids or something. You look at the water and think, “What the hell is going on here?”

Because you have a mixture-of-experts model, you don’t pull all these trillion parameters out of 5 trillion parameters, or whatever. You just pull out a few tens of billions and say that this is my expert. Now you will be using a different expert than me. So think about it from a data center perspective, right? All these things are random. It’s chaotic.

It’s like the weights are being pulled out in all sorts of places. The weights you need may be in the cluster over there, but then your next request might require weights all the way at this end of the room. Literally, that’s what it is. It’s as if the weights could be stored at different ends of the room you’re asking about. So when you look at the patterns of data movement, it’s chaotic, right?

So we’ll see all these things like divergence. Inference now unlocks, because of the nature of the problem, a lot more architectures in chips that might be better suited for this. This is not for training. That’s why you don’t see too many companies trying to build training chips, because these big GPUs are great for training, from AMD and NVIDIA, et cetera.

But when you go into the world of inference, it’s the Wild West. You have Cerebras, you have Groq, you have so many approaches. For example, if you go to Hot Chips, you’ll see so many different things. The Maia inference accelerator—Microsoft Maia—is completely different. Meta’s MTIA is a completely different way to do this.

Of course, OpenAI’s Jalapeño itself was completely different from all these other things. So many different things happen in inference. This is simply incredible. Obviously, there’s a reason why everyone’s focused on this, because there’s a lot of monetary value to be uncovered there as we move into the future.

Logan Jastremski

Yes, it’s super fun. I would say—and my world is now in blockchain, actually—if you were to boil it down to 1 key thing, it has historically been quite bandwidth-constrained. And I think the interesting thing about modern AI data centers that you usually talk about is, as you say, the weakest link in the racks is about 200 gigabits per second.

When you look at some things from the cryptocurrency side, you laugh. This is measured in megabytes, and I think that's so cool. We are pushing the boundaries of what is possible—for example, the physical limits of these different materials. Apparently, they're scaling to terabits and beyond, which is very interesting.

I hope that cryptocurrency and these backend blockchains will reach gigabytes and continue to scale, but so far, it's taken a little longer than I would have liked. That's crazy, isn't it? The speed is crazy. Most people don't even have gigabit fiber at home, you know? Most people don't have gigabit fiber-optic cable. Some people might get it, I think, but that's about it.

Think about the speed in a data center. This is madness. The bandwidth is very high. I think even conventional data centers have 10-gigabit lines and then scale to 100 gigabits. I know I might be a little off, but I think blockchain is pretty much the next evolution of finance.

Like fintech, it's just bringing together all these disparate databases and synchronizing information at a basic level. The bandwidth problem is that you have these different—I call them banks or exchanges—that have relatively low bandwidth, and what you need to do is synchronize the data between them. We've scaled from kilobytes, which was terrible, to megabytes, and now hopefully to gigabytes.

If you take it to the extreme, it's more like high-frequency trading, where you get, for example, 100-gigabit interconnects. You take a bunch of data and do data parsing, and then some interesting things with models of what you want to do. The reason I've been so interested in the data center side, and why I feel like I want to at least try to learn as quickly as possible on the journey from zero to one—and I appreciate your help again—is that a lot of it has to do with various bandwidth issues.

One thing I would like to touch on, in terms of what you mentioned, is moving to scale and also the footprint side. From a rack perspective, it's really interesting. A single server in a rack has a certain amount of bandwidth on the chip itself, and then the rack has some bandwidth with things like NVLink. As you go from rack to rack, the bandwidth varies a lot between those racks.

In simpler terms, can you explain the scaling there and why it's difficult, starting with the bandwidth issue and even scaling beyond GPUs to 144 and so on? Why was that more difficult?

5. Why distance limits copper bandwidth

Vikram Sekar

It's always easy to transmit bits very quickly if the distance is very close. For example, if you look at the bandwidth between HBM and the GPU, you get something like 22 terabytes per second. That's extraordinary.

It's really fast because every time you transfer a bit over copper, the longer it stays in that copper line, the worse the signal gets. You want to get from point A to point B and get rid of it, because the longer the signal stays in the copper, the worse it becomes. The same speed won't work once you go from that level of GPU to HBM, which is very, very close.

But when you go between 2 graphics cards in a rack that are, let's say, a few feet apart, instead of a few millimeters, you have to go a few feet. That's pretty bad. It's a big leap in distance, so you can't achieve the same speed. There's no way—you have to slow down.

It becomes slower compared to the connection between chips. To still increase the speed, you have to use all kinds of crazy design tricks where you compensate for what's going on in the copper on both ends. Sometimes you know, “The copper is going to lose this much energy by the time the signal gets there, so what if I boost the energy first so that, by the time it gets there, I've compensated for it?”

On the other hand, you can get the bits and then count them to see if you can detect some sort of error—for example, how many of these 10 bits are wrong, whether you can tell which 10 bits are wrong, and how to fix them. That's why you sometimes need a digital signal processor, or DSP, on the other end, so you can compensate for all those nuisances that happen on copper lines.

This is a problem because the DSP needs time to count the bits and say, “Bit 7 is wrong. Let's change it from 0 to 1.” It also requires energy because you need to spend silicon and energy on the chip that does all these functions. So this becomes a problem.

With copper, whatever you do, we're now reaching 200 gigabits per second. Maybe the next generation will be 400 gigabits per second. That's exhausted. It's as if physics has reached its limits.

For example, when you put one electrical signal on one end of a copper line, it's almost indistinguishable 5 meters away at that speed. It's useless—you can't even tell what's going on. Now the industry is saying, “Okay, look, copper cable has reached its limit at 400 gigabits.”

Of course, 200 gigabits still works. For example, Rubin racks still have 200 gigabits, and copper cable is good for that. But at 400 gigabits, it's completely dead. Even with 200 gigabits per line, within 1 rack everything is fine. But if you need to move to the next rack, which is about 2 or 3 meters away, the coverage is not enough.

Now everyone is saying, “Wait, so I can't put more GPUs in the rack? Isn't there room there, or would I have to put 400 kilowatts in the rack?” In the cloud era, the energy consumption of a data center per rack was about 20 kilowatts. Now the AI rack consumes 10 times more, and we want to double that because it's very complicated, right?

Cooling becomes difficult. Power supply is getting complicated. So they say, “Okay, the best temporary solution while we figure out all this crazy stuff is to put the individual racks next to each other and then connect them.” Think of it as a giant rack made up of 8 smaller racks, right?

This requires connections that reach across 8 racks. Copper won't do that. That's why the whole industry is saying, “Optics, optics, optics.” We need to move faster. How do we connect the racks? How do we get the reach that copper can't achieve?

6. Co-packaged optics and the push beyond copper

They say, “Wait, I can't use traditional optical modules that plug in.” They're just plugs—you can see them online. They're energy-efficient. NVIDIA never wanted to do that, or they would have gone to optical a long time ago. Why fight copper when there's a solution? It's too much energy.

The solution the industry is proposing is to move this optoelectronic converter right next to the GPU or the switch. That's kind of like co-packaged optics, right? It saves power, but it also solves the reach problem. However, it creates other problems, like reliability. When something goes wrong, how do you replace it? It's right next to the GPU, so there are many other problems.

It's never been done before. Can we deploy this at scale? Can we connect 100,000 GPUs with this technology? Those are the real questions, right?

Logan Jastremski

Yeah, it's very interesting. I really like the physics of all this because, again, you're pushing the envelope of these different materials. As you mentioned with copper, copper has a certain physical limitation in terms of how far you can actually send a signal over longer distances. Potentially, you need repeaters or other devices there, but it's also much more energy-efficient than something like fiber or an optical interconnect.

Optical cable now has much higher bandwidth and potentially uses more energy. It's like, “Okay, how do we deal with optical interconnect? We want the bandwidth to be high, but we don't want the increased energy consumption.”

Vikram Sekar

We don't want energy usage; we want reach. Copper was essentially free. You just ran energy through it, it ran through a copper wire, and it was simple. Optics is much more complicated, right? You have to convert electrical current to optical and do all that.

But now it's a necessity, and that's why you see so much optical content in data centers. You're only going to see more of it. The optics industry has never seen anything like this. In the past, optics was just for long-haul connections, metro networks, submarine cables, and some other applications. But it's never been as important as it is today.

Today, it's a technology that's essential for artificial intelligence and the future of computing, and the whole world is fighting for it because they're saying, “Look, we were never ready for this demand. This industry was never built for this level of demand. You're asking too much.” That's why everyone's running out of capacity.

Logan Jastremski

So you're telling me, Vic, that I'm going to have a terabit of internet connection to my house soon because of the data that's being fed to you?

Vikram Sekar

You don't need that. You can download this podcast unless you're doing homeschooling.

Logan Jastremski

Maybe, maybe. That's funny. Yeah, that's super interesting. Again, I think the coolest thing I've found is that, because the demand is so high, people's willingness to push this has really reached the physical limits of the material.

That's really exciting, because then you have to come up with new, smart engineering solutions to make all these things work. So maybe let's touch a little bit more on the training side, particularly the interconnections, and then we can move on to the inference side, because I think that's also super interesting.

You start to scale these interconnections, and it seems like people want to have almost one giant GPU, if possible. Obviously, if you could avoid bandwidth issues, that would be ideal, but there's a trade-off. So as you start to scale more and more GPUs and more racks, what have you been following? Is it the interconnect between racks—things like Lumentum and Coherent—or are there other things that you think are more interesting to follow as you scale further? You're only as fast as your slowest connection in the data center. We're starting to scale, and we have clusters of up to 1 gigawatt, now up to 2 gigawatts, and again, I think Elon is pushing 10. How much of that is going to be interconnected versus not, and how much will use the old H100 versus the new B300, is still to be determined, but the direction seems clear.

7. New approaches to optical interconnects

Vikram Sekar

Right now, I think the biggest problem today is the interconnect problem. That's why you see these other interconnect technologies, also called “wide and slow.” Instead of running 8 links at 200 gigabits per second each to get 1.6 terabits per second, you can run, I don't know, 50 gigabits per second, but a lot more of them.

Logan Jastremski

How much is that? Like 32, right? Something like that.

Vikram Sekar

So you can run lower data rates but more lanes, like adding more lanes on a highway.

Logan Jastremski

Not like they do today?

Vikram Sekar

No, that's the main discussion right now. We're at this stage where laser technology and optics, when you do 200 gigabits per second per lane, usually combine about 8 lanes together, and you get 1.6 terabits per second per connection.

It's not a 1.6-terabit link; they're sharing it between those 8 lanes. If you have 8 of them running at 200 gigabits per second, that's 1.6 terabits per second. Basically, you need a specific type of optics for that. You need lasers that are specifically made of indium phosphide, which is perfect for lasers and has always been the mainstay of the industry.

The problem is, as I mentioned, nobody has that kind of power. People say, “Look, these types of lasers can connect continents. This is how continents work under submarine cables using this technology. This type of laser can connect data centers kilometers apart, tens of kilometers apart, or even under the ocean. Why would we use this to connect a neighboring rack? Why? Don't you think that's excessive?”

We have no supply. Why don't we make something else that isn't copper? Copper doesn't work. Why do we need to switch to this Ferrari technology that's useful for something completely different? Why don't we just do something else?

Well, we're kind of going back in time, because we're going back to these specific types of lasers called vertical-cavity surface-emitting lasers, or VCSELs. You'll hear most people shorten that to “VCSELs.” These lasers used to do what we're doing now. You can even make them out of gallium arsenide, which is much more affordable than indium phosphide. You can make a GaAs VCSEL, and the supply chain for something like that is affordable.

People say, “Why don't we use this stuff?” The problem is, it doesn't work at 200 gigabits per second. It only works at, I don't know, 20 or 50 gigabits per second. Then people say, “Just put more of them together. Why are you struggling with that? You just tie more cables together, and the speed will come back. Why not just use this technology? We're not constrained by supply chains, and we know that this technology works.”

There's another competitor, LEDs. People say, “LEDs are another light source that can work with optics, so why not use them instead?” But the problem with microLEDs is that they're even slower. You can't push them beyond 5 gigabits per second or 3 gigabits per second. That's the limit. To get that speed, you have to string even more cables together. People say, “No, no, that's a terrible approach, because you have to string hundreds of them together to get that speed.”

Why make hundreds when you can tie dozens together? This is what VCSELs provide. You can just multiply 50 gigabits by 32, and everything is fine. That's the state of the industry right now. There's a debate going on about what's the best way to connect optics in the short-reach realm. Do we really need this Ferrari technology, or is there something else we can use?

Logan Jastremski

Yeah, so everyone is arguing about this. I don't know the correct answer, but this is very interesting. I thought they had a 1.6-terabit link. I didn't know they were sharing it across 8 lanes.

Vikram Sekar

Yeah, because LEDs are also based on gallium nitride, which is another widely available material. You can even create gallium nitride on silicon substrates. You can make really big wafers out of a material like silicon, and you can make a lot of them. It's quite convenient to manufacture. This doesn't require the special material indium phosphide.

The problem with indium phosphide is that it's made on a 3- or 4-inch wafer. I don't know if you know the kind of dough called a stroopwafel. It looks like a small pancake. That's the size of these wafers. It's like a very tiny pancake, and you have to make lasers out of it.

How limited is the supply when you can't even make these Ferrari lasers on large wafers?

Logan Jastremski

So we need more big pancakes?

Vikram Sekar

Yes, exactly. It's also not easy to make these pancakes bigger because they don't provide enough yield. You can't make enough of them to make the process work. There are so many problems, so the industry is saying, “You know what? Forget about it. We'll just do something else.”

The whole industry is in a big mess. What are we going to do about this interconnection problem? This is a very relevant problem right now. To answer your original question about what's happening in scaling, this connectivity problem is the biggest problem that's happening.

Logan Jastremski

The next thing I think concerns the power aspect. How are we going to supply more and more power to these devices? This is a long-term problem, and it's a huge and extremely interesting problem because it starts with energy production. You can talk about everything from nuclear power plants to how that power gets delivered to a GPU.

8. The AI memory hierarchy

There are an infinite number of things in between, from inverters and UPSs to energy storage and capacitors. The supply chain is crazy. I want to go back to energy and even talk about space data centers, but I want to talk about the compute side very quickly, given that this seems to be an area of focus, as you mentioned, and it's more like the Wild Wild West.

Even just looking at the equity market, I think there were a lot of questions about people trying to get from 0 to 1 in memory, just because their stock was one of the best this year. I'd like to ask a little bit about the memory hierarchy that's starting to take shape, with SRAM from Groq and Cerebras, then obviously high-bandwidth memory, potential new tiers with high-speed flash and NAND. How do you think that memory hierarchy unfolds when you have extremely high bandwidth—hundreds of terabytes per second at the top with SRAM—versus the limited capacity that it has?

As you go down that hierarchy, you have much more capacity but limited bandwidth. Where does everything sit in that stack?

Vikram Sekar

This memory hierarchy used to be simple. In the days of processors, you had SRAM in the processor. SRAM was where you stored your L1 and L2 caches. Then you had DRAM, quite simply—how everyone bought DRAM sticks for their gaming PCs and put them in there. It was quite simple.

Then you had flash. An SSD drive is one type. Then there's hard-disk storage and other types, and you can talk about an even larger type of archival storage called tape. You could store everything on tape. This is actually something that's talked about very little when it comes to building AI data centers.

If everyone wants to keep everything permanently, the only way to do it is to switch to tape. It's archival storage, very slowly pulling everything out of there. It's the lowest bandwidth imaginable, but the capacity is almost infinite, and the cost is very low. Tape storage is the other extreme.

It used to be simple, but now you have SRAM. Cerebras wants to put DRAM on top of SRAM, and Qualcomm wants to put DRAM on top of the logic. Then you even have high-bandwidth memory, 3D-stacked memory, and all that stuff in this memory hierarchy.

Then, at the DRAM level, you have to decide where you put DRAM and HBM, or whether you use HBM at all. When you move to flash storage, it's another explosion, because you have all kinds of high-performance flash memories where the performance is tuned toward higher speed at the expense of capacity, like single-level-cell NAND.

Or you can move to higher capacity, like quad-level-cell, or QLC, NAND memory. All these things exist. Even here, there's a whole continuum, because you can have single-level cell all the way to triple-level cell storage.

You can have different kinds of combinations and performance optimizations in all possible styles just within the NAND layer. Then you have hard drives and tape, which still exist.

Now the question is, how do you use all this? For example, what do you do in inference? How do you decide? That's where I think there's still a lot of work to do to figure out how to map the memory hierarchy to the workload that you're doing.

Do you always need the highest bandwidth? I don't think so. Everyone brags about bandwidth, just like everyone used to brag about FLOPS. Now you don't see anyone bragging about FLOPS. Everyone is talking about how we can get 2,000 tokens per second or about 1,000 tokens per second, so it all depends on speed.

9. Useful work per token

Ultimately, the question always arises: What is the useful amount of work being done per token? Do I always need 2,000 tokens per second? No. As you know, I'm talking to you right now, but my AI could be processing some PDFs that I want to read later or something. I have no immediate need for this.

Ultimately, the memory hierarchy should be used in a way that best suits the job it will be doing, as this will control my costs and, therefore, system usage. I'll use this more often if I can control the costs, right? As a business or consumer, I want more work to be done per token.

So run lots of tokens, but run them slowly for the workloads that need that. Run them very quickly if I'm doing certain types of workloads. If I want something coded really fast for an enterprise, then time is money. They want to build a product faster than their competitors, so they are willing to spend more money on tokens because they can get it back by being first to market. It makes business sense for them.

10. How much context does an agent need

That use case requires higher speeds. I think we'll see memory hierarchy being used all over the place to align the use of tokens with the most useful work they can do. I mean, there's a lot of optimization to be done here.

Logan Jastremski

I completely agree. In one of my recent podcasts with Baba Boy [?], I thought about how agent workloads could potentially make a difference.

It seems like the industry has optimized a lot for bandwidth, which makes sense considering you're charging for the token. The more tokens per second, the more you can charge. But, going back to the agent's perspective and what you mentioned earlier about us chatting and potentially your agent being able to go out and do things, it seems like agents don't need maximum bandwidth.

It would be interesting to me if they could have more context. It would seem like you need a larger KV cache. Over time, you forget fewer things. Maybe it goes from 1 million to 10 million. I was expecting the context windows to get bigger, but I don't think it needs to remember everything about me all the time.

When the context arises, it can look for that part of what it knows about me, right? For example, if I'm working out and I want an AI to tell me something like, “Hey, do you know how many steps I took this week?” it can look at my calorie records, steps from my fitness tracker, or anything else. That's enough to store in memory.

It doesn't need to remember anything about my work. It doesn't need to turn to optics and power in data centers when it comes to fitness. I think the context can be broken down. Even for a person, context has sections, right? You don't remember everything all the time.

I think this can be broken down as a problem. What do you think it looks like under the hood? One of my assumptions—and it could be wildly wrong—was that context windows should grow over time because you want your agent to remember more for more specific workloads.

I saw this podcast with Dario Amodei, who said, “Well, eventually I'll have maybe 100 million tokens.” At that point, you can do recursive learning for the context window itself and maybe pin the most important things back into the weights. I thought, “Okay, cool. Let's move in this direction.”

But yes, of course. It seems like a lot of things have been really optimized. According to your point, if one user needs a certain amount of hardware—say, 10 million context windows, and I'm just making up these numbers—you could do 10 people on a smaller model that could potentially provide more logical inferences than one superuser.

People, and I think this is a good thing, are starting to optimize a lot more for the overall user base and to utilize their technology than for extremely experienced users, at least for now.

Vikram Sekar

Yes, that's right. I don't think we know the answer to how much context is enough. We can't answer that question. It seems like more is always better, but at what point does it stop being useful?

Everything that naturally exists in nature has diminishing returns. At what point do we reach diminishing returns in the context window? I don't think we know yet. I think we'll find out soon. That's why I say we're still in the early stages of this. There is so much still unknown that we're still quite early.

Logan Jastremski

What is your guess?

Vikram Sekar

I think about how I think, right? What is my context window? For some things, it's a lot. I remember my children's entire lives—not every moment, but I can trace them back to when they were born. It's still in my brain. But I can't remember what I had for breakfast yesterday.

What context is useful context? Not all context is useful. What is the best type of context to preserve? Not all of it, right? That sets the basis for a useful boundary of context, which has a decreasing value here.

I don't want to forget the youth of my children and everything that is joyful for me. That's why I remember it. But I don't really care what I had for dinner 2 nights ago. That's normal. If it was a delicious dinner, I would probably remember it, but otherwise I don't know.

That information doesn't need to be stored. It all depends on what context you want to preserve and what is important.

Logan Jastremski

I think you also wrote a great article about this with DeepSeek and some of the unique techniques they use with KV-cache caching. When you hit the cache, their costs essentially decrease because you're using the cache instead of doing a prefill, which reduces costs significantly. That was extremely interesting to me because it tries to optimize for as many cache hits as possible.

You can either potentially pass the savings on to the consumer or simply increase your margin because it costs much less than an expensive prefill.

Vikram Sekar

Yes, it's the equivalent of writing in a notebook. This is my NAND memory. It takes longer to search for information in a notebook, but there's a lot more context than I can fit in my head. This is a useful technique.

11. DRAM, HBM, and memory on logic

DeepSeek is essentially writing information into a notebook. This is perhaps the last item in the memory hierarchy.

Logan Jastremski

I appreciate your time. If I have to move on to the next stage, let me know and we can complete it. Regarding the memory hierarchy, is there any part of the hierarchy that interests you more than the others?

Are you really interested in SRAM and what they do there? Is it the interconnect, for example, high-speed memory to have larger weights? Is it NAND and a notebook, so to speak? Is there any part of the stack that interests you the most, or is it perhaps the area where the industry is really doubling down on its efforts?

Vikram Sekar

DRAM. DRAM is the best area to focus on, as HBM, or high-bandwidth memory, is required. At the moment, you can't do without HBM. In terms of technology, this is the most important layer because it provides the best balance between bandwidth and capacity.

But stacking becomes a problem. It becomes too expensive and has questionable future value. How high can you stack? We are already at 16-high. Are we going to go to 20? What next—24? This becomes a bit illogical at a certain point. Maybe we'll keep stacking things.

I think one useful way to do this is to put the memory on top of the logic, and you can see a lot of companies doing that. Cerebras is going to bond an entire DRAM wafer on top of its wafer-scale engine. You'll have DRAM on top of SRAM, fully bonded wafer to wafer. I think Groq will have a version of that as well.

D-Matrix already has its Raptor engine, which has logic on top of memory. Qualcomm has also announced high-speed computing, which places perhaps 2 or 4 levels of DRAM stacks on top of the logic.

Now you're suddenly opening up the entire area of the chip to connections, not just the edge of the chip, right? You have a lot more bandwidth. You get SRAM-like bandwidth with DRAM-like capacity. This is a very important area to work on.

12. The next generation of inference chips

Logan Jastremski

This is super interesting. If you're interested in the investment side of private business, are there any companies that you find uniquely interesting in that regard? As you mentioned, on the funding side, it's kind of the Wild West. People are experimenting with a lot of different things, like memory hierarchy. We're trying to see where everything fits best, figure out different arrangements of things, and work out the connections.

As you continue to scale these things, are there any companies that you find uniquely interesting? I know companies like Etched just raised a big round of funding. Are there other Weka-like companies? I know you guys mentioned Weka on the podcast as well. They're extremely interesting in how they do data sharding.

Are there any companies that you find interesting and that you think are worth double-clicking on?

Vikram Sekar

There are so many of them.

That’s the whole point of tracking all these private companies—startups and other things that are doing logical deduction. It’s very interesting. I like the 3D DRAM approach, so I mentioned that D-Matrix is one of them. I had previously considered SambaNova.

I think their architecture gives you a different kind of optimization. It’s more of a TCO-based optimization, where they can provide services to a different end user. So it’s not always about maximum productivity. I think they have a different customer profile in mind for their hardware.

Etched is very interesting. Although I don’t really understand their low-voltage pin architecture, it seems very interesting. I saw their equipment at Hot Chips, and that was really cool. I would be really excited to see what comes out of Etched.

Then we have MatX, we have Fractile—we have so many companies doing inference that are all doing slightly different things. I thought it was the Wild West of inference chips. This is really fun.

In terms of lasers, you have companies making lasers so that now we can connect inter-chip connections using optics. Why use copper at all, right? There are so many interesting things in this field. The company that does this is called Nubis. They have these lasers, which they call nanolasers for scaling. I find it really interesting.

13. AI's biggest engineering bottleneck

Logan Jastremski

Yes, there are many different things happening. So maybe, if you could push for an answer, what do you find most uniquely exciting about inter-chip interconnects? Do you think it’s ongoing innovation on the memory side, from a stacking perspective? Or perhaps, if you could close your eyes and fast-forward 5 years, what do you think are the most pressing engineering challenges in the data center stack as we continue to scale?

Vikram Sekar

Memory is the biggest problem. Literally, an LLM is a method that requires a lot of memory, and memory is everything. One of the biggest problems people have to solve is how to get the same performance without using as much memory. Once you solve this problem, a lot of things will fall into place, right?

If you don’t have a memory constraint, then you probably don’t need to switch between memory as often, and your interconnect problem will become easier. You might not have any memory-bandwidth issues because of the way it computes everything. Right now, we’re leaving so much compute on the table. GPUs have compute capacity that we don’t use because the network doesn’t transfer data fast enough.

If we can figure out a way to do LLMs that aren’t memory-intensive, that would be very, very interesting, because that would eliminate the memory bottleneck, and therefore the memory-bandwidth bottleneck, and therefore the network bottleneck. The real transformational change happens at the top of the stack, because it’s this bunch of esoteric LLM equations that are driving this whole AI industry, right?

The “Attention Is All You Need” paper from 2017—we implement these equations into hardware. Now, in 5 years, let’s say some really smart people at some university or company research lab say, “Hey, why don’t we do these equations instead?” We can do the same thing, but without all this madness. This will change everything. This is the real turning point that we should pay attention to.

Every time something comes out, like TurboQuant or something like that, everyone thinks, “Oh my God, the need for memory has plummeted, hasn’t it?” Everyone is afraid of memory. So yes, that’s something to pay attention to.

Then we’ll go back to Jevons’s paradox again, increase overall usage, and then we’ll go back to where we started. That’s right, isn’t it? I don’t know when this Jevons’s paradox will cease to exist. Until now, every time there is some improvement, you think, “We’re back to square one. Sorry, we still don’t have enough computing resources. We still don’t have enough.”

I think this is a useful exercise for everyone listening to this, and I would like to do it myself. The exercise is this: find a time in history when Jevons’s paradox didn’t work. As a challenge, your listeners can write back to you and say, “Hey, I found one. This is where Jevons’s paradox didn’t work, and here’s why it didn’t work.”

I would love to have a case like that, because it would tell us when this thing would stop constantly consuming all of our computing resources. I guess the moral of this podcast is: the show goes on, and it’s a deep dive into memory. Let’s try to figure out ways to use memory more efficiently, but it looks like memory is still a bottleneck.

Memory is the foundation of all AI inference and learning. There’s no way to change that, so—

Logan Jastremski

Yes, exactly. Well, I want to respect your time, Vikram, so thank you very much for joining us. I feel like we could keep talking for another hour about all of these things. Again, I really appreciate your clarity of thought, because I think that’s a really distinctive trait of an expert: to explain complex things simply, and you do it wonderfully.

Vikram Sekar

Absolutely. Thank you, Logan. It was a lot of fun to be on the pod. Thank you. Thank you for inviting me.