[BidClub_]
SemiAnalysis · · 63 分钟

第034期 - 快速 Token 之争、TPU v7、Vera Rubin 与 Engrams(AI 供应链,InferenceX)

Jordan NanosCam QuiliciBryan ShanAlec Ibarra

半导体AI与软件技术
YouTube ↗
TL;DR
  • Engrams 将模型架构与内存层级合并为一个协同设计问题。 它在哈希表中学习 2-gram 和 3-gram 的 token embedding,让模型直接取回熟悉的语义,而不是交给后续 Transformer 层重新构建;消融实验显示,Engram 与 MoE 层的混合架构优于全 MoE/无 Engram 或全 Engram/无 MoE。由于 token ID 在计算开始前就已知,从 DRAM 或 SSD 中取数可以与计算重叠,相当于获得“免费层”;但放置得越靠前,模型收益越大,能够用来隐藏取数延迟的时间也越少。
  • 这一设计反映了中国模型架构面临的结构性约束:HBM 越少,就越依赖更激进的压缩、分层与异步卸载。 嘉宾将 Engrams 与 DeepSeek 早期压缩 KV cache、使用 HBM 的工作联系起来,也提到国内加速器的带宽落后于美国系统。更广泛的判断是,中国实验室会继续把 KV cache 及相关模型状态向更低的内存层级下移,而美国实验室过去受到这一约束的驱动较小。
  • AgentX 说明,在真实编码 Agent 工作负载中,KV-cache 管理会成为决定性因素。 通过 SemiAnalysis 内部代理捕获并重放的会话显示,P50 请求接近 100,000 个输入 token,理论缓存命中率超过98%,因为每次工具调用都会向既有长对话追加内容。低并发交互场景下,DRAM 卸载价值有限;但当并发 Agent 达到数百个、吞吐在 40–60 tokens/s 左右,或进入离线批处理、强化学习 rollout 时,其价值就会显现。
  • 在工作负载与定价假设成立的前提下,当前推理经济学可能让最新系统看起来像“印钞机”。 团队的计算器将硬件、电力、数据中心成本、吞吐、并发和 OpenRouter 定价结合起来;讨论中的情景包括 GB200 NVL72 约 40.7%的利润率,以及 GB300/Moonshot 在按牌价的 30%和60%计价时约 50%的利润率。结果仍高度依赖利用率、节点可用性、软件和流量。
  • TPU V7 对外开放,构成了 NVIDIA 与 AMD 之外最清晰的潜在第三家竞争者。 团队早期测试显示,TPU V7 几乎占据 Pareto 曲线低成本一侧的全部位置,只有一个点胜过它,外部 TCO 输入约为 $1.21。软件仍不成熟;原生 Torch 支持、TPU vLLM、megakernel、V8,以及 Google 内部环境之外的 GCP 托管,都可能进一步改善其位置。
  • Vera Rubin 在运营层面看起来不如 Blackwell 激进,但经济性可能强得多。 在约 100 tokens/s 的运行点上,早期讨论认为性能提升可能接近 3×;在高吞吐场景,提升幅度则可能更接近约 1.4×。HBM3 到 HBM4 约 3× 的内存带宽进步,有助于解释曲线中段的内存受限特征。一条硬件加速的查找表量化路径值得关注:8-bit 表值、3-bit key,有效位宽约 3.125 bit,但目前尚未发现重大应用。
  • 关于快速 token 的争论尚未定论,因为 Agent 工作流同时包含对延迟敏感和对延迟不敏感的阶段。 Bryan、Cam 和 Alec 认为,在多个 Agent 并行运行,或 30 秒工具调用占主导时,快速解码的回报有限;尤其当价格接近每100万输出 token 300美元时更是如此,但 Alec 也提醒未来工作流可能改变。TensorRT 在 NVIDIA GPU 上实现每秒数百个 token,强化了专用推理软件的逻辑;不过 batch size=1 的 megakernel、不断变化的 prefill/decode 比例和困难的部署经济性,都让定制芯片的论点更加复杂。
摘要 · 为研究而整理的核心内容

1. Engrams 用可寻址内存替代重复语义计算

  • Cam 的框架是:普通 token embedding 在穿过 Transformer 层时不断累积语义;Engrams 则为 2-token 和 3-token 序列学习 embedding,将其哈希到表中,再把取回的表示注入模型。

  • 讨论中提到 DeepSeek 的相关工作:比较不同隐藏状态深度对 next-token prediction 的能力,以及相对于最终层的 KL divergence。Bryan 的解释是,Engram 将词义信息前置,后续层就不必再动用那么多参数去重建这些信息。

  • 关键消融并不是“用内存替代一切”。损失曲线在 Engram 模块与传统 MoE 层结合的中间位置有所改善,而全 MoE/无 Engram 和全 Engram/无 MoE 两个端点的效果都更差。因此,团队将收益描述为获得了若干“免费层”(free layers),而不是取消 Transformer 计算。

  • 模块放置位置带来了核心的协同设计权衡。更早进行取数,意味着更多下游层可以访问这部分内存,报告中的消融结果对模型帮助也更大;但这样一来,可用于隐藏 DRAM 或 SSD 取数的计算就更少,模块放得更靠后,则有更多机会与取数异步重叠。

2. 内存稀缺正在塑造中国模型架构

  • token ID 在 forward pass 开始前就已知,因此 Engram lookup 可以通过统一虚拟寻址和 CPU 上的固定页主机内存立即启动。它不像依赖中间 hidden state 的检索那样需要等待前序计算,因此可以与模型其他部分重叠执行,使 DRAM、甚至 SSD,成为实际可用的内存层级。

  • 嘉宾将 Engrams 与 DeepSeek 早期压缩 KV cache、使用 HBM 的工作放在一起讨论。更广泛的判断是,凡是可以提前取回并与计算重叠的信息,都天然适合卸载到更低层级。

  • 中国加速器进一步强化了这一动力。国内系统可能使用中国制造的 HBM 类技术,但带宽更低,因此中国实验室会被推向更小的 KV 表示、更强的压缩、更低层级的内存以及各种架构技巧;相比之下,美国实验室拥有更多 HBM,相对于这一约束,Cam 形容其算力几乎是无限的。

3. AgentX 让 KV-cache 问题变得可测量

  • Cam 描述的 Agent trace 是:用户发送任务,模型生成文本和工具调用,工具输出被追加进上下文,增长后的对话再返回给无状态模型。每一轮对话都会让 KV cache“以惊人的速度增长”。

  • AgentX 使用 SemiAnalysis 员工通过代理捕获的编码会话,去除不适合的请求,再将剩余 trace 重放到开源 serving stack 上。团队承认,这并不代表全部编码流量,但认为其基本形态——任务、工具循环、新增上下文、重复前缀——具有普遍性。

  • 最突出的数字是:P50 请求约包含 100,000 个输入 token,理论缓存命中率超过98%。大多数轮次都会复用此前的大量上下文,因此,将请求路由到正确的 prefill server、在 prefill 和 decode 之间传输 KV,以及维持缓存驻留,都需要精细的系统设计。

  • 卸载的价值取决于运行点。低并发、高交互场景下,需要溢出的 KV 不多,而更慢的二级内存价值有限;当吞吐达到 40–60 tokens/s、并发会话达到数百个,或进入离线批处理、强化学习 rollout 时,DRAM 卸载就变得重要。Bryan 另表示,他希望看到 HBM 而不是 VRAM 的缓存命中率处在 20–30%区间。

4. Serving 利润更奖励利用率,而非基准测试领先

  • Bryan 解释,计算器先从每小时持有成本出发,包括加速器、数据中心成本、电力和相关基础设施;随后将这些数字与模型吞吐、并发、利用率假设和 OpenRouter 定价结合起来。它的目标,是把基准测试曲线转换成特定工作负载下不同芯片能够产生的利润。

  • 按讨论中的假设,GB200 NVL72 部署可以实现约 40.7%的利润率。团队还展示了 Kimi K2 和 DeepSeek V4.1 Flash 的其他利润率情景。Jordan 也提到,在 GB300/Moonshot 情景中,按牌价的 30%和60%计价,利润率约为 50%。

  • 质疑本身已经嵌入模型:按60%利用率计算,除非节点停机,系统实际上就相当于接近饱和;而不同服务商可能面临较低利用率、不同软件或不同算力价格。Bryan 表示,部分服务商在 30–40%利用率下仍然可以盈利。Jordan 的意思是,服务商希望在数月内收回机架资本,而不是说每个情景都能保证这一结果。

5. TPU 对外开放使 Google 成为平台型竞争者

  • Alec 的早期 TPU V7 结果显示,这一公有云产品几乎占据整条价格-性能 Pareto 曲线最便宜的一侧,只有一个点胜过它。团队使用了约 $1.21 的外部 TCO 输入,并强调这是第一天的结果,尚未计入后续软件优化。

  • 这一比较并不完全对等:TPU 的芯片、编程模型、可访问性和社区支持,都与 B200 或 B300 存在实质差异。固定尺寸的脉动阵列可能因为 padding 浪费算力,而 Google 的内部模型过去是围绕自身约束设计的,开源模型未必遵循这些约束。

  • 软件是最明确的追赶路径。嘉宾预计,TPU vLLM、更多 megakernel 工作以及即将推出的 Torch-TPU backend,将降低编程模型摩擦;V8 属于 TPU 下一阶段的发展方向,而 V7 本身也会随着 kernel 和公有模型支持成熟而继续改善。

  • Bryan 更大的判断是“大规模对外开放”(large-scale externalization):Google 仍会在内部使用大量 TPU,但也会出售其中一部分。如果 GCP 能在 Google 自有环境之外托管 TPU,生态就可能向更多客户开放;Cam 则认为,这会让 AMD 和 NVIDIA 面对一个真正的第三家竞争者。

6. Rubin、快速解码与定制芯片重新切分工作负载

  • 从 H100/H200 HGX 系统转向 GB200 和 GB300,引入了 Arm CPU、scale-up NVLink switch、新网络和液冷;Blackwell 的 SM100 还引入了新的 FP4 数据类型。Vera Rubin 在运营层面看起来更像渐进式升级:没有重大的新 scale-up 网络或 Arm CPU,液冷也已经部署,但内存带宽大幅提升。

  • 性能曲线取决于工作负载。在约 100 tokens/s 的位置,早期讨论认为 Rubin 的性能或相同交互水平下可服务的用户数可能接近提升 3×;在最大吞吐附近,优势则可能更接近约 1.4×。HBM3 到 HBM4 约 3× 的带宽提升,与曲线中段的内存受限特征高度吻合。

  • Rubin 还暴露出一条新的硬件加速查找表量化路径:使用 8-bit 表值和 3-bit key,有效位宽约为 3.125 bit。团队尚未看到这条路径的专门应用,也明确表示开源生态会在多大程度上采用它仍待观察;与 MXFP4 类似的 scaling 方式以及准确率评估,都是后续关注点。

  • 对快速 token 的需求引发了最尖锐的分歧。Bryan 更偏好多个 Agent 并行运行;Cam 指出,一次 30 秒的工具调用就可能抵消每秒 1,000 个 token 的解码速度;Alec 则质疑投入产出比,认为价格过高。Alec 的历史性谨慎在于,1–2年前他不会预测自己今天的 AI 工作流,因此未来需求很难预测。Cam 的判断是,快速模式会对部分任务有用,而可以批处理的工作能够容忍更高延迟。

  • TensorRT 展示了软件优化的极端形态:将模型的大部分融合进一个巨大的 dataflow 风格 megakernel,在 NVIDIA GPU 上实现每秒数百个 token;一张图表中的点位接近 500 tokens/s,对应的是一个规模相当大、参数量接近1万亿的模型,而且只使用 NVIDIA GPU。代价则是巨大的工程投入、缓慢的模型接入速度,以及偏向交互而非高吞吐 serving 的 batch size=1 特性。

  • 因此,Jordan 并不担心加速器初创公司找不到第一批客户,因为“需求完全领先于供给”(demand is completely ahead of supply);但团队并不认同这就能证明经济性具有持久性。如今做出能工作的芯片已经变得可行,真正更难的商业问题仍包括取得良率、完成服务器组装、 확보分销渠道、运营数据中心、适应不断变化的 prefill/decode 比例,以及与 NVIDIA 的软件和受控 GPU 供给竞争。

完整逐字稿
Jordan Nanos

It’s been about 2 months since the last episode. This week, the AgentX team joined me. We’re talking about Engrams, AgentX, and TPU offload, and we have plenty to discuss. Joining me are Cam, Bryan, and Alec. Welcome to the show, everyone.

Cam Quilici

Hello. Happy to be here.

Bryan Shan

Good to be here.

Alec Ibarra

Good.

1. Engrams

Jordan Nanos

A new article came out on September 18, about 2 weeks before this episode. Let’s start with that. It’s about Engrams: offloading some of the model components, or KV caches, to SSD, and what this means for model architecture co-design. What are Engrams? Can you give us a short introduction?

Cam Quilici

Engrams are basically a way for models to transfer more information into embeddings. The idea has existed for a while, I think. Bryan can touch on more of the history, but, basically, in a Transformer architecture, each token has an embedding. An embedding is what gives a token meaning in a latent space.

As tokens pass through the Transformer’s layers, each layer adds more information to that representation. The tokens are related to one another, and that information flows through the layers into a hidden state.

The Engram idea is basically that, during training, instead of learning all of this meaning across multiple layers, you learn some 2-gram and 3-gram token embeddings. These are stored in a hash table. The lookup is essentially free—not completely free, but close.

If you look at the 2 graphs on the right, the experiments are ordered by depth. On one side, there are all MoE layers and no Engram layers; on the other, there are all Engram layers and no MoE layers. The loss is actually reduced at the midpoint. When there are a few Engram layers and slightly fewer MoE layers, the performance is very good.

That’s shown in Qwen3-8B, I think.

Jordan Nanos

Bryan, could you speak a little more about the history of Engram and its inspirations?

Bryan Shan

This year, DeepSeek introduced a paper arguing that, for language models, representations of word meaning are built in the later layers. They compare the next-token prediction capacity of hidden states at different layers. There was an earlier paper—I forgot what it was called.

Basically, if you calculate the KL divergence between the hidden states in the previous layers and the last layer, the later hidden states for Engrams have a low KL divergence. This shows that, instead of using the later layers of the model to predict the next token, Engram is moving that information earlier. That’s useful because the model doesn’t have to waste parameters trying to reconstruct the meaning of words. They put the second Engram layer in the second layer, because if you put it in the first layer, you don’t have enough time for the retrieval from DRAM or SSD to return, so there’s very little overhead. But in the other Flash model, the interesting part is that they put the layer in layer 1. There was still room for overlap, but not as much.

One really special thing about this model’s co-design is that, when you’re considering the architecture at inference time, the Engram modules can be asynchronous. They can overlap with other parts of the model because Engram module inputs are just token IDs.

During inference, the model receives token IDs. You can use those IDs to retrieve the hidden-state additions, so overlapping the Engram lookup is very natural. In the original paper, they mention a trade-off between putting Engram layers earlier versus later. Putting them earlier helps the model more, based on the ablations, because earlier Engram layers can bring in previous semantic memory for as long as possible.

But the problem with putting Engram layers early is that there’s nothing to overlap with. You have to wait for the Engram retrieval to finish before continuing. Jordan, can you pull up the chart?

Jordan Nanos

Yes, I’m trying to find it. Don’t worry.

Bryan Shan

While Jordan finds the chart, let’s talk about how offload works. Offload uses UVA—unified virtual addressing, or unified virtual address, sorry. Basically, this moves the Engram to pinned host memory on the CPU.

Jordan Nanos

Here, we’re comparing the DRAM and SSD offload versions.

Bryan Shan

Yes, thank you, Jordan. We also tried HBM for Engram. We didn’t discuss it in the article because the cause-and-effect relationship wasn’t very interesting, but HBM is one step in this direction.

Jordan Nanos

Go back, Bryan. The main point is that there’s an HBM controller, and presenting Engrams as a control mechanism around DRAM- or SSD-based offload makes it easier for people to innovate, right?

Bryan Shan

Yes. Over time, this is a natural progression. A lot of this goes back to last year, when HBM was limited. DeepSeek’s work was about reducing the KV cache for the whole model and using HBM.

I don’t know if Engram was specifically an HBM co-design. It was partially a control mechanism. The idea is that the model can retrieve the information it needs from memory without requiring everything to remain in HBM.

Jordan Nanos

Over the past year, a lot of the conversation around 12-high, 8-high, and 4-high Rubin D-spec has been about American labs optimizing HBM bandwidth, capacity, or bandwidth per bit. Chinese labs already have that limitation because of the chips they can access. If bandwidth isn’t available, you need to do something to improve high-concurrency serving or high-interactivity workloads.

Cam, you’re also an expert in this area. Maybe you can explain it.

Cam Quilici

Chinese chips are very interesting. In our previous DeepSeek V4 day-zero article, we talked about the Ascend 950 chip design. They use HBM, although they don’t call it HBM. It’s their own Chinese-built version. Of course, the bandwidth is lower, but they have very creative hardware solutions to deal with it.

At a high level, the trend is that the United States has more HBM, while China has limited HBM and limited chip-to-chip bandwidth. Therefore, Chinese labs are trying to make KV caches smaller and compress them more efficiently, then offload them to lower levels of memory.

Engram looks like a natural extension of that because it effectively gives you free layers. As Bryan mentioned, you don’t need to increase the representation, and it can be offloaded to DRAM. Engram can also overlap with other computations because it starts with token IDs.

It seems like Chinese labs are trying to be more efficient. In the United States, labs have effectively unlimited compute and are collecting more HBM because they aren’t forced to make these trade-offs.

Jordan Nanos

That’s very interesting. It seems like a lot of innovation is coming from China.

2. AgentX Benchmark

Let’s talk about the workload implications. There’s an obvious desire to store KV because many requests can be served at the same time, and the requests are very large. Agentic coding and other coding-for-agent workloads need high throughput for long-context serving.

Agentic coding was a great inspiration for developing a benchmark that more accurately represents this workload pattern. We applied it to all the new models and released the initial dataset and results. This is what AgentX is called.

Cam, what is AgentX? Compared with traditional A/K1K, what changes does AgentX represent? And how does the model architecture differ in terms of efficiency? Let’s use this benchmark to show that clearly.

Can you tell me?

Cam Quilici

Good question. Before, we were doing A/K1K with prefix-only, no caching, and testing at the chip level. Now we’ve changed to AgentX. Basically, this is agent traffic only, isn’t it?

In this setup, you’re looking at an agentic harness using code. What happens is that you send a message, add the output, maybe call some tools, add the tool output to your messages in a streaming fashion, and then keep sending that back to the model. The model is stateless, isn’t it? But as we know, the KV cache grows at every turn at a crazy rate. When you have models with millions of tokens of context length, you need to store a lot of KV cache.

The idea behind AgentX is that we’re not testing chip-level performance alone. We’re testing the whole system and looking at how KV cache moves between chips, how prefill and decode disaggregation work in different circumstances, and how offloading works.

Jordan Nanos

Can you pull up some results on inferencex.com?

Cam Quilici

Yes, of course. We replay the traces, and the results are really interesting. The cache-hit rate—the theoretical cache-hit rate—is very high. Many people are surprised that it’s above 98%.

Jordan Nanos

What does that 98% mean for your system? If you click on “B300 vLLM,” for example, what does that show?

Bryan Shan

The system is shown with offloading enabled. For high-end, highly interactive workloads, this is not very important, because offloading to secondary-tier memory is slow. When you want low concurrency and high interactivity—sorry, I said that incorrectly—you don’t need to store that many KV cache entries, so offloading is not especially relevant.

But when you’re in the range of 40 to 60 tokens per second, with very high throughput, offloading becomes relevant. That could be an offline batch assumption, hundreds of concurrent agent sessions, or reinforcement-learning rollouts. In those cases, offloading is very relevant. For HBM rather than VRAM, I’d like to see cache-hit rates in the 20% to 30% range.

Cam Quilici

I see. So, yes, AgentX is intended to show what an actual workload looks like. I think, Bryan, you may want to talk about the profit calculator.

Bryan Shan

The hourly cost is why this benchmark is so interesting. For Cerebras, our TCO model calculates the total cost of ownership for the accelerator chips. Another way to think about it is: how much does it cost to run the system for 1 hour?

That includes data-center costs, electricity, chip costs, and so on. For InferenceX users, we’re giving you a number that will be shown soon. This is the cost basis. On the router, depending on the assumptions—how many users you have, what the concurrency is, and what throughput you’re achieving—we combine all those numbers and calculate the profit you can make using different chips.

For example, with a GB200 NVL72 serving one of the highest-end models, I believe you can achieve approximately a 40.7% profit margin. With Kimi K2, depending on the parameters, there’s another margin profile. With DeepSeek V4.1 Flash, you can see the trend change under different settings, and in some cases you can’t make a profit. That’s very interesting because I believe the methodology is going to be useful.

Jordan Nanos

The biggest realization is that, under the open-source assumption, you can achieve real profit margins by running these models. Compute cost is normal at this level, but these systems are essentially money printers.

To break this down, with Dynamo and vLLM on an existing GB300, I’m seeing a 60% reduction—or, sorry, we’re assuming 60% utilization. That’s probably too much. At 60% usage, the system is always completely saturated, unless some nodes are down. For Moonshot, at 30% and 60% of the list price, you can save on GB300 and achieve about a 50% profit margin. This is just a money printer, isn’t it?

Cam Quilici

In AgentX, let’s use the assumed cache-hit rate and then look at the OpenRouter data. This is very practical.

Jordan Nanos

Let’s go back for 1 minute and talk about the dataset. The agent-coding request distribution is interesting. We discussed how the P50 request is still about 100,000 input tokens, and that has implications for the cache-hit proportion. Some people might say, “That’s real? How do you know?”

Cam Quilici

We replayed our own data, which is why it’s real. All SemiAnalysis employees have agent-coding sessions intercepted by a proxy. We take those requests, remove the bad ones, do a little post-processing, and then replay them against the original open-source servers.

This does not completely represent all agent-coding traffic, but the actual workload shape is very simple. Everyone has the same basic process: you have a task, you send it to the model, the model suggests something, it decides to use some tools, collects more context, and then writes. You repeat that self-agent loop until the task is complete.

The key point is that, most of the time, the cache is a hit. So how you store and handle the KV cache, how you transfer it, and how you route requests between different prefill servers all have to be handled intelligently.

As you can see, there are a lot of parameters you can set. We use NIXL as the KV-transfer engine. For CPU KV offload, we use vLLM, and we’re only using DRAM offloading. We’re also using the Dynamo router. None of this is trivial; there are many ways to configure it.

But to answer the question, the data represents actual captured coding sessions. That’s why it’s useful.

Jordan Nanos

That’s very interesting. Let’s explore all of this at inferencex.com.

Bryan Shan

When you ask how the per-gigabyte profit is calculated, we take our original traffic, apply the discounted usage rate, and then divide by the original hardware-cost numbers. If people are using the newest chips and the best models, it can be unbelievably profitable.

If you’re listening to this podcast and you have a few million dollars, you might think about buying a GB300 rack.

Jordan Nanos

No, I’m serious. You want to get your return on investment as quickly as possible, ideally getting your money back within a few months.

Bryan Shan

For inference providers, they have their own driver kernels, and they work on OpenRouter. Even if you go back to the page, you don’t need to; even at a 30% or 40% utilization rate, it can still work. For some inference providers on OpenRouter, there’s no change in the result.

Their utilization might be very poor, but it can still be profitable. I don’t know what the compute price will be. I’m not a finance expert, but I think this is a normal level. It’s still very profitable.

Jordan Nanos

At the beginning, people had a lot of questions about this. Alec, can you speak to that? I want to bring you in and encourage you to take this one.

Alec Ibarra

Two things are changing. First, the models themselves are becoming more efficient and higher quality. We discussed how new models can charge higher prices because of their quality, but as the cost of serving them goes down, they can earn more money, especially with KV-cache optimizations.

3. TPU v7

Hardware is also getting better, so there’s more competition. You recently wrote another article about TPUs. Maybe the most exciting announcement is TPU racks. Here’s the picture. It looks like you can rent them on GCP. What’s missing from the assumption? Everyone talks about NVIDIA and AMD, but what happened with TPUs?

Jordan Nanos

Alec, speak to us at a high level. What has happened with TPUs over the years?

Alec Ibarra

We tested TPU v7. Google’s own accelerator appears to be available through the cloud, which makes it the first accelerator you can buy that way. We wanted to do an apples-to-apples comparison with the B200 and B300 and see what the performance looked like.

On the Pareto curve, the TPU v7 points are blue. They’re below the others, which is the important part. Except for 1 point, it’s the cheapest option across the board. That’s amazing.

It depends on the TPU accelerator’s TCO. For the external TPU TCO, I believe it’s $1.21. That’s very good. They’ve extracted amazing performance from it.

This is still very early. These are first-day performance assumptions, so software improvements will make it cheaper. This is the open-source version, though. It’s not a major assumption among providers because they’re running it internally, I think. For public access, this is what’s available.

Jordan Nanos

Yes, let’s talk a little about the trend. Obviously, the TPU is very different from the B200 or B300 in terms of chip design and programming model. How much access do people have to it? How much of a community is there, how open source is it, and so on?

When you compare TPU v7—or perhaps the current generation—with TPU v8, which is coming very quickly, it’s like comparing apples to apples and then apples to bananas. There are a lot of differences. Google’s highest priority is to optimize certain things and make models open source, as we’ve discussed. vLLM-TPU is very new, so what is the trend for TPUs in the future? It seems like they’re going to get better from here, right?

Recently, we’ve seen that TPUs and megakernels can provide a major performance boost. It’s exciting to see what comes from that. Google has produced many TPUs over the years, but they were all used internally, with the TPU engineering team providing full support for those workloads. Now customers will be able to use them as well.

vLLM-TPU is in preview. How is that looking?

Bryan Shan

Two things are becoming clear. First, as we know, we’re seeing the large-scale externalization of TPUs, with v7 and, later, v8 becoming more mature. Google is moving away from keeping all of this internal for its own workloads. It will still use a ton of TPUs, but it will sell some of them.

Second, you’ll see major progress in the open-source ecosystem. Right now, TPU doesn’t have a native Torch backend, right? It passes through intermediate representations in a different way. Torch-TPU is very close to coming out, and that will provide a native Torch backend for TPU and make the programming model easier from a software perspective. You’ll immediately see improvements in the open-source ecosystem. I think that will be a trend for some time.

Jordan Nanos

Bryan, is there anything you’d add?

Bryan Shan

In the past, TPUs were built around systolic arrays, which prioritize data movement. One problem with a systolic array is that it has a fixed starting point and a fixed size. With smaller workloads, you have to pad them, which wastes compute.

Because TPUs were mainly used internally, many models were built around that limitation. In open source, that wasn’t necessarily the case. This was a limitation for inference workloads, and it held TPUs back. As a result, the full power of the TPU wasn’t really being used.

A year ago—maybe two years ago—everyone at Google and DeepMind was bullish because TPUs gave them more compute. I don’t remember the exact number; perhaps 3 million TPUs had been manufactured. Google also counted the amount of research being conducted on TPUs, which was another reason people were so bullish.

If you attend any machine-learning conference, you’ll see Google and DeepMind papers everywhere. Among prominent people in machine learning, almost everyone has had some period of their career at DeepMind.

Cam Quilici

Yes, the fact that TPUs are actually being externalized is very important. AMD and NVIDIA will have a real third competitor, and it’ll be interesting to see how things develop over the next few years.

I want to run private inference support in vLLM—or a private fork of the original public repository—on some permanent clusters. That’s really exciting.

Jordan Nanos

Another thing is cloud TPU deployment. For the first time, GCP is going to host TPUs outside Google’s own environment. That will open up the ecosystem, and more customers will use TPUs through GCP for workloads beyond Google’s internal use.

Cam Quilici

People will actually be able to reduce their costs. If they don’t have to build everything themselves and can go through GCP, they’ll be happier.

Jordan Nanos

Alec, tell us about the GCP console. What are your top 3 cloud consoles? What do you think?

Alec Ibarra

My top 3? I haven’t used many of them, so I can’t really say. To be honest, I prefer the command line. You don’t need to navigate all those crazy user interfaces.

What about Codex or Claude Code? Which is your favorite? Of the 3 major cloud CLIs, which one is best?

GCP, bro. When you’re updating GCP from the command line, you can give it 14 prompts and it’ll do it. It installs a lot of things. It’s free, I think—the software is free.

I think you can now install the whole thing with Claude Code. Or maybe you opened Claude—

4. Vera Rubin

Jordan Nanos

Okay, moving on from TPUs, some new hardware is hitting the inference wires: Vera Rubin. Let’s talk a little about the difference between Vera Rubin and Grace Hopper.

Jensen is sandbagging again, isn’t he?

Cam Quilici

Of course he’s sandbagging. Look, this is continuous progress. People are asking you to serve interactivity, and this is no longer the case. I don’t think that older setup is used anymore, but the current systems are actually serving interactivity in this range, and it’s 3 times better.

That means you can serve 3 times as many users at the same level of interactivity. You can charge for 3 times as many users at the same price, although you’re not necessarily worth that. If you go to the profit estimator, you can see the effect.

Jordan Nanos

Bryan, can you talk more about the architectural improvements?

Bryan Shan

There are more FLOPs and more memory bandwidth.

Jordan Nanos

Yes, that’s excellent. One thing I mentioned before on the show is the transition from Grace Hopper to Vera Rubin. In terms of the physical system alignment, there’s a small improvement going from GB200 to Vera Rubin.

When H100 and H200 HGX systems move to GB200 and GB300, they introduce an Arm CPU. For the first time, they also introduce scale-up NVLink switches. For GB300, there’s a new 400-gigabit or 800-gigabit network. Blackwell’s SM100 is a complete redesign; in the past, there wasn’t a new FP4 data type.

Vera Rubin looks like an even bigger change. There’s no major change in the scale-up network and no new Arm CPU. Liquid cooling has already been physically deployed, so everyone should be familiar with it. On one hand, that makes it easier to adopt. On the other hand, you may look at it and say, “Compared with GB300, I’m not getting a major performance improvement.” There may be less inspiration to move to the new system.

Bryan, can you talk about the initial software-porting experience with Vera Rubin?

Bryan Shan

To back up what Cam said, compared with GB200, Vera Rubin should have token-based performance benefits and cost benefits. Vera Rubin isn’t as big a jump over the previous generation as Blackwell was. I agree with your comment.

We’re still early, but some interesting things have been said about Vera Rubin. We’re about 8 months away. Because of the shortage of VRAM, scale is important. Since the network prioritizes memory bandwidth, labs may be able to use model sparsity and similar techniques to make more GPU memory bandwidth available to users.

That means the capacity of each individual GPU’s memory may not be as important, because aggregate memory bandwidth is what matters. We can’t say much more about the complexity of Vera Rubin yet.

Jordan Nanos

Another interesting new feature of Vera Rubin is its NVFP4 implementation. We’re seeing that for the first time, and it’s very interesting.

For those who haven’t read our article, Vera Rubin’s internal documents show a new hardware-accelerated path for NVFP4. This is already confirmed in the PTX—I forget which version, but we’ve discussed it before. There are also open-source Triton branches and forks implementing NVFP4.

At a high level, they’re doing quantization with a lookup table. If I’m not mistaken, they’re using 8-bit table values with 3-bit keys. If you remember, that comes out to approximately 3.125 bits. NVIDIA provided a special hardware-accelerated path for this in Rubin.

We haven’t seen any applications yet, so it’ll be interesting to see what comes out of it. It may be that there isn’t much specialized usage in the open-source ecosystem so far, but there’s a lot to watch.

Cam Quilici

What is that—MXFP4, or an MXFP4-like scaling quantization? How does it compare? Jordan, thanks for showing that.

What is it specifically used for? It’ll be interesting to see. You can actually emulate it. In Blackwell, you can emulate it to measure how much efficiency you get with high-accuracy weights. You can use evaluations to check whether you’re losing things like perplexity.

A lot of work is coming out of this. Rubin is very interesting in some respects. Even if there isn’t a major architectural difference in every area, performance itself is the point. On the performance side, we’ve seen a remarkable increase.

Jordan Nanos

These are somehow very early results. It’s like, the faster we are, the better results we’ll get, I think. This is before Blackwell comes out, so maybe we’ll get better results. This is how it seems to me, isn’t it? That’s it. It’s crooked.

So, with the base comparison, what is the multiplier? To see which one is on the curve, the point you choose depends. If you take 200 tokens per second for people to serve, that’s impossibly good, because now Vera Rubin and NVL72 can do that. But in the middle, at 100 tokens per second, making it 3 times better is a different story. In high-throughput cases, 3 times better may not be the best; it may be greater than 1.4. So you’re asking whether a 40% improvement is good for throughput.

Cam Quilici

For me, it’s interesting. Oh, sorry, Jordan, but this is CPX for me. I’m reminding you that the sound on CPX is too much decreased, but because of the CPX workload, with the prefill processing and the decode phase, it will do that. Our inference benchmark has no CPX. We hope to get CPX results soon, maybe to check how CPX improves disaggregated workloads.

This is different at the ends of the curve: failures, bandwidth trade-offs, and how to optimize. It makes people think. For me, it’s clear: one GPU with Rubin is a big step in memory bandwidth. When you move from HBM3 to HBM4, you’ll get about 3 times the progress in memory bandwidth. Assuming the curve is memory-bandwidth-limited, the big middle parts of the curve seem to match that 3-times progress.

Jordan Nanos

What do you think, Bryan? This is an early result looking at lookup-table quantization schemes, and that exploration hasn’t been done on the GPU. There are some other hardware-acceleration features that haven’t been used so far. On a per-dollar or per-watt basis, or even a dollar-watt basis, Rubin should be more efficient than Blackwell. We’re seeing evidence of that here.

Bryan Shan

It’s a great chip. The price hasn’t been determined, but it will be priced aggressively. By aggressively, I mean that people are going to buy it because they know it’s expensive. If you get 3 times the performance for 2 times the price, that’s not a bad trade. That’s what everyone wants to do next year.

Going back to the previous profit-estimator calculations, people are thinking, “Yes, I can carry the calculation forward.” It’s clear: “Yes, I want the newest and the best.” NVIDIA has done that again. Jensen’s sandbagging performance again. There are really parts of the curve that are memory-bandwidth-limited, and NVIDIA is telling you, “It should be possible.”

5. Fast Tokens

Jordan Nanos

For the last question, with coding agents—you use many of them—can you speak about your experience using fast tokens? The consequence of these performance curves is that people are going to pay a premium for more than 100 tokens per second. We’ve had some heated discussions on the phone, but I won’t share mine. What are your initial thoughts on fast versus ultra-fast? Bryan, can you start? Are you using Super, Max, or bare? What are you thinking?

Bryan Shan

No, no, I’m not doing that quickly. My personal use case is that I like working with multiple agents at the same time. I use multiple VS Code windows, and in each VS Code terminal I’m using multiple agents, basically doing different jobs.

How quickly the agent finishes isn’t a problem for me because, by the time I return to the agent, the other things I was spending time on should already be finished. But, Cam, you have a different perspective on ultra-fast models.

6. Profit Calculator

Cam Quilici

Yes, I do. It’s very strange because my thought is that, as tasks become more agentic, you have these powerful models that can call tools very well, don’t you? They can bring in more context to complete things and call tools. But tools take time to run. Cerebras can serve 1,000 tokens in a second, but tool usage can take up to 30 seconds. In production, I don’t understand that speed because you’re still interrupted by the tool-usage time.

If you’re creating HTML pages all day—100,000 lines of HTML—that could be very useful, I think. These models, in terms of my use, give me a good hint. No, no, I’m kidding. Alec, what’s your comment?

Alec Ibarra

Sometimes I need to do something very quickly or get an answer suddenly, and for things like that, fast mode would be good. But most of the time, I have multiple agents. While they’re working, I can do something else. I have the code, and it doesn’t need to be immediate, I think. It doesn’t need to be merged instantly, you know? That can happen within 1 or 2 hours.

So I’m not using it that much anymore. It’s very expensive. Is there an ROI? I think so. I’m not sure.

Jordan Nanos

That’s what I mean, especially when you’re interrupted by tool-usage time. This is very fast stuff. I’m not paying $300 for 1 million output tokens, right?

Alec Ibarra

Okay, then let me give you the opposite fact. I agree about my current use case, but 1 or 2 years ago, I wouldn’t have predicted that I would use AI in this style today. So what will it be 1 year from now? I can’t confidently predict that fast tokens of any kind wouldn’t be useful. It’s always difficult for me to predict.

Jordan Nanos

Let’s push this to 2 extremes. On one side, there’s a model that’s really big. On the other, if I may say so, the network is really big. Where does it slow down—the network, the device you use, or the endpoint? If the model is where it slows down, then nothing comes back because it isn’t giving you anything. You’ll be disappointed.

Can models that are 10 times bigger but 10 times less efficient still be more useful? Can we imagine a world where they are? We need them to hurry up.

Cam Quilici

This is a very interesting question. As you scale, you get better results; the scaling laws show that. So the approach to scaling is interesting. That’s a bitter lesson, isn’t it? Scaling computation allows computers to take over.

Recently, I was thinking about why scaling can’t deliver usable token output quickly at the right price. But what is the industry actually doing? For example, Astra 6.1 Loop Transformers are rumored, although that isn’t confirmed. That’s very interesting, isn’t it? They’re showing where the industry is going. It’s scaling, but it doesn’t seem like they’re bringing architectural changes. It’s not a new method, but it is a very interesting innovation.

Bryan Shan

Yes, even Loop Transformers require a lot of computation. How much computation you do for reasoning isn’t free, right?

Cam Quilici

Yes, exactly. On reasoning, that’s a good point. If you’re using Ultra Max, or if you like reasoning, that completely affects output speed. For example, if you need to decode very quickly, that’s important.

To be honest, my prediction is that, for some tasks, it will be possible to route to a fast mode, while tasks where low latency is unnecessary can be batched. If you’re bootstrapping a React app, having fast tokens may be useful. But as Bryan said, if you’re running a performance benchmark all night to climb the hill, and the benchmark takes 50 minutes, then it takes 54 minutes and costs $40. Spending a few thousand tokens on decoding and then waiting 50 more minutes is just a waste of money.

So that’s my prediction: it will be very specialized, but it may still be useful for some things.

Bryan Shan

The other question is intensity. What about models that are really small and really efficient while still being high quality?

Cam Quilici

Let me give you an example. If you’re someone who bootstraps a React app and you’re using OpenAI, Manus, Operator, or Perplexity Computer, you can get instant responses. You think, “Wow, this is really interesting.” Do you have that experience?

It’s an interesting direction, but personally, I don’t think we’re going to arrive at a world of small models. In the history of computing, if you make something more efficient, someone else adds more things. CPUs and RAM have improved a lot, but websites with libraries have also become much slower and more inefficient.

So I don’t think a world using small models—a true small-model world—will really come true. But I could be wrong. If I’m proven wrong, that will be interesting. It would be good for the environment.

Alec Ibarra

That’s my short tangent. Congratulations. When you’re using all these different agents in parallel, the return time may not really matter to you. Even if it’s 10 times slower, you can still get the result eventually. For me, though, I have doubts. If you go to sleep, wake up in the morning, and still haven’t received a response, that’s not good.

So, in the future, there’s a place for every kind of fast token. Maybe not all of them need to be fast, right? Also, Jordan, in America, even Big Macs don’t have a fixed price, do they? So would you pay up to $100 for 1 million output tokens? I don’t know. I really don’t know.

Jordan Nanos

Of course. I agree. Speed is important, to some extent. Jensen has always emphasized this, and he also said in the keynote that fast tokens are excellent tokens. Obviously, he is inclined to say that because he sells GPUs.

But, to some extent, that's right. If you can get instant output, I think that can be useful in some cases because you get an answer quickly. But I don't know. Time will tell.

Speaker 2

Who knows? Let's see. Okay, let's talk about this—we need to talk. That was a perfect lead-in, man. I was going to conclude this, but let's do this. Let's end it.

7. TileRT

No, maybe we should talk about TensorRT. I don't think so, because there are many people talking about ultra-fast models. When OpenAI's DevDay announcement came out, everyone said, “Oh, okay, OpenAI and Cerebras have a deal.” Over the next few years, they are promising a megawatt-scale deployment.

Jordan Nanos

Obviously, Cerebras is impressive, but TensorRT is also very capable. From our perspective, if you look at the models and the optimization TensorRT is showing, they are using similar techniques. We're seeing 200 tokens per second here, and, as shown on the screen, a reasonably large model approaching 1 trillion parameters can reach up to 500 tokens per second on NVIDIA GPUs alone, right?

Speaker 2

Yes. For those who don't know, TensorRT is basically an inference engine optimized for ultra-low latency. It is designed to achieve as much interactivity as possible without needing specialized SRAM architectures. These results are bounded by physics, of course, but they show what is possible with highly optimized inference engines.

There is also a result from AMD. Going forward, as you look at different use cases, inference engines will become more specialized. I think that's the trend we're already seeing.

TensorRT is very cool because you don't need a new chip. You're building on the best chips. How does that work? Can you speak a little about it? NVIDIA has SRAM in its GPUs, obviously, and there are many ways to configure the serving setup, write special kernels, and provide model-serving support. Technically, let's talk about it. For high interactivity, this is great. Bryan, how does it work? But for high throughput, this is bad, isn't it?

Bryan Shan

Yes. With TensorRT, you can potentially see the entire model as one big megakernel. Therefore, there is less of an interactivity problem because the data flows through the model. Companies such as Cerebras built dataflow hardware—dataflow machines. That is a little more difficult to do with GPUs.

With TensorRT, you can run at a batch size of 1. Model releases are actually very slow because getting that megakernel to run requires a great deal of software effort, and GPUs are not particularly friendly to this kind of dataflow. But GPUs are very common, which makes it interesting. They can use specialized hardware to replace some of that, in fact.

For an AI agent to read anything and then enter a giant megakernel seems impossible. How much of this is actually possible? But it works that way: you get good accuracy and high interactivity, right?

Cam Quilici

Yes, but people were doing this. That's for sure. Of course, it is important for agents because there is a lot to do, and you want to be able to do it easily. I personally believe in it because it can be checked very easily. We've seen many other kernel competitions on GPU leaderboards, and we've seen agents do this.

This is still very new, though. Recently, I've seen these kinds of megakernel ideas coming out on MLX. Cerebras recently released SLS on MLX. There have also been other developments—for example, TPU megakernels. The idea is the same: use the entire model as one megakernel. That's very interesting.

8. Chip Startups

Cam Quilici

Jordan, I have a question for you. You don't have to name names, but if you're a chip startup whose whole thesis is fast tokens—like the megakernels you're talking about, such as Jalapeño—are you worried about interactivity in this area changing faster than general-purpose GPUs overall? What do you think?

Jordan Nanos

I'm not worried, because demand is completely ahead of supply. It is possible to build chips, and anyone can find a customer. I have faith in that. People need all kinds of chips.

In the long term, these chip startups will be able to deploy what they're working on, and that will make or break their businesses. If you have a large business, you should be able to handle the supply chain and build a group of chips. You should be able to get good yield, build servers, build a data center, and train people to operate all of it.

In some ways, getting something to work well is much more difficult than building it. If the interactivity is high enough, and you have both high interactivity and high throughput, I'm here for it. I think that's the case.

If you don't need a lot of dynamic batching, the ratio between throughput and interactivity can still work. If you reach 1,000 in interactivity, you'll get 1,000 in throughput. If you reach 2,000, your throughput is 2,000. Therefore, I will favor custom silicon that provides as much high interactivity as possible.

But that will be very expensive. For example, if you use a batch size of 1 and get 1,000 tokens per second, each user is paying for the entire system. You can do one thing per batch, but if batch size 1 is really fast, you can pipeline everything. You can achieve high effective throughput even without dynamic batching. So, by that definition, fast tokens do not necessarily mean inefficient tokens. It doesn't seem that way to me.

Cam Quilici

I think it depends on the hardware design. As they say, it is expensive: HBM is caviar, while SRAM is even more expensive. I don't understand the energy-consumption implications. That will be interesting.

Another concern is that many alternative chip startups, without naming names, rely on general-purpose GPUs for prefill. That is fine, but your argument is that chips are limited and people need more of them. If a chip cannot run efficiently, though, it still seems that you need an NVIDIA GPU. I don't know.

Jordan Nanos

Yes, but I think the meaningful setup is to have prefill and decode as a pool. As the pool size increases, the traffic can be divided between prefill and decode in a disaggregated setup.

If you know the prefill-to-decode ratio for the entire data center in advance, then that is absolutely possible. But that is a problem, because the workload is going to change, and inefficient silicon will go to waste. If the workload pattern changes and you need more decode or less decode, you have to account for that.

Chips are monolithic. If you can change the setup, that makes more sense.

Jordan Nanos

In all circumstances, if you rely on NVIDIA, one of your partners has to be a GPU provider. If you buy NVIDIA GPUs and more of them, business decisions have more influence on your ability to win than technical decisions.

NVIDIA is clearly taking over. All of these chip startups need some technical justification for what they are doing. Otherwise, why would someone own $20 billion worth of chips that are not operating? It is impossible to build something that works for less than the price of an NVIDIA GPU. Anyone else can do the same thing.

The question is whether this is technically beneficial to someone, or whether it creates demand signals for NVIDIA. This is a business issue. Do these companies have enough capacity to build and deploy enough chips to run the workloads? NVIDIA's GPU exports are controlled, and data centers want to buy those GPUs and build around them. Stopping that is a very large challenge.

Cam Quilici

Yes, I think so too. At this stage, chip startups generally say they can tape out a chip. It has been built. It is proven. CPU, GPU, and custom silicon have been constructed by everyone. MTIA and Maia—and even the hyperscalers—have proven that they can build general-purpose silicon.

You don't need to build this specific silicon for decode. If necessary, build a second chip specifically for prefill. Build a prefill-and-decode chip, right? Habana Labs has done it, I think—or at least claims to have done it.

Yes, the design is not free. RTL is much cheaper than it used to be, though. The software moat—NVIDIA's CUDA software moat—takes time, effort, and money to develop. If you spend that time, effort, and money on alternatives, you may still not reach NVIDIA's performance standard.

I hope chip startups will do the pre-silicon work and then build silicon for decode. Whether they work with NVIDIA, Broadcom, Marvell, or another vendor, they will have to figure out the trade-offs involved in working with multiple vendors and removing some margin.

I hope they can deploy enough of these systems and produce enough chips, but that will be very difficult because distribution, networking, and data center operations are extremely important.

Jordan Nanos

But there is excitement, man. During the conversation that somehow got buried, I know, TPU inference, for me, is the most exciting, as a matter of fact. Because it is a one-way street. It’s like, did you know? They are TPUs. They are now starting to sell externally, and then they’ll say, “We stopped everything for us to keep going.”