第001期 - Claude Code、内存狂热、CPU回来了 | Jordan Nanos、Doug O'Laughlin、Myron Xie
- 内存周期已从1996年以来最糟糕的下行——当年“让行业三分之一破产”——剧烈反转为历史性短缺,短期内几乎看不到缓解。 Myron Xie解释了其中的机制:HBM每片晶圆产出的比特数只有传统DRAM的几分之一;每增加一颗GPU/XPU,都会把更多晶圆拉向HBM;而后疫情时代几乎没有新增晶圆投片或洁净室空间——“需求正在爆发,但我们几乎没有增加任何供给”。Doug O'Laughlin:内存和NVMe价格单日上涨约20%,环比上涨约100%,而且“至少还会紧张几年”。
- 经典的周期见顶信号——Samsung最后投降、最后增加产能——这次可能失灵。 Doug表示:“我认为他们会扩产,但仍然不够。供需缺口是有史以来最大的,而且根据我们看到的情况还在扩大。”他的“银河系大脑式判断”是:EUV瓶颈可能把原本预计约2035年出现的3D DRAM提前到约2030年,但不会早于2030年。
- CPU回来了,甚至 Intel 仍然——或者大概率将会——售罄,“这说明供给很紧”。 讨论把 Intel 定位为边际供应商;Arm“构成了前所未有的更大威胁”,而 AMD“仍在持续碾压”。Phoenix CPUs尤其值得关注,因为它可能改变 Arm 商业模式的适应方式。
- 据 Doug 抓取的每日序列,Claude Code如今生成了公开 GitHub 提交的约4.7–4.8%;Anthropic还在其300亿美元融资、3800亿美元估值的新闻稿中,链接了包含这一估算的文章。 这一占比从1月20日的约2%升至2月5日的约4%,目前仍在增长;Doug是在一条推文提醒他 Claude Code 会签署自己的提交后,才搭建了这套追踪器。
- Opus Fast 和运行于 Cerebras 的 OpenAI GPT-5.3-Codex-Spark,是定价在每年2400美元 Max 级别之上的价格弹性实验——“鱼子酱交易”。 Opus Fast按token收费约为6倍,速度约为2–2.5倍;Jordan Nanos形容 Cerebras 这笔交易是“花10倍的钱换4倍性能,或者更准确地说,花20倍的钱换10倍性能,差不多就是这样”。Jordan说 Fast 模式已经对他产生了效果:“现在我已经对 Fast tokens 上瘾了……我觉得自己回不去了。”
- 中国模型周带来了真正的突破:MiniMax 2.5“达到 Opus 性能……价格只有1/10”、GLM-5,以及最重要的 ByteDance Seedance/CDance 2 视频模型——“动画搞定了,熟透了”。 Doug称这是首次“彻底击败”Google Veo/Genie长期领先的视频模型,并提出一个阴谋论:“监控国家让所有人都拥有更好的训练数据”——中国监控占全部硬盘需求的10%。
- Jordan的存储判断是:日志、合成数据和生成视频将推动SSD/HDD需求爆发,而不是KV cache卸载;后者在生成视频中的留存比例可能“低于1%”。 他还警告,不要直接套用 Jensen 每颗GPU配16TB存储的估算;按此计算,ICMS只占硬盘出货量的不到0.7%——“这是对 ICMS 的正确结论,但不是对硬盘需求增长的正确结论。”
1. Claude Code的提交占比逼近5%——相关文章还进了 Anthropic 新闻稿
- Doug讲述了事情的起点:他看到一条推文——“你们这些蠢货,提交信息全都写着是 Claude Code 提交的,你们可以把这个关掉”——于是询问 Claude Code 能否抓取 GitHub,并建立了一套每日序列,统计由 Claude Code 完成的公开提交。从1月20日核心研究图表中的约2%,到2月5日公开文章中的约4%;“现在已经几乎5%了……4.7%或4.8%”,而且仍在增长。随后,Anthropic在其300亿美元融资、3800亿美元估值的新闻稿中,链接了包含这一估算的文章。
- Doug的紧迫感,用他原话说是:“只要有人在关注,就能看出这一点……我们必须成为第一个。”付费数据会继续向机构订阅者提供,未来可能还会推出公开仪表盘。
- 谈到AI辅助写作,两人都拒绝让模型完全生成文章。Doug用 Claude Code 做提纲和删改——“这一段里最不相关的部分是什么?”——但坚持“用手指敲出人工token”;“这是我们最后的手工艺……我们手里只剩这块手工搅打的黄油了。”Myron只负责润色措辞。值得保留的一点行业内幕是,很多精彩文章其实出自 Myron 之手——“你觉得这些文章是 Dylan 写的吗?不是,兄弟。这是一个团队。”
- 一个关于agent使用的实时数据点:Doug曾在一项任务上同时运行7个agent,打开7个窗口,其中一个窗口又运行了7个子agent。
2. CPU回来了:“甚至 Intel 都卖光了”
- CPU讨论的关键证据在于边际供应商:没人会想到,全球最大、最成熟的算力装机基础会出现短缺,Intel 反而获得需求支撑。节目嘉宾表示,Intel 目前已经售罄,或者大概率即将售罄——“这说明供给很紧。”
- Doug的判断是,Arm“构成了前所未有的更大威胁”;Phoenix CPUs正在改变 Arm 商业模式可能的适应方式;与此同时,AMD“仍在持续碾压”。即使竞争者更多、市场更加碎片化,Intel仍很可能售罄。
- Jordan的结构性判断是:AI增长体现为通过API、以token形式消耗算力,而不是像云计算时代那样由财富500强企业自行部署虚拟机。根据 Microsoft Fairwater 园区的照片,CPU数据中心支撑着低成本GPU数据中心。Hock Tan接手后,VMware已经被“掏空”;至于 Pat Gelsinger 手臂上那枚 VMware 纹身,依然无济于事。
3. 内存狂热:从1996年以来最糟糕的周期,剧烈反转为最好的周期
- Myron解释了供给机制:HBM每片晶圆产出的比特数只有传统DDR DRAM的几分之一——TSV禁布区侵蚀裸片面积,HBM要达到目标性能时良率“非常糟糕”,而8层或12层堆叠在封装环节还会进一步损失良率。每增加一颗GPU/XPU,都会把晶圆拉向HBM并压缩全行业总比特数;CPU增长则进一步收紧传统DRAM。与此同时,新增晶圆投片几乎为零,因为后疫情余波过后,“实际上已经没有多少洁净室空间可以真正放进设备了。”
- Doug回顾称,这轮下行确实是1996年以来最严重的一次,内存行业还做了一件历史性的事:不是承受设备闲置成本,而是直接关掉机器——“不,不如想办法说服所有人把它关掉。”如今已经没有地方可以重新打开供给水龙头。内存和NVMe价格单日上涨约20%,环比上涨约100%,“有一点恐慌性买入,真的只有一点点。”
- NAND供给比DRAM更差,利润率也更低,因此厂商会优先把新增产能分配给DRAM——“至少几年内都会疯狂。”
4. Samsung投降是见顶信号——但这次仍然不够
- Jordan先介绍历史规律:通常当 Samsung 在 Hynix 和 Micron 之后最后增加产能时,周期就见顶了。Doug说,上一个周期里,Samsung顶住要求减产的压力,较平常多坚持了大约三个季度;最终减产标志着“周期的绝对底部”。他现在的判断是:“我认为他们会扩产,但仍然不够。”根据 SemiAnalysis 的研究,供需缺口已是有史以来最大,而且还在扩大。
- 晶圆厂上方真正的约束来自 ASML:ASML每年能交付的 EUV 设备数量有限,逻辑芯片产能同样紧张,而设备应如何在 TSMC 和内存厂商之间分配,仍是内部激烈争论的问题——“这是 ASML 战略团队要回答的问题,他们或许应该买我们的模型。”Myron认为,既然逻辑芯片需要内存、内存也需要逻辑芯片,系统就应该以完整系统的配比为目标进行优化;但“这个交换比例本身说实话非常难……我不认为有很多人知道答案。”
- Doug现场估算,并保留了他的犹疑——“我是在凭感觉估”:先进逻辑芯片每10万片晶圆投片对应的资本开支,可能是DRAM的约2–3倍;但如果DRAM的交换比例约为3倍,HBM可能就是全球最昂贵的晶圆——“我认输了。我觉得HBM最难。”
- 更具“银河系大脑”色彩的推论是:EUV挤压“可能”把原本预计约2035年出现的3D DRAM提前到约2030年——“明确说,不会早于2030年。”原因在于3D DRAM不使用EUV,能够释放整个系统层面的吞吐量。
5. 鱼子酱交易:Opus Fast 和 GPT-5.3-Codex-Spark 是定价实验
- Jordan提到仅使用SRAM的加速器:花约6倍的钱换取约2.5倍性能。Myron接着给出本期最精彩的比喻:“这就像说鸡蛋贵了一点,所以我要早餐吃鱼子酱。”SRAM的成本远高于DRAM,而逻辑芯片产能本就紧张,但“市场确实有付费意愿”。
- 市场表现是:Opus Fast的token价格约为普通版本的6倍,速度约为2–2.5倍;OpenAI GPT-5.3-Codex-Spark运行在 Cerebras 上,Jordan形容这是“花10倍的钱换4倍的性能,或者更准确地说,花20倍的钱换10倍性能……差不多就是这样”。Jordan认为,各家实验室正在突破每年2400美元 Max 级别,沿着需求曲线继续向上定价——“一定存在更高的价格层级……他们正在摸清价格弹性曲线。”他的亲身体验是:“现在我已经对 Fast tokens 上瘾了,我不觉得自己回得去,老兄……你把我从海洛因推到了芬太尼。”
- Myron的平衡判断是:目前 HBM 仍然拥有最优的成本—带宽—密度组合,而 GDDR 的每GB价格也正在逼近 HBM;非 HBM 架构“从来都有取舍”。
6. 中国模型周:“我人生中最中国的时代”
- Doug基于11月或12月读到的一篇 SCMP 文章,推测 DeepSeek V4 会在美国总统日发布。Jordan认为总统日在中国并不知名,Doug则坚持自己的发布时间理论:DeepSeek专门选择美国节假日发布——感恩节、平安夜——“就是为了狠狠地恶心美国。”他还说:“如果他们真的想在我们头上耀武扬威,就应该放在7月4日发布。”
- 这周发布的模型包括:MiniMax 2.5,“达到 Opus 性能……价格只有1/10”,激活参数约10B;GLM-5,由智谱推出的“一次不可思议的发布”;以及一家类似中国 Indeed/LinkedIn 的公司训练基础模型,在 Soy Bench 上达到约70%。此外还有 ByteDance 的视频模型,讨论中先称为 Seedance,后来又称为 CDance 2,能够生成角色保持一致、时长达1分钟的场景。Doug说:“这已经不再是乱码了。”它带来了类似 Opus 4.5 突破时的能量,“至少动画搞定了,熟透了”,并首次“彻底击败”了 Gemini 长期以来在视频领域领先的 Veo/Genie。
- Doug关于原因的阴谋论是:“监控国家让所有人都拥有更好的训练数据。”中国监控占全部硬盘需求的10%,这一规模足以让他认为,疫情封控进一步伤害了 HDD 厂商,因为这部分需求本身就很重要。
7. 存储需求来自视频和日志,而不是KV cache——还要警惕配套率
- Jordan一直在回应业内人士提出的问题:模型处理的所有内容都会写入日志,“他们不会丢掉数据”;合成数据的输出会被保存;每台摄像机都会写盘。数据源的规模是模型活跃参数或输入的10倍到数百倍。生成视频中,KV cache占比“低于1%”。
- Doug用 iPhone 存储作类比:打开你的 iPhone,约20%是应用,约70%是照片和视频。“这就是原因。”像 CDance 这样的海量视频生成,以及更重要的发送和观看,才是真正的存储驱动力。
- 关于经验法则,SemiAnalysis重建了完整BOM——每个端口、东西向流量和南北向流量——用来检验1.5倍收发器配套率,结果确实是1.5倍。“一些非常强大的经验法则,足以让你在人生中走得很远。”到目前为止,资本开支,或者至少是每GW对应的收入变现,仍是最干净的衡量标准。但市场流行的配套率捷径——Jensen在 CES 上提出的每颗GPU配16TB存储——意味着 ICMS 只占硬盘出货量的不到0.7%,这“是对 ICMS 的正确结论,但不是对硬盘需求增长的正确结论。”
- Jordan说 Jensen“很擅长在恰当的时间说出正确的话”,把市场对内存的恐慌引向“连接在他的GPU上的内存”。Doug认为,1000亿美元 OpenAI 交易“后来发生,是为了压制 AMD”;他还提到 Cerebras 发布 Spark 时宣称:“NVIDIA 仍然是我们所做一切的核心。”最后一句是:“这是 Jensen 的世界,我们都生活在其中。”
All right. Hello, everyone. Welcome to episode number 1 of the SemiAnalysis weekly podcast. We don't know exactly what the name is or if this is ever gonna be public, but we're testing things out, and we're gonna talk about what happened in the week that was. Today's Thursday, February 12th.
We had 3 articles that came out on the free-tier newsletter: Claude Code Rising, CPUs Are Back, and Memory Mania. We'll try to talk about all 3 and other stuff that's going on. By way of introduction, I'm Jordan Nanos. I'm here with Doug O'Laughlin and Myron Ji—I don't even know how to pronounce your last name properly. Ji?
Xie.
Yeah.
Have you ever heard xie xie?
Yeah.
Thank you in Chinese?
Yeah.
My last name is literally that.
Thank you for coming on this podcast.
I'm a very grateful person.
Thank you for coming on this podcast today.
Thank you for having me.
Yeah.
So is it like “thank thank” or “thank you, thank you”?
If you say xie xie, that's “thank you.”
Okay. Let's keep it on topic. This is not Transistor Radio Lite. So let's go through each of the things.
Okay, great.
1. Claude Code Goes Mainstream
I think first, Claude Code, bro. I don't wanna talk about Claude Code. I've talked so much about Claude Code. But I do wanna talk about the fact that we're on the press release. I think it's kind of a big deal. It's very rare to get linked on a press release, anyway, and we're on Anthropic's raise press release. It's kind of cool.
Yeah.
Yeah.
We meaning you. So, Doug, you made a chart that made it first into our Core Research product a few weeks ago, and then made it into the free-tier newsletter. It tracks how many contributions to public GitHub repos are being authored by Claude Code—how many commits are being made by Claude Code—because it signs off as a contributor when it sends a PR. That rate went from about 2% to 4%, and then they included that in the announcement release.
Yeah.
So walk us through that. How'd you get the idea to do that in the first place?
I read a tweet where it said, “You idiots, it all says it's committed by Claude Code. You could turn this off.” And I was like, “Wait, can you scrape GitHub?” Then I just asked Claude Code, and that was it. I was like, “Oh my God, can we make a daily series?”
Then I realized, wow, this is like—I’ve definitely been rushing to make sure we were the first to put that chart out, because I was like, dude, anyone who's paying attention can figure this one out, especially with Claude Code. But I was like, “We need to be the first.”
I don't think this is exactly news. We're definitely gonna be publishing the data for our paid subs, or our institutional research tier, but we'll probably have some kind of public dashboard as well, or some kind of public-facing information. It's already almost 5%, right? At least according to the most recent one we did, it's 4.7% or 4.8% or something like that. It's still growing.
That's how I got to the idea: just doomscrolling, and I was like, “Wait, wait, wait, wait.” Then Claude Code and I made the chart and, obviously, wrote the article.
How much of the article did Claude write, in addition to making the chart about itself?
I don't think—I personally still want to be writing with a human voice, right? We gotta train these models on something. We can't just be sloshing all this stuff back around.
I do like it for outlining more than anything else. I think the very first round of it was when I was still at peak excitement about Claude Code, so I dictated most of it. Then I had the dictation cleaned up by AI, and then I wrote, “Okay, make a slop outline based on this,” and wrote over it using my voice—just manual tokens, with my fingers.
Afterward, we sent it to the powers that be, AKA the internal SemiAnalysis brain squad. People wrote in sections and gave feedback. At some point, I think the most heavy lifting that Claude Code did—and I did all this specifically on Claude Code, screw vanilla chat or anything anymore—was mostly the outline stuff.
When you write something so much, you just get lost as hell. You're like, “Okay, this all makes sense in your weird little brain,” but I'm just like, “Hey, just do a fresh go-through of each piece and walk me through this.”
I do sometimes have it do some editing by asking, “What's the least relevant portion of this paragraph?” I usually like to rip stuff out. But that's just me. Sometimes SemiAnalysis loves to add stuff in.
Are you writing with Claude or with ChatGPT or Gemini?
No. I agree with Doug. I like having a human voice. Although sometimes, for certain paragraphs, I'm like, “How do I express this concept?” If I use an LLM and it's like, “Okay, that's a really good way to express it,” that's helpful.
But I mostly use, to Doug's point, a human voice. I wouldn't generate a whole piece from an LLM. It's just specific ideas I'd flesh out with an LLM. If I want to make this sound more elegant, I'd use an LLM.
Yeah. Makes sense.
Maybe I'm too much of a bootlegger.
I think it's our last artisanal point, you know? It's our hand-churned butter. That's all we have. This hand-churned butter is all we got.
I also don't think people realize that Myron writes so much of the bangers at SemiAnalysis—essentially anything accelerated. Do you think Dylan's writing these articles? No, brother. It's a team.
People still believe, even as late as 2024, that it's just Dylan writing everything: “Dylan, I just love everything you write.” And I'm like, “We just put out 3 10K reports this week—like 5K-word reports.” No, there are a lot of people on these author bylines, and there's definitely a lot of heavy lifting.
Obviously, Dylan's giant brain is behind so much of everything we do here. But Myron's an OG SemiAnalysis person, and he's written some of the greats.
When you read something like one of our big 20K- or 10K-word bangers that obviously has a lot of contributions from various people, can you tell whose voice it is? Who wrote this? Did Dan write this? Did Jeremy write this?
Yes, I can kind of tell, at least.
I'm not at Doug's level. I can tell certain voices, or that certain people use certain jargon, or choose to use the ampersand instead of saying A-N-D, or choose to reference it.
There's also British spelling. British spelling definitely pops up.
Oh, yeah. That's a big giveaway.
That's a Myron thing—a Myron giveaway.
I'm pro on that.
That's kind of his British spelling.
Most of it. Actually, our Prime Minister, Mark Carney, is going on a crusade right now, making all of his employees use the British spelling of many common words. So it's interesting.
Wow.
It's a return back to the—
That's true, though. That's like—
Long run.
The reverse of Dylan.
Yeah. The data center, centre, or whatever. And my favorite is, I think Jeremy has the most typos, because it's very clear that his entire keyboard is configured for French.
You think his autocorrect is gonna catch the English typo? No way, bro. So he definitely slips the most typos, I think.
Well, the other thing for Jeremy is that if an entire sentence is highlighted with a link to an article, followed by the next sentence also being highlighted with a link to the article, so that the whole block of text is just links to previous articles, that's a Jeremy play for sure.
Yeah, that's definitely a Jeremy play. There are so many people who go into all these things. It really does help.
Actually, I think that's how all these articles start. Most of an article is like a 2K-word thing, a 1K-word thing, and then we're like, “Hey, it'd be really nice if you helped with this section.”
Someone helps with the section and adds 1K words, and then each incremental thing gets added—another 700 words or whatever. Then all of a sudden we're like, “Wait, this is 6,000 words.”
Whenever someone says, “Oh, let’s do a brief article on this,” it ends up becoming a Frankenstein.
Dude, I think I’m proud. I feel like the Claude Code one was really short because I was the quarterback and I just rammed it out. I was like, “We need to send this chart out. We need everyone to use our number.” I needed that so badly, so I definitely rammed that one out.
Yeah.
So whenever—
Let me try something here and show a bit of the backstory, just because we’re talking inside baseball about how SemiAnalysis works right now. Here’s the first chart, which went out on January 20th in Core Research, and then here’s the one that went out publicly. So just 2 weeks later—this one’s February 5th—
Yep.
We’re all the way up at around 4%.
Yeah.
And—
Yep.
Here’s the Anthropic press release we were referring to, where they announced their $30 billion raise at a $380 billion valuation, and a link to the article that makes that estimate. So, just for reference, some articles are going to the free tier and not to subscribers first. In other words, the CPU article that we just put out about the return of CPUs and their importance is a very different article. It’s a bit of a history, as opposed to a comment about what’s happening right now, or at least it leads in with history or something like that.
Yeah. I think it’s mostly because CPUs are history, right? I mean, sorry, that’s—
Not anymore. They’re back.
Not anymore. They’re back. Yeah. I definitely think that—
Even Intel.
Yeah, even Intel.
Even—
I mean—
Intel is actually—
That’s the—
Yeah.
Yeah.
That tells you things are tight.
Yeah. That’s actually been kind of the story across the entire accelerator space, right? The marginal supplier actually benefits massively, and I just don’t think anyone would have ever guessed that we would hit the largest, most mature install base of compute CPUs and be like, “Yeah, actually, we’re going to run out of supply.” The marginal supplier, Intel, is actually going to catch a bid. It’s been kind of incredible.
I think part of the story that’s interesting, too, is that Arm is such a bigger threat than it’s ever been, and that’s been the other part of the story that’s very interesting. AMD continues to crush; it’s just kind of fragmenting. But the Phoenix CPUs are definitely an interesting change, especially in terms of how Arm’s business model is going to adapt. Even with more competition, it doesn’t matter: Intel is still sold out, or probably will be sold out. So—
Yeah.
I don’t know.
Just in terms of the article, did you guys see that they pulled up a picture of Pat Gelsinger getting a massive tattoo of VMware on his arm—
Yeah.
In the article.
Yeah, dude, I did see that. I think it’s sick. I also think it’s so— You know how gutted VMware is post–Hock Tan? The company he was CEO of is dead. It’s gone.
There’s Pat on display. I don’t know if we want to do the eulogy of VMware right now. I’m sure it’s going to continue to be used in the future. But all of the growth in AI is about people consuming compute through an API via tokens, not every single Fortune 500 company spinning up a bunch of virtual machines in the cloud the way it was in the cloud era.
The growth is in the form of CPU data centers supporting cheap GPU data centers. They had a great picture in that article about the Fairwater campus from Microsoft, showing the GPU building versus a CPU building. Let me throw that up on screen as well for the people who are watching online.
I like how this is going to be the dummy version of these articles. You don’t have to read the whole thing; you can just listen to our weekly. That’s actually probably the best thing we could possibly do, and it would force me to actually read all of our articles. Boom.
Are you saying you didn’t read this article and I’m just showing you this picture for the first time?
I super-skimmed the hell out of it, if I’m being honest with you.
Yeah. Well—
There’s too much content.
I guess this is the concept that we write—
We ship so much we can’t even read it.
Yeah, we ship so much. Seriously, I think there are 3 articles going out today, smaller ones and—
Yeah.
I mean, the Core Research Weekly is the definitive register of everything that we write. It’s becoming time-consuming to make the table. The table is what everyone hates most about the weekly now. They figured this out literally today—I watched them figure it out—but up until then, you had to manually do it. There was no way in WordPress to avoid manually creating this table.
In the beginning, it was 7 or 8 things a week, and now we’re at 15 or 20 things or whatever. Someone is essentially just manually typing all this stuff out, and they’re like, “I hate the table.” But the table’s in salt. Anyway, keeping up with this is becoming a real deal.
Sorry, burning question: How many agents do you have running right now while we’re recording this?
I have an agent team going right now. I have 7 in 1 task, and then I have 7 windows open, 1 of them with 7 subagents, or whatever the agent team is. 3 of the 7 windows are really the hot ones. Some of them I’m waiting for later.
The star players on the team.
Yeah, the star players. I kind of want to get into tmux, but I guess I just hate the Enter, the Ctrl+B arrow thing. It’s a pain, so I’m a boomer.
Ctrl+B is asking too much.
Well, no, it’s Command. No, it’s just like—it gets really— Then there’s the scrolling. I don’t know. I’m too much of a boomer on that one. But yeah, just 7 right now.
Okay. Let’s move on to the next one. Obviously, the topic that everybody’s been talking about recently is the increase in memory prices—just the absolute shortage. I think it’s kind of been this game of whack-a-mole, where you figure out the GPU supply chain to ship tens of millions of H100 equivalents per year. Then HBM and CoWoS capacity is what drove that.
Now, where’s the next long pole in the tent? Everybody’s turned to either SSDs—the drives—or even hard drives, or now memory. I think the questions on people’s minds are, first, why is this happening, and when will it let up? And second, what’s next after memory? You guys have thoughts on that?
Yeah.
So, basically, again, it started with GPUs being the main driver of incremental memory demand, and that's mostly important for HBM. I think what's really different about this cycle is that, in the past, we haven't seen a new memory product that is so much more difficult and intensive to manufacture. With HBM, the amount of bits you get per wafer is far less than, say, conventional DDR DRAM.
There are a few reasons. Number 1, you need to dedicate more area to your TSV keep-out zone, with TSVs being the wires built through each stack so that you can have a vertical stack where you deliver power and signal. Also, because of the performance requirements of HBM, the yield is terrible, or much worse than conventional DDR DRAM. Then you have to stack these things in an 8-high or 12-high stack, and you lose yield on the package.
All in all, you get several times more bits out of a conventional DDR DRAM wafer than an HBM wafer. As HBM demand grows, as you ship more GPUs or XPUs, the more wafers the industry has to dedicate to HBM, and that reduces total industry supply in bits. Conventional DRAM demand is ticking up because, again, with all CPUs, you need DRAM. So that's driving conventional DRAM to be tight.
Then you combine that with very little incremental wafer starts. After the COVID hangover, the memory manufacturers decided, “Okay, we're not going to invest more.” There's literally not much cleanroom space available where you can actually put equipment. We're adding virtually no additional supply while demand is booming, and that's what's causing things to be so tight. People just realized this about 5 or 6 months ago, and now that's driving things to be pretty wacky in the DRAM market.
It's similar for SSDs and hard drives. NAND supply has been much worse than DRAM. I think it peaked about half a year ago, and all the industry players were barely making a profit—or I don't think they were profitable. No one added supply, of course. Then, all of a sudden, a lot of that incremental AI demand started hitting, and now things have just become tight.
The problem with NAND is that it's less profitable than DRAM. The DRAM manufacturers are probably going to add more incremental capacity to DRAM first before thinking about NAND. It's going to be crazy for the next year, for a few years at least.
Yeah, for a few years. One aspect I wanted to add on the historical side is that it really was literally the worst cycle since 1996. It was the worst cycle of all time in memory—literally number 2. 1996 bankrupted a third of the industry, pretty much. So it was pretty much the worst of all time, and then it whipsaws to the best of all time.
One example from the last cycle that was historically different is that they literally shuttered machines. They were like, “Hey, we're just going to stop.” Usually what happens is bit demand growth ramps the entire time, and they're like, “Yeah, yeah. Actually, no, no. We're not even going to do supply. We're just going to turn off supply. We're going to take the charges.” Usually, they just have underutilization charges.
They were like, “We're going to save the cleanroom space.” That's historic. No one has ever been—you know, when you have something fully deployed, usually it's cheaper to just keep it on. But NAND was so bad that they were like, “No, no, no, no. Let's try to convince everyone to turn it off because there's so much supply.”
So we go from that market of being like, “We're turning off the spigot in order for us to fix the problem,” to, “There's not even a place for us to turn the spigots on anymore.” It's kind of crazy. There's a giant cleanroom shortage, and there's no other way to put it: it's historic. It's crazy to watch.
I think we mentioned this in our Slack the other day, but I want to say that memory and NVMe prices for SSDs went up 20% in a day recently. Prices are also going up 100% month on month, and there's a little bit of panic buying—just a wee bit. But I think it's kind of crazy. This is definitely historic.
So, I talked to some other guys about this, and somebody made the point about expansion: how do you actually add more capacity? It requires the memory manufacturers to put in some CapEx to build new fabs in order to build more memory and alleviate the shortages here.
You said that previously, when these cycles happen, they get caught because they pour in too much CapEx, build too much capacity, and then suddenly there are no buyers, prices drop, and they go bankrupt. Then people start asking, “When's the top? When are we going to see the top in the current memory cycle?”
It seems like, based on history, the top is going to be as soon as Samsung capitulates and decides to start building out more capacity, because historically they're the last one to do it after SK Hynix and Micron, who are currently building out capacity. Samsung is still sticking with the plan and not going to take the bait. At what point do you think they capitulate and decide to actually start building out capacity here?
It's—
—they capitulate and decide to actually start building out capacity here?
I think they're going to—
How much higher do we need to go?
I think they're going to build capacity, and it's still not going to be enough. I mean, it's going to be—
I think they're going to build capacity, and it's still not going to be enough. I mean, it's going to be the largest supply-demand gap we've ever seen, and it gets worse from the work that we see.
What's interesting is, if you want to talk specifically about Samsung, during the memory down cycle, everyone was like, “Please, Samsung, stop making memory.” SK hynix and Micron, on every earnings call, were like, “We really think the industry needs to stop making memory.” Samsung would be like, “No, we are going to keep making memory.”
When they finally stopped making memory, everyone's like, “Oh, that was the bottom. That was the absolute bottom, the pico bottom.” What was crazy is that I think Samsung did it for 3 extra quarters longer than they historically have.
I expect this to happen this cycle, but I still think Samsung's going to add capacity. It's just not enough.
Yeah. I think they want to add capacity. It's just that there are constraints. They don't have many cleanrooms. They need to build some cleanrooms, and that's going to take time.
Then, on the tooling side, ASML can only output so many EUV tools every year. One of the debates we're having internally is how many EUV tools the industry can get, with TSMC needing a lot and the memory guys needing a lot, and how they're allocated. Logic is also tight.
Yeah. And you made a comment earlier that logic is effectively less complicated to manufacture than memory in the current state.
No. No. Logic is probably more complicated to manufacture, especially from a manufacturing perspective.
Well—
—logic is probably more complicated to manufacture.
Then DRAM. Doug's on mute, but I mean, I was maybe referring to the HBM part of the process, where you're doing advanced packaging with the logic and the memory.
Oh, yeah. It's hardcore, I think.
Logic's worse.
Yeah. I think it depends. I'd say the packaging, in terms of stacking things—stacking 8 DRAM dies on top of a base die or 12 DRAM dies on top of a base die, which is what's most common now—is much more difficult, just because if you mess up this one layer on the way, the whole thing is done.
The cumulative yield loss of doing something 8 or 12 times really kills you. That's something that TSMC doesn't have to deal with. Of course, they have to do really big CoWoS modules that are difficult to yield. They do stuff like hybrid bonding and SoIC, which is also very complicated.
But in terms of just doing such a high stack, a large 3D IC assembly, it's really hard for the guys to do it.
But, okay. I want to put some hype on the logic name. Let's not forget the—
Another way to think about this—and I don't actually have—I mean, actually, dang, I think I came up with this on the fly—is the CapEx per wafer start. Per 100K wafer starts, that's one way to keep it levelized. I think it's 2 or 3 times higher for leading-edge logic versus DRAM, right? So that would mean, hey, leading-edge logic is, let's say, 2.5 times higher.
I'm vibing it out. I think we have this number internally, but I have no idea off the top of my head. It's on me.
It would be 2.5 times more, but you can argue that the 3 times trade ratio of DRAM means that logic or HBM is the most expensive per 100K wafer starts in the entire world. So, yeah, I guess I'm defeated. I think HBM's the hardest.
Right. Does that affect the allocation of the EUV tools to the buyers? In other words, the people who are willing to pay more are the ones who can make more money out of the output from the tool.
It's a good question. I think the most sustainable thing is that, because logic needs memory and memory needs logic within a system, you want to make sure everyone gets their fill so that the industry as a whole can maximize the chips that go into completed systems. Otherwise, you'd get some weird balances. That's a question for ASML's planning and strategy team, and they should probably buy our models to help them make that decision.
Yeah, to help them make the decision, they need to buy our models. That's the right answer.
I think the industry itself is not massively zero-sum. One of my favorite parts of the semiconductor industry is that I don't perceive it to be massively zero-sum, meaning that one random person is like, "Yeah, screw everyone else. We're just going to gouge the shit out of whatever." Usually, what happens is a technological innovation gives them great margins. But at the end of the day, it is about making more chips, and this memory problem is becoming such an issue where it's like, "Hey, we're not going to be able to sell the right ratio of GPUs to DRAM."
For ASML, what they're maximizing is essentially selling the most chips possible, and so they're going to try to make the right trade ratio. But the trade ratio itself is honestly a very hard question to answer, and I don't think many people know. All in all, even the hyperscalers and NVIDIA—everyone's aware of this memory problem. We've got to get some EUV machines into the DRAM processes.
Actually, my galaxy-brain take is that it's going to pull forward 3D DRAM. I mean, it's not going to be before 2030, to be clear. But 3D DRAM is supposed to be 2035 or whatever, right? I think it might actually be closer to 2030 because 3D DRAM doesn't use any EUV, and that's a galaxy-brain way to increase total throughput of the entire system. But we're pretty far from there.
2. SRAM Trades Cost For Speed
How about alternative accelerators that use no DRAM or HBM, and instead go all SRAM and get the trade-off of 6 times more money for 2.5 times the performance? How about that as a segue to talk about Opus Fast and Codex 5.3 Spark?
I think that's like saying, "Eggs are a bit more expensive. I'm going to eat caviar for breakfast instead."
That's pretty good. Dang, dude.
I've had that before.
That was sort of my thing. Myron produces the SemiAnalysis bangers that he just comes—
Oh.
—with one on the spot here.
Dang.
Yeah, because SRAM is just so much more expensive than DRAM. And if everything was SRAM-heavy—I mean, logic's already tight, right? So that's just going to make—
Yeah.
—logic—
We're literally seeing people go for the caviar trade because they're not thinking about calories. They're thinking about their enjoyment of the meal, which is like Opus Fast releasing with a price tag that's 6 times higher per token and is 2 times faster—2.5 times or whatever at the top speed.
Then OpenAI responds with GPT-5.3-Codex-Spark, which they say runs on Cerebras, and we know that's 10 times more money for 4 times the performance, or really more like 20 times the money for 10 times the performance or something like that. That's an interesting trade, right? They're trying to do this—
Yeah.
—caviar trade right now.
The true galaxy-brain take is that we're so compute-starved that they're like, "I'm going to eat caviar so I don't die." I'm mostly joking because I definitely think you're right in terms of Cerebras and really fast inference stuff. I think it's experimenting because clearly one of the issues with the entire token-consumption model is that the fixed usage on Claude Max or whatever is pretty high. That's $2,400 per year.
But there has to be some higher tier. There has to be a higher tier of pricing, and it's pretty clear that Fast mode exists. So now it's time to figure out what that pricing elasticity curve was. If you think about it, they're like, "Hey, first we're just going to give it away for free. Okay, now we're going to sell it to you for cheap, and you have unlimited usage. But maybe you can't use our best models or blah, blah, blah. It's really slow. Now we're going to give it away for less cheap, and you have even more usage."
And then it's like, "Whoa, whoa, what if you gave it away—what if I sold it to you in the fastest mode possible and you spend 10 times?" They're finding out the demand curve, the price elasticity curve. From what it seems like, at least, now that I'm addicted to Fast tokens in Claude Code, it's working. I don't think I could go back, dude. It's—
Now it's blue cheese, man. Now you swear you're addicted to blue cheese, man.
Yeah.
Yeah, bruh.
No, dude. You moved me from heroin to fentanyl. That's what happened.
No.
I don't think I can go back.
This is a PG podcast here.
Oh, sorry.
PG-rated.
Okay. Lord, forgive me for my old ways, okay? You moved me from—
You know, a CJ Whoopty reference. We'll see who in the audience actually understands that one. But, okay. Bringing it back, Myron, do you think it's possible that there's a need to explore alternative architectures because there's just so little incremental DRAM, HBM, or logic, or whatever the constraint is?
I think there's always been that search for different architectures. It's funny now that HBM, at least in the short term, because of how pricing dynamics and contracting work, the price of conventional GDDR RAM on a per-gigabyte basis is actually closing in on HBM.
But back when HBM was much more expensive on a relative basis, a lot of new and established accelerator companies, like NVIDIA with CPX, as well as newer companies, explored different ways of doing non-HBM-based architectures. Those always had trade-offs. But it does seem like HBM, for now, has probably the best balance in terms of cost, bandwidth, and density.
But to Doug's point, the SRAM-based architectures do offer really, really fast tokens. They're more expensive in terms of getting much less throughput per dollar, but there's an appetite to pay, right?
Yeah.
And that's the caviar upgrade, right? It's much more expensive. It might be twice as delicious as eggs, but people are going to be willing to pay for it.
Makes sense. Yeah, yeah. Do you guys see any of the other model releases, like CDance from ByteDance generating videos? What, like—
3. China Models Challenge The Leaders
Can we talk about Asian Model Week? I've been calling it China Model Week, because it's actually China Model Week.
It's before Chinese New Year.
Who made the call? Was it Dylan or you?
No, we've all been saying there are going to be model releases on the New—
This is DeepSeek.
No, dude.
Yeah.
Whoa, whoa, whoa. I was the first person. I want you to say I was the first person. I literally was—well, because I read some SCMP thing on my ginormous whatever, and I was like, "I bet you it's going to be Presidents' Day for V4." That was like— But I did that in November or December or something like that. And then—
I think every DeepSeek release has corresponded with a Chinese holiday.
Yeah.
They cook them up.
No, not with a Chinese holiday—with an American holiday. It's on Thanksgiving; it's on Christmas Eve. It's always on an American holiday to troll the shit out of America. Sorry, you got it wrong. It's on an American holiday. Presidents' Day just happens to be Monday, okay? And it's near Chinese New Year, and it's kind of all bunched together.
You think the Chinese people are so aware of American holidays that they're going for Presidents' Day? Where was the Martin Luther King Jr. Day release?
DeepSeek is.
DeepSeek, I think, is—if they really wanted to mog us, they would really cook and put it out on July 4. That would be like the total victory.
Yeah. They're more—
They put out like a—
They're well aware of July 4. But if something comes out on Shrove Tuesday or Pancake Day or one of these—
The bank's closed, bro. The bank's closed.
Presidents' Day is not well known in China.
They put it out. Well, the timing just lines up. I don't know if it'll actually be Presidents' Day. That's my shitty speculation. Maybe they'll do Valentine's Day as a little love poem from the U.S. to China. But it's on a Saturday.
Sorry, does that holiday—does Valentine's Day exist in China? Come on.
I don't know. I think a lot of American holidays are really frowned upon. Christmas is banned in China. I don't know if you knew this.
No, I didn't know that.
Christmas is definitely perceived as Westernization. But anyway, sorry. I wanted to go back to China Model Week because China Model Week is actually really good. I feel like I'm the only one who's hyping up the China models right now.
Not the only one—a lot of people are on Twitter or whatever. But I feel like ever since Kimi 2.5, you're finding me in the most Chinese era of my life because I believe in these China models. MiniMax 2.5 is really good. It's at Opus performance, but it's 1/10 the price.
It's 10 billion active parameters. It's really, really—way, way cheaper. At least the way they're pricing it is cheaper. GLM-5 was an incredible release. You've got the Zhipu guys doing this. This is like the Chinese version of Indeed or LinkedIn is training foundation models that were hitting 70% on Soy Bench.
Yeah. No, they're really—well, you say that, but Zhipu just went public as an A—well, I think they've pivoted away from that. They're now an AI company, but they started as a—
Everybody's an AI company. But the biggest one is Seedance, right? ByteDance's CDance 2 is putting out a video-generation model that can generate minute-long scenes with consistent characters.
There are videos going around the internet that are still clearly AI-generated as you watch them. It's not actually the real actors, but it's a scene from The Sopranos that never happened, or Breaking Bad, or Brad Pitt fighting Tom Cruise or something like that.
Yeah, but it's consistent. I don't know how to explain it, but I feel like CDance 2 hit some level of breakthrough where you're like, "Oh, this isn't gibberish out of three anymore. This is actually really good—out of four," right? Or like Opus 4.5, where everyone's freaking out.
Opus 4.5 feels like some kind of breakthrough where I'm like, "Wow." Software engineering is automated. I think at least animation is done—cooked. I feel like if they make CDance just a little bit better, you could actually just generate all video.
And also, I think the thing that's notable is it is definitively the state of the art. The Gemini team, with Veo and Genie, has meaningfully been ahead for a long time, and I think this is the first time in a long while where I'm like, I think it's just total victory. Maybe this is just because we're between model releases or something like that, but it feels that way.
You would argue that the gap has always been much closer in China on the video side. My conspiracy theory is they have more video-data generation. The surveillance state gives everyone better training data. I mean, I'm not joking. Did you know 10% of all hard-drive demand is Chinese surveillance? It's a meaningful amount.
During the COVID shutdown, it actually hurt the hard-disk-drive companies even more because of how meaningful—
4. Storage Demand Goes Beyond AI
Yeah, we gotta find that stat. I gotta ask you for that stat after this, because when people ask, "Is ICMS gonna drive the future of drives—the demand for storage?" I gotta respond and be like, "No, actually, traditional data storage—the training data or logs—is where all the storage is going."
Saying that the Chinese surveillance state is much bigger than your KV-cache offload would be a pretty good point, I think.
Well, okay, but I do think KV-cache offload is going to be a big part of the market. It's not going to be zero. It's going to be double-digit—
It's not going to be zero, no. But—
It's going to be double-digit whatever, but yeah.
A hard drive.
Video wins.
Yeah.
Yeah.
So, okay.
Well, okay.
First of all—
Anything that comes through the model needs to get stored in logs, and they don't throw away data right now. So that means there's always going to be some multiple more stored in logs than there is active KV cache, right?
Okay, then you'd look at synthetic data that they're generating with these models, so all the output they're storing as well, which is not part of KV cache, right? And then you've got to look at all of the training datasets that come from existing sources. So every time somebody turns on a video camera, as you're saying, it gets stored down to disk.
This is much bigger—multiples bigger, 10 to hundreds of times bigger—than these models are in terms of active parameters, or these inputs to the models are in terms of active parameters. So, yeah, it's really bothering me that everybody keeps asking, "Is there going to be a massive incremental market-wide demand for KV-cache offload that gets driven by demand for SSDs for KV, that is driven by KV-cache offload?"
And the answer is, there's going to be massive demand for SSDs and hard drives, and it's not going to be driven by KV-cache offload.
I mean, I'm going to push back on "some massive demand." I'm going to put air quotes around it. I do think it will be double-digit total demand in bits over some period of time. But, yeah, 100%.
Let me use the analogy from before. This is like worrying that we're going to run out of trees or something because we're printing too many books when we're building houses or something. Video just takes so much more.
Video generation and images—I mean, that's my favorite analogy, and I've said this all along. And then, obviously, I feel like we were too slow on the memory side of things. Open up your iPhone and see what takes up most of the memory.
They're going to be like, "Okay, whatever," and I'm going to be like, "Yeah, it's probably 20% apps and 70% photos and videos." And they're like, "How'd you know?" It's because that's literally the reason why. That's the single biggest driver.
So CDance, for example, is actually going to be a huge demand driver because it's one thing to generate a lot of videos and then obviously give them to your user and have them share that back and forth. It'd be crazy if we're generating tons and tons and tons and tons of video. That is a lot of storage demand.
It'd be—
And that's gonna—
How much of those generated videos do you think are going to be cached in the internal representations of the model in a KV cache? A very small fraction, right? Less than 1%.
No, no, no, I know. I agree. It's a very small fraction. The part that's going to take the most storage is actually sending it and watching it—it's the consumption of it, you know? Not the—
Well, of the raw file that was generated for the user.
Yes, agreed.
Not the—
Yeah, I don't—
The stuff that's in memory that's getting evicted while the user is editing the video. This is a much smaller fraction than just the total storage of all generated videos.
Yeah. People forget there's a lot of compression going on. That's kind of the name of the game, if you think about it. But, yeah, no, I feel like you've been fighting a little bit of a war. People have been doing some silly stuff, man. The stuff I get these days—
They love an attach rate, man. How many ASICs—
Well, you—
—are going to go with these Rubin GPUs?
It's hard to explain as a former finance bro, but it does simplify your life to really focus on attach rates.
Former finance bro.
Whoa, I—
Current Claude Code babysitter.
There you go.
Current Claude Code vibe god. It just really helps to simplify. Actually, we went through this conversation where I think a lot of people were like, “Yeah, 1.5 attach for a transceiver,” or whatever.
We did a lot of this work. We did the P × Q. We did the whole thing. We did the whole BOM. We made sure every port—east, west, north, south—all this stuff—and it was 1.5 times.
Some really powerful heuristics get you pretty far in life. You’d be really surprised. For better or for worse, I think the CapEx—or at least the revenue monetization per gigawatt—has so far been the cleanest way to figure all this out because there are a lot of different ways, and maybe you’re off a little bit, but it’s such a big number that it ends up being directionally correct.
As long as your ratio isn’t totally messed up, that rule of thumb has been winning more than trying to be really precise. Big rules of thumb—mental models, so to speak—really crush in some areas. That’s why you want attach ratio, bro. Tell me, what’s the attach ratio?
The current attach ratio that people are throwing around is Jensen’s comment on stage at CES, which is 16 terabytes per GPU. Then all you do is pass that through to calculate how many exabytes are shipped from the fab every year. You see that that’s less than 0.7% of total drive shipments, and so people are like, “Oh, ICMS is a small incremental.” That’s the right conclusion for ICMS, but it’s not the right conclusion for the growth of drive demand for everything. It’s going to grow very much.
Yeah, 100%. My favorite part about that, too, is that the attach rate is so crazy. If you think about it, Jensen has a good nose for saying the right thing at the right time. People are freaking out about memory, and he’s like, “Whoa, whoa, whoa. Let me tell you about memory—the memory that’s attached to my GPU.”
He’s like, “No, no, no, no. AI’s not going to kill software. We’re actually going to have way more software.” In the same way that you have to be a man of the moment, he finds a way to bring it back to what matters to Jensen.
Yeah, like the—
Yeah.
$100 billion investment in OpenAI—whoa, whoa, whoa. We were invited to invest up to $100 billion. We’re not going to invest $100 billion, but we might invest more than $100 billion.
My understanding is there’s definitely—
Invest more.
I feel like that deal happened because of AMD. It happened afterward to screw over AMD, you know. Definitely—
They made the Cerebras announcement about GPT-5.3-Codex-Spark, and in the announcement they included the language, “NVIDIA is still the heart of what we do.”
Dude, the amount of NVIDIA press releases where it’s like, “This has been done with GB200,” and then Greg Brockman being like, “It’s incredible what you can do with GB200.” It’s like Jensen has the mouth puppet, and he’s like, “Say the words. Say the words.”
So, yeah. No, dude, it’s Jensen’s world we’re all living in. Okay, anything else for our first weekly? This is actually very helpful. I feel like I’ve caught up on our own content. It’s great to hear.
Okay, anything else for our first weekly? This is actually very helpful. I feel like I’ve caught up on our own content. It’s great to hear.
No, thanks everybody for listening. We’ll—
Yeah.
Good job, guys, and we’ll talk to you again next week.
Yeah.
Nice job, Myron. Nice job, Tuck. Yeah.
Cool. See you guys.