Jordan Nanos
All right. Hello, everyone. Welcome to episode number 1 of the SemiAnalysis weekly podcast. We don't know exactly what the name is or if this is ever gonna be public, but we're testing things out, and we're gonna talk about what happened in the week that was. Today's Thursday, February 12th.
We had 3 articles that came out on the free-tier newsletter: Claude Code Rising, CPUs Are Back, and Memory Mania. We'll try to talk about all 3 and other stuff that's going on. By way of introduction, I'm Jordan Nanos. I'm here with Doug O'Laughlin and Myron Ji—I don't even know how to pronounce your last name properly. Ji?
Speaker 1
Xie.
Doug O'Laughlin
Yeah.
Speaker 1
Have you ever heard xie xie?
Doug O'Laughlin
Yeah.
Speaker 1
Thank you in Chinese?
Doug O'Laughlin
Yeah.
Speaker 1
My last name is literally that.
Doug O'Laughlin
Thank you for coming on this podcast.
Speaker 1
I'm a very grateful person.
Doug O'Laughlin
Thank you for coming on this podcast today.
Speaker 1
Thank you for having me.
Doug O'Laughlin
Yeah.
Jordan Nanos
So is it like “thank thank” or “thank you, thank you”?
Speaker 1
If you say xie xie, that's “thank you.”
Doug O'Laughlin
Okay. Let's keep it on topic. This is not Transistor Radio Lite. So let's go through each of the things.
Jordan Nanos
Okay, great.
1. Claude Code Goes Mainstream
Doug O'Laughlin
I think first, Claude Code, bro. I don't wanna talk about Claude Code. I've talked so much about Claude Code. But I do wanna talk about the fact that we're on the press release. I think it's kind of a big deal. It's very rare to get linked on a press release, anyway, and we're on Anthropic's raise press release. It's kind of cool.
Jordan Nanos
Yeah.
Doug O'Laughlin
Yeah.
Jordan Nanos
We meaning you. So, Doug, you made a chart that made it first into our Core Research product a few weeks ago, and then made it into the free-tier newsletter. It tracks how many contributions to public GitHub repos are being authored by Claude Code—how many commits are being made by Claude Code—because it signs off as a contributor when it sends a PR. That rate went from about 2% to 4%, and then they included that in the announcement release.
Doug O'Laughlin
Yeah.
Jordan Nanos
So walk us through that. How'd you get the idea to do that in the first place?
Doug O'Laughlin
I read a tweet where it said, “You idiots, it all says it's committed by Claude Code. You could turn this off.” And I was like, “Wait, can you scrape GitHub?” Then I just asked Claude Code, and that was it. I was like, “Oh my God, can we make a daily series?”
Then I realized, wow, this is like—I’ve definitely been rushing to make sure we were the first to put that chart out, because I was like, dude, anyone who's paying attention can figure this one out, especially with Claude Code. But I was like, “We need to be the first.”
I don't think this is exactly news. We're definitely gonna be publishing the data for our paid subs, or our institutional research tier, but we'll probably have some kind of public dashboard as well, or some kind of public-facing information. It's already almost 5%, right? At least according to the most recent one we did, it's 4.7% or 4.8% or something like that. It's still growing.
That's how I got to the idea: just doomscrolling, and I was like, “Wait, wait, wait, wait.” Then Claude Code and I made the chart and, obviously, wrote the article.
Jordan Nanos
How much of the article did Claude write, in addition to making the chart about itself?
Doug O'Laughlin
I don't think—I personally still want to be writing with a human voice, right? We gotta train these models on something. We can't just be sloshing all this stuff back around.
I do like it for outlining more than anything else. I think the very first round of it was when I was still at peak excitement about Claude Code, so I dictated most of it. Then I had the dictation cleaned up by AI, and then I wrote, “Okay, make a slop outline based on this,” and wrote over it using my voice—just manual tokens, with my fingers.
Afterward, we sent it to the powers that be, AKA the internal SemiAnalysis brain squad. People wrote in sections and gave feedback. At some point, I think the most heavy lifting that Claude Code did—and I did all this specifically on Claude Code, screw vanilla chat or anything anymore—was mostly the outline stuff.
When you write something so much, you just get lost as hell. You're like, “Okay, this all makes sense in your weird little brain,” but I'm just like, “Hey, just do a fresh go-through of each piece and walk me through this.”
I do sometimes have it do some editing by asking, “What's the least relevant portion of this paragraph?” I usually like to rip stuff out. But that's just me. Sometimes SemiAnalysis loves to add stuff in.
Jordan Nanos
Are you writing with Claude or with ChatGPT or Gemini?
Speaker 1
No. I agree with Doug. I like having a human voice. Although sometimes, for certain paragraphs, I'm like, “How do I express this concept?” If I use an LLM and it's like, “Okay, that's a really good way to express it,” that's helpful.
But I mostly use, to Doug's point, a human voice. I wouldn't generate a whole piece from an LLM. It's just specific ideas I'd flesh out with an LLM. If I want to make this sound more elegant, I'd use an LLM.
Jordan Nanos
Yeah. Makes sense.
Speaker 1
Maybe I'm too much of a bootlegger.
Doug O'Laughlin
I think it's our last artisanal point, you know? It's our hand-churned butter. That's all we have. This hand-churned butter is all we got.
I also don't think people realize that Myron writes so much of the bangers at SemiAnalysis—essentially anything accelerated. Do you think Dylan's writing these articles? No, brother. It's a team.
People still believe, even as late as 2024, that it's just Dylan writing everything: “Dylan, I just love everything you write.” And I'm like, “We just put out 3 10K reports this week—like 5K-word reports.” No, there are a lot of people on these author bylines, and there's definitely a lot of heavy lifting.
Obviously, Dylan's giant brain is behind so much of everything we do here. But Myron's an OG SemiAnalysis person, and he's written some of the greats.
Speaker 1
When you read something like one of our big 20K- or 10K-word bangers that obviously has a lot of contributions from various people, can you tell whose voice it is? Who wrote this? Did Dan write this? Did Jeremy write this?
Doug O'Laughlin
Yes, I can kind of tell, at least.
Jordan Nanos
I'm not at Doug's level. I can tell certain voices, or that certain people use certain jargon, or choose to use the ampersand instead of saying A-N-D, or choose to reference it.
Doug O'Laughlin
There's also British spelling. British spelling definitely pops up.
Speaker 1
Oh, yeah. That's a big giveaway.
Doug O'Laughlin
That's a Myron thing—a Myron giveaway.
Jordan Nanos
I'm pro on that.
Speaker 1
That's kind of his British spelling.
Jordan Nanos
Most of it. Actually, our Prime Minister, Mark Carney, is going on a crusade right now, making all of his employees use the British spelling of many common words. So it's interesting.
Speaker 1
Wow.
Jordan Nanos
It's a return back to the—
Speaker 1
That's true, though. That's like—
Jordan Nanos
Long run.
Speaker 1
The reverse of Dylan.
Doug O'Laughlin
Yeah. The data center, centre, or whatever. And my favorite is, I think Jeremy has the most typos, because it's very clear that his entire keyboard is configured for French.
You think his autocorrect is gonna catch the English typo? No way, bro. So he definitely slips the most typos, I think.
Jordan Nanos
Well, the other thing for Jeremy is that if an entire sentence is highlighted with a link to an article, followed by the next sentence also being highlighted with a link to the article, so that the whole block of text is just links to previous articles, that's a Jeremy play for sure.
Doug O'Laughlin
Yeah, that's definitely a Jeremy play. There are so many people who go into all these things. It really does help.
Actually, I think that's how all these articles start. Most of an article is like a 2K-word thing, a 1K-word thing, and then we're like, “Hey, it'd be really nice if you helped with this section.”
Someone helps with the section and adds 1K words, and then each incremental thing gets added—another 700 words or whatever. Then all of a sudden we're like, “Wait, this is 6,000 words.”
Speaker 1
Whenever someone says, “Oh, let’s do a brief article on this,” it ends up becoming a Frankenstein.
Doug O'Laughlin
Dude, I think I’m proud. I feel like the Claude Code one was really short because I was the quarterback and I just rammed it out. I was like, “We need to send this chart out. We need everyone to use our number.” I needed that so badly, so I definitely rammed that one out.
Speaker 1
Yeah.
Doug O'Laughlin
So whenever—
Jordan Nanos
Let me try something here and show a bit of the backstory, just because we’re talking inside baseball about how SemiAnalysis works right now. Here’s the first chart, which went out on January 20th in Core Research, and then here’s the one that went out publicly. So just 2 weeks later—this one’s February 5th—
Doug O'Laughlin
Yep.
Jordan Nanos
We’re all the way up at around 4%.
Doug O'Laughlin
Yeah.
Jordan Nanos
And—
Doug O'Laughlin
Yep.
Jordan Nanos
Here’s the Anthropic press release we were referring to, where they announced their $30 billion raise at a $380 billion valuation, and a link to the article that makes that estimate. So, just for reference, some articles are going to the free tier and not to subscribers first. In other words, the CPU article that we just put out about the return of CPUs and their importance is a very different article. It’s a bit of a history, as opposed to a comment about what’s happening right now, or at least it leads in with history or something like that.
Doug O'Laughlin
Yeah. I think it’s mostly because CPUs are history, right? I mean, sorry, that’s—
Speaker 1
Not anymore. They’re back.
Speaker 2
Not anymore. They’re back. Yeah. I definitely think that—
Speaker 0
Even Intel.
Speaker 2
Yeah, even Intel.
Speaker 1
Even—
Speaker 2
I mean—
Speaker 1
Intel is actually—
Speaker 2
That’s the—
Speaker 1
Yeah.
Speaker 2
Yeah.
Speaker 1
That tells you things are tight.
Speaker 2
Yeah. That’s actually been kind of the story across the entire accelerator space, right? The marginal supplier actually benefits massively, and I just don’t think anyone would have ever guessed that we would hit the largest, most mature install base of compute CPUs and be like, “Yeah, actually, we’re going to run out of supply.” The marginal supplier, Intel, is actually going to catch a bid. It’s been kind of incredible.
I think part of the story that’s interesting, too, is that Arm is such a bigger threat than it’s ever been, and that’s been the other part of the story that’s very interesting. AMD continues to crush; it’s just kind of fragmenting. But the Phoenix CPUs are definitely an interesting change, especially in terms of how Arm’s business model is going to adapt. Even with more competition, it doesn’t matter: Intel is still sold out, or probably will be sold out. So—
Speaker 0
Yeah.
Speaker 2
I don’t know.
Speaker 0
Just in terms of the article, did you guys see that they pulled up a picture of Pat Gelsinger getting a massive tattoo of VMware on his arm—
Speaker 2
Yeah.
Speaker 0
In the article.
Speaker 2
Yeah, dude, I did see that. I think it’s sick. I also think it’s so— You know how gutted VMware is post–Hock Tan? The company he was CEO of is dead. It’s gone.
Speaker 0
There’s Pat on display. I don’t know if we want to do the eulogy of VMware right now. I’m sure it’s going to continue to be used in the future. But all of the growth in AI is about people consuming compute through an API via tokens, not every single Fortune 500 company spinning up a bunch of virtual machines in the cloud the way it was in the cloud era.
The growth is in the form of CPU data centers supporting cheap GPU data centers. They had a great picture in that article about the Fairwater campus from Microsoft, showing the GPU building versus a CPU building. Let me throw that up on screen as well for the people who are watching online.
Speaker 2
I like how this is going to be the dummy version of these articles. You don’t have to read the whole thing; you can just listen to our weekly. That’s actually probably the best thing we could possibly do, and it would force me to actually read all of our articles. Boom.
Speaker 0
Are you saying you didn’t read this article and I’m just showing you this picture for the first time?
Speaker 2
I super-skimmed the hell out of it, if I’m being honest with you.
Speaker 0
Yeah. Well—
Speaker 2
There’s too much content.
Speaker 0
I guess this is the concept that we write—
Speaker 1
We ship so much we can’t even read it.
Speaker 2
Yeah, we ship so much. Seriously, I think there are 3 articles going out today, smaller ones and—
Speaker 0
Yeah.
Speaker 2
I mean, the Core Research Weekly is the definitive register of everything that we write. It’s becoming time-consuming to make the table. The table is what everyone hates most about the weekly now. They figured this out literally today—I watched them figure it out—but up until then, you had to manually do it. There was no way in WordPress to avoid manually creating this table.
In the beginning, it was 7 or 8 things a week, and now we’re at 15 or 20 things or whatever. Someone is essentially just manually typing all this stuff out, and they’re like, “I hate the table.” But the table’s in salt. Anyway, keeping up with this is becoming a real deal.
Speaker 0
Sorry, burning question: How many agents do you have running right now while we’re recording this?
Speaker 2
I have an agent team going right now. I have 7 in 1 task, and then I have 7 windows open, 1 of them with 7 subagents, or whatever the agent team is. 3 of the 7 windows are really the hot ones. Some of them I’m waiting for later.
Speaker 0
The star players on the team.
Speaker 2
Yeah, the star players. I kind of want to get into tmux, but I guess I just hate the Enter, the Ctrl+B arrow thing. It’s a pain, so I’m a boomer.
Speaker 0
Ctrl+B is asking too much.
Speaker 2
Well, no, it’s Command. No, it’s just like—it gets really— Then there’s the scrolling. I don’t know. I’m too much of a boomer on that one. But yeah, just 7 right now.
Speaker 0
Okay. Let’s move on to the next one. Obviously, the topic that everybody’s been talking about recently is the increase in memory prices—just the absolute shortage. I think it’s kind of been this game of whack-a-mole, where you figure out the GPU supply chain to ship tens of millions of H100 equivalents per year. Then HBM and CoWoS capacity is what drove that.
Now, where’s the next long pole in the tent? Everybody’s turned to either SSDs—the drives—or even hard drives, or now memory. I think the questions on people’s minds are, first, why is this happening, and when will it let up? And second, what’s next after memory? You guys have thoughts on that?
Speaker 1
Yeah.
So, basically, again, it started with GPUs being the main driver of incremental memory demand, and that's mostly important for HBM. I think what's really different about this cycle is that, in the past, we haven't seen a new memory product that is so much more difficult and intensive to manufacture. With HBM, the amount of bits you get per wafer is far less than, say, conventional DDR DRAM.
There are a few reasons. Number 1, you need to dedicate more area to your TSV keep-out zone, with TSVs being the wires built through each stack so that you can have a vertical stack where you deliver power and signal. Also, because of the performance requirements of HBM, the yield is terrible, or much worse than conventional DDR DRAM. Then you have to stack these things in an 8-high or 12-high stack, and you lose yield on the package.
All in all, you get several times more bits out of a conventional DDR DRAM wafer than an HBM wafer. As HBM demand grows, as you ship more GPUs or XPUs, the more wafers the industry has to dedicate to HBM, and that reduces total industry supply in bits. Conventional DRAM demand is ticking up because, again, with all CPUs, you need DRAM. So that's driving conventional DRAM to be tight.
Then you combine that with very little incremental wafer starts. After the COVID hangover, the memory manufacturers decided, “Okay, we're not going to invest more.” There's literally not much cleanroom space available where you can actually put equipment. We're adding virtually no additional supply while demand is booming, and that's what's causing things to be so tight. People just realized this about 5 or 6 months ago, and now that's driving things to be pretty wacky in the DRAM market.
It's similar for SSDs and hard drives. NAND supply has been much worse than DRAM. I think it peaked about half a year ago, and all the industry players were barely making a profit—or I don't think they were profitable. No one added supply, of course. Then, all of a sudden, a lot of that incremental AI demand started hitting, and now things have just become tight.
The problem with NAND is that it's less profitable than DRAM. The DRAM manufacturers are probably going to add more incremental capacity to DRAM first before thinking about NAND. It's going to be crazy for the next year, for a few years at least.
Doug O'Laughlin
Yeah, for a few years. One aspect I wanted to add on the historical side is that it really was literally the worst cycle since 1996. It was the worst cycle of all time in memory—literally number 2. 1996 bankrupted a third of the industry, pretty much. So it was pretty much the worst of all time, and then it whipsaws to the best of all time.
One example from the last cycle that was historically different is that they literally shuttered machines. They were like, “Hey, we're just going to stop.” Usually what happens is bit demand growth ramps the entire time, and they're like, “Yeah, yeah. Actually, no, no. We're not even going to do supply. We're just going to turn off supply. We're going to take the charges.” Usually, they just have underutilization charges.
They were like, “We're going to save the cleanroom space.” That's historic. No one has ever been—you know, when you have something fully deployed, usually it's cheaper to just keep it on. But NAND was so bad that they were like, “No, no, no, no. Let's try to convince everyone to turn it off because there's so much supply.”
So we go from that market of being like, “We're turning off the spigot in order for us to fix the problem,” to, “There's not even a place for us to turn the spigots on anymore.” It's kind of crazy. There's a giant cleanroom shortage, and there's no other way to put it: it's historic. It's crazy to watch.
I think we mentioned this in our Slack the other day, but I want to say that memory and NVMe prices for SSDs went up 20% in a day recently. Prices are also going up 100% month on month, and there's a little bit of panic buying—just a wee bit. But I think it's kind of crazy. This is definitely historic.
Speaker 0
So, I talked to some other guys about this, and somebody made the point about expansion: how do you actually add more capacity? It requires the memory manufacturers to put in some CapEx to build new fabs in order to build more memory and alleviate the shortages here.
You said that previously, when these cycles happen, they get caught because they pour in too much CapEx, build too much capacity, and then suddenly there are no buyers, prices drop, and they go bankrupt. Then people start asking, “When's the top? When are we going to see the top in the current memory cycle?”
It seems like, based on history, the top is going to be as soon as Samsung capitulates and decides to start building out more capacity, because historically they're the last one to do it after SK Hynix and Micron, who are currently building out capacity. Samsung is still sticking with the plan and not going to take the bait. At what point do you think they capitulate and decide to actually start building out capacity here?
Doug O'Laughlin
It's—
Speaker 0
—they capitulate and decide to actually start building out capacity here?
Doug O'Laughlin
I think they're going to—
Speaker 0
How much higher do we need to go?
Doug O'Laughlin
I think they're going to build capacity, and it's still not going to be enough. I mean, it's going to be—
Doug O'Laughlin
I think they're going to build capacity, and it's still not going to be enough. I mean, it's going to be the largest supply-demand gap we've ever seen, and it gets worse from the work that we see.
What's interesting is, if you want to talk specifically about Samsung, during the memory down cycle, everyone was like, “Please, Samsung, stop making memory.” SK hynix and Micron, on every earnings call, were like, “We really think the industry needs to stop making memory.” Samsung would be like, “No, we are going to keep making memory.”
When they finally stopped making memory, everyone's like, “Oh, that was the bottom. That was the absolute bottom, the pico bottom.” What was crazy is that I think Samsung did it for 3 extra quarters longer than they historically have.
I expect this to happen this cycle, but I still think Samsung's going to add capacity. It's just not enough.
Speaker 1
Yeah. I think they want to add capacity. It's just that there are constraints. They don't have many cleanrooms. They need to build some cleanrooms, and that's going to take time.
Then, on the tooling side, ASML can only output so many EUV tools every year. One of the debates we're having internally is how many EUV tools the industry can get, with TSMC needing a lot and the memory guys needing a lot, and how they're allocated. Logic is also tight.
Speaker 0
Yeah. And you made a comment earlier that logic is effectively less complicated to manufacture than memory in the current state.
Speaker 1
No. No. Logic is probably more complicated to manufacture, especially from a manufacturing perspective.
Speaker 0
Well—
Speaker 1
—logic is probably more complicated to manufacture.
Speaker 0
Then DRAM. Doug's on mute, but I mean, I was maybe referring to the HBM part of the process, where you're doing advanced packaging with the logic and the memory.
Speaker 1
Oh, yeah. It's hardcore, I think.
Doug O'Laughlin
Logic's worse.
Speaker 1
Yeah. I think it depends. I'd say the packaging, in terms of stacking things—stacking 8 DRAM dies on top of a base die or 12 DRAM dies on top of a base die, which is what's most common now—is much more difficult, just because if you mess up this one layer on the way, the whole thing is done.
The cumulative yield loss of doing something 8 or 12 times really kills you. That's something that TSMC doesn't have to deal with. Of course, they have to do really big CoWoS modules that are difficult to yield. They do stuff like hybrid bonding and SoIC, which is also very complicated.
But in terms of just doing such a high stack, a large 3D IC assembly, it's really hard for the guys to do it.
Doug O'Laughlin
But, okay. I want to put some hype on the logic name. Let's not forget the—
Another way to think about this—and I don't actually have—I mean, actually, dang, I think I came up with this on the fly—is the CapEx per wafer start. Per 100K wafer starts, that's one way to keep it levelized. I think it's 2 or 3 times higher for leading-edge logic versus DRAM, right? So that would mean, hey, leading-edge logic is, let's say, 2.5 times higher.
Jordan Nanos
I'm vibing it out. I think we have this number internally, but I have no idea off the top of my head. It's on me.
Doug O'Laughlin
It would be 2.5 times more, but you can argue that the 3 times trade ratio of DRAM means that logic or HBM is the most expensive per 100K wafer starts in the entire world. So, yeah, I guess I'm defeated. I think HBM's the hardest.
Doug O'Laughlin
Right. Does that affect the allocation of the EUV tools to the buyers? In other words, the people who are willing to pay more are the ones who can make more money out of the output from the tool.
Speaker 1
It's a good question. I think the most sustainable thing is that, because logic needs memory and memory needs logic within a system, you want to make sure everyone gets their fill so that the industry as a whole can maximize the chips that go into completed systems. Otherwise, you'd get some weird balances. That's a question for ASML's planning and strategy team, and they should probably buy our models to help them make that decision.
Jordan Nanos
Yeah, to help them make the decision, they need to buy our models. That's the right answer.
I think the industry itself is not massively zero-sum. One of my favorite parts of the semiconductor industry is that I don't perceive it to be massively zero-sum, meaning that one random person is like, "Yeah, screw everyone else. We're just going to gouge the shit out of whatever." Usually, what happens is a technological innovation gives them great margins. But at the end of the day, it is about making more chips, and this memory problem is becoming such an issue where it's like, "Hey, we're not going to be able to sell the right ratio of GPUs to DRAM."
For ASML, what they're maximizing is essentially selling the most chips possible, and so they're going to try to make the right trade ratio. But the trade ratio itself is honestly a very hard question to answer, and I don't think many people know. All in all, even the hyperscalers and NVIDIA—everyone's aware of this memory problem. We've got to get some EUV machines into the DRAM processes.
Actually, my galaxy-brain take is that it's going to pull forward 3D DRAM. I mean, it's not going to be before 2030, to be clear. But 3D DRAM is supposed to be 2035 or whatever, right? I think it might actually be closer to 2030 because 3D DRAM doesn't use any EUV, and that's a galaxy-brain way to increase total throughput of the entire system. But we're pretty far from there.
2. SRAM Trades Cost For Speed
Jordan Nanos
How about alternative accelerators that use no DRAM or HBM, and instead go all SRAM and get the trade-off of 6 times more money for 2.5 times the performance? How about that as a segue to talk about Opus Fast and Codex 5.3 Spark?
Speaker 1
I think that's like saying, "Eggs are a bit more expensive. I'm going to eat caviar for breakfast instead."
Doug O'Laughlin
That's pretty good. Dang, dude.
Speaker 1
I've had that before.
Jordan Nanos
That was sort of my thing. Myron produces the SemiAnalysis bangers that he just comes—
Speaker 1
Oh.
Jordan Nanos
—with one on the spot here.
Doug O'Laughlin
Dang.
Speaker 1
Yeah, because SRAM is just so much more expensive than DRAM. And if everything was SRAM-heavy—I mean, logic's already tight, right? So that's just going to make—
Doug O'Laughlin
Yeah.
Speaker 1
—logic—
Jordan Nanos
We're literally seeing people go for the caviar trade because they're not thinking about calories. They're thinking about their enjoyment of the meal, which is like Opus Fast releasing with a price tag that's 6 times higher per token and is 2 times faster—2.5 times or whatever at the top speed.
Then OpenAI responds with GPT-5.3-Codex-Spark, which they say runs on Cerebras, and we know that's 10 times more money for 4 times the performance, or really more like 20 times the money for 10 times the performance or something like that. That's an interesting trade, right? They're trying to do this—
Speaker 1
Yeah.
Doug O'Laughlin
—caviar trade right now.
Jordan Nanos
The true galaxy-brain take is that we're so compute-starved that they're like, "I'm going to eat caviar so I don't die." I'm mostly joking because I definitely think you're right in terms of Cerebras and really fast inference stuff. I think it's experimenting because clearly one of the issues with the entire token-consumption model is that the fixed usage on Claude Max or whatever is pretty high. That's $2,400 per year.
But there has to be some higher tier. There has to be a higher tier of pricing, and it's pretty clear that Fast mode exists. So now it's time to figure out what that pricing elasticity curve was. If you think about it, they're like, "Hey, first we're just going to give it away for free. Okay, now we're going to sell it to you for cheap, and you have unlimited usage. But maybe you can't use our best models or blah, blah, blah. It's really slow. Now we're going to give it away for less cheap, and you have even more usage."
And then it's like, "Whoa, whoa, what if you gave it away—what if I sold it to you in the fastest mode possible and you spend 10 times?" They're finding out the demand curve, the price elasticity curve. From what it seems like, at least, now that I'm addicted to Fast tokens in Claude Code, it's working. I don't think I could go back, dude. It's—
Doug O'Laughlin
Now it's blue cheese, man. Now you swear you're addicted to blue cheese, man.
Jordan Nanos
Yeah.
Doug O'Laughlin
Yeah, bruh.
Jordan Nanos
No, dude. You moved me from heroin to fentanyl. That's what happened.
Doug O'Laughlin
No.
Jordan Nanos
I don't think I can go back.
Doug O'Laughlin
This is a PG podcast here.
Jordan Nanos
Oh, sorry.
Doug O'Laughlin
PG-rated.
Jordan Nanos
Okay. Lord, forgive me for my old ways, okay? You moved me from—
Doug O'Laughlin
You know, a CJ Whoopty reference. We'll see who in the audience actually understands that one. But, okay. Bringing it back, Myron, do you think it's possible that there's a need to explore alternative architectures because there's just so little incremental DRAM, HBM, or logic, or whatever the constraint is?
Speaker 1
I think there's always been that search for different architectures. It's funny now that HBM, at least in the short term, because of how pricing dynamics and contracting work, the price of conventional GDDR RAM on a per-gigabyte basis is actually closing in on HBM.
But back when HBM was much more expensive on a relative basis, a lot of new and established accelerator companies, like NVIDIA with CPX, as well as newer companies, explored different ways of doing non-HBM-based architectures. Those always had trade-offs. But it does seem like HBM, for now, has probably the best balance in terms of cost, bandwidth, and density.
But to Doug's point, the SRAM-based architectures do offer really, really fast tokens. They're more expensive in terms of getting much less throughput per dollar, but there's an appetite to pay, right?
Doug O'Laughlin
Yeah.
Speaker 1
And that's the caviar upgrade, right? It's much more expensive. It might be twice as delicious as eggs, but people are going to be willing to pay for it.
Jordan Nanos
Makes sense. Yeah, yeah. Do you guys see any of the other model releases, like CDance from ByteDance generating videos? What, like—
3. China Models Challenge The Leaders
Doug O'Laughlin
Can we talk about Asian Model Week? I've been calling it China Model Week, because it's actually China Model Week.
Speaker 1
It's before Chinese New Year.
Speaker 1
Who made the call? Was it Dylan or you?
Jordan Nanos
No, we've all been saying there are going to be model releases on the New—
Jordan Nanos
This is DeepSeek.
Doug O'Laughlin
No, dude.
Speaker 1
Yeah.
Doug O'Laughlin
Whoa, whoa, whoa. I was the first person. I want you to say I was the first person. I literally was—well, because I read some SCMP thing on my ginormous whatever, and I was like, "I bet you it's going to be Presidents' Day for V4." That was like— But I did that in November or December or something like that. And then—
Jordan Nanos
I think every DeepSeek release has corresponded with a Chinese holiday.
Jordan Nanos
Yeah.
Doug O'Laughlin
They cook them up.
Doug O'Laughlin
No, not with a Chinese holiday—with an American holiday. It's on Thanksgiving; it's on Christmas Eve. It's always on an American holiday to troll the shit out of America. Sorry, you got it wrong. It's on an American holiday. Presidents' Day just happens to be Monday, okay? And it's near Chinese New Year, and it's kind of all bunched together.
Jordan Nanos
You think the Chinese people are so aware of American holidays that they're going for Presidents' Day? Where was the Martin Luther King Jr. Day release?
Doug O'Laughlin
DeepSeek is.
Jordan Nanos
DeepSeek, I think, is—if they really wanted to mog us, they would really cook and put it out on July 4. That would be like the total victory.
Doug O'Laughlin
Yeah. They're more—
Doug O'Laughlin
They put out like a—
Jordan Nanos
They're well aware of July 4. But if something comes out on Shrove Tuesday or Pancake Day or one of these—
Doug O'Laughlin
The bank's closed, bro. The bank's closed.
Jordan Nanos
Presidents' Day is not well known in China.
Doug O'Laughlin
They put it out. Well, the timing just lines up. I don't know if it'll actually be Presidents' Day. That's my shitty speculation. Maybe they'll do Valentine's Day as a little love poem from the U.S. to China. But it's on a Saturday.
Jordan Nanos
Sorry, does that holiday—does Valentine's Day exist in China? Come on.
Doug O'Laughlin
I don't know. I think a lot of American holidays are really frowned upon. Christmas is banned in China. I don't know if you knew this.
Jordan Nanos
No, I didn't know that.
Doug O'Laughlin
Christmas is definitely perceived as Westernization. But anyway, sorry. I wanted to go back to China Model Week because China Model Week is actually really good. I feel like I'm the only one who's hyping up the China models right now.
Not the only one—a lot of people are on Twitter or whatever. But I feel like ever since Kimi 2.5, you're finding me in the most Chinese era of my life because I believe in these China models. MiniMax 2.5 is really good. It's at Opus performance, but it's 1/10 the price.
Jordan Nanos
It's 10 billion active parameters. It's really, really—way, way cheaper. At least the way they're pricing it is cheaper. GLM-5 was an incredible release. You've got the Zhipu guys doing this. This is like the Chinese version of Indeed or LinkedIn is training foundation models that were hitting 70% on Soy Bench.
Doug O'Laughlin
Yeah. No, they're really—well, you say that, but Zhipu just went public as an A—well, I think they've pivoted away from that. They're now an AI company, but they started as a—
Jordan Nanos
Everybody's an AI company. But the biggest one is Seedance, right? ByteDance's CDance 2 is putting out a video-generation model that can generate minute-long scenes with consistent characters.
There are videos going around the internet that are still clearly AI-generated as you watch them. It's not actually the real actors, but it's a scene from The Sopranos that never happened, or Breaking Bad, or Brad Pitt fighting Tom Cruise or something like that.
Doug O'Laughlin
Yeah, but it's consistent. I don't know how to explain it, but I feel like CDance 2 hit some level of breakthrough where you're like, "Oh, this isn't gibberish out of three anymore. This is actually really good—out of four," right? Or like Opus 4.5, where everyone's freaking out.
Opus 4.5 feels like some kind of breakthrough where I'm like, "Wow." Software engineering is automated. I think at least animation is done—cooked. I feel like if they make CDance just a little bit better, you could actually just generate all video.
And also, I think the thing that's notable is it is definitively the state of the art. The Gemini team, with Veo and Genie, has meaningfully been ahead for a long time, and I think this is the first time in a long while where I'm like, I think it's just total victory. Maybe this is just because we're between model releases or something like that, but it feels that way.
You would argue that the gap has always been much closer in China on the video side. My conspiracy theory is they have more video-data generation. The surveillance state gives everyone better training data. I mean, I'm not joking. Did you know 10% of all hard-drive demand is Chinese surveillance? It's a meaningful amount.
During the COVID shutdown, it actually hurt the hard-disk-drive companies even more because of how meaningful—
4. Storage Demand Goes Beyond AI
Jordan Nanos
Yeah, we gotta find that stat. I gotta ask you for that stat after this, because when people ask, "Is ICMS gonna drive the future of drives—the demand for storage?" I gotta respond and be like, "No, actually, traditional data storage—the training data or logs—is where all the storage is going."
Saying that the Chinese surveillance state is much bigger than your KV-cache offload would be a pretty good point, I think.
Doug O'Laughlin
Well, okay, but I do think KV-cache offload is going to be a big part of the market. It's not going to be zero. It's going to be double-digit—
Jordan Nanos
It's not going to be zero, no. But—
Doug O'Laughlin
It's going to be double-digit whatever, but yeah.
Jordan Nanos
A hard drive.
Doug O'Laughlin
Video wins.
Jordan Nanos
Yeah.
Doug O'Laughlin
Yeah.
Jordan Nanos
So, okay.
Doug O'Laughlin
Well, okay.
Jordan Nanos
First of all—
Anything that comes through the model needs to get stored in logs, and they don't throw away data right now. So that means there's always going to be some multiple more stored in logs than there is active KV cache, right?
Okay, then you'd look at synthetic data that they're generating with these models, so all the output they're storing as well, which is not part of KV cache, right? And then you've got to look at all of the training datasets that come from existing sources. So every time somebody turns on a video camera, as you're saying, it gets stored down to disk.
This is much bigger—multiples bigger, 10 to hundreds of times bigger—than these models are in terms of active parameters, or these inputs to the models are in terms of active parameters. So, yeah, it's really bothering me that everybody keeps asking, "Is there going to be a massive incremental market-wide demand for KV-cache offload that gets driven by demand for SSDs for KV, that is driven by KV-cache offload?"
And the answer is, there's going to be massive demand for SSDs and hard drives, and it's not going to be driven by KV-cache offload.
Doug O'Laughlin
I mean, I'm going to push back on "some massive demand." I'm going to put air quotes around it. I do think it will be double-digit total demand in bits over some period of time. But, yeah, 100%.
Let me use the analogy from before. This is like worrying that we're going to run out of trees or something because we're printing too many books when we're building houses or something. Video just takes so much more.
Video generation and images—I mean, that's my favorite analogy, and I've said this all along. And then, obviously, I feel like we were too slow on the memory side of things. Open up your iPhone and see what takes up most of the memory.
They're going to be like, "Okay, whatever," and I'm going to be like, "Yeah, it's probably 20% apps and 70% photos and videos." And they're like, "How'd you know?" It's because that's literally the reason why. That's the single biggest driver.
So CDance, for example, is actually going to be a huge demand driver because it's one thing to generate a lot of videos and then obviously give them to your user and have them share that back and forth. It'd be crazy if we're generating tons and tons and tons and tons of video. That is a lot of storage demand.
Jordan Nanos
It'd be—
Doug O'Laughlin
And that's gonna—
Jordan Nanos
How much of those generated videos do you think are going to be cached in the internal representations of the model in a KV cache? A very small fraction, right? Less than 1%.
Doug O'Laughlin
No, no, no, I know. I agree. It's a very small fraction. The part that's going to take the most storage is actually sending it and watching it—it's the consumption of it, you know? Not the—
Jordan Nanos
Well, of the raw file that was generated for the user.
Doug O'Laughlin
Yes, agreed.
Jordan Nanos
Not the—
Doug O'Laughlin
Yeah, I don't—
Jordan Nanos
The stuff that's in memory that's getting evicted while the user is editing the video. This is a much smaller fraction than just the total storage of all generated videos.
Doug O'Laughlin
Yeah. People forget there's a lot of compression going on. That's kind of the name of the game, if you think about it. But, yeah, no, I feel like you've been fighting a little bit of a war. People have been doing some silly stuff, man. The stuff I get these days—
Jordan Nanos
They love an attach rate, man. How many ASICs—
Doug O'Laughlin
Well, you—
Jordan Nanos
—are going to go with these Rubin GPUs?
Doug O'Laughlin
It's hard to explain as a former finance bro, but it does simplify your life to really focus on attach rates.
Jordan Nanos
Former finance bro.
Doug O'Laughlin
Whoa, I—
Jordan Nanos
Current Claude Code babysitter.
Doug O'Laughlin
There you go.
Doug O'Laughlin
Current Claude Code vibe god. It just really helps to simplify. Actually, we went through this conversation where I think a lot of people were like, “Yeah, 1.5 attach for a transceiver,” or whatever.
We did a lot of this work. We did the P × Q. We did the whole thing. We did the whole BOM. We made sure every port—east, west, north, south—all this stuff—and it was 1.5 times.
Some really powerful heuristics get you pretty far in life. You’d be really surprised. For better or for worse, I think the CapEx—or at least the revenue monetization per gigawatt—has so far been the cleanest way to figure all this out because there are a lot of different ways, and maybe you’re off a little bit, but it’s such a big number that it ends up being directionally correct.
As long as your ratio isn’t totally messed up, that rule of thumb has been winning more than trying to be really precise. Big rules of thumb—mental models, so to speak—really crush in some areas. That’s why you want attach ratio, bro. Tell me, what’s the attach ratio?
Doug O'Laughlin
The current attach ratio that people are throwing around is Jensen’s comment on stage at CES, which is 16 terabytes per GPU. Then all you do is pass that through to calculate how many exabytes are shipped from the fab every year. You see that that’s less than 0.7% of total drive shipments, and so people are like, “Oh, ICMS is a small incremental.” That’s the right conclusion for ICMS, but it’s not the right conclusion for the growth of drive demand for everything. It’s going to grow very much.
Jordan Nanos
Yeah, 100%. My favorite part about that, too, is that the attach rate is so crazy. If you think about it, Jensen has a good nose for saying the right thing at the right time. People are freaking out about memory, and he’s like, “Whoa, whoa, whoa. Let me tell you about memory—the memory that’s attached to my GPU.”
He’s like, “No, no, no, no. AI’s not going to kill software. We’re actually going to have way more software.” In the same way that you have to be a man of the moment, he finds a way to bring it back to what matters to Jensen.
Jordan Nanos
Yeah, like the—
Doug O'Laughlin
Yeah.
Jordan Nanos
$100 billion investment in OpenAI—whoa, whoa, whoa. We were invited to invest up to $100 billion. We’re not going to invest $100 billion, but we might invest more than $100 billion.
Doug O'Laughlin
My understanding is there’s definitely—
Jordan Nanos
Invest more.
Doug O'Laughlin
I feel like that deal happened because of AMD. It happened afterward to screw over AMD, you know. Definitely—
Doug O'Laughlin
They made the Cerebras announcement about GPT-5.3-Codex-Spark, and in the announcement they included the language, “NVIDIA is still the heart of what we do.”
Doug O'Laughlin
Dude, the amount of NVIDIA press releases where it’s like, “This has been done with GB200,” and then Greg Brockman being like, “It’s incredible what you can do with GB200.” It’s like Jensen has the mouth puppet, and he’s like, “Say the words. Say the words.”
So, yeah. No, dude, it’s Jensen’s world we’re all living in. Okay, anything else for our first weekly? This is actually very helpful. I feel like I’ve caught up on our own content. It’s great to hear.
Doug O'Laughlin
Okay, anything else for our first weekly? This is actually very helpful. I feel like I’ve caught up on our own content. It’s great to hear.
Jordan Nanos
No, thanks everybody for listening. We’ll—
Doug O'Laughlin
Yeah.
Jordan Nanos
Good job, guys, and we’ll talk to you again next week.
Doug O'Laughlin
Yeah.
Jordan Nanos
Nice job, Myron. Nice job, Tuck. Yeah.
Doug O'Laughlin
Cool. See you guys.