# Why Memory Is AI's Biggest Bottleneck with Vikram Sekar | EP 170

Frictionless · 2026-10-03 · 73 min · https://www.youtube.com/watch?v=13bFAvjTsyA

## Transcript

Vikram Sekar

Memory is the biggest problem. LLMs are literally memory-intensive models, and memory is everything. One of the biggest problems people have to solve is how to get the same performance without using as much memory. Once you solve this problem, a lot of things will fall into place.

The optics were the same as for long-haul connections, subsea and submarine cables, and some other applications. But optics have never been as important as they are today. Today, it is a technology that is essential for AI and future infrastructure. The whole world is fighting for it because everyone thinks, “Look, we were never prepared for this level of demand. This industry was never built for this level of demand. You’re asking for too much.” Because of this, everyone is facing shortages.

Logan Jastremski

What is the useful amount of work performed per token? Do I always need 2,000 tokens per second? I don’t. I’m talking to you right now, but my AI can process some PDFs that I want to read later or something. I don’t have an immediate need for this. Perfect.

### Vikram Sekar and Semi Doped

Well, Vic, thank you very much for joining me. I appreciate you participating in the podcast. I’ve been following your work for a while now, and I’m really impressed. It helped me personally get from 0 to 1 as quickly as possible.

I wanted to reach out to you and see if I could invite you to the podcast because I think you’re really good at explaining technical concepts in a way that people can understand. Knowing many engineers, I know that this is a difficult skill. I appreciate all the work you put into teaching others.

Vikram Sekar

Thank you. Thank you for inviting me. I’m always happy to sit down and talk about technology and how things work. That’s basically what I usually think about.

Logan Jastremski

Yes, I’m always happy to talk about different things.

Vikram Sekar

It’s even more gratifying to hear that my explanations, articles, and conversations on SemiAnalysis and my Substack newsletter, The Vik Newsletter, have resonated with people who care about this subject. I’m very happy to hear that.

Logan Jastremski

Yes, and for those who don't know, I'll link to Vic's entire podcast on Substack, everything he's worked on, including his institutional research. Please contact him because I think he does a great job. And then I'll have to contact you to better set up the podcast studio. Maybe I can get my own seamless logo in the background.

Vikram Sekar

Yes, why not? This is great.

### AI agents and hardware demand

Logan Jastremski

Perhaps a great place to start would be that the industry has really changed, it seems, from 0 to 1—or perhaps it’s better to say from 0 to 100—since Opus 4.5, which was in Q4 last year. We started moving from OAuth subscriptions to API-based usage, and that seemed to give the whole race a boost.

Now everyone is talking about agents, which I think is really exciting: just using a computer and being able to do things for ordinary people outside of software engineering and coding. We have a new race starting.

Maybe it’s worth starting the podcast with a new question: How do you see the situation with agents, and what are your general thoughts?

Vikram Sekar

I really like the current agent situation more than at any time since the beginning of this year, when agents like this really took off. At the beginning of the year, it was more a question of, “How do we get away from the chat interface and maybe, for the first time, make it work on the phone?”

When OpenClaw came out, you could send text messages from your phone and receive replies, scheduled events, notifications, and more. But the problem with that approach, at least at the beginning of this year, was that it was very difficult for regular people to go and set it up.

I’m not just saying this as a regular nontechnical person. I tried setting up OpenClaw. I still have a home server that runs it. It’s not that simple.

I had to spin up my own virtual machine. There were all these security concerns about giving OpenClaw access to your entire computer. People would say, “No, no, no, don’t install it on your laptop. It’s too intrusive. Go and buy another computer.”

So everyone went and bought a Mac mini. Do you remember this whole story?

Logan Jastremski

Yes. It was very funny to me.

Vikram Sekar

Considering that most of the models people used, in my opinion, were not local models, even if people were buying multiple Mac minis and trying to get as much unified memory as possible, they would still be making API requests to data centers.

Logan Jastremski

Yes, the whole idea was that you could run some kind of local model that would at least handle your personal tasks or something like that.

Vikram Sekar

But again, this is a very expensive proposition because models become cheaper to use over time, while you have to make an initial investment of several thousand dollars, hoping that the local model you deploy on your memory allocation will actually do the job. This is not a very long-term bet. You would be better off betting on token prices dropping so that frontier models can be used more cheaply over time.

Logan Jastremski

Yes, completely. It might be interesting to start this topic by saying that agents are starting to gain popularity. We had OpenClaw, which started it all, and then you saw the Hermes model come out. Each iteration was a little more user-friendly than the last.

Then the Grok bot came out, and then you had Muse. What do you think it actually takes to run these things? It’s interesting that people were trying to buy a bunch of Mac minis, and in your opinion, the spending limit for the average person was high. It doesn’t seem realistic to ask the average person to spend $5,000 to $10,000 on a personal assistant.

As these things start to scale, it obviously affects data centers and their buildout. What do you see overall under the hood if you deploy this—if these agents really start to scale to tens of millions, hundreds of millions, and billions of people?

Vikram Sekar

The evolution is very interesting because when people were buying these Mac minis, it was important to ask the question: Why were they buying them? The reason was not always that they could run local models. They just wanted a separate environment where they didn’t need to keep their personal information.

On this Mac mini, you could store what the agent had access to. Even if you were using cloud models, you could make sure never to log into your email on this computer, for example. Never access your bank accounts. You don’t know what this agent is going to do, so you just want to give it its own sandbox environment.

It’s like, “Go crazy. Your damage will be limited. You’re not going to blow up my life. You’re going to stay in this box.” If something goes completely wrong, I just clean this box and everything is fine again.

But then people immediately said, “Look, you don’t have to spend hundreds of dollars to buy a Mac mini. Just buy a cloud virtual machine.” So people said, “Okay, let me just spin this up.”

All this talk about companies like DigitalOcean and other providers came about because everyone wanted to spin up virtual machines in the cloud. However, all of this is difficult to do. First, you need to buy a machine and install OpenClaw. Then you need to set up a virtual machine and install OpenClaw on that virtual machine.

The next step in the evolution was for these companies to say, “Wait, why don’t I do all this for you? I’ll give you a virtual machine. I’ll install the agent, but I don’t want to install OpenClaw. I’ll install something that my company is developing.”

Grok bot was the first one I tried. As soon as it came out, I signed up for $200 and said, “I’m going to use it right away.” It was fun because you could literally launch agents by talking to it. It had these cute icons, and the GUI was beautiful.

I thought, “This is cool. I don’t need to do any of the things I tried to do with OpenClaw. That’s great.” In a matter of minutes, I had downloaded it. My credit card was already on file, so I thought, “Goodbye, $200.” Then I had my agents working. That’s great, isn’t it?

Now it’s more accessible—much more so than ever before. This has implications for exactly what kind of hardware we’re going to need to run all of these things, right? People are now buying into companies like Grok, and now also Muse.

Muse supposedly runs on 2 processors and about 8 GB of RAM. People say it works for every agent you sign up to use. But all of this is stored somewhere. There’s a computer somewhere in the world that works on your behalf, right? That does your work for you.

All of this translates into demand for processors, demand for RAM, and demand for more inference, because all of these machines will now be performing inference without your asking them to. This is the whole point of the agent. It will follow its own logic and tell you what is important and when it is important.

The number of these things will increase. The real question is all about relative amounts, right? More or less, yes, but they all have to grow.

Logan Jastremski

Yes. This is interesting. When Grok got its own virtual machine and was able to start using the computer on the backend instead of using my own, it felt like a magical moment. Even though you could set it up yourself, it was, to your point, just a bunch of extra steps that were a little more cumbersome.

Vikram Sekar

Yes. Yes. I think what’s interesting to me—and I think Elon talked about this quite a bit—is that, on the computing side, there’s a bunch of software that has been invented in the world that doesn’t have APIs or MCP connections. If you can use it like a normal person, it unlocks a lot of functionality.

And it was even more amazing recently, just with Project Astra and certain advances in computer use. I was pleasantly surprised. It’s not yet obvious at a human level, but it seems like agents in each iteration of the model are getting better and better at simply doing things on your behalf. This is amazing. I have a story about this.

There’s another agent thing that we haven’t mentioned yet called Instinct, which is getting a lot of attention right now. It has received a lot of funding, too. For example, even GOTO posted about Instinct on its X account today. This is very exciting technology, and I managed to get an invitation pretty early.

So I thought, “Okay, I guess I have to try it. For the sake of science, let me in. If my stuff gets broken, I’ll deal with it later. YOLO all of this right now.” So I signed up and was granted access. Then I got this conference registration email: “Hey, the Open Compute Project Global Summit is open, and you can register here. Here’s the code and everything.”

I was doing something, and Instinct said, “Yeah, okay. Do it. Sign me up.” I forgot about that and returned to my work. Then I came back to it and got a notification on my phone because I was running it on my iPhone. It said, “Oh, yes, you’ve been registered. Here is the QR code. And, by the way, I ordered you a T-shirt—medium size.”

I thought, “This should be normal, right?” I was like, “How the hell did you know my T-shirt size?” That worked great. It seems like everything had been done. Even my T-shirt size was correct, so maybe it was a good guess. But I thought, “Here I am on the team.” If anything can do all this for me, then I’m on the team.

Logan Jastremski

Yes, it seems that, from zero to one on the software engineering side, it was a pretty slow iteration over time. I’m glad that more people can now interact with artificial intelligence beyond chatbots and LLMs. It seems like agents are moving in that direction—hopefully in a good way, like in the movie Her—where the agent just knows everything about you and can do things for you. Maybe you have a personal relationship with your agent, but it certainly seems like that’s the direction we’re going.

Now, how do you actually make this work under the hood, based on how many gigawatts and what model size you need, versus how much memory you need and whether you need one processor or many? I think this is also an interesting question for people who follow this space and try to dive deeper into some of the more complex aspects of things.

Vikram Sekar

Yes, I have my own opinion on this matter. I don’t know if I’m right, but I’ll say it. For most day-to-day tasks, I don’t think we need frontier intelligence. We don’t really need to run heavy models if it’s something like an additional set of considerations and so on.

Overall, I think a lot of these models used for personal assistance could use Sonnet-level models. Even if you go open source, the cost of launching these things could be even lower. It follows that the number of tokens will increase, but not everyone will use the frontier model, so that has its consequences, right? Maybe we can provide more users who consistently use the Sonnet model than users who are trying to consistently run the frontier models.

There is a big difference in infrastructure needs between these 2 cases, and I think personal agents mostly don’t need the frontier model. You’re not going to ask your personal agent to solve this giant math problem. That’s not the job of a personal agent, is it? These kinds of things will continue to happen on this side, so these will be mid-range models that will do just fine.

The other thing is that, in the whole story of processor allocation, it would be wrong to assume that everyone who runs an agent would have their own processor all the time, right? It’s not the same as buying a Mac mini, but in the cloud. That’s not the case, because when your agent isn’t using it, this equipment will be used to serve other people. This is a shared resource for everyone.

The real question is, how often will sharing occur? Do you know how often you need it? Do you need your agents to constantly work for you and constantly do things for you? Or maybe, in about an hour, it might come to you and say, “Hey, this is what happened.”

I already have my ChatGPT that does this for me. Every hour, I get a notification in ChatGPT: “Hey, you missed these important emails. This one needs your attention immediately, but you can do the rest later tonight.” I already set it up, so that’s about an hour every day. I have other things that I don’t always need an agent working for me.

So, when it comes to hardware requirements, I think a lot of people will get on this train now because it’s easy to use. You can write text messages. Everyone knows how to write to someone, right? You just need to write messages to an agent. That’s all you need to do, so many people can do it.

But I think the hardware requirements themselves won’t be as advanced as the technology all the time. These will be shared processors with lower-level logical inference. Memory will also be shared. Maybe they’ll take your memory out of commission when you’re not using it and spin up your compute just in time to use it. A lot of this will be shared.

There’s a lot of sentiment that everything will blow up because of agents, and it may not actually happen.

Logan Jastremski

Is it because you think it’s more of a shared resource, and that advanced intelligence isn’t necessarily needed for everyday tasks, so you can stay consistent and just bring in additional users?

Vikram Sekar

Yes, you can. Now, you don’t have to stay constant. I think it will continue to grow. We’ll need more computation. We always need more computation. Even for training, the scaling laws haven’t ended, so we can build bigger and bigger clusters and install more and more GPUs. These things are getting smarter, which is amazing.

We’ll continue to build bigger GPUs, and that path has never changed. But the growing demand for these kinds of agents comes from a much broader population that would like to use these things. Of course, it depends on the cost, because right now this GPT dots thing in chat will only be for prosumers, right? Implementation at this stage is questionable. I think Muse is much cheaper to use.

Implementation is questionable, but assuming that a lot of people implement it, they will use more computing resources than we did in the past. But we have to be careful to limit our expectations that this doesn’t completely blow up and we’re suddenly on a whole new wave.

This is good because now AI has become useful for the average person doing normal things. You can monitor your kids’ school calendar and find out when their trips are, when their soccer lessons are, and whether there are any schedule changes. Everyone can use this thing. I think this will be much more useful than it was in the past, and that’s good for the development of AI, right?

### AI spending and the risk of overbuilding

Logan Jastremski

100%. It was really interesting for me to try to understand what exactly the broad adoption curve is, because I think everyone is trying to predict, at least from my perspective, how many gigawatts are going to come into service. Is it 10 this year? Is it 20? Is it 30 next year? Is it 40?

Is it for pretraining, to build bigger and bigger models, or just for inference? Since potentially more of the workload goes to the inference side, maybe it has more to do with memory than GPUs. I’m trying to put all the pieces together. One thing that I’ve personally tried to track at a high level is just the number of gigawatts, because I feel like it’s downstream of everything else—or upstream, so to speak—and everything else falls back on that.

Maybe the space on the agent side obviously increases overall adoption, but maybe it’s not that straight a line, so to speak. What are the general things that you’ve actually been following as a general trend? I think people are divided into 2 camps right now. It’s like an artificial-intelligence bubble: the buildup is crazy, and it’s not sustainable. Or, on the other hand, it’s like it’s completely devoid of AI.

You think, “Of course, all these gigawatts are being built. The demand for inference is insatiable. We’ll never be able to get enough.” Everything else in between is a bit complicated. So, I guess, maybe not on that exact spectrum, but as you see the next year or even 6 months, where are you?

Vikram Sekar

That’s a good question. I think I have a long-term and a short-term view. The long-term view is much clearer to me. I wouldn’t consider myself completely devoid of AI or anything like that, but I really think AI is a useful technology that has beneficial consequences in the long run.

In 10 years, no matter what happens, it’s going to move up and to the right. When the internet came along and we had the dot-com era, we certainly had a crash. A lot of things happened, but it was a fundamental technology that still exists today, and it’s a long-lasting technology that will last a long time.

Now, we can argue about whether this is what LLMs in AI are actually capable of. I don’t know. People say, “No, this probabilistic, statistical approach to intelligence isn’t really true intelligence, and something else has to come along.” Maybe. I don’t know. There are much smarter people than me who know these things better. But considering what AI can do today, I think it’s a useful tool.

I use it for so many everyday things, which has opened up so many things that would have been impossible to implement otherwise, right? Otherwise, even what I’m doing now—between Substack, the podcast, and the institutional things I’m running—is actually too complicated for one person. So I have ways to optimize so many things: meetings, rescheduling. All the administrative work disappears.

### Training versus inference

In the long term, I see this as a very useful technology that’s here to stay. In the short term, do I think we’re overdoing it? Well, that’s a difficult question to answer, because something new comes out about every 6 months. If you had asked me this question last year, in December or November, I would have said, “What, actually? Coding agents—that’s good, right? So coding—is that all we’re going to do with this thing? Okay, what about this chatbot?”

This chatbot used to give you half-answers. It’s not that you ever knew whether it was saying the right things or hallucinating. I think many of those fears have disappeared today. So even in 6 months, it’s very difficult to predict where the trajectory of AI development will go.

Then we had agents. Now we have agents in our phones, in our pockets, right? And now the next question is: What are these phones supposed to do? Is this a game about whether edge AI will finally be useful? Because even though you make some inferences on your phone, like the Instinct to text or the Muse, why do you have to go to the cloud every time? What if you could do some things on your phone, right?

This is the whole next step: What do we do with edge AI? So in the short term, if I look at the financial side, it makes me very nervous. I’m like, “Oh my God, look at the number, the volume of the deployment.” Look at Anthropic's S1 and see how much they’re obligated to deploy in capital and compute resources. This is incredible.

The figure seems to exceed 500 billion. They committed to creating that much. From a financial perspective, this is scary. But from a purely technical perspective, I think a year from now we’ll be in a much better position than we are now, even with this technology, because we haven’t even started implementing it yet, surprisingly. I think there are still many good things ahead.

So I’m optimistic about the technology, but don’t ask me if the financial issues are a bit too much. Are people ordering too much memory? Is everyone buying too much, and will we oversell, fall into surplus, and collapse? All of this can happen, okay? I don’t rule this out at all. But in the meantime, I think we’re still moving forward; we still have a long way to go.

So that’s a long answer to your short question.

Logan Jastremski

No, that’s a great answer. In terms of deployment, what were you most interested in? Because it seems like the leading labs are doing or investing the most, and maybe the hyperscalers are investing over $1 trillion in capital expenditures right now. Obviously, they’re creating smarter and smarter models.

At least from my perspective, it seems that as these models scale, the clusters should be coherent, in the sense that they’re located next to each other. We’ve seen some attempts to do decentralized learning, and for the most part, I don’t think it’s worked. So we’re continuing to focus on coherent clusters that have extremely high throughput, and it seems Elon even talked about building a data center in Memphis later this year.

I think they’re targeting 2.5 gigawatts, then about 8–10 next year. But it seems that as models generally grow in terms of the total number of parameters over time—and Jensen seems to have said this on Brad Gerstner’s podcast—the logical conclusion would be something like X million or X billion. And that seems true, considering that inference is the source of all revenue in general, and training is actually the loss leader for inference.

So how do you see things like this on the ground floor? I know you’re very deep into the technical side of different developments, both in inference and training, and how that shifts the workload a little bit from more GPUs to more memory, which obviously memory stocks have done pretty well this year.

Vikram Sekar

Yes. So the scaling in training, I think, will continue. We’re still trying to create larger and larger world sizes. We want to combine 144 GPUs and scale, then 576 and scale, and find 1,152 GPUs and scale. Will all this continue now? Seventy-two? Yes. Today it’s 72 in the rack.

But we want to go beyond that and put 144 GPUs in a rack, and that becomes very difficult because even a 72-GPU rack today—a Blackwell rack—is like a 200-kilowatt rack. Now, if you want to double that capacity, you will also increase a significant portion of the network capacity. So even realistically, if you just double the number of GPUs and say you have over 400 kilowatts, that’s a lot of power for one rack. Cooling is an issue, and connectivity is an issue, so there’s a lot going on when you put that many GPUs in one rack.

Now imagine you put 576 in one rack or 1,152 in one rack. That’s just impossible, isn’t it? So now you want to put them in multiple racks, right? And this creates many other problems, such as how to connect them. Because it’s no longer 1 or 2 feet, but 1 or 2 meters. To cover about 8 racks, you need almost 10 meters of reach, and copper can’t do that at the current speed of 200 gigabits per second.

Copper can’t handle a 10-meter reach at all—no way. So that’s the problem we’re facing right now. We’re going to continue to move to larger and larger domains and scale, because training is going to continue and is very important as we push the boundaries of what’s possible. We haven’t seen any signs of the scaling laws slowing down yet.

But inference, I think, is a really big source of income, because everyone needs inference, right? We all need a model, but do we all need a model at the edge of possibility? Maybe, overall, it’s a good thing to have. This is a longer-term project, because everyone needs research projects that give you the next best thing, right? Otherwise, what progress is there? Progress must continue, so it will happen.

But inference will reveal much more useful things, right? Agents are just one use case. Chips and architectures are constantly changing from the ones previously used for training. Even Google TPU v8 has a chip for inference and a chip for training. They’ve been doing this since version 4, actually, but this is the first time they seem to have released them together, and they have different chip architectures.

You can even look at data movement in inference and data movement in training; they’re completely different. There are all these weights in training. I like to think of them as waves, like in the ocean: they come in sets and go in sets. That’s how data moves. You have all these operations that are completed, and then there’s this thing called “all-reduce.”

So the whole wave of data comes back, and the next iteration completes, and then the wave of data comes back. This is a very familiar movement of the data flow. But at the level of inference, it’s complete chaos. It’s like looking not at the beach, but at the rapids—Grand Rapids or something. You look at the water and think, “What the hell is going on here?”

Because you have a mixture-of-experts model, you don’t pull all these trillion parameters out of 5 trillion parameters, or whatever. You just pull out a few tens of billions and say that this is my expert. Now you will be using a different expert than me. So think about it from a data center perspective, right? All these things are random. It’s chaotic.

It’s like the weights are being pulled out in all sorts of places. The weights you need may be in the cluster over there, but then your next request might require weights all the way at this end of the room. Literally, that’s what it is. It’s as if the weights could be stored at different ends of the room you’re asking about. So when you look at the patterns of data movement, it’s chaotic, right?

So we’ll see all these things like divergence. Inference now unlocks, because of the nature of the problem, a lot more architectures in chips that might be better suited for this. This is not for training. That’s why you don’t see too many companies trying to build training chips, because these big GPUs are great for training, from AMD and NVIDIA, et cetera.

But when you go into the world of inference, it’s the Wild West. You have Cerebras, you have Groq, you have so many approaches. For example, if you go to Hot Chips, you’ll see so many different things. The Maia inference accelerator—Microsoft Maia—is completely different. Meta’s MTIA is a completely different way to do this.

Of course, OpenAI’s Jalapeño itself was completely different from all these other things. So many different things happen in inference. This is simply incredible. Obviously, there’s a reason why everyone’s focused on this, because there’s a lot of monetary value to be uncovered there as we move into the future.

Logan Jastremski

Yes, it’s super fun. I would say—and my world is now in blockchain, actually—if you were to boil it down to 1 key thing, it has historically been quite bandwidth-constrained. And I think the interesting thing about modern AI data centers that you usually talk about is, as you say, the weakest link in the racks is about 200 gigabits per second.

When you look at some things from the cryptocurrency side, you laugh. This is measured in megabytes, and I think that's so cool. We are pushing the boundaries of what is possible—for example, the physical limits of these different materials. Apparently, they're scaling to terabits and beyond, which is very interesting.

I hope that cryptocurrency and these backend blockchains will reach gigabytes and continue to scale, but so far, it's taken a little longer than I would have liked. That's crazy, isn't it? The speed is crazy. Most people don't even have gigabit fiber at home, you know? Most people don't have gigabit fiber-optic cable. Some people might get it, I think, but that's about it.

Think about the speed in a data center. This is madness. The bandwidth is very high. I think even conventional data centers have 10-gigabit lines and then scale to 100 gigabits. I know I might be a little off, but I think blockchain is pretty much the next evolution of finance.

Like fintech, it's just bringing together all these disparate databases and synchronizing information at a basic level. The bandwidth problem is that you have these different—I call them banks or exchanges—that have relatively low bandwidth, and what you need to do is synchronize the data between them. We've scaled from kilobytes, which was terrible, to megabytes, and now hopefully to gigabytes.

If you take it to the extreme, it's more like high-frequency trading, where you get, for example, 100-gigabit interconnects. You take a bunch of data and do data parsing, and then some interesting things with models of what you want to do. The reason I've been so interested in the data center side, and why I feel like I want to at least try to learn as quickly as possible on the journey from zero to one—and I appreciate your help again—is that a lot of it has to do with various bandwidth issues.

One thing I would like to touch on, in terms of what you mentioned, is moving to scale and also the footprint side. From a rack perspective, it's really interesting. A single server in a rack has a certain amount of bandwidth on the chip itself, and then the rack has some bandwidth with things like NVLink. As you go from rack to rack, the bandwidth varies a lot between those racks.

In simpler terms, can you explain the scaling there and why it's difficult, starting with the bandwidth issue and even scaling beyond GPUs to 144 and so on? Why was that more difficult?

### Why distance limits copper bandwidth

Vikram Sekar

It's always easy to transmit bits very quickly if the distance is very close. For example, if you look at the bandwidth between HBM and the GPU, you get something like 22 terabytes per second. That's extraordinary.

It's really fast because every time you transfer a bit over copper, the longer it stays in that copper line, the worse the signal gets. You want to get from point A to point B and get rid of it, because the longer the signal stays in the copper, the worse it becomes. The same speed won't work once you go from that level of GPU to HBM, which is very, very close.

But when you go between 2 graphics cards in a rack that are, let's say, a few feet apart, instead of a few millimeters, you have to go a few feet. That's pretty bad. It's a big leap in distance, so you can't achieve the same speed. There's no way—you have to slow down.

It becomes slower compared to the connection between chips. To still increase the speed, you have to use all kinds of crazy design tricks where you compensate for what's going on in the copper on both ends. Sometimes you know, “The copper is going to lose this much energy by the time the signal gets there, so what if I boost the energy first so that, by the time it gets there, I've compensated for it?”

On the other hand, you can get the bits and then count them to see if you can detect some sort of error—for example, how many of these 10 bits are wrong, whether you can tell which 10 bits are wrong, and how to fix them. That's why you sometimes need a digital signal processor, or DSP, on the other end, so you can compensate for all those nuisances that happen on copper lines.

This is a problem because the DSP needs time to count the bits and say, “Bit 7 is wrong. Let's change it from 0 to 1.” It also requires energy because you need to spend silicon and energy on the chip that does all these functions. So this becomes a problem.

With copper, whatever you do, we're now reaching 200 gigabits per second. Maybe the next generation will be 400 gigabits per second. That's exhausted. It's as if physics has reached its limits.

For example, when you put one electrical signal on one end of a copper line, it's almost indistinguishable 5 meters away at that speed. It's useless—you can't even tell what's going on. Now the industry is saying, “Okay, look, copper cable has reached its limit at 400 gigabits.”

Of course, 200 gigabits still works. For example, Rubin racks still have 200 gigabits, and copper cable is good for that. But at 400 gigabits, it's completely dead. Even with 200 gigabits per line, within 1 rack everything is fine. But if you need to move to the next rack, which is about 2 or 3 meters away, the coverage is not enough.

Now everyone is saying, “Wait, so I can't put more GPUs in the rack? Isn't there room there, or would I have to put 400 kilowatts in the rack?” In the cloud era, the energy consumption of a data center per rack was about 20 kilowatts. Now the AI rack consumes 10 times more, and we want to double that because it's very complicated, right?

Cooling becomes difficult. Power supply is getting complicated. So they say, “Okay, the best temporary solution while we figure out all this crazy stuff is to put the individual racks next to each other and then connect them.” Think of it as a giant rack made up of 8 smaller racks, right?

This requires connections that reach across 8 racks. Copper won't do that. That's why the whole industry is saying, “Optics, optics, optics.” We need to move faster. How do we connect the racks? How do we get the reach that copper can't achieve?

### Co-packaged optics and the push beyond copper

They say, “Wait, I can't use traditional optical modules that plug in.” They're just plugs—you can see them online. They're energy-efficient. NVIDIA never wanted to do that, or they would have gone to optical a long time ago. Why fight copper when there's a solution? It's too much energy.

The solution the industry is proposing is to move this optoelectronic converter right next to the GPU or the switch. That's kind of like co-packaged optics, right? It saves power, but it also solves the reach problem. However, it creates other problems, like reliability. When something goes wrong, how do you replace it? It's right next to the GPU, so there are many other problems.

It's never been done before. Can we deploy this at scale? Can we connect 100,000 GPUs with this technology? Those are the real questions, right?

Logan Jastremski

Yeah, it's very interesting. I really like the physics of all this because, again, you're pushing the envelope of these different materials. As you mentioned with copper, copper has a certain physical limitation in terms of how far you can actually send a signal over longer distances. Potentially, you need repeaters or other devices there, but it's also much more energy-efficient than something like fiber or an optical interconnect.

Optical cable now has much higher bandwidth and potentially uses more energy. It's like, “Okay, how do we deal with optical interconnect? We want the bandwidth to be high, but we don't want the increased energy consumption.”

Vikram Sekar

We don't want energy usage; we want reach. Copper was essentially free. You just ran energy through it, it ran through a copper wire, and it was simple. Optics is much more complicated, right? You have to convert electrical current to optical and do all that.

But now it's a necessity, and that's why you see so much optical content in data centers. You're only going to see more of it. The optics industry has never seen anything like this. In the past, optics was just for long-haul connections, metro networks, submarine cables, and some other applications. But it's never been as important as it is today.

Today, it's a technology that's essential for artificial intelligence and the future of computing, and the whole world is fighting for it because they're saying, “Look, we were never ready for this demand. This industry was never built for this level of demand. You're asking too much.” That's why everyone's running out of capacity.

Logan Jastremski

So you're telling me, Vic, that I'm going to have a terabit of internet connection to my house soon because of the data that's being fed to you?

Vikram Sekar

You don't need that. You can download this podcast unless you're doing homeschooling.

Logan Jastremski

Maybe, maybe. That's funny. Yeah, that's super interesting. Again, I think the coolest thing I've found is that, because the demand is so high, people's willingness to push this has really reached the physical limits of the material.

That's really exciting, because then you have to come up with new, smart engineering solutions to make all these things work. So maybe let's touch a little bit more on the training side, particularly the interconnections, and then we can move on to the inference side, because I think that's also super interesting.

You start to scale these interconnections, and it seems like people want to have almost one giant GPU, if possible. Obviously, if you could avoid bandwidth issues, that would be ideal, but there's a trade-off. So as you start to scale more and more GPUs and more racks, what have you been following? Is it the interconnect between racks—things like Lumentum and Coherent—or are there other things that you think are more interesting to follow as you scale further? You're only as fast as your slowest connection in the data center. We're starting to scale, and we have clusters of up to 1 gigawatt, now up to 2 gigawatts, and again, I think Elon is pushing 10. How much of that is going to be interconnected versus not, and how much will use the old H100 versus the new B300, is still to be determined, but the direction seems clear.

### New approaches to optical interconnects

Vikram Sekar

Right now, I think the biggest problem today is the interconnect problem. That's why you see these other interconnect technologies, also called “wide and slow.” Instead of running 8 links at 200 gigabits per second each to get 1.6 terabits per second, you can run, I don't know, 50 gigabits per second, but a lot more of them.

Logan Jastremski

How much is that? Like 32, right? Something like that.

Vikram Sekar

So you can run lower data rates but more lanes, like adding more lanes on a highway.

Logan Jastremski

Not like they do today?

Vikram Sekar

No, that's the main discussion right now. We're at this stage where laser technology and optics, when you do 200 gigabits per second per lane, usually combine about 8 lanes together, and you get 1.6 terabits per second per connection.

It's not a 1.6-terabit link; they're sharing it between those 8 lanes. If you have 8 of them running at 200 gigabits per second, that's 1.6 terabits per second. Basically, you need a specific type of optics for that. You need lasers that are specifically made of indium phosphide, which is perfect for lasers and has always been the mainstay of the industry.

The problem is, as I mentioned, nobody has that kind of power. People say, “Look, these types of lasers can connect continents. This is how continents work under submarine cables using this technology. This type of laser can connect data centers kilometers apart, tens of kilometers apart, or even under the ocean. Why would we use this to connect a neighboring rack? Why? Don't you think that's excessive?”

We have no supply. Why don't we make something else that isn't copper? Copper doesn't work. Why do we need to switch to this Ferrari technology that's useful for something completely different? Why don't we just do something else?

Well, we're kind of going back in time, because we're going back to these specific types of lasers called vertical-cavity surface-emitting lasers, or VCSELs. You'll hear most people shorten that to “VCSELs.” These lasers used to do what we're doing now. You can even make them out of gallium arsenide, which is much more affordable than indium phosphide. You can make a GaAs VCSEL, and the supply chain for something like that is affordable.

People say, “Why don't we use this stuff?” The problem is, it doesn't work at 200 gigabits per second. It only works at, I don't know, 20 or 50 gigabits per second. Then people say, “Just put more of them together. Why are you struggling with that? You just tie more cables together, and the speed will come back. Why not just use this technology? We're not constrained by supply chains, and we know that this technology works.”

There's another competitor, LEDs. People say, “LEDs are another light source that can work with optics, so why not use them instead?” But the problem with microLEDs is that they're even slower. You can't push them beyond 5 gigabits per second or 3 gigabits per second. That's the limit. To get that speed, you have to string even more cables together. People say, “No, no, that's a terrible approach, because you have to string hundreds of them together to get that speed.”

Why make hundreds when you can tie dozens together? This is what VCSELs provide. You can just multiply 50 gigabits by 32, and everything is fine. That's the state of the industry right now. There's a debate going on about what's the best way to connect optics in the short-reach realm. Do we really need this Ferrari technology, or is there something else we can use?

Logan Jastremski

Yeah, so everyone is arguing about this. I don't know the correct answer, but this is very interesting. I thought they had a 1.6-terabit link. I didn't know they were sharing it across 8 lanes.

Vikram Sekar

Yeah, because LEDs are also based on gallium nitride, which is another widely available material. You can even create gallium nitride on silicon substrates. You can make really big wafers out of a material like silicon, and you can make a lot of them. It's quite convenient to manufacture. This doesn't require the special material indium phosphide.

The problem with indium phosphide is that it's made on a 3- or 4-inch wafer. I don't know if you know the kind of dough called a stroopwafel. It looks like a small pancake. That's the size of these wafers. It's like a very tiny pancake, and you have to make lasers out of it.

How limited is the supply when you can't even make these Ferrari lasers on large wafers?

Logan Jastremski

So we need more big pancakes?

Vikram Sekar

Yes, exactly. It's also not easy to make these pancakes bigger because they don't provide enough yield. You can't make enough of them to make the process work. There are so many problems, so the industry is saying, “You know what? Forget about it. We'll just do something else.”

The whole industry is in a big mess. What are we going to do about this interconnection problem? This is a very relevant problem right now. To answer your original question about what's happening in scaling, this connectivity problem is the biggest problem that's happening.

Logan Jastremski

The next thing I think concerns the power aspect. How are we going to supply more and more power to these devices? This is a long-term problem, and it's a huge and extremely interesting problem because it starts with energy production. You can talk about everything from nuclear power plants to how that power gets delivered to a GPU.

### The AI memory hierarchy

There are an infinite number of things in between, from inverters and UPSs to energy storage and capacitors. The supply chain is crazy. I want to go back to energy and even talk about space data centers, but I want to talk about the compute side very quickly, given that this seems to be an area of focus, as you mentioned, and it's more like the Wild Wild West.

Even just looking at the equity market, I think there were a lot of questions about people trying to get from 0 to 1 in memory, just because their stock was one of the best this year. I'd like to ask a little bit about the memory hierarchy that's starting to take shape, with SRAM from Groq and Cerebras, then obviously high-bandwidth memory, potential new tiers with high-speed flash and NAND. How do you think that memory hierarchy unfolds when you have extremely high bandwidth—hundreds of terabytes per second at the top with SRAM—versus the limited capacity that it has?

As you go down that hierarchy, you have much more capacity but limited bandwidth. Where does everything sit in that stack?

Vikram Sekar

This memory hierarchy used to be simple. In the days of processors, you had SRAM in the processor. SRAM was where you stored your L1 and L2 caches. Then you had DRAM, quite simply—how everyone bought DRAM sticks for their gaming PCs and put them in there. It was quite simple.

Then you had flash. An SSD drive is one type. Then there's hard-disk storage and other types, and you can talk about an even larger type of archival storage called tape. You could store everything on tape. This is actually something that's talked about very little when it comes to building AI data centers.

If everyone wants to keep everything permanently, the only way to do it is to switch to tape. It's archival storage, very slowly pulling everything out of there. It's the lowest bandwidth imaginable, but the capacity is almost infinite, and the cost is very low. Tape storage is the other extreme.

It used to be simple, but now you have SRAM. Cerebras wants to put DRAM on top of SRAM, and Qualcomm wants to put DRAM on top of the logic. Then you even have high-bandwidth memory, 3D-stacked memory, and all that stuff in this memory hierarchy.

Then, at the DRAM level, you have to decide where you put DRAM and HBM, or whether you use HBM at all. When you move to flash storage, it's another explosion, because you have all kinds of high-performance flash memories where the performance is tuned toward higher speed at the expense of capacity, like single-level-cell NAND.

Or you can move to higher capacity, like quad-level-cell, or QLC, NAND memory. All these things exist. Even here, there's a whole continuum, because you can have single-level cell all the way to triple-level cell storage.

You can have different kinds of combinations and performance optimizations in all possible styles just within the NAND layer. Then you have hard drives and tape, which still exist.

Now the question is, how do you use all this? For example, what do you do in inference? How do you decide? That's where I think there's still a lot of work to do to figure out how to map the memory hierarchy to the workload that you're doing.

Do you always need the highest bandwidth? I don't think so. Everyone brags about bandwidth, just like everyone used to brag about FLOPS. Now you don't see anyone bragging about FLOPS. Everyone is talking about how we can get 2,000 tokens per second or about 1,000 tokens per second, so it all depends on speed.

### Useful work per token

Ultimately, the question always arises: What is the useful amount of work being done per token? Do I always need 2,000 tokens per second? No. As you know, I'm talking to you right now, but my AI could be processing some PDFs that I want to read later or something. I have no immediate need for this.

Ultimately, the memory hierarchy should be used in a way that best suits the job it will be doing, as this will control my costs and, therefore, system usage. I'll use this more often if I can control the costs, right? As a business or consumer, I want more work to be done per token.

So run lots of tokens, but run them slowly for the workloads that need that. Run them very quickly if I'm doing certain types of workloads. If I want something coded really fast for an enterprise, then time is money. They want to build a product faster than their competitors, so they are willing to spend more money on tokens because they can get it back by being first to market. It makes business sense for them.

### How much context does an agent need

That use case requires higher speeds. I think we'll see memory hierarchy being used all over the place to align the use of tokens with the most useful work they can do. I mean, there's a lot of optimization to be done here.

Logan Jastremski

I completely agree. In one of my recent podcasts with Baba Boy [?], I thought about how agent workloads could potentially make a difference.

It seems like the industry has optimized a lot for bandwidth, which makes sense considering you're charging for the token. The more tokens per second, the more you can charge. But, going back to the agent's perspective and what you mentioned earlier about us chatting and potentially your agent being able to go out and do things, it seems like agents don't need maximum bandwidth.

It would be interesting to me if they could have more context. It would seem like you need a larger KV cache. Over time, you forget fewer things. Maybe it goes from 1 million to 10 million. I was expecting the context windows to get bigger, but I don't think it needs to remember everything about me all the time.

When the context arises, it can look for that part of what it knows about me, right? For example, if I'm working out and I want an AI to tell me something like, “Hey, do you know how many steps I took this week?” it can look at my calorie records, steps from my fitness tracker, or anything else. That's enough to store in memory.

It doesn't need to remember anything about my work. It doesn't need to turn to optics and power in data centers when it comes to fitness. I think the context can be broken down. Even for a person, context has sections, right? You don't remember everything all the time.

I think this can be broken down as a problem. What do you think it looks like under the hood? One of my assumptions—and it could be wildly wrong—was that context windows should grow over time because you want your agent to remember more for more specific workloads.

I saw this podcast with Dario Amodei, who said, “Well, eventually I'll have maybe 100 million tokens.” At that point, you can do recursive learning for the context window itself and maybe pin the most important things back into the weights. I thought, “Okay, cool. Let's move in this direction.”

But yes, of course. It seems like a lot of things have been really optimized. According to your point, if one user needs a certain amount of hardware—say, 10 million context windows, and I'm just making up these numbers—you could do 10 people on a smaller model that could potentially provide more logical inferences than one superuser.

People, and I think this is a good thing, are starting to optimize a lot more for the overall user base and to utilize their technology than for extremely experienced users, at least for now.

Vikram Sekar

Yes, that's right. I don't think we know the answer to how much context is enough. We can't answer that question. It seems like more is always better, but at what point does it stop being useful?

Everything that naturally exists in nature has diminishing returns. At what point do we reach diminishing returns in the context window? I don't think we know yet. I think we'll find out soon. That's why I say we're still in the early stages of this. There is so much still unknown that we're still quite early.

Logan Jastremski

What is your guess?

Vikram Sekar

I think about how I think, right? What is my context window? For some things, it's a lot. I remember my children's entire lives—not every moment, but I can trace them back to when they were born. It's still in my brain. But I can't remember what I had for breakfast yesterday.

What context is useful context? Not all context is useful. What is the best type of context to preserve? Not all of it, right? That sets the basis for a useful boundary of context, which has a decreasing value here.

I don't want to forget the youth of my children and everything that is joyful for me. That's why I remember it. But I don't really care what I had for dinner 2 nights ago. That's normal. If it was a delicious dinner, I would probably remember it, but otherwise I don't know.

That information doesn't need to be stored. It all depends on what context you want to preserve and what is important.

Logan Jastremski

I think you also wrote a great article about this with DeepSeek and some of the unique techniques they use with KV-cache caching. When you hit the cache, their costs essentially decrease because you're using the cache instead of doing a prefill, which reduces costs significantly. That was extremely interesting to me because it tries to optimize for as many cache hits as possible.

You can either potentially pass the savings on to the consumer or simply increase your margin because it costs much less than an expensive prefill.

Vikram Sekar

Yes, it's the equivalent of writing in a notebook. This is my NAND memory. It takes longer to search for information in a notebook, but there's a lot more context than I can fit in my head. This is a useful technique.

### DRAM, HBM, and memory on logic

DeepSeek is essentially writing information into a notebook. This is perhaps the last item in the memory hierarchy.

Logan Jastremski

I appreciate your time. If I have to move on to the next stage, let me know and we can complete it. Regarding the memory hierarchy, is there any part of the hierarchy that interests you more than the others?

Are you really interested in SRAM and what they do there? Is it the interconnect, for example, high-speed memory to have larger weights? Is it NAND and a notebook, so to speak? Is there any part of the stack that interests you the most, or is it perhaps the area where the industry is really doubling down on its efforts?

Vikram Sekar

DRAM. DRAM is the best area to focus on, as HBM, or high-bandwidth memory, is required. At the moment, you can't do without HBM. In terms of technology, this is the most important layer because it provides the best balance between bandwidth and capacity.

But stacking becomes a problem. It becomes too expensive and has questionable future value. How high can you stack? We are already at 16-high. Are we going to go to 20? What next—24? This becomes a bit illogical at a certain point. Maybe we'll keep stacking things.

I think one useful way to do this is to put the memory on top of the logic, and you can see a lot of companies doing that. Cerebras is going to bond an entire DRAM wafer on top of its wafer-scale engine. You'll have DRAM on top of SRAM, fully bonded wafer to wafer. I think Groq will have a version of that as well.

D-Matrix already has its Raptor engine, which has logic on top of memory. Qualcomm has also announced high-speed computing, which places perhaps 2 or 4 levels of DRAM stacks on top of the logic.

Now you're suddenly opening up the entire area of the chip to connections, not just the edge of the chip, right? You have a lot more bandwidth. You get SRAM-like bandwidth with DRAM-like capacity. This is a very important area to work on.

### The next generation of inference chips

Logan Jastremski

This is super interesting. If you're interested in the investment side of private business, are there any companies that you find uniquely interesting in that regard? As you mentioned, on the funding side, it's kind of the Wild West. People are experimenting with a lot of different things, like memory hierarchy. We're trying to see where everything fits best, figure out different arrangements of things, and work out the connections.

As you continue to scale these things, are there any companies that you find uniquely interesting? I know companies like Etched just raised a big round of funding. Are there other Weka-like companies? I know you guys mentioned Weka on the podcast as well. They're extremely interesting in how they do data sharding.

Are there any companies that you find interesting and that you think are worth double-clicking on?

Vikram Sekar

There are so many of them.

That’s the whole point of tracking all these private companies—startups and other things that are doing logical deduction. It’s very interesting. I like the 3D DRAM approach, so I mentioned that D-Matrix is one of them. I had previously considered SambaNova.

I think their architecture gives you a different kind of optimization. It’s more of a TCO-based optimization, where they can provide services to a different end user. So it’s not always about maximum productivity. I think they have a different customer profile in mind for their hardware.

Etched is very interesting. Although I don’t really understand their low-voltage pin architecture, it seems very interesting. I saw their equipment at Hot Chips, and that was really cool. I would be really excited to see what comes out of Etched.

Then we have MatX, we have Fractile—we have so many companies doing inference that are all doing slightly different things. I thought it was the Wild West of inference chips. This is really fun.

In terms of lasers, you have companies making lasers so that now we can connect inter-chip connections using optics. Why use copper at all, right? There are so many interesting things in this field. The company that does this is called Nubis. They have these lasers, which they call nanolasers for scaling. I find it really interesting.

### AI's biggest engineering bottleneck

Logan Jastremski

Yes, there are many different things happening. So maybe, if you could push for an answer, what do you find most uniquely exciting about inter-chip interconnects? Do you think it’s ongoing innovation on the memory side, from a stacking perspective? Or perhaps, if you could close your eyes and fast-forward 5 years, what do you think are the most pressing engineering challenges in the data center stack as we continue to scale?

Vikram Sekar

Memory is the biggest problem. Literally, an LLM is a method that requires a lot of memory, and memory is everything. One of the biggest problems people have to solve is how to get the same performance without using as much memory. Once you solve this problem, a lot of things will fall into place, right?

If you don’t have a memory constraint, then you probably don’t need to switch between memory as often, and your interconnect problem will become easier. You might not have any memory-bandwidth issues because of the way it computes everything. Right now, we’re leaving so much compute on the table. GPUs have compute capacity that we don’t use because the network doesn’t transfer data fast enough.

If we can figure out a way to do LLMs that aren’t memory-intensive, that would be very, very interesting, because that would eliminate the memory bottleneck, and therefore the memory-bandwidth bottleneck, and therefore the network bottleneck. The real transformational change happens at the top of the stack, because it’s this bunch of esoteric LLM equations that are driving this whole AI industry, right?

The “Attention Is All You Need” paper from 2017—we implement these equations into hardware. Now, in 5 years, let’s say some really smart people at some university or company research lab say, “Hey, why don’t we do these equations instead?” We can do the same thing, but without all this madness. This will change everything. This is the real turning point that we should pay attention to.

Every time something comes out, like TurboQuant or something like that, everyone thinks, “Oh my God, the need for memory has plummeted, hasn’t it?” Everyone is afraid of memory. So yes, that’s something to pay attention to.

Then we’ll go back to Jevons’s paradox again, increase overall usage, and then we’ll go back to where we started. That’s right, isn’t it? I don’t know when this Jevons’s paradox will cease to exist. Until now, every time there is some improvement, you think, “We’re back to square one. Sorry, we still don’t have enough computing resources. We still don’t have enough.”

I think this is a useful exercise for everyone listening to this, and I would like to do it myself. The exercise is this: find a time in history when Jevons’s paradox didn’t work. As a challenge, your listeners can write back to you and say, “Hey, I found one. This is where Jevons’s paradox didn’t work, and here’s why it didn’t work.”

I would love to have a case like that, because it would tell us when this thing would stop constantly consuming all of our computing resources. I guess the moral of this podcast is: the show goes on, and it’s a deep dive into memory. Let’s try to figure out ways to use memory more efficiently, but it looks like memory is still a bottleneck.

Memory is the foundation of all AI inference and learning. There’s no way to change that, so—

Logan Jastremski

Yes, exactly. Well, I want to respect your time, Vikram, so thank you very much for joining us. I feel like we could keep talking for another hour about all of these things. Again, I really appreciate your clarity of thought, because I think that’s a really distinctive trait of an expert: to explain complex things simply, and you do it wonderfully.

Vikram Sekar

Absolutely. Thank you, Logan. It was a lot of fun to be on the pod. Thank you. Thank you for inviting me.
