# Naveen Rao: 4D Computing, AI's Energy Wall & Beating Biology

All-In · 2026-09-21 · 23 min · https://www.youtube.com/watch?v=yAsrMA_ADPc

## Transcript

Speaker 1

Naveen Rao, co-founder and CEO of Unconventional AI, which is an AI chip startup. Best known for building and selling 2 deep-tech companies, Naveen is kind of a definitionally outlier founder. “When I came there, we had about a $20 million business, and it was $700 or $800 million when I left.” “I don't think you really understand something until you can build it.” “Just because something is tried does not mean it's wrong.” “I'm the opposite of an AI doomer. I think AI is the next evolution of humanity. We need innovation on the hardware substrate to actually build true intelligence.” Please welcome Naveen Rao.

Naveen Rao

Hey, everyone. Great to be here. Switching gears a little bit to AI now, which you may have heard a little bit about. It's super exciting to be at this conference specifically because, as was said in the intro, I'm the opposite of a doomer. I think AI is one of the most transformational technologies that humanity has ever created and will enable us to get to that next level of evolution, which I'm here for. This is sort of the anti-doomer conference, so let's go.

Before we get going, I'll tell you a little bit about myself. It's kind of weird: I'm really right where I wanted to be my whole life. This was me at about 5 or 6 years old, something like that. We had a computer very early on, so I'll date myself: this was in 1978. We got a computer. This is probably in the early '80s. I learned to program when I was a little kid. I just thought it was like a puzzle.

I became an electrical engineer, really because I enjoyed science fiction and always wanted to think about how I could make an intelligent machine. Then, after a career in building computers, I went back to school and got a PhD in neuroscience. The idea was, “Let's go back to that thing. How do we make computers intelligent?” Fortunately, the whole world kind of moved in this direction. As a technologist, it's sort of the dream right now.

A little bit about me from a tech entrepreneurship standpoint: I actually founded the first AI chip company, Nervana Systems. This was in 2014. If anyone remembers back then, there was no AI, or at least not in the common vernacular, and it was really hard to convince people that this was important, much less to build hardware around it.

You heard from Jensen up here—the largest company in the world, a hardware company because of AI. So we were early on. I think I sold the company way too early to Intel, but I started and ran the AI group at Intel. After I was done with that in 2020, I started thinking about the next problem: How do we build bigger models, like the large language models we talk about today? How do I build the infrastructure to build those models?

We started platformizing GPUs and enabling them to scale, making that easy to use for other people. After ChatGPT happened in 2022, we were kind of the best game in town for people to start building their own models. It took off really fast. We decided to join forces with Databricks. That was in 2023, and that's a quarter of the total revenue of Databricks today. A lot of fun doing that whole thing with Ali and the team at Databricks.

Now I want to tell you about Unconventional AI, which is rethinking the foundations of how a computer works. We're going back to first principles here, really trying to build a new machine. Computers have worked a certain way for a long time. We want to rethink that for the singular purpose of making something very power-efficient.

The goal was initially to get to a 1,000× power efficiency within 5 years. I've actually revised this to 3.5 years because things have gone faster than we anticipated. We've actually solved very deep scientific problems more quickly because of AI, interestingly enough.

Just a little bit about how we're organized: We're truly a top-to-bottom company. We start with theorists. These are people with math PhDs and backgrounds in theoretical neuroscience, that kind of thing. They come up with concepts that we think would effectively give us more power efficiency from the perspective of moving less information around.

We then translate that into models that do real things, trained on real data and evaluated against real criteria. So it's kind of the rubber hitting the road for these concepts. Then eventually we have to actually build something physical. These are people who architect a physical circuit, actually design those circuits, model them, see if they work, and try to connect this whole stack together.

### Is energy really the problem? The cost of a token, power contracts & the gap to close

Eventually, we have to build a system and a board and all that kind of stuff, and build a product. Is energy really a problem? I'm not sure how much everyone in this audience has thought about this, but, interestingly enough, I'll give you some data points here.

This is one company. This is just Google. I'm using Google because Google has actually talked about this publicly. Per month, they cross 3.2 quadrillion tokens. It's a crazy number. I never even think in quadrillions, but that's the world we're in today.

If I just take 10 joules per token of energy—this is actually on the lower end of the energy spectrum for models—but let's just take that number and multiply it out, this is 12 gigawatts. The US puts about 40 gigawatts of energy into data centers today, and we're about half of the data center capacity of the world. We're under 100 gigawatts of data center energy in the world today.

Twelve gigawatts is going into one company just for AI services. You can imagine that if models get bigger, that energy goes up, and if demand grows, which it is, that energy goes up. We're going to run out of energy pretty fast—in about 3 years or so, is my estimate.

To put it graphically, this is what we have. We have this huge market that's growing exponentially—call it a $1 trillion market in 2030. Maybe it's bigger than that. Then we have this kind of linearized energy at the bottom. You've heard a lot about this today, but this gap is the problem. We want to solve that gap with technology.

I don't know if people are aware of this, but the way we think about data centers has shifted over the last several years. It used to be about floor space. Can I get the floor space? Can I get the rack space? Then it was about networking equipment. Then it became about GPUs. Today, it's about energy.

First, you think about energy. I get the energy contract, and then I have to figure out how to fill it and basically create infrastructure out of GPUs and things like this. About 50% of the cost of serving a token is energy. Every time you try something on ChatGPT, 50% of that cost is energy. The rest of it is the capex of the hardware, the floor space, and all that kind of stuff.

Today, we sort of think about it as: I get a power contract; I need to monetize every watt. Simply put, our business case is pretty easy: We're going to monetize that 1,000× better than existing hardware.

Then the question becomes, okay, great, that all makes sense, but can we actually do it? How do we build a better, more efficient computer? Biology actually provides some proof for us here. The human brain, you may have heard this, runs on about 20 watts of energy.

What's even more remarkable to me is animal brains. That red number is how many neurons are in the brain. If you scale it linearly down to a monkey's brain, it runs on 1 watt. To put that in perspective, the cellphone in your pocket runs on about 1 watt.

Other animals, like rats and bats and things like this, run on milliwatts of energy. Something that's pretty relatable to everyone is a squirrel. You've probably watched how accurate they can be. They jump between branches, and they do it perfectly 1,000 times out of 1,000. Their brain runs on 8 milliwatts of energy. You could run over 100 squirrel brains on your phone, and they have very precise and accurate behavior.

Biology created something quite incredible. In fact, it's the right kind of physical substrate for intelligence. This is a quote I love: “I don't feel like we truly understand something until we can create it.” We've gotten a lot better at creating intelligence systems. However, they do it in a kind of inefficient way.

What kind of inefficiency is there? As I hinted at the beginning, most of the energy in a computing system goes into moving information around. Just to put some numbers on it, the human cortex—the squiggly part of your brain, the outside of it—only moves about 16 billion bits per second. That's actually kind of a small number if you think about it, because there are some 13 or 14 billion neurons in that cortex.

A GPU, or a high-end computing system, moves nearly 30 trillion bits in and out of memory per second. That's outside the chip. Inside the chip, it's probably 10–100× more than that. We're moving a lot more bits in these synthetic systems than the brain does, and that's actually what drives the energy demand.

How did we get here? Computers have been around for hundreds of years, actually, in some form. They were mechanical. They became analog around the turn of the century, and they became digital back in the 1930s and 1940s. The operation of that computer in 1940–1945 is actually very similar to how they operate today. There's not a huge paradigm shift.

You actually have this memory on the outside, some kind of computing, and you move bits back and forth. That operation creates a machine that just requires a lot of movement, but it's very fast.

We built computers to be fast. They were always faster than the alternative. Incidentally, that computer in 1945, called ENIAC, was built to do artillery trajectory calculations, and it was built to do them faster than the alternative. The alternative was humans, who actually did the calculations.

Now the alternative typically is some other computer: this computer is twice the speed of that computer. That's how we sell computers, but it doesn't contemplate energy efficiency. And that's what we're changing. These trends have been going on for a long time, where the number of transistors kept going up, but we couldn't keep scaling the frequency. We couldn't keep scaling single-thread performance. And now we're actually not scaling efficiency any longer.

### Cutting out the middleman: abstractions, dynamical systems & a new kind of machine

Moore's law, if you may have heard of this, is that making transistors smaller has largely ended. So we're not seeing efficiency gains just from making transistors smaller. We need to rethink the problem a bit. So how do we do it? A good intuition is that we cut out the middleman.

Computers have been built up with these abstractions. I mentioned digital. Digital means 1 and 0. That itself is an abstraction of the physical world. We don't actually have systems that behave as 1 and 0. A transistor actually has states in the middle, but we engineer it to behave that way. That's an abstraction.

We kept building these abstractions up, and eventually we started creating neural networks and learning machines on top of them. Each one of these abstractions is lossy. It means it's inefficient. It doesn't contemplate all the complexity underneath it. That's why it's an abstraction.

What we're doing is simplifying this in some sense. We find an abstraction of the physics of the semiconductor and connect that to the neural network. If you think about your brain for a moment, it has a bunch of neurons in it, but there's no linear algebra. There's no floating-point math. It's actually the physics of the neurons that gives rise to intelligence. We want to mimic some of that with a semiconductor.

This is also not a new concept. There's computation all throughout nature. Birds in a flock—you may have seen things like this—where a bird does a simple behavior: it looks left, looks right, and figures out where the next bird is going. When they do that, they actually create these interesting flocking behaviors, this emergent behavior. We see that with ant colonies. Ant colonies actually do intelligent things just by following very simple rules.

This study is called dynamical systems theory. It's basically how I get these emergent properties from very simple behaviors of individual components. Our brain actually works this way as well. We're taking these ideas and starting to build circuits out of them.

Let me give you an example of such a system. Everyone here probably knows what a metronome is. When you're learning to play the piano or something like that, it's just tick-tock, tick-tock—a physical thing moving back and forth. If you put multiple ones of them on a rigid plank, and that plank can roll back and forth, you'll actually see them start to synchronize.

That's just due to the physics of the system. They each push against the plank just a little bit, and even if they're off by a little bit from each other, they'll all synchronize to exactly the same phase. You can actually scale this up to hundreds of metronomes on a physical system, and they'll all synchronize. This is a form of a dynamical system that goes through some starting point of all these different phases and always synchronizes.

You can imagine a slightly more complicated version of this where maybe they don't all synchronize. Maybe half of them are synchronized with each other, and the other half are synchronized in an opposite pattern or something like this. The idea is that this is a physical system that behaves the way it does just by the inherent interconnection of the system itself.

We asked the question: Can we actually use such a system, like that metronome system, to do computation? We want to connect that to generative AI. That's what we really care about. So we released a model we called UNO, which actually demonstrated this. It's an image-generation model built on a set of oscillators like that. We simulated it and made it open source so people can play with it.

This was the first demonstration that I could actually scale something up, train it, and get useful output like image generation. This is some of the analysis that we provided in that write-up. What you're seeing here is what we call a state-space trajectory. It's basically how the system evolves in time.

You can imagine characterizing the state of the system as all the phases of those oscillators, then looking at how it evolves through time and conditioning that on the output. You say, “I want to generate an airplane, a car, or a bird.” It will actually go through different state-space trajectories. That's what we're seeing here: an analysis of this. These are actual images that were generated by it.

It turns out we started to build a lot of that fundamental science up over the last couple of months, and we found other things that enable this to work even better. This is a concept we call sparsity. Sparsity means that if I have a bunch of elements all connected to each other, like we have on the left-hand side there, and I have 10 elements that I want to connect to each other, I have 10 × 10 elements. So I have 100 connections. That's okay.

But if I have 1,000, now I have 1,000 × 1,000, which becomes 1 million. This doesn't scale very well. We call this n-squared scaling. The more I add, the harder it becomes to scale. Sparsity allows us to say, “Can I throw away some of those connections?” If I throw them away, can I actually preserve the behavior of the whole system?

It turns out you can not only throw away some of the connections, but you can actually get better behavior out of the whole system. It becomes more trainable. There are a lot of theoretical reasons for this, but we're able to do this not only in simulated systems, but in real physical systems.

It's one of these rare things where you get something that's more efficient, more scalable, and actually gives you more performance. This is kind of a holy grail. It's been a problem for a long time, but we had to frame the problem the right way to actually find this solution.

This is the first time I'm talking about this publicly. I wanted to do it at this venue because I think it's a really big deal. This is the first physical dynamical computer ever built. We did this in 5 months. This company started in earnest in January. We didn't even have a team, but we said we were going to build this first physical prototype and do it this year.

We taped out the design—meaning we sent it to the fab—on June 1. The chip is back in our lab, and we actually have results from it. These are the first-ever images generated from such a computer. Thank you.

Now, great, cool—but does it do anything useful beyond just images? You can actually do any kind of task, like sequence modeling or language models. The interesting thing here is that it's only 500 or so nanojoules per image. To put that in perspective, a normal computer, like a GPU, is on the order of millijoules. A nanojoule is 10^-9 joules. It's really, really small, so it's many orders of magnitude more efficient than a standard computer, and it's because it just doesn't move information around.

This is proof positive that it works. This is literally the first time we're talking about it publicly. Thank you.

What's cool here is that this is really the emergence of something new. Computers have gone from CPUs to GPUs, becoming more and more parallel, to compute-in-memory, which is even more parallel and fine-grained. But all of these are what we call von Neumann architectures. They have memory and compute, and we move information back and forth.

What we built is what's called a dynamical computer, which actually has compute and memory in one thing. We don't have a memory interface. Each individual computing element is a memory. It's a completely different way to look at the problem.

We call this 4D computing, where we use the time dimension in the dynamics, and we use the physical 3 dimensions of die stacking—putting things vertically—as well as in a planar form. We have 3 dimensions from the physical structure, and we have 1 dimension in time. This is a new way of thinking about a computer, and it really is proving to work for efficiency's sake.

What are the implications if we build something that's 1,000 times more power-efficient? I think this is pretty cool. Intelligence per watt is what we care about. Can we optimize this and make it better over time?

There is actually a thermodynamic limit that you can never exceed. Mammalian brains—animal brains—are somewhere within 1 or 2 orders of magnitude of that. Today we're on the far left of this graph, and we're about 10 billion times away. That's 10 billion—1 with 10 zeros after it—from that thermodynamic limit.

We think in 3.5 years we can hit the limits of 2D lithography. The overarching goal of this company is to beat biology. We want to make something better and enable compute everywhere, including compute in new robotic forms and things like this, in the next decade or so.

I think what'll be interesting is that we'll see the shift from big data centers with gigawatts to many small data centers all over the place. I think this is a good thing. It actually makes things more environmentally friendly, more local, and more adaptive. As I said, I think enabling the ability to build billions of robots that dynamically assemble to solve problems in our world is actually really cool. This is something that will enable us to think about bigger problems and do even more.

I talked about AI being a trillion-dollar market. If we disrupt it by 1,000×, there's a concept called Jevons's paradox: when you drop the underlying cost of an asset, you actually consume more than the drop in the cost of that asset. If you make something half the price, you'll consume more than 2×. If you make something 1,000th the price, you'll consume more than 1/1,000th of it. I think this will create the largest market that humanity's ever seen.

Chamath Palihapitiya

I mean, that was extremely unexpected. I've got to say, that was pretty amazing.

### Chamath joins: the path to product, porting existing models & building the team

Let me start with probably the thing that's on everybody's mind. To the extent that this works, Nav—and you probably saw Jensen earlier—there just needs to be an entire ecosystem of people beside you and around you, whether it's the fabs, packagers, et cetera. What does it take to get from this early version to something that sits in somebody's hand or that people use? How do you see that path in terms of time and complexity? What does that look like?

Naveen Rao

Yeah, timewise, we're within 2 years of getting it to a full product.

Chamath Palihapitiya

And what is the product?

Naveen Rao

Yeah, it's a VM that sits somewhere that you guys manage, and effectively we're building a new data center product. So it's a whole rack as a system, right? The idea is that we'll run those models on it. Tokens in, tokens out through a network cable, but the inner guts are completely different from an existing computer.

Chamath Palihapitiya

Do you expect that you'll have to move to support the existing model families and existing architectures? Will this work in a world where we've spent all of this time thinking, okay, KV cache, and this is all just so mechanically reductive based on, as you said, these abstractions that we've lived on, right? How do you expect the rest of us to move toward this? I think you see that efficiency curve. We'd all want it. So how do we take advantage of it?

Naveen Rao

Yeah. I think there's a sliding scale between how much better something is and how much pain you'll take to move to it. I basically took the tack of, "Let's make it really, really compelling to move." There is going to be some work to port things over. We actually don't port at the operations layer; you port the model layer. So yes, the existing models will work, but there's a fair bit of compute required to make that transition happen.

Chamath Palihapitiya

And very basic elements like matmul—does that exist in yours?

Naveen Rao

I mean, you can characterize it as matmul, but it doesn't implement it as matmul. It implements it as a sort of time-varying behavior. But each one of those time steps, you can analyze as basically a matrix of the current state times a transition matrix.

Chamath Palihapitiya

And when you're building a team like this, who are these people? These are biologists plus physicists plus what are these people?

Naveen Rao

They're sort of theorists. Dynamical systems theory has been around for 100 years.

So we got people from that world, and then we got people who actually build chips.

Chamath Palihapitiya

And they don't talk to each other.

Naveen Rao

They don't talk to each other. So we had to facilitate that. That's actually one of the most challenging things about this company: the span of talents that we have is so big that getting them to all coordinate and build one thing is actually pretty hard.

Chamath Palihapitiya

And what is this CUDA-like equivalent, if you will, just to use a bad analogy, that allows these people up here to talk to these people down there?

Naveen Rao

Yeah. We actually built a set of libraries in Python. It's Python. It's not CUDA, but it's a language of sorts that allows you to express time-varying elements that have stochastic behavior.

Chamath Palihapitiya

I mean, it's incredibly impressive. It's so ambitious. Thank you very much. It's great to see you. Great to see you. Amazing.
