[BidClub_]
The a16z Show · · 22 min

How Fei-Fei Li Is Rebuilding AI for the Real World

Erik TorenbergFei-Fei LiMartin Casado

YouTube
TL;DR
  • World Labs’ core bet is that language is a powerful encoding of thought but a lossy, insufficient encoding of physical reality, leaving spatial intelligence as a foundation-model frontier. Fei-Fei Li calls language “a lossy way to capture the world”; a world model must understand 3D structure, shape, and compositionality so machines can act, not merely describe.
  • Martin Casado sees the sequencing as the surprise: language became “unit economic positive” almost immediately while autonomous navigation absorbed roughly $100 billion over 20 years. LLMs did not settle the spatial problem; their generative success offered clues for tackling the older, harder layer.
  • The proposed platform reconstructs a complete 3D scene from one or more 2D views, including geometry the camera cannot see. Once a model can fill in “the back of the table,” software can measure, move, stack, and manipulate objects—supporting architecture, design, robotics, games, and other horizontal markets.
  • Depth is operational data, not a visual embellishment, because physics and interaction happen in 3D. After temporarily losing stereo vision, Li could not drive on the highway at speed and drove near 10 miles per hour locally because she could not reliably judge the distance between her car and parked vehicles.
  • World Labs pairs reconstruction with generation, opening embodied-machine applications and synthetic environments. Li’s expansive call is that AI can create “infinite universes”—for robots, creativity, socialization, travel, and storytelling—and “enable us to live in the multiverse.”
  • The execution thesis rests on combining AI, computer vision, graphics, optimization, data, and established 3D techniques inside one concentrated team. The ingredients include NeRF, Gaussian splatting, GAN-era image generation, and style transfer; the company-level unlock is bringing “compute, data, talent” together to productize one north-star problem.
Digest · the substance, structured for research

1. World Labs began with a shared conviction that AI lacks a world model

  • Li was seeking more than financing: her “unicorn investor” needed to endure deep-tech ups and downs while serving as an intellectual partner across computer science, product-market fit, and go-to-market.

  • At an AI dinner dominated by LLM enthusiasm, Li leaned toward Casado and asked, “You know what we’re missing?” Her answer—“We’re missing a world model”—matched the conclusion he had reached independently through image investing.

  • Li later tested whether that agreement was substantive. Casado defined a world model as AI that truly understands the world’s 3D structure, shape, and compositionality—unlike the “polite nod” she received from most people she spoke with.

2. Language’s rapid victory exposed the unsolved spatial layer

  • Li’s honest surprise is that data-hungry models have produced such powerful emergent behavior, despite her own role in making data central to AI. Yet language remains “a lossy way to capture the world.”

  • Her distinction is structural: language is purely generative and does not exist in nature, while the physical and perceptual world is already there. Animal intelligence developed through perception and embodied interaction; humans use that intelligence not only to survive but to construct and change the world.

  • Li frames World Labs as a North Star problem rather than a company or papers project. After years of academic research, she concluded that an industry-grade effort concentrating compute, data, and talent was needed to bring it to life.

  • Casado had thought the field’s path was to solve world navigation first. Industry invested roughly $100 billion in autonomous vehicles; after the 2006 DARPA Grand Challenge seemed to signal completion, practical arrival still took about 20 years. LLMs appeared “out of nowhere” and solved many language tasks almost immediately.

  • Li’s qualification matters: “We’re not here bashing language.” ChatGPT and foundation-model breakthroughs instead convinced her that the moment for world models was closer: “I didn’t need an LLM to convince me an LWM is important.”

3. A machine cannot act on depth information it never received

  • Casado’s blindfold analogy makes the gap concrete: a verbal description of a room gives someone a very low chance of completing a physical task, while sight lets the brain reconstruct the space and manipulate what is inside it.

  • Li traces spatial intelligence through animal evolution: trees do not need eyes because they do not move. Perception became essential because animals navigate and interact; human spatial reasoning later enabled discoveries such as DNA’s 3D double helix and the buckyball’s carbon structure.

  • Torenberg’s pushback—why not remain in 2D?—draws a categorical answer: “Physics happens in 3D and interaction happens in 3D.” Humans infer depth from video, but a robot receiving only 2D output lacks the Z-axis information needed to judge distance or grasp an object.

  • Li experienced the failure personally after a corneal injury temporarily removed her stereo vision. Even on familiar neighborhood roads, she slowed to nearly 10 miles per hour because known object sizes could not replace reliable distance measurement.

4. Reconstruction and generation make the opportunity horizontal

  • Li places the first demand clusters in visual creativity—film, architecture, industrial design, machinery—and embodied machines, a category extending far beyond humanoids and cars. Every such machine must understand its 3D environment and sometimes collaborate with people inside it.

  • Casado’s concrete product primitive starts with a single 2D view and reconstructs the full 3D scene, including unseen surfaces such as the back of a table. That representation can then be manipulated, moved, measured, and stacked; the same system can fill in what was never there and create a 360-degree representation.

  • Li imagines “infinite universes” for robots, creativity, socialization, travel, and storytelling. Casado emphasizes the opportunity’s horizontal scope, analogous to one LLM serving emotional conversation, code, to-do lists, and self-actualization.

5. The technical stack joins 3D research with industrial concentration

  • Li calls world-model research newer than LLM research, but not brand new. Computer vision already supplied “bits and pieces,” including Ben Mildenhall’s NeRF work, which made deep-learning-based 3D reconstruction widely influential roughly four years earlier.

  • She also cites Christoph Lassner’s pioneering work as part of Gaussian splatting’s renewed popularity as a representation of 3D volumetric data, plus Justin Johnson’s foundational image-generation work before transformers. Li says GAN-based image generation and style transfer helped popularize ingredients of the current effort.

  • World Labs’ strategy is to concentrate specialists in computer vision, diffusion models, graphics, optimization, AI, and data around one north-star problem. Casado’s outside assessment is that solving the problem requires both model intelligence and graphics expertise: the data and models must be paired with a workable computer representation.

Fei-Fei Li

Space—the 3D space, the space out there, the space in your mind’s eye—spatial intelligence is a critical part of intelligence. Suddenly, we can actually create infinite universes. Some are for robots, some for creativity, some for socialization, some are for travel, and some are for storytelling. It suddenly will enable us to live in the multiverse.

Erik Torenberg

Martin, why don’t you briefly brag on behalf of Fei-Fei a little bit and share how you would summarize her contributions to AI for people who are unfamiliar?

Martin Casado

She’s someone who doesn’t need a lot of introduction, and she’s done so many things that I can’t fit them all in. So maybe I’ll just do the ones that are appropriate to this.

Of course, she was on the Twitter board. She was a Google executive, founder and CEO of World Labs. But very importantly, as we all know, AI—we all talk about neural networks, and there are a number of people who focused on making those effective—but Fei-Fei really singularly brought data into the equation, which now we’re recognizing is actually probably the bigger problem, the more interesting one. And so she truly is the godmother of AI, as everybody calls her.

Erik Torenberg

And, Fei-Fei, why did you have to have Martin as the first investor?

Fei-Fei Li

Well, first of all, I knew Martin for more than a decade. I joined Stanford in 2009 as a young assistant professor, and Martin was finishing his PhD there. Martin’s adviser, Nick Matune, was a good friend, and I knew Martin would go on to become a very successful entrepreneur and a very successful investor.

We would see each other and talk about things. But as I was formulating the idea of World Labs, I was looking for what I would call my unicorn investor. I don’t know if that’s a word, but that’s how I think about it: someone who is not only an established and successful investor who can be with entrepreneurs on this journey through the ups and downs, who can be very insightful, and who can bring the kind of knowledge, advice, and resources, but I was also particularly looking for an intellectual partner.

What we are doing at World Labs is very deep tech. We are trying to do something no one else has done. We know with a lot of conviction that it will change the world, literally. But I need someone who is a computer scientist, who is a student of AI, who understands product-market fit, go-to-market, and consumers, and who can be on the phone or in person with me every moment of the day as an intellectual partner.

Martin Casado

The origin story of us first connecting is actually pretty interesting. Fei-Fei has clearly been thinking about this idea for a very long time, well before it started—maybe for years, even—and she’ll talk about it. She has this very deep intuition of what AI needs in order to navigate the world.

But we were at one of Mark’s fancy dinners or lunches, and there were a bunch of AI people. Everybody was so excited about LLMs, and they were talking about language. I’d come to this independent conclusion, just because I’d actually done a lot of image investing, that that wasn’t the end of the story.

Fei-Fei, at the end of this table with all these people talking about it, leans over to me. She’s like, “You know what we’re missing?” And I said, “What are we missing?” She said, “We’re missing a world model.” And I’m like, “Yes.”

It kind of fell into place then, because I’d been thinking about this at a high level. She just perfectly articulated it, as she does, and she had a year’s worth of thinking about this and had talked to people. In some way, we had arrived at a very similar intuition through our own crooked paths. Hers was way more filled out. Mine was just kind of this fancy thing.

After that, we had a number of conversations where we both agreed that we were aligned on this idea.

Fei-Fei Li

Actually, I don’t know if you know this. During that lunch, we hit it off on this world-model idea, but I was at that point already talking to various people—not just computer scientists and technologists, but also investors and potentially business partners.

To be honest, most people didn’t get it. When I said “world model,” they nodded, but I could just tell that was a polite nod. So I called Martin. I said, “Do you mind coming over to the Stanford campus and having coffee with me?”

I said, “Martin, can you define your world model to me?” I really wanted to hear if Martin actually meant it. The way he defined it—as an AI model that truly understands the 3D structure, shape, and compositionality of the world—was exactly what I was talking about. I was like, “Wow, he’s the only person so far I’ve talked to who actually meant it.” It wasn’t just nodding.

Erik Torenberg

Okay, so we’re going to get to World Labs and the specifics of this, but maybe first let’s take you both back to your PhD days and your professor days and reflect on this: If you could go back in time with knowledge of what’s happened in the preceding 10 years in AI, what do you think would have been the biggest surprise? What was the thing you didn’t see coming that would have shocked your younger self?

Fei-Fei Li

Yeah. It’s ironic to say because, as Martin said, I was the person who brought data into the AI world, but I still continue to be so surprised emotionally that data-hungry models and data-driven AI can come this far and genuinely have incredible emergent behaviors of thinking machines.

Why start another foundation-model company? Why aren’t LLMs enough? My intellectual journey is not about a company or papers; it’s about finding the North Star problem. It’s not like I woke up and said, “I have to do a company.” I wake up every day, day after day, thinking that there is so much more than language.

Language is an incredibly powerful encoding of thoughts and information, but it’s actually not a powerful encoding of the 3D physical world in which all animals and living things live. If you look at human intelligence, so much is beyond the realm of language. Language is a lossy way to capture the world.

Another subtlety of language is that language is purely generative. Language doesn’t exist in nature. We look around; there’s not a syllable or a word. Whereas the entire physical, perceptual, visual world is there, and animals’ entire evolutionary history is built upon so much perceptual and eventually embodied intelligence.

Humans not only survive, live, and work, but we build civilization by constructing the world and changing the world. That’s the problem I want to tackle.

In order to tackle that problem, research was important, and I spent years doing that as an academic. It’s still fun. But I do realize, especially talking to Martin, that the time has come for a concentrated, industry-grade, focused effort in terms of compute, data, and talent. That is really the answer to bringing this to life. That’s why I wanted to start World Labs.

Martin Casado

Yeah, Erik, you can do a very simple thought experiment that highlights the difference between language and space. If I put you in a room and blindfolded you, and I just described the room and then asked you to do a task, the chances of you being able to do it are very low. I’m like, “10 feet in front of you is a cup; on the left is this.” It’s just a very inaccurate way to convey reality, because reality is so complex and exact.

On the other hand, if I took off the blindfold and you could see the actual space, what your brain is doing is actually reconstructing 3D. Then you can go and manipulate things and touch things.

One way to think about it is that we do a lot of language processing, and we use that to communicate high-level ideas. But when it comes to navigating the actual world, we really rely on the world itself and our ability to reconstruct it.

Erik Torenberg

And how and when did you realize that language might not be enough? Because it seems like it’s not super widely known. I don’t hear about this all the time.

Martin Casado

If you ask me what was surprising, the breakthrough was that language went first, because we’ve worked so hard on robotics. Even looking at autonomous vehicles, as an industry, we’ve invested like $100 billion in it. I remember when Sebastian Thrun actually won the DARPA Grand Challenge in 2006, and we were like, “Hooray, AV is done.” Then, 20 years later, we’re finally there, after $100 billion, et cetera. This is a 2D problem.

That was the path we were going on: Do you actually solve world navigation? It’s hard. Then, out of nowhere, come these LLMs, and they’re unit-economics-positive. They solve all of these language problems basically immediately.

It just took me a moment. Actually, Fei-Fei said it beautifully: The part of our brain that deals with language is pretty recent, and so we’re actually pretty inefficient at it. The fact that a computer does it better is not super surprising.

But the part of the brain that actually does navigation—the spatial part—has been around for millions of years. Maybe the reptilian brain is about four million years old.

It’s even more than that. It’s trial and error. Right, trial and error, right? 500 million years. It’s almost like we’re unrolling evolution, right? The language part is very, very important for high-level concepts and laptop-class work, which is what it’s impacting right now.

But when it comes to space—and this is everything from robotics to anything where you’re trying to construct something physical—you have to solve this problem. We know from autonomous vehicles that it’s a very tough problem, and maybe this is worth talking about: the generative wave gave us some insight into how you might want to do it. So it really felt like that was the time.

Fei-Fei Li

Well, my journey is very different because I’ve always been in vision, right? So I feel like I didn’t need an LLM to convince me that an LWM is important. I do want to say we’re not here bashing language. I’m just so excited. In fact, seeing ChatGPT, LLMs, and these foundation models having such breakthrough success inspires us to realize that the moment is closer for world models.

But Martin said it so beautifully: the 3D space, the space out there, the space in your mind’s eye—the spatial intelligence that enables people to do so many things beyond language—is a critical part of intelligence. It goes from ancient animals all the way to humanity’s most innovative findings, such as the structure of DNA, right? That double helix in 3D space. There’s no way you can use language alone to reason that out.

That’s just one example. Another one of my favorite scientific examples is the buckyball, a carbon-carbon molecule structure that is so beautifully constructed. That kind of example shows how incredibly profound space and the 3D world are.

Erik Torenberg

Let’s paint even more of a picture. When World Labs has achieved its vision, or large world models have achieved their vision, what are some applications or use cases that we can present to the audience to help make it concrete?

Fei-Fei Li

Yeah, there is a lot. For example, creativity is very visual. We have creators from design to movies to architecture to industrial design. Creativity is not just for entertainment; it could be for productivity, machinery, and many things. That alone is a highly visual, perceptual, spatial area of work.

And, of course, we mentioned robotics. Robotics, to me, is any embodied machine. It’s not just humanoids or cars. There’s so much in between. But all of them have to somehow figure out the 3D space they live in, be trained to understand that 3D space, and do things sometimes even collaboratively with humans—and that needs spatial intelligence.

Of course, I think one thing that’s very exciting for me is that, for the entirety of human civilization, we have all collectively lived in one 3D world, and that is the physical Earth—the 3D world. A few of us went to the Moon, but that’s a very small number. That’s one world, but what makes the digital virtual world incredible with this technology—which we should talk about—is the combination of generation and reconstruction.

Suddenly, we can actually create infinite universes. Some are for robots, some are for creativity, some are for socialization, some are for travel, and some are for storytelling. It suddenly enables us to live in a multiverse-like way, and the imagination is boundless.

Martin Casado

These conversations can sound abstract, but they’re actually not. The reason they sound abstract is because it’s truly horizontal, just like LLMs are, right? If you ask what LLMs are good at, the same LLM we use for an emotional conversation, we use to write code. We use it for to-do lists. We use it for self-actualization, right?

With these models, you can take a view of the world, like a 2D view of the world, and then you can actually create a full 3D representation, including what you’re not seeing—like the back of the table, for example—within the computer. Given just a 2D view, you have the full thing, and then you ask, “Okay, well, what can you do with that thing?”

For example, you can manipulate it, move it, measure it, and stack it. So anything that you would do in space, you could do. That means you could do architecture and design. But it turns out the ability to fill in the back of the table means that you can fill in stuff that was never there to begin with, right?

Let’s say that I just had a 2D picture of this. I could create a 360-degree representation of everything. And so now you have something fully generative. What does that mean? That means video games and creativity. It’s a super, super horizontal piece that takes basically a computer with a single view of the world—or maybe multiple views of the world—and creates a full 3D representation that the computer can then act on.

You can see that that’s a very concrete, pivotal thing for everything from robotics to video games to art and design.

Erik Torenberg

It seems like we haven’t fully been appreciating the 3D components until now. Is that fair to say?

Fei-Fei Li

It is fair to say. In fact, I think evolution took a long time. 3D is not an easy problem, but I always come back to the fact that I had a conversation with my 6-year-old years ago about why trees don’t have eyes, right? The fundamental thing is trees don’t move. They don’t need eyes.

The fact that the entire basis of animal life is moving, doing things, and interacting gives life to perception and spatial intelligence. In turn, spatial intelligence is going to reinvent horizontally, as Martin said, so many of the ways of work and life that humans are engaged in.

Erik Torenberg

Does it need to be 3D, or can you just use 2D?

Fei-Fei Li

Physics happens in 3D, and interaction happens in 3D. Navigating behind the back of the table needs to happen in 3D. Composing the world, whether physically or digitally, needs to happen in 3D. So fundamentally, the problem is a 3D problem.

Martin Casado

One way to think about it is that if it’s a human being looking at, say, a 2D video, the human being can reconstruct the 3D in their head, right? But if you need a computer—let’s say I’ve got a robot that has the output of the model—if that’s 2D and then you ask the robot to measure distance or grab something, that information is missing.

You’ve got the XY plane; the Z plane just isn’t there at all, right? And so for many things that are spatial, you need to provide that information to the computer so that you can actually navigate in 3D space. A 2D video is great if it’s a human, because we already can turn it into 3D, but for any computer program, it’ll need to be 3D.

Fei-Fei Li

Actually, I want to tell you a personal story. About 5 years ago, ironically, I lost my stereo vision for a few months because I had a cornea injury. That means I was literally seeing with 1 eye. And like Martin said, my whole life had been trained with stereo vision.

So even if I was seeing with 1 eye, I kind of knew what the 3D world looked like. But it was a fascinating period, as a vision scientist, for me to experiment with what the world is like. One thing that truly drove home—literally—was that I was frightened to drive.

First of all, I couldn’t get on the highway at that speed. I could not. But I was just driving in my own neighborhood, and I realized I didn’t have a good distance measure between my car and the parked cars on a local, small road, even though I had a perfect understanding of how big my car is and almost how big my neighbors’ parked cars are. I had known the roads for years and years, but just driving there, I had to be so slow—almost 10 miles an hour—so that I didn’t scratch the cars.

And that was exactly why we needed stereo vision.

Martin Casado

That’s a great articulation of why 3D is just essential if you’re doing some processing, right? I don’t recommend it, but if you park your car and drive it with 1 eye, you’ll feel it yourself.

Erik Torenberg

With LLMs, a lot of the research was done at the big companies. What’s the state of the research here?

Fei-Fei Li

This is definitely a newer area of research compared to LLMs. It’s not totally fair to say it’s new, because in computer vision, as a field, we have been doing bits and pieces.

For example, one important revolution that happened in 3D computer vision was Neural Radiance Fields, or NeRF. That was done by our co-founder Ben Mildenhall and his colleagues at Berkeley. It was a way to do 3D reconstruction using deep learning that was really taking the world by storm about 4 years ago.

We’ve also got a co-founder, Christoph Lassner, whose pioneering work was part of the reason Gaussian splatting started to become really popular again as a way to represent 3D volumetric data. And, of course, Justin Johnson, who was my former student and is also a co-founder of World Labs, was among the first generation of deep-learning computer-vision students who did so much foundational work in image generation before transformers were out.

We were using GANs to do image generation and then style transfer, which really popularized some of the components or ingredients of what we’re doing here. So things were happening in academia, and things were happening in industry.

At World Labs, we just have the conviction that we’re going to be all in on this one singular, big north-star problem, concentrating the world’s smartest people in computer vision, diffusion models, computer graphics, optimization, AI, and data—all of them coming into this one team to try to make this work and productize it.

Erik Torenberg

I will say, from an outsider standpoint—and I’m not an expert in any of these spaces—it really feels like, to solve this problem, you need experts both in AI, which is the data and the models, including the actual model architecture, and in graphics, which is how you represent these things in memory in a computer and then on the screen.

It takes a very special team to crack this problem, which Fei-Fei has managed to put together.

How Fei-Fei Li Is Rebuilding AI for the Real World | BidClub