Fei-Fei Li 如何重建面向真实世界的 AI
World Labs 的核心押注是:语言是强大的思想编码方式,却是有损且不足以表达物理现实的编码方式,因此空间智能仍是基础模型的前沿。 Fei-Fei Li 称语言是“捕捉世界的一种有损方式”;世界模型必须理解3D结构、形状和组合性,让机器能够行动,而不只是描述。
Martin Casado 认为,真正令人意外的是发展顺序:语言几乎立刻实现了“单位经济为正”,而自主导航在20年间吞噬了约1000亿美元。 LLM 并没有解决空间问题,但生成式 AI 的成功为攻克这一更早、也更困难的底层问题提供了线索。
这套平台试图从一个或多个2D视角重建完整的3D场景,包括摄像头看不见的几何结构。 一旦模型能够补全“桌子的背面”,软件就能测量、移动、堆叠和操控物体,从而服务于建筑、设计、机器人、游戏及其他横向市场。
深度信息是操作数据,而非视觉装饰,因为物理规律和人机交互都发生在3D空间。 在短暂失去立体视觉后,Li 无法高速驶上高速公路;在本地道路上,她只能以接近每小时10英里的速度行驶,因为无法可靠判断自己的车与路边停放车辆之间的距离。
World Labs 将场景重建与生成结合起来,打开具身机器和合成环境的应用空间。 Li 更宏大的判断是,AI 可以创造“无限宇宙”(“infinite universes”)——用于机器人、创意、社交、旅行和叙事——并“让我们生活在多元宇宙中”。
执行层面的核心,在于将 AI、计算机视觉、图形学、优化、数据和成熟的3D技术集中到一支团队中。 技术组成包括 NeRF、Gaussian splatting、GAN 时代的图像生成和风格迁移;公司层面的突破,则是把“算力、数据、人才”聚集起来,将一个北极星问题产品化。
1. World Labs 始于一个共同判断:AI 缺少世界模型
Li 寻找的不只是融资:她需要的“独角兽投资人”能够穿越深科技行业的周期起伏,同时在计算机科学、产品市场匹配和市场进入策略上成为思想伙伴。
在一场被 LLM 热情主导的 AI 晚宴上,Li 转向 Casado,问道:“你知道我们缺什么吗?”她的答案——“我们缺一个世界模型”——与 Casado 通过图像领域投资独立得出的结论不谋而合。
Li 后来进一步验证这种共识是否足够深入。Casado 将世界模型定义为:AI 真正理解现实世界的3D结构、形状和组合性;这不同于她与多数人交谈时得到的“礼貌性点头”。
2. 语言的快速胜利暴露了尚未解决的空间层
Li 坦言,数据饥渴型模型竟然展现出如此强大的涌现能力,仍让她感到意外——毕竟正是她本人推动了数据在 AI 中的重要地位。但语言依然只是“捕捉世界的一种有损方式”。
她的区分是结构性的:语言完全是生成出来的,天然界并不存在语言;而物理世界和感知世界本来就在那里。动物智能通过感知和具身交互发展起来;人类运用这种智能,不只是为了生存,也为了建造和改变世界。
Li 将 World Labs 定义为一个北极星问题,而不是一家公司或论文项目。经过多年的学术研究,她得出结论:必须以产业级方式集中算力、数据和人才,才能把这个问题真正做出来。
Casado 原本认为,行业应先解决世界导航问题。产业界在自动驾驶上投入了约1000亿美元;2006年的 DARPA Grand Challenge 似乎宣告问题已经解决,但真正实现大规模落地仍花了约20年。LLM 则“凭空出现”,几乎立刻解决了大量语言任务。
Li 的限定很重要:“我们不是来抨击语言的。”ChatGPT 和基础模型的突破反而让她相信,世界模型的时机已经更近:“我不需要 LLM 来说服我,LWM 很重要。”
3. 机器无法利用从未接收到的深度信息
Casado 用蒙眼类比把差距说得很具体:仅凭对房间的语言描述,一个人完成物理任务的概率极低;而视觉能让大脑重建空间,并操控其中的物体。
Li 从动物演化的角度追溯空间智能:树木不需要眼睛,因为它们不会移动。感知之所以变得关键,是因为动物需要导航和交互;后来,人类的空间推理能力又促成了 DNA 3D 双螺旋结构、巴克球碳结构等发现。
Torenberg 追问,为什么不能停留在2D?答案是绝对的:“物理发生在3D中,交互也发生在3D中。”人类可以从视频中推断深度,但只接收2D输出的机器人缺少判断距离或抓取物体所需的 Z 轴信息。
角膜受伤、立体视觉暂时丧失后,Li 亲身经历了这种失败。即使在熟悉的社区道路上,她也把车速降到接近每小时10英里,因为物体的已知尺寸无法替代可靠的距离测量。
4. 重建与生成让机会覆盖横向市场
Li 认为,最早的需求将来自视觉创意领域——电影、建筑、工业设计和机械——以及具身机器;后者远不止人形机器人和汽车。所有这类机器都必须理解所处的3D环境,有时还要与其中的人协作。
Casado 描述的具体产品原语,是从单个2D视角出发,重建完整的3D场景,包括桌子背面等不可见表面。随后,这套表示可以被操控、移动、测量和堆叠;同一系统也能补全原本不存在的部分,生成360度表示。
Li 想象,机器人、创意、社交、旅行和叙事都可以拥有“无限宇宙”。Casado 则强调,这个机会具有横向覆盖能力,类似于一个 LLM 同时服务于情感对话、代码、待办事项和自我实现。
5. 技术栈将3D研究与产业化集中能力结合起来
Li 称,世界模型研究比 LLM 更新,但并非从零开始。计算机视觉已经提供了“零件”,其中包括 Ben Mildenhall 的 NeRF 工作;大约4年前,这项基于深度学习的3D重建技术开始产生广泛影响。
她还提到 Christoph Lassner 的先驱工作,以及 Gaussian splatting 作为3D体积数据表示方式重新走红的背景;在 transformer 出现之前,Justin Johnson 也完成了奠基性的图像生成研究。Li 表示,基于 GAN 的图像生成和风格迁移,帮助普及了当前这项工作的部分技术要素。
World Labs 的策略,是让计算机视觉、扩散模型、图形学、优化、AI 和数据领域的专家围绕一个北极星问题集中协作。Casado 的外部判断是,解决这一问题既需要模型智能,也需要图形学能力:数据和模型必须与一种可用的计算机表示结合起来。
Space—the 3D space, the space out there, the space in your mind’s eye—spatial intelligence is a critical part of intelligence. Suddenly, we can actually create infinite universes. Some are for robots, some for creativity, some for socialization, some are for travel, and some are for storytelling. It suddenly will enable us to live in the multiverse.
Martin, why don’t you briefly brag on behalf of Fei-Fei a little bit and share how you would summarize her contributions to AI for people who are unfamiliar?
She’s someone who doesn’t need a lot of introduction, and she’s done so many things that I can’t fit them all in. So maybe I’ll just do the ones that are appropriate to this.
Of course, she was on the Twitter board. She was a Google executive, founder and CEO of World Labs. But very importantly, as we all know, AI—we all talk about neural networks, and there are a number of people who focused on making those effective—but Fei-Fei really singularly brought data into the equation, which now we’re recognizing is actually probably the bigger problem, the more interesting one. And so she truly is the godmother of AI, as everybody calls her.
And, Fei-Fei, why did you have to have Martin as the first investor?
Well, first of all, I knew Martin for more than a decade. I joined Stanford in 2009 as a young assistant professor, and Martin was finishing his PhD there. Martin’s adviser, Nick Matune, was a good friend, and I knew Martin would go on to become a very successful entrepreneur and a very successful investor.
We would see each other and talk about things. But as I was formulating the idea of World Labs, I was looking for what I would call my unicorn investor. I don’t know if that’s a word, but that’s how I think about it: someone who is not only an established and successful investor who can be with entrepreneurs on this journey through the ups and downs, who can be very insightful, and who can bring the kind of knowledge, advice, and resources, but I was also particularly looking for an intellectual partner.
What we are doing at World Labs is very deep tech. We are trying to do something no one else has done. We know with a lot of conviction that it will change the world, literally. But I need someone who is a computer scientist, who is a student of AI, who understands product-market fit, go-to-market, and consumers, and who can be on the phone or in person with me every moment of the day as an intellectual partner.
The origin story of us first connecting is actually pretty interesting. Fei-Fei has clearly been thinking about this idea for a very long time, well before it started—maybe for years, even—and she’ll talk about it. She has this very deep intuition of what AI needs in order to navigate the world.
But we were at one of Mark’s fancy dinners or lunches, and there were a bunch of AI people. Everybody was so excited about LLMs, and they were talking about language. I’d come to this independent conclusion, just because I’d actually done a lot of image investing, that that wasn’t the end of the story.
Fei-Fei, at the end of this table with all these people talking about it, leans over to me. She’s like, “You know what we’re missing?” And I said, “What are we missing?” She said, “We’re missing a world model.” And I’m like, “Yes.”
It kind of fell into place then, because I’d been thinking about this at a high level. She just perfectly articulated it, as she does, and she had a year’s worth of thinking about this and had talked to people. In some way, we had arrived at a very similar intuition through our own crooked paths. Hers was way more filled out. Mine was just kind of this fancy thing.
After that, we had a number of conversations where we both agreed that we were aligned on this idea.
Actually, I don’t know if you know this. During that lunch, we hit it off on this world-model idea, but I was at that point already talking to various people—not just computer scientists and technologists, but also investors and potentially business partners.
To be honest, most people didn’t get it. When I said “world model,” they nodded, but I could just tell that was a polite nod. So I called Martin. I said, “Do you mind coming over to the Stanford campus and having coffee with me?”
I said, “Martin, can you define your world model to me?” I really wanted to hear if Martin actually meant it. The way he defined it—as an AI model that truly understands the 3D structure, shape, and compositionality of the world—was exactly what I was talking about. I was like, “Wow, he’s the only person so far I’ve talked to who actually meant it.” It wasn’t just nodding.
Okay, so we’re going to get to World Labs and the specifics of this, but maybe first let’s take you both back to your PhD days and your professor days and reflect on this: If you could go back in time with knowledge of what’s happened in the preceding 10 years in AI, what do you think would have been the biggest surprise? What was the thing you didn’t see coming that would have shocked your younger self?
Yeah. It’s ironic to say because, as Martin said, I was the person who brought data into the AI world, but I still continue to be so surprised emotionally that data-hungry models and data-driven AI can come this far and genuinely have incredible emergent behaviors of thinking machines.
Why start another foundation-model company? Why aren’t LLMs enough? My intellectual journey is not about a company or papers; it’s about finding the North Star problem. It’s not like I woke up and said, “I have to do a company.” I wake up every day, day after day, thinking that there is so much more than language.
Language is an incredibly powerful encoding of thoughts and information, but it’s actually not a powerful encoding of the 3D physical world in which all animals and living things live. If you look at human intelligence, so much is beyond the realm of language. Language is a lossy way to capture the world.
Another subtlety of language is that language is purely generative. Language doesn’t exist in nature. We look around; there’s not a syllable or a word. Whereas the entire physical, perceptual, visual world is there, and animals’ entire evolutionary history is built upon so much perceptual and eventually embodied intelligence.
Humans not only survive, live, and work, but we build civilization by constructing the world and changing the world. That’s the problem I want to tackle.
In order to tackle that problem, research was important, and I spent years doing that as an academic. It’s still fun. But I do realize, especially talking to Martin, that the time has come for a concentrated, industry-grade, focused effort in terms of compute, data, and talent. That is really the answer to bringing this to life. That’s why I wanted to start World Labs.
Yeah, Erik, you can do a very simple thought experiment that highlights the difference between language and space. If I put you in a room and blindfolded you, and I just described the room and then asked you to do a task, the chances of you being able to do it are very low. I’m like, “10 feet in front of you is a cup; on the left is this.” It’s just a very inaccurate way to convey reality, because reality is so complex and exact.
On the other hand, if I took off the blindfold and you could see the actual space, what your brain is doing is actually reconstructing 3D. Then you can go and manipulate things and touch things.
One way to think about it is that we do a lot of language processing, and we use that to communicate high-level ideas. But when it comes to navigating the actual world, we really rely on the world itself and our ability to reconstruct it.
And how and when did you realize that language might not be enough? Because it seems like it’s not super widely known. I don’t hear about this all the time.
If you ask me what was surprising, the breakthrough was that language went first, because we’ve worked so hard on robotics. Even looking at autonomous vehicles, as an industry, we’ve invested like $100 billion in it. I remember when Sebastian Thrun actually won the DARPA Grand Challenge in 2006, and we were like, “Hooray, AV is done.” Then, 20 years later, we’re finally there, after $100 billion, et cetera. This is a 2D problem.
That was the path we were going on: Do you actually solve world navigation? It’s hard. Then, out of nowhere, come these LLMs, and they’re unit-economics-positive. They solve all of these language problems basically immediately.
It just took me a moment. Actually, Fei-Fei said it beautifully: The part of our brain that deals with language is pretty recent, and so we’re actually pretty inefficient at it. The fact that a computer does it better is not super surprising.
But the part of the brain that actually does navigation—the spatial part—has been around for millions of years. Maybe the reptilian brain is about four million years old.
It’s even more than that. It’s trial and error. Right, trial and error, right? 500 million years. It’s almost like we’re unrolling evolution, right? The language part is very, very important for high-level concepts and laptop-class work, which is what it’s impacting right now.
But when it comes to space—and this is everything from robotics to anything where you’re trying to construct something physical—you have to solve this problem. We know from autonomous vehicles that it’s a very tough problem, and maybe this is worth talking about: the generative wave gave us some insight into how you might want to do it. So it really felt like that was the time.
Well, my journey is very different because I’ve always been in vision, right? So I feel like I didn’t need an LLM to convince me that an LWM is important. I do want to say we’re not here bashing language. I’m just so excited. In fact, seeing ChatGPT, LLMs, and these foundation models having such breakthrough success inspires us to realize that the moment is closer for world models.
But Martin said it so beautifully: the 3D space, the space out there, the space in your mind’s eye—the spatial intelligence that enables people to do so many things beyond language—is a critical part of intelligence. It goes from ancient animals all the way to humanity’s most innovative findings, such as the structure of DNA, right? That double helix in 3D space. There’s no way you can use language alone to reason that out.
That’s just one example. Another one of my favorite scientific examples is the buckyball, a carbon-carbon molecule structure that is so beautifully constructed. That kind of example shows how incredibly profound space and the 3D world are.
Let’s paint even more of a picture. When World Labs has achieved its vision, or large world models have achieved their vision, what are some applications or use cases that we can present to the audience to help make it concrete?
Yeah, there is a lot. For example, creativity is very visual. We have creators from design to movies to architecture to industrial design. Creativity is not just for entertainment; it could be for productivity, machinery, and many things. That alone is a highly visual, perceptual, spatial area of work.
And, of course, we mentioned robotics. Robotics, to me, is any embodied machine. It’s not just humanoids or cars. There’s so much in between. But all of them have to somehow figure out the 3D space they live in, be trained to understand that 3D space, and do things sometimes even collaboratively with humans—and that needs spatial intelligence.
Of course, I think one thing that’s very exciting for me is that, for the entirety of human civilization, we have all collectively lived in one 3D world, and that is the physical Earth—the 3D world. A few of us went to the Moon, but that’s a very small number. That’s one world, but what makes the digital virtual world incredible with this technology—which we should talk about—is the combination of generation and reconstruction.
Suddenly, we can actually create infinite universes. Some are for robots, some are for creativity, some are for socialization, some are for travel, and some are for storytelling. It suddenly enables us to live in a multiverse-like way, and the imagination is boundless.
These conversations can sound abstract, but they’re actually not. The reason they sound abstract is because it’s truly horizontal, just like LLMs are, right? If you ask what LLMs are good at, the same LLM we use for an emotional conversation, we use to write code. We use it for to-do lists. We use it for self-actualization, right?
With these models, you can take a view of the world, like a 2D view of the world, and then you can actually create a full 3D representation, including what you’re not seeing—like the back of the table, for example—within the computer. Given just a 2D view, you have the full thing, and then you ask, “Okay, well, what can you do with that thing?”
For example, you can manipulate it, move it, measure it, and stack it. So anything that you would do in space, you could do. That means you could do architecture and design. But it turns out the ability to fill in the back of the table means that you can fill in stuff that was never there to begin with, right?
Let’s say that I just had a 2D picture of this. I could create a 360-degree representation of everything. And so now you have something fully generative. What does that mean? That means video games and creativity. It’s a super, super horizontal piece that takes basically a computer with a single view of the world—or maybe multiple views of the world—and creates a full 3D representation that the computer can then act on.
You can see that that’s a very concrete, pivotal thing for everything from robotics to video games to art and design.
It seems like we haven’t fully been appreciating the 3D components until now. Is that fair to say?
It is fair to say. In fact, I think evolution took a long time. 3D is not an easy problem, but I always come back to the fact that I had a conversation with my 6-year-old years ago about why trees don’t have eyes, right? The fundamental thing is trees don’t move. They don’t need eyes.
The fact that the entire basis of animal life is moving, doing things, and interacting gives life to perception and spatial intelligence. In turn, spatial intelligence is going to reinvent horizontally, as Martin said, so many of the ways of work and life that humans are engaged in.
Does it need to be 3D, or can you just use 2D?
Physics happens in 3D, and interaction happens in 3D. Navigating behind the back of the table needs to happen in 3D. Composing the world, whether physically or digitally, needs to happen in 3D. So fundamentally, the problem is a 3D problem.
One way to think about it is that if it’s a human being looking at, say, a 2D video, the human being can reconstruct the 3D in their head, right? But if you need a computer—let’s say I’ve got a robot that has the output of the model—if that’s 2D and then you ask the robot to measure distance or grab something, that information is missing.
You’ve got the XY plane; the Z plane just isn’t there at all, right? And so for many things that are spatial, you need to provide that information to the computer so that you can actually navigate in 3D space. A 2D video is great if it’s a human, because we already can turn it into 3D, but for any computer program, it’ll need to be 3D.
Actually, I want to tell you a personal story. About 5 years ago, ironically, I lost my stereo vision for a few months because I had a cornea injury. That means I was literally seeing with 1 eye. And like Martin said, my whole life had been trained with stereo vision.
So even if I was seeing with 1 eye, I kind of knew what the 3D world looked like. But it was a fascinating period, as a vision scientist, for me to experiment with what the world is like. One thing that truly drove home—literally—was that I was frightened to drive.
First of all, I couldn’t get on the highway at that speed. I could not. But I was just driving in my own neighborhood, and I realized I didn’t have a good distance measure between my car and the parked cars on a local, small road, even though I had a perfect understanding of how big my car is and almost how big my neighbors’ parked cars are. I had known the roads for years and years, but just driving there, I had to be so slow—almost 10 miles an hour—so that I didn’t scratch the cars.
And that was exactly why we needed stereo vision.
That’s a great articulation of why 3D is just essential if you’re doing some processing, right? I don’t recommend it, but if you park your car and drive it with 1 eye, you’ll feel it yourself.
With LLMs, a lot of the research was done at the big companies. What’s the state of the research here?
This is definitely a newer area of research compared to LLMs. It’s not totally fair to say it’s new, because in computer vision, as a field, we have been doing bits and pieces.
For example, one important revolution that happened in 3D computer vision was Neural Radiance Fields, or NeRF. That was done by our co-founder Ben Mildenhall and his colleagues at Berkeley. It was a way to do 3D reconstruction using deep learning that was really taking the world by storm about 4 years ago.
We’ve also got a co-founder, Christoph Lassner, whose pioneering work was part of the reason Gaussian splatting started to become really popular again as a way to represent 3D volumetric data. And, of course, Justin Johnson, who was my former student and is also a co-founder of World Labs, was among the first generation of deep-learning computer-vision students who did so much foundational work in image generation before transformers were out.
We were using GANs to do image generation and then style transfer, which really popularized some of the components or ingredients of what we’re doing here. So things were happening in academia, and things were happening in industry.
At World Labs, we just have the conviction that we’re going to be all in on this one singular, big north-star problem, concentrating the world’s smartest people in computer vision, diffusion models, computer graphics, optimization, AI, and data—all of them coming into this one team to try to make this work and productize it.
I will say, from an outsider standpoint—and I’m not an expert in any of these spaces—it really feels like, to solve this problem, you need experts both in AI, which is the data and the models, including the actual model architecture, and in graphics, which is how you represent these things in memory in a computer and then on the screen.
It takes a very special team to crack this problem, which Fei-Fei has managed to put together.