[BidClub_]
No Priors · · 36 分钟

No Priors 第117期|对话 Stanford HAI 联席主任、World Labs 创始人 Dr. Fei-Fei Li

Sarah GuoElad GilFei-Fei Li

YouTube
TL;DR
  • Fei-Fei Li 正押注 World Labs,认为空间智能是 AI 缺失的基础层。她将空间智能定义为理解、推理、交互并生成 3D 世界的能力;没有它,“AI 将是不完整的”。World Labs 称其为已知首家攻克 3D 生成基础模型问题的公司,即便生成的是奇幻世界,输出也必须服从自洽的几何与物理规律。

  • 近期最清晰的商业切入口,是 AI 辅助 3D 创作。Li 预计设计师、视觉特效艺术家、游戏开发者、营销人员及其他创作者会像开发者使用 Cursor 和 Windsurf 一样,与模型协作。生成式空间模型也可能解决元宇宙、XR、AR 和 VR 普及的一大约束:除了改善硬件,“我们还在寻找内容创作”。

  • 机器人需要的,不只是扩大视觉能力或模仿学习数据。Li 预计训练将依赖多种数据形态的混合,包括视频、仿真与合成数据、远程操作以及具身数据;她认为仿真被低估了,但也指出许多专家和机器人公司已经在研究。触觉对于操作能力“确实被严重低估”。她也不认同只有人形机器人的未来:任务和能耗经济学应当催生多样形态——水下机器人应当像鱼,而飞机正变得越来越像机器人。

  • 3D 模型的核心挑战始于数据,而产品化仍未解决。不同于 NLP 和 LLM 领域的同行,Li 表示 World Labs 未必拥有丰富的互联网原生素材,而是需要越来越复杂的数据采集、处理、工程和合成能力。产品化是另一项挑战,因为 3D 是交互式体验,而非被动消费:“没有人早上醒来会说,‘我就坐在这里看 3D。’”

  • Li 的职业经历提供了一个具体先例:逆向押注数据,最终催生新的模型类别。2003年前后,她在导师提出100个类别后,建立了一个101类别数据集,随后将同样的信念扩展为 ImageNet:覆盖数千个类别、1500万张标注图片——尽管当时有人告诉她,这可能让她拿不到终身教职。AlexNet 最终取得突破,验证了“那个没人相信的猜想”。

  • Li 给出的研究建议很简单:“无所畏惧”(Be fearless)。她认为有生产力的野心位于“有点妄想、有点疯狂”和“理性地大胆”之间;过度理性会让人找不到足够大的问题,而彻底疯狂则可能让很多事情失控。因此,World Labs 招募来自图形学、视觉、数据、生成式 AI、基础设施、优化、工程和产品等领域的人才,而不是把空间智能当成一个同质化问题。

  • Li 的终局仍是以人为中心的增强,尤其是在社会缺少专业能力与照护资源的领域。医疗行业缺少药物发现、诊断、精准医疗、可及治疗、慢性病支持和更好的老龄化方案,而不是缺少人的价值。她的核心判断是“AI 是帮助人的工具”,同时应保留爱、关系、繁荣与正义这些人类价值,不让机器夺走它们。

摘要 · 为研究而整理的核心内容

1. 空间智能是缺失的基础层

  • Li 创办 World Labs,是因为“我内心想要去构建”,也是因为 3D 世界模型可以解锁创作、导航、仿真以及 AR/VR。人类和动物本来就拥有空间智能;它与进化深度纠缠,因此是一项核心能力,而不是另一种媒体格式。

  • 她对空间智能的定义涵盖在根本上属于 3D 环境的理解、推理、交互和生成。World Labs 是“我们所知的第一家”探索 3D 生成基础模型的公司,但这一判断保留了边界:即便世界是奇幻的,也应当合理可信,因为其几何结构和物理规律必须自洽。

  • 神经科学视角解释了问题为何困难:动物通过眼睛收集光线,在内部重建 3D 世界,再在其中导航、行动和交互。但即使是人类,如果没有训练,也很难闭上眼睛,生成周围环境极其复杂的 3D 模型。

  • Li 希望实现的跃迁,是把这种经过训练的能力“放到你的指尖上”,实现流畅的交互与编辑。她认为,语言在很大程度上已经“解决了”;3D 同样关键且困难,而情感智能则是第三个主要领域,她还不知道该从何处开始解决。

2. 创作与机器人揭示最先形成的经济市场

  • 近期创作领域最接近的软件工程:Cursor 和 Windsurf 已经展示了不同技能水平的人如何与 LLM 协作。Li 预计,设计师、3D 艺术家、视觉特效艺术家、营销人员和游戏开发者也会建立类似关系,因为即使受过训练的专业人士,面对 3D 这一媒介仍然会感到困难。

  • XR 的约束不只是硬件。Li 认为,元宇宙、AR 和 VR 仍然需要内容创作,而这“天然适合”生成式空间模型。她也认为设计天然适合强化学习场景,因为人们可以围绕美感、效率及其他目标进行优化。

  • 谈到机器人时,Li 预计人类将与机器人“共同生活”,但反对把机器人等同于人形机器人。她引用 Stanford 某实验室关于形态智能的论文:智能体的形态可以随着任务优化而改变。她的假设是,任务要求过于多样,少数几种形态不可能始终保持高效:水下机器人应当像鱼,而非人形结构更适合飞行——“我们的飞机正变得越来越像机器人”。

  • Elad 将讨论扩展到宏观和微观尺度的物理仿真,包括材料科学,并指出机器人不仅是算力问题,也同样是系统集成问题。Li 补充说,智能在一个生物体内可能更加分布式,而不是集中于单一神经系统。

  • 训练这些机器需要一座混合式数据金字塔,其中包括仿真和合成数据。Li 认为仿真被低估了,同时指出许多专家和机器人公司已经在研究。更明显的缺口是触觉:对于操作而非单纯导航,触觉信息必须与视觉、感知和空间数据融合,而 Li 认为这一要求“绝对关键”。

3. 稀缺的 3D 数据与尴尬的交付方式定义产品挑战

  • Sarah Guo 的问题击中了核心约束:图像和视频唾手可得,但结构化 3D 世界并不存在。Li “羡慕” NLP 领域的同行,因为 World Labs 需要的是越来越复杂的数据采集、工程、处理和合成能力。

  • 第二个瓶颈是产品形态。人们持续地体验 3D,但语言更容易交到用户手中;3D 要求主动交互,而不是被动观看。Li 留下了一个令人印象深刻的测试:“没有人早上醒来会说,‘我就坐在这里看 3D。’”

  • 她对产品的期待,比抽象描述更能说明其价值:放大进入微观世界,进入发动机内部观察其工作原理,或站在洗碗机内部。如果世界模型能够表达“任何东西”,陌生的尺度和无法抵达的视角就会变得可以探索,而不只是可以描述。

4. ImageNet 证明了非主流数据论如何产生复利

  • 2003年前后,Li 的博士研究聚焦于物体识别,当时数据集规模很小,也没有 scaling law。Li 与 Pietro 讨论过15个还是30个类别,随后她的导师提出100个。Li 已经意识到,足够的数据在数学上对推动模型泛化至关重要。

  • 一本用于她学习英语的词典,意外提供了类别来源:词典随机用花、自行车和狗等物体配图。她选出101个词,比要求多1个;随后,她不得不应对早期 Google Image Search 结果质量低下的问题,最终让母亲通过一个简单界面点击清理图片。

  • ImageNet 将这一信念扩展为覆盖数千个类别的1500万张标注图片。整个过程经历了早期挣扎、有人警告 Li 拿不到终身教职、Amazon Mechanical Turk“前来救场”、AlexNet 获胜,以及 Geoffrey Hinton 后来公开承认这个数据集的定义性作用。

  • 第二次加速改变了 Li 对研究时间的预期。她曾认为,视觉叙事可能耗尽整个职业生涯;但在2013或2014年前后,Andrej Karpathy 和 Justin Johnson 的工作将 LSTM 与 CNN 结合,“彻底打开了”图像描述领域,让她开始思考余下的65或70年该做什么。

5. 无所畏惧是 Li 的研究处方

  • 当被问到 AI 进展是否如今只属于100亿美元级别的训练运行时,Li 给出的唯一建议是:“无所畏惧”(Be fearless)。有价值的位置处于“有点妄想、有点疯狂”和“理性地大胆”之间——过度理性会让人无法识别足够大的问题,而彻底疯狂则可能让很多事情出错。

  • World Labs 通过学科多样性落实这一原则。在 AI 这一标签下,团队汇集了计算机图形学、计算机视觉、数据、生成式 AI、基础设施、优化、工程和产品人才,因为“空间智能这样困难的问题,不是一个同质化的问题”。无所畏惧体现在候选人的经历、提问、创造力以及面对不确定性的从容中。

6. 以人为中心的 AI 是增强,而非替代

  • 她创办 Stanford 人类中心人工智能研究院的经历,塑造了一条清晰边界:机器不应夺走爱、关系、繁荣或正义。Li 希望 AI 与人协作、增强人的能力,同时保留这些人类价值。

  • 医疗领域缺少的是帮助——从药物发现、诊断,到精准医疗、慢性病、心理健康、医疗交付和老龄化——因此她对终局的乐观判断是协作:“AI 是帮助人的工具。”

Sarah Guo

Today's guest is Dr. Fei-Fei Li, a pioneer in computer vision and deep learning. She created ImageNet, the groundbreaking data set that helped spark the deep learning revolution. Fei-Fei is a Stanford professor and the co-director of the Stanford Institute for Human-Centered AI. She's also led AI at Google Cloud, advised international policymakers, and recently co-founded World Labs, a company dedicated to developing spatially intelligent AI. Fei-Fei, thank you for joining us today.

So, you have made extraordinary contributions to science and policy over the past 2 decades. I'll start with the biggest question: Why start a company now?

Fei-Fei Li

Because, in my heart, I want to build. I see this as such a critical, fun, and exciting moment to build some extraordinary technology that everybody can use. I believe so much in spatial intelligence and the kind of 3D world models that can empower so many people, as well as so many use cases. I think it's going to be really exciting, and I can do that with an extraordinarily brilliant group of young technologists.

Sarah Guo

I want to come back to the people you're working with, because I know some of your co-founders and was trying desperately to convince them to start a company a while back. Then they were like, “Oh, no, we have a bigger mission now with Fei.” What is spatial intelligence? Can you define it for a broader audience?

Fei-Fei Li

Spatial intelligence, to me, is the ability to understand, reason about, interact with, and generate 3D worlds. Our world, fundamentally—no matter how we project it—is 3D, because physically it's 3D. Digitally, if there is a true 3D representation, then we can make a lot of things happen more easily, whether it's design and creation, navigation, simulation, or the experience of AR and VR. All of this, to me, is part of spatial intelligence.

Humans have spatial intelligence. It's part of our core intelligent capabilities. Animals have spatial intelligence. The entire journey of evolution is deeply intertwined with the evolution of spatial intelligence. It's so fundamental that without spatial intelligence, AI would be incomplete.

Sarah Guo

How does that translate into what you're doing with your company? Is there anything you can share in terms of what that means relative to what you're building?

Fei-Fei Li

We're tackling one of the hardest problems in AI, which is making world models that are fundamentally 3D. Once you can crack that problem, you can unlock a lot of spatial intelligence problems. We are the first company we know of that is solving this 3D generative foundation model problem.

Sarah Guo

I have many questions, but since you're describing this first as the criticality of 3D to understanding the world, does that imply you feel that the world models that World Labs will create—or others in academia or in companies will create—will someday be realistically accurate, represent the physics and understanding of the world, and allow us to do many more things with them?

Fei-Fei Li

It should be realistically accurate or plausible. You can create a fantastical world, but it should be plausible, because the geometry and the physics of it need to be plausible. That is fundamental to spatial intelligence.

Sarah Guo

Does that imply you have a particular point of view, from a neuroscience perspective, on how fundamental visual intelligence is? I mean, you've always been a leader in computer vision, right? How important is visual intelligence versus, let's say, large language models and textual intelligence?

Fei-Fei Li

I actually do. I think, from a neural and cognitive science point of view, that spatial intelligence is a really hard problem that evolution has to solve for animals. What's really interesting is that I think animals have solved it to an extent, but haven't fully solved it.

What is the problem animals have to solve? Animals have to evolve the capability of collecting light in something we call eyes, mostly. Then, with that collection of light, they have to reconstruct a 3D world in their mind somehow so that they can navigate, do things, and interact.

For humans, we're the most capable animal in terms of manipulation. We can do a lot of things, and all of this is spatial intelligence. To me, that's rooted in our intelligence.

What's interesting is that it's not a fully solved problem even in animals. For example, if I ask you to close your eyes right now and draw or build a 3D model of the environment around you, it's not that easy. We don't have that much capability to generate an extremely complicated 3D model until we get trained.

There are some of us—whether they're architects, designers, or just people with a lot of training and talent—for whom that's a hard thing to do. Imagine doing it at your fingertips much more easily, with much more fluid interactivity and editability. That would just be a whole different world for people—no pun intended.

Sarah Guo

Are there other big areas, like spatial intelligence, that you feel haven't been as developed as they could be from a model perspective, or other missing gaps that you think, in general, as we build this AI future, we should focus on over time? I was wondering, in addition to 3D and world generation, whether there are other big problems like that, because it feels like there are a few big things we've solved over time and other things we're working on. We're sort of solving language.

Fei-Fei Li

I would say language is solved to a huge extent, and 3D, to me, is as critical and difficult as language. So what else isn't solved? The entire space of emotional intelligence is something that I don't even know how to begin to solve. I know a lot of people who haven't solved it. So, when AGI is achieved, I can tell you the training data for that is not going to come from Silicon Valley people.

Sarah Guo

Don't underestimate Silicon Valley.

Elad Gil

I'll put myself in this bucket, but I think we probably need a broader set of people.

Fei-Fei Li

Yeah, no, I agree. But these are the 3 big buckets, to be honest. I don't know. What do you think, Elad and Sarah?

Elad Gil

I think it depends a lot on what you encapsulate in each model. I agree with your framework in terms of those 3, and then certain things like spatial intelligence. I'm assuming it also delves into different types of physics simulation and simulations of the world. Those are big areas that I think a lot of people aren't working on, but that I think are really interesting or important. There's the macro and the micro scale of that.

The microscale eventually becomes materials science and other very different types of things from what you're talking about, where it's more molecular modeling. It also somewhat goes outside the current definition of AI, which I do think will be empowered by it, of course. There's robotics, but robotics is very much a system integration problem as much as a compute problem. Even if you look at animals, it's not just the compute in the brain per se, right?

Fei-Fei Li

Yeah, a lot of these things seem much more distributed in terms of spatial intelligence relative to the specific systems that animals have. In some cases, it's not as centralized as one would think. So it's very interesting to start thinking in terms of those models of more distributed intelligence across an organism versus a central nervous system. I think it's very interesting stuff.

Sarah Guo

You've also done work in the field of robotics and physical intelligence. I think of the data hierarchy for robotics foundation models and actuation this way: People, of course, want to use video because that's what is available to us. There's a big question about simulation and how much you can get from that today. Perhaps people do not see the future quality and physics that are going to be available to us. Then there's close-to-embodied data, like different forms of teleoperation, and then embodied data collection. Is that the hierarchy you have in your mind, or do you think people underestimate simulation and world models for the future?

Fei-Fei Li

First of all, I like that you say I do work in robotics, especially in my lab at Stanford. I have no doubt that humanity will move into an age where we cohabit with robots. The word “robot” is not synonymous with humanoid; robots take all kinds of forms and shapes.

Actually, a few years ago, my lab wrote a really fun paper about morphological intelligence, where the morphology of an agent can change by optimizing for the tasks it's trying to achieve. We should be a little more imaginative than just humanoids.

Having said that, you mentioned this whole data hierarchy. Some people call it data pyramids, data cakes, or whatever. I agree that it's going to be a hybrid of many different forms of data.

I also think simulation is underrated. It's not underrated by a lot of experts and people in the field. If you look at a lot of robotics companies, they are working on simulation and synthetic data.

I also think we have to be aware that, unlike language models or even spatial intelligence foundation models, robotics is a highly multimodal system. What is truly underappreciated, in my opinion, is haptics, especially if we want to do manipulation, not just navigation.

I think haptics data and the ability to really integrate haptics into vision, perception, and spatial data is absolutely critical.

Sarah Guo

One thing that you said that I thought was really interesting is: What are the different morphological forms that a robot may adopt? There are sort of 2 counterarguments people make in terms of the potential future. One argument is that, from a supply chain perspective and managing builds and the scale of manufacturing, you're going to have many fewer form factors. The other argument is that the economic value of specialization is very high, and therefore there'll be thousands and thousands of different form factors as we move to a robot-driven future. Do you have a point of view on where we're likely to land between those 2 viewpoints?

Fei-Fei Li

I think we're going to gradient descent into optimization of productivity and efficiency. My hypothesis is that the requirements of different tasks are so vast that having very few forms, or sticking with 1 form, is energy-inefficient, and a lot of tasks can be done and should be done by much more energy-efficient form factors.

Just an extreme and trivial example: If we put robots underwater, they should not be in the shape of humans. They'd better be in the shape of fish, right? Just think about energy efficiency. And the same with flying. I don't think the human form is right for that; our airplanes are becoming more and more robotic. I do think there's going to be diversity. Robotics is one potential application for the future.

Sarah Guo

You're a scientist first, but you've also been on the Twitter board and been involved in startups. What are the near-term commercial applications that you can imagine for generating 3D worlds?

Fei-Fei Li

I believe creativity is a vastly exciting area where humans can be superpowered by AI and by spatial intelligence. Here I draw an analogy with software engineering. If you look at today's success of LLMs in software engineering, including applications like Cursor and Windsurf and all that, what you see is a lot of collaboration between AI and humans, and that collaboration comes in different levels of skill sets and all that. I think creativity will be similar, whether we're talking about designers, 3D artists, VFX artists, or even marketing talent and game developers.

There's so much need for collaboration in designing and creating 3D space, and this is fundamentally such a hard problem, even for trained, skilled people, that having a collaborator will be extremely fun if we do it right. And so I see creativity as an area that is really exciting.

I also think that a lot of what we're waiting for with the metaverse or XR, AR, and VR is content creation. I understand the hardware itself needs to continue to evolve, but I also think software—we're looking for content creation, and that lends itself so naturally to 3D modeling and 3D or generative spatial models. That's another interesting area to look into.

Sarah Guo

Do you have a strong point of view on whether or not world models are an interesting answer to scalable RL for more generalizable agents?

Fei-Fei Li

I actually do think this is—as I said, AI is not complete without spatial intelligence, because humans interact in 3D worlds, and in the digital world we need all kinds of interaction. Take design as an example. When we're thinking about design, there's so much we're optimizing for in our mind's eye, whether it's beauty or efficiency or optimization or whatever it is, and that lends itself pretty naturally to RL settings.

Sarah Guo

What are the biggest challenges in trying to go down this path of designing and training world models? I imagine one is that you've worked on images and you've worked on video, but we have images and we have video, and we don't have lots of 3D worlds in the format I assume you're building.

Fei-Fei Li

Yeah, data is absolutely a challenge. You're totally right about that. To create world models and 3D foundation models, we require more and more sophisticated data engineering, data acquisition, data processing, and data synthesis. I am envious of my NLP and LLM colleagues, that the data is so abundant on the internet and we don't necessarily have that luxury. So that's definitely 1 challenge.

Another challenge is that 3D—this is kind of ironic, right? Every one of us uses 3D every day, in so many settings. Basically, you open your eyes, and the whole life that you experience is 3D, even when we type on the computer or stare at a screen all the time. Yet it's still not as easy a form factor to deliver into the hands of people compared to language.

Language is just so easy, and 3D is a very active form of engagement; it's not a passive consumption through viewing. Nobody wakes up and says, “I'm just going to sit here and watch 3D.” So that creates challenges for productization and how to do it in the right way.

Sarah Guo

Were you ever a Second Life player or anything?

Fei-Fei Li

I'm not a gamer, but my kids love Minecraft.

Sarah Guo

I was going to ask you if there was a world that you wanted to experience or imagine.

Fei-Fei Li

That's a great question, Sarah. I would love to see worlds I don't see—for example, zooming in and into microscopic worlds, or going inside an engine, knowing how the actual engine works. Of course, I know theoretically how it works, but seeing it with my own eyes, experiencing it—or even, you might laugh at this, I want to be inside a dishwasher and just experience what that is.

All this can be done in a virtual way if we manage to create world models of anything.

Sarah Guo

Okay. I think both Elad and I want to talk a little bit about your past career and maybe some insights for anyone doing research or trying to have an impact within AI. Right before this, I asked Andrej Karpathy what I should ask you, and he said, “Fei is really magic about ambition and thinking about data. You should ask her about her PhD and the creation of that Caltech 101 dataset with Pietro, because it's instructive.” So I have to ask you about that.

Fei-Fei Li

First of all, I have to say it's always really the greatest thing when your student is more well-known and achieving so much more than you can imagine. It makes me so proud—so very proud of Andrej. I was surprised he remembers my PhD work.

Well, gosh, it goes back to 2003-ish, and the world was just barely scratching the surface of the internet, and data was not much of a thing. My PhD work was really trying to get object recognition to work. That's the problem of calling out cats and dogs and microwaves and chairs and all that when you're presented with a picture, and we were beginning to hypothesize that data matters. But we had no idea—there was no scaling law. We had no idea how far data could go.

All we wanted was, if we have a machine-learning algorithm—whether it's a neural network or a Bayes net, which at that time was very popular, or a support vector machine—we need some data to train. And there was no data to train on. As a PhD student, you want to graduate, and Pietro was like, “Well, curate a dataset.” I was thinking, “Yeah, I do need to curate a dataset, because every dataset out there is so tiny. I'm just not convinced.”

Pietro and I were just talking: Is it 15 different things or 30 different things? And then, God forbid, the PhD advisor said the 3-digit number 100. I was like, “That's a lot of work.” But deep in my heart, I knew he was right. From a mathematical point of view, to push the model to generalize, we needed enough data, at least.

I wrote about this process in my book, The Worlds I See. I stumbled upon a dictionary somehow, and it really was for my own English study. I think it was Webster's Dictionary, if I'm not wrong. It just randomly had visual depictions of some words. I don't even know what rule they followed to be in there, to be honest. Some are flowers, some are bicycles, and some are dogs. And I was like, “Okay, this is actually—you can call it a cheat or a tool.” I grabbed 101 of those words.

That really made my PhD advisor chuckle, because he was like, “Ah, yeah, you just want to do 1 more than I asked for, you know, to dare me.” So that's what I did. I still remember I downloaded—or tried to download—images from Google, and Google was so new at that point, and Google Image Search was so terrible compared to today. I had to do so much cleaning. At some point I got so desperate, I just asked my mom to clean the images, because I wrote a little interface on the computer. She doesn't know computers, but at least she knows click, click. So she helped me do some of that.

Sarah Guo

I mean, you've had one of the most storied careers in AI, and to your point, many of your students have similarly gone on to do really great things across the field, across industry, and across the world. What are 2 or 3 moments that you think of when you think back on your career today? Obviously, there's still a lot of career to come, but I'm just curious. I mean, obviously, there's a lot of things that you did in terms of image- and visual-recognition-related systems and all sorts of things, but I'm curious: When you think of the last 20 years, what stands out the most, given everything that you've done?

Fei-Fei Li

Thank you for asking that question. Of course, ImageNet is one of those projects that consists of multiple moments, from the early struggles and being told I would not get tenure, to actually realizing Amazon Mechanical Turk comes to the rescue, to the moment of AlexNet winning. Also, a couple of years ago, I was at an event in Toronto with Geoff Hinton, and he said publicly how that was so defining. He was almost a little bit apologetic that ImageNet was not as recognized as neural networks.

So that journey is very validating. For scientists, the validation is not about recognition or awards. It's that you made a difference—that conjecture that no one believed in, that hypothesis that no one believed in, we were able to make it happen. So that's one thread.

Elad Gil

Just to make sure, for people from the business world who aren't familiar with it, ImageNet was a large-scale dataset with millions of labeled images across thousands of categories—not just 10 and 1, right? Fifteen million labeled images. Thank you, Fei. That led to amazing breakthroughs in deep learning, in particular AlexNet, and lots of progress in the field, driving machine vision forward.

I actually remember, in 2016 or 2017, I used to show a slide that was the history of AI. Back then it was CNNs and RNNs, and GANs were just getting going. I had ImageNet and AlexNet as one of the seminal moments in this very small number of events that really defined AI progress. Obviously, now we have transformers as part of that, and maybe diffusion models or something, but it was such a big breakthrough.

Fei-Fei Li

Yeah, thank you. Another moment I'm very proud of was Andrej Karpathy and Justin Johnson and their dissertations. It was, in my opinion, the first time that language and images converged through captioning and writing stories of the visual world.

It was significant for me for 2 reasons. I literally thought—I kid you not—that at the end of my Ph.D., if I could live to 100 years old, that was the problem we might be able to solve: storytelling of pictures. So I entered my career, in my first year as an assistant professor, thinking, “Okay, I'm going to do ImageNet to solve object recognition, and then I'm going to spend the rest of my entire career solving this problem of storytelling.”

By the time Andrej and, a little later, Justin Johnson entered my lab, around 2013 or 2014, at the beginning of deep learning, suddenly the combination of a sequential model—at that point, it was an LSTM, not a transformer model—and a CNN had just blasted open image-captioning work. My work, together with Google's, was the first out of the door, and that was really, to me, almost—I was so proud. I almost had a crisis: “What am I going to do for the rest of my 70 years, or 65 years?”

That was really exciting, how fast the field has evolved.

Elad Gil

Can I ask you one more question about this? You have made this amazing progress very efficiently, right? You and I have talked offline before about how you feel it's really important for there to be moonshots and creativity in AI research beyond very large, well-funded corporate labs, let's say. You pointed to several moments that came from creativity and research in academia. What advice do you have for people about whether there's still opportunity for that, or whether it's all just 10-billion-dollar training runs from here?

Fei-Fei Li

My singular advice—and I still say that in my company and my lab—is: be fearless. I think scientists, technologists, and entrepreneurs have to be fearless. Eventually, you have to figure out: do you need 10-billion-dollar runs, or do you come to Sarah to ask for funding?

Sarah Guo

Probably a lot for both. Yeah, yeah.

Fei-Fei Li

Or you have to figure out data. Sometimes fearlessness is this very interesting position where you're somewhat delusional and crazy, but somewhat just rationally bold, and it's in between. If you're too rational, it's not courageous enough. You're not identifying problems that are big enough. But if you're completely crazy, then—I don't know—there are so many things that can go wrong.

So be fearless. Be courageous. To me, that is really important. Even as old as I am, that's how I feel I started my startup, World Labs. I want to be fearless and solve this problem of spatial intelligence.

Sarah Guo

As part of problem-solving, you've worked with some of the best AI researchers in the world over time and the best engineers. How do you think about that in the context of your company? What sorts of people are you trying to hire? Are there open roles currently? It's an amazing team, and I'm just curious what sorts of folks you want to add and how you're thinking about that over time.

Fei-Fei Li

Yes, we have open roles, and we would love to hire the best engineers as well as product thinkers at this point for our company. So if you're an engineer, AI researcher, or product talent out there who's passionate about joining the most talented team and solving this problem, please join us.

Who do we hire? First of all, we really do hire for diversity of thinking. You call us an AI company, but if you look under the hood, we've got computer graphics experts, computer vision experts, data experts, generative AI experts, machine-learning infrastructure experts, and optimization experts.

It's really important to hire a diverse group of really talented people, because a problem as hard as spatial intelligence is not a homogeneous problem. It takes talent from all kinds of backgrounds to solve it.

I also look for fearlessness. How do you do that? How do you identify whether somebody has fearlessness in their background or in their thinking process? It's in their background. You talk to them; you can sense when someone is fearless. You can sense what drives them. You can sense the questions they ask.

If they start asking you a lot of things about, “I don't know how to get this done”—of course, you have to ask those questions because you want to get it done—but if you sense that it comes from the point of view of being scared of solving it, then that's not fearlessness. Those fearless people are creative and ambitious, and they're not afraid of uncertainty or the unknown. I really love that.

Sarah Guo

Elad and I try to make a business of doing business with fearless people, and hopefully those who are technically creative.

One last broader question for you: An important part of your work has also been thinking about how to bring more people into AI, including co-directing the Stanford Institute for Human-Centered Artificial Intelligence. If you picture the world several years out from your last set of predictions—not to use a pun on the book—what's your most optimistic view of what human-centered AI looks like?

Fei-Fei Li

Thanks for asking. In fact, that's another point of my career that I feel very proud of: the founding of the Stanford Institute for Human-Centered Artificial Intelligence, HAI, and the continued movement toward that way of thinking.

I want to build a world where AI collaborates with and superpowers people. I still believe our world, our human world, needs to be human-centered, where love, relationships, and prosperity across all communities are really important. Justice and all these other values are really important, and I don't think any piece of machinery—whether it's AI, an airplane, or biotech—should take those away.

With those critical values in mind, having AI superpower us is really important, because there are so many unsolved problems. One application area I've worked on is healthcare, for example, at Stanford. If you look at healthcare, from drug discovery and curing diseases to diagnosis that can reach all people in the world, to treatment that can be accessible to all people in the world, to the whole healthcare delivery system—how to make aging better, how to take care of chronic diseases, how to deal with mental health—all of this, we do not have an issue of excess humans or anything. We're lacking help.

We're lacking scientific discovery. We're lacking diagnosis. We're lacking precision medicine. We're lacking safer and more effective ways of healthcare delivery, aging health, and all that. That's what I believe. I think AI is a tool to help people.

Sarah Guo

Elad and I are collectively invested in a series of companies that I hope will be useful here, from Abridge to OpenEvidence to Elation. But as you said, there's a huge spectrum of problems, and honestly, I've been less optimistic about the adoption of technology in healthcare for the last 15 years. It does feel like this time it's different, and it's just massively net good here.

Elad Gil

Yeah, I actually started a digital health company before this, and my hope is that finally a lot of the things that people have been talking about for decades will come to fruition. It seems like AI is a great delivery mechanism for that. So, totally, totally.

Well, thank you so much, Fei. It was fantastic. This has been inspiring, and it's great to hear a little bit more about World Labs as well.

Fei-Fei Li

Thank you. Thank you a lot. Thank you, Sarah.