[BidClub_]
Invest Like the Best · · 67 分钟

AI、LLM 与机器人智能领域的全球顶尖研究者

Patrick O'ShaughnessySergey Levine

YouTube
TL;DR
  • Physical Intelligence 的核心押注,是打造能够“基本控制任何具身系统完成任何任务”的机器人基础模型。原因在于,追求完整通用性“从长期看可能反而更容易”,就像 LLM 借助弱标注网络数据,先建立对世界的理解,最终击败为机器翻译等狭窄应用量身定制的系统。Sergey Levine 的结论是:“问题只有一个,不是很多个”(“there is one problem, not many different problems”)——不是人形机器人问题、汽车问题和推土机问题各自独立存在。
  • 最值得交易的研究更新是:瓶颈已经上移到更高的软件栈。大约6个月前,PI 发现,只要加入带有语义指令标注的数据,模型就会提升——不需要新增遥操作动作——这意味着物理灵巧性不再是主要约束,场景理解才是;而且“人可以直接和机器人说话”,更好地指导它。中层推理表征是 PI 下一步最明显的问题。
  • 不要先算数据集规模,而要先把飞轮搭起来。“我不认为有人真正知道需要多少机器人数据……我们其实不需要知道”;目标是把系统做到足够有用,能够部署并自行收集数据——“Tesla 不会担心自己的汽车能收集多少数据。真要说的话,顺序恰恰相反。”与此同时,硬件成本已从40万美元的 PR2,降到约3万美元的实验室机器人,再降到“可能只有这个价格十分之一”的机械臂。
  • 对人形机器人的怀疑,仍然是温和保留。人形机器人只是“众多可能的机器人类型之一”;行业尚未解决的分歧在于,人形机器人的杂技式动作依赖重度仿真,且“实际上往往完全没有真实世界数据”,而操作任务则依赖大规模真实世界数据和大型基础模型——究竟哪条路线胜出,Levine 表示“我不知道答案”。
  • 在时间表上,Levine 相比资深机器人研究者更乐观,但相较机器人创业者更悲观。最后被攻克的任务将是养老照护,而“给孩子换尿布会非常非常难”——这是“Moravec 悖论的顶峰”。如果2050年仍没有厨房机器人,最可能的罪魁祸首是社会—技术长尾;从纯技术角度看,关键风险在于开放世界的广度。
  • 对准备行动的运营者来说,关键不确定性在于示范数据与自主强化学习之间的权重——“究竟是90/10还是10/90?”两种答案会把正确的系统配置完全翻转;而劳动力范式是编码工具,而不是替代:人与机器人之间会形成一种共同演化的生产力“共舞”。
摘要 · 为研究而整理的核心内容

1. 赌注:一个基础模型,适配任何身体、完成任何任务

  • Levine 对这项使命的定义是:打造能够“基本控制任何具身系统完成任何任务”的机器人基础模型;公司的核心判断是,通用性比狭窄专用性更容易实现,“就像语言模型后来证明,在某些方面,完整解决自然语言任务,比单独瞄准机器翻译或情感分析更容易”。
  • 这套类比背后的机制是:LLM 借助弱标注数据取胜——从网络挖掘的文本建立起“对世界理解的基础”,在此之上开发不同应用会更有效。机器人没有互联网规模的数据集,但这套逻辑反而更有力:与其训练“洗碗专家”或“叠衣服专家”,不如训练一个真正“理解物理交互”的模型;人之所以能快速掌握新技能,正是因为拥有这样的底层基础。
  • Patrick 解释这件事为何重要时,先披露自己是 PI 的投资者:机器人行业存在“稻草人问题”——身体形态越来越酷,“但它们真正需要的东西都是智能,也就是一颗大脑”。

2. 泛化能力会制造糟糕演示——要据此判断行业

  • 这是 Levine 整个职业生涯的烦恼:“有效泛化其实并不是做出真正精彩演示的最优方式。”演示的套路是挑一项酷任务、控制环境、把场景做到一尘不染。泛化看起来平淡无奇——PI 4月发布的视频里,机器人清理的是一个“完全没有训练数据”的厨房,脱离上下文看,“就是,嗯,它在捡盘子”。公开演示只是“能力边界上的展示”;真正的信号在论文里,或者在和单个研究者交流时听到“内部故事”。
  • Patrick 质疑 Boston Dynamics:几十年来一直很酷,却没有给客户带来什么有用的东西——“公平地说,这也是对很多机器人公司的合理问题。”Levine 为正确使用演示、服务于使命辩护:“只要诚实地设定挑战就行。”PI 自己的规则是:“在有用这一约束下,让它尽可能酷。”浓缩咖啡和洗衣服不是目标,而是压力测试。

3. 赌注所在:机器人的个人电脑时刻

  • 成功意味着释放想象力,而不是制造金属人:就像个人电脑在90年代触发了“令人惊叹的应用程序寒武纪大爆发”,一个可提示、可微调的机器人基础模型,能够让任何人开发机器人应用,而不必自己解决智能问题,也不必搭建一套“庞大得吓人的技术栈”。
  • “我们有时会觉得机器人会是一种东西……金属人。但我不认为事情会这样发展,因为没有任何技术是那样的。”他的工具箱愿景是:5条机械臂,其中一条从天花板垂下;或者“1,000架四旋翼无人机组成的蜂群”建造一栋房子。长期看,手术机器人也不应受限于人类能够实时遥操作的范围。
  • 目前,AI 挑战扼杀了形态创新——每一种新身体都意味着系统辨识和动力学特性刻画。“如果你能在车库里组装一个机器人,加载一个机器人基础模型,然后告诉它做一堆事情……那可能会成为一台非常强大的发动机。”因此,PI 已经开源其模型,让更多人动手实验。

4. 多模态 LLM 提供了一条路径

  • 历史可以浓缩成3个时间点:端到端控制并不新鲜——Alvin(很可能是 ALVINN)在“1986或87年”用一个很小的神经网络驶上高速公路;2010年代初的深度 RL 打开了“超越人类水平表现”的道路;最近的进展则是多模态 LLM,它们“知道很多东西”,为处理长尾场景中的常识问题提供了路径,尽管“它们还不擅长把这些知识落到物理情境中”。
  • 他对常识的工作定义是:“肌肉记忆的反面”——把在其他地方获得的知识(读到、看到或听到的)应用到当前情境并完成落地,比如围绕一个从未见过的漏气标志临场发挥。
  • Moravec 悖论支配着能力地图:我们以为捡杯子很容易,是因为对人来说很容易——“那些不太擅长发现老虎的人,都被老虎吃掉了。”机器学习改变了这个方程:只要数据容易收集,即便物理上很复杂,也会落入容易解决的类别;需要多层推理、数据又难以收集的领域,则仍然困难。
  • 机器人“大脑”中更棘手的部分是物理类比。“那家公司势头很足”——你完全知道它是什么意思,但“这个词背后承载了很多推论”。Feynman 关于“自旋”的类比“确实能引导粒子物理中的推断”。“我不知道 LLM 能不能做到这一点。”相关问题还有组合泛化:有学生让 LLM 用国际音标写三明治食谱,而这种格式只会在词典里逐词出现,LLM 却写出了整段文字;机器人同样应该能够重新组合已学技能,处理新问题。

5. 方法论,以及孕育它的研究文化

  • Levine 的路径是:先做计算机图形学,随后在2014年跟随 Pieter Abbeel 做博士后,当时“完全没有机器人经验”,只追逐一个想法——让系统“做得越多,变得越好”。从零开始练技能只在有限环境中有效;Google 的“机械臂农场”(2015年,由二十多台机器人收集数据)实现了泛化,但产出的是“某一项任务的天才”。
  • 他至今仍在追求的综合,是 AI 的两项伟大成果:生成式 AI(LLM——复现人类所做的事)和深度 RL(AlphaGo 的 Move 37——做出“人类此前没有想到的事”)。把互联网规模的知识与超越人类的改进能力结合起来,就是目标——“我还没有解决它”,但他认为已经取得了一些进展。
  • 具体路径是视觉—语言—动作模型:先用文本训练,再用网络图像适配,最后加入多样化机器人数据;再叠加思维链,让机器人“真的和自己说话”(“捡起盘子”),从而让中间推断受益于互联网规模预训练,之后再通过实践强化学习。浓缩咖啡演示正是靠重复练习提升了鲁棒性、速度和吞吐量。
  • 背后的文化是自下而上的:机械臂农场由一名刚入职的“4级研究科学家”发起,他向 Jeff Dean 和 Vincent Van Houcke 要一座堆满闲置机器人的仓库,Jeff 回答:“好,做吧——你需要什么?”他受那些支持个人项目的组织启发:“ChatGPT 曾经基本上只是 John Schulman 的个人实验。它一度并不是一项有大量表格和饼图支撑的公司整体战略。”至于伟大研究者,“唯一不变的是没有任何东西是不变的”——决定性能力在于知道什么时候该继续死磕,什么时候该转身看看周围。

6. 瓶颈上移:现在可以指导机器人

  • 这项发现大约6个月前“有点随手一试”地出现:机器人在新厨房里失败时,只增加带有语义指令标注的经验——不新增遥操作的底层动作——就能改善泛化。瓶颈已经从物理执行任务,转移到“解释场景并选择正确下一步的能力”,而这可以用语言监督。Patrick 说:“基本上就是指导。”Levine 回答:“没错——只要和它说话,就能让它变好。”
  • 下一个明显问题是中层推理表征。LLM 让文本到文本变得便利,“但这不一定是具身系统行动所需的最佳表征——有时它需要空间思考,有时需要语义思考。”如何组织这种内部思考“可能是一个非常重要的问题”,答案也许不同于 LLM 世界里的答案。

7. 不要先算数据集规模,要把飞轮搭起来;硬件正在降价以匹配它

  • “我不认为有人真正知道需要多少机器人数据……我的感觉是,我们其实不需要知道。”关键在于把系统做到足够有用,让它能够被广泛部署并持续带回数据:“直白地说,Tesla 不会担心自己的汽车能收集多少数据。真要说的话,顺序恰恰相反。”
  • Patrick 追问:为什么不先造一个有用、由人操作的机器人,跑起来 Tesla 式飞轮?Levine 认为“这是个好主意”,但取决于领域:“有些人可能不希望家里有一台持续由异地人员控制的机器人,但在某些应用里,这或许无关紧要。”
  • 促成这一切的硬件组合正在形成:10年前 PR2 的成本约为40万美元;他在 Berkeley 实验室里的机器人约3万美元;现在每条机械臂“可能只有这个价格的十分之一”,而且还会继续下降。这些低成本机械臂若采用传统的精度依赖型控制,在工业环境中并不好用;学习能力可以弥补传感不足——该平台配备3个摄像头,没有触觉或力传感器,而“腕部摄像头本质上是伪装起来的触觉传感器,因为你能看到局部形变”。

8. 灵巧性出乎意料,机器人奥运会几乎横扫

  • 创办公司以来最大的意外是:灵巧性几乎不需要特殊处理。高度灵巧的行为以及跨具身迁移——包括多指手和不同自由度——需要微调数据,但不需要修改模型;“甚至不需要通过任何提示告诉它机器人是什么。我原本以为需要一些很复杂的技术。”
  • Robot Olympics 的起点是:前 Everyday Robots 的 Benji Holson 写了一篇博客,列出十几项 Moravec 式日常任务——开门、清洗沾满油脂的煎锅、用塑料袋捡狗屎。PI 按标准任务接入流程把它们作为运营测试跑了一遍——“我们没有为此开发任何特殊东西”——结果几乎全部解决,只在把衬衫翻到反面时失败(夹爪伸不进袖口),剥橙子则是“严格说来”失败(手指力量不够,最后用了小刀)。这说明了通用性的力量。
  • 机器人可以超越人类的地方是插接电缆。人类会停下来处理对齐,遥操作员停得更多。“结果发现,只要找到所有这些停顿并把它们移除,就相当直接。”强化学习是通用路线;若目标只是纯粹提速,一些简单技巧也能奏效。

9. 当前争议:苦涩的教训,以及仿真与现实的二分

  • 行业争论已经从“学习在机器人 AI 中有没有位置”——这个问题曾多年充满争议——转向端到端学习是否会胜出。Levine 认为,“苦涩的教训”还没有得到普遍接受:不要按照你认为机器应该思考的方式去编程,而要让它从数据中学习。他为反方提出最强论证:在复杂的开放世界里,“你不可能不用已有的物理世界知识,而这些东西我们都写进教科书了”。
  • 他认为真正尚未解决的二分是:人形机器人杂技式动作的流程“高度依赖仿真……实际上往往完全没有真实世界数据”;而操作任务很少使用仿真,大量使用真实数据和非常大型的基础模型。“可能其中一条路线最终会胜出……也可能两者会有某种综合。我不知道答案,我有自己的主观看法。”
  • 不依赖特定具身形态的路线有生理学依据:在猴子使用工具的研究中,追踪手部位置的神经元,实际编码的是工具尖端的位置,而不是手的位置——“工具作为身体延伸,是一种真实的生理现象。”他的结论是,一个好的基础模型应该能够操控它所控制的任何身体——“问题只有一个,不是很多个”。

10. 时间表、最后被攻克的任务,以及如何布局

  • 在整个判断分布中,他处于“资深机器人研究者里的乐观端、机器人创业者里的悲观端”。机器人历史上“成功寥寥无几”,他的联合创始人 Carol 有句话很贴切:“只有爬上这座山之后,你才会看到后面是不是还有另一座山。”时间之所以格外不确定,是因为部署本身是一个冷启动事件——必须“跨过激活能垒”,让机器人能够大规模收集开放世界数据。
  • 如果2050年仍没有厨房机器人,他怀疑罪魁祸首会是社会—技术长尾,但也指出开放世界广度是最大的技术风险——就像早期 Tesla 自动驾驶一样,问题在于:“人们是否能接受这种程度的不完美?”“你能接受机器人偶尔打碎盘子吗……在一个有小孩的家里?也许不能。这没关系。”纯技术层面的风险就是开放世界广度:酒店房间和餐厅厨房,他“非常清楚如何把它们控制住”;而在“几乎任何事情都可能发生”的家庭里,机器人需要“基本上在每种情况下都做出某种合理的事”。
  • 最后被攻克的任务将是:“给孩子换尿布会非常非常难。”养老照护和儿童照护——帮助一个人下床——是“Moravec 悖论的顶峰”,因为人类在与他人进行身体互动方面经历了高度进化;这些任务“可能会比人们想象的更难”。
  • 运营者应准备好面对的核心不确定性,是示范数据与自主经验之间的比例——“究竟是90/10还是10/90?”这会彻底改变正确的投资方向:遥操作设备、任务改造,还是新硬件。“机器学习需要数据,所以让我收集数据”是错误的默认选项——“你需要正确类型的数据”,而这取决于领域和具体判断。劳动力范式更像编码工具:不会突然替代人,而是形成一种共同演化的生产力“共舞”。PI 计划在2026年“尝试这个领域里的不同做法”,探索产品形态。
Patrick O'Shaughnessy

This is going to be a real treat and a blast to learn about possibly the most exciting, impactful area of technology being developed. Just to set the stage before we go back in time, maybe you could just define physical intelligence as you see it.

Sergey Levine

Fundamentally, the goal of Physical Intelligence is to develop robotic foundation models that can control basically any embodied system to do any task. Broadly speaking, you could imagine that, in the same way that a language model is rapidly evolving toward a system that can do any task that can be expressed in language, what we would like is to build a new class of models that can do any task that can be done by a physical, actuated device.

Part of the thesis of this company is that we believe doing it at the full level of generality might actually, in the long run, be easier than trying to specialize in very specific, narrow application domains. Again, in much the same way that, for language models, it turned out to be easier in some ways to solve natural-language tasks in their full generality than to narrowly target machine translation or sentiment analysis or whatever.

Patrick O'Shaughnessy

It may not be obvious why you would make that bet versus a robot that just does your dishes or something. What are the key trade-offs to understand, and why make the decision that you made?

Sergey Levine

Maybe I can give you a 2-part answer to this. The first is how it relates to the analogy to language models, and the second is what that means in the robotics world.

The first one is a little bit more informed by evidence. In the world of natural language, we saw that there were a lot of efforts to develop domain-specific solutions that tackle specific problems. Somebody would spend a lot of time thinking about how English differs from French and then build a machine-translation system.

1. Defining Physical Intelligence

The reason that language models took over for all of those different application domains is because they can leverage much broader sources of data. It's not even as simple as saying, “We had this data for this application, this data for this application, and we merged everything.” It is actually more than that. When you can leverage weakly labeled data—data that, in the case of language models, you just mine from the web—you actually learn more about the world. You establish a foundation of world understanding, and then building out different applications on top of that foundation turns out to be much more effective.

To bring this into robotics, obviously that calculus doesn't look quite the same because, in robotics, we don't have an internet-sized data set that we can just draw on. But this notion of understanding the world, if anything, is actually more important in robotics, because if you have many different tasks, maybe even many different physical systems, then you can go from training an individual dishwashing specialist or laundry-folding specialist and instead train a model that actually understands physical interaction.

People can master new skills very, very rapidly because we understand physical interaction. We can intuitively grasp what's going to happen in a new, unfamiliar situation, and that bootstraps things really, really quickly. So, if we can draw on data from many sources, many applications, and many robots, then we can have a model that has a physical understanding, and it'll be much easier to put new applications on top of that platform.

Patrick O'Shaughnessy

What is the hardest part about building in this way for you when you see other approaches that are more legible to the average person? You see a robot moving around doing one specific thing, and it looks a certain way. What's the hardest part about this approach as you're doing it?

Sergey Levine

I think this has actually been an issue in my whole career, because when you work on robotic learning—and the more general it is, the more this becomes important—effective robotic learning and effective generalization aren't actually the optimal way to have a really exciting demo.

The way to have a really exciting demo is to pick a really cool task, control everything else in the environment, set it up so that it's perfectly clean and pristine, and just make it work in that one setting. That's the way you make a robot demo.

With generalization, you can't just show it in one spot. The point of generalization is that it does something relatively mundane that any human could do, but it does it in any situation. We had some demos that we released last April where we showed our robot cleaning kitchens. I think it's kind of cool, but if you watch an individual video out of context, it's just like, “Okay, it's picking up plates. Anybody can pick up plates.”

2. The Challenge of Building General Models

Except that we put it into that home just for that demo, and it had never had training data from that setting. So, obviously, you have to understand what's going on to appreciate why this is actually pushing the frontier.

Patrick O'Shaughnessy

What is your model for the stakes of what you're doing? If you are successful, I'm curious for you to define what that would mean—successful other than just, “We crossed this chasm of general physical intelligence,” or something. But if you cross that line, then what?

Sergey Levine

One of the things that I think would be really, really exciting, and that would be enabled by a general-purpose embodied foundation model, is the ability to unlock people's imagination in how they build robots and other embodied systems.

Personal computers were a really big deal in my mind because they made it possible for lots of people to hack together all sorts of really cool stuff. There was this Cambrian explosion of amazing applications that started in the '90s and so on, and then was further accelerated by the internet.

I think something like that might happen in the world of robotics, but it can't happen today because if you want to put together some cool new robotics application or some cool new robotics idea, you're going to have to build this monstrous stack, and you need to basically solve the intelligence problem.

3. The Stakes and Future of General Purpose Robotics

But if there is a solution that someone can build on top of—if there's a foundation model that you can prompt that'll provide basic functionality, and then you can maybe fine-tune it a little bit or adjust it in some way to your application—it actually makes it a lot more tractable for lots of people, lots of companies, and lots of individuals to try out all sorts of different things.

Sometimes we think that robots are going to be one thing. It's just like, “Yeah, there's people, and now we're going to make metal people, and that'll be robots.” But I don't think that's how it's going to be, because no technology has been like that.

It's going to be more like a toolkit where you can put together all sorts of really cool applications and get really creative with it. Maybe I'm going to make a robot with 5 arms, and this one is going to hang from the ceiling. You can figure out the right thing to tackle your domain and maybe also experiment with software. But you need the right platform on top of which to do that, and I think the foundation model can be that thing.

Patrick O'Shaughnessy

What are the pros and cons, in your mind, of the humanoid approach to robotics?

Sergey Levine

One pro of it is that it's really cool. You can show it to somebody, and they're like, “Yeah, I get it.”

Patrick O'Shaughnessy

Yeah, you hear a lot about the Optimus hand.

Sergey Levine

Yeah, but it is cool. I think there's a lot of value to that. There's a lot of value to capturing the imagination, and there's a lot of value to getting people to think about what the future might look like in a way that's understandable.

But in my mind, it's one of many possible kinds of robots that we're likely to have. Fundamentally, I think the intelligence challenge looks very similar for all these different robots, and I don't think we should be tackling intelligence in the context of one specific body. I think we should handle it in a general way, because otherwise it's just really hard to get a handle on this. We need lots of data.

The cool thing about being able to build robots is that, ultimately, they don't have to be constrained to look like humans at all. You can build the right tool for the job. You could imagine that you're building a house with a robot that is a swarm of 1,000 quadcopters.

4. Pros and Cons of Humanoid Robots

I think that in the future we'll have a robotic foundation model that can then be adapted to all sorts of applications, and it might really run the gamut from bulldozers to humanoids to robotic arms like this thing.

Maybe it would need to be adapted to each one. Maybe it would need to be fine-tuned. Maybe it would need something in context to understand how that body works. But the fundamentals of how you interact with objects, how things move in the world, and how causality works—that's all conserved for all these different systems.

Patrick O'Shaughnessy

Do you have a favorite example of what might be possible with true general intelligence that might not be possible with, say, humanoid-only intelligence or something?

Sergey Levine

There are a few things that I think are worth thinking about. One is that we can make machines that are very big and machines that are very small. I think that in the long run—this is not by any means a short-term thing—there are lots of really exciting applications in medicine and surgery where we not only might, in the long run, not be limited to robots that look like humans. We might not be limited to robots that can even be controlled by humans, right?

Because currently, for example, robotic surgery is done entirely through teleoperation. So you need something that a person can control in real time with the right level of dexterity. And of course, that limitation holds for current learning-enabled systems, too. But in the long run, we could imagine addressing that.

Patrick O'Shaughnessy

If you think about the most important hash marks on the timeline of robotics research that have gotten us to here, I always think it's super helpful to set the historical context before we talk about what the state is today and where we're going. What could you walk us through? What are the relevant hash marks on the history timeline?

5. Historical Milestones in Robotics Research

Sergey Levine

At some level, doing end-to-end control for robotic systems is a very, very old idea, right? The first, for example, autonomous driving systems that used end-to-end learning existed in the 1980s. Alvin, I think, was from 1986 or 1987, and that was a driving system that was demonstrated to drive on highways, controlled by a neural network from a camera. The neural network was tiny, right? So I think that there are some very venerable concepts.

But historically, what has been really difficult in robotic learning is that you need a system that handles the application you want to address and is cost-effective to train for that application. That means you don't need a huge amount of data for every single application you want to tackle. It also needs to handle long-tail scenarios with common sense, so if something weird takes place in the world, it needs to have a reasonable response to it. And then, for the thing that it's actually supposed to do, it needs to be robust, fast, and reliable.

Getting all those things together is very, very hard because, with machine learning, it works best when there's a lot of data. So if you naively approach a robotic problem and say, “I want to do washing dishes,” right? You'd obviously proceed to collect an enormous amount of data on washing dishes, but that's not cost-effective, because then you go on to the next application and you go through that process all over again. So being able to train general-purpose models that can handle many tasks is essential to this, because now you need a lot less data for each new task.

But then, even further—and this is the thing that has probably changed the most in the last few years—you also need to handle the unusual scenarios. And for the unusual scenarios, you are probably not going to have experience. What you need to rely on is knowledge that you've acquired from other sources that you can ground in that new situation.

People are extremely good at this. So if you're driving a car and there's something going on in the middle of the road and someone put up a sign saying, “Don't go here. There's a gas leak,” or something, you've probably never experienced that before, but you can put these things together and figure out what you're supposed to do in that unusual situation because you have common sense. And this has been a huge mystery in the robotic learning world: where do you get that common sense?

This is what's changed in the last few years, because it turns out that multimodal language models are really good at pulling in knowledge and trying to articulate that knowledge. They're not very good at grounding that knowledge in physical situations, but they know stuff. So now there actually is a path to get that kind of common sense by essentially leveraging the knowledge that is contained in multimodal LLMs. But there's also a challenge, because you have to somehow plug into that knowledge in the right way.

You can't just show it a picture and say, “What would you do here?” because it doesn't have the context. It doesn't know that you are a robot, this is what you look like, and this is what's going on. So that's a technological challenge, and we've made some headway on addressing that technological challenge. The research community in general has, too. But most important is that light at the end of the tunnel: “Okay, now we have this way of pulling in lots of knowledge which can help us handle those long-tail scenarios.”

Patrick O'Shaughnessy

Are there hash mark equivalents on the timeline, like AlexNet or the Transformer? Are there big, major events that you think everyone will point to when writing the history books about this?

Sergey Levine

That's a good question. I think it's very early on right now to answer that definitively. I mean, certainly, I would say that you're going to have to look back at least 10 years down the line, something like that. But probably the first end-to-end learning systems, which were in the '80s, that's definitely a milestone. I think that the first deep reinforcement learning systems, which were in the early 2010s, are probably a milestone because deep reinforcement learning gives us a way to go beyond human-level performance, which I think will be essential for robotic systems.

And then there's the more recent stuff, but that's in the last few years. I don't know how that's going to shake out as far as whether that's something that people will point to, but I do think that the advent of multimodal LLMs that can be adapted to robotic control to bring in that common sense is a really important advance. I think we're probably going to see quite a few important advances in the next few years. Maybe those will be the things people point to.

Patrick O'Shaughnessy

Can you tell us your own personal history of approaching the problem? So maybe the origin of when you first became interested and why, and then how you've decided what to spend your personal time and attention on ever since then?

Sergey Levine

I started working in robotics in 2014, and that was actually after I finished my graduate degree and started a postdoc with Professor Pieter Abbeel at UC Berkeley. I actually hadn't worked on robotics before, but I figured I should get a little bit more education after finishing my degree, and his lab worked on robots, so I tried to apply what I had learned previously to robotics. Before that, I worked on computer graphics, actually.

I think that the thing that I've always wanted to really figure out is how to get AI systems that get better and better the more they do things, because I think that's tremendously powerful. If you can have a system that gets better and better the more it does something, and it just keeps getting better, then there's sort of no limit: it can master all the skills you'd want it to master.

Initially, I tried to approach it in a very blank-slate way, meaning you start with nothing, you practice a particular skill, and you get better at that skill. That was kind of okay: you can do that in a limited setting, and you get something that works, but it's very hard to turn that into a general system that can work in open-world settings. Because if I practice something over here and then it goes over there, now something is different, and it needs to practice all over again.

6. Combining Generative AI and Deep RL

The next thing I tried—and I did this when I worked at Google afterward—was to see if we could do that, but now parallelize it across many robots. So, collective learning: can you put 20 robots in a room and have them all learn together? And that works. It generalizes. But it's very hard for that to handle these tail cases, these edge cases, right? Because now it becomes this kind of savant of this particular task, and that's all it knows in the world.

So what I think is the next step is what I mentioned before: combining this ability to practice skills with lots of prior knowledge. And that's actually a really, really hard problem. You know, it's not just in robotics where it's a hard problem. I think it's a hard problem in all of AI because, arguably, the 2 big impressive results in AI over the last few decades have been generative AI and deep reinforcement learning.

If you want a single example to epitomize generative AI, that's LLMs. Deep reinforcement learning: AlphaGo, right? They're both very, very impressive, and they're very impressive for very different reasons. Generative AI is impressive because it can reproduce some of the things that humans can do. It can draw pictures that look like human pictures and write text.

Deep RL is impressive for the opposite reason. It does things that humans hadn't thought of. Like, it can make Move 37.

Patrick O'Shaughnessy

Exactly.

Sergey Levine

So I think the big challenge—and this is kind of what I'm leading up to and what I hope we'll figure out here at Physical Intelligence—is how to combine those threads.

How to bring in all that knowledge that you get with generative AI, but also go beyond just human-level performance with reinforcement learning—I haven't figured that out yet, but I think we've made some good progress on that.

Patrick O'Shaughnessy

So what literally have you done and are you doing to make that happen?

Sergey Levine

So in the past few years, we started off first by developing the basic foundations. The basic foundation is what's called a vision-language-action model. A vision-language-action model is something you can think of as an LLM that has been adapted for robotic control.

The way these things are trained is they're first trained on text data, then they're adapted with lots of image data from the web to understand images, and then they're adapted to robots with lots of very diverse robot data. Now, that's a starting point. That's a way to take all of that web knowledge, get it into a model that can control robots, and get some interesting behaviors out of it.

From there, we studied two threads: how to get this thing to handle unusual situations with common sense, and how to get it to improve with reinforcement learning. The way you get common sense is by essentially using chain of thought. So the robot enters a scene, and instead of directly starting to move, it thinks about what it was asked to do.

If it was told, “Clean up the kitchen,” it looks at the scene and says, “Okay, based on this, I should pick up the plate.” It literally talks to itself. It says, “Pick up the plate,” and then it goes and does it. That unlocks all this prior knowledge because those intermediate inferences benefit from the web-scale pretraining.

That handles edge cases. Then the reinforcement-learning part comes in after you've practiced it a few times, and you keep getting better and better at the task directly through your experience. For example, we had this demo on making espresso. That system practiced making those espressos many, many times and used that to improve robustness, improve speed, and improve throughput. We're not done with that. I think there's a lot more to do there, but we have the starting point.

Patrick O'Shaughnessy

Sorry to be thick about it, but the robot data itself is the right way to think about it. I'm looking at the Gen 1 of these things. I see a camera here; maybe there's some sensors somewhere else. Is the data effectively being gathered by various sensors strategically placed on the robot at different parts?

Sergey Levine

Yeah. Something I'll say about sensors is that I think you can actually get away with less than one might think and still do quite a lot. So this platform here has 3 cameras, 1 on each wrist and a base camera. It doesn't have touch sensing or force sensing. It's very bare-bones and very low-cost.

I'm sure that more sensors could make it better, but a good learning method can actually compensate for deficient sensing fairly well. The wrist cameras are essentially a touch sensor in disguise because you can see local deformations when you touch something.

7. Moravec's Paradox

Patrick O'Shaughnessy

If I think about the analogy to the expert systems of the ’80s and ’90s in basic AI, the lesson that scales all you need, and the counterintuitive nature of that—that you're not teaching it any specific thing, just blasting it with data, and there's this reservoir of internet data—talk about the reservoir and how to create the reservoir of data needed for this.

Sergey Levine

Yeah. I don't think anybody really knows how much robot data is needed to have truly generalizable and powerful embodied AI. But my sense is that we actually don't need to know. What we need to do is get to the point where these systems are useful enough that they can go out to the world and gather more data themselves.

Like, to put it bluntly, Tesla doesn't worry about how much data their cars can collect.

Patrick O'Shaughnessy

Right.

Sergey Levine

If anything, it's the other way around. That's a little too much data, right? So I think that the key is not so much to quantify, like, here is exactly the price tag of getting the ultimate robot dataset. The key is to get a system that can go out to the world, that's useful enough, that does a wide variety of different things, and that can keep pulling in more data.

Patrick O'Shaughnessy

You brought up the example of Tesla—the beautiful system of a thing that's useful without the AI to begin with, because a human drives it and it gathers data. Why then not start with your best guess at something that's useful as a single robot to have the same sort of flywheel thing happen?

Sergey Levine

Yeah.

Patrick O'Shaughnessy

I think it's a good idea.

Sergey Levine

Yeah.

Patrick O'Shaughnessy

And do you think that's an approach that you'll pursue?

Sergey Levine

I don't think that there's one right answer, right? I think there are some domains where deploying a system under human control makes a lot of sense. There are some domains where deploying a partially autonomous system is very reasonable. It's kind of domain-dependent, right? Because robots aren't just one thing.

Maybe some people might not want a robot in their home that is constantly being controlled by a person off-site, but maybe for some applications that doesn't matter.

Patrick O'Shaughnessy

Yeah. If you mark the start of Physical Intelligence through today, what has been the most surprising thing to you that you've discovered, or about the nature of how the research has gone?

Sergey Levine

One of the things that's been surprising to me is that I think we've made a lot more progress on dexterity than I thought we would. I had an expectation based on my prior work that generalization—being able to handle all sorts of different scenes and all sorts of different objects—would steadily get better if we just collected more and more data.

But what was surprising is that we could also get these systems to perform very dexterous behaviors without really doing anything particularly special for that. The same, by the way, also applied to getting systems to work on different embodiments, where we could get our models to work on all sorts of other robots, including robots with multi-fingered hands and robots with different numbers of degrees of freedom.

Obviously, we needed to get data and we needed to fine-tune the model, but the model itself didn't need to change. It didn't even need to be told through any kind of prompt what the robot was. That was also surprising to me because I would have thought that we would need some fancy techniques to adapt the system to faster, more dexterous, more complex tasks and also to different kinds of embodiments. But it actually seems to generalize pretty well across those.

Patrick O'Shaughnessy

I'm always interested in the spectrum of capabilities, and especially where the systems today are more advanced than you think people would probably expect and where they're less advanced than people might expect.

Sergey Levine

This is something that's always been very tricky to understand in robotics. There's this idea that roboticists always talk about called Moravec's paradox. It's actually true in all areas of AI, but especially in robotics, it's a big deal.

We kind of have a cognitive bias to think that things that are easy for us will be easy for the machine, right? Solving calculus problems is difficult for most people, but picking up a cup is easy for most people, so we think, “Machines should be able to do this.”

But it's actually the other way around: there are things that are easy for us because they have to be, otherwise we wouldn't survive. We're very good at spotting the tiger in the jungle because the people who weren't so good at it got eaten by the tiger, and they're not around anymore.

Because of that, we have this cognitive bias and think that there are things that should be very easy, but they're actually very difficult engineering challenges. However, something that is changing is that machine learning slightly changes that equation. Programming something by hand to pick up any cup anywhere is difficult. Getting a machine learning system to do it, if you have data for it, is actually not that difficult.

8. Kitchen Robots

I think increasingly what we'll see is a shift where domains where collecting data is straightforward actually end up falling into the easy bucket over time, even if they are physically intricate. But there will be domains where collecting data is difficult, where you need to use more common sense, where you need to reason at multiple levels of abstraction, and connect physical skills that you've learned in other areas to knowledge that you got from the web. Those will be tough, and that's where we'll need more technology advances.

Patrick O'Shaughnessy

What is the science of common sense? You mentioned common sense before. When we say that, what does that mean?

Sergey Levine

For the purpose of robotic learning, we can think of it as applying semantic inferences using knowledge learned from other domains to the current physical task at hand. You can think of common sense as sort of the opposite of muscle memory. If you play a sport and practice something a lot, you hardly think about it. You just do it on autopilot.

Common sense, in my mind—I don't know if this is the conventional definition, but I think it's a reasonable definition—is when you know something to be true because you saw it, read about it, or heard it, and now you're in a situation where that fact is highly pertinent to what you need to do. You're able to make that connection, apply it to your situation, ground it in the environment that you're in, and make the right decision.

Patrick O'Shaughnessy

One of the other differences that's so interesting to me is that everyone's used a chatbot now. You query it, you get an answer; you query it, you get an answer. We're now seeing what happens with Claude Code and other things, where you give it something complicated and it's able to do a very long task. The measure is how long it can go without failing. What's the similar long-range thing in robotics?

Sergey Levine

It's something that we're working on quite a bit right now, and in fact, the methodology is not that different at some level. The way that our models work now, as I mentioned, is that they use this kind of chain-of-thought process to reason about the task. When you have that, you can actually do very long-horizon tasks. You can have a robot that takes out all the dishes from the dishwasher, puts them in the correct cabinets, wipes down the counter, and does all that kind of stuff.

The interesting thing here is that we found, maybe about 6 months ago, that our models had gotten to the point where they could be improved just from supervising them with high-level instructions. You take a robot, put it in a new kitchen, and ask it to clean the kitchen. It gets to work and then fails somewhere. So, what do you do?

Traditionally, what we would do in that situation is add more teleoperation data to cover a wider range of kitchens. But what we tried, kind of on a whim, was to see what would happen if we didn't add more teleoperation data. What if we just added more data labeled with the semantic command? Basically, just take whatever the robot experienced and label it with some semantic commands, but don't add any more low-level actions.

That actually helps. It improves the robot's ability to generalize. What that means is that the bottleneck had shifted from the lowest level—meaning the robot's ability to physically do the task—to this middle level, where the system is now more bottlenecked by its ability to interpret the scene and select the correct next step, which can be supervised with language. That's a big deal, because now someone can literally talk to the robot.

Patrick O'Shaughnessy

Coaching, basically.

Sergey Levine

Yeah, exactly, and make it better just by talking to it.

Patrick O'Shaughnessy

If we're in 2050 and there's no robot in my kitchen doing my dishes for me, what do you think the most likely explanation is for it not having gotten there by that point?

Sergey Levine

My suspicion is that there is a long tail of challenges that has to do with the interaction of technology and people. In some ways, autonomous cars aren't that different in this regard. Getting to a level of comfort with deploying autonomous vehicles on the road was a significant challenge that ran in parallel with getting the technology to that level.

9. Simulation vs. Real-World Data

For example, Tesla's early self-driving was a bit controversial because it wasn't perfect. There was a question of whether people were comfortable with that level of imperfection. So, probably there are some tasks for robots where people will be comfortable with something that's not perfect, something that needs to learn from its mistakes. There are some areas where people will not be comfortable.

Patrick O'Shaughnessy

Are you comfortable with occasionally breaking your dishes? Maybe in a few years it will stop breaking those dishes, but maybe in the meantime it's not quite there. Are you comfortable with a robot like that in a home where there are small children?

Sergey Levine

Maybe not, and that's okay, right? I think that figuring out how those factors interact and what that means for the timeline, and for how these systems get better with experience, is a tricky question. I think it needs to be approached very carefully, with a lot of sensitivity. There may be some domains where it makes a lot more sense for these systems to be deployed, bootstrapped, and used to collect more data. Maybe other domains require more care.

Patrick O'Shaughnessy

Could you imagine a purely technical explanation for why something might not work?

Sergey Levine

I think the place where I would see the biggest technical risk is dealing with the breadth of different situations. If we were talking about a well-defined but slightly chaotic environment, like cleaning hotel rooms or assisting human cooks in a restaurant, I think I have a very good sense for how to get that under control.

10. The Robot Olympics

If you're imagining a robot going into a home, one place where I can anticipate a challenge is that there are a lot of other unexpected things that can happen, and you need a system that's very good at inferring what's going on and adapting to it or reacting intelligently. We have a lot of ideas for how we can approach it, but that's the hardest part of the problem.

When you're in a situation where just about anything could happen and you're controlling a physical device that affects the world around it, you really need to get things right, at least at some level, pretty much in every case. It doesn't mean that you always have to succeed, but it does mean that you always have to do something sensible that people are okay with. I think there are a lot of really good ideas for how to do that, but that is probably the most challenging part of the equation.

Patrick O'Shaughnessy

If I go back to thinking about the right model for the Physical Intelligence approach to doing this whole exercise, help me make it as simple as possible. One way to think about it might be that we're going to build a whole variety of different kinds of things, with different kinds of form factors, to do a whole variety of different kinds of things, mash all this data together, and start to experiment with how we can make it better on evals. Is that just the simplest way of doing it, or is there an even simpler way?

I'm asking because I'd love to then contrast it with some other approaches that you're interested in, that you're not doing, that others are doing.

Sergey Levine

I see. In my mind, the most important thing to get right is to get the system to be general, and in particular, to get it to be general with respect to how it can be improved. Hand-designed robotic controllers are not very general with respect to how they can be improved, because it requires a human engineer to go in and improve them.

A learning-based perception system is more general because all it requires is human labelers to go in and label more data. A system that learns autonomously from data that it gathers through its own experience is even more general because you don't even need the human labelers. So, the key is this generality, particularly with respect to improvement, and the decisions we make are, to a very large extent, centered around that.

I don't know if the correct design for a robot is to have 3 cameras. I don't know if it needs a touch sensor. I think we're very agnostic to that. We'll try a lot of those different choices. I'm not even sure if, in the long run, it's going to have a language model. Maybe it will have some other kind of model that's trained on very diverse data. But the key is this level of generality.

Patrick O'Shaughnessy

What other approaches are the most interesting to you?

Sergey Levine

I think one thing that's a very important question in this area, and something that I think the research community and the tech community have not fully answered, is the dichotomy between different data sources, particularly with respect to real data and simulation. It's a very controversial topic. I have a very strong opinion about it, but I think it's worth acknowledging that if we look, for example, at humanoids—if you've seen videos of humanoids doing all these acrobatics—there's a particular pipeline that makes that work, which is very heavily reliant on simulation and actually very light on real-world data.

Often, there is actually zero real-world data. And then there are the approaches that work well for robotic manipulation, which are often the opposite. They often use very little simulated data, large amounts of real-world data, and very large foundation models. It is surprising that in these 2 robotic domains, the dominant approaches look so different.

Now, it may be that 1 will win out and there is a particular approach that can handle everything in the long run, or maybe there’s some sort of synthesis of these ideas that’s important. I don’t know the answer. I have my own subjective opinions. I think the approach we’re taking is a very good one, but I think it’s interesting to look at that and see why these things are so different.

Patrick O'Shaughnessy

Can you talk about the contrast between cool and useful? The Boston Dynamics robot is very cool. The backflip is super cool. I don’t know what I need that requires a robot to do a backflip, so I’m curious how you think about optimizing around cool versus useful.

Sergey Levine

I think the strategy we’ve taken—I don’t know if it’s the right strategy—is subject to the constraint that it’s useful: make it as cool as possible. That’s reflected in our blog posts and our videos. We make decisions first and foremost based on our assessment of what will drive the technology forward toward this truly general, broadly applicable robotic foundation model.

But in doing that, we try to stress-test it against the toughest challenges we can throw at it. The toughest challenges are often ones that look cool. We didn’t set out, for example, to build a robot that can make espresso or fold the laundry, but in the process of building these general systems, we figured these would be particularly challenging, particularly exciting things to try with them to see how far we can push them.

Patrick O'Shaughnessy

Can you talk about the robot Olympics?

Sergey Levine

Yeah. There was a gentleman named Benji Holson who used to work at Everyday Robots, part of Alphabet before it dissolved, and he spends a lot of time thinking about tasks that robots can do. He wrote a really interesting blog post a while back where he basically said, “Hey, there was this robot Olympics held in China where robots would run around on a track and jump and so on, but maybe these aren’t the real challenges we should worry about. How about a robot Olympics centered around everyday tasks that people do?”

11. The Physiological Reality of Embodiment

That’s more of a paradox thing, where the tasks are things that people find really easy but that robots struggle with. He had things like opening a door, washing a frying pan with grease on it, and using a plastic bag to pick up dog poop. Those are things that people don’t find particularly challenging, but no current robotic system can do them. He listed maybe a dozen of these things.

We wanted to give this a shot. This wasn’t actually part of a concerted research project. It was more that we had developed processes and systems for ingesting new tasks that we wanted to use for all sorts of tasks, and we figured a good way to test this was to say, “Here’s a big list of tasks. Let’s just go through this process that we’ve developed and see if it works, basically.” It’s almost like a test of our internal operations and model-training system.

Then we tried these things, and it turned out that we could solve almost all of them. There was 1 we couldn’t do: turning a dress shirt inside out, because the grippers on this thing wouldn’t fit inside the sleeve. We probably need to change the gripper. On a technicality, we didn’t succeed at peeling an orange because they said, “Do it with the fingers,” and our fingers weren’t strong enough, so we had to use a little tool, basically a little knife. But everything else we could do.

What was interesting to me is that, obviously, it’s cool—the videos are nice—but if anybody watches those videos, 1 thing that’s important to keep in mind is that we didn’t develop anything special for this. We literally used this as a test of our task-onboarding process. I think that’s interesting because it suggests the power of generality: when you have this kind of general system, you can onboard all these crazy tasks without really doing anything particularly sophisticated.

Patrick O'Shaughnessy

I was curious before, when you said superhuman ability, like on dexterity or something like that, where we’re limited by what we can do or maybe by what we can control, even if it gets smaller. What are some of the other dimensions like that where we might surpass human ability in terms of physical ability? What are the other trend lines that are most interesting to you?

Sergey Levine

Here’s a fun 1. We were working on a task where our robot had to plug in cables, like power cables or Ethernet cables or something like that. When a person does this without having practiced a lot, they pause frequently, right? You have to process what’s going on. You have to make sure that it’s all aligned and all that stuff, so you do it very slowly.

12. Controversies in the Robotics Community

If you’re teleoperating a robot, you do it even more slowly because there’s this level of indirection. It turns out to be pretty straightforward to find all those pauses and remove them. You can speed things up further, so you can get a task where a person demonstrates what it means to succeed, and then you can have the robot practice the task and succeed in the same way, but a lot more quickly and a lot more efficiently.

The most general way to do this is with reinforcement learning, but there are also some simple tricks you can do if you just want speed. That’s 1 example of something where you can have a machine that does it a lot better. At some level, you have a processing bottleneck. That’s why a person does it slowly, because they have to process what’s going on, but speeding up processing is something that people understand quite well in computer science.

Patrick O'Shaughnessy

There’s this amazing Michael Crichton novel called Prey. Is it a question about form factor? It seems like for a given problem, there may be an optimal shape, or set of optimal shapes, for the robot to perform the task, and that what you should do is analyze the problem, then have something that can almost morph or transform into the form factor. How do you think about that—the innovation on the form-factor side rather than the data and model side?

Sergey Levine

I think that, in general, in robotics, the ability to innovate on form factors has been very constrained because of the AI challenge. If you have a traditional AI pipeline—you’re doing some motion planning and stuff like that—it’s hard to just cobble together some new robot. When you do that, you have to characterize the dynamics of the system. You have to do system identification. You have to build up all this stuff.

If you could just put together a robot in your garage, load up a robotic foundation model, and tell it to do a bunch of stuff, maybe it won’t be perfect out of the box. Maybe it needs more data to really perfect it, but you can at least get the thing moving. I think that could be a really powerful engine just to get everybody to experiment with this stuff.

I don’t think that I’m the right person to design the perfect robot. There are people here, of course, who are a lot better at that. But in general, I think it’s just like with personal computers. The key is to let people experiment and play around with it and radically lower the barrier to entry for that. I think then we’ll see a lot more creativity.

When people first started using personal computers, there were a limited number of form factors. Now you can have a computer in your phone, a computer in your car, or embed a computer in your refrigerator. They’re everywhere, and they’re very different. Generality, good software, and a good foundation on which you can build applications—those are key to enabling that.

Patrick O'Shaughnessy

Professor Andy Clark once described to me the feeling of physical intelligence for a human as being like learning how to ride a bike. There’s that moment when you didn’t know how to do it, and then you do know how to do it, and that feeling is physical intelligence—that snap of understanding.

Sergey Levine

There’s actually a physiological explanation for this. There were studies that were done in monkeys using tools, and you can actually find which neurons activate for the monkey to figure out where its hand is. It turns out that if it’s using a tool, they activate based on the location of the tool tip, not based on the location of the hand.

The tool being an extension of your body is a real physiological thing. Your brain literally does that.

Patrick O'Shaughnessy

So, knowing that, what does that do to impact the approach to your research?

Sergey Levine

I think, to me, it says that physical intelligence should be, at some level, agnostic to embodiment. A good foundation model should figure out how to manipulate whatever body it’s controlling and whatever tools it has at hand—that there’s basically 1 problem, not many different problems.

There isn't a humanoid problem and a car problem and a bulldozer problem and a robot bolted to the table problem. There is one problem, and if you solve it at this full level of generality, that's really, really powerful.

Patrick O'Shaughnessy

We're in the early stages of seeing some of the job and other sorts of transformation in businesses and the economy that LLMs make possible. Certainly, we've seen it in engineering. How do you think about what might happen, or what you hope will happen, when we're at a similar stage, whenever that happens to be, for robotics, where all of a sudden we have this thing that's general and useful?

People are creative. Where do you think the world's very efficient at deploying these things? Where do you expect to see the world start to change most in the early days post-physical intelligence?

Sergey Levine

Yeah, that's a really interesting question. I really don't know. I don't think anybody would have been able to predict how the LLM stuff evolves. People would have guessed, but this is why I keep coming back to this idea that maybe the key is to let people try lots of things.

13. What Makes a Great Researcher

One of the really amazing things about applications of LLMs is that they are really accessible. Somebody could put together a really cool new prototype that, under the hood, is just prompting ChatGPT or something, but they can experiment with it, try it out, and see what it does. There's an amazing power to having lots of smart people rapidly iterating and prototyping lots of things.

That's why Physical Intelligence has really put a premium on engagement. We've open-sourced our models. We would like to engage with lots of other companies that are building robots, because we all see a lot of power in this effect of having many people trying out lots of things.

Patrick O'Shaughnessy

What are the major controversies in the robotics community?

Sergey Levine

Obviously, I'm an academic, so to me, a controversy is someone getting in an argument with me at a conference. But I can tell you that the kinds of arguments that I've found myself in have followed an interesting trajectory.

In the early days, the main argument I would have with people was whether learning has a place in robotic AI. Part of why that was often a controversial point is that, in a traditional engineering pipeline, robots look very different from software artifacts. They're physical, they can affect stuff around them, there are safety considerations, and there are a lot of weird situations they can get into.

It took a really long time for the robotics research community to really internalize that you don't necessarily need to program in things like knowledge of physics. You don't necessarily need a physics simulator inside your robot when it's planning. We can actually have a learning system figure all that stuff out, and that was a very controversial thing for a very long time.

I think at this point there's a lot of acceptance that learning is a really important part of robotics, but I don't think there's still universal acceptance that end-to-end learning is the right way to go. I don't think there's universal acceptance of the bitter lesson.

The bitter lesson says that you should not program the machine to think the way you think it should think, but you should let it learn from data. That is not a universally accepted idea. I think there are good arguments against it, but I think that in the long run, if we want that generality—especially generality in the machine's ability to improve—then we need it to primarily learn from data.

Patrick O'Shaughnessy

What is the good argument against it?

Sergey Levine

My best attempt at steel-manning this is that if you want something reliable in a really complicated, open-world setting, then you can't afford not to use what you already know about the physical world. We've got textbooks full of this stuff, so why don't we just plug in what we know from the textbooks?

Patrick O'Shaughnessy

What is compositional learning? Can you describe that?

Sergey Levine

There's an example I can give you, which is maybe the more vivid way to communicate this. This is an example that is due to one of my students, actually.

He had this idea where he asked a language model to provide a recipe for how to make a sandwich in the International Phonetic Alphabet. The International Phonetic Alphabet is made up of symbols that they use in a dictionary to explain how to pronounce a word. It's very peculiar because it only ever appears for individual words in a dictionary. You never see free-form text written in the International Phonetic Alphabet.

But if you ask a good language model, it will write paragraphs in IPA for you. That is compositional generalization. It means that you have never seen this particular language or this particular alphabet used to write paragraphs, but you understand paragraphs. You understand that it's compositional with different alphabets, so you can solve the problem.

You can imagine the same thing coming up in robotics: you've learned a repertoire of skills, and now you can combine and mix those skills and apply them to solve new problems.

Patrick O'Shaughnessy

It makes me wonder what the last types of tasks you think will be possible for a robotic system to achieve.

Sergey Levine

I think changing a child's diaper will be really, really hard.

Patrick O'Shaughnessy

Say more.

Sergey Levine

I think this really is just Moravec's paradox all over again. People are extremely good at certain things. We're very good at physical things, and we're also very good at interacting with other people. That makes sense. We have to be. That's a lot of our systems.

14. How Businesses Should Prepare for Robotics

Things that involve behaviors that interact with other people, where you actually have to help somebody—like helping somebody get out of bed or something like that—I think are a lot harder than people appreciate. I think elderly care and taking care of small children are going to be hard, and they're probably going to be harder than people think.

Mhm.

Patrick O'Shaughnessy

And the stakes are very high. I want my baby to be safe.

Sergey Levine

Not just that. The stakes are high in many places. It's just that this is probably the pinnacle of something that fools us into thinking that it's easier than it really is, right?

We are so evolved for interacting with people and doing things physically. If you're helping somebody get up the stairs or get out of bed or something like that, you don't have to think very carefully about how you're going to do that. You kind of know. So I think it's really the pinnacle of Moravec's paradox.

Patrick O'Shaughnessy

If I think about an LLM as a brain, and now it's effectively studied everything—I don't know how else to put it—and then I think about a robotics model's brain instead, what are the dark parts of the brain? What has it not been able to study or penetrate or learn? What are the areas that have just been really difficult, that matter but have been hard to—

Sergey Levine

Mhm.

Patrick O'Shaughnessy

—for us to get into?

Sergey Levine

One of the things that people are remarkably good at is using physical analogies to understand other situations. I don't know whether this is something that LLMs can or can't do, but it is something that people use a lot. They use it in everyday life, and they also use it for very sophisticated problems.

For example, you could say, “That company has a lot of momentum.” That's a physical analogy. You know exactly what it means. I don't have to explain that statement to you. But if you actually think about that, it is quite a complex thing. There's a lot riding on that word “momentum.”

There is an interview with Richard Feynman where he talks about analogies that he makes regarding subatomic particles. He says, “Okay, we use the word ‘spin.’ The thing is not really spinning. It's not like a spinning top.” But all those kinds of analogies really help us make sense of it, and not just in a way that allows us to explain concepts—they actually lead to conclusions. They lead to inferences, and those inferences actually make sense.

It's kind of remarkable that we are so primed to interact with the physical world, so primed to have physical intelligence, that you can use it in everyday speech by saying that a company has a lot of momentum, and you can use it when advancing fundamental theoretical physics.

That's kind of remarkable. I don't know if LLMs can do that. Maybe they can, but I think that really understanding physical interactions, causal structures, and all that kind of stuff is something kind of special about that, and it's clearly something that people get a lot of mileage out of.

Patrick O'Shaughnessy

I'd love to talk about the role of researchers and the actual people doing the research. In the LLM world, it's fairly shocking how few people on a global scale are responsible for basically all the progress in LLMs. Someone like Ilya Sutskever is an example. What is that like in robotics? How many people in the world are truly impacting this trajectory?

And then I want to ask about what good research means.

Sergey Levine

I think those kinds of questions are often very hard to answer about science, because we sometimes have a tendency, especially when we look at history, to underline particular milestones. Certainly in machine learning, this is the case. You can say, "Okay, AlexNet was a big step forward. This was a big step forward." That's true.

But I think it's also important to remember that these advances happen because lots of people are trying lots of things, and even some of the failures are actually very instructive. I complained before, in a low-key way, about the controversy around end-to-end robotic learning, but I don't know if robotic learning would have advanced the same way if it were not for the controversy, so to speak.

It is true that you can look through the list of successes and mark down that these folks have had a history of repeatedly hitting home runs. But I think, in reality, in the scientific community, it's not just the home runs that are responsible for progress. Even some of the failures and some of the bad ideas are very instructive in pushing toward the good ideas.

Patrick O'Shaughnessy

Yeah, it's fascinating to think about. The example you gave before is so interesting, where the research insight was, "Just give it some coaching and it gets better." It seems like that sort of insight can be very powerful and high-leverage, which makes me wonder: What have you learned about what makes for a great researcher?

Sergey Levine

Research is definitely different from engineering because, in research, the important thing is to get to an answer to a question, which often requires cutting some corners. One of the most delicate decisions in research is when to try new things and when to stick with what you're already trying. That's very delicate. It's very hard to figure that out.

If you get it wrong and you don't stick with something for long enough, you might be right there. You might be about to get to the answer, and then you stop just short of it. That's terrible. Or you could get stuck hammering against something that's never going to give way for years.

So deciding when to turn a little bit and look this way and that, to open yourself up to more opportunities, versus when you should keep hammering on the thing because you're about to get the solution—that's often the most important decision. Some people have an instinct for getting that right, and that counts for a lot.

Patrick O'Shaughnessy

You've obviously been in and around, and are one of, great researchers. What are these people like as people? How do they tend to be distinctive from the average person?

Sergey Levine

I think they're just the same. I'm thinking about the people that I deeply respect, who are really good at this stuff, and I have a very hard time thinking of a single set of personality traits. The one constant is that there's no constant, basically.

15. Tracking Progress Through Research Papers

There might be a commonality in that, to do effective science, you have to be very passionate about it. But even that passion can come from many different places. I've worked with people who were remarkably effective and were driven purely by the desire for novelty. They don't give a damn about what their technology does. They don't give a damn about whether it's useful. They just want cool new ideas.

I've also worked with other people who really want to solve a particular thing. They're just as happy building stuff, testing out experiments, or hammering away at things—whatever it takes. All those types can be very effective.

Patrick O'Shaughnessy

You mentioned the distinction between research and engineering, which also makes me think of manufacturing. Elon would be fond of saying that the factory is the product. The hardest part of this whole equation is actually the scale-up of whatever this thing ends up looking like—making 100 million of those.

How do you think about that part of the equation, or is it too remote at this stage to spend too much of your time on?

Sergey Levine

No, I think it's an important part of the equation. I'm not sure it's the part of the equation that we most need to figure out right now, but it's certainly part of it.

As you might have guessed from my answers to the other questions, I prefer to think about this by figuring out the hard part and then enabling a lot of experimentation on the other parts, right? So, yes, making a robot at scale is difficult. Making a robot at scale is even more difficult if you don't know what kind of software is going to run it afterward and you're not even sure whether it's the right kind of robot.

I think one of the really valuable things we can get out of general-purpose AI tools like robotics foundation models is the ability to get a lot of the other stuff figured out, so that at least some of the uncertainty goes away. When you scale things up, you have some confidence that this is really going to work.

Patrick O'Shaughnessy

A lot of people who listen to this are entrepreneurs and people who run companies. A very popular question has become: How should a traditional company begin to think about using LLMs or preparing itself for the ongoing improvement of these models? How would you answer the same question for robotics?

Sergey Levine

It's a very good question. It's also a very difficult question because the technology is changing so rapidly. I want to illustrate why this question is difficult with an example.

Here's a particular uncertainty about the technology. It's going to be a little bit specific, but it's an example: Will the robots rely more on demonstrations or on reinforcement learning from autonomous data? We're working on both of those things, and they're clearly both important. But how somebody should prepare for the technology will be pretty different if they're expecting that they need lots of teleoperation to produce lots of demonstrations and a little bit of autonomous experience versus the opposite—a tiny number of demonstrations and huge amounts of autonomous experience.

16. The Next Step: Mid-Level Reasoning

Is it 90/10 or 10/90? That's something we're hopefully going to learn about over the next few years, but it does change the correct approach pretty dramatically. So that's a case study of how changes in technology will dramatically alter this.

Patrick O'Shaughnessy

From a business standpoint, is the right way to think about it just to get really clear on the economics of the labor in your business or something? I'm curious how you think about the way that this will change the nature of labor itself.

Sergey Levine

I think coding tools are a really nice example to look at as a template for how this might work. It's not like coding tools came on the scene and suddenly we don't need software engineers anymore. It's that the coding tools increase the productivity of individual software engineers.

There's some amount of work that needs to be done to make sure that people are able to use them. There's some amount of technology development that needs to be done to make them useful for the appropriate use case, and these things are co-evolving. They're also still changing. Coding agents are different from code-completion tools, and so on.

But I think it's a nice template for us to look at to see how AI tools combine with people doing a job, increase their productivity, and also raise new challenges. I think we'll actually see something like that with robotics, too.

A more realistic template is not that the humanoid goes in and the people just leave. It'll be more like there are some aspects of the job that can be done by a robot, some that can be done with a robot working together with a person, some where the person needs to do something special to make the robot more productive, and some where it's the other way around, where the robot does something that makes the human more productive.

It'll be this kind of dance that we've seen with coding tools.

Patrick O'Shaughnessy

Do you have a favorite robot that's not part of what Physical Intelligence is doing?

Sergey Levine

I do really like the Boston Dynamics robot, especially the new version of the Atlas, because it is, in some ways, very human-like and, in some ways, very not human-like.

They made some interesting decisions about how they want more range of motion on the joints, so it can do some pretty cool things. It's also a very agile robot, which is really cool. It makes those awesome demos. So I'm a big fan of that. I'm generally a big fan of everything that Boston Dynamics has done.

Patrick O'Shaughnessy

Should or could anything be read into the fact that Boston Dynamics has been doing very cool demos for a very long time and doesn't actually do anything useful for people or customers?

Sergey Levine

Yeah, it's a fair question.

I think it's also a fair question for lots of robotics companies, to be fair. What I'll say in general terms is that there is a lot of value in demos that serve to illustrate challenges on the road to something useful and productive. Obviously, you can also do a demo without being on the road to something useful and productive, but I think there is value in demos.

I think that demos that are used correctly in service to a mission can provide people with an illustration of what to expect, and they also provide a challenge. You just have to be honest in setting up that challenge.

Patrick O'Shaughnessy

How much do you think about the business endpoints? I think, to this point, Roomba is the best-selling robot of all time in the consumer category, which is kind of surprising. Of course, we might be on the edge of some sort of Cambrian explosion, but how much of your cycles do you spend thinking about, “This is the shape of a product that might result from this, and maybe this is the way we bootstrap our way to all this data”?

Sergey Levine

I certainly spend some time thinking about it. I think it's just something that's very hard to reduce to a very concrete answer right now. But it's not too bad to think about a space of possibilities. A lot of what we're doing when we develop our models, experiment with different tasks, and do demos like the Robot Olympics is prototyping what it looks like when we try to do something real with this, obviously at different degrees of reality, and seeing what goes wrong.

It is something we think about a lot. It's not something that I have even close to a concrete answer to, but there is a space of possibilities, and a lot of what we're actually planning to do in 2026 is also experiment with different things in that space.

Patrick O'Shaughnessy

When you study the history of general-purpose technologies—which this certainly would be a major one if it comes to fruition—you often find this constellation of things happening around that thing that enable it. Obviously, LLMs are a direct complement to what you're doing. Are there any other surprising technology areas or trends that help you do what you do but are different?

Sergey Levine

One interesting thing is that robotics hardware has become dramatically more affordable over the last few years. When I started working in robotics about a decade ago, I worked with a robot called a PR2, which I believe had a cost of about $400,000. When I started my lab at UC Berkeley, I used a robot that was in the ballpark of $30,000. Now, each arm on this thing is maybe a tenth of that. We think they can be even less.

17. The Kindest Thing

That's not due to any one single technology. It actually involves both hardware and software. The low-cost arms that we have here wouldn't be useful in an industrial setting because traditional control methods that rely on a great deal of precision wouldn't be able to use them. So, there's a wide range of different factors—a cluster, a constellation, as you said—that have pushed down the price point of these things. I think that does make it a lot more practical to think about general-purpose robotics today.

Patrick O'Shaughnessy

For people who would want to be fairly technical about following major milestones that are happening in this field, where does that information show up? A lot of it shows up in research papers.

Sergey Levine

Research papers, unfortunately, are not a very accessible source of information because it takes a bit of care to sort through everything and figure out what the signal is and what something really means. Research results are intended for an audience that already understands the starting point from all the past research results, but that's a big one.

I think robotics, and technology in general, is one of those things where the public-facing artifacts, like the demos and videos that somebody might post on social media, are often not very good for providing a sense of the true underlying state of things because they're meant more as a demonstration at the edge of capability. Grounding what the demo really means requires digging deeper.

So, probably research papers are the way to go. Sometimes, even worse than that, you have to actually go talk to the individual people and find out what the inside story really is. Maybe that's not a great situation to be in, but that's kind of how science works.

Patrick O'Shaughnessy

As we look forward to the future in your mission, what feels the most uncertain?

Sergey Levine

I do think the timeline is uncertain. If anything, my sense of the timeline has gotten more optimistic since we started, but it's uncertain because of the nature of the technology. There's a bootstrap challenge: getting to a particular level of usefulness so that robots can be deployed, do useful tasks, and start collecting data from open-world settings at scale. Getting past that activation energy is such a sudden kind of event, so I think there is a lot of uncertainty about the timing of that.

That's exacerbated by the fact that the timeline looks different depending on what kind of technology is deployed. The example I gave before was whether data collection should be through teleoperation, with autonomous systems, or something in between—maybe shared autonomy, maybe this coaching kind of thing. Those all change the picture in terms of how deployments work and how in-the-wild data collection works. Because of that, I do think there's quite a bit of uncertainty about the timing.

Patrick O'Shaughnessy

You're in such an interesting position because you're at the center of research, and lots of different kinds of people are talking to you and asking you questions. What are questions that you're surprised people don't ask you? What are the things that people don't ask you about that they should?

Sergey Levine

I think the question you asked earlier about how somebody should prepare actually has a variant that would be something like, “Okay, if I want to start using autonomous robots for a thing, what should I start setting up?” Should I set up teleoperation? Should I modify my task in some way so that it's more accessible? Should I design new hardware? Maybe I should design new hardware so I can plug your software into it.

I think people make a lot of assumptions about that. For example, one assumption is, “Well, machine learning requires data, so let me just figure out something that will collect data.” That's not often the best assumption because you need the right kind of data. Maybe some data is easy. It's easy to get videos of people doing something, but that doesn't mean it's the right kind of data.

It might be domain-dependent. It might depend on your thesis about the technology that will succeed. I think people make a lot of assumptions about that. Not that I necessarily have a better answer for them even if they ask me, but it's something where there's a big space of possibilities.

Patrick O'Shaughnessy

We talked about these big, uncertain, long-term timelines. What is the very next thing you're trying to solve that's extremely visible?

Sergey Levine

Without giving too much away, what I can say is that a big focus for us right now is better understanding this kind of middle-level reasoning part of the problem. We think that we have a pretty good sense for how to acquire low-level physical behaviors, but getting those low-level physical behaviors to generalize requires bringing to bear a lot of this kind of common-sense knowledge. The representation of that might be really important.

LLMs make certain kinds of representations very convenient. They make it very convenient to basically turn text into other text. But that's not necessarily the best representation for what an embodied system needs to do. Sometimes it needs to think about things more spatially, sometimes semantically, and sometimes in other representations.

Trying to figure out exactly how to structure that kind of internal thinking process might be a very important question. The answer to that question might be different in the world of embodied foundation models than it is in the world of LLMs. That's a concrete thing that we're working on now.

Patrick O'Shaughnessy

Where do you think you fall? If I could somehow get the 100 most informed and active robotics researchers in the room at once and poll them on how certain they are that things will have unlimited capabilities, and how soon that might happen, where do you fall in that distribution?

Sergey Levine

I'm on the optimistic end when it comes to established robotics researchers and on the pessimistic end relative to robotics entrepreneurs.

Patrick O'Shaughnessy

Interesting. I understand the entrepreneur part for sure. They're optimistic by nature. Why are you on the optimistic end of the researcher community?

Sergey Levine

Robotics has a very long history with precious few successes, especially when it comes to robotic AI. Let me put it this way: if we're being honest about it, most robots that are out there doing useful work are still running, and that's because the robotics problem is hard. Maybe that's not our fault. It's just a difficult problem.

Because of that, I do think there's good reason for caution. Maybe we've made a lot of headway on this part of the problem, but there are many other problems that still remain.

Right? Now, part of why I'm optimistic about this is that I have a sense of what has proven tough for me before, and I can see a lot of the puzzle pieces that I'm imagining could be slotted in to address many of those things. But, as my co-founder Carol likes to say, when you've climbed the mountain, only then do you see if there's another mountain after it, right? And in robotics, there's been a lot of experience of lots of mountains.

Patrick O'Shaughnessy

Given that endurance is required, who or what most inspires you?

Sergey Levine

I am actually quite inspired by Boston Dynamics. Again, I think there are a lot of things that we can debate on the technology side, but I think there is a lot of value in repeatedly showing something that people wouldn't have thought possible, even if there are all sorts of caveats and assumptions and so on. Certainly in robotics, whatever we might say about demos and whatnot, I think it's very fair to say that people have revised their thoughts about what's possible from seeing some of that stuff. So, I think that's definitely one.

I think I'm also inspired by organizations that create an atmosphere for experimentation. There are some research labs that have done a very good job of this. I think, actually, OpenAI has historically done a great job of creating an atmosphere where individual researchers can experiment with things and be empowered to see those things through. ChatGPT was basically John Schulman's pet experiment for a while; it wasn't a concerted corporate strategy with lots of spreadsheets and pie charts. It was a pet project.

I think there's something pretty inspiring about organizations that empower people to have pet projects turn into world-changing successes. Certainly, one of the aspirations that I and my co-founders have here at Physical Intelligence is to provide some of that to the best of our ability. It's hard to do. It's very hard to have an organization that has that kind of capability.

Patrick O'Shaughnessy

I feel like Google used to have that one-day-you-can-do-whatever-you-want thing. Is that the spirit of it?

Sergey Levine

I was absolutely shocked when I started working at Google at the level of leverage that I felt I could have. One of the projects that I did with many of my colleagues there in 2015 was colloquially called the arm farm. We took a couple dozen robots, put them in a lab, and had them collect data. That was a very bottom-up thing.

I found out from somebody that they had a warehouse full of robots that nobody was using. I asked Jeff Dean and Vincent Van Houcke if we could stick them in a lab. I was just thinking, “Okay, they're not going to take me seriously.” I had just started; I was a Level 4 research scientist. Jeff was like, “Yeah, let's do it. What do you need?”

I just remember feeling like, wow, I had never in my life thought that I could have that kind of leverage. Obviously, I was very young at the time. I think that's very special. Getting to a place where people can unlock their creativity and have that kind of agency can make for a very remarkable place.

Patrick O'Shaughnessy

My friend Jesse has this great question, which is: for companies that you're not involved with, which one do you most hope succeeds, and why? People used to say Boom a lot because they want to fly places faster. Increasingly, as I've asked this question, people have said BYD because the sheer impact that it might have if you're successful is massive on such a global scale. It's been really fun just to hear about all the ins and outs of how you're thinking about the problem and attacking it. When I do these interviews, I have the same traditional last question for everyone: What is the kindest thing that anyone's ever done for you?

Sergey Levine

It's a tough question to answer because I do think there are many moments in my career where I felt like I got a leg up on something. I think I have the kind of personality where I sometimes don't appreciate things in the moment and only reflect on them afterward. I don't think I have a single answer to that, but probably the 3 moments in my career that stand out. Actually, one of them I'd already mentioned to you, which was the arm farm thing. I'm especially grateful to Jeff and Vincent for being willing to take that bet on me and my colleagues.

There are a couple other moments. Certainly, when I started my postdoc with Pieter Abbeel at Berkeley, I had zero robotics experience. I had done virtual character animation and computer graphics, right? I felt like that was sort of a bet on my potential, more so than my actual accomplishments. There was another moment even earlier that maybe is even more minor, where, when I was in college, I got an internship at NVIDIA that really got me to experience some cool stuff when I was just a sophomore. I think the hiring manager for that also took a bet on me.

I think these kinds of things really matter in a person's career. Maybe at the moment I should have been more grateful, but certainly in hindsight, it's something that made a big difference. Hopefully I can make that difference in other people's careers as well.

Patrick O'Shaughnessy

Well, I've learned so much from you and your co-founders, and so much today. Thank you so much for your time.

Sergey Levine

Thank you.