[BidClub_]
The a16z Show · · 42 min

Fei-Fei Li is Solving the Hardest Problem in Robotics | World Labs with a16z

Martin CasadoFei-Fei LiYunzhu Li

YouTube
TL;DR
  • World Labs is extending its spatial-intelligence thesis into robotics by bringing SpAItial, initially a Marble customer, into the company rather than becoming a robot manufacturer. The combined stack pairs World Labs’ generative modeling and 3D reconstruction with SpAItial’s robotics, simulation, and hardware expertise. Yunzhu Li’s north star is blunt: “I want the robot to work.”
  • The key bottleneck is the absence of scalable robotics data and evaluation. Unlike language models, robots cannot harvest abundant internet data; physical testing is slow, costly, and dangerous because “atoms have to move through space.” SpAItial’s real-to-sim-to-real pipeline aims to replace much of the data and evaluation work with aligned digital environments.
  • World Labs argues that consistent world models offer something video-only approaches still struggle to guarantee. A useful environment must remain consistent across space, time, viewpoints, and interactions: if a robot pushes an object and it “just magically disappears,” the prediction supplies a poor learning signal. Marble generates geometrically consistent worlds from text or images, represented as Gaussian splats or meshes.
  • Simulation and real-world data are complementary stages of a flywheel, not competing doctrines. Early systems may lean more heavily on physics, geometry, and randomization; accumulated customer and robot data can progressively move modeling toward learned dynamics. Fei-Fei Li’s key distinction is that simulation enables “counterfactual reasoning” about events that have not happened, cannot happen, or lack enough real-world data.
  • Evaluation may be the platform’s sharpest near-term wedge because iteration speed governs robotics development. SpAItial wants to distinguish a 90% checkpoint from a 92% checkpoint, or measure 95% versus 99.9% reliability, without repeating every trial physically. The pitch is “scalable, safe, and much faster evaluations” whose results remain aligned with real-world performance.
  • Commercial deployment should advance from structured factories to semi-structured warehouses, hotels, and restaurants before reaching homes. Robustness comes from sufficient coverage of scenarios, and controlled environments make that coverage tractable; fully unstructured homes remain the “grand challenge.” Martin argues that this favors specialized embodiments over prematurely general humanoids, while SpAItial remains model- and embodiment-agnostic.
  • Human-level robotic efficiency is not presented as a five-year inevitability. Yunzhu expects it to take “a very long time” because a reliable robot is an integrated system spanning hardware, software, the robot’s brain, dynamics, and details such as fingertip friction; Martin notes that even language models do not match a roughly 30-watt human brain. The nearer two-year objective is measured: prove value in a small number of verticals and turn those customers into “lighthouse examples.”
Digest · the substance, structured for research

1. World Labs is making spatial intelligence actionable

  • Fei-Fei defines World Labs as a two-year-old frontier-model lab pursuing “AI that has the ability to generate, understand, reason with, and interact with spaces,” physical or virtual. Large world models are the means; spatial intelligence is the goal.

  • Acting within spaces already includes VFX, gaming, and design—not just robotics. Her broader thesis is a “multiverse” in which builders can create and operate across different spaces, while physical action remains one of AI’s most consequential destinations.

  • The combination began through product pull: after World Labs released Marble last winter, around November or December, SpAItial signed up without Fei-Fei realizing it was Yunzhu’s company. Fei-Fei also says World Labs was already receiving inbound interest from early-stage robotics companies and practical downstream use cases, though it could not yet serve them.

  • Marble accepts text, one image, or several images and generates a geometrically consistent world represented through 3D geometry such as Gaussian splats or meshes. SpAItial contributes full-stack robotics and simulation; World Labs supplies generative modeling, computer vision, and sparse 3D reconstruction to make SpAItial’s heavier real-world capture process more efficient.

2. A robotics foundation model must understand actions and consequences

  • SpAItial’s real-to-sim-to-real pipeline reconstructs appearance, geometry, and dynamics—how an environment changes when an action is applied. The desired alignment is that behavior observed digitally is also likely to occur physically, allowing simulated data to support both training and evaluation.

  • A robotics foundation model would likely be an omni-model spanning frames, text, images, depth, and actions. Yunzhu’s distinction is architectural: actions as inputs create a forward simulator predicting the next world state; actions as outputs create a policy selecting what moves toward a goal.

  • Martin contrasts this with the popular video-model approach. Yunzhu’s answer is consistency “over space, over time, over different viewpoints, and over different types of interactions”; the eventual system may sit between physics and learning, then improve as executed policies return fresh real-world data.

3. Simulation supplies the counterfactuals real data cannot

  • Martin preserves the central objection, citing Sergey Levine’s view that simulation will eventually diverge from reality and real-world collection remains critical. Yunzhu says the positions do not conflict: start with stronger physical structure, ingest real outcomes, and shift toward learned modeling as the flywheel accumulates data.

  • Fei-Fei’s philosophical case is that humans constantly simulate because reality cannot provide every relevant experience. Martin uses planning for World Cup games as an analogy; Fei-Fei calls the resulting capability “counterfactual reasoning.” She cites Waymo’s stated use of billions of simulation hours and adds that cars are among the simplest robots.

  • Fidelity need not mean reproducing every snowflake or bush. Quadrupeds and bipeds can traverse snow and vegetation without precisely simulating each element; the model must capture “the essential structure of the problem,” then randomize enough variation to transfer reliably.

  • Simulation offers reliability through systematic coverage of lighting, friction, geometry, objects, and physical parameters. It also offers efficiency: teleoperation can collect data more slowly than a human performs the task, while customers may need faster-than-human robots—and simply accelerating a policy fails because “gravity doesn’t change.”

4. Faster evaluation is the immediate platform wedge

  • Robotics evaluation asks how reliably a checkpoint performs—95% or 99.9%—and how much wall-clock time is needed to distinguish 90% from 92%. Physical trials make iteration “multiple orders of magnitude slower” than language-model development, while adding danger and cost.

  • When digital and physical results are aligned, a checkpoint that performs better in simulation is highly likely to perform better in reality. Training gains the same controllability: teams can specify which state distributions were covered and where they can reasonably expect the robot to work.

  • This is infrastructure for worlds in which other companies’ robots learn, not a plan to manufacture robots. It supports fixed arms, bimanual robots, mobile manipulators, grippers, and other embodiments, while generated data can train models from scratch or post-train vision-language-action and world-action models.

5. Commercial rollout starts where the world is constrained

  • Robotics has historically moved from fully structured factories into semi-structured warehouses, restaurants, and hotels, then toward unstructured homes. Yunzhu favors that progression because robustness requires scenario coverage; in a survey of roughly 1,000 desired robot tasks, about one-third involved cleaning.

  • Martin argues that humanoid anatomy evolved for general survival in unstructured environments, not optimal performance at any single task. A specialized body may solve a narrower commercial problem better, making SpAItial’s embodiment-agnostic infrastructure more useful than a single humanoid bet.

  • Asked whether robots can reach human power efficiency in five years—or ever—Yunzhu answers that it will take “a very long time.” Progress has repeatedly outrun his expectations, but reliability depends on the whole system, down to finger friction; Fei-Fei calls the required posture “measured optimism.”

  • Integration will be deliberate: the teams are “not rushing to blend” into a “full salad bowl,” though SpAItial already uses Marble internally and collaboration has begun around simulation and action-conditioned models. Two-year success means validated automation value in a few important verticals, producing lighthouse customers from which the business can scale. Current clients are close to deployment and are targeting practical tasks with dozens or hundreds of candidate situations to automate.

Fei-Fei Li

We’re building the next frontier of AI, which is what we call spatial intelligence.

Yunzhu Li

At SpAItial, we are developing what we call a real-to-sim-to-real pipeline. We can replace all the data and all the evaluation we need in the real environment by using data that we can generate at scale in our digital world.

Fei-Fei Li

Think about human intelligence. We do a lot of simulation in our head. Why? There’s a very important role simulation plays that real-world data doesn’t play, which is counterfactual reasoning.

Yunzhu Li

What we are building is a consistent world, consistent over space, over time, over different viewpoints, and over different types of interactions. My north star is: I want the robot to work.

Fei-Fei Li

The world we live in can be a multiverse, and we can create technology to allow people—builders, developers—to act within different spaces.

Martin Casado

Do you believe we’ll ever be able to build robots that have the power efficiency of a human being? How far away are we from this? Is this 5 years away, or is this never?

All right. Well, it’s great to have you both here. Fei, for the listeners who may not have the background, maybe you can give an overview of what World Labs does.

Fei-Fei Li

World Labs is a 2-year-old startup. I think we should just recognize that it’s a frontier model lab. We’re building the next frontier of AI, which is what we call spatial intelligence.

Spatial intelligence is about creating AI that has the ability to generate, understand, reason with, and interact with spaces, whether they’re physical or virtual. Of course, a means to an end toward spatial intelligence is building large world models. That’s what World Labs is mostly focused on.

Martin Casado

You’ve been saying this since the very beginning: the machine’s ability to perceive and reason about spaces and act on spaces. But I always had the assumption that acting on spaces was some long-distant-future thing. Now you’re acquiring a robotics company, so maybe talk a little bit about the timeliness of this and the intentions.

Fei-Fei Li

First of all, it doesn’t just take robotics to act within spaces or to interact. Look at the creative field, whether it’s VFX, gaming, or design. There are many use cases where you can create and act within virtual spaces. World Labs’ thesis has always been that the world we live in can be a multiverse, and we can create technology to allow people—builders, developers—to act within different spaces.

Having said that, the ability to act within physical space is one of the most exciting and profoundly important capabilities of the future AI world. Robotics is very much that. World Labs has always believed that robotics is an important application, as well as a use case, of spatial intelligence and world modeling.

So, inviting the SpAItial team to join World Labs is part of our long-term vision and mission. We’ve always been committed to that.

Martin Casado

Amazing. Yunzhu, you’re the co-founder of SpAItial. Maybe provide everyone with a quick overview of your background and what SpAItial does.

Yunzhu Li

I’m Yunzhu. I’m currently a co-founder of SpAItial and also an assistant professor at Columbia University.

Martin Casado

Wow.

Yunzhu Li

My research started with my PhD at MIT and then a postdoc with Fei-Fei at Stanford University.

Martin Casado

That’s great.

Fei-Fei Li

The world is small.

Martin Casado

It is.

Yunzhu Li

Throughout my career, my goal has been very simple: trying to help robots better perceive and interact with the physical world. I’m a very practical person. I want my robot to work in real physical environments.

For SpAItial, the unique opportunity we see is that there have been a lot of bottlenecks faced by the development of general-purpose robots, especially around training and evaluations. We’re developing what we call a real-to-sim-to-real pipeline.

We want to map real environments into the digital world that has the best alignment with the real environments. By alignment, we mean that whatever happens in the digital world is also going to happen in the real environments, such that we can replace all the data and all the evaluation we need in the real environment by using data that we can generate at scale in our digital world.

That’s how everything started at SpAItial. We put together a very, very strong team around robotics, robot learning, simulation, and rendering, trying to build this real-to-sim-to-real stack to solve some of the key bottlenecks.

Martin Casado

It’s amazing that you two work together.

Fei-Fei Li

There’s a funny story here, because you would think that, because we worked together—he was my amazing postdoc—we’d been talking about this World Labs integration for a long time. It’s actually not true. They came into World Labs as a customer.

Martin Casado

Really?

Fei-Fei Li

When we released the 1st version of our generative model, called Marble, last winter, around November or December, SpAItial just signed up.

Martin Casado

No kidding—as a customer.

Yunzhu Li

Yes.

Fei-Fei Li

I didn’t even know what it was. Then I realized this was Yunzhu’s company. I called Yunzhu. I was like, “Wow, this is your company.” Then we realized there was so much synergy.

Martin Casado

Maybe Fei could just quickly describe what Marble is.

Fei-Fei Li

Marble is the codename for the base model that World Labs has been training and iterating on. The fundamental capability of Marble right now, which is publicly released, is to take a prompt—it can be an image, a few images, or text—and turn that into a geometrically consistent world that can be represented in 3D geometry, whether it’s Gaussian splats or a mesh.

Martin Casado

The SpAItial team is trying to solve this extremely difficult problem in robotics, which is the lack of data—the lack of data in training and the lack of data in evaluation. This is very, very different from language models, where data is abundant on the internet.

Fei-Fei Li

In order for robotics to work, we have to somehow unlock the power of scaling laws. But where does that come from? This is a profound problem that everybody is battling with in robotics.

Martin Casado

It’d actually be great to talk about this synergy. You have put together a very talented team, and I’d like to understand to what extent there’s overlap and to what extent this is an extension. Maybe talk a little bit about that.

Yunzhu Li

The TL;DR is that it’s very complementary, with a shared mission. I’m 1 of the 3 technical co-founders. The other 2 are Changi Jan, another Columbia professor who has been a world-class technologist in simulation, and Sunonni, a phenomenal engineering leader.

Changi Jan has a background in VFX. He worked at Weta, he worked at Tencent, and he’s been an entrepreneur. Sunonni was also in a startup that was acquired by Amazon many years ago. He worked in many different tech stacks in the computer vision field at Amazon.

Fei-Fei Li

When we started talking more seriously, I recognized that a couple of things SpAItial has, from a talent point of view, are extremely complementary to World Labs.

One is obviously incredible thought leadership and just technical prowess in robotics—from hardware and full-stack robotics to modeling. Even when Yunzhu was my postdoc at Stanford, he already had his faculty offer, so he was there for only 1 year. I wanted him for more than 1 year, but he had to go become—have a real job.

He was a full-stack researcher in robotics, from modeling to hardware. And of course, Yunzhu and his students at SpAItial were that pool of talent World Labs hadn’t had yet.

Then, on the Changi Jan side, there’s just incredible simulation capability. He’s such a senior researcher and technologist in simulation, and what World Labs is doing is very much interfacing with the world of simulation.

I think what they don’t have, obviously, is the generative model side, as well as the computer vision and 3D reconstruction side. We’re also very strong at World Labs. That’s technology SpAItial needs. So, these 2 sides come together and make it much more complete.

Martin Casado

This is an extension and a complement to get into robotics. Having been in your situation—which is deciding when to sell a company—it would be great to hear from you on how you think about joining World Labs, the fit there, and why you made the decision to do it.

Yunzhu Li

At the very beginning, we were deciding, “Okay, we want to just keep going.” But after chatting with Fei and seeing all the synergies that were happening, it just made perfect sense for the forces to join each other.

In essence, as SpAItial, what we’ve been doing is real-to-sim-to-real: reconstructing the environment. We capture the appearance of the environment, the geometry of the environments, and also the dynamics of the environment—meaning how the environment is going to change when you apply actions.

This reconstruction right now is still a little bit on the heavier side, and what World Labs has been doing involves a lot of profound capabilities around sparse reconstruction and generation. So, we see a lot of opportunities to leverage Marble and other capabilities at World Labs in order to do very efficient reconstruction and modeling of the environments.

Martin Casado

Can we expect a foundation model for robotics from World Labs?

Fei-Fei Li

World Labs is building a foundation model. As you know, Martin, we’re building a base model, and as the technology has been evolving, some of the most exciting base models are omni-models, right? They take multimodal input and have multimodal outputs.

Yunzhu Li

And what is a foundation model for robotics? It’s very likely going to involve actions.

Martin Casado

It’s very likely going to involve the output of actions in addition to the state of the world, and we’re definitely not ruling this out.

Fei-Fei Li

Yeah. Great.

Yunzhu Li

For example, a foundation model essentially needs to be a multimodal model. It has to take into account frames, text, images, depth, and different kinds of modalities, and action is a very, very important part of those modalities.

If you think about frames and actions as inputs, that is essentially a forward simulator that is going to predict how the environment is going to change when you apply a specific action. When the action is the output, this is essentially a policy model that is trying to predict, given a specific goal, what action you should take in the real environment to get closer to that goal.

These kinds of omni models can benefit a lot and provide a huge amount of value for the robotics community in trying to understand how to model environments and, at the same time, how to act in those environments. This can also act as a backbone that you can fine-tune for specific robotic applications, making sure it really lives up to the reliability and efficiency expected by clients.

Martin Casado

You know, if you don’t mind a kind of lay-investor question, I see a lot of robotics companies, and a very popular approach right now is, “We’ll use a video model.” That’s the predominant method. This is 3D and simulation, which is a very different approach, so maybe you could contrast the popular approach of using video only with what the ambition here is.

Yunzhu Li

In order to create worlds where the robot can learn, as I mentioned, you have to capture the essential structure of the problem. One of the most important requirements for those worlds is consistency. That’s where I see very strong synergies with Marble, because what we are building is a consistent world—consistent over space, over time, over different viewpoints, and over different types of interactions.

The generated world from Marble also provides an infrastructure component of that entire world that we believe is necessary for the robot to learn. Imagine if a robot pushes an object forward and the object just magically disappears, which has been a problem with many existing video-prediction models. This won’t provide a good enough signal for the robot to know what the right thing to do is.

Obviously, there has been a lot of investigation into building better and stronger video models. We see a way that some of the infrastructure we’re building can provide initial momentum for going through this data flywheel: moving from more simulation-driven models into robot-policy models, which execute in the real environment and collect new data.

That data will come back in, and the model doesn’t necessarily have to be physics-only or learning-only, but somewhere in the middle. It will be able to capture the essential structure of the problem while also scaling and becoming better and better as you accumulate more data.

Martin Casado

You know, I’ve worked with you very closely for a while, and you’ve always had this north star that has driven all of this. You’ve articulated it variously as 3D and in a number of other ways. I’m just wondering: is there also a similar philosophical north star for you, or are you more pragmatic, like I am—build the system, do the thing?

Yunzhu Li

My north star is to make robots work.

Martin Casado

Amazing. Yeah.

Yunzhu Li

In the real environment. I’m a very practical person. I want the robot to work. One interesting thing that’s actually coming from my collaborations with Fei-Fei during my postdoc is that we’re building this kind of benchmark. We actually sent out surveys asking the general public what they want their robots to do for them.

Among the 1,000 tasks we collected, 1/3 of the tasks are about cleaning. People just don’t like to do those dull and dirty tasks. Those are the scenarios where we really want to make sure we have robotic solutions to deal with them.

Fei-Fei Li

One thing I really like about SpAItial, Martin, especially continuing your question, is that there are a lot of robotics companies building models and all that. One thing I truly like about SpAItial is that Yunzhu and his co-founders have such an incredibly pragmatic approach to robotics.

They especially come from academia—well, one of them doesn’t, but the other two come from academia. Their first instinct, though, is to work with design partners and customers in real industry, whether it’s industry labs, warehouses, or electronics assembly. That is such a refreshing way of approaching robotics, and that really made me very excited to work with them.

Martin Casado

Maybe this is for you, but this is just personal curiosity. It seems to me that, for robotics, you have to be pretty exact—not perfect, but pretty close. But for the creative use cases that World Labs has done a lot of, you kind of don’t need to be, because sometimes being wrong is stylistic or intentional or whatever.

From a technical perspective, what is the challenge of reconciling these two things? Or do they never get reconciled? Will there always be two points in the design space?

Yunzhu Li

They will be reconciled in the long term, of course. Modeling the environments doesn’t have to be perfect. The model doesn’t have to be perfect in robotics.

Martin Casado

And, by the way, is there a more formal way to say that? This is pure curiosity, but what does it mean not to be perfect? It has to be pretty close.

Yunzhu Li

Let me put it this way. For the development of all different kinds of robotic applications, models have been a very important cornerstone. If you look at all the existing robotic applications—planes, drones, Roomba, quadruped robots, or even bipedal robots—models have been the way for them to actually work and be able to transfer from simulation to the real environment.

Martin Casado

I see.

Yunzhu Li

But if you look at locomotion robots, like quadruped robots and bipedal robots, they can walk on snow and they can walk on bushes. You don’t need to have a simulator that can simulate all the bushes and snow very precisely. You need to have a simulation that captures the essential structure of the problem and does a whole range of different kinds of randomizations inside the digital environment.

That is what we’re aiming for. Basically, with SpAItial and together with World Labs, we’re trying to investigate what level of fidelity we need to model the massive worlds, in addition to the robots, so we’ll be able to transfer robotic systems trained in the simulated environment—the digital worlds—back into real scenarios.

Martin Casado

As an investor, I’ve heard other researchers, like Sergey Levine, say that simulation will always eventually deviate from the physical world, and real-world data collection is absolutely critical. Maybe talk a little bit about the viability of this approach, where simulation is a cornerstone, as opposed to some other approach.

Yunzhu Li

They don’t contradict each other. If you think about simulation, simulation is essentially trying to predict how the environment is going to change when you apply actions. This is essentially a model of the world. It doesn’t necessarily have to be pure physics; it can be a combination of physics and learning.

We are collecting real-world data, and we will be using that real-world data. It’s just that it comes in at different stages of the data flywheel. Maybe at the very beginning we place a stronger emphasis on physics to make sure we have the right consistency and the right structure for us to learn the world and train the robot policies.

But as we accumulate more and more data, both through data collection and through collaboration with our clients, we’ll have data that moves us toward more learning-based modeling of the environments. This kind of transition and data flow is really an enabling factor for getting the best of both physics, geometry, and consistency, as well as all the power and magic from data and compute.

Fei-Fei Li

I want to add to this and be slightly philosophical. There isn’t a binary choice between simulation and no simulation. All of this comes together to make robotics work.

Think about human intelligence. We do a lot of simulation in our heads. There’s a very important role simulation plays that real-world data doesn’t, which is counterfactual reasoning. You play out events that haven’t happened, or cannot happen, or for which you don’t have enough data to make them happen in the real world. While you play them out, you learn how to act in those situations. Humans do this all the time.

Martin Casado

I know you were at World Cups.

Fei-Fei Li

I was at the World Cup.

Martin Casado

Congratulations to Spain for winning. I’m sure that, in the planning of every game, there is simulation, whether it’s digital, on a whiteboard, or whatever. The role simulation plays is counterfactual reasoning, and that’s really important in robotics because we simply cannot possibly have enough real-world data for that.

Here’s a real-life example: the self-driving-car industry. Waymo has officially said they use billions of hours of simulation.

Yunzhu Li

And actually, Waymo is more simulation-heavy than just real-world-data-heavy.

Fei-Fei Li

So these are real examples, and, as you know, autonomous cars are the simplest kind of robots.

Martin Casado

Yeah.

Yunzhu Li

Yeah. So clearly, simulation plays a huge role in robotic learning. I also want to add to that. More specifically, simulation can provide 2 levels of benefits. The first one is reliability, and the second one is efficiency.

For reliability, if you're thinking about a robotic system working reliably in real environments, you need data to provide systematic coverage of all the state space and the variations that robots might encounter. That's how you can learn how to make the system robust. With simulation, you can do systematic randomizations and control the variations in lighting, friction, geometries, object types, and all different kinds of physical parameters to make sure you have sufficient coverage of the state space. This is what can give robotic systems reliability.

The second is efficiency. Right now, many people are doing teleoperation, and if you look at many of the teleoperation devices, imagine all the actual exoskeletons you're using: you're collecting data at a speed that's actually slower than a human doing the task.

Martin Casado

But for many of our clients, human speed is not good enough for them. They want faster-than-human speeds.

Yunzhu Li

So for a robot to move faster, it's not as simple as just driving the robot faster, because gravity doesn't change. But in simulation, you can systematically speed up the robot's behaviors to train the robot such that it considers all the dynamics changes of the environments. This is what can give our clients efficiency.

For both reliability and efficiency, there are some very unique values that simulation can provide.

Martin Casado

You've talked about the technology and the platform and what it does. Maybe talk about the specific use cases people use it for.

Yunzhu Li

There are essentially 2 specific use cases, especially around both training and evaluations.

Martin Casado

Okay.

Yunzhu Li

Starting with evaluations, evaluation is something people often overlook in robotics. But if you're tuning robotic models, you have to know how well they work, and that is the only source of information for you to iterate.

Martin Casado

A lot of AI people really understand what evals are and use them all the time. For non-AI people, it often means something a little different, so maybe it's even worth describing specifically what you mean by evaluation.

Yunzhu Li

Okay. What I mean by evaluation is that you'll be able to understand, for a specific checkpoint, how well it performs. Does it perform, for example, 95% of the time or 99.9% of the time?

The key criterion people use in industry is: How long does it take? How much wall-clock time does it take for you to distinguish between a checkpoint that's 90% accurate and a checkpoint that's 92% accurate? If you only do that in the real environment, it just takes so long for you to make that distinction.

If you really think about the robotic evaluations people are doing right now in real environments, the iteration speed is multiple orders of magnitude slower than the iteration speed of those language models.

Martin Casado

Yeah.

Yunzhu Li

Not only are robotic tasks very varied and very diverse—

Martin Casado

Oh, yeah, because you actually have to do the thing. The atoms have to move through space. [laughter]

Yunzhu Li

Exactly.

Martin Casado

The laws of physics have to be obeyed. Have you watched those robotics videos? Every video has 10× or 8× speed-up because it moves so slowly.

Yunzhu Li

Exactly. So not only is it slow, it's dangerous and costly, but at the same time, the speed is also multiple orders of magnitude slower.

Some of our clients actually need a digital environment that can be used to evaluate their robotic systems, and our digital environment has proven alignment with the real world. Whatever happens in the simulation is also likely to happen in the real environment. If a checkpoint is working better in the simulation, it's also highly likely to work better in the real environment, as we've also discussed in the blog post.

That actually gives our clients very strong confidence in using the data and the signal from the digital environment to do scalable, safe, and much faster evaluations of their robotic systems.

Martin Casado

Great. So that's the evaluation. Then, on the training—

Yunzhu Li

On the training side, as I also mentioned, it's about controllability. You want to control all the different possible variations of states, parameters, lighting, friction, physical parameters, and even object geometry and object types. You want to make sure you have sufficient coverage of all different kinds of scenarios, such that you'll be able to generate informative data for your robots to be robust.

This is just going to be so hard to do in real environments, as we discussed. If you do teleoperation, the speed at which you're collecting data is slow. You're also limited by how many robots you have and how many teleoperation devices you have. There are all different kinds of challenges around the data operations involved.

But in simulation, everything can be controllable, everything can be systematic, and everything can be understood at a level where you know exactly what distribution you have covered. You can develop confidence that, within that distribution, the robot will work. That confidence, efficiency, and scalability are things our clients also value when they use our digital worlds for training robotic systems.

Fei-Fei Li

Here's the crazy thing: even before Yunzhu and I were talking, our inbound customers for Marble were already showing this kind of demand. We just couldn't serve these customers. But we were already getting a lot of phone calls from early-stage robotics companies, all the way to downstream, very pragmatic use cases, and we were seeing these needs.

When people hear you're going into robotics, what they're going to envision is that you're pulling out a 3D printer, making hardware, programming the brain of a robot, sticking it in the robot, and then you've got a robot. I don't think that's what you guys are talking about here. So maybe talk about where this fits in the life cycle of creating a robot, where you will end, and where the rest of the ecosystem will begin.

Yunzhu Li

What we've been building, you can imagine, is infrastructure, with the software around this infrastructure, for people to build worlds such that robots can learn and be evaluated. This infrastructure is naturally model-agnostic and embodiment-agnostic.

Martin Casado

I just want to be very clear, because this is actually a subtle point. It's obvious to you, but it's a subtle point: from what you said, you're not building a robot. You're building an environment in which another company can place its robot brain—

Yunzhu Li

Exactly.

Martin Casado

—to navigate and learn.

Yunzhu Li

Yeah. So our customers right now have all different kinds of robots. Some are using a single robot arm, some are using bimanual robots, some are using a fixed arm, and some are using mobile manipulators. Some are using grippers, and some are using more elaborate versions of end effectors.

Our platform is naturally embodiment-agnostic. We can very easily integrate different kinds of robotic embodiments and put them into the worlds we generate with Marble, such that we'll be able to give those individual robots the capability to do the right tasks at the right levels of reliability and efficiency in real environments.

We're also model-agnostic. We can use the data generated by our worlds to train different models, either from scratch or through post-training of existing foundation models, like vision-language-action models or world-action models. To us, it doesn't matter. We just want to make sure we have the infrastructure and all the worlds such that the robot can work reliably in the real environment.

Martin Casado

You've told me that you think a lot of the predictions around humanoids were a little bit aggressive, and we're likely to see more constrained rollouts, like warehouses or whatever. Can you talk a little bit about that and how it impacts what you're going to be tackling here at World Labs?

Yunzhu Li

That's a very good question. If you look at the progress of robotic applications in real environments, it has always followed the trend of going from fully structured environments into semistructured environments and then into unstructured environments.

Martin Casado

For fully structured environments, what do we mean? That you have knowledge and control over all the configurations within the environments?

Yunzhu Li

Like factories—

Martin Casado

Like factories.

Yunzhu Li

—or car manufacturing. Those have been automated for decades.

Martin Casado

Yeah. Yeah. Yeah.

Yunzhu Li

Then you have semistructured environments, in which you have certain control over the environment—for example, Amazon warehouses, restaurants, and hotels. You have certain control over the environment to make the task easier for your robots, but there are obviously many other objects—for example, clothes—that you don't control.

For unstructured environments, it's your home and my house.

Martin Casado

Those are, I would say, the grand challenge. Especially my house, trust me: 3 dogs. [laughter]

Fei-Fei Li

5-year-old.

Speaker 1

Yes, dogs.

Speaker 2

Exactly.

Yunzhu Li

If you're thinking about where robustness comes from, robustness comes from sufficient coverage of the scenarios that robots might encounter. It's so much easier and more approachable, at least right now, to focus more on semistructured environments before we move on to fully unstructured environments. We will move in that direction; we just want to take a more sustainable and realistic approach toward it.

Martin Casado

I think your point here is that humanoids mimic the human body, and evolution has optimized the human body for unstructured environments. Our fingers and legs are not the best apparatus for doing one thing. For example, if our only goal as a species were to climb trees, we wouldn't have this body necessarily, right? We'd have different kinds of fingers. But what humans ended up evolving into is this body shape that can be very general, but not necessarily the best at everything. That is for survival in unstructured environments.

From a business point of view, and from a pragmatic technology point of view, this unstructured environment and a generalized body are actually the hardest problems to solve. It's not necessarily even the right way to solve the problem. We specialize, so we take a more specialized body to solve a narrower problem. But the challenge for World Labs is to be more body-agnostic, so that its infrastructure can serve different bodies and different semistructured environments.

You know, a common lens to look at exactly this question is an economic lens, right? You compare it to generative LLMs, where they can create prose or code 10,000 times faster than a human being and a bunch cheaper than a human being. So the economic case makes sense because our brains aren't very efficient at that. However, our brains and our bodies are very efficient at 3D navigation—moving through the world or picking things up.

This is just a prediction question, but do you believe we'll ever be able to build robots, at least in the foreseeable future, that have the power efficiency of a human being when it comes to menial tasks? Let's say minimum wage or something like that. How far away are we from this? Is this 5 years, or is this never?

Yunzhu Li

I think it's going to take a very long time. If you really think about robots in the real environment, in the end it will always be a system. Every working robot in the real environment is a system. You need to be very mindful and thoughtful about how the system comes together: the hardware, the software, the brain, and even details such as the friction coefficients of your fingers. There are a lot of things you have to consider to make these things a reality, and it will take iterations.

What I am excited about is that I have always been at the state of the art of robot learning and trying to push the state of the art forward. The state of the art is always moving faster than I expected. What I'm focusing on and trying to investigate right now is very different from what I was focusing on when I started my PhD. This speaks to how fast the whole ecosystem has been evolving and how all the moving pieces are starting to come together to build these robotic systems. But we also have to be calibrated about our predictions. We will see a lot of progress, but to achieve, for example, human-level efficiency and capabilities, it will take longer.

Fei-Fei Li

Martin, the hardest thing in today's AI is to have the right measured optimism. [laughter]

Martin Casado

Right? I mean, even LLMs do not have human-brain efficiency. The human brain operates on 30 watts.

Fei-Fei Li

Yeah, that's true.

Martin Casado

But performance-to-power may be close, right, in narrow tasks like software engineering, like generating an image or software engineering. I don't think we're anywhere close when it comes to robotics.

Does this change how you think about, strategically, the level of ambition that your team can go after? Has it changed that, or is it still very much in line with what you expected to do when you started?

Yunzhu Li

It definitely changed the trajectories in a very profound manner. We see a lot of unlocks in being able to do this whole process—the modeling of environments—much more efficiently and at a much more scalable level, especially in partnership with World Labs.

I also want to add that, if you think about the current state of language models, those are models with incredible capabilities. But you don't just blindly trust them to book your flight tickets or make your hotel reservations. Hopefully, there is still a person who reads the output from those language models.

Martin Casado

Yeah.

Yunzhu Li

But that is very different from how people will be using robotic models, because robotic models have to work reliably in the real environment out of the box. We don't even have the data or all the necessary infrastructure around robots for them to work reliably out of the box in real environments.

For that reason, being able to create these digital worlds—scalable digital worlds where the robot can learn and evaluate within them—is going to unlock so much more potential by replacing costly and unsafe data in real environments with data generated from these worlds, allowing robots to learn and evaluate at scale.

Martin Casado

I've seen many of these kinds of integrations. They actually work very well at this stage when they have this much alignment, which is great. But there's always this question of whether you integrate now into what's happening now, or whether you keep things quite separate and provide a long-term trajectory that will be realized in a year-long time frame.

How are you thinking about this, Fei? Is this something that integrates right away, or is this a separate, longer-term effort?

Fei-Fei Li

This is a great question. I think at this point Chang, Sonny, Justin, Ben, and I have been talking about this. We are going to take it thoughtfully. We're not rushing to integrate everything, from codebases to teams, because SpAItial does have a very well-thought-out—and I wouldn't call it completely standalone, but fairly contained—tech stack, as well as its customers and the products it's building.

We're going to take time. We definitely will integrate; we're already starting to talk on the simulation side, as well as on the potential base-model and action-conditioned-model side. They are also using Marble as an internal customer. We will be integrating, but we're not rushing to blend the team into a full salad bowl.

Martin Casado

Yeah. How are you thinking about geographies with SpAItial's move? Is it going to stay in the same place?

Fei-Fei Li

Yunzhu is going to move.

Martin Casado

Oh, well, welcome here. [laughter] Moving to San Francisco. Florence and the Renaissance. Perfect.

Fei-Fei Li

I think World Labs is officially becoming a bicoastal company, where the headquarters is in San Francisco. I've—you know, I live in Palo Alto. I feel like I'm in a different state. [laughter]

We're actually excited that we're going to have an office in New York that can help us attract talent on the East Coast. We've also been talking about making sure that in both offices we set up the robots, so that we can basically test out and mature our engineering stack and work with robots remotely, because we have to do that for our customers anyway.

Martin Casado

So maybe, just to be very concrete, Fei, let's pencil out what the perfect success case is in 2 years. What product do you have? Who's engaging with it? How do they use it? Just the crisp version—what is it?

Fei-Fei Li

We're very happy that the SpAItial team and World Labs team will have validated customers in a small number of important vertical use cases where our system and infrastructure have proven to be truly beneficial to their automation needs. These customers will become our lighthouse examples to scale our business.

Martin Casado

And how early—let's say someone listening to this is running a robotics company—at what stage do they engage with World Labs? Is it really early on, or is it somewhere in the middle?

Fei-Fei Li

Right now, for our customers, we're building these real-to-sim-to-real pipelines, where the simulation is essentially the world we're going to provide as the training and evaluation grounds. Some customers need only the real-to-sim part. They want to digitalize the tasks they care about and be able to evaluate their robotic systems.

Some customers need this entire real-to-sim-to-real pipeline, so they will be able to have policies running on their hardware. Our platform is designed in a way that is flexible, depending on what our clients need.

At the same time, the clients we are working with are actually pretty close to the deployment stage. Basically, they are working on very practical tasks—tasks that, when replaced with robotic solutions, can create value immediately—and they have at least dozens or hundreds of these kinds of situations they are thinking about automating.

Together with World Labs, we'll be able to develop reliable solutions for those scenarios. As we've already shown, we have a number of scenarios instantiated in our blog post, and we'll be able to further our investigation to see how they can actually solve the key requirements and constraints faced by real-world deployments.

Martin Casado

Great. I want to be very specific about this. Is it ever too late or too early to call World Labs if you’re a robotics company?

Fei-Fei Li

No. We want everybody to call us. We want to learn about your use case.

Martin Casado

Wonderful. If you’re listening to this and you’re anywhere close to a robotics project or robotics company, please contact World Labs.

Yunzhu Li

Yes. Thank you.

Martin Casado

Definitely open for business.

Fei-Fei Li

We are open for business.

Yunzhu Li

Not too early.

Martin Casado

All right. If you’re doing robotics, call World Labs. Thank you both very much for coming.

Fei-Fei Li is Solving the Hardest Problem in Robotics | World Labs with a16z | BidClub