[BidClub_]
No Priors · · 39 min

No Priors Ep. 141 | With Sunday Robotics Co-Founders Tony Zhao and Cheng Chi

Sarah GuoTony ZhaoCheng Chi

YouTube
TL;DR
  • Tony Zhao places robotics “in between the GPT moment and the ChatGPT moment”: the field appears to have a scalable recipe, but has not yet converted scale into a great consumer product. Classical sense-plan-act systems required bespoke interfaces for every task and environment; newer learning methods aim to scale data and models instead of repeatedly rebuilding task-specific systems. His bet is that robotics will follow other AI fields: more relevant data should improve usefulness.

  • Sunday’s central asset is an in-the-wild data engine with almost 10 million long-horizon trajectories, collected by more than 500 people. Cheng Chi’s UMI work replaced lab-bound teleoperation with a GoPro and gripper, while Diffusion Policy stabilized multimodal imitation, ALOHA made dexterous data collection more intuitive, and ACT/action chunking plus transformers helped make bimanual tasks scalable. A three-student collection produced 1,500 espresso-serving clips in roughly two weeks.

  • The apparent simplicity of glove-based collection hides full-stack execution difficulty. Sunday iterates the data device, robot controls, automated filtering and cleaning, calibration, training pipeline, and mechanical design together; the glove has gone from V0 to V5 with around 20 iterations per version. As Zhao recalls, the founders initially feared someone could “just take our glove,” but learned that “things are so much harder than we thought.”

  • Sunday rejects humanoid completeness when simplification can produce a useful robot faster and more cheaply. Its friendly, cartoon-like robot uses three fingers because people often use those fingers together for chores such as grasping handles or opening dishwashers. Perception lets the robot correct inexpensive, compliant, imprecise hardware, supporting a mechanically safe and compliant design while retaining sufficient accuracy.

  • The founders currently see imitation learning as more sample-efficient for manipulation, while reinforcement learning works well for locomotion. Ground contact is comparatively tractable to simulate, but reproducing a deforming hand, transparent cup, orange juice, reflections, and physical properties is extremely difficult and expensive. For manipulation, the behavior—put the hand in front of the cup and close with suitable force—can be easier to demonstrate than the world is to simulate.

  • Commercialization begins with a 2026 in-home beta, with shipping to the masses in 2027 or 2028 possible but explicitly dependent on reliability, capability, safety, and price. Prototypes currently cost $6,000-$20,000; at a few thousand units, Sunday expects material cost likely under $10,000 as costly low-volume cladding shifts to injection molding, implying a selling price around that level. Zhao envisions more than 1 billion home robots within a decade, while Cheng frames one possible future as the marginal cost of labor in homes approaching zero.

  • The founders’ rule for evaluating robotics videos is “make zero assumptions. No Priors.” Verify autonomy, then treat the exact object, person, environment, and demonstrated sequence as the proven capability—not evidence the robot can generalize. Sunday’s evidence includes long-horizon table cleanup, fragile glass handling, zero-shot trials across about six Airbnbs, espresso operation, and sock folding, though Cheng says the team shattered many glasses during development.

Digest · the substance, structured for research

1. Robotics has a scaling recipe, but not yet its ChatGPT product

  • Zhao’s framing: robotics sits “in between the GPT moment and the ChatGPT moment.” Researchers increasingly agree on promising manipulation methods, but nobody yet knows what the jump from robotics’ equivalent of GPT-2 to GPT-3 will yield because the field has only recently obtained data at meaningful scale.

  • Classical robotics moved through human-designed sense-plan-act modules. Every task and environment demanded new interfaces—effectively “for every task, that means a paper”—so researchers and companies repeatedly discarded task-specific engineering instead of accumulating general capability.

  • Tony describes Diffusion Policy as stabilizing imitation learning when demonstrations contain multiple valid responses to the same observation. That made it possible for multiple, sometimes untrained operators to contribute without training diverging or the robot behaving strangely.

2. Better interfaces unlocked dexterity, then wild data unlocked scale

  • ALOHA made teleoperation feel more like playing a video game by greatly reducing the delay between human motion and robot response. Once the demonstrations became smooth and dextrous, transformers—which robotics had struggled to use after years of relying on 3-layer MLPs and ConvNets—began working well.

  • ACT and action chunking predict a trajectory rather than a single millisecond-scale action. Zhao’s intuition is biological: humans perceive, then move for a while without looking again, so chunking produces more consistent motion and better overall performance.

  • Chi’s escape from lab-bound teleoperation was UMI, using a 3D-printed gripper and GoPro to capture paired video and hand motion. Three students carried it into restaurants and gathered roughly 1,500 espresso-serving clips in two weeks, producing an unusually large dataset and an end-to-end policy that served drinks around Stanford in unseen locations.

  • The failure case was equally informative: the policy broke under direct sunlight because its collection period had been rainy. “In order for a robot to work in a sunny environment, it must have seen sunny environments” illustrates why wild-data breadth, not merely trajectory count, governs generalization.

3. Full-stack iteration turns collection scale into execution difficulty

  • Sunday grew from two founders clamping a robot to a desk in Chi’s apartment, to about eight people by late 2024, to around 30-40. Building a product rather than a demo required mechanical engineering, controls, software, AI, supply chain, and operations to optimize one system together. Because no general-purpose home robot exists, the right interfaces and standards are still unknown; the founders say this also makes outside partners difficult to use as their standard of “good” keeps changing.

  • Nearly 10 million wild trajectories now include navigation and extended tasks, not merely isolated cup pickups. More than 500 people using the gloves expose every failure mode; the V5 device followed around 20 iterations per version from V0 through V5, while automated calibration and fault detection protect data quality without requiring a human to inspect every clip.

4. A useful home robot wins by simplifying the humanoid

  • Sunday’s mission is to put a home robot in everyone’s home and remove chores that contribute little to what makes people “intrinsically human,” returning time for family, hobbies, and passions. If robots become cheap, safe, and capable, Zhao envisions more than 1 billion in homes within a decade.

  • Sarah Guo’s pushback—why not simply build a complete human form?—draws the core design answer: simplify wherever usefulness survives. Three fingers capture most benefits of grasping handles or opening dishwashers without multiplying actuator cost to separate fingers that ordinarily move together.

  • Industrial robots must be fast, stiff, and precise because they are “blindly following a trajectory.” Perception changes that constraint: cheap, compliant actuators may be mechanically inaccurate, but the AI system can correct hardware inaccuracies and deliver sufficient household accuracy while remaining “mechanically inherently safe and compliant.”

  • Commercialization remains conditional. A selected-user beta in 2026 will put real robots into homes; results determine whether shipping to the masses happens in 2027 or 2028. Prototypes cost $6,000-$20,000, but at a few thousand units Sunday expects material cost likely under $10,000, chiefly as CNC-machined, hand-painted cladding becomes injection-molded.

5. Manipulation favors imitation where simulation is hardest

  • The founders initially expected glove data to trail perfectly distribution-matched teleoperation. Instead, the form factor elicited more natural, dextrous behavior; after roughly 20 engineering iterations and full-cycle hardware/software work to convert human imagery into robot-like data, they no longer see a meaningful quality gap.

  • Reinforcement learning works well in environments that are easy to simulate, especially locomotion, where the relevant physical model is largely rigid-body dynamics and ground contact. Chi notes that a perfect simulator would make any task possible, so the practical question is which method gets there faster. Manipulation flips the equation: the behavior itself can be easier to capture, while accurately simulating transparent vessels, liquid color, reflections, deforming hands, contact, and force is extraordinarily expensive.

  • Data quality became more important with scale, not less. Hardware failures and uncontrolled wild behavior require continuous monitoring and repeatable cleaning; meanwhile, Sunday began serious research only about three months before the interview because “cute, fancy research ideas” developed under data scarcity might not scale into products.

  • Chi identifies two remaining challenges: finding a robust training recipe at scale and making hardware reliable. The learning team keeps pushing hardware toward its limits, so parts break; having mechanical and learning teams under one roof lets the company feed failures back into design quickly.

6. Demo credibility lives in autonomy, generalization, and sequence length

  • Zhao’s test begins with whether a demo is autonomous or teleoperated. If it shows one cup handed to one person, assume only that exact interaction works; viewers instinctively imagine different cups, people, laundry, and dishes, but should “only index on the things that are demonstrated strongly.”

  • Sequence length matters because every interaction introduces another failure probability. Sunday’s cleanup demo spans mobile manipulation, dumping food waste, loading dishes, and operating the dishwasher, and includes force-sensitive handling of two transparent wine glasses in one hand. Chi says, “We shattered a ton of glasses” during experimentation.

  • For generalization, the team booked around six Airbnbs and attempted tasks zero-shot: collect utensils into a caddy and load plates into dishwashers. The robot received no household-specific demonstrations yet handled reflective silverware, millimeter-sensitive grasps, and even a transparent table—coverage attributed to the diversity captured from more than 500 people in the dataset.

  • Espresso operation and sock folding test fine-grained force control. Conventional teleoperation leaves the operator’s hand effectively “numb,” allowing large forces to be applied without awareness; glove users can feel contact naturally. Sock folding also creates a force-closure loop where a stiff grip can apply effectively unlimited force without visible change, making the glove’s natural force feedback valuable for dexterity.

Tony Zhao

Nobody wants to do their dishes. Nobody wants to do their laundry. People would love to spend more time with their family and loved ones. So what we believe is that if the robot is cheap, safe, and capable, everyone will want our robot. We see a future where we have more than 1 billion of these robots in people's homes within a decade.

Thanks, Memo.

Sarah Guo

Today we're here with Tony Zhao and Cheng Chi, co-founders of Sunday and makers of Memo, the first general home robot. We'll talk about AI and robotics, data collection, building a full-stack robotics company, and a world beyond toil. Welcome, Cheng. Tony, thanks for being here.

Tony Zhao

Thanks for having us.

Sarah Guo

First, I want to ask: where are we here? Classical robotics has not been an area of great optimism over time, or of massive velocity of work, and now people are talking about a foundation model for robotics or a ChatGPT moment. Can you contextualize the state of AI robotics and why we should be excited?

Tony Zhao

I would say we're kind of in between the GPT moment and the ChatGPT moment. In the context of LLMs, what it means is that it seems like we have a recipe that can be scaled, but we haven't scaled up enough to have a great consumer product out of it. That's what I mean: GPT, which is a technology, and ChatGPT, which is a product.

We're seeing that across academia, there's consensus around what the method is for manipulation, but everybody's talking about scaling up. We know there's a sign of life in the algorithms people are picking, but people don't know what will happen if we have more data—like what happened with GPT-2 and GPT-3. We see a clear trend that there's no reason to believe that robotics doesn't follow the trajectory of other AI fields, and that scaling up is going to improve performance.

Sarah Guo

Maybe if you took a step back, what was the process for deploying a robot into the world 10 years ago? What was the set of generalizable AI algorithms, and why was it so slow as a field?

Tony Zhao

Previously, classical robotics had this sense-plan-act modular approach, where a human designed the interface between each of the modules. Those interfaces needed to be designed for each specific task and each specific environment. In academia, that means for every task, there's a paper. You design a task, design an environment, design the interfaces, and then produce engineering work for that specific task.

But once you move on to the next task, you throw away all your code and all your work and start over again. That's also what happened in industry. For each application, people built a very specific software and hardware system around it, but it wasn't really generalizable. It just felt like we were running in loops: we built one system, and then we built the next one, but there was no synergy between them. As a result, progress was somewhat slow.

Sarah Guo

I feel like that's a good segue into some of the amazing research work that you guys have contributed over the last 5 years to the field. Should we start with Diffusion Policy? What was the impact of that?

Tony Zhao

Diffusion Policy is a specific algorithm for a paradigm called imitation learning. That's really the most intuitive way of using machine learning for robotics. You collect paired action-and-observation data of what the robot should do, use that to train a model with supervised learning, and then the robot does the same thing.

The problem is that, in the field, it's known to be very finicky. When I talked to researchers after I started in the field, the researchers themselves—the specific researcher—needed to collect the data so that there was exactly one way to do everything. Otherwise, either the model training would diverge or the robot would behave in some weird way.

A diffusion model really allows us to capture multiple modes of behavior for the same observation while preserving training stability. That really unlocked more scalable training and more scalable data collection.

Sarah Guo

So it doesn't have to be you personally wearing a teleoperation headset in order to make a robot learn?

Tony Zhao

Yep. We can have multiple people, sometimes even untrained people, collecting data, and the result will still be great.

Sarah Guo

Where do ALOHA and ACT play into this?

Tony Zhao

These 2 papers are actually super close to each other. They're 1 or 2 months apart. That's actually how Cheng and I know each other: we were looking at each other's papers, and we met on Twitter, I think, when Cheng was back in Colombia, before ALOHA.

The typical way people collect data is with a teleoperation setup with a VR headset. It turns out to be very unintuitive, and it's hard to collect data that is actually dexterous. ALOHA is a very simple and reproducible setup, so it's very intuitive.

Sarah Guo

Sorry—in terms of most people who haven't worn a teleoperation setup, is it the lag? How should I compare it to playing a video game or something?

Tony Zhao

ALOHA makes it feel more like playing a video game. Normally, it feels disconnected: you're moving in free air, and the robot is moving with some delay. But ALOHA reduces that delay by a lot, and that contributes to the smoothness and how fast humans can react.

Once we get that really dexterous data, what it allows us to do is investigate algorithms that are actually solving difficult problems. In this case, it's through introducing transformers in robotics. There was a long period of time when I think robotics was stuck with 3-layer MLPs and ConvNets, and as you made them deeper, they worked worse. But once you have very strong and dexterous datasets, you can just throw a transformer at them, and it works quite well.

Sarah Guo

Actually, just in terms of the progress of the industry over time, transformers didn't make sense without a certain level of data-collection capability.

Tony Zhao

And also an entire system around it—for example, action chunking, which is predicting a trajectory as opposed to predicting single samples of actions. All these things combined to make dexterous bimanual tasks more scalable.

Sarah Guo

Why is chunking important here? If I think about the analogy to LLMs and text-sequence prediction—

Tony Zhao

I think it just throws your model off if you're trying to force it to react every millisecond. That's not how humans act. We perceive, and we can actually move quite a bit without looking at things again. That turns out to make the motion a lot more consistent and overall performance a lot better.

Sarah Guo

And you discovered that transformers architecturally did apply to robotics. Cheng, you felt then that data collection was still a problem. So enter UMI.

Cheng Chi

After ALOHA and Diffusion Policy, I was super excited about imitation learning, but at the time, both of us were still doing teleoperation, and that just felt super limiting. The problem is that in a teleoperation setup, it takes a PhD student a couple of hours to set it up in a lab. That pretty much restricts data collection to a lab.

But in order for the robot to actually work as a product, it needs to work in the wild, in unseen environments. That requires data to also be collected in the wild. At the time, I was thinking: Is there a way we can collect robotic data without actually using a robot? That forced me to think: What's the most essential part of robotics data?

After Diffusion Policy and ACT, the paradigm is simple: you just need paired observation and action data. In our case, the observation is a video clip; the action is the movement of your hand, plus how your fingers move. I realized that you can get all this information from a GoPro. You can track the movement of the GoPro in space, and you can track the motion of the gripper and the fingers through images as well.

That's why I built this UMI gripper, which was 3D-printed at the time. The project had 3 students, and we just took the grippers everywhere. Every time we went to a restaurant, before the waiter came in, we collected some data. Very quickly, we got 1,500 video clips of this espresso-cup-serving task.

That turned out to be one of the biggest datasets in robotics, simply by 3 people. That's where the power really shines. With that amount of data, we were able to train the first end-to-end model that could generalize to unseen environments. We could push the robot around Stanford's campus—Tony was there as well—and anywhere, the robot could serve you a drink.

Tony Zhao

Yeah, I think that was the moment when I was like, “Hey, maybe we should start a company. This is actually working so well.”

Sarah Guo

I remember just following Cheng, and at times it didn't work well.

Cheng Chi

Yes.

Sarah Guo

I think the only exception I saw was when it was under direct sunlight, right? I think the reason was that, over that whole 2- or 3-week period of data collection—

Cheng Chi

Those 2 weeks were all rainy, so there was no sunlight data. It fails. That also demonstrates the importance of distribution matching. In order for a robot to work in a sunny environment, it must have seen sunny environments during training.

Sarah Guo

Yeah. This is really interesting because I remember when I first met you guys, you spent, I don't know, $200,000 across all of your academic research, and yet the scale of data collection, as translated to model capability, is leading. Right? So it's very interesting that we look at where we are—maybe going back to Tony's point of scaling and massive capital deployment—but that entire paradigm actually wasn't relevant before people realized you should train on all of the internet data, and we just don't have that in robotics. So the entire field is just blocked on having any scale of data that's relevant.

Tony Zhao

Yeah. I think these days there are so many debates about what is even the right way to scale. There are world models, there are simulations, there is teleoperation, and there are all these new ideas. I think this is the area that we really want to innovate in, that we want to differentiate in, that we want to find out something that is both high-quality and scalable.

Sarah Guo

And then you guys decide to start a company pushing this cart around Stanford. Tell me about that decision, and congratulations on the launch and the direction and team you've built.

Tony Zhao

Yeah, it's a very interesting journey. I remember in the beginning it was actually the 2 of us, in Cheng's apartment, on his desk. We clamped a robot there and tried to do some tasks, and it soon became, I think, an 8-person team towards the end of 2024, and now we're at around 30 to 40 people.

We're not the best at everything, right? But starting a company allows us to find people who we really love working with and then bring all the expertise together—from mechanical engineering and supply chain to software engineering and controls—to build a system together that is not a demo but a real product.

Sarah Guo

You've built this amazing team. What are people actually signing up for? What's the mission of Sunday?

Tony Zhao

Yes, it is to put a home robot in everyone's home. I think there are a lot of AI systems trying to make you more efficient during work. But there is not enough AI that actually helps you with all these mundane things that are not creative, that really have nothing to do with what's making us intrinsically human.

What's ideal for people to spend more time on is actually their hobbies and passions, as opposed to spending more time doing chores.

Sarah Guo

So if you guys are going from these amazing research breakthroughs to actually shipping a home robot, that's a product. You have to talk about cost and capability and robustness. What's the design philosophy?

Tony Zhao

As these AI models become more capable and as hardware costs continue to go down, home robots—or all kinds of robots—will be everywhere. So if we start from the most surface level, which is the design of the robot, when we design it, we think about what a robot should look like if it is ubiquitous. You need to see it every single day. What should it look like?

What we end up with is that we really think the robot should have a face. It should have a cute face, and it should be very friendly. So instead of a Terminator doing your dishes, we want the robot to feel like it's out of a cartoon movie.

And then a huge decision is how many arms should the robot have? Should it have 4 arms? Should it have 1 arm? Should it have legs? Should it have 5 fingers, 2 fingers, or 3 fingers? It's a huge space.

Sarah Guo

Why isn't the obvious answer that it should just be a full human?

Tony Zhao

I think the core motivation for us is: how can we build a useful robot as soon as possible? So whenever we see something that we can accelerate with simplification, we'll go simplify that.

One example of that is the hand that we designed, which has 3 fingers. We combine the 3 fingers that we have together. The reasoning there is just that most of the time when we use those fingers, we use them together, whether it's grasping a handle or opening the dishwasher.

So it really doesn't make sense to multiply the cost by 3× to separate it into 3 when we can do 1 with most of the benefits. So this is how we think about the whole robot as well. It's with the constraint that we are building a general-purpose robot that can eventually do all your chores, and we'll simplify everything we possibly can so that the robot can be as low-cost and as easy to repair as possible.

Cheng Chi

Yeah, I just want to add a little bit more to the actuation and mechanical design. Traditionally, most robots are designed for industrial use cases, and the robots are very fast, very stiff, and very precise. The reason is that all the industrial robots are blind, so they're blindly following a trajectory that's programmed by someone.

Sarah Guo

It's not reacting to perception.

Cheng Chi

Correct. But because of the breakthroughs we had in AI, now the robot has eyes. So it can actually correct its own mechanical and hardware inaccuracies. That kind of opened up a new, different space of design.

Sarah Guo

Intuitively, it should be like, I can't tell you exactly what the distance is here on a millimeter scale, but I'm going to get to the cup because I can stop.

Cheng Chi

Yeah, exactly. So that allows us to use these low-cost actuators that are cheap and compliant, but imprecise. But because of the AI algorithms and systems we build, it allows us to build a robot that's mechanically inherently safe and compliant while simultaneously being able to achieve the sufficient accuracy we need for the home tasks.

Sarah Guo

Where are we in that timeline? You said we're between GPT and ChatGPT. And so, when do consumers get ChatGPT, and when will you guys ship something?

Tony Zhao

Yeah. It's actually a really exciting time because we have so many prototypes internally. What we will do next year, 2026, is actually start doing beta programs. We'll have these robots—all kinds of different ones—in people's homes and see how they react to them.

That will be when we learn the most about what people like. Do people want to talk to a robot? Do people want to have the robot maybe teach their kids some new knowledge about the world? This will inform us what the eventual product should look like.

Internally, we just have an extremely high standard for what the minimal consumer product we want to ship is. It needs to be extremely safe. It needs to be extremely capable and low-cost.

Sarah Guo

Do you feel like you know something now that you didn't when you started the company?

Tony Zhao

Absolutely. At the beginning, I would describe it as seeing light at the end of the tunnel. There are 2 axes: there's dexterity, and there's generalization. When we add more data, things work better. What is this company about? It's the cross product of these 2: how can we scale and have both dexterity and generalization?

This is something we're able to show in our generalization demo, which is that we can pick up these very precise forks—actual metallic forks—on ceramic plates with very high success rates. Honestly, this is not something that we thought would work so easily just by having so much more data.

Cheng Chi

Yeah. I actually just want to expand a little bit. The process was long and painful. There were so many issues; scaling up a robotic system is very, very hard. There are mechanical issues, reliability issues, and data-quality issues that come out of it.

In the beginning, I actually thought it was going to be much easier than this. Compared to teleoperation, it's much harder to get this system scaled up, but once it's scaled up, it's very powerful and very repeatable.

Sarah Guo

So it is both harder than you thought it would be to get to here, and you are further than you thought you would be.

Tony Zhao

Yes. And I remember in the beginning we were having this funny conversation. We were like, if we build this, someone can just take our glove and build the same thing. What moat do we have? Are we worried about that?

I think in the beginning, actually, we were a little bit worried because we thought they could probably just replicate it. But as we go along the path, it turns out things are so much harder than we thought they were.

Sarah Guo

Yeah. And when you say scaling up the robotic system, you mean the data-collection-to-training pipeline and the hardware itself?

Cheng Chi

Yeah. For this to work at all, you need the data-collection system. You need the robotics and control systems to be able to deliver the hand to where we want it to go, and you also need the data-filtering pipeline, data-cleaning pipeline, and training pipeline. All these things need to be iterated together.

We've actually gone through several loops of these. It's kind of hard to imagine, without having a full-stack team in-house, how this can even be done. The glove we're using right now is—we call it V5.

For V0 to V5, each version has around 20 iterations.

Sarah Guo

Okay. So 100.

Cheng Chi

Yes. Also, when you make these at scale, right now we have more than 500 people using these gloves in the wild. All the things that could go wrong will go wrong.

Sarah Guo

They did.

Cheng Chi

Yes. For example, how things are assembled. If you don't specify exactly how it should be done, people will assemble it in creative ways. And the creativity doesn't help us here because we really want the data-collection device to be extremely precise.

Sarah Guo

You obviously can’t know everything that’s happening in every company, in academia, and in industry. But from what you know, how would you compare the scale of the training data you have today relative to the industry?

Tony Zhao

At this point, we have almost 10 million trajectories being collected in the wild. Those trajectories aren’t just, “Oh, pick up a cup.” These are long trajectories involving walking, navigation, and long-horizon tasks.

Elad Gil

Tony, as you mentioned, it’s an open question what the right way to scale data up is. There are strong theories around teleoperation, pure RL, video, and world models. How have you thought about all of these?

Tony Zhao

From our perspective, it was somewhat surprising. In the beginning, we worried that data from the glove would have higher quantity but lower quality compared to teleoperation, because with teleoperation you’re using exactly the same hardware and software stack between training and testing, so it’s perfectly distribution-matched.

What we realized is that the glove form factor encourages people to make more dexterous and natural movements. Those movements actually result in more intelligent behavior on the modeling side. In terms of data quality, we don’t really see a gap between teleoperation and glove data after we did the 20 engineering iterations.

Cheng Chi

Yeah. Apparently, there is a mismatch: in the camera frame, there’s a human instead of the robot. There are a lot of things we need to do to convert human data one-to-one into robot data, as if it were robot data, and make sure the model can’t tell the difference. That relies on the full-cycle iteration between hardware and software.

Elad Gil

What about RL? We see a lot of promise for RL in locomotion, and we think that will continue to be true for locomotion.

Tony Zhao

From what we see, RL as a method is very powerful, but it’s much less sample-efficient compared to imitation learning. We see it working well in environments where it’s easy to simulate. For locomotion, you only need to worry about rigid-body dynamics and rigid-body contact between the robot and the ground. Because you engineered the robot, you know everything.

For manipulation, it’s hard for us to imagine having the same amount of diversity and the same distribution of real objects, in terms of matching both appearance and physical properties. We think that’s going to be challenging compared to glove data collection and teleoperation.

Cheng Chi

I think it’s really about which method can get us there faster. There might be different methods that will eventually get there. For example, simulation and world models. It’s almost a tautology to say that if I have a perfect world simulator, anything can be done there. As long as you can do it in the real world, you can do it in a simulation. You could cure cancer in a simulator.

But what it turns out for robotics is that some things are harder than others, and it really depends on the problem itself. In the case of locomotion, as I mentioned, all we need to model in a simulator are point contacts with somewhat flat ground, like feet. But the behavior we want out of it is actually very difficult to model.

It’s all these reactive behaviors: when you feel like your leg is hitting something, you should retract and step again. These are very hard to describe or learn from demonstrations directly. In the case of manipulation, I think the difficulty is flipped: it’s a lot easier to capture the behavior itself, and it’s a lot harder to simulate the world.

For example, if you were to grasp a transparent cup with some orange juice in it, it’s ridiculously hard to simulate how your hand deforms around the cup, how all those ripples and the color of the juice affect the rendering, and what the policy ends up seeing. Simulating that is very expensive and difficult.

But all we need to learn is to get your hand in front of the cup and then close it with the appropriate amount of force. That’s actually very easy to learn. That’s why we see so many successes with imitation learning in the case of robotic manipulation: the behavior itself is not as hard as simulating the world, and that’s why we see faster progress there.

Sarah Guo

Is there anything that you’ve changed your point of view on in data over the last year?

Tony Zhao

One thing I wouldn’t say changed is that data quality really matters. I always knew data quality mattered, but once you scale up, it really matters. The diversity of behavior that you experience in the wild is very hard to control, and hardware failures are hard to control as well. You need to constantly monitor them and spend a huge amount of engineering effort just to make sure that the data is clean.

Cheng Chi

Also, building all those automatic processes is important. We have our own way of calibrating the glove before we ship it out, and we have this whole software system to catch when something is broken on the glove so we can detect it automatically. The importance of data quality translates into all these repeatable processes, so we don’t need a human staring at the data to know that something is wrong.

Sarah Guo

When you described the beta for next year, a lot of it sounded like, “We just want to understand behavior—how people actually want to use it—and we can make some design decisions for the actual product.” What technical challenges do you still see?

Cheng Chi

To me, there are 2 kinds. Number 1 is really figuring out the training recipe at scale. As a field, we’ve just entered the realm of scaling, and we’ve just gotten the amount of data that we need. I think now is the perfect time to start doing research and actually figure out what exact training recipe we need to get robust behaviors. We’re in a unique position because of the amount of data and the entire pipeline we’ve built around data.

The second point is that hardware is hard. We’re still pushing the performance envelope of hardware. It’s not really clear what is needed for the hardware to be reliable, because whenever the mechanical team builds hardware, the learning team will try harder to push it against the boundary, and then it will break at some point.

What’s interesting about this company is that everybody’s under the same roof. Immediately after something breaks, it goes straight back into mechanical design, and then we have another iteration—for the hand parts, for example—very quickly. Hardware is hard, but it is important. It’s the hard but right thing to do, and we as a field shouldn’t avoid doing the hard things just because they’re hard.

Tony Zhao

I want to echo Cheng’s point about research. When there is data scarcity, it’s really easy to come up with cute, fancy research ideas that don’t end up scaling very well. This is why, when we built the company, we focused on the infrastructure, a scalable data pipeline, and operations before we started to really dive into research, which we only started doing 3 months ago. We want to avoid doing research that doesn’t scale and focus on things that contribute to the final product.

The second point is that robotics is intrinsically a system. Right now, there isn’t an existing general-purpose home robot out there, and we don’t really know what the interface between different systems is or what is even good. If you’re working with a partner, it’s actually really hard for them to understand your standard of good, because your standard of good is changing all the time.

This is why we’re building everything in-house, taking a more full-stack approach. We build our own data collection device, which is co-designed with the robot. We build our own operations team to figure out how to get the most high-quality data out as efficiently as possible, and of course our own AI training team to make the best use of that data.

These things are really not easy. They make a company a lot harder to build, because you suddenly need so many teams and they need to work together. But we believe it’s the right thing to do.

Sarah Guo

Okay. I’m going to ask you a few questions that are uncomfortable guesses now. When will people be able to buy robots commercially for the home?

Tony Zhao

This is something we’re really excited about, because we have so many prototype robots in our office and we really want to get them out there. The next step of our plan is to have a beta program in 2026. For people who sign up and are selected, they will have a real robot in their home, and it will start doing chores for them.

It’s going to be a really interesting learning experience for us, because we’ll see how humans interact with the robots and what kinds of things people really want the robot to do. I think this will happen before we actually ship it to the masses, because we have an incredibly high standard for what we’re willing to ship from a consumer-experience standpoint.

We want the robot to be highly reliable. We want it to be capable. We want it to be cheap. I think it really depends on the results of the beta program, which will decide when is a good time to ship it. Is it 2027? Is it 2028? All of those are possible.

Sarah Guo

But it's not a decade away.

Cheng Chi

No, it's definitely not that far away.

Sarah Guo

How much do you think it could cost?

Tony Zhao

Right now, the prototype robots we have in-house range from around $6,000 to something like $20,000. What's interesting is that the big difference here isn't that we found a better actuator. We're using the same actuators, which are very low-cost. The real difference is the cladding of the robot. When you're trying to make them at low scale, it's just really expensive. The claddings cost a few thousand dollars to make.

But these are the types of things that, as we scale up, become dirt cheap. Instead of doing CNC and hand-painting them, we'll use injection molding. As we get to a scale of a few thousand units, we can drastically reduce the material cost, likely to under $10,000. What that implies is that when we sell the robots, the price will be somewhere around that.

Sarah Guo

Okay. Fast-forward 2 or 3 years. If you look 5 years and beyond, and home robots are ubiquitous, what does life look like? How does it change for your average person?

Cheng Chi

This is a different answer for everyone. For me, I really hate dishes. In my sink, there are always 4 or 5 dishes that are somewhat dirty, and it kind of stinks. After a long day of work, it really doesn't feel good to come home and see a home like that. I think the world we'll live in is—

Sarah Guo

It's going to be cleaner.

Cheng Chi

It seems like it will be cleaner. I was just thinking about it as the marginal cost of labor in homes goes to zero.

Sarah Guo

The last thing I want to make sure we do is talk about demos. There are a lot of robotics launch videos today. It's been years since you saw an Optimus serving drinks at a bar. Why aren't those robots available, and what is actually hard?

Tony Zhao

I think the way I would put it is: make zero assumptions. No Priors.

Sarah Guo

Okay.

Tony Zhao

If you see a robot handing one drink to one person, first ask whether that's autonomous or teleoperated. That's the first thing. We should look at the tweet and see what they say about it. Then ask whether it shows the robot giving another, slightly different-colored cup to the same person. If they didn't show it, it means that the robot may literally only be able to pick up that single cup and give it to that same person.

When we look at demos, we tend to put our human instincts into them. If you can hand a cup to that person, it must be able to hand a different cup to another person. Maybe it can also do my dishes. Maybe it can do my laundry. There are a lot of visual inferences we can make, which is what's great about robotics: there are a lot of possibilities. But when we look at demos, we should only index on the things that are demonstrated strongly, and assume that's likely the full scope of the task.

Another aspect is that, at least as a researcher, I appreciate the number of interactions that happen in a demo. Usually, the more interactions you have, the more opportunities there are for failure. The longer the sequence is, the harder it actually is. That's something we really emphasize here, and it's somewhat uniquely easy for us because the glove-based method of data collection is so intuitive for people. It's really about generalization and reliability.

Sarah Guo

Can you explain the demos that you guys are showing?

Cheng Chi

Yeah, of course. We're showing 3 categories of demos. The first one, as you saw, is a messy table. The robot cleans up the whole table, dumps the food into the food-waste bin, loads the dishes into the dishwasher, and then operates the dishwasher.

What makes this demo really hard is that it's a mix of very fine-grained manipulation with these super-long-horizon, full-range tasks. You need to go up, and you also need to go down a lot.

Sarah Guo

It's mobile manipulation.

Cheng Chi

Exactly. The reason we can show this is that our data-collection process is so nimble and easy that we can make long-horizon dexterous demos possible. It's also about the forces involved. You might have seen that we're trying to pick up 2 wine glasses with 1 hand.

Sarah Guo

I struggle with this, but yeah.

Cheng Chi

It's actually really hard. Because they're transparent objects, we also need to load them very precisely into the dishwasher. A lot of it is about how much force you apply. If you're trying to grasp 2 glasses in 1 hand and you squeeze a little bit harder, you're going to break one of them. When you load them into a dishwasher, if you're pushing in the wrong direction and they hit something, they're going to shatter.

We shattered a ton of glasses when we were experimenting with this. These tasks are really high-stakes. It's not just about recovering from mistakes; it's about not making those mistakes in the first place. That's generally the case in a lot of home tasks: you're just not allowed to make any mistakes.

Then we get into the generalization demos. We book around 6 Airbnbs and get there zero-shot to see if the robot can do part of the task. The 2 tasks we use are, first, going around the table and collecting all the utensils into the caddy. The other is grabbing a plate and loading it into the dishwasher.

What makes these demos very interesting is that we don't need any data when we enter that home. It's pure generalization. This is as close to a real product as you can get, because when someone buys our home robot, we really don't want them to have to collect a huge set of data themselves just to unbox it.

In addition to generalization, those 2 tasks are also really precise. We're using the exact silverware in the home, and you need a few millimeters of precision to grasp it properly. Those forces are also hard to perceive because the objects are reflective, and the lights look weird on them. We have a transparent table in one home. I think the table looks like nothing, and the robot still reacts very well to it.

Again, the reason we can do it is that we have more than 500 people in our data set, and we've seen so many glass tables. The robot is able to handle it.

The last set of tasks we did is about pushing what's possible in terms of dexterity. The 2 tasks we chose were operating an espresso machine and folding socks.

Sarah Guo

Mhm.

Cheng Chi

What makes these hard is that they require very fine-grained force, which is hard to get with teleoperation. These days, there isn't a good teleoperation system that lets you feel how much force the robot is feeling. Basically, when you're teleoperating, your hand is numb, and sometimes you're applying a huge amount of force to the robot without knowing it. That can result in very low-quality data, with the robot also doing things in that aggressive way that we really want to avoid.

The sock is a very good example. When you're trying to fold it, your 2 fingers can touch, and that forms what we call a force closure. You have a closed loop for the force, and if your grip is stiff, you can apply an infinite amount of force to it and it doesn't look like anything. But because we're using the glove to collect the data, the human collecting it can naturally feel that. It's very intuitive.

I think we're the first in the industry to do sock folding and use end-to-end learning to operate an espresso machine.

Sarah Guo

One of the things that you will also need to scale as you scale up the company is the team. What are you hiring for? What are you looking for?

Cheng Chi

One thing I'm really looking for is full-stack roboticists and people who aspire to become full-stack roboticists. What you learn in this company is that robotics is such a multidisciplinary field. You need to know a little mechanical engineering, a little electrical engineering, a little code, and a little bit of data to actually fully optimize the system.

We have a couple of examples of training full-stack software engineers to become roboticists and training ML engineers to become roboticists. If you want to learn about robotics and learn the whole thing, rather than just be boxed into your small little cubicle, let us know.

Sarah Guo

You told me that you didn't write code until you got to college or something.

Cheng Chi

I was super enthusiastic about robotics, but before that I was mostly doing mechanical and electrical design. Then I realized that the bottleneck was actually how the robot would move, and that there was something called programming. The more I got into it, the deeper it got.

Toward the end of college, I realized there was a thing called machine learning, and that you could train models. The field just goes on and on. I think it's very natural for me to gradually expand my skill set because I'm always looking to build a robot.

Sarah Guo

Well, I hope you discover the next field, because you're no longer doing dishes, too.

Cheng Chi

It's a very fun place to work. Whatever you can imagine about robotics, consumer products, and machine learning, you can find it here, because we're fundamentally such a full-stack company.

We're not just about the software. We're not just about the hardware, but we're about the whole experience, the whole product, and making sure that product is general and scalable in the future.

Sarah Guo

Awesome. Congratulations. It's really exciting.

No Priors Ep. 141 | With Sunday Robotics Co-Founders Tony Zhao and Cheng Chi | BidClub