[AIEWF预览] CloudChef:你的机器人厨师——12美元/小时的米其林星级美食(附厨房参观!)
- CloudChef卖的是按小时计价的厨房劳动力:机器人每小时12美元,无需资本开支,并宣称首日即可实现ROI。 Nikhil Abraham称,这约为“全包人工成本的40%”,即把一次性资本采购转化为餐厅本来就有的工资预算。
- 需求逻辑建立在餐饮业异常依赖劳动力且人员不稳定之上。 Nikhil称,餐饮业每创造100万美元收入需要13名全职员工,而医院只需要4名;餐厅员工平均流失率约为130%——“到第10个月,实际上整个团队都换了一遍”——与此同时,人工成本仍在上升。
- 真正具备投资差异化的是软件,而不是定制硬件。 CloudChef将现成机器人与VLM、语音模型和机器人基础模型结合,再叠加自研热力学感知与烹饪逻辑,用来判断上色程度、菜谱状态和温度;其目标是让机器人“像厨师一样,在真实烹饪流程中推理并作出决策”。
- Nikhil提出了一项颇为惊人的性能主张:在全球具备商业价值的菜系中,已有40–50%的流水线烹饪可以由机器人处理。 在这一范围内,他称机器人能够稳定胜过教会它菜谱的专业厨师;目前客户已包括米其林星级厨师、新鲜快餐餐厅和航空餐饮供应商。
- “一次演示”确实成立,但它指的是配置一个模块化系统,而不是从零开始教会机器人一项新的运动技能。 以煎蛋卷为例,系统会提取厨师作出的视觉或热力学判断、调用的搅拌或煎炒技能及其参数,再将这些信息转化为可跨厨房和设备迁移的中间菜谱。这套方法的前提是机器人已经具备所需基础技能;若没有这些工程化中间系统,就必须在不同背景、尺寸和设备条件下反复演示。
- CloudChef帕洛阿尔托办公室厨房售出的食物,是一个现实验证点,但其自主化系统仍有明确边界。 烹饪决策实现“100%自主”,动作执行达到90%;机器人可以完成粗粒度操作,但需要超过2或3根手指的任务大概率无法处理,而实际部署还依赖安全过滤器、能自动旋转的炉具旋钮和称重秤。Nikhil开始介绍如何用QR码处理食材盒,但这部分访谈录音在此中断。
- Nikhil称,未来1年内,只有少数几家应用机器人公司有望部署超过100台机器人。 另外,他将CloudChef归入“可能只有2或3家公司”组成的交集:既能为现实世界客户创造价值,又掌握前沿通用模型,还有快速扩张的部署管线。真正待验证的执行问题是,其软件驱动的架构能否在不同厨房、设备以及最终“不同机器人形态”之间扩张时,继续保持厨师级产出。
1. CloudChef卖的是一名工人,而不是一台机器
- Nikhil的核心承诺,是通过机器人取代商业厨房中几乎所有非管理岗位的工作,让“每个人都能获得高质量、营养均衡的食物”;这些机器人要“像人一样行动、像人一样学习、像厨师一样工作”。
- CloudChef的第一款机器人配备移动底座和两只机械手,进入设施后,只需一名厨师演示一次就能学会菜谱,随后重复制作或加入工作流。客户按小时支付工资,而不是购买硬件。
- 最终目标是压低餐饮价格:“在McDonald’s的价格区间内,你应该能吃到自己这辈子吃过的最美味食物。”
2. 烹饪智能,而非定制硬件,才是切入口
- 公司创立时确立的原则,是只解决“能够被建模为软件问题”的问题。随着通用机器人零部件和机器人基础模型不断进步,CloudChef可以把电机和制造交给生态伙伴,自己持续迭代软件。
- Nikhil称,这些机器人已经被米其林星级厨师、新鲜快餐餐厅、航空餐饮供应商及其他商业设施使用。他认为,全球具备商业价值的菜系中,约40–50%的流水线烹饪都可以由机器人处理;在这些菜系里,机器人能够稳定做出比教会它菜谱的专业厨师更好的食物。
- CloudChef的烹饪层将自研热力学感知与VLM、语音模型和机器人基础模型结合起来。系统需要判断洋葱的上色程度,跨设备推断菜谱状态,选择火力,与同事交流,并在过程中修正操作。
- 当前能力边界是“粗粒度操作”:如果任务需要超过2或3根手指,机器人可能就无法完成;不过Nikhil认为,大多数厨房工作都可以用2根足够有力的手指完成。
3. 一次性学会菜谱是真的,但边界经过精心限定
- 主持人的质疑值得保留:考虑到食物处理的复杂性,“一次学会”是否只是营销话术?Nikhil给出的明确回答是:“这不是营销。这是真的。”但前提在于系统架构。
- CloudChef的流程并不是一个端到端的“从像素到动作”模型。神经网络子系统和硬编码路径会把演示转化为工程化的中间形式;“学习本质上是在配置这套AI系统”,前提是所需的基础技能已经存在。
- 以煎蛋卷为例,系统会识别厨师的判断是基于视觉还是热力学,判断其调用的是搅拌还是煎炒技能,并提取相应参数。这种菜谱形式可以跨厨房、设备,最终甚至跨不同机器人形态迁移。
- 如果没有这些工程化中间系统,Nikhil称,厨师就必须在不同背景、尺寸和设备条件下分别演示这道菜。
4. 餐饮经济更偏好工资预算,而不是资本开支
- Nikhil量化了餐饮业的劳动力问题:餐饮业每创造100万美元收入需要约13名全职员工,而医院只需要4名;餐厅员工平均流失率约为130%,意味着10个月后几乎整个团队都已换新。
- 餐饮业没有充裕现金,也没有专门为机器人试验预留的固定预算,但它本来就有劳动力预算。他的类比是:雇主不会替员工支付大学学费,而是支付工资。
- CloudChef的价格是每小时12美元、无需资本开支,约为全包人工成本的40%,公司因此宣称“首日即可实现ROI”。从长期看,机器人还可能让餐厅提供此前无法制作的菜谱。
5. 办公室厨房把宣传主张变成部署测试
- CloudChef自营的外卖厨房原本并不是核心业务。由于创始人想念印度菜,他们邀请孟买和德里的心仪餐厅录制菜谱,换取在加州提供服务的版税;Nikhil称,这项业务在DoorDash上的表现出乎意料地好,主持人还提到Uber Eats并称赞了食物。
- 厨房参观进一步厘清了自主化程度:烹饪决策“100%自主”,动作执行达到90%。如果机器人已经偏离轨道,或偏离轨道的概率超过90%,系统就会启用安全过滤器;Nikhil认为,这些过滤器是实际部署所必需的。
- 设备兼容性通过把普通旋钮替换为可自动旋转的旋钮实现,由此建立统一的执行界面。食材通过称重秤计量;Nikhil开始介绍如何用QR码处理食材盒,但访谈录音在此中断。
- Nikhil对人才的吸引力在于生产环境经验:一台正在工作的机器人、早期客户价值,以及1年内部署超过100台机器人的潜在路径。他将公司描述为处于“利用前沿技术为客户交付价值的效率前沿”,同时拥有快速扩张的部署管线。
Okay, we are in the remote studio with this very special podcast. We recorded a tour of CloudChef's Kitchen a while ago, but we wanted to record a little bit of an intro in our remote studio so that we at least get a nice audio podcast intro to the company. We're here with my friend and co-host swyx, as well as Nikhil, who's a founder of CloudChef. Welcome, Nikhil.
Thanks for having me, swyx.
Okay, so yeah, welcome back, swyx. By the time this launches, people will have heard the Anthropic podcast that we did. I think the headline people will see when they see CloudChef is that it is an AI chef. You have this pretty viral video on Twitter from when you launched and told everybody about it, but people also don't know that this is a real restaurant: you actually run a real restaurant. You can order food on Uber Eats. It's really good. What is CloudChef? What is the scope of it? How do you pitch the company?
At a very high level, what we're trying to do is make high-quality, nutritious food available to everyone. The reason it's possible to even think of a future like that is because you can automate practically all non-managerial work inside a commercial kitchen with culinary intelligent robots. Culinary intelligent robots are basically robots that act like human beings, learn like human beings, and work like human chefs.
The video that Sean was talking about is the launch video of our first robot. It's a robot that has a mobile base and 2 hands, goes around, and does work inside a kitchen. It's just like how you would hire a human employee or a human chef: you would hire a robot, and the robot would come to your facility, cook, and learn recipes from the chefs inside the facility with a single demonstration. It would cook that dish over and over again, or participate in that workflow over and over again, like a human employee would. Then you pay the robot an hourly wage, just as you would pay a human.
So far, our robots are used by Michelin-star chefs, fresh fast-food restaurants, airline caterers, and a whole bunch of commercial facilities. They use our robots as hourly-wage labor instead of buying a robot. The thing that makes our robot special is the fact that it can execute at a chef level, or the fact that it has culinary understanding better than even the best chefs in any single cuisine.
What that means is that if a robot is cooking, it needs to know how brown the onions are, how far along you are in the cooking process, what happens if you're cooking in a slightly different appliance, what state the recipe is in, and how much heat to give it. All of this requires thermodynamic modeling of cooking and a visual understanding of what's going on. This is what we call culinary intelligence, and it wasn't possible until recently, when multimodal models got good enough and we built out some thermodynamics modeling to aid that. The end result is a robot that can reason and make decisions in real-world cooking processes like a chef would.
More robot foundation models are coming up and getting better. These robots are finally also able to do real actions and real motions inside a kitchen. Right now, they're good enough to do things like gross manipulation. If a human requires more than 2 or 3 fingers to do a task, the robot is probably not able to do it. The good part is that most tasks inside a kitchen can actually be done with just 2 fingers. If you go around any commercial kitchen and your 2 fingers had enough strength, you could probably do most tasks inside that kitchen.
We start with line cooking, which is the biggest labor cost for restaurants and other food-producing facilities. Our robot is able to do line cooking for about 40% to 50% of the world's commercially valuable cuisine, to the point that if we put our robot against an expert chef in that cuisine, our robot can consistently make the food better than even the chef whose recipe it is. Computers just do some things inherently much better than the human brain.
That's a quick overview. We've trained our in-house models to do thermodynamic perception, and we leverage current VLMs and voice models for perception. We also use them to enable the robot to do tasks, talk to human beings, interact with other coworkers in the facility, and course-correct its goals and whatnot.
The high-level goal, as I said, is to replace all non-managerial work inside commercial kitchens with culinary intelligent robots. When that plays out, we think we'll all live in a future where we have access to really high-quality food at fast-food price points. At McDonald's price points, you should be able to eat the tastiest food that you've ever had in your life. That's the thing that we want to create, and now that the robots have started to work in the real world, we see a future in which we'll make that possible soon enough.
We actually started experimenting with these robots in our own facility. We built an in-house delivery kitchen at our office in Palo Alto. We weren't expecting it to do this well; it picked up really well on DoorDash. My co-founder and I had moved from India to Palo Alto, and we were missing really high-quality Indian food here. We went to our favorite restaurants in Bombay and Delhi and asked, “Can you record your recipes? We'll serve them in California and give you a royalty.”
That's not the core business we're focusing on. It's just something we use to validate our technology. The fact that it's doing so well and the ratings are so good is a testament to how good the technology is and how well the robot is functioning right now.
I can confirm we tried the food. We'll see it later in the video. It's really, really good food. I like the term you use there: artificial culinary intelligence. ACI has been achieved internally. It's interesting because when you frame it like that, you guys are doing something pretty different from robotics, right? You mentioned that the robots are more like off-the-shelf parts. It's not specialized robotics; it's actually the software underneath. So yeah, we can talk a bit more about that.
Correct. When we started the company, we had 1 core ideology: we would only solve problems that could be modeled as software problems. Culinary intelligence and decision-making were the first big open problems that we could model in software and solve.
But when we started, robots were still very much an electromechanical problem. They weren't really a software problem. Now, with robot learning models and robot foundation models, it has gotten to a point where you can start modeling the physical actions that somebody does in software, solve them in software, and have software iteration cycles.
We didn't want to build any hardware. We didn't want to be a hardware company because that wasn't our strength, and we didn't think the hardware iteration cycles would be beneficial for a company like this. Very recently, it has gotten to a point where you can take off-the-shelf parts, put a bunch of robot intelligence—quote-unquote—on top of them, and get the robots to work. That is a software iteration cycle. You don't have to build your own hardware, get into how to manufacture the motors, manufacture the robots, or design everything. We rely on the ecosystem to solve all those open questions for us.
We source general-purpose robot parts, or general-purpose robots, and write software on top of them. We take general-purpose robots and leverage general-purpose intelligence, such as LLMs, VLMs, and robot foundation models. Then we build this proprietary culinary layer on top, which has everything to do with thermodynamic modeling of cooking, custom evaluations for manipulation, and understanding through perception what stage of the cooking process you're in.
We've built all of those things, and we are hoping to ride the tide of advances in both multimodal models and robot foundation models, using general-purpose robots as the vehicle to make that happen. That's, in a nutshell, how our approach works. We are very focused on modeling every part of the workflow as a software process.
Now that we're able to do it, we can do full-stack work inside a kitchen. We're not just an assistant robot that can be prompted by somebody on-site or a guidance system that tells humans what to do. It's now able to do full-stack work because the entire workflow can be modeled as a software process.
I know that we'll see how the robots work and what's happening under the hood later, but the overall question I'm sure a lot of people have is: What's the business model? How do people hear about this? How do you think about renting a robot for $12 an hour versus hiring a chef? How do you arrive at this hourly rental rate and all this stuff? It's very cool to see, and it's good to see that it works.
Yeah, yeah. I’m just curious: how does that side of the business work?
From a business perspective, the main thing to keep in mind here is that food prep is the most labor-intensive industry of all labor-intensive industries. Quantitatively, the way you measure it is by how many full-time employees you need per $1 million of revenue generated. Food requires about 13 people per $1 million of revenue generated, and the second-most-labor-intensive industry is hospitals, which require 4 people per $1 million of revenue generated.
So food is more than 3 times as labor-intensive as the second-most-labor-intensive industry. Labor costs are going through the roof and have been increasing year over year. Depending on what Trump does with illegal immigration, they could go even higher, and staff turnover is really high. The average restaurant is operating at around 130% staff turnover. By the end of 10 months, practically your entire staff is new.
So you have high turnover, very high costs, and the most labor-intensive industry. The reason we ended up at this price point, or with this sort of pricing model, is that food service is not a very profitable industry. They don’t have free cash just lying around to do experiments, and there aren’t fixed budgets set aside for buying new robots or testing things out. If it doesn’t work, they don’t take the attitude of, “It doesn’t work.” There is a very readily available labor budget that we can tap into.
Just like when you hire somebody, you don’t pay for their college tuition; you just pay them a salary. We thought, why should that be any different for robotics? These robots are now cheap enough that you can put that business model out there without losing money on every robot you sell. The robot costs have gotten to a point where an hourly labor-pricing model works, and the robots are also good enough to do the entire chain of work.
At $12 an hour, it’s around 40% of what a loaded human would cost, so our customers get their ROI on day 1. The robot starts working from day 1, and over time these robots just get better. The hope is that at some point they also start making better food at any given facility—not just by cooking the same thing, but by enabling the facility to make recipes that it wasn’t able to make before.
Yeah, awesome. I think the last part—we’ll cut right into the kitchen walkthrough video later—but there’s this general goal of demonstration learning: learning from experts, learning from Michelin-starred chefs. How realistic is this? Is it a marketing promise, or do you really just learn from 1 example? Obviously, food is messy. Food needs a lot of different demonstrations. How realistic is it?
I want to clarify 2 things. One, it is not a marketing thing. It’s actually true. Two, the reason it might feel counterintuitive is because our entire pipeline is not 1 end-to-end model.
If you had 1 end-to-end model and you had to train it to do a new thing, being a one-shot learner would be a very big deal. But in our case, we have many AI subsystems that work with each other. Some of those are end-to-end neural networks, and some of them are hard-coded software pathways. We use the best of both worlds to function, and that architecture choice means we don’t have to go directly from the pixels of what a chef is doing and text or whatever to a generalizable recipe that can be cooked across any robot at any time or scale.
There are software workflows and pathways before that. They take the chef’s demonstration and convert it into an intermediate format that is easily digestible by different parts of our system.
One example would be making an omelet. If you want to teach it how to make an omelet, we’re not learning a new omelet-making skill while you’re showing the robot how to make one. If we had that capability, we would be a robot foundation model, and we would already be doing single-shot learning. It is, from a single demonstration—assuming that we have all the base skills for the robot to do it—extracting what kinds of decisions the chef is making.
Is it visual? Is it thermal? What kind of skill is the chef invoking? Is it stirring or sautéing? What are the parameters for those skills? For us, learning is basically configuring this AI system, not using an end-to-end model that goes directly from pixels to robot actions.
We wouldn’t be able to do single-shot recipe learning if we didn’t have those systems. We would have to have the chef cook the recipe in various backgrounds, at various sizes, and with various appliances. Because we have these engineered midpoints that go from 1 expert demonstration to a recipe form, the recipe can then be recreated across different kitchens, different appliances, and, in the future, different robot morphologies.
Oh, other robot morphologies. But yeah, now you only do the 2 arms, right?
Okay, so we'll get people to a call to action, and then we'll cut to the video. You are going to be at the AI Engineer World's Fair next week. If people want to see the robot live, they can see it there, probably taste some food. Although I don't know how much food we can serve. We'll see. We'll see.
Obviously, I think part of the reason you’re doing this is that you’re trying to hire, right?
Yes. This is an immediately applicable use case. There are only a handful of applied robotics companies that have a path to deploy more than 100 robots in the next year. Now that our robot is working and we have early signs of it being super helpful to customers, you’ll actually be working on a robot that is in production.
There are people in robotics who are working on more complex hardware problems and more complex software problems, but I think we are at the efficient frontier of value being delivered to the customer using cutting-edge techniques and having a rapid scale-up pipeline. I don’t think a lot of companies can say that they have all 3 of those things.
And cooking, like I said, is a very powerful mission. Today, the food that you’re eating is fast food. Most cheap food is fast food. Fast-forward 10 years, and you can eat very high-quality food. In fact, if you work with us, you can already eat very high-quality food in our office, as Sean and Vivu can confirm.
Coming back to it, I think the mission is very powerful. If you are excited by creating value in the real world while also doing it in a way that serves all the current capabilities of state-of-the-art general-purpose models, we are probably 1 of maybe 2 or 3 companies at that intersection. If that excites you, you should come talk to me.
Yeah, I think that’s actually a very strong pitch. I would say that you can actually just try the food on Uber Eats. You can order it here. It's one of those virtual cloud kitchens, and it looks so good. We've tried it. We'll cut to the video later, but thanks for jumping on and sharing your journey with us. I think this is very exciting. I think you’ve somehow found a way toward the most immediately applicable industrial use case of robots. There’s obviously a lot of scope for vision-language models and for solving a lot of hard engineering problems.
One thing you didn’t say is that all of this is within very tight engineering parameters, which I think is pretty hard. You have to run it at a lot of frames per second and do a lot of that on-device.
Yeah, cool. I'm looking forward to seeing you next week at the conference, and we'll cut to the video.
Now, all the culinary decision-making is 100% autonomous, and the actions are 90% autonomous.
If the robot has gone off, or if the probability of the robot going off is more than 90%, we have safety filters like that. That’s what makes it deployable. Otherwise, it’s not very deployable.
Basically, appliances are controlled in 1 of 2 ways: any appliance in any kitchen is controlled either using a knob or a touchscreen. We remove the knobs that control the appliance and put in knobs that can turn themselves.
That gives us an actuation surface across all appliances, so we don’t need to teach our robot to use every appliance.
Huh?
Yeah. So, $12 an hour is what you pay for this. No capex.
What?
Yeah.
Oh, is there anything else going on, like any other equipment? For example—
The other thing is that all the ingredients that are required basically get measured in these weighing scales. Regardless of which kitchen we go to, all kitchens store their ingredients in boxes. We just slap a bunch of QR...