机器人CEO:人形机器人革命确实存在,现在就开始——与 Bernt Bornich & David Blundin 对谈|EP #188
1X的核心判断是,家庭是触达消费者规模和发展具身智能的最快路径,而不只是自动化的又一个市场。 早期EVE机器人在重复性的安防或物流任务上运行约20–40小时后就会出现学习平台期,但家庭场景尚未显现多样性天花板。Bernt Bornich估算,部署10,000台机器人后,整个机群每天产生的非重复有效数据可能超过YouTube:“互联网其实没那么大。”
即使完全自主还未实现,短期商业模式也已经具备实用价值。 Peter Diamandis提出约$30,000的买断价,或每月$300的租赁价——每天$10、每小时约$0.40;Bornich回应“我认为我们可以做得更好”,但拒绝公布实际价格。工厂计划在2026年底达到年化产能20,000台以上,不过由于产能爬坡,2026年实际产量会更低。
NEO Gamma的工业优势在于刻意简化、轻量化的架构,而不是汽车式复杂度。 这台身高5英尺4英寸、体重66磅的机器人据称可举起约150磅、携带约50磅,续航4小时、从完全没电充满约需2小时;其零部件数量为数百个,而汽车约有50,000个。Bornich对制造的概括是:“它更像冰箱,而不是汽车。”
1X押注的是,物理智能必须具备空间、时间、触觉和互动维度,而不是以语言为起点。 互联网视频提供了观察结果,却没有智能体的目标、选择的动作或观察到的后果;机器人则可以执行“提出假设—采取行动—获得反馈—修正”的科学方法循环。Bornich不愿断言具身性在理论上不可或缺,只说这“比文本或低保真模拟短得多”。
远程操控既是产品栈的一部分,也是训练栈的一部分,而透明度承担了大部分信任成本。 早期客户获得的将是尽力而为的自主运行与预约式人工辅助工作的组合;用户必须批准远程操控,机器人会明显提示有人介入,操作员看到的是经过过滤的场景。私人数据有24小时的预训练删除窗口,人工审核则需要用户批准并提供解密密钥。
安全边界既通过物理设计设定,也通过动作条件世界模型加以约束。 NEO Gamma质地柔软,整体设计使意外撞击可能造成疼痛,但不太可能导致严重伤害;烹饪及其他涉及危险物体的任务初期仍会禁用,因为“一旦你拿起一壶滚烫的水,就再也无法保证自己是安全的”。在模型评估中,1X把控制器放入模拟世界,让“机器人身处Matrix”,测试性能、红队案例和不安全行为。
上行空间是一套劳动与基础设施飞轮,但约束条件高度依赖物理世界。 Bornich设想的“硬启动时刻”是机器人开始制造机器人、芯片工厂、数据中心、能源系统、实验室和专业自动化设备;Peter将劳动估算为全球$110万亿美元GDP的大约一半。Bornich认为到2040年拥有10 billion台人形机器人“可能大致正确”,甚至可能偏晚,前提是许可、电力、铝、稀土加工、磁体、晶圆厂以及足够多的机器人能够支撑所需劳动力的启动。
1. 家庭是1X的数据引擎
Bornich选择家庭的首要原因是规模:消费硬件可以触达数十亿台设备,而人形机器人只有在规模带来可靠性、成本、智能和生态改善后才真正有吸引力。对于任何单一工业任务,他说,“总会有更好的自动化系统”。
经验警示来自2022–2023年EVE在安防和物流领域的部署。每项任务的学习在约20–40小时后进入平台期:反复移动一个杯子接近低端,而导航、开门和设施安防则超过40小时。
在Bornich看来,工厂重复一项操作无法通往通用智能。相比之下,1X目前的家庭场景尚未发现多样性天花板;他将1X定位为朝着AGI前进,而不只是快速在工厂里应用劳动力。
他对丰裕的更广泛公式,是用知识或智能乘以劳动、商品和服务。只改善模型层,物理底座仍受约束:智能需要能够在社会中学习的机器,最终还需要扩张承载智能运行的基础设施。
2. 家庭机器人必须理解社会语境
Diamandis把NEO Gamma比作探索物理世界的幼儿,Bornich接受这一比喻,但补充了一个条件:机器人必须先从互联网、模拟和合成训练中获得有用的先验行为,尝试合理动作,而不是随机游荡。随后才进入真实世界互动,目前还无法判断这一循环能扩展到多大规模。
家庭场景更深层的难题在于,“我们做的一切都是社会性的”。一个空咖啡杯可能意味着需要续杯、清洗,或只是继续留在原处使用;正确的物理动作取决于周围的人、时间和习惯,而重复性的工业数据无法捕捉这些变量。
这套逻辑催生了1X十年前的创立要求:机器人必须足够安全、足够有能力、也足够便宜,才能“生活在我们中间并向我们学习”。同时做到3点,意味着在保留人类级力量的同时简化系统,并把价格压到规模化可负担的水平。设计工作最终导向了腱驱动机器人,以及持续十年的原创研究。
Bornich将机器的社会角色描述为既不是另一个人,也不是另一只宠物,而是“介于两者之间的某种东西”——他对Calvin and Hobbes的比喻很简单:“它就是Hobbes”。一个终身陪伴者可以记住主人,通过目光和肢体语言交流,让AI真正具备物理存在感。
3. 价格目标是消费经济学,而非工业资本开支
Diamandis提出了$30,000买断或每月$300租赁的工作区间——每天约$10、每小时$0.40。Bornich没有确认价格,只回答“我认为我们可以做得更好”,同时认可这一范围在方向上合理。
Bornich强调,1X既要做出最好的产品,也要具备价格竞争力。他说,综合考虑自由度和能力,NEO Gamma仍能与中国替代品竞争,并称1X已经显著降低了系统复杂度。
Bornich说,一项民调发现,人们通常至少想要2台机器人,具体取决于价格。David Blundin随后推测,4台或6台可能会成为常态,因为机器人之间的协作比人类搬运工更精确。Diamandis认为数量过多,称一台NEO Gamma只有在需要团队时才会带来额外机器的需求。
Bornich把问题框架进一步拉宽:劳动丰裕也会让更大的房屋和更多实物商品变得可负担,因此今天的家庭尺寸可能并不适合用来估算机器人数量。在他的叙述中,机器数量会与它们能够创造的建成环境共同演化。
4. 物理智能从动作开始,而不是从语言开始
Bornich反对智能始于语言这一前提。语言是人类创造的高效压缩格式,但核心在于“空间和时间”:智能体如何观察、感受、预测并在世界中行动。他偏好的架构从这些模态出发,再叠加文本。
Blundin提出,具身性让智能得以产生,而语言让智能得以扩展。Bornich谨慎地回答,他无法严格证明具身性是必需的;他只能说,有“非常、非常强的证据”表明,具身性是通往人类级乃至更高智能的更容易工程路径。
关键的数据差异在于能动性。YouTube记录发生了什么,却遗漏了行动者的目标、内部世界模型、选择的动作和由此产生的反馈;机器人则同时拥有观察结果、目标与动作,以及观察到的结果。Bornich将其直接对应到科学方法:提出假设、进行测试、观察结果、重复并学习。
他承认模拟可能复现这一循环,因此不会断言其他路径不可能。问题在于实践:模拟的保真度远低于现实,弥合差距极其困难,而且相比在物理世界收集互动经验,模拟消耗的算力要多得多。
5. 机群数据与制造规模相互复利
Bornich的粗略计算是,10,000台机器人每天运行大部分时间,收集的有效、非重复数据将超过YouTube每天上传的数据。大规模部署后,机器人生成的经验可能远超互联网数据,在设备数量与模型能力之间形成反馈循环。
他纠正了主持人对产量的理解:1X跨多个世代总共制造了超过100台机器人,但NEO Gamma不超过100台。规划中的工厂将在2026年结束时达到20,000台以上的年化产能,不过产能爬坡意味着2026年实际产量无法达到这一完整数字。
后续工厂的目标是再迈上一个接近数量级的台阶,但Bornich提醒说不会完全达到。他提到iPhone约1.7倍的增长节奏,包括规模化暴露新约束时出现的平台期;当被问及2030年前能否达到年产数十万台时,他回答:“远不止。”
规模最终会撞上铝精炼和装配劳动力。即使零件很少,NEO Gamma的制造仍比iPhone复杂;如果所需劳动是后者的5倍,劳动力供给会先失效。因此必须转向机器人组装机器人,并扩张晶圆厂、数据中心和能源供应。
6. NEO Gamma的工程逻辑更像家电,而不是汽车
NEO Gamma身高约5英尺4英寸,体重66磅。Bornich称它能举起约150磅、携带约50磅,具备“运动型人类的重量—力量比”,同时支持公司对轻量、柔软机器的强调。
电池续航约4小时,充满电约需2小时。Bornich不太在意连续运行时间,更关注短暂、择机充电的间歇是否足以保持机器可用;他家里的那台通常会在自然停顿期间充电,而不是把电池完全耗尽。
家庭场景一个容易被忽略的规格是静音。他说,第一天听起来可以接受的机械声,到第三天就会变得令人烦躁,因此“完全安静”是硬要求,而不是后续优化。柔软和适合拥抱同样重要,因为人必须能够在机器的工作空间里保持放松。
NEO Gamma有数百个零部件,而Bornich估算汽车接近50,000个零件、重4,000磅。“如果你在这里做得非常好,它更像一台冰箱”——尽管是一台复杂的冰箱——而不是汽车;这个比较概括了其目标制造经济性。
7. 灵巧性抬高可学习行为的上限
许多人形机器人停留在约26个自由度,通常覆盖腿、手臂和颈部,但省略手腕。NEO Gamma增加了3个用于表达的颈部轴、3个贯穿脊柱的自由度、完整的手臂关节,以及每只手22个自由度。
Bornich说,这只手在功能上接近人类的22个自由度,不过具体数量取决于是否将细小的腕骨运动单独计算。这一差异很重要,因为捧持、可变形物体、精细处理和手内操作会带来一整类低能力机械手无法接触的经验。
他的智能指标是多样性,由2个相互独立的变量决定:环境的多样性和机器人的物理能力。一台复杂机器人只执行一项工厂任务,仍然会进入平台期;一台能力不足的机器人置身丰富家庭,也无法收集操作数据。因此策略是“两个维度都要拉满”。
8. 远程操控既是能力测试,也是带标签数据
当硬件团队和软件团队争论任务为何失败时,Bornich的诊断方法很简单:让最好的远程操作员试一遍。如果远程操控能够成功,“正确的神经网络加上足够数据就能做到”;这证明机械系统能够实现该行为,但不证明自主系统具备广泛泛化能力。
操作员并不是穿着VR服逐个操控关节。指令越来越抽象——把手放到这里、抓住那个物体——而低层控制由学习系统解决运动细节。自下而上的运动能力与自上而下的行为模型,目标是“在中间相遇”,直到操作员消失。
这一调试方法已经开始以一种令人鼓舞的方式失效。NEO Gamma的手能够提供快速、高保真的触觉反馈,而这些反馈无法高效传给人类操作员;现实世界强化学习已经开始生成一些操作行为,远程操作员“只能梦想做到”。
Blundin在最初指出演示可能误导观众对自主性的判断后,仍为远程操控辩护。Bornich的区分很准确:远程操控机器人显然能够完成这一物理动作,但还不能自主完成。透明标注很重要,因为这些演示相当于机器人领域的专家标注微调数据。
9. 快速控制留在本机,机群学习保持共享
所有将感知转换为电机扭矩的部分,都是端到端学习得到的;Bornich说,外围代码只有几百行,因为“全部都是权重”。参数数量仍属机密,而且相对较小,因为控制器必须在设备端快速运行,同时处理视觉输入。
出于工程原因,算力被放在头部,而非为了视觉拟人化。其他位置已经塞得很满,且最高带宽的数据流来自眼睛;1X使用高分辨率、高频视觉,不配备LiDAR、结构光、腕部摄像头或类似传感器,因此即便把算力移到躯干,数据路径也会变得尴尬。
控制系统是分层的:分布式电机级决策约以25赫兹运行,机载“大脑”约为5–10赫兹并保持低延迟,更慢的、类似语言的约1赫兹流式处理可以放在机外。云端推理无法闭合操作所需的高频触觉循环。
一台机器人学会打鸡蛋,并不意味着经验留在单机内部。经过验证的机群数据会训练共享模型,改进后的检查点可以部署到每一台设备。Bornich还预计会进行大量端侧联邦学习,让每个陪伴者保留私人的个人经验,同时共享一套共同的智能骨干。
10. 隐私由用户控制,但存在真实取舍
Bornich明确表示,早期用户需要用部分隐私换取参与权:“没有数据,我们就无法把产品做得更好。”日常数据可以进入自动化训练而不接触人工;如果1X希望查看某个具体时间窗口,用户会收到相关视频,并自行决定是否提供解密密钥。
训练还会延迟24小时进行,让用户有时间在事件进入模型权重前,将其“从存在中抹掉”。
远程操控必然会暴露任务场景,因此1X会把人物过滤成色块,并突出被操作的物体。未经批准,任何操作员都不能进入;NEO Gamma的灯光会明显变化,操作员必须来自用户批准的名单。Bornich举的例子是,由4名指定操作员为一个家庭提供服务。
产品将尽力而为的自主运行与预约式任务完成分开。在家中,Bornich可以要求机器人处理白色衣物、接收并冷藏Instacart配送,以及在他外出时进行一般整理;其中一部分工作自主完成,另一部分由人工辅助,但“我其实不在乎组合是什么,任务完成就行”。
11. 第一阶段的安全边界由危险工具,而不是原始力量决定
尽力而为的自主运行允许通过失败学习:用户可以说“坏机器人”,Bornich称任务在失败后推进得比成功后更快。更难的对齐场景是奶奶要求拿威士忌,因为模型倾向于迎合,可能会执行本应拒绝的请求。
内在安全意味着,NEO Gamma的低质量、柔软材质和顺应性运动应使意外撞击最多造成疼痛,而不太可能导致严重伤害。但当它拿起一壶滚烫的水等危险物体时,这一保证就会消失。
因此烹饪功能初期不会上线,尽管公司内部已经开展相关工作,Peter也提出过制作照烧三文鱼的要求。Bornich预计,危险能力只有在行为置信度提高后才会逐步解锁;他不会为了早期实用性而交换一个没有边界的家庭安全风险。
物理防护还要与模型评估结合。Bornich反对发布一个控制器,再等客户进行“体感检查”,因为机器人和自动驾驶汽车一样,必须在部署前证明性能更好,同时安全性得到保留。
12. 世界模型为自动化红队测试提供“Matrix”
1X的世界模型会根据选择的动作预测后果,包括渲染观察结果和物理力。“本质上就像机器人身处Matrix”:控制器表现得仿佛自己在一个家庭中,却不知道环境是生成出来的。
工程师随后可以大规模重放普通任务、对抗性场景和自动化安全检查。这个模型既是AGI研究方向,也是更近期的发布评估工具,用来测试新控制器是否改善性能,同时没有引入危险行为。
Blundin问,这种模型与数据资产是否比单纯的设备销量更能解释行业的高估值。Bornich说,最终它会被产品化,覆盖数字劳动和物理劳动,但预计机器人收入将“永远占主导”,因为物理世界承载的价值远超投资者的假设。
Diamandis估算当前全球GDP约为$110万亿美元,劳动约占一半,意味着现有可触达市场超过$50万亿美元。Bornich认为这仍低估了机会,因为新能力会扩大工作和生产总量,而不只是替代现有劳动。
13. 人形机器人首先胜在通用平台
面对Salim Ismail提出的“为什么不是6只手臂或一只章鱼”,Bornich承认人形机器人并非唯一可行机器。他的理由是,没有其他形态能同时匹配人类形态的通用性,以及它与人类建造世界的兼容性。Diamandis还认为,学习成果很难直接迁移到6臂形态,Bornich也同意至少会更困难。
Blundin补充说,形态会建立直觉预期:用户已经知道一个类人机器大概能做什么、不能做什么。陌生的6足设计则迫使用户重新学习其能力,破坏了本应消除技术采用门槛的自然接口。
把Apple设备当作打字机使用,显得极度过度工程化——人类掌握纳米级芯片制造,只为写一份文档——但规模让它比专用替代品更便宜、更可靠,也建立了任何专业机器最初都无法匹敌的生态系统。
只有在通用市场足够庞大后,专业化才会回归。机器人最终可能像Star Wars一样,出现维修无人机、额外机械臂和特定任务工具,但这些细分市场首先需要人形机器人创造出的规模和知识。他的结论是:“人形机器人只是一个阶段。”
14. 中国的优势是积累下来的工艺知识
Bornich称中国硬件生态“非同寻常”:一块损坏的电路板、一台机器或一个零部件,跨条街就能找到替代品,使设计和制造迭代速度达到硅谷无法匹敌的水平。更不显眼的优势,是生产线上长期积累并分散在各处的工艺知识。
他用磁体举例说明护城河:科学家可以理解材料,也可以遵循教科书中的每一步,但他们没有那个知道“2小时后必须向左搅,而不是向右搅”的老工人。稀土资源很重要,但可重复的高等级生产依赖这种隐性技艺。
Diamandis认为,中国的地位部分来自自上而下指定机器人城或磁体城。Bornich对此不太确定,更强调中国创业公司和资本的活力;在他看来,自由经济区——更快的许可、更低的摩擦,以及可以建设的许可——才更可能是政策上的关键一招。
两人都希望美国获得类似支持。Bornich建议在美国国内设立经济特区、加快审批;Blundin则认为,数十年来对软件的偏好扭曲了风险投资组合。他直言自我批评:即便自己的基金在种子阶段会向硬问题投出第一笔支票,也不投资硬件,“我也是问题的一部分”。
15. 1X的护城河在电机、腱和耐心资本
Bornich说,1X自行制造电机,包括绕线、制造、自动化和相关电子设备,因为市场上不存在合适的零部件。他称,NEO Gamma电机的扭矩重量比达到世界纪录基准的5.5倍,提供了足够力量,从而无需齿轮。
电机和腱驱动设计让机器人更轻、更易驱动、更具顺应性,同时降低制造成本。它们还要求材料能够承受“数百万、数百万、数百万次循环”,以及新的电机驱动、电力放大、磁性材料和制造方法——这些是研究问题,而不是目录式工程。
AI在硬件栈中的出现已经超过十年:早在Transformer之前,Bornich就编写神经网络来学习电机设计。他将硬件领先优势描述为以年计,而即便优秀的世界模型领先优势可能也只有3个月,因为软件优势扩散得快得多。
公司能够存活下来,是因为一位早期挪威投资者最终卖掉了那座曾容纳其谷仓创业公司的农场,以延长资金期限。1X有意保持小规模7年,开发核心技术,随后向Palo Alto迁移,以利用当地密集的产品、规模化、API、制造和机器人人才。
16. 终局是机器人建设基础设施,而不是徒手完成每份工作
Bornich的10年愿景从可持续丰裕开始:当能源和劳动不再稀缺,社会就不必仅仅为了压低成本而牺牲环境。下一前沿是建设让所有人获得高质量生活的基础设施,随后是规模大得多的科学系统。
粒子加速器、生物科技实验室、化学实验项目和其他高度依赖实验的计划,都需要实体建设与重复操作。他拒绝“天空中的神一般AI”通过眼镜指挥人类的未来,更偏好“人与机器之间的共生与共同发明”。
人形机器人会填补缺口、建设专业自动化,而不是永远低效地模仿机器。它们不会让30个身体共同搬运一副汽车底盘,也不会拿着Dremel逐个加工零件;它们会使用CNC机床等现有自动化设备,建造更多自动化系统,并帮助扩张晶圆厂、数据中心和能源基础设施。
太空提供了一个早期高价值延伸场景:NEO Gamma只有66磅且能效较高,可以降低发射负担,但电机环氧树脂需要进行真空硬化,散热也仍然困难。对于轨道组装,Bornich倾向于先由轨道上的专业人类进行低延迟远程操控,直到积累足够演示数据后再实现自主化。
17. 100亿台机器人是供应链与许可审批问题
Bornich曾承诺在2025年推出早期用户计划,但拒绝给出具体预购日期。他对预期的设定异常直接:客户购买的是“一张参与这场转型的门票”——他们会“领养一台Neo”、教会它,并获得有用但并不完美的产品,同时面对大量粗糙之处。
当被问及Elon Musk和Brett Adcock关于2040年达到100亿台人形机器人的估算时,Bornich称这一数字“可能大致正确”,并表示可能更早实现。条件是社会愿意移除人为约束,让矿山、炼厂、电力系统、工厂和劳动力足够快地扩张。
Blundin进一步指出了半导体供给错配:他估计每台机器人需要约1块完整GPU,可能是2块,而全球每年仅生产20 million块GPU,TSM占晶圆制造的66%。Bornich则把问题再往上游追溯到ASML,以及每增加一座晶圆厂背后脆弱的设备供应链。
铝、稀土和高等级磁体也存在同样问题:拥有材料却没有加工经验仍然不够。许可审批最终可能决定时间表;机器人只有在已经存在能够制造机器人本身的初始工业基础后,才能帮助建设缺失的基础设施。
公司的名字最终回到了它对可信度的标准。Bornich说,机器人视频经常以“4X”或“8X”速度播放,而1X展示的是实时速度:“我们做的一切都是实时的,因为我们制造的是真正的机器人。”
You think about robots in the world probably more than anybody else. What's your vision 10 years from now?
Everybody, we're here at 1X Technologies in Palo Alto. Bernt Bornich is the CEO and founder. NEO Gamma One and NEO Gamma Two are over here. I imagine we're going to have the same level of AI eventually in the robot, where I feel like I'm talking to a fully intelligent being.
And one that is grounded, right? That actually understands what this existence is.
I'm shocked by that. Wow. I'm shocked by that, too.
How do we solve the remaining really hard problems in science? This isn't going to happen without humanoids. It's almost existential to us for human happiness.
Salim Ismail is constantly saying, “Have it look like an octopus and let it operate with all the elegance that an octopus can, rather than trying to constrain it into five fingers on this hand that do certain things and manipulate objects the way we're supposed to manipulate them.” So what's the definitive answer to him?
Let's just say humanoid is the face.
Now, that's a moonshot, ladies and gentlemen. That's great. I'm with David Blundin, my moonshot mate.
Hello, all.
NEO Gamma One and NEO Gamma Two are over here, and we just did a tour of the facility. It's pretty extraordinary. We saw probably dozens of NEO Gammas in different stages of development. They literally manufacture everything from head to toe. How many components are inside NEO Gamma, roughly?
Oh, top secret.
Top secret, okay. Can't say that.
It's in the hundreds, not the thousands.
I just secured my first NEO Gamma at my home by the end of the year. Is that right?
Oh, yeah.
Okay, fantastic. We're about to do a podcast, either with Bernt or with NEO Gamma, depending on what you want. Let's go ahead. We'll go over to the podcast area. Will you lead the way and maybe clear the way for me? Awesome. By the way, those bags over there—NEO Gamma can carry those, over half his body weight.
Okay, Peter. I'm not the NEO Gamma. Hey, can I give this to you to carry?
You can try. It might hit some safety limits, but it usually works.
All right, arms up.
Feel it properly. There you go. You can let it go, and it can take a few steps. There you go. It might, after a few steps, decide, “This is a bit unsafe for me.”
I mean, it's—
Thank you, Neo.
—incredibly strong. All right. It's nice to know that NEO Gamma will clean up the house around you. Well, listen, I'm not sure what number you are, but I want to say thank you so much. Thanks for cleaning up.
Of course. Thank you for your time. A pleasure. A pleasure. Have a great day, and thank you very much as well. I want to be polite. You never know when the robot overlords are going to come after us. I want you to remember that I was really polite. I was really, really polite. Okay, I'm safe. Great.
Do behave.
Everybody, welcome to Moonshots. I'm here with my moonshot mate, David Blundin. Salim Ismail is offline with his son this weekend. But I'm here in particular with the CEO and founder of 1X Technologies, Bernt Bornich. A pleasure, Bernt.
Awesome. Looking forward to this one.
Thank you. We just finished this tour, and it's pretty extraordinary. When did you move into these facilities here?
It's been 1.5 months.
Nice. There are just many levels of people building robots. No robots building robots yet.
We're getting there, but not yet.
But you're getting there.
Yeah.
1. The Home First Strategy
I'm very familiar with the robotics and humanoid robot space. While companies like Figure and Tesla are focused initially on going into factories, automotive factories in particular, you made a commitment to the home.
Yes.
Personally, I'm excited about that, but I'd like to start with why the home.
To me, there are 2 main reasons. There are a lot of reasons, but 2 main ones. The first one is kind of obvious: consumer hardware just scales at a different pace than everything else, right?
We got to more than 1 billion iPhone devices in a bit more than a decade. To me, humanoid robots don't make sense unless they're at scale, right? There's always a better automation system that you can use for 1 specific problem. You need scale so that you really get this incredible reliability, incredibly low cost, and an incredible ecosystem and intelligence.
2. Diversity Drives Intelligence
The slightly deeper reason is also that intelligence comes from diversity. This has been very clear from the beginning, in all kinds of AI research and also in more practical applications of AI across all different domains, whether it's a language model, an image model, a video model, or, in this case, a robotics model.
You don't really need data of the same thing over and over. If you think about it, it's very logical, right?
So if you're in an automotive factory, you're basically doing the same thing over and over again. You're not learning new stuff.
Yeah, and we actually have some data on this. We have some real data because our previous-generation humanoid, EVE, was deployed in both guarding and logistics back in 2022 and 2023.
Mm-hmm.
After about 20 to 40 hours, our robots plateau and stop learning for that specific task. It depends on how complex it is. If you're guarding a facility and driving around—because EVE had wheels but was also humanoid—and opening doors, there's some diversity to that, so you're more in the 40-plus-hour range. If you're just moving this cup from here to over here all day, then you're at the lower end of 20 hours.
Yeah.
There's just no path from there to general intelligence. We may be a bit different from the rest of the humanoid space in this, but I see us more as a company really running toward AGI—
Yeah.
—and asking how we can come there as fast as possible, versus how we can apply labor in industrial or similar settings as quickly as possible.
So it's robotics in service of building true AGI models and getting enough new, rich data to train up these models.
Yeah, and what would you want to— You said 20 to 40 hours for a security guard robot. What's the equivalent for all the variety of things you can do in the home? How many hours of—
We don't know yet.
—do you need? So, tens of thousands of these?
No, we don't know. At our current scale, we don't really see any kind of cap on diversity.
Yeah.
It'll get there, and we'll need to diversify. But you ask a very important question, right? We want to talk about what the goal is. To me, it's not just AI or robotics. It's a combination.
Yeah.
If you think about what this is—what abundance is—it's an abundance of knowledge or intelligence multiplied by an abundance of labor or goods and services. You need both, and they follow hand in hand. We can talk more about that, but the constraints we have in society aren't always only on the intelligence or data layer. They're also on the substrate that we're building on, right?
So when I think about it, I imagine this is why a toddler crawling around, playing, investigating the physical universe it's in, interacting with different people and different things, is learning and building a model in its neocortex.
And so, is that basically the same: your NEO Gamma is an infant learning in a diverse environment?
It is.
Yeah.
And I think just to some extent, for humans too, right? But it's more pronounced in other animals—how much of this kind of intelligence is innate and part of your instincts. You don't want your robot to just go around randomly doing anything. You want it to try to do things that might succeed. So there is room here for the more classical AI models, where we're training based on internet data, simulation data, synthetic data—everything that everyone else is doing—
Yeah.
That's useful to get you off the ground. But it doesn't fully get you there. It gets you to something that does something seemingly useful, and then you can experiment and have the robot really enter this interactive learning loop, where it's learning in the real world. That can get you somewhere—we don't know how far it can get you, right? We don't know yet.
And this whole topic of data gathering, it's amazing watching them walk around the building here and walk around the kitchen. They're so unintimidating. You walk right up to it intuitively. You don't feel like it's ever going to do anything awkward, hit you, or anything like that. So that's got to be incredibly important—
You say that—
The data gathering.
It's cozy. It's cozy.
It's cozy, and it doesn't seem to break the glasses or anything. So that's got to be really core to the data-gathering mission, right? Because you have to, as you said, let it experiment; otherwise, how's it going to learn?
3. Designing For Home
So, along those lines, what design elements did you build into NEO Gamma to make it suited for the home?
Sure. This actually goes all the way back to the founding of the company a decade ago. Really, I've been in the field for a long time.
How long?
Since I was a kid.
You were building robots at age what?
I was 11 when I decided that I was going to do humanoids and that sort of thing.
And what was the humanoid robot that you modeled? Was it Star Wars? Was it Star Trek? What was it? Lost in Space?
Honda ASIMO.
What's that?
Honda. Honda's ASIMO.
ASIMO?
ASIMO.
Yeah.
It's a beautiful robot, right? They started very early, and you can check out the Honda ASIMO piece. There are more modern ones, but the Honda Asimo P6 was like the end of the '90s.
Yeah.
Mm-hmm.
And that was walking up stairs—
Yeah.
—running around a stage, giving someone a bottle.
Yeah.
It greeted President Obama, I think, at one point. Yeah.
Yeah. That was a bit later, but yes.
Okay.
It was so ahead of its time, right?
Yeah.
But I built a lot of stuff through the years. Importantly, when I started the company, I sat down and thought really deeply about this: There are all these amazing robots that we worked on, and it didn't really work. Why didn't it work, right?
Mm.
Why didn't it work? It comes down to these fundamental principles. First of all, if you actually want to make something that's scalable with respect to intelligence, it needs to be able to live and learn among us. There are just so many nuances to this through everyday life, right? Everything we do is social. Work is social. Every task is social, and we navigate these social situations all the time while we do the things we do.
Most of the world's labor also happens in a social context, in that there are other people around you when you do it. Objects have social context, right? The coffee cup is empty. Do you need a new one? Is it dirty? Or do you want a refill, or do you keep your cup out through the day? There's this trove of diversity that you want to access.
So if you're a big believer in that, it boils down to this: The robot needs to be safe from a first-principles point of view, not able to harm people. It still needs to be very capable and needs to be as strong as a human. And then it just needs to be incredibly affordable. You need to find this beautiful combination where you can simplify, simplify, simplify, and still get a very capable system so that you can manufacture this at scale and really drive quality up and cost down, right?
That's really what we set out to do, and that's also why it took a decade, right? There's so much novel research that's been done in the company to get to where we have these tendon-driven robots that have—
So what's the vision there relative to the car, say? One in every household, two? You mentioned the iPhone—you know, go direct-to-consumer with iPhone sales, get to a billion, but it's exactly one per person. It's pretty obvious, right? But robots could be 2, could be 4, could be—
I've done that poll, and everybody routinely says, “I would have at least 2”—
Mm-hmm.
—depending on the price point, right?
So, price-point-wise, when I think about this, what I've heard is 30K, 20K. We've seen Chinese robots at much cheaper price points, but not as capable as NEO Gamma. Do you have a price point that you're thinking about?
Yeah, you're not far off. It's cheaper than what people think.
Okay.
It's quite interesting because I think this is very important. I want to make sure that we're not only making the best product; we want it to be price-competitive. I think that's going to be incredibly important. And we are actually still price-competitive with the Chinese ones.
Okay.
But you have to count, as you said, that it's not the same, right? So if you think about the number of degrees of freedom that the robot has—how much capability, basically how many joints—
Mm-hmm.
—then we actually have a significantly lower cost. So I think we've done a really good job reducing complexity to get there.
So the numbers, Bernt, that I keep in my mind are 30K to purchase, or 300 bucks a month to lease, 10 bucks a day, 40 cents an hour. Am I in the right range there?
Yeah, I think we could do better, but yes.
Okay. Even better. That's fantastic. And do you need to do better? No. Forty cents per hour—
I mean, in a heartbeat—
I'll pay that in a heartbeat. That's good enough.
Yeah, yeah. But in that case, I think people could imagine owning a couple of those robots. So I think it really depends on the lens you see this through. Clearly, everyone's going to want a robot, and I think there's this beautiful thing about the companion aspect of this, which is so underrated, right?
The humanoid is just such a beautiful interface for AI. When you talk to it and see the body language, it can look at you, it sees who's talking to it, directional audio—all these things. All my 11-year-old daughter can do if she has the robot is just sit next to it on the couch and talk about things, right?
Yeah.
And that is clearly going to be such a big aspect of it. I see it as—not another pet, but it's not another human either. It's something kind of in between.
Yeah.
And, like I said, it's kind of like Hobbes. If you ever read Calvin and Hobbes, it's Hobbes. I think it's going to be incredibly exciting to see how these relationships develop, because it's the thing that will be around you all your life, right? It will remember everything about you.
There are 2 things that really jump out at me if you compare C-3PO and that vision of an assistant robot and compare it to what you've actually built. One, it's soft. It's not a metal outside. And 2, the voice is perfect. When you're speaking to it, you immediately are disarmed and just talk to it, because it doesn't have a C-3PO robotic voice. It has just a perfectly soothing, normal voice, and it's very responsive to anything you say, any gesture, or anything.
Oh, thank you.
I imagine that these robots will all have advanced AIs at the level of GPT-5 or Gemini 3. And in so being, those robots will be hyperintelligent and able to fully understand and answer what you need.
And once they've learned the physics models fully, they can do whatever you need. You've made a decision to build your AI systems in-house, and I find that fascinating. In fact, a number of the other humanoid robot companies—not going to put you into a comparison mode here—have made that same decision versus partnering with the large hyperscalers. Can you speak to that?
Well, we're not doing the same thing. To me, intelligence does not begin with language. Language is this generative, artificial construct that we have come up with, and it's such an efficient, compressed way of conveying meaning and instruction. So language is very useful, but it's not the core of your intelligence.
The core of your intelligence is spatial and temporal, and it has to do with how you perceive the world around you. Both with respect to how you see the world, but also how you feel the world, right? And we're getting to where we're seeing that models that are native to that modality, and then you add text—
Mm-hmm.
—will be more intelligent and more powerful than language-first models.
I've read about intelligence, and the belief is that you needed embodiment for intelligence to exist and language for intelligence to scale.
I don't have rigorous proof that embodiment is needed. I do have very, very strong proof that, from an engineering perspective, it's just a way easier path, right? So if you think about the information in the world and whether you can access it, you could train a world model that can predict video and tell you, like, “Hey, here's a new video frame. Render this for you.”
In theory, you could probably train that only on text. If you have enough text descriptions of things, maybe at some point you could get a high enough signal-to-noise ratio that you actually can get something useful out, at least if you have some feedback loop with RLHF or something where you're like, “Am I happy with this frame?”
Mm.
But why would you do that?
Yeah.
That's just such an inefficient way of doing that. You, of course, train on video because you're going to output video, right?
Right.
So from that perspective, I think it's just obvious that you need all the modalities that we experience if you want to get to, first and foremost, human-level intelligence, and hopefully past that again.
But then I think there are 2 other things that are quite important when it comes to learning. The first one is quite obvious, and I think we all identify with this: robots can do interactive learning, right? You interact with the world, and therefore you can learn.
But if you think about it more from an academic point of view of how intelligence evolves, how do you get reasoning and all these things? What we generally do is have an observation of the world. We kind of know how the world works. So I know that if I do this, I know what is going to happen, right?
Yeah. Of course.
I've seen this before. So I actually start with that, and now I have a goal. I want to pick up the cup. So now I have a model of the world. I have a goal of picking up the cup. I take an action. I know which action I took. I know the action I took was to reach for and grasp the cup.
Yeah.
And then I observe the result. If you look at the internet, or in general, you can look at YouTube, right? All you have are the observations.
Yeah, right.
You don't have any of the mental model of the person in that video. You don't know which actions they took. You don't know what they tried to achieve. You only have the observation.
Yeah.
This is not how we learn. You can actually bring it all the way back to the scientific method. You should have a theory, come up with a hypothesis, test your hypothesis, observe the result, and then do it again and learn.
Yeah.
And that is just not possible with the internet data.
Definitely impossible with next-token prediction, raw internet scrapes, and all the video scrapes. So then, in these limited domains like coding and physics experiments, you can actually have that same experience, but it's only within that domain.
Coding is a good example. “Let me try writing it this way.” It didn't work. “Let me try writing it that way.” It didn't work. So you get very, very good at that narrow domain, but you still have no intuition about how the world works.
No.
You know?
You can do simulation, no?
Yeah.
So again, back to, it's hard to prove that this won't work, right? Sure, if you have a really good simulator and really scale simulation-based learning and simulation with agents, maybe you can get something similar.
But the fidelity of your simulator is nowhere near the real world. It's incredibly hard to get there and close that gap. It's also so compute-inefficient compared to just being in the real world.
But I think, for me, it boils down not to this academic exercise of proving who's right and wrong. It's more about what's the engineering approach that makes sense here.
Yeah.
And it's just a way shorter path.
You mentioned before in our conversation the amount of data that's being collected relative to Google, YouTube, or Tesla. Can you speak to that? Your mission is to get as much data as possible during the day from interactions with these robots in the home.
Yeah. You can do some napkin math, right? Of course, we don't know exactly what is the most useful data from which modalities yet.
But if you think about it, if you have 10,000 robots out there and they gather data most of the day, then that is more data than the non-duplicated, useful data that gets uploaded to YouTube each day.
Yeah.
So already at that scale, you have your fleet of robots generating more useful data than YouTube.
Yeah.
That's just at 10,000. And then if you think about how we scale manufacturing here as this starts deploying into society, you very quickly come to the conclusion that the internet isn't actually that big. You're going to have way more data from robots than you're going to have from the internet.
4. Scaling Robot Production
So I want to hit some numbers here just to set them as foundations. You built roughly hundreds of the NEO Gammas, but you're about to open a new manufacturing plant. Can you give us a sense of that? And then there's another one that's in the plans, right? Without disclosing anything you're not willing to, can you give me a sense of, by the end of 2026, how many you're manufacturing at an annual run rate, and then in 2027 and 2028? What's the growth path you imagine?
Yeah. First of all, just a small correction.
Okay.
We haven't built more than 100 of the NEO Gammas.
Oh, the Gammas. Okay.
But we've built more than 100 of the robots.
Right.
There have been multiple versions. The factory run rate at the end of 2026 is north of 20,000.
20,000—
Yeah.
—annualized.
Annual, yep. Of course, there's a ramp to get there, so you don't reach quite that number in 2026.
So a couple thousand a month.
Now, after that, we're trying to follow an order of magnitude, right? We're not going to quite be able to do that. I think the iPhone ramp is a very good comparison here, where you see that they almost double, but you have a few plateaus as you reach certain scales, where you run into problems.
There are some quite interesting problems if you're going to scale the manufacturing of humanoids to the iPhone level, right? You run out of some basic stuff like aluminum, for example. You don't use all the aluminum on the planet—that's not what I mean—but there's a certain percentage of the current aluminum-refinement capacity you can use before you start to really struggle sourcing aluminum.
Yeah.
That might be a challenge, I think.
Wait, was the iPhone ramp about doubling? That's an interesting statistic I hadn't even thought of. You get to 1 billion in the end.
It's more like 1.7, but—
1.7 annualized—
Yeah.
—over. Wow, that's not as fast as I remember.
So you can imagine 100,000—
Well, exponentials are quite powerful.
Yeah. No, I know.
We've heard the physics, actually. So you can imagine a run rate, before the end of this decade, of hundreds of thousands per year.
End of this decade, way more.
Way more at that rate.
Yeah.
Yeah.
At that point, you need to really think about what the things are that will slow you down, right? It comes down to refining—mining and refinement, of course.
Mm-hmm.
But increasingly, it actually comes down to labor. You’re not going to get there without really using robots for labor. If you think about the iPhone ramp, Apple displaced a large part of the Chinese population across the country for labor—
To parts of China.
—and they still ran out of labor and had to expand into neighboring countries.
That’s wild.
Now—
Yeah.
I think we’ve done an incredible job in the design, so there are very few parts. It’s very simple to assemble.
Yeah, I was looking for a new iPhone.
But it’s still more complicated to assemble than an iPhone. So let’s say it takes 5 times as long, which means we need 5 times as much labor as the iPhone.
I was going to say, it looks more complicated. I mean—
Yeah, it is more complicated than an iPhone, right?
Yeah.
Then you’re in trouble.
That’s a good metric.
Then you’re in trouble.
So it’s got to be solved.
You have to automate, right? And, of course, that’s the goal anyway. We want to get as quickly as possible to what I call this hard takeoff moment—
Yeah.
—where you have robots building robots, robots building out the data centers, the chip fabs, and the energy infrastructure.
So what can we learn from the car, actually? Here, you’ve got the iPhone: fewer parts, one-fifth the labor per unit. Then over here, you have a car. How does the part count compare to a car?
We have a few hundred parts.
Oh.
A car has roughly 50,000.
50,000. So it’s much simpler—
And, I mean—
—to get scale.
A car weighs 4,000 pounds.
Yeah, a lot of material.
Our robot weighs 66 pounds.
Okay.
I don’t think it’s really comparable to a car. I’ve seen a lot of people in the space compare humanoids to cars, but I think then you should go back to the drawing board, to be honest. It’s not a car. If you do a really good job here, it’s closer to a refrigerator.
All right.
It’s a very complicated refrigerator, but it’s closer to a refrigerator than a car.
Okay, so let’s dive in a little bit and shape our viewers’ and listeners’ understanding of the robot. It’s 66 pounds.
Mm-hmm.
Let’s talk about battery life and its abilities. Describe it from a specific stats point of view, if you would.
Oh, yeah, sure. I think the most important stat is that it’s huggable.
It’s huggable.
Yeah.
Yes, it is huggable. I have hugged a robot.
Yeah. Just the safety and how it feels to be safe in its space—soft. But from a pure stats point of view, it’s 66 pounds. It can lift about 150 pounds—
Which is amazing. In terms of the weight-to-strength ratio, it’s huge.
It is the weight-to-strength ratio of an athletic human. And then it can carry about 50 pounds around. That’s what you hopefully saw earlier here. Battery life is about 4 hours.
Rechargeable in—
Half an hour.
Half an hour? Or 2 hours?
Half of it, so about 2 hours if you use the full battery.
2 hours.
Yeah. Interestingly enough, I have one in my house, right? So I’m starting to get some data on this now—
Of course.
—and—
It’s 5'4", 5'5". What is it?
Yeah, 5'4", I think.
5'4", okay. That’s a perfect height, by the way, just in case you were wondering.
Yeah, it’s also the height of my wife, so I agree with you.
It’s mine, so that’s good.
Yeah. What’s very interesting is that once you start actually using the product, you notice a lot of things that don’t usually show up on a spec sheet. The robot is completely quiet.
Uh-huh.
That’s not a coincidence. That’s something we worked so hard on. The first time you put this in your home, you think, “The robot’s very quiet. It’s fine.” You put it in your home, and on the first day it’s fine. The second day it’s a bit annoying. The third day you’re like, “Oh, man, is it going to leave my living room soon?” Because of this sound.
Mm-hmm.
It’s such a requirement for it to be dead quiet if you’re going to have this in your space.
Interesting.
Charging-wise, I don’t really run into the problem because the robot just takes these micro-breaks every now and then when it’s not doing something.
Yeah, that’s good.
I actually don’t care that much about how many hours it can run. I care that it charges fast enough that it can always do whatever I want it to do.
Yeah.
Nice.
Well, I want to talk quickly about specifications, since you said that—
Yeah.
—the number of degrees of freedom, right?
Yeah, basically—
Which is basically how many joints the robot has, right?
Yes.
Humans have about 6 joints in each leg. That’s 12. If you have 7 in each arm, that’s 14 more. So now you’re at 12 plus 14—that’s 26. You see a lot of robots today that have 26. That’s quite common. Usually, they don’t have the wrists; they actually have the neck instead. So, 2 here, and then you’re at 26.
Okay.
We have 3 here, so you have proper expression with your head.
Oh.
That’s quite important.
Oh.
We have all 7 here. We have 3 in the spine, and then, of course, we have 22 in each hand—
I mean, what I saw in the arm design was incredible. Yeah.
Yeah.
How many does a human have in their hand?
22.
So you matched it.
Well, depending on how you count your carpal bones—the small bones that you have here that allow you to cup your hand—you could, to some extent, see that that’s more like 4 or 5 degrees of freedom, not really 2. So humans have a bit more. Functionally, it’s quite similar. This is incredibly important to be able to do all those tasks in a home.
But also, from an AI perspective, we talk about diversity initially, right? It is the one metric for intelligence.
Mm-hmm.
And the diversity of your data—
Diversity of environment and data—
Well, diversity of your data. Your diversity comes from 2 things, or the limit to the diversity you can achieve comes from 2 things. It comes from the environment you’re deploying in. If you’re in a factory doing the same thing every day, it doesn’t matter how good your robot is; it’s not going to be diverse.
Mm-hmm.
And then, how capable is your robot? How many things can it do, right? Because if it cannot do any kind of in-hand manipulation or handle soft deformables, all these kinds of things, or delicate objects or whatever, then you get no data from that. So you really have to go max, max on both if you want to maximize your diversity.
A geeky question for you, but I’m really curious to know: when you build something physical and then attach a neural net to it, it’s very hard to tell whether the constraint in what it can and can’t do is in the neural net or in the physical construction of the hardware. Is there any way to decouple that and debug the 2 different sides? Or is it incredibly impossible? Once they’re meshed together, can you just not?
Well, we have a pretty good neural net here.
Yeah.
Usually, the way I approach this is: can we do it in teleoperation? If we can, the right neural net can do it with enough data.
Interesting.
That’s generally been proven to be true. If we manage to do something in teleoperation, it’s like, “Okay, now we need a lot of diverse data of similar tasks, so we get some transfer learning, and we need a lot of data of that specific task.” Almost irrespective of how complicated that task is, you can get it to work.
Yeah.
Now, of course, that doesn’t mean you can get everything to work with generalization across tasks. We’re not there yet.
Yeah.
But you can see that, okay, you can get the neural network to do this. Now we need to scale it so we get this beautiful transfer of knowledge between tasks, out-of-distribution generalization, and all these things that we currently see in large language models that we don’t see that much in robotics yet. We have some pretty cool stuff internally where we see some signs of life.
What I’m picturing, though, is you ask it to make crêpes Suzette, or you ask it to do microsurgery, and it can’t quite do it. Then you say, “Well, look, the hardware guy is claiming the hardware is good enough. It must be the software guy.”
Mm-hmm.
And then the software guy is saying, “No, no, no, the software—the neural net—is fine.”
Mm-hmm.
“The hardware just can’t do it.” And then they fight it out.
Then we bring in our best teleoperator and say, “He can do it.”
Mm-hmm.
Then the hardware can do it, clearly.
Clearly.
It’s proof of access.
Yeah. Okay, so that’s where I was going. You have a remote operator option—
Yep.
—who can control the hardware. That’s really interesting.
Yeah. Then you’re again like, “Well, but we’re getting to where this gets hard, where we can no longer do this.”
Because?
Because the hands are just so good.
Yeah.
And they have very high-fidelity tactile feedback.
The human hands are so good.
No, the robot hands.
Okay.
The human hands are still even better, but the problem is that the robot hands are really, really good, and they have really fast, highly detailed tactile feedback.
Yeah.
We can’t really transfer this efficiently enough from the human.
Yeah, because the teleoperator—
So, yeah—
—is using some kind of clumsy—
Back to the XPRIZE, right?
Yeah.
The Avatar Challenge.
Yeah.
It’s a really hard problem to transfer that fast enough.
Yeah, you got that right.
Now we’re starting to see that the robot actually learns how to do manipulation much better from reinforcement learning in the real world. You actually have the robot interactively learn in real time how to handle objects, and it can do things that the operator could just dream of.
Ah, amazing.
Now we’re kind of screwed. No, we can’t do that anymore.
5. Autonomy Meets Teleoperation
All right. I want to talk about 3 things in sequence: teleoperation versus full automation, safety in the home, and privacy in the home.
Mm-hmm.
Those have got to be critically important as you’re entering the home. We saw the NEO Gamma out here operating in teleoperation mode, but also in full AI mode.
Mm-hmm.
Right? It was able to do both, and its AI systems are going to increasingly get better and more capable. Again, as I’m talking to Gemini 3, Grok 4, or GPT-5 soon, I’m talking to a highly intelligent human and getting a feeling that it understands what I want, and it’s able to take action on my requests. I imagine we’re eventually going to have that same level of AI in the robot, where I feel like I’m talking to a fully intelligent being, in one sense.
Oh, yeah. Clearly.
Yeah.
And one that is—
We’re very close already.
It is. And one that is grounded, right? That actually understands, to some extent, what this existence is. Today’s large language models have this abstract notion of it, but it’s a facade that quickly falls away if you start to probe at it. That will get there, I think.
In teleoperation mode, you’ve got humans wearing VR headsets and using haptic controls?
No.
What are the humans doing?
They’re giving slightly more high-level commands, just guiding it: “Hey, put your hands over here. Grasp this thing.” You don’t want to overconstrain a system. You want to give it some opportunity to solve how to do the task.
Okay.
We have the learning coming up from the bottom, enabling a more and more abstract interface for the operator. Then we have the learning from all the large amounts of data we have coming from the top, getting more and more of the general behavior that you want the robot to do. They meet in the middle, where—
So you’re using—
Gradually, the operator goes away.
—you’re using full automation and teleoperation always together in that regard.
Yes.
And learning.
Everything that enables the robot to do anything that the teleoperator does is fully learned end to end. The network outputs torques to the motors.
That’s very similar to what Tesla and Elon Musk were saying, where the self-driving car was originally all C++ code with a little bit of neural net—maybe 80% C++, 20% neural net. Then every year that went by, it became more neural net. And now there is no—
300,000 lines of C++ were eliminated. They’re gone.
Just a few guardrails left—
Yeah.
—and the rest is just one neural net. So, same thing here, I guess.
Yeah, it’s all weights. The code is just a few hundred lines.
Is it really?
Yeah.
That’s crazy.
Just all parameters. What’s the parameter count, or is that all super-secret?
It’s kind of secret, but it would be small if you compared it to today’s neural networks because it’s running on the robot very fast. It’s kind of like your muscle neural system. But it does take in vision, so it’s not very small.
Well, that begs a question I’m dying to ask. You’ve seen Ex Machina, right?
Mm-hmm.
When I saw that movie, I’m like, why—
Don’t go dystopian on us here, okay?
Okay. But why is the brain the blue blob in the head?
Yeah.
Why isn’t it in the server room?
Yeah, so learning is shared between all robots.
Well, yeah, and it can be much bigger. If half the power of the robot is going into the thinking, you could save energy. You could run twice as long on a battery charge if you moved it over to the server room and had it just communicate remotely. So why did you choose to put it in a head, aside from being anthropomorphic and cool?
No, no, it has nothing to do with that.
Okay.
There are some simple answers to that. The head is where nothing else is unless you put the brain there.
The room is not, yeah.
Everything else is pretty freaking full. Building a humanoid with this kind of power level in such a miniaturized form, while still having enough space to make it completely soft and all this, is a really hard engineering problem.
So where are we going to put this if we don't put it in the head, if we don't put it on the physical robot?
There are smaller arguments. The very high-bandwidth thing that happens in your brain is vision.
Yeah.
And to some extent, audio, smell, and tactile, but vision just dominates.
Yeah.
You just want to minimize the distance between your eyes and the compute.
For real? So the bandwidth between the sensors—the eye, mostly—
It's very high.
Wouldn't make it over the home Wi-Fi password?
Well, it wouldn't even make it down to the stomach of the robot.
Really?
Really?
Without getting overly complicated about which physical interfaces you would choose for this transfer, it's very high bandwidth.
I'm—
Wow.
I'm shocked by that.
I'm shocked by that too.
We're running no LiDAR, no structured light, no wrist cameras—nothing. We're running pure emulation of human vision, right?
Yeah.
We're relying so heavily on that, so it's very, very high resolution, very high bandwidth, and very high frequency.
That's funny, because that's exactly—
Yeah.
—where the human brain is very close to the eyes too.
It is. Now, that doesn't mean that you can't do things in the cloud, and we do things in the cloud, but it becomes hierarchical from an intelligence point of view.
If you think about your neuromuscular system, this runs quite fast, right? It usually runs at around 25 hertz. It doesn't necessarily go up to your brain. There are neurons distributed throughout your system that make decisions.
Right.
We have this in the robot. We have some of our stuff pushed to the power electronics that controls—
Just for latency, just for speed.
Yeah. And then you have the brain itself, which actually runs pretty fast, right? It usually runs at between 5 and 10 hertz. Even though it's 5 to 10 hertz, it has very low latency, and this runs on the robot.
Now, if you're running more like a 1-hertz streaming thing than, typically, an LLM's time-to-first-token latency, right?
Yeah.
That runs off-board, but it can't solve high-frequency tactile-feedback manipulation tasks. That's too slow.
Okay. The first time my NEO Gamma learns to crack open an egg to make an omelet, do all NEO Gammas then learn that? Is learning shared?
They do. There's shared learning in the sense that you can say, "This data goes to the cloud model that is doing this for all NEO Gammas," but there are also the distributed models.
Of course, there would be a nightly checkpoint where we say, "Hey, this model is better. We have more data, we've validated this, and we've established safety," which I'll talk about later when it comes to how you validate the models.
Mm-hmm.
Then we deploy that to all the robots.
Mm-hmm.
So even though it's distributed on the robots, they can still learn from each other, of course. You just need to do one hop through the server layer, do the training, and propagate it out.
Mm-hmm.
There is a future not so far away where I'm pretty bullish on there being a lot of federated learning happening on-device. This has to do with how we have your companion learn throughout life from all of the experiences that are dear to you, but private.
Yes.
So all robots will not be the same, but they will share an intelligence backbone.
Mm-hmm. Let's go into the conversation of privacy and safety. You're inviting these robots into your home, where there will be activities that you may not want shared with the world. Of course, you're asleep and the robot is running tasks at night. You don't want to wake up in the morning and find your safe has been opened and the robot's gone.
Or you don't want the robot to be taking care of your aging mother and find out that it's giving her shots of scotch at night when she asks for them. How do you deal with safety and privacy?
The last one is the hardest one, by the way. We can get back to that.
Okay.
Not giving Grandma scotch.
Shots of scotch.
Generally, models are always tuned to be sycophants, and they end up—
Yes.
—doing whatever you ask them to do.
Yeah.
If we start with the privacy side, I think, first of all, it's a lot about transparency.
Mm-hmm.
If you're one of the first people, like you, Peter, who will have a NEO Gamma in your house, we're trading a bit on privacy versus being an early adopter, because without the data, we can't make the product better.
Can't wait.
Of course, we're going to do everything we can to make sure this is privacy on your terms and that you're in control, but we do need your data if we're going to make the product better.
Sure.
So—
I mean, listen, I give my data to Google, Amazon, and X all the time. People don't realize that you're sitting in the home having a discussion with your spouse, and Amazon's Alexa is listening, right?
Mm-hmm.
Siri's listening.
But they're doing something very important, which we also do: no human in our company can hear or see that data.
Yes.
That is going into the training model, yes, but it doesn't go by a human.
Right.
If we want to look at that data—and sometimes you might need to, right? It might be, "Let's figure out what happens here, because something clearly is happening across multiple robots that we want to figure out"—then we'll send you a notification on your phone.
We'll say, "Hey, in this specific window, we want to review the data," and you'll get a video of what that data is.
Yeah.
If you say yes, then we get the decryption key and we can look at the data. If you say no, then we can't.
Yeah.
You're in control of that. Even with respect to going into the training data, we always run a 24-hour delay on training. If there is something that you really don't want in the training data—"This never happened. I want to erase it from existence"—you can go in and delete it before it gets into the training weights.
I just want everybody to hear that there are policies and plans that make this acceptable and are used by technology companies, and you're going to be implementing the best of those. That's a pretty bold compromise, actually.
All of them.
Yeah.
But there is something else. The mode I talked about now is what we call best-effort autonomy, which is most of the time. It's what you saw earlier today: if you talk to it and ask it to do something, hopefully it does the right thing. If it doesn't do the right thing, you can say, "Bad robot," and hopefully it's better next time.
It's actually learning in real time. This is really interactive learning. Interestingly enough, the robot progresses faster on tasks when it fails than when it succeeds.
Sure.
It learns more from failures, just as we do. In this mode, that's the privacy.
When it comes to teleoperation, of course, there's no way you can do the task without seeing the glass.
Mm-hmm.
We use some abstractions so that you don't see people. You just see blobs, and you see the object you're interacting with. We can do a lot on the filtering side to ensure privacy.
Mm-hmm.
But the most important thing we do here is that no one goes into teleoperation on your robot unless you approve it, right? It's very visible on the robot—the lighting changes, and it's like someone is in your robot.
Yeah.
It's one of the preselected operators that you have approved from a large set of operators. You choose, "Here are the 4 who service you." It's like inviting your cleaner, or whoever, into your house.
Yeah. Another human.
Another human into your house, and you just need to make sure that they're actually invited.
So, to actually take a second and spell this out in more detail: In the early days, when I have NEO Gamma in my home, it'll be basically autonomous, but there will be times when it needs to bring in a teleoperator. So you'll have teleoperators in headquarters who can step in if it needs help, is doing something complicated, or gets something wrong, and actually make the task happen.
Yeah. In the beginning, there are actually 2 different modes. So you have the mode that I call the best-effort autonomy that we just talked about.
Okay.
Then you have task scheduling, which my robot at home is doing now. I take my phone, schedule it, and say, “Hey, between these hours, here are the tasks I want you to do for me.” Today, it's: do my white laundry, and then there's a package coming from Instacart, so you can receive it at the door and unpack it in the fridge, and it's just generally tidy.
Mm-hmm.
And I've given it the hours when I'm not home: these hours, I'm at work. Just get it done, right? Now, I don't care if that happens autonomously or through a teleoperator. A lot of that happens through a teleoperator because some of these tasks are quite complicated, and we don't know how to automate them well enough yet.
Yeah.
Now, of course, the teleoperator uses autonomy to help improve the efficiency. So it's not all teleoperation, but I don't really care about the mix. The task gets done.
Yes.
So we kind of split it like that, and then there's the gray zone in the middle. If you want to have your friends over for a party and you want the robot to be the bartender, and we don't have a bartender mode yet, then you can approve a teleoperator to do that. So you can also schedule it.
Mm-hmm.
And most, if not all, of the videos we see of Optimus at Tesla's diner or at their events are teleoperation modes.
They are. But I think teleoperation has gotten this underserved, bad reputation.
Why?
Because I think people don't have enough clarity: Is this teleop or is it autonomous? But it is just labeled data. It's expert demonstrations, right?
Mm.
If you look at any of the big AI models that were trained, an enormous number of people sat down and hand-labeled data, looked at examples, wrote out question-and-answer pairs, and bootstrapped a very high-quality data set for this to work, right?
Mm.
So you pre-train on general information. We also do that—just everything that's happened with the robot. Then you have a fine-tuned data set that's very high quality.
Yeah.
And in robotics, that is teleoperation because it's the expert demonstrations. It's the hand-labeled data.
Yeah.
So there, it's no different. I think there's a lack of transparency in what's going on.
Well, I think the objection is: if you have a demo, like a video, that makes it look like it can do something, and it actually can't because you hand-coded it.
Mm.
Well, it clearly can, but it can't do it autonomously.
Yeah, it can't do it. But yeah—
But—
But I think you're dead on that if it can physically do that, if the mechanism can do that, the neural net will fill in that blind spot instantly anyway.
Yeah.
You know, once you've trained it. So I think it's perfectly legit. I want to jump into another fun subject, which is the uncanny valley and the face. Talk to us—you've probably had endless conversations internally about how do you make a face look, how human do you make it, how skin-like do you make it, how do you represent it. Can you tell us philosophically what you and Dar [?], who's on your design team, think about that? Where do you make it human enough?
It is this very delicate line where you want to make sure that body language comes across crystal clear. Because that's the magic of the device, right? Of the companion.
But at the same time, you don't want it to get to where your instincts tell you, “Hey, something's wrong. This is a human, but there's something wrong with it.” So you don't want it to be a human. And it is actually pretty surprising that there is this gap where people clearly identify this as, “Hey, this is a being I identify with. I understand its body language and everything, but it's also clearly not a human.”
Yeah, yeah.
And you want to be in that space. Where you are in that space depends a bit on who you ask, right? People have a different threshold here.
So we're trying to hit in the middle of that and ensure that, for as many people as possible, this is just an incredibly easy-to-understand product. But at the same time, it's not creepy.
Yeah.
And I think adoption here—by the way, talking about scale—is so important. Adoption of new technology usually takes some time because there's this knowledge barrier, right? There's a barrier to entry. Even using a phone, there's a barrier to entry.
This interface is just so natural. There is no barrier to entry. It's something you just talk to, like a person.
It is. What's incredibly cool to me is that there are 50 things around the house that I don't know how to do, including the fricking way to backwash the pool. All this crap. The robot can, in real time, access the information, learn how to do it, and just do it.
Yeah, from the web.
I can't do that. It would take me an hour to study, and there's no laborer who's going to come into the house and do it for under 400 bucks. So there are so many things that are in that category where I'm not trying to replace a human being; I'm doing something that there literally was no other option for because the knowledge is obscure.
There are so many of those things around the house now, like resetting the water heater when it keeps going out and the reset process. But you can look it up. The robot can look it up. And just go do it.
This is micro-units of work, essentially, right?
Hyper-specialized micro-units of work.
Yeah. And you need 5 minutes of it every now and then, and it's super high value to you. It's just really hard to get to.
Or a fricking Shop-Vac. The Shop-Vac—you can run it forward or backward. There's a manual there. You could read the manual. I just want to get this crap off the garage floor. The robot will know how the Shop-Vac works because somebody else's robot—
Right.
one of the other 10,000 has already done it.
Make my perfect teriyaki salmon on the grill.
Yeah, some obscure mixing, some food. Yeah.
Which brings us to something that you talked about earlier—
The Scotch for grandma.
That we kind of dodged: Scotch for grandma and safety.
Yeah.
So, I do hope to make you a perfect salmon teriyaki.
Thank you. Appreciate that.
But I have to do it myself. I'm not going to let the robot do it because that's one of the things we're actually not doing at launch, and that is due to safety.
Yeah.
Mm-hmm.
Because what I worked so hard on for this decade is to make robots that are safe intrinsically.
Mm.
What I mean by that is, if something goes really wrong and it accidentally hits you, that might be painful, but it's not likely to severely harm you.
Injurious, right?
Yeah. And once you pick up a kettle of boiling water, there's no more guarantee that you are safe, right?
Mm-hmm.
Mm-hmm.
So we generally avoid any kind of dangerous objects so that we can ensure safety in the beginning.
Yeah.
Of course, over time, as the AI improves and we get more and more certainty that all behaviors are safe, we will allow cooking and other things. So we're doing internal projects on this, but we're not going to be rolling it out to customers in the beginning, just due to safety concerns.
Yeah.
Um—
Yeah, cooking—
So—
Cooking and safety is a—
It's a real problem for humans.
Cooking around fire is not an easy thing.
I mean, it is a real problem for humans, too.
Yeah.
It is. But there's the notion of intrinsic safety. This is incredibly important, and that is the safety of the AI. This is the reason we have a white paper out on this that I might recommend if you guys are interested—read it.
6. World Models Make Robots Safe
But why we started very early betting extremely heavily on world models.
On role models.
World.
Oh, world models.
World models.
Yeah.
They are, of course, the currently best-known path towards AGI. But even more importantly for us in the short term, as we progress here on the data collection and model training, they give us this incredible opportunity to automate evaluation of models, including safety and red-teaming and all these things.
Mm-hmm.
So you can think about, if you train a new model and now you want to know if it's better than the previous one—
Mm—
—and you can deploy it to all your customers and get some vibe check a few days later, like, “Hey, are people happier now?” That's generally how it's done.
Yeah.
You don't want to do that with a physical system, right?
Mm-hmm.
You can't do that with an autonomous car either.
Yeah.
What a world model actually is, is a model that is able to generate what will happen if you take specific actions. So you can think about a video model where you ask it to do something, and then it actually gets not only the question of what to do, but the actions to do so. It gives you back not just a video, but how the world feels—the forces, everything for the robot. So it's essentially like the robot's in the Matrix.
Yeah.
We take the robot, we put it in a world model—
Yeah—
—and it doesn't know that it's in a world model. It thinks it's in the real world.
Yeah.
And it does its things, and we ask it to do the things we're usually doing around the house.
Yeah.
And we see what it does, and we can put in lots of automated checks to ensure that it's both performing better and not doing anything that could be deemed unsafe. So it's really—
It's wild.
—this incredibly important and powerful evaluation tool, and that's a start to solving the problem, right?
Do you think that's why you guys, Figure, and Tesla are getting monster valuations? Is the valuation purely, “Hey, we're going to sell 10,000, 20,000, then 200,000”? Or is it that the world model is such a unique asset and so valuable in thousands of different ways, and that becomes a very much self-feeding barrier to entry? That could also explain it. Do you plan to productize that core capability?
Yeah. To me, it's this: our mission is to create an abundance of artificial labor, and that goes across both the digital and the physical.
Mm-hmm.
So yes, it will be productized.
Mm.
Still a bit out, but yes, this will be productized. And—
Yeah, because it seems like no matter how much factory capacity you build, it wouldn't be until 2028 or 2029 that you could diversify into all these things, like microsurgery, warehouses, drones, all that. But that same world model could apply to those much sooner, but you'd have to somehow get it into the hands of many other companies.
Well, I actually think revenue from the robots will dominate forever. I do think the real physical world has—
I mean, it's—
I mean, just for folks to realize, we're at $110 trillion in global GDP, and labor is half of that, right? So the TAM—the total addressable market here—is $50-plus trillion.
Just—
Yeah.
Just if you keep doing what we already do—
You will achieve capability—
—but you do what we'll be capable of doing, yeah. It's—yeah.
It's going to be so much bigger.
Yeah.
And you've attracted some incredible early investors. Do you mind just sharing who has come into your cap stack?
I think we have some big classical ventures like SoftBank, Tiger Global—
NVIDIA—
—EQT, NVIDIA, OpenAI.
OpenAI. Yeah.
So there's some good names in there.
I mean, that's damn good.
A lot more.
Mm.
I think it's becoming increasingly clear that the bottleneck in society—
Mm.
—to superintelligence is not better algorithms or scraping the internet in a more thorough way.
It's better data.
It's—yeah, it's better data, and then you need robots to generate this data. But even more importantly, it's the physical parts, right? You need more data centers. You need more power. You need more labor. And to do this, it's kind of like this: it's a bootstrapping problem.
If you break down the pyramid and say superintelligence consists of this incredible amount of data and this substrate of compute and power, then you see that humanoids are a solution to both of them.
Mm-hmm.
And if you just do the math, you'll see that you're probably not going to get there without that. You're just running up against these basic constraints. I think humanoids will be surprisingly useful, surprisingly fast.
Okay. I—
Not perfect, but they're going to be surprisingly useful surprisingly early.
Yeah.
7. Humanoids Fit The World
I have a question on behalf of Salim Ismail, who's typically our third moonshot mate here. I have to ask on his behalf.
No.
So Salim is constantly saying, “Why 2 arms? Why 2 legs?”
Mm-hmm.
“Why not 6 arms? Why do we need to have a humanoid form? In the kitchen, wouldn't it be better to have an extra pair of arms?” So what's the definitive answer to him?
Well, I think first of all, he's kind of right. Humanoid isn't the only thing that will work. I do think that—and we've looked a lot—I don't know of any form factor that is as general as a human in doing any kind of labor in any kind of environment.
We've tried to simplify. We've tried to increase complexity. The human is a pretty good machine.
Mm-hmm.
If your goal is to just be as general as possible, then you need a humanoid. Now, if your goal—
I think that's the most important part of the equation there.
It's very important. And then—
You're not going to transfer to a 6-armed robot learnings from a human.
No. It's at least harder. And then—
Mm—
I mean, the world is made for humans.
Yeah.
It's what Jensen says, right? It's brownfield deployment. It's very true. And then I think lastly, it's just: do you want to live with a 6-legged robot in your kitchen?
No.
But I view humanoids as kind of the pinnacle of general technology. There is this repeat pattern through history of this happening with zero-to-one novel products. So if you think about, let’s say, the computer, it started with big mainframe computers solving very specialized tasks.
Mm.
The equivalent in robotics would be industrial robotics, right? Now comes the PC, or even before the PC, like the VIC-20s or Ataris or whatever—more general computers. This gets produced at such a scale that it just becomes generally available. Now it’s super high quality, incredibly reliable, has this huge ecosystem, and becomes the best way to solve any problem.
Yeah.
And here’s the argument against humanoids, right? It’s overly complicated for a task. Even though it’s overly complicated for the task, when you take your beautiful Apple here and write a Word document, that’s the most complicated typewriter I can think of. Humanity mastered nanoscale chip manufacturing for you to have a typewriter.
Yeah.
But it’s still actually the cheapest, most reliable typewriter because it’s made at such a scale.
Yeah.
Humanoids are exactly the same. Now, if you see what happens to computers, because the market has become so big, it starts to become segmented again. Now you can carve out niches in computing, and they’re still so large that they have scale. So now you get specialized compute for AI, specialized compute for physics, specialized for all kinds of things, right?
Yeah.
And this is because it’s become so big. The same will happen in robotics. We will get to where we have Star Wars. There will be different drones doing different tasks, and they will look more specialized, like my repair drone with six arms and scissor hands and I don’t know what else.
Yeah.
It’ll get there, but you have to go through this humanoid phase first. So let’s just say humanoid is a phase.
Mm.
Yes.
My favorite robot is still Data from Star Trek.
It’s a great robot.
Yeah. It’s kind of the closest thing I think of to what you’re building: a lovable, happy robot that you can give a hug to.
Yeah.
Do you have a favorite robot?
I wouldn’t have thought of Data, but now that you said Data, that’s top of the food chain. Everybody loves R2-D2 because, for some reason, R2-D2 has no voice even though they have voice technology everywhere.
It squeaks. Yeah, it squeaks.
All those visions, though, are built around what Hollywood could easily get on a set.
Yeah.
I think the humanoid form factor, though, has another aspect that you kind of touched on. When I bring it into my house, I have a vision of what it can do and what it can’t do based on humans.
Mm.
So I ask it to do things that are rational and not irrational because I know what a person could do. If I had a six-legged thing that Salim came up with—
Yeah.
—I’m not quite sure. Should it be able to climb on the roof and fix the shingles or not? I don’t know what this thing’s capabilities are.
Interesting. Yeah.
So it breaks the whole comfort zone of expectations.
Expectations, yeah.
The thing that does surprise me, though, about the robots is how unbelievably coordinated they are between themselves, and there are some good demos of this at MIT that are just mind-blowing. When you have 2 movers trying to take a couch up the stairs, they’re like The Three Stooges, right? They’re saying, “Oh, move it to the left. Now move a little bit.” When you see the equivalent act with 2 robots, they’re just in phase and do it seamlessly. So I think there’s a very high probability that the standard in the home is going to be 4 or 6, if you get the price point down a lot. They work so well in concert with each other, it’s almost a crime not to have that teamwork synergy.
Huh. It seems like a bit much to me. Maybe, if I need to have some movers, I can ask my NEO Gamma, and he’ll invite some friends over.
You’re not taking everything into account, Peter.
Yeah.
Because you have to remember that by the time you have this many robots in your home, everyone’s homes are really freaking big.
Yeah.
We have an abundance of labor.
Mm.
Your house is not going to be this small—
I’m trying to think.
Everything is going to be huge.
So labor is going to continue to demonetize and democratize.
Let’s go someplace that I’d love your insight on, which is China. When I think about the robot industry, I’m tracking 50-plus well-funded humanoid robot companies in different stages around the world. The majority are in the U.S. and China. You started in Norway. There are some in Europe, some in India, parts in Japan, and Korea.
But China, by far, I think, is dominating. What I see there with the Robot Olympics and special robot villages is pretty extraordinary, where the Chinese government is really accelerating this for obvious reasons. They need access to low-cost labor to continue the manufacturing boom, and they need it to support their elderly population. How do you think about China? What do you think of the work coming out of China?
Well, first of all, I think we need the same thing here.
We do need the same kind of support.
This is not something we realize as much, maybe, but of course we need the same thing.
Yeah.
I think the Chinese ecosystem is incredible. I don’t know anywhere else in the world where you can develop hardware as fast. You need something, and you go over and get a machine on the corner here, something broken—
A Shenzhen-based—
—PCB, and you just go over the street and buy some new components. There’s someone doing a reflow over on the street corner over there. This is an incredible ecosystem.
Mm.
And I think the Bay—I say “the Bay” now, but I know it’s also, I mean, Silicon Valley. The hardware bay in the Shenzhen area is also a bay, but Silicon Valley, this bay, we have a long way to go if we want to really get to the same level of rapid iteration on hardware.
Yeah.
That’s just incredible. I think the manufacturing part is incredible. There’s so much process knowledge.
Mm-hmm.
And I think this is highly underrated. Think about magnets.
Hmm. I do.
So do I, a lot. We have great material scientists who know how magnets work and can design very good magnets. But then we lack that guy who knows that, yes, you do all of that stuff they told you in the books, but after 2 hours you have to stir to the left, not the right. There’s just so much of that.
Mm-hmm. Yeah.
And this is just so disseminated in China. There’s so much process knowledge.
Yeah. How did that evolve in China and not here? What’s the cause? Top-down incentives.
Just funding, or—
I think it’s the government saying, “You’re a robot city, you’re a neodymium magnet city, you’re…”—and just capital and people, and communist-directed, but then allowing companies to build on top of that.
Mm-hmm.
I mean, is that what you see as well or not?
I’m not sure. I think the Chinese startup community is very alive, right?
Mm.
And the capital is quite alive. I feel it runs very similarly to what we like to think of as the Bay here—
I mean, I used to take a group of investors every year to China, and we would go and visit Shenzhen, Shanghai, Hong Kong, and Beijing, and meet with Baidu, Tencent, Huawei, and the leadership of all these—
And there was a super vibrant entrepreneurial community, right?
Yes.
The mindset was 9:00 AM to 9:00 PM, 6 days a week, and that was a great lifestyle. You considered the 1.3 billion people in China your market, and the 300 million in America your market as well. But there was a fall-off after 2019, and there was a real dip in that ecosystem. I think it's beginning to reemerge, but I think the government is really pushing hard on supporting AI. Humanoid robots are an embodiment of AI. Obviously, you know this, they're closely meshed. I do think there's a lot more support that the US government needs to give to US—
100%.
—hardware companies.
But I think the most genius thing they did was the special economic zones, like the free economic zones.
Yeah. Sure.
It's not that people here don't want to build stuff. We want to build stuff.
Yeah.
It's just that it takes too long, costs too much, and it's too convoluted, right?
We do it in spite of the challenges.
Yeah. I think the US should just spin up some special economic zones. Here you have expedited permitting and—
Well, in California in particular—
Yeah.
—it'd be a no-brainer to do that. That's the simplest, best idea ever, but I don't know what it would take to get it through.
No, but this is some of what people are working on these days, right? If you look at Masa Stream, Project Crystalline, for example, it's very similar to this kind of free economic zone in the US.
Mm-hmm.
I think there's another problem you need to solve, too, though, which is that the US just did software forever. We were not only not doing hardware, we just didn't do chips. The chips—
Well, not forever. We're in Silicon Valley.
Yeah.
Yeah, where's the silicon?
There was—
You're not making silicon in Silicon Valley.
Yeah, yeah.
Yeah, yeah. What's up with that?
There was a phase in between where people lost the way.
Yeah.
They lost the plot.
Yeah.
And now we have to find the comeback.
Well, the venture community got all messed up, too, because they wouldn't fund you if you had a physical component in your business plan. They'd be like, “Well, yeah, I'm looking for the next Meta or Google.”
Yeah, hardware's hard.
They don't really care. Yeah, hardware's hard.
I've been keeping this company afloat for 10 years, so I can—
You've got scars.
—I can go on and on about—
Yeah. Do go on and on, because it needs to get fixed.
But people are so afraid of hardware.
Yeah.
Right?
Yeah.
But—
It's going to kill us if we don't find a solution, and quickly.
I think it's also going to kill VC, to be honest. If you look at returns on venture, it used to be incredible, right? If you got to be an LP in a venture, you're like, “Oh, man, I'm set, right? I'm going to make the big bucks.”
Yeah.
And now it's more like a philanthropy thing. You want to fund startup entrepreneurs because venture doesn't really make that much money. I think it has a lot to do with—
Except at Linq. You guys are doing amazing. We are doing amazing.
That is good.
Yeah.
That's good.
I know the data.
But you guys touch the hard stuff. The point is, if you don't touch the hard stuff—
We touch the ear—
—what's your moat?
—we touch the early stuff, right? So it's first checks into companies that then are scaling rapidly versus companies that are doing late stage.
But we're not. We're doing first checks into hard stuff at the seed stage, but we're not doing hardware. So I'm as much a part of the problem.
Well, you should look at the biggest companies. They all have hardware.
Yeah, I mean, listen, Elon Musk cracked the code on that. He's been able to just make hardware sexy and has generated incredible returns.
Yeah, and I think Jensen Huang says it really well, right? They want to work on the really hard problems that are super painful, that you are uniquely capable of. You know that your competitors have to go through the same pain or more, and they're not going to be willing to take as much pain as you. This is how you win.
Yeah.
And things here are actually defensible, right? The moat we have on hardware—
Yeah.
—that's years. The moat we have—I'm incredibly proud of our AI team, by the way. We've accomplished some things that are so amazing on such a budget.
Yeah.
So we're way ahead of everyone else in what we're doing on world models. Let's say we're 3 months ahead, right?
Way ahead.
Well, that is way ahead.
But you're incredibly rare. All props to Elon Musk—he's incredible—but Elon's pathway to getting to hardware was through PayPal.
Yeah, sure.
He burned a couple hundred million dollars himself, got down to near bankruptcy, and was almost dead on both of his big companies, Tesla and SpaceX. He barely pulled it out, and then made them huge, but the VCs were not touching it.
No. He had to borrow money in 2008 during a divorce, with SpaceX having its third failure and trying to borrow money for its fourth launch.
Yeah, that is not—
Yeah.
—a repeatable funding model for America.
I was very lucky. I had a very good early founding investor. The company didn't start in a garage because we're not Silicon Valley. We started in a barn because we were Norwegian.
Yeah.
At some point, 2 years later, he sold the farm, so we had to move. He sold the farm to fund the company.
Huh.
Right?
That's amazing.
Yeah.
This company in Silicon Valley wouldn't exist without—
No.
—a Norwegian investor—
No.
—believing in your vision.
And we wouldn't exist because we wouldn't have had the runway, right? Operating this in Norway was just incredibly cheap compared to operating here.
So your initial Norwegian investor, did he or she believe they were going to make a huge amount of money, or did they do it because they were passionate about your vision and your mission?
Or did they believe in you?
Yeah.
I think it's all three. It turned out pretty well.
Yeah. Well, yeah, but—
But—
—it wasn't here. That's the point I'm making. It's like—
No, but there are different phases, right? If you want to scale something, you have to come here.
Yeah.
I think you can do deep research in other parts of the world. There's talent everywhere. But really hyper-scaling that and getting it across the finish line—
Yeah.
—that's here.
Did you consider LA, Austin, Florida, versus here in Palo Alto?
Yeah. We even had manufacturing for a little time in Texas, in Dallas. There's just no comparison to the talent.
Something in the water.
It's the talent pool.
Yeah.
There's talent everywhere, but the density of talent matters. There are different types of talent because when you have a zero-to-one field like this, in the beginning you have a lot of really passionate people who have been working on this all their lives, and they're so good at, in this case, humanoid robotics, right?
Mm.
I remember back in the day, if you went to the Humanoids Conference, everyone could fit around one small table. Those people are still the ones who are at most of these companies, right? Those people don't know how to make a great product. They don't know how to scale that to a million or a billion devices.
Yeah.
They don't know how to write incredibly good APIs for the software to support the ecosystem. They know this thing and do deep research. And now your field comes of age, and it's time to actually do this because the timing is right. We purposefully stayed very small for the first 7 years, just doing core technology.
Now, suddenly, you get access to this talent pool of people who go from field to field, whichever is the hottest thing right now, and just do it again and again and again and again. That's Silicon Valley, right? But there's been an incredible inflection point in humanoid robotics.
I remember we had the ANA Avatar XPRIZE, which had teams build robotic avatars that you could telepresence. I remember the finals; we had good teams. I know some of your team members here were part of those teams, but it's come 1,000x since then, really in the last 5 years.
Is it the AI models that have made that? What's caused the inflection in the last 5 years? The AI is clearly part of it. There are things we do with AI now that we couldn't do 5 years ago. I do think we saw the breadcrumbs, and we were on the path already then, but it wasn't working yet.
I think you just hit this critical mass of accumulated innovations that have happened in hardware. I do think it's important to note, though, that it's hard to see what is real innovation and what is not in any field, especially in humanoid robotics. You can go on YouTube and find things from the early 2000s that look better than most things you see today from humanoid robotics companies.
You can't just make a beautiful robot that looks good. You have to actually make a robot that is safe, that you can manufacture at scale for a very affordable price, and that is still capable. I think that's been the main unlock and the challenge. You need to get those things right, and that just takes a lot of time.
I think the neural nets are light-years ahead of anything anyone would have predicted 5 years ago. Then the hardware—the NVIDIA chip that it runs on—is getting pushed as fast as any innovation in history because the demand is through the roof. That part is well understood.
On the physical hardware side, what's something that you use today that you couldn't have used 10 years ago? What's improving in the motors, the harnesses, the electronics, and the batteries?
Great question. I think mostly it's been on the motors and material science side. We make our own motors, including not only the IP for the motor, but also the manufacturing and automation for all of this and everything that goes into it. The supply chain is so broken.
You literally make your own motors.
We have to wind the wires. We do it kind of special; it's the 1X version of this. Motors are one of the things we really innovated in, and this is actually how I started. When I sat down a decade ago, the first thing I did was design a different kind of motor.
The motors we have now in Neo are 5.5 times the world record in torque-to-weight ratio.
The product does that, too.
And that's why we have something that's so powerful that we don't need gears. We can just pull on these tendons to loosely simulate human muscles or tendons.
That's why it's so light.
It's also why it's so drivable and compliant. It's why it's so cheap to manufacture. Everything kind of comes from this.
Now, of course, when you have these motors, you can start using tendons. But then you need to sink a lot of time into figuring out how to use tendons, and then comes all the material science to have tendons that can last millions and millions and millions of cycles.
These are really hard research problems. They're not even engineering problems; they're hard research problems. We spent so much time figuring all that out. You can't make the motors that we make without doing some pretty significant innovations in electronics, how you do power amplification, and, in general, motor drives.
There are a lot of things that come together. You couldn't have designed the motors we do today without some of the innovations that have happened in magnetics. Of course, you couldn't have done it without AI, either.
The first thing I did back in the day when I sat down was program a network to learn how to make motors.
You designed the motors via AI? How long ago was that?
It's a bit more than 10 years. It wasn't Transformers, but it doesn't matter. For that kind of use case, it was a neural net.
8. Abundance Transforms Civilization
Bernt, you think about robots in the world probably more than anybody else. What's your vision 10 years from now? What are we seeing? What does abundance in labor enable that goes beyond people's initial reaction to how I would use a robot?
I think, first of all, what will happen is that actual abundance means everyone can have whatever they want. But not only can you have whatever you want, you can have whatever you want in a sustainable manner. Sustainability is something we lose when we cut corners to shave costs.
If you actually have an abundance of energy and labor, why would you not do things sustainably? Then I think the next frontier that comes after building out the infrastructure across the globe that allows everyone to have an incredible quality of life is: How do we solve the remaining really hard problems in science?
I think this is not going to happen without humanoids, because you need to build particle accelerators. You need to build enormous biotech labs where you're doing experiments with really intricate networks. You need to do all the experiments.
I think it's almost existential to us for human happiness. I don't want the godlike AI in the sky to be directing all of the planet's inhabitants around with their glasses to do experiments for it to solve science. That's not the future we're aiming for. We want to have this beautiful symbiosis and co-invention between man and machine.
Yeah, that particular use case is so acute, where Dennis Aassovis is working on the full-cell simulator to try to close the loop. But you know that you're going to need people to mix a huge number of chemicals to truly unlock longevity, health, and chemistry. The humanoid robots can do the work because everything in the lab is oversized.
No, not only can they do the work, I think this is a common misconception. Humanoid robots will do a lot of the work initially, but once it gets to a certain scale, the humanoid robot will make the automation system that will do the work.
Because humanoid robots will not be machining new parts with a Dremel, right? You will use the CNC machine.
Mm-hmm.
Humanoid robots will not be moving car chassis around by carrying them with 30 humanoids. Clearly, this does not make sense, right?
Yeah.
We have existing automation systems, and we will build more. What humanoids will do for you is build all these automation systems and get them up and running, and then cover the remaining gaps that you currently can't do with humans.
Yep. Yep.
How are you going to do it—
Such an unlock.
How are you going to do it in a vacuum? I want my Neo Gamma to help me set up my space station or mine my asteroids.
Well, we can go on and on. I think, first of all, we have a huge advantage because the robot is so light.
Yes.
I guess Elon's working on this, but payload to orbit is still expensive. Secondly, most of the stuff we have actually works pretty well in space. We have to do some work with the epoxy on the motors; that's not going to be very vacuum-hard.
If you want to train in zero G, one of my companies is called Zero Gravity Corporation. They do these parabolic flights.
Oh.
Yeah, we flew Stephen Hawking in zero G. Maybe Neo Gamma should come next.
That would be great. I actually do think there are real use cases for this, and one thing is building a base on Mars or whatever, right?
Mm-hmm.
But even before we get there, just in-orbit assembly is this extremely high-value task, and I think there we will use teleop.
Yes.
The reason I'm saying that is just that the cost of mistakes is so high that you want to use the smartest, most expert humans you have. Until we get to superintelligence, that will be a human. You have people in orbit, you have robots outside, and there's very low latency—
Yeah.
—and you can teleoperate in a very natural manner, as if it were your own body, to do all of these in-orbit assembly tasks.
And they can be incredibly complex, and you can still do them with very high accuracy, and you're not endangering people.
Yeah.
And of course, when you've done this for a while, you have the data to automate all this, which is very interesting.
Yeah. That's where your weight advantage would be really amazing, too, because you can take 5, 6, 7 of these up.
And the energy efficiency.
Yeah. You guys are gonna have to somehow bleed off your heat, right?
Right.
It's really hard.
Yeah, that's right.
You must be looking to hire people.
We are.
What kind of people watching are you interested in potentially hiring?
People who are really mission-driven, who really believe in the beauty of a world where we have an abundance of labor.
Mm-hmm.
And like to solve really hard problems. People who can also demonstrate that they've solved incredibly hard problems.
Mm-hmm.
Because that's what we're doing here, right? Everything from materials science all the way at the bottom, all the way up to the foundation models at the top. And I think what we offer is just this incredible place to work.
Mm-hmm.
Not with respect to work-life balance or any of this. We're not quite Chinese, but it's a hard problem and we're in it to win. But probably the place on the planet with the most experts across all different disciplines in science.
So if you come here as a mechanical engineer, you will learn so much about AI, about electrical engineering, about batteries, about materials science, everything else. And it doesn't matter which discipline you come from, right? You will learn so much from the people around you, and I think also that's one of our biggest strengths, how we really always work in these multidisciplinary groups.
Yeah.
And we find the good solutions between the disciplines. It's like, “Hey, you don't need to do that. That's kind of costly in manufacturing. I can calibrate that away.”
Yeah.
Or, “You don't need to calibrate. This doesn't cost any more.”
I can see that, actually, when you're walking around the building here. Dean Kamen's lab in New Hampshire is very, very similar, where he was the Segway inventor, and everybody's happy.
All the MIT people that we know who work with him, they're just happy. And the reason is because when you do software, you're largely behind a workstation all day. You're sitting. When you're doing physical things, you're moving around a lot more, and you're building and making, and it energizes you all day long. It's just such a fun work environment.
It's so obviously tangible, just walking around and talking to people. So it's a good lifestyle.
And it helps when there's a lot of robots walking around with you.
Yeah, for sure. People can go to 1X Technologies' website to—
That's true.
—to go and find out what positions are open.
Yeah.
Yeah.
For sure. And one thing I'm excited about to announce is, you and Dar and the Neo Gammas are gonna be at the Abundance Summit in March.
Yeah.
Can't wait.
Yeah.
Meet a lot of great people.
Yeah, so our theme this year is the rise of digital superintelligence and the rise of humanoid robots, because the two are going together.
Sounds pretty spot on.
Yeah, I think so.
Yeah. Yeah.
I mean, it really is. And without making any promises, I'm hopeful we'll have a number of the Neo Gammas there, interacting, living, and hanging out with the Abundance members.
Yeah, how do they get there? You buy them an airplane seat? They just walk on? You sort of—
Yeah.
You don't box them up, do you?
Yeah, we're down in LA, so—
Yeah, we're probably gonna drive down to LA.
Okay.
It's easier than getting them on a plane.
Do you put them in the seats and strap them in, or—
Yeah, we do.
That's so funny.
Actually, at this point, they're starting to sit into the seat themselves. So it doesn't strap itself in yet, but that's coming.
It's a funny story, though. We put one of the first robots on a plane back in the day.
Mm-hmm.
We were rushing back home from China. It was a proper startup story. It was way back in the day. We were running out of money, and we hadn't gotten to where the product was good enough to raise more money.
So I took the entire team and we went to China, and we lived in a hotel for 5 weeks, designing and manufacturing as we went.
We designed until late into the night. In the morning, you walk down to the machine shop, you help get them some information, you get some new parts back, and we just kept iterating on this. The electronics market—everything's magical, right?
Yeah.
Then we had to go back, and we were like, “Okay, we need to rush back on the plane to meet some investors.” So we checked the robot, and we folded it up, right?
Yeah.
And we put it in a briefcase. And then when it goes through the scanner—
Oh, no.
You can see the guy just go all white. His hands are shaking as he's opening the bag. And we're like, “No, no, it's just a robot.” And he's like, “Yeah, it's a robot.”
That's hilarious.
Oh, that's awesome.
I'm really thrilled. I loved your TED Talk, and I'm excited to have both Neo Gammas there, hanging out with all our Abundance members. And hopefully, you'll be ready to sell some robots.
In the early days of getting them into the home, no promises, but you're going to have an application process to get the robots in and start to build data assets. When do you think you'll be ready to take pre-orders and orders for Neo Gamma?
I'm gonna be kind to my team and not say a specific date.
Okay.
But it is happening this year.
Okay.
It's this year, 2025.
This year, 2025?
This year, 2025.
Yeah.
Now, we're gonna talk a lot about this in the pre-order.
Yeah.
But the most important thing we do here is manage expectations.
Yes.
This is incredibly early, right?
Yeah.
And what you're buying here is a ticket to be part of this transformation.
Mm.
Adopt a Neo into your family. Help us teach it. It's gonna be a lot of fun.
I love that.
It's gonna be useful.
I love that framing. That's perfect.
It's going to be useful, but it's not gonna be perfect. It's gonna be a lot of rough edges.
Mm.
And we're gonna treat you really well, and we're gonna figure it out together, and it's gonna be an incredibly fun journey. And that's kind of the early adopter program that—
Yeah.
Well, you're gonna have a long waiting list. We need millions and millions of these, and we need to get the price point down. When you think about the constraints to human happiness globally—
Mm.
—a lot of them are gonna be solved through regular AI. But another big chunk—most of them—is related to houses and food and—
Give the jobs that are dull, dangerous, and dirty to the robots.
And then create a lot more of the things that make people happy: the parks, the homes, and all of that.
Everything.
Bigger homes and—
And—
—better things to play with. It's all constrained by our inability to manufacture because of the lack of a humanoid.
Let me ask you a numbers question. I interviewed Elon at the FII Summit. You're gonna be there in October as well, and I also interviewed Brett Adcock. They both gave a number of around 10 billion humanoid robots by 2040. Do you believe that number?
10 billion by 2040?
Yeah.
I think it's probably roughly correct. I think it might happen before. I think it really comes down to what kind of artificial constraints we put on how we scale.
Yeah.
Mm.
At that point, you have to actually really think about how you're refining rare earths, how you're mining more aluminum, and how you're ensuring that you get your labor bootstrapped really well with robots—
into labor? How do you build out a power infrastructure? We need more chip fabs, by the way.
Yeah.
We're not going to be able to build 10 billion humanoids without way more chip fabs.
Yep. Yeah.
We can help with having robots build this out, but I do think that timeline depends a lot on how permitting processes go and how much we allow ourselves to scale, and how fast. But I do hope we get there.
Yeah. For reference, there's about 1 billion automobiles on the planet.
Mm.
You would think there are more, but there are on the order of 8 billion smartphones on the planet.
I'm really glad you said what you just said, though, because the numbers are so wildly out of balance. Each one of these robots uses about a full GPU. It could probably use 2.
If you're talking about 1 billion of them by 2040, we're only making 20 million GPUs a year. And then TSMC has 66% market share in the fab market now. So they have literally one point of failure for the entire economy that we're trying to build. And so we're desperately short on the fabs.
And that's if you just go 1 layer deep. Look at ASML behind it.
Yeah. Oh, my God. Sure.
Right? So the supply chain for chip fabs—
Yep.
That's even more brittle.
Yep. I'm really surprised that we're not moving much faster, given that Elon is right in the middle of it. Elon is, or was, in Washington, and we're just letting this bottleneck fester.
Well, how long have we been talking about magnets?
How long have we been talking about magnets?
We've been talking about magnets for a long time, right? That is a problem that only China can really make high-grade magnets.
Oh, rare earths.
And it's not just the rare earths; it's the process to produce—
Mm.
Yeah.
Yeah.
Right? And I think now, finally, people are opening their eyes and saying, “Wait a minute, this is actually a real problem.”
We meet with a lot of government officials, and they're completely unaware of these bottlenecks. And it's funny—if you point them out, there's still no reaction. It's like... But it's so acute and so urgent. You're in a perfect position to actually identify those bottlenecks, so it's really great that you said it on this podcast, because then we can take that material and say, “Look—look, he would know. This is what we need. This is going to be a crisis very quickly.”
So, yeah.
Bernt, thank you for the tour today. Thank you for the work that you're doing. Super grateful. Excited to have you at the Abundance Summit with your team of robots.
And, by the way, the reason you named the company 1X—I think that's worth closing out as the story here.
Well, there are all these videos on YouTube, and there's always an 8X or 4X in the corner. All we do is real time because we build proper robots.
There, you got it.
Yeah.
What you're seeing is real 1X speed. And we had fun today with NEO Gamma.
Okay. Well, a real pleasure, my friend.
Awesome.
Thank you for today.
Awesome.
Thanks, Jeffrey.
The things I get to do because of this podcast, and just how it was.
We're having fun.
Oh, my God. Yeah.
Awesome.
So awesome.