[BidClub_]
Gradient Dissent · · 46 分钟

Nvidia 和 Uber 支持的86亿美元自动驾驶 AI | Alex Kendall, Wayve

Alex KendallLukas Biewald

YouTube
TL;DR
  • Wayve 持续10年的逆向押注是:自动驾驶应通过一个通用 AI 驾驶员实现规模化,而不是依赖 HD maps、改装传感器和基础设施的车队。 Kendall 起步时只有150万美元,而行业正沿着他认为需要1000亿美元以上的“AV 1.0”范式前进;Wayve 最终打造的模型如今已“能够在任何地方驾驶任何车辆”,但尚未推出无人驾驶服务。

  • Wayve 瞄准的规模机会是消费级汽车,而不只是 robotaxi。 全球每年大约生产1亿辆汽车,相比之下,“全世界的 robotaxi 不到1万辆”;Kendall 认为,将系统原生集成进3万-5万美元的大众车型,在每辆车最终都具备无人驾驶能力之前,就能创造数千万个车辆机会。

  • 瓶颈已从证明模型会驾驶,转向验证、集成和商业化部署。 Kendall 表示,“性能已不再是关键路径”,但谨慎补充说,性能仍有“数量级”的提升空间;近期约束包括安全评估、监管证明、持续数年的汽车销售周期、车辆集成和部署运营。

  • Wayve 最有力的技术证据,是在500多个城市、超过10种涵盖电动车、厢式车和 SUV 的车型上实现零样本运行,其中超过一半城市没有提供先前训练数据。 测试范围从北极圈以北持续22小时的黑暗,到导致东京列车和公交停运的台风;在一天的媒体试驾中,Kendall 称“没有任何人工接管,整天都是完美自动驾驶”,但在获得规模化安全证据前,他拒绝将 Wayve 称为超人类。

  • 语言成为性能输入,而不只是对话界面。 2021-22年的一个项目训练了一个视觉-语言-动作模型,让它“看见世界、驾驶汽车并理解语言”;语言预训练帮助模型应对一辆闪灯示意让行的迎面车辆,也让它能够向消费者和汽车品牌解释行为并选择驾驶风格。

  • Kendall 认为,包括 Wayve 在内的自动驾驶行业应拒绝电车难题强行制造的二元选择,而应始终保留“第三个更好的选项”。 这种最低风险操作会降低动能、保持可预测性,并停车或靠边;Kendall 表示 Wayve 自2018年开始运营以来没有发生事故,同时强调,困难故障通常是黑暗、高速、对抗性驾驶、恶劣天气以及可能的传感器失效共同叠加的结果。

  • Kendall 最尖锐的复盘结论是,深度学习的执行力更多取决于系统,而不是新奇性:“算法占1%,基础设施占99%。” 干净的数据、可靠的学习闭环、测量、评估、仿真和稳健的车辆平台,比“闪亮性感的算法创新”更重要;Wayve 曾在一些方案上投入数百万美元和数月时间,后来发现更好的方法后将其删除。

摘要 · 为研究而整理的核心内容

1. Wayve 的端到端模型挑战了地图技术栈正统

  • Wayve 起步于 Kendall 的判断:驾驶是一个“决策型复杂推理挑战”,最适合用端到端学习解决。目标是训练一个能够推理、驾驶并泛化的单一模型,不依赖限制第一代自动驾驶的改装传感器、算力、HD maps 和基础设施。

  • 他的博士研究已经产出用于定位、立体视觉、语义分割和不确定性估计的深度学习系统,能够把数百万维的图像压缩成有用的世界表征。DeepMind 的 AlphaGo 则独立证明,自我博弈可以解决巨大的决策空间。尚未回答的问题是:这些思路能否在物理世界中汇合。

  • 团队拿着150万美元、几个朋友、一个租来的房子和车库里的一辆车,最初尝试了在策略强化学习。车辆通过摄像头输入,以行驶距离最大化为目标;每当车辆试图驶离道路,就由人类介入。仅经过“10比特的干预”,系统就稳健学会了车道跟随,而且事先没有任何世界知识。

  • Biewald 的追问值得保留:这真的安全吗?能否规模化?Kendall 承认做不到,于是 Wayve 转向模仿学习、离线强化学习和世界模型,让系统通过“做梦”学习,而不是在全球道路上进行不安全的试错。

2. 伦敦迫使模型泛化,但部署如今成了瓶颈

  • 团队先后测试了计算机视觉先验、模仿学习、从仿真到现实的迁移和想象状态学习,随后才确定可规模化的方法,并投资建设基础设施和车队。伦敦的训练环境反而成为优势:拥挤且不规则的街道迫使模型处理“不确定性和复杂性”,而不依赖地图基础设施。

  • Kendall 将市场分为昂贵的“AV 1.0”和“AV 2.0”。前者采用地图和规则驱动的方法,需要1000亿美元以上的投入,最终集中到1到2家具备继续投入意愿和资本实力的公司;后者则由 Tesla FSD 和 Wayve 推动成熟,并从5年前的演示走向全球汽车产品。

  • Kendall 指出,消费者愿意为 FSD 付费,也愿意选择 robotaxi,即使等待时间更长、存在地理围栏,有时价格还高于人工驾驶出租车,这构成了商业验证。与此同时,汽车制造商正在生产数百万辆软件定义汽车,配备集中式算力、环视传感器、OTA 更新和车队数据能力。

  • 当 Biewald 注意到算法改进没有出现在关键路径清单中时,Kendall 给出了谨慎回答:Wayve 仍没有无人驾驶服务,性能也必须继续提升。但科学、数据、算力和车辆集成的复合积累意味着,评估、监管证明、持续数年的销售周期和部署如今更直接地限制着进展。

  • 监管也是部署路径的一部分。Kendall 表示,Wayve 正与英国、欧洲和日本合作,并共同主持联合国一个关于自动驾驶系统 DCAS 监管的委员会;他称相关规则预计将在今年晚些时候出台,从而使这些产品合法化。

3. 语言教会驾驶模型社会惯例

  • 当 Vijay 在2021年或2022年提出将语言用于驾驶时,Kendall 最初回答:“不可能,老兄。”他担心语言会在产品交付前造成干扰。但“文本包含有用知识”的论点最终占了上风,团队训练出了一个统一的视觉-语言-动作模型,而不是把 LLM 外接到驾驶系统上。

  • 早期的解释——“我要绕过一辆违停的公交车”或“我要在红灯前停车”——可能很讨喜,但也会产生幻觉。持续训练让模型更加稳健,也让语言成为核心的性能、表征、可解释性和产品组件。

  • 决定性案例是一名迎面驶来的司机闪灯,示意在名义上的路权优先规则下让行。最初的策略会停车、让行,然后等待那辆车通过;经过语言预训练后,模型会选择转弯。同一能力如今还允许用户提示驾驶风格,解决过于谨慎、令人沮丧的乘坐体验与快到让人不适的驾驶体验之间的差距。

4. 授权模式让 Wayve 在更大的汽车市场中保留选择权

  • Wayve 选择授权模式,是因为 Kendall 认为 AI 是最难的问题,因此与合作伙伴分别负责汽车、云、基础设施和保险。自己打造单一车型可能会把优化方向锁定在错误的车身形态上——SUV、面对面座舱、双座电动车,或其他形态——同时压缩可触达的市场份额。

  • 这种灵活性让 Wayve 能够跟随需求变化:在具备集成条件的汽车出现前先做改装,在疫情期间服务杂货配送和车队合作伙伴,软件定义汽车制造商开始接受后转向车企,如今则进入 robotaxi 机会。当前优先事项,是用大众市场硬件为消费级汽车和 robotaxi 提供城市及高速公路点到点驾驶。

  • Kendall 对规模的比较非常直接:全球每年生产约1亿辆汽车,而全球运营的 robotaxi 不到1万辆。因此,原生集成到消费级汽车中将带来数千万辆车,并且“短期和中期都将远远超过 robotaxi 的规模”。

  • 初始市场包括欧洲及英国、日本和北美,分别由伦敦、斯图加特、东京和湾区提供支持。Kendall 表示,北美和日本推进更快,欧洲紧随其后。中国对 Wayve 而言仍很难进入;Kendall 拒绝讨论地缘政治,但表示自动驾驶如今已是中国消费者购车时最看重的前三大理由之一,与价格和安全并列。

5. 500座城市验证了零样本驾驶,但尚未证明超人类安全性

  • 一个模型已经在欧洲、亚洲和北美的500多个城市实现零样本驾驶,并在超过10种涵盖电动车、厢式车和 SUV 的车型上运行。其中超过一半城市没有贡献任何先前训练数据,使这场全球路演真正检验了 Wayve 最初关于泛化能力的判断。Kendall 称,Wayve 是第一家达到这一规模的公司。

  • 最好的压力测试发生在一场严重到导致当地列车和公交停运的东京台风期间。挡风玻璃雨刷持续清理前置摄像头,Wayve 在东京市中心开展了一整天的媒体试驾,并称“没有任何人工接管,整天都是完美自动驾驶”。

  • Biewald 称这几乎是超人类表现;Kendall 拒绝这一标签。他援引 Swiss Re 对 Waymo 的分析:与同等条件下的人类驾驶员相比,Waymo 的事故和安全事件更少,这说明自动驾驶可能达到超人类水平;但 Wayve 仍必须通过规模化部署证明自己的主张。他还认为,既然超过99%的事故由人为错误造成,那么一个始终保持警觉、具备360度视野、每秒决策超过10次的系统,可能将事故率降至接近于零。

  • 剩余的边缘案例通常不是单一因素,而是多重因素叠加:黑暗环境加高速交通、对抗性行为、恶劣天气,以及可能的传感器失效。Wayve 正在增加韧性、冗余、数据和风险评估,同时从有人监督的消费级汽车和 robotaxi 试验开始。

6. 安全来自降低风险,以及将知识与策略分离

  • Kendall 重新定义了电车难题的处理方式:采取最低风险操作,降低动能,保持可预测的行驶路线,然后停车或靠边。系统需要具备故障可运行能力,找到那个“第三个更好的选择”;因为要进入被迫在两种伤害中二选一的局面,通常需要此前发生多次故障。

  • 如果这种极不可能出现的“两害相权”状态确实发生,Kendall 表示,系统行为将取决于训练数据偏差和系统设计。相关原则是管理风险、让系统保持可预测,并最大化安全性。

  • Biewald 追问:如果 Wayve 几乎看不到碰撞事故,碰撞知识如何进入端到端训练?Kendall 提到行车记录仪和汽车制造商的数据合作,以及 GAIA——一个能够模拟那些过于罕见或不适合直接采集的事件的世界模型;团队可以重放、增强或修改场景,建立系统韧性。

  • Wayve 表示,自2018年开始运营以来没有发生事故,但自有车队并不是唯一证据来源。车队学习会持续吸收更广泛道路数据中观察到的“相当离奇”且不安全的行为。

  • Kendall 将学习世界与选择驾驶方式分开:世界模型应吸收尽可能广泛的好驾驶和坏驾驶数据,而策略可以使用更小的精选分布、专家示范或类似 RLHF 的反馈,并与当地法律、OEM 品牌特征和客户偏好保持一致。

7. 基础设施和评估比聪明算法更重要

  • Kendall 的复盘修正异常直接:“相较于基础设施,我过度押注了算法创新。”数据清洁度、学习闭环的可靠性、迭代速度、自省、测量和评估,才是真正的倍增器——“算法占1%,基础设施占99%”。

  • 一个稳健的机器人平台带来了最清晰的跃升:“我们一旦拥有稳健的机器人平台,就像给原本失明的 AI 戴上了眼镜。”平台能够可靠感知并驾驶后,性能大幅提升;Kendall 认为 Wayve 获得这一基础的时间太晚。

  • 评估经历了一条代价高昂的演化路径,先后使用生成模型、程序化游戏引擎、NeRF、高斯泼溅和生成式世界模型。团队曾在后来删除的系统上投入数百万美元和数月时间;一旦出现更强的方法,就会刻意拒绝保护沉没成本。

  • 领导方式也发生了同样激进的变化:联合创始人在公司约有25到30人、刚完成 A 轮融资时离开,Kendall 随后从编写第一个代码仓库,转向搭建互补的高管团队。Wayve 目前约有1,000名员工,下一项挑战是首次产品发布、大规模公开道路曝光和全球客户执行。Kendall 还表示,时机同样重要——如果推进速度快2到3倍,可能会早于 OEM 的准备阶段——并遗憾公司没有更早招募专家、强化基础设施。

Alex Kendall

Autonomous driving is all about looking at the AV problem with an AI approach. There is an opportunity there to drive accidents down to near zero.

What we did was raise $1.5 million, get some friends together, rent a house, put our car in the garage, and start hacking away. We took it from expensive retrofit vehicles, which relied on compute, HD maps, and infrastructure, to mass-market vehicles that you can buy or manufacture for $30,000, $40,000, or $50,000 each, with built-in hardware that’s in global supply chains, doesn’t need HD maps, and can drive anywhere.

We’re now the first company to have driven zero-shot in over 500 cities throughout Europe, Asia, and North America. It’s also driven in over 10 different cars, from electric vehicles to vans to SUVs. We went north of the Arctic Circle to test the driving in 22-hour darkness and snow. We were even in Tokyo during a typhoon, when local trains and buses shut down. There were no disengagements—perfect autonomy the whole day. It demonstrated that we truly can generalize to any vehicle anywhere.

I think now the question is: How do you bring it to global scale? That’s the shift that we want to take the industry through, where every vehicle is capable of driverless operation, which is clearly the steady state of where we’re going to go.

Lukas Biewald

All right. Here we are at the Core Weave House at NVIDIA GTC. I just got done interviewing Alex Kendall, who's the CEO of Wave, a company that I've admired for a long time and has done phenomenally well in self-driving. And I think you can actually see him behind me messing with the Core Weave F1 car. So, it was a fun interview. I hope you enjoyed it.

1. The story of Wayve

Well, I’m honored to be here with you, Alex, at GTC. This is the first podcast I’ve recorded anywhere except my basement or the office, so this is pretty exciting. I have to start by asking you to tell me the story of Wayve, because I think it’s an especially interesting one.

Alex Kendall

Yeah, actually, it’s fun. I just got done giving a speech here at the event, and I put my face up on the big screen in front of a 50,000-person stadium about 5 minutes before Jensen came on. It reminded me that I’ve only spoken in front of 50,000 people twice in my life.

The last time was at a conference called Web Summit back in 2018, where we had this seed-company pitch competition. I got to the final and had to give, I think, a 3-minute pitch in front of the crowd. Long story short, there was a crowd vote that I ended up coming last in. Yet the judges gave the competition to me because they liked what I said.

2. End-to-end deep learning for autonomous vehicles

It just gives you an idea of what we were doing 10 years ago when I started the company: It was completely contrarian. Wayve started with the thesis that autonomous driving is all about looking at the AV problem with an AI approach. It’s a decision-making and complex-reasoning challenge, and we took the approach that the best way to tackle it was with end-to-end learning.

A decade ago, we started working on building a single model that could reason, drive, and generalize at the largest scale possible, without the AV 1.0, or first-generation, approach, which relied on retrofit sensors, compute, HD maps, infrastructure, and really limited scaling.

We started that a decade ago, when this was really laughed at and dismissed by the industry. We’ve been quietly building away, and we’re at the point now where not only is the industry excited about AVs again, but we have an AI model that has driven zero-shot in over 500 cities. This means the model has driven throughout Europe, Asia, and North America. It’s also driven in over 10 different cars, from electric vehicles to vans to SUVs. It’s now emerging as an AI model capable of driving any vehicle anywhere.

We’ve got a sprint ahead of us over the next year to get this deployed in both robotaxis and consumer vehicles you can buy.

3. Driving zero-shot in 500+ cities

Lukas Biewald

I remember having a really amazing experience driving with you in London, where you let me decide where I wanted the model to go. I hadn’t spent much time in London, and I was astonished by how challenging it seemed compared to San Francisco, where people, I think, obey the traffic laws much more thoroughly than in London, as I understand it. There are just many more surprising things coming at you, and I couldn’t believe how well your model worked at the time.

4. Growing up on a farm in New Zealand

But before we get to the quality of your model, I want to go back in time a little bit and ask you about your history. You grew up on a farm, I think, in New Zealand. Is that right? And how do you think that’s affected your leadership or your thinking about autonomous vehicles?

Alex Kendall

Yeah, in Christchurch, in the South Island. I think there are a few things that were really important to me growing up, on reflection.

I spent half my life building things, whether it was video games, different Lego projects, robotics, or other contraptions. I spent a lot of time messing around with that kind of stuff. The other half was in the mountains: climbing, surfing, mountain biking, and all this kind of stuff.

5. Why the company is called Wayve

I called Wayve “Wayve” because I wanted riding in a car to feel as good as surfing a wave. But I really think that, when it comes to pushing forward frontier technology, it’s all about an adventure. It’s all about doing something where you have to plan, risk-manage, be ambitious, and push boundaries. I think I got a bit of that from those experiences and those values that we can take forward today.

What we’ve been doing is one of those situations where, when we started, this wasn’t possible. But I thought and hoped that, in the future, building this technology would become possible. I think we’ve been lucky that so many things have come together to make that a reality.

I was fortunate to get a scholarship that took me from New Zealand to Cambridge University. Turning up to that place was like arriving at a castle from the Middle Ages, full of the most amazing, brilliant minds, where I could learn and explore. Again, optimizing for adventures has somewhat led me to where we are today.

6. Building the first end-to-end prototype

Lukas Biewald

You’ve got to tell me about building the first end-to-end prototype, because I would imagine that would be particularly hard with the way you’re approaching the problem. You can’t really cheat, right, if you’re trying to do end-to-end machine learning? Tell me about that experience and how you made it work.

Alex Kendall

In 2017, I had just finished my PhD thesis, and I had been able to build the first deep-learning model for different perception tasks: localization, stereo vision, semantic segmentation, uncertainty estimation, and all of these kinds of systems. They could take high-dimensional, multi-million-dimensional images and turn them into compressed representations of the world, learning end to end.

At the same time, friends down the road at DeepMind had just produced AlphaGo. That’s a system where there are more states of the game of Go than there are atoms in the universe—extraordinarily complex. They showed that, through self-play, you could learn in simulation and learn an agent that could beat the world champion.

I saw these things and thought, “Okay, we can now understand the complexity of the real world through computer vision. If you have unlimited data through a simulator in a very low-dimensional state, you can solve the hardest game there is, Go. Now it’s time to bring these together into the physical world,” which I’ve always been excited about: physical mobility, adventure, and robotics.

7. First lane following with 10 interventions

So I thought we could go try it. We raised $1.5 million, got some friends together, rented a house, put our car in the garage, and started hacking away.

The first thing we tried was on-policy reinforcement learning, where we put the car on a road and said, “Okay, go drive with the reward of driving as far as you can without any human intervention.” Your reward is distance traveled. The car starts driving randomly, and essentially, over a period of months, we developed a Q-learning algorithm that lets you intervene and grab the wheel every time it tries to drive off the road.

8. Waymo's impact on San Francisco

The big breakthrough was the moment when, with just 10 bits of intervention, we got it to learn how to lane-follow. I remember getting that breakthrough on a Saturday or Sunday. I was pushing really hard, and nothing was working. These kinds of algorithms tended to work 1 out of 10 times.

Lukas Biewald

Wait, let me make sure I understand. You only intervened? It had no prior information about the world. It was just looking at a camera. After 10 interventions, it understood that you wanted it to follow lanes, and it could follow lanes?

Alex Kendall

That’s right. The input was a front camera. There was a GPU onboard doing optimization, and it was essentially trial and error. After a few different trials, it was able to learn a policy.

There’s a video on YouTube that shows this. We muted the audio because the audio is me whooping and cheering as it starts to drive, but it was essentially just able to lane-follow quite robustly after 10 bits of feedback with no prior knowledge of the world.

Lukas Biewald

Were you afraid when you got in the car the first time and it had no prior information? Were you actually on a road?

Alex Kendall

It was going slowly, so you had time to intervene before it drove off the road.

9. Scaling beyond on-policy learning

That brings up a good question. That method obviously doesn’t scale. You can’t do on-policy tests at scale around the world. That’s pretty unsafe. So we quickly moved to an imitation and offline reinforcement learning paradigm and started scaling up world models so we could learn through dreaming and understand through their experience.

Later in 2018, that’s where we started to take things. The growth there has brought us to this point today, where we can drive any vehicle anywhere, all around the world.

Lukas Biewald

What happened after you got the lane following? What was the next step?

Alex Kendall

The next step was starting to scale it up. We tried so many different things. We tried injecting priors around computer vision. We tried building world models and learning in dreaming and in imagination states. We tried imitation learning, and we tried sim-to-real learning—transferring from simulation to the real world. We tried everything.

We wrote all these blog posts that ended up being industry firsts in all those different areas. Then we settled on a recipe that seemed to work and scale, and started really pushing. We had a couple of years of building infrastructure, fleet buildout, and actually bringing up this technology to a point where it could start to drive in London.

10. Why London was a great training ground

Retrospectively, learning in London was great because it forced us to build something that could deal with the chaos and complexity of your average London street. You have about 100 pedestrians and cyclists around you at any one time. So it forced us to build something that didn’t rely on mapping infrastructure or anything like that, but could actually drive in a city that’s 2,000 years old.

11. Handling regional driving norms

Lukas Biewald

I remember 10 years ago, when you were getting started, the self-driving vehicle space was super crowded, with lots of different approaches and lots of VC money flowing in. It seems like it’s consolidated at this point, but it also seems like it’s really working. I use Autopilot in my Tesla all the time. I ride Waymos every day. I feel like it’s kind of quietly snuck up and become a technology that, at least here in San Francisco, we use all the time.

I’m curious how you think about the competition at this point. Is it becoming a commodity? Is there something different about what Wayve can do or how you approach this?

12. The autonomous vehicle market then and now

Alex Kendall

That’s a good question, because I think, broadly speaking, when we started in 2017, that was probably the peak of the hype cycle. Raising $1.5 million to go tackle what was about 10 companies that had just announced billion-dollar ambitions was a bit lopsided. But it’s fair to say we saw a trough of disillusionment, and I think largely that was from the sheer expense—the $100 billion-plus that’s required to build that AV 1.0 paradigm of mapping, rules, and infrastructure, and to get that to maturity.

There was a lot of consolidation there, where really only 1 or 2 companies had the appetite and capital to push through and make that work. At the same time, Tesla FSD and ourselves matured an AV 2.0 approach—a next-generation stack that’s more affordable and scalable—and have now gone from a demo, which we had 5 years ago, to a point where it’s actually ready for a global automotive product.

I think the most exciting thing today is that we’re also seeing huge validation of the product. You and many others are now buying FSD and paying for it, and it’s actually loved by a lot of customers. When you go to Shanghai and San Francisco, places like this, people will pay—even pay more than for a human taxi experience—for a robotaxi, even though it has longer wait times and is limited by a geofence. It’s a product that people really love.

13. Mass market vehicles versus expensive retrofits

So now the question is, how do you bring this to any vehicle anywhere? I think that’s the step change that we’ve built and can offer: taking it from expensive retrofit vehicles to mass-market vehicles that you can buy or manufacture for $30,000, $40,000, or $50,000 each, that have inbuilt hardware in global supply chains, don’t need an HD map, and can drive anywhere. They have the flexibility to be in a fleet that you ride-hail or in a car that you own.

By the way, this ambition and this business model are only available because we’ve built this generalizable AI driver. But I think now the question is, how do you bring it to global scale? That’s the shift that we want to take the industry through.

To me, I’m thrilled that we’re now seeing the conditions in the market come together. People are now paying, and there’s real commercial validation of this as a product. Regulation is in place, not just in the US, but we work really closely with the UK, Europe, and Japan. We co-chair the UN committee on DCAS regulation for autonomous systems, which later this year is going to come through and actually legalize these products.

14. Algorithm improvements versus logistics

Then, of course, manufacturers are now building vehicles at a scale of millions that have centralized compute, surround sensors, and the ability to update over the air and get data off them, which makes a fleet-learning product possible. All of these factors together give me a lot of optimism for going through a massive commercial inflection point in the coming year.

Lukas Biewald

It’s kind of interesting that you didn’t mention improving the algorithm in that list. It’s very hard to judge the quality of an algorithm or the safety of a self-driving car when you’re in it. But I have to say, when I was inside one of your cars—and this was, I think, 2 years ago—it really felt ready for prime time.

I think it felt safer than the Cruise cars that were on the street in San Francisco, which I don’t know if I should trust my own judgment or not. But are there still things to do with the algorithm to get it to the point where you feel good about using it globally, or is that no longer the challenge and you’ve moved on to logistics and regulation and things like that?

Alex Kendall

Absolutely, there are still things to do. We don’t yet have a driverless service today, but this is no longer the critical path. We’re seeing that the quality of innovation coming from our frontier embodied AI science team, plus the data and compute growth that we have, is just driving compounding scale. Then you layer on deeper integration into vehicles with our partners, and you get sensing and compute on the vehicles. These things compound. Performance is no longer the critical path, but we still have work to do there.

We’ll keep pushing that and pushing it into a generalized state. I think a large challenge is evaluation and validation. How do you prove the level of performance is sufficient to yourselves, to regulators, and to consumers?

Then, of course, there’s commercially deploying and setting up these systems. The last couple of years for us have been a big period of growth for the company: building team strength, partnering with regulators, automotive companies, and fleets in major markets like Germany, Japan, and the US, getting the integration going, and navigating the multi-year sales cycle of automotive. Now it’s going to be the deployment and bring-up of these systems.

I think that’s where the critical path is, but we’re not taking our foot off the accelerator in terms of performance, which clearly has orders of magnitude still to go in terms of the opportunity ahead.

15. Plugging perception into an LLM

Lukas Biewald

One of my favorite demos I’ve ever seen was someone on your team who showed me plugging your perception model into an LLM, so that you could actually see what the car was thinking as it was driving, which seemed surprising and delightful. Was that a toy thing that you realized you didn’t need, or is that something that you’ve continued to invest in?

Alex Kendall

In 2021 or 2022, way before ChatGPT, one of my team, Vijay, came to me and said, “Hey, Alex, I want to start working on language for driving.” My first reaction was, “No way, man. We’re going to stay focused. We’re never going to product-ship this. Come on, stay focused on driving.”

16. Adding language to the driving model

He laid out the argument that the knowledge you can get from text data is enormous, and it might produce new, interesting products and interaction modalities. By bringing in language, we can improve the representation, the performance, and the product experience. So we said, “Okay, let’s start exploring this.”

Very quickly, we got up a prototype. It wasn’t plugging together an LLM, because LLMs didn’t really exist at that large scale yet. We trained a vision-language-action model—a single model that could see the world, drive a car, and understand language.

It started off as a gimmick. You could drive along and have it explain its driving: “I’m going around a double-parked bus. I’m stopping for a red light.” It would just explain it. You could cherry-pick some really delightful things, but then again, there were also some things where it was just hallucinating and wrong.

17. Learning social cues like flashing lights

From that point, we started to train it up. We actually got it to a point where it became very robust. We saw one interesting thing: if you’re driving and you have an oncoming car, and you want to make a turn across that oncoming car to go into a side street, if that oncoming car stops and flashes its lights at you, it’s a signal for you to go in front, right? Even though you don’t have the right of way.

Our AI wouldn’t do that. It would stop and yield and just wait for that car because it had the right of way.

When we started bringing language pretraining into the VLA model, it actually demonstrated that behavior. When it started flashing its lights, it would now turn. I saw that for the first time, and that was an example of how we could improve the representation and reasoning capabilities by bringing language into the system.

18. Personalization and driving styles

Today, it’s a core part of our training. It gives us a boost in performance, opens up new interpretability capabilities, and, of course, enables product experiences involving personalization and interaction with the vehicle. You can now get in our car and prompt it to drive in different driving styles, and I think that’ll be a neat experience.

When you get into an ADAS or a robotaxi, if that car is driving too conservatively, you’re going to get really frustrated. Or if it’s driving too fast for you, you’re going to get freaked out. Different people have quite a wide variety of expectations there, as do different brands. You can go to some brands that want to be sporty, or some brands that want to be safe and reliable. I’m not saying the sporty ones won’t be safe and reliable, but you know what I mean. Providing that personalization is quite a nice outcome of this work.

Lukas Biewald

It’s a funny moment for self-driving, at least in San Francisco, where we suddenly have Waymos everywhere. I love it. I think it’s an amazing experience, but we’re also witnessing all kinds of crazy traffic issues being caused. People, maybe for fun or maybe maliciously, are messing with them in ways that you wouldn’t have expected. It does seem like that capability might actually be completely critical to making a self-driving car a system people trust.

Alex Kendall

Yes, although a couple of things. Firstly, our cars, over time, you won’t notice anything different because they’re just normal cars with built-in sensors. I hope that people don’t behave differently around them compared to other traffic.

Secondly, our system is designed to drive in a very human-like way. The way you see it picking up and dropping off in a robotaxi experience in busy cities, or the way you see it nudging through crowds of pedestrians and doing things that are very human-like instead of just being stuck—I think these kinds of things are really going to improve public acceptance.

19. Human-like driving behavior

We design the system to operate in the messy cities we live in and not require infrastructure or behavior change. I think that’s really important for adoption of embodied AI. You need to use mass-market hardware, you need to be able to leverage the availability of off-policy data, and you need to be able to operate as a drop-in to existing infrastructure, not requiring massive capex changes. If you get these things right, then embodied AI is a delightful deployment experience.

20. Why Wayve doesn't build its own cars

Lukas Biewald

How did you make the decision not to build your own car and to license your technology to other car companies?

Alex Kendall

That was very clear to us from the beginning. I think the mindset I’ve always had, and we’ve always had as a company, is that self-driving is an AI problem. We want to work on the hardest problem first.

We started working on AI, and we started aggressively partnering on cloud, on cars, on infrastructure, on insurance, and working with the best companies around us so that we could focus on the critical-path problem. By always focusing on the hardest problem first, we’re not going to stumble into some glass ceiling that we hit down the road. I think that was the mindset.

Retrospectively, it’s become important because I think it’s a mistake to focus on one vehicle form factor, because it’s really unclear what will have product-market fit. Is it a normal SUV like today? Is it a bidirectional vehicle where you sit facing each other? Is it a 2-seater electric vehicle? Are there other applications that really take off? I think that’s unclear.

At the end of the day, if you focus on one form factor, you’re also going to get a small part of the market. We’ve tried to focus on how we build a generalizable driver that can address everything and then follow where the market is most advanced at any one stage.

We started off with retrofit when there wasn’t really anything for us to integrate into. Then, during COVID, grocery and fleet started skyrocketing in profits, and so we started partnering with that sector. Then things changed. Automotive companies, which would never speak to us and said, “This is unsafe. It would never work,” all of a sudden started talking to us and building software-defined vehicles so we could integrate with them.

Now, today, we’ve got the excitement of robotaxis. We’ve followed the market while being open to integrating into anything. The end state, of course, is consumer vehicles, delivery, robotaxis, trucking, and non-automotive forms of robotics. We want to license our embodied AI and make all of them intelligent, safe, and possible.

21. Stack ranking business opportunities

Lukas Biewald

Where you sit right now as CEO, do you have a stack ranking of those different modalities and which one is most top of mind for you?

Alex Kendall

Yes. Today, it’s all about consumer vehicles and robotaxis. It’s about point-to-point driving in urban and highway situations with mass-market consumer-vehicle hardware. This is an aligned product that we can now commercialize with great partners. This is the real focus for us today, and it can generate an extraordinary business.

You look at the scale, whether it’s data, revenue, number of vehicles, or exposure: about 100 million cars are produced each year. People forget that. Often, they only think of robotaxis, but there are fewer than 10,000 robotaxis in the world. Getting this natively integrated into consumer vehicles opens up tens of millions of vehicle opportunities, which is going to far outscale robotaxis in the short and medium term.

Before long, of course, every vehicle will be capable of driverless operation, which is clearly the steady state we’re going to reach, given the safety, convenience, and product benefits that it can bring.

Lukas Biewald

What geographies do you plan to launch in first?

Alex Kendall

We’re really looking to make sure the system is global, but we’re focused today on Europe, including the UK, as well as Japan and North America. We have fleets and offices in London, Stuttgart, Tokyo, and the Bay Area. We’ve driven around most cities throughout Europe and North America, and we’re growing in Japan. Those are the markets we’re starting with, given the customers we work with.

Lukas Biewald

Do you have a sense of where you will go mass market first?

Alex Kendall

It’s those markets. In general, we’re seeing faster movement in North America and Japan, closely followed by Europe. Of course, China is unfortunately very hard for us to work in, and vice versa. I won’t get into geopolitics, but it’s a bit of a shame that’s where the world’s at.

22. The AI 500 road show results

In China, this technology is moving so quickly, and the rate of innovation is very impressive. I think it’s worth looking at if you haven’t, from an automotive-technology perspective. It’s another proof point that when this technology is deployed, consumers really love it. Alongside price and safety, it’s now a top-three reason why people buy a car in China.

Lukas Biewald

Wow. There’s autonomy. You had amazing results at the AI 500 roadshow last year. I don’t know if you want to talk about it, but you took one single AI model and deployed it all over the place, with many of those geographies not having any prior training data. Were you surprised that it worked so well?

Alex Kendall

No, because that was the core thesis: by building a general-purpose AI model, we could drive anywhere. But to actually see it was huge validation.

We drove around the world. There were some very cool anecdotes. We drove in places like around the Arc de Triomphe in Paris. Have you driven around there?

Lukas Biewald

I have, yeah. Well, not driven, but I’ve gone there.

Alex Kendall

It’s chaos, right? We went north of the Arctic Circle in Finland and Sweden and tested driving in 22-hour darkness, snow, and things like that.

We were even in Tokyo during a typhoon, when local trains and buses shut down. That’s actually how we were giving a bunch of media drives. We had a day of journalist drives with no disengagements—perfect autonomy the whole day—in literally the heaviest rain I’ve ever seen.

Our cameras are situated just behind the windshield, including the front camera. The windshield wiper was constantly wiping water off it, but it drove reliably around central Tokyo, which was quite remarkable. There were fewer pedestrians on the road that day, but it was still busy traffic.

We’ve seen all of these things, and it demonstrated that, as the first company to do something of that scale, we truly can generalize to any vehicle anywhere, given that over half of these cities had no prior training data. It was a true zero-shot generalization test at scale.

23. Is this superhuman performance?

Lukas Biewald

That almost sounds like superhuman-level performance.

Alex Kendall

I’d be careful with that word. You and I, if we jump into a new city, can get a rental car and drive there. But to claim superhuman—I mean, there are some great reports that I’ve read, including one from Swiss Re, where they measured the impact of an autonomous system, Waymo’s operation, over its operational history and found that it had much fewer accidents and safety events than equivalent human drivers.

I think there’s some robust evidence that’s really showing autonomy can be superhuman.

24. Introduction

We have work to do to demonstrate that, but we're making good progress there. Sadly, over 99% of accidents are caused by human error. Even just making a self-driving experience that can never be distracted, drunk, or impaired, that can see 360° all at once and make decisions over 10 times a second, presents an opportunity to drive accidents down to near zero. That's what we want to see, but I think you're going to see scale deployment really show that. We want to go do that in the coming years.

25. Scenarios that still challenge the model

Lukas Biewald

Are there situations that still make you nervous at this moment? What would be the kind of configuration that would put you on the edge of what the model can do?

Alex Kendall

I think it's very hard to give general statements to that, because where the models today tend to struggle is when you have a culmination of many failure modes. If it's dark, you have high-speed traffic, and you have adversarial behavior from someone else driving on the wrong side of the road or not yielding, maybe combined with bad weather, that's when you have multiple confounding factors. Then maybe you have a sensor failure. When you get multiple confounding factors, that's when you tend to find challenges and failure modes.

We need to keep addressing that with better system resiliency or redundancy. We're building that with some next-generation vehicle architectures. We need more data to improve generalization and risk assessment, and we need to make sure we deploy responsibly in areas where we're confident we meet the required bar for safety. We'll do that responsibly with our partners as we grow.

I think the other nice thing about being deployed in consumer vehicles and robotaxis, starting with supervised trials, is that we can incrementally grow this with communities and regulators to the end state where every vehicle is autonomous at a global scale.

26. The trolley problem in self-driving

Lukas Biewald

I remember there was a moment in ethics in AI, maybe when self-driving was hot, when people were asking questions—maybe toy questions—about whether a car should prioritize the safety of the passenger versus a pedestrian. End-to-end, you're kind of sidestepping those questions, and maybe you're not inserting yourself. Is that true, or do you need to push the data in certain ways to get certain types of ethical behavior that you want?

Alex Kendall

This is the trolley problem. We've got two bad decisions, and you have to choose bad decision A or bad decision B. There are many ways of framing it. The practical reality is that, as an industry—and I think I speak on behalf of most AV companies—we need to design a system so there's always a third good option.

It's called a minimal risk maneuver. If you're in a bad scenario and you see risk, there's always a third option you can take, which is to minimize risk. That typically involves reducing kinetic energy, maintaining a predictable course of action, and stopping or pulling over the vehicle. We always need to prepare to do that if any systems fail. There needs to be a fail-operational ability to do that. We design these systems to always have a third good choice, rather than being stuck between two bad choices.

27. Minimal risk maneuvers explained

The second point is that so many things have to go wrong to get into that two-bad-choice state. The probability of those kinds of scenarios is so small. Even if you are in that position, the probability that you can accurately sense those two things in those scenarios makes this more of a theoretical thought exercise than an actual practical decision the industry needs to take. We focus—and the industry focuses—on building these minimal risk maneuvers that take us away from those situations.

28. Training data and crash scenarios

Having said that, if you do get into that unlikely situation, then of course the behavior will depend on how you bias it from your training data and the system that you design. We need to be thoughtful with the principles and the knowledge that we distill into the system, and with the behavior that we generate. There are some good principles around managing risk, making the system predictable, and, of course, maximizing safety.

Lukas Biewald

I would imagine that in your end-to-end training, you're not seeing a lot of crashes. If there are crashes, you're probably generating them in the dream data somehow. I think that the types of crashes that you're feeding into your system would give it a certain point of view.

Alex Kendall

I'm really proud of our safety record, actually. We've had no incidents since we started operating in 2018, so we have a really fantastic safety record around the world.

There are a few things that are important here. Firstly, we don't just learn from the data of our vehicles driving. We have enormous data partnerships that give us quite general-purpose knowledge. Unfortunately, there are still a lot of road accidents today, and we partner with dashcam providers and car companies that get data from their consumer vehicles. Sadly, there are still a lot of quite frightening scenarios that they experience. The good thing is that we can use that data to learn from it.

Secondly, we have our world model, GAIA, which can simulate scenarios that are too rare or unsafe to see in the real world and make sure that we're robust to them through synthetic data and simulation, too. Ultimately, with those strategies, we need to build robustness to those events.

We still see some pretty bizarre stuff on the roads around the world, including some pretty unsafe behavior. We can resimulate, augment, or modify these scenes to build resiliency to them. It's a constant game of fleet learning and data iteration to make sure we're grinding up performance over time.

Lukas Biewald

I remember when Tesla launched what I believe they said was end-to-end dream full self-driving. You really felt it, I think, when that version came out, at least in my Tesla. One of the things that was really notable about it was that it stopped at every stop sign in San Francisco, which San Francisco drivers don't really do. It's kind of annoying if you're really following that rule to a T.

It made me think that the average of everybody's driving is probably not what you want. How do you deal with that in your end-to-end training?

Alex Kendall

It's a great question. I was just in Japan last week, and there everyone drives about 20% above the speed limit. It opens up an interesting question around policy, regulation, and liability. If you're an eyes-off or driverless system where the system's taking liability for driving, it should meet the road rules. But if you're a consumer, a lot of products today in a driver-assistance setting allow you to set the speed above the speed limit.

It's a really interesting question. I hope that self-driving improves road safety to the point where we can actually increase speed limits because it's safe enough to do so. But we're probably some distance away from being able to have those kinds of arguments.

29. World models versus policy learning

Today, the first thing to realize is that when you're training a world model, you want to see as diverse data as possible. You want to see good driving and bad driving. You want to really understand the full spectrum of dynamic events in the world.

Then, when it comes to actual policy learning of how you drive, there's a data distribution that matches whichever customer's preferences. The OEM driving style and behavior, as well as the rules and policies of the country they're operating in, become really important. We can separate those two data distributions. The first one should be as diverse and as large as possible. The second one could be highly curated and small, whether it's RLHF-style feedback or expert demonstrations. There are many ways to do it to make sure that you drive in a safe and considered way for that application. You can separate these data distributions.

30. Transitioning from CTO to CEO

Lukas Biewald

Actually, one thing I wanted to ask you, somewhat as a fellow entrepreneur, is that you started, I think, as CTO with another co-founder who then left, and you became CEO. I wonder what that experience was like. I guess it was earlier in the journey, but if you want to share about that.

Alex Kendall

Yeah, it was early on. At that stage, I'm sure you remember, everyone does a bit of everything. As we were getting started, for a number of reasons that unfortunately made sense for my co-founder, he left the company to pursue other things. We're still in good touch today.

The transition was my first real experience of change-management communications for the company. I think we were about 25 or 30 people at the time and had just raised Series A. It has matched my experience over the last decade, where every year has been a completely new challenge.

I started the code repo and was coding a lot of the first demos. Then I started working on hiring the team, finding commercial partners, and growing investments. Now I'm managing customers and managing through an executive team, where I'm so thrilled to be working with people who are truly at the top of their game. They're industry legends, whether it's in science, engineering, product, finance, or commercial. Across the board, we've got truly world-class leaders.

I think it's probably a few things: the importance of building complementary teams, working with people that you can learn from, and also getting used to outgrowing a skill pretty quickly and having a new challenge. As soon as I get good at something, it's no longer the relevant skill for me to be working on because it's usually me plugging a gap for the organization and then growing the organization to absorb it. Then there's a new challenge.

I enjoy that, though. I think that's part of the goal. It's the whole thing—it's part of the adventure of building a company. I think the next year for me is now this new phase, where I'm starting to live on a plane and work with global customers. We'll be 10 years in, and it'll be our first product launch.

31. Advice to your younger self

We're going to start to have major public road exposure through that, while keeping up strong and consistent execution velocity. This is a completely new challenge that I'm really quite excited for. Let's see what comes after that.

Lukas Biewald

If you could send a 30-second message of advice back to yourself when you were just starting out, what would you put in there?

Alex Kendall

Oh, man. I feel like if I built Wayve again, I don't know.

Lukas Biewald

If you built things again, how much faster do you think you could build them the second time around?

Alex Kendall

I would totally say it's like 2 or 3×, right?

Lukas Biewald

Yeah.

Alex Kendall

I took so many shortcuts, did things that I was hesitant to do, and made mistakes that I think I could just take the tactical playbook from.

Lukas Biewald

You had to distill it to a higher level because you only have a 30-second window—not just to give the loss-function recipe and relationships and partnerships, but just really tactical things.

Alex Kendall

I always thought the interesting question was: I feel like we'd be a bit useless if we went back in time 200 years with today's knowledge. What would you do? I don't know what you've learned through Weights & Biases, but you've built on top of stuff that just didn't exist 200 years ago. We're standing on the shoulders of giants.

There's a bit about timing. In many ways, if I developed our product 3 times as fast, we may have been ahead of the industry, when automotive didn't have the right infrastructure or products to actually accept our product. I think timing is everything. In many ways, self-driving has been a question of staying alive and keeping the company capitalized while making progress on our next-generation approach.

If I were to distill it down, I've always had conviction in this approach. It's not like I'd say, "Keep going, Alex," because I've been kind of resilient toward that one. I've been fine with how bold and brave I've been in building this.

Honestly, the biggest thing has been—two things actually genuinely come to mind. One is talent growth at Wayve. I've seen so many great, fantastic, brilliant people and the growth of the expertise we have in the company. I think moving faster and harder, and starting out of a PhD and growing from London compared to the scale of the amazing people we can work with today, there's a great opportunity there that I could have leaned more into.

Lukas Biewald

So, do you mean like try to get the experts earlier?

32. Algorithms versus infrastructure

Alex Kendall

Yeah, yeah, I think so. You know, we started off with a generalist mindset really trying to figure things out and had to build a lot ourselves, but I think maybe it's a mixed feeling on that one. The bigger thing that comes to mind for me is that I over-indexed on algorithmic innovation compared to infrastructure. The big learning I have with deep learning is that it's 1% algorithms and 99% infrastructure. Whether it's the cleanliness of your data, the reliability and iteration speed of your learning loop—we've been very happy adopters of Weights & Biases over the years, and the acceleration you get with it—or introspection, measurement, and evaluation tools, I think we've been on a journey there and underinvested in it.

I think going harder and faster there, rather than pursuing shiny, sexy algorithmic innovation, is something I wish we'd done earlier. The thing is, we were held back by not having a robust platform for many years. As soon as we had a robust robot platform, it was like giving our AI, which was blind, glasses. All of a sudden, it could see and drive in a remarkable way.

33. Evaluation and simulation challenges

Focusing on that earlier and harder, and building that maturity, was quite key to bringing up performance. We were too late in that. In simulation and evaluation, I think both large language models and embodied AI present a real challenge in evaluating the performance of models. It's an open-ended problem. It's not like you can have a clear metric around that today.

Developing these techniques has been a journey. We've consistently tried different things, from generative models to procedural game engines, to NeRF and Gaussian splatting, to now generative world models. Throughout those times, we haven't been afraid to forego sunk costs. We've spent millions of dollars and months of time building something, and then just deleted it and moved on because we found something better.

That culture and that journey have been important. Today, I think we have something on the cusp of a major breakthrough there that will really transform the industry. But that's another example where getting that right earlier would have helped. It's chicken and egg: if you figure out how to measure it, then you can just optimize against that target. Getting that right is something we should have moved faster on.

34. Taking Bill Gates for fish and chips

Lukas Biewald

Okay, some quick questions to end with. When you took Bill Gates for fish and chips in your self-driving car, were you nervous that something would happen? What did you talk about?

Alex Kendall

Yeah, that was one of the first times we'd gotten the system reliable enough. Actually, I think there were a few interventions on that drive, so I think Bill was still excited at the time.

We talked about how, when he was commercializing Microsoft and Windows, he had to get a bunch of design wins with PC OEMs. In a similar way, we've had to get design wins with the biggest car companies in the world. There are some interestingly similar dynamics there that he was interested in.

He was in London for a day when we went to get fish and chips. One of his team members went in to get the fish and chips, handed it through the window, got in the car, and Bill just put it down and sat there. He wasn't going to eat it. I said, "Come on, Bill, we've got to eat it. We can't not eat the fish and chips."

He said, "All right, then." He grabbed a piece of fish in his hand, brought it up, broke it in half, and just started eating this battered piece of fish in the back seat of the car with me, which I loved. He clearly has a soft spot for some greasy fish and chips.

35. Parties at the first prototype house

Lukas Biewald

Yeah, nice. Nice. All right, what about when you built the first prototype in your rented house and said you threw one of the biggest parties of your life? What was that like?

Alex Kendall

We invited everyone we knew in Cambridge to the house. The living room was where we worked with a bunch of computers. The small bedroom was our boardroom. We had another bedroom full of servers, and it was constantly 40°C because of the 60 GPUs that were running there.

Then we just turned the house into a bit of a party house. We had a few musicians on the team, and we were jamming. That was a really fun vibe. Those days are very special, and I really enjoyed the lasting memories of a team that was really in the trenches together during that time.

But that was quite a big one.

Lukas Biewald

I don't know about you, but I feel like I'm always chasing that feeling of the early days, when everyone's cranking together. Such a good feeling.

Alex Kendall

I just visited our Japanese office last week, which is about 20–30 people. It really feels like that. That's super exciting.

I think there are pros and cons to operating at a different level of scale. Today, we're about 1,000 people globally, but the culture you set at that scale really drives through to where we are today. I think it does require different skills operating at that scale, but the impact you can generate is just tremendous.

36. Wrap-up

Holding on to disruptive innovation and not falling for the innovator's dilemma—these principles are still things we really fight for, and I think they still hold true today at Wayve.

Lukas Biewald

Well, you've had some real success. It's been a joy watching you guys grow.

Chris Sacca. Who said it? Chris Sacca. Thanks so much for listening to this episode of Greymatter Decent. Please stay tuned for future episodes.

Nvidia 和 Uber 支持的86亿美元自动驾驶 AI | Alex Kendall, Wayve — 文字稿与摘要 | BidClub