[BidClub_]
Latent Space · · 96 分钟

Runway押注视频之外:世界模型、机器人与神经操作系统——Anastasis Germanidis

Anastasis GermanidisswyxVibhu

AI与软件机器人技术企业经营
YouTube ↗
TL;DR
  • Anastasis Germanidis认为,没有证据表明世界模型需要绕道采用JEPA式架构。 他对Yann LeCun“理解世界的运行规律,与生成可爱视频截然不同”的说法回应得很直接:“它们是同一件事。”他指出,在Physics IQ基准中,固体力学、流体动力学和光学得分会随着模型规模和算力扩大而“可预测地提升”,但也提醒,视频模型可能作弊,当前表现很容易被过度解读。
  • Germanidis认为,实时生成不是可选项,而是必然趋势:“如果2年后我们还不是主要使用实时模型,我会非常惊讶。” Runway的characters模型是一个以24 FPS运行、经过步数蒸馏的自回归头像模型;在他看来,这是“实时视频模型最大规模的部署”。扩散模型提供两条蒸馏路径:缩小模型,或将推理步数从50步降到4步,让许多场景下的前沿模型也达到实时表现。更低的服务成本叠加更好的用户体验,使这条成本曲线值得关注。
  • 本期最具扩张性的判断是:界面世界模型会把软件渲染为生成像素——“这里没有HTML、CSS、React……为什么要生成用于生成像素的代码?直接生成像素就好。” 终局是“完全神经化的操作系统”;同一类模型还可以充当合成数据生成器和计算机使用Agent的实时RL环境,让这项技术突破创意工具的边界。
  • Runway“意外地仅靠扩大视频模型规模,就造出了机器人领域的SOTA模型”,而其数据论也颇具分量:第三人称视频的数量比第一人称数据多3个数量级。 基于Gen 4.5构建的GWM-1,仅使用数百小时机器人微调数据,而机器人预训练通常需要数十万至数百万小时。Runway称,在RoboArena上,真实世界与仿真结果的相关性非常好,这意味着策略评估可以在仿真中扩展,而不必依赖硬件。
  • 2024年2月的Sora时刻让Runway在“Runway完了”的喧嚣中经历了“数小时的存在主义危机”。 随后公司用3个月把模型规模和训练算力扩大10倍,从零开始建立模型并行能力,并在当年夏天推出Gen-3。这段韧性故事,加上最初在Series B阶段决定建设1,000张A100集群的“略显不理性”押注,构成了同一套模式:在证据出现之前下注,然后快速适应。
  • 谈到竞争格局,Germanidis承认西方在视频领域落后:“排行榜前10名、前20名里,大多数模型都是中国模型。” 他与NVIDIA共同发起Cosmos Coalition,希望通过开放权重、物理基准和基础设施,把世界模型研究带入开放生态。他质疑长期依赖中国模型蒸馏的策略,认为这可能锁定性能上限;Runway则认为,不靠蒸馏也能训练出强模型。
  • 创意AI的采购标准已经从质量转向可控性(单次生成最多50个参考素材),如今又开始转向延迟。 企业采用非常活跃,Runway目前可能拥有公司历史上最多的开放岗位,尤其集中在机器人领域。近几个月艺术家态度也明显变化——戛纳导演和Ron Howard都公开表达了对AI的支持;随着工具越来越可控,AI更像“一把画笔”,而不是“一个能替你想好整部电影的魔法实体”。
摘要 · 为研究而整理的核心内容

1. 创业论点:把生成模型交给艺术家,它们就会走出分布

  • Germanidis在加入Runway前把艺术实践与机器学习工程结合起来:2015年做过一件画廊作品,为访客分配Mad Libs式身份(“你是建筑师,你30岁,你喜欢运动”),再通过模板和马尔可夫链模拟闲聊,时间上早于可用语言模型很多年。贯穿始终的问题是:“通过创造这些极其简单的互动模型,我们能对人类了解什么?”
  • 他与联合创始人Chris基于NVIDIA的pix2pixHD打造Uncanny Road(2016–17年;这是最早一批1K分辨率图像生成模型之一,训练数据来自自动驾驶),由此得到关键验证:用户用一份“非常无聊的数据集”生成了“100多万名行人、100多万个交通标志”和“巨型人类”。他的总结是:“你可以拿同一批生成模型,从另一个方向看它们……把它们交给艺术家,他们会做出你意想不到的东西。”
  • 2016–17年的原始推断是“问题不在于是否发生,而在于何时发生”:分辨率和质量会以可预测的方式提升,因此“总会有一个时点,大多数内容都会由模型生成”,创意工具也必须随之重构。

2. 7年压缩:从抠像工具到2022年的阶跃

  • Runway的第1版只是把开源模型封装起来,服务于不会机器学习的艺术家;由于当时的模型“还没有好到可以产品化”,公司在第1年就成立了研究组织。多年间,旗舰产品一直是Green Screen,一款机器学习抠像工具——“没人喜欢手动做这个”——曾被用于《瞬息全宇宙》。
  • 他对2018–2022年社群的怀旧颇有启发:那是一个规模很小的圈子,任何项目都可能走红。比如顾问Gene Kogan制作的片段,把纽约地铁重新风格化成梵高或另一位画家的风格——他认为,这种东西“对人们来说完全是一次启示……也许放在10年前就是如此”,足以说明行业发展有多快。
  • 2022年的潜空间扩散和DALL-E 2构成了阶跃变化;swyx补充说,他之所以进入这个领域,是因为Stable Diffusion可以“在消费级硬件上运行”。

3. 不理性的A100押注——以及Gen-2的周末黑客项目

  • 2022年年中,Runway处于Series B阶段,Germanidis相信语言模型的规模定律会迁移到视频,于是签下了1,000张A100集群的订单——“也许是一个略显不理性的决定”——目标是在2022年秋天实现视频领域的“Stable Diffusion时刻”;当时SOTA还只是256×256分辨率的CogVideo。
  • 2023年1月推出的Gen-1有意采用视频到视频路径:“当你有更强的条件约束时,这基本上是一个更容易的问题”,比文生视频简单。重度用户把Blender渲染结果、用书本搭建的城市等素材交给它,再用手机拍摄并转换成照片级真实感。
  • Gen-2的起源此前很少被讲述:它在Gen-1发布2个月后公布,甚至早于Gen-1全面开放,最初只是Germanidis的“周末项目想法”——把文生深度模型串接到Gen-1的深度到RGB模型上。他提到,2阶段管线如今正在回归:Reve的文生图模型会先由规划器生成边界框,再交给扩散Transformer;Vibhu补充说,Ideogram当天也发布了类似思路,“这个社群很小……大家彼此交流”。

4. Stable Diffusion的源头故事

  • Germanidis的说法是:当时在Runway任职的Patrick Esser于2021年末与Robin Rombach及CompVis的其他研究人员合作潜空间扩散;Stable Diffusion“基本上是同一个模型,只是用了更多算力训练”,并加入无分类器引导和LAION的审美子集,训练运行由Stability出资。“它原本就是一个开源研究项目”,因此Runway继续发布各个版本,包括Stable Diffusion 1.5。
  • 对于后续风波,他的说法相当克制:“可能有一天出现了一点沟通误会,但最终很快就解决了,几个小时内就解决了……现在都过去了。两家公司走上了各自的道路。”

5. 文生视频从来不是答案——镜头控制如何埋下世界模型种子

  • 早期最大的教训是:“很早就已经清楚,文生视频不会是答案。”Gen-2主要被用于探索,因为“没有任何东西可以把它锚定住”;因此2023年连续推进条件控制:Motion Brush通过绘制箭头指定运动,此外还有输入帧控制和镜头控制。
  • 镜头控制的意义超出了电影制作:“那是你第1次感觉到,自己不是在生成一段短视频,而是在世界内部移动……也许这就是我们后来关于世界模型的一些想法的种子。”
  • 关于提示词,他说:“每一个投入生产的视频生成模型,底层都在使用复杂的提示词补全管线。”他把这条脉络从DALL-E 3论文中的合成描述追溯到Sora;Vibhu的概括是:“人类不擅长写提示词。”

6. 世界模型论与LeCun之争

  • Runway在6月公告中提出的框架——“人类心智不再是AI的中心,我们的世界才是”——建立在对DeepMind使命的反转之上:不是先解决智能、再解决其他问题,而是从世界本身那些尚未被书写的复杂性出发。他举例说:“如果我让你用语言描述如何系鞋带,那是很难描述的事,但示范起来却显而易见。”这正是Moravec悖论在训练数据上的体现。
  • 被问及这一范式是否需要类似JEPA的中间层时,他说:“我们是一个非常务实的研究实验室……我们还没有看到任何迹象,表明视频预测本身无法扩展。”对于LeCun在社交媒体上称理解世界运行规律不同于“生成可爱视频”的说法,他的回答是:“它们是同一件事。”但他也补充,模型“可能会作弊”,比如通过切换场景、挑选样本来规避难题,因此“不能被当前视频模型的表现过度欺骗”。
  • 证据来自Physics IQ:该基准通过首帧滚动物理现象,再对固体力学、流体动力学和光学进行评分;结果随模型规模和算力扩大而“可预测地提升”。Gen-2“就像GPT-2”:只能生成勉强连贯的视频,但“没有理由认为它在规模扩大后不应该奏效”。
  • swyx的反驳值得保留:柏拉图洞穴式的疑问是,模型从结果中学习,同时“丢掉”整整一支物理学分支,“感觉不太对”。Germanidis的回应则很有力:“纵观整个机器学习史,很多事情都让人感觉不太对。”

7. Sora改变叙事,Gen-3在3个月内回应

  • 2024年2月,Sora发布,明显优于Gen-2;Twitter上出现“Runway完了”的判断,而OpenAI“正处在最佳状态……没有人甚至能接近它”。Germanidis承认,当时“经历了数小时的存在主义危机”。
  • Runway把过去3年从大语言模型中学到的经验压缩到几个月内:从ConvNet转向扩散Transformer,通过“大量失败尝试”搭建分布式训练基础设施,把模型规模和训练算力扩大10倍,并在“完全没有经验”的情况下学习模型并行。他说,如果问Runway的人最喜欢哪个时期,很多人会选那段冲刺。Gen-3在当年夏天上线;几个月后,Gen-3 Alpha Turbo跟进,他认为这可能是第1个投入生产的步数蒸馏模型。

8. 实时生成不可避免:蒸馏的2条路径

  • 扩散模型提供2个杠杆:像大语言模型那样蒸馏成更小的模型,或者蒸馏推理步数——从50步降到4步,而且输出往往“具有可比性”。很多时候,Runway会直接把前沿模型进行步数蒸馏,使其达到实时表现,而不是从更小的模型起步。
  • characters模型由会说话的头像组成,采用步数蒸馏、自回归架构,运行在24 FPS;它是“实时视频模型最大规模的部署”。他的判断没有保留:“如果2年后我们还不是主要使用实时模型,我会非常惊讶。实时视频生成就是不可避免的——用户体验好得多,服务成本也低得多。”随着蒸馏改进,质量差距会继续收窄。
  • 对于swyx认为一致性模型的热潮已经退去,他回应:“我不会这么确定地说它没有留下来……我们仍然落后于语言模型2-3年。”

9. 界面世界模型:生成像素,跳过代码

  • 现场演示是一个完全由实时视频模型渲染的UI:“没有HTML、CSS、React在驱动这个界面。”它可以接受点击、拖拽和滚动,元素行为通过提示词指定(“如果我按下这个按钮,我希望发生这件事”),还可以生成音频或音效。背后的理念是:“为什么要生成用于生成像素的代码?直接生成像素就好——这是端到端理念在前端的应用。”
  • 终局是Karpathy所说的完全神经化操作系统。在他看来,如今用一个僵硬的聊天框包裹通用LLM“有点奇怪”;最终“界面本身会变得可学习”,独立应用的概念会消失,个性化也更容易通过提示词工程实现,比如为无障碍需求放大文字,或针对不同访客生成不同版本。
  • 每个世界模型都服务于两个方向——“面向人类的世界模型,以及用来训练Agent的世界模型”。因此它可以成为合成数据生成器,也可以成为计算机使用Agent的“实时RL环境”。他坦率承认经济账:运行视频模型显然比渲染HTML更贵;第1步是先找到“明显更有吸引力”的交互,成本问题随后再解决。

10. 长上下文是他等待的研究突破口

  • 核心约束是自回归误差累积:生成帧不断被反馈回模型,微小误差会逐步叠加。“这不是一个新问题”,LLM已经解决了对应版本;“所以它是一个可解决的问题,但无疑仍然是挑战。”
  • 按他自己的数据,目前characters模型可以自回归持续生成最长30分钟;开放式探索模型GWM Worlds在明显退化前能生成“大约几分钟”。swyx提到,Genie大约只能维持1分钟。“如果每隔几分钟就要重启一次,这不是理想的游戏体验。”

11. 机器人:意外的SOTA与第三人称数据论

  • “随着我们扩大视频模型规模,意外地造出了机器人领域的SOTA模型。”最初各实验室来找Runway要合成数据,后来开始要模拟器;GWM-1是在Gen 4.5基础上改成自回归、动作条件化并经过步数蒸馏的模型,形成一个可以“走进去”的模拟环境,用于探索反事实动作。
  • 验证工作使用了RoboArena场景。Runway称,世界模型内部的策略表现与真实世界结果之间存在非常好的相关性,因此评估可以在仿真中扩展,而不必在硬件上进行。相比Isaac Sim、MuJoCo等传统模拟器,布料、湿滑表面和其他复杂交互很难、很耗时,可能还需要数字孪生以及对每个物体进行3D扫描;而这套方法是“只要拍一张环境照片,就能在首帧内展开策略”。
  • 数据层级本身就是这套论点:遥操作最难规模化;UMI数据是人类第一人称视频配合机器人夹爪;GoPro第一人称采集更容易,但第一人称数据仍比第三人称视频“少3个数量级”,而人类很大一部分能力“来自观看其他人怎么做”。结论是:“最终胜出的,是最丰富的数据源。”因此,视频预训练加上数百小时机器人微调,就可能支撑良好表现,所需数据远少于机器人预训练通常使用的数十万至数百万小时。
  • 目前模型会针对合作方的具体形态进行后训练,包括单臂、双手和人形机器人;他预计1-2年内,GWM的不同变体——操作、导航、characters——会统一成“一个实时视频模型”,模拟“身处这个世界是什么感觉”。再加上动作头,世界模型就会变成策略本身,也就是“世界动作模型”的方向。

12. 清醒梦测试:为什么失败必须能够被仿真

  • 他为世界模型设定的图灵测试是:戴上VR头显进入一间房,在里面走动并与物体互动,然后回答自己是在透视模式中,还是在一个生成世界里。“如果你无法确定……就说明模型已经足够好了。我们还远未达到这一点。”
  • 差距在于反事实:让视频模型生成一个人射门得分或射失的画面时,它更擅长生成进球,因为训练数据偏向成功。对于机器人评估和未来的在线RL,“你希望失败能够被非常好地仿真”。这引出了swyx的调侃:这是“唯一一个成功样本太多的领域”。

13. Omni模型与跨模态迁移押注

  • “从外部编排迁移到模型内部”是反复出现的模式:思维链提示最终变成原生推理,单镜头视频加外部编排最终变成原生多镜头生成。把这些能力内化确实有信号:“LLM并不是很好的视频剪辑师——那会让人感到诡异。”与此同时,Runway的视频Agent已经可以处理从简报到分镜再到视频的工作流,以及广告效果反馈闭环。
  • 最大化的世界模型会纳入越来越多模态,押注能力迁移:Riffusion通过在频谱图上微调,把Stable Diffusion变成“一个相当可规模化的音乐生成器”;Runway则在The Well上微调视频模型,把从天体物理到原子尺度的数值物理模拟当作RGB帧处理,获得了“比从零训练快得多的合理表现”。长期目标是:不同于AlphaFold为数据有限的问题定制专门架构,把分散的科学数据汇聚到一个模型中。至于为什么现在不直接做Omni模型,答案是:“我们需要先解决机器人问题。”

14. 开放生态、中国排行榜与艺术家态度的显著变化

  • 竞争层面的承认是:“排行榜前10名、前20名里,大多数模型都是中国模型”,西方模型只有少数。与NVIDIA共同发起的Cosmos Coalition是开放生态方案,重点包括开放权重、物理基准和基础设施。面对swyx关于从中国模型蒸馏的挑衅,他质疑这是否是长期策略:“你的上限会被它的性能所限制”;Runway相信自己不靠蒸馏也能训练出强模型。
  • 他对艺术家反弹的“热辣判断”是:问题在于文生视频的叙事本身把反应推得过热。一个魔法般的单一提示词容易被理解为替代人类,而今天的模型需要“最多50个参考素材”,更像“一件工具,而不是一个能替你想好整部电影的魔法实体”。艺术家态度已经明显变化:戛纳导演和Ron Howard都曾在Runway电影节上表达支持。
  • 用户优先级经历了几个阶段:先是质量,再是可控性,如今转向延迟。即时生成10个结果的迭代速度,重新带回了“过去创意工具所拥有的一些魔力——Photoshop就是即时的”。9月30日的旧金山峰会将把争论搬到现场,参与者包括NVIDIA、Physical Intelligence以及DeepMind的嘉宾;讨论主题涵盖VLA与世界模型、遥操作与大规模视频数据、像素与JEPA与3D,而他“非常强烈地认为”最终会是像素胜出。
完整逐字稿
swyx

Okay, we're here with Anastasis from Runway, with me and Vibhu in the studio. Welcome.

Anastasis Germanidis

Good to be here.

swyx

Congrats on all your success and progress with Runway. You're opening offices all over the world. Did you envision this when you first started out?

Anastasis Germanidis

Not quite. I think even when we started, we had this idea that it was more a matter of when, not if. We were seeing the early generative models of 2016 and 2017 and just extrapolating, assuming resolution and quality would increase predictably over time. There was going to be a point where most content would be generated, and that was maybe the initial thesis of Runway: as a result of those generative models, we would need to rethink how creative tools are made.

As we built out the research behind our generative models, it then became clear that they were useful far beyond that as well.

swyx

And it is more obvious now with the real-world stuff and the world models that we'll talk about later. I'm just curious how you go from a background in Zocdoc and computer vision into Runway. Take us back to those early conversations with Chris and whoever else was on your founding team.

1. Art and Machine Learning

Anastasis Germanidis

I was always split into those 2 worlds. One was my own art practice. I was making a lot of interactive art for a long time. On the other side, I was working in startups, as an ML engineer and as a backend engineer at different companies.

I've always been interested in coding and computation, and especially interested in simulation. I brought it back into my early artwork as well. At the same time, I was interested in—

swyx

His personal site has a few, right?

Anastasis Germanidis

Yeah.

swyx

Is there one that we should pull up, just in case there's something that's like... I'd just like to go—

Anastasis Germanidis

Yeah.

swyx

—down memory lane here.

Anastasis Germanidis

For sure.

swyx

Okay, what is this?

Anastasis Germanidis

This was a project that I made, I think, back in 2015, where I built software that would give voice instructions to people in a gallery space. It would basically coordinate interactions between people.

It would first give you an identity, like, “You're an architect, you're 30 years old, and you like sports.” Then it would match you with another person, and you would have this completely generated interaction. Obviously, language models were not quite there at the time, so it was a mix of some templates and some Markov chain-generated text.

It would completely simulate these small-talk conversations between everyone in the gallery space. I was always very fascinated, on the one hand, with generative models and the early machine-learning work that was being done at the time. At the same time, there was a separate thread of simulation and what it means. What can we learn about humans by creating those very simple models of their interactions and their behavior?

Vibhu

Did you generate the prompts or the 30-year-old whatever? Was it you generating them? How'd you—

Anastasis Germanidis

Exactly. The program would just generate—

Vibhu

Okay.

Anastasis Germanidis

—those. A lot of it would be kind of Mad Lib-style. You have lists of different professions, lists of different—

Vibhu

Hobbies.

Anastasis Germanidis

—you know, personality types, lists of different ages, things like that. Then it would just combine those things together.

2. Runway Creative Thesis

Maybe the next project we can go to is Uncanny Road, which was—

Vibhu

GANs.

Anastasis Germanidis

—one of the first projects that we built with one of my 2 co-founders, Chris. This was taking pix2pixHD, which was one of the early image-to-image models that NVIDIA released back in 2016 or 2017. It was a model that would take a semantic map of a scene and then generate a photorealistic output.

Obviously, it was very early days, so it did not produce very high-fidelity outputs. But I think it was the first image-generation model that would generate at 1K resolution. It was all trained on self-driving datasets, so the semantic categories it would support were only things you would encounter on the road: pedestrians, traffic signs—

Vibhu

Stoplights.

Anastasis Germanidis

—you know, bikes, stoplights. That was actually one of our first indications. We built this, and people were making all this surreal imagery: a million-plus pedestrians, a million traffic signs, or gigantic humans.

It was an indication that you could take a model that was trained on this very boring dataset—essentially, not that many interesting things happen when you're on the road—and repurpose it, go very out of distribution, and make something that was artistically compelling. That was a summary of the thesis of Runway in some ways: you can take the same generative models, and if you look at them from another direction, build interesting tools around them, and give them to artists, they're going to do things that you don't expect.

Vibhu

Very cool. I like the UX of it, basically. You're just given an empty canvas: try whatever, do whatever. Then, with the other one, you see everyone with wired headphones. That's a sign that it's very easy—

Anastasis Germanidis

The Apple—

Vibhu

The Apple AirPods.

Anastasis Germanidis

—original ads. Yeah.

Vibhu

Yeah, yeah, yeah. Take us to today. You've been doing this for 7 years at Runway. How did we get to this? How did we go from driving-simulator data to all this? You kind of cover the whole stack of generative media.

3. Building Runway Research

Anastasis Germanidis

Interestingly, we're almost back in— we're full circle, where we're now applying our models beyond creative tools and into real-world scenarios. But it was a long journey.

Very early on, we realized the first version of Runway was a way to easily use all the open-source models of the day, things like pix2pix, and give them to artists. That was the initial idea: those models are too difficult to use if you're not a machine-learning engineer. What happens when you give them to artists?

Very quickly, we realized we needed to build a research organization inside Runway, and that happened maybe in year 1. A lot of the mandate there was that the image-generation models of the time, and the video-generation models of the time—or there were barely any video-generation models at the time—were not quite at the point where they could be productionized and brought into tools that would be part of creative workflows. We needed to push the frontier of the research.

The first 4 years of Runway Research were almost happening in the background, until there was a moment in 2022 with latent diffusion and DALL-E 2, where there was that step-function change. You guys may remember, at the time—

swyx

I started in this space because of, basically, latent diffusion and Stable Diffusion. I was like, “Wow, this is not only feasible, it is actually doable on consumer hardware.”

Anastasis Germanidis

Exactly, yeah.

Vibhu

I think the delta is also huge. I learned pix2pix. This was introductory ML. The TensorFlow Jupyter and Google Colab notebooks were like this, and then you had a sudden step-function change with diffusion and whatnot.

Were there any other changes since then? There were clear examples of what early diffusion was to get to here. Were there any other changes in key technology or research?

Anastasis Germanidis

Between 2018, when we started, and 2022?

Vibhu

Yeah.

Anastasis Germanidis

One of the early pieces of work that we did at Runway was solving segmentation—image and video segmentation. It was a very important problem because most VFX involves essentially separating—

swyx

Rotoscoping.

Anastasis Germanidis

—subjects. Yeah, rotoscoping. It's an extremely manual process. Nobody enjoys doing that.

A lot of the early days of Runway were spent building a tool called Green Screen, and for a long time it was the main thing that people were using Runway for. It ended up being used in Everything Everywhere All at Once and a bunch of other high-visibility films and series.

Essentially, Runway was a post-production tool for a long time, until latent diffusion and Generative Gen-1 and Gen-2 happened.

swyx

Cool. Let's go past that moment. You've come a long way, and then you started releasing your own models. Maybe describe that journey as well.

4. The Video Model Bet

Anastasis Germanidis

We go to the point, in mid-2022, when it became clear that we were doing research at a fairly small scale of compute. It became clear that scaling laws would apply to image and video generation in the same way they were applying to language generation.

We made a big bet. At the time, we signed this deal to build a cluster of 1,000 A100s. At the time, we were a Series B startup. That was almost a slightly irrational decision, maybe, but we really believed that if we trained a video model at large scale, we would get a great model at the end.

At the time, we set the goal around fall 2022: What does the Latent Diffusion, Stable Diffusion moment look like for video? At the time, the best model was called CogVideo. It was one of the early video models, with a 256-by-256 resolution and not very high quality.

We decided we were going to build out this cluster and invest in building our own video model. It became clear as we were training Gen-1 that it was difficult to get to fully— We wanted to build text-to-video, but it became clear to us that an easier starting point would be to start from video-to-video. When you have stronger conditioning, it’s basically an easier problem to restylize an existing video versus generating a video from scratch.

We released Gen-1 first in January 2023.

Vibhu

It’s just a fun visual podcast, honestly. You can see February 2023 was the state of things.

Anastasis Germanidis

It’s so interesting because, at the time, when you see those results, you think this is so incredible. It’s almost like image generation or video generation is solved, and then you look back a few years later and it’s obviously— You get used to results very quickly with those models. But at the time when we started seeing those results, it felt quite incredible, and the level of quality you could get was remarkable.

Gen-1 was a depth-conditioned video model. It would take an input video, first convert it into a depth map, and then generate pixels with a latent diffusion model.

swyx

Yeah. Very effective.

Vibhu

Yeah. I didn’t realize how distracting the blog post would be. Sorry.

Anastasis Germanidis

One of my favorite examples of Gen-1 was—if you go up to mode three or mode two, there was this storyboard use case where people would basically—

Vibhu

You can mess around with the room.

Anastasis Germanidis

make a city out of books or out of boxes, and then shoot a video with their phone and translate it into a photorealistic output. There were all these ways in which those models were starting to be used for storyboarding and for really—

If you go to mode four, you could take untextured 3D scenes and turn them into photorealistic output. We saw a lot of use cases early on where people who were familiar with the tools—power VFX editors—would take a Blender render and get it translated with Gen-1, or create a scene in Unity, capture a video of it, and then translate it, restylize it.

I still think video-to-video is powerful. We had a recent video-to-video model as well, and it’s one of my favorite ways of using these models: essentially using them to use ground-truth video as the initial inspiration and then translate it into different styles or different outputs.

swyx

I think we’re going to go into the rest of Runway and catch people up to speed today. I did want to cover the Stable Diffusion controversy—or what happened with Stability AI, whatever. I think there were sort of 2 sides to the story. Part of that is a normal thing of people joining and leaving companies. But what’s the retrospective now that there have been some years behind it?

Anastasis Germanidis

Yeah, it’s a very long story—

swyx

Which I remember you actually wrote a really long post about.

Anastasis Germanidis

We would probably cover the whole hour to go into it in more detail. Essentially, there was the Latent Diffusion paper that came out at the end of 2021. Patrick Esser, who was one of the researchers behind Latent Diffusion, worked at Runway at the time. He built Latent Diffusion in collaboration with Robin Rombach and a few other people at CompVis, which was a lab in Germany.

swyx

Like a research group, yeah.

Anastasis Germanidis

After releasing the early Latent Diffusion model, the goal was to keep working on versions of the model, scale it up, incorporate new data, and incorporate new tasks. Stable Diffusion was basically the same model trained on more compute, with a few more tricks. A classifier-free guidance paper came at some point, I think in early 2022, and that—

swyx

Which was a big prompting improvement.

Anastasis Germanidis

Yeah.

swyx

You know.

Anastasis Germanidis

That improved results. The model was trained on better data, like the aesthetic subset of LAION, but it was effectively the same underlying architecture. There was that big training run that happened on Stability’s cluster. Stability financed that run.

Looking back at that story, I think the work to build and train that model was done as a research project. It was a continuation of the Latent Diffusion work. I think the model became very successful, and as a result of its success, other companies tried to figure out the commercialization path for it.

For us, it was very important that we make sure that— It was meant to be an open-source research project. We decided that we should continue releasing versions of it, since that was the original goal of Stable Diffusion. That led to releasing Stable Diffusion 1.5. There was maybe a day of a bit of miscommunication there, but ultimately that was resolved very quickly, within hours.

swyx

Okay. I just wonder—you were actually one of the main players in that journey, and so it’s nice to hear from the source what happened.

Anastasis Germanidis

Yeah. I think it’s all in the past now.

swyx

Yeah.

Anastasis Germanidis

Both companies took their own path. Stability took its own path, and Runway took its own path.

swyx

Yeah. There’s still— I mean, James Cameron is backing the new Stability, whatever they’re doing with the Hollywood studios.

Anastasis Germanidis

Right.

swyx

I don’t know what they’re doing. I think one thing that impresses me, and I’m happy to move on, is that back in 2021 and 2022, there was this community of people that you were involved in researching all this stuff. From everyone I talked to who was active then, it seemed fairly obvious that somebody would do the hero training run that would—

Anastasis Germanidis

Mm-hmm.

swyx

produce Stable Diffusion. You had made investments and had the foresight. Is it accurate to say that this reflected what people were thinking at the time? Or was it still very much, “We’ll use it as a post-production tool or something. I don’t know”? Where was the sentiment back then? Maybe you can think back to what the community was like.

Anastasis Germanidis

I reminisce, and I think very fondly of those early years, from 2018 to 2022, because it was a very small community that, as you said, was very convinced that this was going to be a big thing. At the time, because it was such a small circle, everyone who would be part of that circle and make projects with it would immediately go viral.

swyx

And you don’t even know who they are, right? They’re just some name on GitHub or Hugging Face somewhere.

Anastasis Germanidis

Exactly.

swyx

Yeah.

Anastasis Germanidis

I remember one of the first big viral moments of creative AI. There was the Neural Style Transfer paper—

swyx

Uh-huh.

swyx

There’s something called DeepDream?

Anastasis Germanidis

I think it was called Neural Style Transfer.

swyx

Okay.

Anastasis Germanidis

There was also DeepDream, the puppy slugs, which was also really cool. There was this project by Gene Kogan, who was an early advisor of Runway and one of those—

swyx

Marketing guys.

Anastasis Germanidis

big creative AI folks. He literally showed a video of himself taking the New York subway and going over the Williamsburg Bridge, and then stylized it, I think in the style of Van Gogh or another painter.

At the time, that was so cool, and it went viral. It was a complete revelation to people that you could do this with generative models.

And that was only maybe 10 years ago, just as an indication of how quickly things have gone.

swyx

It’s pretty crazy. Even since then, you’ve got people at every level of the stack: devs, creatives, artists, hobbyists. Everyone’s using it. For people who tried stuff early, they’ll remember how hard regular diffusion was to use, right? Nowadays, you can use your favorite ChatGPT image gen or whatever, give it a sentence, and get a beautiful output. But diffusion was the whole Ultra HD, 4K, high-resolution thing; prompting these models was very different. Anything you learned on the tooling side from the offerings you guys have now—for creatives, devs—you really took the research and brought it to everyone to use. Anything interesting there to share?

5. Controllability Beats Text

Anastasis Germanidis

We had to build the entire model-serving infrastructure for video diffusion models. There was nothing else already available because Gen-2 was, I think, the first text-to-video model out in the market.

There were so many things that we learned over time. I think the biggest one was that it was very clear early on that text-to-video was not going to be the answer. People wanted a lot more control than that, so we invested in controllability on top of those models very quickly. How do you use the camera trajectory as control? How do you use an initial input frame as control? That was a very early learning for us.

For text-to-video, Gen-2 was an amazing step-function improvement in the quality of video models, but it was used much more in an exploratory way because there was nothing to ground it to. There was no reference that you could bring into it. You couldn’t really control the camera motion or the object motion. So the first year, in 2023, was really all about figuring out all the interesting ways in which we could condition those models. It was a lot of post-training rounds on top of the base model to figure out how people actually wanted to control them.

There was this quick succession of updates. One was called Motion Brush, where you could basically draw arrows and dictate where things should move in the scene.

swyx

That’s so cool.

Anastasis Germanidis

There was camera control, where you could describe how you wanted the camera to move in the scene. Because we work with filmmakers from most of Runway’s history, we immediately got this feedback and decided that this was worth investing in. Controllability became a big theme very early on as we were building those models.

Something fun that I haven’t really talked about too much was how Gen-2 came to be out of Gen-1. It was a bit strange because we announced Gen-2 2 months after Gen-1.

swyx

We’re accelerating.

Anastasis Germanidis

It was before Gen-1 was even generally available. Gen-1 was a depth-to-video model, so it would take a depth map and convert it into RGB. We couldn’t get text-to-video or image-to-video to work directly, and that’s why we started from depth-to-video.

We had discussions like, “Okay, we need to spend the next 6 months investing in text-to-video, maybe increasing the compute scale or the model scale, training a larger model.” I had this weekend project idea: what if I take a model that starts from text input and converts to depth maps, and then use Gen-1 to convert the depth maps—

swyx

It’ll probably work.

Anastasis Germanidis

—into RGB? And so Gen-2 was basically that.

swyx

Oh.

Anastasis Germanidis

Yeah.

swyx

The hackathon pipeline.

Vibhu

The weekend hackathon pipeline.

Anastasis Germanidis

Yeah.

Vibhu

But it looks good.

Anastasis Germanidis

It worked pretty well. With the knowledge that it has this 2-stage pipeline, you can tell in some cases that the structure of the video looks a bit off because you had to generate the depth first before you got to the output video. But it worked, and it allowed us to bring this to our users very quickly.

It’s interesting now because people are coming back to this almost 2-stage approach. If you look at the Reve text-to-image model that came out a few months ago, it had this planning model that would generate bounding boxes before feeding them into a diffusion transformer.

Vibhu

Yeah, Ideogram did the same thing that same day. I remember it was very strange that both of them came out the same day with the exact same innovation.

Anastasis Germanidis

It’s a small community, I think.

Vibhu

I’m like, this is—

Anastasis Germanidis

People—

Vibhu

This is completely coincidental, right?

Anastasis Germanidis

People talk. So, yeah, there’s definitely something to this approach. Obviously, now every single video-generation model in production uses a complex prompt-completion pipeline under the hood. I think it’s no secret that there is definitely—

swyx

Humans are terrible at prompting.

Vibhu

I think across the board.

Anastasis Germanidis

Yes.

swyx

But, yeah, I think the original Sora 1 blog post even told you that what happens after your input is that it rewrites your prompt and makes it much more descriptive of what you would want.

Anastasis Germanidis

Exactly, yeah. And there was the DALL·E 3 paper beforehand that was kind of the first public description of the fact that synthetic captions and really detailed captions work really well. Then Sora built on that.

6. From Video Models to Worlds

In 2023, we were releasing all these updates to Gen-2, like camera control and Motion Brush. There was actually something very interesting about camera control because it was the first time that you felt that instead of creating a video, or creating a short video, you were actually navigating inside the world. I think camera control was maybe the seed of some of the ideas that we had around world models and really opening up that research direction.

So this is not the original camera control. This was the updated camera control on top of Gen-3. But yeah, I think it made those models usable to filmmakers. I would say camera control was very popular. We realized that there was one way of seeing those models, which is that you’re just using them as content-creation machines. There’s another way, which is that as you’re predicting video, in order to predict video well, you need to simulate the world with increasing capacity. If scaling laws apply to video, just like they apply to language models, then as we scale the compute that we put into those models, they’re going to be able to simulate physics, human actions, and dynamics increasingly well and predictably well. That was the thesis behind our efforts on world models, and we spun up this research group to focus on world models and on how to turn the video-generation models that we’re building into something broader—something that would be useful beyond content creation as well.

Vibhu

And that was roughly when?

Anastasis Germanidis

That was in late 2023.

swyx

Interesting. A lot of people have been saying a lot of video-generation model companies have all pivoted to world models these days. But in 2023, you were already posting—

Vibhu

I mean, it’s debatable whether it’s a pivot.

swyx

Yeah.

Vibhu

Arguably—

swyx

Yeah, you—

Vibhu

—that’s what you always had to do anyway, right?

Anastasis Germanidis

Yeah. In a way, it’s an expansion—

swyx

Yeah.

Anastasis Germanidis

—of the applications—

swyx

Yeah.

Anastasis Germanidis

—of the models as they become more capable.

swyx

The early signs, it seems like the original models you guys had—people would say they were not “bitter lesson”-pilled, right? You’re adding rewritten prompts, you’re having all these one-off things, but that’s just the state of the technology as it was. Versus the future, as you said, you can scale it up to world models.

Anastasis Germanidis

Yeah. If you looked at the outputs of Gen-2, it was not obvious to people that this would scale to become a general simulator of the world. You had very limited movement, very low fidelity or low resolution, obvious mistakes in human anatomy, and all kinds of limitations.

But the idea was that this was just GPT-2. GPT-2 could barely generate coherent sentences. Similarly, Gen-2 could barely create coherent video, but if you scale it up, there’s no reason why it shouldn’t work in a way.

That’s always the mindset of Runway. Even when we started in 2018 and you looked at the results of the day, you needed to look more at the trend—where we were in 2018 versus where we were when the first GAN came out in 2014 or 2015.

And you started with 32-by-32 images of faces, and by 2018 you could generate street images at 1K resolution. It was kind of the same with world models: very early signs of something much bigger.

swyx

Yeah. I was almost thinking that it's kind of diffusing into focus. If you look at our visible output from year to year, it looks like a diffusion process itself.

Vibhu

Yeah.

Especially watching the early blog posts, you can really see the choppiness—the

Anastasis Germanidis

Yeah.

Vibhu

...details.

swyx

Like human civilization starting from randomized and then—

Vibhu

Yeah.

swyx

...denoising into—

Yeah, yeah.

Anastasis Germanidis

Future civilization.

swyx

Yeah.

Vibhu

That's how you know you're on track; you're still noising, right?

swyx

Yeah. I like the way that you guys phrased it when you announced it in June, which was, “The human mind is no longer the center of AI. Our world is.” Let's call the past 5 years of LLM-based AI very much about trying to emulate human preferences and human speech, but now that's mostly solved. I think that's some of the context of your essay, which you also wrote around that time, and now the focus is on modeling the world accurately.

Anastasis Germanidis

Exactly. The way we see it is, there is that initial mission statement of DeepMind, which is, “Solve intelligence and then use it to solve everything else.” But I think starting from everything else could be valuable. There is just so much complexity and detail in the world that it's hard to learn directly from just human descriptions of the world.

We're assuming that language models learn from everything that humans have written about the world—our own understanding as of the 2020s. There is just so much that we don't know and so much that's not captured by existing text about both the low-level dynamics of the world. We're not describing those in detail. If I tell you to describe how to tie your shoes, that's a very difficult thing to describe in words, but it's a very obvious thing to demonstrate.

There's Moravec's paradox: we're constantly underestimating all the complexity that goes into very basic things that we do subconsciously as humans, and we don't even necessarily always have the words to describe them. In my mind, simulating the world, simulating physics, and simulating the dynamics of the world have always been underestimated. We place too much emphasis on the things that are easy to talk about, but there's all this complexity and richness in the world that, if we just try to train directly on that observational data instead of training on how people describe the world, we would learn something new that we wouldn't otherwise know.

swyx

You think that the present architectural paradigm is fine? You don't need another layer like JEPA, like another famous New York AI leader would say?

Anastasis Germanidis

We're a very pragmatic research lab. If we have evidence that an approach works better than the approach we're taking, then we have no qualms about taking it. We just haven't seen any indication that video prediction itself doesn't scale.

Even if you look now, not just at our work but at the work of others, you're seeing that, in robotics, some of the most promising work starts from video-prediction models and then adapts them to also be action models, for example. So there's very little evidence that you need something else, and that your time is better spent on a novel architectural change compared with improving data and scaling the current approach.

We don't have any indication that there's merit to that counterargument. I think there was a tweet by Yann LeCun a few days ago saying that understanding the dynamics of the world is very different from generating cute videos.

swyx

And your answer is no, they're the same thing.

Anastasis Germanidis

Yeah, they're the same thing.

swyx

My cat videos are the same as understanding physics.

Anastasis Germanidis

Right. Because if you want to generate—obviously, video models can cheat. They could give you successive different shots of the scene in a way that doesn't require you to actually simulate difficult physics. There are always all these different ways in which you can hide the deficiencies of the model, and it's important not to be too tricked by the performance of current video models. It's easy to cherry-pick examples—

swyx

Mm-hmm.

Anastasis Germanidis

...and think that video models are further advanced than they actually are. There's a lot more work that we need to do to improve those models.

In my mind, it's very similar to language. You go from barely coherent sentences to something that could hold a conversation with a human, to something that could operate autonomously for a day and create entire code bases. There are obviously some architecture improvements along the way, but the main thing is scale.

It's the same bet for video, and we have no indication that this is not scaling. We have benchmarks that we use for measuring the physics of those models, and we see those predictably improve as we scale those models. If you want to Google up Physics IQ, that's one of those benchmarks that measures how well the model performs at solid mechanics, fluid dynamics, or optics.

swyx

I'm curious if you've seen any emergence, any scaling law around this.

Anastasis Germanidis

Yeah.

swyx

He's saying there is a scaling law, right?

Anastasis Germanidis

Exactly.

swyx

Yeah, I mean—

Anastasis Germanidis

The way those models and benchmarks work is, the researchers have gone and captured a few videos that are representative of different physical phenomena. You can take the first frame, pass it through an image-to-video model, and then generate a rollout that shows what should happen next.

You have a ball hanging from the ceiling, and then you use that as input, and the model predicts how the ball should fall to the ground. This measures the intuitive understanding of physics that we have. You can imagine what will happen next if I drop this bottle. So it's measuring that same intuitive physics understanding in those models.

We've measured that at different model and compute scales, and we see that the score in Physics IQ predictably improves. There are other tricks and techniques that you can use to improve the score even further, but even scale alone helps the model learn better physics.

swyx

My main sympathy with Yann LeCun is the Plato's cave allegory. You're learning from the output of a thing, not the internal process of a thing, and it's very, very noisy. If only you could observe the internals of a thing. It's hard to observe the internals of a human mind, but you can very much observe—or at least, we have a whole branch of science and physics that we're ignoring—how to model physics, movement, gravity, and other interactions.

We're just throwing all that away and saying, “Just scale data,” which is very much the lesson of unsupervised learning, but it feels wrong. That's the main idea.

Anastasis Germanidis

I think the history of machine learning, at large, feels wrong.

swyx

Yeah. It's a bit of a lesson right now. It's the simple answer to that.

Vibhu

I guess the question is, how much can you scale? Even on, let's say, the video-generation side, there's one side of video understanding and video generation. Are we still going to have tools where it's like, “I want to generate 2 hours, 20 hours”?

There's an infra way to do it in batches and stitch it together, but do we just keep scaling? Do we just continue with long-generation consistency, all that, with scale? Tying it into where we're at now, from looking at Gen-2 to Gen-4.5, technically, what advancements have we made to date, and where do you see things still going?

7. Scaling Beyond Sora

Anastasis Germanidis

Part of the answer is definitely scale. We learned that lesson in a big way with Gen-3. Gen-3 was the model we released the year after, in 2024. That was a few months after Sora was released. There is an interesting story of how that came to be as well.

Gen-3 was, for us, the first time that we really needed to build. Basically, we had to learn all the lessons that the language-model world learned in 3 years in the span of a few months. One of the biggest changes with Sora was using diffusion transformers instead of ConvNets. A lot of the early latent-diffusion models were all ConvNets for the diffusion-model part.

And the Diffusion Transformer paper came out at some point in 2023. It basically showed scaling laws for image diffusion transformers. We realized at that point that we needed to invest in infrastructure for model parallelism, to really scale training beyond a few billion-parameter models. We spent maybe most of the fall of 2023 building out our infrastructure for distributed training. We had a lot of false starts and failures trying to scale image and video diffusion transformers.

At that point, in February 2024, Sora came out, and the results were very much superior to what Gen-2 could produce. There was a lot of chatter on Twitter about Runway being done, that there was no way Runway would catch up. If you remember OpenAI in early 2024, it felt very much like they were—

swyx

To the moon.

Anastasis Germanidis

They were a formidable opponent. At that point, they were on top of their game, and nobody could even get close to them. Gemini—maybe the first version of Gemini—had just been released. So when OpenAI came out with Sora and it was such a big jump in quality, it gave me an existential crisis for a few hours.

But I think the amazing thing about Runway—and we’ve been around for 8 years now, which is almost being a dinosaur in AI—is that we’ve had a lot of those moments where we had to learn, adapt very quickly, and build out skill sets in the team that we didn’t have. If you ask anyone who was at Runway during that time what their favorite period was, it was that push over 3 months to get to a model better than Sora. We scaled the model size and the compute we were training on by 10×. We figured out model parallelism, even though we had zero expertise in it. Then we came out with Gen-3 that summer.

That was a big turning point for the company, where the research organization grew very quickly and we really started pursuing this vision of the general world model in earnest after Gen-3 came out.

swyx

Yeah. That’s the amazing thing about building when you’re building: there’s no stack. You have to invent everything yourself. You have to be completely full-stack. Now, there are inference specialists like Fal or whatever that can help with model serving, and I think you guys work with them as well. But at the time, it was just very interesting to think about what you do when Sora comes out and people are questioning whether your company should still exist.

Anastasis Germanidis

There was no vLLM for diffusion models. We had to build the whole model-serving infrastructure and make things efficient. A few months after we released Gen-3, we released Gen-3 Alpha Turbo, which I think was the first step-distilled model in production at the time.

swyx

That was a whole trend that we covered as well.

Anastasis Germanidis

So that allowed us to serve those models at larger scale, because I think the first version of Gen-3 was quite expensive to serve.

swyx

I think the whole trend in consistency models, Lightning, Turbo, and all these things somehow didn’t really stick around. I don’t know if you have any reflections on this, because at the time I was like, “Obviously everything should start with a distilled model first and then you can upscale, right?” Basically, your bigger models just turn into fancy upscalers. You should always draft with a smaller, faster model, because you can get it so quickly, almost in real time.

Anastasis Germanidis

I wouldn’t be so sure that it didn’t stick around. There are a lot of step-distilled models that are actively used in production. There’s still obviously a gap in quality compared to the non-distilled model, but in my mind, we’re still 2 to 3 years behind language models. It’s just a matter of time before there are better distillation techniques.

Right now, we have a real-time model called Character, which I think is the largest deployment of real-time video models. It’s a step-distilled model, and it’s actively being used. It’s a very specific use case compared to a general video model.

swyx

This is avatar stuff, right?

Vibhu

Consistent character.

Anastasis Germanidis

Yeah. This is a talking-avatar model. We optimized the hell out of it, and it generates at 24 FPS. It’s a step-distilled autoregressive video model.

If we look at our world-model direction, a big component of it is starting from bidirectional diffusion, which basically generates an entire video at once, and making it autoregressive. So you generate 1 frame or a few frames at a time.

There’s a lot that goes into that pipeline to get to a real-time model. First, you need to make it a causal, autoregressive model. Then you need to do some additional step distillation to get it to actually be real time. I think that part is just starting.

I’ll be very surprised if, 2 years from now, we don’t primarily use real-time models. To me, real-time video generation is inevitable. It has a much better user experience, it’s much cheaper to serve, and the quality gap between the base model and the real-time model is only going to close as we figure out better distillation techniques. We’ve made a lot of progress internally on maintaining the quality of the base model when we distill it.

Vibhu

How much of this is transferable? Is it the same base model? If you’re doing diffusion across the whole sequence and converting it to step-autoregressive distillation, is this a case where you still need to train both? Can you use the same base model and convert it? What’s that process like, going from a regular model to something that’s real time, on a technical level?

Anastasis Germanidis

The nice thing about diffusion models is that you have 2 axes of distillation. You can distill to a smaller model, which resembles what you do with LLMs, or you can distill in terms of taking fewer diffusion steps.

You could take a model that generates in 50 steps and make it generate in 4 steps. You get some performance degradation, but very often you get comparable outputs. You can even take a large frontier model, distill it with step distillation, and get to real-time performance. That’s what we’ve seen.

Depending on the use case, in some cases we might also start with a smaller model, but in a lot of use cases we actually just use the—

Vibhu

Step distillation.

Anastasis Germanidis

—the frontier model, and we’re able to make it work in real time.

swyx

I think this might be a good time to cut over to your laptop to show off some of the real-time stuff that you’re doing.

8. Neural Software Interfaces

Anastasis Germanidis

This is one of the research updates that we did recently. We’ve been working on getting our general world models into different applications. One of the applications that we think is very compelling is using general world models as essentially a universal interface to software.

This is a version of our world model called an interface world model. The idea is that it essentially replaces the front end of a software application. It renders the pixels directly for an interface, and it’s trained to predict what happens next as a result of a click or another interaction with the interface.

This is all pixels. There’s no HTML, CSS, or React powering this interface. This is directly the output of our real-time video-generation model, and it takes clicks directly as input.

swyx

And drags. Click and drag.

Anastasis Germanidis

Right. It supports clicks. It supports drags. It also supports scrolling.

The amazing thing about this is that you can effectively describe in the prompt how you want different elements to behave. You can turn an interface from a markup-language description of an HTML interface into something where you can just describe the interface: “If I press this button, I expect this to happen. If I press this button, this should happen.”

It’s useful, we believe, both for prototyping and for testing what different interactions would feel like. You can also add audio to it, so it’s a video-and-audio generation model. You can describe both what the visual outcome of your click should be and what sound effect should come out of it, if there is one.

We believe that’s going to be a much more flexible way of building software. Just render it. Why generate the code that generates the pixels? Just generate the pixels directly. It’s the end-to-end philosophy applied to front ends.

So we think there are a few interesting use cases. You can build creative tools on top of it. For any kind of use case that involves a lot of exploration, or an educational use case where you want to learn about a new concept and want some kind of visualization and open-ended exploration, we think this is a very powerful approach. You can imagine new forms of design in industrial-design software that could emerge as a result of those models. This is all generated in real time as well.

You can build a lot of interesting camera transitions and forms of interaction that are very difficult to build otherwise. One way in which we evaluate this is: what if you try to generate the same interface with Claude by prompting Claude, “Here’s an image reference of my interface that I made in Figma or that I created somewhere else. Create this particular interaction,” which in this case is dragging that object upwards? Beyond it being slower, it’s also very difficult to capture some interactions with just LLMs. So we think this is likely to be the way that a lot of software in the future will be created.

One of the additional benefits is that personalization might be a lot easier with those models. You can essentially try out different prompts based on who’s visiting the interface. You can more easily prompt-engineer the interface to have larger text for accessibility reasons, or if you have particular study preferences. We’re very excited about this approach. It’s obviously early days, and I think we’ll need to make it more cost-effective to serve those models as well. Running a real-time video model versus purely rendering HTML, the computational needs are obviously much higher. But we do see a lot of potential in this approach to building front-end interfaces.

swyx

So we covered a similar thing with Flipbook before, with our Ethan Ha episode about Grok video. I think it’s very engaging visually. I think it may be very good for education, but it does sound expensive. I think there’s an upper bound to how expensive it will be, though, right? The inference cost will go down over time. You’ll figure out ways to optimize it. Effectively, when it pauses, you’re not receiving human input; you don’t have to generate anything, right?

Anastasis Germanidis

Yeah. In this case, you also have ambient motion, so there are parts of the screen that might—if you want to visit Paris, let’s say, and then you get this interface that—

swyx

People walking.

Anastasis Germanidis

…allows you to explore.

swyx

Yeah.

Anastasis Germanidis

You have people walking or things happening. But it obviously makes it more expensive because you need to run the model all the time. Maybe you have some looping mechanism, so you don’t need to do that. But all those things are stuff we’ll need to figure out.

swyx

Yeah.

Anastasis Germanidis

I think our first consideration is: let’s find some use cases where it’s clearly a much more compelling interaction compared to traditional interfaces. Then it’s a matter of time before it becomes more cost-effective to serve.

swyx

Yeah. When it comes to the people walking, I think the approach that makes the most sense to me is basically MUGEN.

Anastasis Germanidis

Nik.

swyx

I keep messing up their name. With Chris Manning and Yunzhen Feng. I don’t know if you’ve come across them, where they basically mapped onto some kind of game engine. I think it’s Unity or something, or Godot, and they can obviously script some NPC behavior behind that and train on that.

Whereas here, you can really imagine whatever you want. That is a UI, right? It feels more tractable, I guess, to create a world model of software that is interactable because we have many examples of that. You can do your fancy RL environment stuff on that, rather than scaling up to embodied, real-world physical use cases. But this is a nice first step.

Vibhu

Or there’s the opposite: you have 1B models, 350-million-parameter language models. It just gets so small that they’re just predicting fishes moving and—

swyx

I mean, small models are now 120B.

Vibhu

Ultra-mini, on-device. But I think it puts it into perspective, at least the car one for me, in terms of the applications, right? The amount of work to do that—sure, you only make one model of your car per year—but applying this is also a cost saving, since you don’t have to manually make all this, right? So it opens up a lot of possibilities, too. I’m curious: if you extend this out 2–3 years, where do you see things going even further?

Anastasis Germanidis

Effectively, the end game of something like interface world models is that you have a fully neural operating system. I think Andrej Karpathy has written about that quite a while back.

To me, it’s a bit odd that, with an interaction with an LLM today, you have this LLM that can basically talk to you about anything. You can take the conversation in any direction. It’s very general, so it can solve all those different tasks, but you interact with it through a very rigid interface. To me, it’s just a matter of time before the interface itself becomes learnable and becomes part of the whole loop.

You’re delivering an application end to end, and that means you’re delivering the language model, but you’re also delivering the renderer and the pixels, and that’s also a learnable component. The concept of applications might not necessarily remain the same. I think we’ll need to figure out new abstractions for software.

The concept of applications comes from this idea that you need separate codebases to describe and power each individual tool and each individual application. But you might think of something a lot more unified if you have a video model that’s actually generating the interface as you go. It can take context from an LLM and allow you to combine different functionalities that traditionally would live in different applications. So it’s a way to solve software end to end, effectively.

We also see this as a powerful way to train computer-use agents. In general, with world models, there are those 2 directions: world models for humans and world models for agents.

Vibhu

Agents.

Anastasis Germanidis

For every new piece of work on world models that we do, both of these uses become possible. This is a powerful synthetic-data generator for training computer-use models. It could become a live RL environment that you could use to do online RL with a computer-use agent. You can get a wide diversity of interactions and kinds of interfaces generated on the fly to improve how robust your agent becomes. That’s the same with the world models that we’re working on for robotics use cases as well.

swyx

Is there a research breakthrough that you’re waiting for that would unlock the next set of use cases that you really want to pursue?

Anastasis Germanidis

Long context is a very important one: being able to maintain consistency for long periods of time. That depends on the use case. For our characters model, for example, or for the interface world model, it’s easier to maintain long sessions of interaction. If you go into more open-ended worlds that you navigate and take arbitrary actions in, the context and duration for which you can generate become limited much more quickly.

swyx

Yeah.

Anastasis Germanidis

We see more degradation and error accumulation happening. The biggest challenge with autoregressive models is error accumulation. You’re feeding generated frames back into the model to generate the next frames, and if there are any small errors, they accumulate over time. That’s not a new problem. It’s a problem that LLMs also have, and we’ve seen the ability to generate really long outputs now. So it’s a solvable problem, but it’s definitely still a challenge.

swyx

Yeah. What is the state of the art? For Grok, it would be 10–20 seconds of context going in there for video.

Anastasis Germanidis

With our characters models, we’re able to generate up to 30 minutes of video autoregressively.

swyx

Yeah.

Anastasis Germanidis

Um—

swyx

But that’s just for the avatars.

Anastasis Germanidis

Yeah.

swyx

Yeah.

Anastasis Germanidis

So if we look at GWM Worlds, which is more our open-ended world-exploration model, it’s on the order of a few minutes—

swyx

Mm-hmm.

Anastasis Germanidis

…which is still—

swyx

It’s probably enough for people because you have to cut to the next scene anyway, right?

Anastasis Germanidis

Yeah. It’s not the ideal game experience if you have to restart every few minutes.

For certain kinds of game experiences, you can work around it. Ideally, you are able to just generate forever and it doesn't degrade, and I think it's a matter of time before we get there.

swyx

Yeah. Genie has, like, 1—max 1 minute, you know?

Anastasis Germanidis

Right.

swyx

Yeah. Shoot.

Vibhu

You did a study on robotics. I think I also have just your Runway Robotics page, though. Is this better?

9. World Models Meet Robotics

Anastasis Germanidis

Last year, we released Gen 4.5, so that was our latest video base model. As I mentioned, we've been doing all this work in world models, and a lot of our approach to world models is: How do you take a bidirectional diffusion model, make it autoregressive, and make it accept actions? Instead of being a video you watch, it becomes a simulation that you step in, and you can control it every step of the way. You can explore counterfactuals, like what happens if I take this action versus if I take that action.

GWM-1 was the world model that we built on top of Gen 4.5. We did all this autoregressive distillation and then step distillation on top of Gen 4.5. One of the biggest use cases that we saw for GWM-1 was in robotics. One thing we like to say is that, as we scaled video models, we accidentally created a state-of-the-art model for robotics just by scaling video models.

We realized at some point mid-last year that robotics labs were coming to us and asking to use video models for synthetic data. They were asking us to post-train our video models to work really well for robotics, so that they could use them to generate variations. That was the first use case that we saw.

It increasingly became clear that the models would be useful beyond just creating synthetic data to train robotic policies. They would also be very useful as simulators. That means that you can use a video model online to test how your robotic action model performs. You can take an action, roll it out, get the outcome of the action inside the world model, and then continue that loop as a closed-loop simulation. You can use that to evaluate how well your robotics model works.

I think the biggest thing that you need to solve if you want to build a simulator is establishing real-to-sim correlation: If you take an action inside the world model, and you take the same action in the real world, you get a similar outcome. That was the goal of some work that we did earlier this year.

If you go to the first link, that was essentially what we wanted to establish: real-to-sim correlation for our world model, so that if you do a series of actions inside the world model and do the same actions in the real world, you get similar outcomes. We took our GWM-1 model and used some benchmark data. There is this RoboArena benchmark that's very commonly used to evaluate how well different action models perform.

We used the same scenarios, settings, and embodiments inside our world model, and we measured the correlation between how well the action model performed inside the world model versus in the real world. We saw that we could get very good correlation between our world model and reality. That means that if you want to evaluate how well your robotic policies perform, you can scale that much faster inside simulation instead of having to do it with actual physical hardware.

That was the first indication that our models could be quite useful in robotics. We saw, as we were working with robotics labs, that this became the first use case where they could use video models in a way that fit into their training pipeline.

Vibhu

Can I ask what—

Anastasis Germanidis

Yeah.

Speaker 2

—the difference was from 4.5 to solving that? The sim-to-real gap has always been the issue, right? You train a robotics model on video data, it doesn't generalize to the real world, and the simulation had an issue. It seems like you solved it, but how?

Anastasis Germanidis

A big problem with simulators is that if you're trying to simulate rigid objects, it works quite well if you can describe the physics of objects very accurately. Then you're able to use Isaac Sim, MuJoCo, or one of the traditional simulators. But for more complex interactions with cloth, for example, or slippery surfaces, with all the complexity that you want to be able to solve in manipulation with an action model that solves manipulation tasks, it's very difficult and time-consuming to build a simulated version of each of those environments and each of those tasks—a digital twin of that environment.

With a world model, you just need to provide the first frame, and then you can roll out the policy inside that first frame. We compared it to methods that required 3D-scanning an environment and then 3D-scanning each individual object before you can bring that into simulation. With a world model, you just take a picture of the environment, and then you're able to test how your policy performs.

Our general thesis on robotics is that there are companies leveraging a lot of teleoperation data to train robotics action models. There are now companies using UMI data, which is essentially human egocentric video where humans use robotic grippers to perform different manipulation tasks. Then there are companies focusing on egocentric data, where you strap a GoPro onto someone's head and capture them performing a task.

We think that all of those are great sources of data for training robotics models, but the most plentiful source of video data is third-person video data. How do we as humans learn how to perform different tasks? A lot of it is by observing others perform those tasks. We don't learn from first-person experience alone. We obviously do some trial and error to learn different things, but ultimately, a lot of what we learn how to do in the world, we learn by watching other people do it.

That's what you're doing when you're pre-training a video model. It's a lot of third-person video footage of people performing different tasks in the world: people doing sports, people doing household tasks. Our main thesis is that once you do video pre-training, you can then adapt a model to be useful in robotics use cases with way fewer hours of actual robotic data.

You require much less teleoperation data, which is very difficult to scale. Even if you look at egocentric data, which is a bit easier to scale compared to teleoperation data because teleoperation requires actual hardware, it's still 3 orders of magnitude less than the third-person video data that exists in the world.

Our thesis is that the most plentiful source of data will ultimately win. Third-person video data pre-training is the right starting point for models that you want to generalize and be able to deal with new environments, new tasks, and things that you haven't seen during training. That's the motivation for why we think our models are especially useful in robotics settings, and we've seen that to be the case as well.

swyx

You said pre-training. Third-person pre-training, first-person SFT—is that a curriculum that you can introduce?

Anastasis Germanidis

Exactly. So GWM Robotics starts from Gen 4.5.

Vibhu

It's the same video diffusion backbone, right?

Anastasis Germanidis

Exactly, yeah. You start from the base video model, the one you're using to generate cats and dogs and other interesting stuff, and then you fine-tune on a very small number of hours of robotic data.

It's on the order of hundreds of hours, compared to—if you were to pre-train a robotics model—the current pre-training runs go up to 100,000s or millions of hours of data. You're able to get quite good performance quickly because the model leverages everything it has learned about the world, physics, human dynamics, and the tasks that people care about from pre-training.

Ultimately, you want those models to generalize. You don't want to just be able to perform the tasks that it has been trained on. The diversity of actions and environments that you have with a pre-training video dataset is much larger than what you can realistically capture manually.

Vibhu

What does the scale look like for post-training? Do you still want to do—is it roughly 90% of the compute in a regular video diffusion model and then scale up a lot? Or is it more like, we want different robotic models for different tasks, or just the one really good base world model that can also apply to robotics?

Anastasis Germanidis

Currently, we are post-training our models for specific embodiments for particular partners. If they have a particular single-arm robot, bimanual robot, or humanoid robot, we would post-train our GWM Robotics model on their particular dataset.

Over time, we see the different variants of GWM unifying. I would expect that if, 1 year from now or 2 years from now, you have a single world model that can simulate manipulation tasks, it can simulate navigation—which a lot of the gaming world models do, since they're navigational world models where you're moving around the space—and it will also simulate human behavior. Those would be the character models.

Instead of having 3 different models, you have a single model that's able to—ideally, you're able to simulate what it's like to be in the world. You're moving around an environment, you may be performing different tasks, you're talking to other people, and that happens with the same single, real-time video model that's generating that.

Vibhu

Do you think you can solve self-driving? If you're learning to drive a car in a simulator, you have a world model. Your robot is basically a car, which can manipulate so many axes. How far off are you from something like that?

Anastasis Germanidis

So world models—

Speaker 2

Or a really good ADAS system?

Anastasis Germanidis

World models are definitely being applied to self-driving research right now, mainly for evaluation use cases, but our focus has been more on robotic manipulation. We've done some work on AV world models as well. We do think that world models and video models are the best starting point for both simulators and policy and action models.

That's the other side to this: once you have a great world model, you can just add an action head and it can predict actions as well. One way to think about it is if you take the starting frame of a scene with a robotic arm and ask the model, “Generate the arm picking up an object.” If it generates an accurate enough video, then it should also be able to generate the exact poses in 3D that the arm should take to perform the same action.

This is the direction that's now popularly called world action models. You're starting from a video model, then adding an action head to predict the actions, and it becomes a policy, essentially.

Speaker 0

One thing I'm also impressed by is how much data you actually need to train these kinds of models. You probably can't say exactly how much, but the original diffusion models and, from what I know, even the open-source Chinese models don't use that much data. Isn't it surprising?

Anastasis Germanidis

What do you define as much data?

swyx

And does it just come down to whether the token count is still relevant?

Anastasis Germanidis

It's a bit more complicated.

swyx

I mean, just gigabytes, right?

Anastasis Germanidis

Yeah. Hours of video, right?

swyx

Yeah.

Anastasis Germanidis

Yeah.

swyx

Yeah, yeah. I feel like something that's interesting is that the token-to-parameter ratios in language models have really—maybe they're 3 years ahead or whatever—seem to be a lot higher than in video models still, even though technically video has more information per bit. I don't know if that seems intuitive, or maybe there's just a lot of variability between 1 pixel and the next pixel that's not that high. So maybe there's just a lot of information that's repeated.

Anastasis Germanidis

My answer would be that it's still very early. Training video models will scale way further than they currently do, and you'll have capabilities that go much further than the current models can do.

One thought experiment that I like to use—it's almost like the Turing test of video models, or the Turing test of world models—I call it the lucid dream test. You have a—

swyx

You mean the actual person lucid-dreaming?

Anastasis Germanidis

It comes from this idea of—

swyx

Lucid dreams, right?

Vibhu

Lucid dreaming is realizing you're dreaming while you're in a—

Anastasis Germanidis

Yeah, exactly. So lucid dreaming is when you realize you're—

swyx

In a dream.

Anastasis Germanidis

—inside a dream, and then you basically—

Vibhu

Play around.

Anastasis Germanidis

—are able to control what happens inside the dream.

swyx

No, there's also an inference guy called lucidrains. Yeah. Or a quantization guy.

Anastasis Germanidis

Oh.

swyx

He's very prolific.

Anastasis Germanidis

Yeah.

swyx

Yeah.

Anastasis Germanidis

So let's say you have a VR headset, and you're in a room wearing it. Most of today's VR headsets have a passthrough mode, so you can see directly what's in front of you in the world, or you can obviously render something inside the VR headset.

There's going to be a point where those interactive, real-time video models become good enough that you wear the headset and you're in the same room, walking around and interacting with objects. You're able to move freely in that room and interact with any object. At the end, someone asks you, “Did you use passthrough mode, or was this rendered or generated footage?”

If you cannot tell for sure whether what you were seeing as you interacted with and moved around the world was generated or was passthrough mode—just what was happening in front of you—that's an indication that the models have become good enough. We're not close to that yet.

A lot of it is just this idea of really simulating dynamics and counterfactuals well. If you ask a video model to generate a person scoring a goal versus a person failing to score a goal, it would do a better job at scoring the goal because there's a bias in the training distribution. There are a lot more videos of a person succeeding at scoring the goal.

But if you have an interactive model, you want it to be able to generate counterfactuals. If I take this action versus that action, you want it to generate equally realistic outcomes. That's, I think, the big gap between video models and world models: the idea of counterfactual generation.

If you want a great model for robotics, you want to simulate failure very well because whether you're using it for evaluation or using it in an online RL loop in the future, you want to be able to have the model try and fail to do things and improve. In order to do that, you need to be able to simulate things failing.

swyx

This is the only domain where you have too many successful examples and not enough bad examples.

Anastasis Germanidis

Yeah.

swyx

It should be easy to generate failure.

Vibhu

Oddly enough, I think early image and video models weren't good at being human-realistic, right? You see a lot of high-res 4K professional photography, but not everyday life—normal pictures, right? Everything looks like it's professionally generated, like professional pictures, but not just normal, messy cables on a desk.

swyx

Okay. So there's this stuff. One thing we also covered was that you guys have video agents that you launched. I guess how does the traditional, let's call it frontier LLMs—autoregressive LLMs—feed into all this? Are they driving your robotics models or your video-agent production? Anything where you see the overlap between autoregressive and diffusion, let's call it?

Anastasis Germanidis

Harnesses are really important across all those different use cases. We have this video agent, which is essentially an LLM that's very effective at tool use with different image models and video models, and it helps you create a project end to end.

Very often, in a traditional advertising flow, you have a brief, then you generate a storyboard, and then you generate the video. A video agent, or Runway Agent, helps you through that whole process, and it also helps you analyze performance data. For example, how well did this ad perform versus this ad? Then generate more based on those learnings and figure out what to generate.

We think the harness is a very important piece of the pipeline. All production video models use some prompt completion, and we expect that to become more and more complex. You generate longer and more detailed descriptions before you use the diffusion transformer.

I do think eventually there is increasingly this unification into omni models, where you're training the models end to end to both do autoregressive text prediction and diffusion as well. So you're predicting the next token, maybe using some reasoning and planning of the scene, and then passing it into the diffusion head that's actually generating the pixels.

swyx

Yeah. I think currently maybe only Gemini and Qwen do it. I'm not sure which of the Chinese models are omni, but it's not a very well-popularized modality, I guess.

Vibhu

It's an interesting use case when you think about it, right? Not only do you have a language model reasoning and a diffusion head generating, but you don't have to stop with that output. You can feed that output back into the same model, reason again about improvements, and it can do a lot of loops on its own.

I guess the question is: do we need that, or can we just use an agent scaffold and do it outside the model? Is there a big benefit to doing it in?

Anastasis Germanidis

I think there's generally a trend: something is first done by a harness, and then—

swyx

Yeah.

Anastasis Germanidis

—it becomes part of the model, right? So you had chain-of-thought prompting, where you had to do this—

swyx

Mm-hmm

Anastasis Germanidis

—super-detailed system prompts to get that output.

swyx

Yeah, it's thinking step by step.

Anastasis Germanidis

And now, the model basically generates the reasoning trace by itself before it gives you an answer. In video models, similarly, a lot of the early video models were single-shot video models, and you had to use some kind of orchestrator to generate multiple shots in parallel and then turn them into an actual video.

swyx

Well, in ComfyUI, it's just all over the—

Anastasis Germanidis

Yeah, like a—

swyx

All these nodes.

Anastasis Germanidis

Spaghetti workflow. And now you have multi-shot video generation where you directly generate multiple shots.

There is a benefit to that because the video model learns something. To generate a single shot well, you obviously need to figure out a lot of stuff about the world. To generate multi-shot video well, you also need to get some video-editing instincts. You need to figure out the right pacing of shots.

LLMs are not that good at it. They're not that great video editors. If you ask an LLM to take some videos and auto-create an edited video out of that, it would feel uncanny. So I don't think LLMs are actually that good yet at being video editors, and I think there's a benefit to learning that end-to-end.

I would expect that, generally, the things that you need the harness for eventually get injected into the model itself, and you learn that end-to-end.

swyx

Do you find that you need to hire engineers who can—or researchers who are also artists—to infuse that taste, or do you have artists in residence to distill it?

Anastasis Germanidis

We have a large creative team that's very actively involved in training those models, in every part of the process. How do you caption videos well so that you capture the things you need for the cinematography, the aesthetics, and the camera direction in as much detail as possible, so that you're able at inference time to elicit that through the model?

Our creative team also does a lot of evaluation of what constitutes a usable video out of those models. They're very involved through every part of the process. I think that's one of the special things at Runway: that mix between creatives and researchers sitting side by side and working together to build the next generation of our models. I think that's been a really important piece of how we've operated as a company.

swyx

Yeah. In some senses, though, you can only do this in New York.

Vibhu

I mean, it's probably—

swyx

Maybe. I mean, you have other offices, but I'm trying to find some poetic significance in the fact that you are a big New York company.

Anastasis Germanidis

There are a few parts to being in New York. Obviously, there is that intersection of all those different industries, like media and advertising—

swyx

Yeah, this is very advertising.

Anastasis Germanidis

The art scene is New York. Not to say anything bad about San Francisco, but there is more going on. There is that component, and I think we also benefit from being outsiders and thinking of things a bit differently, like not being in the same hive mind of—

swyx

B2B SaaS—

Anastasis Germanidis

—of AI, of the Bay Area, and also taking our time to get where we are today. Growing the team intentionally and bringing people who are both on the creative side and on the engineering and research side. There's obviously a huge talent pool of amazing people in New York, so that hasn't really been a problem.

swyx

I mean, congrats on everything. What are you hiring for? What should people look forward to for the future of Runway?

Anastasis Germanidis

We're hiring across the board. I think this is probably the most open roles we've ever had in the history of Runway. We're growing our research team quite significantly.

If you're excited about video models, world models, or especially robotics, the robotics team is hiring across software, hardware, and research. So definitely reach out.

swyx

A lot of people don't have a direct robotics background, but what should they have if they want to be useful in robotics?

Anastasis Germanidis

Ideally, some experience with learned policies—

swyx

Yeah.

Anastasis Germanidis

—that would be—

swyx

Just RLs. Yeah.

Anastasis Germanidis

—good for robotics. We tend to hire generalists as a philosophy, and people who learn really quickly. But some experience and domain expertise in robotics is something we're definitely looking for in the coming months.

We're also scaling the go-to-market team significantly. There is very active enterprise adoption happening around video models at the moment, and we're really trying to respond to all the demand.

swyx

Yeah, great. You want to talk about the open-source robotics stuff?

Vibhu

Sure. These were just random notes we had. NVIDIA launched Cosmos, which I guess is interesting. You're a founding member of the Cosmos Coalition to build open-source world models in physical AI. Anything else to talk about here in terms of open research?

Anastasis Germanidis

The biggest thing is that, as I mentioned, world models are still early. There is still so much that we can scale in those models, so much more advancement and so many things we can figure out about how to improve them further.

I think it's important that some of this research happens in the open, and that we figure out what incentives there are for different companies to come together and bring some of that research into the open and open source. The Cosmos Coalition was an initiative that we co-founded with NVIDIA to bring some of that research into open source.

That could mean open-weight model releases. It could mean benchmarks that measure physics and the things that people care about when building world models. It could mean infrastructure. Really, how do we grow the ecosystem of world models and make it easier for a developer or researcher who's just starting out and excited about world models to contribute to the field?

swyx

I think there's some element of this being a response to the Chinese world models that are being released. Is that part of the consideration?

Anastasis Germanidis

I do think it's important for US video models. If you look at the leaderboards for video models, I would say that right now the majority of models in the top 10 or top 20 are Chinese models. Only a handful of companies from the US or the West have made it onto the leaderboard.

swyx

Yeah. We're doing better with images, but with video we're very behind, right?

Anastasis Germanidis

I think it's definitely important that we invest more broadly as a community to make sure that we have competitive models out there.

swyx

What's to stop us from distilling from them?

Anastasis Germanidis

I don't know if that's the best long-term approach.

swyx

Not the kind of data that you want.

Anastasis Germanidis

You're bounded by the performance that you can—

swyx

I mean, you know, like—

Anastasis Germanidis

It's almost a bit of a pessimistic view that you can get better.

swyx

It's free data. You might as well. If they're doing it on the text and language side, they might as well do it for the video side the other way.

Anastasis Germanidis

Yeah. I do think we're quite capable of training great models—

swyx

Okay.

Anastasis Germanidis

—without distillation at the moment.

swyx

Yeah, yeah.

Vibhu

Anything you have to say on benchmarks and evals? I feel like what I'm hearing is that a lot of people really like arenas for video and image models. Customers and whatnot only want the best on the leaderboard, and they refer to arenas a lot more than language models seem to. Any notes on benchmarks? What's lacking? How does the average person compare?

These both look really hyperrealistic. Beyond that, outside of the robotic simulation, the physics, and all that, is there anything to say?

Anastasis Germanidis

I actually think it's the opposite in some ways. Creatives, artists, marketers, and other people who are using our platforms generally rely less on arena scores. It's so easy to generate with a bunch of different models and then compare the results visually.

One nice thing about image and video models is that you can immediately tell with your eyes what feels good from an aesthetic standpoint. Any artifacts or issues with the physics of those models, you can immediately tell. So it's actually easier, I would say, to evaluate as a human.

There are also cases, as in language models, where you have very complex math and coding tests. It becomes a lot harder for humans to evaluate and discriminate between the performance of frontier models at the time.

So I think in practice people just test the same prompt with a bunch of different models and see what the results look like. Right now, in Runway, you can use our models and third-party models as well. So it’s very easy to do that.

swyx

Amazing. We’re going to end with the Runway AI Summit. The last societal issue, I guess—I don’t know if this is a thing—is the tension between artists, creatives, and AI. A lot of people in that community hate AI. Obviously, people who are in the Runway community don’t mind using tools; it’s just another brush. How have you seen the sentiment change?

10. AI Becomes Another Brush

Anastasis Germanidis

I mean, our perspective is, yes, it’s just another brush. It’s just another camera. It’s the latest of a long generation of tools—

swyx

Technology in art.

Anastasis Germanidis

Technology.

swyx

Yeah.

Anastasis Germanidis

And art and technology have kind of evolved together. I think there’s been a pretty significant shift over the past few months. Some of it came from a lot of public figures speaking out in favor of AI. In Cannes, you saw a few directors speaking in favor of AI. We had Ron Howard in our film festival. Mark Duplass is also adopting AI models.

So you have more of those stories coming out every day of a well-known figure speaking in favor of AI. It’s just a matter of, in my mind, those models becoming more and more demystified. I also have a bit of a hot take: one of the things that made the initial response to those models maybe a bit more heated than it needed to be was this idea of text-to-video, where you have a single text description and get back a full—

swyx

Demo tools.

Anastasis Germanidis

—full video. There was a misconception. Obviously, you can’t generate an entire feature-length film, but the models of today now take a lot of references. They’re very controllable. I think when people see a tool that allows for many degrees of freedom and control, they respond to it differently. It matters less that it’s a generative model than the fact that you can actually steer it in the direction that you want.

When people look at complex workflows on top of those models, when they look at all the ways in which you can steer them, and when you can provide, with some of the latest models, up to 50 references, the conversation becomes a bit different because it feels much more like a—

swyx

Storyboard.

Anastasis Germanidis

A tool—

swyx

Yeah.

Anastasis Germanidis

—rather than something magical that figures out your entire film for you.

swyx

Any notes on workflows changing for people in the field? I think engineering, at least, has had a lot of people whose expectations have changed. They’re 10 times, 100 times more productive, and they can get a lot more done. It’s the same thing as you’re making dev tools for creatives. Any notes there? Some people don’t want to adopt, and some do.

Anastasis Germanidis

Yeah. So I think, in terms of what people care about, we’ve gone through a few stages. We started from the stage where the main thing people were looking for was quality. As we scaled those models, the quality improved dramatically. That’s something people still care about, but it’s now in addition to controllability: being able to steer those models with references, different kinds of inputs, and storyboards.

Now my sense is that people are increasingly going to care about latency. As those models become better, the ability to iterate very quickly becomes more important. If you can generate 10 different outputs with a single prompt almost instantly, you can explore way faster than before, and you get some of the magic that characterized the creative tools of the past. Photoshop was instant.

We lost some of that with generative models. You’re waiting for 2 minutes to get back a video, and I think we’re going to bring some of that back now with those—

swyx

Real-time.

Anastasis Germanidis

—real-time models.

swyx

Yeah. Exciting. And—

Anastasis Germanidis

Exciting.

swyx

You’re finally doing this in San Francisco.

Anastasis Germanidis

The summit. Yeah, we’re very excited about this. In late September, on September 30th, we’re doing a summit primarily focused on physical AI and real-time video generation. We have panelists from NVIDIA, Physical Intelligence, Botco, and DeepMind. It’s going to be a very interesting series of conversations.

We try to make the panels really technical and elicit actual substantive discussion, and hopefully some interesting disagreements and debates on things.

swyx

Since you mentioned it, what kind of disagreements and debates should people think about? What do you expect?

Anastasis Germanidis

There’s one debate right now in the robotics world: VLAs versus world models.

swyx

Yes. Okay.

Anastasis Germanidis

There are labs that are really betting on one of those two directions. There’s also the question of what the best source of data is to train robotics models.

swyx

There’s just the third-party, first-party thing that we talked about.

Anastasis Germanidis

Yeah. There are the people who really believe in further scaling teleop data versus leveraging more large-scale video data. Those are some of the debates. And then there are the world-model debates: predicting pixels directly versus something like JEPA versus a more 3D-based approach.

I think we’re at a nice time in world models because there’s still that active debate about what the best long-term direction is. I feel very strongly that predicting pixels directly in video and scaling video-generation models is the right approach. But I think there’s a lot of interesting debate happening among researchers about the best path to take.

swyx

It’s interesting that it’s all on, let’s call it, the policy layer and the data-model layer. Is the physical side completely solved? All the sensors, all the actuators—do we have everything that we need?

Anastasis Germanidis

I don’t think that’s solved either. I think there are definitely—

swyx

Different problems. I want to dream about all these things, and then I buy a robot or try to assemble my own, and I can’t even get the motors to work right. You’re dealing with very sensitive equipment that has voltage, power, heat, and all these things. Abstracted away, we’re sitting here talking about software and models, but really you have to deal with those kinds of things too.

Anastasis Germanidis

Yeah. I’m generally also not opposed to incorporating other modalities into our models, as we’ve seen.

swyx

Yes.

Anastasis Germanidis

The simplest case is that they can generate video and audio at the same time. They can generate RGB, and they can also generate sound and audio. But I’ve written about this: What does the maximalist version of a world model look like? You’re incorporating more and more modalities from the universe—

swyx

Mm.

Anastasis Germanidis

—and you’re training a model on different scales of observation as well.

swyx

X-rays.

Anastasis Germanidis

And so, yeah.

swyx

You’ve got a good essay that people should read on real-world stuff. Meta released a model that was six modalities in one, right? I forget what the name of the thing was, but depth is one of them, and depth is a transformation of RGB in some sense.

Anastasis Germanidis

ImageBind.

swyx

ImageBind, yes.

Anastasis Germanidis

Yes.

swyx

What other modalities? They had thermal.

Anastasis Germanidis

Audio, depth, thermal, text—

swyx

Whatever IMU is. I do think you might as well do ultraviolet. You might as well do whatever other modality you feel like, because it’s all data to the model.

Anastasis Germanidis

Yeah. A big bet is also that there’s transfer between all those modalities. One of my favorite examples, which is quite old at this point, is a fine-tune of Stable Diffusion called Riffusion—

swyx

Mm-hmm.

Anastasis Germanidis

—which was—

swyx

The music one. Yeah.

Anastasis Germanidis

—yeah, just fine-tuning Stable Diffusion on spectrograms.

swyx

Spectrogram.

Anastasis Germanidis

And it became a pretty scalable music generator, right? So there are probably some spatiotemporal patterns that emerge at different scales and in different modalities. We’re talking about video, and there is some degree of meta-learning that the model has done that allows it to learn faster if you start from a model trained on images and train it to predict audio than if you train from scratch on just audio.

There are some other interesting examples. There is this project called The Well. It's a dataset of physics—numerical simulations in physics and biology and a bunch of other domains. It's essentially different physical systems across very different scales of space and time, from astrophysics to low-level atomistic interactions. We've done some work on this, and we've seen that we can take our video model—real-world video looks nothing like this—and actually fine-tune it on those numerical simulations and just treat them as RGB frames. You actually get reasonable performance much quicker than if you just train from scratch.

swyx

Hmm.

Vibhu

Yeah, I think we've seen this across languages.

swyx

Yeah, DeepSeek-OCR as well.

Vibhu

Yeah, DeepSeek-OCR is cross-modal—

swyx

You don't have to tokenize text. You can just throw it in as images.

Vibhu

There's a lot that happens in that base pre-training. There was an argument a long time ago of people saying, “Oh, humans have so many sensory representations, right? Smell, touch. Models have a whole 2 more modalities that we don't even have data for.” And it's like, okay, you take a cheap sensor, you can try this stuff, but actually, there's so much happening in just the base training that you don't get as much from these little things.

Anastasis Germanidis

Yeah, exactly. I think what it solves is data scarcity. You have so much video data available, but you don't have olfactory data.

Vibhu

The cool thing is it goes the other way too, right? So if you want to do physics, if you want to measure this, or if you want to have a diffusion model do audio, it transfers really well. In your case, the little bit of post-training for robotics gets a video model to use its fundamentals in another domain, so we can apply that to other stuff too.

Anastasis Germanidis

Right.

Vibhu

Yeah.

Anastasis Germanidis

If we look at how to make those models more useful for scientific domains, and if you look at AlphaFold, because of the data—the limited amount of data that it needs to be trained on—it basically has a very fine-tuned architecture just to solve protein structure prediction. But if you take all those disparate sources of scientific data and bring them together under a single model, I think that's an approach that can help us solve new kinds of problems across science by leveraging all the learnings from one modality or one set of data to another. It's very early days for that direction, but I do think that's ultimately what the end game of simulating the world is. You're not just using RGB; you're using RGB as a starting point, but you can incorporate more and more modalities of the universe and leverage the transfer that happens from learning from one to the other.

Vibhu

I guess the follow-up there is: what's the drawback of omni? Why isn't everything an omni model? Why not now? Would you start from a language backbone or an image/video backbone and then go omni from there? Does it matter?

Anastasis Germanidis

Yeah. We need to take it one step at a time. We need to solve robotics first, and then we can go into—

Vibhu

Solve everything now.

Anastasis Germanidis

Yeah. I do think there is a lot of open-ended research that needs to happen for those omni models. There are a lot of things that require careful consideration when you're bringing multiple modalities into a single model for prediction. But I expect those to be solvable.

swyx

Wonderful. You've been very generous with your time. Congrats on all your success, and I'm excited for the AI Summit—or Physical AI Summit.

Anastasis Germanidis

Yeah, thanks for having me.

swyx

And people should check out the film festival if it's in town, right?

Anastasis Germanidis

Yeah.

swyx

You're going to be touring all over the place.

Anastasis Germanidis

Yeah, next year we're probably going to do that. We do film festivals every May or June, and we did the last one in New York, LA, Tokyo, and at the AI Engineer World's Fair—

swyx

Yeah.

Anastasis Germanidis

—fair.

swyx

Yeah.

Anastasis Germanidis

Hopefully even more places next year.

swyx

No, I think someday you'll be hosting the Oscars of AI video, and I think people should take this very seriously as a potential career they can have.

Anastasis Germanidis

The Oscars of AI video will be called the Oscars.

swyx

All right. Thank you.

Anastasis Germanidis

Thank you.