Luma Labs 的扩散革命:从 Dream Machine 到多模态世界模拟——Amit Jain、Jiaming Song
- Luma 的“产品到模型”闭环是一套反复奏效的策略:围绕当下模型搭建用户需要的能力,验证需求、收集数据,再将其内化进下一代模型。 Ray 1 需要大量外部支持;Ray 2 所需支持“不及其10%”,但角色理解在 Ray 2 中仍位于模型外,计划在 Ray 3 中内化。Amit Jain 的判断是绝对的:“凡是能在模型里完成的,都将优于在模型外完成的。”
- 分布外生成的关键,不在于寻找不可能的训练样本,而在于构建一个能将其拆解为可复用概念的基础模型。 数据里没有“坐在牛油果椅子上的腌黄瓜”,但有椅子、坐姿、拟人化物体和腌黄瓜;Amit 认为,“这些模型不是在记忆行为”,而是在提炼可重新组合的基础能力。这意味着 Luma 的核心命题落在数据质量、高效学习和表征深度上,而不只是原始数据集规模。
- Concept School 是 Luma 连接缓慢预训练与专业用户新电影控制需求的桥梁。 它能让 Ray 2 用1到几个样本学会一种动作、姿势或调色,同时不损害既有能力,并支持概念组合;Bolt Cam、dolly 和 reverse-dolly 控制就是早期案例。商业目标十分明确:教会模型“一切电影制作知识”,快速发布新概念,最终让客户自己“把模型带到学校”。
- Luma 把视频生成视为 AGI 计划,而不只是创意软件品类。 Jiaming Song 称视频模型是通往通用智能的“关键路径”,因为故事要求模型跨时间理解因果顺序、角色弧光和后果;多模态模型必须联合推理语言、图像、视频和音频。Amit 的逆向判断是,创意工作之所以最早浮现,恰恰因为“那正是我们需要智能的地方”。
- Luma 最看重的可解释性工作发生在训练过程中,因为此时干预仍能改变模型最终成为的样子。 Amit 把 Golden Gate Claude 这类事后特征发现比作“考古”;Luma 则研究表征形成过程中的信息流、频率学习、课程安排和超参数。他说,基础路径在最初20,000-30,000次迭代中已大体确定,因此早期数据分布和学习率选择的重要性被显著放大。
- Jiaming Song 的 Inductive Moment Matching 瞄准生成模型的经济学三难题:高质量、少推理步数和稳定训练。 它把一致性模型从逐点匹配推广到分布层面的匹配,并用最大均值差异规避 GAN 中训练判别器的内循环。在 Luma 的消融实验中,1个样本不稳定,2个样本后期变得不稳定,4个或更多样本则让训练稳定下来;这项研究可能延续推理成本的快速下降。
- Luma 最大的战略押注是:把图像和视频嫁接到语言骨干上,不会得到原生的多模态智能。 Jiaming 指出,当前多模态语言模型在物体识别、分割和长程因果链条上仍可能落后于专用视觉系统;Luma 希望构建统一潜空间,因为自然界“不区分视频、图像和音频”。Nathan 的框架则补充道,在 scaling law 主导的竞争中与超大规模云厂商正面较量会很难,但上行空间在于一条差异化的多模态 AGI 路径。
1. 强大的基础模型能把不可能场景还原为熟悉的原语
Stephen Parker 提出、Amit 也认同的一点是,越来越多的图生视频流程从生成图像开始,因为图像迭代成本低:创作者可以生成100个候选,找到理想构图,再把胜出者做成动画。这一流程凸显了陌生输入的动画化难题,包括混合物体、不熟悉的角色,以及现实世界中没有直接先例的场景。
Amit 给出的代表性测试是“坐在牛油果椅子上的腌黄瓜”。训练视频里没有这一确切事件,但模型见过腌黄瓜、椅子、人坐下以及拟人化物体;成功的关键在于理解每个原语代表什么,再将它们重新组合。Amit 说:“这些模型不是在记忆行为”,而是在提炼基础能力。
更强的基础模型还能在更长的生成过程中保留初始图像的身份和美学。一些可靠性来自系统在幕后检查并翻译图像,但 Amit 将这类系统称为暂时性的“拐杖”:它们可以暴露需求,直到后续模型直接吸收这项能力。
2. Luma 不断把智能从脚手架迁入模型
Amit 将“模型内还是模型外”视为核心设计选择。外部软件可以分别构建光照、叙事和角色行为,但潜空间包含更丰富的信息,能够同时协调外观、因果、动作和时序。“凡是能在模型里完成的,都将优于”其外部等价方案。
最初的 Dream Machine 模型,即后来追溯命名的 Ray 1,依赖脚手架处理运动、词汇翻译和其他控制。Ray 2 所需支持已不到原来的10%,但更难的客户需求又促成了新一层外部系统。例如,角色理解在 Ray 2 中大部分仍位于模型外;Luma 计划在 Ray 3 中将其内化。
Amit 用“血脑屏障”比喻用户在模型外的智能与生成模型在模型内的智能之间的隔阂。可控性意味着用更高层级的指令穿透这道屏障,然后让模型在内部整合角色弧光、事件顺序、光照及其他交互,而不是通过工作流逐项微操。
Jiaming 所设想的终点类似于和另一个人沟通:用户不应被迫分别编程图生视频、关键帧和镜头运动流程。一个智能的多模态模型应能接收媒体和意图,推断任务,自然完成部分来回沟通,并“无需……把它当成另一种类型的任务”就执行下去。
3. Concept School 把稀缺样本转化为电影控制
视觉模型还不具备语言模型那样稳健的上下文学习能力,但专业人士不断创造网络上不存在的新动作、姿势和调色。Luma 的“概念”旨在用1到几个样本学习这类能力,同时保留基础模型的其他技能,并彼此组合,避免传统微调或 LoRA 带来的能力退化。
Luma 将内部教学工具称为 Concept School:“你把模型带到学校,而你是老师。”团队用它从稀疏样本中较快教会模型 Bolt Cam、dolly、reverse dolly 及其他镜头运动。Amit 说,另一批概念原定当周发布,再下一批则安排在下周。
Ray 1.6 已提供镜头运动控制,但 Jiaming 表示,新实现更加灵活,不需要对镜头坐标进行如此字面的指定。下一个目标介于 Bolt Cam 预设和精确轨迹编程之间:用户可以交互式调整速度、角度变化或主体焦点,同时由模型维持场景一致性。
Stephen 的类比更清楚地说明了产品价值:《人生切割术》第二季开场的 Bolt Cam 据称耗时数月并需要机械臂,而 Luma 将这一想法暴露为一种生成控制。Amit 希望同时提供日常工具——跟拍、向左或向右移动——以及“荒诞又好玩”的效果,但他强调:“我们是为专业人士打造的。”
4. 叙事是 Luma 从视频软件走向通用智能的路径
Amit 将 Luma 的使命概括为“多模态通用智能”。这种智能具体化后,可能类似“地球仪里的世界”:一个必然不完整的近似体,包含物理现象、相互作用的智能生命,以及随时间展开的后果。他更愿意使用“世界模型”,因为“模拟”听起来无法覆盖目标范围。
创意工作之所以重要,是因为它要求系统超越流程。故事要求事件 A 导致 B、B 导致 C,并让角色沿时间线发生变化;这些依赖关系比单纯生成结构化输出更接近智能。Amit 借用了语言模型的类比:当结构化输出从外部强迫变成原生能力后,“突然之间你就有了智能体”。
“视频模型是通往通用智能的关键路径,”Jiaming 说。一个能够联合推理语言、视频、音频和图像的模型,可以成为帮助人们做梦、想象和追踪后果的伙伴,而不只是制造像素或沿着预设生产图执行的工具。
Amit 有意颠倒了历史预期:人们原以为机械劳动会最先被自动化,但 AI 最早的帮助却出现在创意领域。他的解释是,艺术需要抽象能力和一般性思维:“那正是我们需要智能的地方。”因此,艺术是构建 AGI 的重要组成部分,而非偏离主线的消遣。
5. 虚构知识可以改善现实世界中的行动
Nathan Labenz 提出了一个潜在冲突:机器人需要可靠的物理知识,但 Luma 的训练数据也包含龙、魔法和不可能的电影物理。Jiaming 承认,短期内用于家务活动的系统,可能应更多偏向物理和可应用领域,以便“少产生一些幻觉”。
但即便是日常指令也可能需要虚构知识:“拿起那件印着龙的衣服”,如果机器人无法识别龙,就会失败。这个想象生物本身就是现实观察的重新组合——类似蜥蜴的形态、爬行动物鳞片,以及蝙蝠或鸟的翅膀——因此,虚构概念可以成为现实环境中物体的有效识别抓手。
Jiaming 的长期答案是让模型根据上下文定位自身。一个模型应能判断自己是在物理房间、虚拟界面还是想象世界中行动,就像人可以先把印着龙的 T 恤放进洗碗机,再转到电脑上回复邮件。专业化目前可能有帮助,但不应永久成为必要条件。
6. 视觉理解的标准是效用,而非类人的表征
Jiaming 指出,模型已经具备隐性的3D知识:它们能从真实和幻想图像中推断深度,并表征布料、波浪和头发。研究人员还将预训练生成模型作为深度估计的先验,包括 Jiaming 所称的最先进方法;相比定位一个离散的“深度神经元”,把这类知识激发出来可能更容易。
Stephen 的反驳值得保留:也许视频模型只是对2D像素场进行了越来越专业的预测,并没有内部的三维空间理解。Jiaming 将问题重新定义为视觉版图灵测试——完美预测一个场景,并不能证明模型使用了人类定义的物理学,但从效用角度看,功能等价可能已经足够。
Amit 质疑人类是否真的维护着明确的3D表征。大脑接收视觉信息以及本体感觉等其他信号,但主观上确定“这是3D”,并不能说明背后的机制。因此,生成模型不必包含一个网格;只要能生成一致的结果,隐性表征就可能足够。
他的类比将现象与机制分开:飞机产生升力却不拍打翅膀,人形机器人以不同于人的方式运动,洗碗机无需双手也能清洁。20瓦功率的有机大脑与 GW 级计算集群处于不同基底上,要求它们使用相同表征,可能是在限制机器,而不是判断机器是否具备能力。
7. 可解释性在模型形成过程中最具行动价值
Nathan 接受 AI 不必像人类一样思考,但质疑 Amit 对 Golden Gate Claude 可解释性研究的否定:研究人员分离出一个疑似金门大桥特征,放大它并改变了模型行为。概念控制是否也能以类似方式识别并调动动漫、电影风格或其他用户要求能力的潜在方向?
Amit 区分了两件事:假装一个大型模型完全可读,和对黑箱运行经验性实验。“科学不是采访上帝”;物理学从观察中提取模式,而量子力学甚至可以在不解释现实为何如此的情况下,产生经过6、8、10乃至有时23个 sigma 验证的测量结果。
因此,事后特征研究“像考古”:对于陌生模型,它在智识上有用,但只能重建已经形成的东西。训练过程提供了更强的机会,因为研究人员可以见证“这颗行星的形成,以及它经历的所有亿万年”,并仍能改变它的数据、课程、信息流架构和优化机制。
Luma 跟踪高频和低频信息发生了什么、哪些内容在早期学习、哪些在后期学习,以及课程何时应当改变。Amit 说,信息路径大约在最初20,000-30,000次迭代中建立;后续训练会强化或压制这些路径,就像田野中形成小径和麦田怪圈,之后又被反复走过或抹去。
8. 精选数据胜过无差别扩张,但不存在干净利落的配方
Nathan 问 Luma 是否会设计每个 batch,以避免有害的 loss 峰值。Amit 澄清,团队并不是手工挑选每一个 batch,而是通过大规模筛选、过滤和剔除坏样本来处理问题。“这种技术……不是把10亿个垃圾数据样本扔给它”,而是用高质量样本告诉模型应该学习什么。
Jiaming 指出一个更深层的不匹配:训练算法通常假设样本相互独立且同分布,但是否存在一个连贯的“世界分布”,以及该如何定义,仍不清楚。例如,人口占比未必是决定语言模型语料库中英语应占多少的正确规则。
因此,数据集设计是目标驱动且经验性的:改变数据混合,观察质量和目标能力是否改善,然后继续迭代。Jiaming 提醒,确切解法“可能非常混乱”,其中可提炼的简洁统计原则,远少于通用机器学习的简化描述所暗示的数量。
9. 扩散模型凭借让互联网规模的生成训练变得可靠而胜出
Jiaming 将已知最早的扩散方法追溯到 Jascha Sohl-Dickstein 于2015年发表的 NIPS 论文。它在小数据集上没有形成多少势头,因为 GAN 可以一步生成且在规模化时表现更好,而扩散模型速度慢、难以证明其价值。2020年 Jonathan Ho 及合作者完成的 DDPM 工作,在选定任务上让扩散模型具备了竞争力。
DDPM 仍需要约1,000次神经网络评估,但它消除了 GAN 不可预测的不稳定性。对实践者而言,这是一次范式改变:“你启动任务,晚上可以安心睡觉”,预期模型会继续变好,而不是醒来发现一次运行毫无预警地崩溃。
Jiaming 的 Denoising Diffusion Implicit Models 工作攻克了采样成本,当时实现了约20-50倍加速。与此同时,Yang Song 的 score-based 研究将离散去噪连接到连续时间随机微分方程和1980年代的数学;之后在 ImageNet 上的实验则显示,扩散模型在更大规模下可以与 GAN 竞争。
随后,控制从 classifier guidance 发展到 classifier-free guidance;约在2022年前后,GLIDE、Imagen 和 Stable Diffusion 相继出现。Jiaming 用贝叶斯定理解释 classifier-free guidance:条件和无条件模型输出提供了将生成方向引向 P(X given Y) 所需的项,无需单独训练分类器,同时保留条件控制方向。
10. 蒸馏与矩匹配用分布替代僵硬路径
Progressive distillation 训练一个模型步骤,用来替代相邻的2个去噪步骤,然后重复压缩,获得4倍及更高加速。它确实有效,但 Jiaming 称实现过程“很棘手”。Consistency models 则试图让去噪轨迹上不同位置得到相同的最终预测,有望实现单阶段或极少步数生成。
在实践中,一致性训练的稳定性低于其简单形式所暗示的水平,因此,从既有扩散模型初始化的一致性蒸馏更受欢迎。这个领域重新面对旧有权衡:一致性技术通常更容易训练,而 GAN 方法在极少步数下可能有更高质量,但会继承判别器不稳定的问题。
Inductive Moment Matching 放松了一致性模型的逐点约束:生成样本不必遵循某一条精确映射,只要其分布与目标匹配即可。它在再生核希尔伯特空间中使用最大均值差异,其最优比较函数无需反复训练神经网络判别器即可表示,从而消除了 GAN 不稳定的内循环。
消融实验支持这一机制:1个比较样本表现得像不稳定的退化一致性案例;2个样本仍有一定不稳定性,但后期失败;4个或更多样本则带来稳定训练。Jiaming 将其视为同时实现3个目标——高样本质量、少推理步数和稳定训练过程——的进展。
11. 原生多模态需要的不只是语言骨干
Jiaming 在结尾的判断是,行业目前所称的多模态“其实并不是多模态”。微调后能够接收图像、音频或视频的语言模型,在语言任务中可能很有用,但在物体识别、分割和追踪长程因果链条上,仍可能不如专用视觉模型。
警示信号来自架构:如果一个所谓的多模态基础模型在基础感知能力上仍落后于专用系统,那么把所有信号嫁接到语言骨干上,可能并不是终局方案。bag-of-tokens 方法仍有价值,Jiaming 也承认语言领域还有大量工作要做,但他认为行业需要“超越语言骨干”。
Luma 正在推进“统一的单一潜空间”,将语言、图像、视频和音频视为同一事件的不同侧面。“自然界不区分”这些信号;一个行动会在同一环境中同时产生它们。Luma 的押注是,对这套共享结构进行推理,最终会产生真正的多模态智能,而当前的拼接式整合只能近似实现这一点。
Today I'm speaking with Amit Jain and Jiaming Song, CEO and chief scientist at Luma Labs, makers of Dream Machine and the new Ray 2 video-generation model. I'm also joined for this episode by my friend Stephen Parker, creative director at Waymark and one of the few creators who has logged a proper 10,000 hours with video and image-generation models, dating back to the original DALL·E over the last few years.
Our conversation begins with a discussion of how the Luma team trains models to create fantastical and other fundamentally out-of-distribution visuals for which there is little to no relevant training data available. Considering the force of intellect that both Amit and Jiaming display, their belief that video models are on the critical path to AGI, their ambition to create multimodal AGI at Luma Labs, and the range of novel and occasionally hot takes they share, I think this episode should be of interest to anyone, regardless of whether you're particularly interested in video-generation models specifically.
Keys to Luma's model-development success, as you'll hear Amit explain, include a relentless focus on dataset curation, frontier advances in efficient learning algorithms, and a strong drive to understand what their models are actually learning as they go through the training process. These fundamentals create base models that can learn new concepts, including Bolt Cam and many other camera-motion concepts they've recently introduced, in a highly sample-efficient way.
Meanwhile, for things that existing models can't learn so quickly, we also discuss Luma's outer loop of product development, which consists of building scaffolding and other behind-the-scenes systems that unlock new model capabilities and also validate customer demand. With that done, they then seek ways to internalize those capabilities in the next generation of the model. Then they repeat this process for each generation as customers continue to apply new and better models to harder and more valuable challenges.
For me, the most interesting part of this conversation was the discussion of model interpretability. Emphasizing that we should not expect AIs to process, represent, or understand information like we humans do, or even to do so in a way that's generally human-graspable, Amit likens current interpretability techniques to archaeology, in the sense that they're fundamentally limited to piecing together what models have already learned in the past. More interesting from his perspective is the study of training dynamics and the engineering of datasets needed to teach models what they most need to know.
In the last 15 minutes or so, Jiaming offers an intellectual history of diffusion models. This gets pretty technical and, for most people, myself included, will require some additional study to fully understand. But I would summarize it by saying that the generative AI era really began with the realization that, with the right problem formulation, unsupervised learning can work on web-scale datasets. For text, this was simple next-token prediction. For images, it was gradually adding noise to real images and then training models to remove that noise one step at a time.
Since then, there have been a mix of practical tricks, theoretical insights, and model-enabled dataset improvements that have unlocked far more precise steering of outputs and breathtaking efficiency gains. These range from distillation techniques, which amount to training a model to perform multiple denoising steps in a single pass, to consistency models, which try to ensure that a model will generate the same output regardless of where it begins on its denoising path; to flow-matching models, which use theoretical connections to differential equations to take a more direct path through the latent space; to Jiaming's latest inductive moment-matching technique, which optimizes the model in distribution space and performs generations in a small number of optimized steps. All of this should, at minimum, give you a sense of how inference prices have fallen so precipitously even as quality has dramatically improved.
While we didn't have time to go as deep into the philosophical underpinnings of Luma's multimodal strategy as I might have wished, I left this conversation with the sense that Luma Labs is definitely a company to watch. It won't be easy for any model-development startup to compete with the big-tech hyperscalers in the scaling-laws era. But Luma's mix of product-market fit, vision and ambition, and research prowess gives them as good a chance as any I've seen. I absolutely look forward to having them back again in the future.
Amit Jain and Jiaming Song, CEO and chief scientist at Luma Labs, makers of Dream Machine and Ray 2, welcome to The Cognitive Revolution.
Thanks for having us. It's very exciting to be here.
Yeah, thanks for having us. I'm excited for the conversation as well.
I'm also excited to have my good friend and longtime teammate Stephen Parker here as well. Stephen occasionally co-hosts when we do an episode on creative models, especially in the image and video domain, because he is the creative director at Waymark and has logged more hours than anyone I know with these kinds of products. He really has an excellent handle on the exploding array of options in the market today.
I thought that, to broadly structure this conversation, we might start by discussing the latest and greatest stuff that you guys have launched in the products—the latest models, the camera motion, all that new, cool stuff—and how it fits into the broader picture and what you're seeing in terms of usage. Then I want to get into some of the more technical stuff, because you've also put out a very mathematical paper recently on an advance in pretraining for diffusion models and a really interesting position paper on how you think we should be thinking about pretraining going forward in general.
I'm excited to get into all that as well. Then maybe at the end, if there's time, we can get a little bit more speculative and talk about world models, the future of multimodality, what superintelligence looks like, and all that great big-picture stuff. We've got a lot of ground to cover, and I'm excited for it.
Stephen, awesome. Kick us off with a few reflections on your use of Ray 2 recently and what that has you thinking about as we begin today.
Yeah, thank you. Thanks for having me. Amit and Jiaming, it's an honor and a pleasure to speak to both of you, so thank you for taking the time.
I've been playing a lot with Luma Labs recently and really enjoying it. It's one of many video-generation tools that I love to use in my arsenal of possibility when I'm working on various projects. For my own work, I tend to find tremendous value in using several different models at the same time, as they all have different strengths and different weaknesses—really, just a whole blend of capabilities. Luma is right up there at the tippy top, especially at the cinematic end of the models that I like to use, so I just really want to give a shout-out to you guys there.
My first question is: I think more and more of what people want from an image-to-video scenario starts with AI images. That's just a hypothesis on my part, but is that correct?
I think a lot of workflows do start on the image side, because images are just much easier to iterate through. Iteration cycles are fast and generation times are low. You can generate 100 images and find exactly the sort of things that you're thinking about. So, yeah, I think a lot of people actually lean into image workflows as a significant part of their work today.
Okay, so that makes sense with my own workflow and especially what I see out there on social media. I think what I'm driving at is that a lot of those images, to me, seem like they can be strange or new. I'm thinking of avocado-chair-type combinations, weird characters, and all of these sorts of things. What I really want to know is: What has it been like to push these models toward a greater understanding of what I presume is a more novel subject?
That's really interesting. Currently, we are designing a feature, and that is one of the more important problems because there's no data. By necessity, these are out-of-distribution things and stuff that you would just not find in regular use cases.
The singular answer there is that you need to really work on a very strong base model with strong capabilities for understanding what is happening and then dealing with these very unrealistic, out-of-distribution scenarios. So, let's say you have a pickle sitting on an avocado chair, and you start with the pickle standing up and want it to come and sit down. You've never seen that scenario before, right? But you've seen chairs, you've seen people, and you've seen anthropomorphic things, and you've seen them doing things like sitting down, right? So this again comes down to the idea that these models are not memorizing behaviors.
These are not memorizing how this thing is done, how this is done, how this is done. They're generally distilling out base ideas and capabilities—the core things. What does it mean to be anthropomorphic? What does it mean to sit? What does it mean to be a chair? That kind of thing, right?
Then, when you combine them together, the better your model is at this foundational understanding, the better it's going to be able to do when it's presented with these entirely out-of-distribution, funny, uncharacteristic scenarios.
If I'm understanding you correctly, improving the base model is just giving you better and better performance with things that are out of distribution. But another thing I seem to notice is that, especially with image-to-video, the inherent characteristic qualities of that initial image also seem to be carried through better and more effectively with longer and longer generations. Do you think that's aligned with improvement of the base model as well, or are you guys doing more stuff behind the scenes to check against that image more regularly as the generation is occurring? What's going on behind the scenes there?
There's a lot that goes on behind the scenes in trying to understand what is in the image. These things are not just models; they are systems, and so we definitely take care of many things in the back to try to break some of these down. But generally, they are crutches for the model to be able to understand, and in the next iteration of the model, we make it so that we don't need that system.
The initial model, which was called Dream Machine and is now retroactively renamed Ray 1, had many of these crutches to make it feel and work much better—to understand motion, understand different vocabularies, translate things into language it understands, and all this kind of stuff. Ray 2 doesn't need even 10% of that. But now that people are pushing it to do more things, we have to build some of these systems again: “Okay, I see that's what someone is trying to do. How do we understand this part?”
For instance, characters: all the character understanding in Ray 2 is very much external to the model. But then, as we design Ray 3, it's all going to be internal to the model. This is a trend we have seen in language models as well. When people push the current model to doing things it was not designed to do, labs like us build systems around it to at least address those things. We gather the data and then bring it to the model.
This also applies to what we call the application layer and that kind of stuff. Application layers come up and build these specific use cases, but then we see that and are able to just build it out in the next model. The general rule of thumb in this industry—or at least on the technical side of it—is that anything that can be done in the model is going to be better than what is done externally to the model. We can talk a lot about that, but that's generally the situation.
That's great. I'm happy to double-click there as much as you want. I know we're early in, but I do think this is a super-fascinating topic. It kind of skews toward secret sauce, company proprietary information, that sort of thing. And so there isn't, frankly, a lot of conversation about these kinds of systems built behind the models that you then try to reincorporate into the model moving forward.
If there's anything else you could say, maybe more specifically about—beyond data as a loose term—making sense of these systems and how you teach a model from there to appreciate what was inherent about that system that was helping you, I think that would be fascinating to know.
That's a big conversation, and Jiaming, you should also chime in at any time. But to go a little bit deeper into inside the model versus outside the model—intra-model versus extra-model—that's a really important thing.
Whatever you can do in the latent space—all the thinking you can do about, for instance, video—is important. If you're able to reason about the actions that are being performed, the characters that are there, what that character would do, the lighting of the character, and all these kinds of things inside the model, rather than through a system that actually generates the right lighting, then tries to piece together the narrative arc, then tries to piece together all these things, that's much better.
Of course, that system can work. That's not the problem. But people tried to do that. If you remember, what was the name of the company? It was a scriptwriting company, or a copywriting company—Jasper, right? They had some really great ideas, to be honest with you. But the thing is, when you can do it inside the model, inside the latent space, you're just able to work with a lot more information. You're able to do the kind of edits that, outside the model, you just don't have the control mechanisms for the model to be able to do.
It's like this thing where there's a blood-brain barrier between a generative model and the outside world, right? And when people say, “Hey, we want controllability,” they want to penetrate the blood-brain barrier better and be able to tell the latent space exactly what they want. It will get there. There's no question about it.
Inside, there's a level of intelligence that's really, really great. Outside, there's a level of intelligence that's really, really great, which is you or me who is using it. Then there's a barrier. You want to communicate from outside in at a higher level and let the brain inside do as much as possible internally, collating information.
Think about sequencing of events in the video. Think about what happens to a character: their arc, their timeline, their causality—all these kinds of things. The more you can do inside the model, the richer the outputs you're going to get.
This is what chain-of-thought also looks like in language models, right? Instead of forcing the model to structure its thinking into, “Do this, do this,” no, chain of thought is—if you've read the recent Anthropic paper, where they claim chain of thought is actually fictitious—just a way to seed the latent space of the model, to direct it. Whatever it's outputting is not a realistic representation of what it's actually thinking about, right?
That is actually an indication of what we've seen now. Instead of trying to force the model to output JSON through external coercion, let's just make the model very good at structured output, and then suddenly you have agents, right?
You're going to see that in multimodality especially, because if you're able to simultaneously think in audio, video, language, and images all together, and combine reasoning from language, appearance from video, all the aesthetics you've seen in images, and the audio that you're able to hear in all the videos, you're going to have much better outputs than through external systems. But yes, we do build these systems, and then we obviate them in the next model training.
I tend to agree with Amit here. Instead of trying to program this into complicated workflows, just imagine how you would communicate with another human. You don't actually need to program their brains for them to understand that you want to achieve this level of workflow. Of course, there might be some back and forth, but the process itself is pretty much natural.
I think this can only be done if you build these capabilities into the model, in the sense that the models need to be more intelligent. Whereas, right now, even the current generation of models—especially video models—feels less intelligent in the sense that you have to tell them, or program them, to achieve some particular type of task, like image-to-video, for instance.
Of course, there are practical reasons that this is done, but I think eventually they'll be intelligent enough to take in whatever task, in the multimodal sense, you tell them. For example, “I want to do image-to-video, keyframes, camera motion,” and all these different tasks. The model should be able to do it without even thinking about it as a different type of task.
Could you guys maybe give a couple of examples of things that presumably wouldn't be too sensitive to say—things that were essentially scaffolding in the Ray 1 generation that are now handled internally by the model—and then, if you're willing and open to it, things that are currently scaffolding on Ray 2 that you hope to be built into the model when we get to Ray 3?
I guess one example would be camera motion. In Ray 1.6, we had camera-motion features, but that was obviously harder to actually control and lower quality than what we recently released.
In this iteration, you can see that there are a lot more different types of flexible camera control, such as Bolt Cam movements, without the user having to prompt for exactly the extreme X and the extreme Y of the camera. It just works. Of course, you can also give it more explicit controls; that's what some users would want.
But I guess for other users, maybe what they really want is just, “Convert this scene—give me this effect in the scene,” without having to specify the very detailed controls. I think that's one of the features that we have in the model.
I think what would be really interesting is to build into the model—now it's one level ahead of being able to represent the scene, like a boom camera. Again, you want slightly more control than just generating a boom camera, but not as much control as having to specify the exact camera coordinates.
So maybe somewhere in between, you can have something like, “I’ve given it a camera trajectory scene. I want this Bolt Cam to be moving faster, moving slower, having more angle changes, or focusing on some other subject matter.” So this kind of more interactive editability of the scene is something that we are thinking about maybe having in the model in the next generation.
Cool. That was actually my next question. So maybe just to restate a little bit here for our audience who might not be super familiar with Bolt Cams: that’s bringing us to Apple TV+’s Severance season 2, right? It made a huge splash this year with its opening shot, which is a Bolt Cam robot-arm shot. I think it was super impressive. It famously took them months and months to create and was just a big wow moment for audiences.
And now, just a little bit later this year, we have it as one of your key camera motion concepts already available for people like me to use when generating. It’s one of a range of motions that you have added to the model capability. Super useful, and I imagine very highly in demand from your editor-type user approaching the model, as well as novices and everybody else, but it’s got to be high on the professional list. What has it been like to develop that feature specifically with an outlook on the professional user, or maybe customer requests coupled with modern trends like Severance season 2 and the Bolt Cam?
So basically, there are models, right? The things you teach them during pretraining, and then there are many capabilities that people want to teach them. Language models have this really special ability, which is called in-context learning: you give or show some examples, and it becomes that. And I think that’s a truly emergent intelligent capability. Visual models aren’t there yet. They’ll get there soon enough, but they’re not there just yet. But that doesn’t remove the need for teaching these models specific things you want in the moment.
So we have been working on this idea. We published a post about this, a white paper, whatever you want to call it. We call them concepts, and we realized that in visual, especially creative use cases, there are many things people want to teach—many things that they come up with, like a specific motion or particular kind of color grading or a particular human pose, which is just really suited for the story you’re trying to tell or the ad you’re trying to make, whatever it is.
But there’s no way to get it out of the model because it’s a new thing you just came up with. There’s no data for it on the internet, and there’s no way to generate large samples of it to even be able to fine-tune the model. So we designed this idea of concepts to make it so that models can learn from 1 to just a few examples. Their capabilities don’t degrade like they do when you’re fine-tuning or creating LoRAs, and they can be composed together.
How closely can it represent a capability that the model had at pretraining but users actually teach it? We call them concepts. The tool that we’re using to build them is called Concept School, right? Like, you go to school, you take the model to school, and you’re the teacher, and you’re going to teach it concepts or lessons, whatever you want to call it.
So the camera motions you’re talking about—the Bolt Cam, the dolly, and the reverse dolly, all these kinds of things that we taught—we were able to teach them to the model very quickly, relatively speaking, with very few examples. And you’re going to see the next batch come out this week, and the next batch come out next week, so on and so forth. Our goal is to basically teach our models everything about filmmaking this way and eventually also give people the ability to teach.
Right now, it’s a little bit finicky, as new technologies tend to be, so we haven’t made it open source—or, sorry, open access—yet, but we will in the future. So these camera things, right? The way we are building them, to answer your question on feedback and these kinds of things, there are some of these which are basic storytelling tools: being able to track a shot, being able to move left, being able to move right. These are done by people a thousand times a day whenever you’re shooting something, like, “Oh, the camera moves left or right,” or you’re tracking an actor who’s doing these kinds of things.
So you want that, and then you want to also balance it with some things which are absurd and funny, or which people just can’t do in real life very easily. Like, for instance, Bolt Cam. If you want to do Bolt Cam really well with that smooth tracking, you actually need a robot.
Yeah. MKBHD has one. I’m sure James Cameron has a few, but most of us don’t. So can we just have it in the model? We try to balance this great utility with some things which are just absurd and fun.
But ultimately, we’re building this for professionals. We’re building this for people who want to tell stories, who are already telling stories, or who want to become professionals in that world. And we can talk about the changes that are happening in the AI industry very quickly—not the AI industry, but the moviemaking industry very quickly. But, yeah, we’re designing for people who want to tell stories, and these are all storytelling tools.
So, 2 things there that just brings to mind for me. One of them, selfishly, is talking about those repeat actions that the pro user takes. One of the—selfishly, one of the pro actions I take all the time is leveraging your audio generation feature, which is great. For people who don’t know, after you generate a video, you can just press the audio button, and it will give you another prompt opportunity and give you the ability to generate audio for that clip.
However, one thing I do all the time is reverse the playback direction of my video when I’m editing. So, just a tiny little plug here: I would love to have the ability to quickly flip the playback direction before generating that audio so that I’m not dealing with reverse audio when I take that clip out. But I’ll just leave that as a footnote.
I think really what this is getting to is you’re about storytelling. You have all kinds of users telling all kinds of stories, and that skews towards fantasy, Hollywood cinema, anime—that’s everything, right? Yeah. How do you imagine that translating to multimodal understanding? Is it just an attempt to understand everything everywhere all at once, or is there a particularly unique insight that you feel like you gain in the pursuit of art first that helps with the overall mission or multimodality?
See, what we are trying to do is—the mission of Luma is to build multimodal intelligence, right? Multimodal general intelligence. If you were to not personify it, but if you were to materialize it in front of you and ask, “What does that look like?” the intelligence that LLMs embody looks very much like abstract intelligence that humans have with language and things like that.
When you think about multimodal intelligence, of course it has that abstract part, but what does that actually look like? It starts to look very much like a world, a globe. You have this world in front of you, and it has all the physical properties and all the physical phenomena that happen day to day in our physical world. It also has those intelligent beings inside it that do things that increase the entropy of the universe, that interact with each other, that do all these kinds of things.
So a term a lot of people use is world simulator, right? But I think simulation is a weaker term here than it should be. It’s basically just a world model. It’s like a physical manifestation of the universe that you have outside.
Of course, it's a weak facsimile. It's an approximation, but yeah, it's a model of the world and of the processes that we are having.
Now, coming to the creative side of it and why that is important for this mission: storytelling. If you think about it, with language models, we were like, “Oh, this is only good for JSON, and then we're going to produce just JSON.” That's not very good. When you force intelligent systems to play games, to come up with abstract new things, or to follow instructions—“Oh, no, no, I want this, then I want this”—that's when you get intelligent systems rather than just procedural systems that are following a set of rules that you have created.
When you force them to deviate from that by making movies and telling stories, that's a very critical part of it. It's a part of human existence too, right? How good someone is at storytelling is generally a very good barometer of IQ, right? Can they actually think beyond just the most physical thing that is in front of them? What is the consequence of A? What is the consequence of that? What is the consequence of that? What does it lead to?
That's what stories are, right? Event A happens, then B happens, then C happens, and then D happens. Video models are on the critical path to that general intelligence. Once we're able to combine video, audio, and language all together, they will become really, really good at storytelling—being a partner for us who can actually sit next to us and help us dream, help us imagine, and help us think through those kinds of things in creative pursuits.
It's not a surprise, by the way, that people thought the first things to be automated would be the mechanical things. It's not a surprise that the first things where AI is able to help are actually creative pursuits, because that's where we require intelligence. That's where these things come into play. So, yeah, I think art has always had a very significant role to play in general thinking and general intelligence. Art is also going to have a very significant role to play in building artificial general intelligence.
Can I ask a couple of world-model questions?
Yeah, I think this is super interesting. One big thing that jumps out at me, though, is that if I were trying to create a generally useful AGI to go out and do stuff in the world—to maybe control robots and have one walk into my house and make me a coffee, to take a famous example, and vacuum up my living room after my kids have thrown toys all over it, whatever the case may be—I would want this sort of world model.
I'm certainly a big—I wouldn't say “believer” at this point is the right term, because I think it's pretty well demonstrated that the large foundation models are learning these higher-order concepts, and that has, I think, become pretty much indisputable at this point. So I'm well convinced by the evidence in the literature that this is happening, but I wonder about the sort of fictional side of it, the magical side. There are no dragons in the real world, but your models have learned to also represent dragons, right?
So, in a way, you have a real-world model plus something that sort of goes beyond what is real and into the imaginative, the fictional, et cetera. I wonder, obviously, that's good for storytelling because we want to tell stories that are not bounded by reality, but does that have downsides for the practical utility of making multimodal intelligence that I just want to come into the world and do useful work for me?
Yeah. I think there's a short-term answer and there's the answer for the long term. I guess the answer for the short term is that if you want to use this type of model for your daily activity tasks right now, then yes, it might be better to angle toward the more physical, applicable verticals, so to speak, for this type of model.
However, in your model, there are many cases where you ask the model to do some task. For example, say, “Pick up these clothes with a dragon on them and put them in the washing machine.” If the model doesn't really know about the concept of a dragon, you can't actually pick up the clothes with the dragon pattern and put them into the washing machine.
So even this kind of knowledge that is fictional can still be very useful for real-world physical tasks, because, again, even the concept of a dragon by itself is related to other physical concepts that people see in the world. For example, it looks like a lizard, it has scales that are exhibited in reptiles, and it has wings, which you see in bats and birds.
So even these concepts, if you think about it, are human interpretations of what they see in the real world. The dragon is generated by humans looking at the real world as part of their training data, and they generate something that is out of the ordinary.
But back to the original question: yes, for the short term, I do think maybe focusing on more physical, realistic things can be better for these kinds of tasks, so the model doesn't hallucinate as much. But in the long term, I believe that the model should be able to tell for itself what kind of world it is acting in, because there are also other types of worlds it needs to be able to act in, and it's not just a physical world.
For example, in a virtual world, the things people are currently doing with agents are also a perfect example. The model doesn't need world knowledge, but it's a world that is different from the physical world that we are acting in. So having a unified AI that can do both tasks could be useful.
For example, eventually your robot could be like, “Okay, put this dragon T-shirt into my dishwasher, and then go to the computer and type an email responding to my friend,” or something like that. This requires the AI to be able to reason between physical and nonphysical, imaginary, or virtual worlds, just like humans can do. So I think eventually this will be a capability that does not have to be specialized into any sort of model.
I don't know if you're doing interpretability work on your models internally, or if you have any partnerships, academic or otherwise, that allow you to do that kind of stuff. Maybe you know the answer because you may have done this work, or maybe you could speculate as to whether you would expect to find features in the model that would represent, “This is a realistic physics simulation,” versus, “This is a magical-realism-type scenario,” versus, “This is a Minecraft environment that we're in right now,” or we're playing Pokémon, or whatever the case may be.
It seems like that probably would happen at sufficient scale, but I don't know if we are there yet or if anybody has really had the opportunity to look. Up front, I'm not an expert on deep neural network interpretability, and I haven't delved too deeply into this space. My very shallow understanding is that it might be good to seek more alignment, or try to get the model to align with what you want, and reason from there.
For example, one feature that we have discussed heavily in the first world model is this implicit knowledge about 3D. You can ask it to generate things that reason about depth, and even in fantasy or unrealistic art, it has a reasonable interpretation of depth and other physical features, such as cloth simulation, waves, and hair and that kind of stuff. So the question you're asking is whether there is a neuron or substructure within the model that shows that. The honest answer is that it would be very difficult to find that within the model, but it might be easier to prompt the model and ask it to do the task for us.
For example, people have shown that these models are very good priors for things like 3D vision tasks, such as depth estimation. A lot of the state-of-the-art depth-estimation methods are actually based on these pretrained models, which have a very good understanding of the world. So I do believe that, to some degree, the model has great internal knowledge about the world, but it is also up to humans to determine how interpretability exists. Currently, it seems like the API between human neurons and model neurons is not very well-defined just yet. We will probably still have to use language that both sides understand to communicate this interpretability.
Yeah, there are definitely some broad challenges there. Yeah, go ahead.
This is an interesting philosophical question that Nathan and I go back and forth on all the time. We tend to hear from people a lot that these models have physics knowledge inherent in the training. They have a great, rich 3D understanding of the world; we see it in this or that. I personally am not entirely convinced that they aren't just seeing more and more movement across a 2D space and developing greater and greater pixel understanding.
I can understand somebody arguing against that from, perhaps, a systems-based training regime where we're training first on 3D scaffolding or something like that. But my own naive understanding is that sort of thing isn't happening in the training. I just think it's a really interesting question: Are we actually seeing physical appreciation for a 3D space, or are we just seeing more of an expert interpretation of a 3D pixel space?
I would relate this ability to predict or reason about physical scenes in a representation of 2D pixel space to a Turing test. In the Turing test, your goal is not to determine whether the model knows grammar, or whether it internally represents the scene using the grammar that humans define. That's not really part of the question we're asking. What we're asking is whether, when talking to the model, it feels the same as talking to a human. In this case, it's similar.
On the one hand, being able to render a prediction of the next frame of a scene perfectly over the next few seconds or minute doesn't actually mean that, inside, it has the same physical knowledge that we, as humans, currently define. But on the other hand, from a utility standpoint, it might constitute passing the Turing test for visual generation.
Yeah, and I haven't had a lot of time to think about this problem. The question is basically philosophical, in line with what Jiaming is saying: It borders on the definition of understanding, right? What does it mean to understand? People make the argument that humans have something very special or deep where we understand at some level. But it's very hard to argue about that as well, because how would you know?
You can say, “Oh, yeah, but I know this is 3D. That's why I understand it. It clicks in my brain that this is 3D and not 2D.” But the brain is this really interesting organ that is simultaneously thinking and telling you—or itself—that it is thinking, and that's so weird, right? Think about it: It's self-aware. The self-awareness part is the most interesting part of it.
Coming back to 3D understanding for just a second, what does it mean for humans to have 3D understanding? I'll take a slightly different stance than most computer vision researchers, which is, “Oh, humans actually perceive 3D,” and things like that. We don't really perceive 3D. Yes, we have 2 eyes and stereo vision, but stereo dies out at about 20 cm. Outside of that, the disparity is almost nothing, right?
Case in point, if you hurt 1 eye and have an eye patch, you can still drive. You have some initial trouble trying to grab a piece of glass or something close up, but you can still drive. Of course, you have hands and proprioception. The brain gets other signals that teach it depth and some of these things, but it's hard to really argue that the brain actually maintains some sort of 3D representation. It might not. We just don't know. We have no way of understanding it.
So, for a generative model, why is it any more special that it has an explicit 3D representation inside it, like a mesh or whatever have you, than just understanding these concepts more implicitly or in terms of 2D space and time? As long as it's able to generate something that looks really consistent, why do you care how it actually did it? Planes fly, but they don't flap their wings. The phenomenon that birds use and the phenomenon that planes use are actually very similar, right? Birds generate lift by moving their wings; planes do that by pushing themselves through the air.
So while the phenomenon is the same, the mechanisms are very different, right? Humanoids right now—we are trying to build them. They move, but their movements seem very different from how we do it. It's going to happen more and more with all machines that do things humans do. For dishwashing, we wash dishes very differently from a dishwasher, but are the dishes any less washed because the dishwasher doesn't have hands? I don't think so.
Philosophically, coming back to it, I don't think machines have any prerogative to think exactly as we do or to have representations inside them like we do. It doesn't make them any less intelligent. It doesn't make them any less capable. In fact, they should do it differently because the substrate they're on is very different. The human brain is a 20-watt piece of organic tissue. On the other hand, here we are running things on gigawatt-scale clusters. Why should they think the same way?
So that's my answer to that problem, actually. I think people who are focused on making machines think exactly as humans do, or making them exactly as interpretable as humans are, have misguided goals. They are not only reducing the capabilities of machines in that process, but also wasting their own time. They should use that time to scale attention better. They should use that time to design regimes that can do more efficient learning, all these kinds of things, right? I think it's a waste of time.
All right. I agreed with the first 85% of that, but the last 15% I want to challenge.
Go for it.
When it comes to interpretability—why would you care?—the 85% does include the idea that I don't expect AIs to be representing things in the same way that we are, or thinking broadly in the same way that we are. I don't think that should be used to discount them. I have a funny series of tweets where it's like, “It's only a concept if it comes from the concept region of the human brain. Otherwise, it's just sparkling notions,” or whatever. So I'm with you in terms of viewing that as a straw man.
But I do think when you look at something like Golden Gate Claude, for example, they were able to say, “Okay, we were able to go in and isolate what very strongly appears to be the Golden Gate Bridge concept.” Now, when we artificially turn that up, we get Golden Gate Claude. Yes, that's just a curiosity and a cute demo, but I actually thought when you were talking about the concepts feature that you were building, maybe you were doing it that way.
I could imagine that you might say, if you had taken—and I'd be interested to hear how you are doing it, because it seems like it's not this—but what I had imagined you might be doing there is taking a similar approach, trying to identify directions in the latent space and then injecting them or turning them up in order to enable these different concepts. It seems like you probably could do that and be like, “Turn on anime,” or “Turn on Toy Story style,” or whatever. Stephen has a much better vocabulary for the different styles than I do.
But I assumed that your concept work would be similar to Golden Gate Claude. I guess the questions there would be: Doesn't that seem like it would be a very useful thing to study and potentially be able to marshal? And if you're not doing that, maybe can you tell us how you are doing it? There's a difference between treating something as a black box and being able to do empirical experiments on it.
But we are in the philosophical land again. Think about physics, right? A lot of people think our knowledge of the world is completely interpretable, and because we have an equation, we understand how the systems work. That's not how any of the laws of physics actually are, right? Science is not the act of interviewing God, if that existed—or nothing one way or the other. Science is basically: we have an observation. Can we derive a pattern out of it? That's about it. Science is empirical. Theoretical physics is also, again, very much about building a mathematical model, an approximation, a representation.
Machine learning is very much the same. The entire universe is very statistical, right? The best theory that we have of how the universe functions right now, quantum mechanics, is entirely uninterpretable. Today, we don't understand why things are the way they are at all. The measurements get verified up to 6 sigma, 8 sigma, sometimes 10 sigma—as accurate as it gets. There are some measurements that are up to 23 sigma, right? But we don't understand why the world is that way, right? That's what you say when you're asking for interpretation, right?
I'm saying ML models, because of their sheer scale, are just not grokkable by the human brain. You're not going to come up with this coherent, consistent model of how the ML model functions, because it's just not that kind of a system. We can do empirical experiments on it after the fact, like the Golden Gate experiment that you're talking about: we found a cluster of this; we found this. It's like archaeology. We can do archaeology, all right.
But here's something really powerful that we can't do in the physical world. In the physical world, we can only do archaeology. Here, we are actually involved in the formation of the planet and all the eons it went through. That's called training. So we know what the model is going to do.
Instead of people spending their time on archaeology, I think a much better use of time is understanding data and training processes. You actually learn so much during training. At the risk of Jiaming giving something away here, when we train our models, we are working on actually designing these models with a great degree of efficiency. In that process, we do so much information-theoretic work on our models to try to understand what sort of information-flow architecture is in there: what is happening to high-frequency details, what is happening to low-frequency details, what is it learning in the earlier stages and later stages, and when should we actually change the curriculum of what it is actually learning?
This is all basically what people talk about as interpretability. This actually has real consequences. This changes what the model learns, and some of the things we have learned are really interesting. The information-theoretic pathways are set up in the first 20,000–30,000 iterations. The model starts out as this scalar field.
Really, what is a good example? Think about a field of corn that's grown. It's just plants everywhere; they look uniform, all these kinds of things. As you train, pathways keep forming in between them, like someone walking and putting the plants down, right? And then there are crop circles. When you look from the top, you're like, "Oh, there's a pattern. There's a circle," and things like this.
Somewhere, these happen very early, and the distribution of data you have at that time and the kind of learning-rate regimes you've used decide a lot of what is going to happen in the later stages. But in the later stages, a lot of different things happen, right? These pathways are emphasized or deemphasized. Someone goes and undoes the crop circles; someone decides we're not going to take that path as much. This is the kind of interpretability that allows you to design the models that give you the output you want.
This post-facto interpretability work, I think, is pretty interesting. No, don't get me wrong: intellectually, this is empirical science, right? It's very good, and if you have an alien model in front of you, or some model someone else made, it's very useful to understand what's going on. But if you want to actually control the models, you want to control the data and the training process—and I mean the actual hyperparameters of the training process. There's so much that goes into that. So that's where actual interpretability comes into play, in my view, and that's where we do a lot of work.
Yeah, that's super interesting. So you could expand on any number of dimensions of that. One that I've been thinking about quite a bit recently is what you might call batch strategy. It sounds like you're probably getting pretty intentional in those early phases of training about making sure that you're feeding, batch by batch, the right mix of data.
Because I can imagine if you had—I've always looked at these loss curves, and you see these occasional spikes, and you're like, "What's happening there?" My one interpretation is maybe that's a bad batch, some sort of cluster of data that the model wasn't prepared for. But then also, if my interpretation is right, that batch is probably sending a bad signal to the model as well, which is not actually constructive for general learning purposes.
So are you actually doing this sort of batch-by-batch construction of data to try to get the mix right at that very granular level, or am I misinterpreting what you're saying?
I think Jiaming can answer this a bit better, but I can clarify my own statement. I don't mean handpicking it that much. I think we come at it from the other direction, which is a lot of curation and filtering and removing garbage examples. The technique to training really great models is not to throw 1 billion samples of garbage data at it and have it figure it out from the garbage, right? The technique is to show it what you want it to learn and show it good examples from there, at large scale.
So, yeah, you do a lot of this preprocessing, filtering, that kind of stuff. But Jiaming, I don't know if you have a different take on that.
Yeah, I think it's a very interesting question. Basically, what is the right data set? Interestingly enough, the kind of statistical methods that we're using for training these machine-learning algorithms are actually quite different from what the real world actually is. So if you run this algorithm, whether it is on language models or a diffusion model, you are mostly making i.i.d. assumptions.
Basically, you have a data point and assume that these were drawn independently and identically from the world distribution. But the real question is: is there a world distribution? And how do you even define the world distribution? That's a real question.
Regarding the question of what is the best data set, it's actually very hard to have the right answer. I think people probably take a less statistically driven, but more objective-driven, approach. It's like, how do I control the data set such that I get a better quality or better outcome?
This has also been deeply studied in the realm of language modeling, for example: how much English do you want to have in your entire corpus? To be honest, it is very hard to reason about these on a theoretical level, because you wouldn't want to say, "We just base how much English to put into the data set on the number of people who speak English." So that's probably not the right way of doing it.
Instead, having a more empirically focused approach is probably the right way of finding it. But the exact solution is probably very messy, and there's not really a lot of key, concise principles behind it, despite how easy it is to explain the idea of machine learning to the general public.
Jiaming, would it be too much to ask for you to give us a short master class in the intellectual history of diffusion models? I think you're probably the best person in the world, maybe, to do that.
Yeah, yeah, of course.
So I sketched it out, and the mantra that I always say to myself about the original diffusion models is: there are all these images out there. It's really easy to just programmatically add noise to them step by step until you get to pure noise. Then the insight is, if you just train the model to do the reverse and learn a denoising step, then if you run a bunch of those steps in a row, you can start from pure noise and eventually get to some image.
The original versions of that were totally undirected and just generated an image out of nowhere. Then we got classifier guidance, and that sort of was enough to steer us in the direction of particular images. Then I want you to take over and tell us the major updates in the field that have brought us to where we are in terms of efficiency and control.
Yeah, of course. So I guess the first known method of diffusion models was actually reported in this 2015 NIPS paper from Jascha Sohl-Dickstein. He described the actual algorithm for how to train, or how to formulate, the forward and reverse process of a diffusion model.
However, that never really caught on, because back then the experiments were done on very small MNIST data sets, and it still had the same problems that plagued diffusion models very early on back then. To researchers in that field, it seemed like, "Okay, I have GANs that work in one step, and they also work much better at large scale. Why am I bothering myself with this method that seems very hard to grasp and very slow to generate samples?"
So the first real breakthrough in this field came in 2020 with a paper called “Denoising Diffusion Probabilistic Models” by Jonathan Ho et al. The basic idea is that, first of all, it validated the validity of this idea and made it as performant as GANs on certain tasks. It was still very inefficient, in the sense that you needed maybe 1,000 steps to converge to the right image, but it was one of the first algorithms that did not have the unstable-training problem that GANs give you.
For practitioners like us, having an algorithm that trains stably is very important. There is a huge difference between launching a job, going to sleep at night, and waiting until the next few hours for the result to get better, versus the GAN case, where things just become unstable out of nowhere and you have to do a lot of digging. The training algorithm being stable is a very big deal here. This is actually why people got interested in diffusion models, even though they were very slow at the time.
I was part of the people who worked on these general directions. I had worked on other types of models, including GANs, before, so to me, diffusion models seemed like a very fresh perspective. Before the existence of that paper, there really wasn’t any method that could train stably and generate high-quality samples.
Then I was trying to think more about this: Now we have a model that generates high-quality samples, and at the same time, this model is very stable to train. What is the problem with the model? The problem is that it generates very slowly. It takes about 1,000 neural-network iterations to get to a high-quality sample.
So I was looking more into solving that particular angle of the problem, and this is how I got to my work on Denoising Diffusion Implicit Models, or DDIM, which was able to accelerate the model’s sampling time by 20 to 50× back in the day. Of course, simultaneously, Yang Song, who was also from our lab at Stanford, was working on the more generalized idea of diffusion models.
His idea was that, instead of having a fixed, discrete number of time steps, we could make the time step continuous and wrap it more into the math that came from the 1980s, from a person called Anderson. It was all within this stochastic differential equation framework, along with score matching and denoising score matching. Yang was a very early advocate of score matching and denoising score matching, so it naturally fit into his framework.
Then he found this very interesting connection between diffusion models and denoising score matching. Those were the first 2 breakthroughs: one on the theoretical level and the other on the practical level. Then OpenAI did the work on training these models on ImageNet, which was again a big breakthrough at the time. It showed that diffusion models were competitive with GANs in these cases, as well as introducing classifier guidance, which people are still using today.
Then people started to use classifier-free guidance. They realized that we didn’t actually need to train an additional classifier. You could add, say, an unconditional signal to the model such that you could replace the role of a classifier. This also came from Jonathan Ho. After that, there were a few papers around 2022 that tried to scale diffusion models beyond ImageNet.
In early 2022, there was a paper called GLIDE from OpenAI, which basically tried diffusion models for text-to-image generation. A few months afterward, the Imagen paper from Google came out, and a few months after that, Stable Diffusion was released. That is basically how we got to Stable Diffusion.
Of course, after Stable Diffusion, there was an explosion of these techniques and, similarly, open-source models. The next things people cared about were how to get higher-quality models, how to make them work on videos, and how to make them more efficient. I’ll talk more about the efficiency side of things.
There have been a lot of efforts, initially and still today, to do distillation. Previously, the way people did distillation was kind of hacky. The idea was that I had this 1,000-step process, or maybe an even longer process, and I tried to use a model to represent 2 contiguous steps so that I could reduce the number of steps by 2. Then I tried to train another model to simulate those 2 time steps, so that overall I had a 4× acceleration, and repeated this again.
This is called progressive distillation. It also came from Jonathan Ho. But it is pretty tricky to train because the implementation gets a bit hairy.
In early 2023, there was another paper called “Consistency Models,” which at the time aimed to be a replacement for diffusion models, in the sense that it was both efficient and could be trained in a single stage. It also described a method called consistency distillation, which is a distillation-based method for consistency models. Basically, you use an existing diffusion model as the base and try to distill it using consistency-model ideas.
What actually became more popular in the field was consistency distillation, because when people tried training consistency models for other use cases, it turned out not to be as easy as it seemed to train them very stably. People came up with new methods to stabilize the training process for consistency models, which usually still involved initializing the model from a diffusion model, so to speak.
Then came these types of distillation techniques, and there is another set of techniques that are based more on GANs, so to speak. The difference between consistency-distillation techniques and GAN-based methods is that the consistency-distillation techniques are, in some sense, more stable to train, while the GAN-based methods are less stable to train but may have higher quality if you run them for much fewer steps.
Again, this is the interesting trade-off between how stable it is to train the model and how easy it is to get high-quality samples. That takes us back to the original story between GANs and diffusion models. That is basically the current status quo on diffusion models: the history of diffusion models, how people have scaled them up, and the key problems in making diffusion models even faster through this path of distillation.
Could we do just a little bit more on classifier-free guidance, and maybe also on the intuition behind the consistency model? When we’re doing guidance, what exactly are we doing to guide the model? I get it in the classifier sense, or at least I have an intuition where I would say, “Okay, feed it to the classifier and take its feedback.” But when we’re classifier-free, I think it’s a little less intuitive. The consistency-model concept is also a little less intuitive than simple distillation.
Sure. Before we talk about classifier-free guidance, we can talk a little bit more about classifier-based guidance. The idea is that diffusion models try to represent a score, or the gradient of the log probability, of this distribution called P(X). You try to guide it with some conditional signal. Let’s just call it Y.
Instead of trying to sample from P(X), you want to sample from P(X given Y). You can train on this, of course, but another way to treat P(X given Y) is to apply Bayes’ rule. Basically, P(X given Y) is equal to the joint P(X and Y) divided by P(Y). You can think about P(X) as the score of the diffusion model without any condition, and the P(X and Y) part as the score of the model with the condition, because the other term, P(Y), which is the condition part, is something you don’t actually care about during sampling, so you can just drop it.
Basically, that is why in classifier-free guidance we have an unconditional model, which represents the denominator, and a conditional model, which represents the numerator. All in all, it is basically an application of Bayes’ rule.
A consistency model is basically something like this. Maybe it’s easier to explain consistency distillation first. The idea is that I want to have a 1-step model to generate the right solution, but I want to find a way to bootstrap it from a regular diffusion model.
Suppose I have a perfect 1-step model and it follows the trajectory of the diffusion model. The idea is that we want to learn a model to distill the process that a diffusion model would normally go through with many steps. We are just distilling this function.
What a consistency model would do, in principle, in the distillation case is that you can compute the 1-step prediction at different time steps. Time step is a concept that is correlated with how much noise you add. The more noise you add, the higher the time step in this particular case.
In a consistency model, the idea is to build a connection between 2 quantities. The first quantity is: At a given time step, what prediction are you going to make in 1 step? The second quantity is: Suppose you are using 2 steps, and 1 step is a regular diffusion-model step. You run 1 regular diffusion-model step to a time step that is closer to your original state, and then you run this same consistency model to reach the final step.
So basically, the way consistency distillation is trained is that it tries to minimize this loss function. Because your consistency model at the earlier time—the time step closer to the clean signal—has an easier time predicting what the real signal is, it allows the model to build this connection and train this function. But essentially, what a consistency model tries to do is use a model to distill the otherwise hard-to-compute process that is run by diffusion models.
There are so many directions we could go here. Let’s do your latest contribution, Inductive Moment Matching, which I basically take as being the best of all of these prior approaches combined into one. There’s sort of an echo of the consistency model idea and an echo of distillation. What jumped out to me most about it was the idea that the model is being optimized in distribution space, and at more of a—I mean, I guess it’s always batch-level—but at a higher level than just example by example.
Like I mentioned, there are 3 things we want to achieve in generative modeling algorithms. One is high sample quality, two is stability to train, and three is that it is relatively efficient when you are trying to draw samples from it. Most of the—not most, but all—of the existing methods suffer from 1 of the 2 drawbacks.
For example, generative adversarial networks, or GANs, are not that stable to train. The same kind of thing goes, to some degree, for consistency models as well. Diffusion models are stable to train and have high sample quality, but they can’t generate high-quality samples in very few steps, so the inference cost is high.
What we want is to find an algorithm that satisfies all 3 of them: high-quality generation, fast sampling, and a stable training process. In this case, we try to reason about a generalization of what consistency models are doing. Instead of trying to match pointwise samples—basically, I have this function and I want to exactly match the samples—what we can do is match only the distributions, because we don’t actually care that this function has to match exactly.
For example, if you have a bunch of samples and you are trying to push them to another set of samples, there are many different solutions you can use. But in consistency models, you are forced to follow 1 type of solution. That is a little bit more restrictive for the model, and maybe your model needs more capacity to achieve what it is being asked to do.
This is possibly 1 of the reasons why these GAN-based methods have an advantage, because GAN-based methods are actually comparing samples at a distribution level. So we started thinking: instead of trying to match the samples pointwise, we can just match the samples at a distribution level. What is another algorithm that can match distributions based on samples that is not a GAN?
It turns out that this idea has been discussed in the statistics community at least 15 years ago, and this idea is called maximum mean discrepancy. The idea can sound a bit scary, but what you can think of is that it has a very interesting relationship with GANs.
In GANs, you have the discriminator trying to maximize the distance between the prediction of a real sample and the prediction of a generated sample. In maximum mean discrepancy, you do the same thing, except that the discriminator is no longer a neural network. It is a function defined on a space called an RKHS, or reproducing kernel Hilbert space.
You can think about it as a feature representation of the function—a simple function based on complicated, infinite-dimensional features. It is a more closed-form function. What is interesting about MMD, or maximum mean discrepancy, or this RKHS choice in general, is that you don’t actually have to optimize the discriminator.
Once you define a particular type of space to optimize for, the optimal solution can already be represented. That means you skip the inner loop of optimizing the discriminator, which makes this whole process less unstable. You end up with a very stable optimization process that still minimizes the distance between distributions as you are trying to learn with generative models.
Of course, the downside is that you chose this space of functions a priori, so you don’t have the space to optimize it for the best-case scenario. But in our experiments, we didn’t find this to be a huge problem.
Essentially, the easier way to interpret how our approach to Inductive Moment Matching works versus what consistency models are doing is that it is a generalization of consistency models, in the sense that it does distribution-level matching. In consistency models, the distribution matching is based on a single point, and of course you can’t easily represent a distribution with a single point, so it becomes a more degenerate case.
This kind of explains why, in certain cases, consistency models are unstable to train. We actually did ablation studies controlling the number of samples we use to compare distributions. With 1 sample, the consistency model is unstable. With 2 samples, it is also a bit unstable, but it becomes unstable later. With 4 samples or more, the training process actually becomes more stable.
That’s how we get into this 1-stage process that has high-quality generation.
Maybe just 1 last question, because I know we’re at time. Looking forward, you guys are obviously deeply invested in multimodality. I’d love to understand how you think about multimodality broadly and where you think it’s going.
We probably don’t have time to explore this answer, to be honest with you. This is the entire foundation of the company. But I would say this: currently, what people think of as multimodality is not really it.
These are language models that have been fine-tuned to work with images, audio, and video, and they show some capabilities that are really beneficial in the context of whatever the language model is doing. But as you’re seeing individual capabilities—like understanding objects, object recognition, segmentation, or just being able to understand what’s going on, following long, long threads of things people are doing, causality, and all these kinds of things—this understanding in multimodal models is still worse than in dedicated computer vision models.
That tells you that this might not be the approach, right? We’re not quite there. It would be like if we built all these language models, but they were still worse than RNNs at interpreting language or natural language, right? But that’s not the case. Language models are fantastic at that.
What we’re seeing is that this is a very promising direction, and obviously the bag-of-tokens approach is extremely helpful, but we need to look beyond just language backbones and trying to retrofit everything to them.
Luma’s approach is very different. We are coming at it from the direction of a unified, singular latent space, where we can think and reason about all these different pieces of information as if they were one. Technically, they are one, right? Nature doesn’t make a distinction between video, image, audio, and things like that. These are just signals. They just happen to all be part of the same simulation, and they’re all part of the same environment that we are in.
An action produces these signals in all of these different modalities, exposing different facets of its existence and occurrence. We need to think about it in that same way.
Currently, we see that most of the industry is extremely shortsighted when it comes to thinking about multimodality. For good reason, by the way, right? There is so much to be done in language, and people should continue to do that work. But there’s a new approach that is necessary to actually solve multimodality, and that’s what we are pursuing.
Okay, cool. I’m looking forward to part 2 already.