[BidClub_]
The Cognitive Revolution · · 77 min

Luma Labs' Diffusion Revolution: from Dream Machine to Multimodal Worldsim - Amit Jain, Jiaming Song

Nathan LabenzAmit JainJiaming SongStephen Parker

YouTube
TL;DR
  • Luma’s product-to-model loop is a recurring strategy: scaffold a demanded capability around today’s model, validate customer demand, gather data, then internalize it in the next generation. Ray 1 needed extensive external support; Ray 2 does not need “even 10% of that,” while character understanding remains external in Ray 2 and is planned inside Ray 3. Amit Jain’s rule is categorical: “Anything that can be done in the model is going to be better than what is done external to the model.”
  • Out-of-distribution generation depends less on finding impossible training examples than on building a base model that decomposes them into reusable concepts. A pickle sitting on an avocado chair is absent from the data, but chairs, sitting, anthropomorphic objects and pickles are not; “these models are not memorizing behaviors,” Amit argues, but distilling foundations that can be recomposed. That places data quality, efficient learning and representational depth—not raw dataset size alone—at the center of Luma’s thesis.
  • Concept School is Luma’s bridge between slow pre-training and professional demand for new cinematic controls. It teaches Ray 2 a motion, pose or grade from one to a few examples without degrading existing capabilities and allows concepts to compose; Bolt Cam, dolly and reverse-dolly controls are early specimens. The commercial ambition is explicit: teach the model “everything about filmmaking,” release new concepts rapidly, and eventually let customers “take the model to school” themselves.
  • Luma sees video generation as an AGI program, not merely a creative-software category. Jiaming Song calls video models “on the critical path” to general intelligence because stories require causal sequencing, character arcs and consequences across time; a multimodal model must jointly reason through language, image, video and audio. Amit’s contrarian claim is that creative work surfaced early precisely because “that’s where we require intelligence.”
  • The interpretability work Luma values most happens during training, where intervention can still change what the model becomes. Amit compares post-facto feature discovery such as Golden Gate Claude to “archaeology”; Luma instead studies information flow, frequency learning, curricula and hyperparameters while representations form. He says foundational pathways are substantially set in the first 20,000-30,000 iterations, making early data distribution and learning-rate choices disproportionately important.
  • Jiaming Song’s Inductive Moment Matching targets the economic trilemma of generative models: high quality, few inference steps and stable training. It generalizes consistency models from pointwise matching to distribution-level matching, using maximum mean discrepancy to avoid GANs’ learned-discriminator inner loop. In Luma’s ablations, one sample was unstable, two became unstable later, and four or more stabilized training—research that could continue the sharp decline in inference cost.
  • Luma’s largest strategic bet is that retrofitting images and video onto language backbones will not yield native multimodal intelligence. Jiaming notes that current multimodal language models can still trail dedicated vision systems on objects, segmentation and long causal threads; Luma instead wants a unified latent space because nature “doesn’t make a distinction between video and image and audio.” Nathan’s framing adds that competing with hyperscalers in a scaling-law regime will be difficult; the upside is a differentiated route to multimodal AGI.
Digest · the substance, structured for research

1. Strong base models turn impossible scenes into familiar primitives

  • Stephen Parker hypothesizes—and Amit agrees—that many image-to-video workflows increasingly begin with generated images because images are cheap to iterate: a creator can produce 100 candidates, locate the desired composition, then animate the winner. That workflow foregrounds the challenge of animating strange inputs—hybrid objects, unfamiliar characters and scenes that have no direct real-world precedent.

  • Amit’s representative test is “a pickle sitting on an avocado chair.” No training clip shows that exact event, but the model has seen pickles, chairs, people sitting and anthropomorphic things; success comes from understanding what each primitive means and recombining them. “These models are not memorizing behaviors,” he argues—they are distilling foundational capabilities.

  • Better base models also preserve the initial image’s identity and aesthetics over longer generations. Some reliability comes from systems that inspect and translate the image behind the scenes, but Amit characterizes those systems as temporary “crutches”: useful ways to expose a need until a later model can absorb the capability directly.

2. Luma repeatedly moves intelligence from scaffolding into the model

  • Amit distinguishes “inside the model versus outside the model” as a central design choice. External software can separately construct lighting, narrative and character behavior, but the latent space contains richer information and can coordinate appearance, causality, action and timing together. “Anything that can be done in the model is going to be better” than its external equivalent.

  • The original Dream Machine model, retroactively Ray 1, depended on scaffolding for motion, vocabulary translation and other controls. Ray 2 does not need even 10% of that support, yet harder customer requests have prompted another outer layer of systems. Character understanding, for example, is largely external to Ray 2; Luma intends to internalize it in Ray 3.

  • Amit’s metaphor is a “blood-brain barrier” between the user’s intelligence outside and the generative model’s intelligence inside. Controllability means penetrating that barrier with higher-level instructions, then letting the model internally collate the character arc, event sequence, lighting and other interactions rather than micromanaging each through a workflow.

  • Jiaming’s endpoint resembles communication with another person: users should not have to program separate image-to-video, keyframe and camera-motion pipelines. An intelligent multimodal model should accept the media and intent, infer the task, handle some back-and-forth naturally and execute it “without…thinking about it as a different type of task.”

3. Concept School converts scarce examples into cinematic controls

  • Visual models do not yet possess language models’ robust in-context learning, but professionals continually invent motions, poses and grades absent from the web. Luma’s “concepts” are designed to learn such capabilities from one to a few examples, preserve the base model’s other skills and compose with one another instead of producing the degradation associated with conventional fine-tuning or LoRAs.

  • Luma calls its internal teaching tool Concept School: “You take the model to school and you’re the teacher.” The team used it to teach Bolt Cam, dolly, reverse dolly and other camera motions relatively quickly from sparse examples. Amit says another batch was due that week and another the following week.

  • Ray 1.6 offered camera-motion controls, but Jiaming says the newer implementation is more flexible and needs less literal specification of camera coordinates. The next target lies between a Bolt Cam preset and exact trajectory programming: users might alter speed, angle changes or subject focus interactively while the model maintains the scene.

  • Stephen’s comparison sharpens the product value: the Bolt Cam opening of Severance season 2 reportedly took months and a robotic arm, whereas Luma exposes the idea as a generation control. Amit wants both everyday tools—tracking, moving left or right—and effects “which are just absurd and fun,” while emphasizing, “We’re building this for professionals.”

4. Storytelling is Luma’s route from video software to general intelligence

  • Luma’s mission, as Amit states it, is “multimodal general intelligence.” Materialized, that intelligence would resemble a “world in a globe”: a necessarily weak approximation containing physical phenomena, interacting intelligent beings and consequences unfolding through time. He prefers “world model” because “simulation” sounds weaker than the intended scope.

  • Creative work matters because it forces a system beyond procedure. Stories demand that event A lead to B, B lead to C and characters change across a timeline; those dependencies are closer to intelligence than merely emitting structured output. Amit draws the language-model analogy: once structured output became native rather than externally coerced, “suddenly you have agents.”

  • “Video models are on the critical path to that general intelligence,” Jiaming says. A model reasoning jointly through language, video, audio and images could become a partner that helps people dream, imagine and trace consequences—not merely a tool that fabricates pixels or follows a predetermined production graph.

  • Amit’s historical reversal is deliberate: people expected mechanical work to be automated first, yet AI’s earliest assistance appeared in creative pursuits. His explanation is that art requires abstraction and general thought: “That’s where we require intelligence.” Accordingly, art is an important part of building AGI, not a diversion from it.

5. Fictional knowledge can improve real-world action

  • Nathan Labenz presses on a potential conflict: a robot needs dependable physical knowledge, while Luma’s training also represents dragons, magic and impossible cinematic physics. Jiaming concedes that a short-term system for household activity might benefit from angling toward more physical, applicable domains so it “doesn’t hallucinate as much.”

  • Yet even a mundane command can require fiction: “Pick up the clothes with a dragon on it” fails if the robot cannot recognize a dragon. The imagined creature itself recombines real observations—lizard-like form, reptile scales, bat or bird wings—making fictional concepts useful handles for objects that exist in the physical environment.

  • Jiaming’s longer-term answer is contextual self-location. One model should distinguish whether it is acting in a physical room, a virtual interface or an imaginary world, just as a person can put a dragon T-shirt into a dishwasher and then move to a computer to answer an email. Specialization might help now, but should not be permanently necessary.

6. Utility—not human-like representation—is the test of visual understanding

  • Jiaming points to models’ implicit 3D knowledge: they can infer depth in realistic and fantastical imagery and represent cloth, waves and hair. Researchers also use pretrained generative models as priors for depth estimation, including methods he describes as state of the art; eliciting that knowledge may be easier than locating a discrete “depth neuron.”

  • Stephen’s pushback is worth keeping: perhaps video models possess increasingly expert prediction over a 2D pixel field, not an internal appreciation of three-dimensional space. Jiaming reframes this as a visual Turing test—perfectly predicting a scene does not prove the model uses human-defined physics, but functional equivalence may be enough from a utility standpoint.

  • Amit challenges the premise that humans maintain an explicit 3D representation either. The brain receives visual information plus other signals such as proprioception, but subjective certainty that “this is 3D” does not reveal its mechanism. A generative model therefore need not contain a mesh; if it creates consistent outcomes, an implicit representation can be sufficient.

  • His analogies separate phenomenon from mechanism: planes generate lift without flapping, humanoid robots move differently from people, and dishwashers clean without hands. A 20-watt organic brain and gigawatt-scale computing clusters occupy different substrates, so demanding identical representations could restrict machines rather than establish whether they are capable.

7. Interpretability is most actionable while the model is forming

  • Nathan accepts that AI need not think like humans but challenges Amit’s dismissal of interpretability through Golden Gate Claude: researchers isolated an apparent Golden Gate Bridge feature, amplified it and changed the model’s behavior. Could concept controls similarly identify and marshal latent directions for anime, a cinematic style or another user-requested capability?

  • Amit’s distinction is between pretending a large model is fully legible and running empirical experiments on a black box. “Science is not the act of interviewing God”; physics extracts patterns from observations, while even quantum mechanics can produce measurements validated at six, eight, ten and sometimes 23 sigma without explaining why reality has that form.

  • Post-facto feature work is therefore “like archaeology”: intellectually useful, especially for an unfamiliar model, but limited to reconstructing what already formed. Training offers a stronger opportunity because researchers witness “the formation of the planet and all the eons it went through” and can still alter its data, curriculum, information-flow architecture and optimization regime.

  • Luma tracks what happens to high- and low-frequency information, what is learned early versus late and when the curriculum should change. Amit says information pathways become established during roughly the first 20,000-30,000 iterations; later training emphasizes or suppresses them, like paths and crop circles forming in a field and then being retraced or erased.

8. Curated data beats indiscriminate scale, but no clean recipe exists

  • Nathan asks whether Luma constructs each batch to avoid damaging loss spikes. Amit clarifies that the team is not hand-selecting every batch; it approaches the problem through large-scale curation, filtering and removal of bad examples. “The technique…is not to throw 1 billion samples of garbage data at it”; it is to show the model what it should learn through good examples.

  • Jiaming identifies a deeper mismatch: training algorithms commonly assume samples are independent and identically distributed, but it is unclear whether a coherent “world distribution” exists or how one would define it. Population share, for example, is not necessarily the correct rule for deciding how much English belongs in a language-model corpus.

  • Dataset design is consequently objective-driven and empirical: alter the mixture, observe whether quality and desired capabilities improve, then iterate. Jiaming cautions that the exact solution is “probably very messy,” with fewer concise statistical principles than simplified descriptions of general-purpose machine learning imply.

9. Diffusion won by making web-scale generative training dependable

  • Jiaming traces the first known diffusion method to a 2015 NIPS paper by Jascha Sohl-Dickstein. It attracted little momentum on small datasets because GANs generated in one step and worked better at scale, while diffusion was slow and hard to motivate. The decisive 2020 DDPM work by Jonathan Ho and collaborators made diffusion competitive on selected tasks.

  • DDPM still required around 1,000 neural-network evaluations, but it removed GANs’ unpredictable instability. For practitioners, that was transformative: “You launch the job, you can go to sleep at night” and expect improvement, instead of waking to a run that had collapsed without warning.

  • Jiaming’s Denoising Diffusion Implicit Models work attacked sampling cost, delivering roughly 20-50× acceleration at the time. In parallel, Yang Song’s score-based research connected discrete denoising to continuous-time stochastic differential equations and 1980s mathematics, while later ImageNet experiments showed diffusion competing with GANs at larger scale.

  • Control then advanced from classifier guidance to classifier-free guidance, followed around 2022 by GLIDE, Imagen and Stable Diffusion. Jiaming explains classifier-free guidance through Bayes’ rule: conditional and unconditional model outputs supply the terms needed to steer toward P(X given Y), avoiding a separately trained classifier while retaining conditional direction.

10. Distillation and moment matching trade rigid paths for distributions

  • Progressive distillation teaches one model pass to replace two adjacent denoising steps, then repeats the compression to obtain fourfold and greater acceleration. It works, but Jiaming calls implementation “hairy.” Consistency models instead seek the same final prediction from different points on the denoising trajectory, potentially enabling one-stage or very-low-step generation.

  • In practice, consistency training proved less stable than its simple formulation suggested, so consistency distillation initialized from an existing diffusion model became more popular. The field recovered its old trade-off: consistency-style techniques are generally easier to train, while GAN-based methods may produce higher quality at extremely low step counts but inherit discriminator instability.

  • Inductive Moment Matching relaxes consistency models’ pointwise constraint: the generated samples need not follow one exact mapping if their distribution matches the target. It uses maximum mean discrepancy in a reproducing kernel Hilbert space, whose optimal comparison function can be represented without repeatedly training a neural discriminator, eliminating GANs’ unstable inner loop.

  • The ablation supports the mechanism: one comparison sample behaved like the unstable degenerate consistency case; two samples remained somewhat unstable but failed later; four or more produced stable training. Jiaming presents this as progress toward all three desired properties at once—high sample quality, few inference steps and a stable training process.

11. Native multimodality requires more than a language backbone

  • Jiaming’s closing position is that what the industry currently calls multimodality “is not really it.” Language models fine-tuned to accept images, audio or video can be useful within language tasks, yet may remain worse than dedicated vision models at object recognition, segmentation and following long causal threads.

  • The warning sign is architectural: if a supposed multimodal foundation model still trails specialist systems on fundamental perception, retrofitting every signal onto a language backbone might not be the final approach. The bag-of-tokens approach remains helpful, and Jiaming acknowledges substantial language work remains, but he argues that the industry needs to “look beyond just language backbones.”

  • Luma is pursuing a “unified singular latent space” that treats language, images, video and audio as facets of the same occurrence. “Nature doesn’t make a distinction” among those signals; an action produces all of them within one environment. Luma’s wager is that reasoning over that shared structure will produce the multimodal intelligence that present integrations only approximate.

Nathan Labenz

Today I'm speaking with Amit Jain and Jiaming Song, CEO and chief scientist at Luma Labs, makers of Dream Machine and the new Ray 2 video-generation model. I'm also joined for this episode by my friend Stephen Parker, creative director at Waymark and one of the few creators who has logged a proper 10,000 hours with video and image-generation models, dating back to the original DALL·E over the last few years.

Our conversation begins with a discussion of how the Luma team trains models to create fantastical and other fundamentally out-of-distribution visuals for which there is little to no relevant training data available. Considering the force of intellect that both Amit and Jiaming display, their belief that video models are on the critical path to AGI, their ambition to create multimodal AGI at Luma Labs, and the range of novel and occasionally hot takes they share, I think this episode should be of interest to anyone, regardless of whether you're particularly interested in video-generation models specifically.

Keys to Luma's model-development success, as you'll hear Amit explain, include a relentless focus on dataset curation, frontier advances in efficient learning algorithms, and a strong drive to understand what their models are actually learning as they go through the training process. These fundamentals create base models that can learn new concepts, including Bolt Cam and many other camera-motion concepts they've recently introduced, in a highly sample-efficient way.

Meanwhile, for things that existing models can't learn so quickly, we also discuss Luma's outer loop of product development, which consists of building scaffolding and other behind-the-scenes systems that unlock new model capabilities and also validate customer demand. With that done, they then seek ways to internalize those capabilities in the next generation of the model. Then they repeat this process for each generation as customers continue to apply new and better models to harder and more valuable challenges.

For me, the most interesting part of this conversation was the discussion of model interpretability. Emphasizing that we should not expect AIs to process, represent, or understand information like we humans do, or even to do so in a way that's generally human-graspable, Amit likens current interpretability techniques to archaeology, in the sense that they're fundamentally limited to piecing together what models have already learned in the past. More interesting from his perspective is the study of training dynamics and the engineering of datasets needed to teach models what they most need to know.

In the last 15 minutes or so, Jiaming offers an intellectual history of diffusion models. This gets pretty technical and, for most people, myself included, will require some additional study to fully understand. But I would summarize it by saying that the generative AI era really began with the realization that, with the right problem formulation, unsupervised learning can work on web-scale datasets. For text, this was simple next-token prediction. For images, it was gradually adding noise to real images and then training models to remove that noise one step at a time.

Since then, there have been a mix of practical tricks, theoretical insights, and model-enabled dataset improvements that have unlocked far more precise steering of outputs and breathtaking efficiency gains. These range from distillation techniques, which amount to training a model to perform multiple denoising steps in a single pass, to consistency models, which try to ensure that a model will generate the same output regardless of where it begins on its denoising path; to flow-matching models, which use theoretical connections to differential equations to take a more direct path through the latent space; to Jiaming's latest inductive moment-matching technique, which optimizes the model in distribution space and performs generations in a small number of optimized steps. All of this should, at minimum, give you a sense of how inference prices have fallen so precipitously even as quality has dramatically improved.

While we didn't have time to go as deep into the philosophical underpinnings of Luma's multimodal strategy as I might have wished, I left this conversation with the sense that Luma Labs is definitely a company to watch. It won't be easy for any model-development startup to compete with the big-tech hyperscalers in the scaling-laws era. But Luma's mix of product-market fit, vision and ambition, and research prowess gives them as good a chance as any I've seen. I absolutely look forward to having them back again in the future.

Amit Jain and Jiaming Song, CEO and chief scientist at Luma Labs, makers of Dream Machine and Ray 2, welcome to The Cognitive Revolution.

Amit Jain

Thanks for having us. It's very exciting to be here.

Jiaming Song

Yeah, thanks for having us. I'm excited for the conversation as well.

Nathan Labenz

I'm also excited to have my good friend and longtime teammate Stephen Parker here as well. Stephen occasionally co-hosts when we do an episode on creative models, especially in the image and video domain, because he is the creative director at Waymark and has logged more hours than anyone I know with these kinds of products. He really has an excellent handle on the exploding array of options in the market today.

I thought that, to broadly structure this conversation, we might start by discussing the latest and greatest stuff that you guys have launched in the products—the latest models, the camera motion, all that new, cool stuff—and how it fits into the broader picture and what you're seeing in terms of usage. Then I want to get into some of the more technical stuff, because you've also put out a very mathematical paper recently on an advance in pretraining for diffusion models and a really interesting position paper on how you think we should be thinking about pretraining going forward in general.

I'm excited to get into all that as well. Then maybe at the end, if there's time, we can get a little bit more speculative and talk about world models, the future of multimodality, what superintelligence looks like, and all that great big-picture stuff. We've got a lot of ground to cover, and I'm excited for it.

Stephen, awesome. Kick us off with a few reflections on your use of Ray 2 recently and what that has you thinking about as we begin today.

Stephen Parker

Yeah, thank you. Thanks for having me. Amit and Jiaming, it's an honor and a pleasure to speak to both of you, so thank you for taking the time.

I've been playing a lot with Luma Labs recently and really enjoying it. It's one of many video-generation tools that I love to use in my arsenal of possibility when I'm working on various projects. For my own work, I tend to find tremendous value in using several different models at the same time, as they all have different strengths and different weaknesses—really, just a whole blend of capabilities. Luma is right up there at the tippy top, especially at the cinematic end of the models that I like to use, so I just really want to give a shout-out to you guys there.

My first question is: I think more and more of what people want from an image-to-video scenario starts with AI images. That's just a hypothesis on my part, but is that correct?

Amit Jain

I think a lot of workflows do start on the image side, because images are just much easier to iterate through. Iteration cycles are fast and generation times are low. You can generate 100 images and find exactly the sort of things that you're thinking about. So, yeah, I think a lot of people actually lean into image workflows as a significant part of their work today.

Stephen Parker

Okay, so that makes sense with my own workflow and especially what I see out there on social media. I think what I'm driving at is that a lot of those images, to me, seem like they can be strange or new. I'm thinking of avocado-chair-type combinations, weird characters, and all of these sorts of things. What I really want to know is: What has it been like to push these models toward a greater understanding of what I presume is a more novel subject?

Amit Jain

That's really interesting. Currently, we are designing a feature, and that is one of the more important problems because there's no data. By necessity, these are out-of-distribution things and stuff that you would just not find in regular use cases.

The singular answer there is that you need to really work on a very strong base model with strong capabilities for understanding what is happening and then dealing with these very unrealistic, out-of-distribution scenarios. So, let's say you have a pickle sitting on an avocado chair, and you start with the pickle standing up and want it to come and sit down. You've never seen that scenario before, right? But you've seen chairs, you've seen people, and you've seen anthropomorphic things, and you've seen them doing things like sitting down, right? So this again comes down to the idea that these models are not memorizing behaviors.

These are not memorizing how this thing is done, how this is done, how this is done. They're generally distilling out base ideas and capabilities—the core things. What does it mean to be anthropomorphic? What does it mean to sit? What does it mean to be a chair? That kind of thing, right?

Then, when you combine them together, the better your model is at this foundational understanding, the better it's going to be able to do when it's presented with these entirely out-of-distribution, funny, uncharacteristic scenarios.

Nathan Labenz

If I'm understanding you correctly, improving the base model is just giving you better and better performance with things that are out of distribution. But another thing I seem to notice is that, especially with image-to-video, the inherent characteristic qualities of that initial image also seem to be carried through better and more effectively with longer and longer generations. Do you think that's aligned with improvement of the base model as well, or are you guys doing more stuff behind the scenes to check against that image more regularly as the generation is occurring? What's going on behind the scenes there?

Amit Jain

There's a lot that goes on behind the scenes in trying to understand what is in the image. These things are not just models; they are systems, and so we definitely take care of many things in the back to try to break some of these down. But generally, they are crutches for the model to be able to understand, and in the next iteration of the model, we make it so that we don't need that system.

The initial model, which was called Dream Machine and is now retroactively renamed Ray 1, had many of these crutches to make it feel and work much better—to understand motion, understand different vocabularies, translate things into language it understands, and all this kind of stuff. Ray 2 doesn't need even 10% of that. But now that people are pushing it to do more things, we have to build some of these systems again: “Okay, I see that's what someone is trying to do. How do we understand this part?”

For instance, characters: all the character understanding in Ray 2 is very much external to the model. But then, as we design Ray 3, it's all going to be internal to the model. This is a trend we have seen in language models as well. When people push the current model to doing things it was not designed to do, labs like us build systems around it to at least address those things. We gather the data and then bring it to the model.

This also applies to what we call the application layer and that kind of stuff. Application layers come up and build these specific use cases, but then we see that and are able to just build it out in the next model. The general rule of thumb in this industry—or at least on the technical side of it—is that anything that can be done in the model is going to be better than what is done externally to the model. We can talk a lot about that, but that's generally the situation.

Nathan Labenz

That's great. I'm happy to double-click there as much as you want. I know we're early in, but I do think this is a super-fascinating topic. It kind of skews toward secret sauce, company proprietary information, that sort of thing. And so there isn't, frankly, a lot of conversation about these kinds of systems built behind the models that you then try to reincorporate into the model moving forward.

If there's anything else you could say, maybe more specifically about—beyond data as a loose term—making sense of these systems and how you teach a model from there to appreciate what was inherent about that system that was helping you, I think that would be fascinating to know.

Amit Jain

That's a big conversation, and Jiaming, you should also chime in at any time. But to go a little bit deeper into inside the model versus outside the model—intra-model versus extra-model—that's a really important thing.

Whatever you can do in the latent space—all the thinking you can do about, for instance, video—is important. If you're able to reason about the actions that are being performed, the characters that are there, what that character would do, the lighting of the character, and all these kinds of things inside the model, rather than through a system that actually generates the right lighting, then tries to piece together the narrative arc, then tries to piece together all these things, that's much better.

Of course, that system can work. That's not the problem. But people tried to do that. If you remember, what was the name of the company? It was a scriptwriting company, or a copywriting company—Jasper, right? They had some really great ideas, to be honest with you. But the thing is, when you can do it inside the model, inside the latent space, you're just able to work with a lot more information. You're able to do the kind of edits that, outside the model, you just don't have the control mechanisms for the model to be able to do.

It's like this thing where there's a blood-brain barrier between a generative model and the outside world, right? And when people say, “Hey, we want controllability,” they want to penetrate the blood-brain barrier better and be able to tell the latent space exactly what they want. It will get there. There's no question about it.

Inside, there's a level of intelligence that's really, really great. Outside, there's a level of intelligence that's really, really great, which is you or me who is using it. Then there's a barrier. You want to communicate from outside in at a higher level and let the brain inside do as much as possible internally, collating information.

Think about sequencing of events in the video. Think about what happens to a character: their arc, their timeline, their causality—all these kinds of things. The more you can do inside the model, the richer the outputs you're going to get.

This is what chain-of-thought also looks like in language models, right? Instead of forcing the model to structure its thinking into, “Do this, do this,” no, chain of thought is—if you've read the recent Anthropic paper, where they claim chain of thought is actually fictitious—just a way to seed the latent space of the model, to direct it. Whatever it's outputting is not a realistic representation of what it's actually thinking about, right?

That is actually an indication of what we've seen now. Instead of trying to force the model to output JSON through external coercion, let's just make the model very good at structured output, and then suddenly you have agents, right?

You're going to see that in multimodality especially, because if you're able to simultaneously think in audio, video, language, and images all together, and combine reasoning from language, appearance from video, all the aesthetics you've seen in images, and the audio that you're able to hear in all the videos, you're going to have much better outputs than through external systems. But yes, we do build these systems, and then we obviate them in the next model training.

Jiaming Song

I tend to agree with Amit here. Instead of trying to program this into complicated workflows, just imagine how you would communicate with another human. You don't actually need to program their brains for them to understand that you want to achieve this level of workflow. Of course, there might be some back and forth, but the process itself is pretty much natural.

I think this can only be done if you build these capabilities into the model, in the sense that the models need to be more intelligent. Whereas, right now, even the current generation of models—especially video models—feels less intelligent in the sense that you have to tell them, or program them, to achieve some particular type of task, like image-to-video, for instance.

Of course, there are practical reasons that this is done, but I think eventually they'll be intelligent enough to take in whatever task, in the multimodal sense, you tell them. For example, “I want to do image-to-video, keyframes, camera motion,” and all these different tasks. The model should be able to do it without even thinking about it as a different type of task.

Nathan Labenz

Could you guys maybe give a couple of examples of things that presumably wouldn't be too sensitive to say—things that were essentially scaffolding in the Ray 1 generation that are now handled internally by the model—and then, if you're willing and open to it, things that are currently scaffolding on Ray 2 that you hope to be built into the model when we get to Ray 3?

I guess one example would be camera motion. In Ray 1.6, we had camera-motion features, but that was obviously harder to actually control and lower quality than what we recently released.

Jiaming Song

In this iteration, you can see that there are a lot more different types of flexible camera control, such as Bolt Cam movements, without the user having to prompt for exactly the extreme X and the extreme Y of the camera. It just works. Of course, you can also give it more explicit controls; that's what some users would want.

But I guess for other users, maybe what they really want is just, “Convert this scene—give me this effect in the scene,” without having to specify the very detailed controls. I think that's one of the features that we have in the model.

I think what would be really interesting is to build into the model—now it's one level ahead of being able to represent the scene, like a boom camera. Again, you want slightly more control than just generating a boom camera, but not as much control as having to specify the exact camera coordinates.

Jiaming Song

So maybe somewhere in between, you can have something like, “I’ve given it a camera trajectory scene. I want this Bolt Cam to be moving faster, moving slower, having more angle changes, or focusing on some other subject matter.” So this kind of more interactive editability of the scene is something that we are thinking about maybe having in the model in the next generation.

Nathan Labenz

Cool. That was actually my next question. So maybe just to restate a little bit here for our audience who might not be super familiar with Bolt Cams: that’s bringing us to Apple TV+’s Severance season 2, right? It made a huge splash this year with its opening shot, which is a Bolt Cam robot-arm shot. I think it was super impressive. It famously took them months and months to create and was just a big wow moment for audiences.

And now, just a little bit later this year, we have it as one of your key camera motion concepts already available for people like me to use when generating. It’s one of a range of motions that you have added to the model capability. Super useful, and I imagine very highly in demand from your editor-type user approaching the model, as well as novices and everybody else, but it’s got to be high on the professional list. What has it been like to develop that feature specifically with an outlook on the professional user, or maybe customer requests coupled with modern trends like Severance season 2 and the Bolt Cam?

Amit Jain

So basically, there are models, right? The things you teach them during pretraining, and then there are many capabilities that people want to teach them. Language models have this really special ability, which is called in-context learning: you give or show some examples, and it becomes that. And I think that’s a truly emergent intelligent capability. Visual models aren’t there yet. They’ll get there soon enough, but they’re not there just yet. But that doesn’t remove the need for teaching these models specific things you want in the moment.

So we have been working on this idea. We published a post about this, a white paper, whatever you want to call it. We call them concepts, and we realized that in visual, especially creative use cases, there are many things people want to teach—many things that they come up with, like a specific motion or particular kind of color grading or a particular human pose, which is just really suited for the story you’re trying to tell or the ad you’re trying to make, whatever it is.

But there’s no way to get it out of the model because it’s a new thing you just came up with. There’s no data for it on the internet, and there’s no way to generate large samples of it to even be able to fine-tune the model. So we designed this idea of concepts to make it so that models can learn from 1 to just a few examples. Their capabilities don’t degrade like they do when you’re fine-tuning or creating LoRAs, and they can be composed together.

How closely can it represent a capability that the model had at pretraining but users actually teach it? We call them concepts. The tool that we’re using to build them is called Concept School, right? Like, you go to school, you take the model to school, and you’re the teacher, and you’re going to teach it concepts or lessons, whatever you want to call it.

So the camera motions you’re talking about—the Bolt Cam, the dolly, and the reverse dolly, all these kinds of things that we taught—we were able to teach them to the model very quickly, relatively speaking, with very few examples. And you’re going to see the next batch come out this week, and the next batch come out next week, so on and so forth. Our goal is to basically teach our models everything about filmmaking this way and eventually also give people the ability to teach.

Right now, it’s a little bit finicky, as new technologies tend to be, so we haven’t made it open source—or, sorry, open access—yet, but we will in the future. So these camera things, right? The way we are building them, to answer your question on feedback and these kinds of things, there are some of these which are basic storytelling tools: being able to track a shot, being able to move left, being able to move right. These are done by people a thousand times a day whenever you’re shooting something, like, “Oh, the camera moves left or right,” or you’re tracking an actor who’s doing these kinds of things.

So you want that, and then you want to also balance it with some things which are absurd and funny, or which people just can’t do in real life very easily. Like, for instance, Bolt Cam. If you want to do Bolt Cam really well with that smooth tracking, you actually need a robot.

Yeah. MKBHD has one. I’m sure James Cameron has a few, but most of us don’t. So can we just have it in the model? We try to balance this great utility with some things which are just absurd and fun.

But ultimately, we’re building this for professionals. We’re building this for people who want to tell stories, who are already telling stories, or who want to become professionals in that world. And we can talk about the changes that are happening in the AI industry very quickly—not the AI industry, but the moviemaking industry very quickly. But, yeah, we’re designing for people who want to tell stories, and these are all storytelling tools.

Nathan Labenz

So, 2 things there that just brings to mind for me. One of them, selfishly, is talking about those repeat actions that the pro user takes. One of the—selfishly, one of the pro actions I take all the time is leveraging your audio generation feature, which is great. For people who don’t know, after you generate a video, you can just press the audio button, and it will give you another prompt opportunity and give you the ability to generate audio for that clip.

However, one thing I do all the time is reverse the playback direction of my video when I’m editing. So, just a tiny little plug here: I would love to have the ability to quickly flip the playback direction before generating that audio so that I’m not dealing with reverse audio when I take that clip out. But I’ll just leave that as a footnote.

I think really what this is getting to is you’re about storytelling. You have all kinds of users telling all kinds of stories, and that skews towards fantasy, Hollywood cinema, anime—that’s everything, right? Yeah. How do you imagine that translating to multimodal understanding? Is it just an attempt to understand everything everywhere all at once, or is there a particularly unique insight that you feel like you gain in the pursuit of art first that helps with the overall mission or multimodality?

Amit Jain

See, what we are trying to do is—the mission of Luma is to build multimodal intelligence, right? Multimodal general intelligence. If you were to not personify it, but if you were to materialize it in front of you and ask, “What does that look like?” the intelligence that LLMs embody looks very much like abstract intelligence that humans have with language and things like that.

When you think about multimodal intelligence, of course it has that abstract part, but what does that actually look like? It starts to look very much like a world, a globe. You have this world in front of you, and it has all the physical properties and all the physical phenomena that happen day to day in our physical world. It also has those intelligent beings inside it that do things that increase the entropy of the universe, that interact with each other, that do all these kinds of things.

So a term a lot of people use is world simulator, right? But I think simulation is a weaker term here than it should be. It’s basically just a world model. It’s like a physical manifestation of the universe that you have outside.

Jiaming Song

Of course, it's a weak facsimile. It's an approximation, but yeah, it's a model of the world and of the processes that we are having.

Now, coming to the creative side of it and why that is important for this mission: storytelling. If you think about it, with language models, we were like, “Oh, this is only good for JSON, and then we're going to produce just JSON.” That's not very good. When you force intelligent systems to play games, to come up with abstract new things, or to follow instructions—“Oh, no, no, I want this, then I want this”—that's when you get intelligent systems rather than just procedural systems that are following a set of rules that you have created.

When you force them to deviate from that by making movies and telling stories, that's a very critical part of it. It's a part of human existence too, right? How good someone is at storytelling is generally a very good barometer of IQ, right? Can they actually think beyond just the most physical thing that is in front of them? What is the consequence of A? What is the consequence of that? What is the consequence of that? What does it lead to?

That's what stories are, right? Event A happens, then B happens, then C happens, and then D happens. Video models are on the critical path to that general intelligence. Once we're able to combine video, audio, and language all together, they will become really, really good at storytelling—being a partner for us who can actually sit next to us and help us dream, help us imagine, and help us think through those kinds of things in creative pursuits.

It's not a surprise, by the way, that people thought the first things to be automated would be the mechanical things. It's not a surprise that the first things where AI is able to help are actually creative pursuits, because that's where we require intelligence. That's where these things come into play. So, yeah, I think art has always had a very significant role to play in general thinking and general intelligence. Art is also going to have a very significant role to play in building artificial general intelligence.

Nathan Labenz

Can I ask a couple of world-model questions?

Yeah, I think this is super interesting. One big thing that jumps out at me, though, is that if I were trying to create a generally useful AGI to go out and do stuff in the world—to maybe control robots and have one walk into my house and make me a coffee, to take a famous example, and vacuum up my living room after my kids have thrown toys all over it, whatever the case may be—I would want this sort of world model.

I'm certainly a big—I wouldn't say “believer” at this point is the right term, because I think it's pretty well demonstrated that the large foundation models are learning these higher-order concepts, and that has, I think, become pretty much indisputable at this point. So I'm well convinced by the evidence in the literature that this is happening, but I wonder about the sort of fictional side of it, the magical side. There are no dragons in the real world, but your models have learned to also represent dragons, right?

So, in a way, you have a real-world model plus something that sort of goes beyond what is real and into the imaginative, the fictional, et cetera. I wonder, obviously, that's good for storytelling because we want to tell stories that are not bounded by reality, but does that have downsides for the practical utility of making multimodal intelligence that I just want to come into the world and do useful work for me?

Jiaming Song

Yeah. I think there's a short-term answer and there's the answer for the long term. I guess the answer for the short term is that if you want to use this type of model for your daily activity tasks right now, then yes, it might be better to angle toward the more physical, applicable verticals, so to speak, for this type of model.

However, in your model, there are many cases where you ask the model to do some task. For example, say, “Pick up these clothes with a dragon on them and put them in the washing machine.” If the model doesn't really know about the concept of a dragon, you can't actually pick up the clothes with the dragon pattern and put them into the washing machine.

So even this kind of knowledge that is fictional can still be very useful for real-world physical tasks, because, again, even the concept of a dragon by itself is related to other physical concepts that people see in the world. For example, it looks like a lizard, it has scales that are exhibited in reptiles, and it has wings, which you see in bats and birds.

So even these concepts, if you think about it, are human interpretations of what they see in the real world. The dragon is generated by humans looking at the real world as part of their training data, and they generate something that is out of the ordinary.

But back to the original question: yes, for the short term, I do think maybe focusing on more physical, realistic things can be better for these kinds of tasks, so the model doesn't hallucinate as much. But in the long term, I believe that the model should be able to tell for itself what kind of world it is acting in, because there are also other types of worlds it needs to be able to act in, and it's not just a physical world.

For example, in a virtual world, the things people are currently doing with agents are also a perfect example. The model doesn't need world knowledge, but it's a world that is different from the physical world that we are acting in. So having a unified AI that can do both tasks could be useful.

For example, eventually your robot could be like, “Okay, put this dragon T-shirt into my dishwasher, and then go to the computer and type an email responding to my friend,” or something like that. This requires the AI to be able to reason between physical and nonphysical, imaginary, or virtual worlds, just like humans can do. So I think eventually this will be a capability that does not have to be specialized into any sort of model.

Nathan Labenz

I don't know if you're doing interpretability work on your models internally, or if you have any partnerships, academic or otherwise, that allow you to do that kind of stuff. Maybe you know the answer because you may have done this work, or maybe you could speculate as to whether you would expect to find features in the model that would represent, “This is a realistic physics simulation,” versus, “This is a magical-realism-type scenario,” versus, “This is a Minecraft environment that we're in right now,” or we're playing Pokémon, or whatever the case may be.

Jiaming Song

It seems like that probably would happen at sufficient scale, but I don't know if we are there yet or if anybody has really had the opportunity to look. Up front, I'm not an expert on deep neural network interpretability, and I haven't delved too deeply into this space. My very shallow understanding is that it might be good to seek more alignment, or try to get the model to align with what you want, and reason from there.

For example, one feature that we have discussed heavily in the first world model is this implicit knowledge about 3D. You can ask it to generate things that reason about depth, and even in fantasy or unrealistic art, it has a reasonable interpretation of depth and other physical features, such as cloth simulation, waves, and hair and that kind of stuff. So the question you're asking is whether there is a neuron or substructure within the model that shows that. The honest answer is that it would be very difficult to find that within the model, but it might be easier to prompt the model and ask it to do the task for us.

For example, people have shown that these models are very good priors for things like 3D vision tasks, such as depth estimation. A lot of the state-of-the-art depth-estimation methods are actually based on these pretrained models, which have a very good understanding of the world. So I do believe that, to some degree, the model has great internal knowledge about the world, but it is also up to humans to determine how interpretability exists. Currently, it seems like the API between human neurons and model neurons is not very well-defined just yet. We will probably still have to use language that both sides understand to communicate this interpretability.

Nathan Labenz

Yeah, there are definitely some broad challenges there. Yeah, go ahead.

Stephen Parker

This is an interesting philosophical question that Nathan and I go back and forth on all the time. We tend to hear from people a lot that these models have physics knowledge inherent in the training. They have a great, rich 3D understanding of the world; we see it in this or that. I personally am not entirely convinced that they aren't just seeing more and more movement across a 2D space and developing greater and greater pixel understanding.

I can understand somebody arguing against that from, perhaps, a systems-based training regime where we're training first on 3D scaffolding or something like that. But my own naive understanding is that sort of thing isn't happening in the training. I just think it's a really interesting question: Are we actually seeing physical appreciation for a 3D space, or are we just seeing more of an expert interpretation of a 3D pixel space?

Jiaming Song

I would relate this ability to predict or reason about physical scenes in a representation of 2D pixel space to a Turing test. In the Turing test, your goal is not to determine whether the model knows grammar, or whether it internally represents the scene using the grammar that humans define. That's not really part of the question we're asking. What we're asking is whether, when talking to the model, it feels the same as talking to a human. In this case, it's similar.

On the one hand, being able to render a prediction of the next frame of a scene perfectly over the next few seconds or minute doesn't actually mean that, inside, it has the same physical knowledge that we, as humans, currently define. But on the other hand, from a utility standpoint, it might constitute passing the Turing test for visual generation.

Amit Jain

Yeah, and I haven't had a lot of time to think about this problem. The question is basically philosophical, in line with what Jiaming is saying: It borders on the definition of understanding, right? What does it mean to understand? People make the argument that humans have something very special or deep where we understand at some level. But it's very hard to argue about that as well, because how would you know?

You can say, “Oh, yeah, but I know this is 3D. That's why I understand it. It clicks in my brain that this is 3D and not 2D.” But the brain is this really interesting organ that is simultaneously thinking and telling you—or itself—that it is thinking, and that's so weird, right? Think about it: It's self-aware. The self-awareness part is the most interesting part of it.

Coming back to 3D understanding for just a second, what does it mean for humans to have 3D understanding? I'll take a slightly different stance than most computer vision researchers, which is, “Oh, humans actually perceive 3D,” and things like that. We don't really perceive 3D. Yes, we have 2 eyes and stereo vision, but stereo dies out at about 20 cm. Outside of that, the disparity is almost nothing, right?

Case in point, if you hurt 1 eye and have an eye patch, you can still drive. You have some initial trouble trying to grab a piece of glass or something close up, but you can still drive. Of course, you have hands and proprioception. The brain gets other signals that teach it depth and some of these things, but it's hard to really argue that the brain actually maintains some sort of 3D representation. It might not. We just don't know. We have no way of understanding it.

So, for a generative model, why is it any more special that it has an explicit 3D representation inside it, like a mesh or whatever have you, than just understanding these concepts more implicitly or in terms of 2D space and time? As long as it's able to generate something that looks really consistent, why do you care how it actually did it? Planes fly, but they don't flap their wings. The phenomenon that birds use and the phenomenon that planes use are actually very similar, right? Birds generate lift by moving their wings; planes do that by pushing themselves through the air.

So while the phenomenon is the same, the mechanisms are very different, right? Humanoids right now—we are trying to build them. They move, but their movements seem very different from how we do it. It's going to happen more and more with all machines that do things humans do. For dishwashing, we wash dishes very differently from a dishwasher, but are the dishes any less washed because the dishwasher doesn't have hands? I don't think so.

Philosophically, coming back to it, I don't think machines have any prerogative to think exactly as we do or to have representations inside them like we do. It doesn't make them any less intelligent. It doesn't make them any less capable. In fact, they should do it differently because the substrate they're on is very different. The human brain is a 20-watt piece of organic tissue. On the other hand, here we are running things on gigawatt-scale clusters. Why should they think the same way?

So that's my answer to that problem, actually. I think people who are focused on making machines think exactly as humans do, or making them exactly as interpretable as humans are, have misguided goals. They are not only reducing the capabilities of machines in that process, but also wasting their own time. They should use that time to scale attention better. They should use that time to design regimes that can do more efficient learning, all these kinds of things, right? I think it's a waste of time.

Nathan Labenz

All right. I agreed with the first 85% of that, but the last 15% I want to challenge.

Amit Jain

Go for it.

Nathan Labenz

When it comes to interpretability—why would you care?—the 85% does include the idea that I don't expect AIs to be representing things in the same way that we are, or thinking broadly in the same way that we are. I don't think that should be used to discount them. I have a funny series of tweets where it's like, “It's only a concept if it comes from the concept region of the human brain. Otherwise, it's just sparkling notions,” or whatever. So I'm with you in terms of viewing that as a straw man.

But I do think when you look at something like Golden Gate Claude, for example, they were able to say, “Okay, we were able to go in and isolate what very strongly appears to be the Golden Gate Bridge concept.” Now, when we artificially turn that up, we get Golden Gate Claude. Yes, that's just a curiosity and a cute demo, but I actually thought when you were talking about the concepts feature that you were building, maybe you were doing it that way.

I could imagine that you might say, if you had taken—and I'd be interested to hear how you are doing it, because it seems like it's not this—but what I had imagined you might be doing there is taking a similar approach, trying to identify directions in the latent space and then injecting them or turning them up in order to enable these different concepts. It seems like you probably could do that and be like, “Turn on anime,” or “Turn on Toy Story style,” or whatever. Stephen has a much better vocabulary for the different styles than I do.

But I assumed that your concept work would be similar to Golden Gate Claude. I guess the questions there would be: Doesn't that seem like it would be a very useful thing to study and potentially be able to marshal? And if you're not doing that, maybe can you tell us how you are doing it? There's a difference between treating something as a black box and being able to do empirical experiments on it.

Amit Jain

But we are in the philosophical land again. Think about physics, right? A lot of people think our knowledge of the world is completely interpretable, and because we have an equation, we understand how the systems work. That's not how any of the laws of physics actually are, right? Science is not the act of interviewing God, if that existed—or nothing one way or the other. Science is basically: we have an observation. Can we derive a pattern out of it? That's about it. Science is empirical. Theoretical physics is also, again, very much about building a mathematical model, an approximation, a representation.

Machine learning is very much the same. The entire universe is very statistical, right? The best theory that we have of how the universe functions right now, quantum mechanics, is entirely uninterpretable. Today, we don't understand why things are the way they are at all. The measurements get verified up to 6 sigma, 8 sigma, sometimes 10 sigma—as accurate as it gets. There are some measurements that are up to 23 sigma, right? But we don't understand why the world is that way, right? That's what you say when you're asking for interpretation, right?

I'm saying ML models, because of their sheer scale, are just not grokkable by the human brain. You're not going to come up with this coherent, consistent model of how the ML model functions, because it's just not that kind of a system. We can do empirical experiments on it after the fact, like the Golden Gate experiment that you're talking about: we found a cluster of this; we found this. It's like archaeology. We can do archaeology, all right.

But here's something really powerful that we can't do in the physical world. In the physical world, we can only do archaeology. Here, we are actually involved in the formation of the planet and all the eons it went through. That's called training. So we know what the model is going to do.

Instead of people spending their time on archaeology, I think a much better use of time is understanding data and training processes. You actually learn so much during training. At the risk of Jiaming giving something away here, when we train our models, we are working on actually designing these models with a great degree of efficiency. In that process, we do so much information-theoretic work on our models to try to understand what sort of information-flow architecture is in there: what is happening to high-frequency details, what is happening to low-frequency details, what is it learning in the earlier stages and later stages, and when should we actually change the curriculum of what it is actually learning?

This is all basically what people talk about as interpretability. This actually has real consequences. This changes what the model learns, and some of the things we have learned are really interesting. The information-theoretic pathways are set up in the first 20,000–30,000 iterations. The model starts out as this scalar field.

Really, what is a good example? Think about a field of corn that's grown. It's just plants everywhere; they look uniform, all these kinds of things. As you train, pathways keep forming in between them, like someone walking and putting the plants down, right? And then there are crop circles. When you look from the top, you're like, "Oh, there's a pattern. There's a circle," and things like this.

Somewhere, these happen very early, and the distribution of data you have at that time and the kind of learning-rate regimes you've used decide a lot of what is going to happen in the later stages. But in the later stages, a lot of different things happen, right? These pathways are emphasized or deemphasized. Someone goes and undoes the crop circles; someone decides we're not going to take that path as much. This is the kind of interpretability that allows you to design the models that give you the output you want.

This post-facto interpretability work, I think, is pretty interesting. No, don't get me wrong: intellectually, this is empirical science, right? It's very good, and if you have an alien model in front of you, or some model someone else made, it's very useful to understand what's going on. But if you want to actually control the models, you want to control the data and the training process—and I mean the actual hyperparameters of the training process. There's so much that goes into that. So that's where actual interpretability comes into play, in my view, and that's where we do a lot of work.

Nathan Labenz

Yeah, that's super interesting. So you could expand on any number of dimensions of that. One that I've been thinking about quite a bit recently is what you might call batch strategy. It sounds like you're probably getting pretty intentional in those early phases of training about making sure that you're feeding, batch by batch, the right mix of data.

Because I can imagine if you had—I've always looked at these loss curves, and you see these occasional spikes, and you're like, "What's happening there?" My one interpretation is maybe that's a bad batch, some sort of cluster of data that the model wasn't prepared for. But then also, if my interpretation is right, that batch is probably sending a bad signal to the model as well, which is not actually constructive for general learning purposes.

So are you actually doing this sort of batch-by-batch construction of data to try to get the mix right at that very granular level, or am I misinterpreting what you're saying?

Amit Jain

I think Jiaming can answer this a bit better, but I can clarify my own statement. I don't mean handpicking it that much. I think we come at it from the other direction, which is a lot of curation and filtering and removing garbage examples. The technique to training really great models is not to throw 1 billion samples of garbage data at it and have it figure it out from the garbage, right? The technique is to show it what you want it to learn and show it good examples from there, at large scale.

So, yeah, you do a lot of this preprocessing, filtering, that kind of stuff. But Jiaming, I don't know if you have a different take on that.

Jiaming Song

Yeah, I think it's a very interesting question. Basically, what is the right data set? Interestingly enough, the kind of statistical methods that we're using for training these machine-learning algorithms are actually quite different from what the real world actually is. So if you run this algorithm, whether it is on language models or a diffusion model, you are mostly making i.i.d. assumptions.

Basically, you have a data point and assume that these were drawn independently and identically from the world distribution. But the real question is: is there a world distribution? And how do you even define the world distribution? That's a real question.

Regarding the question of what is the best data set, it's actually very hard to have the right answer. I think people probably take a less statistically driven, but more objective-driven, approach. It's like, how do I control the data set such that I get a better quality or better outcome?

This has also been deeply studied in the realm of language modeling, for example: how much English do you want to have in your entire corpus? To be honest, it is very hard to reason about these on a theoretical level, because you wouldn't want to say, "We just base how much English to put into the data set on the number of people who speak English." So that's probably not the right way of doing it.

Instead, having a more empirically focused approach is probably the right way of finding it. But the exact solution is probably very messy, and there's not really a lot of key, concise principles behind it, despite how easy it is to explain the idea of machine learning to the general public.

Nathan Labenz

Jiaming, would it be too much to ask for you to give us a short master class in the intellectual history of diffusion models? I think you're probably the best person in the world, maybe, to do that.

Jiaming Song

Yeah, yeah, of course.

Nathan Labenz

So I sketched it out, and the mantra that I always say to myself about the original diffusion models is: there are all these images out there. It's really easy to just programmatically add noise to them step by step until you get to pure noise. Then the insight is, if you just train the model to do the reverse and learn a denoising step, then if you run a bunch of those steps in a row, you can start from pure noise and eventually get to some image.

The original versions of that were totally undirected and just generated an image out of nowhere. Then we got classifier guidance, and that sort of was enough to steer us in the direction of particular images. Then I want you to take over and tell us the major updates in the field that have brought us to where we are in terms of efficiency and control.

Jiaming Song

Yeah, of course. So I guess the first known method of diffusion models was actually reported in this 2015 NIPS paper from Jascha Sohl-Dickstein. He described the actual algorithm for how to train, or how to formulate, the forward and reverse process of a diffusion model.

However, that never really caught on, because back then the experiments were done on very small MNIST data sets, and it still had the same problems that plagued diffusion models very early on back then. To researchers in that field, it seemed like, "Okay, I have GANs that work in one step, and they also work much better at large scale. Why am I bothering myself with this method that seems very hard to grasp and very slow to generate samples?"

So the first real breakthrough in this field came in 2020 with a paper called “Denoising Diffusion Probabilistic Models” by Jonathan Ho et al. The basic idea is that, first of all, it validated the validity of this idea and made it as performant as GANs on certain tasks. It was still very inefficient, in the sense that you needed maybe 1,000 steps to converge to the right image, but it was one of the first algorithms that did not have the unstable-training problem that GANs give you.

For practitioners like us, having an algorithm that trains stably is very important. There is a huge difference between launching a job, going to sleep at night, and waiting until the next few hours for the result to get better, versus the GAN case, where things just become unstable out of nowhere and you have to do a lot of digging. The training algorithm being stable is a very big deal here. This is actually why people got interested in diffusion models, even though they were very slow at the time.

I was part of the people who worked on these general directions. I had worked on other types of models, including GANs, before, so to me, diffusion models seemed like a very fresh perspective. Before the existence of that paper, there really wasn’t any method that could train stably and generate high-quality samples.

Then I was trying to think more about this: Now we have a model that generates high-quality samples, and at the same time, this model is very stable to train. What is the problem with the model? The problem is that it generates very slowly. It takes about 1,000 neural-network iterations to get to a high-quality sample.

So I was looking more into solving that particular angle of the problem, and this is how I got to my work on Denoising Diffusion Implicit Models, or DDIM, which was able to accelerate the model’s sampling time by 20 to 50× back in the day. Of course, simultaneously, Yang Song, who was also from our lab at Stanford, was working on the more generalized idea of diffusion models.

His idea was that, instead of having a fixed, discrete number of time steps, we could make the time step continuous and wrap it more into the math that came from the 1980s, from a person called Anderson. It was all within this stochastic differential equation framework, along with score matching and denoising score matching. Yang was a very early advocate of score matching and denoising score matching, so it naturally fit into his framework.

Then he found this very interesting connection between diffusion models and denoising score matching. Those were the first 2 breakthroughs: one on the theoretical level and the other on the practical level. Then OpenAI did the work on training these models on ImageNet, which was again a big breakthrough at the time. It showed that diffusion models were competitive with GANs in these cases, as well as introducing classifier guidance, which people are still using today.

Then people started to use classifier-free guidance. They realized that we didn’t actually need to train an additional classifier. You could add, say, an unconditional signal to the model such that you could replace the role of a classifier. This also came from Jonathan Ho. After that, there were a few papers around 2022 that tried to scale diffusion models beyond ImageNet.

In early 2022, there was a paper called GLIDE from OpenAI, which basically tried diffusion models for text-to-image generation. A few months afterward, the Imagen paper from Google came out, and a few months after that, Stable Diffusion was released. That is basically how we got to Stable Diffusion.

Of course, after Stable Diffusion, there was an explosion of these techniques and, similarly, open-source models. The next things people cared about were how to get higher-quality models, how to make them work on videos, and how to make them more efficient. I’ll talk more about the efficiency side of things.

There have been a lot of efforts, initially and still today, to do distillation. Previously, the way people did distillation was kind of hacky. The idea was that I had this 1,000-step process, or maybe an even longer process, and I tried to use a model to represent 2 contiguous steps so that I could reduce the number of steps by 2. Then I tried to train another model to simulate those 2 time steps, so that overall I had a 4× acceleration, and repeated this again.

This is called progressive distillation. It also came from Jonathan Ho. But it is pretty tricky to train because the implementation gets a bit hairy.

In early 2023, there was another paper called “Consistency Models,” which at the time aimed to be a replacement for diffusion models, in the sense that it was both efficient and could be trained in a single stage. It also described a method called consistency distillation, which is a distillation-based method for consistency models. Basically, you use an existing diffusion model as the base and try to distill it using consistency-model ideas.

What actually became more popular in the field was consistency distillation, because when people tried training consistency models for other use cases, it turned out not to be as easy as it seemed to train them very stably. People came up with new methods to stabilize the training process for consistency models, which usually still involved initializing the model from a diffusion model, so to speak.

Then came these types of distillation techniques, and there is another set of techniques that are based more on GANs, so to speak. The difference between consistency-distillation techniques and GAN-based methods is that the consistency-distillation techniques are, in some sense, more stable to train, while the GAN-based methods are less stable to train but may have higher quality if you run them for much fewer steps.

Again, this is the interesting trade-off between how stable it is to train the model and how easy it is to get high-quality samples. That takes us back to the original story between GANs and diffusion models. That is basically the current status quo on diffusion models: the history of diffusion models, how people have scaled them up, and the key problems in making diffusion models even faster through this path of distillation.

Nathan Labenz

Could we do just a little bit more on classifier-free guidance, and maybe also on the intuition behind the consistency model? When we’re doing guidance, what exactly are we doing to guide the model? I get it in the classifier sense, or at least I have an intuition where I would say, “Okay, feed it to the classifier and take its feedback.” But when we’re classifier-free, I think it’s a little less intuitive. The consistency-model concept is also a little less intuitive than simple distillation.

Jiaming Song

Sure. Before we talk about classifier-free guidance, we can talk a little bit more about classifier-based guidance. The idea is that diffusion models try to represent a score, or the gradient of the log probability, of this distribution called P(X). You try to guide it with some conditional signal. Let’s just call it Y.

Instead of trying to sample from P(X), you want to sample from P(X given Y). You can train on this, of course, but another way to treat P(X given Y) is to apply Bayes’ rule. Basically, P(X given Y) is equal to the joint P(X and Y) divided by P(Y). You can think about P(X) as the score of the diffusion model without any condition, and the P(X and Y) part as the score of the model with the condition, because the other term, P(Y), which is the condition part, is something you don’t actually care about during sampling, so you can just drop it.

Basically, that is why in classifier-free guidance we have an unconditional model, which represents the denominator, and a conditional model, which represents the numerator. All in all, it is basically an application of Bayes’ rule.

A consistency model is basically something like this. Maybe it’s easier to explain consistency distillation first. The idea is that I want to have a 1-step model to generate the right solution, but I want to find a way to bootstrap it from a regular diffusion model.

Suppose I have a perfect 1-step model and it follows the trajectory of the diffusion model. The idea is that we want to learn a model to distill the process that a diffusion model would normally go through with many steps. We are just distilling this function.

What a consistency model would do, in principle, in the distillation case is that you can compute the 1-step prediction at different time steps. Time step is a concept that is correlated with how much noise you add. The more noise you add, the higher the time step in this particular case.

In a consistency model, the idea is to build a connection between 2 quantities. The first quantity is: At a given time step, what prediction are you going to make in 1 step? The second quantity is: Suppose you are using 2 steps, and 1 step is a regular diffusion-model step. You run 1 regular diffusion-model step to a time step that is closer to your original state, and then you run this same consistency model to reach the final step.

So basically, the way consistency distillation is trained is that it tries to minimize this loss function. Because your consistency model at the earlier time—the time step closer to the clean signal—has an easier time predicting what the real signal is, it allows the model to build this connection and train this function. But essentially, what a consistency model tries to do is use a model to distill the otherwise hard-to-compute process that is run by diffusion models.

Nathan Labenz

There are so many directions we could go here. Let’s do your latest contribution, Inductive Moment Matching, which I basically take as being the best of all of these prior approaches combined into one. There’s sort of an echo of the consistency model idea and an echo of distillation. What jumped out to me most about it was the idea that the model is being optimized in distribution space, and at more of a—I mean, I guess it’s always batch-level—but at a higher level than just example by example.

Jiaming Song

Like I mentioned, there are 3 things we want to achieve in generative modeling algorithms. One is high sample quality, two is stability to train, and three is that it is relatively efficient when you are trying to draw samples from it. Most of the—not most, but all—of the existing methods suffer from 1 of the 2 drawbacks.

For example, generative adversarial networks, or GANs, are not that stable to train. The same kind of thing goes, to some degree, for consistency models as well. Diffusion models are stable to train and have high sample quality, but they can’t generate high-quality samples in very few steps, so the inference cost is high.

What we want is to find an algorithm that satisfies all 3 of them: high-quality generation, fast sampling, and a stable training process. In this case, we try to reason about a generalization of what consistency models are doing. Instead of trying to match pointwise samples—basically, I have this function and I want to exactly match the samples—what we can do is match only the distributions, because we don’t actually care that this function has to match exactly.

For example, if you have a bunch of samples and you are trying to push them to another set of samples, there are many different solutions you can use. But in consistency models, you are forced to follow 1 type of solution. That is a little bit more restrictive for the model, and maybe your model needs more capacity to achieve what it is being asked to do.

This is possibly 1 of the reasons why these GAN-based methods have an advantage, because GAN-based methods are actually comparing samples at a distribution level. So we started thinking: instead of trying to match the samples pointwise, we can just match the samples at a distribution level. What is another algorithm that can match distributions based on samples that is not a GAN?

It turns out that this idea has been discussed in the statistics community at least 15 years ago, and this idea is called maximum mean discrepancy. The idea can sound a bit scary, but what you can think of is that it has a very interesting relationship with GANs.

In GANs, you have the discriminator trying to maximize the distance between the prediction of a real sample and the prediction of a generated sample. In maximum mean discrepancy, you do the same thing, except that the discriminator is no longer a neural network. It is a function defined on a space called an RKHS, or reproducing kernel Hilbert space.

You can think about it as a feature representation of the function—a simple function based on complicated, infinite-dimensional features. It is a more closed-form function. What is interesting about MMD, or maximum mean discrepancy, or this RKHS choice in general, is that you don’t actually have to optimize the discriminator.

Once you define a particular type of space to optimize for, the optimal solution can already be represented. That means you skip the inner loop of optimizing the discriminator, which makes this whole process less unstable. You end up with a very stable optimization process that still minimizes the distance between distributions as you are trying to learn with generative models.

Of course, the downside is that you chose this space of functions a priori, so you don’t have the space to optimize it for the best-case scenario. But in our experiments, we didn’t find this to be a huge problem.

Essentially, the easier way to interpret how our approach to Inductive Moment Matching works versus what consistency models are doing is that it is a generalization of consistency models, in the sense that it does distribution-level matching. In consistency models, the distribution matching is based on a single point, and of course you can’t easily represent a distribution with a single point, so it becomes a more degenerate case.

This kind of explains why, in certain cases, consistency models are unstable to train. We actually did ablation studies controlling the number of samples we use to compare distributions. With 1 sample, the consistency model is unstable. With 2 samples, it is also a bit unstable, but it becomes unstable later. With 4 samples or more, the training process actually becomes more stable.

That’s how we get into this 1-stage process that has high-quality generation.

Nathan Labenz

Maybe just 1 last question, because I know we’re at time. Looking forward, you guys are obviously deeply invested in multimodality. I’d love to understand how you think about multimodality broadly and where you think it’s going.

Jiaming Song

We probably don’t have time to explore this answer, to be honest with you. This is the entire foundation of the company. But I would say this: currently, what people think of as multimodality is not really it.

These are language models that have been fine-tuned to work with images, audio, and video, and they show some capabilities that are really beneficial in the context of whatever the language model is doing. But as you’re seeing individual capabilities—like understanding objects, object recognition, segmentation, or just being able to understand what’s going on, following long, long threads of things people are doing, causality, and all these kinds of things—this understanding in multimodal models is still worse than in dedicated computer vision models.

That tells you that this might not be the approach, right? We’re not quite there. It would be like if we built all these language models, but they were still worse than RNNs at interpreting language or natural language, right? But that’s not the case. Language models are fantastic at that.

What we’re seeing is that this is a very promising direction, and obviously the bag-of-tokens approach is extremely helpful, but we need to look beyond just language backbones and trying to retrofit everything to them.

Luma’s approach is very different. We are coming at it from the direction of a unified, singular latent space, where we can think and reason about all these different pieces of information as if they were one. Technically, they are one, right? Nature doesn’t make a distinction between video, image, audio, and things like that. These are just signals. They just happen to all be part of the same simulation, and they’re all part of the same environment that we are in.

An action produces these signals in all of these different modalities, exposing different facets of its existence and occurrence. We need to think about it in that same way.

Currently, we see that most of the industry is extremely shortsighted when it comes to thinking about multimodality. For good reason, by the way, right? There is so much to be done in language, and people should continue to do that work. But there’s a new approach that is necessary to actually solve multimodality, and that’s what we are pursuing.

Nathan Labenz

Okay, cool. I’m looking forward to part 2 already.

Luma Labs' Diffusion Revolution: from Dream Machine to Multimodal Worldsim - Amit Jain, Jiaming Song | BidClub