[BidClub_]
All-In · · 64 min

Open Source Wins, AGI Is Here, and Scorsese’s AI Toolkit with CEOs of Cerebras & Black Forest Labs

Andrew FeldmanRobin Rombach

YouTube
TL;DR
  • Cerebras says AI infrastructure is serving booked demand, not betting that customers will eventually appear. Andrew Feldman disclosed a $25 billion backlog and described customers as “trying to capture yesterday’s demand,” with football-field-sized facilities drawing more power than midsize cities. For investors, the operational challenge is keeping customers while demand outstrips the ability to build and fill data centers.

  • Reasoning models make inference speed and token availability part of product capability. Rombach says models such as “Fable” and “5/6” increasingly infer intent without a “prompt whisperer,” while his Hermes-agent experiment used a Bittensor project running Z.ai’s GLM-5.2 and debated its own research strategy before acting. Long-running jobs can already produce “amazing things”; Rombach’s hypothetical was that 15-times-faster Cerebras inference could compress “weeks or months’ worth of thinking” into a day.

  • Cerebras expects its young architecture to improve substantially faster than conventional Moore’s law in the next 18 months. Feldman’s view is “way over 2x,” because a new architecture still offers workload-specific optimization unavailable to a 20-year-old GPU design reliant on smaller fabrication geometries. Hyperscalers’ custom chips do not invalidate merchant silicon demand: their primary objective is avoiding total dependence on any supplier.

  • Open models are becoming the cost, control, and sovereignty layer beneath frontier AI. Feldman’s analogy is that users should not “take your Ferrari to the grocery store”: frontier models handle the hardest problems, while open models absorb routine G&A and the “cutting and pasting economy.” Regulated enterprises also want domestic, on-prem deployment; Cerebras therefore runs gpt-oss-120B, GLM, Kimmy, the “Quincy” set of models, proprietary OpenAI models, and customer-built models.

  • Staged release and red-teaming become defensible when a model can expose serious software flaws in hours. Rombach relayed that Palo Alto Networks found previously unknown bugs and paused other work for six weeks of patching; Feldman said giving government defenders “2 or 3 weeks to patch any obvious holes” is reasonable. He simultaneously warned against reflexive regulation and treated a future massive data breach as inevitable: “Something will happen.”

  • By every AGI definition used 20 years ago, Feldman believes the threshold has already been crossed; the open question is where recursive improvement stops. Repeatedly asking models to learn, retry, and broaden their search yields “not a little bit better answers, but vastly better answers,” compressing thousands of human-style generations into machine-speed iteration. His ledger includes real labor dislocation, but also a chance that children and their loved ones do not die of cancer and that every child receives an adaptive tutor.

  • Black Forest Labs sees generative media as the foundation of multimodal world models, not merely an image-and-video tool business. Robin Rombach traced FLUX back to latent diffusion, then toward models pretrained on images, video, and audio and combined with action prediction—the same model could help make a movie or become “a brain on a robot.” Near-term value remains human-directed: Martin Scorsese used the models to externalize a scene in his head, while IP owners can combine controlled generation, customized models, and potentially licensed fan creativity.

Digest · the substance, structured for research

1. AI infrastructure is chasing yesterday’s demand

  • Feldman’s scale claim is deliberately physical: upcoming data centers may consume more power than the previous 50 years on Earth, while individual football-field-sized buildings receive more electricity than midsize cities. Construction spans the US, Canada, the Nordics, France, Europe, the Middle East, Kazakhstan, Tajikistan, Georgia, and Armenia.

  • Cerebras has a $25 billion backlog, but Feldman says it is not alone. OpenAI, Anthropic, Google, Microsoft, and AWS also want more data centers, while Rombach names SpaceX and xAI among the insatiable capacity buyers. These customers are not pursuing “if you build it, they will come”; they are “trying to capture yesterday’s demand.”

  • On token-maxing, Feldman says experimentation can waste some resources, but the net value is enormous. He compares early token access to inexperienced Costco shoppers buying “four things you didn’t need and each was $22” before learning to shop strategically.

  • Enterprise behavior is already maturing from “Everybody, as many tokens as you want” toward differentiated allocation: highly productive teams get what they need, while cheaper or open models serve other work. Rombach says individuals “start playing with the tool, and then the tool starts playing with them,” forcing users to articulate goals, systems, and requirements.

2. Reasoning turns token volume into a hardware problem

  • Early computers and prompts did “exactly what you tell them”; tiny wording changes could transform the output. Feldman says “Fable” and “5/6” now increasingly infer the desired chart, layout, or solution without requiring users to be a “prompt whisperer”—a material shift from summarization toward understanding intent.

  • Rombach’s Hermes-agent experiment used Z.ai’s GLM-5.2 through a Bittensor project with unlimited capacity. Asked to become the world’s best trend hunter, the agent debated whether to search Hacker News, Reddit, social media, or Instagram: “You were watching a reasoning model work out.”

  • Unlimited tokens might mean “unlimited reasoning,” because models can spend 25 or 48 hours exploring a problem. Rombach’s explicitly hypothetical Cerebras case—15-times-faster execution sustained for 24 hours—could produce “weeks or months’ worth of thinking,” making inference latency part of answer quality rather than merely user experience.

  • Feldman says Cerebras broke from the traditional 18-month doubling curve and expects “way over 2x” improvement during the next 18 months. His logic: new architectures retain large workload-specific optimization opportunities, whereas a 20-year-old GPU architecture depends more heavily on shrinking fabrication geometry.

3. Open models become the minivan of enterprise AI

  • Hyperscalers building silicon does not mean every chip must beat the merchant frontier. Feldman’s explanation is institutional memory: cloud providers once depended on Intel, while GPU vendors depended on a few hyperscalers. “You just can’t be entirely dependent on other people’s chips.”

  • Model routing follows the same diversification logic. “You don’t want to take your Ferrari to the grocery store”: OpenAI, Anthropic, and Gemini may serve frontier problems, while routine Workday manipulation, G&A, and the “cutting and pasting economy” need reliable open-source capability rather than “gold medal math.”

  • Sovereignty adds a second demand driver. Finance, healthcare, and other regulated industries may require domestic, on-prem systems with tighter control over leakage and intelligence; Feldman calls gpt-oss-120B a good move but argues the US needs “more domestic open-source models” to offer an alternative to Chinese models.

  • Feldman presents Cerebras as able to run a wide range of models: GLM, Kimmy, the “Quincy” set, closed OpenAI models, GlaxoSmithKline’s internally developed model, and models from G42 and MBZUAI. Government actions regarding “Fable” and “56” were, he says, particularly in Europe, a wake-up call.

4. Staged releases look reasonable once models find serious holes

  • Asked whether a powerful model should roll out gradually, Feldman offers an honest limitation: “I hadn’t seen it before,” so he cannot judge that specific release. In principle, though, staged access resembles precautions for powerful pharmaceuticals and could give government red teams “2 or 3 weeks to patch any obvious holes.”

  • His balancing point is that polarization “hurts clear thinking.” The US should preserve fierce competition among Dario, Sam, and their peers and avoid becoming a region whose first instinct is regulation, while recognizing that developers are inventing both the technology and its guardrails without a playbook.

  • Guardrails themselves add latency. Cerebras discovered during the previous six weeks that faster chips can make those protections “less painful,” illustrating how safety and performance are coupled rather than separate product layers.

  • Rombach’s sharpest evidence came from Palo Alto Networks’ Nikesh: a tested model reportedly found bugs the company did not know about, forcing six weeks of patching. Feldman expects a massive breach eventually—an unspecified “unknown unknown”—and argues institutions should prepare their response before “something will happen.”

5. Old AGI tests are obsolete; recursive learning is the live question

  • Feldman agrees that “by any definition we had 20 years ago we’ve hit it.” Earlier Turing tests and science-fiction waypoints have been surpassed; the harder task is generating questions appropriate to capabilities that previous observers did not anticipate.

  • The recursive mechanism is simple but potentially exponential: ask, learn from the answer, retry with more information, and broaden the search. These loops produce “not a little bit better answers, but vastly better answers”; what remains unknown is whether improvement stops at compute, token, or budget limits—or keeps climbing.

  • His large-project analogy carries the argument: Versailles and other major structures accumulated learning across three or four generations of specialist families, while humans often wait 20 or 40 years for paradigms to change because “paradigms don’t die. People do.” AI instead approaches Drosophila-like iteration—“two a day”—compressing thousands of equivalent generations.

  • Feldman does not deny economic dislocation; cars were bad for horseshoers and carriage builders. His counterweight is a chance that cancer research might spare children and their loved ones, plus adaptive tutors that teach each child individually rather than today’s “factory farming” classroom model. “Thoughtfully and fairly” written, he thinks the ledger can come out okay.

6. Black Forest Labs is building multimodal world models

  • Rombach traces Black Forest Labs’ foundation to latent diffusion, developed by him and collaborators as PhD students in Munich. The method compresses natural data into efficient representations—analogous in principle to JPEG or MP3—then trains transformers over them; the team subsequently built Stable Diffusion and FLUX on that base.

  • The roadmap now combines models pretrained on images, video, and audio with action prediction. Rombach distinguishes intuitive visual intelligence from a deep reasoning layer: a complete intelligence needs both, but video pretraining can already provide implicit knowledge of physics and real-world interaction.

  • Control has expanded progressively from text-to-image, to text-plus-image editing, to multiple images combined semantically through a prompt. Applying the same pattern across video, audio, and action makes every modality a potential input or output, exposing more “manipulation layers” to users and developers.

7. Scorsese’s use case is communication, not one-click cinema

  • Rombach rejects prescribing how generative models should be used: “These AI models, they are a medium.” With Martin Scorsese, the team explored an Eastern European village, iterating from the director’s description until images communicated the mental scene more directly than language could.

  • The mechanism matters more than celebrity endorsement. Language is a “lossy communication medium,” whereas an image or video contains unusually rich signal; the technology can “parallelize your brainstorming,” extending the familiar storyboard and miniature workflow rather than removing the filmmaker.

  • Rombach is unsure that generating an entire movie is the ultimate goal. He favors a human-in-the-loop workflow and will not predict when high-end cinema arrives, though capability has progressed from 64-by-64-pixel images during his PhD to high-resolution, multi-minute videos.

  • Robin supplied the production-economics case: Gal Gadot described a Bitcoin film shot on a soundstage with generative scenery, reportedly costing $30 million rather than the $150 million physical sets would have required. Rombach confirms some comparable production use, but stresses that the technology remains on a rapidly improving trajectory.

8. The same visual stack reaches robots, IP libraries, and customers

  • Rombach’s most expansive thesis is that “the same kind of AI model” could make a movie and serve as a robot’s brain. Generation becomes simulation; action prediction requires perception and understanding. Today, varied factory hardware still needs a few hours of task-specific fine-tuning data, while the longer-term goal is natural in-context instructions such as picking up a glass.

  • On intellectual property, Black Forest Labs blocks generation of certain IP in its public tools. It also works directly with IP holders on customized models, sometimes built from open-source foundations and sometimes from more powerful proprietary systems—a control-plus-capability proposition for owners of large libraries.

  • Robin’s proposed consumer model is licensed fan creation, extending Star Wars fan fiction and films into AI-made untold stories. Rombach agrees that a model acceptable to rights holders while enabling “super-creative customization” could let viewers visualize alternate events and interpretations they already imagine.

  • The business was two years old and had just crossed 100 employees. Black Forest Labs is hiring in Germany and San Francisco for large-scale and diffusion-model researchers, customer-facing engineers building physical-AI or IP solutions, and infrastructure specialists focused on stable training and maximizing MFU.

Robin Rombach

We are in the race for superintelligence, and Andrew Feldman is back. He's obviously the CEO and founder of Cerebras, which is doing inference chips, pioneered the space, and had a successful IPO. We've talked about this a couple of times. We got to see each other in January at Davos. The IPO happened, and the boys and I got to sit with you recently.

Andrew Feldman

That was fun.

Robin Rombach

At Liquidity.

Andrew Feldman

That was really fun.

Robin Rombach

We had a great discussion with the boys, but I wanted to deep dive with you about a couple of topics. The first one is the buildout of AI. We've never seen a buildout like this since the Great Wall of China.

Andrew Feldman

Right.

Robin Rombach

Who knows—since the pyramids? It feels like the amount of capital, time, and intelligent people on the planet dedicating themselves to the buildout of something—I can't think of anything in our lifetimes, but perhaps before our lifetimes, the war effort.

Andrew Feldman

Right.

Robin Rombach

This is a mobilization at a scale that we read about and hear about, but you're actually doing it. You have customers who are building data centers, and you're a key piece of that.

AppLovin started with an $8 domain and no VC funding and became one of the largest ad platforms in the world. Now that same engine powers AppLovin ads for e-commerce. Your ads run inside mobile games reaching over a billion people with full screen distraction free attention. The platform finds buyers and optimizes for profit. You set the target, it does the rest. One cookware brand went from $4 million to $16 million turned profitable and is on pace for 80 million this year. Visit applovin.com/allin to launch your first campaign today.

Maybe you could just enlighten us: In 2026, what is Cerebras doing, and what is happening with this buildout in Texas? These are some gigantic efforts.

Andrew Feldman

The size and scope of what is being built—physical size and scope—are usually not what we talk about when we talk about software or hardware. We're talking about chips and boxes, and they don't have the same sort of physical enormity.

Robin Rombach

Right.

Andrew Feldman

Right. What we're talking about now are data centers that, in the next several years, are going to use more power than the previous 50 years on Earth used.

Robin Rombach

Wow.

Andrew Feldman

Right. We're talking about individual buildings the size of football fields that have more power coming into them than midsize cities. They're being built across the U.S., in Canada, throughout the Nordics, here in Paris, and throughout France, Europe, and the Middle East—in nations that weren't front and center in anybody's mind previously. Kazakhstan, Tajikistan, Georgia, and Armenia are all building out data centers of size. Everybody's focused on huge data centers.

Robin Rombach

And every state in America obviously feels it needs to participate in this. The people who are buying the capacity—OpenAI, Anthropic, SpaceX, xAI, and Google—are insatiable right now.

Andrew Feldman

Yeah. How many years out are they building? When you talk to them, they were ordering chips from Cerebras before you were finished with the chips. They're putting orders in ahead of time.

Andrew Feldman

The irony is that, unlike many exciting times in technology, they're trying to capture yesterday's demand. The demand is way outstripping our ability to build data centers and fill them with hardware. We have a $25 billion backlog.

Robin Rombach

$25 billion backlog.

Andrew Feldman

We are not alone in that. OpenAI and Anthropic—you go through this list. Google wants more data centers, Microsoft wants more data centers, AWS wants more data centers. All of these players are not chasing, "If you build it, they will come." They're chasing demand that is already booked.

Robin Rombach

Right. How do we keep them from leaving?

Andrew Feldman

Right. That's extremely unusual.

Robin Rombach

It's very unusual. Now we have people who are token-maxing.

Andrew Feldman

Yeah. You know, what I liken this to is when we first started with AWS, and it was so good to get around your own IT organization that you told every engineer, "Go ahead, put it on your credit card. Sign up."

Robin Rombach

Yeah.

Andrew Feldman

Right? A lot of it was really useful, and some of it was, "God, I wish we didn't do that."

Robin Rombach

Yeah.

Andrew Feldman

For sure, there's experimentation. But it doesn't mean that the net value isn't enormous. It means some of it is going to go nowhere.

I remember when Costco opened up in the Palo Alto area in 1988. People used to shop at Costco like they shopped at Safeway. They'd go down every aisle.

Robin Rombach

Yes.

Andrew Feldman

That's a horrible way to shop at Costco, because you end up with 4 things you didn't need, and each was $22. As people got more accustomed to it, they'd go to the back, get the chicken, get 18 cupcakes for the kids' birthday party, and be out.

Robin Rombach

Bang, you were out. Strategic. Strategic.

Andrew Feldman

And it's exactly the same. I think at first people opened up and said, "Everybody, as many tokens as you want." In enterprises, there's no open loop. We don't give any resource unconstrained to people. Now we're jumping on and saying, "Whoa, all right. These guys should have as much as they need. They're enormously productive. Over here we can use maybe an open-source model, maybe a cheaper model over here." Now we're running like a business.

Robin Rombach

We're really seeing a certain type of person emerge who knows how to deploy this technology: systems thinking, which developers tend to have innately. CEOs tend to be great strategists and understand systems.

The intelligence is getting so much better every step along the way that I'm watching individuals—typically startup founders, but also venture capitalists and associates who work at my venture firm—start playing with the tool, and then the tool starts playing with them. They start to go, "Oh, I haven't clearly defined what my goal is. I don't understand what a system is. I've never heard about making a requirements document."

The software is like, "Do you have a requirements document? What's your goal?" The AI starts telling people, "You're token-maxing, and you need to get a little more focused here."

Andrew Feldman

One of my colleagues, 20 years ago—a really smart computer scientist—said, "Computers are really dumb. They do exactly what you tell them."

Robin Rombach

Yeah. At first, prompting was like that, right? You modified your prompt a little bit, and it changed the answer.

Andrew Feldman

Dramatically.

Robin Rombach

Dramatically. Increasingly, it's understanding what your intent was.

Andrew Feldman

Right.

Robin Rombach

Right. If you have a chance to play with Fable or 5/6 from OpenAI, increasingly you don't have to get the prompt just right. You don't have to be a prompt whisperer. Instead, you ask it, and it says, "Well, here's some things. By the way, maybe you wanted the chart to go two ways. You wanted it aligned in a bar." And you're like, "Well, that's exactly what I wanted. I didn't ask for it, but that is better." It's understanding intent, and that's a huge leap.

If we were sitting here 2 years ago, we never would have been able to predict that, in a short 24 months, it would go from being a great summarizer and researcher of web results—

Andrew Feldman

Right.

Robin Rombach

—to actually understanding your intent, providing a solution, and abstracting it all from you.

Andrew Feldman

That's right. That's right.

Robin Rombach

Which is a very weird thing. I don't know if you've played with the Hermes agent yet. Have you played with it yet? I asked it just this morning, and I was given a secret Bittensor project that has the new Z.ai model, GLM 5/2.

Andrew Feldman

And they gave me GLM 5/2, yeah.

Robin Rombach

GLM-5.2. Somebody in Bittensor—I don't think you understand Bittensor. You've heard of it, the distributed crypto project. They have all this extra capacity. A whisperer told me there was probably some capacity in China that has free energy. Okay, fine. They gave me unlimited capacity, so I started having it do some really crazy jobs. I said, "Every hour, I want you to tell me what the trends in the world are that nobody else has identified yet. You can do whatever you want to do to accomplish that, but my goal is to be the smartest trend hunter in the world."

I watched what it was doing in the background, and it started debating itself on where it should find things. It said, "We should probably go to Hacker News and Reddit." Then it was like, "Yeah, but there's also social media, and trends tend to manifest on Instagram."

Andrew Feldman

That's a reasoning model. You were watching a reasoning model work things out.

Robin Rombach

Yeah.

Andrew Feldman

Isn't that interesting? I mean, that's amazing.

Robin Rombach

It was collapsed. So, as a civilian—

Andrew Feldman

Right.

Robin Rombach

—who doesn't get the uncollapsed moment, if you were using ChatGPT 3.5 or 4.8, whatever it was, and you haven't used this new level of reasoning, inference, and essentially unlimited compute—

Andrew Feldman

Right.

Robin Rombach

—it opened my eyes just this morning to what a world of unlimited tokens might look like.

Andrew Feldman

Right.

Robin Rombach

Because unlimited tokens, I believe, means unlimited reasoning.

Andrew Feldman

It does. What does that mean?

Robin Rombach

Yeah. If you run these for 25 or 48 hours, you get amazing things now. What if, by using Cerebras, we were 15 times faster and then you ran it for 24 hours?

Robin Rombach

Right.

Andrew Feldman

And you got weeks or months’ worth of thinking.

Robin Rombach

Yeah.

Andrew Feldman

And I mean, it is extraordinary. I think one of the things is that people like Ilya and Sam, in the early days, were saying this was coming.

Andrew Feldman

Right. And I think when you look back, you say to yourself, “Holy crap. Those guys—”

Robin Rombach

They knew.

Andrew Feldman

Saw it.

Robin Rombach

Yeah, they could see around the corner.

Robin Rombach

That’s right. And the rest of us were like, “What? I’m not sure.”

It’s when we had Sam on All-In at one point. He said, “Oh, you know, I’d love to come on at some point.” I said, “Sure, come on.” He was talking about it, and I said, “What’s next?” He said, “Reasoning.” I said, “Unpack that. What does it mean?” He was like, “Well, understanding what your intent was, just as you’re saying, and then figuring out a strategy, and then maybe talking to other agents and other threads about, like, is this the right thing to do, and vetting each other’s work.” And I’m like, “Wow, we have come a long way from ‘guess the next word.’”

Andrew Feldman

Right. Right. Right. Fill the sentence in. Summarize this PDF.

Robin Rombach

Now, Cerebras is at the center of this because this reasoning is inference.

Andrew Feldman

This reasoning is inference, and it’s computationally intensive.

Robin Rombach

Right.

Andrew Feldman

Right. And so fast compute makes this sort of work fast and sort of tractable. It doesn’t have to take a huge amount of time to get a good answer. And so it’s exactly the fact that this reasoning consumes a huge amount of tokens internally that allows a blisteringly fast machine like ours—and I brought one—

Robin Rombach

Oh, you—

Andrew Feldman

I’m never far without one. When one costs half a billion to make, you bring it everywhere with you.

Robin Rombach

And we were tossing this back and forth at Davos. What’s the model number of this one?

Andrew Feldman

This was in the first 8 or 10.

Robin Rombach

Got it. So this has a special place.

Andrew Feldman

This has a special place. My wife says it’s like I’m a kid with a dirt bike for his 8th birthday. He’s in his bedroom at night. I carry him with me.

Robin Rombach

I mean, when you have your next party at the house, I highly recommend just a little hors d’oeuvres—something. I think it’d be a great fit. It would be a great fit if you had some—

Andrew Feldman

That’s right. [laughter]

Robin Rombach

But what we’re looking at here is the ability to do that reasoning at scale. What is Moore’s law for inference, and for Cerebras? Do you have something internally you discuss as, “We’re going to double this every X time period”?

Andrew Feldman

All chips prior to us in the processor world followed Moore’s law.

Robin Rombach

Got it.

Andrew Feldman

And we broke it—

Robin Rombach

Doubling every 18 months.

Andrew Feldman

Doubling about every 18 months.

Robin Rombach

Got it.

Andrew Feldman

And we crushed it with this chip, and we’ve carved out a whole new trajectory. My view is, in the next 18 months, we’ll be way over 2x.

Robin Rombach

Interesting.

Andrew Feldman

And so I think that early in an architecture, you have room to do much better than what was traditionally Moore’s law. Now, if you’ve got a 20-year-old architecture like the GPU, it’s much harder.

Robin Rombach

Right.

Andrew Feldman

You have to rely on things like smaller geometry.

Robin Rombach

Mhm.

Andrew Feldman

Right, going to the next fab node. But in a newer architecture, you have a huge amount of room still to learn about the work that is being presented and make optimizations that give you huge gains.

Robin Rombach

How do you run the company? Just being the CEO now in the age of AI—you have $25 billion in demand. You have to deploy at an incredible, blistering pace. You have to hire people. You have to create a road map. I don’t mean to give you a panic attack here. You have to keep up with somebody like OpenAI, which is moving so unbelievably quickly.

Andrew Feldman

Yes.

Robin Rombach

Right? And they’re competing. You’ve got to keep up.

Andrew Feldman

Right.

Robin Rombach

Right? Your hardware, your software, your deployments have to keep up with some of the fastest-moving organizations in history.

Andrew Feldman

They’re demanding customers. They are not—

Robin Rombach

They’re not pushovers, for sure.

Andrew Feldman

Yeah, and also potentially competitors down the road. Look, I think there’s so much demand right now that there is no silicon that will go unused.

Robin Rombach

Right. But why isn’t OpenAI releasing Jalapeño? Why is Amazon making its own chips? You see this recurring trend. Is it a way to let you know, to let Jensen and NVIDIA know, “Hey, we can do this too, so we need good pricing”? Is it a little bit of a flex that way, or is that the future—that they’re going to be in your business?

Andrew Feldman

No. I think nobody likes being dependent. And I think some of the lessons learned by the hyperscalers of the x86 world is that they were dependent on Intel.

Robin Rombach

Mhm.

Andrew Feldman

And some of the lessons learned by the GPU makers was that they were dependent on a small number of hyperscalers.

Robin Rombach

Yeah.

Andrew Feldman

And they wanted more customers. And so they set about to help fund these neoclouds.

Robin Rombach

Mhm.

Andrew Feldman

And so I think mostly it’s about an opportunity to control at least an important part of your destiny.

Robin Rombach

Got it.

Andrew Feldman

And I think that’s a very reasonable thing. You don’t have to make the fastest chip. You just can’t be entirely dependent on other people’s chips.

Robin Rombach

And that dependency has become a hot topic. I’m not sure if you caught the episodes over the last 2 weeks, but we’ve been talking over the last year about open source. I’ve been championing that a lot, just because I was early into open claw and quickly started using Kimmy and was like, “Wait a second. I’m blowing out my claw tokens, but this Kimmy, I can’t tell the difference.” And then we started smart-routing it, and suddenly this open source started to figure out reasoning, and the gap—

Andrew Feldman

Well—

Robin Rombach

—suddenly closed this year.

Andrew Feldman

Well, you don’t want to take your Ferrari to the grocery store. There are times you want to drive your fun car, and there are times you want to throw the kids in and not worry if there are Cheerios on the floor.

Robin Rombach

Minivan time.

Andrew Feldman

Right. There’s minivan time. And I think that as the sophistication of the user grows, you’re going to have hard problems, and those are going to be frontier-model problems. They’re going to be OpenAI problems, Anthropic problems, Gemini problems. And behind that, there are going to be a lot of ordinary problems.

Robin Rombach

Right? I mean, if you think about a company, you know how much time is spent cutting things out of Workday and getting it in a different cell for—

Andrew Feldman

Yeah. Right.

Robin Rombach

The cutting-and-pasting economy is real. Yeah.

Andrew Feldman

That’s right. And this doesn’t need gold-medal math.

Robin Rombach

No.

Andrew Feldman

What this needs is rock-solid open-source capabilities.

Robin Rombach

Yeah.

Andrew Feldman

And if you think about what I mean—well, we’ve been thinking a lot about it in G&A, but a huge amount of G&A, all right, is not invention.

Robin Rombach

No.

Andrew Feldman

Right? And you may not need the most sophisticated agents for this.

Robin Rombach

And another card that’s turned over recently is that some folks may have concerns with the ambition of the frontier models, and maybe sharing their data—data leakage and sovereignty of intelligence. They’re saying, “Hey, our company is going to choose—maybe we’re in a regulated industry: finance, health care, HIPAA, FINRA, all kinds of different regulations. We need to have this on-prem—”

Andrew Feldman

Yeah, but on-prem—

Robin Rombach

—domestically, and we’d like an open-source version where we have a little bit more control.

Andrew Feldman

Yeah.

Robin Rombach

And I think—are you seeing that now?

Andrew Feldman

I’m seeing that for sure, and I think OpenAI made a good call releasing gpt-oss-120B some months back. That was a good open-source model. But I think in the U.S. we need more domestic open-source models. We need to give the world a choice. Right now, if they want to run open source, it’s gpt-oss-120B or Chinese models.

Robin Rombach

NVIDIA has some.

Andrew Feldman

NVIDIA has seen the same opportunity to push open-source models. I think giving them more power might be sort of—

Robin Rombach

Well, I was about to— that was you. You cut me off at the pass. My understanding was Jensen was like, “Hey, we don’t even want to talk about these open-source models we have because our customers—”

Andrew Feldman

Right.

Robin Rombach

—are now going to be competing with Sam, Dario, Elon, Sergey. Do we want to be in that position?

Andrew Feldman

Right.

Robin Rombach

But we do need some more champions here, and it’s open source, so people can fork it. But that puts you in a more neutral position.

Andrew Feldman

That’s right. We run GLM, we run Kimmy, we run the Quincy set of models, and we run OpenAI’s models, the closed-source ones. We run models for, say, GlaxoSmithKline, which they wrote and developed. We run models for our partner in the UAE, G42 and MBZUAI, that are their models that they designed. So we have a wide variety.

Robin Rombach

So sovereignty is a trend.

Andrew Feldman

Sovereignty is a trend, and I think the government's actions with regard to Fable and 56, where they said, “Whoa, whoa. Let’s think, and then we can act.” I think particularly here in Europe, it was a bit of a wake-up call.

Robin Rombach

And when you saw this going down, there’s a layer of partisanship in our country right now.

It's pretty fervent. Dario is pretty explicitly not part of this administration.

Andrew Feldman

They've been very adversarial. Both sides have admitted that. They're starting to work it out now. So it's hard, I think, for us, not being in the room with these parties, to understand what's partisanship and what's gamesmanship here.

Robin Rombach

But do you believe that what they released was truly dangerous for cyberwarfare, for cyberattacks? And if you were to rate Dario's communication—not his communication, because he's a very effervescent communicator; I think that's a diplomatic way to say it—but to have a scheduled rollout release, we'll put aside the government's control of it, do you think that is a wise thing for us to do at this point? And do you think there was actually a major threat there?

Andrew Feldman

So what's interesting is I hadn't seen it before.

Robin Rombach

Mm-hmm.

Andrew Feldman

Right. And I think, if we just step back and say, "Is it reasonable?" I don't know whether this was the right time, but at a time—

Robin Rombach

Mm-hmm.

Andrew Feldman

—that a model is sufficiently creative in its thinking that it poses a meaningful threat for the government to say, "We'd like you to roll it out in steps"—

Robin Rombach

Yeah.

Andrew Feldman

Now, this doesn't seem unreasonable to me.

Robin Rombach

Not at all.

Andrew Feldman

Right? I mean, we do this with powerful pharmaceuticals. We like—I mean, we're certainly not encouraging 7 years of trials and the amount of paperwork and all the garbage that has accrued to the FDA. But with a powerful new technology, it certainly doesn't seem unreasonable to say, "Hey, guys, let's at least do some red-teaming at the government so we know our defenses can block this."

Robin Rombach

Yeah. Have we checked the infrastructure of the country, like—

Andrew Feldman

NSA? Have we checked the infrastructure of—right? [Clears throat.] And can you give us 2 or 3 weeks to patch any obvious holes that are found?

Robin Rombach

Right. This doesn't seem to me—

Andrew Feldman

—an unreasonable thing for the government to ask.

Robin Rombach

Right. But in this very polarized time—

Andrew Feldman

Yeah.

Robin Rombach

—you put on top of it, "Oh my God, it's President Trump doing it," and then you have to think, "Well, what if it was President AOC or President anybody in between the two extremes?"

Andrew Feldman

I think the polarization hurts a great deal. It hurts clear thinking.

Robin Rombach

It does.

Andrew Feldman

It hurts clear thinking, and both sides are going to do some dumb things and some really smart things.

Robin Rombach

Right.

Andrew Feldman

Right. And in fact, what I found is that the people in the government are trying really hard.

Robin Rombach

The rank and file.

Andrew Feldman

The rank and file are trying really hard. And this is moving fast. And I think that our ability to set aside some of the polarization and say, "How do we do this in a reasonable manner?" I mean, we want Dario and Sam competing like crazy.

Robin Rombach

100%. It's been awesome to watch.

Andrew Feldman

Yeah.

Robin Rombach

Right. It's good for the technology. It's good for entrepreneurs to see, even with thousands of people, that this is what you can continue to achieve.

Andrew Feldman

Right.

Robin Rombach

Right. This is a pain—

Andrew Feldman

—in the ass. It makes them get sharper. Amazon did a better job at that.

Robin Rombach

Everybody got better because of that. We want that. And we certainly don't want to become a region where the first thing we want is regulation.

Andrew Feldman

Right.

Robin Rombach

Right. But as it gets more powerful—

Andrew Feldman

The industry really should do a better job of regulating itself, perhaps. And it did seem like they were starting that process, but then the communication was lacking, maybe.

Robin Rombach

You know, I think not only are they racing hard, but they're inventing this as they go, too.

Andrew Feldman

Yeah. Right. There's not a playbook.

Robin Rombach

No.

Andrew Feldman

Right. They're inventing that. We say, "Oh, just put on guardrails." Well, they have to design the guardrails.

Robin Rombach

Sure.

Andrew Feldman

Right. The guardrails have an impact. One of the things that fast does is make the guardrails less painful, and we discovered that in the last 6 weeks.

Robin Rombach

Yeah.

Andrew Feldman

It is that the very guardrails can add time and make it feel slower, and so fast chips like ours can really help that. But—

Robin Rombach

Right.

Andrew Feldman

Right. They're racing against competition. They're racing against their own sense of greatness, which is maybe even the biggest driver here. And I think they're earnestly trying to think about how to do the right thing. All of those are mixed in this bucket, and sometimes you're on one side rather than the other.

Robin Rombach

Yeah, and as you're saying, this is a first time, right?

Andrew Feldman

That's right.

Robin Rombach

When GPT-3.5 came out, it wasn't taking down networks.

Andrew Feldman

Right.

Robin Rombach

But in talking to Nikesh from Palo Alto Networks, I asked him, "Hey, how would you grade this?" And he said, "We put it against our software, and we found bugs we were not aware of."

Andrew Feldman

Yes, it killed them.

Robin Rombach

Yeah, he said, "We had to stop everything we're doing and do patches for 6 weeks."

Andrew Feldman

Right. And that's when you know, right? I mean, Nikesh leads a—maybe the leading security software firm, right? And when it finds, in an hour—

Robin Rombach

[Laughter.]

Andrew Feldman

—tens of critical holes, you're like, "Whoa, this is a powerful tool."

Robin Rombach

Yeah, I mean—

Andrew Feldman

And we need to think. And maybe you show it to a group first, right? Maybe you—I don't know what the right thing is, but—

Robin Rombach

Red-teaming. We've always had that. When you were releasing the new version of an operating system, you know, when you have your iPhone, you can say, "I want to be part of the beta."

Andrew Feldman

That's right.

Robin Rombach

You know.

Andrew Feldman

Right. And there's like 2 other betas that you don't even get the chance to opt into as consumers.

Robin Rombach

Right.

Andrew Feldman

Those ones are for security. Those ones are for making sure you don't lose your data or that deleted data—

Robin Rombach

That's right. Doesn't disappear or leak or—

Andrew Feldman

—get corrupted.

Robin Rombach

Right.

Andrew Feldman

I think we also know that there will be a massive data leak.

Robin Rombach

Of course.

Andrew Feldman

We know this, right? And it's like Warren Buffett talked about the reinsurance industry: You know something bad's going to happen. You don't know when—

Robin Rombach

Yeah.

Andrew Feldman

—but you've got to save up for it. You put money away for insurance.

Robin Rombach

But—

Andrew Feldman

There will be a tornado. There will be a massive earthquake. I mean, we know this. And we can do our best to plan. But there'll be a massive breach. And we have to steel ourselves in advance. And we have to think about it and think about the right response at the time. And so, to prepare ourselves for a future that is, in specifics, unknown, but in general, we're pretty sure something's going to happen.

Robin Rombach

Something—

Andrew Feldman

[Snorts.]

Robin Rombach

—will happen.

Andrew Feldman

And yeah, it's typically a black swan, right? I mean, by definition, it's going to be something we didn't consider or a question we didn't know to ask.

Robin Rombach

Right. But even knowing that there are some unknown unknowns is a useful place to start.

Andrew Feldman

Yeah. What are we not asking ourselves?

Robin Rombach

That's right.

Andrew Feldman

With reasoning, the AI is going to be able to tell us, "Hey, schmuck humans."

Robin Rombach

That's right.

Andrew Feldman

"By the way, here's what you're not thinking about." This is now my closing sentence when I do my prompting: "I need you to make me a prompt that will help me do this trend scouting," for example. And then I always say at the end, "Please check your work, and then tell me what I haven't considered in terms of my goals. And ask me some questions every time you run the job." And that has changed everything, because it's like, "I checked my work. By the way, this was incorrect."

Robin Rombach

Right.

Andrew Feldman

And I'm wondering, "Hey, would you like me to also do this?" Some of the tools, like Perplexity, do that automatically to give you your next 3 prompts.

Robin Rombach

Right.

Andrew Feldman

But if you give it explicit instructions, my lord, is it good at that. So, you know—

Robin Rombach

Over the course of the last 10 years, as I was raising money, I thought one of the smarter questions I got at the end of a conversation was when someone asked, "What was the smartest question you heard that wasn't covered by what I asked?"

Andrew Feldman

Incredible.

Robin Rombach

Right. Now, that's somebody who's curious and thinking and humble and trying to use this to get a picture of—

Andrew Feldman

—of the space. And to the extent that you can ask the AI that and that it can broaden your view, maybe what questions should I have asked?

Robin Rombach

Yeah.

Andrew Feldman

To be an expert in this, what would a PhD-level questioner ask?

Robin Rombach

Right.

Andrew Feldman

Or a gold medal mathematician? I mean, I think those are questions that you don't even know how to ask, which—

Robin Rombach

You know, if we start thinking about AGI and superintelligence, they're just definitions, but they're important definitions, I think, to keep in mind because they're waypoints.

Andrew Feldman

That's right. And AGI—I suspect you'll agree with me that we've hit it. We just haven't exactly deployed it fully. We have artificial general intelligence now. It feels like, when we're talking about these reasoning moments, the ability for it to be as smart as any human, but—

Robin Rombach

By any definition we had 20 years ago, we've hit it.

Andrew Feldman

Yes.

Robin Rombach

Right. I mean, if you think about it, all those Turing tests blew it away. I mean, you think about any period of time—10, 15, 20, 30, 40, 50 years ago—we would have put forward any definition, and it would have previously—

Andrew Feldman

We've blown past it.

Robin Rombach

Which goes back to our previous point: Do we know the questions to ask?

Andrew Feldman

That's right.

Robin Rombach

20 years ago, science fiction authors had their say, and we answered all their questions. If they were to look at this today, they'd be like, “Well, I'm out of questions. Sorry.”

Andrew Feldman

That's where listening to people who sometimes sound like they're on the fringe comes in.

Robin Rombach

Yeah.

Andrew Feldman

When Ilya Sutskever was talking 8 or 10 years ago about the need for safety, you're like, “What?” Dead right.

Robin Rombach

Yeah.

Andrew Feldman

Right, when Elon was talking about building rockets and driving the cost of a launch vehicle to near zero, you're like, “What?” And there it is. Now you can see it. That's why it's really fun to be a technologist now.

Robin Rombach

With these tools specifically, we're talking about building all of these tools, and then the tools are starting to build themselves in this recursive loop.

Andrew Feldman

That's right.

Robin Rombach

We're just starting to see people apply these loops. In fact, loop maxing became a buzzword for me. When I was doing my training, it kept picking up looping, and it kept picking up the maxing stuff, and it created this buzzword for me: loop maxing.

Andrew Feldman

Right.

Robin Rombach

Then, magically, people started talking about loop maxing, and I was like, “Wow, this is really weird.” It anticipated that other humans would come up with this word. But talk a little bit about recursion and the road to superintelligence. Do you have a way, Andrew, that you think about superintelligence, what it will mean for humanity, how we'll define it, and how we'll experience it?

Andrew Feldman

Yeah. I think let's begin with loop maxing, or recursive learning. I think what Sam Altman and Ilya Sutskever, and then later Dario Amodei and Demis Hassabis, saw 6 years ago, or 5 years ago, was that powerful recursive gains are exponential.

Robin Rombach

Mhm.

Andrew Feldman

Right? You get better, you do it again, and if you continue to get gains, the slope of that curve is so steep.

Robin Rombach

Yeah.

Andrew Feldman

We're just beginning to see that now.

Robin Rombach

You ask it a question, you learn from the results, and you ask it to do it again. The results get better, and more information is added. Your answer gets better. You ask it to do it again, and it covers more material. These sorts of loops are producing not a little bit better answers, but vastly better answers.

Andrew Feldman

Yeah. And that is enormously powerful because we don't quite know where it ends. Keep throwing compute at it. How much better does the answer get? We run out of tokens or our budget, but holy cow. When does the exponential stop, or does the answer keep going up and up and up to the right? That's an enormously interesting intellectual question right now.

Robin Rombach

Yeah, like when do we run out of problems to solve?

Andrew Feldman

Well, that's right. And when are the problems no longer intellectual problems, and they're now people problems? Right. How to organize people to get done what the AI asked for. As you know, in running your company, a lot of your problems aren't hard intellectual problems. They're people-working-together problems.

Robin Rombach

Yeah.

Andrew Feldman

Right.

Robin Rombach

And your motivation—

Andrew Feldman

Motivation. You spend a lot of time as a leader spraying WD-40 on your team.

Robin Rombach

Right.

Andrew Feldman

Right. It just reduces friction. So how do we learn about those things from AI? How do we get behavioral insights from AI? I think that's one of the things world models are going to bring us as they begin to watch human behavior.

Robin Rombach

Yeah, we didn't even get to that. This is going to be for another interview, but when these things jump off the screens—

Andrew Feldman

Right.

Robin Rombach

—and they're in the real world, and the recursiveness starts not by trying to solve math problems and humanity's most difficult ones, but by saying, “Hey, there's an incredible world out here, and here's the Palace of Versailles.” You're just like, now we're saying, “Make me a new version of Salesforce,” and we're like, “Hey, you know what? I'd like a Palace of Versailles. I've got 100 acres somewhere out in West Texas or Nevada. I'll just send a thousand optimists out there. Make me the Palace of Versailles.”

Andrew Feldman

Right.

Robin Rombach

Sounds fantastical, but the Palace of Versailles would seem fantastical to people who lived 1,000 years before it.

Andrew Feldman

And it was fantastical, I think, to the people who built it. I think they were awed by it as they built it.

Robin Rombach

Yeah, they're compounding recursive learning—

Andrew Feldman

That's right.

Robin Rombach

—and generations, as we talked about. You had such a great insight: in building this place, you had generations of masons.

Andrew Feldman

I think in all these large projects, often there were families who were specialists, and you apprenticed under your father or your uncle. When you had a project that took 50, 70, or 100 years, you might have 3 or 4 generations of the same family—the same stonemason family—working on the same structure.

Robin Rombach

And passing on the learning.

Andrew Feldman

On the learnings.

Robin Rombach

New innovations.

Andrew Feldman

Right.

Robin Rombach

Which is what we've modeled with these new models and what you're building in the infrastructure. It's pretty incredible when you think about it.

Andrew Feldman

But Versailles—and that's what I mean—I think the problem with human learning is that it often moves at the pace of a generation.

Robin Rombach

Huh.

Andrew Feldman

Like elephants and other large mammals, we don't have generations but every 15 or 20 years. If you want to move really quickly across generations, you want them happening more like Drosophila, like fruit flies. You want 2 a day.

Robin Rombach

Yeah.

Andrew Feldman

And you see that in genetics. That's why we study them in genetics, because with learning encoded in the DNA, you can study it over thousands of generations. I think what we're getting is that equivalent in AI. We're getting learning so quickly over the equivalent of thousands of generations.

Robin Rombach

People would be in awe of this pace of evolution.

Andrew Feldman

That's exactly right. You think about it as—

Robin Rombach

I remember when I was getting my psychology degree, and they were teaching us about paradigms. I was trying to understand how the paradigms shifted, and the professor said to me, “What you have to understand is that paradigms don't die.”

Andrew Feldman

They don't.

Robin Rombach

People do.

Andrew Feldman

That's right. That's how it was with Freud, Skinner, and Jung: it took them dying.

Robin Rombach

That's right. It took the new generation to question it, and that was 20 years.

Andrew Feldman

Yeah, sometimes 40 years.

Robin Rombach

Right. Their students maintained positions of leadership until someone said, “Maybe we can do it differently.” And I think what you're seeing is that this iteration is shortening the intergenerational gap, and the learning is so fast.

It's always so great to talk to you because, one, it's just intellectually—your approach to it is so intellectually rigorous. But also, with so much p(doom) in the world, I feel so good that you're such an optimist about this technology and that you're building it with such thoughtfulness. I think people who are hearing these horror stories about AI and job loss and everything need to understand that there are people like yourself who are building this in an incredibly thoughtful way, and this is going to be a net benefit for humanity that just is unimaginable.

Andrew Feldman

Yeah. We have a shot with this technology so that neither our children nor anyone they know dies of cancer. Right? I mean, say it like that. There will be some dislocation in the economy. There will be. There was dislocation when cars came, and it was a bad deal to be a guy who shod horses or built carriages.

Robin Rombach

Yeah.

Andrew Feldman

But you've got to also, against that, make your tally of the cons and the pros.

Robin Rombach

Yeah.

Andrew Feldman

Right? There's a shot that our children—none of them, nor the people they love—will die of cancer. And that's 1 thing—

Robin Rombach

That's funny.

Andrew Feldman

—that we can work on with this technology, and we will have great purchase on. I think you begin listing those, and then it's a more thoughtful discussion.

Robin Rombach

Yeah. Unlimited energy, unlimited calories, unlimited knowledge, unlimited education, unlimited housing.

Andrew Feldman

And how we do it. We imagine—we know how to teach children, and we don't do it, right? Aristotle was a tutor to Alexander the Great. Socrates was his tutor. We know that if you give a child a tutor and the tutor modifies the teaching for the child, they learn better.

Robin Rombach

Adapt. That's not how we do it, teaching classes.

Andrew Feldman

No, factory farming.

Robin Rombach

That's right.

Andrew Feldman

We teach to some sort of middle level. Imagine if we built agents that taught children for their way of learning.

Robin Rombach

Right.

Andrew Feldman

Right? Here's the way we've been doing it: the same way for 1,000 years. During that entire time, we knew how to do it better, and we chose not to.

Robin Rombach

Yeah.

Andrew Feldman

And here's the way we can do it. Put that on the pro side. As long as we're thoughtfully and fairly writing the good and the bad, I think it'll come out okay.

Robin Rombach

Get out there, Andrew, and keep communicating your version of the world, because some people see around the corner and get a little nervous. Okay, fair enough, but I think the ledger, as you describe it, is heavily weighted toward abundance.

Andrew Feldman

I think it'll create abundance for sure.

Robin Rombach

Abundance. Andrew, pleasure always. I'll see you in 6 months for our checkup. That'll be great.

Andrew Feldman

Robin Rombach is the co-founder and CEO of Black Forest Labs. You are based in Germany, in the Black Forest, which is a city in Germany.

Robin Rombach

A mountain range, actually.

Andrew Feldman

A mountain range.

Robin Rombach

Yes.

Andrew Feldman

Where you grew up.

Robin Rombach

Well, I grew up there.

Andrew Feldman

And you are working on open-source image and video models. You worked at stable diffusion for a little bit. Cut your teeth on that, and you're known for the open-source model FLUX, and maybe also for some closed-source models. Tell us about the business of Black Forest Labs. What is the business, and what is the goal?

Robin Rombach

That's correct. One quick addition: We are based in the Black Forest, in a town called Freiburg, and in San Francisco.

Andrew Feldman

Oh, and in San Francisco, of course, yeah.

Robin Rombach

I'm splitting my time to a certain degree. We started the company 2 years ago. My co-founders and I, as you said, have worked on Stable Diffusion in the past. Before that, we invented an algorithm called latent diffusion, which is basically the fundamental algorithm behind all of the generative models that are being deployed for image generation, video generation, and even physical AI now.

Andrew Feldman

Yeah.

Robin Rombach

It basically makes use of this principle that you can compress natural data, such as images, video, and audio, into a much more efficient representation and then train a transformer model on that. This is the stuff where JPEG, MP3, and all of that works. We basically translated that into a neural algorithm a few years ago, when we were still PhD students in Munich, actually. Then, on top of that, we built Stable Diffusion, and on top of that, the generative models that we are developing today.

Of course, the technology has advanced. But we are now tackling, I would say, models that are really made for understanding the whole world around us: multimodal visual models pretrained on images, video, and audio data at the same time. And we are now entering a new paradigm, which is combining that with something called action prediction, such that you can actually use the same model to make images, make videos, make audio, and predict actions—which means you can ultimately deploy it on a robot in the real world.

Andrew Feldman

Wow. So from the image to the video, the audio, and then eventually the real world with robotics and a real-world model. Because if you can make the image and train the model, that means by default you understand the world. In order to make a video of the world, you have to understand the world, yeah? And the objects in it.

Robin Rombach

I think that's a really good way to think about it. It's an intuitive way to interact with the world, right? I would say there are these complementary forms of intelligence, ultimately. There's intuitive intelligence, and then there's a deep reasoning layer. Ultimately, for a complete form, you need both. And you need them to interact, and I think we've been approaching it more from the intuitive side.

Images are a very natural way to approach this whole field because it's not as computationally intensive as, let's say, video, right? But now, I think we're combining it. It's converging into a more multimodal model, and we see exactly that pretraining on videos gives implicit understanding of the physics of interactions with the real world. Then you can get stuff like action prediction, like robotics, out of the same model.

Andrew Feldman

And with these models and the training, there has kind of been a limitation in creating videos and creating images, where the criticism of generative AI is that it's a bit of a slot machine. I give a prompt, and it gives me something back. But how did it come up with that? The training data. But maybe I want a different style. Maybe I want a different color. Maybe I want a different aesthetic.

Robin Rombach

Yep.

Andrew Feldman

It has that. How does that problem get solved? And do you actually understand what's happening when the image is being made under the hood?

Robin Rombach

Yeah, I think ultimately it's about exposing as many manipulation layers as possible to a user or developer that builds on top of this model, right? And I think we've seen that in the past. Image models basically started from simple text-to-image systems. Then they expanded into text-plus-image-to-image systems, which meant you could suddenly take an image, like a real image or a generated image, and iterate on that based on a text prompt—edit it, modify it, right?

Then this expanded into taking multiple images and a text prompt and combining them in a semantic way and producing new content. The same principle now applies to video, and I think it becomes even more interesting when all of these modalities are combined as inputs and outputs of the same model.

Andrew Feldman

Let's talk about video. There's an announcement that you're working with the greatest director of all time, or living director, Martin Scorsese. We'll talk about that in a second, yeah?

Robin Rombach

Fantastic.

Andrew Feldman

But in a movie, this promise of being able to make a movie in which the camera angle and the sound could be something that Martin Scorsese would be proud to release to his fans—how close are we? Maybe tell us a little bit about this partnership. The technology being able to make an actual movie like Goodfellas, or a scene from Goodfellas, versus where it is today, where you can make interesting 5- or 10-second clips, and then maybe people struggle making 10 of them. Then they use some post-editing software to put them together, but you immediately understand this is not that. It's not a movie. It's AI slop. It's kludgy. It doesn't pass the uncanny valley.

Robin Rombach

Well, I think it's important—and that's at least the view that we have—that these AI models are a medium. We don't want to set any way in which they are supposed to be used. We don't want to tell anyone, especially not someone like Martin Scorsese, how he is supposed to use the model. He's one of the greatest filmmakers ever. It was insane sitting in the same room with him multiple times and actually seeing him explore our models. As one of the core researchers behind it, it was just an insane feeling, right? At the same time, I'm also a big fan.

Andrew Feldman

So you sat in a room with Marty Scorsese and showed him your tools.

Robin Rombach

Exactly, yeah.

Andrew Feldman

And what was his reaction? What did he key off of? What was the thing that he found most inspiring or interesting?

Robin Rombach

I think it was really this idea that he clearly has a vision in his head of a scene or scenery where maybe a new movie will be shot. He's trying to explore that, and we basically looked at the scenery of a village in Eastern Europe somewhere. He was describing it, we saw some outputs, and we iterated on the outputs.

Ultimately, I think—and that's what he said in the end—getting the mental picture of something out of your head and communicating it in a visual way by making these images, or a series of images, just makes it easier to communicate and convey an idea of what is actually in your head. I think that's one of the very interesting and powerful ways to use this technology. And I think ultimately—

Andrew Feldman

It is to get the inspiration, to get the vision out of his head and onto an image.

Robin Rombach

Yeah, language ultimately is a little bit of a lossy communication medium, right? It's also interpreted in different ways, but visual information is so rich. An image or a video has so much signal in it, and it's just another way of communicating. I think that's one of the beautiful things that this technology ultimately enables.

To your question of making full movies with, I don't know, a video-generation model, for example, I'm not sure if that is the ultimate goal. Maybe it's interesting to plug this into some kind of generative workflow and make a very long video, and I think that's really cool to explore. But ultimately, the really interesting use cases come when you have a human in the loop who interacts with it and uses it as a medium. This is at least the perspective that I take. That makes it interesting, and this is most often when the most interesting outputs arrive or are actually being made.

Andrew Feldman

Production-level is so obviously a huge win for GenAI.

Robin Rombach

Parallelize your brainstorming, basically.

Andrew Feldman

Yeah, and I like that: parallelize your brainstorming. They have an analogy for this. They do storyboards. Some of the great directors—Ridley Scott of Aliens and Gladiator—were known for making their own. I also believe Spielberg liked to sketch Raiders of the Lost Ark and some of these. George Lucas was known for collaborating with many amazing artists, even making miniatures and storyboards for the Star Wars franchise. He had those people on full-time helping him with that.

So that's the obvious place to start. But if we look at startups, startups always want to try to figure out how to do something cheaply. People used to make a launch video for their startup for $100,000 or $250,000. So they take their $10 million venture raise and spend $250,000 on a launch video. I've seen it with a lot of the startups I'm investing in now. They'll just spend a week or two working with a director to make a launch video. You've probably seen this trend, yeah? And I'm sure people use FLUX and some of your models for this. Have you seen this?

Robin Rombach

Yeah, of course.

Andrew Feldman

Yeah, what's your take on that? Because that feels like the early stage of storytelling. You're trying to communicate a product or service in a fun, engaging, punchy 30-second or 90-second way, yeah?

Robin Rombach

I think we support this exploration based on these tools, right? Ultimately, it's great to see all different kinds of launch videos and products being built on top of the same kind of base model or the same technology. I think that's what's making it so interesting and also so powerful.

Andrew Feldman

Yeah. And what else are people using the technology for? I understand there's a Bitcoin movie coming out. Instead of using a green screen in this Bitcoin movie, I was talking to Gal Gadot, the woman who played Wonder Woman.

Robin Rombach

Yeah, Gal Gadot. Gal, of course.

Andrew Feldman

I was talking to her at an event. It was the Breakthrough Prize, Yuri Milner's event, and she was telling me she just did a Bitcoin movie. They did it on a soundstage without green screens. All the actors just worked on a soundstage, and all of the scenery behind them was being done by generative AI. That's a real movie. That's a $30 million-budget movie. She said it would have cost $150 million if they had to build sets, and the film would have never been greenlit. Are you starting to see people use that in production, not just in the back end and the animation phase, but actually in production yet with your tools?

Robin Rombach

Yeah, we see some use cases like that in production. I think high-end film production is one of the most demanding use cases, and I'm glad that it's being explored. But I also really want to say that it's important to see that this technology is on a trajectory and it's improving. It's improving rapidly. If I look back at where we started a few years ago, when I was doing my Ph.D. in this field, the only thing that you could do was generate images that were 64 by 64 pixels. Now you can do multi-minute videos at high resolution.

But it's not going to stop there. It's going to continue to improve, and I think then it's going to unlock even more of these high-end use cases. But I think the main thing—

Andrew Feldman

How do we get to that? Yeah.

Robin Rombach

Ultimately, I think you still want to have the tool that enables this human-in-the-loop kind of—

Andrew Feldman

Of course, yeah.

Robin Rombach

—production workflow, right? But when I look at multimodal generative models as a whole, what really excites me is that you can use the same kind of AI model to make a movie and deploy that as a brain on a robot. I think this is so interesting. There are some thoughts around trying that in the digital world, right? Whether, for example, computer use actually works remains to be seen. But the technology is so powerful and so versatile. It's just moving into that. All the talk around world models, world action models, all of that—it's basically all the same. I think that's what's making it so interesting and what I find most exciting.

Andrew Feldman

So do you believe that the technology will be used to analyze—or primarily to analyze—the real world? Here is a video of somebody making a sandwich. Now we have the robot study it and make the sandwich. Or do you think there'll be a lot of synthetic data made that then the robots will just study? Or are they going to, in some way, innately know based on all this massive amount of training data?

Robin Rombach

I think it's a combination of prediction, right? Prediction is a way—you can think about it as simulation, as generation. It's predicting actions, which means you have to understand the input, the visual inputs, in order to actually predict a reasonable next action. And it's about perception. You can only do that if you understand—if you perceive—the content. You can only transform it into a new piece of content, predict an action, or describe what you actually see in that scene.

The combination of all of that is, I think, what's driving it. There's not a single one of them. It's a combination of these tasks.

Andrew Feldman

And what's the best way to get that training data? Do you need to have people put on glasses, get a first-person perspective, and have them put on gloves so you have that fidelity of understanding? “Hey, this glass is moving. I'm pouring this glass. I'm putting ice into it.” Here's how that works—the splashing and the condensation, the water—so I can pick it up and not drop it because it's wet on the outside.

Or is it going to be just, “Hey, take the corpus of YouTube videos, and the robots know exactly what to do because they'll find 1,000 videos of people pouring drinks?”

Robin Rombach

Ultimately, I think you would want to go to a place where you could prompt a robot in context, right? As you can do with a language model. Basically, just tell it, “Hey, go and pick up this glass with the orange juice or whatever it is.” Yeah, exactly. We're not there yet, but I think this is one of the goals.

How these models are deployed currently is that there are a lot of different types of hardware and different robots running in factories, and they all have some different kind of action representation that you need to tune the models toward, right? So in practice, what you do is you have all this visual understanding in the models, and then you need only a little bit—a few hours—of fine-tuning data to adjust the model for that specific task.

I think the goal would be to move away from that toward as much in-context learning as possible. But it is a bit of a research problem.

Andrew Feldman

In-context learning is kind of having a moment right now. We've been discussing it on the podcast a whole bunch recently. People are also talking about sovereignty. You have companies that own incredible IP libraries. I mentioned Star Wars before. Disney owns an incredible library. What would your advice be to a company like Disney? Should they take your open-source software, train their own models, or work with you to train their own models to control it and then say, “Hey, this is our IP”?

They've already made a point of working with ChatGPT and saying, “Hey, you can and cannot use certain characters.” In fact, OpenAI had a relationship with them—that's for sure no longer happening. But they officially licensed some characters for the output. So how do you think about those major IP holders? What's your advice to them? Are you in discussions with them? We know about the Martin Scorsese or tour deal. But how do you think about content libraries?

Robin Rombach

Look, I think the most interesting use cases of this, if you think about content creation, are in generating something, making something that hasn't been there before, right? That's a fundamental, interesting aspect of this technology. When it comes to IP, what we implement, for example, on our public-facing tools is that you cannot generate certain IP with these models, right? I think that's a sensible approach.

And then, yes, we do work with certain IP holders to develop models together with them. Some are based on our open-source models; some are based on our more powerful proprietary models. But I think that is a very attractive value proposition.

Andrew Feldman

What do you think that will look like for consumers in another couple of years? What would potentially happen when you open up Disney+?

Robin Rombach

That's a good question. I'm not in Disney, right? So it's up to them to decide that. But I think we want to enable them to build all kinds of stuff that they envision, and I think we can support them. We can support other companies in that space to integrate the technology in the best possible way.

One of the very interesting angles of it is that it's becoming much faster; it's becoming more interactive.

I can envision a whole bunch of very interesting interactive content creation tools that you could host on Disney+ or elsewhere.

Andrew Feldman

I think the most interesting thing I've seen in this regard is fan films.

Robin Rombach

Right.

Andrew Feldman

So, there was a category before generative AI: fan fiction. People would write their own Star Wars stories. Then there came fan films, where people would dress up as Jedi Knights and record their own films.

And George Lucas said, “As long as you're not doing it commercially, you're not selling it, I give you permission to go make Jedi movies.” And they even released how-tos on how to make a lightsaber, or sound files for how to make a lightsaber sound.

Now, people are taking the stories that haven't been told from the Star Wars universe and recreating them using AI. For the fans, they're becoming quite popular on YouTube. Star Wars Stories Untold is, I think, the biggest one. It's getting millions of views per video already.

And I think that's really the future: letting the customer base pay a licensing fee or pay a fee, maybe rent software, or maybe pay based on the output, and let them be creative with the characters. Let them make their own stories. And you could be in a unique position to empower that.

Robin Rombach

Well, 100%. I think if you find a model that works for the IP owners, but then also can enable the super-creative customization use cases, I think that's great.

And for myself, when I read a book or watch a movie, I have so many ideas about how it could be done differently or what could have happened, right? And it's so nice that you can actually enable people to visualize these ideas.

Andrew Feldman

Yeah, it's going to be incredible. Continued success with it. You have an office in San Francisco. You're hiring people, yeah?

Robin Rombach

We do, yeah. We just—

Andrew Feldman

A bunch of money.

Robin Rombach

We raised a bunch of money. We just crossed 100 people. We're hiring in Germany and in San Francisco.

Andrew Feldman

Fantastic. Who are you looking for? What's the right type of person, the right type of skill?

Robin Rombach

On the one hand, we are always looking for researchers who have experience in large-scale model training, experience in diffusion model training, and flow-matching training.

We're looking for engineers who want to be working with customers to develop these customized physical AI solutions, or, for example, with an IP owner, develop these models jointly with them. We're looking for engineers who have experience in large-scale compute infrastructure, managing that, and making sure that the training runs smoothly, that we maximize our MFU, and all of that.

And we're looking for people who have an interest in getting the technology out there—

Andrew Feldman

Yeah, in the hands of people. The forward deployment of this—there are just so many great ideas and so many great partners for you.

I think you're going to do well with open source specifically. It seems like the corporates really want to have some additional level of control, but they also need the frontier models, or your proprietary ones, for some of those refined features. So, I think you have a very bright future ahead of you.

All right, continued success. Thank you so much for doing the show.

Robin Rombach

My pleasure.

I'm going all in.

[music]

I'm going all in.

Open Source Wins, AGI Is Here, and Scorsese’s AI Toolkit with CEOs of Cerebras & Black Forest Labs | BidClub