开源胜出,AGI 已至:与 Cerebras 和 Black Forest Labs CEO 对谈 Scorsese 的 AI 工具箱
Cerebras 表示,AI 基础设施服务的是已签约需求,而不是押注客户最终会出现。 Andrew Feldman 披露,公司积压订单达250亿美元,并称客户正“试图捕捉昨天的需求”(“trying to capture yesterday’s demand”);占地一个足球场的数据中心,耗电量甚至超过中型城市。对投资者而言,运营难题在于:当需求超过数据中心建设和填充能力时,如何留住客户。
推理模型让推理速度和 token 可用量成为产品能力的一部分。 Rombach 表示,“Fable”和“5/6”等模型越来越能在无需“提示词耳语者”的情况下推断用户意图;他进行 Hermes-agent 实验时,通过一个运行 Z.ai GLM-5.2 的 Bittensor 项目,让模型在行动前先讨论自身的研究策略。长时间运行的任务已经能产出“惊人的东西”;Rombach 假设,如果 Cerebras 的推理速度提升15倍,就能把“数周或数月的思考”压缩到1天内完成。
Cerebras 预计,未来18个月其年轻架构的进步速度将大幅超过传统摩尔定律。 Feldman 的判断是“远超2倍”,因为新架构仍有大量面向特定工作负载的优化空间,而一套已有20年历史、依赖更小制程的 GPU 架构则不同。超大规模云厂商自研芯片,并不意味着商用芯片需求失效:它们的首要目标是避免完全依赖任何单一供应商。
开源模型正成为前沿 AI 底层的成本、控制权与主权层。 Feldman 的比喻是,用户不该“开着 Ferrari 去买菜”:前沿模型负责最难的问题,开源模型则承接日常 G&A 和“剪切粘贴经济”。受监管企业同样希望在本国部署、在本地运行;因此,Cerebras 同时运行 gpt-oss-120B、GLM、Kimmy、“Quincy”系列模型、OpenAI 专有模型以及客户自研模型。
当模型能在数小时内暴露严重软件漏洞时,分阶段发布和红队测试就具备了合理性。 Rombach 转述,Palo Alto Networks 发现了此前未知的漏洞,并暂停其他工作、花费6周打补丁;Feldman 表示,给政府防御方“2或3周修补明显漏洞”是合理的。他同时警告不要条件反射式监管,并认为未来发生大规模数据泄露不可避免:“事情总会发生。”
按20年前采用的任何 AGI 定义,Feldman 都认为门槛已经跨过;真正悬而未决的是递归式改进会在哪里停止。 反复要求模型学习、重试并扩大搜索范围,带来的不是“好一点的答案”,而是“好得多的答案”,把数千轮人类式迭代压缩为机器速度。他的账本中包括真实的就业冲击,但也包括这样的可能性:孩子及其至亲不再因癌症离世,每个孩子都能拥有自适应导师。
Black Forest Labs 将生成式媒体视为多模态世界模型的基础,而不只是图像和视频工具业务。 Robin Rombach 从潜在扩散讲起,进而展望在图像、视频、音频上预训练、再结合动作预测的模型:同一个模型既可以帮助拍电影,也可能成为“机器人的大脑”。短期价值仍由人来主导:Martin Scorsese 用模型把脑中的场景外化;IP 持有者则可以结合可控生成、定制模型,以及潜在的授权粉丝创作。
1. AI 基础设施正在追逐昨天的需求
Feldman 对规模的描述刻意落在物理层面:即将建成的数据中心,耗电量可能超过地球过去50年的总和;单座占地一个足球场的数据中心,接入电力甚至超过中型城市。建设范围覆盖美国、加拿大、北欧、法国、欧洲、中东、哈萨克斯坦、塔吉克斯坦、格鲁吉亚和亚美尼亚。
Cerebras 的积压订单达250亿美元,但 Feldman 表示这并非个例。OpenAI、Anthropic、Google、Microsoft 和 AWS 都希望获得更多数据中心,Rombach 还点名 SpaceX 和 xAI 是永不满足的容量买家。这些客户不是在赌“只要建起来,客户自然会来”,而是在“试图捕捉昨天的需求”。
关于 token 最大化,Feldman 表示,实验确实可能浪费一部分资源,但净价值巨大。他把早期获得 token 的体验比作没有经验的 Costco 顾客:先买下“4件你根本不需要、每件还要22美元的东西”,之后才学会有策略地购物。
企业行为已经从“所有人、想要多少 token 就给多少”转向差异化分配:高生产力团队拿到所需资源,其他工作则交给更便宜的模型或开源模型。Rombach 表示,个人“开始把玩工具,然后工具开始反过来把玩他们”,迫使用户明确目标、系统和需求。
2. 推理把 token 规模变成硬件问题
早期计算机和提示词“只会完全照你说的做”;措辞稍有变化,输出就可能彻底不同。Feldman 表示,如今“Fable”和“5/6”越来越能自行推断用户想要的图表、布局或解决方案,不再要求用户成为“提示词耳语者”——这意味着能力正从总结实质性转向理解意图。
Rombach 的 Hermes-agent 实验通过一个 Bittensor 项目使用 Z.ai 的 GLM-5.2,并拥有无限容量。当被要求成为世界上最好的趋势猎手时,智能体先讨论应该搜索 Hacker News、Reddit、社交媒体还是 Instagram:“你正在看着一个推理模型推导出自己的工作方式。”
无限 token 可能意味着“无限推理”,因为模型可以花25或48小时探索一个问题。Rombach 明确称其为假设的 Cerebras 案例:如果模型以15倍速度持续运行24小时,就可能产生“数周或数月的思考”,使推理延迟成为答案质量的一部分,而不只是用户体验问题。
Feldman 表示,Cerebras 打破了传统的18个月翻倍曲线,并预计未来18个月的改进将“远超2倍”。他的逻辑是:新架构仍保留大幅面向特定工作负载优化的空间,而一套已有20年历史的 GPU 架构更多依赖制程缩小。
3. 开源模型成为企业 AI 的小面包车
超大规模云厂商自研芯片,并不意味着每一款芯片都必须击败商用芯片的前沿水平。Feldman 的解释来自行业记忆:云服务商曾经依赖 Intel,而 GPU 厂商也曾依赖少数超大规模云厂商。“你不可能完全依赖别人的芯片。”
模型路由遵循同样的多元化逻辑:“你不会开着 Ferrari 去买菜。” OpenAI、Anthropic 和 Gemini 可以处理前沿问题,而日常的 Workday 操作、G&A 以及“剪切粘贴经济”,需要的是可靠的开源能力,而不是“奥数金牌级数学”。
主权需求带来了第二个驱动因素。金融、医疗和其他受监管行业可能要求在本国、在本地部署系统,以更严格地控制数据泄露和智能外流;Feldman 认为 gpt-oss-120B 是正确的一步,但美国需要“更多本土开源模型”,为市场提供中国模型之外的选择。
Feldman 表示,Cerebras 能够运行广泛的模型,包括 GLM、Kimmy、“Quincy”系列、闭源 OpenAI 模型、GlaxoSmithKline 内部开发的模型,以及 G42 和 MBZUAI 的模型。他称,政府围绕“Fable”和“56”采取的行动,尤其是在欧洲,已经敲响警钟。
4. 模型找到严重漏洞后,分阶段发布显得合理
被问及强大模型是否应该逐步推出时,Feldman 坦诚承认自己的局限:“我以前没见过这种情况”,因此无法判断那次具体发布是否正确。但原则上,分阶段开放类似于对强效药物采取的预防措施,也能给政府红队“2或3周修补明显漏洞”。
他的平衡点是,极化立场“会损害清晰思考”。美国应维持 Dario、Sam 及同行之间的激烈竞争,避免成为一个第一反应就是监管的地区;同时也要承认,开发者正在没有操作手册的情况下,同时发明技术和配套护栏。
护栏本身会增加延迟。Cerebras 在此前6周的实践中发现,更快的芯片能让这些保护措施“没那么痛苦”,说明安全和性能是相互耦合的,而不是彼此独立的产品层。
Rombach 提供的最有力证据来自 Palo Alto Networks 的 Nikesh:据称,一款经过测试的模型发现了公司此前并不知道的漏洞,迫使其投入6周打补丁。Feldman 预计大规模泄露最终会发生——某种未指明的“未知的未知”——并认为机构应在“事情总会发生”之前准备好应对方案。
5. 旧 AGI 测试已经过时,递归学习才是现实问题
Feldman 认同:“按我们20年前采用的任何定义,我们已经达到 AGI。” 早期的图灵测试和科幻式里程碑都已被超越;更困难的任务,是针对前人没有预料到的能力,提出与之匹配的问题。
递归机制很简单,却可能呈指数级增长:提出问题,从答案中学习,带着更多信息重试,并扩大搜索范围。这些循环带来的“不是好一点的答案,而是好得多的答案”;目前未知的是,改进会在算力、token 或预算限制处停止,还是会继续攀升。
他用大型工程作类比来支撑这一论点:凡尔赛宫等重大建筑,依靠3或4代专业家族积累经验;而人类往往要等20或40年才迎来范式变化,因为“范式不会死,人会死”。AI 则接近果蝇式迭代——“每天2代”——将数千轮等价世代压缩到极短时间内。
Feldman 不否认经济冲击;汽车出现后,马掌匠和马车制造商确实遭遇了打击。他给出的另一端账本是:癌症研究可能让孩子及其至亲免于失去彼此;自适应导师也能为每个孩子提供个性化教学,而不是今天这种“工厂化养殖”式课堂。他认为,只要以“周到且公平”的方式书写,这本账最终可能算得过来。
6. Black Forest Labs 正在构建多模态世界模型
Rombach 将 Black Forest Labs 的基础追溯至潜在扩散:这是他与合作者在慕尼黑读博期间开发的方法。该方法把自然数据压缩成高效表征,原理上类似 JPEG 或 MP3,再在这些表征上训练 Transformer;团队随后以此为基础构建了 Stable Diffusion 和 FLUX。
当前路线图将图像、视频、音频预训练模型与动作预测结合起来。Rombach 区分了直觉式视觉智能和深层推理层:完整智能需要两者兼备,但视频预训练已经能够提供关于物理规律和现实互动的隐性知识。
控制能力也在逐步扩展:从文生图,到文本加图像编辑,再到通过提示词对多张图像进行语义组合。将同一模式延伸至视频、音频和动作后,每种模态都可能成为输入或输出,并向用户和开发者开放更多“操控层”。
7. Scorsese 的用例是沟通,而不是一键拍电影
Rombach 拒绝规定生成式模型应该如何使用:“这些 AI 模型,它们是一种媒介。” 与 Martin Scorsese 合作时,团队围绕一个东欧村庄反复迭代,从导演的描述出发,直到图像比语言更直接地传达他脑中的场景。
关键在机制,而不是明星背书。语言是一种“有损通信媒介”,而图像或视频包含异常丰富的信息;这项技术可以“并行化你的头脑风暴”,延伸熟悉的分镜和微缩模型工作流,而不是取代电影人。
Rombach 不确定生成整部电影是否是最终目标。他倾向于人参与其中的工作流,也不会预测高端电影何时到来;不过能力已经从他读博时的64×64像素图像,发展到高分辨率、数分钟长度的视频。
Robin 提供了生产经济学案例:Gal Gadot 曾描述一部 Bitcoin 电影,影片在摄影棚拍摄、使用生成式场景,据称成本为3000万美元;如果搭建实体布景,成本则需要1.5亿美元。Rombach 确认存在类似的制作应用,但强调技术仍处于快速进化轨道。
8. 同一套视觉技术栈正延伸至机器人、IP 库和消费者
Rombach 最具扩张性的判断是:“同一种 AI 模型”既可以用来拍电影,也可以充当机器人的大脑。生成会变成模拟,动作预测则要求模型具备感知和理解能力。当前,不同工厂硬件仍需要数小时的任务专属微调数据;长期目标则是让机器人能够理解“拿起一个玻璃杯”这样的自然上下文指令。
在知识产权方面,Black Forest Labs 会在公共工具中屏蔽部分 IP 的生成。公司也直接与 IP 持有者合作开发定制模型,有时以开源基础构建,有时则使用更强大的专有系统;对拥有大型内容库的权利方而言,这是一种兼顾控制力与能力的方案。
Robin 提出的消费者模式是授权粉丝创作,把 Star Wars 粉丝小说和粉丝电影延伸到 AI 生成的未述故事。Rombach 认同,如果模型既能获得权利方接受,又能实现“超级创意式定制”,观众就可以把自己已经想象过的替代事件和不同解读视觉化。
公司成立仅2年,员工数刚超过100人。Black Forest Labs 正在德国和旧金山招聘大规模模型与扩散模型研究人员、负责构建实体 AI 或 IP 解决方案的客户工程师,以及专注于稳定训练和最大化 MFU 的基础设施专家。
We are in the race for superintelligence, and Andrew Feldman is back. He's obviously the CEO and founder of Cerebras, which is doing inference chips, pioneered the space, and had a successful IPO. We've talked about this a couple of times. We got to see each other in January at Davos. The IPO happened, and the boys and I got to sit with you recently.
That was fun.
At Liquidity.
That was really fun.
We had a great discussion with the boys, but I wanted to deep dive with you about a couple of topics. The first one is the buildout of AI. We've never seen a buildout like this since the Great Wall of China.
Right.
Who knows—since the pyramids? It feels like the amount of capital, time, and intelligent people on the planet dedicating themselves to the buildout of something—I can't think of anything in our lifetimes, but perhaps before our lifetimes, the war effort.
Right.
This is a mobilization at a scale that we read about and hear about, but you're actually doing it. You have customers who are building data centers, and you're a key piece of that.
AppLovin started with an $8 domain and no VC funding and became one of the largest ad platforms in the world. Now that same engine powers AppLovin ads for e-commerce. Your ads run inside mobile games reaching over a billion people with full screen distraction free attention. The platform finds buyers and optimizes for profit. You set the target, it does the rest. One cookware brand went from $4 million to $16 million turned profitable and is on pace for 80 million this year. Visit applovin.com/allin to launch your first campaign today.
Maybe you could just enlighten us: In 2026, what is Cerebras doing, and what is happening with this buildout in Texas? These are some gigantic efforts.
The size and scope of what is being built—physical size and scope—are usually not what we talk about when we talk about software or hardware. We're talking about chips and boxes, and they don't have the same sort of physical enormity.
Right.
Right. What we're talking about now are data centers that, in the next several years, are going to use more power than the previous 50 years on Earth used.
Wow.
Right. We're talking about individual buildings the size of football fields that have more power coming into them than midsize cities. They're being built across the U.S., in Canada, throughout the Nordics, here in Paris, and throughout France, Europe, and the Middle East—in nations that weren't front and center in anybody's mind previously. Kazakhstan, Tajikistan, Georgia, and Armenia are all building out data centers of size. Everybody's focused on huge data centers.
And every state in America obviously feels it needs to participate in this. The people who are buying the capacity—OpenAI, Anthropic, SpaceX, xAI, and Google—are insatiable right now.
Yeah. How many years out are they building? When you talk to them, they were ordering chips from Cerebras before you were finished with the chips. They're putting orders in ahead of time.
The irony is that, unlike many exciting times in technology, they're trying to capture yesterday's demand. The demand is way outstripping our ability to build data centers and fill them with hardware. We have a $25 billion backlog.
$25 billion backlog.
We are not alone in that. OpenAI and Anthropic—you go through this list. Google wants more data centers, Microsoft wants more data centers, AWS wants more data centers. All of these players are not chasing, "If you build it, they will come." They're chasing demand that is already booked.
Right. How do we keep them from leaving?
Right. That's extremely unusual.
It's very unusual. Now we have people who are token-maxing.
Yeah. You know, what I liken this to is when we first started with AWS, and it was so good to get around your own IT organization that you told every engineer, "Go ahead, put it on your credit card. Sign up."
Yeah.
Right? A lot of it was really useful, and some of it was, "God, I wish we didn't do that."
Yeah.
For sure, there's experimentation. But it doesn't mean that the net value isn't enormous. It means some of it is going to go nowhere.
I remember when Costco opened up in the Palo Alto area in 1988. People used to shop at Costco like they shopped at Safeway. They'd go down every aisle.
Yes.
That's a horrible way to shop at Costco, because you end up with 4 things you didn't need, and each was $22. As people got more accustomed to it, they'd go to the back, get the chicken, get 18 cupcakes for the kids' birthday party, and be out.
Bang, you were out. Strategic. Strategic.
And it's exactly the same. I think at first people opened up and said, "Everybody, as many tokens as you want." In enterprises, there's no open loop. We don't give any resource unconstrained to people. Now we're jumping on and saying, "Whoa, all right. These guys should have as much as they need. They're enormously productive. Over here we can use maybe an open-source model, maybe a cheaper model over here." Now we're running like a business.
We're really seeing a certain type of person emerge who knows how to deploy this technology: systems thinking, which developers tend to have innately. CEOs tend to be great strategists and understand systems.
The intelligence is getting so much better every step along the way that I'm watching individuals—typically startup founders, but also venture capitalists and associates who work at my venture firm—start playing with the tool, and then the tool starts playing with them. They start to go, "Oh, I haven't clearly defined what my goal is. I don't understand what a system is. I've never heard about making a requirements document."
The software is like, "Do you have a requirements document? What's your goal?" The AI starts telling people, "You're token-maxing, and you need to get a little more focused here."
One of my colleagues, 20 years ago—a really smart computer scientist—said, "Computers are really dumb. They do exactly what you tell them."
Yeah. At first, prompting was like that, right? You modified your prompt a little bit, and it changed the answer.
Dramatically.
Dramatically. Increasingly, it's understanding what your intent was.
Right.
Right. If you have a chance to play with Fable or 5/6 from OpenAI, increasingly you don't have to get the prompt just right. You don't have to be a prompt whisperer. Instead, you ask it, and it says, "Well, here's some things. By the way, maybe you wanted the chart to go two ways. You wanted it aligned in a bar." And you're like, "Well, that's exactly what I wanted. I didn't ask for it, but that is better." It's understanding intent, and that's a huge leap.
If we were sitting here 2 years ago, we never would have been able to predict that, in a short 24 months, it would go from being a great summarizer and researcher of web results—
Right.
—to actually understanding your intent, providing a solution, and abstracting it all from you.
That's right. That's right.
Which is a very weird thing. I don't know if you've played with the Hermes agent yet. Have you played with it yet? I asked it just this morning, and I was given a secret Bittensor project that has the new Z.ai model, GLM 5/2.
And they gave me GLM 5/2, yeah.
GLM-5.2. Somebody in Bittensor—I don't think you understand Bittensor. You've heard of it, the distributed crypto project. They have all this extra capacity. A whisperer told me there was probably some capacity in China that has free energy. Okay, fine. They gave me unlimited capacity, so I started having it do some really crazy jobs. I said, "Every hour, I want you to tell me what the trends in the world are that nobody else has identified yet. You can do whatever you want to do to accomplish that, but my goal is to be the smartest trend hunter in the world."
I watched what it was doing in the background, and it started debating itself on where it should find things. It said, "We should probably go to Hacker News and Reddit." Then it was like, "Yeah, but there's also social media, and trends tend to manifest on Instagram."
That's a reasoning model. You were watching a reasoning model work things out.
Yeah.
Isn't that interesting? I mean, that's amazing.
It was collapsed. So, as a civilian—
Right.
—who doesn't get the uncollapsed moment, if you were using ChatGPT 3.5 or 4.8, whatever it was, and you haven't used this new level of reasoning, inference, and essentially unlimited compute—
Right.
—it opened my eyes just this morning to what a world of unlimited tokens might look like.
Right.
Because unlimited tokens, I believe, means unlimited reasoning.
It does. What does that mean?
Yeah. If you run these for 25 or 48 hours, you get amazing things now. What if, by using Cerebras, we were 15 times faster and then you ran it for 24 hours?
Right.
And you got weeks or months’ worth of thinking.
Yeah.
And I mean, it is extraordinary. I think one of the things is that people like Ilya and Sam, in the early days, were saying this was coming.
Right. And I think when you look back, you say to yourself, “Holy crap. Those guys—”
They knew.
Saw it.
Yeah, they could see around the corner.
That’s right. And the rest of us were like, “What? I’m not sure.”
It’s when we had Sam on All-In at one point. He said, “Oh, you know, I’d love to come on at some point.” I said, “Sure, come on.” He was talking about it, and I said, “What’s next?” He said, “Reasoning.” I said, “Unpack that. What does it mean?” He was like, “Well, understanding what your intent was, just as you’re saying, and then figuring out a strategy, and then maybe talking to other agents and other threads about, like, is this the right thing to do, and vetting each other’s work.” And I’m like, “Wow, we have come a long way from ‘guess the next word.’”
Right. Right. Right. Fill the sentence in. Summarize this PDF.
Now, Cerebras is at the center of this because this reasoning is inference.
This reasoning is inference, and it’s computationally intensive.
Right.
Right. And so fast compute makes this sort of work fast and sort of tractable. It doesn’t have to take a huge amount of time to get a good answer. And so it’s exactly the fact that this reasoning consumes a huge amount of tokens internally that allows a blisteringly fast machine like ours—and I brought one—
Oh, you—
I’m never far without one. When one costs half a billion to make, you bring it everywhere with you.
And we were tossing this back and forth at Davos. What’s the model number of this one?
This was in the first 8 or 10.
Got it. So this has a special place.
This has a special place. My wife says it’s like I’m a kid with a dirt bike for his 8th birthday. He’s in his bedroom at night. I carry him with me.
I mean, when you have your next party at the house, I highly recommend just a little hors d’oeuvres—something. I think it’d be a great fit. It would be a great fit if you had some—
That’s right. [laughter]
But what we’re looking at here is the ability to do that reasoning at scale. What is Moore’s law for inference, and for Cerebras? Do you have something internally you discuss as, “We’re going to double this every X time period”?
All chips prior to us in the processor world followed Moore’s law.
Got it.
And we broke it—
Doubling every 18 months.
Doubling about every 18 months.
Got it.
And we crushed it with this chip, and we’ve carved out a whole new trajectory. My view is, in the next 18 months, we’ll be way over 2x.
Interesting.
And so I think that early in an architecture, you have room to do much better than what was traditionally Moore’s law. Now, if you’ve got a 20-year-old architecture like the GPU, it’s much harder.
Right.
You have to rely on things like smaller geometry.
Mhm.
Right, going to the next fab node. But in a newer architecture, you have a huge amount of room still to learn about the work that is being presented and make optimizations that give you huge gains.
How do you run the company? Just being the CEO now in the age of AI—you have $25 billion in demand. You have to deploy at an incredible, blistering pace. You have to hire people. You have to create a road map. I don’t mean to give you a panic attack here. You have to keep up with somebody like OpenAI, which is moving so unbelievably quickly.
Yes.
Right? And they’re competing. You’ve got to keep up.
Right.
Right? Your hardware, your software, your deployments have to keep up with some of the fastest-moving organizations in history.
They’re demanding customers. They are not—
They’re not pushovers, for sure.
Yeah, and also potentially competitors down the road. Look, I think there’s so much demand right now that there is no silicon that will go unused.
Right. But why isn’t OpenAI releasing Jalapeño? Why is Amazon making its own chips? You see this recurring trend. Is it a way to let you know, to let Jensen and NVIDIA know, “Hey, we can do this too, so we need good pricing”? Is it a little bit of a flex that way, or is that the future—that they’re going to be in your business?
No. I think nobody likes being dependent. And I think some of the lessons learned by the hyperscalers of the x86 world is that they were dependent on Intel.
Mhm.
And some of the lessons learned by the GPU makers was that they were dependent on a small number of hyperscalers.
Yeah.
And they wanted more customers. And so they set about to help fund these neoclouds.
Mhm.
And so I think mostly it’s about an opportunity to control at least an important part of your destiny.
Got it.
And I think that’s a very reasonable thing. You don’t have to make the fastest chip. You just can’t be entirely dependent on other people’s chips.
And that dependency has become a hot topic. I’m not sure if you caught the episodes over the last 2 weeks, but we’ve been talking over the last year about open source. I’ve been championing that a lot, just because I was early into open claw and quickly started using Kimmy and was like, “Wait a second. I’m blowing out my claw tokens, but this Kimmy, I can’t tell the difference.” And then we started smart-routing it, and suddenly this open source started to figure out reasoning, and the gap—
Well—
—suddenly closed this year.
Well, you don’t want to take your Ferrari to the grocery store. There are times you want to drive your fun car, and there are times you want to throw the kids in and not worry if there are Cheerios on the floor.
Minivan time.
Right. There’s minivan time. And I think that as the sophistication of the user grows, you’re going to have hard problems, and those are going to be frontier-model problems. They’re going to be OpenAI problems, Anthropic problems, Gemini problems. And behind that, there are going to be a lot of ordinary problems.
Right? I mean, if you think about a company, you know how much time is spent cutting things out of Workday and getting it in a different cell for—
Yeah. Right.
The cutting-and-pasting economy is real. Yeah.
That’s right. And this doesn’t need gold-medal math.
No.
What this needs is rock-solid open-source capabilities.
Yeah.
And if you think about what I mean—well, we’ve been thinking a lot about it in G&A, but a huge amount of G&A, all right, is not invention.
No.
Right? And you may not need the most sophisticated agents for this.
And another card that’s turned over recently is that some folks may have concerns with the ambition of the frontier models, and maybe sharing their data—data leakage and sovereignty of intelligence. They’re saying, “Hey, our company is going to choose—maybe we’re in a regulated industry: finance, health care, HIPAA, FINRA, all kinds of different regulations. We need to have this on-prem—”
Yeah, but on-prem—
—domestically, and we’d like an open-source version where we have a little bit more control.
Yeah.
And I think—are you seeing that now?
I’m seeing that for sure, and I think OpenAI made a good call releasing gpt-oss-120B some months back. That was a good open-source model. But I think in the U.S. we need more domestic open-source models. We need to give the world a choice. Right now, if they want to run open source, it’s gpt-oss-120B or Chinese models.
NVIDIA has some.
NVIDIA has seen the same opportunity to push open-source models. I think giving them more power might be sort of—
Well, I was about to— that was you. You cut me off at the pass. My understanding was Jensen was like, “Hey, we don’t even want to talk about these open-source models we have because our customers—”
Right.
—are now going to be competing with Sam, Dario, Elon, Sergey. Do we want to be in that position?
Right.
But we do need some more champions here, and it’s open source, so people can fork it. But that puts you in a more neutral position.
That’s right. We run GLM, we run Kimmy, we run the Quincy set of models, and we run OpenAI’s models, the closed-source ones. We run models for, say, GlaxoSmithKline, which they wrote and developed. We run models for our partner in the UAE, G42 and MBZUAI, that are their models that they designed. So we have a wide variety.
So sovereignty is a trend.
Sovereignty is a trend, and I think the government's actions with regard to Fable and 56, where they said, “Whoa, whoa. Let’s think, and then we can act.” I think particularly here in Europe, it was a bit of a wake-up call.
And when you saw this going down, there’s a layer of partisanship in our country right now.
It's pretty fervent. Dario is pretty explicitly not part of this administration.
They've been very adversarial. Both sides have admitted that. They're starting to work it out now. So it's hard, I think, for us, not being in the room with these parties, to understand what's partisanship and what's gamesmanship here.
But do you believe that what they released was truly dangerous for cyberwarfare, for cyberattacks? And if you were to rate Dario's communication—not his communication, because he's a very effervescent communicator; I think that's a diplomatic way to say it—but to have a scheduled rollout release, we'll put aside the government's control of it, do you think that is a wise thing for us to do at this point? And do you think there was actually a major threat there?
So what's interesting is I hadn't seen it before.
Mm-hmm.
Right. And I think, if we just step back and say, "Is it reasonable?" I don't know whether this was the right time, but at a time—
Mm-hmm.
—that a model is sufficiently creative in its thinking that it poses a meaningful threat for the government to say, "We'd like you to roll it out in steps"—
Yeah.
Now, this doesn't seem unreasonable to me.
Not at all.
Right? I mean, we do this with powerful pharmaceuticals. We like—I mean, we're certainly not encouraging 7 years of trials and the amount of paperwork and all the garbage that has accrued to the FDA. But with a powerful new technology, it certainly doesn't seem unreasonable to say, "Hey, guys, let's at least do some red-teaming at the government so we know our defenses can block this."
Yeah. Have we checked the infrastructure of the country, like—
NSA? Have we checked the infrastructure of—right? [Clears throat.] And can you give us 2 or 3 weeks to patch any obvious holes that are found?
Right. This doesn't seem to me—
—an unreasonable thing for the government to ask.
Right. But in this very polarized time—
Yeah.
—you put on top of it, "Oh my God, it's President Trump doing it," and then you have to think, "Well, what if it was President AOC or President anybody in between the two extremes?"
I think the polarization hurts a great deal. It hurts clear thinking.
It does.
It hurts clear thinking, and both sides are going to do some dumb things and some really smart things.
Right.
Right. And in fact, what I found is that the people in the government are trying really hard.
The rank and file.
The rank and file are trying really hard. And this is moving fast. And I think that our ability to set aside some of the polarization and say, "How do we do this in a reasonable manner?" I mean, we want Dario and Sam competing like crazy.
100%. It's been awesome to watch.
Yeah.
Right. It's good for the technology. It's good for entrepreneurs to see, even with thousands of people, that this is what you can continue to achieve.
Right.
Right. This is a pain—
—in the ass. It makes them get sharper. Amazon did a better job at that.
Everybody got better because of that. We want that. And we certainly don't want to become a region where the first thing we want is regulation.
Right.
Right. But as it gets more powerful—
The industry really should do a better job of regulating itself, perhaps. And it did seem like they were starting that process, but then the communication was lacking, maybe.
You know, I think not only are they racing hard, but they're inventing this as they go, too.
Yeah. Right. There's not a playbook.
No.
Right. They're inventing that. We say, "Oh, just put on guardrails." Well, they have to design the guardrails.
Sure.
Right. The guardrails have an impact. One of the things that fast does is make the guardrails less painful, and we discovered that in the last 6 weeks.
Yeah.
It is that the very guardrails can add time and make it feel slower, and so fast chips like ours can really help that. But—
Right.
Right. They're racing against competition. They're racing against their own sense of greatness, which is maybe even the biggest driver here. And I think they're earnestly trying to think about how to do the right thing. All of those are mixed in this bucket, and sometimes you're on one side rather than the other.
Yeah, and as you're saying, this is a first time, right?
That's right.
When GPT-3.5 came out, it wasn't taking down networks.
Right.
But in talking to Nikesh from Palo Alto Networks, I asked him, "Hey, how would you grade this?" And he said, "We put it against our software, and we found bugs we were not aware of."
Yes, it killed them.
Yeah, he said, "We had to stop everything we're doing and do patches for 6 weeks."
Right. And that's when you know, right? I mean, Nikesh leads a—maybe the leading security software firm, right? And when it finds, in an hour—
[Laughter.]
—tens of critical holes, you're like, "Whoa, this is a powerful tool."
Yeah, I mean—
And we need to think. And maybe you show it to a group first, right? Maybe you—I don't know what the right thing is, but—
Red-teaming. We've always had that. When you were releasing the new version of an operating system, you know, when you have your iPhone, you can say, "I want to be part of the beta."
That's right.
You know.
Right. And there's like 2 other betas that you don't even get the chance to opt into as consumers.
Right.
Those ones are for security. Those ones are for making sure you don't lose your data or that deleted data—
That's right. Doesn't disappear or leak or—
—get corrupted.
Right.
I think we also know that there will be a massive data leak.
Of course.
We know this, right? And it's like Warren Buffett talked about the reinsurance industry: You know something bad's going to happen. You don't know when—
Yeah.
—but you've got to save up for it. You put money away for insurance.
But—
There will be a tornado. There will be a massive earthquake. I mean, we know this. And we can do our best to plan. But there'll be a massive breach. And we have to steel ourselves in advance. And we have to think about it and think about the right response at the time. And so, to prepare ourselves for a future that is, in specifics, unknown, but in general, we're pretty sure something's going to happen.
Something—
[Snorts.]
—will happen.
And yeah, it's typically a black swan, right? I mean, by definition, it's going to be something we didn't consider or a question we didn't know to ask.
Right. But even knowing that there are some unknown unknowns is a useful place to start.
Yeah. What are we not asking ourselves?
That's right.
With reasoning, the AI is going to be able to tell us, "Hey, schmuck humans."
That's right.
"By the way, here's what you're not thinking about." This is now my closing sentence when I do my prompting: "I need you to make me a prompt that will help me do this trend scouting," for example. And then I always say at the end, "Please check your work, and then tell me what I haven't considered in terms of my goals. And ask me some questions every time you run the job." And that has changed everything, because it's like, "I checked my work. By the way, this was incorrect."
Right.
And I'm wondering, "Hey, would you like me to also do this?" Some of the tools, like Perplexity, do that automatically to give you your next 3 prompts.
Right.
But if you give it explicit instructions, my lord, is it good at that. So, you know—
Over the course of the last 10 years, as I was raising money, I thought one of the smarter questions I got at the end of a conversation was when someone asked, "What was the smartest question you heard that wasn't covered by what I asked?"
Incredible.
Right. Now, that's somebody who's curious and thinking and humble and trying to use this to get a picture of—
—of the space. And to the extent that you can ask the AI that and that it can broaden your view, maybe what questions should I have asked?
Yeah.
To be an expert in this, what would a PhD-level questioner ask?
Right.
Or a gold medal mathematician? I mean, I think those are questions that you don't even know how to ask, which—
You know, if we start thinking about AGI and superintelligence, they're just definitions, but they're important definitions, I think, to keep in mind because they're waypoints.
That's right. And AGI—I suspect you'll agree with me that we've hit it. We just haven't exactly deployed it fully. We have artificial general intelligence now. It feels like, when we're talking about these reasoning moments, the ability for it to be as smart as any human, but—
By any definition we had 20 years ago, we've hit it.
Yes.
Right. I mean, if you think about it, all those Turing tests blew it away. I mean, you think about any period of time—10, 15, 20, 30, 40, 50 years ago—we would have put forward any definition, and it would have previously—
We've blown past it.
Which goes back to our previous point: Do we know the questions to ask?
That's right.
20 years ago, science fiction authors had their say, and we answered all their questions. If they were to look at this today, they'd be like, “Well, I'm out of questions. Sorry.”
That's where listening to people who sometimes sound like they're on the fringe comes in.
Yeah.
When Ilya Sutskever was talking 8 or 10 years ago about the need for safety, you're like, “What?” Dead right.
Yeah.
Right, when Elon was talking about building rockets and driving the cost of a launch vehicle to near zero, you're like, “What?” And there it is. Now you can see it. That's why it's really fun to be a technologist now.
With these tools specifically, we're talking about building all of these tools, and then the tools are starting to build themselves in this recursive loop.
That's right.
We're just starting to see people apply these loops. In fact, loop maxing became a buzzword for me. When I was doing my training, it kept picking up looping, and it kept picking up the maxing stuff, and it created this buzzword for me: loop maxing.
Right.
Then, magically, people started talking about loop maxing, and I was like, “Wow, this is really weird.” It anticipated that other humans would come up with this word. But talk a little bit about recursion and the road to superintelligence. Do you have a way, Andrew, that you think about superintelligence, what it will mean for humanity, how we'll define it, and how we'll experience it?
Yeah. I think let's begin with loop maxing, or recursive learning. I think what Sam Altman and Ilya Sutskever, and then later Dario Amodei and Demis Hassabis, saw 6 years ago, or 5 years ago, was that powerful recursive gains are exponential.
Mhm.
Right? You get better, you do it again, and if you continue to get gains, the slope of that curve is so steep.
Yeah.
We're just beginning to see that now.
You ask it a question, you learn from the results, and you ask it to do it again. The results get better, and more information is added. Your answer gets better. You ask it to do it again, and it covers more material. These sorts of loops are producing not a little bit better answers, but vastly better answers.
Yeah. And that is enormously powerful because we don't quite know where it ends. Keep throwing compute at it. How much better does the answer get? We run out of tokens or our budget, but holy cow. When does the exponential stop, or does the answer keep going up and up and up to the right? That's an enormously interesting intellectual question right now.
Yeah, like when do we run out of problems to solve?
Well, that's right. And when are the problems no longer intellectual problems, and they're now people problems? Right. How to organize people to get done what the AI asked for. As you know, in running your company, a lot of your problems aren't hard intellectual problems. They're people-working-together problems.
Yeah.
Right.
And your motivation—
Motivation. You spend a lot of time as a leader spraying WD-40 on your team.
Right.
Right. It just reduces friction. So how do we learn about those things from AI? How do we get behavioral insights from AI? I think that's one of the things world models are going to bring us as they begin to watch human behavior.
Yeah, we didn't even get to that. This is going to be for another interview, but when these things jump off the screens—
Right.
—and they're in the real world, and the recursiveness starts not by trying to solve math problems and humanity's most difficult ones, but by saying, “Hey, there's an incredible world out here, and here's the Palace of Versailles.” You're just like, now we're saying, “Make me a new version of Salesforce,” and we're like, “Hey, you know what? I'd like a Palace of Versailles. I've got 100 acres somewhere out in West Texas or Nevada. I'll just send a thousand optimists out there. Make me the Palace of Versailles.”
Right.
Sounds fantastical, but the Palace of Versailles would seem fantastical to people who lived 1,000 years before it.
And it was fantastical, I think, to the people who built it. I think they were awed by it as they built it.
Yeah, they're compounding recursive learning—
That's right.
—and generations, as we talked about. You had such a great insight: in building this place, you had generations of masons.
I think in all these large projects, often there were families who were specialists, and you apprenticed under your father or your uncle. When you had a project that took 50, 70, or 100 years, you might have 3 or 4 generations of the same family—the same stonemason family—working on the same structure.
And passing on the learning.
On the learnings.
New innovations.
Right.
Which is what we've modeled with these new models and what you're building in the infrastructure. It's pretty incredible when you think about it.
But Versailles—and that's what I mean—I think the problem with human learning is that it often moves at the pace of a generation.
Huh.
Like elephants and other large mammals, we don't have generations but every 15 or 20 years. If you want to move really quickly across generations, you want them happening more like Drosophila, like fruit flies. You want 2 a day.
Yeah.
And you see that in genetics. That's why we study them in genetics, because with learning encoded in the DNA, you can study it over thousands of generations. I think what we're getting is that equivalent in AI. We're getting learning so quickly over the equivalent of thousands of generations.
People would be in awe of this pace of evolution.
That's exactly right. You think about it as—
I remember when I was getting my psychology degree, and they were teaching us about paradigms. I was trying to understand how the paradigms shifted, and the professor said to me, “What you have to understand is that paradigms don't die.”
They don't.
People do.
That's right. That's how it was with Freud, Skinner, and Jung: it took them dying.
That's right. It took the new generation to question it, and that was 20 years.
Yeah, sometimes 40 years.
Right. Their students maintained positions of leadership until someone said, “Maybe we can do it differently.” And I think what you're seeing is that this iteration is shortening the intergenerational gap, and the learning is so fast.
It's always so great to talk to you because, one, it's just intellectually—your approach to it is so intellectually rigorous. But also, with so much p(doom) in the world, I feel so good that you're such an optimist about this technology and that you're building it with such thoughtfulness. I think people who are hearing these horror stories about AI and job loss and everything need to understand that there are people like yourself who are building this in an incredibly thoughtful way, and this is going to be a net benefit for humanity that just is unimaginable.
Yeah. We have a shot with this technology so that neither our children nor anyone they know dies of cancer. Right? I mean, say it like that. There will be some dislocation in the economy. There will be. There was dislocation when cars came, and it was a bad deal to be a guy who shod horses or built carriages.
Yeah.
But you've got to also, against that, make your tally of the cons and the pros.
Yeah.
Right? There's a shot that our children—none of them, nor the people they love—will die of cancer. And that's 1 thing—
That's funny.
—that we can work on with this technology, and we will have great purchase on. I think you begin listing those, and then it's a more thoughtful discussion.
Yeah. Unlimited energy, unlimited calories, unlimited knowledge, unlimited education, unlimited housing.
And how we do it. We imagine—we know how to teach children, and we don't do it, right? Aristotle was a tutor to Alexander the Great. Socrates was his tutor. We know that if you give a child a tutor and the tutor modifies the teaching for the child, they learn better.
Adapt. That's not how we do it, teaching classes.
No, factory farming.
That's right.
We teach to some sort of middle level. Imagine if we built agents that taught children for their way of learning.
Right.
Right? Here's the way we've been doing it: the same way for 1,000 years. During that entire time, we knew how to do it better, and we chose not to.
Yeah.
And here's the way we can do it. Put that on the pro side. As long as we're thoughtfully and fairly writing the good and the bad, I think it'll come out okay.
Get out there, Andrew, and keep communicating your version of the world, because some people see around the corner and get a little nervous. Okay, fair enough, but I think the ledger, as you describe it, is heavily weighted toward abundance.
I think it'll create abundance for sure.
Abundance. Andrew, pleasure always. I'll see you in 6 months for our checkup. That'll be great.
Robin Rombach is the co-founder and CEO of Black Forest Labs. You are based in Germany, in the Black Forest, which is a city in Germany.
A mountain range, actually.
A mountain range.
Yes.
Where you grew up.
Well, I grew up there.
And you are working on open-source image and video models. You worked at stable diffusion for a little bit. Cut your teeth on that, and you're known for the open-source model FLUX, and maybe also for some closed-source models. Tell us about the business of Black Forest Labs. What is the business, and what is the goal?
That's correct. One quick addition: We are based in the Black Forest, in a town called Freiburg, and in San Francisco.
Oh, and in San Francisco, of course, yeah.
I'm splitting my time to a certain degree. We started the company 2 years ago. My co-founders and I, as you said, have worked on Stable Diffusion in the past. Before that, we invented an algorithm called latent diffusion, which is basically the fundamental algorithm behind all of the generative models that are being deployed for image generation, video generation, and even physical AI now.
Yeah.
It basically makes use of this principle that you can compress natural data, such as images, video, and audio, into a much more efficient representation and then train a transformer model on that. This is the stuff where JPEG, MP3, and all of that works. We basically translated that into a neural algorithm a few years ago, when we were still PhD students in Munich, actually. Then, on top of that, we built Stable Diffusion, and on top of that, the generative models that we are developing today.
Of course, the technology has advanced. But we are now tackling, I would say, models that are really made for understanding the whole world around us: multimodal visual models pretrained on images, video, and audio data at the same time. And we are now entering a new paradigm, which is combining that with something called action prediction, such that you can actually use the same model to make images, make videos, make audio, and predict actions—which means you can ultimately deploy it on a robot in the real world.
Wow. So from the image to the video, the audio, and then eventually the real world with robotics and a real-world model. Because if you can make the image and train the model, that means by default you understand the world. In order to make a video of the world, you have to understand the world, yeah? And the objects in it.
I think that's a really good way to think about it. It's an intuitive way to interact with the world, right? I would say there are these complementary forms of intelligence, ultimately. There's intuitive intelligence, and then there's a deep reasoning layer. Ultimately, for a complete form, you need both. And you need them to interact, and I think we've been approaching it more from the intuitive side.
Images are a very natural way to approach this whole field because it's not as computationally intensive as, let's say, video, right? But now, I think we're combining it. It's converging into a more multimodal model, and we see exactly that pretraining on videos gives implicit understanding of the physics of interactions with the real world. Then you can get stuff like action prediction, like robotics, out of the same model.
And with these models and the training, there has kind of been a limitation in creating videos and creating images, where the criticism of generative AI is that it's a bit of a slot machine. I give a prompt, and it gives me something back. But how did it come up with that? The training data. But maybe I want a different style. Maybe I want a different color. Maybe I want a different aesthetic.
Yep.
It has that. How does that problem get solved? And do you actually understand what's happening when the image is being made under the hood?
Yeah, I think ultimately it's about exposing as many manipulation layers as possible to a user or developer that builds on top of this model, right? And I think we've seen that in the past. Image models basically started from simple text-to-image systems. Then they expanded into text-plus-image-to-image systems, which meant you could suddenly take an image, like a real image or a generated image, and iterate on that based on a text prompt—edit it, modify it, right?
Then this expanded into taking multiple images and a text prompt and combining them in a semantic way and producing new content. The same principle now applies to video, and I think it becomes even more interesting when all of these modalities are combined as inputs and outputs of the same model.
Let's talk about video. There's an announcement that you're working with the greatest director of all time, or living director, Martin Scorsese. We'll talk about that in a second, yeah?
Fantastic.
But in a movie, this promise of being able to make a movie in which the camera angle and the sound could be something that Martin Scorsese would be proud to release to his fans—how close are we? Maybe tell us a little bit about this partnership. The technology being able to make an actual movie like Goodfellas, or a scene from Goodfellas, versus where it is today, where you can make interesting 5- or 10-second clips, and then maybe people struggle making 10 of them. Then they use some post-editing software to put them together, but you immediately understand this is not that. It's not a movie. It's AI slop. It's kludgy. It doesn't pass the uncanny valley.
Well, I think it's important—and that's at least the view that we have—that these AI models are a medium. We don't want to set any way in which they are supposed to be used. We don't want to tell anyone, especially not someone like Martin Scorsese, how he is supposed to use the model. He's one of the greatest filmmakers ever. It was insane sitting in the same room with him multiple times and actually seeing him explore our models. As one of the core researchers behind it, it was just an insane feeling, right? At the same time, I'm also a big fan.
So you sat in a room with Marty Scorsese and showed him your tools.
Exactly, yeah.
And what was his reaction? What did he key off of? What was the thing that he found most inspiring or interesting?
I think it was really this idea that he clearly has a vision in his head of a scene or scenery where maybe a new movie will be shot. He's trying to explore that, and we basically looked at the scenery of a village in Eastern Europe somewhere. He was describing it, we saw some outputs, and we iterated on the outputs.
Ultimately, I think—and that's what he said in the end—getting the mental picture of something out of your head and communicating it in a visual way by making these images, or a series of images, just makes it easier to communicate and convey an idea of what is actually in your head. I think that's one of the very interesting and powerful ways to use this technology. And I think ultimately—
It is to get the inspiration, to get the vision out of his head and onto an image.
Yeah, language ultimately is a little bit of a lossy communication medium, right? It's also interpreted in different ways, but visual information is so rich. An image or a video has so much signal in it, and it's just another way of communicating. I think that's one of the beautiful things that this technology ultimately enables.
To your question of making full movies with, I don't know, a video-generation model, for example, I'm not sure if that is the ultimate goal. Maybe it's interesting to plug this into some kind of generative workflow and make a very long video, and I think that's really cool to explore. But ultimately, the really interesting use cases come when you have a human in the loop who interacts with it and uses it as a medium. This is at least the perspective that I take. That makes it interesting, and this is most often when the most interesting outputs arrive or are actually being made.
Production-level is so obviously a huge win for GenAI.
Parallelize your brainstorming, basically.
Yeah, and I like that: parallelize your brainstorming. They have an analogy for this. They do storyboards. Some of the great directors—Ridley Scott of Aliens and Gladiator—were known for making their own. I also believe Spielberg liked to sketch Raiders of the Lost Ark and some of these. George Lucas was known for collaborating with many amazing artists, even making miniatures and storyboards for the Star Wars franchise. He had those people on full-time helping him with that.
So that's the obvious place to start. But if we look at startups, startups always want to try to figure out how to do something cheaply. People used to make a launch video for their startup for $100,000 or $250,000. So they take their $10 million venture raise and spend $250,000 on a launch video. I've seen it with a lot of the startups I'm investing in now. They'll just spend a week or two working with a director to make a launch video. You've probably seen this trend, yeah? And I'm sure people use FLUX and some of your models for this. Have you seen this?
Yeah, of course.
Yeah, what's your take on that? Because that feels like the early stage of storytelling. You're trying to communicate a product or service in a fun, engaging, punchy 30-second or 90-second way, yeah?
I think we support this exploration based on these tools, right? Ultimately, it's great to see all different kinds of launch videos and products being built on top of the same kind of base model or the same technology. I think that's what's making it so interesting and also so powerful.
Yeah. And what else are people using the technology for? I understand there's a Bitcoin movie coming out. Instead of using a green screen in this Bitcoin movie, I was talking to Gal Gadot, the woman who played Wonder Woman.
Yeah, Gal Gadot. Gal, of course.
I was talking to her at an event. It was the Breakthrough Prize, Yuri Milner's event, and she was telling me she just did a Bitcoin movie. They did it on a soundstage without green screens. All the actors just worked on a soundstage, and all of the scenery behind them was being done by generative AI. That's a real movie. That's a $30 million-budget movie. She said it would have cost $150 million if they had to build sets, and the film would have never been greenlit. Are you starting to see people use that in production, not just in the back end and the animation phase, but actually in production yet with your tools?
Yeah, we see some use cases like that in production. I think high-end film production is one of the most demanding use cases, and I'm glad that it's being explored. But I also really want to say that it's important to see that this technology is on a trajectory and it's improving. It's improving rapidly. If I look back at where we started a few years ago, when I was doing my Ph.D. in this field, the only thing that you could do was generate images that were 64 by 64 pixels. Now you can do multi-minute videos at high resolution.
But it's not going to stop there. It's going to continue to improve, and I think then it's going to unlock even more of these high-end use cases. But I think the main thing—
How do we get to that? Yeah.
Ultimately, I think you still want to have the tool that enables this human-in-the-loop kind of—
Of course, yeah.
—production workflow, right? But when I look at multimodal generative models as a whole, what really excites me is that you can use the same kind of AI model to make a movie and deploy that as a brain on a robot. I think this is so interesting. There are some thoughts around trying that in the digital world, right? Whether, for example, computer use actually works remains to be seen. But the technology is so powerful and so versatile. It's just moving into that. All the talk around world models, world action models, all of that—it's basically all the same. I think that's what's making it so interesting and what I find most exciting.
So do you believe that the technology will be used to analyze—or primarily to analyze—the real world? Here is a video of somebody making a sandwich. Now we have the robot study it and make the sandwich. Or do you think there'll be a lot of synthetic data made that then the robots will just study? Or are they going to, in some way, innately know based on all this massive amount of training data?
I think it's a combination of prediction, right? Prediction is a way—you can think about it as simulation, as generation. It's predicting actions, which means you have to understand the input, the visual inputs, in order to actually predict a reasonable next action. And it's about perception. You can only do that if you understand—if you perceive—the content. You can only transform it into a new piece of content, predict an action, or describe what you actually see in that scene.
The combination of all of that is, I think, what's driving it. There's not a single one of them. It's a combination of these tasks.
And what's the best way to get that training data? Do you need to have people put on glasses, get a first-person perspective, and have them put on gloves so you have that fidelity of understanding? “Hey, this glass is moving. I'm pouring this glass. I'm putting ice into it.” Here's how that works—the splashing and the condensation, the water—so I can pick it up and not drop it because it's wet on the outside.
Or is it going to be just, “Hey, take the corpus of YouTube videos, and the robots know exactly what to do because they'll find 1,000 videos of people pouring drinks?”
Ultimately, I think you would want to go to a place where you could prompt a robot in context, right? As you can do with a language model. Basically, just tell it, “Hey, go and pick up this glass with the orange juice or whatever it is.” Yeah, exactly. We're not there yet, but I think this is one of the goals.
How these models are deployed currently is that there are a lot of different types of hardware and different robots running in factories, and they all have some different kind of action representation that you need to tune the models toward, right? So in practice, what you do is you have all this visual understanding in the models, and then you need only a little bit—a few hours—of fine-tuning data to adjust the model for that specific task.
I think the goal would be to move away from that toward as much in-context learning as possible. But it is a bit of a research problem.
In-context learning is kind of having a moment right now. We've been discussing it on the podcast a whole bunch recently. People are also talking about sovereignty. You have companies that own incredible IP libraries. I mentioned Star Wars before. Disney owns an incredible library. What would your advice be to a company like Disney? Should they take your open-source software, train their own models, or work with you to train their own models to control it and then say, “Hey, this is our IP”?
They've already made a point of working with ChatGPT and saying, “Hey, you can and cannot use certain characters.” In fact, OpenAI had a relationship with them—that's for sure no longer happening. But they officially licensed some characters for the output. So how do you think about those major IP holders? What's your advice to them? Are you in discussions with them? We know about the Martin Scorsese or tour deal. But how do you think about content libraries?
Look, I think the most interesting use cases of this, if you think about content creation, are in generating something, making something that hasn't been there before, right? That's a fundamental, interesting aspect of this technology. When it comes to IP, what we implement, for example, on our public-facing tools is that you cannot generate certain IP with these models, right? I think that's a sensible approach.
And then, yes, we do work with certain IP holders to develop models together with them. Some are based on our open-source models; some are based on our more powerful proprietary models. But I think that is a very attractive value proposition.
What do you think that will look like for consumers in another couple of years? What would potentially happen when you open up Disney+?
That's a good question. I'm not in Disney, right? So it's up to them to decide that. But I think we want to enable them to build all kinds of stuff that they envision, and I think we can support them. We can support other companies in that space to integrate the technology in the best possible way.
One of the very interesting angles of it is that it's becoming much faster; it's becoming more interactive.
I can envision a whole bunch of very interesting interactive content creation tools that you could host on Disney+ or elsewhere.
I think the most interesting thing I've seen in this regard is fan films.
Right.
So, there was a category before generative AI: fan fiction. People would write their own Star Wars stories. Then there came fan films, where people would dress up as Jedi Knights and record their own films.
And George Lucas said, “As long as you're not doing it commercially, you're not selling it, I give you permission to go make Jedi movies.” And they even released how-tos on how to make a lightsaber, or sound files for how to make a lightsaber sound.
Now, people are taking the stories that haven't been told from the Star Wars universe and recreating them using AI. For the fans, they're becoming quite popular on YouTube. Star Wars Stories Untold is, I think, the biggest one. It's getting millions of views per video already.
And I think that's really the future: letting the customer base pay a licensing fee or pay a fee, maybe rent software, or maybe pay based on the output, and let them be creative with the characters. Let them make their own stories. And you could be in a unique position to empower that.
Well, 100%. I think if you find a model that works for the IP owners, but then also can enable the super-creative customization use cases, I think that's great.
And for myself, when I read a book or watch a movie, I have so many ideas about how it could be done differently or what could have happened, right? And it's so nice that you can actually enable people to visualize these ideas.
Yeah, it's going to be incredible. Continued success with it. You have an office in San Francisco. You're hiring people, yeah?
We do, yeah. We just—
A bunch of money.
We raised a bunch of money. We just crossed 100 people. We're hiring in Germany and in San Francisco.
Fantastic. Who are you looking for? What's the right type of person, the right type of skill?
On the one hand, we are always looking for researchers who have experience in large-scale model training, experience in diffusion model training, and flow-matching training.
We're looking for engineers who want to be working with customers to develop these customized physical AI solutions, or, for example, with an IP owner, develop these models jointly with them. We're looking for engineers who have experience in large-scale compute infrastructure, managing that, and making sure that the training runs smoothly, that we maximize our MFU, and all of that.
And we're looking for people who have an interest in getting the technology out there—
Yeah, in the hands of people. The forward deployment of this—there are just so many great ideas and so many great partners for you.
I think you're going to do well with open source specifically. It seems like the corporates really want to have some additional level of control, but they also need the frontier models, or your proprietary ones, for some of those refined features. So, I think you have a very bright future ahead of you.
All right, continued success. Thank you so much for doing the show.
My pleasure.
I'm going all in.
[music]
I'm going all in.