[BidClub_]
The a16z Show · · 54 分钟

Google DeepMind Developers:Nano Banana 是如何打造出来的

Oliver WangNicole Brichtova

YouTube
TL;DR
  • Nano Banana 将 Gemini 的对话式多模态智能与 Imagen 的视觉质量结合起来,一举破圈。 在 LMArena 的流量反复超过基于既有模型预估的查询容量预算之前,团队并未预料到它会走红——即使用户只有一定概率拿到 Nano Banana。它释放出的实用性信号异常清晰:用户“会想方设法”去访问这个模型。

  • 保留身份的编辑跨过了实用性门槛,让个性化真正具备吸引力,并支撑起叙事、广告和视频工作流。 过去,一张参考照片无法在零样本条件下生成足够可辨认的肖像,往往需要多张图片、LoRA 微调、部署和等待;现在,这些步骤都不再必要。随后,角色一致性成为倍增器:当同一个人或物体能够在不同画面中持续存在,创作者就能构建故事、漫画和分镜,最终甚至拍出电影。

  • 生产力叙事正在从“90%的时间用于编辑”反转为“90%的时间用于创作”,但用户会选择不同程度的自主权。 Oliver Wang 认为,专业人士可能把繁琐操作交给模型,同时保留意图和控制权;消费者可能只想制作一张家庭图片;知识工作者则可能把整套演示文稿交给智能体。关键的产品问题在于,用户究竟想“不断调整并与模型协作”,还是只需说明最终结果。

  • 界面市场很可能分化为对话式工具、庞大的专业消费者中间层,以及高度技术化的专业工作流。 Chat 对普通用户有效,是因为无需学习新的 UI;ComfyUI 式节点图则能让专家把模型组合成稳定、数百步的流水线。真正的机会位于两者之间:比聊天机器人拥有更多控制力,但不需要学习“100种东西”。

  • 没有单一模型能够满足所有创作目标,因为遵循指令与自由构思之间可能相互冲突。 一个被优化为完全照办的模型,对于希望它“放开来玩”的用户可能反而更差;而像素级专业用户仍会要求明确控制。这为多模型、编排层、垂直界面,以及栅格、矢量和代码混合表示方式留下了持久空间。

  • 最大的扩张机会可能不是娱乐,而是事实性视觉推理。 模型已经能够处理几何题、根据 HTML 图片渲染网页,并推断学术图表中缺失的结果;更长期的愿景是个性化、多语言教材和“视觉深度研究”。只有当文字渲染、事实性、比例图表和最差情况下的可靠性得到改善,这个市场才会真正打开,而不只是靠精选出来的惊艳图片。

  • 下一个基准是质量下限:“我们现在处在挑柠檬阶段。” 每个领先模型都能生成一张完美的精选图片,因此差异化正转向10秒级迭代、对150页品牌指南等长文档的长上下文遵循,以及交付前在推理阶段进行自我批评。抬高最差输出的质量,可能让图像生成从偶尔使用的创意工具,变成教育、品牌和生产力场景的可靠基础设施。

摘要 · 为研究而整理的核心内容

1. Nano Banana 融合 Gemini 智能与 Imagen 质量

  • Oliver Wang 回溯了 Nano Banana 的技术谱系:一端是 Imagen 家族,另一端是 Gemini 内部更早的图像生成能力。随着团队将重心转向交互式、对话式编辑,双方最终合并,打造出后来被称为 Nano Banana 的模型——它也被称为 Gemini 2.5 Flash Image,但所有人都同意,Nano Banana 这个昵称“酷得多”,也更容易说出口。

  • Nicole Brichtova 将模型形容为“集两者之长”:Imagen 曾提供顶级视觉质量,Gemini 2.0 Flash 则展示了同时生成文字与图像、讲故事,以及通过对话完成编辑的魔力。但 Gemini 2.0 Flash 的视觉质量还未达到团队期望;Nano Banana 把这种交互模式与 Imagen 的保真度结合了起来。

  • Wang 没有预料到产品会病毒式传播。转折发生在 LMArena:团队按照既有模型的水平配置查询容量,但需求不断攀升,他们“不得不一再上调这个数字”。用户愿意反复访问一个只有一定概率提供 Nano Banana 的网站,让稀缺性发挥了类似可变奖励机制的作用。

  • Brichtova 的内部突破带有强烈的个人色彩:只需放入一张照片,她就能被可信地放进童年梦想中,比如成为宇航员或走上红毯。过去,要实现相近的肖像还原,需要多张图片、LoRA 微调、时间,以及承载结果的部署环境;这次则以零样本方式完成。很快,内部演示文稿“铺满了我的脸”,随后又扩展到配偶、孩子、狗,以及1980年代风格的变装。

2. 创作收益是夺回时间,而不是自动化审美

  • Wang 对专业创作者的判断很具体:模型可以让创作者从“90%的时间”都花在编辑和繁琐手工操作上,转向将90%的时间用于创作。他预计这会带来“创意大爆发”,但也承认,不同任务和用户对模型自主程度的需求存在巨大差异。

  • 他的光谱一端,是家长制作一张万圣节服装图片;另一端,是智能体完成整套演示文稿。作为前咨询顾问,Wang 记得过去需要花几个小时让幻灯片变得美观、连贯;智能体可以完成故事线布局并生成合适的视觉内容,除非用户本身就想参与协作和迭代。

  • Nicole 先把艺术视为一个分布外样本,随后又说这个定义“有点过于严格”:相对于早期艺术作品,大量优秀艺术仍然处于分布内。她更看重的标准是意图。模型提供工具,创意人士带来想法、判断和目的;即使获得同一个界面,她也无法复现这些东西。

3. 一致性与迭代,让艺术家重新获得缺失的控制力

  • 艺术家告诉团队,早期 AI 工具缺乏专业创作所需的控制力。Nano Banana 的角色和物体一致性,让同一主体能够跨图片延续;多张参考图则支持这样的指令:把一张图片的风格应用到另一个角色上,或把指定物体插入既有场景。

  • Wang 强调,艺术本质上是迭代过程,因此多轮对话是核心能力,而非附属功能。模型可以吸收连续修改,成为“创作伙伴”;但他也承认,在极长的对话中,模型的指令遵循能力会下降。提升这种持续工作能力,是明确的开发重点。

  • 当被问及艺术家为何抵触 AI 时,Brichtova 指向了早期的一次性系统:用户输入文字,接受主要由模型和训练数据做出的决定,然后把结果称为自己的艺术。这种方式可能会“让人有点不舒服”。更强的可控性则为创作者留下表达空间,而不只是从模型输出中挑选一个结果。

  • 单纯的新奇感也已经不再是有效差异化。Brichtova 说,人类“很快就会厌倦”;一张明显由单次提示词生成的图片,不会仅仅因为出自机器就继续让人感到有趣。标准正在回到技艺本身:AI 作品必须承载足够的意图和控制力,让另一位艺术家能够辨认出人的贡献。

4. 最终胜出的界面,取决于用户对控制力的重视程度

  • 结合自己在 Adobe 的经历,Wang 将设计张力归纳为:一端是手机或语音界面,另一端是专业人士要求的细粒度调整。团队还没有同时解决这两个问题。他希望未来的工具能够理解上下文并主动推荐有用的下一步,从而不再要求用户学习每一个传统控制项。

  • Nicole 的反驳是,专业人士首先在乎结果,也愿意承受相当程度的复杂性。编程工具已经提供上下文控制和不同模式,而不是一个神奇提示词。类似地,Wang 也称赞了 ComfyUI 的节点式工作流:它们确实复杂,但足够稳定,可以把 Nano Banana、其他模型和视频工具组合起来,制作分镜或关键帧。

  • Brichtova 将市场分为3层:对她父母这样的消费者而言,聊天机器人足够好用——上传一张图片、直接对话即可;专业人士需要深得多的控制力;两者之间,则是过去因专家软件而望而却步的潜在创作者。“中间层存在巨大机会。”

  • Wang 否定了“一个模型统治一切”的思路。优化精确遵循指令,可能会削弱模型在构思阶段的价值,因为用户有时希望模型接管任务并“放开来玩”。不同目标、受众和审美偏好,将继续支撑多样化模型与组合式工作流。

5. 视觉推理让模型从艺术工具扩展为学习系统

  • Brichtova 不希望 AI 把每个5岁孩子的涂鸦都变成精致图片:“我们可能会失去一些东西。”她设想的更好方向,是让 AI 成为一个能够点评、讲解步骤,或提供图像自动补全的伙伴。具有讽刺意味的是,有意生成儿童蜡笔画仍然很难,因为这种抽象程度极高,团队因此还在设计专门的评测。

  • Wang 指出,很多人通过视觉学习;Brichtova 则补充说,目前的 AI 导师大多通过对话或文字提供帮助。多模态导师可以用专门生成的图片和图表配合解释,让概念更有用、更易理解。对于以人为中心的智能体,她认为视觉沟通至关重要:“100%。我绝对这么认为。”

  • 更长期的概念被称为“视觉深度研究”。模型可能花2小时研究一次住宅改造,搜索合适的家具,然后返回多个方案或一份10页演示文稿。对于操作手册或 IKEA 式说明,先把困难任务拆解成一系列视觉中间步骤,可能比直接生成一张最终图片更有价值。

6. 二维生成可以学习世界,但机器人仍需要3D

  • Nicole 对明确的3D世界模型给出了坦率的非答案。3D模型的优势在于持续的几何一致性,但训练数据受到限制,因为“我们不会随身带着3D采集设备到处走”;现有的大多数观察结果,都是空间世界投射到二维平面上的图像。

  • 她个人更倾向于从这些投影中学习潜在的世界表征。视频模型已经展现出足够的3D理解能力,可以支持准确的重建算法;与此同时,人类艺术起源于洞穴墙壁上的线条,现代界面也仍然以二维为主。人类尤其擅长通过空间世界的平面投影进行推理。

  • Oliver 追问了导航问题:一张桌子的画面可能看起来具有立体感,但具身系统必须知道自己不能穿过它。Nicole 将高层规划与运动控制区分开来——识别一栋建筑并向左转,或许可以依靠记忆中的投影完成;但“机器人,是的,它们可能需要3D”。

7. 评测关注的是不均衡感知与有意识的取舍

  • 早期的角色一致性评测使用陌生面孔,结果“什么也说明不了”。团队后来改用自己和熟悉的人,因为一旦对主体足够熟悉,即使是轻微的面部错误也会变得明显,随后又在不同年龄和群体上进行测试。整个过程至今仍有相当部分依赖有经验的人进行目测,而不是依靠一个自动化分数。

  • Wang 解释了为什么排行榜无法把图像质量压缩成一个总分:一次编辑可能更擅长保留身份,另一次则更擅长迁移风格,究竟哪个更好取决于用户意图。Brichtova 补充说,“不存在唯一正确答案”,而且已经发布的模型,都明显带有各家研究实验室的偏好。

  • Brichtova 表示,团队如今拒绝回退的能力包括角色一致性,以及用户要求照片时的逼真度,尤其是在产品和广告场景中。但本次发布的文字渲染仍未达到理想标准;团队接受了这一限制,因为模型的其他优势足以支持上线。

  • 对于 ControlNet 或姿态输入等结构化控制,Brichtova 认为,模型未来可能越来越能根据用户意图判断一次编辑应当保留结构还是采取自由形式,但始终会有用户要求更多控制。像素级用户仍然可以在传统工具中把某个元素向左移动一点,或把它调得更蓝。Wang 也提出,或许一张参考图就能更容易地传达目标姿态,而不必要求用户先提取结构化数据。

8. 可靠性、平台覆盖与人类审美定义下一条前沿

  • Brichtova 将 Gemini app 定位为入口:“乐趣某种程度上是通往实用性的门户。”用户可能因为制作玩偶而来,最终留下来做作业或写作。公司也在探索 Flow 这类与特定场景深度耦合的界面,服务 AI 电影创作者;至于 API 和企业分发,则会把建筑软件等垂直产品留给专业开发者。

  • 日本很快出现了生态层面的证据:开发者已经构建出 Easy Banana 等扩展,服务漫画和动漫工作流。Brichtova 认为,一旦质量越过最低门槛,延迟就会成为倍增器:10秒出一帧会鼓励用户迭代,等待2分钟则会导致放弃。更好的事实性和文字能力,同样可能打开个性化、多语言视觉教材的市场。

  • 模型涌现出的推理能力甚至让构建者自己感到意外。Brichtova 提到几何题,以及根据 HTML 图片渲染网页的案例;Wang 则提到一张学术图表中缺失的结果,模型能够推断出来,并在多个应用中补全。这些案例暗示,视觉问题求解、场景理解,以及长多模态上下文中的状态延续,还存在尚未被发现的“零样本或少样本提示”用法。

  • 团队最终关注的基准已经不再是最好的图片,而是最差的图片:“我们现在处在挑柠檬阶段。”Brichtova 认为,教育和事实性尤其重要。Wang 设想,模型可以吸收150页品牌指南,发现第52页的规定使草稿失效,然后自主修改。背后的共同机制,是在推理阶段进行批评,让控制和合规变得可靠。

  • 艺术家仍然不可或缺,因为模型没有数十年积累形成的审美。Brichtova 讲到与 Ross Lovegrove 合作的案例:模型先基于他的草图进行微调,再用于开发新作品,最终为一款实体椅子的原型提供了灵感。整个结果依赖长时间对话和设计专业能力,而不是“一条提示词、2分钟”。

  • Wang 警告说,如果以所有人的平均偏好为优化目标,最终得到的东西可能广泛讨喜,却很少真正改变视角。他还强调了一个尚未充分利用的能力:交错生成。一次提示词可以产出一系列内容,比如让同一个角色贯穿多张图片的睡前故事。一致性、连续生成与有意识的审美结合起来,指向的将不再是孤立的图片,而是交互式视觉媒体。

Nicole Brichtova

These models are allowing creators to do less tedious parts of the job, right? They can be more creative, and they can spend 90% of their time being creative versus 90% of their time editing things and doing these tedious manual operations.

Oliver Wang

I’m convinced that this ultimately really empowers the artists, right? It gives you new tools. It’s like, hey, we now have watercolors for Michelangelo. Let’s see what he does with it, right? Amazing things come out.

Nicole Brichtova

Maybe start by telling us about the backstory behind the Nano Banana model. How did it come to be? How did you all start working on it?

Oliver Wang

Sure. Our team has worked on image models for some time. We developed the Imagen family of models, which goes back a couple of years. There was also an image-generation model in Gemini before the Gemini 2.0 image-generation model. What happened was that the teams started to focus more on the Gemini use cases, like interactive, conversational, and editing. Um, and essentially what happened is we teamed up and we built this model, which became what’s known as Nano Banana. So yeah, that’s sort of the origin story.

Nicole Brichtova

Yeah, and I think maybe just some more background on that. Our Imagen models were always top of the charts for visual quality, and we really focused on these specialized generation and editing use cases. Then, when Gemini 2.0 Flash came out, that’s when we really started to see some of the magic of being able to generate images and text at the same time, so you can maybe tell a story. Just the magic of being able to talk to images and edit them conversationally. But the visual quality was maybe not where we wanted it to be. And so that became Nano Banana, or Gemini 2.5 Flash Image.

Oliver Wang

Nano Banana is way cooler.

Nicole Brichtova

It’s easier to say. It’s a lot easier.

Oliver Wang

It’s the name that stuck.

Nicole Brichtova

Yes, it’s the name that stuck. But it really became the best of both worlds in that sense: the Gemini smartness and the multimodal, conversational nature of it, plus the visual quality of Imagen. I feel like that’s maybe what resonates a lot with people.

Oliver Wang

Wow. So I guess when you were testing out the model as you were developing it, what were some wow moments that made you think, “I know this is going to go viral. I know people will love this”?

Nicole Brichtova

I actually didn’t feel like it was going to go viral until we had released it on LMArena. What we saw was that we budgeted a comparable amount of queries per second as we had for our previous models that were on LMArena. We had to keep upping that number as people were going to LMArena to use the model. I feel like that was the first time when I was really thinking, “Oh, wow. This is something that’s very, very useful to a lot of people.”

It surprised even me. I don’t know about the whole team, but we were trying to make the best conversational editing model possible. Then it really started taking off when people were going out of their way and using a website that would only give you the model some percentage of the time. Even that was worth using—going to that website to use the model. So I think that was really the moment, at least for me, that I thought, “Oh, wow. This is going to be bigger.”

Oliver Wang

That’s actually the best way to condition people: only give them a reward partially [laughter], not all the time, by design.

Nicole Brichtova

I had a moment earlier. I’ve been trying some similar queries on multiple generations of models over time. A lot of them have to do with things I wanted to be as a kid, like an astronaut or explorer, or with putting me on the red carpet. I tried it on a demo that we had internally before we released the model. It was the first time when the output actually looked like me.

You guys play with these models all the time. The only time that I’ve seen that before is if you fine-tune a model, using LoRA or some other method to do that. You need multiple images, it takes a really long time, and then you have to actually serve it somewhere. So this was the first time when it was zero-shot: one image of me, and it looks like me. I was like, “Wow.” And then there became these like we have decks that are just like covered in my face as I was trying to convince other people that it was really cool. Um, and really I think the moment more people realized that it was like a really fun feature to use is when they tried it on themselves because it’s kind of fun when you see it on another person, but it doesn’t really resonate with people emotionally. It makes it so personal. It’s like your kids, your spouse and I think that’s your dog,

Oliver Wang

your dog.

Nicole Brichtova

Your dog. And that’s really what started resonating internally. People started making all these ’80s makeover versions of themselves. That’s when we really started to see a lot of internal activity, and we were like, “Okay, we’re onto something.”

Oliver Wang

It’s a lot of fun to test these models when we’re making them because you see all these amazing, creative things that people make. “Oh, wow. I never thought that was possible.”

Nicole Brichtova

So it’s really fun.

Oliver Wang

No, we’ve dealt with the whole family, and it’s a crazy amount of fun.

Nicole Brichtova

So, think a bit about the long term. Where does this lead? We built these new tools that I think will change visual arts forever, right? We suddenly can transfer style. We can generate consistent images of a subject. I have what used to be a very complex, manual Photoshop process. Suddenly, I type 1 command and it magically happens.

But what’s the end state of this? Do we have an idea yet? How will creative arts be taught in a university 5 years from now? Want to take that?

Oliver Wang

So I think it’s going to be a spectrum of things. On the professional side, a lot of what we’re hearing is that these models are allowing creators to do less tedious parts of the job. They can be more creative, and they can spend 90% of their time being creative versus 90% of their time editing things and doing these tedious manual operations. So I’m really excited about that. I think we’ll see an explosion of creativity on that side of the spectrum.

For consumers, there are probably 2 sides of the spectrum. One is that you might just be doing some of these fun things, like Halloween costumes for my kid, right? The goal there is probably just to share it with somebody—your family or your friends.

On the other side of the spectrum, you might have tasks like putting together a slide deck. I started out as a consultant. We talked about that at the beginning. You spend a lot of time on very tedious things, like trying to make things look good and trying to make the story make sense. I think for those types of tasks, you probably just have an agent that you give the specs of what you’re trying to do, and then it goes out and lays it out nicely for you. It creates the right visual for the information that you’re trying to convey.

It really is going to be this spectrum, depending on what you’re trying to do. Do you want to be in the creative process and actually tinker with things and collaborate with the model, or do you just want the model to go do the task and be as minimally involved as possible?

Nicole Brichtova

So, in this new world, then, what is art? Somebody recently said art is if you can create an out-of-distribution sample. Is that a good definition, or is it aiming too high?

Oliver Wang

Do you think art is out of distribution or in distribution for the model?

Nicole Brichtova

There we go. [laughter] I think that out-of-distribution sample is a little bit too restrictive. I think a lot of great art is actually in distribution for art that occurred before it. What is art? I think it’s a very philosophical debate, and there are a lot of people who discuss this. To me, the most important thing for art is intent.

What is generated from these models is a tool to allow people to create art. I’m actually not worried about the high end and the creatives and the professionals, because I’ve seen—if you put me in front of one of these models, I can’t create anything that anyone wants to see.

Oliver Wang

But I’ve seen what people can do who are creative and who have intent and ideas.

Nicole Brichtova

That’s the most interesting thing to me: the things they create are really amazing and inspiring for me. I feel like the high-end professionals and creatives will always use state-of-the-art tools, and this is another tool in the tool belt for people to make cool things.

I think one of the really interesting things that I kept hearing about this model in particular from creatives and artists was that a lot of them felt like they couldn’t use a lot of AI tools before because those tools didn’t allow them the level of control that they expected for their art. On one side, that was character or object consistency. They really used that to have a compelling narrative for a story, and so before, when you couldn’t get the same character over and over, it was very difficult.

The second thing I hear all the time from artists is that they love being able to upload multiple images and say, “Use the style of this on this character,” or, “Add this thing to this image,” which is something that I think was very hard to do even with previous image-editing models.

I guess I'm curious: Was that something you guys were really optimizing for when you trained this one, or how did you think about that?

Oliver Wang

Definitely. Customizability and character consistency are things that we closely monitored during development, and we tried to do the best job we could on them. I think another thing is the iterative nature of an interactive conversation. Art tends to be iterative as well, where you make lots of changes, see how it's going, and make more.

This is another thing that I think makes the model more useful, and that's an area where I also feel like we can improve the model greatly. I know that once you get into really long conversations, it starts to follow your instructions a little bit worse. But this is something that we're planning to improve on and make the model more of a natural conversation partner, or a creative partner, in making something.

Nicole Brichtova

One thing that's so interesting is that after you guys launched Nano Banana, we started to hear about editing models all the time, everywhere. It was like after you launched, the world woke up: There were editing models—it's great, everyone wants it. Then, obviously, it kind of goes into the customizability and personalization of it.

And, Oliver, I know you used to be at Adobe, and there was also software where we used to manually edit things. How do you see the knobs evolve now at the model layer versus what we used to do?

Oliver Wang

Yeah. I think one thing that Adobe has always done, and professional tools generally require, is lots of control and lots of knobs. There's always a balance: We want someone to be able to use this on their phone.

Nicole Brichtova

Maybe with just a voice interface.

Oliver Wang

And we also want someone who's a really professional art creator to be able to make fine-grained adjustments. I think we haven't exactly figured out how to enable both of those yet.

Nicole Brichtova

But there are a lot of people building really compelling UIs, and I think we're figuring out different ways it can be done. You have thoughts?

Oliver Wang

Well, I also hope that we get to a point where you don't have to learn what all these controls mean, and the model can smartly suggest what you could do next based on the context of what you've already done. That feels like it's prime for someone to tackle.

What do the UIs of the future look like in a way where you probably don't need to learn 100 things that you had to before, but the tools should be smart enough to suggest what they can do based on what you're already doing?

Nicole Brichtova

That's such an insightful take. I definitely had moments when I used Nano Banana where I was like, “I didn't know I wanted this, but—”

Oliver Wang

But I didn't even ask for this style. I don't even know the words for what that style is called. This is very insightful about how image embeddings and language embeddings are not one-to-one: We can't map all editing tasks to language. Go ahead.

Nicole Brichtova

Yeah, let me take a little counterpoint just to see where this goes. The other question of how complex the interface can be is limited by what we can express in software and how easy we can make something in software, which to some degree is also limited by how much complexity a user is willing to tolerate. If you have a professional—

Oliver Wang

They only care about the result. They're willing to tolerate a vast amount of complexity. They have the training, the education, and the experience to use that.

Then we may end up with lots of knobs and dials. It's just very different kinds of dials. Today, if you use Cursor for coding, it's not that it has a super-easy, single-text-prompt interface. It has a good amount of context: Here, add context; here, different modes; and so on.

So will we have the ultra-sophisticated interface for the power user, and how would that look? I'm a big fan of ComfyUI and node-based interfaces in general.

Nicole Brichtova

And that is complex.

Oliver Wang

And it's complex, but it's also very robust, and you can do a lot of things. After we released Nano Banana, we saw people building all these really complicated ComfyUI workflows where they were combining a bunch of different models and tools together. That generated some really amazing outputs—for example, using Nano Banana as a way to get storyboards or keyframes for video models. You can plug these things together and get really amazing outputs.

So I think that at the pro or developer level, these kinds of interfaces are great. In terms of the prosumer level, I think it's very much unknown what it's going to look like in a couple of years.

Nicole Brichtova

Yeah. I think it really depends on your audience, right? For the regular consumer, I use my parents as an example. The chatbot is actually kind of great.

Oliver Wang

Oh, yeah.

Nicole Brichtova

Because you don't have to learn a new UI. You just upload your images and then talk to it. It's kind of amazing that way.

Then for the pros, I agree that you need so much more control. Somewhere in between are probably people who may want to be doing this, but were too intimidated by the professional tools in the past. For them, I do think there's a space where you need more control than the chatbot gives you, but you don't need as much control as the professional tools give you. What's that kind of in-between state?

Oliver Wang

There's a ton of opportunity there.

Nicole Brichtova

There's a ton of opportunity there. It is interesting that you mentioned ComfyUI because it's on the other far end of the workflow spectrum. A workflow can have hundreds of steps and nodes, and you need to make sure all of them work, whereas on the other side there's Nano Banana. You describe something with words, and then you get something out.

I don't know much about model architecture and things like that, but is your view that the world is moving toward an ensemble of models hosted by one provider doing it all, or do you think the world is moving more toward everyone building workflows, with Nano Banana as one of the nodes in ComfyUI?

Oliver Wang

I definitely don't think that the broad range of use cases will be fully satisfied by one model at any point. I think there will always be a diversity of models.

I'll give you an example. We could optimize for instruction-following in our models and make sure they do exactly what you want, but it might be a worse model for someone who's looking for ideation or inspiration, where they want the model to take over and do other things, go crazy.

I just think there are so many different use cases and so many types of people that there's a lot of space in this area for multiple models. That's where I see us going. I don't think this is going to be a single model to rule them all.

Nicole Brichtova

Makes complete sense. Let's go to the very other end of the spectrum from the professional. Do you think kindergarteners in the future will learn drawing by sketching something on a little tablet and then having AI turn that into a beautiful image? Is that how they'll get in touch with art?

Oliver Wang

I don't know if you always want it to turn into a beautiful image, but I think there's something there about the AI being, again, a partner and a teacher to you in a way that you didn't have.

Nicole Brichtova

I didn't know how to draw, and I still don't. I don't have any talent for it, really. But I think it would be great if we could use these tools in a way that actually teaches you the steps, helps you critique, and maybe shows you an autocomplete for images—what the next step could be.

Or maybe it could show me a couple of options and explain how to actually do this. I hope it's more in that direction. I don't think we all want every 5-year-old's image to suddenly look perfect.

Oliver Wang

We would probably lose something in the process. As someone who struggled the most in high school out of all my classes in art and sketching, I actually would have preferred it. But I know a lot of people want their kids to learn to draw, which I understand.

Nicole Brichtova

It's funny because we've been trying to get the model to create childlike crayon drawings, which is actually quite challenging.

Oliver Wang

Ironically, sometimes the things that are hard to make are hard because the level of abstraction is very large, right?

Nicole Brichtova

So it's actually quite difficult to make those types of images. Your dedicated pre-K fan. [laughter]

Oliver Wang

We do have some evals right now.

Nicole Brichtova

To try to see if we're getting better.

Oliver Wang

In general, I'm very optimistic about AI for education. Part of the reason is that I think most of us are visual learners.

Nicole Brichtova

Right? So right now, AI as a tutor can basically only talk to you or give you text to read, and that's definitely not how students learn. I think these models have a lot of potential as a way to help education by giving people visual cues.

Imagine if you could get an explanation for something where you get the text explanation, but you also get images and figures that help explain how they work.

Oliver Wang

I think it’ll just make everything much more useful and much more accessible for students. So I’m really excited about that.

Nicole Brichtova

On that point, one thing that’s very interesting to us is that when Nano Banana came out, it almost felt like part of the use case was the reasoning model. You have a diagram. Absolutely. Right? You can explain some knowledge visually. So the model isn’t just doing an approximation of the visual aspect; there’s a reasoning aspect to it, too.

Oliver Wang

Do you think that’s where we’re going? Do you think all the large models will realize that, to be a good LLM or VLM, we have to have both image and language and audio, and so on and so forth?

Nicole Brichtova

100%. I definitely think so. The future for these AI models that I’m most excited about is where they’re tools for people to accomplish more things. If you imagine a future where you have these agentic models that just talk to each other and do all the work, then it becomes a little bit less necessary that there’s this visual mode of communication. But as long as there are people in the loop, and as long as the motivation for the task they’re solving comes from people, I think it makes total sense that visual modality is going to be really critical for any of these AI agents going forward.

Nicole Brichtova

Will we get to a point where, you know, I’m asking you to create an image, it sits for 2 hours, reasons with itself, has drafts, explores different directions, and then comes back with a final answer?

Oliver Wang

Yeah, absolutely, if it’s necessary.

Oliver Wang

And maybe not just for a single image, but to the point where maybe you’re redesigning your house and you actually really don’t want to be involved in the process, right? You’re like, “Okay, this is what it looks like; this is some inspiration that I like.” And then you send it to a model the same way that you would send it to a designer.

Nicole Brichtova

It’s the visual deep research.

Oliver Wang

The visual deep research. I really like that term. Then it goes off and does its thing and searches for maybe the furniture that would go with your environment, and then it comes back to you and presents you with options. Maybe you don’t want to sit for 2 hours on one thing—an art book, you know, a 10-slide deck.

I also think that if you think about instruction manuals or IKEA directions, breaking down a hard problem into many intermediate steps could be really useful as a way to communicate.

Nicole Brichtova

So when can we generate LEGO sets?

Oliver Wang

Yeah, soon, maybe. Do we at some point need 3D as part of it?

Nicole Brichtova

Right.

Oliver Wang

I mean, there’s a whole debate around world models and image models and how they fit together. Thoughts? Enlighten us here. What’s the short summary of where we’ll end up there?

Nicole Brichtova

I mean, I don’t know the answer. Obviously, the real world is in 3D. So if you have a 3D world model, or a world model that has explicit 3D representations, there are a lot of advantages. For example, everything stays consistent all the time.

The main challenge is that we don’t walk around with 3D capture devices in our pockets. In terms of the available data for training these models, it’s largely the projection onto 2D. So I think that both viewpoints are totally valid for where we’re going.

I come a bit from the projection side. I think we can solve almost all, if not all, the problems by working on the projection of the 3D world directly and letting the models learn the latent world representations. We see this already: the video models have very good 3D understanding. You can run reconstruction algorithms over the videos you generate, and they’re very accurate.

In general, if you look at the history of human art, it starts as the projection, right? People drawing on cave walls. All of our interfaces are in 2D, so I think humans are very well suited for working on this projection of the 3D world into a 2D plane. It’s a really natural environment for interfaces and for viewing.

Oliver Wang

That is very true. I’m a cartoonist in my spare time, and drawing in 2D is just light and shadow. Then you present yourself with 3D. Can we trick ourselves into believing it’s 3D, even though it’s on a piece of paper? What a human can do, and what a drawing or a model can do, is let us navigate the world. We see a table; we can’t walk past it. I guess the question becomes: if everything is 2D, how do you solve that problem?

Nicole Brichtova

Well, I don’t think—yeah. So if we’re trying to solve the robotics problems, I think maybe the 2D representation is useful for planning and visualizing at a high level. I think people navigate by remembering 2D projections of the world. You don’t build a 3D map in your head. You’re more like, “Oh, I see this building. I turn left.”

Oliver Wang

Yeah.

Nicole Brichtova

So I think that for that kind of planning, it’s reasonable. But for the actual locomotion around the space, 3D is definitely important.

Oliver Wang

Robotics, yeah. They probably need 3D. [Laughter] That’s the saving grace.

Nicole Brichtova

So, character consistency, which you previously mentioned—I really love the example of when a model feels so personal, because people are so tempted to try it. How did you unlock that moment? The reason why I ask is that character consistency is so hard. There’s a huge uncanny valley to it. If it’s someone I don’t know and I see their AI generation, I’m like, “Okay, it’s maybe the same person.” But if it’s someone I know, and there’s just a little bit of a difference, I’m actually very turned off by it because I’m like, “This is not a real person.”

In that case, how do you know what you’re generating is good? Is it mostly by user feedback—“I love this”—or is it something else? You look at faces, you know. [Laughter]

Nicole Brichtova

No. So, not even before you ever released this, right? When we were developing this model, we actually started out doing character consistency evaluations on faces we didn’t know, and it doesn’t tell you anything, right?

Then we started testing it on ourselves and quickly realized, “Okay, this is what you need to do,” because this is a face that I’m familiar with. There are a lot of eyeballing evaluations that happen, with the team testing it on themselves and generally on people they know. Oliver probably knows my face at this point well enough to be able to tell whether or not it’s actually me when it’s generated.

We do a lot of that. Ideally, you test it on different sets of people, different ages, and different groups of folks to make sure that it works across the board.

Oliver Wang

Yeah, I think they’re right. That touches a little bit on this bigger issue, which is that evaluations are really difficult in this space because human perception is very uneven in terms of the things that it cares about. It’s very hard to know how good the character consistency of a model is and whether it’s good enough.

I think there’s still a lot of improvement we can make on character consistency, but for some use cases, we got to a point—and we weren’t the first image-editing model by any means—where once the quality gets above a certain level for character consistency, it can just take off because it becomes useful for so much more. As it gets better, it’ll be useful for even more things, too.

I think one of the really interesting things we’re seeing across a bunch of modalities, of which image editing and generation are obviously examples, is that the arenas and benchmarks are awesome. But especially when you have multidimensional things like images and video, it’s very hard, as all of the models get better and better, to condense every quality of a model into one judgment.

You’re judging, “Okay, you swap a character into an image and change the style of the image.” Maybe one model did the character swap and consistency much better, and the other did the style much better. How do you say which output is better? It probably comes down to what the person cares most about and what they want to use it for.

Are there certain characteristics of the model that you value more than other things in making those trade-offs when deciding which version of the model to deploy, or what to really focus on during training?

Nicole Brichtova

Yes, there are. One of the things I like about this space is that there’s no right answer. There’s quite a lot of—I don’t know if it’s taste, but it’s preference that goes into the models. You can see the differences in the preferences of the different research labs in the models that they release.

When we’re balancing 2 things, a lot of it comes down to, “Oh, well, I don’t know. I just like this look better,” or, “This feature is more important to us.”

Oliver Wang

I’d imagine it’s hard for you guys, too, because you have so many users. Google, being in the Gemini app, means everyone in the world can use that, versus many other AI companies that think, “We’re only going for the professional creatives,” or, “We’re only going for the consumer meme-makers.” You have the unique and exciting—but challenging—task that literally anyone in the world can do this.

How do we decide what everyone would want?

Nicole Brichtova

Yeah. Sometimes we do make these trade-offs. We do have a set of things that are super high priority, that we don't want to regress on. Right? Because character consistency was so awesome and so many people are using it, we don't want our next models to get worse on that dimension. Right? So we pay a lot of attention to it.

We care a lot about images looking photorealistic when you want photos, and this is important. 1, I think we all prefer that style, too. [laughter] 2, for advertising use cases, a lot of it is photorealistic images of products and people. We want to make sure that we can do that.

Sometimes there are just things that fall by the wayside. For this 1st release, the model is not as good at text rendering as we would like it to be, and that's something that we want to fix in the future. But it was 1 of those things where we looked at it and thought, okay, the model's good at XYZ, it's not as good at this, but we still think it's okay to release, and it will still be an exciting thing for people to play with.

Oliver Wang

If you look at the past, for previous model generations, there were a lot of things we did with sidecar models, like ControlNet or something like that, where we basically figured out a way to provide structured data to the model to achieve a particular result. It seems like these newer models have taken a step back just because they're so incredibly good at just prompting or giving a reference image and picking things up from there. Where will this go long term? Do you think this will come back to some degree?

From the creator's perspective, having OpenPose information so I can get a pose exactly right for multiple characters seems very tempting, right? Or, to rephrase it a little bit, does the bitter lesson hold here—that at the end of the day, everything's just 1 big model and you throw things in—or is there a little structure we can offer to make this better?

Nicole Brichtova

I think there will always be users that want control that the model doesn't give you out of the box. But I think we tried to make it so that—because really, what an artist wants when they want to do something is for their intent to be understood. And I think these AI models are getting better at understanding the intent of users. So often, when you ask text queries now, the model gets what you're going for.

So, in that sense, I think we can get pretty far with understanding the intent of our users. Maybe some of that is personalization: we need to know information about what you're trying to do or what you've done in the past. But I think once you can understand the intent, then you can generally do the type of edit—is this a very structure-preserving edit, or is this a free-form kind of edit? We can learn these kinds of effects, I think.

But still, of course, there's 1 person who's going to really care about every pixel: this thing needs to be slightly to the left and a little bit more blue. Those people will use existing tools to do that.

Oliver Wang

I think it's like, I want an image with 26 people spelling out every letter of the alphabet or something like that. That's the kind of thing where I think we're still quite a bit away from getting that right in the 1st try.

Nicole Brichtova

But then the question, I guess, is: do you really want to be the 1 who's extracting the pose and providing that as information, or do you just want to provide some reference image and say, “This is actually what I want. Model, go figure this out,” right?

Oliver Wang

There are 26 people, every 1 in a different style. Fair enough. Yeah, I think in that case I wouldn't spend a ton of time building a custom interface for making this picture of 46 people. That seems like the kind of thing that we can solve.

Nicole Brichtova

Just transfer.

Oliver Wang

Do you think the representation of what AI images are will change? The reason I ask the question is that, as artists, there are different formats we play with. There are SVGs, where we have anchor points and Bézier curves.

On the other side, there's Procreate or Fresco, what have you. There are layers that we can also play with. There's the other parameter, which is the brush you use—the brush, the texture of it. For every 1 of those parameters, you can write a script and actually do something very personal with it.

Nicole Brichtova

Mhm.

Oliver Wang

Do you think pixel is the right representation—the endgame—for an image-generation model, or do you think there's a net-new representation that we haven't invented yet?

Nicole Brichtova

That's an easy question. [laughter] I'll say that everything is a subset of pixels.

Oliver Wang

That's true.

Nicole Brichtova

Text is a subset of pixels because I could just render all the text as an image. So how far can we get with just pixels is an interesting question. I think if the model is really responsive and handles multi-turn interactions well, then you can probably get pretty far, because the primary reason I think you would want to leave the pixel domain is for editability.

In cases where you need to have your font, or you want to change the text, or you want to move things around just like with control points, it could be useful to have a kind of mixed generation, which consists of pixels and SVGs and other forms. But if we can do it all—if the multi-turn interaction is enough—then I think you can get pretty far with pixels.

I will say that 1 of the things that's exciting about these models that have native capabilities is that you now have a model that can generate code and images.

Oliver Wang

So there's a lot of interesting things that come in at that intersection, right? Maybe I wanted to write some code and then make some things be rasterized, some things be parametric.

Nicole Brichtova

Yeah.

Oliver Wang

Like, stick it all together—

Nicole Brichtova

Train it together. This would be very cool. That's such a good point, because I did see a tweet of someone asking Claude Sonnet to replicate an image on an Excel sheet where every cell is a pixel. [laughter]

Oliver Wang

Which is a very fun exercise. It was a coding model that doesn't really know anything about images, yet it worked.

Nicole Brichtova

Yeah, there's the classic pelican-riding-a-bicycle test.

Oliver Wang

Yeah, totally. I have 1 on models and interfaces, if that's okay. Sorry if I'm bringing up too much product stuff, guys. I'm just very curious on the product front.

I guess I'm curious how you think about owning the interface where people are editing or generating images with Nano Banana versus really just wanting a ton of people to use the model for different things in the API. We've talked about so many different use cases: ads, education, design, architecture. Each of those things could have a standalone product built on top of Nano Banana that prompts the model in the right way or allows certain types of inputs or whatever.

Is your vision that the product in the Gemini app is a playground for people to explore, and then developers will build the individual products that are used for certain use cases? Or is that something you're also interested in owning?

Nicole Brichtova

I think it's a little bit of everything. I definitely think that the Gemini app is an entry point for people to explore. The nice thing about Nano Banana is that I think it shows that fun is kind of a gateway to utility: people come to make a figurine image of themselves, but then they stay because it helps them with their math homework or helps them write something, right? And so I think that's a really powerful transition point.

There are definitely interfaces that we're interested in building and exploring as a company. You may have seen Flow from Josh's team in Labs. That's really trying to rethink what's the tool for AI filmmakers, right? For AI filmmakers, image is actually a big part of the iteration journey, because video creation is expensive. A lot of people think in frames when they initially start creating, and a lot of them even start in the LLM space for brainstorming and thinking about what they want to create in the 1st place.

There's definitely a place that we have in that space, just trying to think about what that looks like. We have the advantage of it sitting close to the models and the interfaces, so we can build that in with tight coupling.

And then there's definitely the—you know, we're probably not going to go build software for an architecture firm. My dad is an architect, and he would probably love that. But I don't think that's something that we will do. Somebody should go and do that.

And that's why it's exciting, because we do have the developer business and the enterprise business. People can go use these models and then figure out what's the next-generation workflow for this specific audience, so that I can help them solve a problem. So I think the answer is kind of like, yes, all 3.

Oliver Wang

Yeah. I brought that up. I don't know if you guys have been following the reception of Nano Banana in Japan, but I'm sure you've heard—it's been insane.

It’s so funny. Half of my X feed is now these really heavy Nano Banana users in Japan who have created Chrome extensions. There’s one called Easy Banana that’s specifically for using Nano Banana for manga generation and specific types of anime and things like that. They go super deep into prompting the model for you and storing the outputs in various places, using, obviously, your underlying model to generate these amazing anime that you would never guess were AI-generated because the level of precision and consistency and that sort of thing is just beyond what I’ve seen any single model be able to do today.

To Justin’s point, what are some force multipliers that you guys have seen in the model? What I mean by this is, for example, if you unlock character consistency, you can generate different frames, and then you can make a video, and then you can make a movie, right? These are the things that, if you get them right, have so many more downstream tasks that can derive from them. Just curious: how do you think about the force multipliers that you want to unlock?

Nicole Brichtova

What’s the next big one?

Oliver Wang

What’s the next big wave of people who can just use Nano Banana as the base model for all the downstream tasks?

Nicole Brichtova

I think one current one actually is also the latency point, right? It makes it really fun to iterate with these models when it just takes 10 seconds to generate the next frame. If you had to sit there and wait for 2 minutes, you would probably just give up and leave. It’s a very different experience. I think there has to be some quality bar, because if it’s just fast and the quality isn’t there, then it also doesn’t matter, right? You have to hit a quality bar, and then speed becomes a force multiplier.

I think this general idea of just visualizing information—to your education point from earlier—is another one, right? And that needs—

Oliver Wang

Good text.

Nicole Brichtova

It needs factuality, right? Because if you’re going to start making visual explainers about something, it looks nice, but it also needs to be accurate, right? I think that’s probably the next level, where at some point you could also just have a personalized textbook for you. It’s not just the text that’s different, but also the visuals.

Oliver Wang

Yeah, The Diamond Age, basically.

Nicole Brichtova

And then it should also internationalize really well, right? A lot of the time today, you might actually be able to find a diagram that explains the thing you’re trying to learn about on the internet, but it’s maybe not in the language that you actually speak. I think that becomes another way to improve and open up accessibility of information to a lot more people, and again visually, because a lot of people are visual learners.

Oliver Wang

Interesting. How do you think about images being generated? The reason why I ask is that there’s another very cool example I’ve seen someone making work with Nano Banana. He wrote a script and then kept prompting the model to say, “Generate the frame 1 second after this,” and then it became a video. When I saw it, I thought, well, is every image just 1 frame in a continuum? You always know about the continuum in a parallel universe. You could have generated any 1 of them.

Nicole Brichtova

Video and images are very closely related. I also think what we’re seeing in these kinds of what’s-coming-next, or sequence-predicting, use cases is the model’s generalization in world knowledge as well. Where do I think it’s going? I think video is an obvious next domain. When you have editing, a lot of times what you’re asking is, “What happens if I do this?” And that’s what video has: the time sequence of actions. It’s like we have a low-frames-per-second video that you can interact with, but obviously making something fully interactive and real-time is the direction this field is headed.

Oliver Wang

So you are probably in the 0.001% of the most experienced people in the world using image models. What are your personal favorite use cases? How do you use it day to day if you’re not just testing an existing model?

Nicole Brichtova

I’m not sure I am in the very top [laughter], but I’ll tell you what. The personalization aspect is the thing that totally drives it home for me. I have 2 young kids, and the best things that I do with the model are the things I do with my kids. We can make their stuffed animals come to life in these types of applications, and it’s so personal and gratifying to see.

We also have a lot of people taking old pictures of their family, for example, and restoring them. I think that’s the real beauty of the edit models: you can make it about the 1 thing that matters most to you. So that’s what I use it for—my kids, basically.

Oliver Wang

Very nice. You’re basically making content that you probably would have never made before, and it’s for the consumption of 1 person, right? Or 1 family. You’re telling these stories that you would have never told before.

Similarly, I do a lot of family holiday cards and birthday cards and whatnot. Now, anytime I make a slide deck, I force myself to generate some images that are contextually relevant and then try to get the text right and all of those things. Then we try to push the boundaries around whether you can make a chart in the pixel space. That’s another question, right? Because you also want the bars in the bar chart to be accurately positioned relative to one another.

I think we do a lot of these things. I’m actually really impressed with the people we work with on the team who are just very creative. We have a team that works really closely with us on models that we’re developing, and they just push the boundary. They’ll do crazy things with the models.

What’s the most surprising thing you’ve seen here? Like, “I didn’t know our model could do this.”

Nicole Brichtova

Even simple things, like texture transfer. You take a portrait of a person and ask, “What would it look like if it had the texture of this piece of wood?” I would never have thought of this as a use case, because my brain just doesn’t work that way. But people push the boundaries of what you can do with these things.

Oliver Wang

That is an interesting example of world knowledge, because texture technically is 3D. There’s the whole 3D aspect of it, with light and shadow, but this is a 2D transfer. That’s very cool.

Nicole Brichtova

I think for me, the thing I’m most excited by and maybe most impressed by is the use cases that test the reasoning abilities of the models. Some people on our team figured out that you could give geometry problems to the model and ask it to solve for X here, fill in this missing thing, or present this from a slightly different view.

These types of things that really require world knowledge and the reasoning ability of a state-of-the-art language model are the things that make me go, “Wow, that’s amazing. I didn’t think we would be able to do that.”

Oliver Wang

Can it generate and compile code on a blackboard yet? If I take a picture of my code on the laptop, would it know if it compiles from the image model?

Nicole Brichtova

I’ve seen examples where people give it an image of HTML code and have the model render the webpage, and it can do that.

Oliver Wang

That’s very cool. The coolest example I saw—since I came from academia, I spent a lot of time writing papers and making figures—was that 1 of our colleagues took a picture of 1 of the result figures from 1 of their papers. It had a method that could do a bunch of different things—this 1 had a bunch of different types of applications in the paper—and asked the model to sort of erase the results.

You have the inputs, and you ask the model to solve all of these in picture form, in a figure of a paper, and it was able to do that. It could figure out what problem the figure was asking for, find the answer, put it in the image, and do that for a bunch of different applications at the same time, which was really amazing.

Nicole Brichtova

That’s very cool. Has anyone built an application on top of that capability yet? What’s the application that will come out of that?

Oliver Wang

I think there are a lot of very interesting, I would say, zero-shot capability and problem-solving-type things that we don’t even know the boundary of yet. Some of these are probably quite useful. If you want to have a method that solves some problem X—I don’t know, finds the normals of the scene or the surface orientations or something—you probably can prompt the model to give you a reasonable estimate.

Nicole Brichtova

So I think there are lots of problems, like understanding problems and other types of things, that we could maybe solve with zero- or few-shot prompting that we don't know yet.

Oliver Wang

Yeah, there's one thing you mentioned I found super interesting, which is world knowledge transfer. But in a lot of world models or video models, there always is something that keeps the state. Just because you look away doesn't mean that the chair should disappear or change color, because that's not what the state of the world is. How do you see that? Do you think there's relevance there in image models? Is that something you even consider optimizing for?

Nicole Brichtova

Yeah, I mean, if you think about an image model that has a long context where you can put other things in that context, like text, images, audio, and video, then I think you're definitely reasoning over the context of things you have to produce a final output image or video. So, yeah, I think there's definitely some model capability to do this type of stuff already.

Oliver Wang

Got it. I haven't tested it out yet for this big use case, but I'll let you know. [laughter] That's one of my favorite things about these models: just finding—and I'm sure it's really fun for you guys, and you probably have much more of a hint than we do about what they can do—but sometimes you'll just see some crazy X or Reddit post, or wherever, about some incredible thing that someone has figured out how to do that you would never expect the model might be able to do, necessarily. Then other people kind of build on that and say, “Oh, and then I tried the next iteration of this thing,” and suddenly you have this almost entirely new space that's been discovered in terms of what the models are capable of. It must be fun, as people much more deeply involved in building these models and building the interfaces, to watch that happen.

Nicole Brichtova

Yeah.

Oliver Wang

So, if you talk to visual artists today, I personally love this stuff. I post about it on the internet. You can get some very skeptical answers. People are like, “Oh, this is terrible.” Right? Do you have any idea what triggers this reaction? I'm convinced that this ultimately really empowers the artists. It gives you new tools. It's like, “Hey, we now have, I don't know, watercolors for Michelangelo. Let's see what he does with it,” and amazing things come out. It's a similar thing, but what triggers this strong reaction against it?

Nicole Brichtova

So, I think it's something to do with the amount of control over the output. In the beginning, when we had these kinds of text-to-image models, they would be very much like a one-shot: you put in some text, you get an output, and people would be like, “Oh, this is art. This is the thing I made.” I think that maybe rubs people a little bit the wrong way when they come from the creative community, because most of the decisions that were made were made by the model and by the data that was used to train it.

Oliver Wang

You can't express yourself anymore physically, right?

Nicole Brichtova

Yeah, exactly. As a creative person, you want to be able to express yourself, so I think as we make the models more controllable, a lot of these concerns—like, “Oh, the computer is doing everything”—may go away. The other thing is that there was a period of time where we were all so amazed by the images these models could create that we were pretty happy to see just, “Oh, this stuff comes out of these models.” But I think humans get bored fast of this type of thing.

There was a big rush, and now if you see an image that was just a single-prompt image—someone didn't think about it much—you can kind of tell, like, that's an AI-generated image. It's not that interesting. So I think there's still this boundary: now you need to be able to make interesting things with the AI tools, which is hard, but this will always be a requirement. We need someone to be able to do this. And I think—

Oliver Wang

We still need artists.

Nicole Brichtova

We still need artists. And I think artists will be able to also recognize when people have actually put a lot of control and intent—

Oliver Wang

And still not be an artist. [laughter]

Nicole Brichtova

Maybe, but there's a lot of craft and a lot of taste that you accumulate, sometimes over decades. I don't think these models really have taste, and I think a lot of the reactions that you mentioned maybe also come from that.

We do work with a lot of artists across all the modalities that we work with—image, video, and music—because we really care about building the technology step by step with them and trying to figure out how they really help us push the boundary of what's possible. A lot of people are really excited, but they really do bring a lot of their knowledge and expertise, and 30 years of design knowledge.

We just worked with Ross Lovegrove on fine-tuning a model on his sketches so that he can then create something new out of that, and then we designed an actual physical chair that we have a prototype of. So, there are a lot of people who want to bring the expertise that they've built and the rich language that they use to describe their work, and have that dialogue with the model so that they can push their work to the frontier.

It doesn't happen in 1 prompt and 2 minutes. It does require a lot of that taste and human creation and craft that goes into building something that actually then becomes art.

Oliver Wang

At the end, it's still a tool that requires the human behind it to express the feelings and the emotions and the story and everything.

Nicole Brichtova

Yeah, absolutely. Absolutely.

Oliver Wang

And that's what resonates with you when you probably look at it, right? You will have a different reaction when you know that there's a human behind it who has spent 30 years thinking about something and then poured that into a piece of art. I think there's also a bit of this phenomenon: most people who consume creative content—and maybe even people who care a lot about it—they don't know what they're going to like next. You need someone who has a vision and can do something that's interesting and different, right? Then you show it to people, like, “Oh, wow. That's amazing,” but they wouldn't necessarily think of that on their own.

Nicole Brichtova

Right?

Oliver Wang

So when we're optimizing these models, one thing we could do is optimize for the average preference of everybody.

Nicole Brichtova

But I don't think you end up with interesting things by doing that. You end up with something that everyone kind of likes, but you don't end up with things that people are like, “Oh, wow. That's amazing. I'm going to change my whole perspective of art because I saw that.”

Oliver Wang

There's the avant-garde edition of the model. If I use the term, there's the—I don't know what's the other end of the spectrum—the marketing edition, where it's very predictable and—

Nicole Brichtova

Very straightforward.

Oliver Wang

Yeah. Well, since we're coming up on time, last couple of questions. One is, what's one feature that you know the model is capable of that you wish people asked you about more?

Nicole Brichtova

Interleaving.

Oliver Wang

Yeah, I think we've always been amazed that nobody ever posts anything about in-story generation, which is what we call the model's ability to generate more than 1 image for a specific prompt. You can ask for something like, “I want a story, like a bedtime story or something,” and generate the same character over a series of images. I think people haven't really found it useful yet, or haven't discovered it. I don't know.

Nicole Brichtova

Oh, interesting.

Oliver Wang

Well, if you're listening to the podcast, go try this out.

Nicole Brichtova

Try. [laughter]

Oliver Wang

Yeah. And what's the most exciting technical challenge that you look forward to tackling in the next, I don't know, months or years?

Nicole Brichtova

So, I think there's really a high ceiling in terms of quality for where we're going. People look at these images and say, “Oh, it's almost perfect. We must be done.” For a while, we were in this cherry-pick phase where everyone would pick their best images. You look at those and they're great. But actually, what's more important now is the worst image. We're in a lemon-picking stage, because every model can cherry-pick images that look perfect.

So now I think the real question is: how expressive is this model, and what's the worst image you would get given what you're trying to do? By raising the quality of the worst image, we really open up the amount of use cases for things we can do. There are all kinds of productivity use cases beyond these immediate creative tasks that we know the model can do, and I think that's a direction we're headed. If these models can do more things reasonably, then the use cases will be far greater.

Oliver Wang

So that's the moral equivalent of the monkeys on typewriters, basically: any model, given enough tries, will eventually make an amazing image.

Nicole Brichtova

But the other way around, it's hard.

Oliver Wang

Yeah, the other way around is hard. One monkey writing a book would be very hard.

Nicole Brichtova

It would be a good monkey for that one. [laughter]

Oliver Wang

What are the applications you think would come out when we reach the lower bound?

Nicole Brichtova

So, the one I'm most interested in—we mentioned this before—is education and factuality.

Oliver Wang

I don't know how many times a month I want to use these models for creative purposes, but I have way more use cases for information-seeking, factuality, learning, and education-type use cases. So I think once that starts working, then it'll open up all these new areas. Amazing.

Nicole Brichtova

There's also something about, I think, taking more advantage of the model's context window. You can input a really large amount of content right into these LLMs. Some companies—you mentioned a few before—will have 150-page brand guidelines on what you can and cannot do, right? And they're very precise: colors and fonts, right?

Oliver Wang

And the size of a LEGO brick, maybe. Being able to actually take that in and follow it to a T when you're doing generation—that's a whole new level of control that we just don't have today, to make sure that you're actually following that to a T. I think that will build a lot of trust with very established brands, where we have a second creative compliance review model that then double-checks everything that I could do against the guidelines.

The model should do it on its own, right? It should have this loop: I generate this, but then page 52 says that I shouldn't have done that, and I'm going to go back and try again. Then 2 hours later, it will come back to you with that respect.

Nicole Brichtova

Yeah.

Oliver Wang

And we saw with the text models how much this inference-time scaling can help, right? Being able to critique your own work. Yep.

Nicole Brichtova

So this feels really important.

Oliver Wang

Boy, an incredibly, amazingly exciting future for image models.

Nicole Brichtova

Yes. And congrats on all the amazing work.

Oliver Wang

Thank you.

Nicole Brichtova

Thanks for having us.

Oliver Wang

Well, thank you so much for coming on the pod.