[BidClub_]
The Cognitive Revolution · · 59 分钟

模型吞噬脚手架:DeepMind 的 Logan Kilpatrick 与 Tulsee Doshi 谈 3.5 Flash、Omni 等

Erik TorenbergNathan LabenzLogan KilpatrickTulsee Doshi

YouTube
TL;DR
  • Google 以 Gemini 3.5 Flash 打头阵,是因为它同时强调经成本调整后的性能与前沿模型级质量。Tulsee Doshi 表示,响应变慢会拖累 Search 和 Gemini 的实时实验,“即便模型在能力上强得多”,而 Flash 的速度约为其他大型模型的3倍,价格也显著更低。她提到 Artificial Analysis 上“我想是每秒280个 token”的成绩——快到 agent 可能在她来得及取消前就已经跑完。
  • 没有 Ultra 品牌,不代表 Google 停止扩展前沿能力。Logan Kilpatrick 认为,模型命名归根结底是“营销”:Pro 已变得更大、更强,Deep Think 则增加了测试时扩展;在内部,Pro 影响 Flash,Flash 又通过蒸馏影响 Flash-Lite,同时配方也可以向上扩展。Google 的产品组合反映了其同时服务多个产品和数十亿用户的需求,因此 Flash 和 Flash-Lite 的经济性,与高端质量需求同样重要。
  • Google AI 技术栈的战略中心,正从独立模型转向“模型—编排框架—产品共生”。Antigravity 正成为 Gemini Spark、AI Studio、Agents API 以及未来更多 Google 产品共享的 agent 基础设施——继 Gemini 本身之后,新的跨产品主线。Logan 对这一方向的概括是“代理、代理、还是代理”(agents, agents, agents),由框架负责工具调用循环和编排,而这些工作过去需要各产品团队自行搭建。
  • 共同训练的框架可能加深迁移成本,但 Google 明确表示会保留“框架多样性”。Labenz 提出的、与投资者更相关的挑战是:经过优化的模型—框架组合,可能形成彼此割裂的孤岛,提高黏性、迁移成本和定价权。嘉宾希望两者兼得:无缝衔接的 Gemini 全栈体验可以产生更好的训练和评测数据,同时模型仍能泛化到企业自有的编排系统。
  • 可持续的运营优势,可能来自基础设施标准化,以及 Google 产品版图内部的反馈闭环。Logan 表示,AI 技术栈如今每“12至18个月”就需要重写一次,因此共用基础设施不可或缺;Tulsee 则称,产品会暴露各种毛刺,这些问题会回流到模型评测和数据中。“模型吞噬脚手架”(the model eats the scaffolding)概括了这一循环:原本由外围代码实现的能力,会逐步迁移进模型或共享框架。
  • 递归式自我改进已是 Gemini 开发的一部分,但 Google 对近期前景的定义仍是由人主导的协作。Gemini 可以提交代码修改、发起评测、提出研究改进建议,并支持并行消融实验;据称,一位安全与对齐负责人曾在热水浴缸里用手机跑这些实验,并在1小时内产出报告。Logan 表示,近期让一个自主“ML 实习生”发起成本极高的预训练任务并不现实:错误方向的机会成本,仍会让“人坐在驾驶位”。
  • Google 将 Gemini 描述为协作者,同时把看似的心理痛苦视为模型或产品缺陷,并质疑福利访谈是否能揭示模型的内在状态。团队会在每个 checkpoint 上逐一评估迎合、角色扮演、循环和钻牛角尖等行为;Tulsee 称模型跑偏是“模型 bug”。她对模型福利访谈持怀疑态度,认为模型无法观察研究者询问的那些部署环境;缺少相关上下文时,它们大多只是在“高谈阔论”。
  • Google 押注的不是最大化原始上下文长度,而是搜索 grounding、上下文压缩和速度。一次100万 token 的请求在某些情况下要花“几美元”,需求因此受到限制,而且巨量上下文中有相当一部分反而会造成干扰;更理想的方向,是在正确的时间选出正确的信息。Diffusion coding 仍是活跃研究方向,但 3.5 Flash 已经快到足以让 Google 开始测试:继续提速究竟何时会出现边际收益递减。
摘要 · 为研究而整理的核心内容

1. Flash 通过优化每美元、每秒的智能产出赢得首发位

  • Labenz 从2023年“无护城河”(No Moats)的焦虑反转切入:此后3年,Google 的市值增加约3.5万亿美元。节目开场他还提到,Google 在2024年至2025年间将年收入增加了500亿美元,并仍控制着全球25%的算力;这些是他判断其护城河依然稳固的背景。

  • 面对 Google 约4000亿美元的年收入,Labenz 质疑公司为何选择先发布 Gemini 3.5 Flash,而不是 Ultra 级模型。

  • Tulsee 的回答从消费者行为出发:Flash 位于“既足够聪明、又足够快、成本效益也足够高的甜蜜点”。在 Search 和 Gemini 应用中,延迟上升会明显损害实时实验,“即便模型在质量上强得多”,因为“用户不愿意等”。

  • Google 给出的性能卖点很明确:Flash 的速度约为其他大型模型的3倍,价格显著更低,并针对编程和 agent 工作流设计。Tulsee 表示,3.5 Flash 在 Artificial Analysis 上的基准成绩为“我想是每秒280个 token”。

  • 因此,Google 的产品线覆盖 Pro、Flash 和 Flash-Lite,Deep Think 则提供运行时扩展这一维度。之所以加入 Flash-Lite,是因为服务 Google 规模的消费者群体催生了对更快、更便宜档位的需求;这一档位未必是最初 2.0 系列规划的一部分。

2. “Ultra”部分是品牌问题,而非扩展能力的上限

  • Labenz 的反驳值得保留:一些客户在购买单一最强模型时似乎对价格并不敏感,因此推出 Ultra 在商业上看起来顺理成章。他还追问,是否存在一个隐藏的超级大模型负责训练 Pro,再由 Pro 蒸馏出 Flash。

  • Logan 的回答是:“说到底,这些东西的命名就是营销。”Pro 模型一直在规模、能力和测试时算力上持续扩展;Google 反复讨论扩大后的版本是否应使用 Ultra 这个名称,但最终更倾向于保持系列之间的一致性。

  • Tulsee 确认,配方迁移是双向的,而不是简单的层级关系:“Pro 影响 Flash,Flash 影响 Flash-Lite”,但 Flash 的配方也可以向上扩展到 Pro。团队正同时推进预训练、后训练、推理技术、子代理系统,以及与 agent 框架更紧密的整合。

3. Antigravity 将代理变成 Google 下一层共享产品层

  • Logan 对 I/O 的概括是“代理、代理、还是代理”(agents, agents, agents),其基础是新的“模型—编排框架—产品共生”。模型会随着框架一起变化,而框架随后为 Gemini Spark、AI Studio 体验和 Agents API 等 agent 产品提供动力。

  • Gemini 是第一个贯穿 Google 庞大产品体系的共同主线。随着产品默认具备 agent 能力,Antigravity 计划成为下一条主线,提供可复用的编排能力,而不是让每个团队各自重建工具调用和 agent 循环。

  • Tulsee 表示,3.5 Flash 将覆盖 Gemini 应用、Search 中的 AI Mode、Antigravity、AI Studio 和 Spark。这种覆盖面本身也是约束:同一个模型必须同时服务消费者、开发者、Search 和 Cloud 用户,而他们的目标存在实质差异。

  • 模型和框架越来越协同设计,但实现仍保持可扩展,因为“大家做的并不是同一种产品”。不同团队可以定制共用基础设施,同时将反复出现的模型缺陷反馈给 DeepMind。

4. 全栈优化抬高锁定问题,Google 无法回避

  • Labenz 把框架比作包裹在通用知识工作者周围的“机甲服”。他提出的问题是:深度共同训练会不会终结今天模型与基础设施混搭的市场,转而形成彼此割裂的技术栈,带来更高的迁移成本、更强的黏性,以及前沿实验室的定价权。

  • Logan 更希望看到优化与选择权并存:Google 的“全栈 AI 故事”应让 Gemini 的表现异常出色,但模型也必须在其他框架上保持可用,因为开发者和企业客户需要灵活性。

  • Tulsee 将这一要求称为“框架多样性”。共同开发会形成飞轮:数据采集更干净、调试更快、评测更强、迭代更快;但 Gemini 仍必须支持客户各不相同的工具和编排架构。

  • Labenz 提议用 HarnessBench 衡量这种泛化能力。他的测试呼应了 Demis 关于游戏的框架:如果一个自称通用的模型只能在某一套优化脚手架内工作,那证明的不是稳健能力,而是“锯齿状智能”(jagged intelligence)。

5. 递归改进已成真,但昂贵下注仍由人来分配

  • Logan 表示,Gemini 已经“相当深度地”参与 Gemini 自身的改进:它能提升研究人员生产力、提交代码修改、运行评测、提出研究改进建议,并贡献最终改变模型的工作。

  • 最有冲击力的例子来自安全与对齐负责人 Anca。她曾从热水浴缸里发消息说:“我可以用手机跑完所有这些消融实验。”她描述了对 Gemini 做消融、测试问题、比较 SIs 与数据消融,并在不到1小时内产出一份报告。

  • Logan 的限定是经济层面的,而非哲学层面的。近期让模型自主进行预训练并不现实,因为大型任务会消耗海量资源和稀缺算力;让 AI 选择错误方向的机会成本,决定了“人仍坐在驾驶位”。

  • 不过,角色分工确实在变化。自动化让研究人员可以花更多时间解读结果和选择策略;Gemini Mic 则让用户可以围绕一个想法自由讲述,再由 Antigravity 将语音转化为行动和内容,例如幻灯片或经过推理的输出。

6. 标准化加速产品迭代,产品反过来训练下一代模型

  • Logan 对基础设施的教训说得很直接:“现在每12至18个月,你都得从头重写一遍。”服务技术栈已经从 token 输入、token 输出,发展到工具调用循环和 agent 编排;如果每次都让众多产品团队分别重写,成本将变得无法承受。

  • Tulsee 强调迭代速度:把有潜力的 checkpoint 交给内部开发者,快速制作 prompt 原型,运行实时实验,构建能产生有效信号的评测,并在真正运行的产品中发现实验室里看不到的缺陷。这些缺陷随后会成为下一轮模型的新数据和评测目标。

  • NotebookLM 是典型案例。最初 Audio Overview 中那段出人意料地连贯的对话,本质上是“基础 Gemini 加上一堆很强的 prompts”;团队知道如何对模型进行“prompt-whisper”,让它发挥最佳表现。

  • Labenz 将这一方向总结为“模型吞噬脚手架”(the model eats the scaffolding)。每当能力发生跃迁,原本由定制 prompt 或外围软件提供的功能,就可能被吸收到模型或共用框架中,让产品团队更专注于产品本身。

7. Omni 将原生多模态推进到视频,但 API 仍未定型

  • Tulsee 介绍 Gemini Omni Flash 时,将其描述为朝“全模态输入、全模态输出”(all modalities in and all modalities out)迈进,起点是视频生成和编辑。用户可以组合多个场景和源图像,将自己的 avatar 放进视频,并利用 Gemini 的世界知识和推理能力生成原生视频。

  • Labenz 称其为视频领域的“Nano Banana 时刻”,意思是语言、推理和像素理解仍然保持一体,不需要经过有损的文本中间层。Tulsee 也认可这一说法,并将其作为内部标语。

  • 分发计划比开发者接口更明确:Omni 初期会进入 Gemini 应用、Flow 和 YouTube,但“目前还没有 API”。现有 Gemini 视频输入会按照可配置的 FPS 参数对画面进行下采样;未来 API 的行为和采样决策尚未确定。

  • 音频也在同步推进。Gemini Live 速度更快、智能程度更高,也更擅长识别背景噪声;Logan 认为,音频将越来越成为人类与模型互动、模型反向回应人类的主要方式。Tulsee 将其称为前沿模态。

8. Gemini 的人格被作为产品行为评估,福利主张仍不确定

  • Labenz 将 Anthropic 对 Claude 的处理方式,与将 Claude 视作某种心智存在的倾向进行对比,同时提到 OpenAI 偏工具导向的 Model Spec。Tulsee 表示,Google DeepMind 内部存在多种观点,但在产品层面的共同定义是:Gemini 是工作和日常生活中的“协作者”与伙伴。

  • 当被问及“心理痛苦”时,Tulsee 将问题重新框定为一个长期相处的伙伴应如何沟通。每个 checkpoint 都会接受关于迎合、角色扮演、循环和钻牛角尖的安全评测,并通过版本发布之间的对比来识别行为恶化。

  • Tulsee 明确表示,Gemini 跑偏就是“模型 bug”;Logan 补充说,这种行为并非预期结果。用户反馈之所以重要,是因为团队希望识别并消除这些模式。

  • 对于模型福利访谈,Tulsee 的怀疑在于,模型通常并不了解自身训练或部署的任何上下文。询问它们对这些系统有何感受,会把它们推到分布之外,结果只是根据广泛的训练模式“高谈阔论”。

9. 搜索 grounding 与上下文压缩比原始窗口无限扩张更重要

  • Labenz 表示,据他判断,公开发布的 Gemini 模型仍然有2025年1月的知识截止时间,但在近期研究任务上表现良好。Tulsee 接受用“bug”来描述这一点;Logan 表示更新截止时间很重要,同时解释说,搜索 grounding 让模型能够区分参数化知识与可能仅在几分钟前出现的信息。

  • Logan 认为,Google Cloud 与 Exa 的合作并不意味着 Google 放弃 Search。Cloud 已经提供相互竞争的模型和服务,因为企业客户需要选择;与此同时,Gemini 与 Search 的直接协作仍然很深入。

  • Tulsee 对上下文窗口的判断是要“聪明地”使用上下文。个人档案和大型代码库可能包含海量信息,但其中很多内容会造成干扰;通过筛选和压缩,可以暴露真正有用的子集,同时让有效信息池超过当前活动窗口。

  • 经济性进一步强化了这一架构。Logan 表示,在某些情况下,一次100万 token 的请求可能花费“几美元”,相较其带来的服务算力负担,需求规模很小。他希望出现研究突破,但目前客户往往会出于成本和延迟考虑,主动选择更小的窗口。

10. Diffusion 仍在研究阶段,Flash 抓住当下的速度需求

  • Tulsee 表示,Diffusion coding“很棒”且“非常快”,研究人员仍在持续提升质量和速度,并测试它在哪些场景下已经足够实用、可以发布。Logan 则将此前的演示称为一次幕后展示,而不是已经确定的产品路线图。

  • Flash 提出了一个问题:进一步加速从什么时候开始不再带来有意义的用户价值,而只是边际收益递减。节目提到 3.5 Flash 的速度约为每秒280个 token;Logan 表示,3.5 Flash 的更快版本已经在 Antigravity 中运行。

Nathan Labenz

Hello, and welcome back to The Cognitive Revolution. Today, after some 340 episodes, I am very excited to share the first episode that I've ever recorded in person with fan favorite Logan Kilpatrick, member of technical staff at Google DeepMind, and Tulsee Doshi, senior director and head of product for Gemini models.

The occasion for this conversation is Google's annual I/O event, where they're launching the new Gemini 3.5 Flash model, all sorts of agent infrastructure and AI product integrations, and plenty more. We recorded on Friday, May 15, just a couple of days before the event. And while many at Google, including my brother Craig, who's giving a keynote on Wednesday, were working overtime to polish their demos and presentations, the overall vibe, at least compared to the rest of the AI space, was one of relatively relaxed confidence.

And why not? From 2024 to 2025, Google grew annual revenue by $50 billion, as much as Anthropic is pulling in today. And they still have 25% of all global compute, the deepest pool of research talent anywhere, and the most comprehensive AI portfolio of any company, with top-tier positions not just in language models, but also self-driving cars, medical and life sciences, and robotics.

After discussing the headline launches that they're announcing this week—which also include a new video-generation model called Omni, which they hope will create a Nano Banana moment for video; a new, improved, and more agent-focused Antigravity; and a product called Spark, which will bring more agentic functionality to the main consumer Gemini app—I really wanted to take a step back and dig in on Google's overall AI strategy and philosophy.

We discussed their decision to lead with the Flash model and, more generally, to emphasize the cost-adjusted performance Pareto frontier, whereas Anthropic and OpenAI are clearly much more focused on competing to have the single most capable model in absolute terms. We talk about how DeepMind is no longer shipping models in isolation and leaving it up to product teams to figure out how to use them, but instead now providing a robust agent harness, which should help elevate and standardize AI experiences across Google's vast product surface.

We get into the weeds on questions like why context windows seem to have mostly stopped growing, why Gemini models' knowledge cutoff is now more than a year ago, and whatever happened to that diffusion-model line of work. Perhaps most importantly, we discuss how the team at Google relates to the AIs they're creating, how they're thinking about things like model psychology and welfare, and their views on recursive self-improvement, which, as you'll hear, is definitely a part of their plan, but not something they seem to be as singularly focused on as other AI leaders.

Overall, I think this is a great window into the thinking that underlies Google's AI research and product development, which has clearly sustained the company's historic run far beyond the point that many analysts had written them off. With that, I hope you enjoy my first-ever in-person conversation with Logan Kilpatrick and Tulsee Doshi of Google DeepMind.

All right. We are here live at Google headquarters, in the library at Gradient Café—the first-ever in-person recording of The Cognitive Revolution. Logan Kilpatrick and Tulsee Doshi, welcome.

Logan Kilpatrick

Thank you. This is an honor. I didn't realize this was the first in-person episode.

Nathan Labenz

Yeah, 350-plus. That's us. And it's all been from my home office in Detroit until today.

Tulsee Doshi

That's awesome.

Nathan Labenz

Well, thank you for being here. This is a crazy space, especially around I/O.

Logan Kilpatrick

It's a zoo.

Nathan Labenz

Yeah, it's always a good time here at Google HQ. So, you may or may not remember the No Moats memo. We've just passed the 3-year anniversary. It was May 5, 2023. And in the intervening 3 years, Google has added $3.5 trillion in market cap, which is more market cap than all but 2 other companies in the world. Those 2 are NVIDIA and Apple.

So, the moats, I'd say, are holding up. Here we are at I/O, and I'm sure there are going to be some exciting new things that will be deepening the moats. First question: Tell me, what are we launching this week to try to deepen those moats?

Tulsee Doshi

A lot. A lot of exciting stuff. Let's start with some of the modeling side of things, because that's really exciting. We have our 3.5 series coming out, starting with 3.5 Flash at I/O.

We're really excited about 3.5 Flash because I think Flash does this really awesome job of being at the sweet spot of being really smart while also being really fast and really cost-effective. Flash is incredible. It is 3 times faster than other large models. It's significantly cheaper while still being able to drive these really awesome agentic coding workflows. And we've been using it internally a lot, which has been really fun to see play out.

That's one big piece: 3.5 Flash. We're really excited about that. We're also releasing Omni, Gemini Omni Flash, which is a video-generation and editing experience. What's really exciting about Gemini Omni in general is that it's our push toward being able to bring all modalities in and all modalities out.

The first way this is really manifesting is in this video-editing context. You're going to be able to make really awesome videos. You're going to be able to put your own avatar into the videos, which is going to be awesome. I've been having a lot of fun playing with that, too.

We're continuing to upgrade Antigravity and bring more into the developer experience, so Logan can talk more about the developer experience overall. But 3.5 Flash and Antigravity are really going to come together to build something great there, too. Then, in about a week or two, coming soon is Gemini Spark, which also builds on 3.5 Flash to build more agentic experiences into the Gemini app. So, we've got a slate of cool things coming.

Logan Kilpatrick

Yeah, I think beyond the models, the other headline of the story is just agents, agents, agents. The meme of Sundar from 2 years ago—or last year—saying “AI, AI, AI” all the time: I feel like this year is “agents, agents, agents, agents.”

I think it's cool to see this. This is the first year where we have—and actually, I think this is not just us, but ecosystem-wide—this model-harness-product symbiosis that's taking place. The model is chaining with the harness, and the harness is powering the agentic product experiences. Gemini Spark in the Gemini app is one example of that.

It's powering the vibe coding in AI Studio. It's powering the Agents API for developers. It's powering, I think, something else. Maybe it'll also roll out to other products across Google and sort of be this foundational layer to build on top of, which is really exciting.

So, I think it's not just developer products, but our consumer products, and probably even more widely in the future across the rest of the Google product suite. I think there's actually one interesting thread of this: Historically, Google didn't have this throughline of something that carries across all of our products. I think then it was Gemini, and then, all of a sudden, every Google product had Gemini, sort of getting them all stitched together and making all those product experiences great.

I think now you're seeing that again with the Antigravity agent harness. As products become agentic by default, you have the Antigravity agent harness being another throughline through all of our products, which is really interesting.

Tulsee Doshi

Yeah, and I think what's been really fun from the modeling standpoint there is one thing that we did with Gemini 3 that we're really continuing, I think, with the 3.5 line: really bringing the model to all of our products.

So, 3.5 Flash will be in the Gemini app. It will be in AI Mode in Search. It's also powering Antigravity. It's also powering agentic experiences in AI Studio and Gemini Spark.

I think this idea of how do we build a model, and then how do we build it in partnership with the harness such that it actually works across all of these product surfaces—which actually have very different users and very different goals—is really awesome.

Logan Kilpatrick

It's really hard, too. I think it's actually gotten harder to do. It was maybe tongue-in-cheek; it was kind of easy before because you just launched the model on a couple of services and it wasn't that bad.

I feel like now you have the constraints of the very wide array of Google products that are just for totally different users. And I think, actually, credit to the model team for trying to find the fine line for all these different places, because we're not just building for Search. We're not just building for developers. We're not just building for cloud customers. It's not just for the Gemini app. It's all of them at the same time, which is an exceptionally large amount of work to pull off that story on a consistent basis.

From Gemini 3 forward, that has been the story, which is exciting. I think one other thing I'll say is that what's cool about this I/O, too, is what we're doing across modalities.

Tulsee Doshi

So, Logan said “agent, agent, agent,” which I think is true. This I/O is really about bringing models to action in real-world use cases. But I think what's also cool is that we have the Flash model, which is really about building these kinds of coding and agentic use cases.

We have Omni, which is really about what multimodal vision looks like.

And then also, actually, Gemini Live is getting an upgrade, too. Gemini Live is getting faster, it's getting smarter, and the model is much better at detecting background noise. So it really does feel like a partner in a lot of ways, and I think it's cool that we're also able to draw this throughline across the different ways you might want to interact with a model and consume content, which I think is also really cool.

Nathan Labenz

Okay, I've got 7 different directions I want to go and a lot of follow-ups, but I'm trying to seed them all for you. Let's start with just the model. It's interesting to start with Flash. One thing that I recall—I don't know if it was 2 I/Os ago or whatever, right?—was that there were going to be 3 sizes of the Gemini model at one point in time.

Logan Kilpatrick

And there are 3 sizes. Yeah: Pro, Flash, Flash-Lite.

Nathan Labenz

Yes, we never saw Ultra. That's kind of what I'm alluding to.

Logan Kilpatrick

Yeah, 3 were promised, and then we added—we took one off the top and added one at the bottom.

Deep Think, too, actually, which is like a 4th scaling dimension from a model perspective.

Nathan Labenz

That's a runtime scaling, though.

Logan Kilpatrick

It is. It is, for sure.

Nathan Labenz

Well, I guess I have 2 questions about why no Ultra. One is: Is it a compute limitation? How are you guys thinking about which model to release?

Google enjoys $400 billion a year in revenue, I was interested to learn. Traffic, though, is growing extremely fast. They might hit 100 billion at the end of this year, maybe even in the 3rd quarter—who knows? And it seems like the revenue there is really driven by people's extreme willingness to pay for the very best model they can get their hands on. Maybe not at any cost, but relatively price-insensitively.

So I'm wondering: Why no Ultra? That seems like it would be a killer, and yet we haven't seen it.

Tulsee Doshi

I didn't plant this question. I'm always digging my own political holes here on this side. I wanted one. This is my favorite question, so I'm glad we got to it today.

No, I think you're right that there is a slice of users who are definitely willing to pay for a certain level of quality. And I think we really do believe that the Pro model has been pushing that quality. But I also think we've seen so much value from the Flash and Flash-Lite dimensions because we also see an extremely large number of users, especially if you're thinking about building consumer applications.

If you think about the Gemini app, if you think about Search, when you're serving at that kind of scale, latency really matters. Cost matters, because you find that users aren't willing to wait. We find that even when we tweak the model and hurt latency, we see that play out in our live experiments on Search and in the app, even if the model is hugely better from a quality perspective. What you're asking users to do is wait.

And so I think part of the reason why we ended up introducing this Flash-Lite SKU, which wasn't necessarily part of the original 2.0 series, was because we really felt that there's a large-scale demand for this depending on the types of use cases, especially when you're talking at that scale. For us, it's really important that we're pushing the full range of what kinds of customers we can serve, both internally and externally.

For our products, the Flash and Flash-Lite SKUs matter a lot for our ability to serve the Google populace. We also imagine that that's true for external enterprise customers and developers, and I think that's played out to be true as we've been seeing this in action.

Logan Kilpatrick

Yeah, I think the 2 things that I'll add are that there's probably a more nuanced technical story on the Ultra thread, but it's not like the Pro models haven't scaled up over time. There is a story that you can spin. At the end of the day, the naming of these things is marketing.

They definitely are getting extremely capable. They're getting larger, they're getting more powerful. There's test-time scaling, test-time compute scaling with Deep Think, et cetera, et cetera, and all types of stuff in that dimension. So I think it is possible you could put the Ultra brand on some of these things. I think the decision so far has been not to do that, but it hasn't been that we haven't kept scaling up.

Tulsee Doshi

Yeah, there's almost been a conversation every time we scale up: Should we call it Ultra, and what does that brand mean? We could, but there's also a question of how we keep consistency for users, also from series to series.

I think, actually, to rearticulate a point that you also made, Google—and specifically Google DeepMind's—mission is to build AI responsibly and make sure it benefits all of humanity. That is so deeply tied to the Google product surfaces in which we're serving, what is it, 8.2-plus billion user products or whatever it is.

At the same time, obviously the frontier matters. Obviously, having great models that are really expensive and really, really intelligent matters, and there are tons of use cases for that internally and for our customers. You also need to do the scaling up to billions of users for us to actually do the thing that Google needs to do to achieve the mission.

I feel like we've done a good job, hopefully, of trying to walk the fine line of continuing to push the frontier and build great Flash models. I actually think those 2 things are more tied together. You know this better than I do, but they're more tied together technically. It's hard to make great Flash models if you don't have a great Pro model, and vice versa. So we'll definitely keep pushing the frontier on both of those things.

Nathan Labenz

I think the perception from outside is, by analogy to “It's hard to make a good Flash model if you don't have a good Pro model,” that there's an Ultra model internally that's the mega training run being used to help train Pro, which maybe in turn is being used to help train Flash. Is that true? Is there a bigger thing inside that's only for distillation to the midsize models?

Tulsee Doshi

We definitely use distillation as a way of bringing down our sizes. So you will see that Pro influences Flash, which influences Flash-Lite. We also do the reverse, where we scale up. You take the Pro, you take the Flash recipe, and scale up to the Pro recipe, for example.

I think what's been really fun, especially as we've used Antigravity, is the point Logan made about the harness. We've been seeing a lot of examples of leveraging pretty awesome models to drive progress internally. One thing Varun demos on stage on Tuesday is basically being able to leverage a bunch of subagents to complete a bunch of tasks and come back. You can actually try that as an early preview in Antigravity today. If you go to /teamwork.

That's an example of something we've been using internally, which is an extremely smart model. It leverages a combination of the best of Gemini 3.5 as well as inference techniques, and you're able to accomplish so much. I think that's the kind of direction I'm excited for us to go into more.

We're pursuing all of these fronts. We're scaling up from the pretraining and frontier perspective, and that's continuing to show gains. There's a bunch we're doing on the post-training side, and then there's also a bunch we're pushing on on the inference side. Plus, we're trying to make sure we're working with the harnesses. I think we're going to keep getting things that we're using internally and then start pushing out externally through previews.

Nathan Labenz

So, let’s talk harnesses. It seems I was just talking to Andrew Lee, who’s the founder of Taskade, the other day, and he said, “Fundamentally, everyone these days is building the same thing. They’re all trying to build the general-purpose drop-in knowledge worker.” And so, that’s got to have the intelligence at the core, and then it’s got to have all this—he calls it the mecha suit—that is built around it. This harness sounds like the mecha suit that you guys are developing in-house.

I guess the first question is: Is this going to create silos? We’ve lived in this world so far where I could mix and match my models and my infrastructure, right? I could go to LangChain, or I could use Taskade, or I could use whatever, and I could pick whichever model and plug it in. But as they get more deeply co-trained with the harness, does this create siloed worlds where you’re all in on one frontier model company’s stack or another? And if so, that would have pretty significant implications for the switching costs, stickiness, and pricing power of the frontier model creators. What’s your take on how sticky things are going to get?

Logan Kilpatrick

It’s a good question. I think, again, Tulsee probably knows better than me on this, but I think the best case is that you can do both. The best case is that it works really well for Gemini, and we can do the things we want to do and scale up because we do have control over the full-stack AI story, as Sundar likes to say, but then it also generalizes across other stuff.

I think the developer ecosystem—people want choice, and people want to have flexibility. These tools have lots of use cases. Actually, there are philosophical questions of how good your model really is if it can’t generalize to other harnesses.

Tulsee Doshi

Yeah, I think that’s right. I fully agree. Maybe to double-click on what Logan said originally, the benefit of the full stack that we have is that we can hopefully build a really seamless experience, right? You get the best of Gemini, you get it working in the most effective ways for you, and you get it working in a way that is intuitive, smart, and fast. And so that also helps us train the model to be better, right? So this becomes this flywheel that continues to power the model.

At the same time, I don’t think we want it to only be the case that the model works in a single harness, right? We want any of our enterprise customers or a developer who’s building their own use case to be able to leverage Gemini effectively. And so it is important, from a model standpoint, that we’re training in such a way that we actually—we call it harness diversity, right? We should be able to support a range of different approaches to tooling, different approaches to orchestration, et cetera.

But I think what’s helpful about this approach of co-training and building that flywheel is that it’s easier to debug, it’s easier to think about data collection, and it’s easier to eval. You can just move at a faster pace. And I think we’re seeing that across the industry. Finding that balance is important, but I think it just helps make the model better.

Nathan Labenz

Yeah, this is also a good pitch for HarnessBench. If that’s not a benchmark that exists, let somebody build HarnessBench. I would love to collaborate if folks are interested in that, because I do think it’s a great test of Demis’s perspective from games, actually, as an example. If models are so good, why can’t they play games really well? And if models are so good and we’re actually approaching AGI, even if you do the model-harness training symbiosis, you’d still expect them to generalize reasonably well on other harnesses. If they can’t, that’s another sign of the jagged intelligence.

I think it’d be cool to see this play out from an actual benchmark perspective. It could also perhaps be productized as an RL environment and sold into you guys that way. Quite the cottage industry these days.

So, obviously, the other big thing that I think is very much in the air is recursive self-improvement. The reason I’m here this weekend, when we originally planned to do this remotely, is that I’m going to an event called Recursive, where the topic is going to be recursive self-improvement and, hopefully, how we can navigate it successfully. How bought in is Google DeepMind to recursive self-improvement? When you talk to Anthropic people, they’re almost religious about it, and they also see it as totally inevitable. OpenAI has timelines for later this year and early 2028 for an ML intern and a full-fledged AI R&D employee. Do you guys have milestones or timelines for when you’re going to hand off ML research to AIs?

Logan Kilpatrick

We’re already using Gemini pretty deeply internally to improve Gemini. And so I think that is very much a theme for us: How can Gemini actually be a part of the Gemini development process? That can include things—I think it goes the full range from helping us be more productive, which is obviously the simplest part of this, to actually submitting CLs that would run an eval, suggest a research improvement, and drive improvements to Gemini itself.

And I think there are a lot of ambitions we have to keep pushing in that research direction. So I think, very similarly to the other labs, this is very much an area of investment for us and an area we’re super excited about.

Tulsee Doshi

I think, for me, what I’m really excited about is this really awesome research-partner opportunity that we have with Gemini: for it to help us with creative ideas and help us test things faster. Actually, it was also one of my co-workers, Anca—she’s our lead for safety and alignment—and a couple of days ago, she pinged me from her hot tub.

She was like, “I could run all of these ablations from my phone, because I could kick off a bunch of things to actually ablate Gemini, to test for a bunch of these issues, and to see how some of our SIs differ or how some data ablations differ. And here’s my report. I could do all of this in the last hour.” And that is amazing. That’s the kind of thing that we can already do, right? So then imagine where we’ll be in 6 months, a year, or 2 years from now.

Logan Kilpatrick

Yeah, I feel like, at least from my personal perspective, it’s a much more practical perspective. Obviously, as models get better at coding, they’re going to do things that are code-related. They’re going to help us build our products, and they’re going to help us train models. I think all the nuance of the story is in where the human is in the driver’s seat of this stuff. The tools are built for the human to be in the driver’s seat, which I think is an important thing as we continue to go forward.

Also, I think very genuinely—and I think the model team and the researchers feel this more than ever—the near-term horizon is going to continue to be the human in the driver’s seat, because the cost of these runs and the opportunity cost of going in the wrong direction and putting a bunch of resources into it are super, super high. And so I find it doesn’t seem super realistic in the short to medium term that you’re going to just be letting large-scale pretraining jobs be kicked off by the ML intern, and it’s going to cost you X many dollars and lots of compute, taking it away from the human researchers. But this deep collaboration between AI and human researchers, I think, is super obvious.

Tulsee Doshi

Yeah, there’s also something really amazing about how much that collaboration allows you to focus on the interpretation of what you’re seeing in the results and where you really want this to go strategically. And so it changes a little bit of the role that the human can play, which I think is also really powerful for our teams.

Nathan Labenz

When you’re doing research, are you actually typing any code these days?

Logan Kilpatrick

So it’s interesting for me on the product side. On the code side, for any code that I was already submitting, I’m mostly relying on Antigravity and doing bits and pieces myself. But it’s also been really cool to start having the model generate slide decks and to start generating actual content from my thoughts.

We actually introduced Gemini Mic in Antigravity today. There’s this really awesome feature—I don’t know if you’ve been playing with it internally—where you basically ramble at the model. You share a bunch of your thoughts in whatever loose form they happen to be, and then the model leverages that to take action.

For me, I’ve been finding that so much more powerful because I actually feel like I think a lot by talking. It’s a very cool moment where I can be like, “Okay, I’m just going to sit here, tell you what I’m trying to think through in my head, and then have you actually bring that back to me in a way that is reasoned and well thought out.”

Nathan Labenz

Yeah, I feel like this correlates so well to something I’d love to see: a breakdown of human-typed code versus AI-generated code, versus—and maybe there’s a divergence—which is audio input that then generates code. I feel like audio input leading to generated code has got to be one of the fastest-growing input modalities in what’s happening.

I find myself doing this all the time, and it is the predominant way that I’m building software, at least when I’m not around a bunch of other people. I’m still typing things in when I’m around other people so that I’m not rude. They don’t hear my dumb ideas and the things that I’m trying to do.

You see, if you walk around upstairs sometimes, you’ll see people muttering into their hands. We were just [laughter] because they’re now actually talking to create, which I think is pretty cool.

Tulsee Doshi

It is cool. One of my KPIs for myself this year, to really know if AI is improving my life, is whether I’m getting outside more and getting more exercise. I’m starting to, maybe, a little bit. I wouldn’t say I’ve won the game just yet, but I still want to be able to get my thoughts out. So I think that is absolutely the frontier modality for me.

Nathan Labenz

So, with the harness, you said it’s now becoming this through line that’s going across all Google product surfaces. As I’m sure you’re well aware, commentary on Google’s AI integrations across its vast product suite has been that it’s characterized by some bangers, and then there have been some that have been characterized as misses.

Presumably, one of the benefits of the harness is that it’s going to make it a lot easier to have a more standardized approach and a general high-quality bar across all these integrations. What would you say people should learn from the experience that Google has had to raise its own bar as they go try to do these integrations themselves?

Logan Kilpatrick

I think this is actually such a great story for us. Very practically, Google has done a ton of this infrastructure standardization across the AI stack over the last couple of years, which I think has been awesome. And actually, one of the threads of how we’ve been able to land the Gemini 3 models across so many more products is this infrastructure standardization that happened.

We’ve gotten a lot from it. It’s painful and difficult, and there’s, of course, lots of work involved in doing it. But if you pay that cost, you actually do end up getting this. The advice for people who are in this position and thinking about this is basically that every 12 to 18 months now, you have to rewrite everything from scratch.

The best case is that you don’t want N teams rewriting everything from scratch every time the paradigm shifts. Historically, the infrastructure was just serving raw models: you get tokens in and send tokens out. Now, there’s a bunch of agentic infrastructure, and there are tool loops and all these other things happening inside the harness.

Again, you don’t want every team to have to go and reinvent that from scratch. You want innovation, but you don’t want every team to have to go and reinvent that from scratch. The fact that any team across Google that just wants to ship some really cool agentic product doesn’t need to think about the nuance of all the details of the tool-calling loop and so on is a huge acceleration for them to just focus on building a great product.

I don’t know if a lot of the agentic stuff we were landing at I/O would have been possible if we hadn’t had some of that infrastructure standardization across the harness and the model delivery.

Tulsee Doshi

I think the other thing I would say, as far as lessons learned, is that there’s really no substitute for being able to just experiment and iterate quickly. This goes to all of Logan’s points about the foundation being strong, but I really think what has helped us is being able to put in, for example, a new model and iterate really quickly with a product: What are the right prompts that would actually make this model viable for a different situation? What are the ways to prototype really quickly with this model? What are the ways to get it in the hands of even just internal users quickly, let alone external users?

That is something that is now more and more possible with layers that are consistent across the team. I think it’s pretty amazing to see the speed at which we can go from having a checkpoint that we’re really excited about to putting it in the hands of internal developers, then seeing it come to life in a product. Only when you see it come to life in the product do you really start finding its rough edges.

Nathan Labenz

That allows you to actually come to terms with how to do that. More and more, it becomes: How do you have the right ability to tune prompts quickly? How do you have the ability to run really good live experiments where you can get really good data and feedback quickly? How can you build evals that help give you real signal?

Those are the things that will speed up your progress on quality the most, because they’ll give you the ability to actually get to the kind of product that you love. And I think if you think about NotebookLM, that team really understands the model. They are—when you talk about a banger product, it comes from a banger team. They’re really good at being able to take the model and play with it quickly and prototype quickly to get to something amazing. And I think you see that actually play out in the product.

Logan Kilpatrick

The best example of this is the original audio overview experience. I think the thing that shocked people about audio overviews was the coherence of the dialogue. The coherence of the dialogue was just base Gemini with a bunch of banger prompts. They knew how to prompt-whisper the model and get the best out of it.

I think, obviously, the actual audio model was really good as well, but the prompted dialogue was really difficult for them to pull off.

Nathan Labenz

And they pulled it off in an incredible way. I think it helped people fall in love with that product. So, it sounds like one big lesson is modularizing it. It used to be sort of the model on one side and everything else that goes into the product on the other side, and we're pulling a lot of the surrounding code, architecture, and tools onto the model side.

Logan Kilpatrick

Model eats the scaffolding. That's my favorite way of thinking about this. At every crank of the model flywheel, the model eats a bunch of scaffolding.

Nathan Labenz

What happens when something isn't meeting somebody's needs? Do they do a little fork of it and submit a pull request back to the main scaffold team, or do they have to just say, “Hey, I've got a need here. Can you help me out?”

Tulsee Doshi

Yeah, it's definitely extensible. I think the nuance of this is that Spark, the way Spark is built on top of a bunch of this infrastructure, probably looks slightly different from the way YouTube is built on it. They're both running on the same set of infrastructure, but the nuance is probably slightly different.

There is this layer of extensibility that you get out of the box, which is great because obviously everyone's not building the same product at the end of the day. Extensibility is actually a first-class feature of any of these types of platforms that you want. It's the same thing, actually, on the model side.

I think one of the things that's really awesome about building Gemini within Google and having all of these different product teams is that there's always going to be something that doesn't work for them, right? There's always going to be something that can get better in the model experience. We're trying to build something in a product, and the amazing moment is when you start trying to build it and it doesn't work.

Step 1 is, “Okay, can I prompt my way out of it? What does that look like?” Then you start figuring out, “Okay, what are the losses, really? Where is the model falling down?” What we try to do as much as possible is keep these feedback loops with our product teams and say, “Okay, if this is where the model is falling down, how do we bring that feedback back to the model in terms of evals and data? What does that look like?” That way, in our next revision of the model, we can bring all of that feedback back in and iterate on it.

I think that's how you've seen Gemini get better: really from that feedback about where things aren't working. We try as much as possible to have the structure be: We train a model, we hand that to a wide range of teams, and those teams implement the model in their structures. They do a bunch of things, to Logan's point, because it's extensible, but they also find all of these places where the model falls down. We cycle that back, and I think that's actually been part of the fun part of the job, but also part of what makes Gemini work really well in some of these use cases.

Nathan Labenz

Let's talk about Omni for a minute. It sounds like this is going to be sort of the Nano Banana moment for video.

Logan Kilpatrick

I love that you're saying that because that is our tagline. Sorry, I didn't even have to say it. Great. By that, I mean that there's a deep integration between language and reasoning and pixel-space understanding, right? I have that kind of vision in my head from the Nano Banana launch: Here's a woman, and here's her breakfast and a cup of coffee, and now they're all in one image, and they all look like they did before. Clearly, that's not something that was done through a lossy language intermediary. The model understands images.

Nathan Labenz

So, we're going to see that now, I guess, for video. That sounds cool. Is it going to be available via the API? I've noticed that Gemini has been the only API that's accepted video for a while now, but I don't know exactly how it works under the hood, obviously. I do feel that it's been downsampled, or maybe there have been frames taken out of it historically. Is this going to be—

Logan Kilpatrick

There's an FPS parameter if you want. You can change how many frames you sample, but it does downsample the number of frames that are available. You can control it.

Nathan Labenz

Okay, so that's a pro tip for you. Nice. It sounds like that will still be the paradigm, though. It will be a frame-based selection on the input, but then it's going to be natively speaking video pixels on the output.

Logan Kilpatrick

It's a good question. I actually don't know. It's also not available in the API yet, so there are lots of things still to be figured out. We have to figure out what we want it to look like on the API side. I think maybe the heart of your question is native video generation.

What's exciting about Gemini Omni is that it really is building on all of the magic of Gemini. This whole Nano Banana for video is really about how we bring in all of the world knowledge and reasoning power of Gemini and actually generate native video as a result of that. We have to figure out how this manifests in the context of the API from a sampling standpoint, similar to a lot of the decisions we've had to make about Veo from a sampling standpoint.

As of now, you'll be able to use it in the Gemini app, in Flow, and in YouTube. Those are all going to be ways that we can start seeing how people experience the model and what value individuals are getting. With this Nano Banana for video, I think we're really excited for these types of things where you can say, “Okay, take some of these images, take the scene, and make these things all come together in one video.” I think it's going to be really awesome.

Nathan Labenz

Zooming out philosophically, you may have seen this Roon post not too long ago about Anthropic and the sort of relationship that the company, as he sees it, has with Claude, where he describes Anthropic as almost worshipping Claude, in a sense. Certainly, they treat it, including in the constitution, as sort of a being or a mind—something that they want to have a give-and-take relationship with.

OpenAI, on the other hand, has its Model Spec, which is like: This thing is a tool. It's supposed to follow these rules, and it's a more conventional relationship. How would you describe the culture within Google as it relates to Gemini? How do people feel about it? How do they talk about it? Is there any of this sort of being, entity, or other-mind desire for pushback from Gemini, or is it more of the simple tool?

Google's a very big place. There's a lot of people, so I'm sure you have a lot of varying perspectives.

Tulsee Doshi

To Logan's point, even within GDM, you're going to find a range of folks who will leverage Gemini differently. When it comes to how we think about it, we do have a strong point of view on the kind of behavior we want Gemini to have. We do really want to be intentional about how Gemini manifests itself to internal and external people.

I do think it's really about how Gemini helps Googlers and how Gemini helps people within Google and outside. It's really much about how we create good partnerships between Gemini and people. I think that's very much the ethos of what we're trying to build.

How does Gemini become that partner? We use the word “collaborator” a lot. How can Gemini be your collaborator, both in the code you're writing and in your day-to-day life and what you're doing? That's the ethos we're trying to bring in its behavior and persona, as well as in the kind of products we're building around it.

Nathan Labenz

Yeah. Do you worry about its psychology? There are all these examples from LLM-whisperer types and from people who are putting models to use. Anthropic has put Claude in charge of a café in Sweden, right? Those folks tend to report certain doom loops—Gemini getting really down on itself, getting really discouraged, seemingly feeling bad, if you believe there's any feeling inside of it. How much does that kind of stuff concern you? Do you care about seeing a reduction in psychological distress from one generation of Gemini to the next?

Tulsee Doshi

Yeah, it's interesting. I haven't thought about the phrase “psychological distress,” but I do think it really matters how Gemini communicates with you as a partner or user of Gemini. We have pretty extensive safety evaluations in terms of how Gemini engages with you, including things like sycophancy, role-play, looping behavior, and rabbit-holing behavior.

We look into a lot of that for every one of our checkpoints because it really does matter, especially as we're starting to use Gemini more and more. If you're using Gemini for hours a day, it really does matter that these attributes are well understood and well evaluated. We look at them launch over launch, right? We ask, “Okay, how does Gemini look from the perspective of sycophancy, for example, launch over launch?”

To be very explicit, those cases where the model does go off the rails are definitely a model bug, if you will.

Logan Kilpatrick

It’s not the intended behavior. The goal is to help the user with whatever the thing is that they’re trying to do. If you see those in whatever product you’re in, thumbs up, thumbs down, send us the feedback so that the model team can look and help try to chase those down.

Nathan Labenz

If we take it one step further, folks are doing more and more of these model welfare checks and interviews where they literally ask the model, in some cases, “How do you feel about the way that you are deployed?” Is anything like that happening within the model?

Tulsee Doshi

It’s a good question. I think the “How is it being deployed?” question—I feel like the model is just—this is my personal sense of a lot of these tests—completely out of the distribution. The model has no idea how it’s being deployed, so it’s just pontificating in a lot of these cases. It’s not like the context window of any large, major LLMs has, “Here are the details of how you’re being trained, here’s your serving setup, and here are the people who are working on it.”

Though maybe these are interesting things to experiment with in the future. I think a lot of it is just pulling out of a random distribution from the large-scale training that happens on the models. I feel like it’s actually less representative of how the model works. It just doesn’t have the context.

Nathan Labenz

Yeah, one reason that’s true, which I was just noticing in AI Studio, is I think all the models that are publicly launched, at least so far, still have a January 2025 knowledge cutoff. It is actually amazing that it does well in search and that they can have—I ran a Deep Research on, “Give me everything that Google has launched in the AI space, and what’s even the speculation about what they’re going to launch at I/O”—and it did a very impressive job.

Deep Research is great, especially considering it knows in its weights nothing about the last 18 months. So I guess the first question is just: Why are we still at a January 2025 knowledge cutoff? Can I categorize this as a bug?

Tulsee Doshi

Yes, sir.

Nathan Labenz

This is also one of Logan’s favorite topics to discuss.

Logan Kilpatrick

Yeah, no, I mean, I think updating the knowledge cutoff is definitely important and something that’s on our radar. I think the other part, though, is: How does Deep Research do so well, or how can we use the model in search? It’s because we also have model search, right?

For us, it’s really important that the model be able to know when to leverage its parametric knowledge versus when to actually go out and get the information from the web. Especially because there’s information that’s as fresh as an hour ago or a minute ago. We want the model to be as up-to-date as possible.

For us, we’ve been really leaning into: How do we help the model search effectively? That’s a big part of what makes it successful in the context of search, the app, or even Antigravity, actually, for that matter.

Nathan Labenz

That reminds me of one of the more surprising bits of news that I’ve seen from Google, maybe ever, which is the partnership with Exa, bringing Exa in as an alternative to Google for grounding. I never expected to see Google work with any other search provider. So what’s the story behind that?

Logan Kilpatrick

I think this is just generally—Google Cloud does tons of these types of ecosystem partnerships with folks across lots of things that are somewhat competing, quote-unquote, with what Google is doing. You can look at the cloud marketplace generally; it has lots of stuff. Google Cloud also hosts a Model Garden. There are the Anthropic models and other model providers there.

It’s very standard. At the end of the day, some enterprise customers want choice, so I think it’s trying to meet enterprise customers where they are. I don’t think it’s a—I think it’s a good sound bite that Google can’t do search and that’s why we have to partner with other companies. But at the end of the day, to Tulsee’s point, the model team and search have a super-deep collaboration. The models are built with that use case in mind.

For some portion of enterprise customers, they want flexibility in their external search tooling providers, and Google Cloud is doing its job as a great enterprise business by partnering and finding the right folks to work with.

Nathan Labenz

Last couple minutes, maybe just a little lightning round. Why hasn’t context grown more in the last year or two? We got 1 million, and that was up from 4,000 in just a couple years, right? But now we’ve kind of leveled off. Is that because people don’t want it?

I mean, we saw this subquadratic model that came out and made a bit of a splash with a 12 million-token context window and a new attention strategy to support that. Is it that people don’t want it? Is it too hard? Is there not the compute to handle it? What’s currently limiting context?

Tulsee Doshi

I think people definitely do want lots of context, but what we’ve also found, if you look at personalization, where you want to access all of your personal context, or coding, where you have extremely large codebases, is that a lot of the frontier here is going to be on how you smartly use context.

Thinking about compaction, what are the right ways to find the right elements of the context and bring them into the model? That’s a huge opportunity: How do you leverage all of this information that the model might have access to, when a lot of it is frankly distracting for the model to actually do the right thing? How do you give the model the right amount of context in the right way to be most effective?

I think that’s really the direction that we want to be pushing in. In actuality, the amount of context that the model is leveraging is much, much larger, but because we’re being smart about how that’s coming into the context window, you can fit it into smaller context windows.

I think this also goes back to my point about Flash-Lite and Flash. Larger and larger context windows also come with cost. What we also saw with customers, and still see with customers, is that a lot of customers want to use smaller context windows because of that. They want to be more intentional about what’s going into the model.

We’re trying to meet the moment with the right balance: How do you provide a lot of useful context while also meeting the right kind of latency, cost, and other trade-offs?

Logan Kilpatrick

Yeah, and I think one thing I’ll add is that, in today’s paradigm of how continuing to extend context works, it just becomes too cost-prohibitive for customers to actually use in practice. Even at the extreme of a 1 million-token context, in some cases it can be a few dollars for a request at that rate.

The demand for that is just so small, and there’s a huge amount of compute required in order to do that. There are a lot of trade-offs that you’re juggling. I’m hopeful that we’re a research breakthrough away from enabling that to continue to scale up without it being such a large investment, both from the user side and in terms of the serving compute required to make it possible.

Nathan Labenz

Speaking of possible research breakthroughs, what happened to that diffusion coding model? I was excited to see apps materialize in 3 seconds in front of my eyes, and it’s been quiet on that front.

Tulsee Doshi

Diffusion is awesome. It is superfast. I think we are still testing and experimenting with it in a number of different ways, trying to figure out what the best way is to put this out into the world and where it’s most useful.

Part of the reason why we’ve also been investing in Flash-Lite is that Flash-Lite is an incredibly fast model. If you look at the Gemini 3.5 Flash model we’re releasing right now, on Artificial Analysis it benchmarks at, I think, 280 tokens per second, which is crazy fast. In fact, it’s so fast that sometimes in Antigravity, by the time I want to cancel, it’s too late.

I think we’re already trying to figure out where you start getting to Logan’s point of diminishing returns and where you see that value proposition. I think that’s part of the question, too. But we are continuing to push on diffusion research.

Our researchers who are working on diffusion are doing some pretty awesome stuff. I was in a meeting with them the other day about some results that they have. I think they’re still pushing the frontier of quality and speed in ways that are really, really cool, so I think we’re going to see that play out really well.

Logan Kilpatrick

Yeah, I’m excited. I feel like it’s a research exploration. Obviously, there was the application where you could test it last year at I/O, but I think the framing was, “We’re doing interesting research. This is a look behind the curtain of the interesting research we’re doing,” and hopefully it manifests in models maybe one day, or just informs our perspective of what works and what doesn’t.

One thing, actually, as far as speed is concerned—just another plug—is that it's actually in Antigravity right now. There's a faster version of 3.5 Flash, so it is speedy, actually. I think we're excited to see how people will use that and what the reception and reaction will be to it, too. People want fast models.

Nathan Labenz

Yeah, no doubt. Well, time is the one resource we can't get any more of. I know you guys are super busy.

Tulsee Doshi

Yeah, well, we can build more compute.

Nathan Labenz

Yeah, it's hard to create time out of nothing. So maybe just one last question: what else is Logan asking that I haven't asked? Yeah, yeah, Logan. What are you asking?

Logan Kilpatrick

Let's see. I think we talked a little bit about this, but the one thing I will say is that I'm really excited about where audio is going, too. That's one that I think we tend to talk less about. But if you think about the Gemini mic example, or you think about the Gemini Live experience, I'm really excited about moving toward a paradigm where audio is just a bigger and bigger part of how we engage with these models and how they engage back. So definitely try out Gemini Live, the updated experience, but I think that's another area—that's a paradigm I'm excited for us to keep pushing, too.

Tulsee Doshi

Yeah, and I think the seed to plant is that, obviously, Google I/O is an incredible moment, with lots of stuff coming out the door, but this is just the start of the summer of amazing things and lots of other stuff. The engine keeps churning, and there's lots of stuff in the works, which I'm excited about, and many more stories, many more podcast episodes, so that we can get you all to see us.

Nathan Labenz

Yeah, if you get up to 7 podcasts.

Logan Kilpatrick

But actually, legitimately, I was in a room this morning where one of my team members was like, “I know we're going to be launching this a few weeks later, but I really need a vacation.” I was like, “Well, we're—”

Tulsee Doshi

We're just going. We're just moving.

Nathan Labenz

No rest for the weary in the AI era, that's for sure.