[BidClub_]
The Cognitive Revolution · · 83 分钟

2025年5月15日至22日的十年:Google 50倍 AI 增长与转型——Logan Kilpatrick

Logan KilpatrickNathan Labenz

YouTube
TL;DR
  • Google 的 AI 使用量在1年内增长50倍,从每月约10万亿 tokens 飙升至500万亿 tokens,背后是组织重建、基础设施扩容和产品需求激增。 这意味着在不计入其他供应商的情况下,地球上每个人每月对应超过5万 tokens。Kilpatrick 将其归因于2023年中 Google Brain 与 DeepMind 的合并、反复进行的训练与发布周期、TPU 扩容,以及 DeepMind 从基础研究转向产品化;关键在于这4个方面的“改进斜率”。

  • Kilpatrick 预计,随着容易获得的增益耗尽、结构性优势开始主导,头部模型将出现分化。 研究初期会趋同,因为“有人照亮了道路”,产品团队也在竞速,避免看起来落后;但从现在开始,“低垂的果实已经摘完”,改进难度上升,Google 的算力基础设施将变得更重要。Nathan 的反驳更接近投资者视角:如果 DeepMind 都认为下一前沿很难,那么训练第二梯队基础模型的前景会更加残酷。

  • 即便前沿模型的经济学趋于集中,应用层的经济性仍异常有利。 Kilpatrick 称这是“人类历史上创办创业公司的最佳时点”:拥有10亿用户的巨头解决通用问题,创业公司则可以凭借速度、聚焦和更便宜的工具,攻入狭窄细分市场。Nathan 的证据是切实的——每月约1000美元的 AI 订阅能带来远高于成本的价值;一个 AI 辅助的电台广告项目报价约比原有流程低75%,但按他个人投入的每小时计算,项目收入约为3000美元。

  • Windsurf 事件暴露了供应商风险,但 Kilpatrick 认为 Google 的激励机制会推动其广泛且尽早分发 API。 他认为 Anthropic 在 Windsurf 与 OpenAI 合作后切断其访问权限,可以解释为将算力分配给可能的长期合作伙伴。对 Google 而言,Cloud 的使命、外部评测信号、竞争势能,以及10亿用户产品内部缓慢的模型迁移,都构成了不扣留最佳模型的“多层动机与博弈论”——这是一种预期,并不保证 Gemini 3 一定如此。

  • Gemini 2.5 Pro 最清晰的差异化在于驾驭长上下文,而不只是上下文窗口更大。 Nathan 用40万至50万 tokens 的代码库测试时,感受到了跨越式提升;Kilpatrick 提到 OpenAI 的 MRCR 八针评测,称最新的2.5 Pro“提升大致在20%左右”。他的解释是,推理能力终于让模型能够真正使用完整窗口,推动架构加载更多上下文,尽管 RAG 仍不可或缺。

  • 推理模型正在把智能体架构变成移动靶:今天必需的脚手架,明天可能就会成为技术债。 模型正在“开箱即用地成为系统和智能体”,而 NotebookLM 的音频概览流程已经从14个以 Gemini 为主的阶段缩减至4个。实操建议是:为当前可靠性搭建结构化工作流,同时避免那些无法随着能力进入模型而简化的单向门设计。

  • 扩散语言模型仅凭速度,就可能带来另一轮软件范式断裂。 Kilpatrick 认为其由粗到细的生成过程,比严格自回归写作更符合直觉,并表示如果这一范式在2年内胜出,他不会感到意外。Nathan 怀疑扩散模型最终可能更优,并指出 transformer 并非历史终点。质量、成本和性能之间的权衡仍未确定,但 Kilpatrick 称其已展示出的速度“快得难以置信”,而编辑工作流尤其适合这一范式。

  • Kilpatrick 认为,AGI 会以产品组装的形式被感知,而其中真实的人类视角仍将稀缺。 一个推理能力提升50%、长上下文提升50%,并能在恰当时机调用记忆的模型,可能先带来 AGI 般的体验,而不是等到某次发布让所有人一致同意“这就是 AGI”。但他写邮件和帖子时约95%完全不用 AI,因为能动性和声音很重要:“人们想要的是 Nathan 体验”,即使 NotebookLM 已能按需生成主题明确、制作精良的替代内容。

摘要 · 为研究而整理的核心内容

1. 竞争先制造趋同,再制造差异化

  • Kilpatrick 将研究趋同与简单抄袭区分开来。推理能力就是一个例子:一旦某项创新及其所需的“一个数量级的投资”变得清晰,“有人照亮了道路”;竞争者可以吸收这一洞见,同时叠加各自已经在推进的独立押注。

  • 产品趋同背后是更残酷的激励机制。AI“是当今全世界竞争最激烈的生态”,资本、人才、知识资本和执行速度高度集中;团队必须在长期差异化押注与短期声誉成本之间权衡,因为一旦缺少竞争对手刚发布的功能,就会立刻显得落后。

  • 这种张力直接传导到 Google 的开发者平台:兼容性让其他供应商的客户更容易采用 Gemini,但每投入一单位精力追赶功能 parity,就意味着少一单位精力用于构建下一代 API 和模型能力。

  • Nathan 怀疑密集发布反映了协调后的准备就绪,但现实层面的反驳是:部分时间安排确实出于刻意协调,更多则是偶然。规模超过约20人的组织不可能在1小时通知后可靠地调整重大版本发布;“公司没有那么敏捷”,前沿实验室也一样。

2. Google 的反转来自组织重建,而非一次模型发布

  • Google 过去的分散化有其合理基础。Google Brain 进行广泛研究,产出了包括 transformer 在内的成果;Google Research 更偏应用,并将能力向上游输送到产品;DeepMind 则遵循 Demis Hassabis 对通往 AGI 路径的更具体判断。

  • 一旦某条近期路径明显奏效,分散押注就从资产变成了负担。2023年中,Brain、Research 的部分团队与 DeepMind 合并;按照 Kilpatrick 的说法,这是“让自己处于未来10年能够成功的位置”的起点。

  • 重组不会立即带来前沿执行力,因为“大型人类系统极其复杂”。管理层必须协调不同文化、搭建团队结构、建立可重复的训练与发布周期,并吸取运营层面的教训;在那之前,OpenAI 已更频繁地完成了这些实践。

  • 算力也必须与组织同步扩张:TPU 既用于研究,也用于推理,闲置容量并不是只等需求出现。DeepMind 还成为负责 Gemini app 和开发者业务的产品组织;Kilpatrick 预计,未来3至5年将展现这些“正确但艰难的决策”带来的回报。

3. 使用量增长50倍,使基础设施成为战略约束

  • Google 在约1年内将月度 tokens 使用量从约10万亿提升至500万亿,增长50倍,相当于在世每个人每月超过5万 tokens。由于使用量分布不均,加上其他供应商贡献的额外规模,活跃用户的实际使用强度还要更高。

  • Kilpatrick 认为,激励机制和企业 DNA 会进一步强化这条曲线。早在当前生成式 AI 形态出现之前,Google 就已将基于 transformer 的系统部署到数十亿用户规模的 Search 中;如今更好的模型正在改善 Docs、Sheets、YouTube、Waymo、Cloud 等产品,而不是作为可拆卸的外挂存在。

  • 他的预测是模型将进一步分化:“低垂的果实已经摘完”,每一次后续进步都会更难,Google 的基础设施优势应当变得更加明显。Nathan 进一步明确了含义:如果 DeepMind 都说下一个层级很难,那么成为第二梯队基础模型训练商就尤其缺乏吸引力。

  • 专业化仍是一条可能路径:一家实验室理论上可以选择打造全球最强的代码模型,而不是最强的通用模型。Kilpatrick 以 Anthropic 为假设案例,但随即补充,其广泛使命使得转向纯代码模型的可能性不高;规模较小的模型公司已经在探索能够建立差异化视角的领域。

4. 前沿训练趋于集中,创业公司仍保有速度与聚焦

  • Nathan 的看空逻辑是,大型科技公司可以赢下任何它们选择的品类,包括应用层:当时3家主要前沿开发商都刚推出代码智能体,直接与 Cursor、Windsurf 等同样构建在其模型之上的产品竞争。

  • Kilpatrick 承认,Google 的 Jules“还非常早期”,远远落后于头部代码产品的采用规模;Nathan 提到,其中一家刚宣布 ARR 约为5亿美元。这一差距本身说明,模型供应商拥有现成的模型,并不意味着自动拥有应用分发能力。

  • 训练前沿模型需要深厚资本,但“人类历史上没有比现在更适合创办应用层创业公司”。软件创建速度更快、试验成本更低、变现速度异常快,而面向10亿用户构建的通用解决方案之下,仍有100万个具体问题等待解决。

  • 创业公司的持久优势在于速度和聚焦。大型企业面临安全、隐私和评测约束,采用新工具的速度更慢;创业公司可以解决一个用户的问题,不必在“100万零1件创新事项”之间做协调。Kilpatrick 的原话是:“能够只聚焦一件事,本身就是一种祝福。”

5. Cloud 的经济性不支持扣留 Google 的最佳模型

  • Windsurf 清晰展示了依赖风险:在与 OpenAI 达成协议后,Anthropic 收回了 Claude 的访问权限,而 Claude 此前是 Windsurf 的主力模型。Kilpatrick 对 Anthropic 所称的、希望把稀缺算力分配给潜在长期合作伙伴的理由表示“理解”;Nathan 也承认,不支持竞争对手是常规商业策略。

  • Nathan 随后提出更黑暗的情景:Google 可以先在内部的代码智能体、Gmail 和 Docs 中部署 Gemini 3,然后将公共 API 延后数月。Kilpatrick 认为这很难想象,因为 Google Cloud 是他所称的全球第5大企业业务,其使命是输出 Google 级基础设施,让其他人能够构建产品,而无需重新创造这些基础设施。

  • 外部开发者目前可以比 Google 内部团队更快。拥有10亿甚至1.5亿用户的产品必须避免行为回归、进行广泛评测并维持用户预期;小型创业公司则几乎可以立即切换模型。要求内部先完成部署,反而会显著拉长发布时间。

  • 针对 AI 2027 中内部模型与公开模型差距扩大的情景,Kilpatrick 提到发布前评测信号较弱、展示发展势头的重要性,以及客户的切换成本。每百万 tokens 的定价未来或许会改变,但广泛分发仍能覆盖各种产品的使用场景:“把这个东西提供给其他人,会是一门很好的生意。”

6. Gemini 2.5 Pro 最清晰的差距在于可用的长上下文

  • 默认人格比看起来更难标准化。基础版 Gemini 服务于用户结构截然不同的产品,因此 Google 采取了相对中性的定位,同时让 Gemini app 等产品自行添加个性;否则模型切换可能让用户觉得“那个人……已经不见了”。

  • Nathan 测试的是一个40万至50万 tokens 的研究代码库。Gemini 2.5 Pro 能够掌握整个代码仓库,反复重写长文件并调试由此产生的错误;在他能够使用的产品中,没有其他供应商在这一上下文长度上提供过类似体验。

  • Kilpatrick 的量化证据来自 OpenAI 的 MRCR 基准。在最难的八针检索变体中,最新的2.5 Pro“提升大致在20%左右”;单针检索在 Gemini 1.5 Pro 时就已接近100%,但随着不同目标数量增加,性能过去会明显衰减。

  • 这一机制是在与推理负责人 Jack Ray 的对话中形成的:推理能力与上下文能力相融合,推理让模型真正利用可用窗口。Google 目前已经看到请求长度显著增加,Kilpatrick 预计会有更多工作负载直接将信息加载进上下文——“当然,在某些情况下仍然需要 RAG。”

7. 早期访问依赖关系,Veo 3 仍在等待规模化

  • Kilpatrick 将可信测试者的进入路径讲得非常明确:拥有有趣项目的构建者可以发邮件至 lkilpatrick@google.com。Google 有一套“非常成熟的早期访问计划”,他的筛选标准是开发者反馈是否有用,而不是是否属于某个神秘的核心圈层;保密程度则因项目而异。

  • Veo 3 API 当时尚未向外部用户开放接入。约束来自基础设施:API 需求的数量级与高价消费产品不同,Google 必须先确保容量能够支撑这一规模,再大范围开放。

  • 原生音频改变了 Kilpatrick 对生成视频的看法。他此前持怀疑态度,因为将无声片段变成实用媒体需要大量下游工作;同步的语音和声音让结果真正“活”起来。Nathan 提到 Waymark 的案例:本地商家广告可以用 Veo 2 制作动画素材,再用 Veo 3 添加不同声音的片段,突破传统配音形式。

8. AI 支出已经制造出失衡的经济性

  • Nathan 个人每月 AI 订阅支出已达到约1000美元,覆盖 OpenAI、Claude、Gemini 以及累计约20个工具。其中很大部分是重复测试,但他说,即便将试验成本算在内,生产力增益仍“远远高于”支出。

  • 交互模式现在更像是把工作委托给人类。旅途中,Nathan 在 Replit 上构建了2个应用,智能体几乎完成了全部工作;他不再逐 token 指挥工程实现,而是提供产品说明,只有当结果显示自己的指令可能误导了智能体时才介入。

  • Kilpatrick 更希望未来用每美元带来的经济生产力来评估系统,而不是继续依赖已经饱和的学术基准。真正有用的问题是:一项系统能从20美元订阅中创造多少正向价值;他预计5年后的答案会与今天明显不同。

  • Nathan 最有力的例子是一个本来价值数十万美元的本地电台制作项目,按地点和版本计费,每个只需几百美元。AI 让客户获得约75%的折扣,但按他个人投入的每小时计算,项目收入接近3000美元——并非全部归他所有。Kilpatrick 认为,这类优势还会持续出现,“处于前沿的人很可能获得不成比例的回报”。

9. 更强的推理能力将智能体脚手架折叠进模型

  • Nathan 的“智能体马蹄理论”认为,早期聊天机器人和先进代码智能体都属于回合制交互,只是新一代的每个回合规模更大。真正可靠的无人值守自动化仍通常位于中间位置:结构化、多提示词、分支有限的工作流,而不是给模型50个工具,再允许它自行选择路径。

  • Kilpatrick 预计,更多脚手架会迁移到推理和模型层,搜索、代码执行、沙箱、工具和函数调用将越来越多地开箱即用。脚手架仍有必要,但构建者应避免单向门设计,以免模型原生处理昨日的定制编排后,不得不进行根本性重写。

  • NotebookLM 以数字呈现了这一转变。其音频概览生成最初是一条14步、主要由 Gemini 驱动的流水线,现在已缩减为4步;消除顺序调用不仅简化了系统,也让产品变快,而不只是更易维护。

  • A2A 解决的是生产级智能体的要求,而 MCP 并不涵盖这些要求,包括身份验证以及大规模部署智能体的其他环节。Kilpatrick 对标准的最终走向保持开放:MCP 可能扩展到这些功能,也可能为 A2A 等互补框架保留空间。

10. 扩散语言模型可能让软件变得即时

  • Kilpatrick 认为扩散模型由粗到细的过程更接近他自己的认知方式:思考从模糊和结构化开始,随后拆解成各个部分,直到最后才进入句子层面的 token 排序。他表示,如果这一范式在2年内胜出,他不会感到意外,但没有作出绝对预测。Nathan 后来说,他怀疑扩散模型最终可能更优,并指出 transformer 并非“历史终点”。

  • Kilpatrick 的第一反应是速度——“快得难以置信”。如果质量和成本相当,生成速度快到可以在“人眼一眨之间”完成渲染,个性化生成式界面就能真正实用,而不必要求用户盯着屏幕看 token 流过。

  • 其中的权衡仍未明朗,自回归模型也可能与扩散模型并存。更大的启示是继续探索替代范式;而扩散模型能够修改输出的能力,似乎尤其适合日益重要的编辑工作流。

11. AGI 可能先作为产品组装完成,再作为模型被宣布

  • Kilpatrick 怀疑某一次单独发布会让所有人一致同意“我们显然已经构建出了” AGI,部分原因在于人们对 AGI 的定义本就不同。他预期的 AGI 时刻更偏体验层面:有人将非常强大的模型与包括记忆、上下文在内的正确产品组件结合起来,用户最终认为整个系统具有通用性。

  • 底层跃迁可能没有产品表现得那么戏剧化:也许长上下文提升50%,推理能力提升50%,某个团队终于在正确时刻调出了正确的记忆。最后一步可能同样涉及工程、神经科学和人类心理学,而不仅是模型规模扩张。

  • Nathan 将 Google 的 Titans 工作视为记忆能力独立进步的证据,提到其展示过最高1000万 tokens 且表现强劲,并存在进一步扩展的可能。与大脑类似,最终形成的智能可能仍是模块化的,而不是完全驻留在某个单一的庞大模型中。

12. 合成内容越丰富,人类视角越稀缺

  • Kilpatrick 的世界观仍然“根本上以人为中心”。他写邮件和帖子时约95%完全不使用 AI,因为希望掌控自己的语气和公共身份;即便数字分身替他表达的内容很可信、很像他,也仍会让他感到陌生。

  • 合成内容的泛滥可能提升作者视角的价值。如果有人发送完全由 AI 生成的内容,Kilpatrick 并不太在意,因为这意味着更少的创作投入;但他保留一个重要例外:如果软件运行得足够好,他通常并不关心是不是人写的。

  • Nathan 的反驳是 NotebookLM:它可以即时生成一档围绕某个主题的节目,而此前没有任何人类播客主持人覆盖过这个主题,并且已经占用了他一部分收听时间。Nathan 随后引用 Sundar 与 Search 的类比:AI 聊天用户增长到数亿的同时,Google 查询量仍在增长,说明相关产品可以解决不同意图,而不是一对一替代。

  • 注意力仍是硬约束——Nathan 已将默认收听速度从2倍提高到2.5倍,并怀疑随着模型速度超过人类阅读速度,神经接口最终是否会成为必要。Kilpatrick 的反向押注是关系稀缺:在数千个经过优化的替代品中,受众仍会想要“ Nathan 体验”,以及具有差异化的人类视角。

Hello and welcome back to the Cognitive Revolution. Today's guest, Logan Kilpatrick, needs no introduction. This is his fifth appearance on the show and his tireless work in support of AI application developers previously at OpenAI and now for the last year and change at Google is legendary. In this conversation, with the benefit of at least a little time to process, we're looking back and digesting the overwhelming volume of major new AI models and products that Google and others have recently released. Logan describes his personal experience at Google as their AI usage has grown some 50x from 10 trillion tokens per month just a year ago after he started to 500 trillion tokens per month today, which is notably more than 50,000 tokens per month for every person living on planet Earth. Logan also shares his perspective on Google's incredible organizational transformation from what once was described as a sleeping giant to now an indisputable top tier AI powerhouse with assets headlined by the strongest overall compute infrastructure of any company. top tier and paro frontier models including Gemini 2.5 Pro, highly original and viral products like Notebook LM, gamechanging applications in medicine and science that are starting to ship to trusted users and what I have always and still consider to be the deepest bench of AI research talent and the most diversified well-rounded research agenda to be found anywhere in the world. He also offers thoughtful analysis on whether we'll continue to see convergence among leading AI companies or more divergence as the lowhanging fruit gets picked. Why he believes that startups still have unprecedented opportunities despite big tech's advantages. The implications of anthropic cutting off windsurf after they partnered with OpenAI. How the blinding speed of Google's latest diffusion language model could bring about yet another revolution in software. and why despite all the AI capabilities advances he's seen and helped to popularize, he's still betting that humans will continue to matter and taking a relationshipcentric approach to his work. Speaking of relationships, perhaps the highest alpha part of this episode was when I asked Logan for advice for those who want to break into the early access programs and other support structures that he and people in similar positions can provide. I won't spoil his response here, but it did include his personal email and an invitation to reach out. As always, if you're finding value in the show, we'd appreciate it if you'd share it with friends or leave us a review. We always welcome your feedback, too, either via our website, cognitiverevolution.ai, or by DMing me on your favorite social network. With that, I hope you once again enjoy hearing from Logan Kilpatrick of Google Deep Mind.

Nathan Labenz

Logan Kilpatrick, everybody knows who you are. Welcome back to The Cognitive Revolution.

Logan Kilpatrick

Thank you, Nathan. This is the world record for the most number of times. I feel like you should just make me an independent recurring segment on some regular cadence, because we get to do this a lot, and it’s wonderful to be back.

Nathan Labenz

It’s only your calendar that would prevent that from happening, so be careful what you wish for.

We were joking about titling this podcast “The Decade of the Week of May 15 to May 22, 2025.” Holy moly, we’ve got just an absolute avalanche. I’ve been saying for a long time that my grip on all AI news is slipping, and with this moment, I think it’s officially slipped for everybody. We’ve come a real long way since GPT-4, and a ton of stuff is happening.

I want to run it down, but I also realize at this point we can’t even be comprehensive. I also want to take some strategic opportunities to zoom out a little bit and get your bigger-picture perspective on some things.

The first one in that vein is that over the last however many months, we’ve seen several waves of leading AI companies launching very similar things in pretty short periods of time. This has happened with reasoning models, and most recently it has happened again with coding agents. On a feature level, we’re also seeing quite a bit of connecting to your Gmail, connecting to your Google Docs, and different kinds of context retrieval.

Some of that’s obvious, but some of it is pretty core research-driven, right? Getting the models to reason. How do you understand why the different leading companies seem to have such a similar development trajectory and also launch timelines?

Logan Kilpatrick

Yeah, that’s a great question. I think there are a couple of dimensions to this. One, I think on the research side, there are true innovations that, when people go out and talk about them, become clear in hindsight why something should be done. I think reasoning is maybe that story. Obviously, DeepMind has been working on that reasoning stuff for a long time.

I think some of the particular techniques just became clear: let’s make this level of order-of-magnitude investment, and things ended up working out pretty well. Then you bake in all the other stuff that we had been trying that was independent and different, and you start to see some really interesting things.

I think some of this is people lighting the path, and this is what’s awesome about the ecosystem: other people light a path, you get to benefit from the path they’ve lit, and then you go and bake those innovations into what you’re doing, plus benefit from the bets you were making that were independent of that. I think that’s exciting on the research side.

I think on the product side, AI, both on the model side and the product side, is the most competitive ecosystem in the entire world right now. There is not a more competitive ecosystem, with the amount of money, talent, intellectual capital, speed of execution, and so on, than there is in this AI ecosystem right now.

I think there are a lot of competitive people who are really good at what they do, and they don’t want to be pushed behind by their competitors. There’s this feeling that you have to stay on par with what everyone else is doing.

That actually ends up being a tension I feel as somebody who builds products in the AI ecosystem: the tension between doing what you think the long-term future is versus not trying to look like you’re behind in the present moment. Finding that balance point between the long-term bets that are distinct and actually going to give you a differentiated perspective over time, and the short-term requirement that we just need to have parity, is a tough challenge.

I feel this on the developer side as well, because we obviously provide compatibility layers for people who use other model providers to come and use Gemini. There are always thoughts about how much we invest in that to get feature parity while also developing next-generation API capabilities and model capabilities. It’s a tough trade-off.

As for the dates and timing, I think some of it is intentional. I think a lot of it actually just ends up being relatively happenstance, that things launch around the same times.

It is always fun to see the conspiracies of X, Y, and Z companies just sitting on something and then, with 1 hour’s notice, all deciding to put something out. Anyone who’s worked in any company larger than 20 people knows that’s not actually possible. The amount of operational overhead and burden to do something like that is enormous. Companies are not that nimble. Even your favorite AI lab is not nimble enough to have that level of reaction.

Nathan Labenz

Well, I do have to say Google and DeepMind, and your team specifically, have been pretty nimble, right?

A year ago, and definitely 2 years ago, the outside view of Google was “sleeping giant” turned sclerotic—everybody managing their own little fiefdoms. I don’t know to what degree that was really true. Now the narrative is totally flipped: the giant is wide awake and really remarkably keeping pace with even much younger and smaller companies.

One of the most interesting data points shared at the Google I/O event was 500 trillion tokens per month now being processed across Google’s services. We’re now into the next month, so at the rate of that curve, it might be literally 2 times more already. I don’t know if you’re watching the dials that closely to know, but how has that happened at Google culturally?

What has shifted, or what has your experience been internally—not just going through that curve, but rallying everybody to actually support all the different work that has gone into supporting that curve?

Logan Kilpatrick

Yeah. One of the interesting parts about this story is that, at the core, it’s a people and organizational story. I think the challenge is that it’s just not—that’s not selling front-page New York Times articles, or whatever your favorite analogy is.

Historically, if you look at how Google was set up to do a bunch of this work, I think the reality was that it wasn’t set up for this moment. Google was structured with many different teams doing AI work, and for a lot of the right reasons, actually, because they were pursuing fundamentally different goals in some sense.

Google Brain, as an example, historically had a very wide breadth of truly different research. That’s where a lot of the Transformer and other things came out of—that very varied breadth of research.

At the same time, you had Google Research, which was doing more applied things in some cases and trying to upstream a lot of that into other parts of Google. Of course, Brain also did that sometimes.

Then you had DeepMind, where Demis and the team had a strong opinion about how they thought—and how they think—we’ll get to AGI.

So they were pursuing a very specific research direction, and you've seen a lot of the things that have come out of DeepMind over the last 6 or 7 years around that. I think in a world where there wasn't a clear winner as far as what the technology, at least in the short term, could be to help scale us closer to systems that get closer to AGI, it made sense for Google to have those bets across the board, with very different structural organizations and approaches to doing this. But I think it became clear at a certain point that one of those things was working in the short term, and we should rally resources and get everyone on the same page. I think that happened in the middle of 2023, when Google Brain and part of Google Research actually merged with DeepMind.

I think that was the start of the story of Google putting itself in a position to be successful for the next 10 years with this technology. The challenging part is that large human systems are extremely complex. It's not just that there are a lot of humans involved; I think it's very easy to forget the level of complexity and chaos in human systems, regardless of what you want the outcome to be. I think the DeepMind team, Demis, and the folks on the leadership team there have done a great job of reinventing the culture and all that stuff in order to make an organization out of two very different organizations, bring them together under a single roof, and actually set up the team structure so that Google could be successful building these models and going through the process.

The large-scale training runs don't take a minute, so your iteration cycle isn't super quick. Your iteration cycle for releasing models and doing all the end-to-end work isn't super quick. It took time to actually get the iteration cycle going, and obviously OpenAI and others—or OpenAI specifically—had been doing that iteration cycle a little bit more leading up to some of these moments than I think we had been doing at the time.

As you do that iteration cycle, and at the same time all of a sudden everyone wants AI, the 500 million tokens a month is a 50× increase over the previous year. At the same time that you're setting up the right organizational structure and doing the iteration loop to make sure that you're actually training the world's best models, you also need to scale up hardware. That doesn't happen instantaneously either.

We need TPUs to do research. We need TPUs in order to do inference. The amount of TPUs that you need isn't just sitting there idly waiting. There's work involved and timelines involved and all that.

All things considered, given the constraints, I think we're in an incredibly good position. The thing that gets me most excited is the slope of improvement across all those dimensions. How can we work better as a team and get everyone on the same page? How do we keep making sure that the breadth of research upstreams back into the main Gemini models? How do we make sure we have the world's best infrastructure? How do we make sure that we get that iteration cycle for releasing new models down, and sort of learn the hard lessons and develop rigor around that?

I think we're doing all those things, which has been incredibly exciting to see. I think the last comment I'll make is that DeepMind has also transitioned, and I think we talked about this before, from being an organization that did foundational research to now actually building products. I think that's the last step of this organizational journey: How do we build the Gemini app? How do we think about what we do for developers? How does that actually influence how we train models and what that iteration cycle looks like?

All of that stuff has now happened and the work has been done, and I think we're putting the pieces in the right places. I think now, for the next 3 to 5 years, we get to see the outcome of making good and hard decisions to put things in the right place.

Nathan Labenz

Yeah. If I had to summarize that—and maybe contrast, and I won't ask you to contrast—there's another big AI research organization out there that has a sort of fragmented structure, tons of compute, and researchers who have been pursuing lots of different directions for a number of years. If anybody hasn't already identified it, that is Meta, and the contrast has been pretty strong.

One possible explanation would just be that leadership at DeepMind saw that, yeah, we're getting close. This actually seems like it might be a thing now, and so it's worth going through all that trouble of reorganizing. I'm sure there are many things Demis would rather do than reorganize an organization and redraw lines of reporting—who reports to whom and whatever. But if you're close and you're getting on that wartime footing, so to speak—not that I ever wanted to see an AI war—but it's worth it.

You haven't seen that same kind of thing at Meta and maybe a few other companies, and you also haven't seen the integration of the work. I mean, they do have Meta AI in their apps, but clearly it's not on the same level. They also haven't done the reasoning thing. I'm sure they had Meta researchers at the same San Francisco parties getting those same very few big hints that, hey, this seems to be sort of working, that everybody else seemed to say, “Okay, we better make sure we're on that train,” and thus far they kind of haven't. So I guess for me, the takeaway there was maybe just the importance of conviction in leadership to do whatever it takes and push hard on small hints that seem credible. That seems to maybe matter a lot right now.

Logan Kilpatrick

Yeah. The other piece that I'll add is that I think incentives matter a lot, and for Google, the incentives in the DNA matter a lot. Google has been—and Sundar has said this many times, and I think he's spot-on—an AI company since Sundar took over in 2016, or whatever it was. People joke that Google was sitting on the Transformer and didn't use it. The Transformer was powering Google Search at multibillion-user scale in a bunch of different ways. So the technology was being used. It wasn't in the same incarnation as the current generative AI stuff, but it was being used at that level of scale.

Building models, deploying them, building that infrastructure, and that iteration process had been in Google's DNA. It obviously needed to be reformulated a little bit for the current team that's doing that across Google.

The other piece of this is just the incentives as well. If you look across Google's products, Google, organizationally and in terms of what the future of our products looks like, is so incentivized to make great models because the great models that we make—and this is an interesting thread that we should talk about—are present across all of our products. You're writing in Docs, you're doing things in Sheets, you're in Waymo, you're doing stuff on YouTube with video, you're a Cloud enterprise customer and you're doing something—all of those use cases end up benefiting from this.

It's not just an add-on thing. It is fundamental to the success of those products. So I think there's an interesting angle to this around how incentivized Google is to be successful. I think we're highly incentivized, and it's in the DNA of what the company's been doing for the last 10 years. I think those two things as well—if you don't have them, it makes this moment probably a lot more painful than it would have to be otherwise.

Hey, we'll continue our interview in a moment after a word from our sponsors. In business, they say you can have better, cheaper, or faster, but you only get to pick two. But what if you could have all three at the same time? That's exactly what cohhere Thompson Reuters and Specialized bikes have since they upgraded to the next generation of the cloud. Oracle cloud infrastructure. OCI is the blazing fast platform for your infrastructure, database, application development, and AI needs where you can run any workload in a high availability, consistently high performance environment and spend less than you would with other clouds. How is it faster? OCI's block storage gives you more operations per second. Cheaper, OCI costs up to 50% less for compute, 70% less for storage, and 80% less for networking. And better in test after test, OCI customers report lower latency and higher bandwidth versus other clouds. This is the cloud built for AI and all of your biggest workloads. Right now, with zero commitment, try OCI for free. Head to oracle.com/cognitive. That's oracle.com/cognitive. Build the future of multi-agent software with agency. Agncy. The agency is an open-source collective building the internet of agents. It's a collaboration layer where AI agents can discover, connect, and work across frameworks. For developers, this means standardized agent discovery tools, seamless protocols for inter agent communication and modular components to compose and scale multi- aent workflows. Join Crew AI, Langchain, Llama Index, Browserbase, Cisco, and dozens more. The agency is dropping code, specs, and services, all with no strings attached. Build with other engineers who care about highquality multi-agent software. Visit agency.org and add your support. That's agncy.org.

Nathan Labenz

That 500 trillion tokens per month is 50,000 tokens per month for every human being on the face of the earth, which is a pretty crazy number.

Logan Kilpatrick

That is crazy. It’s grown a lot faster than I expected it might. Obviously, there are other providers out there processing a lot of tokens, too. So we’re now getting into the regime where—we should acknowledge that not everybody’s using it—we’re starting to get some significant inference numbers on a per-capita basis.

Nathan Labenz

As you look ahead, do you think we’re going to continue to see this sort of convergence, where the leaders will mostly be doing stuff that’s measurable on the same bar charts? Or do you think we’ll start to see more divergence, which could mean different form factors and significantly different strengths and weaknesses? Who knows what divergence might look like, but what’s your expectation there?

Logan Kilpatrick

Yeah, I would guess we see more divergence, to be honest. I was actually just at a dinner a couple of nights ago, talking to some founders, and a bunch of them were clearly betting on the fact that there’ll be model convergence. I think it depends on what sort of—what level of abstraction you want, as far as what things are going to converge.

My general sense is that the low-hanging fruit has been captured. So now it’s about what structural advantages you have as a business to train LLMs. I think Google has a really important infrastructure advantage in the ecosystem, along with a bunch of other things like that, where I think you’ll actually see those things shine through. I think getting to this point was not unexpected. Getting to the next level is not going to be easy for a lot of people to do, and that’s where the world-class teams and folks who are really making an order-of-magnitude bet on this are going to see the advantages.

Intuitively, through the lens that it only gets harder from this point, AI innovation does not become easier after this point. Making these models better is going to be more difficult. I would guess a bunch of teams start to—and it’ll be interesting to see what the size of the labs is. Maybe all the labs will keep doing everything, but I think there really will be opportunities to focus on a specific area.

I’m just conjecturing here, but you could imagine Anthropic deciding, “Hey, we actually just want to be the world’s best coding-model company. That’s the only thing we care about, and that’s what success looks like.” I don’t think this is going to be true because they have a very broad mission of what they want to do, and it doesn’t seem like it’s specifically code. But you could imagine that there are companies doing some angle of this where they end up deciding, “Hey, there’s actually value in diverging from this path of something super general, because we could build a really great business by doing that.”

I think it’s, again, at odds with some of the broad missions that these companies have, but I do think there’s real value in that. You can really go deep and start to understand how to build a long-term business and company around some of those things. Maybe the big labs won’t do that, but there are obviously more people training foundation models than just the big labs. A lot of those companies are going down the path of, “Let’s find something that we’re really good at. Let’s build a differentiated perspective on how to solve this problem, whether from a model perspective or an infrastructure perspective.” And I think that makes a lot of sense, honestly.

Nathan Labenz

Yeah, that’s interesting. I don’t know. I should be at least somewhat deferential—you probably know better than I do—but I just look at how fast these core foundation models from the leaders are getting better, and I’m like, I would not want to be a Tier 2 foundation-model trainer in today’s world. Especially if it’s only getting harder from here, and you’re saying that from the DeepMind position, I’m like, boy, that sounds really hard from any other position.

I guess it won’t be for you to sign on to this in this moment, I don’t think. But, transparently, my position is that I just think the big tech companies are going to win anything and everything that they want to win. You’re kind of seeing this bleeding into the application layer as well, right?

I’m interested to hear—I know you’ve been very focused on supporting developers directly via AI Studio and the APIs. You’re also, I’m sure, in regular dialogue with folks like Cursor and Windsurf and anybody who might use Gemini 2.5 Pro as a coding model. But now we’ve also seen all 3 of the big frontier developers in this last wave put out a coding agent, too, right?

How are they feeling, and how are you talking to them about the fact that they’re using it? You want them to use the model, but you also now have a competing product on the market against them, right?

Logan Kilpatrick

Yeah. Well, a couple of things. One, I think the product that we do have—you’re referencing Jules, right?

Nathan Labenz

Yeah.

Logan Kilpatrick

At least for us, Jules is definitely super early. I’m super excited. It’s a great team inside Google that’s working on it, but obviously the level of adoption that some of these other AI coding products have is very much on a different level.

Nathan Labenz

500 million ARR—they just said today.

Logan Kilpatrick

Yeah, I saw that tweet, which is exciting for them. I was talking to someone last night, and the comment that I made—which continues to be true, and I don’t know if I’ve said it on another episode that you and I have talked about—is that there’s no better time in human history than right now to be building a startup. Truly, if you’re building a startup to build language models and compete against all the big labs, you better be very well-capitalized to do that, because that’s a very difficult problem.

If you’re building out the application layer, it’s never been easier. The time to build software, the opportunity to explore new ideas, and the pace at which this current AI moment enables you to potentially scale monetization, build really retentive user products, and build a profitable business—all of these things have never been easier. The barrier to doing all those things has never been lower than it is right now.

As somebody who fundamentally believes in developers changing the world, I think that’s the coolest opportunity ever. Sure, the big tech companies will hopefully be successful as well, and will sell infrastructure and do some things at the application layer. But the real opportunity is that there are a million and 1 different problems to be solved, and some of these large, billion-user products solve things in a really general way.

The cool thing is that you can really go deep for some specific user segment and solve their problem in a unique way. The cost to do that—from building a startup and writing the software to do it—has never been lower. In the startup world, there are a thousand and 1 different AI tools that you get to leverage in order to get to that place.

Larger companies, just because of the level of security, privacy, and enterprise requirements, often don’t use a lot of those tools. It’s a different, bespoke set of tools, and the pace of tooling innovation for large companies often happens a little bit slower than it does for the startup ecosystem. You have all of these entrenched speed advantages, and then you couple in the idea that everyone’s going to have a bunch of agents building stuff for them in the future.

I continue to be super excited for people, even in the coding space at the application layer, who are building stuff. There are so many cool things to be built.

Nathan Labenz

The importance of speed—or the criticality of the advantage of speed for startups—I think is definitely extreme now. It’s always been true, I suppose, but it seems like it’s taken on extreme importance now.

I recently talked to Andrew Lee from Shortwave, who’s building a Gmail on top of Gmail, but also a Gmail competitor. He said, “After really soul-searching deeply, we came to the conclusion that our only advantage is speed.” Focus is the other piece.

Logan Kilpatrick

I think this goes to big companies, and I feel this as well. There’s lots of tension for me because the cool thing about Google is that there are a million and 1 innovative things happening. The challenge is, how do you actually balance that and take action based on the million and 1 innovative things that are happening? It’s a real burden.

The nice thing for startups is that you don’t have a million and 1 innovative things happening. You can just go and do 1 thing. I have a lot of envy for folks who have that because you just don’t need to make a lot of decisions. You can really focus on solving the problem at hand.

I think the speed of execution and the ability to focus on just a single thing is a blessing, so take advantage of that as much as possible.

Nathan Labenz

What do you make of this Windsurf news lately? The brief story is that they agreed to a deal with OpenAI. They had been using Claude as their primary model, and then Anthropic cut them off by virtue of having agreed to this deal with OpenAI. On the face of it, that's all pretty reasonable, but if I am a coding-agent company or whatever that's thinking, “What's my long-term prospect here?” right now, the speed advantage goes to the startups because the big tech companies have been so friendly, I guess, to the rest of the ecosystem as to put the models out before they've implemented them in their own products, in many cases.

But it's not too hard to imagine that flipping, right? If Google wanted to say, “Okay, Gemini 3, we're going to deploy it in our own coding agent, Gmail, and Docs, and then a few months later we'll put it in the API,” that would definitely flip the speed advantage on its head. Do you think startup founders should be worried about that?

Logan Kilpatrick

Yeah, that's an interesting question. I think, on the Anthropic piece, I do think that the byline of Anthropic wanting to sort of invest in who they think will be long-term partners and getting compute to those customers—I think, actually, as somebody who's spent a bunch of time thinking about how we get compute to the right teams that are building products, I have empathy for that argument. I think that could make sense. It's totally defensible from a simple business-strategy perspective, right? Don't support your competitors. I think that's totally defensible in many, certainly in normal business contexts.

As far as our strategy, the great thing for builders is that Google Cloud is the 5th-largest enterprise business in the entire world, and the mandate of Cloud is to bring this infrastructure to the rest of the world—to bring Google's infrastructure to the rest of the world—so that people can build world-class startups and not need to rebuild the level of infrastructure that Google built in order to scale the internet to where it is today. It's such a core and foundational part of the business that I would find it hard to believe the strategy shifting from shipping across our own services, but also shipping across Cloud services.

Interestingly, oftentimes today it's actually even more extreme than the picture that you painted. The external developers often have an even larger speed advantage from a model perspective because, if you think about who the customer of models inside Google is, it's teams that are building billion-user products and teams that are building 150-million-user products. I was just talking to someone about some of the features and products that exist inside Google Workspace, and some of the ones that I've never even thought about have 150 million monthly active users, which is crazy.

That user persona—if you go and talk to enterprise users of LLMs—they don't move the LLMs as quickly. They don't switch models as quickly because, even though they're inside of Google, there's still all of the normal constraints of building a large user product. You don't want the behavior to shift, and you have to do a ton of evals. You have to do all these things, and all of that requires time.

When you have a small product, it's very easy to quickly switch models, and you don't really have to think about it that much. But I think for teams inside of Google, they do have to think about that, and the responsibility to the users that we have is to be really thoughtful about that. So, again, I think the time horizon would shift so dramatically as far as getting these models out the door if the strategy became, “We have to sort of force internal teams to use these and then deploy them before we give them to external developers,” that they'd be so far off from where we are today that I have a hard time imagining that would make sense.

Also, from the business perspective, it's important for us to give LLMs to developers because it's a core part of the Google Cloud business, which is, again, a huge business for Google.

Nathan Labenz

I know Daniel Cocatello. I know you guys—I don't know how well, necessarily—but you overlapped at OpenAI. I'm sure you're familiar with his AI 2027 scenario. One of the interesting things in that is that he projects that, basically, over these next 2 years, model developers are going to start widening the gap.

I saw Roon not too long ago on Twitter. Somebody asked, “How much ahead is what you have internally versus what we see externally?” And he said, “You guys have no idea how good you have it. It's 2 months. You're on the bleeding edge, just behind where we are internally.” But Daniel's projection is that this will change, and that for multiple reasons—including wanting to use the models intensively for their own ML research automation, dreams of recursive self-improvement and takeoff, and who knows what—the gap is going to widen.

He has pretty aggressive scenarios in mind there. He thinks that, basically, the developers are going to start to hold the best models back for themselves. The public will kind of satisfice more often, and really insane stuff will be held very closely and known to few people. It sounds like you don't buy that scenario, basically, or at least don't see any signs of that happening at Google.

Logan Kilpatrick

I think there are 2 dimensions. First, it's hard to get signal on how good models are. Evals are just such a difficult problem, so you often don't really have an intuition as to whether this could be the right model long-term if you don't release it to the world. I think that's just one of many pressures on the idea that you should get the model out the door.

I would underscore the momentum—sort of a quasi-momentum war—that happens. It's really important to project what the external momentum looks like from an AI perspective because, ultimately, I think if you go and talk to developers and people who are building companies, that's actually a really large influence on who they end up building on.

I also think there's a bunch of other threads to this, including that switching stuff is hard. There are so many layers that I have a hard time buying that we're not going to deliver models to the world in the same way that we're doing it now. I think there are just many levels of motivation and game theory that tell me that won't be true.

It's also interesting to think about how you actually capture the most value from this technology from an economic perspective. Maybe it's not us. The cool thing about developers is that, obviously, Google has large distribution, but you get this really wide aperture of distribution across so many different things.

Maybe the economic model looks slightly different, where the unit of intelligence is a token, assuming the models are way better and can do all this economically productive stuff, and you charge people on a per-million-token basis. I could buy that—that changes in the future, and the economic model looks different from how developers do it today.

But I still think fundamentally you would want to build a great business giving that to other people, because how you're going to use this model looks very different from how other people are potentially going to use it. You could build a great business doing that by releasing it to the world.

Nathan Labenz

Yeah, the importance of feedback definitely is not to be missed. That also connects back to the “what's going on at Meta” line of conversation—which, for the record, you're not commenting on, but I'm just tangentially mentioning. They're not getting nearly as much of that, right, as the companies that currently have the best models on the market are getting.

So, yeah, that's interesting. How much do you use other companies' models? Do you go and do your own different vibe checks across different providers? What's your model diet?

Logan Kilpatrick

Yeah, I play around with a bunch of stuff. I think it's interesting. It's fun to see how—I'm also, independent of my job, somebody who loves technology and loves seeing cool AI products—so I spend a lot of time playing around with all the coding models. I think that's probably the thing that I experiment with the most.

But there's tons of cool stuff happening in the audio space right now. We launched our native audio model at I/O, which was one of the threads, and it's available in NotebookLM and a bunch of other products, as well as for developers. It's been really interesting to see that as an emergent space that people are building products and services in.

Nathan Labenz

I've been spending a bunch of time playing around with the products and models. ElevenLabs just landed a new model to do something similar—native, really robust, natural-sounding audio. It's been super cool. Across whatever dimension you want, there's tons of fun stuff to play with.

I still try ChatGPT occasionally to play around with it and see what that experience has evolved into. It's fun to be somebody who likes using this technology. It is interesting, though. I'm not one of those people who sends the same query to 3 different models, examines the differences between them, and does it in 3 different tabs.

These are people I have a lot of respect for who are running companies and all this stuff, and I'm like, "That seems like a pretty cryptic way to be doing that type of experimentation." That has led me to think there's probably an interesting product to build there, where people are really trying to understand the nuances and differences between models and engage with multiple answers. I think the multiple-pieces-of-content thing is a really interesting thread to pull on.

You can imagine that in the future and in product experiences. But I'm not at the level where I have Claude, Grok, ChatGPT, and Gemini open at all times, asking my question in 3 or 4 places. I don't do that all the time by any means, but I do it on occasion.

My general philosophy is always to try to be doing 2 things at once: 1 being whatever the object-level task is, and 2 being learning about AI's ability to help me with that object-level task. If I have any sort of contract review, that would be a great example. If I'm going to take on some advisory agreement or whatever, they'll send me the contract.

I'll send it to at least 3 AIs. If none of them have an issue with it, I'll just sign it without even reading it myself. Usually, they are pretty consistent, but that's also an interesting opportunity to see just how they're presenting things a little differently. Claude is typically the shortest and the least formatted.

I don't know. It is hard to characterize. How would you characterize it? Gemini 2.5, for me—which you just launched a new version of onstage at the AI Engineer World's Fair—I can't claim any deep familiarity with the new one because it just came out yesterday. But the Gemini 2.5 Pro class of models, I think we're on the third date-stamped version, right?

It was one of those hair-raising moments for me because the command of the context window that it has was just so incredible. I dumped a research codebase into it that had 400,000 to 500,000 tokens. No other provider, at least with the level of access that I have, even supports that length of context.

To see the command that it had of it was incredible. I was literally going back and forth debugging problems in AI Studio. Don't tell anybody, because they might cut me off for this behavior, but it's rewriting whole files for me. Then I'm saying, "I got this bug. Please fix it," and it's rewriting another version of these long scripts for me with 500,000 tokens of context.

That was like, "Wow, this feels like that step change." I wonder what other step changes you would highlight that people might not be fully aware of, or what more subtle vibes and tone differences you think distinguish Gemini from other options.

Logan Kilpatrick

Yeah, this whole model-behavior piece has been really interesting to see. Folks have a strong reaction to default personalities, and I think we're very early in coming up with a rigorous point of view as far as how to make a default personality that works well through certain lenses.

If you look at the requirements for building the Gemini model, the baseline Gemini model is used across so many different products, even inside Google. Those products have such varied points of view and products and users that they're building for, so it becomes really difficult to come up with a default personality.

I know Claude and Anthropic have done a ton of stuff with trying to make the model personality feel distinct, and they have a point of view about what that should look like. I think this goes back to the advantage for startups. Anthropic gets to do that because the consumer product is relatively small compared with 1-billion-user products.

It has been interesting to see us take a more middle-of-the-road approach—not try to have too much of a personality, but also make sure that the model can have that if that's what the product people want to build. The best example of this is the Gemini app. The Gemini app probably wants to actually have a personality and do some of that stuff.

I think the challenge becomes how you maintain that. I've seen that time and time again: as you change the models, the personality changes dramatically. If you're intimately conversing with these models, it feels like the person or model you were talking to before is now gone and has been replaced by something else.

I think that's a pretty jarring experience for today's model-iteration process. There is some interesting stuff to happen to make that not be the case. I actually have a tweet queued up that I need to put out about long-context capabilities, because far and away, 2.5 Pro in its current iteration has a really large gap.

I'm in my era of tweeting things live right now, just because it's top of mind. I'll put this out, and then I'll send you in the chat the tweet that I just put out. It's far and away one of the gaps in model performance right now. Long context is clearly one of those gaps. I'll put it in the chat for us.

Nathan Labenz

I just opened Twitter, and there it was—8 seconds ago. This is showing OpenAI's MRCR, which is a long-context eval that OpenAI built. You can see the delta between the models. Far and away, the latest version of 2.5 Pro is on the order of 20% better.

Logan Kilpatrick

This is 8 needles, which is the hardest version. It's not just single-needle context, which is retrieving 1 thing from the context window. Even the Gemini 1.5 Pro model from over a year ago was close to 100% accurate. That was basically a solved problem.

The problem is exponential decay as soon as you start adding more needles. To see this level of progress from a model perspective, with the ability to process and find 8 distinct items, is pretty remarkable—especially given that there wasn't a whole lot of long-context innovation in the last year.

I think it's this combination. I've had this conversation—I was just with Jack Ray yesterday, who leads our reasoning team and our reasoning efforts and originally worked on long context. I've talked to him a lot about this fusion of long context and reasoning.

Really, it's reasoning that enables you to use the full context window that's available. It's cool to see that happen. It's cool to see that happen finally.

Nathan Labenz

Yeah, you can feel it. Of course, benchmarks and practical use are not the same thing. I don't know what I would have said about Gemini 1.5 in terms of whether it felt like it had that depth of command.

I did a few things where we put whole books through it and asked it to find relevant quotes, and it could do that pretty well. But this new thing, if people haven't tried it recently, is worth exploring. Video is one of the best use cases, I think.

I've been doing this, and it's a fun experiment. Take a long video you've watched, then go and ask a bunch of questions. The challenge is that you perhaps already have to watch the video, but that use case tends to shine, which is really interesting.

Logan Kilpatrick

We've also, interestingly, seen a shift in what the distribution of request sizes looks like because of how good the context window has gotten. Historically, we'd been like, "Why does no one really use long context that much?" We were definitely ahead of the technology and the curve from that perspective.

Now, with 2.5 Pro, long-context usage is dramatically higher than it has historically been. It's been awesome to see people coming around and building on it. I think it gets closer to this future of the whole long-context-versus-RAG discussion.

Historically, you could have brushed it off a little bit because the model wasn't really that good and no one was really doing it. But I think the future is going to look more like people putting more and more stuff into the context window. Of course, they'll still need RAG in some cases, but it's awesome to see that corner turning from a long-context perspective.

Nathan Labenz

Yeah, "turn your hyperparameters up" is one of my current mantras. Another good one, for people who want to get a qualitative sense of the command that the new models have of long context: I have a simple Colab notebook for extracting emails from the Gmail API, where I just use “‘from:me’”—basically, just email sent by me—as my simple search.

Nathan Labenz

So filter out all the crap I’m not engaging with, but just threads that I’ve sent a message to, and pull all of those. You can go back, depending on your volume of email, pretty far and get a pretty robust picture of who you are that still fits into 1 million tokens. Then you can start to get a sense for what the model understands of you from those 1 million tokens, and it’s pretty impressive. I can share that Colab notebook if anybody wants to mess around with it.

I put it in that format so that you can do it without your data ever having to leave Google. It’s just going from your Gmail to your own Google Drive via the Colab notebook, so it was the most secure way I could think to make it that I could share with you. I don’t want to have your email. That’s the last thing I need.

Nathan Labenz

Okay, you’ve mentioned a few times being with people—dinners, talking to founders. I guess there are 2 angles on that. One is, what are you looking for in those groups?

Everybody who’s building wants to be in the inner circle of early access programs, the trusted tester rosters, and all that kind of stuff. How do people get into those programs? How do they get into the trusted tester sets? Also, can you enable the Veo 3 API for me, please? How much is networking—how you’re keeping up—versus other ways of keeping up?

Logan Kilpatrick

Yeah, that’s a good question. I think this looks different depending on what you’re doing. For me, maybe—I’m not sure how well this will track across people—but for folks who are listening this far into the conversation, I assume you’re an AI enthusiast, really, after the first 5 minutes. I love that.

So send me an email. Honestly, we have a super-robust early access program. We’d love feedback from people building interesting things. If you’re building something interesting, email me at lkilpatrick@google.com. Send me an email. I’d love to hear about what you’re building, and I’d love to get you into the early access program.

It’s not some big thing. Some stuff is more secret, and some stuff is less secret. We really just love to work closely with developers, get feedback, and be as open and collaborative with people as possible. So email us.

Hopefully, the Veo 3 API is a work in progress. We don’t have one available at the moment that we can onboard people to externally. We’re setting up a bunch of things and also working on how we can make it so that the model’s order of magnitude of scale for the API product, relative to putting it into a consumer product with a high price point, is just different. It’s a lot of different dimensions.

We’re working on ways to make sure the model can actually work at the scale of demand we’re going to see from an API perspective. It’s incredible to see the audio really bringing video to life. Historically, I’d been pretty skeptical of a lot of the video models. It was cool to see the video generated, but the practical use cases were hard for me because the amount of work it would take to do something meaningful with that video was pretty substantial.

I’m curious, for Waymark—I’ve used the product, but I haven’t used it recently—how important audio has been as part of that story. I think it really brings the video to life for me now that it has audio, and the audio feels like it’s actually native to what the video was meant to be saying.

Nathan Labenz

Yeah, it’s incredible. First of all, I’ve been thinking recently that we’re quite lucky. I don’t think anybody planned this, but there was a lot of hand-wringing about deepfakes and fake voices, cloned voices making calls, including a little bit from yours truly around the election that didn’t really come to pass. I think maybe that was mostly because the models weren’t quite there yet.

We’re fortunate that it’s landing early in a cycle where, hopefully, by the time the next election comes around, we’ll have enough reps, people will have built up cultural immunity to it, and there’ll be more guardrails and whatever, such that hopefully we’ll be able to deal with it. But it is getting to the point now where I genuinely don’t always know if something is AI video or real video.

For Waymark in particular, audio has been really important. We have traditionally, and still do, take the approach of just having a voice-over track. We mostly make TV commercials, and we mostly partner with big media companies. YouTube ads are a natural part of that. These are all sound-on environments.

Anytime you see a TV commercial, there’s usually a person talking to you, and then there are visuals, and there might be on-screen text and images and what have you. That’s our usual approach. We use a mix of providers, but ElevenLabs has certainly been a very important provider for us, and their voice quality just continues to climb.

With Veo 3, it kind of opens up a new dimension. In the past, we mostly used images that the businesses have. The next step is, well, if we can bring those images to life by doing image-to-video with even a Veo 2, then that just makes the whole thing more dynamic.

Quality is really important there. I would say Veo 2 mostly hits the mark, but sometimes has a little bit of weird stuff. Honestly, it is pretty damn good. But with Veo 3 now, it’s like, oh, you could even rethink the form factor a little bit. You could imagine having the voice-over talk a bit, but then also flipping over to a clip and having that thing present in a different voice.

It definitely opens up the space of possibilities for us in terms of the sorts of stories we can try to tell. We’re mostly telling small local business stories, but there are a lot of different ways to tell them. We’ve been relatively narrow in that space over time just because the technology could only do so much.

This is the kind of thing that our creative team sees, and it’s up to them to figure out exactly what they would want to make out of this new thing now that you can have all kinds of different voices showing up in a real context like that. Honestly, I think we’re still wrapping our heads around it and are also somewhat limited by the fact that we’re still just testing it in the actual top-tier Gemini app.

My personal AI spend is up to about $1,000 a month, which is also an interesting thing. I’ve been saying that for a while, but I hadn’t actually gotten there. Now I’m pretty much there between OpenAI, Claude, and Gemini, all at the top level, plus 20 other things that I’ve accumulated. It’s amazing to be spending $1,000 a month on AI subscriptions, but some of these things you’ve got to have.

Logan Kilpatrick

Yeah. Do you think you’re getting that level of value out of them? I assume, given the position that you’re in, that some of them are duplicative because you want to test all the different stuff. But if you were to remove the duplicative ones and just had whatever the best was across a bunch of different categories, do you feel like you’re getting that level of productivity boost relative to what you’re spending today?

Nathan Labenz

No question. If I weren’t committed to testing everything and having the earliest point of view on things that I can get, I think I could get a very similar productivity boost for much less. But the productivity boost is still dramatically higher than what I’m paying. No doubt, the acceleration of all sorts of different work is tremendous.

Last week, I was traveling a bit and ended up coding 2 different apps on Replit, just with the agent doing almost everything for me. It’s starting to feel like delegating work to other humans, much more so than a few years ago when we talked about prompt engineering. The original prompt engineering, right, is setting things up so that a natural completion of what you provided would be what you wanted.

Now I literally don’t think that much about the fact that this is even AI. It’s more just, “Here are some product notes,” and if it messes up, then I’m like, “Why did you mess up? Did I mislead you or something?” I have to think a little harder. But I’m really struck by how the communication to the AIs now feels much more natural and much more high-level. No doubt, the boost is tremendous.

Logan Kilpatrick

This is the eval that I think is one of the most exciting ones to me: relative to the amount of money that you’re spending, how much value is that creating? It’s hard to measure, I know, so it is somewhat theoretical. But that is what I think, long-term, as we move away from all the regular academic benchmarks being saturated, et cetera, is the economic productivity that’s created by some of these systems.

You need a broad sort of mandate in order to do something like that, just because there are so many possibilities—an infinite number of possibilities. But it is really interesting to think about, and I do think it’s a cool north star to drive up the amount of value you can create in the world in a very positive way through a $20 subscription.

The value you get today from a $20 subscription relative to what it’s going to be in 5 years, I think, is actually materially different. It’ll be cool to see that play out.

Nathan Labenz

Yeah. I think another dynamic that’s going to be interesting to watch is what, if any, stable equilibrium we ever arrive at. I think right now one of the reasons that there’s so much surplus for me is that we’re not yet in equilibrium.

Nathan Labenz

And so, to a certain degree, I have superpowers that other people don't have, which they could have, but they don't because they're not aware of them or they just haven't developed the habits. A lot of it is honestly just thinking to do it in the moment: go use the AI instead of doing it manually, for whatever version of it you might be considering.

Back in the holidays late last year, there was a project—and I wasn't involved in the business side of this at all—but somebody basically came to me and said, “Hey, I've got an audio production project. This company typically has a big network, and they want to do a ton of local radio ads, and I thought of you. Maybe you could do it.” They usually would pay a couple hundred bucks per location, per version of the ad. This would end up being in the six figures, but they were wondering if I could do it for less.

The discount that we were able to provide to this company relative to what they were used to spending was probably around 75%. Still, the revenue per hour that I actually spent on it was probably around $3,000 an hour—not all of which came to me, by the way. But that sort of disequilibrium, I think, doesn't last forever. A lot of people will figure that stuff out over time. So I do wonder, in that project in particular and in general, whether I'm still being compensated based on assumptions that have not fully taken on board the fact that productivity can—and in some places has—significantly jumped. I think maybe it'll happen in the future. I don't know.

Logan Kilpatrick

Yeah, I think there are still going to be those edges in the future, which is interesting. If anything, I think the pace of innovation is going to go up. Going back to my comment from before, that doesn't mean it's not going to be more difficult, but I do think the pace of innovation is going to continue going up and to the right. Because of that, I think there are going to be a lot of discontinuities and opportunities. Being on the frontier is likely to be disproportionately rewarded because you're using all the tools and stuff like that, which is super interesting to see play out.

But there are so many edges and so many opportunities that are left as the frontier keeps moving forward that, even if you're just showing up today and saying, “I'm not on the frontier,” there are probably 50 things that you could go and explore that end up being super, super interesting and a force multiplier.

Speaking of things that are on the frontier and super interesting, let's talk a little bit about agents. Obviously, everybody's talking about agents in all sorts of different ways. Here's my horseshoe theory of agents: I've found that the latest things, whether it's Claude Code or Jules or any of these more agentic models that take multiple steps and do bigger mini-projects for you, feel much more like the original ChatGPT to me, in that the mode of interacting with them is very turn-based.

What's happening is that the turns are getting bigger. The output is getting bigger and, hopefully, more valuable—hopefully more accurate—in order to be able to do all that stuff and succeed. But you're still on a one-off basis. You're still on the hook as a human for figuring out: Did it do what I wanted? Did I ask it the right thing? Is this actually working for me at all or not? And how do I proceed based on what it did? I have to evaluate that on a step-by-step basis.

And then, in the middle, is where I think people are actually getting scalable automation value. They're not letting the AI choose its own adventure. They're not just saying, “Here's 50 tools and a goal. Go,” which can sometimes create these magic moments, but often doesn't do what you want.

In the middle, it's a much more structured paradigm, whether it's LangChain or whatever, that's like: “We're going to break this thing down into its constituent parts. We're going to have 8 different prompts for the 8 different steps. There might be a couple of little forks or double-back points in there. So we'll give the AI some discretion to choose exactly what route it's going to follow, but it's a pretty on-rails sort of system.” Those seem to be the things, from what I've seen, where people are actually getting to the point where the reliability is high enough that they no longer have to look at the output on a task-by-task basis.

So how would you coach people as they think about the spectrum from the original chatbots that are now familiar to workflows, agents, agentic systems, and autonomous systems? How do you see that spectrum, and where should people be? I'm sure you've got lots of thoughts.

Logan Kilpatrick

Yeah, my take right now is that, with reasoning, it's become very clear that a lot of the scaffolding will move into that layer. You'll send a request and provide a bunch of scaffolding to the model in the reasoning step. Today, it'll have access to search, code execution, a code sandbox, tools, and function calling, but models in general are on this trajectory to become agents out of the box, which is really interesting.

They'll have all these capabilities baked in, and the thing will be able to do a lot of things. Of course, there will be limits to what it does because you don't build everything into it, but by default it will have access to do a lot of things like that. You can imagine having a bunch of other hosted tools and things like that, which then sort of gets the data flywheel spinning, as far as actually being able to build and train the models to go and do that.

Then you can imagine some of those trajectories that you're describing, with the flows that the model goes through and the way that it tries to solve problems. All of that ends up also being upstream of the model. So I do think the models are on that path to be systems and agents out of the box.

But the practical reality is that there will still always be a need for scaffolding. I think it's this balance of how you make the current version of the product that you want work in a way that likely needs to use scaffolding, but you don't build it in a way that ends up being a one-way door. As soon as the model can do that thing, you'd have to fundamentally rework it. Maybe the coding models are good enough that rewriting everything from scratch actually won't be that difficult, and it'll all be fine.

Historically, if you have a larger product, it ends up being really hard. I think this is actually a transition that I've talked to a lot of companies and products about. They're in the middle of this AI 2.0, LLM 2.0 transition moment, where they had actually built a lot of the original tooling around the fact that models weren't good at a lot of things. So they had all this additional scaffolding, all these additional layers and systems, and it was actually a pretty complex system to make LLMs work in production at scale.

Now that the models have become so good and can do a lot of these things natively, you can actually remove a lot of that complexity. Again, depending on the complexity of what you built, that can actually be really, really difficult. I think the folks who built the scaffolding and the complex system did the right thing because they wanted to make that product experience work. They probably benefited from the fact that they were AI-native and powered by it from the beginning, and hopefully won a bunch of customers and business.

But I think you also need to make sure that you can continue to adapt, because I think the models will be able to do more and more and hopefully take on more and more of that burden. I've had an increasing number of conversations with people who are in that boat of going through that transition right now, and it's been specifically because of reasoning that this has become possible for a lot of people.

Nathan Labenz

Yeah. Long context obviously goes hand in hand with that. We've experienced that at Waymark, especially in image processing. I've told this story repeatedly, too, so just again, super briefly: it's kind of crazy how hard I had to work at one point in time just to have any minimal understanding of what a random user-uploaded image was. Now it's like, “Feed 100 images into Gemini Flash,” and it'll just tell you which ones to use. It's really simple.

So what was once a highly scaffolded workflow—and had to be, in order to get to the reliability point—now, for us, is basically just a prompt. That does seem like that sort of cycle will repeat. That sounds like basically what you're describing. Just making sure you're ready to rip out the scaffolding and convert it to a prompt as that moment starts to hit for whatever you're building is the recommendation.

Logan Kilpatrick

Yeah. I had a conversation with Josh Woodward, who runs the Gemini app and Google Labs, and he was saying how this played out for them in NotebookLM as well. Originally, to make those NotebookLM Audio Overviews happen, it was a 14-step process, and there were all these different handoffs and steps in the loop, most of them powered by Gemini.

Today, it's a 4-step process. It's dramatically simplified the level of complexity because the models are just so good at doing a lot of those things now that they don't need to have an entire bespoke system built around writing the transcripts for the Audio Overviews, which is really cool.

You actually feel that in the product experience in some ways, too, where the product experience has become a lot faster, and there are a lot of other things that are possible because you don't need 14 different independent LLM calls that all sort of have to happen in sequence.

Logan Kilpatrick

So it’s been cool to see the product experience actually benefit in a lot of ways from this level of simplicity that’s come as the models have gotten better.

Nathan Labenz

Any other things you think are really interesting, hidden gems, underappreciated, or just strong trends in the agent space? A2A is something I’ve been looking into and honestly haven’t really been able to wrap my head around yet.

Logan Kilpatrick

Yeah, I’ve got a non-agent thing that’s interesting and I’m happy to talk about, but I think on the agent side, at least for A2A, the quick mental model is that there are just parts of the agent-building ecosystem that MCP doesn’t solve. And I think A2A is trying to solve some of those. One of the examples is the auth model and things like that.

So there are parts of the story from putting agents into production at scale that still need to be solved. And I think it’s an open question where MCP is going to go long term. Is it going to do a bunch of those things? Is it going to leave space for other frameworks or standards to solve some of those problems? I don’t think we know yet, so I’m watching closely and interested to see what happens.

Nathan Labenz

Okay. What else is on your mind?

Logan Kilpatrick

Diffusion. Did you see the demo of Gemini?

Nathan Labenz

I did have it in here. I didn’t get to it, but yes. Did you get to play around with it yet?

Logan Kilpatrick

I haven’t used it.

Nathan Labenz

You could. I’d love to have it on that list as well.

Logan Kilpatrick

I’ll get you on the list. Unbelievably, first of all, it does make, in some intuitive sense, a lot more sense to me than the autoregressive model. When I reflect on my own pattern of thinking, I feel like what I’m doing is much more fuzzy and high-level first, and then it gets segmented down into parts. Then I try to do those parts, and at some level I’m writing sentences token by token.

That resonates far more than trying to sit down and write the whole thing linearly from the first token to the last, even with a reasoning model or a scratchpad place to mess around. So, yeah, would it surprise me if, in 2 years, the diffusion paradigm has won because this sort of coarse-to-fine structure turns out to be better? Not really. And damn, is it fast. It’s unbelievably fast. Unbelievably fast.

I’m really excited. I think even if there’s a world where the next-token-prediction paradigm continues, just for people who want to build products that have that level of speed, maybe it’s—I don’t know—it’s unclear at this point what the performance-trade-off characteristics will be. Is the cost going to be the same? All those things. So there are a bunch of open questions.

But assuming you could build product experiences for a similar cost to what they are today, with similar model quality, there are a lot of really interesting product experiences to be built if you have that level of speed. I think that’s actually the thing that could enable this personal generative UI experience to really happen. If the tokens can actually be generated that quickly, rendering on a screen in the blink of a human eye would be really, really cool to see.

So I’m super excited, and I’ll get you on the list for access to that. I think, even if it doesn’t end up working out, it’s just a good reminder that we need to be pushing in different directions, because there are other paradigms that I think could work. Maybe it’s not next-token prediction, and there are a bunch of properties of some of these things, like editing, which is becoming more and more common for a lot of these use cases, that the diffusion model seems to be really well suited for.

Nathan Labenz

Yeah, I suspect in the end it could be quite a bit better for a lot of use cases, too. I mean, it just seems so natural. I always say the transformer is not the end of history. Obviously, there is an attention mechanism in a lot of these diffusion models, too. So attention is still part of what we need.

What do you think we’re missing right now from AGI? Memory is one often-cited candidate. What’s on your list?

Logan Kilpatrick

Memory is definitely one of those. I think AGI is going to end up being much more of a product experience. If I have a hypothesis about how people are going to end up having the AGI moment, my assumption right now—and we’ll see if this plays out—is that someone is going to release a model that ends up being really good. It’s not going to be this thing where everyone says, “We’ve clearly built whatever your definition of AGI is.”

Which is also the problem: everyone now has a different definition of AGI. It’d be easy if we all had the same definition, but we don’t. So that’s the other problem: it’s not going to happen that way. I think it is going to be a product experience. Someone is going to weave together the right components at the product level with a model that’s really smart.

Maybe—I don’t know—the delta in how smart the model needs to be relative to today for this experience to actually work could be just that long context is 50% better and reasoning is 50% better, and then you somehow figure out a way for memory to work. The memory piece is actually a completely different engineering, neuroscience, and human psychology problem: How do you surface the right things at the right time?

I think someone’s going to build that experience, and people are going to say that the feeling of this thing is going to be AGI. Again, it’s really a product experience enabled by a model, but the model itself isn’t able to do all those things. It’s what happens when you take the model and build everything around it, and do it in a really thoughtful way, that people are going to say is sort of the AGI moment for a lot of folks.

So that’s my guess right now. And again, I think the models are doing more and more of this stuff, and you could imagine maybe the models are doing the memory stuff themselves and that gets trained into the model. I think that’s very far out there, but in the short term, it’s definitely going to be a product experience that gets us to AGI. That’s not what I think the AGI narrative is—it’s so model-driven right now. I just don’t think that’s actually how people are going to feel and experience what ends up happening.

Nathan Labenz

Really good memory work is coming out of Google, as you might expect. We recently did an episode on the Titans architecture, and there’s already been a follow-up to that. It’s now up to 10 million tokens of memory and probably able to go beyond that, but they’ve demonstrated up to 10 million tokens with pretty strong memory performance. So, yeah, it could be coming sooner rather than later, but it is kind of a distinct module, right? Obviously, our brains have many modules, too.

Okay, maybe last question, then I’ll let you hit on anything else you want. One of the striking moments from I/O was when Sergey was asked in a little fireside chat what he thinks the future of the web is going to look like in 5 to 10 years. He almost spit out his coffee at that moment, where he was like, “The future of the web?” He’s like, “I don’t think we know what the future of the world is going to look like in 5 to 10 years.”

And that’s a striking reminder that even the people pushing the frontiers of this technology don’t have a crystal ball and don’t really know what exactly we’re getting into. So I wonder what your expectation for the future of your life and your job is in the next, let’s say, 2 to 5 years. Are we going to get a drop-in Logan replacement? I mean, NotebookLM is going to replace me. Do you think you’re on the chopping block in the next 2 to 5 years for AI replacement as well, or how do you see this shaping up?

Logan Kilpatrick

Yeah, I was sitting in the front row of that fireside chat next to Cory, who’s our CTO in DeepMind, and Emanuel Tropa, who drives a bunch of our infrastructure stuff. It was fun to see their reactions as well to some of the conversation.

I have such a fundamentally human-centric view of the world. Even in today’s world, as somebody who builds AI and thinks all the AI products are cool, I write everything I do personally. All of the work that I do, every email that I write, every tweet that I write, is written through my head, and I have, in probably 95% of cases, zero AI assistance involved in that process.

It’s because I have conviction in my worldview and because I have conviction in my tone. I think maybe you can make a loose approximation of someone, but the reality is that I want to be the entity that has agency over the things that come out around who I am. I think that’s fundamentally important.

I think people will end up having this fundamental question for themselves: Who do they want representing them? Even if I have this digital twin that knows all the things, and maybe it could make loose approximations, and I could say, “Yeah, that seems reasonable. I could potentially see myself saying something like that,” do I actually want that thing going and saying those things on my behalf? Probably not. That is a very foreign concept compared to what humans do today.

I think maybe the only exception—not a notable exception to this, but an example against this—is people who run companies. I can imagine that, if you have a large company, you’re like, “Oh, some team or some person is representing Google, as an example.” Someone’s saying, “Maybe I wouldn’t have said it that way, or I wouldn’t have phrased it that way,” and they’re still representing Google as a whole, but they’re sort of an independent agent on behalf of it.

I think that, unless you’ve had that experience, it’s still fundamentally different from the human experience of, “I want some other entity representing me.” It’s not clear to me that people are really going to want that experience. I personally don’t want that experience right now, and that’s my personal opinion. That’s the decision that I’m making.

But I do think it’ll be interesting to see where the balance is, as far as how much people do that. This goes back to—and I have a bunch of these random convictions on my personal website—I think the value of humanity, using you as an example, Nathan, is that this podcast is exponentially more valuable in a world where AI can generate human-sounding things, analyze content, and put together research reports.

The reality is that the next-token prediction coming out of all of those systems isn’t the next-token prediction coming out of your brain. Or, if you want to use that example instead, the diffusion of thoughts coming out of your brain—that’s what I care about. I care about your perspective because you’re another human, and we have shared lived experiences, we’ve done stuff together in person, and all that stuff.

I think there are places where you won’t care about that, maybe because of the type of content. There are certain dimensions where that won’t matter as much. But I really do fundamentally believe that humans are interested in what other humans have to say. When I think about someone sending me AI content that was written or generated by AI, I just care a little less. I’m not really that interested.

I can kind of tell they’re not willing to put in the craft and the time to do something. Why am I that interested in it? Again, there are exceptions to this. Software is a great example: if someone builds great software, do I care whether or not a human wrote it? Not really. Maybe in some cases I’d appreciate it more if a human did it, and maybe not in some cases.

It is interesting. There’ll be a spectrum. I’m not worried. I think the way that people do work will shift in some capacity. I think the value of having a differentiated perspective is also just going to be incredibly beneficial in a world where intelligence isn’t the sort of limiting factor in a lot of ways.

All that to say, I’m excited for another 500 to 600 podcast episodes from you over the next 2 to 5 years.

Nathan Labenz

Yeah, thank you. We’ll see how many I can tick off.

I think there’s a lot to appreciate in your thoughts there, and I’m with you on some portion of it. There’s this whole notion that if people don’t have jobs, they’ll have no meaning, and I’m not on that train. I definitely am on the idea that one of the great things about AI could be that it would allow people to make much more of an effort, or put much more of an emphasis, on making connections with each other.

At the same time, I’m like, I don’t know. LLMs are getting awfully good, and they can handle any topic on demand. I’ve noticed that in myself. NotebookLM isn’t a huge fraction of my listening yet, but it is starting to eat away at my listening to other podcasts.

There are times when I want to know about a specific thing, nobody’s done a podcast on it yet, and NotebookLM will. It has that background knowledge, so even if it’s a little worse in some ways—and maybe it is—the expressiveness of the voice and all that is getting pretty good, too. It hits a spark in some other ways that really matter.

I think the fundamental point here is actually a great example that Sundar gave. This is a great comparison between search and AI chat product experiences. Everyone 2 years ago was saying, “Now that ChatGPT has however many hundreds of millions of monthly active users, search is going to go away.” Yet Google searches are growing, the number of queries on search is growing, and the search business is still growing.

It’s because they’re actually, in some sense, solving fundamentally different problems. That NotebookLM example you gave—“I want on-demand entertainment about this very specific topic, and maybe no one has created that type of content before”—is kind of a different use case.

Maybe this doesn’t fully track across all podcasts, but there are lots of podcasts that I listen to where I’m like, “I just want to hear what this person has to say about this.” That’s the thing I actually care about, and I’m willing to listen to them talk about whatever. But if I’m trying to learn about some very discrete task that I know nothing about, the chance that one of my top 5 favorite podcasts has talked about that is probably slim.

You need some other mechanism in order to do that. Maybe that’s not fully true, so I’m curious what your reaction is to that. But I think that is going to play out across a lot of other domains and dimensions, where this thing is actually net additive.

It puts pressure on things in some capacity because there’s a limited amount of time in the human day, but it doesn’t end up being as disruptive as it would look on paper.

Nathan Labenz

Yeah, I hope not. That time limit is a very hard constraint as it stands right now. I’ve recently started to increase my listening speed; it had been 2x by default.

Actually, YouTube just increased its mobile maximum speed. These are the kinds of things that really move the needle for me. It used to be capped at 2x, and I don’t even know what the maximum is now, but you can go well beyond 2x. Now I can listen to things at 2.5x by default and save myself another 6 minutes.

That saves me 6 minutes on a 1-hour piece of content. This is the way I’m trying to pack more and more in. I don’t think I can go too much farther down that path. At some point, in terms of competition for time and attention, you hit some sort of fundamental limit.

Maybe we get Neuralink working, and that’s the next big unlock, where the bandwidth increases so dramatically that all bets are off. I do think there’s a credible line of thought, honestly, that upgrading human cognition in deeply integrated ways is going to be necessary.

That’s sort of Elon’s brief pitch for why he founded Neuralink: to be able to go along for the ride with the AIs. Especially when you see these diffusion models. The autoregressive ones are already faster; they can write. Gemini can write much faster than I can read, and then the diffusion models are another order of magnitude faster.

The speed of it all is just going to be another wild thing to contend with. Now you’ve got these video models. I just saw one that is generating video dynamically in real time, and you can interact with it.

I share a lot of the excitement and enthusiasm, and I’m also like, I don’t know that I can compete with all this stuff. It just seems like it’s going to get really, really good at everything, be ubiquitous, and be so personalized to each individual user, kind of knowing what they know and what they don’t need to know.

How do I compete as somebody who’s a one-size-fits-all with not a huge audience? There are enough people out there that I certainly can’t customize the podcast for each one. And when AI can do that, I think the point is what people want: the Nathan experience, I think, is the point.

Logan Kilpatrick

That’s at least my fundamental bet and conviction. In the long term, in a world where I could spin up 1,000 podcasts that look something similar to yours, people want your perspective. There’s value in that, even though it’s not the highest-order optimization of the delivery of the content or whatever.

That’s my bet, so we’ll see if that ends up being true. But I have conviction in that bet. So, hopefully it’ll turn out right.

Nathan Labenz

Yeah, I hope you’re right. Well, I know many people want the Logan experience, and I know we’re over time, so I appreciate you sharing so much time and information with us today. I look forward to doing it again in the not-too-distant future.

Logan Kilpatrick

This was great, Nathan. Thank you for having me. It’s fun to chat, and hopefully I’ll see you in person again soon.

Nathan Labenz

Cool. Veo 3 diffusion model. Put me on your list.

Logan Kilpatrick

I will.

And with that, Logan Kilpatrick, thank you for being part of the cognitive revolution. It is both energizing and enlightening to hear why people listen and learn what they value about the show. So, please don't hesitate to reach out via email at tcrturpentine.co or you can DM me on the social media platform of your choice. [Music]

2025年5月15日至22日的十年:Google 50倍 AI 增长与转型——Logan Kilpatrick — 文字稿与摘要 | BidClub