为什么所有人都在复制 Deep Research?
- Gemini Deep Research 瞄准的是那些“从0到50”的工作:这类任务通常要耗掉一个周末和50–60个浏览器标签页。 它运行在经过后训练的 Gemini 1.5 Pro 上,大约用5分钟把多面问题整理成一份有来源的报告;如果用户已经明确知道自己要找什么,普通 Search 仍然更合适。
- 可编辑的研究计划既是用户掌舵的界面,也是昂贵代理任务的一份契约。 用户说“告诉我电池相关的一切”时,Gemini不会连环追问,而是展示拟研究的维度,允许用户通过对话修改——swyx称之为“可编辑的思维链”,Aarush表示认同。大多数用户仍会直接点击Start,但这份计划解释了报告为何会长成现在这样。
- 真正的技术优势在于迭代式规划,而不只是多搜几页网页。 模型会并行探索不同规划分支,读取结果,识别缺口或矛盾,再决定下一步调查什么——比如发现欧盟禁令后,继续核查FDA政策。随后它生成提纲、起草、自我批评并修改,试图从“高层次要点”推进到有依据的二阶结论。
- 长上下文保留正在进行的研究过程,而RAG则负责溢出内容和长期记忆。 近期来源会直接留在上下文中,因为用户可能追问细粒度比较;“10轮之前”的材料则可以转入检索。Sridhar警告称,当查询包含多个属性时,基于余弦相似度的检索效果会减弱;而更新的长上下文模型即便窗口逐渐填满,仍能保持有效。
- 延迟正逐渐变成努力程度的信号,也为研究代理带来了最终的“质量还是表演”问题。 Google测试过15分钟的“hardcore mode”,但最终上线的版本大约耗时5分钟,并将上限设在10分钟以内;出人意料的是,用户并没有一味要求即时答案,反而可能认可看得见的工作过程。团队仍未找到在更广泛探索与更深入核验之间分配算力的最佳方案,而如果提供“max power”,用户大概会直接按下去。
- 评估依旧顽固地依赖人,因为有效输出空间太大,不可能由一个基准分数概括。 自动化检查会监控计划长度、迭代步骤以及行为分布变化,人工评审则从覆盖面、完整性和依据充分性等维度打分,评估范围从广泛发现选项到狭窄而深入的调查。“如果我在HLE上表现很好,并不代表我就是优秀的深度研究员。”
- 更大的机会在于连接专有信息的个性化、多模态研究代理,而不只是更好的网页摘要器。 团队希望针对15岁青少年和PhD提供不同输出,并生成图表、地图、图片或交互式界面,再叠加私有文件和付费订阅内容。swyx面向投资人的判断异常直接:这可能是“第一个真正达到产品市场匹配的代理”,目前对部分用户而言每月200美元已经合理,如果能力大幅提升,价格或许能到2,000美元。
1. Deep Research 从协商工作开始
Selvan给出的产品定位是:Deep Research是一名个人研究助理,帮助用户在新主题上快速“从0到50”。它浏览约5分钟,再返回一份用户可以审阅、追问或重塑的报告。
它与 Google Search 的边界在于用户意图是否明确。用户已经知道自己要找什么时,Search仍然是终点;Deep Research处理的是多面向的探索过程——这类任务会打开“50–60个标签页”,耗掉一个周末,最后往往还会让用户放弃。
在投入5–10分钟和大量算力之前,Gemini会先提出研究计划。“告诉我电池相关的一切”可能指创新、化学原理,也可能指某项具体技术,因此模型先给出初步拆解,而不是让用户经历一连串追问。
swyx称这个界面为“可编辑的思维链”,Aarush表示认同。早期测试显示几乎没人主动编辑,团队于是加了一个显式按钮;即便用户像点击“I’m Feeling Lucky”一样直接按Start,这份计划仍然是透明机制,也是用户接受的任务契约。
2. 迭代式规划把网页浏览变成研究
Sridhar的技术解释是:被接受的计划包含可并行执行的分支,模型主要通过两种能力展开——搜索,以及深入阅读某个选定页面。关键在于,它会先读取此前的结果,再决定下一步行动。
食品监管演示展示了这一机制:如果一次搜索发现欧盟委员会禁止某些添加剂,Gemini可以继续判断是否需要核查FDA是否也采取同样做法。Sridhar认为,没有这种迭代式的依据校验,报告就会不完整,最后只能退化成“高层次要点”。
初始探索通常是广度优先,但团队并没有把这种行为写死。Gemini会先采样计划中的每个维度,再对信息不完整或来源冲突的地方深入追查;研究结束后,它构建提纲、起草报告、自我批评并完成修改。
最终的牛奶和肉类报告不只是罗列规则,还推导出一组哲学上的分歧:欧盟采取预防性路径,即便证据尚不充分也倾向于先行禁止;美国则更偏反应式路径,在危害得到证明前允许其继续存在。Selvan认为,这正是目标中的“二阶洞察”。
3. 报告是工作空间,背后由分层记忆支撑
Selvan把后续操作分成3类:找回已经遇到过的事实、针对实质性新增范围启动新一轮研究,或直接编辑产物——压缩、删除或增加章节。报告与聊天并排呈现,正是为了同时支持这3种行为。
如果用户把比较范围从美国和欧盟扩展到亚洲,Gemini会判断现有研究是否足够,还是需要重新浏览。所有近期读过的网站仍然可用,因此缺少某个细节时,系统可以快速回答,无需重复最初那次5分钟的任务。
当反复执行研究任务开始逼近100万–200万token的上下文窗口时,团队会使用内部检索。经验法则是:复杂比较尽量把近期工作留在上下文中,再把大约“10轮之前”的材料放到RAG之后;相关且连贯的项目可以留在同一条线程里,因为早期发现的细分信息可能会指导后续搜索。
Sridhar对RAG的限定是:当查询本身包含多个属性时,点积或余弦距离检索会遇到困难。新一代模型在更长上下文中也能更深地保留细粒度信息,因此检索何时优于直接放入上下文的临界点正在后移。
4. 网页呈现与多模态仍是现实取舍
Deep Research同时提供Markdown和HTML两种呈现。Selvan表示,Markdown有助于减少纯HTML中的噪声;对话中提到JavaScript和Tailwind CSS就是这类噪声的例子,但嵌入式HTML片段仍可能需要原生处理。
目前描述的产品还不包含视觉能力。Selvan承认,典型问题是关键信息被困在JPEG图片里;但他认为,对今天的主要使用场景而言,渲染页面会增加延迟,而价值只集中在“尾部的一小部分”。
Fanelli反驳称,代理本来就已经拥有数分钟的延迟预算。Selvan表示,模型的VQA能力正在提升,但仍把页面渲染留作未来的取舍,而不是当前能力。
5. 评估从研究行为开始,而不是从垂直领域开始
输出熵让评估“很难做”。自动化评分器可以识别行为漂移——研究计划长度、步骤数量、规划耗时或迭代搜索深度的变化——但这些分布只能说明发生了变化,无法说明是“变好还是变坏”。
因此,人工评审仍然不可或缺,他们会根据产品定义的覆盖面、完整性和依据充分性等质量指标进行评分。
Sridhar的评估本体不使用旅行或购物等垂直领域标签。一端是广而浅的选项探索:找出大量夏令营并分别总结;另一端是狭窄而深入的理解,中间则由比较任务以及不同广度和深度的组合填充。
复合项目会同时检验多种模式。规划一场里斯本婚礼,可能需要在10个子任务中分别研究婚礼策划师、场地和餐饮。系统没有硬性的对话轮数上限,但目前大多数用户并不会深入太多。swyx认为,完成后的文档在视觉上更像终点,而不是“起点”;Mukund也承认,用户体验还可以更积极地邀请用户继续探索。
6. 有用的延迟正在颠覆Google的速度正统
swyx指出了一个“反常激励”:搜索70个网站或运行1小时的代理,即使其中30个来源都无关,也可能看起来更有能力。低效暂时会被解读为勤勉,但他预计,当用户开始追问为什么同样的质量不能更快交付时,这段蜜月期就会结束。
Google最初做了两个版本:一个运行约15分钟的“hardcore mode”,以及最终上线的约5分钟版本。Selvan曾要求工程团队把硬停止时间设在10分钟以内,因为他以为用户会放弃任何更慢的产品。
真正令人意外的是,Jason Calacanis问Google是否其实在10秒内就生成了答案,只是故意延迟展示。这与团队在Assistant及其他Google产品上的经验完全相反——在那里,延迟越低,满意度和留存通常越高;而在这里,看得见的努力本身就有价值。
Selvan把算力决策定义为探索与核验之间的取舍。关于美联储利率变化和中产家庭收入的问题,需要精确回答并引用历史来源;生日餐厅则可以更宽松。理想的代理应当自行推断这项取舍,因为如果设置一个显式的“max power”控制,用户会倾向于把所有任务都调到最大。
7. 掌舵、专业化与持久执行构成产品护城河
swyx最强烈的用户体验批评是:用户应该能在研究运行期间修改计划。Devin会展示实时计划,并接受任务进行中的修正;他认为,如果研究最终需要运行1小时,Gemini就应当像一名实习生:带着发现回来,报告遇到的问题,并请求下一步指示。
Selvan同意,任务越长,执行过程中的掌舵就越有价值:当前几分钟的研究阶段几乎没有干预空间,但1小时的任务可以让代理带着发现和问题回来,等待用户指示。swyx提出的设计是一个实时更新的计划,它可以安排下一项任务,同时不锁死聊天窗口。
这套体验底层是一个平台级的异步系统:用户可以离开、关掉电脑,完成后通过手机收到通知。5分钟或6分钟的任务必然会失败,因此系统会保存状态、选择性重试,并避免丢弃已经完成的研究。Mukund表示,系统已经能够稳定处理数百次LLM调用,也足够灵活,可以支持未来运行1小时甚至数天的任务。
购物场景说明了专业化为何重要。Deep Research不擅长根据外观挑鞋,但在HVAC系统上更有吸引力,因为规格、额定电压以及寻找能够安装的承包商,比外观更关键。Selvan将其概括为“选项探索”,同样适用于商品、奖学金或夏令营。
8. 更好的推理必须带来有来源的新颖结论,而不是基准测试表演
这款产品并非纯粹使用 Gemini 1.5 Pro:Selvan称其为一个经过后训练的版本,并表示并不存在所谓特殊访问权限。他认为,广义上的系统可以通过工具调用和微调复现,包括使用Gemma;但要实现稳定规划和可靠性,仍需要大量后训练。
思考模型在迭代式网页搜索之外,又引入了第二种推理时算力。Mukund表示,模型可以更多调用自身记忆,生成更强的二阶洞察,但事实仍需核验;即使模型记得某个正确的美联储统计数据,没有.gov来源也仍然可疑。难点在于平衡模型记忆与事实依据。至于是否会超越1.5,他的回答是“stay tuned”。
可泛化的迭代式规划是最难的建模问题。为每个领域或研究本体分别训练轨迹将是“噩梦”,因此团队强调高效利用模型记忆、进行数据增强,并通过后训练恰到好处地教会模型这种行为,同时不抹掉预训练继承的能力。
基准测试仍然有助于凝聚研究人员。Selvan回忆,MLPerf竞赛曾推动TPU性能快速提升;但产品有效性是另一回事。团队希望避免针对“科比进入联盟那天,总统的侄子是谁”这类不自然的琐碎问题优化,Sridhar则警告,文本输出熵太高,会让验证和公平比较都变得困难。
要实现超越网页综合的信息发现,既需要二阶推理,也需要能够检验假设的环境。代码和数学有沙盒与验证器,化学却没有等价的合成实验室。Sridhar的条件很明确:代理需要一个试验场、准确反馈和重复实验,只有这样,“新想法”才不至于沦为未经验证的漂移。
Selvan的路线图正转向个性化和生成式UI:15岁青少年和post-doc应当收到不同的研究报告,而图表、地图、图片和交互式结构应取代千篇一律的文本文档。
开放网络最终会成为受限语料库。高价值的产业研究存在于公司文件、付费订阅和私人资料库中;但Sridhar提醒,现在还处于将代理平台化为横向组件的早期阶段。他给构建者的建议是,选择一个使命,然后“把这一件事做到极致”。
This is Alessio, partner and CTO of Decibel Partners, and I'm joined by my co-host, swyx, founder of Small Eye.
Hey. Today we're very honored to have Aarush Selvan and Mukund Sridhar from the original Deep Research team in our studio. Welcome.
Thanks. Thanks for having us.
Thanks for making the trip up. I was fortunate enough to be one of the early beta testers of Deep Research when it came out, and I was very keen on it. Even at the end of last year, people were already saying it was one of the most exciting agents coming out of Google. We previously had Riza and Osama from the NotebookLM team, and I think this is part of an increasing trend where Gemini and Google are shipping interesting user-facing products that use AI. Congratulations on your success so far.
It's been great. Thanks so much for having us here. We're excited.
Thanks for making the trip up. I'm also excited for your talk that's happening next week. Obviously, we have to talk about what exactly it is, but I'll ask you about that toward the end. For now, we have the screen up, so maybe we can start at a high level. For people who don't yet know, what is Deep Research?
1. Deep Research Becomes Your Assistant
Deep Research is a feature where Gemini can act as your personal research assistant to help you learn about any topic more deeply. It's really helpful for queries where you want to go from 0 to 50 very quickly on something new.
The way it works is that it takes your query, browses the web for about 5 minutes, and then outputs a research report for you to review and ask follow-up questions about. This is one of the first times something has taken 5 or 6 minutes to perform research for you, so there are a few challenges that brings. You want to make sure you're spending that time and compute doing what the user wants.
There are also UX design considerations that we can talk about as we go through an example. Then there are challenges in browsing the web. The web is extremely fragmented, and being able to plan iteratively as you pass through that noisy information is a challenge by itself.
This is the first time Google is automating the way you search. You're supposed to be the experts at search, but now you're meta-searching—determining the search strategy.
We see it as 2 different use cases. There are things where you know exactly what you're looking for, and search is still probably one of the best places to go. I think Deep Research really shines when there are multiple facets to your question and you spend a weekend opening 50 or 60 tabs. Many times, I just give up, and we wanted to solve that problem and give people a great starting point for those kinds of journeys.
Do we want to start a query so that it runs in the meantime and then we can chat over it?
Here's one query that we love to test: super-niche, random things where there's no Wikipedia page already about the topic. That's where you'll see the most lift from a feature like this. I've come up with this query—it's actually Mukund's query, which he loves to test: “Help me understand how milk and meat regulations differ between the US and Europe.”
2. Research Plans Come First
What's nice is that the first step is where it puts together a research plan that you can review. This is its guide for how it's going to carry out the research.
This was a pretty decently well-specified query, but let's say you came to Gemini and said, “Tell me about batteries.” That query could mean so many different things. You might want to know about the latest innovations in battery technology, or you might want to know about a specific type of battery chemistry.
If we're going to spend 5 to 10 minutes researching something, we want to understand exactly what you're trying to accomplish and give you an opportunity to steer where the research goes. If you had an intern and asked them this question, the first thing they would do is ask you a bunch of follow-up questions: “Help me figure out exactly what you want me to do.”
We thought, why don't we have the model produce its first stab at the research query—how it would break the question down—and then invite the user to engage with how they want to steer it?
Many times, when you try to use a product like this, you don't know what questions to ask or what things to look for. We made the decision deliberately that instead of asking users follow-up questions directly, we would lay out what we would do and show the different facets.
Here, for example, it could be what additives are allowed and how that differs, or labeling restrictions on products in the US and the EU. The aim is to tell the user a little bit more about the topic and get steered at the same time. We also elicit follow-up questions.
It's kind of like editable chain-of-thought.
Exactly.
We were talking to you about your top tips for using Deep Research, and your number-one tip is to edit the plan. Just edit it, right?
You can actually edit it conversationally. We put in a button here just to draw users' attention to the fact that they can edit it. In early rounds of testing, we saw that no one was editing, so we thought that if we put a button here, maybe people would engage with it.
I just hit Start. I think we see that too. Most people hit Start. It's like the “I'm Feeling Lucky” button.
All right, I can just add a step here, and what you'll see is that it should refine the plan and show you a new proposal.
Here we go. It added step 7: “Find information on milk and meat labeling requirements in the US and EU.” Or you can just go ahead and hit Start.
I think it's still a nice transparency mechanism, even if users don't want to engage. You still understand why you're getting the report you're going to get, which is useful.
While it browses the web, Mukund, you should maybe explain how it browses. We show the websites it's reading in real time.
I'll preface this with the fact that I forgot to explain the roles. You're a PM?
Yes.
Okay.
Just for people who don't know, we maybe should have started with that. We know each other's work sometimes as well, but that's how it is. More or less, that's the boundary.
3. Research Browsing Goes Iterative
What's happening behind the scenes is that we give the model this research plan as a contract—something that has been accepted. If you look at the plan, there are things that are obviously parallelizable, so the model figures out which of the substeps it can start exploring in parallel.
It primarily uses 2 tools. It can perform searches, and it can go deeper within a particular web page of interest. Oftentimes, it will start exploring things in parallel, but that's not sufficient. Many times, it has to reason based on information it has found.
In this case, one of the searches could reveal that the EU Commission has banned certain additives. The model then wants to check whether the FDA does the same thing. This notion of being able to read outputs from the previous turn, ground on them, and decide what to do next was key. Otherwise, you have incomplete information and your report becomes a set of high-level bullet points.
We wanted to go beyond that blueprint and figure out what the key aspects were. This happens iteratively until the model thinks it has finished all its steps. Then we enter analysis mode. There can be inconsistencies across sources, so the model comes up with an outline for the report and starts generating a draft. It then tries to revise that by self-critiquing to finalize the report. That's broadly what's happening behind the scenes.
What's the initial ranking of the websites? When you first started it, there were 36. How do you decide where to start, since it sounds like the initial websites carry a lot of weight because they inform what comes next?
In the initial turns—again, this isn't something we enforce; it's mostly the model making these choices—the model typically explores all the different aspects in the research plan that was presented. We get a breadth-first view of the different topics to explore.
In terms of which ones to double-click on, it comes down to what the model learns every time it searches. It gets some idea of what a page contains, and depending on what it finds, there may be inconsistencies or partial information. Those are the pages it double-clicks on.
It can continuously search and browse iteratively until it feels like it's done.
I'm trying to think about how I would code this. Do you think we could do this with the Gemini API, or do you have some special access that we can't replicate? If I model this with tool calls for search, double-click, and whatever else, would that work?
I don't think we have special access per se. It's pretty much the same model. We of course have our own post-training work that we do, and y'all can also fine-tune from the base model and so on.
I don't know that we can do all this fine-tuning.
Well, if you use our Gemma open-source models, you could fine-tune.
Yeah, yeah.
Yeah, so I don't think there's special access per se, but a lot of the work for us is first defining that there needs to be a research plan and how you go about presenting that, and then doing a bunch of post-training to make sure it's able to do this consistently, well, and with high reliability.
Okay, so Gemini 1.5 Pro with Deep Research is a special edition of Gemini 1.5 Pro?
Yes, so it's not purely Gemini 1.5 Pro; it's post-trained.
This also explains why you can't just toggle on Gemini 2.0 Flash.
Right.
Yeah, but I assume you have the data, and you know it should be doable.
Yep. There's still this question of ranking.
Right. And, oh, it looks like you're already done?
Yeah, yeah, we're done. We can look at it. So, let's see. It's put together this report, and what it's done is sort of broken it down. It started with milk regulation, and then it looks like it goes into meat, probably further down, covering how the U.S. approaches the problem of how to regulate milk, comparing it with the EU, and then, like I said, going into meat production.
What's nice is that it also reasons over why there are differences. I think what's really cool here is that it's showing a difference in philosophy between how the U.S. and the EU regulate food. The EU would adopt a precautionary approach, so even if there's inconclusive scientific evidence about something, it's still going to prefer to ban it, whereas the U.S. takes a reactive approach, allowing things until they can be proven to be harmful.
What's nice is that you also get the second-order insights from what it's putting together. So, yeah, it's kind of nice. It takes a few minutes to read and understand everything, which makes for a quiet period during a podcast, I suppose.
Oh, well, this is fun.
But, yeah, this is kind of how it looks right now. From here, you can keep the usual chat-and-iterate flow. Compared to other platforms, it's kind of like Anthropic Artifacts or a ChatGPT Canvas, where you have the document on one side and the chat on the other, and you're working on it.
4. Reports Become Ongoing Work
This is something we thought a bit about. One of the things we feel is that your learning journey shouldn't just stop after the first report, and so what you probably want to do is, while reading, be able to ask follow-up questions without having to scroll back and forth.
There are broadly a few different kinds of follow-up questions. One type is that maybe there's a factoid you want that isn't in here, but it was probably already captured as part of the web browsing that it did. We actually keep everything in context; all the sites that it has read remain in context. So, if there's a piece of missing information, it can just fetch that.
Another kind is, “Okay, this is nice, but I actually want to kick off more Deep Research.” For example, “I also want to compare the EU and Asia in how they regulate milk and meat.” For that, you'd want the model to recognize that this is sufficiently different and that it needs to do more Deep Research to answer the question; it won't find that information in what it has already browsed.
The third is that maybe you just want to change the report. Maybe you want to condense it, remove sections, add sections, and iterate on the report that you got. We've broadly tried to teach the model to be able to do all 3, and this side-by-side format allows the user to do that more easily.
Yeah. So, as a PM, there's an Open in Docs button there, right? How do you think about what belongs there versus—kind of sounds like the condensing and things should be in Google Docs?
Yeah. Bard extensions is different; it's just an amazing editor. Sometimes you just want to directly edit things, and now Google Docs also has Gemini in the side panel. The more we can help this be part of your workflow throughout the rest of the Google ecosystem, the better, right?
One thing we've noticed is that people really like that button and really like exporting it. It's also a nice way to save it permanently, and when you export, all the citations carry over. In fact, I can just run it now, which is also really nice.
Gemini Extensions is a different feature. That's really about Gemini being able to fetch content from other Google services in order to inform the answer. That was actually the first feature that we both worked on on the team: building extensions in Gemini. Right now, we have a bunch of different Google apps, as well as, I think, Spotify and a couple of others. I don't know if we have any Samsung apps as well.
Who wants Spotify? I have this whole thing about how much I love Spotify. What's that in your Deep Research?
In Deep Research, I think less. The interesting thing is that we built extensions and weren't really sure how people were going to use them, and a ton of people are doing really creative things with them. A ton of people are also doing things they loved on the Google Assistant, and Spotify—playing music on the go—was a huge value.
Oh, it controls Spotify?
Yeah. Deep Research purely uses Search.
But this is Search. Otherwise, you can have Gemini go—you have YouTube, Maps, and Search. There's also Gemini 2.0 Flash Thinking Experimental with Apps, the newest—yeah, longest model name that has been launched. Gmail is an obvious one, and Calendar is an obvious one.
Exactly. You know, those are the ones I want.
Yeah, Spotify. Fair enough.
Then, obviously, feel free to dive in on your other work. You're not just doing Deep Research, right? You were just kind of focusing on Deep Research here. I actually asked for modifications after this first run, where I was like, “Oh, you stopped. I actually want you to keep going. What about these other things?” And then continue to modify it. It really felt like a little bit of a copilot-type experience, but more like an agent that would research. I thought it was pretty cool.
Yeah, I think one of the challenges is that currently we kind of let the model decide, based on your query, among the 3 categories. There is a boundary there: depending on how deep you want to go, you might just want a quick answer versus kicking off another Deep Research. Even from a UX perspective, I think the panel allows for this notion that not every follow-up is going to take you 5 minutes.
Right now, it doesn't do any follow-up search, does it?
It always does. It depends on your question. Since we have the liberty of really long-context models, we actually hold all the research material across turns. If it's able to find the answer in things it has already found, you're going to get a faster reply. Otherwise, it's just going to go back to planning.
A bit of a follow-up: since you talk about the product context, I had 2 questions. One, do you have an HTML-to-Markdown transform step, or do you just consume raw HTML? There's no way you consume raw HTML, right?
We have both versions. The models are getting much better at natively understanding these representations. The Markdown step definitely helps because, as you can imagine, there's a lot of noise with pure HTML.
JavaScript, Tailwind CSS, exactly.
When it makes sense, we don't artificially try to make it hard for the model, but sometimes it depends on the kind of access we get as well. For example, if there's an embedded snippet that's HTML, we want the model to be able to work on that too.
And no vision yet?
Currently, no vision yet.
The reason I ask all these things is because I've done the same, but I haven't done vision.
The tricky thing about vision is that I think the models are getting significantly better, especially if you look at the last 6 months, at natively being able to do VQA stuff and so on. But the challenge is the trade-off between having to actually render it and the added latency versus the value add you get.
You have a latency budget of minutes.
Yeah, yeah, yeah, it's true. In my opinion, the places you'll see a real difference are in a small part of the tail. In this kind of open-domain setting, if you just look at what people ask, there are definitely some use cases where it makes a lot of sense to do it, but I still feel it's not in the head cases. We do it when we get there, I guess.
The classic is that it's a JPEG with some important information, and you can't touch it.
Yeah.
And then the other technical follow-up was just: you have a 1 million–2 million-token context. Has it ever exceeded 2 million? What do you do there?
Yeah, so we had this challenge sometime last year when we started wiring up this multi-turn flow. We said, “Hey, let's see how long somebody on the team can take Deep Research.”
What's the most challenging question you can ask that takes the longest?
No, we keep asking follow-ups. For example, here you could say, “Hey, I also want to compare it with—”
Okay, so you're guaranteed to bust it.
Yeah, yeah, yeah. We also have retrieval mechanisms if required. We natively try to use the context as much as it's available, beyond which we have a RAG setup to figure out.
Okay. Is this all in-house tech?
Yes, yes.
5. Long Context Meets RAG
What are some of the differences between putting things in context versus RAG? When I was in Singapore, I went to the Google Cloud team, and they talked about Gemini plus grounding. Is Gemini plus search kind of like Gemini plus grounding? How should people think about the different shades of “I’m doing retrieval on data” versus “I’m using Deep Research” versus “I’m using grounding”? Sometimes the labels can be hard, too.
Let me try to answer the first part of the question. I’m not fully sure about the grounding offering, so I can at least talk about the first part.
I think you’re asking about the difference between when you would do RAG versus relying on the long context. I think we all get that. I was more curious, from a product perspective, when you decide to do RAG versus not. Do you get better performance just by putting everything in context?
The tricky thing with RAG is that it really works well because a lot of these systems are doing cosine distance, a dot-product kind of thing. That gets challenging when your query has multiple different attributes. The dot product doesn’t really work as well. At least for me, that’s my guiding principle for when to avoid RAG.
The second thing is that, with the initial generations of these models, even though they offered long context, you would see some kind of decline as the context kept growing. But as newer-generation models came out, they became really good at picking out fine-grained information, even if you kept filling in the context. So those are my guiding principles.
Just to add to that, a simple rule of thumb that we use is that if it’s the most recent set of research tasks, where the user is likely to ask lots of follow-up questions, that should be in context. But as stuff gets 10 turns ago, it’s fine if that stuff is in RAG, because it’s less likely that the user needs to do very complex comparisons between what’s currently being discussed and the stuff that they asked about 10 turns ago. That’s just a very simple rule of thumb that we follow.
From a user perspective, is it better to just start a new research instead of extending the context?
I think that’s a good question. If it’s a related topic, there’s a benefit to continuing with the thread, because the model, since it has this in memory, could figure out, “I found this niche thing about milk regulation in the U.S. Let me check if your follow-up country or place also has something like that.” You might not catch those things if you start a new thread.
It really depends on the use case. If there’s a natural progression and you feel like this is part of one cohesive project, you should just continue using it. If my follow-up turn is, “I’m just going to look for summer camps or something,” then I don’t think it should make a difference. But we haven’t really pushed that or tested that aspect of it. For us, most of our tests are more natural transitions.
How do you evaluate Deep Research?
6. Evaluation Needs Human Judgment
Oh boy. This is a hard one. I think the entropy of the output space is so high. People love auto-raters, but they bring their own set of challenges.
For us, we have some metrics that we can automatically generate. When we do post-training and have multiple models, we want to make sure that the distribution of certain statistics—such as how long the model spent planning and how many iterative steps it takes on a dev set—doesn’t change unexpectedly. If you see large changes in the distribution, that’s an early signal that something has changed. It could be for better or worse.
So every time you have a new version, you run it across a test suite of cases and see how long it takes?
Yeah, we have a dev set and automatic metrics that can detect behavior end to end. For example, how long is the research plan? Does a new model produce a much longer plan? Not just in terms of the number of characters, but in terms of the number of steps in the research plan. As we spoke about, the model iteratively plans based on previous searches, so we look at how many steps that goes on average over some dev set.
There are some things like this that you can automate. Beyond that, there are auto-raters, but we definitely do a lot of human evaluation. We’ve defined, with the product team, certain things that we care about, and we’ve been very opinionated about them: Is it comprehensive? Is it complete? Is it grounded? So it’s a mix of these two approaches.
Is this where the other challenge is that sometimes you just have to have your PM review examples?
Yeah, exactly.
Broadly, what we try to do for the evaluation question is think about all the ways in which a person might use a feature like this. We came up with what we call an ontology of use cases. We try to stay away from verticals like travel or shopping and instead focus on the underlying research behavior that a person is engaging in.
On one end, there are queries where you’re going very broad but shallow. Shopping queries are an example of that: “I want to find the perfect summer camp. My kids love soccer or tennis.” You want to find as many different options as possible, explore all the options available, and then synthesize a TL;DR about each one. Those are the kinds of journeys where you open many Chrome tabs but then need to take notes somewhere about what’s appealing.
On the other end of the spectrum, you have a specific topic that you want to go very deep on and really understand. There are all sorts of points in the middle, too. Maybe you have a few options that you want to compare, or you don’t want to go super deep on one topic but want to cover slightly more topics.
We developed this ontology of different research patterns. For each one, we came up with queries that would fall within it, and that became the evaluation set. We then run human evaluations on it to make sure we’re doing well across the board.
You mentioned three things. Is it literally three, or is it three out of 20 things? How long is the conversation?
I basically just told you the full set. No, I told you the extremes, and then we had several midpoints. So it goes from something super broad and shallow to something very specific and deep.
We weren’t actually sure which end of the spectrum users would really resonate with. On top of that, you have compounds of those. You can have things where you want to make a plan. A great example is, “I want to plan a wedding in Lisbon, and I need you to help with these 10 things.” That becomes a project with research enabled. It needs to research planners, venues, and catering.
There are compounds that emerge when you start combining these different underlying ontology types, and we also thought about that when we put together our evaluation set.
What’s the maximum conversation length that you allow or design for?
We don’t have any hard limits on how many turns you can do. One thing I will say is that most users don’t go very deep right now. It might just be that it takes a while to get comfortable, and then over time people start pushing it further and further. But right now, we don’t see a ton of users going very deep.
I think the way that you visually present it suggests that you stop when the document is created. You don’t really encourage ongoing chats. The UI doesn’t encourage ongoing chats, even though it was designed like a project.
I think there are definitely things we can do on the UX side to invite the user to say, “Hey, this is the starting point. Now let’s keep going together. Where else would you like to explore?” There are definitely some explorations we could do there.
In terms of how deep people go, I don’t know. We’ve seen people internally dogfood this and push it quite a long way. I think the other thing that will change with time is people uncovering different ways to use Deep Research.
For the wedding-planning example, that’s not one of the first things that comes to mind when we tell people about this product. As people explore and find that it can do these various different kinds of things, some of that can naturally lead to longer conversations. Even for us, when we dogfooded this, we saw people use it in ways we hadn’t really thought of before.
Yeah. That was because this was new for us, and we didn't know: Would users wait 5 minutes? What kinds of tasks would they try that take 5 minutes? Our primary goal was not to specialize in a particular vertical or target 1 type of user. We just wanted to put this in the hands of a busy-parent persona and various different user profiles, see what people tried to use it for, and learn more from that.
How does the ontology of your use case tie back to Google's main product use cases? You mentioned shopping as 1 ontology, right? There's also Google Shopping. This sounds like a much better way to do shopping than going on Google Shopping and looking at the wall of items. How do you collaborate internally to figure out where AI goes?
When I said shopping, I was trying to boil down what exactly the behavior is underneath. That's really around what I called options exploration: You want to see whether you're shopping for summer camps, a product, or scholarship opportunities. It's sort of the same action: You need to sift through a lot of information to curate a bunch of options for yourself. That's what we tried to distill, rather than thinking about it as a vertical.
Google Search is awesome if you want really fast answers. You've got high intent—you know exactly what you want—and you want super up-to-date information. I still use Google Shopping because it's multimodal, and you can see the best prices and stuff like that. I think creating a good shopping experience is hard, especially when you need to look at the thing. If I'm shopping for shoes, I don't want to use Deep Research because I want to look at how the shoes look.
But if I'm shopping for HVAC systems, great. I don't care how it looks, and I don't even know what it's supposed to look like. I'm fine using Deep Research because I really want to understand the specs, how exactly it works, the voltage rating, and stuff like that. I also need to look at contractors who know how to install each HVAC system. I'd say where we really shine when it comes to shopping is at the more complex end of the spectrum, where it matters less what it looks like. It's perhaps less on the consumer-y side of shopping.
One other thing I've observed about the metrics—or the communication of what value you provide—and this also goes into a latency budget, is that I think there's a perverse incentive for research agents to take longer and be perceived as better. People say, “You're searching 70 websites for me,” but 30 of them are irrelevant. I feel like we're in a honeymoon phase where you get to pass all this off. Being inefficient is actually good for you because people care about quantity and not quality. They're like, “This thing took an hour for me; it's doing so much work,” or, “It's slow.”
7. Latency Changes the Product
That was super counterintuitive for us. The 1st time I realized what you were saying was when I was talking to Jason Calacanis, and he was like, “Do you actually just make the answer in 10 seconds and then make me wait for the balance?” We hadn't expected people to value the work it was putting in because you're not actually worried about it. We were really worried about it.
We had actually built 2 versions of Deep Research. We had a hardcore mode that took 15 minutes, and what we actually shipped was something that took 5 minutes. I even went to Eng and said, “There has to be a hard stop, by the way. It can never take more than 10 minutes.”
Yep. Because I think at that point, users will just drop off.
Yep. But what's been surprising is that that's not the case at all; it's been going the other way. When we worked on Assistant, at least, and other Google products, the metric was always that if you improve latency, all the other metrics go up. Satisfaction goes up, retention goes up, all of that.
When we pitched this, it was like, “Hold on. In contrast to all Google orthodoxy, we're actually going to slow everything right down, and we're going to hope that users will still stay engaged.”
Not on purpose.
I think it comes down to the trade-off: What are you getting in return for the wait? From an engineering/modeling perspective, it's trading off inference compute and time to do 2 things: either explore more, to be more complete, or verify more on things that you probably know already. It's a spectrum, and we don't claim to have found the perfect spot. We had to start somewhere, and we're trying to see where there are probably some cases where you care about verifying more than others.
In an ideal world, based on the query and conversation history, you know what that is. I think it basically boils down to 3 things. From a user perspective, am I getting the right value-add? From an engineering/modeling perspective, are we using the compute to explore effectively, and also to verify and go in depth on things that are vague or uncertain in the initial steps?
The other point about the number of websites is that it also comes with a trade-off. Sometimes you want to explore more early on before you narrow down on the sources or topics you want to go deep on. If you look at how Deep Research works for most queries, initially it goes broad. It tries to explore all the different topics mentioned in the research plan. Then you see the choices of websites getting a little narrower around a particular topic or entity that it has come across, and so on. That's roughly how the number fluctuates. We don't do anything deliberate to either keep it low or try to increase it.
Would it be interesting to have an explicit toggle for the amount of verification versus the amount of search?
I think so. Users would always just hit that toggle. I worry that if you give them a Max Everything button, they're always going to hit it. So the question is: Why don't you just decide from the product point of view where the right balance is?
OpenAI has a preview of this—I think it's either in Anthropic or OpenAI—and they have a preview of a model-routing feature where you can choose intelligence, cheapness, and speed. They're all 0-to-1 values, so you just choose 1 for everything. Obviously, they're going to do some normalization, but users are always going to want 1, right?
We've discussed this a bit. If I wear my pure-user hat, I don't want to set anything. I come with a query; you figure it out. Sometimes, based on the query, there will be different requirements. If I'm asking, “How do rising rates from the Fed affect household income for the middle class, and how has that traditionally happened?” you want to be very accurate and precise about the historical trends.
Whereas there's a little more leeway when you're saying, “I'm trying to find businesses near me to celebrate my birthday,” or something like that. In an ideal world, we figure out that trade-off based on the conversation history and the topic. I don't think we're there yet as a research community, and it's an interesting challenge by itself.
This reminds me a little bit of the NotebookLM approach. We also asked Riza about this, and she was like, “People just want to click a button and see magic.” People just want to hit Start every time, right? Most people don't even want to enter the system.
My feedback, if you want feedback, is that I'm still kind of a champion for Devin. Devin will show you the plan while it's working on the plan, and you can say, “Hey, the plan is wrong,” and chat with it while it's still working. It will live-update the plan and then pick off the next item on the plan. Both also have this; that's the most default experience.
I think you should never lock the chat. You should always be able to chat with the plan and update the plan, and the plan scheduler—or whatever orchestration system you have under the hood—should just pick off the next job on the list. That's my 2 cents.
Especially if we spend more time researching. If you watch that query we just did, it was done within a few minutes. By the time it left the research phase, your opportunity to chime in and steer was less. But imagine a world where these things take 1 hour and you're doing something really complicated. Then your intern would totally come check in with you: “Here's what I found. Here's some hiccups I'm running into in the plan. Give me some steer on how to change that or how to change direction.”
You would do that with them. I could see that, especially as these tasks get longer. We actually want the user to come engage.
Way more, to create a good output. I guess Devin had to do this because some of these jobs take hours.
Right. So, yeah. Yeah, totally. Magic. And it's perverse incentives where they charge by the hour, so they make more money the slower they are. Interesting. Have we thought about that before? I'm calling this out because everyone is like, “Oh my God, it takes hours. It does hours of work autonomously for me,” and they're like, “Okay, it's good.” But this is a honeymoon phase. At some point we're going to say, “Okay, but, you know, it's very slow.”
Anything else? Obviously, within Google, you have a lot of other initiatives. I'm sure you sit close to the NotebookLM team. Any learnings coming from shipping AI products in general?
They're really awesome people. They're really nice and friendly, just as people. I'm sure you met them and realized this with Riza and stuff. They've actually been really, really cool collaborators and people to bounce ideas off.
I think one thing I found really inspiring is that they just picked a problem. Hindsight is 20/20, but in advance they said, “Hey, we just want to build the perfect IDE for you to do work, be able to upload documents, ask questions about them, and just make that really, really good.” I think we were definitely inspired by their ability and their vision to say, “Let's pick a simple problem, really go after it, do it really, really well, be opinionated about how it should work, and just hope that users also resonate with that.” That's definitely something that we tried to learn from.
Separately, they've also been really good at extracting the most out of Gemini 1.5 Pro, and they were really friendly about sharing their ideas about how to do that.
I think you learn a bit when you're trying to do the last mile of these products, and the pitfalls of any given model and so on. So, yeah, we definitely have a healthy relationship and share notes, and we're doing the same for other products.
You'll never merge, right? It's just different teams?
They are different teams. They're in Labs as an organization, and the mission of that is to really explore different bets and explore what's possible.
Even though I think there's a paid plan for NotebookLM now.
Yeah. It's the same plan as us, actually.
It's more than just Labs. That's what I'm saying.
It's more than just Labs because, ideally, you want things to graduate and stick around. But hopefully one thing we've done is not created different SKUs, but just been like, “Hey, if you pay for AI Premium, it's cool.”
Yeah, whatever, you get everything. Good thing.
What about learning from others? Obviously, OpenAI has Deep Research, literally the same name. I'm sure there's a lot of contention. Is there anything you've learned from other people trying to build similar tools? Do you have opinions on what people are getting wrong and what they should do differently? From the outside, a lot of these products look the same: ask for research, get back research. But obviously, when you're building them, you understand the nuances a lot more.
When we built Deep Research, there were a few different bets that we took around how it should work. What's nice is that some of those are actually things where we feel like we took the right approach.
We felt like agents should be transparent about telling you up front, especially if they're going to take some time, what they're going to do. That's really where the research plan we showed in a card came from.
We really wanted to be very publisher-forward in this product. While it was browsing, we wanted to show you all the websites it was reading in real time and make it super easy for you to double-click into those while it was browsing. The third thing is putting it into a side-by-side artifact so that, ideally, it was easy for you to read and ask questions at the same time.
What's nice is that, as other products come around, you see some of these ideas also appearing in other iterations of this product. I definitely see this as a space where everyone in the industry is learning from each other. Good ideas get reproduced and built upon. We'll definitely keep iterating and following our users to see how we can make the future better.
And on the model side, OpenAI has the o3 model, which isn't available through the API—the full one. Have you tried it already with the Gemini 2 model? Is it a big jump, or is a lot of the work in the post-training?
I would say, stay tuned. It currently is running on 1.5. The new-generation models, especially these thinking models, unlock a few things. One is obviously better capability in analytical thinking, like math, coding, and these types of things. But there's also this notion that, as they produce thoughts and think before taking actions, they inherently have the ability to critique the partial steps that they take and so on.
So, yeah, we're definitely exploring multiple different options to provide better value for our users as we iterate.
Yeah. I feel like there's a little bit of a conflation of inference-time compute here. One, you can do inference-time compute within the model—the thinking model. Two, you can do inference-time compute by searching and doing more iterations.
I wonder if that gets in the way. Presumably, you've tested thinking plus Deep Research. Does the thinking actually do a little bit of verification and maybe save you some time, or does it try to draw too much from its internal knowledge and therefore search less? Do they step on each other?
Yeah, no, I think that's a really nice callout. This also goes back to the use case. There are certain things that I can tell you from model memory—for example, last year the Fed did a certain number of rate cuts and so on. But unless I source it, it's going to be hallucinated.
Even if I got it right, as a user I'd be very wary of that number unless I'm able to source the .gov website for it. That's another challenge: there are things that you might not optimally spend time verifying, even though the model is saying, “This is a very common fact. The model already knows it, and it's able to reason over it.”
Balancing that out—trying to leverage the model's memory while also grounding it in some kind of source—is the challenging part. I think, as you rightly called out, with the thinking models this is even more pronounced because the models know more and are able to draw more second-order insights just by reasoning over things.
Technically, they don't know more. They just use their internal knowledge more, right?
Yes, but also, for example, with things like math, they've been post-trained to do better math. I think they probably do a way better job in math than the previous ones, in that sense.
Yeah. Obviously, reasoning is a topic of huge interest, and people want to know what the engineering best practices are. We think we know how to prompt them better, but engineering with them is also very, very unknown. Again, you guys are going to be the first to figure it out.
Yeah, definitely interesting times. And there's no pressure, Mukund. If you have tips, let us know.
While we're on the technical elements and technical bets, I'm interested in other parts of the Deep Research tech stack that might be worth calling out. What hard problems did you solve, more generally?
8. Agent Architecture Remains Hard
I think the iterative-planning one—to do it in a generalizable way—was the thing I was most wary about. You don't want to go down the route of teaching the model how to plan iteratively per domain or per type of problem.
Going back to the ontology, if you had to teach the model, for every single type of ontology, how to come up with these traces of planning, that would have been nightmarish. Trying to do that in a super data-efficient way by leveraging a lot of the model's memory was important.
There's also this very tricky balance when you work on the product side of any of these models: knowing how to post-train it just enough without losing things that it knows from pre-training. Basically, not overfitting in the most trivial sense, I guess. The techniques there, the data augmentations there, and multiple experiments to tune this trade-off—that's one of the challenges.
On the orchestration side, this is basically you're spinning up a job. I'm an orchestration nerd. How do you do that? Is it an internal tool?
Yeah, so we built this asynchronous platform for Deep Research, which is basically—most of our interactions before this were synchronous in nature. All chat things are synchronous, right? Now you can leave the chat and come back.
Exactly. And close your computer.
And now it's on Android and rolling out on iOS.
You know, I saw that you said that. I told you we switch roles sometimes.
Okay, you're reminding him, right?
Yeah, we ramped on all Android phones, and iOS is this week. What’s neat, though, is that you can close your computer, get a notification on your phone, and so on.
So it’s some kind of asynchronous engine that you made?
Yes, yes. The other part is this notion of asynchronicity and the user being able to leave. If you build 5- or 6-minute jobs, they’re bound to have failures, and you don’t want to lose your progress and so on. It’s this notion of keeping state, knowing what to retry, and trying to keep the journey going.
Is there a public name for this, or is it just some internal thing?
No, I don’t think there’s a public name for this.
We can name it now. This is our opportunity.
Yeah, we can name it now. The classic names that I used to work with in this area—which is why I’m asking—are workflows. There’s Durable Functions, like back when you were at Meta before, I think. Apache Airflow and Temporal were both at Amazon, by the way. AWS Step Functions would be one of those, where you define a graph of execution.
Step Functions are more static and would not be as able to accommodate Deep Research-style backends.
What’s neat, though, is that we built this to be quite flexible. You can imagine that once you start doing hour- or multi-day jobs, you have to model what the agent wants to do.
Yeah, you have to model what the agent wants to do.
Exactly. In short, it’s stable for hundreds of LLM calls. It’s boring, but this is the thing that makes it run autonomously.
Right. Yeah. Anyway, I’m excited about it. Just to close out the OpenAI thing, I would say OpenAI easily beat you on marketing, and I think it’s because you don’t launch on benchmarks. Should you care about benchmarks? Should you care about Humanity’s Last Exam or MMLU, or whatever?
I think benchmarks are great. The thing we wanted to avoid is the day Kobe Bryant entered the league, who was the president’s nephew, and weird benchmark friends. These are just weird things that nobody talks that way. Why would we oversolve for some sort of benchmark that doesn’t necessarily represent the product experience we want to build?
Nevertheless, benchmarks are great for the industry. They rally a community and help us understand where we’re at.
No, I think you kind of hit the point. For us, our primary goal is solving the Deep Research user value for the use case. The benchmarks, at least the ones that we’re seeing, don’t directly translate to the product. There are definitely some technical challenges that you can benchmark against, but if I do great on HLE, that doesn’t really mean I’m a great deep researcher.
We want to avoid going into that rabbit hole a bit, but we also feel that benchmarks are great, especially in the whole generative AI space, with models coming every other day and everybody claiming to be SOTA. It’s tricky.
The other big challenge with benchmarks, especially when it comes to models these days, is output-space entropy. Everything is text in and text out, so the notion of verifying whether you got the right answer is difficult. Different labs do it in different ways, but we all compare numbers. There’s a lot of art and figuring out how you verify this or how you run it on a level playing field.
I think there’s definitely value in doing benchmarks. At the same time, from a selfish PM perspective, benchmarks are a really great way to motivate researchers.
Yeah, make the number go up.
Exactly. Or just prove you’re the best. It’s a really good way of rallying the researchers within your company. I used to work on the MLPerf benchmarks, and you’d put a bunch of engineers in a room and, in a few days, they’d make amazing performance improvements on our TPU stack and things like that. Having a competitive nature and pressure really motivates people.
There’s one benchmark that is impossible to benchmark, but I just want to leave you with it: Deep Research. Most people are chasing this idea of discovering new ideas. Deep Research right now will summarize the web in a way that’s much more readable, but what will it take to discover new things from the things that you searched?
First, I think the thinking-style models definitely help here because they’re significantly better at reasoning natively and being able to draw these second-order insights. That’s the premise: if you can’t do that, you can’t think of doing what you mentioned.
The other thing is that it also depends on the domain. Sometimes you can prompt a model for new hypotheses, but depending on the domain, you might not be able to verify that hypothesis. In coding and math, there are reasonably good tools that the model already knows how to interact with, so you can run a test, verify the hypothesis, and so on.
Even if you think about it from a purely agent perspective, you could say, “Hey, I have this hypothesis in this area. Go figure it out and come back to me.” But let’s say you’re a chemist. What are you going to do there? We don’t have synthetic environments yet where the model is able to verify these hypotheses by playing in a playground and having a very accurate verifier or reward signal.
Computer use is another one. In both the open-source systems and elsewhere, there are nice playgrounds coming up. If you’re talking about truly being able to come up with new ideas, my personal opinion is that the model doesn’t just have to do the second-order thinking we’re seeing now with these new models. It also has to be able to play and test that out in an environment where you can verify it and give it feedback so that it can continue iterating.
Yeah, so basically code sandboxes for now.
Yeah, in those kinds of cases, it’s a little bit easier to envision this end to end, but not for all domains.
Physics engines. If you think about agents more broadly, there are a lot of things that go into them. What do you think are the most valuable pieces that people should be spending time on? Things that come to mind that I’m seeing a lot of early-stage companies do are memory—we already touched on emails and tool calls—and the auth piece. Should this agent be able to access this? If yes, how do you verify that? What are things that you want more people to work on that would be helpful to you?
I can take a stab at this from the lens of Deep Research. Some of the things we’re really interested in as we push this agent are, first, personalization, which is similar to memory. If I’m giving you a research report, the way I would give it to you if you’re a 15-year-old in high school should be totally different from the way I give it to you if you’re a PhD or postdoc.
You can prompt it, right?
You can prompt it, right? But the second thing is that it should ideally know where you’re at and everything you know up to that point. It should further customize the report and have an understanding of where you are in your learning journey.
Modality will also be really interesting. Right now, we’re text in, text out. We should go multimodal in, but also multimodal out. I would love it if my reports were not just text, but charts, maps, and images. Make it super interactive and multimodal, and optimize for the type of consumption.
The way in which I might put together an academic paper should be totally different from the way I’m trying to do a learning program for a kid, just in the way it’s structured. Ideally, you want to do things with generative UI and things like that to really customize reports. Those are definitely things I’m personally interested in when it comes to a research agent.
The other part that’s super important is that we will reach the limits of the open web. A lot of the things people care about are in their own documents, their own corpora, or within subscriptions that they personally really care about, especially as you go more niche into specific industries. Ideally, you want ways for people to complement their Deep Research experience with that content in order to further customize their answers.
There are 2 answers to this. One is that, in terms of our approach—or for me, rather—trying to figure out the core mission for building an agent, I feel like it’s still early days for us to try to platformize it or build these 5 horizontal pieces that you can plug and play to build your own agent.
My personal opinion is that we’re not there yet. In order to build a super-engaging agent, if I were to start thinking of a new idea, I would start from the idea and try to do that one thing really well.
Yes, at some point, there will be a time when these common pieces can be pulled out and platformized. There’s a lot of work across companies and in the open-source community to provide these tools to build agents very easily. I think those are super useful for starting to build agents, but at some point, once those tools enable you to build the basic layers, I would try to focus on really curating one experience before going too broad.
Yeah, we have Bret Taylor from Cieran. He’s said that Sierra mostly built everything in-house, which is very sad for VCs—the next great framework and tooling and all that. But the space is moving so fast. The problem I described might be obsolete 6 months from now, and I don’t know; we’ll fix it with one more LLMOps platform.
Okay, so just a final point on plugging your talk. People will be hearing this before your talk. What are you going to talk about? What are you looking forward to in New York?
I would love to actually learn from you guys. What would you like us to talk about? Now that we’ve had this conversation with you, what do you think people would find most interesting?
I think a little bit of implementation and a little bit of vision—kind of 50/50—and I think both of you can sort of fill those roles very well. Everyone looks at you as a very polished Google product, and I think Google always does polish very well. But everyone will want Deep Research for their industry.
Now, he’s invested in Deep Research for finance, and they focus on their thing. There will be Deep Researchers for everything, right? You have created a category here that OpenAI has cloned.
So, let’s talk about the hard problems in this brand of agent, which is probably the first real product-market-fit agent. I would say more so than the computer-use ones. This is the one where people are like, “Yeah, it easily pays for $200 worth of stuff a month, probably $2,000 once you get it really good.”
So, let’s talk about how to do this right from the people who did it, and then where this is going.
Yeah, it’s very simple. Happy to talk about that.
For me as well, I’m always curious to see you interact with the other speakers because then there will be other sorts of agent problems. I’m very interested in personalization and very interested in memory. I think those are related problems: planning, orchestration, all those things.
Auth and security—something that we haven’t talked about. A lot of the web is behind auth walls. How do I delegate my credentials to you so that you can go and search the things that I have access to? I don’t think it’s that hard. It’s just that people have to get their protocols together, and that’s what conferences like that are hopefully meant to achieve.
Yeah, no, I’m super excited. For us, we often live and breathe within Google, and we’re just a really big place, but it’s really nice to take a step back and meet people who are approaching this problem at other companies or in totally different industries. Inevitably, at least where we work, we’re in a very consumer-focused space.
I see. Right. I’m more B2B.
It’s also really great to understand what’s going on within the B2B space and within different verticals.
Yeah, the first thing they want to do is Deep Research for my own docs, right? My company docs.
Yeah, so obviously you’re going to get asked for that.
Yeah, there’ll be more to discuss. I’m really looking forward to your talk, and thanks for joining us.
Yeah, cool. Thanks for having us.
Thanks so much, guys.