[BidClub_]
No Priors · · 28 分钟

No Priors 第109期|Sarah 与 Elad

Sarah GuoElad Gil

YouTube
TL;DR
  • 图像生成正迎来新一轮质量跃迁,但真正具有商业意义的突破在于可控性。 Elad 回顾了从2018年或2019年一件 GAN 艺术品登上 Sotheby’s 拍卖场,到 Midjourney 和早期 Stable Diffusion 经历“人人都有7根手指”的阶段,再到如今风格统一的动漫作品这一轮技术曲线。他认为,约1年后还会出现下一次跃迁,产品将从今天的“横向版本”走向能无缝处理平面设计的垂直产品。
  • 公开市场走弱对早期软件创业公司的经营影响应当很小,除非宏观路径变得具有生存意义。 尽管消费者信心降至多年低位、Nasdaq 下跌8%,Elad 仍称一切照常;即使在2008年金融危机期间,一家只有6人的创业公司也没有太多理由作出反应。Sarah 认为,高质量早期机会仍有充足资本,昂贵的基础模型项目获得的资金也深度超出预期,但 crossover 投资者和 IPO 前项目正面临更高的谨慎程度与流动性压力。
  • 关税是针对具体行业的政策工具,而不是一个一概而论的宏观结论。 Elad 认为,有些关税可能有用,有些是谈判工具,还有些只会造成破坏性的成本上升。Sarah 将中国汽车竞争视为一个案例:欧洲或许会保护自身工业基础;但她认为,真正有生产力的保护主义必须大举投资于本土技能、成本竞争力、防务零部件和汽车产能。
  • 基础模型能力正与产品入口汇合,战略问题因此转向分发和消费者剩余。 Sarah 指出 ArtificialAnalysis.ai 的图表显示模型能力正在收敛,并称 Google 最新的 Gemini 发布说明它仍在牌桌上。Elad 补充说,搜索、研究和推理也正在作为产品入口汇合;基准测试中的模型集群,以及 xAI 约9个月就达到大致 SOTA 模型,说明行业仍可能出现离群者。因此,机会不只是另一个通用 LLM,还包括物理、材料、机器人、科学、健康和生物学等被忽视的模型。
  • 专业模型要建立防御力,起点要么是专有数据生成,要么是真正不同的技术路线。 Sarah 的核心问题是“数据收集引擎”长什么样:化学、生物学和机器人可能需要开展新的物理实验,而现有大型实验室未必会做这些事;这不同于代码拥有丰富的数字语料和可测试的效用函数。Elad 认为,可能成立的技术切口包括状态空间模型、基于 Lean 的形式化推理,以及更好的软件智能体 RL 环境。
  • AI 市场感觉已经“可能进入第3局,而不是第1局”,标准化程度足以支持投资,但尚未形成持久均衡。 模型、基础设施、评估、编排和垂直应用层正变得清晰,MCP 则提供了模型与现有数据系统之间的开放接口。Elad 仍提醒说,“我学得越多,知道得越少”;这种“平静时刻”可能只会持续到下一次发布。
摘要 · 为研究而整理的核心内容

1. 图像生成的下一项突破是控制力,而不是新奇感

  • Elad 将 Ghibli 和动漫风潮放在一条反复出现的技术曲线上:他回忆起2018年或2019年一件 GAN 艺术品登上 Sotheby’s 拍卖场,随后 Midjourney 和早期 Stable Diffusion 即使处于“人人都有7根手指”的阶段,也足以让用户惊叹。如今的系统已经能以惊人的保真度生成风格统一的作品,这正是“最新版本”的公众认知:质量正在以多快的速度复利式提升。

  • Sarah 说,用户很善于感知当前的质量和可控性,但最新一轮浪潮也显示,图像、视频、文本和 logo 仍有巨大提升空间。需求非常原始而直接——“人们想要更多可爱,也想要更多美”。

  • Elad 认为,约1年后还会出现一次类似的时刻,随后是商业上真正无缝的平面设计产品:“我们现在做的是它的横向版本,很快就会有垂直版本。” Sarah 提到 Krea 风格的实时编辑和 HeyGen 的自然语言控制:用户只需几个词,甚至说出“whisper, ASMR”,工具就能直接响应意图。

2. 公开市场动荡几乎传导不到早期软件创业公司

  • Sarah 精确描述了压力情景:消费者信心降至多年低位,Nasdaq 下跌8%,关税则瞄准中国进口商品和汽车。

  • Elad 的回答是“并没有太大压力”。只要没有出现具有生存意义的冲击,软件创业公司只要业务在推进,仍然可以销售产品和融资;硬件受到的直接影响更大。在 Sequoia 2008年的“RIP Good Times”演讲期间,他曾问,一家只有6人的公司为什么要在意这些。一位合伙人的回答是:“你根本不该担心这些。”

  • Sarah 认为,高质量的早期机会仍然资金充足,并称昂贵的基础模型公司所能获得的资本市场深度超出预期。压力集中在经历多年流动性受限后的 crossover 和 IPO 前项目,但复苏中的并购市场,以及准备上市的公司,可能带来帮助。

  • Elad 的关税框架是“逐项分析”:有些保护性关税可能有用,有些可能服务于谈判,或者带来净成本。Sarah 以汽车竞争为例,认为面对竞争力不断增强的中国汽车,欧洲或许会保护自身工业基础。她补充说,保护主义需要配套的积极产业政策,因为重建美国在防务零部件或汽车领域的能力,需要对技能和成本竞争力进行大规模投资。

3. LLM 收敛将注意力转向被忽视的模型市场

  • Sarah 指出 ArtificialAnalysis.ai 的图表显示模型能力正在收敛,并称 Google 最近发布的 Gemini 证明它仍在竞争中。Elad 补充说,产品入口也在收敛:搜索、研究和推理正在变成标准能力,使分发和消费者剩余成为核心问题。

  • Elad 指向 ArtificialAnalysis 的基准测试:许多模型已经聚集在相距很近的区间,同时在编程或推理能力上各自出现尖峰。Grok/xAI 约9个月就达到大致 SOTA 模型,令他觉得“非常惊艳”,说明收敛并不会消灭突然出现的离群者。

  • 覆盖不足的机会在核心语言模型之外:物理、材料、机器人、科学、健康以及其他专业领域。生物学获得了不少关注——“每周都有一个新的生物学模型”——但 Elad 认为,资金和研究者兴趣往往与商业价值脱节,留下了可能规模巨大的未开发市场。

  • 对于“是否存在一枚戒指统治所有模型”的问题,Sarah 的答案是数据引擎。现有的语言和推理能力可以为专业系统提供种子,但机器人、化学和生物学可能需要收集或生成尚不存在的知识;对通用模型实验室而言,运营实体实验室,可能远比在 RL 环境中训练代码更加偏离其核心能力。

4. 专业模型需要结构性理由穿越通用模型的碾压

  • Elad 认为,可信的技术切口包括:在可压缩数据上更高效的状态空间模型;将数学和代码转译为 Lean,以实现形式化推理;以及训练模型,使其能够在软件和互联网环境中可靠行动。最后一类仍缺少具备一致泛化能力的智能体 RL 环境,因此它是真正的研究问题,而不只是又一层封装。

  • Elad 从速度、成本和推理保真度3个维度评估模型。一个缓慢、昂贵但能力极强的系统,可以分析一份100页的最高法院案情摘要;快速的专业模型则可以服务于狭窄任务或特定垂直领域。通用模型提供推理和语言能力,编排层则在不同模型之间分配任务——这实际上就是今天“智能体化”产品在编程、客户成功和其他领域所做的事情。

5. AI 可能已经进入第3局,但平静期或许只能持续1周

  • Elad 描述了一个良性循环:并购重新活跃,模型开发和测试时推理仍然昂贵,而解决数据、规模和延迟约束的公司可以改善整个生态。Sarah 认为现在是一个适合投资的舒适阶段,感觉“可能是第3局,而不是第1局”,并且已经出现了一些有用的标准化。

  • 顶尖人才分布在研究、基础设施效率、面向稀疏性或超大规模混合专家模型的软硬件协同设计、具备领域知识的产品工程、评估和 RL 环境等方向。智能体编排仍处于早期:收集上下文、制定计划、并行调用模型、验证并重试。

  • Elad 将其称为 AI 的“照常经营阶段”。模型栈正变得清晰;RAG 已从新鲜事物变成既有技术栈的一部分;评估实践正在成形;一些垂直领域赢家开始出现。消费者实验仍处于早期。ChatGPT、Perplexity 和 Midjourney 可能会被视为更早的消费级尝试,而新一代消费产品才刚刚开始出现。Elad 预计,今天的清晰格局会在1年内再次被打乱。

  • Elad 将 Anthropic 的 Model Context Protocol 描述为一种开放接口,把模型能力连接到文档、日志、商业工具、IDE 和其他系统;OpenAI 已表示将支持它。MCP 尚不完整,开发者仍必须清晰描述工具,但它可能大幅加速智能体开发。搜索和研究之外,下一批胜出的消费级智能体仍不明确,不过 Sarah 预计今年会出现案例。

Sarah Guo

Hey listeners, welcome back to No Priors. Today you've just got me and Elad again. It's a favorite type of episode. Elad Gil, how are you doing?

Elad Gil

I'm great.

Sarah Guo

I'm so excited. Everything is adorable—cartoons that are also slightly nostalgic and sensitive. Tell me about how you react to Studio Ghibli and also just better image generation.

Elad Gil

I'm a long-standing anime fan, so I think converting the world into everything anime or manga is a very positive step for humanity. I view this as something I've been waiting for for a while.

I feel like every year or 2, there's this moment in the image-generation world where people have a “Wow, that's amazing” moment again. The first version of that was—I think maybe even the GAN wave was the first wave. There was a GAN artwork in 2018 or 2019 that went to Sotheby's for auction, which was one of the first AI-generated artworks, back when people were doing these adversarial-network-based approaches to generating artwork.

It was a cludgy toolchain, but even then people were like, “Whoa, look at what AI can do right now.” It was super bad in comparison to what you can do today. Then there was the Midjourney and early Stable Diffusion wave, where those models came out and people were like, “Oh my gosh, this thing is amazing. Everybody has 7 fingers in the images, but oh my God, it's amazing. Look at all the things we can do with it. It's going to transform society,” and so on.

I feel like we've periodically had these moments, and I feel like this is the latest version of that. Part of it is that we're just on this amazing curve of quality and fidelity in artwork, and the ability to do so much more. Even back in the GAN world, there were style transfers—“Do this in the style of Van Gogh”—and so on.

The degree to which it does that so well and so cohesively, in so many styles and with so much aesthetic beauty and oversight, is really striking. I think we're just hitting another one of those moments where people are like, “Wow, this can really do it for forms of animation and other things.” All this is obviously in the context of ChatGPT and OpenAI, and the GPT-4o models have been incorporating a lot of this stuff directly in.

I think it's fantastic. We're going to see another thing like this in another year, I think, and then there will be the very commercial versions of this, which are already sort of happening. Look, we can use it for graphic design completely seamlessly, versus it kind of works; we can use it for all these different use cases. I feel like we're doing the horizontal version of it, and soon we'll have the vertical versions all come out. Obviously, companies like Recraft and others are working on the vertical versions directly, but I just view this as a super interesting evolution of the technology. I think it's super exciting. What do you think?

Sarah Guo

I think it is funny how much, at least in our little niche of the technology ecosystem—but anime and manga are pretty popular—the world reacts to it. They want more cute; they want more beauty. I think it's really exciting.

One of the interesting things this exposes is that users overall are very good at projecting where we are in terms of quality and controllability, and how much more room we have. Going from 8 bits of grayscale to images that might be perceived as photos of real people was a huge jump. To your point, people were shocked at some point, 2 generations ago, by image generation.

One of the things that Midjourney did was really have an aesthetic point of view and take a bunch of user feedback into account in terms of what was preferred. I actually feel like a lot of people thought of image generation—as end users, not researchers—as a little bit more of a solved problem. I think this is another data point for how much more we're going to get, never mind video and everything.

There’s also text and logos, and there’s a lot that's coming that people haven't done, including these truly integrated things where you can start clicking into images and modifying pieces. There are apps that are doing that, and there are things like Krea that do these real-time modifications as you're working on things. I do think there's so much room still. We're very early, but it's still so striking, so it's a very exciting area.

I think ease of controllability is also going to give people a lot more creative power. One of the things that HeyGen has demonstrated and is going to come out with in the product very soon is the ability to use natural language to describe emotion and voice. You can say, “Whisper, ASMR,” and just say, “I want the whole video with this person in this way,” with 3 words of text description. I think that kind of controllability is going to be really powerful.

You can incorporate it into augmented devices, and then I would just be working through an MR world. That's all I would live in. Is that the ideal? No. Maybe the [inaudible] part, but the rest not so much.

Are you freaked out about the macro?

Elad Gil

You mean the Nasdaq or what?

Sarah Guo

The markets? Yeah, the markets. Tariffs, inflation—which part of it? Consumer confidence is at a multiyear low. The Nasdaq is down 8%. There are tariffs on Chinese imports and on autos. I think investors and companies in the market are talking about how stressed they are about that.

Elad Gil

I'm not very stressed about it. I feel like there's a degree of uncertainty in the world right now, for sure, but from the perspective of people building technology companies, barring something truly existential happening, it's business as usual.

I've been through a few of these cycles now, where markets are way up and everybody's freaking out in a different direction, and markets are way down. The main place where it impacts the venture world or the startup world sometimes is if it soaks money out of the venture-capital ecosystem, and therefore valuations come down or there's less funding for the marginal startup, or things like that.

Other than that, these sorts of cycles tend to wash away unless you're a super late-stage company that's about to go public and there's some issue with your valuation in terms of expectations versus where you just want to go out, or something like that. For day-to-day technology startups, particularly ones that are not doing hardware, which would be impacted by the tariffs—for people who are just writing software—it should really be of minimal actual day-to-day impact.

Especially if your startup is working, you'll be able to get customers to pay you or find funding, or whatever it may be. I've been through a few of these, and every time it's been a bit of a shrug.

I actually remember going to the “RIP Good Times” presentation that Sequoia did in 2008. Back in 2008, there was a great financial crisis, and I was running a startup at the time. I was the CEO of this small company.

Sequoia did this big all-hands where they pulled together all their founders, and they had people come in and tell war stories from when the dot-com bubble collapsed: how it's time to batten down the hatches and do layoffs, the world will never be the same again, and everything's over. They were doing this as a service to the startup community. They were trying to help their founders figure this stuff out.

I remember talking to one of the Sequoia partners during it. I said, “We're like a 6-person startup. Who cares?” And he said, “Yeah, you're right. You shouldn't worry about this at all.” That was as all these financial institutions were collapsing around us.

This strikes me as very small in comparison to that. I think back then, that didn't have that much of a real impact on tech. Maybe Google did its first layoff ever, but other than that, tech just kept humming along.

If anything, the biggest tech companies in the world are now 10 or 20 times bigger than they were back then. I think this is an even more minor blip from a long-term tech perspective. Who cares? Again, barring some unexpected path that's going to come off of this. I don't know. What do you think?

Sarah Guo

It has almost no impact on me. I think especially at the early end of the market, the really high-quality opportunities have plenty of capital for them.

I keep discovering that the capital markets are much deeper than I thought for very expensive, for example, foundation-model plays. I still expect capital availability and a lot of inflow there.

I think it's probably a little different for investors who have more public-equities exposure. I bet pre-IPO crossover investors are getting more cautious. You have those sorts of much more long-term issues, with liquidity having been starved for several years now. But I think a return of M&A and several companies ready to go public will help that somewhat.

The place where tariffs kind of matter, which I think is interesting, is for very specific industries where, to some extent, it's useful for America or the West to protect themselves. Automotive would be a good example.

Some of the Chinese car companies seem to be getting so good that, if I were Europe, for example, and given that the industrial base is so automotive-dependent, I would probably be pushing for tariffs on Chinese imports of cars. The internal car industry may not be as competitive.

Elad Gil

And so, I do think there are some areas where the tariffs may be useful. There will be some areas where they're probably being used as a negotiation tool, and then some areas where they may be either net beneficial or net harmful in terms of actual costs passed on and things like that. But I think there may be a few areas where we should make sure that we actually have some in place. Then there may be some areas where it's going to be net negative or destructive, and some areas where it's just good for negotiating broader policy or relationships with certain external parties. People are using a catch-all for all of them versus looking item by item.

Sarah Guo

Yeah, I agree with that. And I think the productive version of tariffs is that there's a need for a broader industrial policy that is more supportive of the industries that we care about. That's going to be a big investment, right? If we want to make key components for defense or automotive in the United States, we are quite behind in many domains in terms of getting competitive from a skill and cost perspective. Some of those things are worth investing in on both the positive and the protection side.

I guess you mentioned that the depth of funding for models is part of all this. What do you think is happening in the foundation model world? You and I were just talking about these ArtificialAnalysis.ai charts showing convergence—kind of a monotonically more competitive market for capabilities—and amazing improvement over the last 18 to 24 months. But you just had the most recent Gemini release from Google. They're clearly still in the game. I don't know who was doubting that, given they have infrastructure and researchers—not just researchers, but very smart people at the helm. They're competing here as well.

Elad Gil

I think one of the more interesting things is that you have convergence not just on capability, but also in the product surface areas. Most people have search, they have a research product, and they have reasoning in the models. I think a lot of it is going to end up with consumer surplus and distribution being the question.

There's actually a really great website called ArtificialAnalysis.ai that shows different benchmarks they've run against these various models for reasoning or for different aspects of how you test a model for a knowledge base or for other forms of performance, speed of tokens per time unit, et cetera. I think that's really worth taking a look at. You see that for certain areas there is really strong convergence, and then there's almost a cluster of models that seem reasonably within the ballpark. Again, certain things spiked dramatically in one form or another around coding, reasoning, or other things. Then you have sort of a longer tail of other models.

At least for the core language model world, which those benchmarks are for, there definitely seems to be some form of convergence happening. Then there are outliers, right? Grok, or xAI, coming out of nowhere with a roughly SOTA model in 9 months was super impressive. There are also some of the things DeepMind and others have been doing.

What they don't really have benchmarks for are image models, and all of those obviously exist on a variety of sites and other places. But then there's a whole other suite of models that I think are discussed a lot less. Part of that is just the economic value, and part of it is what's in the market today. That's things like physics, materials, robotics, and certain types of science, as well as things that are more specialized in terms of post-training, like health-related data on top of some of these core models.

I do think that there are a lot of other types of models that people spend a lot less time on, some of which are becoming quite interesting. Probably the place that gets the most attention outside of the foundation model world—or the core LLM world, I should say, the language models—is probably biology. I feel like there's a new biology model every week.

But there are all these other fields and disciplines where I actually think there are some very big opportunities. Opportunities are obviously both societal in terms of impact, but also, in some cases, there are actually very big markets behind them. I think often the interest level of people working in the industry to build models is divorced from the economic value of these models, and sometimes that's rightfully so. There may be really interesting scientific applications that aren't very commercially applicable.

Sometimes it's really misaligned, where you're like, "Why are all these things getting funded when there are these wide-open spaces for certain types of models that just nobody's working on?" At least I've been looking a lot at what these alternative models are that are interesting from a market perspective and maybe are getting a little bit ignored right now.

Then I guess there's the other question: How many things get subsumed into these core LLMs versus being their own standalone thing? Do you think it's all one ring to rule them all, or do you think it's going to be a fragmented landscape? Where do you think that fragmentation happens? Is it somewhat too binary a distinction to say it's a model company versus not a model company?

Sarah Guo

Actually, even many of the companies that you and I, in the industry, consider to be model research companies are starting with some base of pretraining of existing knowledge and reasoning that is more and more readily available. In the case of robotics, you start with video pretraining. In the case of other domains, if you're going to start separately focusing on code—and we can talk about whether or not that's a good idea—you want both language and code in terms of being able to interact with the model.

I 100% believe that there are big opportunities in some of these domains, but one of the biggest distinctions to me is: What does the data collection engine for this look like? If you are thinking about physics, chemistry, biology, robotics, and maybe even some more near-term commercial applications, the data you would want—the understanding for the model to learn from—often doesn't exist yet. I think a theory of many of these companies that is interesting is: Our job is to go collect or generate it efficiently and use that to train the model.

In that case, I think the question of whether it needs to be in this single model to rule them all comes down to whether it's reasonable to expect one of the existing large labs to go do that data generation. If you have to set up a physical lab with robotics to do experimentation on new chemicals, that feels more far afield than code generation in RL environments, for example.

Anytime you go into the physical world, it's always harder to generate data. That's one of the reasons that language models, where you just effectively collect the wisdom of the internet digitally, are the first places where we've really seen this scale of breakthrough happen in recent times. Coding is a great example, where you not only have a lot of the data resident either online or digitally, but also you have very clear utility functions, or things that you can test against, in terms of code and its performance. Is it doing what you think it's going to do? Those are always going to be the easiest areas.

Elad Gil

This is an odd pet peeve of mine, but it always annoys me when people who do really well as founders in traditional software and tech start telling everybody else to go and do the hard stuff in biology, materials, and physics. You're like, "Well, you made all your money in [__] software. What are you talking about?"

I feel like there's been a long history of that. I remember interviews with Bill Gates from 20 years ago. He was like, "If I were to start today, I'd go into biology." I feel like sometimes there are the model versions of this.

Sarah Guo

You're so funny. I feel like you're the opposite. You're like, "I actually have a PhD in biology."

Elad Gil

Yeah, that's why I know. That's why I know reality. I think the other distinction I would draw is: Is it some orthogonal, totally different technical thesis? Do I think there's a research advance that is just very different, architecturally quite different? I'll describe categories of companies that could be relevant here.

We had Karan and Albert from Cartesia on the podcast. I think state-space models are an interesting direction that is highly efficient for certain types of data that are compressible. If you look, there are several plays on formalization and translating problems into Lean and taking that as a path to increasing reasoning capability for math and code.

There are a number of companies that are trying to train models that are better at taking actions in software and on the web. This is clearly also in line with the large foundation model labs, but I think they're at least trying to work on a question that doesn't feel fully answered in terms of consistent, generalizable RL environments for agents. There are spaces where I think there is a theory of why the company should exist if true, versus just being straight in line with the OpenAI, Anthropic, xAI steamroller, of course, and the Google steamroller.

Sarah Guo

What did I miss? What else do you draw as a distinction, or where do you think there is opportunity?

Elad Gil

To your point on state-space models, there may be advantages in terms of the speed and size of some of those models on a relative basis for very specialized tasks. Usually, I think of it as a 2-by-2 matrix where you have one axis that is speed, performance, and cost, because those are roughly the same thing for many of these models—it’s inference time, effectively. Then there’s reasoning fidelity, or whatever you want to call it.

Depending on where you are in these different quadrants, you have one quadrant where it’s slow, expensive, and not very smart. Obviously, nobody wants to use those models. There’s the one where it’s very slow and expensive, but it’s very smart and very capable. That’s where you’re like, “I’m going to upload a 100-page Supreme Court brief, and it’ll give me this amazing analysis I can use to argue a case,” or whatever. It’s high value, and it’ll take a while to process.

Then there’s the super-fast, highly performant category, which tends to be these very specialized niche models for specific applications. I think some of the state-space models, or SSMs, tend to work very well for very specific application areas. Then there’s the last quadrant.

Based on which of those quadrants you’re in, I think it really determines the type of things that you can build. The really fast, high-performance models tend to be more vertically focused or more focused on very specific types of tasks. The really slow, expensive ones that are actually very performant—you could imagine OpenAI’s versions—but it seems like the backbone for a lot of those is actually these very generalizable models, where a big chunk of what you’re getting is the reasoning and broader linguistic capabilities that you then apply to a domain.

Then, of course, there’s stuff that people build on top of them in terms of orchestration layers and specialized, bespoke things that route things to different models differentially relative to your use case. It seems like everything that’s quote-unquote agentic right now is basically doing that across customer success, code, and every domain that has a specialized approach. They always have this sort of orchestration layer built on top.

I think it’s super exciting to watch all this stuff, and I do think some of the applications in some of the less purely linguistic domains may be interesting in the short run. Going back to the question of whether the macro is stressing you out, there’s such a virtuous cycle in technology happening right now. This is actually quite dominated by the fact that M&A is alive again, so we’re going to have outcomes.

To your point, there’s exploding surface area of stuff that these models can attack. You have research progress and people making different technical bets. You mentioned DeepSeek. I think model development and the continued increase in the aggressive use of reasoning and test-time compute are quite expensive, and training continues to get more expensive. So I think the fact that there are now people trying to solve data, scale, and latency problems will help everybody, too.

Sarah Guo

Do you know if it’s true that the DeepSeek researchers are not allowed to leave China?

Elad Gil

I do not know if that’s true. I think any country should want to hang on to its best talent, but perhaps not restrict people’s movement.

Sarah Guo

I think we should be trying to attract great talent here. We should keep all the AI researchers in the Mission District and just not let them leave.

Elad Gil

Somewhere between the Mission and Dogpatch.

Sarah Guo

Yeah. We could actually just draw a line between our offices. They all have to go to Atlas Cafe every day.

Let’s talk through the talent categories. For anybody who isn’t thinking about their kids 10 years from now but is just thinking about the next 2 or 3 years, what type of expertise is valued, and where should you stay between my office and Elad’s office in the Mission and the Dogpatch?

Elad Gil

You have researchers, and you have infrastructure scaling and efficiency. We welcome all of you. Hardware-software co-design, right? Designing the next-generation TPU or whatever—there’s a special visa for you to move into that region.

Sarah Guo

Yes, we’re here to sponsor you. Visa program.

Elad Gil

Yes. If you’re ready to design chips to better handle sparsity or massive MoE models or something, I’ve got a visa campaign for you.

Sarah Guo

Kind of what you said, right? Anybody who has deep domain-user understanding combined with the basic product engineering—it’s not basic, but the product engineering sense—for this orchestration and applied-ML area, evals for agents, setting up RL environments, all of that is still a very nascent area.

Gather context, plan, make model calls, parallelize, verify, retry—this orchestration that I already described. We’ve got a visa program for you. We’re thinking about naming it. We’ll hire somebody to run it. It’ll be great. I’ll call it Gillingrow. We’re going to work on the marketing.

Elad Gil

I feel like we’re in the business-as-usual phase of AI. I think the stack is reasonably well defined, and obviously it’ll change and there’ll be new things in it, but the last couple of months have been very clarifying in terms of the consolidation of the things that are short-term crucial.

There’s the model layer and all the various accoutrements around agentic stuff and reasoning, et cetera. Obviously, that will only accelerate and get dramatically better, and it’s on its own scaling curve. Then, on the infrastructure layer, I think that’s solidified a bit, right?

I remember when RAG was a big deal, a new thing. All these things are, I feel like, kind of falling into place—evals and how you do them. I think things are solidifying there with companies like Braintrust and others.

Then, on the application-layer side, I think I’ve bought into the notion we’ve been discussing for a year or two now around AI really starting to impact different service-related industries, vertical applications, and different use cases. I’m starting to finally see some inkling of consumer stuff again. It’s nascent and early, but at least people are trying.

I feel like there were 2 or 3 years where nobody was really trying to do anything similar, although one could argue that Perplexity, ChatGPT, Midjourney, and all these consumer-y things were early consumer forays. Maybe ChatGPT is the world’s biggest new AI consumer product. Google was really the original one in some sense.

It feels like a period of brief consolidation. In a handful of verticals, I think we’re starting to see some of the winners emerge. It’s an interesting, clarifying time. Of course, the thing I say about AI is that the more I learn, the less I know. It’s the only industry where I feel like the more I learn about the market, the more confused I am.

There’s this brief moment of clarity, and then I’m guessing in a year all bets are off and all sorts of things will scramble again. But at least for now, it feels to me like a few things have fallen into place, at least temporarily.

Sarah Guo

This actually feels like a very comfortable time to invest for me because, to your point, it feels more like—I don’t know, maybe it’s inning 3 instead of inning 1—where there’s a little bit of stability in the ecosystem. There’s a real goodness around standardization, some standardization of integration with different things. MCP, I think, is going to accelerate a bunch of development for people.

I’m meeting companies where they’ve set up a data source that is useful to the enterprise in some way, that these models can interact with well, and they’re like, “Oh, MCP server.” Do you want to quickly explain to people Model Context Protocol, or MCP, what it is, and how it works?

Elad Gil

I’m going to fudge this, but I’ll try to describe it. This is an attempt by Anthropic. It came from Ben Mann’s group at Anthropic. It’s called Model Context Protocol, which is an attempt to spec out a standard interface for connecting model capabilities to systems where you already have useful data.

That could be documents, logging, business tools, the IDE, whatever. Sam Altman from OpenAI said they’re going to support it as well. I think this is not a complete solution. It has gotten a lot of popularity with developers over a very brief period of time, but it’s just how you expose your data to the model.

It’s an open standard, so it’s not proprietary. Anybody can use it. It’s a two-way connection between data sources and AI-powered tools.

Sarah Guo

And big companies have done it.

Elad Gil

Yeah, I think there’s still a bunch of work for developers to do in terms of describing their tools and how to use them very specifically and cleanly, but it does make it much easier, and I think it will accelerate agent development a lot.

Going back to this idea of what it means for the ecosystem, the fact that you’re accelerating the ways for models to interact with existing ecosystems means we expect agents to get better. You have a bunch of choices around model availability. As you said, there’s this clear pathway about how to automate certain types of work through the orchestration of these capabilities. I think that’s going to be super fertile.

I do think it’s very unclear what types of winning consumer experiences are possible here.

Sarah Guo

There aren't consumer agents that don't look just like search or research in the large-model products that are really working yet, at least that I've seen. But I expect to see them this year. I'm excited about it.

Elad Gil

Yeah, I think cool stuff is coming.

Sarah Guo

When everything destabilizes, Elad and I will be back on No Priors. We'll talk to you all then.

Elad Gil

It's going to get destabilized again, but I think it's a moment of calm, and calm is all relative, right? There's enormous innovation, huge changes coming, big technology waves, new things every week. But at least there's a little bit more of a view of who are going to be some of the main players in some of these areas and how all these things fit together. So I think we should enjoy the calm while it lasts—for the next week or whatever it is, the next few hours before the next thing drops.

Sarah Guo

All right, signing off, y'all. Good to see you. Find us on Twitter at No Priors Pod. Subscribe to our YouTube channel if you want to see our faces. Follow the show on Apple Podcasts, Spotify, or wherever you listen. That way you get a new episode every week. And sign up for emails or find transcripts for every episode at no-priors.com.

No Priors 第109期|Sarah 与 Elad — 文字稿与摘要 | BidClub