[BidClub_]
No Priors · · 19 分钟

No Priors 第138期|2025年最佳(截至目前):Sarah Guo 与 Elad Gil

Sarah GuoElad Gil

YouTube
TL;DR
  • 最强的创业信号来自对工作流的盲测验证,而不是模型的炫技。 Winston Weinberg 和 Gabe 用早期思维链提示测试 GPT-3,向 100 个 r/legaladvice 房东-租客问题提问;3名律师在完全不知道 AI 的情况下判断,其中 86 个回答可以不经修改直接发出。OpenAI 的总法律顾问甚至回复:“我完全不知道这些模型在法律方面这么强。”
  • 当相邻基础设施移除了原有约束,AI 就能让“坟场”市场重新活跃起来。 Arvind Jain 认为,企业搜索失败是因为 SaaS 之前的数据无法访问;标准化、可互操作的 SaaS 系统和 API 让开箱即用的产品成为可能,而企业内部内容也恰好开始爆炸式增长。某个 Glean 客户拥有超过 10亿份文档——这相当于 Jain 回忆中 2004 年 Google 所面对的整个互联网规模。
  • 当推理模型能够识别不确定性并把工作交给工具时,效率会得到实质性提升。 在视觉任务中,模型可以承认自己看不清某个东西,再裁剪或处理图像,从而获得“非常明显不同”的测试时扩展斜率。对于估值,一位研究人员给出的例子更简单:让模型写代码运行计算,“弄清楚实际答案是什么”。
  • 数字劳动力的替代可能比实体自动化或政治适应来得更快。 Brendan Foody 预计,客服、招聘等岗位会迅速且痛苦地被替代;如果超智能带来的收益遵循幂律增长,社会可能出现“一场大型民粹运动”,同时面临艰难的财富分配问题。他预计实体自动化会更慢,潜在工作范围则可能从机器人数据生成,到餐饮和治疗。
  • 超智能可能把共同脆弱性从威慑因素变成先发制人的触发器。 Dan Hendris 将先进 AI 战略与核报复相提并论,但他认为,自动化 AI 研究可能变得极度不稳定,以至于中国或美国可能对对方的数据中心发动网络攻击。他设想俄罗斯重新评估力量平衡,通过间谍活动监控相关项目,并威胁进行报复。
  • 技术上的确信并不能消除模型的意外表现或脆弱性。 Isa Fulford 原本预计在浏览任务上训练模型会奏效,却发现第一个可用模型强得出人意料。Sarah Guo 将这种体验比作一条“铺满草莓的道路”,Fulford 表示认同。同一个模型可以做出异常聪明的事情,随后又犯下一个让人不禁问“你为什么要这么做?停下”的错误。
  • 创业与医疗结果,为纪律化建设提供了不同维度的证据。 Flagship 的演讲者质疑把创业当成随机的“射门次数”,尤其是在医疗、气候、农业和粮食安全领域投入来之不易的资本时。在 Abridge,产品批评是“氧气”;一名医生说自己不会退休,以及一个家庭表示这款工具让母亲能够回家吃晚饭,这些反馈则体现了 Rao 所说的更深层目的。
摘要 · 为研究而整理的核心内容

1. 工作流证据在市场察觉前揭示了能力

  • Harvey 的起点,是 Winston Weinberg 惊讶于“没人谈论 GPT-3”。他和 Gabe 用早期思维链提示,测试 100 个 r/legaladvice 上的房东-租客问题;他们没有告诉 3名律师任何 AI 相关信息,只询问每个回答是否合乎伦理、能否不经修改直接使用。律师们对其中 86 个回答给出了肯定答案。
  • Dr. Fay Lee 将空间智能描述为进化过程中的艰难转换:把收集到的光线转化为用于导航、操作和互动的内部3D世界。即便是人类,闭上眼睛重建周遭环境仍然困难;如果能让丰富的3D内容在指尖实现流畅创作和编辑,就会创造一个“完全不同的世界”。

2. 基础设施变化可以把坟场变成市场

  • Arvind Jain 对企业搜索“坟场史”的诊断,核心在于基础设施:SaaS 出现之前,供应商无法可靠地在不同服务器和存储系统之间定位、连接数据。统一的软件版本、互操作性和 API,最终让统一的开箱即用搜索成为可能。
  • Rubrik 暴露了这种需求:信息散落在 300 个 SaaS 系统中,员工抱怨什么都找不到,但 Jain 发现“根本没有可买的东西”。规模进一步放大了机会:一家大型 Glean 客户拥有超过 10亿份文档,Jain 将其比作 2004 年整个互联网的规模。

3. 工具让推理的扩展曲线变陡

  • OpenAI 的研究人员描述了视觉模型如何识别自身的不确定性——“我确实看不清这个东西”——然后使用工具裁剪或处理图像。这会提高推理 token 的产出效率,并带来明显更陡峭的测试时扩展斜率。
  • 一位研究人员用估值举例,揭示了分工所在:模型可以在上下文中反复拟合系数并自行验证,也可以写一个简单程序运行估值模型,直接确定实际答案。当模型“比较优势”之外的工作交给专门工具时,算力的使用效率就会提升。
  • Fulford 原本预计针对浏览任务训练 Deep Research 会奏效,但实际效果之好仍然让她意外。Sarah Guo 将这种体验比作一条“铺满草莓的道路”,Fulford 表示认同。不过,模型表现仍然参差不齐:它们可能先做出“一些极其聪明的事情”,随后犯下一个无法解释的错误,让人问:“你为什么要这么做?停下。”

4. 数字加速带来劳动力与安全风险

  • Foody 预计,客服、招聘等数字岗位的替代会“非常快”且充满痛苦,并引发重大的政治问题。接近超智能时,更棘手的问题将变成如何重新分配财富——如果收益按幂律集中,答案尤其困难。
  • 主持人追问:“实体世界意味着什么?”这迫使讨论具体化:机器人数据工作、端盘子,以及人们希望与真人互动的治疗工作。Foody 认为,实体自动化会更慢,因为它缺少虚拟世界中自我强化的改进闭环。
  • Hendris 首先用核武器作类比:国家可以通过共同脆弱性遏制彼此发动首次打击。随后他认为,一旦系统能够自动化大部分 AI 研究,快速自举至超智能或制造出“超级武器”可能变得极度不稳定,以至于中国或美国可能对对方的数据中心发动先发制人的网络攻击。他设想俄罗斯重新评估力量平衡,通过间谍活动监控相关项目,并威胁进行报复。

5. 纪律化建设通过结果赢得正当性

  • Flagship 的演讲者回忆说,自己作为一名 24 岁移民于 1987 年创办公司。当时,获得风险资本信任的通常是 Merck 或 IBM 的前高级管理者,每轮融资规模约为 2300万美元。这段经历让他思考:创业为什么不能成为一种职业,而要一直停留在随机、即兴、情绪化、带有赌性(gamy)的状态。
  • 当被追问“带有赌性”是什么意思时,他指向了一种文化:反复失败、偶尔成功,最后变成一块记分牌。当资本用于解决医疗、气候、农业和粮食安全领域“几乎不可能”的问题时,“不能把它看成……一次次射门”。
  • Abridge 将正面反馈汇入一个名为“love stories”的频道,公司内部任何人都可以访问。Shiv Rao 称尖锐的产品批评是“氧气”;正面反馈则包括一名医生说自己不会退休,以及一位乡村医生的家人表示,Abridge 让她能够提前回家吃晚饭。Rao 将超高速增长带来的短暂“多巴胺刺激”,与使命感带来的“催产素刺激”相对比,并表示公司真正追求的是后者。
Sarah Guo

2025 has been another remarkable year in AI. This week on No Priors, we're sharing our favorite moments from the podcast from the year so far. We've talked to visionary leaders at Harvey, OpenAI, Glean, A Bridge, and more. We also talked to legends of science like Dr. Fay Feay Lee and Nubar Fayen. But first, let's start with a moment that captures the magic of leaning into new capabilities at the right time. Harvey CEO Winston Weinberg discovered an extraordinary opportunity hidden in plain sight.

Speaker 1

Gabe and I had met a couple of years before, and I definitely didn't know anything about the startup world and didn't have a plan to do a startup. What had happened was he showed me GPT-3, which at the time was public, and I was incredibly surprised that no one was talking about GPT-3 and no one was using it in any way, shape, or form.

He showed me that, and I showed him my legal workflows. The aha moment was when we went on r/legaladvice, which is basically a subreddit where people ask a bunch of legal questions, and almost every single answer is, “Who do I sue?” Almost every single time. We took about 100 landlord-tenant questions and came up with some chain-of-thought prompts. This was before anyone was talking about chain of thought or anything like that.

We applied it to those landlord-tenant questions and gave it to 3 landlord-tenant attorneys. We said nothing about AI. We just said, “Here's a question that a potential client asked, and here is an answer. Would you send this answer without any edits to that client? Would you be fine with that? Is it ethical? Is it a good enough answer to send?” Eighty-six out of 100 said yes.

We cold-emailed the general counsel of OpenAI and sent him these results. His response was basically, “Oh, I had no idea the models were this good at legal.” We met with the C-suite of OpenAI a couple of weeks after that.

Sarah Guo

Now, from legal reasoning to spatial intelligence, the legendary Dr. Fay Lee opened our eyes to an entirely different dimension of AI capability.

Speaker 2

I think, from a neural and cognitive science point of view, that spatial intelligence is a really hard problem that evolution has to solve for animals. What's really interesting is that I think animals have solved it to an extent, but haven't fully solved it. It's one of the hardest problems, because what is the problem an animal has to solve?

Animals have to evolve the capability of collecting light in something we call eyes, mostly. Then, with that collection of light, they have to reconstruct a 3D world in their mind somehow so that they can navigate, do things, and, of course, interact.

For humans, we're the most capable animal in terms of manipulation, and we can do a lot of things. All this is spatial intelligence. To me, that's just rooted in our intelligence. What's interesting is that it's not a fully solved problem even in animals.

For example, for humans, if I ask you to close your eyes right now and draw out or build a 3D model of the environment around you, it's not that easy. We don't have that much capability to generate an extremely complicated 3D model until we get trained. There are some of us, whether they're architects or designers or just people with a lot of training and a lot of talent, who can do that.

That's a hard thing to do. Imagine you do it at your fingertips much more easily and allow much more fluid interactivity and editability. That would just be a whole different world for people. No pun intended.

Sarah Guo

Data is the beast feeding the AI train. And thus, Merkore CEO Brendan Foody is working with major AI labs on how to build what's next. He gives a clear prediction about what's coming for the workforce.

Speaker 3

I think displacement in a lot of roles is going to happen very quickly, and it's going to be very painful and a large political problem. I think we're going to have a big populist movement around this and all the displacement that's going to happen.

But one of the most important problems in the economy is figuring out how to respond to that. How do we figure out what everyone who's working in customer support or recruiting should be doing in a few years? How do we reallocate wealth once we approach superintelligence, especially if the value and gains of that are more of a power-law distribution?

I spend a lot of time thinking about how that's going to play out, and I think it's really at the heart of—

Speaker 4

What do you think happens eventually? X% of people get displaced from white-collar work. What do you think they do?

Speaker 3

I think there's going to be a lot more in the physical world. I think there's also going to be a lot that's niche.

Speaker 4

What does the physical world mean?

Speaker 3

Well, it could be everything ranging from people who are creating robotics data to people who are waiters at restaurants, or who are just therapists, because people want human interaction—whatever that looks like.

I think automation in the physical world is going to happen a lot slower than what's happening in the digital world, just because of so many of the self-reinforcing gains and a lot of self-improvement that can happen in the virtual world but not the physical one.

Sarah Guo

Which brings us to one of the biggest questions of our time. How do we navigate the geopolitical implications of super intelligence? Dan Hendris, the director of the Center for AI Safety, has an answer.

Speaker 5

Let's think of what happened in nuclear strategy. Basically, a lot of states deterred each other from doing a first strike because they could then retaliate. So they had a shared vulnerability. They were saying, “We're not going to do this really aggressive action of trying to make a bid to wipe you out, because that will end up causing us to be damaged.”

We have a somewhat similar situation later on, when AI is more salient, when it is viewed as pivotal to the future of a nation, and when people are on the verge of making a superintelligence—when they can automate pretty much all AI research.

I think states would try to deter each other from trying to leverage that to develop it into something like a superweapon that would allow the other countries to be crushed, or to use those AIs to do some really rapid, automated AI research-and-development loop that could bootstrap from its current levels to something that's superintelligent, vastly more capable than any other system out there.

I think later on it becomes so destabilizing that China just says, “We're going to do something preemptive, like a cyberattack on your data center,” and the U.S. might do that to China.

Russia, coming out of Ukraine, will reassess the situation, get situationally aware, and think, “Oh, what's going on with the U.S. and China? Oh, my goodness, they're so far ahead on AI. AI is looking like a big deal.”

Let's say it's later in the year, when a big chunk of software engineering is starting to be impacted by AI: “Oh, wow, this is looking pretty relevant. Hey, if you try and use this to crush us, we will prevent that by doing a cyberattack on you, and we will keep tabs on your projects because it's pretty easy for them to do that espionage.”

Speaker 6

The motivation for Flagship stems from what I was doing before, which was that I started a company in 1987, when 24-year-old immigrants didn't start companies in this country. Instead, former Merck senior executives or IBM senior executives were the only ones who were entrusted with the massive amounts of venture capital—namely, $23 million per round—that used to go into venture capital.

This was very early days, and I had the opportunity to start a company right out of graduate school. I ended up raising quite a bit of venture money and eventually went down a path of entrepreneurship.

One of the things that interested me was why the entrepreneurial process was supposed to be random, improvisational, idiosyncratic, almost emotional, and gamey. All of those things I thought were a bit of a put-off when it comes to actually doing things in a serious, professional way.

I used to go around in the very early ’90s saying, “Why isn't entrepreneurship a profession?” And if it was going to be a profession, how could it be a profession?

Speaker 4

What do you mean by “gamey”?

Speaker 6

Because it's supposed to fail most of the time, and once in a while you win and then you celebrate the win. What I mean is, it's random.

Speaker 4

It's random.

Speaker 6

But not only random—there are winners and losers and keeping score. I don't know. It's maybe the wrong word, but people even call it gamification in the software space. There's a version of this.

I don't mind being playful, because if you're overly serious, sometimes you miss things. But it can't just all be play. We take hard-earned money. We deploy it to do things that are damn near impossible. Once in a while, we reduce them to practice so they become not only possible but valuable.

And yet people treat it like, “Oh, well, it didn't work. There are 20 different things we tried. One of them worked.” And that, I don't know—as an engineer by background, as a scientist—I just thought that what we do, especially in healthcare, especially in climate, especially in agriculture and food security, you can't think of this as shots on goal. You've got to say, “Hey, we can get better at this.”

Sarah Guo

Reasoning is the biggest paradigm shift in AI architecture since the transformer. Brandon McKenzie and Eric Mitchell from OpenAI explained a crucial insight about reasoning models.

Speaker 7

I can give very concrete cases for the visual reasoning side of things.

There are a lot of cases where the model can estimate its own uncertainty. You'll give it some kind of question about an image, and the model will very transparently tell you in a thought, like, “I don't know. I can't really see the thing you're talking about very well.” It almost knows that its vision is not very good.

But what's kind of magical is that when you give it access to a tool, it's like, “Okay, well, I've got to figure something out. Let's see if I can manipulate the image or crop around here,” or something like this. What that means is that it's a much more productive use of tokens as it's doing that. Your test-time scaling slope goes from something like this to something much steeper.

We've seen exactly that: the test-time scaling slopes without tool use and with tool use, specifically for visual reasoning, are very noticeably different.

Speaker 2

Yeah. I also say, for writing code, there are a lot of things that an LLM could try to figure out on its own but would require a lot of attempts and self-verification that you could write a very simple program to do in a verifiable and much faster way.

You could say, “Hey, do some research on this company and use this type of valuation model to tell me what the valuation should be.” You could have the model try to crank through that and fit those coefficients or whatever in its context, or you could literally just have it write the code to do it the right way and know what the actual answer is.

I think part of this is that you can allocate compute a lot more efficiently because you can defer things that the model doesn't have a comparative advantage in doing to a tool that is really well suited to doing that thing.

Sarah Guo

Sometimes the most profound moments in AI development aren't the grand theoretical breakthroughs. They're based on taste, data generation, and grinding work—the visceral experience of watching something you hoped would work actually come to life. Isa Fulford from OpenAI captures that moment perfectly. Here she's describing the training that went into Deep Research.

Speaker 3

It really was one of those things where we thought that training on browsing tasks would work. It felt like we had good conviction in it. But actually, the first time you train a model on a new data set using this algorithm and see it actually working and play with the model was pretty incredible, even though we thought it would work.

Honestly, just that it worked so well was pretty surprising.

Mm-hmm.

Speaker 3

Even though we thought it would, if that makes sense.

Sarah Guo

Yeah. It's the visceral experience of, like, “Oh, the path is paved with strawberries,” or whatever.

Speaker 3

Exactly. But then sometimes some of the things that it fails at are also surprising. Sometimes it will make a mistake where it will do such smart things, and then make a mistake where I'm just thinking, “Why are you doing that? Stop.”

So I think there's definitely a lot of room for improvement. We've been impressed with the model so far.

Sarah Guo

One of the biggest surprises of AI, and a core principle for us here at Conviction, is how it can make bad markets suddenly good ones. The right technology can meet the right moment in unexpected ways. Arvind Jain built Glean in what everyone said was a graveyard market: enterprise search.

Speaker 4

It was like a graveyard, with all these companies that tried to solve the problem and didn't. Part of it was just that search is a hard problem in an enterprise. Even getting access to all the data that you want to search was such a big problem in the pre-SaaS world. There was no way to go into those data centers, figure out where the servers were and where the storage systems were, and try to connect with the information in them. It was a big challenge.

SaaS actually solved that issue. Most search products and most of the companies started in the pre-SaaS world. They failed because you just couldn't build a turnkey product. But SaaS actually allowed you to build something.

My insight was that the enterprise world had changed. We had these SaaS systems now, and SaaS systems don't have versions—everybody, all customers, have the same version. They're open, they're interoperable, and you can actually hit them with APIs and get all the content.

I felt that the biggest problem was actually solved: I could easily go and bring all the enterprise information and data into one place and build this unified search system on top. That was a big unlock.

By the way, the origins of Glean go back to Rubrik. We had this problem: We grew fast, we had a lot of information across 300 different SaaS systems, and nobody could find anything in the company. People were complaining about it in our pulse surveys. I always ran those in my startups, and this was a complaint that came to me. I had to solve it.

I tried to buy a search product, and I realized there was nothing to buy. That's really the origin of how Glean got started as a company. SaaS made it easy to connect your enterprise data and knowledge to a search system. That made it possible for us, for the very first time, to build a turnkey product.

But there are a lot of other advances as well. Businesses have so much information and data. One interesting fact is that one of our largest customers has more than 1 billion documents inside their company.

When Larry and I were working on search at Google in 2004, the entire internet actually had 1 billion documents. There's been a massive explosion of content inside businesses. You have to build scalable systems, and you couldn't build a system like that before, in the pre-cloud era.

Sarah Guo

Perhaps no story captures the human impact of this AI moment and its potential better than what's happening in healthcare. Here's Shiv Rao, CEO and founder of Abridge.

Speaker 5

It's pretty heroic, in general, for a doctor to give you feedback like, “Hey, this sucked, and you've got to do better. You didn't recognize the way I said this medication,” or, “I'm a gastroenterologist, and I would never sequence my problems in my assessment and plan section of my note this way. It doesn't serve me well and makes me look terrible as a doctor,” or whatever.

We get that feedback. We love it. It's oxygen. But then we also get feedback like, “Hey, this is amazing, and I'm not going to retire anymore. I've got years, decades left in my career now, thanks to this technology.”

In this love-stories channel, all of that feedback, that positive feedback, gets programmatically funneled. Any one of our people inside the company can always go into that channel, and its purpose. It's fulfillment, immediately. You immediately understand why we're all working so hard and why it makes sense.

Being on this very entrepreneurial journey these last couple of years is obviously new for so many of us. We're all kind of building new muscles, but it's a lot of pressure. This is my favorite bit of feedback.

This love story comes from a doctor at Tanner Health, which is a rural health system. She wrote to us:

“I was sitting at dinner last week, and my son asked me, ‘Mommy, why aren't you working right now?’ I literally took my phone out and explained to him that Abridge is a new tool that lets Mommy come home early and eat dinner with her family.”

I started to tear up and looked over at my husband, who then said, “Mommy's going to be able to eat dinner with us every night now.”

We get feedback like that every day. There are dopamine hits in hypergrowth, and those are awesome, but I think they get us through sprints. I think it's the oxytocin hits like this. It's the purpose. It's the fulfillment. That's, I think, what we're really after in this company.

Everybody's mission-driven out there, but I think this mission hits me at least a little bit differently.

Sarah Guo

These conversations remind us that we're living through a hinge moment in history. Stay tuned as we have more conversations with the builders and thinkers leading the way for the rest of the year. If you like what we're doing, leave us a review on Apple Podcasts or Spotify, comment on YouTube, or let us know who we should have as a guest. Thanks for listening. Find us on Twitter at no prior pod. Subscribe to our YouTube channel if you want to see our faces. Follow the show on Apple Podcast, Spotify, or wherever you listen. That way, you get a new episode every week. And sign up for emails or find transcripts for every episode at no-briers.com.