[BidClub_]
No Priors · · 25 分钟

No Priors 第116期|与 Sarah 和 Elad 对谈

Sarah GuoElad Gil

YouTube
TL;DR
  • Elad Gil 认为,部分 AI 市场正在走向整合,未来两年谁会胜出已初现轮廓——即便今天的领先者未必能赢到第5年。 基础 LLM、医疗记录、编程和客户成功已有较清晰的领跑者;销售效率、金融分析师工具和会计仍未定局,Sarah Guo 还补充了医药。Gil 怀疑,这些尚未定局的市场,要么缺少正确的产品路径,要么仍需要更好的模型。经历多年“我了解得越多,知道得越少”之后,他终于感到“不确定性终于有了喘息空间”。

  • 正在成形的应用打法,是把高价值的垂直工作流、专有数据、分发能力,以及能够基于数据创造、推导或扩展知识的用户结合起来。 Guo 将 Abridge 和 OpenEvidence 归入这一类别,但仍不确定在一个历来高度分散的市场中,销售将如何胜出。编程领域显示,模糊地带正在收窄:目前可行的切入口包括 Cursor、Codeium/Windsurf、Cognition 和 Microsoft Copilot。

  • AI 初创公司的整合,可能成为抵御老牌巨头的战略防线,而不是承认失败。 Gil 建议,一个品类中的前两家初创公司考虑合并,结束初创公司之间的战争,把竞争转向3到4家巨头。Guo 说,创始人、董事会和投资者往往把“考虑合并”本身视为投降,但她称之为“为了赢而投降”。

  • Gil 认为,生物科技有几个规模巨大的市场,商业需求显而易见,却在结构上长期被忽视。 例子包括:用重编程的成人细胞生成精子或卵子,让女性剩余的卵母细胞成熟,治疗皮肤衰老和脱发,恢复近视力或听力,以及通过 USAG-1 等基因实现牙齿再生。细胞衍生生殖最令人震惊的推论是,“任何成年人”理论上都可能与另一个成年人拥有孩子;而通过握手等接触从某人身上交换获得的部分细胞,也可能让涉及该人的生殖成为可能。

  • 生物科技的融资架构,筛选的是药企愿意收购的资产,而不是能长期服务非传统需求的新公司。 Gil 说,孵化器可能投入4,000万美元换取约40%的所有权,随后把项目引向癌症、心血管疾病或神经科学管线,因为这些是收购方想要的方向。监管俘获,以及科学家认为解决皱纹“地位低”,进一步让可用的科学成果滞留在创投漏斗之外。

  • Gil 将世界模型和强化学习视为超越文本预测的潜在路径,而非主要面向近期游戏产品。 智能体需要规划、调用工具、接受反馈、探索替代推理路径并自我评估,而行为克隆一旦离开示范路径就十分脆弱。挑战在于构建“一份宇宙的复制品”:环境既要足够丰富,能够教会系统适应性解决问题,又要足够廉价且多样,避免过拟合。Gil 不确定近期商业应用在哪里;Guo 从围棋和分子进化推演认为,不受约束的搜索可能产出反常但更优的代码、分子或策略。

摘要 · 为研究而整理的核心内容

1. AI 市场的临时赢家正在浮现

  • Gil 观点的转变,是本期节目的主线:AI 曾是少数几个“我了解得越多,知道得越少”的市场,但近几个月来,尽管研究仍在推进,几个品类已经变得清晰起来。

  • 在基础 LLM,以及医疗应用中的医疗记录、编程和客户成功领域,他现在看到了未来2到3年的潜在领跑者——但未必能赢到第5年。他在客户成功领域点名 Sierra 和 Decagon,在编程领域则提到 Cursor、Codeium、Cognition 和 Microsoft Copilot。

  • Guo 将这种理解称为市场的“临时物理学”:找到相关的垂直领域,打造用户真正想要的工作流,加入可检索的专有数据和分发能力,并让用户能够从中创造、推导或扩展知识。Abridge 和 OpenEvidence 符合这一形态;销售仍是未知数。

2. 编程正在预示产品与公司层面的整合

  • 编程领域过去大概有12种可能的胜出路径。如今切入口看起来已大幅收窄,但 Guo 指出,Codestral 等开放模型和更小的模型已经能够用代码完成真正的任务,可能成为专用工程工作流的推动者,只是这些工作流目前还达不到足够高的质量。

  • Guo 还提到 Microsoft 将 Copilot 开源,并将其视为阻止 Cursor 用自己的开源 VS Code 分支反过来抢走 Microsoft 生意的尝试;她说,实际影响仍有待观察。

  • 剩下的分歧在于,同步 IDE 工作与异步云端智能体。Guo 注意到,OpenAI 押注异步模式推出 Codex,同时通过收购 Windsurf 买下 IDE:“两种判断可以同时成立。”

  • Gil 预计产品会趋同,收购也会增加。他给两家领先初创公司的建议很直接:“停止初创公司之间的战争”,像 X.com 和 PayPal 那样合并起来对抗巨头,而不是重演 Uber 与 Lyft 之间多年的相互消耗。

  • 自尊、整合焦虑和私人公司估值阻碍了这类交易。Gil 建议选择一个指标——用户数或收入——按合并后的比例划分所有权,并接受“如果我们最终都能赢,正负几个百分点根本无关紧要”。

  • Gil 不认为每个市场都需要合并。有些市场足够大,可以容纳多家玩家;他以支付为例,Adyen、Stripe、PayPal 以及其他12家左右的处理商能够在一个分散市场中共存。

3. 生殖与长寿生物学藏着显而易见的市场

  • Gil 认为,第一个被忽视的机会是生育。他说,日本已有充分数据表明,细胞可以被重编程为精子或卵子,并且已经用2位父亲的遗传物质培育出存活的小鼠。按照他的设想,最终目标是让“任何成年人都能和任何另一个成年人拥有孩子”。

  • 更简单的版本,是让现有的卵母细胞成熟。女孩出生时有100万—200万个卵母细胞,到了青春期大约只剩30万,但目前没有成熟技术,能够让她们在一生中持续成熟并获取大量卵子。

  • 伦理风险与机会本身无法分开:通过握手或其他接触从某人身上交换获得的部分细胞,理论上可能用于涉及该人的生殖,例如 Elon Musk、LeBron James 或 Taylor Swift。“这类事情的一些后果相当疯狂。”

  • 衰老也提供了另一种需求信号:人们通过 Botox 注射细菌毒素来显年轻。Gil 说,这代表着一家价值400亿美元的公司,每年仅靠美容应用就能产生15亿美元收入。然而,他称没人真正研究皮肤衰老、脱发、白发、近视力减弱、听力损失或通过 USAG-1 等基因让牙齿再生的治疗;据他说,这些基因在某些动物模型中已经有效。

4. 生物科技的结构过滤掉非传统雄心

  • Gil 的历史对比十分鲜明:剔除 Moderna 因新冠疫情推动而崛起的案例后,他认为,上一家从零创建并达到500亿美元以上市值的生物科技公司,可能还是上世纪80年代的 Regeneron。没有年轻创始人驱动的挑战者,生物科技就像一个只有 IBM 和 HP 的科技行业。

  • 风投公司的组建方式进一步强化了这种停滞。一家生物科技基金可能投入4,000万美元孵化一个项目,拿走40%的股份,将其推进到足以吸引准公开市场资金的阶段,随后由交叉基金接力。这种所有权模式“真正的设计,就是把这些公司卖进药企怀里”。

  • 因此,初创公司会沿着药企的收购管线前进——癌症、心血管疾病和神经科学——而不是建设独立市场。Gil 还提到监管俘获,以及 FDA 要求企业为可能并不存在的终点提供证明;同时他承认,一些领域可能仍需要更多基础科学。

  • 第四个约束是地位。科学家可能认为高度商业化的工作不够纯粹。“你怎么敢去研究修复皱纹”体现了这种拘谨,它让有价值的生物学成果被闲置。

5. 世界模型可能教会系统走出人类路径

  • Gil 将世界模型放在大型 LLM 的进展背景下讨论:扩大模型规模和训练数据,已经带来了强大的知识与模式识别基础,但更广义的智能不止于按顺序生成文本。智能体必须能够规划、阅读文档并得出结论、调用工具、接受反馈、探索替代推理路径,并为了实现目标评估自己的工作。

  • 人类任务轨迹可以支持行为克隆——“猴子看,猴子做”——但这种方式仍然脆弱。一旦智能体按下示范者从未碰过的按钮,就会进入分布外领域,可能完全不知道下一步该做什么。

  • 强化学习提供了试错机制,但真实工作缺少国际象棋或围棋那样清晰的规则与奖励。有效环境必须近似宇宙的一部分,缩小与现实的差距,并产生足够多样的情境,让系统学会适应,而不是记住一条路线。Gil 说,他不知道如何抵达“黑客帝国”。

  • Gil 对许多近期商业叙事持怀疑态度,包括生成游戏、游戏资产或机器人训练数据,并表示自己对近期前景没有太多结论。但他仍将世界模型视为“通往更强 AGI 的概念路径”。

  • Guo 随后从围棋推演:在给定效用函数、几乎没有其他约束的情况下,AI 找到了人类此前没有构想过的棋步,随后人类研究并模仿了这些棋步。她想知道,编程和分子进化是否也能产生类似的意外解法——奇特的催化剂、结合蛋白或代码——供人类学习。Gil 表示认同,并设想模型去搜索那些人类从未被教会探索的问题空间。

Sarah Guo

Elad, what's going on?

Elad Gil

How are you doing, Sarah?

Sarah Guo

I'm good. I can't tell whether this is a very stable time in the market—whether it's crystallizing into known businesses and models or whether it's still fluid. What's your take?

Elad Gil

AI is the one market in my career where I've consistently said, "The more I learn, the less I know." Every other market, you learn more, you know more, and you keep advancing. I actually feel like that's shifted in the last couple of months, where, despite the rapid pace of innovation and all the really exciting new models, research findings, and everything else, I feel like a bunch of markets have consolidated. It's clear now who the likely players or winners are in 2 or 3 big areas.

That may change, right? In 3 years, another new startup may launch and displace everybody, or an incumbent may make a bold move, or whatever it may be. But I feel like in the foundation model market, at least for LLMs, there's a clearer view of what's important and what isn't. At the application level, I think it's clear who the winners are going to be in at least the first set of services for healthcare-related things, like medical scribing or other workflows.

In coding, it seems like it's consolidated into 2 or 3 players. Maybe that's Cursor, Codeium, Cognition, and then Microsoft's Copilot, right? There probably aren't 2 dozen companies that are all still competing there. In customer success, it seems like things are consolidating around Sierra and Decagon.

You go market by market and you're like, "Okay, there's a bunch of markets where it's clear who we think some of the winners may end up being," or at least who the important companies will be for the next 2 or 3 years. Then I think there's a set of markets where it's still wide open. You look at sales productivity tooling: there's going to be something really important there. There's going to be some financial analyst tool that's really important, and there's going to be an accounting company that's really important.

The question is, has that not consolidated yet because nobody is doing the exact right product approach? Is it because the models aren't good enough and the capabilities have to get better? It feels like there's a bunch of stuff that's still unknown, but it's way clearer than I think it was a year ago.

For the first time in 2 years or so, I feel like there's more clarity. When I first started investing in generative AI, you just went and backed the things where the people seemed really good and the market seemed interesting, because there wasn't a lot of competition. That's when I led the seed round for Perplexity or invested in Character.AI, Harvey, or some of these other things. That was pre-ChatGPT or pre-Midjourney.

Sarah Guo

Oh, the good old days.

Elad Gil

Yeah, the good old days, when nobody cared. When GPT-3 was out, everybody was like, "This is kind of crappy." But the scaling law was clear, right? I thought a handful of people—you being included—we collectively saw that this stuff was going to be important.

But then there was a period of uncertainty for 2 years or something like that, maybe 3 years, where there was so much innovation, so much change, and so much rapid growth. I think now, finally, we're hitting a period where at least a subset of things are consolidating back down. Again, these may not be the winners 5 years from now, but they definitely seem to be emerging as the winners for the next 2 years.

I think it's a nice breather in terms of uncertainty and having a bit more clarity into what's going to happen. I don't know. What do you think?

Sarah Guo

I feel a little bit like I understand some temporary physics of the market a little bit better. It's like a race to find the verticals of relevance and then get something to work in a way that users actually want. Maybe you have to go get proprietary data sources that you can retrieve against and get distribution, and then ideally have users who can create, derive, or extend knowledge from that, like the companies you just named.

I don't think you explicitly said it, but I put Abridge and OpenEvidence in that category. I think they fit into that shape. One thing you and I have talked about is that I'm actually quite unsure about sales. I don't know how to think about how something wins there. You could go at it from a data perspective or an adoption perspective, but it's been a very fragmented market to date.

I agree with you on finance and accounting. I'd add pharma to that. There are some industries that are really document-driven where you can see something coming there. There are companies in networking and pharma, for example, that are kind of interesting.

Elad Gil

And so, to some extent, it's been clear what markets will be interesting, or at least a subset of them. It just wasn't clear who would win and how. Coding is an interesting analog, where there were probably 4 different approaches to coding that everybody was taking simultaneously.

I think some of those approaches will consolidate over time, but the entry points now seem much clearer in terms of how you actually win in that market. Before, 2 years ago, there were a dozen different ways you could imagine somebody winning. I wonder if the analog there is the sales stuff you're talking about, where it seems a little bit less certain right now, but maybe in 2 years we'll be like, "Of course, it was whatever that workflow was."

Sarah Guo

We had a debate internally at my firm about what it would take for another new entry point to work. I think it would take a lot. I'm open-minded to it, but what is still changing is that you increasingly have open models and little models that can do real things with code. Codestral and this—I think you'll see more there.

Microsoft open-sourced Copilot. We'll see what the impact of that is, but it's like they finally decided they need to fight Cursor from eating its lunch with its own open-source VS Code fork. There's some chance that making specific workflows for engineering work that don't work at sufficient quality today can create enough distribution. That's interesting.

Then it's not clear: you have the synchronous IDE workflow and the asynchronous one, right? One question is how quickly the quality of these asynchronous code agents increases. OpenAI, with Codex, made a bet on an asynchronous, cloud-based software engineering agent, and then they bought the IDE with Windsurf, right?

Elad Gil

It's true, it's true. You can believe both. I think a lot of these things will just consolidate over time. My view is that the market is going to see 2 types of consolidation: product consolidation, and then there will be actual acquisitions. The Codeium/Windsurf acquisition by OpenAI is the first step in that.

If I were number 1 or number 2 in a market and I was a startup, I'd consider merging with the other party if there were 2 main startup players, because the real threat will be fighting the incumbents. I would get ahead of it and say, "Okay, let's stop the startup-to-startup war and just focus on winning against the 3 or 4 incumbents that we have to go up against."

You could just keep fighting and getting distracted by the other party, which is kind of what Uber and Lyft did for a while. There are other precedents. The ones that did merge include PayPal, right? There was X.com, which Musk was running, and then PayPal, which Peter Thiel was running. They decided to merge because they were like, "Why are we competing with each other when there's so much competition?"

I think both paths will happen, but it may be something people should consider as well.

Sarah Guo

What do you think prevents companies from thinking through that or doing that?

Elad Gil

Well, it's 2 things. One is ego. Who's going to run it? They want to subsume me? Sure, I'm number 2, but blah, blah, blah. I'll still beat them. Or what role would I play? Put that aside and just go win. Who cares?

Second, people worry too much about integration. What's the culture, and what's this, and what's that? Often, it's just: merge it, and if it doesn't work, shut down parts of it and move on with life. Whatever parts—either in the buyer or the seller—it doesn't matter. Just merge it. Again, it's a "who cares?" pragmatically. You can fix it all sorts of ways.

Either the cultures mesh or they don't. If they don't mesh, you don't have to keep everybody, honestly, because everybody's going to do very well off the acquisition. You can do all sorts of thank-you packages and move on with life.

Third, sometimes there are dynamics around how you value the things relative to each other for private-to-private companies. Sometimes the easiest way to do that is to choose some metric and say it's divisible by that metric.

For example, years ago when I was at Twitter, I drove an attempt to buy a major social network that was up-and-coming. The way we constructed that offer was that we took their users and our users, added them up, calculated the ratio, and made that offer as a portion of Twitter for the company.

I think you can do that. Take your revenue plus my revenue, add it up, and then what's the ratio? Or maybe it's users. It's whatever the right metric is for your business. I actually think you can do really simple things like that and just say, "Look, fair enough. Plus or minus X% isn't going to matter if we all just win."

So people tend to overthink those things. They overthink role—what am I giving up?—or ego or whatever. Culture—what does the surviving thing look like together? And then, what’s the value, or what’s the relative value, of the 2 pieces? Pragmatically, it’s like: Do you want to fight it out for the next 5 years, or do you want to go win? Then your battleground shifts to the incumbents versus another startup.

Sarah Guo

Yeah, I’ve seen the simple relative metric also work. I also think that founders, board members, and investors are just unwilling to put something like this inside the Overton window. I think people feel like it is capitulating, but it’s capitulating in service of winning. And so I think that’s a big reason people don’t want to look like they’re unwilling to go to war.

Elad Gil

Yeah. The pie basically gets bigger if you do that because you’re focused on just winning the market versus competing with each other, but your pricing dynamics shift as well. You’re not competing on every deal with another startup. A lot of things shift, and so I think there are all sorts of positive characteristics. Again, people will win in these markets without it.

Some markets are really big, and there is room for a number 1 and a number 2 and maybe a number 3, or maybe incumbents. Payments was that way, right? We have Adyen, Stripe, PayPal, and a dozen other payment processors. It’s a very big, fragmented market. Some markets can sustain multiple players, and that’s fine too. I’m just saying sometimes you want to say, “Hey, let’s put aside our differences and go win together.”

Sarah Guo

Okay. Some part of the market is consolidated. Some could be better consolidated in terms of startups winning. There are areas that you and I have talked about where they feel like obvious commercial opportunities, but people are not chasing them sufficiently, I think. We’ve talked about engineering as one that AI will absolutely change. You have a bunch of ideas in biotech. What’s missing?

Elad Gil

Yeah, the biotech stuff I’m interested in honestly isn’t AI-related, although there’s obviously really cool things happening in terms of models. There’s a whole separate thread of stuff I just think is neat. I’m not an active biotech investor; I’m the wrong person to pitch on things, et cetera. I mainly do software, AI, and so on as investments, as well as the companies I’ve started, which have largely been software-driven companies.

I just think there’s some really cool stuff now that the science in biotech—or in basic science—is far enough along, and nobody, or very few people, are working on it. I’ll give you maybe 2 or 3 examples. One is there’s some really good data now for fertility out of Japan where you can basically take a cell and reprogram it to turn into either a sperm or an egg.

They’ve made mice now with 2 fathers, for example. You could differentiate one father’s cells into sperm and one father’s cells into eggs, and then you can have viable offspring. That really opens up the capability for any adult to have kids with any other adult.

So if a woman is over a certain age, she can suddenly produce either sperm or egg. You can do it for different types of couples. There’s stuff like that where you’re like, why are so few people working on this?

An even simpler version is that girls are born with 1 to 2 million oocytes, which are egg cells. By puberty, they end up with about 300,000, and there aren’t good technologies to basically mature those eggs. If you’re a woman, you should be able to mature your oocytes at different points in your life, and you should be able to harvest tons and tons of eggs if you ever want to have lots of kids, right? There’s a lot of stuff like that that just nobody’s doing.

Sarah Guo

Is the outcome of that that people choose the inputs to having kids differently? For example, the sperm or egg donor market is very different. We’re all just having kids with Elon—you and me both.

Elad Gil

The crazy thing about that, honestly, is say that you meet Elon Musk, LeBron James, Taylor Swift, or whoever it is somewhere and you manage to swap some cells off of them—you shake their hand or whatever—you could potentially reproduce them.

Yeah. No, seriously. Some of the ramifications of this stuff are pretty crazy if you think about it, right? But for society, it’s so impactful in terms of what you could do with that. To your point, suddenly anybody could become an egg or sperm donor in any capacity. It just seems like it has such big implications, even if you just say, “We’re going to limit it to women over a certain age,” or people who just aren’t reproductively viable otherwise, right? It’s a pretty big deal in my opinion.

But again, the science is there. They’ve worked through a lot of the pathways to get there, and now it’s like, okay, I know 1 company doing it. But it’s driven by a very good founder; 1 company, that’s it.

Another area would be: You look at Botox. People are injecting a bacterial toxin into their skin to look younger—literally, a toxin—and that was a $40 billion company, with $1.5 billion a year in revenue just for cosmetic applications. Why isn’t anybody doing actual drugs and treatments for aging? There’s all sorts of science around it, all sorts of biology. Nobody’s working on skin aging, balding, gray hair, all that kind of stuff.

Then there’s the stuff that’s really impactful in terms of neurosensory, right? The muscle that holds the lens of your eye gets weaker with time, and so why don’t you rejuvenate that? That’s why everybody ends up with reading glasses in their 40s. Or hearing loss—there are pathways for that. Or tooth regrowth: You have a cavity; why don’t you just grow a new tooth? There are pathways for it.

Again, a lot of the biology is worked out. Maybe there’s more that needs to be done from a basic science perspective. In many cases, for example, for dental stuff, there are genes like USAG-1, which allow for tooth regrowth in certain animal models. So why don’t we do that in people?

Sarah Guo

What’s your hypothesis for why there are areas that, to me, seem like clear demand if the science you suggest exists? Why isn’t it being funded?

Elad Gil

Yeah, it’s massive markets. I think there are 3 reasons. Number 1, the biotech or biopharmaceutical market for founders is very different from the tech market, and the overall market structure is radically different.

If you look at biotech, the last time a $50 billion-plus biotech company was started from scratch, excluding Moderna, which was kind of an accident of COVID, was in the ’80s. I think it was Regeneron. It’s been almost 40 years since we’ve had a de novo, tens-of-billions-of-dollars company created. All these companies are 50 or 100 years old.

Imagine if tech were basically IBM versus HP right now, and you didn’t have any young, founder-driven, aggressive companies. We wouldn’t have the iPhone. We wouldn’t have the internet. We’d just be logging into IBM mainframes off of HP laptops. Do you know what I mean? There’d be no progress, or very little progress.

That’s one issue. The funding models also are ones where a lot of biotech money is either very early-stage or very late-stage, and a lot of the companies are started as incubations by biotech VCs. They load up a company with $40 million, they buy 40% of it upfront, whatever it is, and then they kind of have to make it far enough that they can get almost public-market money, effectively. A lot of the crossover funds then kick in.

The way these funds are set up, because they have so much ownership, they’re really built to flip these companies into the arms of pharma. That means you build against pharma pipelines. If there are 6 or 7 areas that all the pharma companies care about—it’s cancer, cardiovascular disease, and neuroscience—you only build companies in those domains because your goal isn’t to build a big standalone thing. Your goal is to sell it to a pharma company.

A lot of the dynamics are driven by that. And then there’s big regulatory capture that also prevents a lot of innovation. The FDA will ask for—they’ll push hard on—endpoints for certain things that may not exist. There are those 3 main factors that make it kind of hard to do anything else, but all the science is just sitting there, right?

Oh, I guess the last piece, the fourth, is that for some of these things, the scientists who would work on them don’t want to work on something that’s too commercial. It’s kind of the purity of science. It’s low status. How dare you work on fixing wrinkles? As a scientist, you need to be doing something that’s much more pure, et cetera, et cetera. So there’s also a little bit of that—what do you call it?—prudishness around commerciality that exists.

Sarah Guo

I guess, back to our regular programming, I had a question for you on the AI side. In particular, I know you’ve been thinking a bit about world models and RL, and how these things are overall relevant to capability and the scaling of capabilities. Do you want to explain a little bit about what you mean by world models?

People who are in the AI world get all this stuff, but it would be great for a more general-purpose audience if you could walk through your thinking and what you think is interesting and going on there.

Elad Gil

I think it’s very important as an overall area because, if you zoom all the way out, I actually think this is a time of more open research questions than ever. Scaling up model size and training data for big LLMs has given us this really powerful foundation of knowledge and pattern recognition. But everybody talks about agents—what people want to do from here. The way people think about AGI is not just predicting text, right? They want to move toward broader intelligence and taking actions.

I think it’s really important to describe what we mean when we say “reasoning” or “actions” more concretely, because I don’t know that everybody has a great mental model for these things. It could be planning, reading documents and drawing conclusions, using tools, receiving feedback, going down different reasoning paths, or evaluating your own work. It’s taking a series of actions in pursuit of a goal beyond just sequential text generation.

My understanding is that the labs—some labs more than others—have spent a lot of money collecting traces of humans doing sophisticated tasks. This is how Elad looks at Japanese stem cell differentiation research, right? He does these tasks, calls these people, and then they try to do behavior cloning: monkey see, monkey do, but for software engineering or investment research or whatever.

But it tends to be really brittle when you go off the path with the cloning techniques. The model all of a sudden presses some button that the human never touched, lands in out-of-distribution territory, and then has no idea what’s next. It fails; you get stuck.

Then people are trying a new generation of reinforcement learning, which is broadly trial-and-error training. I think a lot of people who are paying attention to AI have seen agents play games, famously chess and Go, or more complicated games with human interaction. You’re taking actions in an environment and getting feedback in the form of a reward or a penalty, and then you play until you’re better at the game.

For games, that’s easy because you have clear rules, so you know very easily how to either reward or penalize an action. That’s very different from real-world tasks in some cases, and this is exactly the problem with using RL more broadly: What is the task if it’s not just winning in chess or Go? How do you make the environment? You’re trying to make a copy of the universe, or at least some little piece of it, that’s rich enough to teach useful problem-solving but cheap enough to run.

I don’t know how we get to the Matrix. It’s very hard to design rewards, and then you have a gap from reality. You also need diversity, or you’re just memorizing a path through your game—even if that game is the game of an agent doing research work or the game of a software engineering project—and you’re overfitting instead of adapting.

I don’t actually have a ton of conclusions here, but I’ve spent a little bit of time trying to understand it. There’s an interesting set of researchers now who are working on creating more universal environments and world models, or just trying to get better trace data. I actually don’t know that I believe any of the more immediate-term commercial applications of these models are interesting. People say, “Oh, we can generate games, or we’ll have gaming assets, or we’ll use the data for robotics training,” or some other thing. But I do think it’s a really interesting conceptual path toward more AGI.

Sarah Guo

One thing I think is intriguing in what you said—and it’s one of the points that I’ll overextrapolate—is that, if you look at the way AI has done certain things, for example in Go, because there is a utility function but no other constraints, it came up with all sorts of crazy moves that a human wouldn’t have come up with, or at least hadn’t come up with to date. Then humans started studying and copying these moves. They were completely out of the box, but they ended up with a superior outcome.

I always wonder what that looks like for other areas of human endeavor. If coding shifted from, “Hey, let’s copy how people write code,” to, “Let’s just solve this problem,” how different is the type of code that’s written? What sort of traditional approaches are just broken that we can then learn from, because you’ve created a utility function with an unconstrained approach to actually figuring it out?

That happens sometimes in biology, right? You’ll do these molecular evolution experiments where you’ll evolve a molecule to do something, and sometimes it’ll do things in a really weird way that you just completely don’t expect. Suddenly you have this catalyst that works in a really weird way, or a binding protein that doesn’t do it the way you’d expect at all. It’s because it’s not designed; it’s evolved. I think this whole notion of evolved systems or self-selecting systems can yield really weird insights, and I’m really excited to see that kind of stuff in terms of the outcomes.

Elad Gil

Me too. One way I visualize this is that a model is looking in a part of the search space that humans have not traditionally been taught by the Go rulebook, the prior games, or whatever. It could be in the shape of a protein or any other problem.

Have you seen the TV show Pantheon?

Sarah Guo

No. What is that?

Elad Gil

It’s a TV show about AI and mind uploading. It’s a kind of niche animated TV show. You should watch it. Everybody should watch it.

Sarah Guo

I think it’s really interesting because the uploaded beings at some point become your full self—or at least, for us, it would be humans learning to think differently. It’s breaking through your constraint of how you might traditionally solve the problem or see yourself. I do think that, thematically, it’s one of the more inspiring things about AI.

Elad Gil

Oh, that’s interesting. I feel like there are a lot of sci-fi books where eventually you have your brain uploaded into the cloud or whatever, and then there are all sorts of controls you suddenly have access to that you didn’t have before. For example, you should be able to fine-tune your emotions or your emotional state and dial it up and down literally with dials.

I think there are always these really interesting meta-questions. If a human upload were to occur, what does a transhuman species look like? What are the capability sets that aren’t a priori obvious that you suddenly expose?

I mean, obviously, you could also spawn instances of yourself and have those things go do things for you and then merge back in. Maybe some of them don’t want to merge back in, and then who’s the real identity? You know, all that stuff. It’s kind of fun.

Sarah Guo

I’m told that the modulation of emotions and attention actually doesn’t require upload. Fred and some professors we know would say it’s just ultrasound devices coming soon to a consumer shelf near you. But we can talk about that on next year’s episode.

Elad Gil

Yeah, sounds good.

No Priors 第116期|与 Sarah 和 Elad 对谈 — 文字稿与摘要 | BidClub