[BidClub_]
Dwarkesh Podcast · · 12 分钟

我们到底在扩展什么?

Dwarkesh Patel

YouTube
TL;DR
  • Dwarkesh指出的核心矛盾是:不可能一边坚持极短时间线,一边看多基于可验证奖励的RL扩展。 整条RL环境公司供应链都在把浏览器和Excel技能预先烘焙进模型——“这些模型要么很快就能自主在工作中学习,让所有预烘焙都失去意义;要么不能,这就意味着AGI并不迫在眉睫。”
  • 机器人是最好的试金石:这本质上是“算法问题”,而不是硬件或数据问题。 人类只需很少训练就能远程操控现有硬件,因此真正具备类人学习能力的模型基本就能解决机器人问题。可实验室却必须在1000个家庭里练习100万次,说明这样的学习者并不存在。
  • 收入缺口就是能力缺口。 知识工作者每年创造数十万亿美元收入,而实验室的收入仍差着几个数量级,因为“模型离人类知识工作者的能力还远得很”。把问题归咎于扩散滞后只是“自我安慰”——真正运行在服务器上的人类会比招聘真人扩散得更快:几分钟内就能读完你的Slack和云盘,还不存在人类招聘的劣币驱逐良币问题。
  • 值得交易的标志性判断是:模型的惊艳程度,正以短时间线派所预测的速度提升;但实用性,只以长时间线派所预测的速度提升。(“Models keep getting more impressive at the rate that the short-timelines people predict, but more useful at the rate that the long-timelines people predict.”)
  • RL扩展没有预训练那样的正当性。 人们试图把干净的预训练损失曲线所带来的光环“洗”到RLVR看多逻辑上,但后者背后没有公开趋势。Toby Board(听辨如此)拼出了O系列数据:“要让RL带来相当于单个GPT级别的提升,RL总计算量可能需要扩大约100万倍。”
  • 眼下不会出现失控式起飞。 未来的主要驱动力不是软件奇点,而是AGI顶端的持续学习;它会像GBT3之后解决上下文学习那样逐步实现,可能还要“再花5到10年打磨”。突破会迅速被复制,三巨头大约每个月轮换一次榜首,而且没有任何飞轮真正削弱过竞争。
  • 对远期的判断是:到2030年,各家实验室会在持续学习上取得实质进展,赚到数千亿美元收入,但仍未实现知识工作的自动化。 不过,他仍预计“未来10年至20年内会出现真正类脑的智能,这相当疯狂。”
摘要 · 为研究而整理的核心内容

1. 极短时间线与RL预烘焙不可能同时成立

  • Dwarkesh开篇提出的谜题是:整条RL环境公司供应链都在搭建训练环境,教模型浏览网页和使用Excel;Baron Millig(听辨如此)博客文章还提到,已有“数十亿美元”支付给博士、医学博士及其他专家——但如果类人学习者已经近在眼前,所有这些预烘焙都注定失败。“这些模型要么很快就能自主在工作中学习……要么不能,这就意味着AGI并不迫在眉睫。”
  • 机器人的证据在于:这根本上是“算法问题”,而不是硬件或数据问题——人类经过极少训练就能远程操控现有硬件,但实验室却必须进入1000个家庭,“练习100万次如何拿起盘子”。
  • “自动化研究员”的反驳方案是:先用拼凑式RL造出超人类AI研究员,再复制100万个这样的“自动化Ilia”,让它们解决从经验中学习的问题;但他最尖锐的回应是:“我们每卖一份都在亏钱,但可以靠规模把钱赚回来。”实验室自身的行为也透露出线索:自动化Ilia并不需要PowerPoint顾问式技能,因此它们的行动“暗示了一种世界观——这些模型在泛化能力和在岗学习上还会持续表现糟糕”。

2. 巨噬细胞问题:工作依赖无法预烘焙的上下文

  • 这场论证的核心来自一次晚餐轶事:一位长时间线派生物学家说,自己要判断切片上的一个点“到底是巨噬细胞,还是只是看起来像巨噬细胞”;AI研究员则反驳说,图像分类正是“教科书深度学习的核心地带”。Dwarkesh的判断是:为每个实验室特有的微任务搭建一套定制训练流水线,净产出并不划算;真正需要的是能从语义反馈中学习、像人一样泛化的AI。
  • 每个工作者每天都要处理“100件需要判断力、情境意识,以及在工作中学到的技能和上下文的事情”,而且日复一日各不相同——不可能只靠预先烘焙一组固定技能,就自动化哪怕一份工作。

3. 扩散滞后是自我安慰,收入缺口衡量能力

  • 如果模型真是运行在服务器上的人类,“它们会极快扩散”:几分钟内读完你的Slack和云盘,从其他AI员工身上提炼技能,同时跳过人类招聘中的劣币市场问题。因此AI劳动力理应比人更容易扩散;它尚未如此扩散,只能说明“这些模型就是缺能力”。
  • 更诚实的标尺是:到了AGI阶段,买方每年会在token上花费数万亿美元,对应知识工作者每年数十万亿美元的工资总额。实验室的实际收入仍差着几个数量级,这本身就是信号。
  • 对于不断移动的目标线,他也承认其中一部分有合理性:“如果你在2020年给我看Gemini 3,我会确信它能自动化一半知识工作。”每解决一个“足够大的瓶颈”——理解、少样本学习、推理——就会发现“智能和劳动远比我此前意识到的复杂得多”。他的2030年预测是:持续学习取得进展,收入达到数千亿美元,但仍无法实现全面自动化。

4. 可验证奖励RL缺乏公开的规模化趋势

  • 预训练曾给出一条跨越多个数量级的干净损失曲线,“几乎像物理定律一样可预测”;虽然它只是幂律,增长强度远不如指数增长。如今人们试图把这条趋势的光环“洗”到RLVR上,但RLVR并没有公开可知的规模化趋势。
  • 当研究者真正把数据点串起来时——Toby Board(听辨如此)横向整理了O系列基准测试——结论反而偏空:要让RL带来相当于单个GPT级别的提升,“RL总计算量可能需要扩大约100万倍”。

5. 持续学习会逐步到来,不会一锤定音

  • 未来增长的主线不是软件奇点,也不是软硬件叠加的奇点,而是AGI顶端的持续学习——人类主要就是“从经验中变得有能力”。Baron Miller(听辨如此)描绘的路径是:专业化代理——Karpathi(听辨如此)所谓的“认知核心”加上特定工作的技能——先执行真实工作,再通过批量蒸馏“把所有学习成果带回蜂群模型”。
  • 解决这个问题会像解决上下文学习一样漫长:GBT3在2020年已经证明上下文学习可以非常强大,GPT3论文的标题也是“Language Models are Few-Shot Learners”,但GPD3发布时,上下文学习仍谈不上“已经解决”。实验室明年会推出某种名为持续学习的东西,并将其算作进展,但达到人类水平的在岗学习可能还要再花5—10年。
  • 因此不会出现失控式增长。正如Satya(可能是他)在播客中所说,如果某种能力突然被彻底解决,“那可能就是比赛结束、胜负已定”——但事情大概不会这样发展。初步突破很快会被逆向拆解并复制;此前的飞轮——聊天参与度、合成数据——“几乎没能削弱竞争不断加剧”的趋势,三巨头仍大约每个月轮换一次榜首。
  • 他坚持认为市场低估的上行空间,不是当前范式的简单延续,而是“服务器上数十亿个类人智能”,它们能够复制并合并所有学习成果;这预计会在“未来10年至20年内”出现,“相当疯狂”。
Dwarkesh Patel

I’m confused why some people have super-short timelines yet, at the same time, are bullish on scaling up reinforcement learning on top of LLMs. If we’re actually close to a humanlike learner, then this whole approach of training on verifiable outcomes is doomed.

1. What are we scaling?

Currently, the labs are trying to bake in a bunch of skills into these models through mid-training. There’s an entire supply chain of companies that are building RL environments that teach the model how to navigate a web browser or use Excel to build financial models. Now, either these models will soon learn on the job in a self-directed way, which will make all this pre-baking pointless, or they won’t, which means that AGI is not imminent.

Humans don’t have to go through a special training phase where they need to rehearse every single piece of software that they might ever need to use on the job. Baron Millig made an interesting point about this in a recent blog post he wrote. He writes:

“When we see frontier models improving at various benchmarks, we should think not just about the increased scale and the clever ML research ideas, but the billions of dollars that are paid to PhDs, MDs, and other experts to write questions and provide example answers and reasoning targeting these precise capabilities.

“You can see this tension most vividly in robotics. In some fundamental sense, robotics is an algorithms problem, not a hardware or a data problem. With very little training, a human can learn how to teleoperate current hardware to do useful work. So if you actually had a humanlike learner, robotics would be in large part a solved problem.

“But the fact that we don’t have such a learner makes it necessary to go out into a thousand different homes and practice a million times how to pick up dishes or fold laundry.”

Now, one counterargument I’ve heard from the people who think we’re going to have a takeoff within the next 5 years is that we have to do all this clunky RL in service of building a superhuman AI researcher. Then, a million copies of this automated Ilia can go figure out how to solve robust and efficient learning from experience.

This just gives me the vibes of that old joke: “We’re losing money on every sale, but we’ll make it up in volume.” Somehow, this automated researcher is going to figure out the algorithm for AGI, which is a problem that humans have been banging their heads against for the better part of a century, while not having the basic learning capabilities that children have. I find it super implausible.

Besides, even if that’s what you believe, it doesn’t describe how the labs are approaching reinforcement learning from verifiable reward. You don’t need to pre-bake in a consultant’s skill at crafting PowerPoint slides in order to automate Ilya. So clearly, the labs’ actions hint at a worldview where these models will continue to fare poorly at generalization and on-the-job learning, thus making it necessary to build into these models beforehand the skills that we hope will be economically useful.

Another counterargument you can make is that, even if the model could learn these skills on the job, it is just so much more efficient to build in these skills once during training rather than again for each user and each company. And look, it makes a ton of sense to just bake in fluency with common tools like browsers and terminals.

Indeed, one of the key advantages that AGIs will have is this greater capacity to share knowledge across copies. But people are really underrating how much company- and context-specific skill is required to do most jobs. And there just isn’t currently a robust, efficient way for AIs to pick up these skills.

2. The value of human labor

I was recently at a dinner with an AI researcher and a biologist. It turned out the biologist had long timelines, and so we asked why she had these long timelines. Then she said, “One part of work recently in the lab has involved looking at slides and deciding if the dot in that slide is actually a macrophage or just looks like a macrophage.”

The AI researcher, as you might anticipate, responded, “Look, image classification is a textbook deep-learning problem. This is dead-center in the kind of thing that we could train these models to do.”

I thought this was a very interesting exchange because it illustrated a key crux between me and the people who expect transformative economic impact within the next few years. Human workers are valuable precisely because we don’t need to build in special training loops for every single small part of their job.

It’s not net productive to build a custom training pipeline to identify what macrophages look like, given the specific way that this lab prepares slides, and then another training loop for the next lab-specific microtask, and so on. What you actually need is an AI that can learn from semantic feedback or from self-directed experience and then generalize the way a human does.

Every day, you have to do 100 things that require judgment, situational awareness, and skills and context that are learned on the job. These tasks differ not just across different people but even from one day to the next for the same person. It is not possible to automate even a single job by just baking in a predefined set of skills, let alone all the jobs.

In fact, I think people are really underestimating how big a deal actual AI will be because they are just imagining more of this current regime. They’re not thinking about billions of humanlike intelligences on a server that can copy and merge all the learnings.

3. Economic diffusion lag is cope

And to be clear, I expect this, which is to say I expect actual brain-like intelligences within the next decade or 2, which is pretty fucking crazy. Sometimes people will say that the reason AIs are more widely deployed right now across firms and already providing lots of value outside of coding is that technology takes a long time to diffuse.

I think this is cope. I think people are using this cope to gloss over the fact that these models just lack the capabilities that are necessary for broad economic value. If these models actually were like humans on a server, they’d diffuse incredibly quickly.

In fact, they’d be so much easier to integrate and onboard than a normal human employee is. They could read your entire Slack and drive within minutes. And they could immediately distill all the skills that your other AI employees have.

Plus, the hiring market for humans is very much like a lemons market, where it’s hard to tell who the good people are beforehand. Obviously, hiring somebody who turns out to be bad is very costly. This is just not a dynamic that you would have to face or worry about if you’re simply spinning up another instance of a vetted AI model.

So for these reasons, I expect it’s going to be much easier to diffuse AI labor into firms than it is to hire a person. Companies hire people all the time. If the capabilities were actually at AGI level, people would be willing to spend trillions of dollars a year buying tokens that these models produce.

Knowledge workers across the world cumulatively earn tens of trillions of dollars a year in wages. And the reason that labs are orders of magnitude off this figure right now is that the models are nowhere near as capable as human knowledge workers.

4. Goal-post shifting is justified

Now, you might be like, “Look, how can the standard have suddenly become that labs have to earn tens of trillions of dollars in revenue a year? Until recently, people were saying, ‘Can these models reason? Do these models have common sense? Are they just doing pattern recognition?’”

Obviously, AI bulls are right to criticize AI bears for repeatedly moving these goalposts, and this is very often fair. It’s easy to underestimate the progress that AI has made over the last decade. But some amount of goalpost shifting is actually justified.

If you showed me Gemini 3 in 2020, I would have been certain that it could automate half of knowledge work. So we keep solving what we thought were the sufficient bottlenecks to AGI. We have models that have general understanding. They have few-shot learning. They have reasoning. And yet we still don’t have AGI.

So what is a rational response to observing this? I think it’s totally reasonable to look at this and say, “Oh, actually, there’s much more to intelligence and labor than I previously realized.”

While we’re really close, and in many ways have surpassed what I would have previously defined as AGI, the fact that model companies are not making the trillions of dollars in revenue that would be implied by AGI clearly reveals that my previous definition of AGI was too narrow. And I expect this to keep happening into the future.

I expect that by 2030, the labs will have made significant progress on my hobby horse of continual learning, and the models will be earning hundreds of billions of dollars in revenue a year, but they won’t have automated all knowledge work. And I’ll be like, “Look, we made a lot of progress, but we haven’t hit AGI yet. We also need these other capabilities. We need X, Y, and Z capabilities in these models.”

5. RL scaling

Models keep getting more impressive at the rate that the short-timeline people predict, but more useful at the rate that the long-timeline people predict. It’s worth asking: What are we scaling with pretraining?

We had this extremely clean and general trend in improvement in loss across multiple orders of magnitude in compute. That was on a power law, which is as weak as exponential growth is strong. But people are trying to launder the prestige that pretraining scaling has—which is almost as predictable as a physical law of the universe—to justify bullish predictions about reinforcement learning from verifiable reward, for which we have no publicly known trend.

When intrepid researchers do try to piece together the implications from scarce public data points, they get pretty bearish results. For example, Toby Board has a great post where he cleverly connects the dots between the different o-series benchmarks, and this suggested to him that:

“We need something like a 1,000,000× scale-up in total RL compute to give a boost similar to a single GPT-level model.”

6. Broadly deployed intelligence explosion

End quote. People have spent a lot of time talking about the possibility of a software singularity, where AI models will write the code that generates a smarter successor system, or a software-plus-hardware singularity, where AIs also improve their successor's computing hardware. However, all these scenarios neglect what I think will be the main driver of further improvements atop AGI: continual learning. Again, think about how humans become more capable than anything else. It's mostly from experience in the relevant domain.

In our conversation, Baron Miller made this interesting suggestion: The future might look like continual-learning agents who are all going out, doing different jobs, generating value, and then bringing back all their learnings to the hive-mind model, which does some kind of batch distillation on all of these agents. The agents themselves could be quite specialized, containing what Karpathi called the cognitive core, plus knowledge and skills relevant to the job they're being deployed to do.

Solving continual learning won't be a singular, one-and-done achievement. Instead, it will feel like solving in-context learning. GBT3 already demonstrated that in-context learning could be very powerful in 2020. Its in-context-learning capabilities were so remarkable that the title of the GPT3 paper was “Language Models are Few-Shot Learners.” But of course, we didn't solve in-context learning when GPD3 came out. Indeed, there's still plenty of progress that has to be made, from comprehension to context length.

I expect a similar progression with continual learning. Labs will probably release something next year which they call continual learning and which will, in fact, count as progress toward continual learning. But human-level, on-the-job learning may take another 5 to 10 years to iron out. This is why I don't expect runaway gains from the first model that cracks continual learning and gets more and more widely deployed and capable.

If you had fully solved continual learning and it dropped out of nowhere, then sure, it might be game, set, match, as [Speaker?] put it on the podcast when I asked him about this possibility. But that's probably not what's going to happen. Instead, some lab is going to figure out how to get some initial traction on this problem, and then playing around with this feature will make it clear how it was implemented. Other labs will soon replicate the breakthrough and improve it slightly.

Besides, I just have some prior that the competition will stay pretty fierce between all these model companies. This is informed by the observation that all these previous supposed flywheels, whether that's user engagement on ChatGPT or synthetic data or whatever, have done very little to diminish the greater and greater competition between model companies. Every month or so, the big 3 model companies will rotate around the podium, and the other competitors are not that far behind. There seems to be some force—potentially talent poaching, the rumor mill, or just normal reverse engineering—that has so far neutralized any runaway advantage that a single lab might have had.

我们到底在扩展什么? — 文字稿与摘要 | BidClub