[BidClub_]
Dwarkesh Podcast · · 20 分钟

下一代训练范式会是什么样?

Dwarkesh Patel

YouTube
TL;DR
  • Dwarkesh 的核心判断是:各实验室押注的研究路线——在“数千个多样化 RL 环境中”把 RLVR 扩展到“数百万个可验证任务”,从而用规模碾平样本低效、甚至让持续学习变得多余——很可能并不完整。 真正的下一次突破可能是 AI 在工作中学习,通过部署经验更新权重。
  • AI 进步中被低估的约束是:可验证性还不够,一个领域还必须“可刷”(grindable),也就是能在确定性、可复现的模拟器上并行运行数千次 rollout。 这也是计算机操作明显落后于 coding 的原因,尽管前者显然可验证:“Andy Jassy 会找到你的机器人,然后把你们的屁股踢出去。”
  • 企业经营、打赢官司或实现盈利交易等技能无法被容器化;外循环验证可能需要数月甚至数年的真实世界行动,因此对大多数经济价值最高的领域而言,样本效率才是决定性瓶颈。
  • 他把 Dario 关于训练与服务阶段上下文长度的引述,解读为 RLVR 泛化能力并非无限强的证据。 如果短时域训练不一定能泛化到长时域,那么白领任务训练凭什么能造出一个在 2002 年拿着 1亿美元“替你创办 SpaceX”的智能体?
  • 一个可交易的算力切口是:实验室 30-50% 的算力用于推理,而这些算力“目前并未在帮助模型改进方面发挥任何生产性作用”。 这是一种巨大的浪费,因为最有价值的训练信号——组织特有的隐性知识——恰恰存在于部署过程中。
  • 两套候选方案是:on-policy self-distillation(OPSD),即把每次会话中的学习以密集监督蒸馏进权重,同时保留稀疏更新;以及更具推测性的“dreaming”,让模型在自建模拟器上训练,成为与预训练、RL 和推理时算力并列的第4条扩展轴。
  • 2027-28 年的情景是:RLVR 先产出一个足以部署的智能体,1周级上下文配合“工作复盘”的点赞触发器,把会话蒸馏回基础模型。 随后能力会从可验证领域向相邻领域大幅扩展——“你每次与 AI 互动时,它都会更聪明……这既令人恐惧,也令人兴奋。”
摘要 · 为研究而整理的核心内容

1. 实验室对 RLVR 的押注

  • Dwarkesh 对自己正在检验的乐观情景的铺陈是:在数百万个可验证任务上训练,就能得到一个通用的问题求解智能体;数据效率不足之类的缺陷,“只要继续扩大训练规模就能直接碾过去”,就像投入足够算力后,自然语言处理中的基础研究问题都相继坍缩。
  • 乐观派对其样本效率质疑的回答是:模型的样本效率只有人类的百万分之一,但训练只是一次性成本,可以摊薄到数十亿次会话中;真正重要的是单次会话内的能力,而 RL 显然正在改善这一点。
  • 只要上下文窗口足够长,持续学习“可能根本没有必要”:员工上岗“6个月甚至更久”后才实现净生产力——“如果能把这6个月直接塞进上下文窗口呢?”

2. 可验证还不够,领域必须“可刷”

  • 他对一个问题的解释是:为什么计算机操作明显慢于 coding,明明“我的税务申报提交成功了吗?”显然可以验证?除了多模态预训练数据薄弱这一因素,更被低估的原因是,必须有确定性、可复现的模拟器,才能从相同起点并行运行 rollout。Coding 可以让1000个智能体在相同容器中工作;“你不能让1000个智能体同时去 Amazon 试同一套结账流程……Andy Jassy 会找到你的机器人,然后把你们的屁股踢出去。”
  • 克隆 Slack 和 Gmail 的确可行,但“劳动密集且无法扩展”——除非 AI 自己把这些克隆品写出来,而这同时又会成为极佳的 coding RL 目标。因此,计算机操作可能很快就会被解决,但它的迟缓揭示了“AI 进步这条河流只能慢慢凿穿的峡谷壁”。

3. 许多高价值技能无法批量刷出,Dario 暗示 RLVR 泛化存在上限

  • 最棘手的领域包括企业经营、诉讼、交易和选举:它们都是无需重置、且不断变化的环境,验证“可能需要数月甚至数年的真实世界行动”,也无法通过扰动条件后的并行 rollout 重新观察。“要造一个政治能力达到 Lyndon Johnson 水平的 AI,RL 环境究竟是什么?”
  • 砸下1万亿美元打造 RL 环境,最终得到一个被投放进 1948 年得州政坛后、给出的建议优于 LBJ 的智能体,这件事“仍是一个经验问题”。但 Dwarkesh 从 Dario 关于训练与服务阶段上下文长度退化的引述中读出了某种信号,并谨慎补充“也许是我过度解读了”:如果短时域 RL 不一定能泛化到长时域,白领任务训练凭什么能让 AI 像 Sam Walton 一样经营企业?

4. 信号存在于部署中,却被直接丢弃

  • 算力浪费在于:实验室 30-50% 的算力用于推理,却没有为模型改进贡献任何东西;而部署恰恰会暴露最有价值的信息——“我在真实世界里通常会犯什么类型的错误?”他的比喻是:这就像一个从未被允许参加真正实习的天才研究生,只能不断学习课堂案例。
  • 上下文无法替代真正的学习:不断膨胀的 KV cache 不具备可扩展性,也不是人类的工作方式——人类不会因为记忆变多就把头骨撑大;那些能以模型般的精度记住无意义音节的 savant,在抽象能力上却可能严重受限。真正的持续学习,是“把正确的直觉重新凿回权重里”。
  • 但梯度更新的样本效率太低,已经上线的在线学习必须让数百万用户共享同一个目标——Cursor Tab 每天要预测 4亿次请求中的编辑接受情况;而真正的持续学习需要针对每次部署的具体性,“根本不可能被塞进某个共享训练任务”。

5. OPSD 与“dreaming”:候选方案

  • 瓶颈可能不在架构:稀疏注意力、KV 压缩方案已经很多,真正的问题也许是损失函数。OPSD 来自他与 Sasha Rush 的黑板讨论:让基础模型匹配带有完整上下文的“资深教师”模型的预测,不需要外循环的可验证奖励,而每个 token 的偏差也比单个最终奖励提供密集得多的信号。
  • OPSD 也比 SFT 更适合这项任务:以“完美保真”的方式回忆 transcript,本身就是错误目标;RL 式的稀疏更新——“只在绝对必要的范围内改变模型”——可以避免覆盖基础模型。他也推翻了自己此前的判断:RL 每个样本学得更少,“可能是好事,而不是坏事”。
  • 更具推测性的押注是“dreaming”:让模型构建自己的模拟器并反复演练。EfficientZero 在一个从未见过的 Atari 游戏中,训练2小时后“很可能已经胜过新手人类”,但前提是每走一步真实操作,都要在脑中“玩上几十局模拟游戏”。如果有效,这会成为第4条扩展轴;不再是 /compact 所谓的“持续学习拟像”,而是直接进入 /dream,“焚烧掉海量算力”。

6. 2027-28 年的交接:部署本身成为训练任务

  • 这一情景中,RLVR 真正带来的礼物,是一个“至少有能力开始获得真实世界经验”的智能体:先把它部署出去,让有效上下文延伸到完整1周的共同工作,再由一次点赞式的“工作复盘”触发通过 OPSD、dreaming 或“我们甚至还不知道的技术”把这段会话蒸馏回基础模型。
  • 这个循环会不断复利:每一轮,AI 都会在与此前在线学习相邻的领域变得更强,能力范围“远远超出可验证领域”,直到改善能力的主要驱动力不再是发布前训练,而是广泛的经济部署。“你每次与 AI 互动时,它都会变得更聪明,不仅因为它从你此前的会话中学习,也因为它从与世界上所有其他用户的互动中学习。而这既令人恐惧、令人兴奋,也不同于 AI 目前的进步方式。”
Dwarkesh Patel

So here's a big research bet that all the labs are making. They think that if we train AIs to accomplish millions of verifiable tasks across thousands of diverse RL environments, then we will have basically built AGI, because this kind of training will have created a kind of problem-solving agent: the kind of thing that can make progress on open-ended tasks for weeks on end in the face of errors, mistakes, and ambiguity.

The people who are optimistic about this vision will say that all these things we talk about as the fundamental deficits in the current training paradigm—for example, the data inefficiency of these models or the fact that they lack continual learning—can just be steamrolled if we scale training more, in the same way that all the fundamental research problems in natural language processing collapsed when we threw enough compute into LLMs.

So, in the previous essay, I talked about how these models are one-millionth as sample-efficient as humans, and the people who are in favor of the current training paradigm will say, “Look, that might be true, but this is only true during training.” Training is this one-time cost that is amortized across billions of sessions that a model will experience.

What really matters is how smart, general, and sample-efficient the model is during a session, and this has clearly been improving as we've been doing more RL training. AI agents are able to solve more and more ambitious problems over longer and longer time spans. Anybody who has used these models for coding knows that.

Similarly, people would say, “Look, continual learning—this capability I keep harping about, where the model's weights get updated based on what it's learning from deployment—may simply not be necessary.” Because if in-context learning gets so good across longer and longer time horizons, then you don't need to distill everything the model is learning on the job back into the weights.

People often say that their employees are not net productive until 6 months or more on the job. So clearly, online learning is necessary for competence. But what if you could just fit those 6 months into the context window?

There have been tons of architectural innovations that dramatically increase the amount of information, or the amount of context, that a transformer can store. Why not think that, with a couple more years of progress, we might have what feels like infinitely large context windows?

1. Grindability is just as important as verifiability

Before we discuss this research bet a bit further, I want to step back and ask a completely tangential question, which I find actually very interesting and confusing about the nature of current AI progress: Why has progress on computer use been so much slower than in other domains?

Computer use is so clearly verifiable. You could ask a question like: Did the desired Etsy item I ordered get delivered? Is the venue for an event I'm trying to organize booked? Have my taxes been submitted?

Isn't it weird that computer use has been making so much slower progress than coding, math, and these other verifiable domains? I'm sure there are many reasons for this, and one of them, of course, is the fact that the models are exposed to far less high-quality multimodal data during pretraining.

But one reason that I think is actually quite underrated by people, and which I think reveals the canyon walls against which this river of AI progress will only slowly chip away, is that it is not enough for a domain to be verifiable. It also has to be very grindable, in the sense that you have to be able to run lots of parallel rollouts against a deterministic and replayable simulator, and you have to run those rollouts from the same starting point.

If you're trying to make a model better at coding, you can define some container that has a software repo with some missing feature that you have tasked the AIs with creating. Then you have 1,000 parallel agents go at the problem, each of which has an identical copy of the container.

But this doesn't work with computer use, at least not trivially. You can't just have 1,000 agents go try the same checkout flow on Amazon to get better at using websites, because Andy Jassy will find your bots and shut your ass down.

You can solve this by making clones of Slack, Gmail, and all the other common applications and websites. But at least currently, this is a very labor-intensive and unscalable way to build environments.

Of course, once AIs get good enough at coding themselves to build these clones with extremely high fidelity, then I'm sure computer use will make quicker progress than it is right now. And you're also killing 2 birds with 1 stone with this kind of procedure, because getting AIs to rebuild whole applications from scratch is also a great RL objective for coding.

So while computer use itself may soon be solved, its current lethargy is telling us the following: Unless you can build a very replayable training target for a domain, the models will struggle to make much progress. And the reason this is true, of course, is that the models are incredibly sample-inefficient during training.

This is the point I was making in my last video essay. So for computer use, we might be able to make up for the sample-efficiency deficit by building these farmable, deterministic simulators. But for so many other different kinds of skills that we need AIs to have, we simply can't do this.

How do we train an AI to get really good at building a business from scratch? How about winning court cases, having a profitable day of trading in the markets, or helping a candidate win an election?

The rollout here requires interacting with the real world, and you can't recreate it from just within a data center. The outer-loop verification here may take months or even years of real-world actions to elicit, and you can't re-observe it by perturbing the model's actions slightly in thousands of parallel rollouts to isolate exactly what the model did that actually worked.

Now, dealing with such reset-free, non-stationary environments is a known open problem in RL. I'm not pointing out anything new. But I really do want to emphasize that because of the idiosyncratic and sparse nature of data in most domains in the world, you need sample efficiency in order to get proficient.

If AIs are to develop all the skills that humans have, and even skills that humans don't have, then they need to be able to learn from information revealed in unstructured, unverifiable, and ambiguous ways from scarce amounts of real-world interaction. Because in many domains, the relevant training information simply doesn't exist in any other way.

What is the RL environment to make an AI that is as good at politics as Lyndon Johnson, or as good at building a space-launch business as Elon Musk?

2. Will RLVR alone generalize?

The labs are betting that RLVR will generalize. That is, if you train on enough containerized, reproducible environments, you will develop a very general agent that can make and execute plans, learn rapidly from new information, and even pick up new skills, all within a single session.

If you drop this endlessly RLVR-trained AI into Texas politics in 1948, it could give you better advice than LBJ about winning the Senate seat. And if you gave it $100 million in 2002 and let it cook, it would build SpaceX for you.

Now, whether RLVR can generalize this well is an empirical question. If the labs went from spending billions of dollars on RL environments to $1 trillion, would you get the kind of thing that is a fully human-like general intelligence within the context window?

Dario gave a telling quote during our podcast together, which I think hints that RLVR generalization is not infinitely strong. When he was explaining why model performance tends to degrade at long context, he said: “There's 2 things. There's the context length you train at, and there's a context length that you serve at. If you train at a small context length and then try to serve at a long context length, maybe you get these degradations.”

Maybe I'm reading too much into this, but it seems like he's saying that short-horizon RL training doesn't necessarily generalize to long-horizon RL performance. And if you can't generalize from short horizon to long horizon, then how are agents supposed to generalize from getting trained at a bunch of white-collar tasks to, say, having the ability to be dropped in the real world and build a business from scratch as well as Sam Walton?

Even if, after enough in-context experience, the AIs could become like Henry Ford or Albert Einstein, all that would be ephemeral and wasted if you couldn't get those learnings back into the weights.

Around 30% to 50% of a lab's compute goes to inference, and that compute is currently not playing any productive role in helping improve the model. This seems like a huge waste. And it's even worse than it sounds, because it is only in deployment that the most valuable bits of information which your model could learn from are actually revealed.

What's actually happening in the organizations where I'm being used? What are they using me for? And what kinds of mistakes do I tend to make in the real world?

We've got some genius grad student who's never been allowed to take a real internship, and we keep giving it more and more classroom case studies in the form of RL training on environments. It's so bizarre that we have AIs that are broadly deployed through the economy already, are participating in so many different kinds of tasks, and are privy to so much domain- and organization-specific tacit knowledge, and they're not able to make use of it.

3. Getting the learning back to the weights

But this kind of continual learning requires going back to the weights. AIs can't just keep building up a bigger and bigger KV cache as they learn from more and more users. That's just not scalable, and that's also not how humans do it.

There's no clean separation in our brain between parameters and activations, and it's not like some part of your skull keeps expanding as you learn more things throughout your lifetime.

When we learn stuff, there's clearly some kind of compression, and this aids our generalization and grokking. There are, in fact, some humans who have this autistic-savant-type ability to recall random tables of numbers or nonsense syllables years later—essentially the kind of fidelity of information that models have in context. And such sheer volume cripples these humans' ability to understand abstractions and metaphors.

Human continual learning is less about having all your observations at the tip of your tongue and more about chiseling the right intuitions and big-picture knowledge back into the weights. But the moment you move into the weights, you have to give up on in-context learning's sample efficiency. Because gradient updates are super sample-inefficient, all of the successfully shipped online-learning models have had to learn the exact same thing across millions of users.

For example, the Cursor Tab model online-learns by predicting the same exact objective for over 400 million requests a day. The objective here is which edits actually got accepted by the user. At least so far, we haven't seen models online-learn different kinds of things for different users, because while a single session may generate more than enough data for a human to learn from, it's not enough to train a more capable AI.

Current online learning can work for a very limited number of use cases. But the whole point of continual learning is that the world is very complicated, and each job and company and problem is different, and you need your intelligence to be able to learn the specific information related to a particular deployment, which simply can't be stuffed into some shared training run. These are all the things we're talking about when we talk about on-the-job learning: things like how everything in your organization works and fits together, how to cooperate with all the infrastructure and the other people around you to make progress on some larger project, what the common failure modes are, and many other things like this.

In this way, sample efficiency and continual learning are actually deeply connected problems. Relatively little data is available to the model on the job. To learn from this data requires sample efficiency, and models can do that in context, but using the fast weights that are built on the fly by attention, which allow for this sample efficiency, scales very poorly in terms of memory.

So we need architectural innovations that allow for some kind of intermediate representation. I talked before about how we already have many different working ideas for this kind of thing, from sparse attention to KV cache compaction. And every week, somebody releases a new paper suggesting some kind of other architectural optimization. It doesn't seem to me that architecture is fundamentally what is bottlenecking continual learning.

So perhaps the bottleneck is the loss function. How do we update the weights, aka how do we improve the model itself, based on information that was learned from 1 particular session? Even here, naively, it seems like there are many ideas that ought to work.

A lot of people are talking about this technique called on-policy self-distillation recently. The idea is that we encourage the base model to make the same predictions when trying to solve some real-world problem as the model with all the context accumulated after a long session would have made. The whole point of this procedure is to distill what the model learned in a session back into the weights themselves.

This is better than RLVR for 2 reasons. 1, OPSD doesn't require us to have some outer-loop verifiable reward. We just need a model that can learn the right things within the context window. And as long as we have that, we can train the base model to match our veteran teacher model, which has built up all this experience during the session.

And 2, OPSD provides a much denser supervision signal than naive RL. Instead of projecting a single reward through the whole trajectory, you can train on the per-token probability discrepancy between the teacher and student.

For continual learning, OPSD is also superior to supervised fine-tuning. The most naive version of SFT for this application that you can imagine is just to train the base model to predict all the tokens that are observed during the session. But this makes no sense if you think about it as a learning target.

The way you get better at your job is not by recalling the transcript of every single thing that happened every day with perfect fidelity. Rather, it's by consolidating the handful of insights and pieces of knowledge that are actually relevant to you getting better at your job. RL training doesn't suffer from this failure mode. RL is great at concentrating the update to only what is relevant to getting the outcome right.

That's why the updates from RL are incredibly sparse. And this is a very important property for continual learning, because as you're learning on the job, you don't want to overwrite and forget all the other things that the base model knows. I wrote a post a few months earlier arguing that RL learns much less information per sample than supervised learning. But this may be a good thing rather than a bad thing.

You only change the model as much as is absolutely necessary to achieve the outcome, and no more. OPSD preserves this property of RL, where instead of slingshotting towards the teacher distribution as supervised learning would have you do, you only extract the knowledge that is necessary to achieve the same results as the teacher on actual real-world tasks.

4. Dreaming

OPSD is 1 way to attack the sample-efficiency problem. You take this scarce real-world experience, and you squeeze all the signal into a tiny, well-targeted update. But there's also another much more speculative idea. Let's call it dreaming.

If the AI can build a good simulation of reality against which to rehearse new skills, or try alternative strategies and reinforce what actually works, then AIs could experience orders of magnitude more simulated samples in the same wall-clock time. Let's go back into history a bit. A couple years after DeepMind released AlphaZero, a group of researchers trained a model called EfficientZero, and the whole point of this model is to be very efficient with data.

So if this model and a human both got 2 hours to play against a simulator of an Atari game that they hadn't seen before, this model would actually probably beat the novice human. Does this mean that the model was more sample-efficient than the humans? Well, that was the goal of the training, but it depends on how you measure sample efficiency.

Because for each step in the real game, EfficientZero is playing dozens of simulated games in its head. In a similar way, future LLMs might be able to consume far less real-world data while practicing endlessly against environments that they build for themselves. The big difference, of course, is that it will be much harder to build a simulation of the whole world than it is to emulate the game of Go.

That's why I said this is a much more speculative idea. If it works, it would become a 4th axis of scaling alongside pretraining, RL, and inference-time compute. You could call it test-time training or dreaming. The model spends compute writing up RL environments and then training against them, and it's rehearsing all the skills that will actually be used in production for a specific user.

So instead of hitting /compact in Codex or Cursor or Claude, which kindles a small amount of compute to write up a summary, and which gives you the simulacrum of continual learning, you hit /dream. And this incinerates huge amounts of compute to build and train against a video-game version of what the model is witnessing in the real world.

5. What 2027 looks like

So what might continual learning look like by 2027 or 2028? And how do we get there? Here's 1 scenario. All of this RLVR training is producing an agent that can get its bearings when it's thrown at an unfamiliar problem, and it can try different strategies, and it can iterate when it hits a roadblock.

This is the crucial thing that RLVR has given you: an AI that is at least competent enough to start getting some real-world experience, if it could learn from it. And once you have that, you send it out into the world to do real work, even on projects that are off the training distribution.

Now let's say at this point, the effective context lengths have expanded such that AIs can jam and co-work with you for a full week of wall-clock time. At the end of a week, you give it a thumbs-up or a thumbs-down, you give it a work review. And if you give it a thumbs-up, the base model distills everything that the AI learned during the session, and it may use OPSD, it may use dreaming, it may use some other technique that we aren't even aware of, or it'll use a combination of all of the above.

And AI can get better at domains that are adjacent to what it was explicitly trained for beforehand with RLVR. In the next round, it gets better at the thing adjacent to what it previously learned online. In this way, the gamut of AI skills, knowledge, and capabilities can expand far beyond the verifiable domains that the model was originally trained against before it was deployed. Just as pretraining created a base intelligence that was smart enough to become a competent agent with enough RLVR on top, so RLVR has created an agent that is competent enough to actually be broadly deployed in the world, and from this broad deployment to learn on the job once the training recipe for continual learning actually arrives.

By this point, the main way that AIs get better is not from the training they have received before they are released to the public. Rather, it’s from all this experience that they’ll be accumulating from being broadly deployed in the economy and engaging in so many different kinds of tasks. Every time that you interact with an AI, it’ll be smarter, not only because it’s been learning from your previous sessions, but also because it’s been learning from all its interactions with all the other users in the world. And that’s very scary and exciting and different from the way that AI improves right now.

This was a narration of a blog post that I also released on my website at dwarkesh.com. Go there if you want to read all the footnotes, or if you want to sign up so you can find out when I release the next blog post. Otherwise, I’ll see you on the next episode.