[BidClub_]
Dwarkesh Podcast · · 12 分钟

AI 核心的数据黑洞

Dwarkesh Patel

YouTube
TL;DR
  • Dwarkesh 的核心论点是:AI 的进步主要来自拓宽数据分布,而不是样本效率出现了明确改善。 强化学习“基本上是一种合成数据生成”,而模型仍是“弗兰肯斯坦怪物”(a Frankenstein's monster)——由10亿个精心构造的样本嫁接而成,再全部缝合在一起。
  • 可交易的结构性判断是:Dwarkesh 认为,真正驱动进步的是数据,而不是超参数、训练技巧或架构。 这也解释了为什么开源模型能在数月内追上前沿模型(Epoch 报告的滞后期为4个月)。数据“可以轻松从公开 API 中蒸馏出来”,这意味着训练技巧构筑的护城河可能比市场原先以为的更薄。
  • 差距大得惊人:人类到成年时接触约2亿个 tokens,而前沿模型接触的是数十万亿至数百万亿个 tokens。 这意味着“接近100万倍的差距”。青少年只需20小时就能学会开车;即便把16年的成长和建立物理直觉也算进去,所需数据仍比 Waymo/Tesla 使用的数据少3-4个数量级。如果 AI 能像人类远程操作员那样学习,机器人产业将成为10万亿美元级产业。
  • 单纯扩大规模无法弥合这一差距:按 Chinchilla 常数,参数量无限增加也只能将数据需求削减10倍,而人类的效率高出数千至数百万倍。 人类根本处在另一条 scaling curve 上。
  • 看多逻辑依然成立:即便训练效率低得离谱,各大实验室仍可能赚得盆满钵满。 技能可以摊销到数十亿次会话上;围绕专家标注和 RL 环境生产数据的产业每年营收已达数十亿美元,很快将达到数百亿美元规模。
  • Dwarkesh 的逆向判断是:2028年对人类软件工程师的需求将高于现在。 主要因为 AI 提供了互补性投入;软件工程很可能是一种分布外工作,尽管它恰恰是 AI 被认为最先取代的职业。
摘要 · 为研究而整理的核心内容

1. 智能就是样本效率——但进展并不明确

  • Dwarkesh 开场给出的定义是:智能,就是在某个领域做到胜任所需的数据量;而且“我们在训练样本效率上到底取得了多大进展,其实并不明确”——进步主要来自加入更多、更优质的数据,RL 则是“一种合成数据生成”,把算力砸向验证器或评分标准。
  • 由于模型必须对正确解保有一定先验概率,因此每项技能都需要“多到令人咋舌的人类专家轨迹”。他的直觉案例来自 Mercor/Surge 的职位列表:Word 文档润色员、M&A 尽调撰稿人、咨询顾问市场研究模板制作员。每项技能至少对应数百名专家;使用 GRPO 时,为解决 credit assignment,每个任务还要跑数百至数千次 rollout。
  • 最具标志性的图景不是一个学会技能的人类,而是“一个弗兰肯斯坦怪物”——由10亿个精心构造的样本嫁接而成,再全部缝合在一起。

2. 数据或许解释了前沿模型的快速追赶

  • 他对 Epoch 发现的解读是:开源模型落后4个月,真正的驱动因素是数据;“数据可以轻松从公开 API 中蒸馏出来,而超参数、训练技巧和架构优化却做不到”。如果后者才是最重要的因素,开源模型追平的难度应远高于现实所见。

3. 近100万倍的差距,量化来看

  • 按每小时约2,000词计算,人类到成年时接触约2亿个 tokens,而前沿模型接触的是数十万亿至数百万亿个 tokens——差距“接近100万倍”。人类几小时就能学会远程操作机械臂;如果 AI 也能达到这种水平,机器人将成为10万亿美元级产业,市场上会有一支“无穷无尽的 Unitree G1 大军”。
  • 开车只需青少年练习20小时;即便把16年的物理直觉积累也算进去,所需数据仍比 Waymo 和 Tesla 使用的数据少3-4个数量级。

4. 3个反对意见,逐一拆解

  • 把进化视为预训练——Dwarkesh 认为 Karpathy 在其播客中提出过这一观点——并不能解释问题。基因组“只有3GB”,其中仅1-2%编码蛋白质,“空间根本不够存下这些参数”。进化找到的是超参数和损失函数,而神经连接组是从零构建的。即便接受这一解释,每项新增技能仍需要海量数据:受过教育的人学习一门新的编程语言,并不需要100名教授。
  • 多模态 tokens 也不是答案:盲人和聋人仍然具备通用智能,因此感官 tokens“并不真正是让人类变聪明的东西”。聋人手语使用者摄入的语言 tokens 可能远少于2亿,所以100万倍的差距“可能还被低估了”。
  • 继续扩大规模也不够:人类拥有100万亿个突触,模型约有5万亿个参数;但 Chinchilla 常数表明,参数量无限增加也只能将数据需求削减10倍。相比之下,人类的效率高出数千至数百万倍——“人类根本处在另一条 scaling curve 上”。

5. 实验室为何仍能赢——以及 OOD 挑战

  • 白领工作的押注是:常见任务本来就更常见,因此实验室可以把它们纳入分布;收入曲线显示,即便没有类人学习能力,这条路线依然能释放“巨量价值”。他的类比是:一个人如果必须读完 GitHub 上每一个公开代码仓库,才能成为合格的软件工程师,“在教育早期就该领 Social Security 了”;但 AI 以数GW级功率持续灌入训练,将成本摊销到数十亿次会话,因此实验室“训练效率低得荒谬,却依然能大幅盈利”。
  • 真正的挑战在分布外(OOD)工作,而且取决于具体职业。软件工程很可能就是这类工作,尽管它恰恰是 AI 被认为最先取代的职业——他的押注是:“2028年对人类软件工程师的需求会高于现在”,主要来自 AI 这一互补性投入。
  • 对于后一类工作,实验室的计划是先自动化 AI 研究,再让自动化的 AI 研究员解决样本效率问题——这一点留待未来文章展开。当前关于智能爆炸的讨论“非常笨拙”:人们要么否认 AI 能加速 AI,要么“假设最后会冒出某种上帝”,却没有去推演如何在 LLM 之上、基于 LLM 特有的智能形态实现快速进展。
Dwarkesh Patel

So one definition of intelligence is sample efficiency. That is to say, how much data do you need in a given domain to operate fluently and competently? And it’s actually not clear that we’ve made that much progress in training sample efficiency over the last few years. It seems more like we’ve just dramatically widened and improved the data distribution.

The main way that AIs have been getting better is from adding more and better data, and scaling the compute required to develop that data in the first place. Obviously, RL is the main way that this has happened. You can think of RL as a kind of synthetic data generation, where you dump a ton of compute against a verifier—or a rubric, if you have an LLM as a judge—in order to find out what the good data is in the first place.

And then you train your model to predict these correct rollouts, much in the same way that you might train that model to predict the next word in internet text. For this process to work, the model must have at least some prior probability of anticipating the correct solution in the first place, which is why you need mind-stretching amounts of human expert trajectories in every single field and skill that you want the model to eventually be competent in.

It’s hard to overstate how task-specific and bespoke this human expert data is. If you want some intuition, I recommend checking out the job descriptions on Mercor or Surge’s websites. There are listings for Word specialists who will convert legacy documents into polished Word files, and legal experts who will write realistic M&A diligence reports or securities filings, and management consultants who will write up template market research.

And it’s not only that the data have to be so domain-specific, but there has to be so much of it. Each skill corresponds to at least hundreds of human experts who are generating example completions, writing rubrics, and explaining their chain of thought. There’s a reason that the data industry producing these expert labels, and the RL environments in which these meticulously cataloged skills can congeal, is earning billions a year in revenue, soon to be deca-billions.

Now imagine if it took a couple of decades’ worth of courses with hundreds of concurrent professors and millions of practice tasks for you to learn how to polish a Word file. Even the task-count difference here understates the gap, because the models have to grind through their far more numerous tasks, each far harder.

Whereas a human student might practice a textbook problem once or twice, with GRPO, these models are generating hundreds to thousands of rollouts per task, and they need to do this to solve the credit assignment problem. The correct way to think about these models is not like a human who has learned all these different skills that you see the models displaying. It’s more like a Frankenstein’s monster that has been built out of a billion grafts of carefully constructed examples, all sewn together.

Epoch recently reported that open models lag state-of-the-art frontier models by 4 months. I think the reason it is relatively easy for open-source and previous laggards to catch up to within months of the frontier is that data is the real driver of progress. And data can be easily distilled from public APIs, whereas hyperparameters, training tricks, and architectural optimizations cannot.

If the latter were driving most of the progress, then catching up would be far harder than we are observing it to be. It is easy to forget how much data these models are trained on, and how much more it is than what we humans see in our lifetimes. We see these AIs as a galaxy glittering with capabilities. But at their center, invisible to the naked eye, holding all the constellations together, is an unimaginably massive black hole of data.

1. Comparing human vs AI sample efficiency

I just want to make a couple of points of comparison to illustrate just how big the sample-efficiency gap is. Here’s one. If a person sees and hears, on average, let’s say generously, 2,000 words an hour, then between the time they’re born and the time they’re an adult, they’ll see about 200 million tokens.

Now, by contrast, these frontier models are trained on somewhere between tens to hundreds of trillions of tokens. That is close to a millionfold difference. Here’s another point of comparison. If you wanted to, you could learn to teleoperate any random humanoid or robot arm within hours.

And if we could get AIs to learn just as fast, robotics would be a deca-trillion-dollar industry, and you’d have an endless army of Unitree G1s doing all kinds of useful work in the world. But the reason we can’t do this is that our AIs learn much less efficiently than we do, and even with the millions of hours of demonstrations that we’ve collected, this is not enough to allow them to perform complex, open-ended tasks.

And a final point of comparison: a teenager can learn to drive a car with about 20 hours of practice. And even if we include their 16 years of growing up and understanding how the world works and building physical intuition, that is still 3 to 4 orders of magnitude less data than Waymo and Tesla are using to train their self-driving car models.

Now I want to deal with a couple of common responses and objections that people have to these kinds of comparisons. One thing people will say, and I think Karpathy said this when he came on my podcast, is that for humans, many billions of years of evolution had to go into pretraining us. And so we’re being unfair when we’re comparing how little data we see within our lifetimes to what these cold-started LLMs, which are just starting off with a totally random initialization, have to learn from.

I think this is not the right way to think about it. Our genome is only 3 gigabytes, and only 1% to 2% of it is protein-coding. There is simply not enough space to store the parameters of this network that evolution supposedly pretrained.

I think the closer analogy is that evolution found the right hyperparameters and the right loss functions, and that within our lifetime, we are still building up the connectome in our brain from scratch. That is to say, the thing analogous to the weights and parameters of the neural net itself.

And even if you granted this comparison and said, “Yes, the hundreds of trillions of tokens these models see to get pretrained is similar to just catching up to evolution,” that still doesn’t explain why any new marginal capability that you want to give these models takes so much data. Once you have been educated, again, you don’t need 100 different professors to teach you how to learn a new programming language.

But these AIs, even once they’re pretrained, still require enormous amounts of data to learn the next marginal skill, and the next marginal skill after that. Another objection to this kind of comparison is that we’re not including the multimodal data that we’re seeing in our lifetimes.

So if we include all this sensory information that we see from birth to adulthood, that’s probably tens to hundreds of billions of tokens of data. And my response to this objection is simply that blind or deaf people, who are cut off from parts of this sensory stream, still have general intelligence. That suggests to me that all these billions of sensory tokens are not really the thing that is making humans smart.

In fact, deaf people who communicate through sign language and reading, and not through hearing, are probably ingesting far less than the 200 million language tokens that we ballparked earlier, which suggests that even the millionfold difference that we calculated earlier might be an understatement.

Okay, the 3rd common objection people make is that we just haven’t scaled enough. We have these scaling laws. They tell us that bigger models are more sample-efficient. The human brain, we know, is about 100 trillion synapses, and we have frontier models that are currently around 5 trillion parameters.

So maybe we could just achieve human-level sample efficiency if we made these models 1 to 2 orders of magnitude bigger. The reason this objection is off-mark is actually quite interesting. If you look at the way the scaling-law equations work, they tell you that the parameter and data terms are added to the loss independently.

Suppose you have a model, and you’ve trained it compute-optimally, and you say, “I want to be sample-efficient. I want to use as little data as possible, and I’ll throw in as many parameters as necessary to make that happen.” Take the constants from the Chinchilla scaling-law paper.

Even if you increased the number of parameters to infinity, that would only decrease by a factor of 10 the amount of data that you need in order to keep the same loss. Humans are somewhere between thousands to millions of times more sample-efficient than these models. So scaling the size of current models simply can't make up for that discrepancy, and this really does suggest that humans are on a different scaling curve altogether.

2. Does sample efficiency matter?

Okay, all these nerdy comparisons aside, you might ask: Why do we even care about sample efficiency? Is this actually necessary for the labs to achieve the 2 overarching objectives they have, which are, 1, to automate white-collar work, and 2, to automate AI research itself?

The bet that the labs are making with white-collar work is that the common tasks that a software engineer, analyst, or accountant needs to do are common, and as a result, you can bring them into the training distribution quite easily. If you look at the revenue curves of these labs over the last few months, it does suggest that there's an enormous amount of value from bringing these kinds of common tasks into the distribution, even if we can't replicate whatever is making human learning so special.

It might be more inefficient to train AIs to do these kinds of tasks than it is to train humans, but so what? Human lifespan simply does not allow for the quantity and breadth of training that these models experience. If you, as a human, had some weird learning disability where you needed to read through every public repository on GitHub before you could be a competent software engineer, then it simply wouldn't make sense to train you up. You'd be on Social Security by the early stages of your education, and even once you were trained, you would only be able to work on 1 project at a time.

But AIs can learn these skills by firehosing gigawatts of training at a time, and what they learn can be amortized across billions of sessions at once. So we can be ludicrously inefficient in training them up and still be wildly in the green.

And then there's a question of how much out-of-distribution thinking white-collar employees need to do that you simply can't train for in advance. This is more a question about the nature of different jobs than it is a question about AI research, and it also depends on which job you're talking about. Some jobs are so mechanical and predictable that we were able to automate them long before the modern era of AI, for example, bank tellers or travel agents. But there are other jobs that require dealing on a daily basis with problems that are quite distant from the data distribution.

I think software engineering is probably one such job. This is the job that AIs are supposed to take first, but I would be willing to bet that there's overall more demand for human software engineers in 2028 than there is right now, largely due to the complementary input of AI.

The labs' plan for this latter category of jobs is first to automate AI research and then have the automated AI researchers solve the sample-efficiency problem. So then the question is: Can AIs, which do not have human-level sample efficiency, nonetheless solve the remaining research problems that stand in the way of human-like intelligence and learning?

This is a very complicated question, and I'll have to address it in a much longer future blog post. But just to tease it a bit, I think that the way people currently think about an intelligence explosion is very clumsy, because either people dismiss the possibility of AIs speeding up AI progress altogether, or they assume that some kind of god pops out the other end. They don't reason carefully about what it looks like to have a period where AI progress is much faster than usual, but to have that happen on top of LLMs and the particular kind of intelligence that LLMs are.

But I'll save that for next time.