[BidClub_]
Dwarkesh Podcast · · 12 min

What are we scaling?

Dwarkesh Patel

YouTube
TL;DR
  • The core contradiction Dwarkesh flags: you can't hold super-short timelines while being bullish on scaling up RL from verifiable rewards. An entire supply chain of RL-environment companies is pre-baking browser and Excel skills into models — "either these models will soon learn on the job in a self-directed way which will make all this pre-baking pointless, or they won't, which means AGI is not imminent."
  • Robotics as the tell: it's "an algorithms problem, not a hardware or a data problem" — a human learns to teleoperate current hardware with very little training, so a true humanlike learner would largely solve robotics. Instead labs must practice a million times in a thousand homes, revealing the learner doesn't exist.
  • The revenue gap is the capability gap. Knowledge workers earn tens of trillions a year; labs are orders of magnitude off because "the models are nowhere near as capable as human knowledge workers." The diffusion-lag excuse is "cope" — real humans-on-a-server would diffuse faster than human hires (read your entire Slack and drive within minutes, no lemons market).
  • Signature line worth trading on: "Models keep getting more impressive at the rate that the short-timelines people predict, but more useful at the rate that the long-timelines people predict."
  • RL scaling lacks pre-training's legitimacy: people are "trying to launder the prestige" of the clean pre-training loss trend to justify RLVR bullishness with no public trend behind it. Toby Board (as heard) pieced the O-series data together: "we need something like a million-x scale-up in total RL compute to give a boost similar to a single GPT level."
  • No immediate runaway takeoff expected: continual learning — not a software singularity — is the main driver ahead, and it'll be solved gradually like in-context learning was after GBT3 (2020), taking "another 5 to 10 years to iron out." Breakthroughs get replicated fast; the big three rotate the podium every month or so, and no flywheel has done much to diminish competition.
  • The out-year call: by 2030 labs make real continual-learning progress and earn hundreds of billions in revenue — but haven't automated knowledge work. Still, he expects "actual brain-like intelligences within the next decade or two, which is pretty crazy."
Digest · the substance, structured for research

1. Short timelines and RL pre-baking can't both be right

  • Dwarkesh's opening puzzle: an entire supply chain of RL-environment companies is building environments to teach models how to navigate a web browser and use Excel, plus "billions of dollars paid to PhDs, MDs, and other experts" per Baron Millig's (as heard) blog post — but if a humanlike learner is close, all this pre-baking is doomed. "Either these models will soon learn on the job in a self-directed way... or they won't, which means AGI is not imminent."
  • The robotics proof: fundamentally "an algorithms problem, not a hardware or a data problem" — a human teleoperates current hardware after minimal training, yet labs must go into a thousand homes and "practice a million times on how to pick up dishes."
  • The automated-researcher counter — do kludgy RL to build a superhuman AI researcher, then a million copies of this "automated Ilia" solve learning-from-experience — gets his sharpest dismissal: "we're losing money on every sale, but we'll make it up in volume." And the labs' own behavior offers a clue: you don't need PowerPoint-consultant skills to automate Ilia, so their actions "hint at a worldview where these models will continue to fare poorly at generalization and on-the-job learning."

2. The macrophage crux: jobs run on unbakeable context

  • The dinner anecdote carrying the argument: a biologist with long timelines describes deciding whether a dot on a slide "is actually a macrophage or just looks like a macrophage"; the AI researcher shoots back that image classification is "dead center" textbook deep learning. Dwarkesh's read: it's not net productive to build a custom training pipeline per lab-specific microtask — you need an AI that learns from semantic feedback and generalizes like a human.
  • Every worker does "a hundred things that require judgment, situational awareness, and skills and context learned on the job," differing day to day — "it is not possible to automate even a single job by just baking in a predefined set of skills."

3. Diffusion-lag is cope — the revenue gap measures capability

  • If models were humans on a server, "they'd diffuse incredibly quickly": read your entire Slack and drive within minutes, distill skills from other AI employees, and skip the lemons-market problem of human hiring. So AI labor should diffuse easier than people — the fact it hasn't means "these models just lack the capabilities."
  • The honest yardstick: at AGI, buyers would spend trillions a year on tokens against tens of trillions in knowledge-worker wages; labs being orders of magnitude off is the tell.
  • On goalposts, a concession: some shifting is justified — "if you showed me Gemini 3 in 2020, I would have been certain that it could automate half of knowledge work." Each solved "sufficient bottleneck" (understanding, few-shot, reasoning) reveals "there's much more to intelligence and labor than I previously realized." His 2030 prediction: continual-learning progress, hundreds of billions in revenue, still no full automation.

4. RL from verifiable reward lacks a public scaling trend

  • Pre-training gave a clean loss trend across orders of magnitude — "almost as predictable as a physical law" (albeit a power law, "as weak as exponential growth is strong"). People are "laundering the prestige" of that trend onto RLVR, which has no publicly known scaling trend.
  • When researchers do connect the dots — Toby Board (as heard) across O-series benchmarks — the result is bearish: "something like a million-x scale-up in total RL compute to give a boost similar to a single GPT level."

5. Continual learning arrives gradually — no game-set-match

  • The main driver of future gains isn't a software or software-plus-hardware singularity but continual learning at the top of AGI — humans get capable "mostly from experience." Baron Miller's (as heard) sketch: specialized agents (Karpathi's (as heard) "cognitive core" plus job-specific skills) doing real work, then "bringing back all their learnings to the hive-mind model" via batch distillation.
  • Solving it will feel like solving in-context learning: GBT3 already demonstrated that in-context learning could be very powerful in 2020; the GPT3 paper was titled "Language Models are Few-Shot Learners," yet in-context learning wasn't "solved" when GPD3 came out. Labs will ship something called continual learning next year that counts as progress, but human-level on-the-job learning may take another 5–10 years.
  • Hence no runaway gains expected: a fully-solved drop from nowhere "might be game, set, match" as Satya (likely) put it on the podcast — "but that's probably not what's going to happen." Initial traction gets reverse-engineered and replicated; prior flywheels (chat engagement, synthetic data) "have done very little to diminish the greater and greater competition," with the big three rotating the podium every month or so.
  • The upside he insists people underrate: not more of the current regime, but "billions of humanlike intelligences on a server which can copy and merge all the learnings" — expected "within the next decade or two, which is pretty crazy."
Dwarkesh Patel

I’m confused why some people have super-short timelines yet, at the same time, are bullish on scaling up reinforcement learning on top of LLMs. If we’re actually close to a humanlike learner, then this whole approach of training on verifiable outcomes is doomed.

1. What are we scaling?

Currently, the labs are trying to bake in a bunch of skills into these models through mid-training. There’s an entire supply chain of companies that are building RL environments that teach the model how to navigate a web browser or use Excel to build financial models. Now, either these models will soon learn on the job in a self-directed way, which will make all this pre-baking pointless, or they won’t, which means that AGI is not imminent.

Humans don’t have to go through a special training phase where they need to rehearse every single piece of software that they might ever need to use on the job. Baron Millig made an interesting point about this in a recent blog post he wrote. He writes:

“When we see frontier models improving at various benchmarks, we should think not just about the increased scale and the clever ML research ideas, but the billions of dollars that are paid to PhDs, MDs, and other experts to write questions and provide example answers and reasoning targeting these precise capabilities.

“You can see this tension most vividly in robotics. In some fundamental sense, robotics is an algorithms problem, not a hardware or a data problem. With very little training, a human can learn how to teleoperate current hardware to do useful work. So if you actually had a humanlike learner, robotics would be in large part a solved problem.

“But the fact that we don’t have such a learner makes it necessary to go out into a thousand different homes and practice a million times how to pick up dishes or fold laundry.”

Now, one counterargument I’ve heard from the people who think we’re going to have a takeoff within the next 5 years is that we have to do all this clunky RL in service of building a superhuman AI researcher. Then, a million copies of this automated Ilia can go figure out how to solve robust and efficient learning from experience.

This just gives me the vibes of that old joke: “We’re losing money on every sale, but we’ll make it up in volume.” Somehow, this automated researcher is going to figure out the algorithm for AGI, which is a problem that humans have been banging their heads against for the better part of a century, while not having the basic learning capabilities that children have. I find it super implausible.

Besides, even if that’s what you believe, it doesn’t describe how the labs are approaching reinforcement learning from verifiable reward. You don’t need to pre-bake in a consultant’s skill at crafting PowerPoint slides in order to automate Ilya. So clearly, the labs’ actions hint at a worldview where these models will continue to fare poorly at generalization and on-the-job learning, thus making it necessary to build into these models beforehand the skills that we hope will be economically useful.

Another counterargument you can make is that, even if the model could learn these skills on the job, it is just so much more efficient to build in these skills once during training rather than again for each user and each company. And look, it makes a ton of sense to just bake in fluency with common tools like browsers and terminals.

Indeed, one of the key advantages that AGIs will have is this greater capacity to share knowledge across copies. But people are really underrating how much company- and context-specific skill is required to do most jobs. And there just isn’t currently a robust, efficient way for AIs to pick up these skills.

2. The value of human labor

I was recently at a dinner with an AI researcher and a biologist. It turned out the biologist had long timelines, and so we asked why she had these long timelines. Then she said, “One part of work recently in the lab has involved looking at slides and deciding if the dot in that slide is actually a macrophage or just looks like a macrophage.”

The AI researcher, as you might anticipate, responded, “Look, image classification is a textbook deep-learning problem. This is dead-center in the kind of thing that we could train these models to do.”

I thought this was a very interesting exchange because it illustrated a key crux between me and the people who expect transformative economic impact within the next few years. Human workers are valuable precisely because we don’t need to build in special training loops for every single small part of their job.

It’s not net productive to build a custom training pipeline to identify what macrophages look like, given the specific way that this lab prepares slides, and then another training loop for the next lab-specific microtask, and so on. What you actually need is an AI that can learn from semantic feedback or from self-directed experience and then generalize the way a human does.

Every day, you have to do 100 things that require judgment, situational awareness, and skills and context that are learned on the job. These tasks differ not just across different people but even from one day to the next for the same person. It is not possible to automate even a single job by just baking in a predefined set of skills, let alone all the jobs.

In fact, I think people are really underestimating how big a deal actual AI will be because they are just imagining more of this current regime. They’re not thinking about billions of humanlike intelligences on a server that can copy and merge all the learnings.

3. Economic diffusion lag is cope

And to be clear, I expect this, which is to say I expect actual brain-like intelligences within the next decade or 2, which is pretty fucking crazy. Sometimes people will say that the reason AIs are more widely deployed right now across firms and already providing lots of value outside of coding is that technology takes a long time to diffuse.

I think this is cope. I think people are using this cope to gloss over the fact that these models just lack the capabilities that are necessary for broad economic value. If these models actually were like humans on a server, they’d diffuse incredibly quickly.

In fact, they’d be so much easier to integrate and onboard than a normal human employee is. They could read your entire Slack and drive within minutes. And they could immediately distill all the skills that your other AI employees have.

Plus, the hiring market for humans is very much like a lemons market, where it’s hard to tell who the good people are beforehand. Obviously, hiring somebody who turns out to be bad is very costly. This is just not a dynamic that you would have to face or worry about if you’re simply spinning up another instance of a vetted AI model.

So for these reasons, I expect it’s going to be much easier to diffuse AI labor into firms than it is to hire a person. Companies hire people all the time. If the capabilities were actually at AGI level, people would be willing to spend trillions of dollars a year buying tokens that these models produce.

Knowledge workers across the world cumulatively earn tens of trillions of dollars a year in wages. And the reason that labs are orders of magnitude off this figure right now is that the models are nowhere near as capable as human knowledge workers.

4. Goal-post shifting is justified

Now, you might be like, “Look, how can the standard have suddenly become that labs have to earn tens of trillions of dollars in revenue a year? Until recently, people were saying, ‘Can these models reason? Do these models have common sense? Are they just doing pattern recognition?’”

Obviously, AI bulls are right to criticize AI bears for repeatedly moving these goalposts, and this is very often fair. It’s easy to underestimate the progress that AI has made over the last decade. But some amount of goalpost shifting is actually justified.

If you showed me Gemini 3 in 2020, I would have been certain that it could automate half of knowledge work. So we keep solving what we thought were the sufficient bottlenecks to AGI. We have models that have general understanding. They have few-shot learning. They have reasoning. And yet we still don’t have AGI.

So what is a rational response to observing this? I think it’s totally reasonable to look at this and say, “Oh, actually, there’s much more to intelligence and labor than I previously realized.”

While we’re really close, and in many ways have surpassed what I would have previously defined as AGI, the fact that model companies are not making the trillions of dollars in revenue that would be implied by AGI clearly reveals that my previous definition of AGI was too narrow. And I expect this to keep happening into the future.

I expect that by 2030, the labs will have made significant progress on my hobby horse of continual learning, and the models will be earning hundreds of billions of dollars in revenue a year, but they won’t have automated all knowledge work. And I’ll be like, “Look, we made a lot of progress, but we haven’t hit AGI yet. We also need these other capabilities. We need X, Y, and Z capabilities in these models.”

5. RL scaling

Models keep getting more impressive at the rate that the short-timeline people predict, but more useful at the rate that the long-timeline people predict. It’s worth asking: What are we scaling with pretraining?

We had this extremely clean and general trend in improvement in loss across multiple orders of magnitude in compute. That was on a power law, which is as weak as exponential growth is strong. But people are trying to launder the prestige that pretraining scaling has—which is almost as predictable as a physical law of the universe—to justify bullish predictions about reinforcement learning from verifiable reward, for which we have no publicly known trend.

When intrepid researchers do try to piece together the implications from scarce public data points, they get pretty bearish results. For example, Toby Board has a great post where he cleverly connects the dots between the different o-series benchmarks, and this suggested to him that:

“We need something like a 1,000,000× scale-up in total RL compute to give a boost similar to a single GPT-level model.”

6. Broadly deployed intelligence explosion

End quote. People have spent a lot of time talking about the possibility of a software singularity, where AI models will write the code that generates a smarter successor system, or a software-plus-hardware singularity, where AIs also improve their successor's computing hardware. However, all these scenarios neglect what I think will be the main driver of further improvements atop AGI: continual learning. Again, think about how humans become more capable than anything else. It's mostly from experience in the relevant domain.

In our conversation, Baron Miller made this interesting suggestion: The future might look like continual-learning agents who are all going out, doing different jobs, generating value, and then bringing back all their learnings to the hive-mind model, which does some kind of batch distillation on all of these agents. The agents themselves could be quite specialized, containing what Karpathi called the cognitive core, plus knowledge and skills relevant to the job they're being deployed to do.

Solving continual learning won't be a singular, one-and-done achievement. Instead, it will feel like solving in-context learning. GBT3 already demonstrated that in-context learning could be very powerful in 2020. Its in-context-learning capabilities were so remarkable that the title of the GPT3 paper was “Language Models are Few-Shot Learners.” But of course, we didn't solve in-context learning when GPD3 came out. Indeed, there's still plenty of progress that has to be made, from comprehension to context length.

I expect a similar progression with continual learning. Labs will probably release something next year which they call continual learning and which will, in fact, count as progress toward continual learning. But human-level, on-the-job learning may take another 5 to 10 years to iron out. This is why I don't expect runaway gains from the first model that cracks continual learning and gets more and more widely deployed and capable.

If you had fully solved continual learning and it dropped out of nowhere, then sure, it might be game, set, match, as [Speaker?] put it on the podcast when I asked him about this possibility. But that's probably not what's going to happen. Instead, some lab is going to figure out how to get some initial traction on this problem, and then playing around with this feature will make it clear how it was implemented. Other labs will soon replicate the breakthrough and improve it slightly.

Besides, I just have some prior that the competition will stay pretty fierce between all these model companies. This is informed by the observation that all these previous supposed flywheels, whether that's user engagement on ChatGPT or synthetic data or whatever, have done very little to diminish the greater and greater competition between model companies. Every month or so, the big 3 model companies will rotate around the podium, and the other competitors are not that far behind. There seems to be some force—potentially talent poaching, the rumor mill, or just normal reverse engineering—that has so far neutralized any runaway advantage that a single lab might have had.

What are we scaling? | BidClub