Dwarkesh Podcast · · 12 min
What are we scaling?
TL;DR
- The core contradiction Dwarkesh flags: you can't hold super-short timelines while being bullish on scaling up RL from verifiable rewards. An entire supply chain of RL-environment companies is pre-baking browser and Excel skills into models — "either these models will soon learn on the job in a self-directed way which will make all this pre-baking pointless, or they won't, which means AGI is not imminent."
- Robotics as the tell: it's "an algorithms problem, not a hardware or a data problem" — a human learns to teleoperate current hardware with very little training, so a true humanlike learner would largely solve robotics. Instead labs must practice a million times in a thousand homes, revealing the learner doesn't exist.
- The revenue gap is the capability gap. Knowledge workers earn tens of trillions a year; labs are orders of magnitude off because "the models are nowhere near as capable as human knowledge workers." The diffusion-lag excuse is "cope" — real humans-on-a-server would diffuse faster than human hires (read your entire Slack and drive within minutes, no lemons market).
- Signature line worth trading on: "Models keep getting more impressive at the rate that the short-timelines people predict, but more useful at the rate that the long-timelines people predict."
- RL scaling lacks pre-training's legitimacy: people are "trying to launder the prestige" of the clean pre-training loss trend to justify RLVR bullishness with no public trend behind it. Toby Board (as heard) pieced the O-series data together: "we need something like a million-x scale-up in total RL compute to give a boost similar to a single GPT level."
- No immediate runaway takeoff expected: continual learning — not a software singularity — is the main driver ahead, and it'll be solved gradually like in-context learning was after GBT3 (2020), taking "another 5 to 10 years to iron out." Breakthroughs get replicated fast; the big three rotate the podium every month or so, and no flywheel has done much to diminish competition.
- The out-year call: by 2030 labs make real continual-learning progress and earn hundreds of billions in revenue — but haven't automated knowledge work. Still, he expects "actual brain-like intelligences within the next decade or two, which is pretty crazy."
Digest · the substance, structured for research
1. Short timelines and RL pre-baking can't both be right
- Dwarkesh's opening puzzle: an entire supply chain of RL-environment companies is building environments to teach models how to navigate a web browser and use Excel, plus "billions of dollars paid to PhDs, MDs, and other experts" per Baron Millig's (as heard) blog post — but if a humanlike learner is close, all this pre-baking is doomed. "Either these models will soon learn on the job in a self-directed way... or they won't, which means AGI is not imminent."
- The robotics proof: fundamentally "an algorithms problem, not a hardware or a data problem" — a human teleoperates current hardware after minimal training, yet labs must go into a thousand homes and "practice a million times on how to pick up dishes."
- The automated-researcher counter — do kludgy RL to build a superhuman AI researcher, then a million copies of this "automated Ilia" solve learning-from-experience — gets his sharpest dismissal: "we're losing money on every sale, but we'll make it up in volume." And the labs' own behavior offers a clue: you don't need PowerPoint-consultant skills to automate Ilia, so their actions "hint at a worldview where these models will continue to fare poorly at generalization and on-the-job learning."
2. The macrophage crux: jobs run on unbakeable context
- The dinner anecdote carrying the argument: a biologist with long timelines describes deciding whether a dot on a slide "is actually a macrophage or just looks like a macrophage"; the AI researcher shoots back that image classification is "dead center" textbook deep learning. Dwarkesh's read: it's not net productive to build a custom training pipeline per lab-specific microtask — you need an AI that learns from semantic feedback and generalizes like a human.
- Every worker does "a hundred things that require judgment, situational awareness, and skills and context learned on the job," differing day to day — "it is not possible to automate even a single job by just baking in a predefined set of skills."
3. Diffusion-lag is cope — the revenue gap measures capability
- If models were humans on a server, "they'd diffuse incredibly quickly": read your entire Slack and drive within minutes, distill skills from other AI employees, and skip the lemons-market problem of human hiring. So AI labor should diffuse easier than people — the fact it hasn't means "these models just lack the capabilities."
- The honest yardstick: at AGI, buyers would spend trillions a year on tokens against tens of trillions in knowledge-worker wages; labs being orders of magnitude off is the tell.
- On goalposts, a concession: some shifting is justified — "if you showed me Gemini 3 in 2020, I would have been certain that it could automate half of knowledge work." Each solved "sufficient bottleneck" (understanding, few-shot, reasoning) reveals "there's much more to intelligence and labor than I previously realized." His 2030 prediction: continual-learning progress, hundreds of billions in revenue, still no full automation.
4. RL from verifiable reward lacks a public scaling trend
- Pre-training gave a clean loss trend across orders of magnitude — "almost as predictable as a physical law" (albeit a power law, "as weak as exponential growth is strong"). People are "laundering the prestige" of that trend onto RLVR, which has no publicly known scaling trend.
- When researchers do connect the dots — Toby Board (as heard) across O-series benchmarks — the result is bearish: "something like a million-x scale-up in total RL compute to give a boost similar to a single GPT level."
5. Continual learning arrives gradually — no game-set-match
- The main driver of future gains isn't a software or software-plus-hardware singularity but continual learning at the top of AGI — humans get capable "mostly from experience." Baron Miller's (as heard) sketch: specialized agents (Karpathi's (as heard) "cognitive core" plus job-specific skills) doing real work, then "bringing back all their learnings to the hive-mind model" via batch distillation.
- Solving it will feel like solving in-context learning: GBT3 already demonstrated that in-context learning could be very powerful in 2020; the GPT3 paper was titled "Language Models are Few-Shot Learners," yet in-context learning wasn't "solved" when GPD3 came out. Labs will ship something called continual learning next year that counts as progress, but human-level on-the-job learning may take another 5–10 years.
- Hence no runaway gains expected: a fully-solved drop from nowhere "might be game, set, match" as Satya (likely) put it on the podcast — "but that's probably not what's going to happen." Initial traction gets reverse-engineered and replicated; prior flywheels (chat engagement, synthetic data) "have done very little to diminish the greater and greater competition," with the big three rotating the podium every month or so.
- The upside he insists people underrate: not more of the current regime, but "billions of humanlike intelligences on a server which can copy and merge all the learnings" — expected "within the next decade or two, which is pretty crazy."