Dwarkesh PodcastAI researchers debate how close we are to recursive self-improvement
At a glance
WHAT IT’S REALLY ABOUT
Recursive self-improvement hinges on objectives, realism, and continual learning
- The guests steelman why recursive self-improvement might not happen soon: models could keep winning benchmarks while remaining bottlenecked by weak judgment, verification, continual learning, and sim-to-real transfer.
- They debate whether today’s transformer-plus-RL paradigm is near the “global optimum” or whether RSI would require a discontinuous paradigm shift that scaled-up current methods might fail to discover.
- They argue the hardest-to-automate part of research is not executing experiments but choosing objectives and asking the right questions—linking this to alignment and to the persistent need for human taste/specification.
- They explain Chinese labs’ rapid progress through distillation dynamics and access to realistic prompt distributions (e.g., router/proxy data), while noting that realistic interactive environments may be harder to copy than benchmark performance.
- They outline a pragmatic path to automated AI researchers: iterative post-training with human feedback plus increasingly realistic, long-horizon RL environments, constrained by verification bottlenecks, sample efficiency limits, and continual-learning failures like forgetting.
IDEAS WORTH REMEMBERING
5 ideasThe most plausible technical failure mode is persistent bottlenecks in generalization, judgment, and sim-to-real.
The panel’s “non-exogenous” reason for no 2036 RSI is that systems keep excelling on benchmarks while failing to translate to messy real-world competence—especially where self-checking, judgment, continual learning, and sim-to-real transfer are required.
RSI may require discontinuities that current methods can’t discover from within the paradigm.
Even with superhuman research agents, progress may stall if major capability gains require a new paradigm that is too far from today’s “transformer + scaling + RL” recipe for gradient-based search to find automatically.
Objective specification (and alignment-adjacent taste) may be the last human moat.
A recurring theme is that AIs can optimize well-specified goals, but struggle with the upstream step: selecting/creating the right objectives, tasks, and rubrics—what the guests call “taste” and specification.
Distillation is a strong decentralizing force, but prompt distribution and realism are the true scarce resources.
The guests argue distillation counteracts winner-take-all dynamics because behaviors learned via RL can often be copied from trajectories—yet doing this well depends heavily on having the right prompt distribution and realistic interaction traces.
Automated AI researchers will be built via continual patching with HF + evolving RL environments, not a single end-to-end leap.
They describe a likely training path to automated AI R&D as iterative patching: mix human feedback to absorb researchers’ taste with increasingly realistic multi-step research environments, “diffing” what broke last generation and turning it into new training tasks.
WORDS WORTH SAVING
5 quotesSo there's this, uh, cycle that keeps repeating where people think, uh, where a new model comes out and people are blown away and they're like, "This is it. This is the-- This is AGI." But then, uh, they use it a bit and, and then it starts to feel dumb after a month or so.
— John Schulman
I remember in the early OpenAI days, uh, having the int-intuition that, um, actually just, uh, do, like, minimizing log loss wasn't gonna get you- ... to intelligence... But then it turned out that it- ... it just worked anyway.
— John Schulman
I would say that the last, um, job for humans, uh, or, or the, the, the role for humans that'll last the longest is, like, defining the objective, uh, and, like, dec- deciding what we actually want.
— John Schulman
I think the success of RL comes down to a bunch of different things... an awful lot of, like, what we see as successes of RL actually comes from, like, very, very good mid-training data... And then what RL does on top of that is it does, like, a lot of, you know, essentially tweaking to the policy.
— Beren Millidge
So, like, the signal just like doesn't exist anywhere in the original data we have. Like no amount of filtering will like get this, this, you know, there's no like hidden proof of like the Millennium Prize Problem sitting in Common Crawl.
— Beren Millidge
High quality AI-generated summary created from speaker-labeled transcript.