Dwarkesh PodcastShane Legg (DeepMind Founder) — 2028 AGI, superhuman alignment, new architectures
CHAPTERS
- 0:00 – 2:11
Defining AGI and why measuring it is inherently hard
Shane Legg explains that AGI is about broad generality, which makes it difficult to capture with any single metric like loss. He proposes that AGI assessment must involve many tasks spanning human cognitive abilities, plus attempts to find remaining gaps.
- •AGI as the ability to do the kinds of cognitive tasks humans can do (and possibly more)
- •Single-number metrics (e.g., loss) don’t directly translate to “how close to AGI”
- •Need a broad test suite across many cognitive domains
- •Practical AGI criterion: when adversarial gap-finding fails to uncover human advantages
- 2:11 – 3:53
What today’s benchmarks miss: video understanding, episodic memory, and rapid learning
Dwarkesh asks what popular LLM benchmarks fail to measure. Legg highlights missing modalities (like streaming video) and missing memory systems—especially episodic memory that supports rapid retention of specific experiences.
- •Current LLM benchmarks under-test multimodal skills such as streaming video understanding
- •Humans have multiple memory systems: working memory, cortical memory, and episodic (hippocampal) memory
- •Episodic memory enables rapid learning of specific facts/events (e.g., remembering a conversation tomorrow)
- •Longer context windows are a partial substitute but don’t replicate true episodic memory
- 3:53 – 6:08
Is LLM sample inefficiency a fatal flaw? Why Legg thinks the blockers are solvable
The conversation turns to whether needing trillions of tokens is a fundamental limitation. Legg argues it’s not; modern foundation models unlocked scalable “understanding,” and the remaining shortcomings (memory, factuality, video, etc.) look addressable with further research.
- •LLMs learn fast in-context and slowly in weights, but miss a “middle” rapid-learning mechanism
- •Legg sees no hard wall—most shortcomings have plausible research paths forward
- •Examples of solvable gaps: hallucinations/factuality, episodic memory, better video understanding
- •Scalability changed the game: we can now build models with meaningful world understanding
- 6:08 – 7:16
What would count as “human-level”: comprehensive suites and adversarial testing
Dwarkesh presses for concrete criteria (Minecraft, perfect MMLU, etc.). Legg reiterates there’s no single decisive task; instead you need a comprehensive battery plus an adversarial search for failures, declaring success when meaningful gaps are hard to find.
- •No single benchmark defines AGI because generality is the point
- •Need a wide suite covering diverse human cognitive competencies
- •Adversarial evaluation: actively seek tasks humans do that the system cannot
- •AGI “for practical purposes” when gap-finding efforts fail
- 7:16 – 11:41
From universal intelligence theory to human-centered reference points
Legg reflects on his thesis-era work defining universal intelligence mathematically. He explains the ‘reference machine’ free parameter in Kolmogorov complexity and why he now views human intelligence and our environment as a natural practical reference point.
- •Universal intelligence framing: performance breadth across many computable environments
- •Kolmogorov complexity depends on a ‘reference machine,’ creating an unresolved free parameter
- •No single privileged universal Turing machine; choices change task weightings
- •Human intelligence is a meaningful reference: known to exist, powerful, and economically transformative
- 11:41 – 14:43
Why memory likely requires new architecture (and why AlphaFold isn’t ‘on the AGI path’)
Dwarkesh asks whether missing capabilities need architectural changes. Legg argues episodic memory and rapid learning likely require explicit architectural mechanisms distinct from slow weight updates, and notes that many DeepMind domain projects (like AlphaFold) aren’t direct steps toward AGI.
- •Current models have ‘working memory’ (context/activations) and slow ‘cortical’ learning (weights) but lack a fast episodic layer
- •Brain separates rapid specific learning from slow general learning due to different optimization targets
- •Episodic memory is likely an architectural add-on rather than mere scaling
- •AlphaFold and similar domain successes may teach lessons but aren’t direct AGI building blocks
- 14:43 – 16:25
Compression → prediction → agency: why LLMs fit older AGI intuitions
Dwarkesh notes Legg’s early idea that compression/prediction can measure intelligence. Legg connects modern foundation models to Solomonoff induction ideas: strong sequence prediction plus added search and reinforcement can produce general agents.
- •Legg’s thesis aligns with today’s LLM paradigm: powerful sequence prediction as “world compression”
- •AIXI framing: prediction + search + reinforcement yields a general agent (theoretical ideal)
- •Foundation models approximate strong predictors; agency can be layered on top
- •Today’s shift: scalable predictors make the next step toward agents plausible
- 16:25 – 19:19
Is search required for real creativity? AlphaGo’s Move 37 as the template
They discuss Sutton’s “Bitter Lesson” and the role of search. Legg argues true creativity requires searching a space of possibilities to find ‘hidden gems,’ not merely blending patterns from training data—citing AlphaGo’s Move 37 as search-driven novelty.
- •Foundation models act like world models, but creativity needs search/planning over possibilities
- •Move 37 didn’t come from imitation; it emerged from exploring unlikely but plausible actions via search
- •LLMs mainly mimic and remix human-generated internet ingenuity
- •To go beyond training data in a deep way, powerful search must be integrated
- 19:19 – 23:52
Superhuman alignment: from ‘System 1’ outputs to ‘System 2’ deliberation
Legg outlines an alignment vision based on deliberation: systems shouldn’t ‘blurt’ first responses, but instead reason step-by-step about consequences and ethics. He frames RLHF/constitutional approaches as shifting a high-dimensional distribution, but not fully robust without a System 2 reasoning layer.
- •Containment/limitation won’t work long-term for highly capable systems; alignment is necessary
- •Current sampling resembles Kahneman ‘System 1’; alignment needs ‘System 2’ deliberation
- •Robust alignment requires: strong world model, strong ethics understanding, reliable reasoning
- •RLHF/constitutional methods help but may be brittle as distribution-shaping in high dimensions
- 23:52 – 29:26
Values and enforcement: specifying ethics, auditing reasoning, and avoiding deceptive training
Dwarkesh challenges how to ensure the system truly cares about the right things. Legg emphasizes training deep understanding of ethics, social choice about which principles to adopt, and ongoing verification via process-level scrutiny—preferring reasoning audits over naive reinforcement that could incentivize deception.
- •Train broad ethics competence (like a top ethicist), then choose which principles to follow (a societal question)
- •Engineering goal: each decision includes explicit ethical analysis using world understanding + reasoning
- •Need to check reasoning processes and decisions continuously, not just outcomes
- •Reinforcement alone risks training deception; process supervision and scrutiny may be more robust
- 29:26 – 29:57
Capability-based safety benchmarks and ‘pause’ frameworks: sensible but difficult
Dwarkesh asks whether there should be concrete safety requirements tied to capability levels. Legg agrees it’s sensible in principle and notes efforts like Anthropic’s, but stresses that specifying actionable thresholds and tests is genuinely hard.
- •Milestone-based safety benchmarks are a reasonable governance idea
- •Hard part: making benchmarks concrete, measurable, and enforceable
- •Industry examples exist (e.g., public “responsible scaling” style proposals)
- •Encourages more work on operationalizing these frameworks
- 29:57 – 34:02
DeepMind’s net effect on safety vs capabilities: counterfactual uncertainty
Dwarkesh asks whether DeepMind accelerated safety or capabilities overall. Legg says it’s hard to judge, describing early difficulty hiring for safety, DeepMind’s role in legitimizing AGI safety, and the uncertainty around what would have happened anyway without DeepMind’s contributions.
- •Early years: AGI safety work was hard to staff and seen as career-risky
- •DeepMind maintained an AGI safety group and published in the area, lending credibility
- •DeepMind did accelerate capabilities (e.g., AlphaGo), but counterfactuals are unclear
- •Many breakthroughs are “in the air” when the time is right, reducing attribution certainty
- 34:02 – 37:38
Timelines to AGI: why Legg still sees ~2028 as a plausible 50% point
Legg revisits his 2009 forecast (modal 2025, EV 2028) and explains the underlying logic: exponential growth in compute and data plus positive feedback loops incentivize scalable algorithms. He maintains a ~50% chance by 2028, while acknowledging uncertainty and potential research surprises.
- •Reasoning rooted in Kurzweil-era expectations: compute and data growth over decades
- •Scalable algorithms become increasingly valuable as compute/data scale
- •Positive feedback loop: better algorithms raise the value of compute/data, driving more investment
- •Claim is probabilistic (50% by ~2028), with explicit openness to delays
- 37:38 – 41:25
What progress looks like on the way: better factuality, less hallucination, and multimodal maturity
Dwarkesh asks what the world looks like if AGI arrives by 2028 and what intermediate progress will be. Legg predicts maturation of current models: improved factuality, recency, broader multimodality, and an explosion of practical applications—alongside some misuse.
- •Near-term evolution: reduced delusions/hallucinations and stronger factuality
- •Models become more up-to-date and useful in real-world workflows
- •Multimodality grows (images/video), enabling more grounded understanding
- •Expect many beneficial applications, while acknowledging potential misuse cases
- 41:25 – 44:18
Next landmark: multimodality as the major shift (and why it hasn’t ‘percolated’ yet)
Legg argues the next remembered milestone will be models that natively understand and integrate text, images, and especially video. He explains why current multimodal tools feel early: deeper grounding and new applications emerge once models digest large-scale video and richer modalities.
- •Future retrospection: text-only chat will feel narrow compared to fully multimodal systems
- •Video understanding is key to grounding models in the physical/social world
- •Early days: multimodal features exist but haven’t unlocked their full application space
- •New modalities expand training data sources and enable applications we can’t yet predict