Skip to content
Dwarkesh PodcastDwarkesh Podcast

Terence Tao on Dwarkesh Patel: How Erdős Problems Exposed AI

Tycho Brahe data let Kepler derive orbital laws by regression on six points; AI solved 50 Erdős problems fast then stalled on cumulative partial progress.

Dwarkesh PatelhostTerence Taoguest
Mar 20, 20261h 23mWatch on YouTube ↗

CHAPTERS

  1. 0:00 – 4:09

    Kepler’s discovery: from Platonic-solid aesthetics to data-driven orbital laws

    Tao retells how Kepler moved from a beautiful but wrong geometric theory to empirically grounded laws of planetary motion using Tycho Brahe’s unprecedented observations. The story sets up the modern tension between abundant hypothesis generation and the need for rigorous verification.

    • Copernicus’s heliocentrism with circular orbits as Kepler’s starting point
    • Kepler’s early (incorrect) Platonic-solid model and the role of aesthetic priors
    • Tycho Brahe’s high-precision dataset as the enabling constraint for real progress
    • Kepler’s iterative data analysis leading to elliptical orbits and the area law
    • Newton later providing a unifying explanatory theory for Kepler’s empirical laws
  2. 4:09 – 9:51

    “Kepler as a high-temperature LLM”: cheap idea generation vs. costly verification

    Dwarkesh frames Kepler as an engine for trying many strange hypotheses against a trusted dataset—analogous to LLM-style exploration. Tao emphasizes that without strong validation, idea generation collapses into slop, and that modern science increasingly flips from ‘hypothesis then test’ to ‘data then hypotheses.’

    • LLMs resemble rapid, high-variance hypothesis generators when verification is available
    • Verification is the real counterweight that makes exploration productive rather than “slop”
    • Science now often starts from big datasets and extracts patterns rather than testing one idea
    • Kepler as an early “data scientist,” though still guided by priors
    • Questioning whether hypothesis generation is even the bottleneck today
  3. 9:51 – 11:44

    When regression misleads: small data, Bode’s law, and tentative empirical “discoveries”

    Tao contrasts Kepler’s lucky six-point regression with a famous near-miss: Bode’s law, which looked predictive until Neptune broke it. The segment highlights how easy it is to overfit patterns and why scientific conclusions must be calibrated to data quantity and reliability.

    • Kepler’s third law as a regression-like fit on only ~6 data points
    • Bode’s law: a compelling fit that seemed confirmed by later discoveries (Uranus, Ceres)
    • Neptune as the falsifier revealing a numerical fluke
    • Why Kepler may have treated the third law more cautiously than the first two
    • Lesson: pattern-finding scales faster than trustworthy inference
  4. 11:44 – 14:13

    The new bottleneck: sorting real breakthroughs from AI-generated scientific slop

    Tao argues AI has driven the marginal cost of idea generation close to zero, shifting the bottleneck to evaluation and filtering. Peer review and existing institutions weren’t built for floods of plausible-looking content, so science needs new scalable mechanisms for triage and validation.

    • AI as a step-change like the internet: abundant generation doesn’t ensure value
    • Verification/validation and prioritization become the limiting factors
    • Existing filters (peer review, journals) are already overwhelmed by AI submissions
    • Hard problem: assessing partial progress and dead ends at massive scale
    • Need for new scientific “structures” to evaluate and curate ideas
  5. 14:13 – 17:31

    How unifying concepts are recognized (late), and why context and culture matter

    Dwarkesh asks how we’d detect a “bit”-level unifying idea amid mountains of mediocre work. Tao responds that recognition is often retrospective: adoption depends on community context, path dependence, and standardization—making objective scoring difficult to automate.

    • Many foundational ideas are underappreciated initially and recognized only over time
    • Alternative paradigms (trits, other architectures) could have won in different worlds
    • Path dependence: standards (decimal system, transformers) persist via inertia
    • Scientific value can’t be graded in isolation from past/future context
    • This resists straightforward reinforcement-learning style evaluation
  6. 17:31 – 21:35

    Progress that initially looks worse: from Copernicus to Darwin and today’s AI “Copernican shift”

    They discuss why early versions of correct theories can appear inferior (less accurate, incomplete, or philosophically troubling) compared to polished wrong ones. Tao draws an analogy to today’s reframing of intelligence as we encounter non-human systems with different strengths.

    • Copernicus initially underperformed Ptolemy’s tuned geocentric model
    • Newton’s gravity raised deep conceptual issues later reframed by Einstein
    • Major advances can come from deleting assumptions (e.g., Aristotelian rest)
    • Darwin’s evolution as a conceptual leap despite limited direct observation
    • AI forces a reordering of what tasks “require intelligence”
  7. 21:35 – 26:09

    Darwin vs. Newton: evidence loops, persuasion, and the social machinery of science

    Dwarkesh wonders why Darwin’s seemingly simpler idea arrived much later than Newton’s. Tao stresses communication: Darwin’s clear narrative accelerated acceptance, while Newton’s Latin, new math, secrecy, and personality slowed diffusion; persuasion and exposition are central and hard to formalize.

    • Darwin as an exceptional communicator who synthesized disparate evidence
    • Newton’s barriers: Latin, invented machinery, secrecy, competitive culture
    • Science requires narrative and persuasion beyond data and equations
    • Credibility and uptake depend on social processes, not just correctness
    • Why “persuasiveness” is difficult to quantify for automated optimization
  8. 26:09 – 30:30

    The deductive overhang and “signal extraction”: astronomy mindset and meta-metrics for science

    Tao explains why astronomy developed extreme skill at extracting maximal information from scarce data, and how that mindset generalizes. He then gestures at using sociological traces (citations, conference mentions, typo propagation) to quantify attention and impact—possible tools for AI-era filtering.

    • Astronomy’s culture of squeezing information because data is expensive
    • Anecdote: finance hiring astronomy PhDs for signal extraction skills
    • Clever inference from metadata (citation typos) to measure real engagement
    • Potential for ‘sociology of science’ metrics to detect fruitful ideas early
    • Open question: turning such signals into robust scalable evaluation
  9. 30:30 – 34:19

    Erdős problems and the “jumping robot” model: breadth wins, but partial progress is hard

    They examine the burst of AI-assisted solutions to Erdős problems and why progress plateaued after low-hanging fruit. Tao’s metaphor: AIs can ‘jump’ over small walls but struggle to create intermediate footholds, making them poor at cumulative partial progress despite strong breadth.

    • ~50 Erdős problems solved with AI assistance; progress now slower
    • One-shot ‘pure AI’ successes decreased despite multiple large-scale attempts
    • Humans+multiple AI tools can solve via iterative collaboration, not one-shotting
    • Metaphor: AIs jump; humans climb with intermediate markers and decomposition
    • Core limitation: evaluating and generating partial progress/waypoints
  10. 34:19 – 36:56

    Redesigning science for AI breadth: mapping easy terrain, then targeting islands of difficulty

    Dwarkesh highlights a bullish implication: once AIs reach a competence ‘waterline,’ they can cover everything at that level in parallel. Tao agrees that breadth and depth are complementary and argues science should be reorganized to exploit broad exploration—using AIs to clear easy observations and flag hard regions for human focus.

    • AIs scale breadth massively; humans (experts) still dominate depth
    • Current scientific workflow is optimized for human depth; it must adapt
    • Use AIs to map fields, clear easy results, and identify hard subproblems
    • Expect a new style of science built around large-scale exploration
    • Long-run goal: systems that integrate breadth and depth seamlessly
  11. 36:56 – 49:19

    Why AI enriches papers but not depth: richer outputs, experimental math, and ‘existing techniques’ limits

    Tao describes how AI changes his work by accelerating secondary tasks—plots, code, formatting, literature search—making papers broader but not necessarily deeper. They discuss how much of math progress is ‘apply known tools broadly’ versus inventing new techniques, and why reported wins suffer selection bias.

    • AI accelerates auxiliary tasks, enabling more code/figures/numerics in papers
    • Depth-critical steps remain largely pen-and-paper for Tao
    • Math at scale: AI enables systematic ‘experimental mathematics’ across many problems
    • Many AI successes come from combining existing obscure techniques on neglected problems
    • Selection bias: systematic sweeps show low per-problem success rates despite viral wins
  12. 49:19 – 52:54

    Artificial cleverness vs. intelligence: adaptivity, cumulative learning, and session-level amnesia

    Tao distinguishes clever trial-and-error from the adaptive, cumulative refinement seen in human collaboration. Current models can imitate dialog but don’t truly build durable internal progress across attempts; they often lack reliable decomposition into intermediate goals and don’t retain new skills within a session-to-session workflow.

    • Human problem-solving involves iterative refinement and mapping what fails/works
    • Current AIs rely more on brute-force trial than cumulative strategy-building
    • They struggle to create useful intermediate milestones and partial achievements
    • Solving a problem doesn’t meaningfully update the model’s understanding in situ
    • At best, outputs become tiny training signals for later model generations
  13. 52:54 – 59:19

    If AI proves big theorems in Lean: will we get understanding, and how do we extract it?

    Dwarkesh asks whether an AI proof (even of something like RH) could be incomprehensible ‘assembly code,’ yielding little insight. Tao notes precedents like the four-color theorem and argues that once a formal artifact exists, we can analyze, refactor, summarize, and ablate it—potentially creating new roles focused on proof interpretation and elegance.

    • Some theorems may only admit brute-force or case-split proofs (four-color)
    • RH is expected (but not guaranteed) to require new conceptual connections
    • Formal proofs are ‘atomically inspectable’: lemmas can be evaluated in isolation
    • Future workflows: ablation/refactoring/elegance optimization as post-processing
    • Having one proof may be enough to recover human-usable understanding afterward
  14. 59:19 – 1:13:06

    Beyond proofs: the need for a semi-formal language of strategies, plausibility, and scientific talk

    Tao argues Lean formalizes deduction well, but science relies heavily on heuristic plausibility, strategy selection, and narrative—areas with subjectivity and vulnerability to “reward hacking.” He sketches the prime-number heuristics (Gauss’s data-driven conjectures, random model of primes, RH belief) as an example of reasoning that is powerful yet not fully formalized.

    • Lean captures formal deduction; it doesn’t capture conjecture-making and strategy
    • We lack robust, non-hackable frameworks for ‘plausibility’ and heuristic reasoning
    • Gauss’s prime number theorem conjecture as early data-driven statistical insight
    • Random-model heuristics underpin belief in twin primes, RH, and crypto assumptions
    • Possible path: study many ‘mini-universes’/small AIs to learn formalizable strategy principles
  15. 1:13:06 – 1:17:05

    How Tao allocates time: serendipity vs. optimization, and the value of productive randomness

    Dwarkesh probes how Tao’s time would be used under ideal optimization; Tao argues over-optimization can kill serendipity and inspiration. He describes how distractions, unplanned encounters, and browsing can create valuable randomness—something remote/scheduled workflows and hyper-efficient search can reduce.

    • Senior responsibilities add obligations, but they also create unexpected opportunities
    • Serendipity from in-person chance encounters vs. fully scheduled remote life
    • Efficient search loses ‘adjacent discovery’ (library browsing effect)
    • Too little distraction can reduce inspiration even in ideal research environments
    • Balance: deliberate scheduling plus room for randomness and exploration
  16. 1:17:05 – 1:23:44

    Will AI ‘replace’ top mathematicians? Hybrids for a long time, and advice for newcomers

    Tao predicts many routine mathematical tasks will be automated within a decade, but that doesn’t equate to full replacement—math will shift to new scales and activities. He expects human–AI hybrids to dominate for a long time and advises early-career mathematicians to stay adaptable, pursue curiosity, and exploit new non-traditional entry points enabled by AI tools.

    • AI already surpasses humans on some ‘frontiers’ (like calculators did), but not the same ones
    • Automation will cover much of what appears in today’s papers; the ‘important part’ may shift
    • Full autonomy requires additional breakthroughs; timelines are stochastic and uncertain
    • Hybrid human+AI workflows likely dominate for a long time
    • Career advice: embrace change, stay adaptable, explore new pathways to contribute early

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.