Skip to content
OpenAIOpenAI

What happens now that AI is good at math? — the OpenAI Podcast Ep. 17

Math is one of the clearest ways to see how far AI has come in a short span. OpenAI researchers Sébastien Bubeck and Ernest Ryu join host Andrew Mayne to explain what changed and what it could mean for the future of research. They reflect on how Ernest used ChatGPT to help solve a 42-year-old open problem, the difference between deep literature search and original mathematical discovery, and what changes when AI can work over longer timelines. Chapters 01:27 The surprising progress of AI’s math capabilities 03:01 Solving an open problem with ChatGPT 06:57 How models went from basic math to research level 11:32 Why math matters for AGI 14:26 AI and the Erdős problems 21:26 Building an automated researcher 28:19 The role of humans as models improve 33:52 Verifying proofs with AI 36:00 The risk of shallow understanding 41:19 Advice for learning math with ChatGPT

Andrew MaynehostSébastien BubeckguestErnest Ryuguest
Apr 28, 202643mWatch on YouTube ↗

CHAPTERS

  1. 0:00 – 0:34

    AI math leap: from “can’t schedule a meeting” to Olympiad and beyond

    The hosts set the stage: math performance has become a surprising headline for language models. The guests frame recent progress as unusually rapid—moving from weak arithmetic and brittle reasoning to systems that can meaningfully assist expert mathematicians.

    • Guests Sébastien Bubeck and Ernest Ryu introduce the episode’s theme: AI’s sudden competence in math
    • Progress characterized as faster than most researchers expected
    • Math as a lens for understanding reasoning and AGI
    • Early claim: models now help working mathematicians, not just students
  2. 0:34 – 1:26

    Who the guests are: optimization, ML theory, and why they care about math+AI

    Sébastien and Ernest describe their backgrounds in mathematics, optimization, and machine learning theory, and how they arrived at OpenAI. Their roles center on evaluating and pushing AI systems toward research-level math capability and scientific usefulness.

    • Bubeck’s path: academia (Princeton) → industry → OpenAI; focus on AI for difficult math
    • Ryu’s path: applied math professor (UCLA) → OpenAI; focus on optimization and theory
    • Why mathematicians are well-positioned to evaluate reasoning progress
    • Math as both a personal craft and a benchmark for AI progress
  3. 1:26 – 3:01

    What changed in LLM math ability—and why it wasn’t just “scaling” or tools

    They unpack the common misconception that language models ‘aren’t for math’ and contrast it with what reasoning models can now do. Bubeck argues the breakthrough can’t be reduced to scaling alone: multiple research advances compounded, and the timeline has been easy to underestimate.

    • Two years ago: no true reasoning models; now: substantive theorem/proof assistance
    • Debate recap: researchers underestimated how quickly research-level math would arrive
    • Progress attributed to multiple innovations, not one trick (and not only tool use)
    • Historical perspective: earlier milestones like Minerva once seemed shocking
  4. 3:01 – 6:21

    Case study: ChatGPT helps solve a 42-year-old open optimization problem

    Ernest recounts testing ChatGPT on a genuine open question about divergence in Nesterov’s accelerated gradient method. Over roughly 12 focused hours across three days, he guided, corrected, and verified the model’s reasoning until a correct proof emerged—then shared it publicly for feedback.

    • Motivation: beyond contest math—can AI do research-level work?
    • Problem: existence of a divergent worst-case for Nesterov accelerated gradient
    • Human role: Ryu as verifier and steering agent, not “one prompt, one proof”
    • Outcome: correct proof found; social media used as an informal review catalyst
  5. 6:21 – 11:16

    From everyday arithmetic failures to “99% of people’s math is covered”

    They calibrate the progress using practical examples—expense splitting and time-zone scheduling—that early models routinely botched. Ryu argues that today, for most STEM users who apply advanced math (but don’t invent new math), the model can handle nearly all needed mathematics—still with verification cautions.

    • Earlier models failed at multi-step real-world math (ledgers, scheduling)
    • Sudden capability jump noticed by external users, not just insiders
    • Current frontier: professional mathematicians and novel research problems
    • Best practice: double-check with simulations/verification; models still err
  6. 11:16 – 14:05

    Why math matters for AGI: long, consistent chains of reasoning

    Bubeck explains why mathematics was the ideal benchmark: clear questions and checkable answers (up to a point). More importantly, math requires extended, error-intolerant reasoning—training the kind of consistency and self-correction that should generalize to other domains.

    • Math as benchmark: unambiguous prompts and objectively verifiable outputs
    • Research-level evaluation gets harder than “answer is correct/incorrect”
    • Core skill: sustained reasoning where a single mistake breaks everything
    • Analogy: humans learn math to develop logical thinking and disciplined reasoning
  7. 14:05 – 19:33

    Erdős problems: literature-search ‘solutions,’ controversy, then genuine novelty

    They introduce Paul Erdős and the culture around his many questions, including a community-maintained list of open problems. Bubeck describes early AI “solutions” that were actually deep cross-literature connections, the public misunderstanding and controversy, and then the later emergence of truly new, publishable results generated with model assistance.

    • Who Erdős was, what an Erdős number is, and why his problems are a ‘treasure trove’
    • Early wins: AI finds solutions by mapping an Erdős question to results in distant literature
    • Public confusion: “solved 10 open problems” vs. “found existing solutions and connected them”
    • Acceleration claim: later results include genuinely new, publishable combinatorics solutions
  8. 19:33 – 21:31

    What counts as ‘discovery’: recombination vs. sparks of genius

    The conversation turns philosophical: is scientific progress mostly assembling existing pieces with reasoning, or does it require rare human insight? They suggest AI progress forces clearer thinking about what discovery really is, and whether ‘genius’ is distinct from high-powered recombination plus validation.

    • AI prompts debate on the nature of research breakthroughs
    • Questioning the myth of lone genius; discovery as cumulative and social
    • Possibility that recombination + reasoning can scale far further than assumed
    • Implication: AI may expand the space of accessible scientific advances
  9. 21:31 – 24:07

    Toward an automated researcher: extending ‘AGI time’ from days to weeks/months

    They outline the shift from interactive “professor–student” workflows to more autonomous systems that can pursue goals over long horizons. Bubeck introduces ‘AGI time’—how long a system can sustain human-like reasoning—and argues the key frontier is extending that duration dramatically.

    • Current workflow: iterative human-in-the-loop research collaboration
    • Automated researcher vision: models working autonomously for long periods
    • ‘AGI time’ framing: seconds → minutes → hours → days → (goal) weeks/months
    • Hard frontier: long-horizon coherence, memory, experimentation, and iteration
  10. 24:07 – 26:25

    Beyond context windows: math workspaces, persistent notes, and Codex-like loops

    Ryu explains that true research often exceeds a single chat context, because the thinking behind papers is far larger than the final write-up. He draws an analogy to coding agents that manage large repositories: math agents may maintain long-lived notes, compress history, and continue work across extended projects.

    • Limitation: typical chat sessions approximate ~‘50 pages’ of usable context
    • Research requires persistent artifacts (notes) and revisitation over months/years
    • Analogy: Codex handles large codebases via iterative instruction and summarization
    • Expectation: similar persistent-workflow agents will emerge for mathematics research
  11. 26:25 – 28:17

    Science acceleration in practice: instant benchmarks, more coding for mathematicians

    Andrew shares a concrete example of agentic assistance—generating a benchmark dataset within minutes during an experiment—highlighting ‘acceleration’ as a daily experience. Bubeck expands the idea: AI enables mathematicians to run computational experiments without relying on scarce coding help, and lets other scientists use more advanced math.

    • Practical acceleration: generating benchmark data quickly while mid-project
    • ‘Early experiments in science acceleration’ framing: time saved changes what gets attempted
    • Democratizing coding for mathematicians via agents like Codex
    • Cross-pollination: non-math scientists gain access to higher-level math through ChatGPT
  12. 28:17 – 33:35

    Humans’ role as models improve: choosing goals, asking questions, staying in control

    They argue models are beginning to exceed humans in some narrow research tasks (e.g., spotting errors, generating interesting questions). The human role shifts toward guiding what matters—science as understanding and control for human ends (health, robustness, engineering), rather than paper-count or novelty for its own sake.

    • Trendline claim: models may replicate most researcher behaviors within years
    • Models can also ask valuable questions, not just answer them
    • Science is about understanding and human goals (curing disease, building better systems)
    • Critical principle: humans must steer objectives and remain accountable
  13. 33:35 – 43:28

    Verification, shallow understanding risks, and advice for learning math with ChatGPT

    They close with two tensions: AI can accelerate verification of long proofs and reduce the lag between publication and trust, but overreliance can erode human depth and produce confident-looking nonsense. Their practical advice: use ChatGPT as a tutor and conversation partner, personalize prompts to your background, and keep human responsibility and rigor at the center.

    • AI-assisted verification can shorten years-long validation cycles and catch mistakes
    • Risk: mental atrophy and shallow understanding if users stop doing the hard work
    • Non-experts may generate long but incorrect proofs; rigor and expertise still matter
    • Learning advice: describe your background, ask follow-ups, request appropriately-leveled questions

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.