Skip to content
OpenAIOpenAI

AGI progress, surprising breakthroughs, and the road ahead — the OpenAI Podcast Ep. 5

How close are we to automating scientific discovery? What do AI competition wins really tell us about progress toward AGI? OpenAI Chief Scientist Jakub Pachocki and researcher Szymon Sidor share inside stories—from gold medals at the International Math Olympiad to surprising leaps in reasoning—that reveal where AI is headed next. Chapters 1:20 – From high school in Poland to AI research leaders 4:50 – Explaining AGI: technical and everyday perspectives 6:30 – Automating scientific discovery with AI 7:50 – Breakthroughs in medicine, AI safety, and alignment 10:30 – Today is a decade in the making 14:30 – Benchmark saturation and its limits 16:50 – Why math competitions matter for AI 18:15 – How models reason without tools 21:45 – Recognizing when a model can’t solve a problem 23:30 – Storytime: AtCoder competition in Japan 26:50 – How reasoning breakthroughs really happen 28:55 – What’s next for scaling and long-horizon reasoning 30:30 – What AGI will look and feel like 36:25 – Balancing trust and personal value 34:00 – Advice to high school students in 2025

Andrew MaynehostJakub PachockiguestSzymon Sidorguest
Aug 15, 202540mWatch on YouTube ↗

CHAPTERS

  1. 0:00 – 1:21

    Setting OpenAI’s research direction: roles, roadmap, and what “chief scientist” means

    Andrew Mayne opens by introducing guests Jakub Pachocki (Chief Scientist) and Szymon Sidor and frames the episode around measuring AI progress and AGI. Jakub explains his responsibility for setting OpenAI’s research roadmap, while Szymon describes his more fluid IC-focused role and prioritizing usefulness.

    • Episode focus: measuring AI progress, defining AGI, and anticipating breakthroughs
    • Jakub’s role: choosing technical bets and long-term research directions
    • Szymon’s role: mostly individual-contributor work with some leadership
    • Early teaser: models recognizing when they’re stuck and org readiness for rapid progress
  2. 1:21 – 2:56

    From a Polish high school to AI leadership: mentorship, competitions, and deep foundations

    Jakub and Szymon recount meeting in the same high school in Gdynia, Poland, and how they became closer after moving to the US. They credit a standout computer science teacher and a competition-driven curriculum for building their technical depth and ambition.

    • Shared origin: same high school, later bonding through moving to the US
    • Influential mentor: Mr. Ryszard Szubartowski and a competition culture
    • Curriculum beyond typical high school (graph theory, matrices, etc.)
    • Competition training as a formative pathway into research and engineering excellence
  3. 2:56 – 4:35

    AI as a learning companion: what it can (and can’t) replace in education

    The conversation shifts to how tools like ChatGPT change learning by enabling interactive explanations and multimedia. Jakub highlights that great teachers provide emotional support and space—something AI alone may struggle to replicate—positioning AI as an amplifier rather than a replacement.

    • Interactive explanation as a new capability (e.g., Monty Hall simulations)
    • Socratic tutoring and concept explanation as strong AI use cases
    • Human mentorship provides emotional support and structure AI may lack
    • Best outcome: teachers using AI become more capable, not displaced
  4. 4:35 – 7:29

    Defining AGI as capabilities diversify: from conversation to real-world impact

    Jakub explains how “AGI” used to feel abstract, but progress has separated distinct capabilities (natural conversation, math, research). He argues pointwise milestones like Olympiad results are becoming less adequate, pushing evaluation toward real-world impact—especially AI’s ability to automate the creation of new technology.

    • AGI once felt monolithic; now capabilities are clearly distinct
    • Conversation competence is already strong across broad topics
    • Olympiad milestones matter, but are imperfect as progress accelerates
    • Impact-based framing: automating discovery and technology creation as the key shift
  5. 7:29 – 9:33

    Automating scientific discovery: general intelligence over narrow deployments

    Andrew probes what “automating science” could look like, and Jakub emphasizes building broadly general intelligence rather than targeting narrow domains. He notes some areas are particularly amenable—especially those requiring heavy reasoning plus domain intuition—highlighting medicine as promising, and stressing the importance of automating AI research and alignment work too.

    • Strategic priority: an “automated researcher” built from general intelligence
    • Generality is seen as the path to the biggest discoveries and breakthroughs
    • Early strong domain signals: medicine and other reasoning-heavy fields
    • Automating AI research and alignment/safety is both feasible and crucial
  6. 9:33 – 14:27

    A decade in the making: why ‘small economic impact’ headlines miss the trajectory

    Szymon offers a historical perspective: deep learning for language barely worked 10 years ago, and progress has been compounding through GPT-1/2/3/4 to today’s reasoning and competition-grade systems. Both guests argue that focusing on today’s measured economic impact ignores how quickly these capabilities have moved from near-zero utility to widely transformative tools.

    • Anecdote: older sentiment models failed basic negation ("not bad")
    • Progress milestones: GPT-2 coherent text, GPT-3 scale, GPT-4 surprise factor
    • From “slightly better Google” to “deep research” usefulness and reduced confabulation
    • Economic impact metrics lag capability growth; compounding progress is the story
  7. 14:27 – 17:01

    Benchmark saturation and what it breaks: human-level scores, distortion, and utility gaps

    Jakub explains why standard benchmarks are becoming less informative: many are saturating at human-level performance, while training methods can disproportionately boost narrow skills (e.g., math) without reflecting broad intelligence. The discussion reframes evaluation around overall utility and the ability to generate genuinely new insights rather than just scoring well on tests.

    • Saturation: standardized tests become less discriminative at high capability
    • Benchmarks can be noisy/limited and sometimes poorly constructed
    • Skill-targeted training can inflate scores without broad intelligence gains
    • Shift in evaluation: usefulness and insight discovery over test-taking prowess
  8. 17:01 – 21:47

    Why math and programming competitions still matter: long-horizon thinking under constraints

    Jakub defends Olympiads (IMO/IOI) as valuable because they test sustained, creative reasoning under time pressure with limited required knowledge. Andrew highlights that the IMO gold-level system succeeded without external tools, emphasizing an important milestone: reasoning competence rather than formula application.

    • IMO/IOI as constrained tests of deep thinking over hours, not memorization
    • Large evidence base: global competition makes difficulty credible
    • Milestone: gold-level performance without calculators/tools
    • Competitions as a proxy for creative reasoning and persistence
  9. 21:47 – 23:41

    Knowing when you’re stuck: problem 6, calibration, and reducing hallucination risk

    Jakub describes how IMO “problem 6” often requires out-of-the-box insight, and how both OpenAI and DeepMind models solved problems 1–5 but stalled there. A key breakthrough is the model correctly identifying it made no progress—an important step toward better self-assessment, calibration, and avoiding confident but wrong outputs.

    • IMO’s “problem 6” as a boundary test for out-of-distribution creativity
    • Observed pattern: perfect performance on 1–5, failure to advance on 6
    • Meta-cognition: model detects lack of progress instead of bluffing
    • Connection to hallucinations and the value of calibrated refusal/uncertainty
  10. 23:41 – 26:29

    Storytime: AtCoder in Japan—10-hour optimization, live competition, and second place

    Jakub recounts entering a model into AtCoder, a prestigious open competition featuring long-horizon heuristic optimization over a single 10-hour problem. He describes watching the model race a top human competitor (Saiho, an OpenAI colleague) on a livestream, with the model finishing second—underscoring progress in extended-horizon problem solving.

    • AtCoder differs from Olympiads: single problem, 10 hours, heuristic optimization
    • No single “correct” answer; success depends on exploration and iterative improvement
    • Personal dynamic: Saiho previously predicted long-horizon contests would resist automation longer
    • Outcome: model takes 2nd; Saiho wins and jokes about being exhausted
  11. 26:29 – 28:45

    How reasoning breakthroughs really happen: hard-earned training, surprise moments, and org readiness

    Szymon pushes back on the idea that “longer chain-of-thought” is a simple tweak, emphasizing the difficulty of making reasoning training work reliably. He describes moments when internal results were shocking enough to trigger leadership-level discussions about whether the organization was prepared for very rapid progress.

    • Reasoning gains are not “free”; making them work required major engineering/research effort
    • Internal surprise: clear inflection points when models started improving with reasoning data
    • Organizational implication: serious conversations about readiness and safety posture
    • Public perception lags: breakthroughs look sudden but are years in the making
  12. 28:45 – 30:22

    The road ahead: scaling, vastly more compute per problem, and persistent long-horizon agents

    Jakub argues scaling remains fundamental and will compound with newer reasoning approaches. He highlights a next step: persistence—spending orders of magnitude more compute on problems that matter (medical research, next-model development), enabling systems that plan and work for long periods on focused tasks.

    • Scaling hasn’t disappeared; it will stack with reasoning-centric methods
    • Today’s per-answer compute increases are modest compared to what’s possible
    • Future: spending massive compute budgets on high-value research questions
    • Persistence as capability: long-duration planning, iteration, and focus
  13. 30:22 – 32:30

    What AGI will look and feel like: automated R&D organizations and more human interfaces

    Jakub paints a picture of AGI as a largely automated “company” of researchers and engineers producing technology artifacts—codebases, designs, experiments—interfacing with people and the world. He also expects major evolution in interfaces: persistence, richer expression modalities, and stronger relational dynamics as systems feel more human-like.

    • AGI as automated R&D: systems that build technology, not just answer prompts
    • Outputs extend to experiments, code, designs, and integrated workflows
    • Interfaces will change: persistence and multimodal expression increase attachment and utility
    • Societal and technical work needed to ensure benefits and manage risks
  14. 32:30 – 33:24

    Balancing trust and personal value: data access, robustness gaps, and exploitation risks

    The discussion turns to the trade-off between giving AI access to personal data (calendar, email) for high utility and the current limits of robustness and security. Jakub emphasizes that while the value is clear, models and systems aren’t yet fully trustworthy against adversarial exploitation, making this an urgent field-wide iteration problem.

    • Personal integrations (calendar/email) signal rising trust and usefulness
    • Trade-off: economic/personal value vs. security and robustness concerns
    • Key risk: models/systems can be exploited by attackers or prompt-based manipulation
    • Need for ongoing iteration across the field, not just one company
  15. 33:24 – 40:23

    Advice to high school students in 2025: learn to code, think structurally, dream bigger

    Szymon advises students to keep learning to code because it builds structural thinking—decomposing complex problems—even if programming changes over time. Jakub adds that many perceived constraints are self-imposed, encouraging ambitious goal-setting (including studying abroad) and believing you can meaningfully affect the world; they close with formative inspirations (Hackers & Painters, Iron Man, AlphaGo).

    • Coding remains valuable as a way to train rigorous problem decomposition
    • Don’t be swayed by claims that coding is obsolete; understanding systems matters
    • Reframe constraints: ambition and unconventional paths are often feasible
    • Personal inspirations: Hackers & Painters, Iron Man’s robotics spark, AlphaGo’s trajectory

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.