OpenAIAGI progress, surprising breakthroughs, and the road ahead — the OpenAI Podcast Ep. 5
CHAPTERS
- 0:00 – 1:21
Setting OpenAI’s research direction: roles, roadmap, and what “chief scientist” means
Andrew Mayne opens by introducing guests Jakub Pachocki (Chief Scientist) and Szymon Sidor and frames the episode around measuring AI progress and AGI. Jakub explains his responsibility for setting OpenAI’s research roadmap, while Szymon describes his more fluid IC-focused role and prioritizing usefulness.
- •Episode focus: measuring AI progress, defining AGI, and anticipating breakthroughs
- •Jakub’s role: choosing technical bets and long-term research directions
- •Szymon’s role: mostly individual-contributor work with some leadership
- •Early teaser: models recognizing when they’re stuck and org readiness for rapid progress
- 1:21 – 2:56
From a Polish high school to AI leadership: mentorship, competitions, and deep foundations
Jakub and Szymon recount meeting in the same high school in Gdynia, Poland, and how they became closer after moving to the US. They credit a standout computer science teacher and a competition-driven curriculum for building their technical depth and ambition.
- •Shared origin: same high school, later bonding through moving to the US
- •Influential mentor: Mr. Ryszard Szubartowski and a competition culture
- •Curriculum beyond typical high school (graph theory, matrices, etc.)
- •Competition training as a formative pathway into research and engineering excellence
- 2:56 – 4:35
AI as a learning companion: what it can (and can’t) replace in education
The conversation shifts to how tools like ChatGPT change learning by enabling interactive explanations and multimedia. Jakub highlights that great teachers provide emotional support and space—something AI alone may struggle to replicate—positioning AI as an amplifier rather than a replacement.
- •Interactive explanation as a new capability (e.g., Monty Hall simulations)
- •Socratic tutoring and concept explanation as strong AI use cases
- •Human mentorship provides emotional support and structure AI may lack
- •Best outcome: teachers using AI become more capable, not displaced
- 4:35 – 7:29
Defining AGI as capabilities diversify: from conversation to real-world impact
Jakub explains how “AGI” used to feel abstract, but progress has separated distinct capabilities (natural conversation, math, research). He argues pointwise milestones like Olympiad results are becoming less adequate, pushing evaluation toward real-world impact—especially AI’s ability to automate the creation of new technology.
- •AGI once felt monolithic; now capabilities are clearly distinct
- •Conversation competence is already strong across broad topics
- •Olympiad milestones matter, but are imperfect as progress accelerates
- •Impact-based framing: automating discovery and technology creation as the key shift
- 7:29 – 9:33
Automating scientific discovery: general intelligence over narrow deployments
Andrew probes what “automating science” could look like, and Jakub emphasizes building broadly general intelligence rather than targeting narrow domains. He notes some areas are particularly amenable—especially those requiring heavy reasoning plus domain intuition—highlighting medicine as promising, and stressing the importance of automating AI research and alignment work too.
- •Strategic priority: an “automated researcher” built from general intelligence
- •Generality is seen as the path to the biggest discoveries and breakthroughs
- •Early strong domain signals: medicine and other reasoning-heavy fields
- •Automating AI research and alignment/safety is both feasible and crucial
- 9:33 – 14:27
A decade in the making: why ‘small economic impact’ headlines miss the trajectory
Szymon offers a historical perspective: deep learning for language barely worked 10 years ago, and progress has been compounding through GPT-1/2/3/4 to today’s reasoning and competition-grade systems. Both guests argue that focusing on today’s measured economic impact ignores how quickly these capabilities have moved from near-zero utility to widely transformative tools.
- •Anecdote: older sentiment models failed basic negation ("not bad")
- •Progress milestones: GPT-2 coherent text, GPT-3 scale, GPT-4 surprise factor
- •From “slightly better Google” to “deep research” usefulness and reduced confabulation
- •Economic impact metrics lag capability growth; compounding progress is the story
- 14:27 – 17:01
Benchmark saturation and what it breaks: human-level scores, distortion, and utility gaps
Jakub explains why standard benchmarks are becoming less informative: many are saturating at human-level performance, while training methods can disproportionately boost narrow skills (e.g., math) without reflecting broad intelligence. The discussion reframes evaluation around overall utility and the ability to generate genuinely new insights rather than just scoring well on tests.
- •Saturation: standardized tests become less discriminative at high capability
- •Benchmarks can be noisy/limited and sometimes poorly constructed
- •Skill-targeted training can inflate scores without broad intelligence gains
- •Shift in evaluation: usefulness and insight discovery over test-taking prowess
- 17:01 – 21:47
Why math and programming competitions still matter: long-horizon thinking under constraints
Jakub defends Olympiads (IMO/IOI) as valuable because they test sustained, creative reasoning under time pressure with limited required knowledge. Andrew highlights that the IMO gold-level system succeeded without external tools, emphasizing an important milestone: reasoning competence rather than formula application.
- •IMO/IOI as constrained tests of deep thinking over hours, not memorization
- •Large evidence base: global competition makes difficulty credible
- •Milestone: gold-level performance without calculators/tools
- •Competitions as a proxy for creative reasoning and persistence
- 21:47 – 23:41
Knowing when you’re stuck: problem 6, calibration, and reducing hallucination risk
Jakub describes how IMO “problem 6” often requires out-of-the-box insight, and how both OpenAI and DeepMind models solved problems 1–5 but stalled there. A key breakthrough is the model correctly identifying it made no progress—an important step toward better self-assessment, calibration, and avoiding confident but wrong outputs.
- •IMO’s “problem 6” as a boundary test for out-of-distribution creativity
- •Observed pattern: perfect performance on 1–5, failure to advance on 6
- •Meta-cognition: model detects lack of progress instead of bluffing
- •Connection to hallucinations and the value of calibrated refusal/uncertainty
- 23:41 – 26:29
Storytime: AtCoder in Japan—10-hour optimization, live competition, and second place
Jakub recounts entering a model into AtCoder, a prestigious open competition featuring long-horizon heuristic optimization over a single 10-hour problem. He describes watching the model race a top human competitor (Saiho, an OpenAI colleague) on a livestream, with the model finishing second—underscoring progress in extended-horizon problem solving.
- •AtCoder differs from Olympiads: single problem, 10 hours, heuristic optimization
- •No single “correct” answer; success depends on exploration and iterative improvement
- •Personal dynamic: Saiho previously predicted long-horizon contests would resist automation longer
- •Outcome: model takes 2nd; Saiho wins and jokes about being exhausted
- 26:29 – 28:45
How reasoning breakthroughs really happen: hard-earned training, surprise moments, and org readiness
Szymon pushes back on the idea that “longer chain-of-thought” is a simple tweak, emphasizing the difficulty of making reasoning training work reliably. He describes moments when internal results were shocking enough to trigger leadership-level discussions about whether the organization was prepared for very rapid progress.
- •Reasoning gains are not “free”; making them work required major engineering/research effort
- •Internal surprise: clear inflection points when models started improving with reasoning data
- •Organizational implication: serious conversations about readiness and safety posture
- •Public perception lags: breakthroughs look sudden but are years in the making
- 28:45 – 30:22
The road ahead: scaling, vastly more compute per problem, and persistent long-horizon agents
Jakub argues scaling remains fundamental and will compound with newer reasoning approaches. He highlights a next step: persistence—spending orders of magnitude more compute on problems that matter (medical research, next-model development), enabling systems that plan and work for long periods on focused tasks.
- •Scaling hasn’t disappeared; it will stack with reasoning-centric methods
- •Today’s per-answer compute increases are modest compared to what’s possible
- •Future: spending massive compute budgets on high-value research questions
- •Persistence as capability: long-duration planning, iteration, and focus
- 30:22 – 32:30
What AGI will look and feel like: automated R&D organizations and more human interfaces
Jakub paints a picture of AGI as a largely automated “company” of researchers and engineers producing technology artifacts—codebases, designs, experiments—interfacing with people and the world. He also expects major evolution in interfaces: persistence, richer expression modalities, and stronger relational dynamics as systems feel more human-like.
- •AGI as automated R&D: systems that build technology, not just answer prompts
- •Outputs extend to experiments, code, designs, and integrated workflows
- •Interfaces will change: persistence and multimodal expression increase attachment and utility
- •Societal and technical work needed to ensure benefits and manage risks
- 32:30 – 33:24
Balancing trust and personal value: data access, robustness gaps, and exploitation risks
The discussion turns to the trade-off between giving AI access to personal data (calendar, email) for high utility and the current limits of robustness and security. Jakub emphasizes that while the value is clear, models and systems aren’t yet fully trustworthy against adversarial exploitation, making this an urgent field-wide iteration problem.
- •Personal integrations (calendar/email) signal rising trust and usefulness
- •Trade-off: economic/personal value vs. security and robustness concerns
- •Key risk: models/systems can be exploited by attackers or prompt-based manipulation
- •Need for ongoing iteration across the field, not just one company
- 33:24 – 40:23
Advice to high school students in 2025: learn to code, think structurally, dream bigger
Szymon advises students to keep learning to code because it builds structural thinking—decomposing complex problems—even if programming changes over time. Jakub adds that many perceived constraints are self-imposed, encouraging ambitious goal-setting (including studying abroad) and believing you can meaningfully affect the world; they close with formative inspirations (Hackers & Painters, Iron Man, AlphaGo).
- •Coding remains valuable as a way to train rigorous problem decomposition
- •Don’t be swayed by claims that coding is obsolete; understanding systems matters
- •Reframe constraints: ambition and unconventional paths are often feasible
- •Personal inspirations: Hackers & Painters, Iron Man’s robotics spark, AlphaGo’s trajectory