OpenAIHow AI is accelerating scientific discovery today and what's ahead — the OpenAI Podcast Ep. 10
CHAPTERS
- 0:00 – 1:15
OpenAI for Science: compressing decades of research into years
Andrew Mayne sets the stage with Kevin Weil and Alex Lupsasca to discuss how frontier AI models are starting to do genuinely novel scientific work. Kevin frames OpenAI for Science as an effort to dramatically speed up discovery by putting the best models in the hands of top researchers.
- •OpenAI for Science mission: accelerate scientific discovery at scale
- •Ambition: “25 years of research in 5 years”
- •Why now: frontier models are starting to do novel science
- •Early “existence proofs” that models can extend known results
- 1:15 – 2:52
From "can’t" to "can barely" to "can’t imagine without AI": the acceleration curve
Kevin explains the pattern he’s seen repeatedly: models move quickly from inability to weak capability to indispensable performance. Science is entering the early phase—small breakthroughs and workflow gains that hint at much larger impact.
- •Frontier science tasks are starting to become feasible with GPT-5-class models
- •Rapid step-change improvements over 6–12 month periods
- •Small breakthroughs today signal much bigger potential
- •Acceleration includes both speed and breadth of exploration
- 2:52 – 5:37
Physics case study: GPT finds obscure math identity to solve a pulsar PDE (with a typo)
Alex recounts a pivotal moment: using an advanced ChatGPT model to simplify an infinite series solution involving Legendre polynomials for a pulsar magnetic-field problem. The model located a niche 1950s identity in a Norwegian math journal, producing an almost-correct derivation that was easy to repair.
- •Real research task: PDE solution expressed as special-function series
- •Model decomposes the problem and retrieves an obscure published identity
- •Near-miss accuracy: correct method, small final typo (human-like error)
- •Impact: reframed Alex’s belief about what AI can do in theoretical physics
- 5:37 – 8:26
Conceptual literature search across fields and languages
Kevin and Alex highlight AI’s ability to perform conceptual, not keyword-based, literature discovery—surfacing relevant work even when terminology differs or sources are in another language. This reduces wasted effort and reconnects modern problems to forgotten or siloed prior art.
- •GPT can match ideas across domains with different jargon
- •Example: finding a related German PhD thesis via conceptual similarity
- •Literature search as a major (often underrated) accelerator
- •Helps researchers avoid reinventing the wheel and import methods from other fields
- 8:26 – 11:19
AI as a 24/7 collaborator: going deeper and broader than specialization allows
They discuss how modern science forces extreme specialization, making it hard to explore adjacent areas. With AI as an always-available collaborator, researchers can revisit abandoned directions, bridge disciplines, and sustain productive “infinite patience” collaboration.
- •Specialization makes adjacent-field exploration costly for humans
- •AI enables breadth (adjacencies) and depth (technical follow-through)
- •Collaboration as a core driver of breakthroughs, amplified by AI
- •Rebuttal to superficial criticisms (e.g., earlier model failures)
- 11:19 – 14:50
Kevin’s “fusion ladder” demo: undergrad to 20-year expert questions
Kevin shares his turning point: a Lawrence Livermore physicist demonstrated model performance on progressively harder fusion/shock-physics questions. The model kept up through increasingly expert-level prompts and even suggested specialized simulation tools, illustrating dramatic time savings for high-end work.
- •Stepwise escalation of technical depth (undergrad → grad → postdoc → veteran expert)
- •Model provides correct guidance across levels in a single workflow
- •Acceleration: tasks that take days can compress to hours/minutes
- •Signals how domain experts can weaponize AI for real scientific throughput
- 14:50 – 18:13
Black-hole symmetries: warm-ups, priming, and frontier reliability
Alex describes challenging GPT-5 Pro with a black-hole symmetry problem: it initially failed, succeeded on a simpler flat-space version, then solved the harder original after being “warmed up.” This becomes a lesson about effective interaction patterns and how to work near the edge of model capability.
- •Hard problem first: model incorrectly reports no symmetries
- •Warm-up problem: model identifies conformal symmetry and generators
- •Primed retry: model solves the full black-hole case correctly after long reasoning
- •Practical takeaway: scaffolding and incremental prompting mirror human research practice
- 18:13 – 24:33
Low pass-rate problems and the hidden frontier: making iteration less painful
Kevin emphasizes that frontier tasks often have low pass rates—models can be capable but inconsistent. The challenge is that researchers may abandon solvable problems too early; OpenAI is interested in reducing the cognitive and operational burden of repeated trials and refinement.
- •Frontier usage is not “one shot”; it’s iterative back-and-forth
- •Low pass-rate tasks are often the most valuable scientifically
- •Users struggle to distinguish “too hard” from “rarely right”
- •Opportunity: tools/workflows that automate retries, verification, and refinement
- 24:33 – 27:22
A snapshot paper: documented workflows plus new mathematical results
They introduce an upcoming OpenAI for Science paper compiling cross-discipline case studies of GPT-5 accelerating real research. It aims to be non-hypey, showing what works and what doesn’t, sharing full chat logs, and including several non-trivial new math results.
- •Collaborators: OpenAI team plus 8–9 external academics across fields
- •12 sections covering literature search, calculations, and other accelerators
- •Transparent evidence: share links showing the actual back-and-forth
- •Includes 4–5 new non-trivial mathematics results; some “paper-worthy”
- 27:22 – 29:53
Advice to students: AI as a pathfinder through the unknown
Asked about the future of scientific careers, Alex argues AI won’t eliminate scientists—it will amplify them, amid broader academic anxieties unrelated to AI. He explains how ChatGPT helps prototype approaches quickly by generating signposted avenues of attack, enabling rapid parallel exploration.
- •Current academic anxiety has multiple causes beyond AI
- •Research is navigation with uncertain routes; AI helps explore options fast
- •Practical workflow: upload notes, ask “what if” routes, get signposts
- •Outcome: higher productivity and more experimentation, especially for younger researchers
- 29:53 – 36:43
12 months vs 5 years: bottlenecks, biology, and turning predictions into reality
Kevin predicts profound near-term shifts in how science is conducted, analogous to recent changes in software engineering. They discuss life sciences constraints—models may generate more hypotheses than labs can test—while noting AI can prune search spaces in drug discovery and even speed regulatory documentation.
- •Rapid progress makes long-range forecasting difficult; 12-month horizons already transformative
- •Risk: experimental bottleneck vs exploding in-silico predictions
- •Drug discovery: AI prunes massive candidate spaces to focus experiments
- •Downstream acceleration: regulatory writing and synthesis; pilots with companies
- 36:43 – 40:31
Compute and thinking time: longer reasoning boosts success on hard problems
They stress that many skeptics are using older or free versions with limited “thinking time.” Kevin argues pass rates keep improving with more reasoning time—minutes to hours—suggesting that allocating more compute to expert scientific users could yield outsized gains even without new model training.
- •Model capability perception is stale; tool evolves unusually fast
- •Free tiers vs pro tiers: longer thinking enables harder problem-solving
- •Longer reasoning (2–24 hours) can significantly raise pass rate on frontier tasks
- •Compute allocation to scientists is itself a lever for accelerating discovery
- 40:31 – 42:18
Scientific benchmarks after saturation: from GPQA to frontier evaluations
Andrew and Kevin discuss how conventional benchmarks become less meaningful as models master them. Kevin cites GPQA’s jump from GPT-4’s ~39% to near 90% for latest models, motivating new “frontier science” evaluations and economically grounded tests like GDPVal.
- •Benchmark saturation forces creation of harder, more realistic tests
- •GPQA: originally hard PhD-level; models now near 90% vs humans ~70%
- •Need for evaluations tied to frontier math/science tasks
- •Economic-value evals (GDPVal) as another lens on real-world capability
- 42:18 – 48:12
Where they want acceleration most: black holes, dark matter, and fusion—then global adoption
Alex advocates for AI-accelerated black hole research and proposes integrating disparate dark matter data and theories to rule out candidates. Kevin highlights fusion as a world-changing target, while both emphasize a bottom-up strategy: release strong general tools so the global scientific community drives breakthroughs.
- •Alex’s priorities: black holes; integrate knowledge to unlock thorny theory
- •Dark matter idea: unify experiments + theories to eliminate inconsistent models
- •Kevin’s priority: fusion scaling and reliability; huge societal upside
- •Strategy: not top-down—enable worldwide scientists; goal: many Nobel-level discoveries with AI