Skip to content
a16za16z

Inside OpenAI’s Breakthroughs in Mathematical Reasoning

a16z Infra Partner Lisha Li sits down with OpenAI mathematicians Mehtaab Sawhney and Mark Sellke to discuss how quickly AI’s mathematical capabilities are advancing, what recent results reveal about model reasoning, and what happens when AI begins making progress on problems mathematicians have struggled with for decades. Mehtaab and Mark unpack several recent results from OpenAI’s models, including advances in sphere packing and the construction of a non-sofic group. They explain why the surprising part isn’t simply that models can search more possibilities or work longer than humans: in many cases, the reasoning traces look remarkably similar to the work of an expert mathematician, including choosing promising approaches, backtracking when they fail, and combining ideas from across the literature. They also explore what this means for mathematics itself: how the role of human taste and judgment may change, whether AI could produce far more mathematics than humans can absorb, and why models that accelerate discovery may also make sophisticated results easier to understand. Timestamps: 00:00 - Intro 00:50 - From Practicing Mathematician to OpenAI: Meet Mark & Mehtaab 02:43 - Why GPT-5 Was the Conversion Moment 04:21 - Beyond Search & Connections: How Recent Progress Goes Deeper 09:51 - Reasoning Traces: Is It Lucky Sampling or Actual Backtracking? 11:44 - Why Math Papers Are a Bad Training Set for Real Mathematics 16:20 - The Astra 10-Problem Set: Favorites & Deep Dives 36:17 - The Harness vs the Model: What Actually Matters? 40:01 - What Even Is "Taste" in a Model? 57:32 - How Should the Math Community Adopt AI? 01:00:01 - Empirical vs Theoretical Math & the Positive Vision Resources: Follow Lisha Li on X: https://x.com/lishali88 Follow Mehtaab Sawhney on X: https://x.com/mehtaab_sawhney Follow Mark Sellke on X: https://x.com/MarkSellke Stay Updated: If you enjoyed this episode, be sure to like, subscribe, and share with your friends! Find a16z on X: https://twitter.com/a16z Find a16z on LinkedIn: https://www.linkedin.com/company/a16z Listen to the a16z Show on Spotify: https://open.spotify.com/show/5bC65RDvs3oxnLyqqvkUYX Listen to the a16z Show on Apple Podcasts: https://podcasts.apple.com/us/podcast/a16z-podcast/id842818711 Follow our host: https://x.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see http://a16z.com/disclosures.

Lisha LihostMehtaab Sawhneyguest
Sep 8, 20261h 5mWatch on YouTube ↗

At a glance

WHAT IT’S REALLY ABOUT

How OpenAI’s Astra models solve hard math with humanlike reasoning

  1. Two OpenAI mathematicians describe how recent reasoning models moved from literature-connection and search advantages to producing mathematician-like multi-step proofs with backtracking and decision-making.
  2. They argue models excel at executing intricate technical details and can restart or parallelize attempts, reducing human limitations like context pollution and sunk-cost persistence.
  3. They walk through Astra’s 10-problem set highlights, including a tight asymptotic characterization of the sphere-packing linear-programming bound and improved bounds for spherical/binary codes via symmetry and representation theory.
  4. They discuss how prompting and “harness” choices affect outcomes, noting models can stop early once a prompt’s goal is met even if further improvements are available.
  5. They anticipate math culture will adapt by valuing explanation, synthesis, and community understanding as theorem-generation becomes less scarce, while still leaving room for enduring grand challenges (e.g., P vs NP).

IDEAS WORTH REMEMBERING

5 ideas

AI’s biggest edge in math is relentless, high-precision follow-through on viable ideas.

The guests argue models are unusually strong at “execution”: once an approach is plausible, they can carry it through dense, error-prone inequalities and bookkeeping that often causes humans to stall or quit. This changes the risk/reward calculus of pursuing finicky approaches that humans would abandon due to time constraints.

Reasoning traces suggest genuine backtracking and probabilistic “path management,” not mere lucky sampling.

They describe the model trying multiple approaches, making mistakes, revisiting earlier steps, and updating which branches seem promising—more like a researcher pruning a search tree than brute-forcing all paths. The ability to restart sessions or run parallel attempts also reduces the ‘context pollution’ and sunk-cost bias humans experience.

Math reasoning improvements are framed as an emergent property of general reasoning training, not “math papers as curriculum.”

The conversation highlights that formal artifacts (papers/textbooks) often omit motivation and struggle, yet reasoning behaviors still appear. Their claim is OpenAI is training general-purpose reasoning abilities (long-horizon planning, backtracking, decomposition) that transfer into mathematics, rather than relying on math-specific datasets alone.

Astra doesn’t just improve bounds; it can fully characterize an optimization framework (achievability + optimality).

In the sphere-packing LP-bound result, the model both constructs an explicit witness function achieving a bound and proves optimality within that LP framework—turning a numerically conjectured asymptotic into an explained, tight statement. The guests emphasize the proof’s surprising brevity and “why didn’t we think of that?” feel.

Prompting and harness design can gate breakthroughs even when the underlying capability is present.

For spherical/binary codes, the model initially delivered an improvement, but further progress required a human to ask an additional “push it further” question. The guests interpret this less as a capability gap and more as task-orientation: models stop once the prompt’s objective is satisfied.

WORDS WORTH SAVING

5 quotes

Often as a practicing mathematician, you, you have an idea, and then you kind of think it might work, then you try for a few hours, a few days, a few weeks. And at some point, you give up, and then a not so uncommon experience is that you find out a year or two later that somebody else got the idea to work that you thought that didn't work.

Mark Sellke

Whereas for GPT, like, okay, I'll, like, a human told me to do this. Let's, let's just do this. And so that's why we're sort of in this renaissance of, like, reachable, uh, results.

Lisha Li

Like, is the model just guessing in some insane way? Like, is it thinking in some totally foreign... Like, what, what's going on? But, but actually it, it's, it's reasoning kind of shockingly like a, an expert human would.

Mark Sellke

But, but I, I think it's... Like, like a year ago, I, I would've been very surprised to learn that like all of these AI proofs are like very short and elegant.

Mark Sellke

The ceiling for difficulty of a math problem is pretty high. Even if, um, you know, kind of, e- even if AI get, you know, continues getting like exponentially better at math, like it might, you know, it's plausible we'll never solve something like P versus NP.

Mehtaab Sawhney

GPT-5 as “conversion moment” for mathematiciansReasoning traces, backtracking, and pruning searchWhy math papers/textbooks are poor ‘thinking’ datasetsSphere packing in high dimensions and LP boundsViazovska, E8/Leech lattices, asymptoticsSpherical and binary codes; error correction; symmetryRepresentation theory vs complex analysis techniques in AI proofs”“Harness vs model; task-orientation; prompting for longer horizon work”“Taste in research: utilitarian definition and division of labor”“Non-sofic groups; approximating infinite groups by finite ones; Cayley graphs”“Aldous–Lyons conjecture relationship and contrast in proof lengths”“Adoption in the math community: understanding, attribution, synthesis

High quality AI-generated summary created from speaker-labeled transcript.

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.