Skip to content
a16za16z

What Today’s Best Models Still Can’t Do in Math

a16z’s Lisha Li sits down with Daniel Litt, Assistant Professor of Mathematics at the University of Toronto, to unpack AI's rapid progress in mathematics, what today's frontier models can actually do, and what they're still missing about the way mathematicians think. Daniel explains why some recent AI-generated results are genuinely impressive, including an autonomous solution to the Erdős unit distance problem, but argues that solving problems is only one part of mathematics. Today's models can grind through calculations, combine known techniques, and search enormous spaces, but still struggle with intuition, theory building, identifying the right questions, and developing the kind of big-picture understanding that drives much of mathematical progress. Lisha and Daniel also explore how AI is already changing mathematical research, why an explosion of AI-generated papers could distort academic incentives, and what happens if researchers outsource the work of thinking rather than use AI to deepen it. Ultimately, they ask a question that extends far beyond mathematics: as AI gets better at intellectual work, how do we make sure humans keep getting better at thinking too? Timestamps: 00:00 - Intro 01:00 - Meet Daniel Litt: A Practicing Mathematician's Evolving Views on AI 02:24 - The Most Impressive Result: The Erdős Unit Distance Problem 06:12 - What AI Is Actually Doing for Working Mathematicians Today 12:12 - Intuition, Taste & Why Math Isn't Just About Proofs 20:26 - Deep Thinking vs Pattern Matching: What Models Are Missing 26:17 - Why the Unit Distance Result Was Actually Creative 33:33 - How Should the Math Community Adapt to AI? 43:26 - Where AI Will Impact Applied Math First 46:24 - Taking Advantage of AI Without Losing the Craft 49:21 - Comparing Anthropic vs OpenAI in Math 51:34 - Why Some Labs Have Gone More Secretive 59:40 - Raising a Mathematician: Teaching Math to a Toddler Resources: Follow Daniel Litt on X: https://x.com/littmath Follow Lisha Li on X: https://x.com/lishali88 Stay Updated: If you enjoyed this episode, be sure to like, subscribe, and share with your friends! Find a16z on X: https://twitter.com/a16z Find a16z on LinkedIn: https://www.linkedin.com/company/a16z Listen to the a16z Show on Spotify: https://open.spotify.com/show/5bC65RDvs3oxnLyqqvkUYX Listen to the a16z Show on Apple Podcasts: https://podcasts.apple.com/us/podcast/a16z-podcast/id842818711 Follow our host: https://x.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see http://a16z.com/disclosures.

Daniel LittguestLisha Lihost
Sep 1, 20261h 3mWatch on YouTube ↗

CHAPTERS

  1. 0:00 – 1:00

    Math as understanding vs papers: why “knowledge in model weights” feels unsatisfying

    Daniel frames mathematics as a pursuit of understanding, not paper production—and worries that even correct results are hollow if the understanding only “lives” inside a model. This sets up the episode’s central tension: capability gains vs the human craft of making sense.

    • Mathematics aims at understanding, not output volume
    • “Understanding in model weights” is personally unsatisfying
    • Curiosity and human comprehension are core motivations
    • Sets up later concerns about incentives and human capital
  2. 1:00 – 2:24

    Meet Daniel Litt: practicing mathematician, rapidly updating views on AI

    Lisha introduces Daniel as a University of Toronto mathematician who has been unusually vocal about how AI is changing mathematical work. They position the discussion around what mathematicians actually do, beyond symbol pushing.

    • Daniel’s role as an active research mathematician
    • The conversation will focus on practice, not just benchmarks
    • Math work includes confusion, intuition, and question-formation
    • Motivation: how the community should respond
  3. 2:24 – 6:19

    Most impressive AI math result so far: the Erdős unit distance problem

    Daniel highlights the AI solution to the Erdős unit distance problem as his favorite fully autonomous result, partly because it felt unexpected and cross-pollinated ideas from another area. They discuss how to judge significance by whether new ideas prove fruitful later.

    • Unit distance result felt more than “last mile” polishing
    • Unexpected outcome + importing classic techniques from elsewhere
    • Post-hoc value test: did it generate useful new ideas?
    • Follow-on: humans used ideas to attack other open questions
  4. 6:19 – 8:25

    Human-like proofs, not “inhuman” math: what the outputs feel like to read

    Contrary to the idea that models win via alien symbol manipulation, Daniel says many AI proofs look recognizably human in structure. The “inhuman” part is more about stamina and breadth of recall than exotic reasoning moves.

    • Released chain-of-thought summaries appear human-like
    • Outputs are understandable (though often poorly written)
    • Models don’t tire and can draw from vast knowledge
    • Still weak at certain mathematical activities (not just fields)
  5. 8:25 – 14:32

    Claude vs ChatGPT in math: similar capabilities, different UX—both weak at theory of mind

    They compare Anthropic and OpenAI models and conclude they largely solve a similar slice of math tasks. Daniel’s preference for ChatGPT is mostly historical inertia (it became useful earlier), and both systems struggle to model what the user knows.

    • Claude and ChatGPT appear neck-and-neck on many tasks
    • ChatGPT improved earlier; usage persists by habit
    • Both are “pretty bad” at theory of mind for expert users
    • Solutions share a recognizable flavor tied to model strengths
  6. 14:32 – 18:36

    What mathematicians actually do: problem solving, theory building, and finding the right questions

    Daniel describes day-to-day mathematical activity as exploring failures of understanding, developing analogies, and even inventing the right conjectures. He emphasizes that much of the process is non-rigorous, experimental, and concept-driven.

    • Problem solving vs theory building as a useful taxonomy
    • Open problems serve as benchmarks of misunderstanding
    • Power of analogies (e.g., cohomology vs fundamental groups)
    • Hard part is often formulating the right conjecture
    • Example: Birch–Swinnerton-Dyer as an early “big data” conjecture
  7. 18:36 – 21:17

    How AI helps (and doesn’t): from “Google substitute” to parallel example-hunting and coding leverage

    For long-running, theory-heavy projects, Daniel finds AI mostly saves time on lookup and background learning rather than doing deep work. But it dramatically changes what he attempts by removing friction in coding and large-scale experimentation.

    • AI helps less when questions are vague or ill-posed
    • On older projects, it’s mainly faster literature/overview help
    • AI unlocks coding-heavy projects for non-coders
    • Massively parallel search for examples becomes practical
  8. 21:17 – 23:16

    Taste, “ugly proofs,” and doing science with concepts (not art)

    They unpack what mathematicians mean by elegance and why “ugly” often means unilluminating grind. Daniel argues against over-indexing on aesthetics, preferring a science-like lens: pursue what is fundamental and opens new understanding.

    • Aesthetics can be a limiting failure mode for young researchers
    • “Ugly” often means grinding calculations without insight
    • Daniel’s stance: “win by any means necessary,” then understand
    • Math as ‘physics with concepts’: prioritize fundamental ideas
    • Progress relies on many people pursuing diverse curiosities
  9. 23:16 – 29:28

    Deep thinking vs pattern matching: why proving true conjectures and building new machinery is harder for models

    Daniel distinguishes between tasks where a counterexample construction exists (often more accessible) and deep conjectures likely requiring new ideas within broad theoretical frameworks. They discuss why theory-building is a fuzzier skill to reward-train and evaluate.

    • Counterexamples can be more tractable than proof of truth
    • Hard conjectures sit inside large frameworks with strong evidence
    • Many frontier problems likely require serious new techniques
    • Theory-building is difficult to reward with intermediate signals
    • Daniel expects continued growth, possibly via better RL environments
  10. 29:28 – 33:20

    Why humans sometimes discover better proofs: a lemma story about grind avoidance and conceptual upgrades

    Daniel recounts a paper where models couldn’t prove a lemma until he reframed it via extensive example work, after which the model proved the improved statement quickly. He notes that if the model had simply produced a huge grind proof, it might have prevented the conceptual improvement from being found.

    • Models failed on an initial lemma; examples suggested a stronger statement
    • Once reframed, AI could prove the better lemma quickly
    • Human reluctance to grind can catalyze conceptual insight
    • Models can now output extremely long, brute-force proofs—often undesirable
    • Informal reasoning aims to transmit understanding, not just correctness
  11. 33:20 – 52:45

    How the math community should adapt: incentives, ‘slot machine’ papers, and mode collapse in exploration

    They turn to institutional impacts: models make it easy to generate many correct-but-low-engagement papers, potentially distorting hiring and publication incentives. Daniel worries about duplicated proofs and reduced diversity of exploration if research becomes subordinated to model priors.

    • Risk: optimizing for paper count rather than understanding
    • ‘Slot machine’ dynamic: prompt models to prove recent conjectures
    • Uptick in low-quality arXiv output (sometimes correct but unengaged)
    • Multiple near-identical papers can appear within days
    • Concern: model-driven exploration may reduce ‘thousand flowers’ diversity
  12. 52:45 – 59:33

    Verification limits and long-horizon failures: why short proofs dominate and how experts actually check work

    They discuss why many showcased AI proofs are short: long arguments are harder to validate, and models can’t reliably “stress test” proofs at a high level yet. Daniel gives an example of a human expert skill—seeing a paper can’t work from its global structure—that models still struggle to replicate.

    • Short proofs are easier to check; long proofs exceed current reliability
    • Models can sometimes detect their own short-proof errors; long ones are riskier
    • Lean formalization provides stronger evidence than informal writeups
    • Example: 800-page AI ‘proofs’ (e.g., major claims) are likely untrustworthy
    • Human validation relies on global-structure reasoning and special-case tests
  13. 59:33 – 1:03:11

    Raising a mathematician in the AI era: early numeracy, shapes, and preserving education’s deeper goals

    The conversation closes on parenting and pedagogy: how to instill love of math and robust thinking skills when AI is everywhere. Daniel emphasizes that learning math (and humanities) remains valuable for clear thinking and understanding the world, regardless of AI capability.

    • Daniel’s three-year-old is learning counting and single-digit addition
    • Encouraging playful math: shapes, solids, and curiosity-led learning
    • AI hasn’t entered the child’s life yet; intentional exposure matters
    • Education’s enduring purpose: clear thinking and world-understanding
    • Humans should use AI to deepen understanding, not outsource it

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.