Skip to content
a16za16z

What Today’s Best Models Still Can’t Do in Math

a16z’s Lisha Li sits down with Daniel Litt, Assistant Professor of Mathematics at the University of Toronto, to unpack AI's rapid progress in mathematics, what today's frontier models can actually do, and what they're still missing about the way mathematicians think. Daniel explains why some recent AI-generated results are genuinely impressive, including an autonomous solution to the Erdős unit distance problem, but argues that solving problems is only one part of mathematics. Today's models can grind through calculations, combine known techniques, and search enormous spaces, but still struggle with intuition, theory building, identifying the right questions, and developing the kind of big-picture understanding that drives much of mathematical progress. Lisha and Daniel also explore how AI is already changing mathematical research, why an explosion of AI-generated papers could distort academic incentives, and what happens if researchers outsource the work of thinking rather than use AI to deepen it. Ultimately, they ask a question that extends far beyond mathematics: as AI gets better at intellectual work, how do we make sure humans keep getting better at thinking too? Timestamps: 00:00 - Intro 01:00 - Meet Daniel Litt: A Practicing Mathematician's Evolving Views on AI 02:24 - The Most Impressive Result: The Erdős Unit Distance Problem 06:12 - What AI Is Actually Doing for Working Mathematicians Today 12:12 - Intuition, Taste & Why Math Isn't Just About Proofs 20:26 - Deep Thinking vs Pattern Matching: What Models Are Missing 26:17 - Why the Unit Distance Result Was Actually Creative 33:33 - How Should the Math Community Adapt to AI? 43:26 - Where AI Will Impact Applied Math First 46:24 - Taking Advantage of AI Without Losing the Craft 49:21 - Comparing Anthropic vs OpenAI in Math 51:34 - Why Some Labs Have Gone More Secretive 59:40 - Raising a Mathematician: Teaching Math to a Toddler Resources: Follow Daniel Litt on X: https://x.com/littmath Follow Lisha Li on X: https://x.com/lishali88 Stay Updated: If you enjoyed this episode, be sure to like, subscribe, and share with your friends! Find a16z on X: https://twitter.com/a16z Find a16z on LinkedIn: https://www.linkedin.com/company/a16z Listen to the a16z Show on Spotify: https://open.spotify.com/show/5bC65RDvs3oxnLyqqvkUYX Listen to the a16z Show on Apple Podcasts: https://podcasts.apple.com/us/podcast/a16z-podcast/id842818711 Follow our host: https://x.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see http://a16z.com/disclosures.

Daniel LittguestLisha Lihost
Sep 1, 20261h 3mWatch on YouTube ↗

At a glance

WHAT IT’S REALLY ABOUT

AI can prove some results, but lacks math understanding and taste

  1. Litt argues that mathematics aims at human understanding, not merely producing papers, and that “understanding in model weights” is unsatisfying as a replacement for human comprehension.
  2. He highlights an autonomous AI solution to the Erdős unit distance problem as unusually creative because it imported older techniques from an unexpected area and then enabled further human progress on related questions.
  3. Today’s best models excel at applying known techniques, grinding computations, coding, and searching examples, but remain weak at intuition, big-picture framing, theory-building, and “what is the right question?” discovery.
  4. Reliability and evaluation are bottlenecks: models tend to produce short proofs partly because long-horizon arguments are hard to check, and current models struggle with expert-style proof stress-testing and global error detection.
  5. AI is already distorting incentives via “slot-machine” theorem-proving and arXiv slop, so the math community will need new norms and incentives that reward understanding, verification, and human capability-building.

IDEAS WORTH REMEMBERING

5 ideas

The most meaningful AI math progress is measured by downstream usefulness, not headlines.

Litt values results like the unit distance breakthrough because the ideas transferred and helped humans find counterexamples and advance other problems; novelty matters less than whether understanding and tools generalize.

Current models are strong “technique appliers,” not autonomous theory builders.

They can combine known methods, compute relentlessly, and pull from broad literature, but they rarely originate the fuzzy, philosophical framing mathematicians use to decide what should be true and what to try next.

Human friction (getting stuck, refusing ugly grind) can produce better mathematics.

Litt describes how models failing on a lemma pushed him to explore examples, discover a better statement, and then let the model prove it—yielding more conceptual understanding than a brute-force proof would.

Short AI proofs are often a symptom of verification limits, not inherent elegance.

Long proofs are harder for models to keep correct and for humans to validate; labs preferentially release Lean-formalized or otherwise checkable results, while long unverified “proofs” are likely unreliable.

Proof evaluation requires global structure checks that models still lack.

Expert mathematicians often detect wrong papers by noticing implausible consequences or structural impossibility before locating the exact bug; Litt says models struggle with this kind of high-level stress testing.

WORDS WORTH SAVING

5 quotes

The goal of mathematics is not to produce mathematics papers. It's to produce some kind of understanding. Maybe some of that understanding resides in model weights. To me, that's, like, pretty unsatisfying.

Daniel Litt

I actually think the argument was very human.

Daniel Litt

But they're not, like they seem like weaker in things like intuition or like having some big picture point of view.

Daniel Litt

But you should like win by any means necessary- uh, in my opinion.

Daniel Litt

A lot of progress in mathematics comes from, like, letting, you know, a thousand different flowers bloom and people pursue their own curiosity.

Daniel Litt

Erdős unit distance problem as an AI milestoneAutonomy vs semi-autonomy in AI math resultsIntuition, taste, and theory-building in real mathematical workNatural-language reasoning vs Lean/formal verificationProof checking, long-horizon reliability, and stress-testingIncentives, arXiv “slop,” and mode collapse in research directionsEducation, human capital pipelines, and raising mathematically literate kids

High quality AI-generated summary created from speaker-labeled transcript.

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.