Skip to content
a16za16z

What Today’s Best Models Still Can’t Do in Math

a16z’s Lisha Li sits down with Daniel Litt, Assistant Professor of Mathematics at the University of Toronto, to unpack AI's rapid progress in mathematics, what today's frontier models can actually do, and what they're still missing about the way mathematicians think. Daniel explains why some recent AI-generated results are genuinely impressive, including an autonomous solution to the Erdős unit distance problem, but argues that solving problems is only one part of mathematics. Today's models can grind through calculations, combine known techniques, and search enormous spaces, but still struggle with intuition, theory building, identifying the right questions, and developing the kind of big-picture understanding that drives much of mathematical progress. Lisha and Daniel also explore how AI is already changing mathematical research, why an explosion of AI-generated papers could distort academic incentives, and what happens if researchers outsource the work of thinking rather than use AI to deepen it. Ultimately, they ask a question that extends far beyond mathematics: as AI gets better at intellectual work, how do we make sure humans keep getting better at thinking too? Timestamps: 00:00 - Intro 01:00 - Meet Daniel Litt: A Practicing Mathematician's Evolving Views on AI 02:24 - The Most Impressive Result: The Erdős Unit Distance Problem 06:12 - What AI Is Actually Doing for Working Mathematicians Today 12:12 - Intuition, Taste & Why Math Isn't Just About Proofs 20:26 - Deep Thinking vs Pattern Matching: What Models Are Missing 26:17 - Why the Unit Distance Result Was Actually Creative 33:33 - How Should the Math Community Adapt to AI? 43:26 - Where AI Will Impact Applied Math First 46:24 - Taking Advantage of AI Without Losing the Craft 49:21 - Comparing Anthropic vs OpenAI in Math 51:34 - Why Some Labs Have Gone More Secretive 59:40 - Raising a Mathematician: Teaching Math to a Toddler Resources: Follow Daniel Litt on X: https://x.com/littmath Follow Lisha Li on X: https://x.com/lishali88 Stay Updated: If you enjoyed this episode, be sure to like, subscribe, and share with your friends! Find a16z on X: https://twitter.com/a16z Find a16z on LinkedIn: https://www.linkedin.com/company/a16z Listen to the a16z Show on Spotify: https://open.spotify.com/show/5bC65RDvs3oxnLyqqvkUYX Listen to the a16z Show on Apple Podcasts: https://podcasts.apple.com/us/podcast/a16z-podcast/id842818711 Follow our host: https://x.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see http://a16z.com/disclosures.

Daniel LittguestLisha Lihost
Sep 1, 20261h 3mWatch on YouTube ↗

EVERY SPOKEN WORD

  1. 0:001:00

    Intro

    1. DL

      The goal of mathematics is not to produce mathematics papers. It's to produce some kind of understanding. Maybe some of that understanding resides in model weights. To me, that's, like, pretty unsatisfying.

    2. LL

      Comparing Anthropic with OpenAI, do you detect any differences in how that is similar to human reasoning?

    3. DL

      They definitely are not good at it autonomously. But with some hints, you can kind of get them to do something interesting. A lot of progress in mathematics comes from, like, letting 1,000 different flowers bloom, and people pursue their own curiosity, and then the boundaries of knowledge expand in some fairly uniform way.

    4. LL

      What has been the most impressive result so far?

    5. DL

      My favorite fully autonomous result by an AI so far is the solution to the Erdős unit distance problem. There was some lemma I wanted to prove. None of the frontier models could do it. So I, like, worked out a ton of examples on my own, and I realized, "Oh, well, maybe, like, here is some reason why it could be true." Once I had that statement, the models were able to very quickly prove that sort of better statement.

    6. LL

      How should, like, the mathematics community kind of best adapt and benefit from this?

    7. DL

      Um-

  2. 1:002:24

    Meet Daniel Litt: A Practicing Mathematician's Evolving Views on AI

    1. LL

      I am so excited to have you on, Daniel. Um, and so Daniel Litt is a professor of mathematics at the University of Toronto. Um, Toronto's my hometown, so, uh, also very exciting. Um, but the thing that is most, um, special here is Daniel's an actual practicing mathematician, and in addition, he's been incredibly vocal about, um, his evolving views of AI in math. And so, um, I feel like every-- If I just don't check in with you, you know, like in, in two weeks, something, you know, different has been revealed, and then you're, you're, um, very kind of like, uh, what do you call it? Um, s- so you-

    2. DL

      I have a lot of opinions.

    3. LL

      You have a lot of opinions, exactly. So I wanna get into that. So I mean, you know, one of the things, um, that I'm most interested in is not just, like, a discussion of how the capabilities have advanced. I feel like in math, it's, you know, that, that's definitely the, the headline, et cetera. But also, um, you've been very thoughtful about how practicing mathematicians should respond. Um, and so that kind of gives us a chance and opportunity to talk about actually what is special about math, you know? Like, it's not just like, "Hey, AI's been really making progress here," but, like, delve into what actually mathmis- mathematicians do. Um, and so maybe, like, with that, um, uh, arc in line, we can start with, um, what has been the most impressive result so far, uh, given all the recent progress, um, uh, for you? And then maybe, yeah, we'll kind of take it from there.

  3. 2:246:12

    The Most Impressive Result: The Erdős Unit Distance Problem

    1. DL

      Yeah. So okay. So there have, you know, now been a lot of results, some of them produced autonomously, some produced, like, semi-autonomously, some who's, like, the, where the AI contribution is just not at all clear. Um, they're in a lot of different areas, so anything I say is kind of, you know, I can only really comment on things that I'm, you know, cl- have some expertise on. So it's quite possible that if you talk to a different mathematician, you'll get, like, different answers here. Uh, so my favorite, uh, like, fully autonomous result by an AI so far is the solution to the Erdős unit distance problem-

    2. LL

      Oh

    3. DL

      ... which I think was announced in mid-May. Um, so at least what I liked about that is it seemed to me that, like, it was in some ways a little bit creative. So, uh, I think some of the results, uh, we've seen have had kind of the flavor of like, you know, you kind of take some known techniques and apply them in maybe a clever way, or you, uh, I don't know, they've kind of been, uh, some kind of results I would characterize as, like, last mile. Like where-

    4. LL

      Yeah

    5. DL

      ... some recent work was done, like quite deep work done by a group of human mathematicians, and then the AI kind of took the final step.

    6. LL

      Yeah.

    7. DL

      Um, but yeah, with this Erdős dis- unit distance problem, I, I think it was something where the result was, like, a little unexpected. So first of all, my sense was, like, that people working in the area thought it was true, and then there was a counterexample found. But then also, it, like, brought in some techniques for, from another area. I think those i- techniques were, like, not especially, uh, you know, like, deep or new. They were sort of classic ideas from the '60s, but they were new to this area of studying, um, you know, point configurations in the plane. And so that was pretty cool, and then afterwards, we got to see, like, it was kind of fruitful. So-

    8. LL

      Hmm

    9. DL

      ... a bunch of mathematicians took those ideas and used them to find counterexamples to a bunch of other interesting open questions. So for example, like the sum-product conjecture for the real numbers. Um, so that's what, at least one way I like to think about how cool-

    10. LL

      Yeah

    11. DL

      ... the result is. Like, you look at it post hoc, and you see, like, "Oh, well, you know, were whatever new ideas that were introduced, if any, like, kind of useful to do other things? Like, did they improve our understanding of something?" And I think that's maybe so far the, the, the main example I know of, of, uh, of a result of that form.

    12. LL

      Yeah, I think it's really meaningful that you're commenting on this because in, um, you know, the, that result came out, to your point, in May, and there's been so much, kind of, like, so many headlines, uh, so far. And it's kind of probably hard for somebody who's not a practicing mathematician to appreciate the differences in these headlines. And so you kind of already started laying out sort of a ta- taxonomy of, like, what is different in that proof. And so it'd be kind of interesting maybe to use as, that as kind of both an excuse to talk about where you sense the model differences are and what, like, mathematicians actually do. So in this case, I mean, the most kind of maybe naive understanding of what mathematicians do is that we're pushing around symbols in a logical manner, and this is why, you know, RL is so successful at this because you can kind of both verify it somewhat cheaply, um, compared to other domains, and then also, you know, because, like, the rules are quite, um, uh, quite legible. And so you can kind of like, you know, if you're, if you're superhuman at that, you might be good at math. Um, but I think that, of course, betrays most of actually what is mis- uh, what is interesting about mathematics, which is perhaps, I think you said this as well, but I think anybody who's tried to do math is sort of like, it's about the understanding and sort of getting at truth, remaining confused, and developing intuitions. And a l- very lit- I mean, the tool to do that stuff is, of course, in having really s- good, uh, v- very, uh, strong abilities to push out logical, [chuckles] you know, implications. Um, but maybe if you can kind of talk to, speak to like, you know, when you say it's most impressive and creative, like decoupling the, um, just the ... inhumane maybe feats of just like logical, um, implication from like where is it being

  4. 6:1212:12

    What AI Is Actually Doing for Working Mathematicians Today

    1. LL

      creative? What is kind of, um, what is it helping, uh, engender in ter- in terms of like mathematical activity as well?

    2. DL

      Yeah. Okay. So first of all, I mean, I, you, you kind of characterized it as like inhuman in some way. I actually think the argument was very human.

    3. LL

      Ah.

    4. DL

      Like even, you know-

    5. LL

      Oh, I'd love to get into that

    6. DL

      ... ChatGPT, or sorry, OpenAI, like released some chain of thought.

    7. LL

      Yeah.

    8. DL

      It was very recognizable. It was like kind of, you know, if you, if I tried to imagine like my chain of thought, uh-

    9. LL

      Yeah

    10. DL

      ... in trying to solve a problem, like it might look kind of like that.

    11. LL

      Yeah.

    12. DL

      Um, you know, we haven't seen the raw chain of thought, so maybe, maybe that would-

    13. LL

      [chuckles] Maybe they, uh, cleansed it a little bit. Yeah.

    14. DL

      Yeah. Or maybe, you know, maybe the model liked to swear a lot in the middle of the chain of thought and they cleaned that up or something. We'd know. But like, at least the summary seemed pretty recognizable. And I would say that's actually like kind of typical of, uh, most of the results that I've studied. Like they don't seem inhuman at all. They seem absolutely like something a human mathematician could produce. Um, and, uh, they're like typically understandable if like not so well-written. If you just like look at raw model output, it's not like there's some, you know, move 37 or whatever. Um, it's like a human-

    15. LL

      Yeah

    16. DL

      ... mathematician doing math. It's like a human mathematician doing certain types of math. So like they're definitely, the models are still like not very good at some mathematical activities, and here I don't mean like by field, but just like certain things you do when you try to solve a problem, the models don't seem to be doing. But certain things they're very good at. So, um, the ways they might be a little bit un- inhuman is like they don't get tired, they know a lot, um, but it, it, if you actually just read the final output, it doesn't seem kind of inhuman.

    17. LL

      It's interesting because, I mean, obviously we can't get too much of, uh, information from the labs who are producing these models of like why, um, uh, you know, how, of, of what the training recipes are or how they're kind of advancing in reasoning. Um, but at least one of the things we do know, um, and I think kind of OpenAI spearheaded this, is just like reasoning in natural language is actually what they-

    18. DL

      Right

    19. LL

      ... happen to scale up, and it's not actually pushing a lot of lean verified proofs as like the training corpus, and that's kind of amazing. Um, on another side, what's kind of interesting is it's not clear that a lot of the mathematical, um, training data, if they use that to a large extent at all, is reflective of how mathematicians think, if that's fair. Because a lot of it is not like-

    20. DL

      Yeah

    21. LL

      ... legible as traces, right? Like most of the papers are crisp and like polished. The textbooks certainly are just n- show very little motivation of how something is developed, which is why it's like usually easier to kind of follow research direction by actually talking to the researchers and how they're thinking about it. So I'm kind of curious, before maybe even going to the taxonomy as an excuse, like when you're examining these, uh, models and their results, comparing Anthropic with OpenAI, do you detect any differences in, in how that is similar to human reasoning? And then also if you have any comments or like insights on perhaps why natural language scales so well that way, even though it's-

    22. DL

      Okay. So, so first of all, I, I really like your point, by the way, that, that they're mostly doing natural language reasoning rather than lean. Like I think that-

    23. LL

      Uh-huh

    24. DL

      ... that suggests to me that like these, you know, you hear a lot of people say like math is a verifiable domain, like that explains these, the progress or whatever. Like my sense is that because they're, you know, primarily scaling informal reasoning, like probably the techniques are gonna generalize to other domains pretty well. That's just my guess.

    25. LL

      Mm-hmm.

    26. DL

      Okay. Um, you asked a little bit about like Claude versus-

    27. LL

      Yes [chuckles]

    28. DL

      ... ChatGPT. Um, my sense is that they're pretty similar in terms of capabilities. I've, I've played a lot, around a lot more with ChatGPT than, than Claude Fable, but, you know, uh, um, it, it seems like there's a lot of cases where, you know, uh, OpenAI will drop a solution to some problem and Anthropic s- will say, "Oh, you know," Fable-

    29. LL

      We also [chuckles] Exactly

    30. DL

      ... can do that too, or... So it actually seems like they're st- they're solving a very similar collection of problems, and it's like kind of a small, you know, uh, it's like a relatively small portion of what human mathematicians do. Um, so yeah, one thing, um, that is interesting is like we see the solutions have a certain flavor, right? Like-

  5. 12:1220:26

    Intuition, Taste & Why Math Isn't Just About Proofs

    1. LL

      Okay. I would love to, so go into the, um, the intuition part and where it, um, sucks at. Um- To, to put it in a very basic way. But, um-

    2. DL

      Sure

    3. LL

      ... you actually, uh, mentioned a small detail, which is that, you know, from the public you've gleaned that Anthropic and OpenAI are probably neck and neck, but you personally are using a lot more, um, ChatGPT. Like what, why is that? Or, uh, like 5.6 now. [chuckles]

    4. DL

      Uh, well, the way, why is... I don't know. I mean, I just te- I, I think it's just inertia. Like I have-

    5. LL

      Creature of habit, yeah.

    6. DL

      Um, it got... W- uh, one thing is that ChatGPT got better at math earlier, so-

    7. LL

      Yes

    8. DL

      ... like for a long time, um, the Claude model, models were just like not useful for research math.

    9. LL

      Yeah.

    10. DL

      And then I think maybe around Opus 4.5 or Opus 4.6, they like more or less caught up.

    11. LL

      Yeah.

    12. DL

      Uh, but you know, some experimentation suggests to me they're pretty neck and neck, and so for my own work, you know, except when I'm just experimenting, I mostly just stick with one.

    13. LL

      Yeah.

    14. DL

      And it's sort of at random.

    15. LL

      A- as people outside of labs like us, it's really interesting just to compare how they differ on the frontier and, I mean, y- you know, to your point, it might be a little bit of momentum. I do think that, at least from my anecdotal experience, 5.6 has been like a lot more clear in exposition, and it just like, it's, there's a little bit, and this might not be true, I mean, obviously the, the models are just incredibly jagged at the frontier. Um, but like it's, in the explanations of, uh, results to me, I, I always find that 5.6 is giving a, a more accurate theory of mind of what it assumes I know and don't know. [chuckles] Whereas Fable might be explaining something very trivial, but then just like jump at like, "Well, you know, obviously you should know these ki- things and um-"

    16. DL

      Yeah. I'm not, I find they're both pretty bad at theory of mind.

    17. LL

      [laughs] Okay, great. So then seeing it from, in your eyes, it's, you're probably asking much deeper questions. Um, okay. So, um, the thread that I really wanted to pull on was when you were talking about, um, the models maybe starting to get better intuitions or even theory building. So maybe before we even dive into that, it would be useful to kind of talk through like what is your primary, you know, activity as a mathematician, um, especially in your area of like algebraic geometry. Probably has a very different flavor than a combinatorialist or, you know, um, some other area. So if you can give a little maybe brief lay of the land and then kind of explain what your ac- what your mathema-math-mathematical activity was pre-AI, and then maybe how it's kind of like changing with AI.

    18. DL

      Yeah. So, so I, I mean, I think there are a lot of different kinds of mathematicians. Um, there's like a lot of different, uh, you know, spectra on which one can put a mathematician. Um, so, uh, definitely a lot of mathematicians like solving open problems.

    19. LL

      Mm-hmm.

    20. DL

      Uh, and then I think, you know, I, I'm one of those. Like I like to solve an open problem. That's, uh, I think of myself as a problem solver as opposed to like one, you know, one other taxonomy you could, you could have is a problem solver versus a theory builder.

    21. LL

      Mm-hmm.

    22. DL

      Um, at least for me, uh, the point of an open problem is it's like supposed to measure your failure to understand something. So it's kind of like a benchmark.

    23. LL

      Mm-hmm.

    24. DL

      Um, right? Like, uh-

    25. LL

      Mm-hmm

    26. DL

      ... you know, uh, the one problem I really like is the Gross and DeCap Speed Curvature Conjecture, whatever that is. It's, it measures something about our, our failure to understand differential equations. So it's like there's some very basic object we would like to understand. If we can't answer this conjecture, we know we don't understand it.

    27. LL

      Mm-hmm.

    28. DL

      Okay. So in practice, like how do you get a, a, a problem which is supposed to be measuring something you don't understand? Well, like, of course, you try to understand the thing better. Um, and, and in practice, what that means is like, well, you try to find the smallest situation where you can't understand something and, and, and fu- fiddle around with it, and then you stare at like what, once you win, you stare at like what you developed to win and try to turn that into some theory. Um, so that's one thing you might do. You might try to like solve a problem and in so doing, develop some kind of new understanding of the situation. Uh, you might also just like have some feeling that like this thing is related to this other thing. Um, and you know, you might start building a table like, oh, this property A is related to property A prime, property B is related to property B prime, and so on and so forth. So, uh, for example, uh, in, in my work, a, a lot of it is motivated by some analogy between cohomology of algebraic varieties and representations of fundamental groups. Okay. So that's some, some fancy stuff. Uh, but it's just that this analogy is like very, very fruitful and like really any phenomenon that appears on one side, you can find an analog on the other side. And so tr- tr- you know, trying to, uh, to realize that dream has led, led to a lot of beautiful mathematics over the last like 30 or 40 years, um, by people like Carlos Simpson and, and Takuro Mochizuki and others. Um, and so here it's really just like someone noticed like here's an analogy, and then that analogy has like led to a huge amount of development. There's not really like an open problem at the end, although of course, like as you develop this, you come up with lots of open problems.

    29. LL

      Mm-hmm.

    30. DL

      There's just like a philosophy that you're trying to, uh, you know, you're trying to realize, and that philosophy is like super not rigorous actually.

  6. 20:2626:17

    Deep Thinking vs Pattern Matching: What Models Are Missing

    1. LL

      solver. Or like what is the thing that is... Because like the deep thinking part, maybe, maybe like, you know, kind of for this audience, it might also be useful to kind of say or explain, you know, uh, most of pure math is, it's not like it's being motivated by... And, you know, nothing on applied math, but applied math is like at least there's some external motivation for why a certain formal structure might be interesting to study. Whereas this one, it's like purely, it seems almost sociological and to some extent I think like Thurston made some comment in, or a point about this in the '70s, that it is a sociological phenomenon of more and more mathematicians start examining something, then you'll maybe, you know, uh, converge on some interesting structures. But it's just like it's not... It, there's some reason why people would prefer to study something or they think it's beautiful. Like what is it that drives maybe you in particular, and then maybe you can make a more general comment about the profession?

    2. DL

      Yeah. I mean, so definitely some people are like motivated by like beauty or, or, or kind of aesthetic considerations.

    3. LL

      Like what is that, right?

    4. DL

      I, I try not to be motivated by that.

    5. LL

      And you try not to. [chuckles] Oh, that's more commercial. Yeah.

    6. DL

      Yeah, I think that's like a bad way.

    7. LL

      Yeah.

    8. DL

      And like, well, one thing, like that sort of limits you, right? Like one-

    9. LL

      Mm

    10. DL

      ... thing, uh, kind of a failure mode I see among young mathematicians sometimes is like you, you have something and you kind of think you know how to prove it, and then-

    11. LL

      Mm-hmm

    12. DL

      ... like the proof feels really ugly.

    13. LL

      Mm-hmm.

    14. DL

      And you decide... But like, okay, I mean, what if you're wrong and it's not ugly? Like what-

    15. LL

      But what's ugly then?

    16. DL

      ... why, why limit yourself? You should-

    17. LL

      For this audience.

    18. DL

      What?

    19. LL

      What is ugly? 'Cause I have an, uh, intuition-

    20. DL

      Yeah, like-

    21. LL

      ... of what's ugly, but like what is that for kind of spelling it out?

    22. DL

      Well, I, I don't, I don't really know. I, I don't really have aesthetic feelings, but people, people sometimes feel this way.

    23. LL

      Yeah.

    24. DL

      Like maybe it involves a lot of grinding-

    25. LL

      Yeah

    26. DL

      ... or some calculation that's not illuminating or, or something like that.

    27. LL

      Exactly.

    28. DL

      But you should like win by any means necessary-

    29. LL

      Yeah

    30. DL

      ... uh, in my opinion. Like I, I like to think of what I'm doing, like doing kind of physics except with concepts.

  7. 26:1733:33

    Why the Unit Distance Result Was Actually Creative

    1. LL

      porting it over. And, and to your point, that's why maybe the unit distance problem was such a more creative result 'cause it was maybe doing more of that on its own.

    2. DL

      It was, like, from another. At least it was, like-

    3. LL

      But it was-

    4. DL

      ... drawing in something from-

    5. LL

      Yeah, it was

    6. DL

      ... an unexpected area. Yeah.

    7. LL

      Exactly, exactly. And so, like, maybe then to kind of ask the question, it's not that, 'cause now we're like, "Okay, fine, AI's getting so good so fast, we can't count it out." But, like, why, what do you think it has to, um... Yeah, I guess it's, it's, uh, you're spelling out what it has to get better at, but maybe some more-

    8. DL

      Yeah

    9. LL

      ... kind of feelings on, like, why it's kind of particularly hard to then develop that new theory and technique.

    10. DL

      Um, yeah, that's a good question. I mean, I think you just need a different... Like, my, my guess actually is it's probably totally doable, um, and it just hasn't been done yet. Like-

    11. LL

      Yeah

    12. DL

      ... maybe you just need a different RL environment. Um, I don't know. But yeah, so I, and at this point-

    13. LL

      Good point. Yeah

    14. DL

      ... my expectation is just, like, that the trajectory will continue upwards. I'm not a, a skeptic of continued capabilities growth.

    15. LL

      Yeah.

    16. DL

      Um, but yeah, you know, I think what is definitely true is that, like, the skill of, like, developing a theory, um, or, like, building your understanding of some poorly understood object is, like, a fuzzier one.

    17. LL

      Mm-hmm.

    18. DL

      Uh, so it m-might be harder, you know, uh, I guess you can try, you can tell it, you know, uh, develop your understanding of zeta functions, and then once it proves through your hypothesis, you give it a reward. But it's like-

    19. LL

      Yeah

    20. DL

      ... kind of harder, I think, to come up with some, like, intermediate-

    21. LL

      No. Yeah. [chuckles]

    22. DL

      ... things that you can reward. Um.

    23. LL

      Yeah.

    24. DL

      Yeah, I mean, that said, you know, uh, I do think, uh, [chuckles] mathematics as a whole provides a lot of conjectures of varying levels of difficulty.

    25. LL

      Mm-hmm.

    26. DL

      So, you know, maybe, uh, maybe this explains why there seems to be a little bit of progress in these areas. Like, presumably, they are trying to, you know, get it to solve lots of problems, and some of those problems develop at least some of the skills of theory building. And humans are able to develop these skills. Um, you know, I guess sometimes they get rewards from their PhD advisors, and their advisor says, "Oh, that's a good idea," or something based on some element of human taste or whatever.

    27. LL

      Yeah.

    28. DL

      Uh, and that, that might be something one can do too. But yeah, my, and my, my expectation is that just, like, as part of continued capabilities growth, we'll see growth in these areas too.

    29. LL

      Yeah, yeah. And I like the framing where you're sort of casting these increasingly difficult conjectures as a, as a, you know, f-form of curricula for both humans-

    30. DL

      Yeah

  8. 33:3343:26

    How Should the Math Community Adapt to AI?

    1. LL

      How shou- how should, like, the mathematics community kind of best, um, adapt and benefit from this? I mean, as a, just kind of saying as somebody who's, you know, doesn't have the time to kind of practice mathematics anymore, this is kind of ver-very great because I can maybe dabble more. There's a lot of-

    2. DL

      Mm-hmm, mm-hmm

    3. LL

      ... results that can come out. But I can also see where, you know, you've made, uh, the more precise point of we can not motivate the right kind of behavior of understanding and development. And so I would love to kind of hear more about, um, your views there.

    4. DL

      Yeah. So, okay, so first of all, like, it is clearly really exciting, like, that there are increasingly capable models that are, like, producing high-quality results. Um, you know, at least, you know, some high-quality results. Also, also a lot of slop.

    5. LL

      Yeah. [laughs]

    6. DL

      Uh, you know, some, some good stuff. Um, and so, you know, as the models get really capable, you know, my hope is that they will answer a lot of the questions that I've been, like, you know, kept up at night thinking about and, like, I'll, I'll get to learn the answers. I think that's really exciting and, you know, a lot of people got into math, uh, like largely 'cause they enjoyed learning math.

    7. LL

      Yeah.

    8. DL

      Um, right, like the first thing you do as a, a math student is you, like, learn stuff that other people did, and you do that for like 20 years-

    9. LL

      Mm-hmm

    10. DL

      ... before you, before you start doing... Well, maybe not quite 20 years, but-

    11. LL

      But, yeah

    12. DL

      ... you know, 15 years.

    13. LL

      If you're lucky, 20 years, yeah. [laughs] If you started early. Yeah.

    14. DL

      Yeah, yeah. Um, yeah. So, um, that said, like the goal of mathematics is not to produce mathematics papers. Um, like it's to produce some kind of understanding. Uh, so maybe, okay, maybe that some of that understanding resides in model weights or something. To me, that's like pretty unsatisfying. Like-

    15. LL

      Mm-hmm

    16. DL

      ... uh, my own, you know, my motivation for doing mathematics is, like, I would like to satisfy my own personal curiosity. Um, I think people should be, uh, have the ca-capability to do that. And, like, that requires a pretty substantial apparatus. Like, it's, like, simply not the case that you can, like, study the questions I think are fundamental unless you've invested a huge amount of time and effort, um, kind of getting to the point where you can meaningfully do so.

    17. LL

      Mm-hmm.

    18. DL

      And then moreover, like that, you know, even the people, like, you know, you have the small group of people doing like fancy research math on the frontier or whatever, um, that relies on like a huge apparatus of like, you know, thousands and millions or billions of people who are trying to learn to think mathematically. Like you need a, an entire mathematical community to support a small group of people who are on the frontier. Like you just, without the pipeline-

    19. LL

      Mm-hmm

    20. DL

      ... the pipeline doesn't exist. Uh, so if you think that's important, like development of human capital that can like meaningfully engage with frontier mathematics, I think that, uh, you know, you still have to incentivize those people to actually invest their time and effort getting to the point where they can engage and then do so in a, in a high quality, meaningful way. Uh, so right now I think like the existing incentive structures for math research do not, uh, do not encourage people to do that. Um, so, you know, right now if you're like a postdoc on the market You wanna get a job maybe for the next couple years before the community adapts.

    21. LL

      [sighs]

    22. DL

      Uh, the best way to do that is, like, if you wanna produce a lot of papers which maybe prove, you know, old conjectures or whatever. And you can do that by playing the slot machine until-

    23. LL

      Yeah

    24. DL

      ... um, the model produces a hopefully correct proof of such a result. Um, so you don't even have to pick the, the p- the, the theorem in advance. So, like, here's an experiment you can do. You can take Codex, you can say, "Go online and find five recent conjectures in algebraic geometry and prove them." And, uh, okay, I've run this experiment, and with some back and forth I was able to, you know, in an hour get, like, three, you know, quite bad papers.

    25. LL

      Mm-hmm. [chuckles]

    26. DL

      But correct papers. Um, which, okay, are now sitting on my hard drive waiting for me to email the relevant people. But, you know, this is not, okay, good use of my time to invest in-

    27. LL

      [chuckles]

    28. DL

      ... these results.

    29. LL

      Yeah.

    30. DL

      Uh, but yeah, you, you definitely see people doing this. Um, so there's, you know, been a huge uptick in post to archive.

  9. 43:2646:24

    Where AI Will Impact Applied Math First

    1. LL

      various things. I mean, this is not why mathematicians do it, um, but as just somebody who like-

    2. DL

      I think it's part of why we do it.

    3. LL

      Okay, great. [chuckles] 'Cause I, it was kind of w- and you know, I, I thought it just really helped give me a very, um, good framework to think about many things, not just mathematics. And as somebody who's also, you know, now a parent of a two-year-old, I kind of think about this, uh, a lot too. You know, it's, it's not about kind of grinding, even though, you know, whatever, it's, uh, still good. [chuckles] But, like, not to-

    4. DL

      Yeah

    5. LL

      ... not to, um, shit on grinding too much. But, you know, just, uh, hearing, uh, I, I used to collaborate a bunch with, uh, some Hungarian mathematicians, um, Balázs Szegedy among them, and I j- I heard that, um, in, in Budapest, they would just teach group theory when you're in primary school. And I'm like, well-

    6. DL

      Mm-hmm

    7. LL

      ... we should definitely do that. We should continue doing that. Now that AI is-

    8. DL

      Yeah

    9. LL

      ... so good at, you know, some- somewhat good at explaining, but it's far more accessible, we should actually, you know, probab- proliferate, uh, that even more. Um, and so maybe that helps with bring more people to the frontier rather than, you know, just-

    10. DL

      Yeah, I mean, this is something I'm concerned about, right? Like, I mean, of course, I mostly talk about math 'cause that's, like, where I live.

    11. LL

      Yeah.

    12. DL

      But, like, I think, you know, one nice thing about thinking about this is that we're one of the first, uh, professions, uh, to, to kind of, you know, have a significant impact of high-quality models. Although I think, um, maybe we're one of the first professions for it to happen so publicly. But, you know, my sense is that there are plenty of other professions that are similarly impacted.

    13. LL

      Coding. [chuckles] 100%.

    14. DL

      Coding, but also just, like... I mean, I think that, you know, anything you do at a computer, like, probably a huge amount is being done by the models at this point.

    15. LL

      Yeah.

    16. DL

      And then, like, there's not been a public reckoning about it, but-

    17. LL

      Yeah

    18. DL

      ... um, you know, because, uh, [chuckles] math capabilities are, are useful for the company, for the mo- labs to talk about. I think-

    19. LL

      Yeah

    20. DL

      ... we've been a bit more publicly than everyone else. Um, but yeah, I mean, I think, uh, you know, one reason to try to maintain, like, human capital in this area is just, like, it's a model for all professions. Like, presumably, we still want people who are, like, meaningfully engaging with the world and, like, experts and, you know, have, like, talents and trained skills and so on. And like, um, yeah, so at least, uh, yeah. It's actually very convenient that the math profession is so entwined with education here, 'cause like, I think we're also seeing, like, you know, uh, some amount of, uh, challenges, um, you know, among, uh, college and, and-

    21. LL

      Mm-hmm

    22. DL

      ... uh, like secondary education coming from AI too. So, as you said, I mean, it's also an amazing tool to learn, so.

    23. LL

      Yeah.

    24. DL

      Uh, you know, people have, have... I, I, I've heard, heard people start talking about a bimodal distribution in their classes, where there's some people who are really, like, figuring out how to take advantage of new tools and other people who are just, like, letting them do their homework and then bombing everything else.

    25. LL

      Yeah. Unfortunately, I don't think that adapts fast enough. But it's like we definitely wanna be living in a world where we're producing better thinkers. I think it's just-

    26. DL

      Yeah

    27. LL

      ... you know, even if we talk about just the pure kind of optimization game, I think that's better for us. [chuckles] And so-

    28. DL

      Right

    29. LL

      ... um, but as, you know, just like a human being, I'm like, that would be pretty, um, inconvenient if we became worse thinkers just as AI ascends, and it's too easy to let that happen. [chuckles]

  10. 46:2449:21

    Taking Advantage of AI Without Losing the Craft

    1. LL

      So we should kind of be thinking hard on how to actually take advantage of this and harness it for our own improvement-

    2. DL

      Right

    3. LL

      ... as well.

    4. DL

      Yeah, I think it's like w- it's sort of interesting to observe, like, the models, um, right at their current level of capabilities let you do a lot more. They let you do a lot of things you wouldn't have done more cheaply than, uh, you know, cheaply enough to do them now. But it's not clear to me that, like in many cases, they're actually improving the quality of outputs.

    5. LL

      Yep.

    6. DL

      Um, and I think this is common, like you have a new technology that's doing something a little bit worse than was previously done, but much cheaper. And so you get a lot of, suddenly a lot of, like, low-quality outputs that are displacing high, previous high-quality outputs. But I think it's possible to use the tools in a way that actually, like, improves, you know, the quality of, among, along every dimension. It just refers, re- requires some thoughtfulness and some, some redesign of institutions to actually incentivize that.

    7. LL

      Yeah. Well, hopefully capitalism works there. I do feel like that the most-

    8. DL

      Right

    9. LL

      ... high-value things do require people to use it effectively. And, and right now the models are not good enough without, like, the human experts to actually participate. Um, but to your point, there's a vast majority of maybe, like, more junior and, and entry level, um, kind of... You know, it's, it's harder for, um, for those, uh, roles to adapt as well. And so the thing that would be a mistake is to use the AI models in a way that doesn't... Basically, you need to be ascending and using the models to-

    10. DL

      Right

    11. LL

      ... deepen your understanding. And it, it's so easy for human nature just to be lazy, and you have to resist that, because that is the moment that you will kind of, kind of lose, [chuckles] basically. And so you kind of, you know, kind of forgive the very co-competitive language, but it really is that, like, it's just so easy to, to kind of relinquish the thinking to the models. The models can't really think. And so as things are ascending so fast, it's, like, critical that, um, you continue developing those facilities and actually leverage it to, to, to improve those facilities rather than relinquish. Yeah.

    12. DL

      Yeah. Yeah, I mean, one, one thing I, I think has been nice about, you know, sort of this vast increase in, in, like, semi-expert attention or, like, model attention on math problems is, like, now there's been... You know, okay, I've been complaining about slot papers or whatever.

    13. LL

      Yeah.

    14. DL

      Like, you know, people who are not, um, you know, producing really high-quality stuff. Often that's actually coming from professionals, like it's not- I'm not saying like, like, you know, there are people who, you know, the, like within academic mathematics, there are incentives to produce like a lot of-

    15. LL

      Yes

    16. DL

      ... a lot of stuff, and, and that's, you know, that's one place that slop comes from. So there's definitely also like slop coming from non-experts, but that I, I kind of actually don't see as a net negative. Like, okay, there's a lot of... Now there's a lot of like documents on the internet one might have to comb through to figure out if a problem has been solved or not.

    17. LL

      Mm-hmm.

    18. DL

      Um, but to me it seems like just the fact that there's lots of people excited about math is like kind of a positive, so that's like a nice thing.

    19. LL

      I totally agree, and I get to talk to you about it.

    20. DL

      And not even kind of a positive, obviously a positive. It's great.

    21. LL

      Yeah, yeah. No, exactly. It's like suddenly there's a spotlight on it and I can, um, nerd out about math more. Um-

  11. 49:2151:34

    Comparing Anthropic vs OpenAI in Math

    1. DL

      Yeah

    2. LL

      ... I, I was actually kind of curious, um, if you have any comments on the, um... I guess this is another constructive, uh, result, but the elliptic curve of rank 30. That just came out yesterday.

    3. DL

      Oh, yeah.

    4. LL

      So if that-

    5. DL

      That's right.

    6. LL

      Yeah, yeah. Uh, not-

    7. DL

      That's cool. I mean, we don't have any details about it yet.

    8. LL

      I know. There's rumors and it's-

    9. DL

      Yeah, we don't know how it, how it... I mean, so it was, it's due to, uh, I guess Claude Fable, um, prompted by Levent Alpoge and, um-

    10. LL

      Right

    11. DL

      ... a collaborator whose name I unfortunately forget. Maybe you can-

    12. LL

      Is it Levan?

    13. DL

      ... add it in post.

    14. LL

      Yeah.

    15. DL

      Um-

    16. LL

      I just know the Twitter handle. [chuckles]

    17. DL

      Yeah.

    18. LL

      Yeah, we can add it in.

    19. DL

      Uh, so yeah. I mean, uh, so a, a lot of these nice recent results have, have come from Levent with unclear amounts of autonomy.

    20. LL

      Mm-hmm.

    21. DL

      Um, so my sense is that, uh, some of them are semi-autonomous rather than fully autonomous. Um, yeah, with this, we don't know anything about the methods. Uh, so yeah, this is a fun construction. Um, without knowing about the methods, it's very hard to say how significant it is.

    22. LL

      Yeah.

    23. DL

      What I would say is that if you want to understand, um, kind of historically how such results have been proved, or sorry, have been, uh, uh, have been understood by the community, they're like cool, but I wouldn't say they're like a big deal. Um, so like the typical place a result like this might go is like someone's website of records.

    24. LL

      Mm-hmm. [chuckles]

    25. DL

      Um, but not-

    26. LL

      It's not an annals level result.

    27. DL

      Yeah, it's not an annals. But, but that said, it's cool and it's like ama- You know, it definitely, there were a few very, very, or are a few very, very talented mathematicians who like these kinds of questions.

    28. LL

      Yeah.

    29. DL

      So Noam Elkies being-

    30. LL

      Yep, yep

  12. 51:3459:40

    Why Some Labs Have Gone More Secretive

    1. DL

      uh.

    2. LL

      Have they been mostly more secretive? I ca- I think they've released some traces for stuff, but-

    3. DL

      Yeah. So for this one, I think they haven't yet, unless I missed it. Yeah. Uh, they, they, they... You know, Levent likes to tweet out his, uh, his results.

    4. LL

      I think.

    5. DL

      But yeah, he has been, you know, slowly releasing some kind of PDF write-ups too. I think it's, you know, he's just having fun on, on, uh, fun on the internet.

    6. LL

      Yeah, yeah. Um, it's, it reminds me of, uh, when you're saying, you know, you have to evaluate how the results came. Um, so, um, I had, um, Mark Zel- Zelky and, um, uh, Mehtab Swani on from OpenAI recently.

    7. DL

      Mm-hmm.

    8. LL

      And, um, they were saying how, you know, what's kind of been the most charming or delightful is just that the proofs have been relatively short. They're not like 200 page. And, you know, maybe, uh, corresponding to your grinding point as well. Um, but maybe this is kind of optimized for, in retrospect, it was picked that it was short? Or maybe do you find that on average, uh, stuff that you throw at, um, GPT, you know, Sol or Fable tends to be shorter and more kind of legible to the human, uh, or versus it might just go haywire and just grind it out?

    9. DL

      Yeah. So, so I mean, I, I, I think it is nice that they'll sometimes produce short, clever proofs.

    10. LL

      Yeah.

    11. DL

      Um, so that's, uh, of course, everyone likes a short, clever proof. Um, I think my sense is that the, uh, reason they're not producing long, complicated proofs is that they cannot.

    12. LL

      Mm.

    13. DL

      Um, like the, just like the, the ability to check correctness is not yet there. Um, so, uh, you know, I mean, even actually, you know-

    14. LL

      Yeah

    15. DL

      ... if you ask the models to produce a short proof, you, you can then often ask them to, you also just ask, "Is that correct?" And they will often say no. [chuckles]

    16. LL

      [chuckles]

    17. DL

      Right? So like they're much more reliable than they were six months ago, for sure.

    18. LL

      Yeah.

    19. DL

      But like it's still, uh, you know, they will still sometimes produce things that are just wrong, and they, they know they're wrong.

    20. LL

      Yes. Yeah, yeah.

    21. DL

      Um, I think the problem with producing a very long thing is they might not know they're wrong. Um-

    22. LL

      Mm-hmm

    23. DL

      ... and so what I wonder if, uh, you know, presumably internally, OpenAI and, and Anthropic have probably solved a lot more problems than they've released.

    24. LL

      Mm.

    25. DL

      And I imagine quite a few of them are, they're just not sure if they're true.

    26. LL

      Mm-hmm.

    27. DL

      Um, so for example, uh, you know, with these, this recent list of 10 problems, uh, released by OpenAI, those were all formalized in Lean.

    28. LL

      Mm-hmm.

    29. DL

      Which is, of course, a very good, good evidence that they're true.

    30. LL

      Yep.

  13. 59:401:02:57

    Raising a Mathematician: Teaching Math to a Toddler

    1. LL

      thought a lot about. Um, I think you also have a toddler, right?

    2. DL

      Yeah.

    3. LL

      Yeah, yeah. So how have you-

    4. DL

      Yeah, I have a three-year-old. Yes.

    5. LL

      Yeah, a three-year-old. Great, great. So you're one year more advanced and probably, you know, have more thoughts on this. Like, how are you thinking about, um, is it her education or his education?

    6. DL

      Her, yeah. Her.

    7. LL

      How are you thinking about her education in, in math and, you know-

    8. DL

      Yeah. Um-

    9. LL

      ... not to grind, but, [chuckles] you know, really just like how, you know-

    10. DL

      Yeah

    11. LL

      ... to pass on the love of it and, and how to, how to react to AI.

    12. DL

      Yeah. So she's, she's three. She's, she's never used, uh, AI.

    13. LL

      Good.

    14. DL

      Um, she, she is starting to add. That's, that's about as far as we are in math.

    15. LL

      Yay. Oh, that's, that's more far along than my two.

    16. DL

      Yeah. So she can, she can add single, single-digit numbers, like, but, you know, by counting on her fingers and-

    17. LL

      Nice

    18. DL

      ... count up to, like, maybe 30 reliably and 50 semi-reliably. So I'm very proud of her.

    19. LL

      That's good. Yeah, yeah.

    20. DL

      So yeah, I definitely enc-encourage that. We talk about shapes and stuff. Um, uh, [chuckles] actually, a couple days ago, uh, I woke her up, and she was, like, hiding under the blankets.

    21. LL

      Aw.

    22. DL

      And I was like, "Oh, what are you up to under there, Sophia?" And she was like, "Oh, I'm doing some math."

    23. LL

      Oh, I saw that. She-

    24. DL

      Uh, it was super cute

    25. LL

      ... I think you tweeted about that. That was adorable. [chuckles]

    26. DL

      Yeah, I did. Yeah, it was great. So, you know, I think she has, like-

    27. LL

      Yeah

    28. DL

      ... some sense that I like math, and she's into it because of that.

    29. LL

      Aw.

    30. DL

      Um, yeah. I mean, I, I would say I, uh, don't know. You know, I think the world is probably gonna look pretty different in, you know, 20 years or whenever she's kind of, uh, you know- ... fully adult and, and doing her own thing. Um, but you know, I think a lot of what we educate people for is, like, pretty robust to changes in the nature of the world. Like, uh, I think the reason to learn math has always been, like, to think clearly and, like, better understand the world, and, like, presumably that's something you wanna do even if, uh, there are sort of extremely capable AIs. And this is also true, you know, I, I personally like math a lot, but also the reason to, like, read a lot of books and do the humanities-

Episode duration: 1:03:11

Install uListen for AI-powered chat & search across the full episode — Get Full Transcript

Transcript of episode tQI35CSNB08

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.