Skip to content
a16za16z

Inside OpenAI’s Breakthroughs in Mathematical Reasoning

a16z Infra Partner Lisha Li sits down with OpenAI mathematicians Mehtaab Sawhney and Mark Sellke to discuss how quickly AI’s mathematical capabilities are advancing, what recent results reveal about model reasoning, and what happens when AI begins making progress on problems mathematicians have struggled with for decades. Mehtaab and Mark unpack several recent results from OpenAI’s models, including advances in sphere packing and the construction of a non-sofic group. They explain why the surprising part isn’t simply that models can search more possibilities or work longer than humans: in many cases, the reasoning traces look remarkably similar to the work of an expert mathematician, including choosing promising approaches, backtracking when they fail, and combining ideas from across the literature. They also explore what this means for mathematics itself: how the role of human taste and judgment may change, whether AI could produce far more mathematics than humans can absorb, and why models that accelerate discovery may also make sophisticated results easier to understand. Timestamps: 00:00 - Intro 00:50 - From Practicing Mathematician to OpenAI: Meet Mark & Mehtaab 02:43 - Why GPT-5 Was the Conversion Moment 04:21 - Beyond Search & Connections: How Recent Progress Goes Deeper 09:51 - Reasoning Traces: Is It Lucky Sampling or Actual Backtracking? 11:44 - Why Math Papers Are a Bad Training Set for Real Mathematics 16:20 - The Astra 10-Problem Set: Favorites & Deep Dives 36:17 - The Harness vs the Model: What Actually Matters? 40:01 - What Even Is "Taste" in a Model? 57:32 - How Should the Math Community Adopt AI? 01:00:01 - Empirical vs Theoretical Math & the Positive Vision Resources: Follow Lisha Li on X: https://x.com/lishali88 Follow Mehtaab Sawhney on X: https://x.com/mehtaab_sawhney Follow Mark Sellke on X: https://x.com/MarkSellke Stay Updated: If you enjoyed this episode, be sure to like, subscribe, and share with your friends! Find a16z on X: https://twitter.com/a16z Find a16z on LinkedIn: https://www.linkedin.com/company/a16z Listen to the a16z Show on Spotify: https://open.spotify.com/show/5bC65RDvs3oxnLyqqvkUYX Listen to the a16z Show on Apple Podcasts: https://podcasts.apple.com/us/podcast/a16z-podcast/id842818711 Follow our host: https://x.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see http://a16z.com/disclosures.

Lisha LihostMehtaab Sawhneyguest
Sep 8, 20261h 5mWatch on YouTube ↗

EVERY SPOKEN WORD

  1. 0:000:50

    Intro

    1. SP

      Often as a practicing mathematician, you have an idea, and then you kind of think it might work. Then you try for a few hours, a few weeks, and at some point you give up.

    2. LL

      Whereas for GPT, like, okay, a human told me to do this, like, let's just do this. And so that's why we're sort of in this renaissance of like reachable results.

    3. SP

      This is the best part about this problem, which is really nobody had any idea.

    4. MS

      Is the model just guessing in some insane way?

    5. LL

      But it doesn't seem like there's a limit so far, but it doesn't have that context yet.

    6. SP

      It'd be nice for the world if applied mathematics went a lot faster.

    7. MS

      The ceiling for difficulty of a math problem is pretty high. Even if AI continues getting like exponentially better at math, plausible we'll never solve something like P versus NP.

    8. LL

      What's the ideal way that this is being taken up by the math community? Probably most at this point are like, "Okay, AI is obviously doing some non-trivial stuff."

    9. MS

      Um, so-

  2. 0:502:43

    From Practicing Mathematician to OpenAI: Meet Mark & Mehtaab

    1. LL

      Well, thank you guys for coming. This is really exciting because I think math has been moving so fast, uh, with AI. I'd just love to, you know, get two both practicing mathematicians and who work at OpenAI to chat on, uh, some of these results. Um, so, you know, we have with us Mark Sellke and, um, Mehtaab, uh, Sawhney. We're connected actually because Yufei, um, was actually your advisor. And so both of you guys, um, have worked much more deeply in math since I've like quit many, many... like over a decade ago. Um, so this is very exciting to kind of hear a download of your thoughts on how OpenAI has been sort of approaching this, and also just like where you think math is going with the incredibly rapid advance of, uh, of how AI's been helping. Um, so yeah, I don't know. Maybe, uh, we can start off with some very basic questions of like, you know, you guys, what, what do you do to the ex- extent that you can, of course, share? Um, and, uh, and how did you come from, you know, being a practicing mathematician to working at OpenAI?

    2. SP

      Yeah, I mean, I guess we, we both, you know, broadly got excited last year when the models started to really take off in math.

    3. LL

      Yeah.

    4. SP

      Um, so I, uh, I joined a little bit, uh, before Mehtaab. I, I saw the IMO Gold Medal last summer basically, and I, I thought, you know, "This is, this is amazing," you know. "I, I want to see what the heck they did." [chuckles]

    5. LL

      Yeah.

    6. SP

      "Let me, let me go see." Um, and then, uh, yeah, I guess in the fall, we, um... Mark gave me a G- Mark gave me a GPT-5 account, and then-

    7. LL

      [laughs]

    8. SP

      ... I started playing with the models and very quickly became convinced that, yeah, it was extremely exciting to play with them, and yeah, I guess-

    9. LL

      And you, you two were collaborating before this.

    10. SP

      I, I, yeah.

    11. MS

      Yeah.

    12. SP

      We've known each other for a while.

    13. LL

      Yeah.

    14. SP

      Yeah.

    15. MS

      We had like one paper we actually wrote jointly.

    16. LL

      Yeah.

    17. SP

      Yeah.

  3. 2:434:21

    Why GPT-5 Was the Conversion Moment

    1. LL

      So GPT-5 was your conversion? [chuckles]

    2. SP

      Yeah, yeah.

    3. LL

      What was the magic that sort of, I don't know, as you're doing, like what question do you throw at it? What process-

    4. SP

      Yeah, so I think actually-- [chuckles] yeah, so I think how this started was, um, at least for me, the starting moment was something like there's a collection of problems called, um... So Er- Paul Erdős is a very famous mathematician. He posed a bunch of problems, and so they've now all been collected on this site. And I, so I specifically work in combinatorics, and a lot of these questions are among the most important, so it's always fun to flick through the site. But one thing that often happened to me that was extremely frustrating was I would look at a question, see that it's marked as open, and then not actually know if it's correct, uh, not actually know if it had... was still unsolved, because the literature is often quite hard to search. Um-

    5. LL

      Yeah.

    6. SP

      And I- one instance, I just plugged it into GPT-5 and like five minutes later it found a reference and, and this was a ca- this was a case where a few p- few of my friends had actually started thinking about the problem on the site. I was talking with them and, I mean, we had spent a few hours. It didn't seem... it wasn't clear if the problem was within reach, and it was just very nice, okay, to be told, "Yes, this is in reach. Here's how you do it." And yeah, GPT-5 told me this, and then I told Mark about this, and this... Yeah. This was sort of, yeah, this for me was quite a surprising moment.

    7. MS

      Yeah. And then we, we looked more into it, and we found like 10 more cases sort of like this.

    8. LL

      Mm. At the time, I feel like, you know, being better at maybe making connections between wide-- or even just like as you're saying, like the, the search for whether there's been a result or a related thing earlier is just like kind of humanely hard, but maybe better

  4. 4:219:51

    Beyond Search & Connections: How Recent Progress Goes Deeper

    1. LL

      for machine. But I imagine as the progress has happened in the last year, what has been impressive has kind of reached beyond that. And, um, maybe, you know, through talking about it more abstractly or if it's more natural to talk about it through like one of the problems that has been recently announced through, you know, Astra, you can kind of enlighten me as to like how, um, how the recent progress has been a lot more than just like, you know, searching through more areas, making these connections between the field, and perhaps just like actually deeper, more mathematical reasoning that's similar to a working mathematician.

    2. SP

      Yeah. I mean, I, I, I think this like search point of being, you know, familiar with everything is still definitely like a relative strength that-

    3. LL

      Yeah

    4. SP

      ... maybe informs like the types of problems that AI is solving now. Um, I, I think there's, there, there are some other relative strengths and weaknesses. Um, a- another relative strength that's pretty noticeable is just like it's very good at executing on some like idea once it, once it has it. Like, you know, when- whenever you have an idea, there's like, there's, there's usually some amount of, you know, getting everything lined up. Like, you know, is epsilon like smaller than delta? This, this kind of, of thing. You, you have to get everything correct. And like for a human, you know, you, you... I- it's easy to get lost in these kinds of details, and the, the AIs just kind of always nail these kinds of arguments, I find.

    5. LL

      Yeah. Is it usually just-- I mean, I feel like you guys will know more detail on this, but like for the unit distance problem, it was just, like the approach, there was definitely contributions from, you know, the OpenAI, but like the approach perhaps was suggested even- You know, originally by Erdos, and then it's just that the actual reasoning was a very, very, like, momentous, um, l- like, um, feat. And so for a human, you're like, "Well, I only have a limited amount of time," and if after so many, you know, steps, it is still not clear... I mean, maybe you're like Andrew Wiles and you actually spend 10 years alone and do something, but, like, it's not clear, then it's just, it, it's, it, it doesn't become... Like, the risk/reward is not good enough. Whereas for GPT, like, okay, I'll, like, a human told me to do this. [chuckles] Let's, let's just do this. And so that's why we're sort of in this renaissance of, like, reachable, uh, results. Um, does that track? And did you feel like with the Astra results, is that sort of, like, where the strengths have been primarily, or there's an extra ingredient or, or magic here?

    6. SP

      I, I feel, I think the unit distance example is somehow... It, it's quite telling. In the sense of maybe the exact construction, you can kind of, you can make it look very similar to what people have tried before. But I think what, what ha- You can often... I mean, often as a practicing mathematician, you, you have an idea, and then you kind of think it might work, then you try for a few hours, a few days, a few weeks. And at some point, you give up, and then a not so uncommon experience is that you find out a year or two later that somebody else got the idea to work that you thought that didn't work. So somehow getting an idea to work is, it can even be a large portion of the battle. And, um, I think what the mo- Especially in the case of the unit distance conjecture, there's just a lot of extraordinarily finicky details. And, um, very often when you're doing mathematics, it's, it's, you're kind of gambling against the problem. You're like, "Maybe I should try this approach, but it seems really unlikely and just not worth my time." And the model, I think in several of these cases, both by combining what it knew and sort of having good taste, kind of made the correct math. And you can kind of see this in, um, the, the summarized chain of thought we released. You can sort of look at it and it, it's reasoning like a mathematician, and because it knows a few very correct bits, it, it makes the right decisions and is eve-eventually able to prune the search tree. It, it's not really trying everything.

    7. LL

      Mm-hmm.

    8. SP

      It tries a lot of different things. It's extremely dogged. But it kind of... I mean, it can't try every idea. It has to try a limited set of ideas, and it's able to kind of use its knowledge plus good, good mathematical judgment and, and find the right path to go along. So I mean, that for me was like... Because this was a problem which a lot of people have thought about, and-

    9. LL

      Yeah

    10. SP

      ... I mean, the idea, the fact that the idea is not so foreign probably indicates that a lot of people have tried it.

    11. LL

      Right.

    12. SP

      Or at least a few very serious mathematicians have tried, and I think that's what made it really interesting to see. I, I think something else that's, uh, that, like, I feel when I see these proofs is like, like, if I have an idea and I'm trying to execute, like, it might be that I have some, like, wrong plan for, like, how to get things to work. And, um, like, as a human, if you have some, like, wrong path you go down for a while, it, it can be hard to, like, rewire your brain to, like, start over and, like, try a different path. Like, your, your... Kind of the initial idea is kind of linked in your brain with these other things that ended up not working. You know, it's sort of like your context window is, like, like, a little polluted and, like, you know, you can't just make another clone of yourself from, like, last week and say, you know, "Don't do this. Try something else. Build your intuition another direction."

    13. LL

      Mm-hmm.

    14. SP

      But, you know, it's very easy to do this with an AI. So I, I think this is another reason-

    15. LL

      Mm

    16. SP

      ... that it's like, um, it... Like, like, getting the details right once you have some good general direction is, like, much less of a barrier all of a sudden.

  5. 9:5111:44

    Reasoning Traces: Is It Lucky Sampling or Actual Backtracking?

    1. LL

      Mm. And when you say it's much easier to do with AI, it's like it's not actually being directed with human interference, too. It's like, as you were saying, in the reasoning traces, it's like making these choices. Maybe it backtracks, but then it's able to not be distracted by maybe, like, the context in which it's thinking about the problem, like, via these machinery, and it's like, go, go back. Like, do you see it kind of go back as well, or is it just, like, making good choices? Like, is it a lucky sample, or is it, like, actually reasoning like a mathematician where, okay, it doesn't, it, you know, doesn't do well on this path, but it goes back, but then it doesn't let that pollute?

    2. SP

      No, I mean, it definitely makes mistakes, and then-

    3. LL

      Yeah

    4. SP

      ... it goes back and thinks about it. I, I think it's somehow very calculating, very correct. I mean-

    5. LL

      Yeah

    6. SP

      ... as, as human mathematicians, you're not always perfect in making these decisions and, like, like, the first time something doesn't work, you, you automatically kind of downgrade how likely this approach is to work, and you keep doing this a few times.

    7. LL

      Yeah, yeah.

    8. SP

      The model somehow is much better able to, like... It seems, for several of the solutions we've seen, it somehow seem, it seems much better able to update the solu- like, how likely the path is to work, like-

    9. LL

      Yeah

    10. SP

      ... ver- versus rejecting a path versus a human doing it. Um, so- But, but I think even if it weren't, the fact that you could just start another model session over means that, like-

    11. LL

      Oh, okay

    12. SP

      ... you know, you kind of... [chuckles] It's always going to be the case that, that-

    13. LL

      Got it. Got it

    14. SP

      ... there's this advantage here.

    15. LL

      So in some sense, it is, like, still leveraging the fact that you could, like, run kind of parallel, you know, agents on the problem.

    16. SP

      Yeah.

    17. LL

      Um, but if it were kind of backtracking, then it does make it seem much more like a human, you know, mathematician. And, and perhaps it is kind of doing some of that stuff too, because, like, obviously, like, we have to, like, make mistakes in, in, in order to, like, even gain intuition for, like, why that solution space is, like, not, you know, not in the, in the, the set of paths that it could be in. Um, yeah.

    18. SP

      I mean, I think this kind of thing happens with humans too, where, like, if you get stuck on some approach, you might tell another human your kind of general idea.

    19. LL

      Yeah.

    20. SP

      And then they'll come back and, like, figure out how to get it to work and, you know, it's just-

    21. LL

      Yeah. Yeah, yeah, yeah

    22. SP

      ... it, it takes more time to do this, like, with humans.

  6. 11:4416:20

    Why Math Papers Are a Bad Training Set for Real Mathematics

    1. LL

      I wonder, I mean, you know, maybe this gets to an extent that you can't actually talk about sort of like... Obviously, don't talk about the training recipes or whatever, but, like, it's, it's interesting that if you're just studying, for instance, from math papers, it's, like, a very poor training set, like, a priority for math because, I mean, maybe math textbooks are even a purer example of this. It's, like, really bad at actually reconstructing the motivation for why things were... You know, it's like, don't I mean, maybe some people like it, but don't learn real analysis from Rudin. [chuckles] Just like it's very clean already and crisp. Um, and I think that- that's bad because it doesn't show the struggle that made us formulate, um, definitions in a certain way. Like, why do we even need to have real numbers be defined in this, like, super abstract way, et cetera. And so, you know, I think papers also, I mean, unless you're... Most people don't write papers with the, the context of, "I need to educate somebody to, to be a mathematician." And so, like, the actual maybe curriculum of, like, learning math is not inherent in, like, a lot of our artifacts as mathematicians. So maybe another way to ask this question is, if the reasoning tracers are actually producing things that it's like, okay, this is actually more close to mathematical thought, like, how does that arise?

    2. SP

      Um, I mean, I, yeah, I guess OpenAI has been, like, the pioneer of reasoning models and-

    3. LL

      Yeah

    4. SP

      ... you know, teaching AI to reason in this way. Um, so, you know, we're, we're doing a lot of work at kind of, kind of all possible directions on, on, you know, teaching models to reason better and for longer and, and, you know, in all kinds of different domains. Um, y- I, I mean, I, I think, I, I think we're training general purpose reasoning models and if, you know, kind of one... A lot of these behaviors that we're describing mathematically, like backtracking or kind of starting again, I mean, these are, these are not really specific to mathema- I mean, we're seeing them specifically in mathematics in these examples, but kind of they're general purpose tools for reasoning. And I think if you work hard at reasoning, you should see these patterns eventually.

    5. LL

      Mm-hmm. So it's just emergent because it's... I mean, I do think that's why the OpenAI approach, um, was so... I mean, it's like it doesn't rely on, you, you know, doing autoformalization in order to, like, guide the reasoning, and I think that's, like, obviously more like us. Um, but it's just, like, so not obvious that if you're just, like, training on, say, a corpus of like math proofs, maybe autoformalized in Lean, that you get, um, uh, get the sort of like projection of, like, how to think well. Like, put another way, like, with maybe, may- maybe if we think about it with code, like code is such a good corpus to train on because it's one of the few, um, data sets that has such large context. You just like... I mean, maybe you see this kind of with books, but they're less structurally interconnected. There's just, like, less structure there, I think, um, it's safe to say, like, on average in a, in a book compared to, like, a piece of code. Um, and so, like, with math papers, I feel like maybe what, what we're still bad at with coding models is stuff that that data set doesn't contain, which is, like, kind of the semantics. Like the syntax is there, but there's a little bit of, like, the higher level semantics of what produce, like, why do I have to write it this way? It's not... I- I'm kind of getting too, too much on the philosophical, but it is just, like, really interesting how it's still emergent, that it's doing good mathematics. And we'll probably get into this in more detail if you guys, you know, wanted to talk in more detail about some of the problems, which is just like it's, it's not just doing, like, the expected, like, we'll push the brute force thing. Like, you clearly are impressed with some of the reasoning traces, and it's just not obvious that's gleaned from, you know, what we would imagine would be the easy training set here.

    6. SP

      Yeah. Absolutely. I mean, I think this kind of thing is one reason we decided it was important to release, like, these summarized chains of thought for-

    7. LL

      Mm-hmm

    8. SP

      ... these kinds of results. Because if you, if you've never seen these and you just see all these proofs coming out, you're kind of... You're not sure what it means. Like, like is the model just guessing in some insane way?

    9. LL

      Yeah, right.

    10. SP

      Like, is it thinking in some totally foreign... Like, what, what's going on? But, but actually it, it's, it's reasoning kind of shockingly like a, an expert human would.

    11. LL

      Yeah. Yeah.

    12. SP

      Yeah. It's, it's very much like reading a colleague's, like, notes. I mean, it's a little more disorganized in some ways. They kind of... Like, especially if you work close enough with a collaborator and sometimes you'll just see them, like, spill out their thoughts in an email to you, and it kind of-

    13. LL

      Yeah

    14. SP

      ... it, it feels like reading a lot of those chained together. So it's, yeah, it's, it's very, it's, it's quite surprising the first

  7. 16:2036:17

    The Astra 10-Problem Set: Favorites & Deep Dives

    1. SP

      few times.

    2. LL

      Um, were you two sort of very involved in choosing the problems to, uh, to release in this, like, last 10 problem set that Astra was applied to?

    3. SP

      Um-

    4. LL

      Which was your favorite? [chuckles]

    5. SP

      Yeah. We were definitely involved. Um, do you wanna start on-

    6. LL

      Yeah, I mean-

    7. SP

      ... sphere packing maybe?

    8. LL

      Yeah, I guess. Yeah. So I guess my personal favorite among these problems is the following. It's, it's extremely simple question, which is just like, it's just about how efficiently can you put a bun... My circles are not very good, and they're not all the same size, but-

    9. SP

      But we're assuming they are. [chuckles]

    10. LL

      Yeah. So the question is just, like, how dense can you place a bunch of, um... So you have a bunch of spheres, uh, you have a bunch of spheres of radius one and D dimensions. Um, so the question is: How densely can they pack? And so, yeah. So, so in two dimensions, it's kind of like the... So, so D equals one, this is not an interesting question, kind of. It's just the real line, and yeah, you can cut it up and, uh, a sphere in dimension one is just a unit segment, so, okay, you can cover everything.

    11. SP

      Mm-hmm.

    12. LL

      Um, so in D equals two, it's kind of the picture that you know, that everybody loves. It's just like, uh, it's just a bunch of spheres which sort of form, like, a hexagonal lattice.

    13. SP

      Mm-hmm.

    14. LL

      Hopefully I've drawn it well enough. But I can draw the hexagon.

    15. SP

      Kind of betraying my naiveté on this problem, is that, like, obvious? Is it, like, a very elegant proof-

    16. LL

      Yeah

    17. SP

      ... that it's a regular lattice?

    18. LL

      Uh, it's not so... Yeah, it's not so obvious that this should work. It was only proved in the '60s, I think. Um, there's a short argument, but it's not, it's not so easy.

    19. SP

      Mm-hmm.

    20. LL

      Yeah. Um...

    21. SP

      What is the intuition? Like, what is kind of like the-

    22. LL

      I mean-

    23. SP

      ... machinery of the argument?

    24. LL

      I mean, it kind of like-

    25. SP

      Yeah. I mean, it definitely looks like it should work. [chuckles] That's why I'm-

    26. LL

      Yeah. So I think-

    27. SP

      But so did... [chuckles]

    28. LL

      Yeah, I think this is the best part about this problem, which is really nobody had any idea how to solve this problem. [both chuckling] Um-

    29. SP

      Yeah

    30. LL

      ... so, yeah. So I mean-

  8. 36:1740:01

    The Harness vs the Model: What Actually Matters?

    1. MS

      coincidence.

    2. LL

      Yeah, yeah. It's like interesting when you're sort of saying the first prompt, which is, you know, maybe so basic, which is like, "Can you push this further?" It does require some judgment from mathematicians. Uh, but like- Eventually, you would imagine by scaling the models, you don't need to do that, or there's another view that the harness actually does matter, and this is kind of part of the harness apparatus. Do you guys have any views on that with your working with Astra? Uh, uh, especially generations of models and how, how much do you have to kind of input or how much the harness matters versus not?

    3. SP

      I mean, I, yeah, I, I guess there have been some, like, funny quirks like this that just come from, like, exactly what you ask the model to do, basically.

    4. LL

      Yeah.

    5. SP

      Like, like in, in this case, what the model was asked to do originally for codes was to improve the bounds-

    6. LL

      Yeah

    7. SP

      ... by, like, some exponential factor. So it, it really, like, shows up in this, like, leading constant up here.

    8. LL

      Mm-hmm.

    9. SP

      Um, and you know, it, it improved the bounds and it, it didn't try to push things too much further. Like sometimes you see it do, but sometimes it just doesn't bother. But yeah, you know, you just ask it again and it, it goes further. So it, it wasn't like a capabilities issue, it just-

    10. LL

      Yeah

    11. SP

      ... kind of didn't feel like it at the time.

    12. LL

      Do you call that judgment or, like, what is the... 'Cause like there, there is a... Yeah, what do you, you call that?

    13. SP

      I mean, it's, it's, models tend to be pretty task-oriented.

    14. LL

      [laughs]

    15. SP

      If you, if you, if you tell it to do a task and it accomplishes the task-

    16. LL

      That's good. I'm training for that. [laughs]

    17. SP

      ... then it's pretty happy.

    18. LL

      So yeah, the, the task-orientedness, it's like, but do we expect that level to kind of ascend up to... It's not that they will be less good at being task-oriented, it's just like w- they'll ascend to the level of like, "Okay, no, let's, let's go in this direction." You'll have the judgment. 'Cause you guys had the judgment. You're like, "Okay, this is pretty promising. Looks like you're using a lot of representation theory. It doesn't seem like there's a limit so far." Um, but it doesn't have that context yet. But, like, I guess what I'm trying to say is like this one it's hard to maybe extr- harder to extrapolate, but from like previous generations, when you had to give it more, maybe prompting more of that harness work, but eventually probably had to give it less. So it probably gives you some confidence that there's this, like, really fast ascension. And, and do you see... Yeah, like what are some promises?

    19. SP

      I mean, somehow, somehow solving a harder math problem is, like, you have to solve many smaller, like, somewhat less hard math problems. And, and the fact that the math problems are getting harder is kind of an indication that you're, the model's able to take on more and more work in, like, a single continuous unit. Um, and I, I think that's, that's the thing that looks very promising. Somehow, like, a- any of these solutions, it's not like one idea, then you're kind of home free.

    20. LL

      Yeah, totally.

    21. SP

      You need several, you need several pieces to kind of interact and talk to each other. And the models, I mean, the, the model doesn't come up with all the ideas at once, right? It, it doesn't pull everything out of, in, in an instance. So kind of the fact that it needs to sort of see how this piece interacts with another piece, that's kind of like solving a problem in itself or piecing together many problems in itself.

    22. LL

      It could just be that, okay, when you're telling it, "Okay, push this even further," that was a- of the same order of, like, magnitude as, like, all the smaller things it's solving as well in, in between. And so you don't think that's as kind of like a privileged direction. It's just sort of like, "Hey, let's give it like one, one more help." Or you actually think that there's... I guess what I'm trying to get at, a bigger question is like, is there a good sense of like, you know, taste? Because, like, when people talk about, for instance, how well the models are getting at like, um, doing research [chuckles] for instance, because that's what we, we know we want a little bit of RSI. Um, and, uh, and sort of like there's surprising things about, um, how that improves, and then there's like the, oh, you know, maybe right now it's at a level of still like a junior researcher. It's like not really asking like the right problems. And so I'm just trying to get like a, maybe a sense of like where you're seeing that progress through the model advancements each generation.

  9. 40:0157:32

    What Even Is "Taste" in a Model?

    1. LL

      I mean, what is taste even?

    2. SP

      Yeah. I, I think somehow I tend to be p- pretty utilitarian in my view of taste. And like if you're able to solve problems faster by making ju- better judgments, like I think that's like the best like general proxy I have for taste.

    3. LL

      Mm-hmm.

    4. SP

      And somehow the fact that it's solving harder problems means it has, kind of by definition means it has better taste. I think there, there, there are these no... Yeah, I think occasionally, occasionally because they are task-oriented, you do occasionally get these, these symptoms of like, oh, it clearly has made a breakthrough. It kind of understands it's made a breakthrough, and then it doesn't kind of push all the way to the limit 'cause that's not what you asked. But that seems, yeah, that seems like, seems rather minor compared to like the set, the, the state of progress we've seen so far.

    5. LL

      Mm-hmm. Mm-hmm. Okay. Yeah. No, I think that's a pretty clear answer.

    6. SP

      Like I, I think it's, it's like maybe you're liable to get confused if you're trying to like do a concrete long horizon task and show taste kind of at the same time. But like, you know, if, if you, if you have like one model that's responsible for taste and one model that's responsible for going out and like, you know, working for a long time at, at solving a hard problem, kind of as the, as the like, uh, you know, uh, underling of the, of the, the supervising AI, I, I feel like that's kind of going to be fine currently.

    7. LL

      Oh, interesting. 'Cause that is like saying that they, these two things are somewhat sep- or if not separate, at least they shouldn't kind of pollute each other's context, which is a, a little bit... I mean, it could be potentially like a strong, stronger statement, um, than... I, I guess, yeah, no, it's just kind of interesting 'cause it's, it, it, it might just be, like to your point, it's, you know, let's take the utilitarian answer. It's, it's solving harder and harder problems. It's doing a lot more than just like, you know, brute forcing something. It's making choices. It's like pruning, you know, a vastly large, uh, space of possible paths into something that's like really, you know, is both tractable but then ends up, um, being like it's a diminishingly small path within that space. Um, but like having, like why would, what would be like a separate [clears throat] model, sep- separate generation or something that's a different version of the model that would contribute to taste? Or maybe that's totally, like it's too abstract, doesn't make any sense. You know, we should just let the actual... This might, like a related question would be like, you know, what is, what is the, the thing that gets us to a better version of intelligence, the harness and the model? Is it just the model? And it's like we, we see this in, you know, at least in, um- Applied AI or, you know, startups where it's like, it's a continual battle of like you need the harness, but then the harness a-adapts very poorly to a new model, 'cause sometimes like a very, very minimal harness is still the best way to expose to the raw power of the model. But then now we also have these like training regimes where we require the harness to be, you know, trained with... Maybe part, part of this is to keep things more proprietary and harder to, harder for other people to use it. But I think partially it's, it's maybe actually that it helps, um, have more control on like the reasoning traces you care about. It's a long rambling way of saying it's like, yeah, I, I don't actually... Like I, I, this is so interesting to see how the models have gotten be-better at math and, um, maybe something that's like very abstract and hard to describe, like taste, is a way to tease out like what is actually necessary here.

    8. MS

      I think my only like nontrivial thought here is that like when you're working, I mean, just when you're doing any task, occasionally you get pigeonholed and you like work really hard and just having a friend look over your shoulder and be like, "What are you doing?" And then just like, just having that one bit of like step back for 10 seconds, like this is often very useful.

    9. LL

      Yeah, yeah.

    10. MS

      I mean, I see no reason why humans would be so different than models somehow.

    11. LL

      Yep, yep.

    12. MS

      Having... Or models would be so different than humans. Having, having a few humans working together is often more powerful than just having one.

    13. LL

      Yeah. It's like in this kind of collaborative thing, you actually, you kind of, yeah, artificially created it, but it's very similar and dynamic.

    14. MS

      But I, I think a lot of taste is also like having a sense of what problems you or like some method you have in mind are going to be good at solving.

    15. LL

      Mm-hmm.

    16. MS

      Like it's... I mean, certainly there's some amount of like absolute aesthetic point, right? But there, there's also just like, you know, having a nose for what, what y-you might want to pursue because you'll be able to make progress. Um, and you know, I, I think for that, like there's, you know, you, you would expect that as a side product of being good at completing tasks, you would, you would get there sort of, right?

    17. LL

      Let me know if we still wanna do like a section on sophic groups, 'cause I think, you know, up to you guys, it's definitely super interesting.

    18. MS

      So maybe the, the first question is what is a group? Let's remind ourselves. So, uh, a group is, uh, a set of elements with some multiplication operation. And, uh, basically this is how mathematicians think about symmetry. So, so you're like... Basically like if G and H are elements of your group, then GH has some, is some other well-defined, uh, element of your group, and you have like, uh, associativity, uh, and you have an inverse. So for every G, there's some inverse, uh, and there's some like specific element in the group that, uh, is, is kind of the identity. Uh, okay. So, you know, it's, it's some like abstraction of like, uh, composing operations. So these could be like numbers, they could be like multiplying matrices, they could be like, like rotating something, which is, you know, a special case of multiplying matrices. Um, and, uh, a group is, uh, sophic if... Well, there's some, uh, you know, precise definition, uh, but, you know, uh, roughly, uh, it means it... Uh, so, so I should say like groups that can be finite or infinite. So like, you know, um, if, if you have like a square, like all the rotations of it form a group with like four elements.

    19. LL

      Mm-hmm.

    20. MS

      If you have like a circle, then the rotations form a group with like uncountably any, many elements. Um, and, uh, so sophic groups are either finite or countable. Uh, you should think of them as being countably infinite, so there's like the same number of elements as like the integers. And, uh, if it's sophic if in some sense, uh, it can be, uh, approximated by finite groups. So we didn't know if there was a non-sophic group. So, so-

    21. LL

      Yeah

    22. MS

      ... the, the, the result that, uh, Astra proved is simply that, uh, there exists a non-sophic group.

    23. LL

      Yeah. And without like, I mean, we can, you know, before going to, to that proof, it is like, you know, it's like I feel like a lot of the programs in math is like, okay, we are such finite creatures. Let's see how well our finite approximations are, you know, uh, do. And in this case, especially for the countable case, it's like, uh, maybe you'll, you'll be relating it to like the Aldis Leon's thing. It's just like, it helps kind of anchor the picture of like, it seems like such a... I mean, it's a nice result if it were true, but it, it's not, and it seems almost like reasonable. And so, yeah, I, I actually didn't, um, didn't, uh, go... I would love to hear the explanation of like how it, it found a counterexample.

    24. MS

      Yeah. I mean, I, I would say that like, you know, the, the hope that there was no non-sophic group, so every group has this kind of approximation. Like maybe this is sort of like people hoping that there's a miracle.

    25. LL

      Mm.

    26. MS

      Because, uh, it turns out that groups like this have a lot of nice properties, um, because, uh, you can run certain proofs for finite groups and then, uh, you know, kind of approximate them in whatever way the definition of being sophic lets you approximate them-

    27. LL

      Mm-hmm

    28. MS

      ... and, uh, get the result. So, so like, um, there's this notion of being a conjunctive group. So there's, uh, there's some fact that, uh, any, uh, group which is sophic is, uh, also conjunctive. Uh, surjunctive is some property of like, um, dynamical systems on the group.

    29. LL

      Mm-hmm.

    30. MS

      And, um, I guess the original question was whether every group is, uh, surjunctive. This is some question of Gottschalk from the '70s. Uh, and this, uh, this fact that, um, follows this, like, pattern of prove it for finite groups and then do this approximation is, is what motivated the question about if there's a non-sofic group.

  10. 57:321:00:01

    How Should the Math Community Adopt AI?

    1. SP

      in this case.

    2. LL

      Yeah. Well, actually maybe that's a great segue into like how, you know, what's the ideal way that this is being taken up by the math community? 'Cause I feel like there's a spectrum of answers from working mathematicians, a sense of like some, you know, uh, probably most at this point are like, "Okay, AI's obviously doing some non-trivial stuff. Um, it would be a disadvantage not to admit that in my workflow." Um, I've definitely heard some stories where people are kind of, you know, would find it hard to either take AI as a co-author or like how do you even do kind of attribution this way? But I don't know, like what, um... But maybe to paint the, the more optimistic picture, so you're saying you want the mathematicians to be building on this, these results. It definitely generates a lot more results to be verified, so it, you know, it puts pressure on, on the community and the profession. Like how do you, how do you kind of expect the evolution of kind of uptake and, and, and collaboration with mathematicians?

    3. SP

      Well, I mean, I mean the, given that the fact that the models can produce sophisticated mathematics means that they can help you understand like sophisticated mathematics. I mean, like I don't know, occasionally like I enjoy looking at the archive, and I want to understand some proof, and like I could read the introduction, but in practice it's just much faster to take the PDF, put it into, put it into my favorite model, and then like get a, get an output of like what is the rough proof strategy.

    4. LL

      Yeah.

    5. SP

      And somehow this like... A- along... I mean, of course, models are going to help us produce exponentially more mathematics, but they also make it much easier to absorb it. Um-

    6. LL

      Yeah

    7. SP

      ... and right now, okay, it's still a bit of a challenge back and forth, but I, I think it's, for me at least, a, it's much, much faster at understanding. Uh, it's much, much faster to understand a piece of mathematics with a model than without it. Um, so it's helping solve the problem it creates anyways.

    8. LL

      Yeah. I feel like that at least... And it, it's, you know, I, I don't, I don't view it as like creating much more problem, but again, like I don't have such, you know, high stakes and like, okay, I'm, I'm gonna get... I'm not gonna get tenure, et cetera. So like I, I, I agree, like making it more accessible, like if I'm not spending so much time absorbing an area, I can like put it into ChatGPT and then expect to... I mean, you guys have an even more powerful model, hopefully releasing, um, for other people to enjoy as well. But like it's, um, I think like the, the positive version of that is actually more people can participate in mathematics. It's like people might be coming with other intuitions, and they could actually maybe generate good mathematics. Is that sort of like closer to the vision of what you're hoping this is, you know, pushing towards?

  11. 1:00:011:05:00

    Empirical vs Theoretical Math & the Positive Vision

    1. LL

      Um, or like what, what, what things do you think mathematicians should be wary of to, to kind of adapt fast enough to take advantage of AI?

    2. SP

      Yeah. I mean, I, I think certainly there will be a lot of changes, right? Like I, I guess in math, like there are, there are a lot of things that are kind of important for, for like a given result, right? You need someone to come up with it, but you also need people to understand and absorb it and like, you know, internalize it enough to, to do more with it and, and like figure out where it fits into like humanity's understanding, right? And like, uh, a, a couple of years ago, uh, like the proving the result was like so hard that kind of the, the other stuff was just kind of coming along for the ride, right? You know, like if you, if you manage to like prove this thing yourself, you're automatically gonna understand it quite well. You're kind of responsible for like maintaining it in, in some sense and like, you know, explaining it to other people. Um, and yeah, now this kind of what was the main bottleneck before is kind of, um-

    3. MS

      Much less of a bottleneck and, you know, these other kind of constraints, uh, come into play. Um, so it... Yeah, the, the, the, like, optimal, um, structuring for, you know, organizing the knowledge could, could look rather different.

    4. LL

      Yeah. How does that look? I mean, d- does this make the field a lot more kind of empirical? Will people do sort of the hard, like the first thing that was scarce, which is like all the reasoning and then more? I mean, not that it's like a bad thing to make it empirical, but it's almost like it functions as a very different discipline. Um, like a lot of the fun stuff is understanding, you know? Um, and so going to Shannon, communicating, maybe assembling, having still the human taste. Does that sort of remain re- rarefied and, and that's how, you know, current mathematicians need to adapt and, and reward, [chuckles] you know, uh, contributions or, or is this too much of a caricature that's like something else?

    5. MS

      I think certainly understanding how to put, as we get more and more mathematics, put it in like a proper framework and sort of how, sort of like being able to explain it to other humans so that they can also appreciate it. I mean, somehow implicitly we valued this, but it was usually because you were the person proving the results that gave everybody else the understanding. But I think increasingly it would be a function of, like, you're sort of helping, y- you're the human who can sort of give this understanding to other people and sort of help them with it. Um, I think that more c- sort of, that communal understanding will, I think, become... It was much more implicit in how we viewed mathematicians, but I think it will be an increasingly more explicit and valuable part of the subject. I mean, a, a nice thing about math is that, um, the, the ceiling for difficulty of a math problem is pretty high. So even if, um, you know, kind of, e- even if AI get, you know, continues getting like exponentially better at math, like it might, you know, it's plausible we'll never solve something like P versus NP.

    6. LL

      Mm-hmm.

    7. MS

      And it, it could be that, like, the, the field kind of becomes more, um, you know, attached to like, like these big mysteries-

    8. LL

      Mm.

    9. MS

      ... and, and less to like smaller mysteries that-

    10. LL

      Mm

    11. MS

      ... are more like routine now.

    12. LL

      Yeah, yeah. I think that's a positive vision of the-

    13. MS

      I mean, also like, I don't know, there are things I spent like months or years of my life wondering about, not getting to know and hopefully I get to-

    14. LL

      And now we get the...

    15. MS

      Yes.

    16. LL

      Yeah, that's like, what a joy.

    17. MS

      Some portion, yeah, some portion of them I'll get to know the answer to. I'm, I'm pretty happy about that.

    18. LL

      Yeah. No, exactly. No, I'm, I'm excited about this like renaissance of results and understanding, and I feel like, I mean, this is such an, an infinite, you know, field. [chuckles] Like, okay, no, no pun intended, but like it's just like, it, it's, it's, it's just there's so much that you can actually create here. Um, so I mean, especially if, for somebody like me who's not gonna have the time to actually like practice mathematics, now there's like a lot more that you can actually do in the activity of math. So yeah.

    19. MS

      Yeah, I think the, the like, the ability of someone who's not working on math as like their literal job all the time to like understand what's going on and like, you know, learn about some of the mysteries they might have wondered about will, will go up quite a lot. Um, also, you know, if you're, if you're like, if you're working on something that requires some math-

    20. LL

      Mm-hmm

    21. MS

      ... you know, suddenly you, you don't need to like find a world expert on this-

    22. LL

      [chuckles] Yeah, yeah

    23. MS

      ... topic to, to be able to, you know, use it in your own work. You can-

    24. LL

      Sorry, mathematicians. [laughs] Yeah, yeah, yeah. No, it, it's true. I mean, I think there was just like a dearth of actual like people who could, could do that, and so I think this is helpful. Maybe it's helpful for theor- theoretical physics, like we'll see. Um, but a lot of other applied areas as well.

    25. MS

      It'd be nice for the world if applied mathematics went a lot faster.

    26. LL

      Yes. [chuckles] I mean, I'm, I'm of that opinion. Uh, well, thank you guys for joining. This was a lot of fun, and I'm, you know, just so excited for how much the models are advancing. So maybe we'll have you guys back soon.

    27. MS

      Yeah. Thanks so much for having us.

    28. LL

      Yeah.

    29. MS

      Yeah, thanks for having us.

Episode duration: 1:05:15

Install uListen for AI-powered chat & search across the full episode — Get Full Transcript

Transcript of episode 1JvyLGd2Sfs

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.