EVERY SPOKEN WORD
65 min read · 13,062 words- 0:00 – 0:50
Intro
- SPSpeaker
Often as a practicing mathematician, you have an idea, and then you kind of think it might work. Then you try for a few hours, a few weeks, and at some point you give up.
- LLLisha Li
Whereas for GPT, like, okay, a human told me to do this, like, let's just do this. And so that's why we're sort of in this renaissance of like reachable results.
- SPSpeaker
This is the best part about this problem, which is really nobody had any idea.
- MSMehtaab Sawhney
Is the model just guessing in some insane way?
- LLLisha Li
But it doesn't seem like there's a limit so far, but it doesn't have that context yet.
- SPSpeaker
It'd be nice for the world if applied mathematics went a lot faster.
- MSMehtaab Sawhney
The ceiling for difficulty of a math problem is pretty high. Even if AI continues getting like exponentially better at math, plausible we'll never solve something like P versus NP.
- LLLisha Li
What's the ideal way that this is being taken up by the math community? Probably most at this point are like, "Okay, AI is obviously doing some non-trivial stuff."
- MSMehtaab Sawhney
Um, so-
- 0:50 – 2:43
From Practicing Mathematician to OpenAI: Meet Mark & Mehtaab
- LLLisha Li
Well, thank you guys for coming. This is really exciting because I think math has been moving so fast, uh, with AI. I'd just love to, you know, get two both practicing mathematicians and who work at OpenAI to chat on, uh, some of these results. Um, so, you know, we have with us Mark Sellke and, um, Mehtaab, uh, Sawhney. We're connected actually because Yufei, um, was actually your advisor. And so both of you guys, um, have worked much more deeply in math since I've like quit many, many... like over a decade ago. Um, so this is very exciting to kind of hear a download of your thoughts on how OpenAI has been sort of approaching this, and also just like where you think math is going with the incredibly rapid advance of, uh, of how AI's been helping. Um, so yeah, I don't know. Maybe, uh, we can start off with some very basic questions of like, you know, you guys, what, what do you do to the ex- extent that you can, of course, share? Um, and, uh, and how did you come from, you know, being a practicing mathematician to working at OpenAI?
- SPSpeaker
Yeah, I mean, I guess we, we both, you know, broadly got excited last year when the models started to really take off in math.
- LLLisha Li
Yeah.
- SPSpeaker
Um, so I, uh, I joined a little bit, uh, before Mehtaab. I, I saw the IMO Gold Medal last summer basically, and I, I thought, you know, "This is, this is amazing," you know. "I, I want to see what the heck they did." [chuckles]
- LLLisha Li
Yeah.
- SPSpeaker
"Let me, let me go see." Um, and then, uh, yeah, I guess in the fall, we, um... Mark gave me a G- Mark gave me a GPT-5 account, and then-
- LLLisha Li
[laughs]
- SPSpeaker
... I started playing with the models and very quickly became convinced that, yeah, it was extremely exciting to play with them, and yeah, I guess-
- LLLisha Li
And you, you two were collaborating before this.
- SPSpeaker
I, I, yeah.
- MSMehtaab Sawhney
Yeah.
- SPSpeaker
We've known each other for a while.
- LLLisha Li
Yeah.
- SPSpeaker
Yeah.
- MSMehtaab Sawhney
We had like one paper we actually wrote jointly.
- LLLisha Li
Yeah.
- SPSpeaker
Yeah.
- 2:43 – 4:21
Why GPT-5 Was the Conversion Moment
- LLLisha Li
So GPT-5 was your conversion? [chuckles]
- SPSpeaker
Yeah, yeah.
- LLLisha Li
What was the magic that sort of, I don't know, as you're doing, like what question do you throw at it? What process-
- SPSpeaker
Yeah, so I think actually-- [chuckles] yeah, so I think how this started was, um, at least for me, the starting moment was something like there's a collection of problems called, um... So Er- Paul Erdős is a very famous mathematician. He posed a bunch of problems, and so they've now all been collected on this site. And I, so I specifically work in combinatorics, and a lot of these questions are among the most important, so it's always fun to flick through the site. But one thing that often happened to me that was extremely frustrating was I would look at a question, see that it's marked as open, and then not actually know if it's correct, uh, not actually know if it had... was still unsolved, because the literature is often quite hard to search. Um-
- LLLisha Li
Yeah.
- SPSpeaker
And I- one instance, I just plugged it into GPT-5 and like five minutes later it found a reference and, and this was a ca- this was a case where a few p- few of my friends had actually started thinking about the problem on the site. I was talking with them and, I mean, we had spent a few hours. It didn't seem... it wasn't clear if the problem was within reach, and it was just very nice, okay, to be told, "Yes, this is in reach. Here's how you do it." And yeah, GPT-5 told me this, and then I told Mark about this, and this... Yeah. This was sort of, yeah, this for me was quite a surprising moment.
- MSMehtaab Sawhney
Yeah. And then we, we looked more into it, and we found like 10 more cases sort of like this.
- LLLisha Li
Mm. At the time, I feel like, you know, being better at maybe making connections between wide-- or even just like as you're saying, like the, the search for whether there's been a result or a related thing earlier is just like kind of humanely hard, but maybe better
- 4:21 – 9:51
Beyond Search & Connections: How Recent Progress Goes Deeper
- LLLisha Li
for machine. But I imagine as the progress has happened in the last year, what has been impressive has kind of reached beyond that. And, um, maybe, you know, through talking about it more abstractly or if it's more natural to talk about it through like one of the problems that has been recently announced through, you know, Astra, you can kind of enlighten me as to like how, um, how the recent progress has been a lot more than just like, you know, searching through more areas, making these connections between the field, and perhaps just like actually deeper, more mathematical reasoning that's similar to a working mathematician.
- SPSpeaker
Yeah. I mean, I, I, I think this like search point of being, you know, familiar with everything is still definitely like a relative strength that-
- LLLisha Li
Yeah
- SPSpeaker
... maybe informs like the types of problems that AI is solving now. Um, I, I think there's, there, there are some other relative strengths and weaknesses. Um, a- another relative strength that's pretty noticeable is just like it's very good at executing on some like idea once it, once it has it. Like, you know, when- whenever you have an idea, there's like, there's, there's usually some amount of, you know, getting everything lined up. Like, you know, is epsilon like smaller than delta? This, this kind of, of thing. You, you have to get everything correct. And like for a human, you know, you, you... I- it's easy to get lost in these kinds of details, and the, the AIs just kind of always nail these kinds of arguments, I find.
- LLLisha Li
Yeah. Is it usually just-- I mean, I feel like you guys will know more detail on this, but like for the unit distance problem, it was just, like the approach, there was definitely contributions from, you know, the OpenAI, but like the approach perhaps was suggested even- You know, originally by Erdos, and then it's just that the actual reasoning was a very, very, like, momentous, um, l- like, um, feat. And so for a human, you're like, "Well, I only have a limited amount of time," and if after so many, you know, steps, it is still not clear... I mean, maybe you're like Andrew Wiles and you actually spend 10 years alone and do something, but, like, it's not clear, then it's just, it, it's, it, it doesn't become... Like, the risk/reward is not good enough. Whereas for GPT, like, okay, I'll, like, a human told me to do this. [chuckles] Let's, let's just do this. And so that's why we're sort of in this renaissance of, like, reachable, uh, results. Um, does that track? And did you feel like with the Astra results, is that sort of, like, where the strengths have been primarily, or there's an extra ingredient or, or magic here?
- SPSpeaker
I, I feel, I think the unit distance example is somehow... It, it's quite telling. In the sense of maybe the exact construction, you can kind of, you can make it look very similar to what people have tried before. But I think what, what ha- You can often... I mean, often as a practicing mathematician, you, you have an idea, and then you kind of think it might work, then you try for a few hours, a few days, a few weeks. And at some point, you give up, and then a not so uncommon experience is that you find out a year or two later that somebody else got the idea to work that you thought that didn't work. So somehow getting an idea to work is, it can even be a large portion of the battle. And, um, I think what the mo- Especially in the case of the unit distance conjecture, there's just a lot of extraordinarily finicky details. And, um, very often when you're doing mathematics, it's, it's, you're kind of gambling against the problem. You're like, "Maybe I should try this approach, but it seems really unlikely and just not worth my time." And the model, I think in several of these cases, both by combining what it knew and sort of having good taste, kind of made the correct math. And you can kind of see this in, um, the, the summarized chain of thought we released. You can sort of look at it and it, it's reasoning like a mathematician, and because it knows a few very correct bits, it, it makes the right decisions and is eve-eventually able to prune the search tree. It, it's not really trying everything.
- LLLisha Li
Mm-hmm.
- SPSpeaker
It tries a lot of different things. It's extremely dogged. But it kind of... I mean, it can't try every idea. It has to try a limited set of ideas, and it's able to kind of use its knowledge plus good, good mathematical judgment and, and find the right path to go along. So I mean, that for me was like... Because this was a problem which a lot of people have thought about, and-
- LLLisha Li
Yeah
- SPSpeaker
... I mean, the idea, the fact that the idea is not so foreign probably indicates that a lot of people have tried it.
- LLLisha Li
Right.
- SPSpeaker
Or at least a few very serious mathematicians have tried, and I think that's what made it really interesting to see. I, I think something else that's, uh, that, like, I feel when I see these proofs is like, like, if I have an idea and I'm trying to execute, like, it might be that I have some, like, wrong plan for, like, how to get things to work. And, um, like, as a human, if you have some, like, wrong path you go down for a while, it, it can be hard to, like, rewire your brain to, like, start over and, like, try a different path. Like, your, your... Kind of the initial idea is kind of linked in your brain with these other things that ended up not working. You know, it's sort of like your context window is, like, like, a little polluted and, like, you know, you can't just make another clone of yourself from, like, last week and say, you know, "Don't do this. Try something else. Build your intuition another direction."
- LLLisha Li
Mm-hmm.
- SPSpeaker
But, you know, it's very easy to do this with an AI. So I, I think this is another reason-
- LLLisha Li
Mm
- SPSpeaker
... that it's like, um, it... Like, like, getting the details right once you have some good general direction is, like, much less of a barrier all of a sudden.
- 9:51 – 11:44
Reasoning Traces: Is It Lucky Sampling or Actual Backtracking?
- LLLisha Li
Mm. And when you say it's much easier to do with AI, it's like it's not actually being directed with human interference, too. It's like, as you were saying, in the reasoning traces, it's like making these choices. Maybe it backtracks, but then it's able to not be distracted by maybe, like, the context in which it's thinking about the problem, like, via these machinery, and it's like, go, go back. Like, do you see it kind of go back as well, or is it just, like, making good choices? Like, is it a lucky sample, or is it, like, actually reasoning like a mathematician where, okay, it doesn't, it, you know, doesn't do well on this path, but it goes back, but then it doesn't let that pollute?
- SPSpeaker
No, I mean, it definitely makes mistakes, and then-
- LLLisha Li
Yeah
- SPSpeaker
... it goes back and thinks about it. I, I think it's somehow very calculating, very correct. I mean-
- LLLisha Li
Yeah
- SPSpeaker
... as, as human mathematicians, you're not always perfect in making these decisions and, like, like, the first time something doesn't work, you, you automatically kind of downgrade how likely this approach is to work, and you keep doing this a few times.
- LLLisha Li
Yeah, yeah.
- SPSpeaker
The model somehow is much better able to, like... It seems, for several of the solutions we've seen, it somehow seem, it seems much better able to update the solu- like, how likely the path is to work, like-
- LLLisha Li
Yeah
- SPSpeaker
... ver- versus rejecting a path versus a human doing it. Um, so- But, but I think even if it weren't, the fact that you could just start another model session over means that, like-
- LLLisha Li
Oh, okay
- SPSpeaker
... you know, you kind of... [chuckles] It's always going to be the case that, that-
- LLLisha Li
Got it. Got it
- SPSpeaker
... there's this advantage here.
- LLLisha Li
So in some sense, it is, like, still leveraging the fact that you could, like, run kind of parallel, you know, agents on the problem.
- SPSpeaker
Yeah.
- LLLisha Li
Um, but if it were kind of backtracking, then it does make it seem much more like a human, you know, mathematician. And, and perhaps it is kind of doing some of that stuff too, because, like, obviously, like, we have to, like, make mistakes in, in, in order to, like, even gain intuition for, like, why that solution space is, like, not, you know, not in the, in the, the set of paths that it could be in. Um, yeah.
- SPSpeaker
I mean, I think this kind of thing happens with humans too, where, like, if you get stuck on some approach, you might tell another human your kind of general idea.
- LLLisha Li
Yeah.
- SPSpeaker
And then they'll come back and, like, figure out how to get it to work and, you know, it's just-
- LLLisha Li
Yeah. Yeah, yeah, yeah
- SPSpeaker
... it, it takes more time to do this, like, with humans.
- 11:44 – 16:20
Why Math Papers Are a Bad Training Set for Real Mathematics
- LLLisha Li
I wonder, I mean, you know, maybe this gets to an extent that you can't actually talk about sort of like... Obviously, don't talk about the training recipes or whatever, but, like, it's, it's interesting that if you're just studying, for instance, from math papers, it's, like, a very poor training set, like, a priority for math because, I mean, maybe math textbooks are even a purer example of this. It's, like, really bad at actually reconstructing the motivation for why things were... You know, it's like, don't I mean, maybe some people like it, but don't learn real analysis from Rudin. [chuckles] Just like it's very clean already and crisp. Um, and I think that- that's bad because it doesn't show the struggle that made us formulate, um, definitions in a certain way. Like, why do we even need to have real numbers be defined in this, like, super abstract way, et cetera. And so, you know, I think papers also, I mean, unless you're... Most people don't write papers with the, the context of, "I need to educate somebody to, to be a mathematician." And so, like, the actual maybe curriculum of, like, learning math is not inherent in, like, a lot of our artifacts as mathematicians. So maybe another way to ask this question is, if the reasoning tracers are actually producing things that it's like, okay, this is actually more close to mathematical thought, like, how does that arise?
- SPSpeaker
Um, I mean, I, yeah, I guess OpenAI has been, like, the pioneer of reasoning models and-
- LLLisha Li
Yeah
- SPSpeaker
... you know, teaching AI to reason in this way. Um, so, you know, we're, we're doing a lot of work at kind of, kind of all possible directions on, on, you know, teaching models to reason better and for longer and, and, you know, in all kinds of different domains. Um, y- I, I mean, I, I think, I, I think we're training general purpose reasoning models and if, you know, kind of one... A lot of these behaviors that we're describing mathematically, like backtracking or kind of starting again, I mean, these are, these are not really specific to mathema- I mean, we're seeing them specifically in mathematics in these examples, but kind of they're general purpose tools for reasoning. And I think if you work hard at reasoning, you should see these patterns eventually.
- LLLisha Li
Mm-hmm. So it's just emergent because it's... I mean, I do think that's why the OpenAI approach, um, was so... I mean, it's like it doesn't rely on, you, you know, doing autoformalization in order to, like, guide the reasoning, and I think that's, like, obviously more like us. Um, but it's just, like, so not obvious that if you're just, like, training on, say, a corpus of like math proofs, maybe autoformalized in Lean, that you get, um, uh, get the sort of like projection of, like, how to think well. Like, put another way, like, with maybe, may- maybe if we think about it with code, like code is such a good corpus to train on because it's one of the few, um, data sets that has such large context. You just like... I mean, maybe you see this kind of with books, but they're less structurally interconnected. There's just, like, less structure there, I think, um, it's safe to say, like, on average in a, in a book compared to, like, a piece of code. Um, and so, like, with math papers, I feel like maybe what, what we're still bad at with coding models is stuff that that data set doesn't contain, which is, like, kind of the semantics. Like the syntax is there, but there's a little bit of, like, the higher level semantics of what produce, like, why do I have to write it this way? It's not... I- I'm kind of getting too, too much on the philosophical, but it is just, like, really interesting how it's still emergent, that it's doing good mathematics. And we'll probably get into this in more detail if you guys, you know, wanted to talk in more detail about some of the problems, which is just like it's, it's not just doing, like, the expected, like, we'll push the brute force thing. Like, you clearly are impressed with some of the reasoning traces, and it's just not obvious that's gleaned from, you know, what we would imagine would be the easy training set here.
- SPSpeaker
Yeah. Absolutely. I mean, I think this kind of thing is one reason we decided it was important to release, like, these summarized chains of thought for-
- LLLisha Li
Mm-hmm
- SPSpeaker
... these kinds of results. Because if you, if you've never seen these and you just see all these proofs coming out, you're kind of... You're not sure what it means. Like, like is the model just guessing in some insane way?
- LLLisha Li
Yeah, right.
- SPSpeaker
Like, is it thinking in some totally foreign... Like, what, what's going on? But, but actually it, it's, it's reasoning kind of shockingly like a, an expert human would.
- LLLisha Li
Yeah. Yeah.
- SPSpeaker
Yeah. It's, it's very much like reading a colleague's, like, notes. I mean, it's a little more disorganized in some ways. They kind of... Like, especially if you work close enough with a collaborator and sometimes you'll just see them, like, spill out their thoughts in an email to you, and it kind of-
- LLLisha Li
Yeah
- SPSpeaker
... it, it feels like reading a lot of those chained together. So it's, yeah, it's, it's very, it's, it's quite surprising the first
- 16:20 – 36:17
The Astra 10-Problem Set: Favorites & Deep Dives
- SPSpeaker
few times.
- LLLisha Li
Um, were you two sort of very involved in choosing the problems to, uh, to release in this, like, last 10 problem set that Astra was applied to?
- SPSpeaker
Um-
- LLLisha Li
Which was your favorite? [chuckles]
- SPSpeaker
Yeah. We were definitely involved. Um, do you wanna start on-
- LLLisha Li
Yeah, I mean-
- SPSpeaker
... sphere packing maybe?
- LLLisha Li
Yeah, I guess. Yeah. So I guess my personal favorite among these problems is the following. It's, it's extremely simple question, which is just like, it's just about how efficiently can you put a bun... My circles are not very good, and they're not all the same size, but-
- SPSpeaker
But we're assuming they are. [chuckles]
- LLLisha Li
Yeah. So the question is just, like, how dense can you place a bunch of, um... So you have a bunch of spheres, uh, you have a bunch of spheres of radius one and D dimensions. Um, so the question is: How densely can they pack? And so, yeah. So, so in two dimensions, it's kind of like the... So, so D equals one, this is not an interesting question, kind of. It's just the real line, and yeah, you can cut it up and, uh, a sphere in dimension one is just a unit segment, so, okay, you can cover everything.
- SPSpeaker
Mm-hmm.
- LLLisha Li
Um, so in D equals two, it's kind of the picture that you know, that everybody loves. It's just like, uh, it's just a bunch of spheres which sort of form, like, a hexagonal lattice.
- SPSpeaker
Mm-hmm.
- LLLisha Li
Hopefully I've drawn it well enough. But I can draw the hexagon.
- SPSpeaker
Kind of betraying my naiveté on this problem, is that, like, obvious? Is it, like, a very elegant proof-
- LLLisha Li
Yeah
- SPSpeaker
... that it's a regular lattice?
- LLLisha Li
Uh, it's not so... Yeah, it's not so obvious that this should work. It was only proved in the '60s, I think. Um, there's a short argument, but it's not, it's not so easy.
- SPSpeaker
Mm-hmm.
- LLLisha Li
Yeah. Um...
- SPSpeaker
What is the intuition? Like, what is kind of like the-
- LLLisha Li
I mean-
- SPSpeaker
... machinery of the argument?
- LLLisha Li
I mean, it kind of like-
- SPSpeaker
Yeah. I mean, it definitely looks like it should work. [chuckles] That's why I'm-
- LLLisha Li
Yeah. So I think-
- SPSpeaker
But so did... [chuckles]
- LLLisha Li
Yeah, I think this is the best part about this problem, which is really nobody had any idea how to solve this problem. [both chuckling] Um-
- SPSpeaker
Yeah
- LLLisha Li
... so, yeah. So I mean-
- 36:17 – 40:01
The Harness vs the Model: What Actually Matters?
- MSMehtaab Sawhney
coincidence.
- LLLisha Li
Yeah, yeah. It's like interesting when you're sort of saying the first prompt, which is, you know, maybe so basic, which is like, "Can you push this further?" It does require some judgment from mathematicians. Uh, but like- Eventually, you would imagine by scaling the models, you don't need to do that, or there's another view that the harness actually does matter, and this is kind of part of the harness apparatus. Do you guys have any views on that with your working with Astra? Uh, uh, especially generations of models and how, how much do you have to kind of input or how much the harness matters versus not?
- SPSpeaker
I mean, I, yeah, I, I guess there have been some, like, funny quirks like this that just come from, like, exactly what you ask the model to do, basically.
- LLLisha Li
Yeah.
- SPSpeaker
Like, like in, in this case, what the model was asked to do originally for codes was to improve the bounds-
- LLLisha Li
Yeah
- SPSpeaker
... by, like, some exponential factor. So it, it really, like, shows up in this, like, leading constant up here.
- LLLisha Li
Mm-hmm.
- SPSpeaker
Um, and you know, it, it improved the bounds and it, it didn't try to push things too much further. Like sometimes you see it do, but sometimes it just doesn't bother. But yeah, you know, you just ask it again and it, it goes further. So it, it wasn't like a capabilities issue, it just-
- LLLisha Li
Yeah
- SPSpeaker
... kind of didn't feel like it at the time.
- LLLisha Li
Do you call that judgment or, like, what is the... 'Cause like there, there is a... Yeah, what do you, you call that?
- SPSpeaker
I mean, it's, it's, models tend to be pretty task-oriented.
- LLLisha Li
[laughs]
- SPSpeaker
If you, if you, if you tell it to do a task and it accomplishes the task-
- LLLisha Li
That's good. I'm training for that. [laughs]
- SPSpeaker
... then it's pretty happy.
- LLLisha Li
So yeah, the, the task-orientedness, it's like, but do we expect that level to kind of ascend up to... It's not that they will be less good at being task-oriented, it's just like w- they'll ascend to the level of like, "Okay, no, let's, let's go in this direction." You'll have the judgment. 'Cause you guys had the judgment. You're like, "Okay, this is pretty promising. Looks like you're using a lot of representation theory. It doesn't seem like there's a limit so far." Um, but it doesn't have that context yet. But, like, I guess what I'm trying to say is like this one it's hard to maybe extr- harder to extrapolate, but from like previous generations, when you had to give it more, maybe prompting more of that harness work, but eventually probably had to give it less. So it probably gives you some confidence that there's this, like, really fast ascension. And, and do you see... Yeah, like what are some promises?
- SPSpeaker
I mean, somehow, somehow solving a harder math problem is, like, you have to solve many smaller, like, somewhat less hard math problems. And, and the fact that the math problems are getting harder is kind of an indication that you're, the model's able to take on more and more work in, like, a single continuous unit. Um, and I, I think that's, that's the thing that looks very promising. Somehow, like, a- any of these solutions, it's not like one idea, then you're kind of home free.
- LLLisha Li
Yeah, totally.
- SPSpeaker
You need several, you need several pieces to kind of interact and talk to each other. And the models, I mean, the, the model doesn't come up with all the ideas at once, right? It, it doesn't pull everything out of, in, in an instance. So kind of the fact that it needs to sort of see how this piece interacts with another piece, that's kind of like solving a problem in itself or piecing together many problems in itself.
- LLLisha Li
It could just be that, okay, when you're telling it, "Okay, push this even further," that was a- of the same order of, like, magnitude as, like, all the smaller things it's solving as well in, in between. And so you don't think that's as kind of like a privileged direction. It's just sort of like, "Hey, let's give it like one, one more help." Or you actually think that there's... I guess what I'm trying to get at, a bigger question is like, is there a good sense of like, you know, taste? Because, like, when people talk about, for instance, how well the models are getting at like, um, doing research [chuckles] for instance, because that's what we, we know we want a little bit of RSI. Um, and, uh, and sort of like there's surprising things about, um, how that improves, and then there's like the, oh, you know, maybe right now it's at a level of still like a junior researcher. It's like not really asking like the right problems. And so I'm just trying to get like a, maybe a sense of like where you're seeing that progress through the model advancements each generation.
- 40:01 – 57:32
What Even Is "Taste" in a Model?
- LLLisha Li
I mean, what is taste even?
- SPSpeaker
Yeah. I, I think somehow I tend to be p- pretty utilitarian in my view of taste. And like if you're able to solve problems faster by making ju- better judgments, like I think that's like the best like general proxy I have for taste.
- LLLisha Li
Mm-hmm.
- SPSpeaker
And somehow the fact that it's solving harder problems means it has, kind of by definition means it has better taste. I think there, there, there are these no... Yeah, I think occasionally, occasionally because they are task-oriented, you do occasionally get these, these symptoms of like, oh, it clearly has made a breakthrough. It kind of understands it's made a breakthrough, and then it doesn't kind of push all the way to the limit 'cause that's not what you asked. But that seems, yeah, that seems like, seems rather minor compared to like the set, the, the state of progress we've seen so far.
- LLLisha Li
Mm-hmm. Mm-hmm. Okay. Yeah. No, I think that's a pretty clear answer.
- SPSpeaker
Like I, I think it's, it's like maybe you're liable to get confused if you're trying to like do a concrete long horizon task and show taste kind of at the same time. But like, you know, if, if you, if you have like one model that's responsible for taste and one model that's responsible for going out and like, you know, working for a long time at, at solving a hard problem, kind of as the, as the like, uh, you know, uh, underling of the, of the, the supervising AI, I, I feel like that's kind of going to be fine currently.
- LLLisha Li
Oh, interesting. 'Cause that is like saying that they, these two things are somewhat sep- or if not separate, at least they shouldn't kind of pollute each other's context, which is a, a little bit... I mean, it could be potentially like a strong, stronger statement, um, than... I, I guess, yeah, no, it's just kind of interesting 'cause it's, it, it, it might just be, like to your point, it's, you know, let's take the utilitarian answer. It's, it's solving harder and harder problems. It's doing a lot more than just like, you know, brute forcing something. It's making choices. It's like pruning, you know, a vastly large, uh, space of possible paths into something that's like really, you know, is both tractable but then ends up, um, being like it's a diminishingly small path within that space. Um, but like having, like why would, what would be like a separate [clears throat] model, sep- separate generation or something that's a different version of the model that would contribute to taste? Or maybe that's totally, like it's too abstract, doesn't make any sense. You know, we should just let the actual... This might, like a related question would be like, you know, what is, what is the, the thing that gets us to a better version of intelligence, the harness and the model? Is it just the model? And it's like we, we see this in, you know, at least in, um- Applied AI or, you know, startups where it's like, it's a continual battle of like you need the harness, but then the harness a-adapts very poorly to a new model, 'cause sometimes like a very, very minimal harness is still the best way to expose to the raw power of the model. But then now we also have these like training regimes where we require the harness to be, you know, trained with... Maybe part, part of this is to keep things more proprietary and harder to, harder for other people to use it. But I think partially it's, it's maybe actually that it helps, um, have more control on like the reasoning traces you care about. It's a long rambling way of saying it's like, yeah, I, I don't actually... Like I, I, this is so interesting to see how the models have gotten be-better at math and, um, maybe something that's like very abstract and hard to describe, like taste, is a way to tease out like what is actually necessary here.
- MSMehtaab Sawhney
I think my only like nontrivial thought here is that like when you're working, I mean, just when you're doing any task, occasionally you get pigeonholed and you like work really hard and just having a friend look over your shoulder and be like, "What are you doing?" And then just like, just having that one bit of like step back for 10 seconds, like this is often very useful.
- LLLisha Li
Yeah, yeah.
- MSMehtaab Sawhney
I mean, I see no reason why humans would be so different than models somehow.
- LLLisha Li
Yep, yep.
- MSMehtaab Sawhney
Having... Or models would be so different than humans. Having, having a few humans working together is often more powerful than just having one.
- LLLisha Li
Yeah. It's like in this kind of collaborative thing, you actually, you kind of, yeah, artificially created it, but it's very similar and dynamic.
- MSMehtaab Sawhney
But I, I think a lot of taste is also like having a sense of what problems you or like some method you have in mind are going to be good at solving.
- LLLisha Li
Mm-hmm.
- MSMehtaab Sawhney
Like it's... I mean, certainly there's some amount of like absolute aesthetic point, right? But there, there's also just like, you know, having a nose for what, what y-you might want to pursue because you'll be able to make progress. Um, and you know, I, I think for that, like there's, you know, you, you would expect that as a side product of being good at completing tasks, you would, you would get there sort of, right?
- LLLisha Li
Let me know if we still wanna do like a section on sophic groups, 'cause I think, you know, up to you guys, it's definitely super interesting.
- MSMehtaab Sawhney
So maybe the, the first question is what is a group? Let's remind ourselves. So, uh, a group is, uh, a set of elements with some multiplication operation. And, uh, basically this is how mathematicians think about symmetry. So, so you're like... Basically like if G and H are elements of your group, then GH has some, is some other well-defined, uh, element of your group, and you have like, uh, associativity, uh, and you have an inverse. So for every G, there's some inverse, uh, and there's some like specific element in the group that, uh, is, is kind of the identity. Uh, okay. So, you know, it's, it's some like abstraction of like, uh, composing operations. So these could be like numbers, they could be like multiplying matrices, they could be like, like rotating something, which is, you know, a special case of multiplying matrices. Um, and, uh, a group is, uh, sophic if... Well, there's some, uh, you know, precise definition, uh, but, you know, uh, roughly, uh, it means it... Uh, so, so I should say like groups that can be finite or infinite. So like, you know, um, if, if you have like a square, like all the rotations of it form a group with like four elements.
- LLLisha Li
Mm-hmm.
- MSMehtaab Sawhney
If you have like a circle, then the rotations form a group with like uncountably any, many elements. Um, and, uh, so sophic groups are either finite or countable. Uh, you should think of them as being countably infinite, so there's like the same number of elements as like the integers. And, uh, if it's sophic if in some sense, uh, it can be, uh, approximated by finite groups. So we didn't know if there was a non-sophic group. So, so-
- LLLisha Li
Yeah
- MSMehtaab Sawhney
... the, the, the result that, uh, Astra proved is simply that, uh, there exists a non-sophic group.
- LLLisha Li
Yeah. And without like, I mean, we can, you know, before going to, to that proof, it is like, you know, it's like I feel like a lot of the programs in math is like, okay, we are such finite creatures. Let's see how well our finite approximations are, you know, uh, do. And in this case, especially for the countable case, it's like, uh, maybe you'll, you'll be relating it to like the Aldis Leon's thing. It's just like, it helps kind of anchor the picture of like, it seems like such a... I mean, it's a nice result if it were true, but it, it's not, and it seems almost like reasonable. And so, yeah, I, I actually didn't, um, didn't, uh, go... I would love to hear the explanation of like how it, it found a counterexample.
- MSMehtaab Sawhney
Yeah. I mean, I, I would say that like, you know, the, the hope that there was no non-sophic group, so every group has this kind of approximation. Like maybe this is sort of like people hoping that there's a miracle.
- LLLisha Li
Mm.
- MSMehtaab Sawhney
Because, uh, it turns out that groups like this have a lot of nice properties, um, because, uh, you can run certain proofs for finite groups and then, uh, you know, kind of approximate them in whatever way the definition of being sophic lets you approximate them-
- LLLisha Li
Mm-hmm
- MSMehtaab Sawhney
... and, uh, get the result. So, so like, um, there's this notion of being a conjunctive group. So there's, uh, there's some fact that, uh, any, uh, group which is sophic is, uh, also conjunctive. Uh, surjunctive is some property of like, um, dynamical systems on the group.
- LLLisha Li
Mm-hmm.
- MSMehtaab Sawhney
And, um, I guess the original question was whether every group is, uh, surjunctive. This is some question of Gottschalk from the '70s. Uh, and this, uh, this fact that, um, follows this, like, pattern of prove it for finite groups and then do this approximation is, is what motivated the question about if there's a non-sofic group.
- 57:32 – 1:00:01
How Should the Math Community Adopt AI?
- SPSpeaker
in this case.
- LLLisha Li
Yeah. Well, actually maybe that's a great segue into like how, you know, what's the ideal way that this is being taken up by the math community? 'Cause I feel like there's a spectrum of answers from working mathematicians, a sense of like some, you know, uh, probably most at this point are like, "Okay, AI's obviously doing some non-trivial stuff. Um, it would be a disadvantage not to admit that in my workflow." Um, I've definitely heard some stories where people are kind of, you know, would find it hard to either take AI as a co-author or like how do you even do kind of attribution this way? But I don't know, like what, um... But maybe to paint the, the more optimistic picture, so you're saying you want the mathematicians to be building on this, these results. It definitely generates a lot more results to be verified, so it, you know, it puts pressure on, on the community and the profession. Like how do you, how do you kind of expect the evolution of kind of uptake and, and, and collaboration with mathematicians?
- SPSpeaker
Well, I mean, I mean the, given that the fact that the models can produce sophisticated mathematics means that they can help you understand like sophisticated mathematics. I mean, like I don't know, occasionally like I enjoy looking at the archive, and I want to understand some proof, and like I could read the introduction, but in practice it's just much faster to take the PDF, put it into, put it into my favorite model, and then like get a, get an output of like what is the rough proof strategy.
- LLLisha Li
Yeah.
- SPSpeaker
And somehow this like... A- along... I mean, of course, models are going to help us produce exponentially more mathematics, but they also make it much easier to absorb it. Um-
- LLLisha Li
Yeah
- SPSpeaker
... and right now, okay, it's still a bit of a challenge back and forth, but I, I think it's, for me at least, a, it's much, much faster at understanding. Uh, it's much, much faster to understand a piece of mathematics with a model than without it. Um, so it's helping solve the problem it creates anyways.
- LLLisha Li
Yeah. I feel like that at least... And it, it's, you know, I, I don't, I don't view it as like creating much more problem, but again, like I don't have such, you know, high stakes and like, okay, I'm, I'm gonna get... I'm not gonna get tenure, et cetera. So like I, I, I agree, like making it more accessible, like if I'm not spending so much time absorbing an area, I can like put it into ChatGPT and then expect to... I mean, you guys have an even more powerful model, hopefully releasing, um, for other people to enjoy as well. But like it's, um, I think like the, the positive version of that is actually more people can participate in mathematics. It's like people might be coming with other intuitions, and they could actually maybe generate good mathematics. Is that sort of like closer to the vision of what you're hoping this is, you know, pushing towards?
- 1:00:01 – 1:05:00
Empirical vs Theoretical Math & the Positive Vision
- LLLisha Li
Um, or like what, what, what things do you think mathematicians should be wary of to, to kind of adapt fast enough to take advantage of AI?
- SPSpeaker
Yeah. I mean, I, I think certainly there will be a lot of changes, right? Like I, I guess in math, like there are, there are a lot of things that are kind of important for, for like a given result, right? You need someone to come up with it, but you also need people to understand and absorb it and like, you know, internalize it enough to, to do more with it and, and like figure out where it fits into like humanity's understanding, right? And like, uh, a, a couple of years ago, uh, like the proving the result was like so hard that kind of the, the other stuff was just kind of coming along for the ride, right? You know, like if you, if you manage to like prove this thing yourself, you're automatically gonna understand it quite well. You're kind of responsible for like maintaining it in, in some sense and like, you know, explaining it to other people. Um, and yeah, now this kind of what was the main bottleneck before is kind of, um-
- MSMehtaab Sawhney
Much less of a bottleneck and, you know, these other kind of constraints, uh, come into play. Um, so it... Yeah, the, the, the, like, optimal, um, structuring for, you know, organizing the knowledge could, could look rather different.
- LLLisha Li
Yeah. How does that look? I mean, d- does this make the field a lot more kind of empirical? Will people do sort of the hard, like the first thing that was scarce, which is like all the reasoning and then more? I mean, not that it's like a bad thing to make it empirical, but it's almost like it functions as a very different discipline. Um, like a lot of the fun stuff is understanding, you know? Um, and so going to Shannon, communicating, maybe assembling, having still the human taste. Does that sort of remain re- rarefied and, and that's how, you know, current mathematicians need to adapt and, and reward, [chuckles] you know, uh, contributions or, or is this too much of a caricature that's like something else?
- MSMehtaab Sawhney
I think certainly understanding how to put, as we get more and more mathematics, put it in like a proper framework and sort of how, sort of like being able to explain it to other humans so that they can also appreciate it. I mean, somehow implicitly we valued this, but it was usually because you were the person proving the results that gave everybody else the understanding. But I think increasingly it would be a function of, like, you're sort of helping, y- you're the human who can sort of give this understanding to other people and sort of help them with it. Um, I think that more c- sort of, that communal understanding will, I think, become... It was much more implicit in how we viewed mathematicians, but I think it will be an increasingly more explicit and valuable part of the subject. I mean, a, a nice thing about math is that, um, the, the ceiling for difficulty of a math problem is pretty high. So even if, um, you know, kind of, e- even if AI get, you know, continues getting like exponentially better at math, like it might, you know, it's plausible we'll never solve something like P versus NP.
- LLLisha Li
Mm-hmm.
- MSMehtaab Sawhney
And it, it could be that, like, the, the field kind of becomes more, um, you know, attached to like, like these big mysteries-
- LLLisha Li
Mm.
- MSMehtaab Sawhney
... and, and less to like smaller mysteries that-
- LLLisha Li
Mm
- MSMehtaab Sawhney
... are more like routine now.
- LLLisha Li
Yeah, yeah. I think that's a positive vision of the-
- MSMehtaab Sawhney
I mean, also like, I don't know, there are things I spent like months or years of my life wondering about, not getting to know and hopefully I get to-
- LLLisha Li
And now we get the...
- MSMehtaab Sawhney
Yes.
- LLLisha Li
Yeah, that's like, what a joy.
- MSMehtaab Sawhney
Some portion, yeah, some portion of them I'll get to know the answer to. I'm, I'm pretty happy about that.
- LLLisha Li
Yeah. No, exactly. No, I'm, I'm excited about this like renaissance of results and understanding, and I feel like, I mean, this is such an, an infinite, you know, field. [chuckles] Like, okay, no, no pun intended, but like it's just like, it, it's, it's, it's just there's so much that you can actually create here. Um, so I mean, especially if, for somebody like me who's not gonna have the time to actually like practice mathematics, now there's like a lot more that you can actually do in the activity of math. So yeah.
- MSMehtaab Sawhney
Yeah, I think the, the like, the ability of someone who's not working on math as like their literal job all the time to like understand what's going on and like, you know, learn about some of the mysteries they might have wondered about will, will go up quite a lot. Um, also, you know, if you're, if you're like, if you're working on something that requires some math-
- LLLisha Li
Mm-hmm
- MSMehtaab Sawhney
... you know, suddenly you, you don't need to like find a world expert on this-
- LLLisha Li
[chuckles] Yeah, yeah
- MSMehtaab Sawhney
... topic to, to be able to, you know, use it in your own work. You can-
- LLLisha Li
Sorry, mathematicians. [laughs] Yeah, yeah, yeah. No, it, it's true. I mean, I think there was just like a dearth of actual like people who could, could do that, and so I think this is helpful. Maybe it's helpful for theor- theoretical physics, like we'll see. Um, but a lot of other applied areas as well.
- MSMehtaab Sawhney
It'd be nice for the world if applied mathematics went a lot faster.
- LLLisha Li
Yes. [chuckles] I mean, I'm, I'm of that opinion. Uh, well, thank you guys for joining. This was a lot of fun, and I'm, you know, just so excited for how much the models are advancing. So maybe we'll have you guys back soon.
- MSMehtaab Sawhney
Yeah. Thanks so much for having us.
- LLLisha Li
Yeah.
- MSMehtaab Sawhney
Yeah, thanks for having us.
Episode duration: 1:05:15
Install uListen for AI-powered chat & search across the full episode — Get Full Transcript
Transcript of episode 1JvyLGd2Sfs
