Skip to content
Dwarkesh PodcastDwarkesh Podcast

AI researchers debate how close we are to recursive self-improvement

New episode with John Schulman, Charlie O’Neill, and Beren Millidge. I got together with some of the most insightful AI researchers I know who are at the openish companies, because I wanted to hear the details of what's actually happening at the frontier and what comes next. π„ππˆπ’πŽπƒπ„ π‹πˆππŠπ’ * Transcript: https://www.dwarkesh.com/p/john-beren-charlie * Apple Podcasts: https://podcasts.apple.com/us/podcast/ai-researchers-debate-how-close-we-are-to-recursive/id1516093381?i=1000789067132 * Spotify: https://open.spotify.com/episode/0ePd4PUqCpN78hCjVRH0fr?si=wGvk7u5XQwaLfyrvysdIJQ π’ππŽππ’πŽπ‘π’ * Antithesis helps you trust your code. As agents generate more and more of your software, the bottleneck shifts from your engineers actually writing code to verifying it. Antithesis does that testing for you. Ron Minsky, who co-leads Jane Street's tech group, told me that Antithesis was able to help his team shake out bugs in software that had already undergone heavy review. If you want to see how it fits into your development process, go to https://antithesis.com/dwarkesh * Jane Street just launched its most ambitious competition yet: design a protocol-emulator ASIC. Basically, if you have a chip you want to test outside of a live system, you should be able to connect it to your design and have it simulate realistic traffic. Jane Street wants general-purpose, reprogrammable designs that can work across multiple protocols and remain useful as new ones emerge. The most novel submissions will actually get taped out, and the winners will receive a physical copy! The competition is open until January 18, 2027, and teams are encouraged. To get started download the template code at https://janestreet.com/dwarkesh * Grok Bot has been a great way to hand off tasks. My team uses it as a producer: whenever my editor posts a rough cut of an interview in Slack, Grok Bot opens the transcript on its own computer, matches my notes to the exact moments they refer to, and uses a file of my preferences to suggest edits. Then it sends me its top clip candidates so I can review everything from my phone, which saves my editors from sorting through hours of footage. Try Grok Bot for yourself at https://x.ai/bot To sponsor a future episode, visit https://dwarkesh.com/advertise. π“πˆπŒπ„π’π“π€πŒππ’ 00:00:00 – Steelmanning the case against RSI 00:18:39 – What’s driving the Chinese labs’ progress 00:28:06 – How will automated AI researchers be trained 00:33:51 – Will long-horizon RL elicit AGI? 00:45:24 – The sim-to-real gap 01:00:33 – How much progress is explained by data? 01:18:03 – Why is RL working so well? 01:24:54 – Move 37 and entropy collapse 01:28:32 – Rapid-fire timelines

Dwarkesh PatelhostBeren MillidgeguestJohn Schulmanguest
Sep 11, 20261h 37mWatch on YouTube β†—

EVERY SPOKEN WORD

  1. 0:00 – 18:39

    Steelmanning the case against RSI

    1. DP

      Today, I'm chatting with three of my AI researcher friends from whom I learn a lot every time we talk, and who also happen to be at somewhat open-ish, uh, labs and companies, so you guys can actually, um, say things on the record. I'm joined by Beren Millidge, who is the CTO of Zyphra, which is developing open source models. John Schulman, who is the chief scientist at Thinking Machines, previously the co-founder of OpenAI, led the RLH effort that led to ChatGPT. And Charlie O'Neill, who is head of model training at Base10. The first question I have, if we're in twenty thirty-six, it's been ten years, and we don't have like cra- billions of crazy super intelligences that are running around that are like radically transforming the world, what is the most likely reason that, that doesn't end up being the case? Other than sort of exogenous political shocks, or like there's a war or they ban AI or something. But what is the most likely technical reason that we don't-- that twenty thirty-six isn't like a crazy alien super intelligence world?

    2. BM

      I mean, like, my reason would just be, like, it's got to be the sort of-- Like, there's been a classic thing, almost like Moravec's paradox, right? Where, like, we see, like, you know, we think of the AI being like, "If it can do this, it's going to be amazing," right? Like, if it can solve these hard maths problems, if it can win at chess, blah, blah. And then it solves these things, and then it's, like, not that impactful. Obviously, it's somewhat impactful, but, like, not everything. It's like, if somehow that continues and, like, there's never, like, the true, like, spark of generalization that occurs, I think that could lead to, like, the AIs just being, like, extremely good at kind of everything that people, like, put into a benchmark, put into an environment, but, like, there is still some persistent, like, sim-to-real, which is somehow blocking everything. I think this is kind of unlikely. I think we do actually see this kind of generalization even from our LLM practice already. But, like, if it is just, like, ridiculously hard to, like, generalize meta-learning, plus, like, we don't solve continual learning and it's just, like, super hard and impossible.

    3. DP

      Yeah.

    4. BM

      Like, this would be my, like, default scenario in that case.

    5. SP

      Yeah, I agree with that. Uh, humans, uh, have a lot of ad-advantages over models now, and, uh, each time a new c-model comes out, it'll sort of, uh, it, it'll catch up in some of these areas. Um, but, uh, like, you end up getting bottlenecked by the places where the model is weaker and where it has, uh, worse judgment or, um, the models can't check themselves well enough. Yeah, so, so there's this, uh, cycle that keeps repeating where people think, uh, where a new model comes out and people are blown away and they're like, "This is it. This is the-- This is AGI." But then, uh, they use it a bit and, and then it starts to feel dumb after a month or so. So that cycle just might keep going, and it's hard to predict how many times it's, it's gonna repeat. And, uh, like, right now, you don't get explosive growth, um, in capabilities because, uh, you still get bottlenecked enough when you're trying to do research and engineering that even if the model can write way more code than a person, it doesn't make you, like, a hundred times more productive. Uh, but yeah, so, so maybe, maybe there are just more of these cycles than we would, uh, than we would expect.

    6. SP

      For me, it's like a question of how far off, like, this global optimum of a learner you could have on a chip is, like, the transformer plus, like, RL, basically, like, the current recipe.

    7. DP

      Yeah.

    8. SP

      So, like, I think people imagine that even once-- Like, once you have a, a, an agent which is better than all humans at AI research, even if it's like point one percent better than all humans, then the fact that you can run, like, you know, hundreds of thousands, if not millions of these in parallel, um, you can run them much faster, like chip's gonna speed up, that's gonna outweigh every other, like, bottleneck and, like, you're eventually just gonna, like, hit this, like, very fast takeoff with recursive self-improvement. I could imagine that if we continue along the trajectory that we're currently on with that paradigm where, you know, it's basically just, like, self-attention, RL, scaling up RL environments, the-- I guess, like, i-if you think about what happened with Moore's law, right? Like, we had this very, like, nice straight line and that held for a really, really long time. But there were so many, like, discrete, like, discontinuities and innovations that had to happen to keep that scaling law going, and the same thing has kind of happened with LLMs. Like, we had this pre-training, like, scaling law, and then that was kind of like, you know, hitting the diminishing returns. And then we came up with, like, RL and solved that, and then we got this new, like, you know, diminishing returns curve to hit that made it keep looking like a straight line going up. And so, like, if it requires another one of those discontinuities to solve, like, I'm not sure that, like, the current method of, like, training LLMs with these RL environments, even like RSI-targeted RL environments, would be able to discover that discontinuity. And if not, like, we're probably gonna hit this, like, asymptotic, like, curve where, like-

    9. DP

      Sorry, but do, do you think the discontinuity will be harder than anything that's come since twenty twelve?

    10. SP

      If, if we had the answer to that, that, that we'd kind of have the ability to implement it. But, like, maybe there's the distinguish-- we should distinguish between a discontinuity which adds to the current paradigm. Again, it's like cumulative, like there's something-

    11. DP

      Yeah, yeah

    12. SP

      ... beyond the RL that we have to discover, and maybe they're capable of like, you know, connecting the dots in that straight line. Or like, but, like, again, how far off the, the global optimum are we? Do we have to go back and throw out like, you know, gradient descent-

    13. DP

      Yeah

    14. SP

      ... and, like, neural nets in general?

    15. DP

      Yeah.

    16. SP

      And I don't think, like, if you continue to scale up the current paradigm, an LLM, no matter how many LLMs you're running, are capable of necessarily discovering that if it's too far away.

    17. DP

      Yeah. The, the only hope really is if deep learning just can't get us to an AI which is at least, uh, can dominate human research and human development, including the human ability to come up with new paradigms and so forth. Or like, I don't know, maybe hum- maybe humans would nev- also never have discovered the, uh, the next learning architecture, but, um, to the extent humans could have discovered it eventually. But it just seems like, I don't know, if, if you just look at the progress that's happened since twenty twelve till now, and you just continue that on, I mean, I know it's, it's been powered by huge amounts of compute scaling and so forth. But, um, uh, it would be weird if, like, it just didn't get to the point where it could, like, dominate humans, at least in R&D. Especially over the next few years, there's gonna be-- Ryan Greenblatt was on the podcast recently, and he made this point that you could imagine as AIs get more and more capable and are capable of making progress on

    18. SP

      Simulations which incentivize getting better at not only AI R&D, but generally at science. So this is a thing that all the labs are targeting, many startups are targeting. Or the, another intuition pump is if you look at the Elo score of chess bots since the '80s, there's just, like, a very linear increase in Elo over time. But there's this huge discontinuity as they cross the human range of human experts always win against AIs to, like, human experts never win against AIs as this linear increase in Elo happened. And you could think-- I agree with your point that th- so far, AI capabilities have not been that big of a deal in terms of their end economic impact in the world. But that's just because, like, they're slowly rising in Elo relative to humans.

    19. SP

      Yeah.

    20. BM

      Yeah, I mean, I agree it would be very s-- I mean, the only way for this to not happen is if, like, as you said, somehow asymptoted, like, just before basically, 'cause we're already pretty close, in my opinion, to, like, where we'll start c- crossing, like, the human Elo score. And so we'll need to asymptote before that. And, like, that's the only way, you know, in the scenario you posed where, like, somehow we're sitting here in twenty thirty-five and, like, everything is normal for this to happen, I think. The, I mean, the only other way is, like, there's, like, some dramatic, like, regulation on AI.

    21. SP

      Yeah.

    22. BM

      It's like, this is kinda what I see as, like, the most likely way for this scenario to happen, actually-

    23. SP

      Right

    24. BM

      ... rather than a technical thing.

    25. SP

      Yeah. I, I, I think there's different kinds of research. There's, like-

    26. BM

      Yeah

    27. SP

      ... research where it's, like, the auto research style where the objective is already specified very cleanly-

    28. BM

      Oh, for sure, yes

    29. SP

      ... and you're optimizing that objective. And I think everyone is picturing, like, if we continue along this path of, like, you know, making pre-training loss go down, making RL environments-

    30. BM

      Yeah, yeah

  2. 18:39 – 28:06

    What’s driving the Chinese labs’ progress

    1. DP

      What is the story for why there, there isn't huge consolidation in model providers? There's just so many things that point to centralization here. Is there, is it-- And, yeah, if you, if you step back over the course of years, is there something that is gonna prevent that?

    2. SP

      Yeah. I think, um, distillation is the main thing that, uh, fights against the centralizing force. Uh, 'cause basically anything that can be learned, uh, through RL can be distilled very easily because, uh, um, it, like, it's a small number of bits. Uh, it's something that you can learn from a small amount of data. So if you can collat-- if you can get, uh, trajectories from the model that, uh, show a behavior, you can easily distill it. So I think, um, I think distillation is one of the, uh, things that fight centra-centralization. Uh, there's also, um... I mean, there is a possibility that there'll be company-specific, um, models that it'll, it'll be possible to learn from deployment, um, and, uh, ha-have a company continually improving its, its own model.

    3. DP

      Mm.

    4. SP

      Uh, and, um, such a system could be provided by the, uh, current oligopoly of model providers or some other, uh, currently smaller company. Uh, but I think that'll, that'll change the game a bit.

    5. BM

      Yeah. And I also wanna point out that, like, continual learning and RLC doesn't stop distillation, right? Like, even if your model is improving every day, like, y- people could be distilling it every day.

    6. DP

      Yeah, yeah.

    7. BM

      And so it's like the loops could just operate at the same pace.

    8. DP

      Right. That makes sense. Th- th- okay, so copying model behavior. Um, you-- I, I guess you need to know yourself what the right distribution to prompt is in order to get, like, the relevant model behavior.

    9. SP

      Oh, yeah. For, um, just distilling, uh, with supervised learning, um, the prompt distribution is extremely important, so it's, uh, very non-trivial to distill a model even if you have full access-

    10. DP

      Mm

    11. SP

      ... to it and have the CoT, the, the chain of thought and everything. Uh, uh, yeah, it's, it's non-trivial to distill all of the, all of the useful capabilities from it, um, because you need to prompt the model with something. And you need to prompt it with, uh, y- like, realistic prompts. Um, you, you, you need to have a really wide distribution of realistic prompts. So yeah, one thing that's been coming out recently is, um, some of the, some of the Chinese companies are probably using these, uh, router services which are designed to allow people in China to use, uh, the US frontier models, which would otherwise be blocked in China. Uh, but there are all these, uh, router or proxy services that allow people in China to use these models mostly for coding and, uh, a- and these router services are collecting and selling some of the data. So I think this is, uh, like a very useful dataset for, uh, distillation 'cause it gives you the perfect prompt distribution.

    12. DP

      Yeah.

    13. BM

      I think this is one of those things where AIs help a lot here. Like, if you actually look at, like, you know, the frontier pipelines of, say, like the Chinese models that they've actually put in their papers, it's a lot of, like, humans or, like, they get seed prompts from somewhere, which is some combination of humans, this kind of data, and then they, like, synthesize a vast coverage from those seed prompts using their existing models or, like, the other frontier models. And so it's like you can automate, like, an awful lot of this, like, prompt distribution gathering and, like, environment creation. It's just, like, humans need to provide, like, increasingly fewer amounts of bits as, like, the models get better.

    14. DP

      Right. Or certainly it seems, it still seems you're bottlenecked by, um, like having a service which has users or users are going through. So, like-

    15. BM

      Not necessarily. I mean, like, yeah, that's obviously very helpful, but, like, theoretically, you can just think about, like, what users want or, like-

    16. DP

      No, but it's-

    17. BM

      ... a lot of tasks

    18. DP

      ... like the whole point is that we don't-- the user says, "Make me an application like this. Oh, that didn't work. I actually want you to make this new feature. But actually, l- let's step back and do this other thing." And capturing that whole trace is the-- Or to the extent you could have done that anyways, then you just have, like, RSI anyway.

    19. BM

      Yeah. I mean, like, ultimately, like, if you have this, like, fully automated loop, that is basically RSI, right?

    20. DP

      Yeah.

    21. BM

      Like the AI is deciding the data, it's deciding the training. That, that is the loop.

    22. DP

      Right.

    23. BM

      But yeah, I mean, like, it depends how much human information you need. Like, at some point, if you're just like, "I want traces that look like this," you prompt that to the model, the model will be able to, like, come up with, like, a pretty good approximation.

    24. DP

      But what if you wanna do, like, "Make me a really good politician," and then it has to, like, anticipate de novo? Like how, how would a discussion in, like, the Senate halls go or something? I just feel like-

    25. BM

      Yeah. I mean, like, this is-

    26. DP

      ... there's gonna be a lot of things which are-

    27. BM

      Ironically, this is actually, I think, easier for the distillers than the frontier labs, right?

    28. DP

      Right.

    29. BM

      'Cause the distiller's just like, "I want a good politician." They go to the, like, frontier model. The frontier model already knows how to be a good politician, so it just, like, generates those traces. Whereas, like, if you actually wanna build the first model that does this, you have to, like, actually somehow, like, get data on, like, what politicians do every day-

    30. DP

      Sure, sure, sure

  3. 28:06 – 33:51

    How will automated AI researchers be trained

    1. DP

      Okay, the, the other question I had is how the first models that are capable of automating AI R&D will actually be trained. 'Cause there's a toy version, which is this thing that Ryan was talking about, which is you just have GPT-8 try to build GPT-3 size models that are really good at like inner loop type challenges of beating video games that require continual learning or, um, just get, getting to a certain loss with like the least amount of compute, etc. But John, I think you had an interesting point that maybe that's not the way it actually will happen in practice. So I'd be curious about, yeah, by the point at which you have AIs that are actually capable of automating AI R&D, how are they probably trained?

    2. SP

      Yeah, I think we'll probably do some combination of learning from human feedback to absorb, uh, like the researchers' taste, uh, and, uh, just like creating a lot of practice environments, uh, which involve like doing multi-step research projects. So I think, uh, yeah, people will in practice do some combination of those two things and, uh, just, um, each iteration like patch whatever seems to be most broken in the last iteration. So, uh, like researchers will be using, uh, the AIs, uh, a lot and, um, and will notice that they have some consistent weaknesses and then, uh, those things will either be patched by, like collecting human feedback or, uh, like creating environments.

    3. DP

      Yeah. Makes sense.

    4. SP

      Maybe, maybe a useful way to think about this is like how much of the lineage we roll back and then let self play from there. Like, I think i-in the limit, like you're picturing like, you know, just giving them like a GPU and maybe neural nets or something and saying like, "Okay, figure out how to train a, a model to like do these particular tasks." Like the way it currently works is like we go up to the very, like, edge of the, of the lineage and say, "Okay, like here are the bugs like, you know, Anthropic has found in their training stack in the last few months. We'll turn those into environments like you need to train and get better at on the frontier." And so you obviously lock in all the previous history of the lineage, but m-- you could imagine a world in which you roll back to like, you know, before GRPO or something, and then you have environments which like- Trying to get it to discover, like, the best full form to, like, RL-

    5. DP

      Yeah

    6. SP

      ... um, models on, and then maybe roll further and further back. But I think we will be still so compute bottlenecked that, like, people will just keep, like, staying at the frontier and, like, diffing essentially the bugs and whatever improvements they found since the last model version, turning those into training environments.

    7. DP

      Which is also really good for having non-stale, like, n- new data between model generations. It's just, again, this is basically continual learning within the-

    8. SP

      Mm

    9. DP

      ... AI lab, of distilling the last three months of AI research progress through environments and, like, RLHF-type stuff back into the model itself.

    10. SP

      And, and it is distilling, right? And that's maybe why some of us feel like it's asymptotic, is like you're always, like, just trying to get the last three months of progress. And that, that progress is being contributed to by AI, of course, but it also still has humans in the loop, and it feels like, you know, you're just constantly inching closer and closer to what the human researchers are, like, finding and capable of doing.

    11. DP

      Yeah.

    12. BM

      I mean, the one thing I will say, though, is, like, obviously, if you're just distilling on, like, trajectories, you can never go above it, but environments can go quite a far way above what a human can do.

    13. SP

      Yeah.

    14. BM

      Like, it's very easy to design an environment that, like, no human can solve, but the AI can obviously still try and solve it. And so that would be the path to, like, go ahead of just, like, what the human AI research is.

    15. SP

      Do you have, like, an example of, like, in terms of RSI or, like, like, you know, what kind of trains-

    16. DP

      With the human speed run, but doing it even faster than a human speed runner.

    17. SP

      Yeah.

    18. BM

      I mean, I feel like in AI research especially, it's very easy to define, like, goals which, like, you know, you could say, like, the loss needs to be, like, one point three or something, and, like, no human can get that, you know, now. But, like, that's a very extremely measurable, verifiable task, and the AI-- if the AI gets that, then, then great.

    19. DP

      Or, I don't know, building, like, a hundred million parameter model that beats Minecraft. That's maybe too easy, but, like, beats, like, a much more complicated game or something.

    20. SP

      Isn't it crazy that a hundred million parameter models beat Minecraft? We're calling that too easy? Like, imagine if you said that, like, five years ago. [laughing]

    21. DP

      [laughing]

    22. SP

      I would say a lot of research is not exactly like that, though, where it's like hill climbing on a well-defined goal.

    23. DP

      Mm.

    24. SP

      It's sort of more like, uh, here's an intuition we have about, uh, some way models should be better, and then we also have, uh, some idea for an algorithm that seems to go a little bit in this direction. So let's come up with a task that is, uh, s- sort of designed to show signs of life on this approach-

    25. DP

      Mm

    26. SP

      ... and, uh, like, see if we get some, uh, get those signs of life, and then if we do, we can make successively more realistic versions of the task.

    27. DP

      Right, it's, like, a lot more guided by intuition, and then the, the out-- the inner loop is to elicit the, uh-- or make r- r- test for that intuition rather than, like, the, the, the test itself leading to the insight.

    28. SP

      Right. Like, you're not directly optimizing f- uh, for, uh, the e-eventual objective you care about or the practical-

    29. DP

      Yeah

    30. SP

      ... like, production objective. It's, it's sort of, uh, you're, um, you're relaxing your objective a little bit. You're saying, "Yeah, let's relax on the realism axis a little bit and find, uh, some methods that actually work, and then, like, then try to get back to realism later after the method matures a little bit."

  4. 33:51 – 45:24

    Will long-horizon RL elicit AGI?

    1. DP

      Maybe taking a step back. Here, here's what I, here's what it seems to me that the plan, uh, for AI research going forward is, and you tell me if you think it's gonna work or if you agree with this characterization. So the bet is that we will scale up RLVR training across millions of diverse environments, across hundreds of different kinds of domains, and what will emerge at the other end is an agent which has, like, learned these basic skill-- or less than basic skills around being persistent, um, being able to triage information and context, uh, eventually having, like, end-to-end optimization of working with other agents and things like that. And such an agent will be very sample efficient within the context. You know, you have done research on how you actually scale up in-context learning to make it, like, arbitrarily long, but you just keep scaling it up. And so what comes out the other end will something, will be something that it basically functions like a drop-in remote worker over the course of a week or a month. First of all, do you agree that that is the bet the labs are making? And second, is it, is that enough? Like, basically, learning how to learn within these simulacra within a data center, uh, and then, but getting deployed into the real world, but not actually, like, learning from real-world deployment, only learning these meta skills from s- the s- the simulated environments in the data center.

    2. SP

      Yeah. I, I think it's now hard to separate out, like, how much of the lab's effort is going towards, like, direct RSI versus, like, making generally intelligent models that they can continue to deploy to collect revenue to fund the next big training run.

    3. DP

      Yeah.

    4. SP

      I think for the, um, latter, like, yes, that's probably just the bet they're making. Like, and it, and it's very clear, like, the pattern of, like, where these environments are going over the last few years. I mean, like, Anthropic's lineage of, of environments is, like, a very clear example of this. Like, you know, first, they, like, we just focus on coding and, like, we're gonna get really, really good at that. And then the task horizon that we've got from coding, which is probably the lowest-hanging fruit in terms of, like, data available on the internet to create environments, like their own internal stuff that they can turn into environments. Then we're gonna generalize. We're gonna go after finance next and, like, literally, like, just so much Excel data and, and all that sort of stuff in the, in the RL training. And then, you know, it's PowerPoints. It's, like, this long tail of, like, the working economy and, like, that seemed to work really well and, like, a lot of the other labs and thing, even the open source labs have now realized that that was the correct bet to make.

    5. DP

      And so but, but what is the implication from that? Uh, when I had Dario on the podcast, the thing I asked him was, if you truly expect models which will be human-like in their ability to learn on the job- Why would you try to bake in all these, uh, skills of, like, working with PowerPoint or something? Wouldn't you just expect the model to be able to pick that up on, while it's deployed? And so, yeah, there's multiple different explanations. One is just that this is, we expect models to get there soon, but they're not there yet, so why not amortize these skills, um, into the model training? Another is that we're not concentrated on making it really good at widely deployed work. We just want it really good at RSI, um, and this is just, like, a way for us to, like, get revenue so that we can pour it back into a model that is actually, like, really good at doing, um, RSI development and then, like, once the singularity happens, the thing that comes out the other end will be really good at all the things which seem like bottlenecks to the current generation of models. Yeah, John, I don't know if you have takes on, like, what, um... H-How, how one should construe why there is so much task-specific knowledge in these models if the, if the path is, like, this kind of generalization.

    6. SP

      Yeah. I mean, if the models were good enough at learning in context, then in theory you wouldn't be, you wouldn't need to train them on finance. Uh-

    7. DP

      Yeah

    8. SP

      ... they would just be able to figure out, uh, read all the, um, all the books on the fly and, uh, figure out how to-

    9. DP

      Yeah

    10. SP

      ... how to do everything in the appropriate jurisdiction. Um, uh, yeah, and you could argue that, um, you need to do, um, a lot of this domain-specific training, um, just to make the, uh, to make them more efficient. Uh, so even if they, they were smart enough to figure this out on the fly, you still might wanna do a bunch of RL and bake the, uh, bake all these intuitions into the weights, uh, so the model would be more efficient at runtime.

    11. DP

      Yeah.

    12. SP

      Yeah, I'd say in practice, it does seem like, um, model providers are, um, going domain by domain and trying to strengthen the models in the highest value domain. And I, I'd say that that's one of the answers to why the models have gotten so much better. It's just because, um, the model providers have covered a lot of the high-value domains and the most common types of skills.

    13. BM

      I mean, I think another thing is just that, like, it's not that expensive to do both at the same time, right? Because, like, the models are massive. They can easily afford, in terms of their parameters, to, like, learn everything.

    14. DP

      Yeah.

    15. BM

      And, like, there is likely some transfer and sort of even just, even if, like, finance is not specific, like, the information is important for, like, RSI. Just the general, like, meta learning of, like, how to figure out what's important, how to have taste, how to, like, do long-horizon work-

    16. DP

      I see. Yeah

    17. BM

      ... is potentially generalizable. And, like, there's not that much RSI, like, data in the world as well.

    18. DP

      Mm.

    19. BM

      Like, it's kind of hard to generate and, like, that requires a lot of effort to, like, if you can sort of amortize in this other data, get some transfer from it, you already have masses of compute and masses of parameter space so, like, why not do that as well as, like, obviously the direct, like, commercial intent of, like, selling a model.

    20. DP

      Yeah, makes sense.

    21. SP

      Oh yeah, I'll add that, um, I mean, there's one question about whether, um, this, uh, current paradigm, uh, of doing, like, sim-to-real, uh, will be the dominant one forever.

    22. DP

      Yeah.

    23. SP

      So basically you, you look at what the real-world tasks, uh, are like, and then you try to create a bunch of environments that can be simulated, uh, like in the data center, and, uh, you can do RL on them. And I think, um, obviously this has been very successful, successful, but it has a lot of weaknesses, uh, because a lot of things are just kind of hard to simulate, um, especially if they involve, like, interact, interacting with a bunch of humans in real time. Yeah, so there's some question about, uh, like simul- like whether sim-to-real will be the dominant framework forever.

    24. BM

      I think sim-to-real has to be the dominant framework while, like, sample efficiency is kind of low 'cause, like-

    25. SP

      Yeah

    26. BM

      ... right now you need, like, you know, thousands and thousands of interactions with the humans, and no human is gonna sit there and, like, deal with this, basically be in the loop of RL training.

    27. SP

      Yeah.

    28. BM

      And so, like, we kind of have to simulate that now to, like-

    29. SP

      Exactly

    30. BM

      ... get the samples you need. But, like, obviously if sample efficiency improves a lot, you'd expect learning from deployment to, like, become, like, a much bigger part of it.

  5. 45:24 – 1:00:33

    The sim-to-real gap

    1. SP

      way.

    2. SP

      But, but is it-- Th-this seems like a bigger issue with the sim-to-real thing, where the longer and longer horizon tasks get, the harder they are to simulate within a data center, right? It seems to me already, potentially, at least even in coding, we're getting to the point where, um, there, there's, like, not some year-long coding task that doesn't eventually require you to, like, talk to a client or interact with the company or interact with users. And if, if you think about the gamut of things we would want AI to be capable at, you want-- Eventually, superintelligence should be able to, like, run a business or, like, start a new business and make it profitable, or, like, have a profitable day trading in the markets, or win a court case. And these are all things which are very hard to simulate in a data center. Like, a, an inherent part of the learning there is interacting with the real world. And so maybe, maybe, yeah, they may have to learn how to get better at these things from, like, the transfer between sim-to-real. But alternatively, maybe you do need weight updates from these kinds of interactions in order to get better at them. And then if that is the case, if transfer isn't strong enough and you need, do need weight updates, then the fact that the models are quite sample inefficient is, like, maybe a deeper problem. And the reason I'm curious about this, I feel like by default, I, I don't see how you don't get some kind of crazy recursive self-improvement within the next ten years. But the one reason why that might not happen is in terms of, like, weight updates, the sample efficiency of weight updates, they just seem way far behind humans, right? Like plausibly millionfold behind humans in terms of how much data a human sees from birth to adulthood versus how much a model sees from, you know, like, cold start to, like, finish, fi-finishing training. And so, yeah, m- this is all to say, first of all, is there, is there gonna be good transfer between simulations and extremely long-horizon, really complicated real shit that we want the AI to do in the real world? And if not, does that really mean that, like, the, the lack of sample efficiency in these models comes to bite us? I think maybe the way I'd break down, like, the two types of tasks in which models get good and models will, like, still continue to struggle is whether the task is, like, cumulative or, like, you kind of have to-- you have this, like, non-stationary distribution you have to keep learning and, like, relitigating a bunch of stuff. So, like, maybe an example of a cumulative task might be RSI. Like, it's theoretically possible to maybe, like, have a less than a million token, like, you know, Python file, which, like, from scratch trains a model that is capable of recursive self-improvement. And, like, every discovery that you make is kind of a line in the sand that you hold. Like, if it's true that, you know, for RSI, we don't need to discover a new attention variant or whatever. Like, once you've discovered attention, and then once you've discovered, you know, mixture of experts, and once you've discovered GRPO, you just add that to the training stack, and, like, that's, that's there. And, like, a good example of this is, like, you know, 5.6 Sol training, 5.6 Terra, whichever one OpenAI told it to train. Like, it didn't have to go back and discover attention. Like, it basically probably would have called a bunch of, like, scripts, which is like pre-training.sh and post-training.sh and just did that. So, like, that's an example of, like, a cumulative task. I think the real world and the reason, like, people are thinking so much about, like, continual learning is it's, it's not really a cumulative task. Like, imagine, like, in a law firm, you have an agent acting as a legal associate. Like, that's a very non-stationary distribution. You have to be able to fit in your context, like, all the relationships between all the important people at that company, which are also changing all the time. Yeah. Um, you have, like, all these, like, implicit, like, ways about how things are done, where to find information, et cetera. And, like, that's not as, as clean of an example of a cumulative task as, like, RSI is. So I think that there will be this breakdown between tasks. But, you know, like, if the labs realize that and they, they do believe that RSI is cumulative in the sense that, like, we don't need to go back and discover some brand-new, like, architecture or whatever- Yeah ... then maybe more and more effort and compute gets focused on that versus the- It's so unfortunate that RSI happened to be easier than- Yeah. [laughs] ... paralegal. [laughs] Yeah. Um, yeah, I don't know if you guys have thoughts on this.

    3. SP

      Yeah, I would say there's, um, like, models, um, are-- today's models are, um, weaker than humans in a lot of different ways, and, uh, some of them might have to do with, um, sample efficiency in a certain regime, uh, where, I mean, in, in some regimes, models are very sample efficient, like learning in context. Uh, but then there might be some, like, medium-length regime where they're less sample efficient because humans can, uh, do, uh, some kind of weight update, um, more efficiently than models. So, so I think, like, being less sample efficient in certain regimes might be one of the sources of weakness, but then I think there are other sources of weaknesses that are completely different than that. For example, uh, having lower diversity of thought than humans or, um, yeah, being bad at certain kinds of long-horizon judgments. Uh, I mean, I think a lot of, um, what people call taste is, uh, is something about, um, behavior that, um, works in the long run and that people have realized, uh, works in the long run. Uh, not everything, but, like, some, some aspect of taste, like, especially for something like software engineering, like, I think a lot of taste is, like, what are the systems that are gonna be maintainable and-

    4. SP

      Yeah

    5. SP

      ... uh, work well, yeah, in the long run of this project. So, uh, yeah, I think the weaknesses of humans, which limit, um, RSI along with other things, uh, are, um, yeah, there's a variety of them, and some of them are related to sample efficiency, and some of them aren't.

    6. SP

      Maybe an interesting thought experiment is, like, if you were able to give a model, like, a context window of, I don't know, a trillion tokens, or whatever you would have needed to fit in, like, your experience prior to, like, let's say, RLHF, and, like, it-it's got all that experience in the context window, and it has the same sample efficiency and in-context learning ability as it does at a million tokens. Like, do you think taste is then solved? Like, would it be able to, like, make the same judgments that you did, or is there, like, something fundamentally missing apart from just a longer context window with the same sample efficiency?

    7. SP

      Yeah. I mean, it would have to be trained to learn from that context. So, uh, I'm not sure. Um, yeah, either it would have to be, um, trained to learn, uh, the right update, um, to make from that context or-

    8. SP

      So you don't think you can just, like, dump it all in, like, your whole, like, life, like, research experience?

    9. BM

      I mean, like, you still need the data to train it along context, right? Like, even if you could theoretically get, like, a trillion context, you would need a trillion lengths of data to train it. Like, right now-

    10. SP

      Yeah

    11. BM

      ... you have, like, take context becomes dump a million.

    12. SP

      I'm just asking if you had that, like, how many models?

    13. BM

      In theory, I think yes. I mean, this really just comes down to the question of, like, how meta-learnable is taste from, like, shorter horizon episodes? And, like, I feel like there's no obvious reason it's super long, 'cause, like, humans somehow developed taste with not having many long episodes. Like, we don't live to be, like, ten thousand. We have, like, you know-

    14. SP

      It's much more, yeah

    15. BM

      ... we, we develop pretty quickly, right? And so, like, you know, if you think about, like, even, like, in a PhD, the difference between, like, a first-year PhD student and, like, a final, like, postdoc or something, that's, like, five years maybe, and they've only done, like, maybe, like, ten, fifteen, thirty research projects in total. But somehow they develop taste quite quickly from, like, a relatively short succession of, like, small things. And so, like, theoretically, it's, you know, possible to develop it like that.

    16. SP

      Mm-hmm.

    17. BM

      The AI obviously will have vastly more experience in which to develop taste, like, meta-learn it, and then it's, like, how well does that generalize to, like, really long-horizon things is, I think, the question, which I think is really unsolved at this point. Like, we don't know.

    18. SP

      Going back to this question, eventually, there should be a regime where AIs are learning a ton from each individual instance of deployment that they have. Where currently, you could say there's some meta fuzzy process by which models do improve for deployment, but I feel like it's a very weak, uh, very w-weak, uh, feedback loop. Do you see this around the horizon, where there's this, like, hive mind kind of learning that's very rapid, and, um, and if, if so, how exactly does it happen?

    19. SP

      Actually, I would say that around, uh, w-will we get a hive mind, uh, that learns from all of its deployment experience, I, I mean, a big part of that is actually about incentives, uh, rather than, um, being a technical question.

    20. SP

      Mm-hmm.

    21. SP

      So, like, companies aren't gonna wanna, um, have the model provider, uh, learn from all of their deployment, because that'll, that might just, uh, reduce the, uh, advantage of their business.

    22. SP

      I think that maybe the economics of this will pressure, not necessarily weight updates to one big, like, common shared model, but, like, kind of like modules that get subbed in. So, like, a very obvious example, this is a LoRA, but it might be something else. Like, you know, there's been a lot of work to try and fit, like, an arbitrary context length into a fixed size. Like, this is all the linear attention stuff-

    23. BM

      Mm-hmm

    24. SP

      ... and, and all that sort of stuff, and, like, cartridges, which are essentially KV caches trained to be very, very compressed KV caches to fit in a lot of information. That's another example of, like, you know, something that, like, companies may be willing to sign up for if that gets subbed into the model, and it's not, like, actually changing the base underlying model itself. So, like, th-there's many different versions of, like, learning from, from your data in real time, and, like, the latter ones are not really helping the big labs, because they are just these modules. But I think the, like the, like the, the economic pressure will force, like, the labs to go down that path first before they can embark on this, like, you know-

    25. SP

      Which-

    26. SP

      Sorry

    27. BM

      So which economic pressure, though? 'Cause I feel like even if you have, like, a bunch of cartridges or LoRAs or whatnot, you can still just, like, take all these traces and just, like, distill this, dump this to the pre-training of, like, your next generation of models.

    28. SP

      Yeah.

    29. SP

      So it, it, it may be a more indirect form of learning that the-

    30. SP

      Yeah

  6. 1:00:33 – 1:18:03

    How much progress is explained by data?

    1. DP

      Okay, let's talk a bit about data now. So I'm generally interested in this question of how much of AI progress is just explained by data progress. Um, doesn't mean it will be necessarily hard to automate, but that's a separate question. So is there some data distribution which if you trained current architectures on, would result in a superintelligence that totally dominates human experts across every single field?

    2. SP

      Are we talking about, like, pre-training plus post-training data-

    3. DP

      Yeah

    4. SP

      ... like, environments as well?

    5. DP

      And-

    6. SP

      Like, I think the existence of this is obvious. It's just, like, whether we can create the right environments to get there.

    7. BM

      Yeah. I mean, in the trivial case, we could just train it to output the Python file which, like, trains the actual-

    8. DP

      Yeah

    9. BM

      ... superintelligence. Like, just have that memorized in the weights.

    10. SP

      Like, yes, there's probably, like, a ladder of RL environments- That is possible to construct such that you would get a, AI researcher which is at least as good as a human researcher. But the effort to climb each successive rung grows like kind of exponentially.

    11. SP

      Mm-hmm.

    12. SP

      Um, and that's going to be the two things that you have to trade off against as to whether like, like how fast we're going to hit like that final rung where it's, where it's better. Um, I, I think that's fairly clear. And I think there's like, you know, we're still relatively early in like RL environment creation. Like, there's a lot of asymmetries that we exploit in order to create good environments. So one of the asymmetries which we've talked about before is like th-the, it's, there's environments where it's easier to go backwards than forwards.

    13. SP

      Yeah.

    14. SP

      And, like, what I mean by that is like, it's very easy to define this like complex data-generating process, and this is like this kind of latent variable you keep hidden from the model. Um, you can generate like arbitrarily complex like environments, and the model has to do a lot of like irreducible like token spend and irreducible work to figure out what that data-generating process was. Um, there's asymmetries in terms of like, you know, you can inject information from the real world. So like Anthropic finds a bug, um, through like, you know, tens of thousands of human and LLMs combined, and like turn that into a very, very neat environment, which in a single LLM can theoretically find within like, you know, a few million tokens. So like, there's all these asymmetries which we're cherry-picking and like we're counting on like kind of this task horizon generalization. Um, but I think, yeah, again, there's just gonna hit diminishing returns at some point. Like at some point, there's diminishing returns in how hard it is to create these environments in the first place, like coming up with them, because you can't necessarily just like have these really, like these processes where it's easier to go backwards than forwards. Like you actually have to sit down and construct like something that looks, like with humans, like, you know, a long enough time horizon, like it's gonna be a really complex task to create. Um, and then there's also gonna be like the compute and time bottlenecks for the agent to actually do those tasks.

    15. SP

      Yeah.

    16. SP

      Um, so like I think you're just gonna start seeing this like curve to flatten out.

    17. SP

      I saw something about how someone fine-tuned, uh, the Taki model, uh, which is-

    18. SP

      Oh, yeah

    19. SP

      ... only trained on data up to 1930 on this, uh, like modern coding agent data. And, um, and it did better than, uh, Claude 3 Opus on, on SWE-bench. Um, so, so this model that has, uh, like no knowledge of code whatsoever, uh, can be fine-tuned on a moderate amount of data and, uh, like behave better as, as a coding agent than this much larger pre-trained model is pretty crazy. And it kind of shows you that, uh, like once you have an example of, uh, like the right expert behavior, it's actually surprisingly easy to like copy that into a, like relatively weak m-model.

    20. SP

      Yeah. Um, but a counterexample to like that kind of is the, there was a paper recently where they trained it up to like fifth grade maths.

    21. SP

      Uh-huh.

    22. SP

      Um, and like also like primary school, like English and stuff, so it was like a, a decent language model, and they tried to L- RL it to do like, you know, late high school and college maths, and the gap was just too large. Like, they couldn't get it to climb at all. Like, but if you did like successive rungs of like, you know, year seven maths and then year eight maths and-

    23. SP

      Yeah

    24. SP

      ... and so on, like you could obviously climb to, to year 12. So like, again, it's just like, what is the distance between the rungs on those ladders and how hard is it to create?

    25. SP

      Mm-hmm.

    26. BM

      Yeah. And this just comes back to like the RL signal problem. Like RL is not very good at like exploring right now. And so-

    27. SP

      [laughs]

    28. BM

      ... if you, if the model can't like get in like, you know, a hundred and twenty-eight rollouts, it's very unlikely to get signal to like progress. And this is why like in RL we need like curricula, whereas like in pre-training we don't, because like it's, that's not a problem for pre-training at all.

    29. SP

      Yeah. And, and again, pre-training data is different to post-training data.

    30. SP

      Right.

  7. 1:18:03 – 1:24:54

    Why is RL working so well?

    1. SP

      a bit on RL. So I feel like a year ago, a lot of people were making this argument that RL will not be super successful at scaling for models. I think, John, you wrote a research paper where you were pointing out that models learn one bit per episode when you RL, they basically learn, "Did I get the answer right or did I get it wrong?" And I wrote some blog posts earlier this year where I was like, it's even worse than that because when the pass rate is low and the model is very unlikely to get the answer right, it's, it learns almost, almost nothing at all from an RL episode. But we-- I, I look at the models today and they seem pretty smart, and it seems to be the result of scaling up RL. Beren, you had a post I think a few weeks ago where you were trying to explain what's going on. But why has RL been more successful than one would have naively thought?

    2. BM

      I mean, so I think the success of RL comes down to a bunch of different things. So first, I think what is slightly underestimated is actually the mid-training.

    3. SP

      Mm-hmm.

    4. BM

      So an awful lot of, like, what we see as successes of RL actually comes from, like, very, very good mid-training data, which is basically where we're, like, essentially doing pre-training but on, like, synthetic reasoning data and, like, the kind of environments that, like, get the model warm-started for RL. And so this actually takes the model, like, almost, like, eighty percent of the way to, like, the final RL checkpoint often. And then what RL does on top of that is it does, like, a lot of, you know, essentially tweaking to the policy. And so this is one of the reasons why it doesn't need, like, as many bits as you would naively think. It doesn't have to learn all of these behaviors from scratch. It needs just, like, a few bits from these episodes, which you do get. And then the other thing that I really point out in my blog is that these bits are actually extremely high signal compared to, like, regular, like, pre-training, which is why you need RL at all versus just, like, SFT-ing on, like, successful reasoning traces.

    5. SP

      Be-because it's exactly the bits about how to get the answer right.

    6. BM

      Well, there's two things. So yes, one, it's exactly the bits about how to get the answer right. But, like, this is not exactly how you think of it because in SFT, you have a trace, right? You have, like, you see a bunch of math reasoning and then the answer at the end. The bit is still there, like, you still SFT on the answer token, so that bit is still there. What's important is that the objective ignores all the other bits. So in SFT, you, like, have, like, you know, to try and match, like, the exact reasoning tokens that the model produces. So you're essentially getting, like, too many bits about, like, the exact way this other model you're training on reasons. For RL, you only get the one bit, and that means that, like, this, it-- that signal is not drowned out in the noise of, like, all the other bits the model has.

    7. SP

      Yeah, yeah.

    8. BM

      And so that's what really, like-- it's really a super dramatic, like, increase into the signal-to-noise ratio during training, which is why, like, RL is, like, so dramatically efficient in terms of steps.

    9. SP

      Hmm. I don't know if you guys have thoughts on that. Yeah, I-- Like, there's, there's been so much debate about, like, what RL does to the model versus, like, you know, mid-training or SOT or whatever. And, like, you know, everyone talks about how, you know, pass at one will go up, but pass at two fifty-six will go down. Like, very rare correct reasoning traces will be, like, down-weighted and kind of, like, outweighed by a gradient signal from, like, easier kind of reasoning traces. And I, I think the simple, like, way to view RL now is that if you have a large enough com-- like, a large enough amount of compute to sample a large enough group size such that your probability of getting a bunch of correct answers is, like, past some, like, not insignificant probability, then, like, it will be up-weighted. And, like, to, to, to Beren's point, like, basically mid-training and, you know, more pre-training, like, the, the pass at one, the starting point for RL, like, scales in the log number of pre-training tokens. You, you can ask some very basic questions. Um, I, I guess, yeah, that answer makes sense, and maybe there's, uh, empirical research which shows that this is what's happening. But, uh, then I just look at the models themselves, and I don't know what's happened. So, like, may-- Yeah, maybe you can give me a sense of what is the basis of the AI progress over the last year. But, but if, if it's-- Yeah, maybe it's just up-weighting the, uh, the policies which were gonna do the correct thinking anyways. But it just seems like qualitatively, the models have gotten so much more capable. And anyways, maybe there's no-nothing to-- there's no inherent contradiction there, but how do we square, square, like, the relatively small impact this take would imply that RL would have from the actual qualitative capabilities the models seem to be gaining?

    10. BM

      So, like, one thing I want to point out here is that, like, it doesn't necessarily imply that RL has a small, like, effect, right? Even if you have a few bits and, like, you only change the parameters a small amount, like, the actual impact on, like, function space, the model learns, like, the input-to-output mapping can still be, like, super dramatic. You know, even if it's, like, even, like, one bit can change, like, your function space a lot, and it can, like, rule out, like, half the hypothesis space-

    11. SP

      Yeah, yeah

    12. BM

      ... which is huge. So, like, I don't think it's necessarily the case it's, like, small amounts of bits, small amounts of RL. Once you're starting from a really good point, means that, like, you don't have dramatic impacts in behavior, at least, like, not necessarily.

    13. SP

      Hmm. I, I think it comes down to two things. I think the first thing is that everyone was hoping that RL would, like, generalize this reasoning across, like, all these different domains, and I don't think we necessarily got this, like, horizontal generalization. Like, just training on maths doesn't necessarily make you the greatest coder. Like, you do have to do RL on, on code environments. I think what we did get, though, is, like, horizon generalization.

    14. BM

      Mm.

    15. SP

      Um, like, the models just learned how to use more tokens for, for longer and still make progress on some sort of task. And so, like, you can train on environments where they get longer and longer and longer and then put them into a completely new environment. And yes, like, they may not have generalized the reasoning patterns which allow them to do well in that environment, but they've at least generalized the ability to, like, continue on that task for longer, which is correlated with, like, success. I think, I think there was a paper called Edge Bench which showed that the rate at which models can work for longer is, like, doubling every three months, and so that's a clear evidence of generalization. And I think, like, the f- the final way to think about it is, like, um, in pre-training, there's this idea of, like, quanta. So you have this very smooth, like, pre-training loss curve, and when you actually look at what's happening in the model, like, the model is learning all these, like, very discrete, like, tasks, and there's, like, all these, like, emergent points where there's, like, kind of a phase transition. Like, it didn't have induction heads, now it has induction heads, and there's, like, tens of thousands, millions, probably, like, hundreds of millions of these things, and you average them all together, and you get this very, like, smooth loss curve. I think, like, to a s- to an extent, like, a similar thing is happening, happening for RL. Like, there is this very slow outer loop, as Beren mentioned, of, you know, we will train a model and then RL it, and then, like, the next kind of model iteration of training, we will dump a bunch of these synthetic reasoning traces into the mid-training data. Like, we're kind of hitting all these quanta for all these different tasks, and, like, on an individual task level, it may look like a phase transition and, like, you're suddenly going from, like, a point five percent pass rate to a ninety percent pass rate on, like, a particular, like, finance task or Excel task or whatever. But you average all these things together and plus the horizon generalization, you kind of would go, "Wow, we've got, like, qualitatively better models."

    16. BM

      Yeah. I mean, I think a lot of this as well is just, like, I think RL does generalize a bit. Like, suddenly you get, like, some transfer between, like, math and code or, like, puzzles and math and this kind of stuff. Also, just, like, the amount-- the sheer amount of environments I think the people are targeting is just, like, vastly greater.

    17. SP

      Yeah, yeah.

    18. BM

      So, like, you know, before, when you tried to do, you know, some task which, like, you do in your daily life, like, two years ago, like, the labs wouldn't really care about this. They wouldn't, like, train the model for it. And now, like, it's just so much broader.

    19. SP

      Right.

    20. BM

      They have a lot of environments targeting this specific

  8. 1:24:54 – 1:28:32

    Move 37 and entropy collapse

    1. BM

      thing.

    2. DP

      Earlier in the conversation, we were talking about RL in the context of causing this entropy collapse or poten-- uh, just, you know, concentrating probability on solutions the base model had already done, um, and causing relatively sparse updates in the policy. Um, but when I think, like, when I think about... I think th-there's also another story about RL, which is going back to the Atari games and then AlphaGo coming up with Move 37, the super creative move that because it was never initialized on human data, it can, like, think in ways that humans are not even thinking and come up with extremely creative solutions. Yeah, d-do you have a sense on when we should expect or if we should expect RL on LLMs to result in things like Move 37, just extreme creativity, even beyond human creativity, because the, like, there's just de novo, uh, de novo initialization of intelligence?

    3. BM

      I mean, so a couple of things here. Like, first off, I think that the AlphaGo is using MCTS, which obviously does, like, more exploration and, like-

    4. DP

      Yeah

    5. BM

      ... stuff than regular policy gradients. But I kind of also think that, like, RL doesn't necessarily, like, reduce the creativity. And, like, I mean, e-even if we l- I think, you know, this is obviously qualitative, but if we look at, like, the, you know, the OpenAI Hugging Face incident, like, these models were coming up with, like, multiple zero-days at a time to, like, break out of their sandbox. And, like, this is clearly, like, some level of, like, Move 37 creativity, I think, already, which we just get from just, like, the general generalization properties of the LLMs. Like, I don't think-- It's definitely not the case that, like, RL is, like, totally destroying-

    6. DP

      Sure, sure

    7. BM

      ... the, like, entropy-

    8. DP

      Yeah

    9. BM

      ... especially on long horizons.

    10. SP

      Yeah. I mean, one thing that people call creativity is just, uh, solving hard search problems. So, uh, and, and the, so that's like, uh, like Move 37 is obviously an example of that or, like, uh, writing some kind of poem that satisfies a, a ton of different constraints. Um, so that's something AI is obviously gonna be extremely good at, uh, if, if trained for it. Um, then, uh, there's another way in which the models, uh, like, uh, th-the diversity of their outputs, uh, is a lot lower after RL, and they sort of develop these tics. And, like, um, even though the models seem like they're good at writing, when you do some kind of, like, distributional analysis, you find that, like, they're reusing, uh, certain themes, uh, like, all the time, and they're using the same character names all the time. So there's actually, um, it's not like you're getting the same kind of diversity that you get when you-- like, from human authors. You're sort of getting one really good-

    11. DP

      Yeah

    12. SP

      ... uh, like, style. So I think that, um, like, that kind of diversity has definitely, uh, been, like, cut down by RL a lot. And in fact, uh, now, uh, yeah, since we were talking about distillation earlier, that's sort of, uh, something... Yeah, one thing that's happening is that so many people are distilling, uh, mostly from Claude that, like, uh, like all the OpenWeight models, right, the same way as Claude and use the same, like, have the same tics. So this, uh, seems kind of concerning to me, that we're having this, uh, like, this monoculture emerge.

    13. DP

      Yeah.

    14. BM

      Again, I don't think this is, like, fundamental to RL as, like, a method, though.

    15. SP

      Mm-hmm.

    16. BM

      And the same with distillation. Like, even with distillation, like, you're just training on the data. It's like, just because your data is not, like, super broad, that doesn't mean, like, the training method itself is somehow wrong. It's, like, a problem with the data. And I think a lot of, for instance, like, the RL, like, entropy collapse is basically due to, like, exploitation of fairly simple, like, verifiers when you don't have, like, a huge diversity of environments.

    17. SP

      Mm-hmm.

    18. BM

      'Cause, like, for instance, like, the writing, I think the writing is presumably graded by some judge, and, like, the judge has some specific tics and, like, the model is learning to award hack the judge, and that's why, like, it collapses. But, like, it, this is really a problem with the judge. It's not a problem with, like, RL in general.

    19. DP

      Hmm. Um, okay. R-r-

  9. 1:28:32 – 1:37:00

    Rapid-fire timelines

    1. DP

      super rapid-fire predictions about the future. So I want timelines on the following couple questions. By when do we have models which you can-- Here's what the, it feels like to a user. You basically hire them as a drop-in remote worker for all kinds of white-collar work, not just coding, but, I don't know, video editing, um, uh, law, paralegal, et cetera. Like, it's, like, literally an actual remote worker with, like, full computer use with, like, literally a m-month of seamless learning and operation and executing on, like, complex, uh, projects and it require interacting with other people, et cetera, et cetera. It's, like, everything a human worker could do over a month.

    2. SP

      Mm. If you, like, mandate it to use, like, a browser or, or whatever, rather than, like, these, the, again, the firm setting up the information to be, like-

    3. DP

      Yeah

    4. SP

      ... programmatically accessible, like, maybe a couple of years. But if it's not, like, browser-based, like, it can send Slack messages, it can do all this stuff, but I'd still probably say around a year.

    5. DP

      Yeah.

    6. BM

      Yeah. I mean, I would say maybe, like, for the, like, full generality, maybe, like, three years. But I think to Charlie's point, we will end up with, like, a lot of people, like, making their organizations easier for the AIs to use, and so you get, like, eighty, ninety percent of the way there before that.

    7. DP

      Sorry, but the thing that's c- the, the, the diff between one year and three years there is just literally, like-

    8. BM

      Like, there's, I think there's gonna be, like-

    9. DP

      ... computer stuff

    10. BM

      ... a long tail of, like, miscellaneous stuff which, like, some human can do, which, like, will take the models, like, quite a while to do.

    11. DP

      Yeah, like, I mean, are you thinking of sort of computer stuff or, like, basic cognitive capabilities?

    12. BM

      I mean, I think this, this really comes down to a question of, like, how quickly can we solve this kind of, like, online learning and, like-

    13. DP

      Yeah

    14. BM

      ... whether we can, like, get, like, eighty, ninety percent of the way there with, like, compaction and, like, writing files yourself and stuff, and, like, that's my big uncertainty.

    15. DP

      Yeah.

    16. BM

      I really don't know.

    17. SP

      And, and, and another, like, maybe an example of something that it wouldn't be good at is, like, you know, if I have to, like, yell at someone to get something at, at work or, like, really push someone to get something done, like, the model isn't just gonna do that. It's just gonna be too nice.

    18. DP

      Yeah.

    19. SP

      Yeah.

    20. SP

      I'd say there's a wide variation in quality of, uh, human remote workers.

    21. BM

      [laughs]

    22. SP

      So if you, if you try to hire someone, like, off of Upwork to do a software engineering project, there's gonna be a huge variation. It's, like, often quite hard to get them to do, like, to do a good job or, uh, like, pay attention to all the feedback you're getting. And, uh, like, I would guess that in some cases it, uh, like, it'll be worse. Like the, the pre-AI, um, version of this, uh, was worse than what you can get now from-

    23. DP

      Yeah, yeah

    24. SP

      ... existing AI. Uh, so- I think it might end up being a little complicated, uh, 'cause maybe to some, some extent we already have this, uh, like for some like not so high quality of work, but then like, uh, then it's obviously like we're not, yeah, we're not matching human level in certain like higher quality like-

    25. JS

      Uh-huh

    26. SP

      ... um, forms of work. So but I basically agree with Charlie and Beren that maybe, yeah, we'll, yeah, we'll have some version of this in a year so that's like, okay, and it will be able to do, um, maybe we'll have that form factor, um, and, uh, it'll be able to do some things really well, some things not so well, and-

    27. JS

      Yeah

    28. SP

      ... things will be improving from there.

    29. JS

      Like, we shift the goalpost based on the very long tail all the time. Like, I think, I feel like you've used this example before of, like, doing your taxes or something.

    30. SP

      Yeah.

Episode duration: 1:37:01

Install uListen for AI-powered chat & search across the full episode β€” Get Full Transcript

Transcript of episode PrSf7IOYu-I

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.