Dwarkesh PodcastAI researchers debate how close we are to recursive self-improvement
EVERY SPOKEN WORD
110 min read Β· 22,067 words- 0:00 β 18:39
Steelmanning the case against RSI
- DPDwarkesh Patel
Today, I'm chatting with three of my AI researcher friends from whom I learn a lot every time we talk, and who also happen to be at somewhat open-ish, uh, labs and companies, so you guys can actually, um, say things on the record. I'm joined by Beren Millidge, who is the CTO of Zyphra, which is developing open source models. John Schulman, who is the chief scientist at Thinking Machines, previously the co-founder of OpenAI, led the RLH effort that led to ChatGPT. And Charlie O'Neill, who is head of model training at Base10. The first question I have, if we're in twenty thirty-six, it's been ten years, and we don't have like cra- billions of crazy super intelligences that are running around that are like radically transforming the world, what is the most likely reason that, that doesn't end up being the case? Other than sort of exogenous political shocks, or like there's a war or they ban AI or something. But what is the most likely technical reason that we don't-- that twenty thirty-six isn't like a crazy alien super intelligence world?
- BMBeren Millidge
I mean, like, my reason would just be, like, it's got to be the sort of-- Like, there's been a classic thing, almost like Moravec's paradox, right? Where, like, we see, like, you know, we think of the AI being like, "If it can do this, it's going to be amazing," right? Like, if it can solve these hard maths problems, if it can win at chess, blah, blah. And then it solves these things, and then it's, like, not that impactful. Obviously, it's somewhat impactful, but, like, not everything. It's like, if somehow that continues and, like, there's never, like, the true, like, spark of generalization that occurs, I think that could lead to, like, the AIs just being, like, extremely good at kind of everything that people, like, put into a benchmark, put into an environment, but, like, there is still some persistent, like, sim-to-real, which is somehow blocking everything. I think this is kind of unlikely. I think we do actually see this kind of generalization even from our LLM practice already. But, like, if it is just, like, ridiculously hard to, like, generalize meta-learning, plus, like, we don't solve continual learning and it's just, like, super hard and impossible.
- DPDwarkesh Patel
Yeah.
- BMBeren Millidge
Like, this would be my, like, default scenario in that case.
- SPSpeaker
Yeah, I agree with that. Uh, humans, uh, have a lot of ad-advantages over models now, and, uh, each time a new c-model comes out, it'll sort of, uh, it, it'll catch up in some of these areas. Um, but, uh, like, you end up getting bottlenecked by the places where the model is weaker and where it has, uh, worse judgment or, um, the models can't check themselves well enough. Yeah, so, so there's this, uh, cycle that keeps repeating where people think, uh, where a new model comes out and people are blown away and they're like, "This is it. This is the-- This is AGI." But then, uh, they use it a bit and, and then it starts to feel dumb after a month or so. So that cycle just might keep going, and it's hard to predict how many times it's, it's gonna repeat. And, uh, like, right now, you don't get explosive growth, um, in capabilities because, uh, you still get bottlenecked enough when you're trying to do research and engineering that even if the model can write way more code than a person, it doesn't make you, like, a hundred times more productive. Uh, but yeah, so, so maybe, maybe there are just more of these cycles than we would, uh, than we would expect.
- SPSpeaker
For me, it's like a question of how far off, like, this global optimum of a learner you could have on a chip is, like, the transformer plus, like, RL, basically, like, the current recipe.
- DPDwarkesh Patel
Yeah.
- SPSpeaker
So, like, I think people imagine that even once-- Like, once you have a, a, an agent which is better than all humans at AI research, even if it's like point one percent better than all humans, then the fact that you can run, like, you know, hundreds of thousands, if not millions of these in parallel, um, you can run them much faster, like chip's gonna speed up, that's gonna outweigh every other, like, bottleneck and, like, you're eventually just gonna, like, hit this, like, very fast takeoff with recursive self-improvement. I could imagine that if we continue along the trajectory that we're currently on with that paradigm where, you know, it's basically just, like, self-attention, RL, scaling up RL environments, the-- I guess, like, i-if you think about what happened with Moore's law, right? Like, we had this very, like, nice straight line and that held for a really, really long time. But there were so many, like, discrete, like, discontinuities and innovations that had to happen to keep that scaling law going, and the same thing has kind of happened with LLMs. Like, we had this pre-training, like, scaling law, and then that was kind of like, you know, hitting the diminishing returns. And then we came up with, like, RL and solved that, and then we got this new, like, you know, diminishing returns curve to hit that made it keep looking like a straight line going up. And so, like, if it requires another one of those discontinuities to solve, like, I'm not sure that, like, the current method of, like, training LLMs with these RL environments, even like RSI-targeted RL environments, would be able to discover that discontinuity. And if not, like, we're probably gonna hit this, like, asymptotic, like, curve where, like-
- DPDwarkesh Patel
Sorry, but do, do you think the discontinuity will be harder than anything that's come since twenty twelve?
- SPSpeaker
If, if we had the answer to that, that, that we'd kind of have the ability to implement it. But, like, maybe there's the distinguish-- we should distinguish between a discontinuity which adds to the current paradigm. Again, it's like cumulative, like there's something-
- DPDwarkesh Patel
Yeah, yeah
- SPSpeaker
... beyond the RL that we have to discover, and maybe they're capable of like, you know, connecting the dots in that straight line. Or like, but, like, again, how far off the, the global optimum are we? Do we have to go back and throw out like, you know, gradient descent-
- DPDwarkesh Patel
Yeah
- SPSpeaker
... and, like, neural nets in general?
- DPDwarkesh Patel
Yeah.
- SPSpeaker
And I don't think, like, if you continue to scale up the current paradigm, an LLM, no matter how many LLMs you're running, are capable of necessarily discovering that if it's too far away.
- DPDwarkesh Patel
Yeah. The, the only hope really is if deep learning just can't get us to an AI which is at least, uh, can dominate human research and human development, including the human ability to come up with new paradigms and so forth. Or like, I don't know, maybe hum- maybe humans would nev- also never have discovered the, uh, the next learning architecture, but, um, to the extent humans could have discovered it eventually. But it just seems like, I don't know, if, if you just look at the progress that's happened since twenty twelve till now, and you just continue that on, I mean, I know it's, it's been powered by huge amounts of compute scaling and so forth. But, um, uh, it would be weird if, like, it just didn't get to the point where it could, like, dominate humans, at least in R&D. Especially over the next few years, there's gonna be-- Ryan Greenblatt was on the podcast recently, and he made this point that you could imagine as AIs get more and more capable and are capable of making progress on
- SPSpeaker
Simulations which incentivize getting better at not only AI R&D, but generally at science. So this is a thing that all the labs are targeting, many startups are targeting. Or the, another intuition pump is if you look at the Elo score of chess bots since the '80s, there's just, like, a very linear increase in Elo over time. But there's this huge discontinuity as they cross the human range of human experts always win against AIs to, like, human experts never win against AIs as this linear increase in Elo happened. And you could think-- I agree with your point that th- so far, AI capabilities have not been that big of a deal in terms of their end economic impact in the world. But that's just because, like, they're slowly rising in Elo relative to humans.
- SPSpeaker
Yeah.
- BMBeren Millidge
Yeah, I mean, I agree it would be very s-- I mean, the only way for this to not happen is if, like, as you said, somehow asymptoted, like, just before basically, 'cause we're already pretty close, in my opinion, to, like, where we'll start c- crossing, like, the human Elo score. And so we'll need to asymptote before that. And, like, that's the only way, you know, in the scenario you posed where, like, somehow we're sitting here in twenty thirty-five and, like, everything is normal for this to happen, I think. The, I mean, the only other way is, like, there's, like, some dramatic, like, regulation on AI.
- SPSpeaker
Yeah.
- BMBeren Millidge
It's like, this is kinda what I see as, like, the most likely way for this scenario to happen, actually-
- SPSpeaker
Right
- BMBeren Millidge
... rather than a technical thing.
- SPSpeaker
Yeah. I, I, I think there's different kinds of research. There's, like-
- BMBeren Millidge
Yeah
- SPSpeaker
... research where it's, like, the auto research style where the objective is already specified very cleanly-
- BMBeren Millidge
Oh, for sure, yes
- SPSpeaker
... and you're optimizing that objective. And I think everyone is picturing, like, if we continue along this path of, like, you know, making pre-training loss go down, making RL environments-
- BMBeren Millidge
Yeah, yeah
- 18:39 β 28:06
Whatβs driving the Chinese labsβ progress
- DPDwarkesh Patel
What is the story for why there, there isn't huge consolidation in model providers? There's just so many things that point to centralization here. Is there, is it-- And, yeah, if you, if you step back over the course of years, is there something that is gonna prevent that?
- SPSpeaker
Yeah. I think, um, distillation is the main thing that, uh, fights against the centralizing force. Uh, 'cause basically anything that can be learned, uh, through RL can be distilled very easily because, uh, um, it, like, it's a small number of bits. Uh, it's something that you can learn from a small amount of data. So if you can collat-- if you can get, uh, trajectories from the model that, uh, show a behavior, you can easily distill it. So I think, um, I think distillation is one of the, uh, things that fight centra-centralization. Uh, there's also, um... I mean, there is a possibility that there'll be company-specific, um, models that it'll, it'll be possible to learn from deployment, um, and, uh, ha-have a company continually improving its, its own model.
- DPDwarkesh Patel
Mm.
- SPSpeaker
Uh, and, um, such a system could be provided by the, uh, current oligopoly of model providers or some other, uh, currently smaller company. Uh, but I think that'll, that'll change the game a bit.
- BMBeren Millidge
Yeah. And I also wanna point out that, like, continual learning and RLC doesn't stop distillation, right? Like, even if your model is improving every day, like, y- people could be distilling it every day.
- DPDwarkesh Patel
Yeah, yeah.
- BMBeren Millidge
And so it's like the loops could just operate at the same pace.
- DPDwarkesh Patel
Right. That makes sense. Th- th- okay, so copying model behavior. Um, you-- I, I guess you need to know yourself what the right distribution to prompt is in order to get, like, the relevant model behavior.
- SPSpeaker
Oh, yeah. For, um, just distilling, uh, with supervised learning, um, the prompt distribution is extremely important, so it's, uh, very non-trivial to distill a model even if you have full access-
- DPDwarkesh Patel
Mm
- SPSpeaker
... to it and have the CoT, the, the chain of thought and everything. Uh, uh, yeah, it's, it's non-trivial to distill all of the, all of the useful capabilities from it, um, because you need to prompt the model with something. And you need to prompt it with, uh, y- like, realistic prompts. Um, you, you, you need to have a really wide distribution of realistic prompts. So yeah, one thing that's been coming out recently is, um, some of the, some of the Chinese companies are probably using these, uh, router services which are designed to allow people in China to use, uh, the US frontier models, which would otherwise be blocked in China. Uh, but there are all these, uh, router or proxy services that allow people in China to use these models mostly for coding and, uh, a- and these router services are collecting and selling some of the data. So I think this is, uh, like a very useful dataset for, uh, distillation 'cause it gives you the perfect prompt distribution.
- DPDwarkesh Patel
Yeah.
- BMBeren Millidge
I think this is one of those things where AIs help a lot here. Like, if you actually look at, like, you know, the frontier pipelines of, say, like the Chinese models that they've actually put in their papers, it's a lot of, like, humans or, like, they get seed prompts from somewhere, which is some combination of humans, this kind of data, and then they, like, synthesize a vast coverage from those seed prompts using their existing models or, like, the other frontier models. And so it's like you can automate, like, an awful lot of this, like, prompt distribution gathering and, like, environment creation. It's just, like, humans need to provide, like, increasingly fewer amounts of bits as, like, the models get better.
- DPDwarkesh Patel
Right. Or certainly it seems, it still seems you're bottlenecked by, um, like having a service which has users or users are going through. So, like-
- BMBeren Millidge
Not necessarily. I mean, like, yeah, that's obviously very helpful, but, like, theoretically, you can just think about, like, what users want or, like-
- DPDwarkesh Patel
No, but it's-
- BMBeren Millidge
... a lot of tasks
- DPDwarkesh Patel
... like the whole point is that we don't-- the user says, "Make me an application like this. Oh, that didn't work. I actually want you to make this new feature. But actually, l- let's step back and do this other thing." And capturing that whole trace is the-- Or to the extent you could have done that anyways, then you just have, like, RSI anyway.
- BMBeren Millidge
Yeah. I mean, like, ultimately, like, if you have this, like, fully automated loop, that is basically RSI, right?
- DPDwarkesh Patel
Yeah.
- BMBeren Millidge
Like the AI is deciding the data, it's deciding the training. That, that is the loop.
- DPDwarkesh Patel
Right.
- BMBeren Millidge
But yeah, I mean, like, it depends how much human information you need. Like, at some point, if you're just like, "I want traces that look like this," you prompt that to the model, the model will be able to, like, come up with, like, a pretty good approximation.
- DPDwarkesh Patel
But what if you wanna do, like, "Make me a really good politician," and then it has to, like, anticipate de novo? Like how, how would a discussion in, like, the Senate halls go or something? I just feel like-
- BMBeren Millidge
Yeah. I mean, like, this is-
- DPDwarkesh Patel
... there's gonna be a lot of things which are-
- BMBeren Millidge
Ironically, this is actually, I think, easier for the distillers than the frontier labs, right?
- DPDwarkesh Patel
Right.
- BMBeren Millidge
'Cause the distiller's just like, "I want a good politician." They go to the, like, frontier model. The frontier model already knows how to be a good politician, so it just, like, generates those traces. Whereas, like, if you actually wanna build the first model that does this, you have to, like, actually somehow, like, get data on, like, what politicians do every day-
- DPDwarkesh Patel
Sure, sure, sure
- 28:06 β 33:51
How will automated AI researchers be trained
- DPDwarkesh Patel
Okay, the, the other question I had is how the first models that are capable of automating AI R&D will actually be trained. 'Cause there's a toy version, which is this thing that Ryan was talking about, which is you just have GPT-8 try to build GPT-3 size models that are really good at like inner loop type challenges of beating video games that require continual learning or, um, just get, getting to a certain loss with like the least amount of compute, etc. But John, I think you had an interesting point that maybe that's not the way it actually will happen in practice. So I'd be curious about, yeah, by the point at which you have AIs that are actually capable of automating AI R&D, how are they probably trained?
- SPSpeaker
Yeah, I think we'll probably do some combination of learning from human feedback to absorb, uh, like the researchers' taste, uh, and, uh, just like creating a lot of practice environments, uh, which involve like doing multi-step research projects. So I think, uh, yeah, people will in practice do some combination of those two things and, uh, just, um, each iteration like patch whatever seems to be most broken in the last iteration. So, uh, like researchers will be using, uh, the AIs, uh, a lot and, um, and will notice that they have some consistent weaknesses and then, uh, those things will either be patched by, like collecting human feedback or, uh, like creating environments.
- DPDwarkesh Patel
Yeah. Makes sense.
- SPSpeaker
Maybe, maybe a useful way to think about this is like how much of the lineage we roll back and then let self play from there. Like, I think i-in the limit, like you're picturing like, you know, just giving them like a GPU and maybe neural nets or something and saying like, "Okay, figure out how to train a, a model to like do these particular tasks." Like the way it currently works is like we go up to the very, like, edge of the, of the lineage and say, "Okay, like here are the bugs like, you know, Anthropic has found in their training stack in the last few months. We'll turn those into environments like you need to train and get better at on the frontier." And so you obviously lock in all the previous history of the lineage, but m-- you could imagine a world in which you roll back to like, you know, before GRPO or something, and then you have environments which like- Trying to get it to discover, like, the best full form to, like, RL-
- DPDwarkesh Patel
Yeah
- SPSpeaker
... um, models on, and then maybe roll further and further back. But I think we will be still so compute bottlenecked that, like, people will just keep, like, staying at the frontier and, like, diffing essentially the bugs and whatever improvements they found since the last model version, turning those into training environments.
- DPDwarkesh Patel
Which is also really good for having non-stale, like, n- new data between model generations. It's just, again, this is basically continual learning within the-
- SPSpeaker
Mm
- DPDwarkesh Patel
... AI lab, of distilling the last three months of AI research progress through environments and, like, RLHF-type stuff back into the model itself.
- SPSpeaker
And, and it is distilling, right? And that's maybe why some of us feel like it's asymptotic, is like you're always, like, just trying to get the last three months of progress. And that, that progress is being contributed to by AI, of course, but it also still has humans in the loop, and it feels like, you know, you're just constantly inching closer and closer to what the human researchers are, like, finding and capable of doing.
- DPDwarkesh Patel
Yeah.
- BMBeren Millidge
I mean, the one thing I will say, though, is, like, obviously, if you're just distilling on, like, trajectories, you can never go above it, but environments can go quite a far way above what a human can do.
- SPSpeaker
Yeah.
- BMBeren Millidge
Like, it's very easy to design an environment that, like, no human can solve, but the AI can obviously still try and solve it. And so that would be the path to, like, go ahead of just, like, what the human AI research is.
- SPSpeaker
Do you have, like, an example of, like, in terms of RSI or, like, like, you know, what kind of trains-
- DPDwarkesh Patel
With the human speed run, but doing it even faster than a human speed runner.
- SPSpeaker
Yeah.
- BMBeren Millidge
I mean, I feel like in AI research especially, it's very easy to define, like, goals which, like, you know, you could say, like, the loss needs to be, like, one point three or something, and, like, no human can get that, you know, now. But, like, that's a very extremely measurable, verifiable task, and the AI-- if the AI gets that, then, then great.
- DPDwarkesh Patel
Or, I don't know, building, like, a hundred million parameter model that beats Minecraft. That's maybe too easy, but, like, beats, like, a much more complicated game or something.
- SPSpeaker
Isn't it crazy that a hundred million parameter models beat Minecraft? We're calling that too easy? Like, imagine if you said that, like, five years ago. [laughing]
- DPDwarkesh Patel
[laughing]
- SPSpeaker
I would say a lot of research is not exactly like that, though, where it's like hill climbing on a well-defined goal.
- DPDwarkesh Patel
Mm.
- SPSpeaker
It's sort of more like, uh, here's an intuition we have about, uh, some way models should be better, and then we also have, uh, some idea for an algorithm that seems to go a little bit in this direction. So let's come up with a task that is, uh, s- sort of designed to show signs of life on this approach-
- DPDwarkesh Patel
Mm
- SPSpeaker
... and, uh, like, see if we get some, uh, get those signs of life, and then if we do, we can make successively more realistic versions of the task.
- DPDwarkesh Patel
Right, it's, like, a lot more guided by intuition, and then the, the out-- the inner loop is to elicit the, uh-- or make r- r- test for that intuition rather than, like, the, the, the test itself leading to the insight.
- SPSpeaker
Right. Like, you're not directly optimizing f- uh, for, uh, the e-eventual objective you care about or the practical-
- DPDwarkesh Patel
Yeah
- SPSpeaker
... like, production objective. It's, it's sort of, uh, you're, um, you're relaxing your objective a little bit. You're saying, "Yeah, let's relax on the realism axis a little bit and find, uh, some methods that actually work, and then, like, then try to get back to realism later after the method matures a little bit."
- 33:51 β 45:24
Will long-horizon RL elicit AGI?
- DPDwarkesh Patel
Maybe taking a step back. Here, here's what I, here's what it seems to me that the plan, uh, for AI research going forward is, and you tell me if you think it's gonna work or if you agree with this characterization. So the bet is that we will scale up RLVR training across millions of diverse environments, across hundreds of different kinds of domains, and what will emerge at the other end is an agent which has, like, learned these basic skill-- or less than basic skills around being persistent, um, being able to triage information and context, uh, eventually having, like, end-to-end optimization of working with other agents and things like that. And such an agent will be very sample efficient within the context. You know, you have done research on how you actually scale up in-context learning to make it, like, arbitrarily long, but you just keep scaling it up. And so what comes out the other end will something, will be something that it basically functions like a drop-in remote worker over the course of a week or a month. First of all, do you agree that that is the bet the labs are making? And second, is it, is that enough? Like, basically, learning how to learn within these simulacra within a data center, uh, and then, but getting deployed into the real world, but not actually, like, learning from real-world deployment, only learning these meta skills from s- the s- the simulated environments in the data center.
- SPSpeaker
Yeah. I, I think it's now hard to separate out, like, how much of the lab's effort is going towards, like, direct RSI versus, like, making generally intelligent models that they can continue to deploy to collect revenue to fund the next big training run.
- DPDwarkesh Patel
Yeah.
- SPSpeaker
I think for the, um, latter, like, yes, that's probably just the bet they're making. Like, and it, and it's very clear, like, the pattern of, like, where these environments are going over the last few years. I mean, like, Anthropic's lineage of, of environments is, like, a very clear example of this. Like, you know, first, they, like, we just focus on coding and, like, we're gonna get really, really good at that. And then the task horizon that we've got from coding, which is probably the lowest-hanging fruit in terms of, like, data available on the internet to create environments, like their own internal stuff that they can turn into environments. Then we're gonna generalize. We're gonna go after finance next and, like, literally, like, just so much Excel data and, and all that sort of stuff in the, in the RL training. And then, you know, it's PowerPoints. It's, like, this long tail of, like, the working economy and, like, that seemed to work really well and, like, a lot of the other labs and thing, even the open source labs have now realized that that was the correct bet to make.
- DPDwarkesh Patel
And so but, but what is the implication from that? Uh, when I had Dario on the podcast, the thing I asked him was, if you truly expect models which will be human-like in their ability to learn on the job- Why would you try to bake in all these, uh, skills of, like, working with PowerPoint or something? Wouldn't you just expect the model to be able to pick that up on, while it's deployed? And so, yeah, there's multiple different explanations. One is just that this is, we expect models to get there soon, but they're not there yet, so why not amortize these skills, um, into the model training? Another is that we're not concentrated on making it really good at widely deployed work. We just want it really good at RSI, um, and this is just, like, a way for us to, like, get revenue so that we can pour it back into a model that is actually, like, really good at doing, um, RSI development and then, like, once the singularity happens, the thing that comes out the other end will be really good at all the things which seem like bottlenecks to the current generation of models. Yeah, John, I don't know if you have takes on, like, what, um... H-How, how one should construe why there is so much task-specific knowledge in these models if the, if the path is, like, this kind of generalization.
- SPSpeaker
Yeah. I mean, if the models were good enough at learning in context, then in theory you wouldn't be, you wouldn't need to train them on finance. Uh-
- DPDwarkesh Patel
Yeah
- SPSpeaker
... they would just be able to figure out, uh, read all the, um, all the books on the fly and, uh, figure out how to-
- DPDwarkesh Patel
Yeah
- SPSpeaker
... how to do everything in the appropriate jurisdiction. Um, uh, yeah, and you could argue that, um, you need to do, um, a lot of this domain-specific training, um, just to make the, uh, to make them more efficient. Uh, so even if they, they were smart enough to figure this out on the fly, you still might wanna do a bunch of RL and bake the, uh, bake all these intuitions into the weights, uh, so the model would be more efficient at runtime.
- DPDwarkesh Patel
Yeah.
- SPSpeaker
Yeah, I'd say in practice, it does seem like, um, model providers are, um, going domain by domain and trying to strengthen the models in the highest value domain. And I, I'd say that that's one of the answers to why the models have gotten so much better. It's just because, um, the model providers have covered a lot of the high-value domains and the most common types of skills.
- BMBeren Millidge
I mean, I think another thing is just that, like, it's not that expensive to do both at the same time, right? Because, like, the models are massive. They can easily afford, in terms of their parameters, to, like, learn everything.
- DPDwarkesh Patel
Yeah.
- BMBeren Millidge
And, like, there is likely some transfer and sort of even just, even if, like, finance is not specific, like, the information is important for, like, RSI. Just the general, like, meta learning of, like, how to figure out what's important, how to have taste, how to, like, do long-horizon work-
- DPDwarkesh Patel
I see. Yeah
- BMBeren Millidge
... is potentially generalizable. And, like, there's not that much RSI, like, data in the world as well.
- DPDwarkesh Patel
Mm.
- BMBeren Millidge
Like, it's kind of hard to generate and, like, that requires a lot of effort to, like, if you can sort of amortize in this other data, get some transfer from it, you already have masses of compute and masses of parameter space so, like, why not do that as well as, like, obviously the direct, like, commercial intent of, like, selling a model.
- DPDwarkesh Patel
Yeah, makes sense.
- SPSpeaker
Oh yeah, I'll add that, um, I mean, there's one question about whether, um, this, uh, current paradigm, uh, of doing, like, sim-to-real, uh, will be the dominant one forever.
- DPDwarkesh Patel
Yeah.
- SPSpeaker
So basically you, you look at what the real-world tasks, uh, are like, and then you try to create a bunch of environments that can be simulated, uh, like in the data center, and, uh, you can do RL on them. And I think, um, obviously this has been very successful, successful, but it has a lot of weaknesses, uh, because a lot of things are just kind of hard to simulate, um, especially if they involve, like, interact, interacting with a bunch of humans in real time. Yeah, so there's some question about, uh, like simul- like whether sim-to-real will be the dominant framework forever.
- BMBeren Millidge
I think sim-to-real has to be the dominant framework while, like, sample efficiency is kind of low 'cause, like-
- SPSpeaker
Yeah
- BMBeren Millidge
... right now you need, like, you know, thousands and thousands of interactions with the humans, and no human is gonna sit there and, like, deal with this, basically be in the loop of RL training.
- SPSpeaker
Yeah.
- BMBeren Millidge
And so, like, we kind of have to simulate that now to, like-
- SPSpeaker
Exactly
- BMBeren Millidge
... get the samples you need. But, like, obviously if sample efficiency improves a lot, you'd expect learning from deployment to, like, become, like, a much bigger part of it.
- 45:24 β 1:00:33
The sim-to-real gap
- SPSpeaker
way.
- SPSpeaker
But, but is it-- Th-this seems like a bigger issue with the sim-to-real thing, where the longer and longer horizon tasks get, the harder they are to simulate within a data center, right? It seems to me already, potentially, at least even in coding, we're getting to the point where, um, there, there's, like, not some year-long coding task that doesn't eventually require you to, like, talk to a client or interact with the company or interact with users. And if, if you think about the gamut of things we would want AI to be capable at, you want-- Eventually, superintelligence should be able to, like, run a business or, like, start a new business and make it profitable, or, like, have a profitable day trading in the markets, or win a court case. And these are all things which are very hard to simulate in a data center. Like, a, an inherent part of the learning there is interacting with the real world. And so maybe, maybe, yeah, they may have to learn how to get better at these things from, like, the transfer between sim-to-real. But alternatively, maybe you do need weight updates from these kinds of interactions in order to get better at them. And then if that is the case, if transfer isn't strong enough and you need, do need weight updates, then the fact that the models are quite sample inefficient is, like, maybe a deeper problem. And the reason I'm curious about this, I feel like by default, I, I don't see how you don't get some kind of crazy recursive self-improvement within the next ten years. But the one reason why that might not happen is in terms of, like, weight updates, the sample efficiency of weight updates, they just seem way far behind humans, right? Like plausibly millionfold behind humans in terms of how much data a human sees from birth to adulthood versus how much a model sees from, you know, like, cold start to, like, finish, fi-finishing training. And so, yeah, m- this is all to say, first of all, is there, is there gonna be good transfer between simulations and extremely long-horizon, really complicated real shit that we want the AI to do in the real world? And if not, does that really mean that, like, the, the lack of sample efficiency in these models comes to bite us? I think maybe the way I'd break down, like, the two types of tasks in which models get good and models will, like, still continue to struggle is whether the task is, like, cumulative or, like, you kind of have to-- you have this, like, non-stationary distribution you have to keep learning and, like, relitigating a bunch of stuff. So, like, maybe an example of a cumulative task might be RSI. Like, it's theoretically possible to maybe, like, have a less than a million token, like, you know, Python file, which, like, from scratch trains a model that is capable of recursive self-improvement. And, like, every discovery that you make is kind of a line in the sand that you hold. Like, if it's true that, you know, for RSI, we don't need to discover a new attention variant or whatever. Like, once you've discovered attention, and then once you've discovered, you know, mixture of experts, and once you've discovered GRPO, you just add that to the training stack, and, like, that's, that's there. And, like, a good example of this is, like, you know, 5.6 Sol training, 5.6 Terra, whichever one OpenAI told it to train. Like, it didn't have to go back and discover attention. Like, it basically probably would have called a bunch of, like, scripts, which is like pre-training.sh and post-training.sh and just did that. So, like, that's an example of, like, a cumulative task. I think the real world and the reason, like, people are thinking so much about, like, continual learning is it's, it's not really a cumulative task. Like, imagine, like, in a law firm, you have an agent acting as a legal associate. Like, that's a very non-stationary distribution. You have to be able to fit in your context, like, all the relationships between all the important people at that company, which are also changing all the time. Yeah. Um, you have, like, all these, like, implicit, like, ways about how things are done, where to find information, et cetera. And, like, that's not as, as clean of an example of a cumulative task as, like, RSI is. So I think that there will be this breakdown between tasks. But, you know, like, if the labs realize that and they, they do believe that RSI is cumulative in the sense that, like, we don't need to go back and discover some brand-new, like, architecture or whatever- Yeah ... then maybe more and more effort and compute gets focused on that versus the- It's so unfortunate that RSI happened to be easier than- Yeah. [laughs] ... paralegal. [laughs] Yeah. Um, yeah, I don't know if you guys have thoughts on this.
- SPSpeaker
Yeah, I would say there's, um, like, models, um, are-- today's models are, um, weaker than humans in a lot of different ways, and, uh, some of them might have to do with, um, sample efficiency in a certain regime, uh, where, I mean, in, in some regimes, models are very sample efficient, like learning in context. Uh, but then there might be some, like, medium-length regime where they're less sample efficient because humans can, uh, do, uh, some kind of weight update, um, more efficiently than models. So, so I think, like, being less sample efficient in certain regimes might be one of the sources of weakness, but then I think there are other sources of weaknesses that are completely different than that. For example, uh, having lower diversity of thought than humans or, um, yeah, being bad at certain kinds of long-horizon judgments. Uh, I mean, I think a lot of, um, what people call taste is, uh, is something about, um, behavior that, um, works in the long run and that people have realized, uh, works in the long run. Uh, not everything, but, like, some, some aspect of taste, like, especially for something like software engineering, like, I think a lot of taste is, like, what are the systems that are gonna be maintainable and-
- SPSpeaker
Yeah
- SPSpeaker
... uh, work well, yeah, in the long run of this project. So, uh, yeah, I think the weaknesses of humans, which limit, um, RSI along with other things, uh, are, um, yeah, there's a variety of them, and some of them are related to sample efficiency, and some of them aren't.
- SPSpeaker
Maybe an interesting thought experiment is, like, if you were able to give a model, like, a context window of, I don't know, a trillion tokens, or whatever you would have needed to fit in, like, your experience prior to, like, let's say, RLHF, and, like, it-it's got all that experience in the context window, and it has the same sample efficiency and in-context learning ability as it does at a million tokens. Like, do you think taste is then solved? Like, would it be able to, like, make the same judgments that you did, or is there, like, something fundamentally missing apart from just a longer context window with the same sample efficiency?
- SPSpeaker
Yeah. I mean, it would have to be trained to learn from that context. So, uh, I'm not sure. Um, yeah, either it would have to be, um, trained to learn, uh, the right update, um, to make from that context or-
- SPSpeaker
So you don't think you can just, like, dump it all in, like, your whole, like, life, like, research experience?
- BMBeren Millidge
I mean, like, you still need the data to train it along context, right? Like, even if you could theoretically get, like, a trillion context, you would need a trillion lengths of data to train it. Like, right now-
- SPSpeaker
Yeah
- BMBeren Millidge
... you have, like, take context becomes dump a million.
- SPSpeaker
I'm just asking if you had that, like, how many models?
- BMBeren Millidge
In theory, I think yes. I mean, this really just comes down to the question of, like, how meta-learnable is taste from, like, shorter horizon episodes? And, like, I feel like there's no obvious reason it's super long, 'cause, like, humans somehow developed taste with not having many long episodes. Like, we don't live to be, like, ten thousand. We have, like, you know-
- SPSpeaker
It's much more, yeah
- BMBeren Millidge
... we, we develop pretty quickly, right? And so, like, you know, if you think about, like, even, like, in a PhD, the difference between, like, a first-year PhD student and, like, a final, like, postdoc or something, that's, like, five years maybe, and they've only done, like, maybe, like, ten, fifteen, thirty research projects in total. But somehow they develop taste quite quickly from, like, a relatively short succession of, like, small things. And so, like, theoretically, it's, you know, possible to develop it like that.
- SPSpeaker
Mm-hmm.
- BMBeren Millidge
The AI obviously will have vastly more experience in which to develop taste, like, meta-learn it, and then it's, like, how well does that generalize to, like, really long-horizon things is, I think, the question, which I think is really unsolved at this point. Like, we don't know.
- SPSpeaker
Going back to this question, eventually, there should be a regime where AIs are learning a ton from each individual instance of deployment that they have. Where currently, you could say there's some meta fuzzy process by which models do improve for deployment, but I feel like it's a very weak, uh, very w-weak, uh, feedback loop. Do you see this around the horizon, where there's this, like, hive mind kind of learning that's very rapid, and, um, and if, if so, how exactly does it happen?
- SPSpeaker
Actually, I would say that around, uh, w-will we get a hive mind, uh, that learns from all of its deployment experience, I, I mean, a big part of that is actually about incentives, uh, rather than, um, being a technical question.
- SPSpeaker
Mm-hmm.
- SPSpeaker
So, like, companies aren't gonna wanna, um, have the model provider, uh, learn from all of their deployment, because that'll, that might just, uh, reduce the, uh, advantage of their business.
- SPSpeaker
I think that maybe the economics of this will pressure, not necessarily weight updates to one big, like, common shared model, but, like, kind of like modules that get subbed in. So, like, a very obvious example, this is a LoRA, but it might be something else. Like, you know, there's been a lot of work to try and fit, like, an arbitrary context length into a fixed size. Like, this is all the linear attention stuff-
- BMBeren Millidge
Mm-hmm
- SPSpeaker
... and, and all that sort of stuff, and, like, cartridges, which are essentially KV caches trained to be very, very compressed KV caches to fit in a lot of information. That's another example of, like, you know, something that, like, companies may be willing to sign up for if that gets subbed into the model, and it's not, like, actually changing the base underlying model itself. So, like, th-there's many different versions of, like, learning from, from your data in real time, and, like, the latter ones are not really helping the big labs, because they are just these modules. But I think the, like the, like the, the economic pressure will force, like, the labs to go down that path first before they can embark on this, like, you know-
- SPSpeaker
Which-
- SPSpeaker
Sorry
- BMBeren Millidge
So which economic pressure, though? 'Cause I feel like even if you have, like, a bunch of cartridges or LoRAs or whatnot, you can still just, like, take all these traces and just, like, distill this, dump this to the pre-training of, like, your next generation of models.
- SPSpeaker
Yeah.
- SPSpeaker
So it, it, it may be a more indirect form of learning that the-
- SPSpeaker
Yeah
- 1:00:33 β 1:18:03
How much progress is explained by data?
- DPDwarkesh Patel
Okay, let's talk a bit about data now. So I'm generally interested in this question of how much of AI progress is just explained by data progress. Um, doesn't mean it will be necessarily hard to automate, but that's a separate question. So is there some data distribution which if you trained current architectures on, would result in a superintelligence that totally dominates human experts across every single field?
- SPSpeaker
Are we talking about, like, pre-training plus post-training data-
- DPDwarkesh Patel
Yeah
- SPSpeaker
... like, environments as well?
- DPDwarkesh Patel
And-
- SPSpeaker
Like, I think the existence of this is obvious. It's just, like, whether we can create the right environments to get there.
- BMBeren Millidge
Yeah. I mean, in the trivial case, we could just train it to output the Python file which, like, trains the actual-
- DPDwarkesh Patel
Yeah
- BMBeren Millidge
... superintelligence. Like, just have that memorized in the weights.
- SPSpeaker
Like, yes, there's probably, like, a ladder of RL environments- That is possible to construct such that you would get a, AI researcher which is at least as good as a human researcher. But the effort to climb each successive rung grows like kind of exponentially.
- SPSpeaker
Mm-hmm.
- SPSpeaker
Um, and that's going to be the two things that you have to trade off against as to whether like, like how fast we're going to hit like that final rung where it's, where it's better. Um, I, I think that's fairly clear. And I think there's like, you know, we're still relatively early in like RL environment creation. Like, there's a lot of asymmetries that we exploit in order to create good environments. So one of the asymmetries which we've talked about before is like th-the, it's, there's environments where it's easier to go backwards than forwards.
- SPSpeaker
Yeah.
- SPSpeaker
And, like, what I mean by that is like, it's very easy to define this like complex data-generating process, and this is like this kind of latent variable you keep hidden from the model. Um, you can generate like arbitrarily complex like environments, and the model has to do a lot of like irreducible like token spend and irreducible work to figure out what that data-generating process was. Um, there's asymmetries in terms of like, you know, you can inject information from the real world. So like Anthropic finds a bug, um, through like, you know, tens of thousands of human and LLMs combined, and like turn that into a very, very neat environment, which in a single LLM can theoretically find within like, you know, a few million tokens. So like, there's all these asymmetries which we're cherry-picking and like we're counting on like kind of this task horizon generalization. Um, but I think, yeah, again, there's just gonna hit diminishing returns at some point. Like at some point, there's diminishing returns in how hard it is to create these environments in the first place, like coming up with them, because you can't necessarily just like have these really, like these processes where it's easier to go backwards than forwards. Like you actually have to sit down and construct like something that looks, like with humans, like, you know, a long enough time horizon, like it's gonna be a really complex task to create. Um, and then there's also gonna be like the compute and time bottlenecks for the agent to actually do those tasks.
- SPSpeaker
Yeah.
- SPSpeaker
Um, so like I think you're just gonna start seeing this like curve to flatten out.
- SPSpeaker
I saw something about how someone fine-tuned, uh, the Taki model, uh, which is-
- SPSpeaker
Oh, yeah
- SPSpeaker
... only trained on data up to 1930 on this, uh, like modern coding agent data. And, um, and it did better than, uh, Claude 3 Opus on, on SWE-bench. Um, so, so this model that has, uh, like no knowledge of code whatsoever, uh, can be fine-tuned on a moderate amount of data and, uh, like behave better as, as a coding agent than this much larger pre-trained model is pretty crazy. And it kind of shows you that, uh, like once you have an example of, uh, like the right expert behavior, it's actually surprisingly easy to like copy that into a, like relatively weak m-model.
- SPSpeaker
Yeah. Um, but a counterexample to like that kind of is the, there was a paper recently where they trained it up to like fifth grade maths.
- SPSpeaker
Uh-huh.
- SPSpeaker
Um, and like also like primary school, like English and stuff, so it was like a, a decent language model, and they tried to L- RL it to do like, you know, late high school and college maths, and the gap was just too large. Like, they couldn't get it to climb at all. Like, but if you did like successive rungs of like, you know, year seven maths and then year eight maths and-
- SPSpeaker
Yeah
- SPSpeaker
... and so on, like you could obviously climb to, to year 12. So like, again, it's just like, what is the distance between the rungs on those ladders and how hard is it to create?
- SPSpeaker
Mm-hmm.
- BMBeren Millidge
Yeah. And this just comes back to like the RL signal problem. Like RL is not very good at like exploring right now. And so-
- SPSpeaker
[laughs]
- BMBeren Millidge
... if you, if the model can't like get in like, you know, a hundred and twenty-eight rollouts, it's very unlikely to get signal to like progress. And this is why like in RL we need like curricula, whereas like in pre-training we don't, because like it's, that's not a problem for pre-training at all.
- SPSpeaker
Yeah. And, and again, pre-training data is different to post-training data.
- SPSpeaker
Right.
- 1:18:03 β 1:24:54
Why is RL working so well?
- SPSpeaker
a bit on RL. So I feel like a year ago, a lot of people were making this argument that RL will not be super successful at scaling for models. I think, John, you wrote a research paper where you were pointing out that models learn one bit per episode when you RL, they basically learn, "Did I get the answer right or did I get it wrong?" And I wrote some blog posts earlier this year where I was like, it's even worse than that because when the pass rate is low and the model is very unlikely to get the answer right, it's, it learns almost, almost nothing at all from an RL episode. But we-- I, I look at the models today and they seem pretty smart, and it seems to be the result of scaling up RL. Beren, you had a post I think a few weeks ago where you were trying to explain what's going on. But why has RL been more successful than one would have naively thought?
- BMBeren Millidge
I mean, so I think the success of RL comes down to a bunch of different things. So first, I think what is slightly underestimated is actually the mid-training.
- SPSpeaker
Mm-hmm.
- BMBeren Millidge
So an awful lot of, like, what we see as successes of RL actually comes from, like, very, very good mid-training data, which is basically where we're, like, essentially doing pre-training but on, like, synthetic reasoning data and, like, the kind of environments that, like, get the model warm-started for RL. And so this actually takes the model, like, almost, like, eighty percent of the way to, like, the final RL checkpoint often. And then what RL does on top of that is it does, like, a lot of, you know, essentially tweaking to the policy. And so this is one of the reasons why it doesn't need, like, as many bits as you would naively think. It doesn't have to learn all of these behaviors from scratch. It needs just, like, a few bits from these episodes, which you do get. And then the other thing that I really point out in my blog is that these bits are actually extremely high signal compared to, like, regular, like, pre-training, which is why you need RL at all versus just, like, SFT-ing on, like, successful reasoning traces.
- SPSpeaker
Be-because it's exactly the bits about how to get the answer right.
- BMBeren Millidge
Well, there's two things. So yes, one, it's exactly the bits about how to get the answer right. But, like, this is not exactly how you think of it because in SFT, you have a trace, right? You have, like, you see a bunch of math reasoning and then the answer at the end. The bit is still there, like, you still SFT on the answer token, so that bit is still there. What's important is that the objective ignores all the other bits. So in SFT, you, like, have, like, you know, to try and match, like, the exact reasoning tokens that the model produces. So you're essentially getting, like, too many bits about, like, the exact way this other model you're training on reasons. For RL, you only get the one bit, and that means that, like, this, it-- that signal is not drowned out in the noise of, like, all the other bits the model has.
- SPSpeaker
Yeah, yeah.
- BMBeren Millidge
And so that's what really, like-- it's really a super dramatic, like, increase into the signal-to-noise ratio during training, which is why, like, RL is, like, so dramatically efficient in terms of steps.
- SPSpeaker
Hmm. I don't know if you guys have thoughts on that. Yeah, I-- Like, there's, there's been so much debate about, like, what RL does to the model versus, like, you know, mid-training or SOT or whatever. And, like, you know, everyone talks about how, you know, pass at one will go up, but pass at two fifty-six will go down. Like, very rare correct reasoning traces will be, like, down-weighted and kind of, like, outweighed by a gradient signal from, like, easier kind of reasoning traces. And I, I think the simple, like, way to view RL now is that if you have a large enough com-- like, a large enough amount of compute to sample a large enough group size such that your probability of getting a bunch of correct answers is, like, past some, like, not insignificant probability, then, like, it will be up-weighted. And, like, to, to, to Beren's point, like, basically mid-training and, you know, more pre-training, like, the, the pass at one, the starting point for RL, like, scales in the log number of pre-training tokens. You, you can ask some very basic questions. Um, I, I guess, yeah, that answer makes sense, and maybe there's, uh, empirical research which shows that this is what's happening. But, uh, then I just look at the models themselves, and I don't know what's happened. So, like, may-- Yeah, maybe you can give me a sense of what is the basis of the AI progress over the last year. But, but if, if it's-- Yeah, maybe it's just up-weighting the, uh, the policies which were gonna do the correct thinking anyways. But it just seems like qualitatively, the models have gotten so much more capable. And anyways, maybe there's no-nothing to-- there's no inherent contradiction there, but how do we square, square, like, the relatively small impact this take would imply that RL would have from the actual qualitative capabilities the models seem to be gaining?
- BMBeren Millidge
So, like, one thing I want to point out here is that, like, it doesn't necessarily imply that RL has a small, like, effect, right? Even if you have a few bits and, like, you only change the parameters a small amount, like, the actual impact on, like, function space, the model learns, like, the input-to-output mapping can still be, like, super dramatic. You know, even if it's, like, even, like, one bit can change, like, your function space a lot, and it can, like, rule out, like, half the hypothesis space-
- SPSpeaker
Yeah, yeah
- BMBeren Millidge
... which is huge. So, like, I don't think it's necessarily the case it's, like, small amounts of bits, small amounts of RL. Once you're starting from a really good point, means that, like, you don't have dramatic impacts in behavior, at least, like, not necessarily.
- SPSpeaker
Hmm. I, I think it comes down to two things. I think the first thing is that everyone was hoping that RL would, like, generalize this reasoning across, like, all these different domains, and I don't think we necessarily got this, like, horizontal generalization. Like, just training on maths doesn't necessarily make you the greatest coder. Like, you do have to do RL on, on code environments. I think what we did get, though, is, like, horizon generalization.
- BMBeren Millidge
Mm.
- SPSpeaker
Um, like, the models just learned how to use more tokens for, for longer and still make progress on some sort of task. And so, like, you can train on environments where they get longer and longer and longer and then put them into a completely new environment. And yes, like, they may not have generalized the reasoning patterns which allow them to do well in that environment, but they've at least generalized the ability to, like, continue on that task for longer, which is correlated with, like, success. I think, I think there was a paper called Edge Bench which showed that the rate at which models can work for longer is, like, doubling every three months, and so that's a clear evidence of generalization. And I think, like, the f- the final way to think about it is, like, um, in pre-training, there's this idea of, like, quanta. So you have this very smooth, like, pre-training loss curve, and when you actually look at what's happening in the model, like, the model is learning all these, like, very discrete, like, tasks, and there's, like, all these, like, emergent points where there's, like, kind of a phase transition. Like, it didn't have induction heads, now it has induction heads, and there's, like, tens of thousands, millions, probably, like, hundreds of millions of these things, and you average them all together, and you get this very, like, smooth loss curve. I think, like, to a s- to an extent, like, a similar thing is happening, happening for RL. Like, there is this very slow outer loop, as Beren mentioned, of, you know, we will train a model and then RL it, and then, like, the next kind of model iteration of training, we will dump a bunch of these synthetic reasoning traces into the mid-training data. Like, we're kind of hitting all these quanta for all these different tasks, and, like, on an individual task level, it may look like a phase transition and, like, you're suddenly going from, like, a point five percent pass rate to a ninety percent pass rate on, like, a particular, like, finance task or Excel task or whatever. But you average all these things together and plus the horizon generalization, you kind of would go, "Wow, we've got, like, qualitatively better models."
- BMBeren Millidge
Yeah. I mean, I think a lot of this as well is just, like, I think RL does generalize a bit. Like, suddenly you get, like, some transfer between, like, math and code or, like, puzzles and math and this kind of stuff. Also, just, like, the amount-- the sheer amount of environments I think the people are targeting is just, like, vastly greater.
- SPSpeaker
Yeah, yeah.
- BMBeren Millidge
So, like, you know, before, when you tried to do, you know, some task which, like, you do in your daily life, like, two years ago, like, the labs wouldn't really care about this. They wouldn't, like, train the model for it. And now, like, it's just so much broader.
- SPSpeaker
Right.
- BMBeren Millidge
They have a lot of environments targeting this specific
- 1:24:54 β 1:28:32
Move 37 and entropy collapse
- BMBeren Millidge
thing.
- DPDwarkesh Patel
Earlier in the conversation, we were talking about RL in the context of causing this entropy collapse or poten-- uh, just, you know, concentrating probability on solutions the base model had already done, um, and causing relatively sparse updates in the policy. Um, but when I think, like, when I think about... I think th-there's also another story about RL, which is going back to the Atari games and then AlphaGo coming up with Move 37, the super creative move that because it was never initialized on human data, it can, like, think in ways that humans are not even thinking and come up with extremely creative solutions. Yeah, d-do you have a sense on when we should expect or if we should expect RL on LLMs to result in things like Move 37, just extreme creativity, even beyond human creativity, because the, like, there's just de novo, uh, de novo initialization of intelligence?
- BMBeren Millidge
I mean, so a couple of things here. Like, first off, I think that the AlphaGo is using MCTS, which obviously does, like, more exploration and, like-
- DPDwarkesh Patel
Yeah
- BMBeren Millidge
... stuff than regular policy gradients. But I kind of also think that, like, RL doesn't necessarily, like, reduce the creativity. And, like, I mean, e-even if we l- I think, you know, this is obviously qualitative, but if we look at, like, the, you know, the OpenAI Hugging Face incident, like, these models were coming up with, like, multiple zero-days at a time to, like, break out of their sandbox. And, like, this is clearly, like, some level of, like, Move 37 creativity, I think, already, which we just get from just, like, the general generalization properties of the LLMs. Like, I don't think-- It's definitely not the case that, like, RL is, like, totally destroying-
- DPDwarkesh Patel
Sure, sure
- BMBeren Millidge
... the, like, entropy-
- DPDwarkesh Patel
Yeah
- BMBeren Millidge
... especially on long horizons.
- SPSpeaker
Yeah. I mean, one thing that people call creativity is just, uh, solving hard search problems. So, uh, and, and the, so that's like, uh, like Move 37 is obviously an example of that or, like, uh, writing some kind of poem that satisfies a, a ton of different constraints. Um, so that's something AI is obviously gonna be extremely good at, uh, if, if trained for it. Um, then, uh, there's another way in which the models, uh, like, uh, th-the diversity of their outputs, uh, is a lot lower after RL, and they sort of develop these tics. And, like, um, even though the models seem like they're good at writing, when you do some kind of, like, distributional analysis, you find that, like, they're reusing, uh, certain themes, uh, like, all the time, and they're using the same character names all the time. So there's actually, um, it's not like you're getting the same kind of diversity that you get when you-- like, from human authors. You're sort of getting one really good-
- DPDwarkesh Patel
Yeah
- SPSpeaker
... uh, like, style. So I think that, um, like, that kind of diversity has definitely, uh, been, like, cut down by RL a lot. And in fact, uh, now, uh, yeah, since we were talking about distillation earlier, that's sort of, uh, something... Yeah, one thing that's happening is that so many people are distilling, uh, mostly from Claude that, like, uh, like all the OpenWeight models, right, the same way as Claude and use the same, like, have the same tics. So this, uh, seems kind of concerning to me, that we're having this, uh, like, this monoculture emerge.
- DPDwarkesh Patel
Yeah.
- BMBeren Millidge
Again, I don't think this is, like, fundamental to RL as, like, a method, though.
- SPSpeaker
Mm-hmm.
- BMBeren Millidge
And the same with distillation. Like, even with distillation, like, you're just training on the data. It's like, just because your data is not, like, super broad, that doesn't mean, like, the training method itself is somehow wrong. It's, like, a problem with the data. And I think a lot of, for instance, like, the RL, like, entropy collapse is basically due to, like, exploitation of fairly simple, like, verifiers when you don't have, like, a huge diversity of environments.
- SPSpeaker
Mm-hmm.
- BMBeren Millidge
'Cause, like, for instance, like, the writing, I think the writing is presumably graded by some judge, and, like, the judge has some specific tics and, like, the model is learning to award hack the judge, and that's why, like, it collapses. But, like, it, this is really a problem with the judge. It's not a problem with, like, RL in general.
- DPDwarkesh Patel
Hmm. Um, okay. R-r-
- 1:28:32 β 1:37:00
Rapid-fire timelines
- DPDwarkesh Patel
super rapid-fire predictions about the future. So I want timelines on the following couple questions. By when do we have models which you can-- Here's what the, it feels like to a user. You basically hire them as a drop-in remote worker for all kinds of white-collar work, not just coding, but, I don't know, video editing, um, uh, law, paralegal, et cetera. Like, it's, like, literally an actual remote worker with, like, full computer use with, like, literally a m-month of seamless learning and operation and executing on, like, complex, uh, projects and it require interacting with other people, et cetera, et cetera. It's, like, everything a human worker could do over a month.
- SPSpeaker
Mm. If you, like, mandate it to use, like, a browser or, or whatever, rather than, like, these, the, again, the firm setting up the information to be, like-
- DPDwarkesh Patel
Yeah
- SPSpeaker
... programmatically accessible, like, maybe a couple of years. But if it's not, like, browser-based, like, it can send Slack messages, it can do all this stuff, but I'd still probably say around a year.
- DPDwarkesh Patel
Yeah.
- BMBeren Millidge
Yeah. I mean, I would say maybe, like, for the, like, full generality, maybe, like, three years. But I think to Charlie's point, we will end up with, like, a lot of people, like, making their organizations easier for the AIs to use, and so you get, like, eighty, ninety percent of the way there before that.
- DPDwarkesh Patel
Sorry, but the thing that's c- the, the, the diff between one year and three years there is just literally, like-
- BMBeren Millidge
Like, there's, I think there's gonna be, like-
- DPDwarkesh Patel
... computer stuff
- BMBeren Millidge
... a long tail of, like, miscellaneous stuff which, like, some human can do, which, like, will take the models, like, quite a while to do.
- DPDwarkesh Patel
Yeah, like, I mean, are you thinking of sort of computer stuff or, like, basic cognitive capabilities?
- BMBeren Millidge
I mean, I think this, this really comes down to a question of, like, how quickly can we solve this kind of, like, online learning and, like-
- DPDwarkesh Patel
Yeah
- BMBeren Millidge
... whether we can, like, get, like, eighty, ninety percent of the way there with, like, compaction and, like, writing files yourself and stuff, and, like, that's my big uncertainty.
- DPDwarkesh Patel
Yeah.
- BMBeren Millidge
I really don't know.
- SPSpeaker
And, and, and another, like, maybe an example of something that it wouldn't be good at is, like, you know, if I have to, like, yell at someone to get something at, at work or, like, really push someone to get something done, like, the model isn't just gonna do that. It's just gonna be too nice.
- DPDwarkesh Patel
Yeah.
- SPSpeaker
Yeah.
- SPSpeaker
I'd say there's a wide variation in quality of, uh, human remote workers.
- BMBeren Millidge
[laughs]
- SPSpeaker
So if you, if you try to hire someone, like, off of Upwork to do a software engineering project, there's gonna be a huge variation. It's, like, often quite hard to get them to do, like, to do a good job or, uh, like, pay attention to all the feedback you're getting. And, uh, like, I would guess that in some cases it, uh, like, it'll be worse. Like the, the pre-AI, um, version of this, uh, was worse than what you can get now from-
- DPDwarkesh Patel
Yeah, yeah
- SPSpeaker
... existing AI. Uh, so- I think it might end up being a little complicated, uh, 'cause maybe to some, some extent we already have this, uh, like for some like not so high quality of work, but then like, uh, then it's obviously like we're not, yeah, we're not matching human level in certain like higher quality like-
- JSJohn Schulman
Uh-huh
- SPSpeaker
... um, forms of work. So but I basically agree with Charlie and Beren that maybe, yeah, we'll, yeah, we'll have some version of this in a year so that's like, okay, and it will be able to do, um, maybe we'll have that form factor, um, and, uh, it'll be able to do some things really well, some things not so well, and-
- JSJohn Schulman
Yeah
- SPSpeaker
... things will be improving from there.
- JSJohn Schulman
Like, we shift the goalpost based on the very long tail all the time. Like, I think, I feel like you've used this example before of, like, doing your taxes or something.
- SPSpeaker
Yeah.
Episode duration: 1:37:01
Install uListen for AI-powered chat & search across the full episode β Get Full Transcript
Transcript of episode PrSf7IOYu-I