Lex Fridman PodcastLeslie Kaelbling: Reinforcement Learning, Planning, and Robotics | Lex Fridman Podcast #15
EVERY SPOKEN WORD
110 min read · 21,941 words- 0:00 – 1:13
GEB, philosophy, and the early spark for AI
- LFLex Fridman
The following is a conversation with Leslie Kaelbling. She's a roboticist and professor at MIT. She's recognized for her work in reinforcement learning, planning, robot navigation, and several other topics in AI. She won the IJCAI Computers and Thought Award and was the editor-in-chief of the prestigious Journal of Machine Learning Research. This conversation is part of the Artificial Intelligence Podcast at MIT and beyond. If you enjoy it, subscribe on YouTube, iTunes, or simply connect with me on Twitter at lexfridman, spelled F-R-I-D. And now, here's my conversation with Leslie Kaelbling.
- LKLeslie Kaelbling
What made me get excited about AI, I can say that, is I read Godel, Escher, Bach when I was in high school. That was pretty formative for me because it exposed, uh, the interestingness of primitives and combination and how you can make complex things out of simple parts and ideas of AI and what kinds of programs might generate intelligent behavior. So...
- LFLex Fridman
So you first fell in love with AI reasoning logic versus robots?
- 1:13 – 3:05
From philosophy at Stanford to robotics at SRI
- LKLeslie Kaelbling
Yeah, the robots came because, um, my first job... So I finished an undergraduate degree in philosophy at Stanford and was about to finish a master's in computer science, and I got hired at SRI, uh, in their AI lab, and they were building a robot that was a kind of a follow-on to Shakey, but all the Shakey people were not there anymore.
- LFLex Fridman
Mm-hmm.
- LKLeslie Kaelbling
And so my job was to try to get this robot to do stuff, and that's really kind of what got me interested in robots.
- LFLex Fridman
So maybe taking a small step back-
- LKLeslie Kaelbling
Yeah.
- LFLex Fridman
...to your bachelor's in Stanford in philosophy-
- LKLeslie Kaelbling
Yeah.
- LFLex Fridman
...did master's and PhD in computer science, but the bachelor's in philosophy. Uh, so what was that journey like? What elements of philosophy do you think-
- LKLeslie Kaelbling
Yeah.
- LFLex Fridman
...you bring to your work in computer science?
- LKLeslie Kaelbling
So it's surprisingly relevant. So the... Part of the reason that I didn't do a computer science undergraduate degree was that there wasn't one at Stanford at the time, but that there's part of philosophy, and in fact, Stanford has a special sub-major in something called now symbolic systems which is logic model theory, formal semantics of natural language. And so that's actually a perfect preparation for work in AI and computer science.
- LFLex Fridman
That, that's kind of interesting. So if you were interested in artificial intelligence, what, what kind of majors were people even thinking about taking? Was it in neuroscience? Was... So besides philosophies, what, what were you supposed to do if you were fascinated by the idea of creating intelligence?
- LKLeslie Kaelbling
There weren't enough people who did that for that even to be a conversation.
- LFLex Fridman
Okay.
- LKLeslie Kaelbling
I mean, I think probably, probably philosophy. I mean, it's interesting, in my class, uh, my graduating class of undergraduate philosophers, probably, maybe slightly less than half went on in computer science-
- LFLex Fridman
Mm-hmm.
- LKLeslie Kaelbling
...slightly less than half went on in law, and, like, one or two went on in philosophy. Uh, so it was a common kind of connection.
- 3:05 – 5:42
What philosophy contributes to AI (and what it doesn’t)
- LFLex Fridman
Do you think AI researchers have a role, be part-time philosophers? Or should they stick to the solid science and engineering without sort of taking the philosophizing tangents? I mean, you work with robots, you think about what it takes to create intelligent beings. Uh, aren't you the perfect person to think about the big picture philosophy of it all?
- LKLeslie Kaelbling
The parts of philosophy that are closest to AI, I think, or at least the closest to AI that I think about are stuff like belief and knowledge and denotation and that kind of stuff. And that's, uh, you know, it's quite formal and it's, like, just one step away from the kinds of computer science work that we do kind of routinely. I think that there are important questions still about what you can do with a machine and what you can't and so on, although at least m- my personal view is that I'm completely a materialist and I don't think that there's any reason why we can't make a robot be behaviorally indistinguishable from a human. And the question of whether it's in- distinguishable internally, whether it's a zombie or not, in philosophy terms, I actually don't... I don't know and I don't know if I care too much about that.
- LFLex Fridman
Right. But there, there is, uh, philosophical notions, they're mathematical and philosophical because we don't know so much, of how difficult it is, how difficult is the perception problem? How difficult is the planning problem? How difficult is it to operate in this world successfully? Because our robots are not currently as successful as human beings in many tasks. The, the question about the gap between current robots and human beings borders a little bit on philosophy. Uh, you know, the, the expanse of knowledge that's required to operate in this world and the ability to, uh, form common sense knowledge, the ability to reason about uncertainty, much of the work you've been doing-
- LKLeslie Kaelbling
Mm-hmm.
- LFLex Fridman
...there's, there's open questions there that, uh, I, I don't know, require to activate a certain big-picture view.
- LKLeslie Kaelbling
To me, that doesn't seem like a philosophical gap at all.
- LFLex Fridman
I see.
- LKLeslie Kaelbling
That's just a t-... To me, it's a, it's a... There is a big technical gap.
- LFLex Fridman
Yes.
- LKLeslie Kaelbling
There's a huge technical gap. But I don't see any reason why it's more than a technical gap.
- LFLex Fridman
Perfect. (laughs) So when... You mentioned AI, you mentioned SRI, uh, and, uh, maybe can you describe to me when you first fell in love with robotics, with robots, or inspired, uh, which... So you sh- me- mentioned, uh, Flaky or Shakey, Shakey/Flaky. And wh- wh- what was the robot that first captured your imagination of what's possible?
- 5:42 – 7:22
Shakey the robot: foundational ideas in planning and navigation
- LKLeslie Kaelbling
Right. Well, the... So the first robot I worked with was Flaky. Shakey was a robot that the SRI people had built. But by the time... I think when I arrived it was sitting in a corner of somebody's office dripping hydraulic fluid into a pan. (laughs)
- LFLex Fridman
Yeah.
- LKLeslie Kaelbling
Uh, but it's iconic and really everybody should read The Shakey Tech Report 'cause it has so many good ideas in it. I mean, they...... invented A-star search and symbolic planning and learning macro operators. They had, uh, low level kind of configuration space planning for their robot. They had vision, they had all, uh, this, the basic ideas of a, a ton of things.
- LFLex Fridman
Can you take a step back?
- LKLeslie Kaelbling
Mm-hmm.
- LFLex Fridman
Does Shakey have arms? The, what was the job? What, what was-
- LKLeslie Kaelbling
Shakey, yeah.
- LFLex Fridman
... the goals that Shakey was trying to-
- LKLeslie Kaelbling
Good. Shakey was a mobile robot, but it could push objects, and so it would move things around.
- LFLex Fridman
With which actuator? With the arms?
- LKLeslie Kaelbling
It, with itself. With its bo-
- LFLex Fridman
Oh, okay.
- LKLeslie Kaelbling
With its base.
- LFLex Fridman
Okay. Great.
- LKLeslie Kaelbling
Um, so it could, but it, and they had painted the baseboards black. Uh, so it used, it used vision to localize itself in a map, it detected objects, it could detect objects that i- were surprising to it. Uh, it would plan and re-plan based on what it saw. It reasoned about whether to look and take pictures. I mean, it really had the basics of, of so many of the things that we think about now, um...
- LFLex Fridman
How did it represent the space around it?
- LKLeslie Kaelbling
So it had representations at a bunch of different levels of abstraction, so it had, I think, a kind of an occupancy grid of some sort at the lowest level. Uh, at the high level, it was, uh, abstract symbolic kind of rooms and connectivity.
- LFLex Fridman
S- so where does Flaky come in. And, and-
- 7:22 – 9:12
Flaky and “situated computation”: learning robotics by reinventing wheels
- LKLeslie Kaelbling
Yeah, okay. So I showed up at SRI and the, we were building a brand new robot. As I said, none of the people from the previous project were kind of there or involved anymore, so we were kind of starting from scratch. And my advisor, uh, was Stan Rosenshine, he ended up being my thesis advisor, and he was motivated by this idea of situated computation or situated automata. And the idea was that the tools of logical reasoning were important, but possibly only for the engineers or designers to use in the analysis of a system, but not necessarily to be manipulated in the head of the system itself, right?
- LFLex Fridman
Mm-hmm. Yeah.
- LKLeslie Kaelbling
So I might use logic to prove a theorem about the behavior of my robot even if the robot's not using logic in its head to prove theorems, right? So that was kind of the distinction. And so the idea was to kind of use those principles to make a robot do stuff. But a lot of the basic things we had to kind of learn for ourselves 'cause I had zero background in robotics, I didn't know anything about control, I didn't know anything about sensors, so we reinvented a lot of wheels on the way to getting that robot to do stuff.
- LFLex Fridman
Do you think that was an advantage or a hindrance?
- LKLeslie Kaelbling
Oh, no. It's, I, I, I mean, I th- I'm big in favor of wheel reinvention actually. I mean, I think you learn a lot by doing it.
- LFLex Fridman
Yes.
- LKLeslie Kaelbling
Uh, it's important though to eventually have the pointers to, so that you can see what's really going on. But I think you can appreciate much better the, the good solutions once you've messed around a little bit on your own and found a bad one.
- LFLex Fridman
Yeah, I think you mentioned reinventing reinforcement learning-
- LKLeslie Kaelbling
Yeah.
- LFLex Fridman
... and referring to, uh, rewards as pleasures, uh, pleasure.
- LKLeslie Kaelbling
Yeah.
- LFLex Fridman
I, or I think.
- LKLeslie Kaelbling
Yeah, that's a-
- LFLex Fridman
Uh, which I think is a nice name for it. (laughs)
- LKLeslie Kaelbling
(laughs) Yeah, that seemed good to me.
- 9:12 – 11:43
AI’s oscillating history: cybernetics, expert systems, and shifting problems
- LFLex Fridman
I think it's, it's more, it's more fun almost. D- do you think you could th- tell the history of AI machine learning, reinforcement learning, and how you think about it from the '50s to now?
- LKLeslie Kaelbling
One thing is that it oscillates, right? So things become fashionable and then they go out and then something else becomes cool and then it goes out and so on. And I think there's, so there's some interesting sociological process that actually drives a lot of what's going on. Early days was kind of cybernetics and control, right? And the idea that, of homeostasis, right? People who made these robots that could, I don't know, try to plug into the wall when they needed power and then come loose and roll around and do stuff. And then I think over time, the thought, well, that was inspiring, but people said, "No, no, no. We wanna get maybe closer to what feels like real intelligence or human intelligence."
- LFLex Fridman
Mm-hmm.
- LKLeslie Kaelbling
And then maybe the expert systems people tried to do that, but maybe a little too superficially, right? So, oh, we get the surface understanding of what intelligence is like because I understand how a steel mill works and I can try to explain it to you and you can write it down in logic and then we can make a computer infer that. And then that didn't work out. But what's interesting, I think, is when a thing starts to not be working very well, it's, not only do we change methods, we change problems, right? So it's not like we have better ways of doing the problem that the expert systems people were trying to do. We have no ways of trying to do that problem. Oh, yeah, no, I think, m- or, or maybe a few. But we kind of give up on that problem and we switch to a different problem and we, we work that for a while and we make progress.
- LFLex Fridman
As a, as a broad community.
- LKLeslie Kaelbling
As a community, yeah.
- LFLex Fridman
And there's a lot of people who would argue you don't give up on the problem, it's just that you, uh, decrease the number of people working on it. You almost kind of like-
- LKLeslie Kaelbling
Mm-hmm.
- LFLex Fridman
... put it on a shelf, say, "We'll come back to this 20 years later."
- LKLeslie Kaelbling
Come back, yeah. Yeah.
- LFLex Fridman
(laughs) That kind of thing.
- LKLeslie Kaelbling
I think that's right. Or you might decide that it's malformed. Like you might say, "It's wrong to just try to make something that does superficial symbolic reasoning behave like a doctor. You can't do that until you've had the sensory motor experience of being a doctor or something," right? So there's arguments that say that that problem was not well-formed, or it could be that it is well-formed but, but we just weren't approaching it well.
- 11:43 – 15:17
Why expert systems hit a wall—and what “symbolic” should really mean
- LFLex Fridman
So you menti- mentioned that your favorite part of logic in symbolic systems is that they give short names for large sets. So there is some use to this, uh, the use, to s- symbolic reasoning, though, i- looking at expert systems and symbolic computing, what, what do you think are the roadblocks that were hit in the '80s and '90s?
- LKLeslie Kaelbling
Ah, okay. So, right. So the fact that I'm not a fan of expert systems doesn't mean that I'm not a fan of some kinds of symbolic reasoning, right?
- LFLex Fridman
Right.
- LKLeslie Kaelbling
So-Let's see. Roadblocks. Well, the main roadblock, I think, was that the idea that humans could articulate their knowledge effectively into, into, you know, some kind of logical statements.
- LFLex Fridman
So it's not just the cost, the effort, but really just the capability of doing it.
- LKLeslie Kaelbling
Right. Because we're all experts in vision, right?
- LFLex Fridman
Yeah.
- LKLeslie Kaelbling
But totally don't have introspective access into how we do that, right? And it's true that, uh, the, the, I mean, I think the idea was, well, of course, even people then would know, of course I wouldn't ask you to please write down the rules that you use for recognizing a water bottle.
- LFLex Fridman
Mm-hmm.
- LKLeslie Kaelbling
That's crazy, and everyone understood that. But we might ask you to please write down the rules you use for deciding, I don't know, what tie to put on- (laughs)
- LFLex Fridman
Right.
- LKLeslie Kaelbling
... or how to set up a microphone-
- LFLex Fridman
Right.
- LKLeslie Kaelbling
... or something like that. But even those things, I think, people maybe... I think what they found, I'm not sure about this, but I think what they found was that the so-called experts could give explanations that, sort of post hoc explanations for how and why they did things, but they weren't necessarily very good. And then they def- th- they depended on maybe some kinds of perceptual things, which, again, they couldn't really define very well. So I think, I think fundamentally, I think the, the underlying problem with that was the assumption that people could articulate how and why they make their decisions.
- LFLex Fridman
Right. So it's almost enco- uh, encoding the knowledge, uh, fro- converting from expert to something that a machine can understand and reason with.
- LKLeslie Kaelbling
No, no. No, no. Not even just encoding, but getting it out of you.
- LFLex Fridman
Just... (laughs)
- LKLeslie Kaelbling
(laughs) Right? Not, not, not writing it... I mean, y- yes. Hard also to write it down for the computer.
- LFLex Fridman
Yeah.
- LKLeslie Kaelbling
But I don't think that people can produce it. You can tell me a story about why you do stuff, but I'm not so sure that's the why.
- LFLex Fridman
Great. So there are still, on the he- hierarchical planning side, places where symbolic reasoning is very useful. So, um, as, as you've talked about. So w- where-
- LKLeslie Kaelbling
Right. So don't-
- LFLex Fridman
Where's the gap?
- LKLeslie Kaelbling
Yeah. Okay, good. So saying that humans can't provide a description of their reasoning processes, that's okay, fine, but that doesn't mean that it's not good to do reasoning of various styles inside a computer. Those are just two orthogonal points. So then the question is, uh, what kind of reasoning should you do inside a computer?
- LFLex Fridman
Right.
- LKLeslie Kaelbling
Uh, and the answer is I think you need to do all different kinds of reasoning inside a computer, depending on what kinds of problems you face.
- LFLex Fridman
I guess the question is, what kinda things can you, uh, encode, uh, symbolically so you can reason about...
- LKLeslie Kaelbling
I think the idea about... And, and even symbolic, I don't even like that terminology-
- LFLex Fridman
(laughs)
- LKLeslie Kaelbling
... 'cause I don't know what it means-
- 15:17 – 18:04
Abstractions for planning: shrinking state, horizon, and complexity
- LKLeslie Kaelbling
... technically and formally. I do believe in abstractions. So abstractions are critical, right? It, you cannot reason at completely fine grain about everything in your life, right? You can't-
- LFLex Fridman
Right.
- LKLeslie Kaelbling
... make a plan at the level of images and torques for getting a PhD.
- LFLex Fridman
Right.
- LKLeslie Kaelbling
So you have to reduce the size of the state space, and you have to reduce the horizon if you're gonna reason about getting a PhD or even buying the ingredients to make dinner. And so, so how can you reduce the spaces and the horizon of the reasoning you have to do? And the answer is abstraction. Spatial abstraction, temporal abstraction. I think abstraction along the lines of goals is also interesting, like you might... Or well, abstraction and decomposition.
- LFLex Fridman
Mm-hmm.
- LKLeslie Kaelbling
Like, goals is maybe more of a decomposition thing. So I think that's where these kinds of, if you wanna call it symbolic or discrete models come in. You, you talk about a room of your house instead of your pose.
- LFLex Fridman
Mm-hmm.
- LKLeslie Kaelbling
You talk about, uh, you know, doing something during the afternoon instead of at 2:54. And you do that because it makes your reasoning problem easier and also because y- you have, you don't have enough information to reason in high fidelity about your pose of your elbow at 2:35 this afternoon anyway.
- LFLex Fridman
Right. When you're trying to get a PhD. That, that, that's-
- LKLeslie Kaelbling
Right. Or when you're doing anything really.
- LFLex Fridman
Oh, re- yeah, okay. Uh-
- LKLeslie Kaelbling
Except for at that moment. At that moment, you do have to reason about the pose of your elbow maybe.
- LFLex Fridman
Right.
- LKLeslie Kaelbling
But then you, maybe you do that in some continuous joint space kinda model. And so I, again, I, m- my biggest point about all of this is that there should be... that dogma is not the thing, right? We shouldn't, it shouldn't be that I am in favor against symbolic reasoning, and you're in favor against neural networks. It should be that just, just computer science tells us what the right answer to all these questions is if we were smart enough to figure it out.
- LFLex Fridman
Well, yeah. When you try to actually solve the problem with computers, eh, the, the right answer comes out. But you mentioned abstractions.
- LKLeslie Kaelbling
Mm-hmm.
- LFLex Fridman
I mean, neural networks form abstractions, or, uh, rather, uh, there's a- there's automated ways to form abstractions.
- LKLeslie Kaelbling
Absolutely.
- LFLex Fridman
And there's expert-driven ways to form abstractions.
- LKLeslie Kaelbling
Mm-hmm.
- LFLex Fridman
And, uh, expert human-driven ways.
- LKLeslie Kaelbling
Mm-hmm.
- LFLex Fridman
And humans just seems to be way better at forming abstractions currently in certain problems. So when you're referring to 2:45 AM, uh, PM versus afternoon, how do we construct that taxonomy? Is there any room for automated construction of such abstractions?
- LKLeslie Kaelbling
Oh, I think eventually, yeah. I mean, I think when we get to be better and machine learning engineers, we'll build algorithms that build awesome abstractions.
- LFLex Fridman
That are useful in this kinda way that you're describing?
- LKLeslie Kaelbling
Yeah.
- 18:04 – 21:45
MDPs and POMDPs: modeling stance, uncertainty, and belief updates
- LFLex Fridman
Yeah. So let's then step from the, the abstraction discussion, and let's talk about, uh, POMMDPs.... partially observable Markov decision processes, so uncertainty. So first, what are Markov decision processes?
- LKLeslie Kaelbling
What are Markov decision processes?
- LFLex Fridman
Uh, how, and maybe how much of our world could be models and, uh, MDPs? How much when, when-
- LKLeslie Kaelbling
Yeah.
- LFLex Fridman
... you wake up in the morning and you're making breakfast, how, do, do you think of yourself as an MDP? Uh, so how do you think about, uh, MDPs and how they relate to our world?
- LKLeslie Kaelbling
Well, so there's a stance question, right? So a stance is a position that I take with respect to a problem. So I-
- LFLex Fridman
(laughs)
- LKLeslie Kaelbling
... as a researcher or a person who designs systems can decide to make a model of the world around me in some terms, uh, right? So I take this messy world and I say, "I'm gonna treat it as if it were a problem of this formal kind," and then I can apply solution, concepts, or algorithms, or whatever to solve that formal thing, right? So of course, the world is not anything. It's not an MDP or a POMDP. I don't know what it is, but I can model aspects of it in some way or some other way. And when I model some aspect of it in a certain way, that gives me some set of algorithms I can use.
- LFLex Fridman
You can model the world in all kinds of ways.
- LKLeslie Kaelbling
Yeah.
- LFLex Fridman
Uh, some have, some are more accepting of uncertainty, more easily modeling uncertainty of the world. Some really force-
- LKLeslie Kaelbling
Mm-hmm.
- LFLex Fridman
... the world to be deterministic. Uh, and so-
- LKLeslie Kaelbling
Right.
- LFLex Fridman
... certainly MDPs, uh, model the uncertainty of the world.
- LKLeslie Kaelbling
Yes. Model some uncertainty. They model-
- LFLex Fridman
Some.
- LKLeslie Kaelbling
... not present state uncertainty, but they model uncertainty in the way the future will unfold.
- LFLex Fridman
Right.
- LKLeslie Kaelbling
Yeah.
- LFLex Fridman
So what are Markov decision processes? Is it-
- LKLeslie Kaelbling
Okay. So a Markov decision process is a model. It's a kind of a model that you can make that says, "I, I know completely the current state of my system." And what it means to be a state is that I, that all the inf- I have all the information right now that will let me make predictions about the future-
- LFLex Fridman
Mm-hmm.
- LKLeslie Kaelbling
... as well as I can, so that remembering anything about my history wouldn't make my predictions any better. Um, and but, uh, but then it also says that, uh, that then I can take some actions that might change the state of the world and that I don't have a deterministic model of those changes. I have a, a probabilistic model-
- LFLex Fridman
Mm-hmm.
- LKLeslie Kaelbling
... of how the world might change. Uh, it's a, it's a useful model for some kinds of systems. I think it's a, I mean, it's certainly n- not a good model for most problems, I think, because for most problems you don't actually know the state. Uh, for most problems, you, it's partially observed. So that's now a different problem class.
- LFLex Fridman
So the, okay, that's where the POMDPs, the partially observable-
- LKLeslie Kaelbling
Yeah.
- LFLex Fridman
... Markov decision processes step in. So how do they address the fact that you, uh, can't observe most, uh-
- LKLeslie Kaelbling
Correct.
- 21:45 – 23:22
Planning under uncertainty is intractable—so approximate intelligently
- LFLex Fridman
And so how difficult is this problem of planning under uncertainty, in your view, in your long experience with modeling the world, trying to deal with this uncertainty in, especially in real world systems?
- LKLeslie Kaelbling
Optimal planning for even discrete POMDPs can be undecidable depending on how you set it up. And for, so lots of people say, "I don't use POMDPs because they are intractable."
- LFLex Fridman
(laughs)
- LKLeslie Kaelbling
And I think that that's a kind of a very funny thing to say because the problem you have to solve is the problem you have to solve. So if the problem you have to solve is intractable, that's what makes us AI people, right? So, uh, we solve, we understand that the problem we're solving is, is complete- wildly intractable, that we can't, we will never be able to solve it optimally, at least I don't... Yeah, right. So later we can come back to an idea about bounded optimality in something. But anyway, I don't, we can't come up with optimal solutions to these problems.
- LFLex Fridman
Mm-hmm.
- LKLeslie Kaelbling
So we have to make approximations, approximations in modeling, approximations in solution algorithms, and so on. And so I don't have a problem with saying, "Yeah, my problem actually it is a POMDP in continuous space with continuous observations and it's so computationally complex I can't even think about its, you know, big O, whatever." Uh, but that doesn't prevent me from... It helps me, gives me some clarity to think about it that way.
- LFLex Fridman
Mm-hmm.
- LKLeslie Kaelbling
And to then take steps to make approximation after approximation to get down to something that's, like, computable in some reasonable time.
- 23:22 – 26:30
From engineering to science: bounded optimality, theory gaps, and what guarantees mean
- LFLex Fridman
When you think about optimality, you know, the, the community broadly has shifted on, on that, uh, I think a little bit in how much they value the idea of, uh, optimality, of chasing as-
- LKLeslie Kaelbling
Mm-hmm.
- LFLex Fridman
... an optimal solution. How has your views of chasing an optimal solution, uh, changed over the years and when you work with robots?
- LKLeslie Kaelbling
That's interesting. I th- I think we have a m- little bit of a methodological crisis actually from the theoretical side. I mean, I do think that theory is important and that right now we're not doing much of it. So there's lots of empirical hacking around and training this and doing that and reporting numbers. But is it good? Is it bad? We don't know. We- it- it's very hard to say things.
- LFLex Fridman
Right.
- LKLeslie Kaelbling
And if you look at, like, computer science theory, so people talked... For a while, everyone was about solving problems optimally or completely and, and then there were interesting relaxations, right? So people look at, "Oh, can I... Are there regret bounds or can I do some kind of, um, you know, approximation? Can I prove something that I can approximately solve this problem, or that I get closer to the solution as I spend more time?" And so on. What's interesting, I think, is that we don't have good approximate solution concepts for very difficult problems. Right? I like to, you know, I like to say that I, I'm interested in doing a very bad job of very big problems. Uh- (laughs) Uh... Good quote (laughs) . (laughing) Right. So very bad job at very big problems. I like to do that. But I would... I wish I could say something... I wish I had a, I don't know, some kind of a, of a f- formal solution concept that I could use to say, "Oh, this, this algorithm actually... It, it gives me something." Like, I know what I'm gonna get. I can do something other than just run it and get out 6.7.
- LFLex Fridman
So that, that notion is still somewhere deeply compelling to you?
- LKLeslie Kaelbling
Hmm.
- LFLex Fridman
The notion that you can say... (laughs) You can drop thing on the table that says, "This... You can expect that this algorithm will give me some good results."
- LKLeslie Kaelbling
I hope there's... I hope science will... I mean, there's engineering and there's science.
- LFLex Fridman
Mm-hmm.
- LKLeslie Kaelbling
I think that they're not exactly the same. And I think right now we're making huge engineering, like, leaps and bounds so that engineering is running way ahead of the science, which is cool and often how it goes, right? So we're making things and nobody knows how and why they work, roughly. But we need to turn that into science, I think.
- LFLex Fridman
There's some form... It's, uh... Yeah. There's some room for formalizing.
- LKLeslie Kaelbling
We need to know what the principles are. Why does this work? Why does that not work? I mean, for a while people built bridges by trying, but now we can often predict whether it's gonna work or not without building it. Can we do that for learning systems or for robots?
- LFLex Fridman
So your hope is, from a materialistic perspective, that intelligence, artificial intelligence systems, robots are ki-... Are just m- fancier bridges. Belief space.
- LKLeslie Kaelbling
Okay.
- 26:30 – 29:20
Belief space control: acting to change what you know, not just the world
- LFLex Fridman
What's the difference between belief space and state space? So we mentioned MDPs, POMDPs, you... Reasoning, uh, about... You sense the world, there's a state. Uh, what, what, what's this belief space idea?
- LKLeslie Kaelbling
Belief space. Yeah. Okay. That's... Yeah.
- LFLex Fridman
That sounds so good.
- LKLeslie Kaelbling
That's... Sounds good. So belief space, that is... Instead of thinking about what's the state of the world and trying to control that as a robot, I think about what is the space of beliefs that I could have about the world? What's... If I think of a belief as a probability distribution over ways the world could be, a belief state is a distribution. And then my control problem, if I'm reasoning about how to move through a world I'm uncertain about, my control problem is actually the problem of controlling my beliefs. So I think about taking actions, not just what effect they'll have on the world outside, but what effect they'll have on my own understanding of the world outside. And so that might compel me to ask a question or look somewhere to gather information, which may not really change the world state, but it changes my own belief about the world.
- LFLex Fridman
That's a powerful way to, to, to empower the agent to reason about the world, to explore the world. Uh, so what kind of problems does it allow you to solve to, to, uh, consider belief space versus just state space?
- LKLeslie Kaelbling
Well, any problem that requires deliberate information-gathering, right? So if... In some problems, like chess, there's no uncertainty, or maybe there's uncertainty about the opponent. Uh, there's no uncertainty about the state. Uh, and some problems there's uncertainty, but you gather information as you go, right? You might say, "Oh, I'm driving my autonomous car down the road and it doesn't know perfectly where it is, but the LiDARs are all going all the time, so I don't have to think about whether to gather information."
- LFLex Fridman
Mm-hmm.
- LKLeslie Kaelbling
But if you're a human driving down the road, you sometimes look over your shoulder to see what's going on behind you in the lane, and you have to decide whether you should do that now.
- LFLex Fridman
Mm-hmm.
- LKLeslie Kaelbling
And you have to trade off the fact that you're not seeing in front of you and you're looking behind you and how valuable is that information and so on. And so to make choices about information-gathering, you have to reason in belief space. Also, also, I mean, also to just take into account your own uncertainty before trying to do things. So you might say, "If I understand where I'm standing relative to the door jamb, uh, pretty accurately, then it's okay for me to go through the door. But if I'm really not sure where the door is, then it might be better to not do that right now."
- LFLex Fridman
The degree of your uncertainty about, about the world is actually part of the thing you're trying to optimize in forming the plan, right?
- LKLeslie Kaelbling
That's right. That's right.
- 29:20 – 35:01
Hierarchical planning in the real world: airports, feasibility leaps, and generalization
- LFLex Fridman
Th- so this idea of a long horizon of, uh, planning for a PhD or just even how to get out of the house or how to make breakfast. You, you show this presentation of the, the WTF, where's the fork?
- LKLeslie Kaelbling
(laughs)
- LFLex Fridman
Uh, of robot looking at a sink. Uh, and, uh, uh, can you describe how we plan in this world? There's this idea of hierarchical planning we've mentioned. Th- so, so yeah. H- how can a robot hope to plan about something, uh, uh, th- is, you know, with such a long hori- where the goal is quite far away?
- LKLeslie Kaelbling
People since probably reasoning began have thought about hierarchical reasoning, the temporal hierarchy in parti... Well, p- there's spatial hierarchy, but let's talk about temporal hierarchy.
- LFLex Fridman
Mm-hmm.
- LKLeslie Kaelbling
So you might say, "Oh, I have this long, uh, execution I have to do, but I can divide it into some segments abstractly." Right? So maybe I have to get out of the house, I have to get in the car, I have to drive and so on. And so-You can plan, if you can build abstractions. So this, we started out by talking about abstractions and we're back to that now. If you can build abstractions in your state space, and abstractions, sort of temporal abstractions, then you can make plans at a high level and you can say, "I'm gonna go to town and then I'll have to get gas, and then I can go here and I can do this other thing." And you can reason about the dependencies and constraints among these actions, again, without thinking about the complete details. What we do in our hierarchical planning work is then say, "All right, I make a plan at a high level of abstraction." I have to have some reason to think that it's feasible without working it out in complete detail.
- LFLex Fridman
Mm-hmm.
- LKLeslie Kaelbling
And that's actually the interesting step. I always like to talk about walking through an airport, like-
- LFLex Fridman
Mm-hmm.
- LKLeslie Kaelbling
... you can plan to go to New York and arrive at the airport and then find yourself in an office building later. You can't even tell me in advance what your plan is for walking through the airport.
- LFLex Fridman
Hm.
- LKLeslie Kaelbling
Partly because you're too lazy to think about it maybe, but partly also because you just don't have the information. You don't know what gate you're landing in or what people are gonna be in front of you or anything. So, there is no point in planning in detail.
- LFLex Fridman
Mm-hmm.
- LKLeslie Kaelbling
But you have to have ... You have to make a leap of faith that you can figure it out once you get there. And it's really interesting to me how you arrive at that. How do you ... So you have learned over your lifetime to be able to make some kinds of predictions about how hard it is to achieve some kinds of sub-goals.
- LFLex Fridman
Mm-hmm.
- LKLeslie Kaelbling
And that's critical. Like, you would never plan to fly somewhere if you couldn't, didn't have a model of how hard it was to do some of the intermediate steps. So one of the things we're thinking about now is how do you do this kind of very aggressive generalization, uh, I mean, to situations that you haven't been in and so on to predict how long will it take to walk through the Kuala Lumpur airport?
- LFLex Fridman
Mm-hmm.
- LKLeslie Kaelbling
Like, you, you could give me an estimate and it wouldn't be crazy. And you have to have an estimate of that in order to make plans that involve walking through the Kuala Lumpur airport, even if you don't need to know it in detail. So I'm really interested in these kinds of abstract models and how do we acquire them. But once we have them, we can use them to do hierarchical reasoning, which is, I think is very important.
- LFLex Fridman
Yeah, there's this notion of go- uh, goal regression and pre-image back chaining, this idea of starting at the goal-
- LKLeslie Kaelbling
Mm-hmm.
- LFLex Fridman
... and just forming these big clouds of states that you, you get, I mean, it, it's almost like saying to the airport, you know, you, you know once you show up to the, uh, the airport that that's, you, you're like a few steps away from the goal. So like, thinking of it this way, uh, is kind of interesting. I don't know if you have sort of, uh, further comments on that-
- LKLeslie Kaelbling
Hm.
- LFLex Fridman
... uh, uh, of starting at the goal, why that's Yeah.
- LKLeslie Kaelbling
cool. I mean, it's interesting that Simon, Herb Simon, back in the early days of AI did, talked a lot about means-ends reasoning and reasoning back from the goal. There's a kind of an intuition that people have that the number of the, with, so state space is big, the number of actions you could take is really big. So if you say, "Here I sit and I wanna search forward from where I am, what are all the things I could do?"
- LFLex Fridman
Right.
- LKLeslie Kaelbling
That's just overwhelming. If you say, if you can reason at this other level and say, "Here's what I'm hoping to achieve. What could I do to make that true?" That somehow the branching is smaller. Now, what's interesting is that, like in the AI planning community, that hasn't worked out. In the class of problems that they look at and the methods that they tend to use, it hasn't turned out that it's better to go backward. Um, it's still kind of my intuition that it is, but I can't prove that to you right now.
- LFLex Fridman
(laughs) Right. I share your intuition, at least for us mere humans.
- LKLeslie Kaelbling
Mm-hmm.
- LFLex Fridman
Speaking of which, uh, y- when you, uh, maybe now we take a, take a, take a little step into that philosophy circle.
- LKLeslie Kaelbling
Uh-oh.
- 35:01 – 41:10
Model-based vs model-free, why perception is harder than planning, and building useful structure
- LKLeslie Kaelbling
W- well, that would be a mistake. I mean, it's not all a planning problem.
- LFLex Fridman
(laughs)
- LKLeslie Kaelbling
Right? I mean, I think it's really, really important that we understand that you have to put together pieces and parts that have different styles of reasoning and representation and learning. I think, I think it's, it's, it seems probably clear to anybody that, that it, it can't all be this or all be that. Uh, brains aren't all like this or all like that, right? They have different pieces and parts and substructure and so on. So I don't think that there's any good reason to think that there's gonna be like one true algorithmic thing that's gonna do the whole job.
- LFLex Fridman
So it's a bunch of pieces together, uh, designed to solve a bunch of specific problem. One specific, uh-
- LKLeslie Kaelbling
Or maybe s- styles of-
- LFLex Fridman
Hm.
- LKLeslie Kaelbling
... problems. I mean, there's probably some reasoning that needs to go on in image space. I think, again, it, there's this model-based versus model-free idea, right? So in-
- LFLex Fridman
Mm-hmm.
- LKLeslie Kaelbling
... reinforcement learning, people talk about, "Oh, should I learn ... I could learn a policy just straight up a way of behaving."
- LFLex Fridman
Mm-hmm.
- LKLeslie Kaelbling
"I could learn, it's popular to learn, a value function that's some kind of weird intermediate ground. Uh, or I could learn a transition model which tells me something about the dynamics of the world." If I take a tra- if, imagine that I learn a transition model and I couple it with a planner and I draw a box around that, I have a policy again. It's just stored a different way.Right?
- LFLex Fridman
Right.
- LKLeslie Kaelbling
It's, and, but it's just as much of a policy as the other policy. It's just I've made... I think the way I see it is it's a time-space trade-off in computation, right? Uh, more overt policy representation, maybe it takes more space, but maybe I can compute quickly what action I should take. On the other hand, maybe a very compact model of the world dynamics plus a planner lets me compute what action to take too, just more slowly. There's no, I don't, I mean, I don't think, there's no argument to be had. It's just like a question of what form of computation is best for us.
- LFLex Fridman
For the various sub-problems.
- LKLeslie Kaelbling
Right. So and, and so like learning to do algebra manipulations for some reason is g- I mean, that's probably gonna want naturally a sort of a different representation than riding a unicycle.
- LFLex Fridman
Mm-hmm.
- LKLeslie Kaelbling
Right? The time constraints on the unicycle are serious. The state spaces may be smaller. I don't know. But so I, uh-
- LFLex Fridman
And it could be the more human sides of falling in love, having a relationship. That might be another, uh-
- LKLeslie Kaelbling
Yeah, it goes-
- LFLex Fridman
... another style of-
- LKLeslie Kaelbling
I have no idea.
- LFLex Fridman
... how to model that. Yeah.
- LKLeslie Kaelbling
No. No.
- LFLex Fridman
Let's (laughs) let's first, uh-
- LKLeslie Kaelbling
(laughs) Right.
- LFLex Fridman
... solve the algebra and the object manipulation.
- LKLeslie Kaelbling
Yeah.
- LFLex Fridman
Uh, s- what do you think is harder, perception or planning?
- LKLeslie Kaelbling
Perception. That's why I've stopped planning.
- LFLex Fridman
Understanding. (laughs) That's why.
- 41:10 – 1:01:08
Human-level robotics, benchmarks, and AI futures: value alignment and research incentives
- LFLex Fridman
I agree. Very compelling description of actually where we stand with the perception problem. Uh, you're teaching a course on embodied intelligence. What do you think it takes to build a robot with human-level intelligence?
- LKLeslie Kaelbling
I don't know. If we knew, we would do it. Uh. (laughs) Uh-
- LFLex Fridman
I- if you were to... I mean, okay. So do you think a robot needs to have a, uh, self-awareness, uh, consciousness-
- LKLeslie Kaelbling
Uh-
- LFLex Fridman
... fear of mortality, or is it, is it simpler than that? Or is consciousness a simple thing? Like, do you th- do you think about these notions?
- LKLeslie Kaelbling
I don't think much about consciousness. Even most philosophers who care about it will give you that you could have robots that are zombies, right, that behave like humans-
- LFLex Fridman
Right.
- LKLeslie Kaelbling
... but are not conscious. And I, at this moment, would be happy enough with that, so I'm not really worried one way or the other.
- LFLex Fridman
So then on a technical side, you're not thinking of, uh, the use of self-awareness, um ?
- LKLeslie Kaelbling
Well, but I, okay. But then what does self-awareness mean? I mean, that you need to have some part of the system that can observe other parts of the system and tell whether they're working well or not, that seems critical. So does that count as, I mean, does that count as self-awareness or not? Well, it depends on whether you think that there's somebody at home who can articulate whether they're self-aware. But clearly, if I have like, you know, some piece of code that's counting how many times this procedure gets executed-... that's a kind of self-awareness, right? So I, there's a big spectrum. It's clear you have to have some of it.
- LFLex Fridman
Right. You know, we're quite far away in many dimensions, but is there a direction of research that's most compelling to you for, you know, trying to achieve human-level intelligence in, in our robots?
- LKLeslie Kaelbling
Well, to me, I guess, the thing that seems most compelling to me at the moment is this question of what to build and, uh, what to learn. Um, I think we're, we don't, we're missing a bunch of ideas. And, and we, you know, people, you know, don't you dare ask me how many years it's gonna be till that happens-
- LFLex Fridman
(laughs)
- LKLeslie Kaelbling
... 'cause I won't even participate in the conversation.
- LFLex Fridman
Yeah.
- LKLeslie Kaelbling
'Cause I think we're missing ideas and I don't know how long it's gonna take to find them.
- LFLex Fridman
So I won't ask you how many years, but, uh, maybe I'll ask you what it, when you'll be sufficiently impressed that we've achieved it. So what's, what's, uh, a good test of intelligence? Do you like the Turing test and natural language in the robotics space? Is there something where you would sit back a- and think, "Oh, that's, that's pretty impressive," uh, a- as a test, as a benchmark? Do you, do you think about these kinds of problems?
- LKLeslie Kaelbling
No. I, I resist... I mean, I think all the time that we spend arguing about those kinds of things could be better spent just making the robots work better. Uh, so-
- LFLex Fridman
You don't value competition? So, I mean, there's-
- LKLeslie Kaelbling
No.
- LFLex Fridman
... the nature of benchmark, uh, uh, benchmarks and datasets or Turing test challenges where everybody kinda gets together and tries to build a better robot 'cause they wanna out-compete each other, like the DARPA challenge with the autonomous vehicles. Uh, do you see the value of that? Or it can get in the way?
- LKLeslie Kaelbling
I think it can get in the way. I mean, and some people, many people find it motivating and so that's good. I find it anti-motivating personally.
- LFLex Fridman
(laughs) Yeah.
- LKLeslie Kaelbling
Uh, but I think what, I mean, I think you get an interesting cycle where for a contest, a bunch of smart people get super motivated and they hack their brains out. And much of what gets done is just hacks, but sometimes really cool ideas emerge, and then that gives us something to chew on after that. So I'm, I, it's not a, a thing for me, but I don't, I don't regret that other people do it.
- LFLex Fridman
Yeah. It's, uh, like you said with everything else, the mix is good. So jumping topics a little bit, you started the Journal of Machine Learning Research and served as its editor in chief. Uh, how did the publication come about?
- LKLeslie Kaelbling
Hm.
- LFLex Fridman
And, uh, what do you think about the current publishing model space in, uh, machine learning-
- LKLeslie Kaelbling
Hmm.
- LFLex Fridman
... artificial intelligence, and of course
- NANarrator
That's interesting.
Episode duration: 1:01:23
Install uListen for AI-powered chat & search across the full episode — Get Full Transcript
Transcript of episode Er7Dy8rvqOc