Skip to content
Lex Fridman PodcastLex Fridman Podcast

Tomaso Poggio: Brains, Minds, and Machines | Lex Fridman Podcast #13

Lex Fridman and Tomaso Poggio on tomaso Poggio on Intelligence, Brains, and the Limits of AI.

Lex FridmanhostTomaso Poggioguest
Jan 19, 20191h 20mWatch on YouTube ↗

EVERY SPOKEN WORD

  1. 0:003:39

    Einstein, thought experiments, and the power of nonconformity

    1. LF

      The following is a conversation with Tomaso Poggio. He's a professor at MIT, and is a director of the Center for Brains, Minds, and Machines. Cited over 100,000 times, his work has had a profound impact on our understanding of the nature of intelligence in both biological and artificial neural networks. He has been an advisor to many highly impactful researchers and entrepreneurs in AI, including Demis Hassabis of DeepMind, Amnon Shashua of Mobileye, and Christof Koch of the Allen Institute for Brain Science. This conversation is part of the MIT course on artificial general intelligence, and the Artificial Intelligence Podcast. If you enjoy it, subscribe on YouTube, iTunes, or simply connect with me on Twitter, @LexFridman, spelled F-R-I-D. And now, here's my conversation with Tomaso Poggio. You've mentioned that in your childhood, you've developed a fascination with physics, especially the theory of relativity, and that Einstein was also a childhood hero to you. What aspect of Einstein's genius, the nature of his genius, do you think was essential for discovering the theory of relativity?

    2. TP

      You know, Einstein was, uh, a hero to me, and I'm sure to many people, because he was able to make, uh, uh, of course, a major, major contribution to physics with, simplifying a bit, just a gedanken experiment, a thought experiment.

    3. LF

      Mm-hmm.

    4. TP

      You know, imagining, uh, communication with lights between a stationary observer and somebody on a train.

    5. LF

      Mm-hmm.

    6. TP

      And, uh, I thought that, um, you know, the, the, the, the fact that just with the force of y- of his thought, of his thinking, of his mind, it could get to some- something so deep in term of physical reality, how time depend on space and speed, is, was something absolutely fascinating. It was the power of intelligence, the power of the mind.

    7. LF

      Do you think the ability to imagine, to visualize as he did, as a lot of great physicists do, do you think that's in all of us human beings? Or is there something special to that one particular human being?

    8. TP

      I think, uh, you know, a- all of us can learn and have, uh, in principle similar b- breakthroughs. Uh, there is lesson to be learned from Einstein. Uh, he was one of five PhD students at ETH, uh, the Eidgenössische Technische Hochschule in, uh, Zurich, in physics, and he was the worst of the five. The only one who d- did not get an academic position when, uh, he, he graduated, when he finished his PhD, and he went to work, as everybody knows, for the patent office. And so it's not so much that he worked for the patent office, but the fact that obviously he was smart, but he was not a top student. O- obviously, he was the anticonformist, he was not thinking in the traditional way that probably his teachers and the other students were doing. So there is a lot to be said about, uh, you know, trying to be... to do the opposite or something quite different from what other people are doing. That's certainly true for the stock market. Never...

    9. LF

      (laughs)

    10. TP

      Never buy if everybody's buying. (laughs)

    11. LF

      And also true for science.

    12. TP

      Yes.

  2. 3:396:14

    Time travel skepticism and the broader dream of building intelligence

    1. LF

      So you've also mentioned, staying on the theme of physics, that you were excited at a young age by the mysteries of the universe that, uh, physics could uncover. Such, as I saw mentioned, the possibility of time travel.

    2. TP

      (laughs)

    3. LF

      So the most out-of-the-box question I think I'll get to ask today, do you think time travel is possible?

    4. TP

      Well, it would be nice if it were possible right now. Uh, you know, in science, you never say no, um...

    5. LF

      But your understanding of the nature of time.

    6. TP

      Yeah, it's very likely that it's not possible to travel in time. Um, you may be able to travel forward in time if we can, for instance, freeze ourselves or, uh, you know, go on some spacecraft traveling close to the speed of light. But in terms of actively traveling, for instance, back in time, I find probably very unlikely.

    7. LF

      So do you still hold the, the underlying dream of the engineering intelligence that we'll build systems that are able to do such huge leaps, like discovering the kind of mechanism that would be required to travel through time? Do you still hold that dream or is- or echoes of it from your childhood?

    8. TP

      Yeah, I, uh, you know, I don't think whether... Uh, there are certain problems that probably cannot be solved depending what, uh, what you believe about the physical reality, like, uh, you know, maybe totally impossible to create energy from nothing or to travel back in time. But, uh, um, about making machines that can, uh, think as well as w- we do or better, or more likely, especially in the short and mid-term, help us think better, which is in a sense is happening already with the computers we have, and it will happen more and more. Well, that I certainly believe, and I don't see in principle why computers at some point, uh, could not become more intelligent than we are. Although the word intelligence...... is a tricky one and one we should discuss- (laughs)

    9. LF

      Yeah, for sure.

    10. TP

      ... what I mean with that. (laughs)

    11. LF

      Uh, in- intelligence, consciousness-

    12. TP

      Yeah.

    13. LF

      ... words like love, is, all these are very, uh-

    14. TP

      Yeah.

  3. 6:148:46

    Why intelligence is the biggest scientific problem

    1. LF

      ... need to be disentangled. So, you've mentioned also that you believe the problem of intelligence is the greatest problem in science, greater than the origin of life and the origin of the universe. You've also, uh, in the talk I've listened to, uh, said that you're open to arguments against, uh, against you. So, uh, what do you think is the most captivating aspect of this problem of understanding the nature of intelligence? Why does it captivate you as it does?

    2. TP

      Well, originally, I think one of the motivation that I had as a, I guess, a teenager, when I was infatuated with theory of relativity, was really that I- I found that there was, uh, the problem of time and space and general relativity, but there were so many other problems of the same level of difficulty and importance that I could ... even if I were Einstein, it was difficult to hope to solve all of them. So, what about solving a problem whose solution en- allowed me to solve all the problems? And this was (laughs) what if we could find the key to an intelligence, you know, 10 times better or faster than Einstein?

    3. LF

      So, that's sort of seeing artificial intelligence as a- as a tool to expand our capabilities, but is there just an inherent curiosity in you in just understanding what it is in our- in- in here that makes it all- all work?

    4. TP

      Yes, absolutely. You are right. So, I was starting- I started saying this was the motivation when I was a teenager-

    5. LF

      Right.

    6. TP

      ... but, uh, you know, soon after, uh, I think the problem of human intelligence became a- a real focus of- f- you know, of- of my sci- of my science and my research because I think is, for me, the most interesting problem is really asking, uh, who we- we are, right? It's asking not only a question about science, but even about the very tool we are using to do science, which is our brain. How does our brain work? From where does it come from? What are its limitation? Can we make it better?

  4. 8:4613:07

    Can we build strong AI without understanding the brain? Lessons from flight

    1. LF

      And that, in many ways, is the ultimate question that underlies this whole effort of science. So, you've made significant contributions in both the science of intelligence and the engineering of intelligence. In a hypothetical way, let me ask, how far do you think we can get in creating intelligence systems without understanding the biological, the understanding how the human brain creates intelligence? Put another way, do you think we can build a strong AI system without really, uh, getting at the core, the functional nat- uh, understanding the functional nature of the brain?

    2. TP

      Well, this is a- a real difficult question. You know, we did, uh, um, solve problems like flying without, uh, really using too much of knowledge about how birds fly. It was important, I guess, to know that you could have, uh, things heavier than- than air being able to fly like- like birds. But beyond that, probably we did not learn very much, you know, some, you know. The Brothers Wright did learn a lot of observation about flo- birds and designing their- their aircraft, but, uh, you know, you can argue we did not use much biology in that particular case. Now, in the case of intelligence, I think that, um, uh, w- it's- it's a bit of a bet right now. If you a- if you ask, uh, okay, um, we- we all agree we'll get, at some point, maybe soon, maybe later, to a machine that is indistinguishable from my secretary, say, in terms of what I can ask the machine to do. I- I think we'll get there, and now the question is, and you can ask people, do you think we'll get there without any knowledge about, uh, you know, the human brain, or that the best way to get there is to understand better the human brain?

    3. LF

      Yeah.

    4. TP

      Okay. This is, I think, an educated bet that different people with different background will decide in different ways. The recent history of the progress in AI in the last, uh, I would say five years or 10 years is- has been the- the main, uh, breakthroughs, the main recent breakthroughs. I really start from neuroscience. I can mention reinforcement learning as one, is one of the algorithms at the core of AlphaGo, which is the system that beat the kind of an official world champion of Go, Lee Sedol, in two, three years ago in, uh, Seoul. Um...That's one, and that started really with the work of Pavlov, um... (laughs)

    5. LF

      (laughs) And his dog.

    6. TP

      ... 1900, Marvin Minsky in the '60s, and many ne- other neuroscientists later on. Um, and deep learning, uh, uh, started, uh, which is at the core again of AlphaGo and systems like, uh, autonomous, uh, driving systems for cars, like the systems that, uh, Mobileye, which is a company started by one of my ex-postdoc, Amnon Shashua-

    7. LF

      Yes.

    8. TP

      ... um, did. So that, that is the core of those things, and deep learning, re- really the initial ideas in terms of the architecture of these layered hierarchical networks started with work of Thorsten Wiesel and David Hubel at Harvard up the river in the '60s. So recent history suggests that neuroscience played a big role in these breakthroughs. My personal bet is that there is a good chance they continue to play a big role, maybe not in all the future breakthroughs, but in some of them.

    9. LF

      At least in inspiration. So you-

    10. TP

      At least in an inspiration, absolutely, yes.

  5. 13:0717:16

    Biological vs artificial neural nets: what’s missing today (labels, data, learning)

    1. LF

      So y- so you studied both, uh, artificial and biological neural networks. You said these, uh, mechanisms that underlie deep learning, deep, uh, and reinforcement learning, but there is nevertheless, uh, significant differences between biological and artificial neural networks as they stand now. So between the two, wh- what do you find is the most interesting, mysterious, maybe even beautiful difference as, as it currently stands in our understanding?

    2. TP

      I must confess that until recently, I found the artificial networks too simplistic relative to real neural networks. But, uh, you know, recently, I've been started to think that, yes, they're a very big simplification of what you find in the brain, but on the other hand, uh, are much closer in terms of the architecture to the brain than other models that we had, that computer science used as model of thinking, which were mathematical logics, you know, LISP, Prolog, and those kind of things. (laughs)

    3. LF

      (laughs) Yeah.

    4. TP

      So in comparison to those, they're much closer to the brain. You have networks of neurons, which is what the brain is about, and the, the artificial neurons in the models are, as I said, caricature of the biological neurons, but they're still neurons, single units communicating with other units, something that is absent in, you know, the traditional, uh, computer-type models of mathematics, reasoning, and so on.

    5. LF

      So what aspect do you th- would you like to see in artificial neural networks added over time as we try to figure out ways to improve them?

    6. TP

      So one of the main differences, and, um, you know, problems, in terms of deep learning today, and it's not only deep learning, and the brain, is the need for deep learning techniques to have, uh, a lot of labeled examples. You know, for instance, for ImageNet, you have like a training set which is one million images, each one labeled by some human-

    7. LF

      Mm-hmm.

    8. TP

      ... in terms of which object is there, and, um, it's, it's clear that in biology, a baby, uh, may be able to see million of images in the first years of life, but will not have million of labels given to him or her by parents or take, take, uh, caretakers. So, uh, how do you solve that? You know, I think that there is this interesting challenge that today, uh, deep learning and related techniques are all about big data, big data meaning a lot of examples labeled by humans-

    9. LF

      Mm-hmm.

    10. TP

      ... um, whereas in, uh, nature you have, uh ... So the, the, th- this big data is n going to infinity, that's the best, you know, n meaning labeled data. But I think the biological world is more n going to one.

    11. LF

      (laughs)

    12. TP

      A- a child can learn-

    13. LF

      It's a beautiful way to put it.

    14. TP

      ... from very small number of, you know, labeled examples. Like you tell a child, "This is a car." You don't need to say, like in ImageNet, you know, "This is a car, this is a car, this is not a car, this is not a car" o- one million times. (laughs)

    15. LF

      So... And of course, with AlphaGo and, or at least the-

    16. TP

      Yeah.

    17. LF

      ... AlphaZero variance, there's ... Because of the, because the world of Go is so simplistic, that you can actually learn by yourself through self-play, you could play against each other.

    18. TP

      Yep.

    19. LF

      And the real world, I mean, the visual system that you've studied extensively, is a lot more complicated than the game of Go. So-

    20. TP

      Right.

  6. 17:1622:36

    Nature vs nurture in learning: evolution, priors, and face-recognition plasticity

    1. LF

      ... uh, on the comment about children, which are fascinatingly good at learning new stuff, how much of it do you think is hardware and how much of it is software?

    2. TP

      Yeah, that's a, a good, a deep question, is, in a sense, is the old question of nurture and nature.

    3. LF

      Yeah.

    4. TP

      How much is in, in the gene and how much is, uh, in the experience of an individual. Obviously, it's both that play a role, and, uh, I believe that the way evolution gives, put prior information, so to speak, hardwired, it's not really hardwired, but, um, uh, that's-Essentially an hypothesis. I think what's going on is that w- evolution, as, um, you know, almost necessarily if you believe in Darwin, is very opportunistic, and, and think about, uh, uh, our DNA and the DNA of Drosophila.

    5. LF

      Mm-hmm.

    6. TP

      Uh, our DNA does not have many more genes than Drosophila. Oh, now I'm-

    7. LF

      The fly.

    8. TP

      The fly.

    9. LF

      Yeah.

    10. TP

      The fruit fly.

    11. LF

      Yeah.

    12. TP

      Now, we know that the fruit fly does not learn very much during its individual existence. It looks like one of these machinery that it's really mostly, not 100%, but, you know, 95%, hardcoded by the genes. But since we don't have many more genes than Drosophila is, evolution could encode in us a kind of general learning machinery, and then had to give very weak priors.

    13. LF

      Mm-hmm.

    14. TP

      Um, like, for instance, let me take, give a, uh, a specific example which is a recent work by a member of our Center for Brains, Minds, and Machines. We know because of work of other people in our group and other groups that there are cells in a part of our brain, neurons-

    15. LF

      Mm-hmm.

    16. TP

      ... that are tuned to faces. They seems to be involved in face recognition. Now, this face area exist, uh, uh, seems to be present in young, uh, children and adults, um, and one question is, is there from the beginning, is hardwired by evolution, or, you know, somehow is learned very quickly?

    17. LF

      So, what's your... By the way, a lot of the questions s- I'm asking, we, the answer is we don't really know. But as a person who has contributed s- some profound ideas in these fields, you're a good person to guess at some of these. So, of course there's a caveat before a lot of the stuff we talk about. But what is your hunch? Is the face, the part of the brain that, that seems to be concentrated on face recognition, are you born with that or are you just, it's designed to, to learn that quickly, like the face of the mother and so on?

    18. TP

      My, my hunch, you know, my, my bias was the second one, learned very quickly. And, uh, turns out that Marge Livingstone at Harvard has done some, uh, amazing experiments in which she raised baby monkeys, depriving them of faces during the first weeks of life.

    19. LF

      Mm-hmm.

    20. TP

      So, they see technicians, but the technician have a mask.

    21. LF

      Yes.

    22. TP

      And, um, and so when they looked, uh, um, at, uh, the area in the brain of these monkeys that, where usually you find faces, they found no face preference. So, my guess is that what evolution does in this case is there is a plastic, an area which is plastic, which is kind of predetermined to be imprinted very easily. But the command from the gene is not a detailed circuitry for a face template.

    23. LF

      Mm.

    24. TP

      Could be. But this will require probably a lot of bits.

    25. LF

      Mm-hmm.

    26. TP

      'Cause you had to specify a lot of connection of a lot of neurons. Instead, the com- the command from the gene is something like imprint, memorize what you see most often in the first two weeks of life, especially in connection with food, and maybe nipples. (laughs) I don't know.

    27. LF

      Right. Well, source of food.

    28. TP

      Yeah.

    29. LF

      And so... And then that area's very plastic at first-

    30. TP

      Yeah.

  7. 22:3627:55

    Is the brain modular or uniform? Cortex as shared “hardware” across functions

    1. LF

      Can you, uh, talk about what are the different parts of the brain, and in your view, sort of loosely, and how do they contribute to intelligence? Do y- do you see the brain as a bunch of different modules and they together come in the human brain to create intelligence? Or is it all one, uh, mush of the same kind of fundamental-

    2. TP

      (laughs)

    3. LF

      ... uh, um, th-

    4. TP

      Right.

    5. LF

      ... uh, architecture?

    6. TP

      Yeah. That's, um, you know, that's, uh, an important question. And, uh, w- there was a phase in, uh, neuroscience back in the 1950 or so, in which, uh, it was believed for a while that the brain was equipotential. This was the term. You could cut out a piece and, um, nothing special happened apart l- little bit less performance. There was a, um, surgeon, Lashley, who did a lot of experiments of this type with mice and rats, and concluded that every part of the brain was essentially equivalent to any other one. It turns out that that's, that's really not true. It's, uh, there are very specific modules in the brain, as you said-

    7. LF

      Mm-hmm.

    8. TP

      You know, people may lose the ability to speak if you have a stroke in a certain region, or may lose control over their legs in another region, or... so they are very specific. Th- the brain is also quite flexible and redundant, so often it can correct things, and, uh, you know, r- uh, kind of, uh, um, take over f- functions from one, uh, part of the brain to the other, but, uh, but, but really, there are s- specific modules. So, the answer that we know from this old work, uh, uh, which was basically on, based on lesions-

    9. LF

      Mm-hmm.

    10. TP

      ... either on animals or very often there were, uh, a mine of, um, well, th- there was a mine of very interesting data coming from, um, from the war, from different types of, um-

    11. LF

      Injuries.

    12. TP

      ... injuries that soldiers had in the brain. And, uh, more recently, um, functional MRI, which allow you to, to check which part of the brain are active when you are doing different tasks as, you know, r- r- can replace some of this. You can see that certain parts of the brain are involved, are active in certain tasks-

    13. LF

      Vision, language. Yeah.

    14. TP

      Yeah.

    15. LF

      Yeah.

    16. TP

      That's right.

    17. LF

      But sort of taking a step back to that part of the brain that discovers, that, uh, specializes in the face, and how that might be learned, what's your intuition behind... you know, is it possible that the, sort of from a physicist's perspective, when you get lower and lower, that it's all the same stuff, and it just, when you're born, it's plastic, and, and it quickly figures out, "This part is gonna be about vision. This is gonna be about language. This is about common sense reasoning"? Do you have any intuition that that kind of learning is going on really quickly or is it really kind of solidified in hardware?

    18. TP

      That's a great question. Uh, so there are parts of the brain, like the cerebellum or the hippocampus, that are quite different from each other. They clearly have different anatomy, different connectivity. They're... then there is, uh, the, the cortex, which is the, the most developed part of the brain in humans. And, uh, in the cortex, you have different regions of the cortex that are responsible for vision, for audition, for motor control, for language. Now, one of the big puzzles of, of this is that in the, the cortex is the cortex, is the cortex, is, is looks like it is the same in terms of, um, hardware-

    19. LF

      Mm-hmm.

    20. TP

      ... in terms of type of neurons and connectivity across these different modalities. So, for the cortex, letting aside these other parts of the brain like spinal cord, hippocampus, cerebellum and so on, for the cortex, I think your question about hardware and software and, uh, learning and so on, it's, it's, I think is rather open. And, uh, you know, it... I find very interesting, for instance, to think about an architecture, computer architecture, that is good for vision and at the same time is good for language. Seems to be, you know, so different problem, uh, areas that you have to solve.

    21. LF

      But the underlying mechanism might be the same, and that's really instructive for-

    22. TP

      It may be.

    23. LF

      ... artificial neural networks.

    24. TP

      Right.

  8. 27:5532:51

    Vision as a gateway to intelligence—and why understanding brains needs many levels

    1. LF

      So, you've done a lot of great work in vision, in human vision, m- computer vision, and you mentioned the problem of human vision is... will be as difficult as the problem of general intelligence, and maybe that connects to the cortex discussion. Uh, can you describe the human visual cortex and how the humans begin to understand the world, uh, through their raw sensory information? What's, uh, for folks who are not f- familiar (laughs) , especially in, on the computer vision side, we don't often actually take a step back except saying, well, the sentence or two that one is inspired by the other. Wh- wha- what is it that we know about the human visual cortex that's interesting?

    2. TP

      So, we know quite a bit. At the same time, we don't know a lot. But the-

    3. LF

      (laughs) .

    4. TP

      ... the, the bit we know, you know, in a, uh, in a sense, we know a lot of the details, and, uh, um, and many we don't know, and, uh, we know a lot of the top level, um, the answer to top level question, but, uh, we don't know some basic ones, even in terms of general neuroscience, forgetting vision. You know, why do we sleep?

    5. LF

      (laughs) .

    6. TP

      (laughs) It's such a basic question, and we really don't have an answer to that. Uh-

    7. LF

      Do you think... so taking a step back-

    8. TP

      Yeah.

    9. LF

      ... on that, so sleep, for example, is fascinating. Do you think that's a neuroscience question? Or if we talk about abstractions, what do you think is an interesting way to study intelligence or most effective on the levels of abstraction? Is it chemical? Is it biological? Is it electrophysical, mathematical, as you've done a lot of excellent work on that side, which... psychology sort of like a... which level of abstraction do you think?

    10. TP

      Well, in terms of levels of, um, abstraction, I think we need all of them.

    11. LF

      All of them.

    12. TP

      It's one... uh, you know, it's like, uh, if you ask me, "What does it mean to understand a computer," right? And, and th- that's much simpler, but in a computer, I could say, "Well, I understand how to use PowerPoint."

    13. LF

      Right.

    14. TP

      ... that's my level of understanding a computer. It's, it tends reasonable, you know, it give me some power to produce slides, and beautiful slides, and ... Now, you can ask somebody else he says, "Well, I, I know how the transistor work that are inside the computer. I can write the equations for, you know, transistor, and diodes, and circuits, uh, logical circuits." And I can ask this guy, "Do you know how to operate PowerPoint?" "No idea." Right?

    15. LF

      Yeah. So, do you think if we discovered computers walking amongst us, full of these transistors that are also, uh, operating under Windows and have PowerPoint, do you think it's ... digging in a little bit more, how useful is it to understand the transistor in order to be able to understand PowerPoint and these higher level-

    16. TP

      Very good. Yes.

    17. LF

      ... intelligent processes.

    18. TP

      So, I think in the case of computers, because they were made by engineers, by us-

    19. LF

      Mm-hmm.

    20. TP

      ... this different level of understanding are rather separate on purpose.

    21. LF

      Mm-hmm.

    22. TP

      You know, you, they are separate modules so that the engineer that designed the circuit for the chips does not need to know what, uh, power, is inside PowerPoint.

    23. LF

      Mm-hmm.

    24. TP

      And somebody can write to the, the software translating from one to the end, uh, to the other end. So, um, in that case, I don't think, uh, uh, understanding the transistor help you understand PowerPoint, or very little.

    25. LF

      Right.

    26. TP

      Um, if you want to s- understand the computer, there is question, you know, I would say you have to understanding at different levels if you really-

    27. LF

      Yeah.

    28. TP

      ... want to s- to build it, one, right? (laughs) But, uh, but for the brain, I think these levels of understanding, so the algorithms, which kind of computation, you know, the equivalent of PowerPoint, and the circuits, you know, the transistors, I think they are more, much more intertwined with each other. There is not, you know, innately a level of the software separate from the hardware. And so, that's wha- why I think, in the case of the brain, the problem is more difficult and more than for computers, requires the interaction, the collaboration between different types of expertise.

    29. LF

      So, it's a big, the brain is a big hierarchical mess-

    30. TP

      Mm-hmm.

  9. 32:5135:47

    Compositionality: when deep networks beat shallow ones

    1. LF

      (laughs) That's a difficult one. That said, you do talk about compositionality-

    2. TP

      Mm-hmm.

    3. LF

      ... and why it might be useful, and when you discuss wha- why these neural networks, in artificial or biological sense, learn anything, you talk about c- uh, compositionality.

    4. TP

      Yeah.

    5. LF

      So, you, there's a sense that nature can be disentangled, or perpe- uh, well, all aspects of our cognition could be disentangled-

    6. TP

      Mm-hmm.

    7. LF

      ... a little, to some degree. So, why do you think, what, first of all, how do you see compositionality, and why do you think it exists at all in nature?

    8. TP

      I spoke about, uh, I used the, the term compositionality when we looked at deep neural networks, multi-layers, and trying to understand when and why they are more powerful than, uh, uh, more classical one-layer networks, like s- linear classifier, uh, kernel machines, so-called. Um, and what we found is that, in terms of approximating, or learning, or representing a function, a, a mapping from an input to an output, like from an image to the label in the image, um, if this function has a particular structure-

    9. LF

      Mm-hmm.

    10. TP

      ... then deep networks are much more powerful than shallow networks to approximate the underlying f- function. And the particular structure is a structure of compositionality. If the function is made up of functions of functions, so that you need to look on, y- when you are interpreting an image, classifying an image, you don't need to look at all pixels at once, but you can compute s- something from, uh, small groups of pixels, and then you can compute something on the output of this local computation, and so on.

    11. LF

      Yeah.

    12. TP

      That is similar to what you do when you read a sentence. You don't need to read the first and the last letter, but you can read syllables, combine them in words, combine the words in sentences. So, this is this kind of structure.

    13. LF

      So, that's as part of a discussion of why deep neural networks may be more effective than the shallow methods.

    14. TP

      Yeah.

    15. LF

      And is your sense for most things we can use neural networks for, th- those problems are going to be compositional in nature? Like, uh, like language, like vision? How far can we get in this kind of way?

  10. 35:4739:17

    Why compositionality exists: physics, brain wiring limits, and evolution

    1. TP

      Right. So, here is almost philosophy.

    2. LF

      Well, let's go there. (laughs)

    3. TP

      You know? (laughs) Yeah, let's go there. So, this friend of mine, Max Tegmark, who is a physicist-

    4. LF

      Yes.

    5. TP

      ... at MIT-

    6. LF

      I've talked to him on this thing, yeah, and he disagrees with you, right?

    7. TP

      Yeah. We-

    8. LF

      A little bit.

    9. TP

      ... you know, we agree on mo- most, but the conclusion is a bit different. He, he, his conclusion is that w-... for images, for instance, the compositional structure of this function that we have to learn or to solve these problems comes from physics, comes from the fact that you have local interactions in physics between m- atoms and other atoms, between particle of matter and other particles, between planets and other planets, between (laughs) stars and other ... It's all local.

    10. LF

      Yeah.

    11. TP

      Um, a- and that's true, um, but you could push this argument a bit farther. Um, not this argument actually. You could argue that, um, you know, maybe that's part of the truth, but maybe what happens is kind of the opposite, is that our brain is wired up as a deep network.

    12. LF

      Mm-hmm.

    13. TP

      So, it can learn, understand, solve problems that have this compositional structure.

    14. LF

      Mm-hmm.

    15. TP

      And it cannot do... it cannot solve problems that don't have this compositional structure. So, the problems we are accustomed to, we think about, we test our algorithms on, are this compositional structure because our brain is made up.

    16. LF

      And that's, in a sense, an evolutionary perspective-

    17. TP

      Yes.

    18. LF

      ... that we've ... so the, the ones that didn't have, uh, th- that weren't dealing with a compositional nature of reality, uh, d- died off?

    19. TP

      Yes-

    20. LF

      Were not able

    21. NA

      to survive?

    22. TP

      But also could be ... maybe the reason why we have this, uh, local connectivity in the brain, like, uh, simple cells in cortex looking only at a small part of the vi- image, each one of them, and then other cells looking at the small number of these simple cells, and so on. The reason for this may be purely that it was difficult to grow long-range connectivity.

    23. LF

      Mm.

    24. TP

      So, suppose it's ... you know, for biology. It's e- possible to grow short-range connectivity, but not long range also because there is a limited-

    25. LF

      Yes.

    26. TP

      ... number of long range that you c- And so, you have this, this limitation from the biology, and this means you build a deep convolutional netwo- this would be something like a deep convolutional network. And this is f- great for solving certain class of problem. These are the ones we are f- we find easy and important for our life. And yes, they were enough for us to survive. (laughs)

    27. LF

      Yeah. And, uh, y- and you can start a successful business on solving those problems.

    28. TP

      (laughs) Yeah.

    29. LF

      Uh, right? Uh, like with Mobileye. Uh, driving is, is, is a compositional problem.

    30. TP

      Right.

  11. 39:1744:47

    Stochastic gradient descent: why it works, and why biology might do something else

    1. LF

      So, on the, on the learning task, I mean, we don't know much about how the brain learns in terms of optimization. But, uh, so the thing that's stochastic gradient descent is what, uh, artificial neural networks use for the most part to, uh, a- adjust the parameters in such a way that it's able to deal, uh, based on the label data, it- it's able to solve the problem.

    2. TP

      Yeah.

    3. LF

      So, what's your intuition about, uh, why it works at all? How hard of a problem it is to optimize a neural network, a artificial neural network? Is there other alternatives? Yeah, just in general, your intuitions behind this very simplistic algorithm that seems to do pretty good, surprisingly so.

    4. TP

      Yes, s- yes. So, I find, um, neuroscience, th- the architecture of, uh, cortex is really similar to the architecture of deep networks. So, th- there is a nice correspondence there between the biology and this kind of local connectivity, hierarchical, um, architecture. The stochastic gradient descent, as you said, is, um, v- is a very simple technique. It seems pretty unlikely that biology could do that from, from what we know right now about t- n- n- you know, cortex and neurons and synapses. Um, so, uh, it's a big question open whether there are other optimization learning algorithms that can replace stochastic gradient descent. And, uh, my, my guess is yes, but nobody has found yet a real answer. I mean, people are trying, still trying, and there are some interesting ideas. The fact that, um, stochastic gradient descent is so successful, this has become clear. It's not so mysterious. And the reason is that, um, it's an interesting fact that, you know, is a change, in a sense, in, uh, how people think about statistics.

    5. LF

      (laughs)

    6. TP

      And, and this is the following, is that typically when you had, uh, data and you had, say, a model with parameters, you are trying to fit the model to the data, you know, to fit the parameter. Typically, the kind of, um, kind of, uh, crowd wisdom, uh, type idea was you should have at least, uh, you know, twice the number of data than the number of parameters, you have.

    7. LF

      Mm-hmm.

    8. TP

      Uh, maybe 10 times is better.Now, the way you train neural network these days is that they have, they have 10 or 100 times more parameters than data. Exactly the opposite. And which s- you know, i- i- it is, it has been one of the puzzles about neural networks. How can you get something that really works when you have so much freedom? In, uh, sin-

    9. LF

      From that little data, it can generalize somehow.

    10. TP

      Yeah, right, exactly.

    11. LF

      Do you think the s- the stochastic nature of it is essential to randomness?

    12. TP

      So, I think we have some initial understanding why this happens, but, um, one nice side effect of having this over-parameterization, more parameters than data, is that when you look for the minimum of a loss function, like stochastic gradi- gradient descent is doing, um, you find... I, I, I made some calculations based on some old, uh, basic theorem of algebra called Bézout's theorem, and that gives you an, a, an estimate of the number of solution of a system of polynomial equation. Anyway, the bottom line is that there are probably more minima for a typical deep networks, um, than atoms in the universe. Just to say, there are a lot. (laughs)

    13. LF

      (laughs)

    14. TP

      Because of the over-parameterization.

    15. LF

      Yes.

    16. TP

      A more global minimum, zero minimum, good minimum. So, it's not too sur-

    17. LF

      More global minima?

    18. TP

      Yeah.

    19. LF

      Oh, okay.

    20. TP

      A lot of them. So we have a lot of solutions, so it's not-

    21. LF

      Okay.

    22. TP

      ... so surprising that you can find them relatively easily. And this is-

    23. LF

      Oh, I see.

    24. TP

      This is because of the over-parameterization.

    25. LF

      The over par- parameterization sprinkles that entire space with solutions that're pretty good.

    26. TP

      Yes, yeah.

    27. LF

      And so you-

    28. TP

      It's not so surprising, right? It's like, you know, if you have a system of linear equation, and you have more unknowns than equations, then you have... We know you have an infinite number of solutions, and the question is to pick one. That's another story. But you have an infinite number of solutions, so there are a lot of, of value of your unknowns that satisfy the equations.

    29. LF

      But it's possible that there's a lot of s- those solutions that aren't very good. Well, what's surprising is they're pretty good.

    30. TP

      Right. So that's separate question. Why can you pick one that generalizes well?

  12. 44:4747:50

    Universal approximation, the curse of dimensionality, and how depth can avoid it

    1. LF

      One, one, uh, theorem that people like to talk about that kind of inspires imagination of the power of neural networks is the universality, uh, universal approximation theorem, that you can approximate any computable function with just a finite number of neurons in a single hidden layer. Do you, do you find this theorem, one, surprising? Do you find it useful, interesting, inspiring?

    2. TP

      No, they, this one, uh, you know, I never found it very surprising. It's, uh, was known since the '80s, since I entered the field, because it's basically the same as Weierstrass theorem, which says that I can approximate any continuous function with a polynomial of sufficiently, with a sufficient number of terms, monomials.

    3. LF

      Mm-hmm, yeah.

    4. TP

      So, basically the same, and the proofs are very similar.

    5. LF

      So y- your intuition was, there was never any doubt that neural-

    6. TP

      Yeah.

    7. LF

      ... networks, in theory, could-

    8. TP

      Right.

    9. LF

      ... could be very strong approximators even?

    10. TP

      Right. The, the, the question, the interesting question is that if this theorem, uh, says you can approximate, fine, but when you ask how many neurons, for instance, or in the case of polynomial, how many monomials I need to get a good approximation, then it turns out that that depends on the dimensionality of your function, how many variables you have. But it depends on the dimensionality of your function in a bad way. It's, for instance, suppose you want an error which is, uh, no worse than 10% in your approximation. You come up with a network that approximate your function within 10%. Then turns out that the number of units you need are in the order of 10 to the dimensionality, D.

    11. LF

      Mm-hmm.

    12. TP

      How many variables. So if you have, you know, two variables is these two, and you have 100 units, and okay. But if you have, say, 200 by 200 pixel images, now this is, you know, tw- 40,000, whatever, and that-

    13. LF

      We again go to the size of the universe pretty quickly.

    14. TP

      Yeah, ra- exact. 10 to the 40,000 or something, and (laughs) -

    15. LF

      Yeah. (laughs)

    16. TP

      And so, uh, this is called the curse of dimensionality.

    17. LF

      Yeah.

    18. TP

      Not, you know, quite appropriately. (laughs)

    19. LF

      (laughs) And the hope is with the extra layers, you can, uh, uh, remove the curse.

    20. TP

      What we proved is that if you have deep layers or hierarchical architecture of the, with the local connectivity of the type of convolutional deep learning, and if you're dealing with a function that has this kind of, um, hierarchical architecture, then you avoid completely the curse.

  13. 47:5051:13

    Unsupervised learning and GANs: impressive outputs vs the real ‘N→1’ challenge

    1. LF

      (laughs) You've spoken a lot about supervised deep learning.

    2. TP

      Yeah.

    3. LF

      Uh, what are your thoughts, hopes, views on the challenges of unsupervised learning, uh, with, uh, with GANs, with, uh, generative adversarial networks? Uh, do you see those as distinct... The, the power of GANs, do you see those as distinct from supervised methods in neural networks, or are they really all in the same representation ballpark?

    4. TP

      GANs is, uh, one way to get, um-... estimation of, uh, uh, probability densities, which is a somewhat new way that people have not done before. I, I don't know whether, um, this will really play an important role in, uh, you know, in intelligence or, um, it's, it's interesting. I'm, I'm less enthusiastic about it than many people in the field.

    5. LF

      Mm-hmm.

    6. TP

      I have the feeling that many people in the field are, um, really impressed by the ability to p- of producing realistic-looking images in a, in this generative way.

    7. LF

      Which describes the popularity of the methods, but you're saying that while that's exciting and cool to look at, it may not be the tool that's useful for-

    8. TP

      Yeah.

    9. LF

      ... for... So, you described it kind of beautifully. Uh, current supervised methods go N to infinity in terms of the number of labeled points, and we really have to figure out how to go to N to one.

    10. TP

      Yeah.

    11. LF

      And you're thinking GANs might help, but they might not be the right-

    12. TP

      I, I don't think in, for that problem, which I really think is important. I think they may help, uh, they certainly have applications, for instance, in computer graphics and-

    13. LF

      Right.

    14. TP

      ... you know, w- I did work long ago, which was a little bit similar in terms of, uh, saying, "Okay, I have an, a network, and, uh, I present images, and, uh, I can..." Uh, so the input is images and the output is, for instance, the pose of the image.

    15. LF

      Mm-hmm.

    16. TP

      You know, a face, how much it's smiling, if it's rotated 45 degrees or not.

    17. LF

      Mm-hmm.

    18. TP

      Uh, what about having a network that I train with the same data set, but now I invert input and output? Now, the input is the pose or the expression, a number, a set of numbers, and the output is the image, and I train it. And we did pretty good, interesting results in terms of producing very realistic-looking images. It was, um, you know-

    19. LF

      That is-

    20. TP

      ... a less sophisticated mechanism, but the, uh, output was pretty, uh, less than GANs, but the output was pretty much of the same quality. So, I think for computer graphics-type application, yeah, f- definitely GANs can be quite useful, and not only for that. For... But, um, for, you know, helping, for instance, uh, on this problem of unsupervised example of reducing the number of labeled examples, um, I think people, uh, it's like they think they can get out more than they put in. You know, it- (laughs)

    21. LF

      There's no free lunch, as-

    22. TP

      Yeah.

    23. LF

      ... you said.

    24. TP

      Right.

  14. 51:1355:05

    How babies learn: bootstrapping weak priors using motion and segmentation

    1. LF

      Uh, so what do you think... What's your intuition, um, of how can we slow the growth of N to infinity in supervise-

    2. TP

      Oh-

    3. LF

      L, uh, uh, N to infinity-

    4. TP

      Oh.

    5. LF

      ... in supervised learning? So, for example, uh, Mobileye, uh, has very successfully, I mean, essentially annotated large amounts of data to be able to drive a car. Now, one, one thought is, so we're trying to teach machines, s- s- school of AI.

    6. TP

      Yeah.

    7. LF

      And w- and we're trying to... So, how can we become better teachers maybe? That's one, one way.

    8. TP

      No, you gotta, you gotta... You know, (laughs) one, I, I like that because when... Um, again, one caricature of, uh, the history of computer science, you could say, i- is, is it begins with programmers-

    9. LF

      Mm-hmm.

    10. TP

      ... expensive.

    11. LF

      Yeah.

    12. TP

      Continuous labelers, cheap.

    13. LF

      Yeah.

    14. TP

      And the future will be schools-

    15. LF

      (laughs)

    16. TP

      ... like we have for kids.

    17. LF

      Yeah. Currently, the labeling methods, we're not selective about which examples we, we teach networks with. So, so I think the focus of making one sh- n- networks that learn much faster is often on the architecture side, but how can we pick better examples with which to learn? Uh-

    18. TP

      Yeah.

    19. LF

      Do you have intuitions about that?

    20. TP

      Well, that's part of the que- uh, the part of the problem, but the other one is, um, you know, if we look at, um, biology, uh, um, a reasonable assumption, I think, is, um, in the same spirit that I said evolution is opportunistic and has weak priors. You know, the way I, I, I think, uh, the intelligence of a child, a baby may develop is, um, by bootstrapping-

    21. LF

      Mm-hmm.

    22. TP

      ... weak priors from evolution. For instance, um, in a... You can assume that you have in most organisms, including human babies, built in some basic, uh, machinery to, uh, detect motion and relative motion. S- and the fact there is, you know, we know all insects from fruit flies to other animals, they have this. Uh, even in the retinas of... In the very peripheral part of... It's very conserved across species, something that evolution discovered early. It may be the reason why babies tend to look in the first few days to moving objects and not to not-moving objects. Now, moving objects means, okay, they're attracted by motion, but motion also means that... Uh, motion gives automatic segmentation from the background.

    23. LF

      Mm-hmm.

    24. TP

      So, because of motion boundaries, you know, either the object is moving-

    25. LF

      Mm-hmm.

    26. TP

      ... or the eye of the baby is tracking the moving object, and the background is moving, right?

    27. LF

      Yeah. So, just on, purely on the visual characteristics of the scene, that seems to be the most useful.

    28. TP

      Right. So, it's like looking at an, at an, a num- an object without background.

    29. LF

      Without background.

    30. TP

      It's ideal for learning the object. Otherwise, it's really difficult, because you have so much stuff. So, suppose you do this at the beginning, first, uh, weeks. Then, after that, you can recognize object. Now they are imprinted, the number of obj-, even in the background, even without motion.

  15. 55:051:02:41

    Limits of today’s AI: scene understanding, existential risk, and AGI timelines

    1. LF

      So, that's at the... By the way, I just w- wanna ask on a object recognition problem. So, there is this being responsive to movement and doing edge detection, essentially. What's the gap between being effectively, effective at visually recognizing stuff, detecting where it is, and understanding the scene? Be- is this a huge gap and many layers, or is it, are we c- is it close?

    2. TP

      No, I think that's a huge gap. I think present algorithm, with all the success that we have, and the fact that are a lot of very useful... It's, I think we are, we are in a golden age for applications of, um, l- low-level vision and low-level speech recognition and so on. You know, Alexa and so on. Um, there are many more things of similar level to be done, including medical diagnosis and so on, but we are far from what we call understanding of a scene, of language, of actions, of people. Uh, that is, despite the claims, uh, that's, I think, are very far.

    3. LF

      We're a little bit off. So, in popular culture and among many researchers, uh, some of which I've spoken with, uh, the Stuart Russell and Elon Musk, uh, uh, in and out of the AI field, uh, there's a concern about the existential threat of AI.

    4. TP

      Yeah.

    5. LF

      (laughs) And how do you think about this concern in the j- th- th- and is it valuable to think about large-scale, long-term, unintended consequences of intelligence systems we try to build?

    6. TP

      I always think it's better to worry first, you know, early rather than late.

    7. LF

      (laughs)

    8. TP

      (laughs) So, uh, it's not-

    9. LF

      So, worry is good?

    10. TP

      Yeah.

    11. LF

      Yeah.

    12. TP

      I'm not against worrying at all.

    13. LF

      Yeah.

    14. TP

      Uh, personally, I think that, um, uh, uh, you know, it will take a long time before there is real reason to be worried. But as I said, I think it's, it's good to put in place and think about possible safety against, uh, uh... What I find a bit misleading are things like, um, that have been said by people I know, like Elon Musk and, uh, what is, Bostrom in particular.

    15. LF

      Yeah, yeah.

    16. TP

      U- what is his first name? I-

    17. LF

      Uh, Nick Bostrom.

    18. TP

      Nick Bostrom, right. Um, you know, and a couple of other people that, for instance, um, AI is more dangerous than nuclear weapons.

    19. LF

      Right.

    20. TP

      Yeah, I think that's really wrong, and (laughs) that can be... it's misleading, right? Because i- in terms of priority, we sh- should still be more worried about nuclear weapons and, uh, you know, what people are doing about it, and so on than AI.

    21. LF

      And, uh, y- you've s- spoken about Demis Hassabis and yourself saying, uh, that you think it'll be about 100 years out before we have s- a general intelligence system that's on par with the human being. Do you have any updates for those predictions?

    22. TP

      Well, I think he said that-

    23. LF

      He said 20, I think you said.

    24. TP

      He said 20, right. This was a couple of years ago. I have not asked him again, so should have.

    25. LF

      Your own prediction, w- what's your prediction about when you'll be truly surprised, uh, and what's the confidence interval on that?

    26. TP

      (laughs) Uh, you know, it's so difficult to predict the future and even the present sometimes. (laughs)

    27. LF

      It's pretty hard to predict, yeah.

    28. TP

      Right. But I would be... But as I said, I, this is completely-

    29. LF

      It's-

    30. TP

      ... is, I would be more like, um, uh, Rod Brooks.

  16. 1:02:411:20:26

    Ethics, consciousness, mortality—and closing advice on science and mentoring

    1. LF

      So, uh, unlearning, for systems to be part of our world, it has a certain ... One of the challenging things that you've spoken about is learning ethics, learning-

    2. TP

      Yeah.

    3. LF

      ... morals. And what, w- how hard do you think is the problem of, first of all, humans understanding our ethics? What is the origin on a neural and low level of ethics? What is it at the higher level? Is it something that's learnable for machines, in your intuition?

    4. TP

      I think, uh, yeah, ethics is learnable, very likely. Um, I, I think I, it's one of these problems where ... I think understanding the neuroscience of ethics. You know, people discuss, there is an ethics of neuroscience.

    5. LF

      (laughs)

    6. TP

      (laughs)

    7. LF

      Yeah, yes.

    8. TP

      You know, how a neuroscientist should or should not behave. Uh-

    9. LF

      Yeah.

    10. TP

      Can you think of a neurosurgeon and the ethics that he r- he really has to obey, or she has to obey. But I'm more interested on the, on the-

    11. LF

      (laughs)

    12. TP

      ... neuroscience of ethics.

    13. LF

      You're blowing my mind right now. The neuroscience of ethics is very meta.

    14. TP

      Yeah. And, uh, you know, I think that will be important to understand also for being able to, to design machines that have, that are ethical machines, in our sense of ethics.

    15. LF

      And you think there is something in neuroscience, there's patterns, tools in neuroscience that could help us, uh, shed some light on ethics? Or-

    16. TP

      Yeah.

    17. LF

      ... is it mostly on the psychologist or sociology, a much higher level?

    18. TP

      No, there is psychology, but there is also, in the meantime, there are, um, there is evidence, fMRI, of, eh, specific areas of the brain that are involved in certain ethical judgment. And not only this, you can stimulate those area with magnetic fields and change the ethical decisions.

    19. LF

      Yeah.

    20. TP

      Okay.

    21. LF

      Wow.

    22. TP

      So, that's the work by a colleague of mine, Rebecca Sax.

    23. LF

      Yeah.

    24. TP

      And there is, uh, other s- researchers doing similar work. And I think, you know, this is the beginning that, um, ideally, at some point, we'll have an understanding of how this works and why it evolved, right? Um ...

    25. LF

      The big why question. Yeah, it must have some, some purpose.

    26. TP

      Yeah. Obviously, it has, you know, some social purposes, is, uh ... probably.

    27. LF

      If neuroscience holds the key to at least illuminate some aspect of ethics, that means it could be a learnable problem.

    28. TP

      Yep, exactly.

    29. LF

      And as we're getting into harder and harder questions, let's go, uh-

    30. TP

      (laughs)

Episode duration: 1:20:20

Install uListen for AI-powered chat & search across the full episode — Get Full Transcript

Transcript of episode aSyZvBrPAyk

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.