Lex Fridman PodcastJeremy Howard: fast.ai Deep Learning Courses and Research | Lex Fridman Podcast #35
EVERY SPOKEN WORD
150 min read · 30,011 words- 0:00 – 1:21
Meet Jeremy Howard and the fast.ai ethos
- LFLex Fridman
The following is a conversation with Jeremy Howard. He's the founder of fast.ai, a research institute dedicated to making deep learning more accessible. He's also a distinguished research scientist at the University of San Francisco, a former president of Kaggle, as well as a top-ranking competitor there. And in general, he's a successful entrepreneur, educator, researcher and an inspiring personality in the AI community. When someone asks me, "How do I get started with deep learning?" fast.ai is one of the top places I point them to. It's free. It's easy to get started. It's insightful and accessible. And if I may say so, it has very little BS that can sometimes dilute the value of educational content on popular topics like deep learning. Fast.AI has a focus on practical application of deep learning and hands-on exploration of the cutting edge that is incredibly both accessible to beginners and useful to experts. This is the Artificial Intelligence podcast. If you enjoy it, subscribe on YouTube, give it five stars on iTunes, support it on Patreon, or simply connect with me on Twitter, @lexfridman, spelled F-R-I-D-M-A-N. And now here's my conversation with Jeremy Howard. What's the first program you've ever written?
- 1:21 – 3:01
First code: using a Commodore 64 to search for better musical scales
- JHJeremy Howard
First program I wrote, that I remember, would be at high school. Um, (sighs) I did an assignment where I decided to try to find out if there were some, like, better musical scales than the normal 12 tone, 12 s- interval scale. So I wrote a program on my Commodore 64 in Basic that searched through other scale sizes to see if it could find one where there were, uh, more accurate, you know, uh, harmonies.
- LFLex Fridman
Like mid-tone, like finding-
- JHJeremy Howard
Like, like you want an actual exactly three to two ratio, whereas with a 12 interval scale it's not exactly three to two, for example. So that's-
- LFLex Fridman
And, and the Common-
- JHJeremy Howard
... well tempered as they say in the... (laughs)
- LFLex Fridman
In Basic on a Commodore 64?
- JHJeremy Howard
Yeah.
- LFLex Fridman
Where was the interest in music from? Or is it just technical?
- JHJeremy Howard
I did music all my life, so I played saxophone and clarinet and piano and guitar and drums and whatever, so...
- LFLex Fridman
How does that thread go through your life? Where's music today? Is it-
- JHJeremy Howard
Uh, it's not where I wish it was. I, for various reasons, couldn't really keep it going, particularly because I had a lot of problems with RSI with my fingers, and so I had to kind of like cut back anything that used hands and fingers.
- LFLex Fridman
Mm-hmm.
- JHJeremy Howard
Um, I hope one day I'll be able to get back to it healthwise.
- LFLex Fridman
So there's a love for music underlying it all?
- JHJeremy Howard
For sure, yeah.
- LFLex Fridman
What's your favorite instrument?
- JHJeremy Howard
Uh, saxophone.
- LFLex Fridman
Sax.
- JHJeremy Howard
Baritone saxophone. Well, probably bass saxophone but they're awkward.
- LFLex Fridman
Well, um, I always love it when, uh, music is coupled with programming.
- JHJeremy Howard
Mm-hmm.
- 3:01 – 8:23
Programming environments that make data feel easy: Access, Excel, and relational thinking
- LFLex Fridman
There's something about a brain that utilizes those, that, uh, emerges with creative ideas. So you've used and studied quite a few programming languages.
- JHJeremy Howard
Mm-hmm.
- LFLex Fridman
Can you give an, an overview of what you've used? What are the pros and cons of each?
- JHJeremy Howard
Uh, my favorite programming environment almost certainly was Microsoft Access, back in like the earliest days. So that w- that was-
- LFLex Fridman
That's like-
- JHJeremy Howard
... Visual Basic for Applications which is not a good programming language, but the programming environment was fantastic. It's like, the ability to create, you know, user interfaces and tie data and actions to them and create reports and all that is... I've never seen anything as good. There's things nowadays like Airtable which are like s- small subsets of that which people love for good reason, but unfortunately nobody's ever, uh, achieved anything like that.
- LFLex Fridman
What is that? If, if you could pause on that for a second.
- JHJeremy Howard
Oh, Access?
- LFLex Fridman
Access. Is it, uh, it's fundamentally a database.
- JHJeremy Howard
Yeah, so it was a, it was a database program that Microsoft produced, uh, part of Office and it kind of withered, you know. But basically it lets you, in a totally graphical way, create tables and relationships and queries and tie them to forms and set up, you know, event handlers and, uh, calculations. And it was a very complete, powerful system designed for not m- massive, scalable things but for, like, useful little applications that I loved
- NANarrator
quickly building.
- LFLex Fridman
So what's the connection between Excel and Access?
- JHJeremy Howard
So very close. So A- Access kind of was the relational database equivalent, if you like. So people still do a lot of that stuff that should be in Access in Excel-
- LFLex Fridman
Excel.
- JHJeremy Howard
... because they know it. Excel's great as well, so, um, but it's just not as rich a programming model as VBA combined with a relational database. I've... And so I've always loved relational databases, but today programming on top of a relational database is just a lot more of a headache. You know, you generally either need to kind of y- y- you know, you need something that connects, that, that runs some kind of database server unless you use SQLite which has its own issues. Then you kind of... Often if you want to get a nice programming model you'll need to like create an... Add an ORM on top and then... I don't know, there's all these pieces-
- LFLex Fridman
Right.
- JHJeremy Howard
... to tie together and it's just a lot more work than it should be. There are people that are trying to make it easier so that... In particular I think of F Sharp, you know, Don Syme who, um, him and his team have done a great job of, um, making, uh, something like a database appear in the type system so you actually get like tab completion for fields and tables and stuff like that. Anyway, so that was kind of... Anyway, so like that whole VBA Office thing I guess was a starting point which I still miss and-... I got into standard Visual Basic, which-
- LFLex Fridman
Well, that's interesting, just to pause on that for a second.
- JHJeremy Howard
Mm-hmm.
- LFLex Fridman
It's interesting that you're connecting programming languages to, um, the ease of management of data.
- JHJeremy Howard
Yeah.
- LFLex Fridman
So in your use of programming languages, you always w- had a love and a connection with data.
- JHJeremy Howard
I've always been interested in doing useful things for myself-
- LFLex Fridman
(laughs)
- JHJeremy Howard
... and for others, which generally means, uh, getting some data and doing something with it and putting it out there again. So that's been my interest throughout. So I also did a lot of stuff with AppleScript back in the early days. Um, so it was kind of nice being able to get the computer and t- computers to talk to each other and to do things for you. And then I think that one I, the programming language I most loved then would have been Delphi, which was Object Pascal, created by Anders Hejlsberg who previously did Turbo Pascal and then went on to create .NET and then went on to create TypeScript. Delphi was amazing 'cause it was like a compiled, fast language that was as easy to use as Visual Basic.
- LFLex Fridman
Delphi, what is it similar to in, in more modern languages?
- JHJeremy Howard
Um, Visual Basic.
- LFLex Fridman
Visual Basic?
- JHJeremy Howard
Yeah. That, a compiled fast version. So I'm not sure there's anything quite like it anymore. If you took, like, C# or Java and got rid of the virtual machine and replaced it with something you could compile a small type binary. I feel like it's where, um, Swift could get to with the new SwiftUI and the cross-platform development going on.
- LFLex Fridman
Mm-hmm.
- 8:23 – 12:59
The “third path” in languages: APL lineage, J, and array-oriented programming
- JHJeremy Howard
... said Pascal's not a nice language. If you wanted to know specifically about what languages I liked, I would definitely pick J as being an amazingly wonderful language.
- LFLex Fridman
Wha- (laughs) What's J?
- JHJeremy Howard
Uh, J, uh, are you aware of APL?
- LFLex Fridman
I am not.
- JHJeremy Howard
Okay, so-
- LFLex Fridman
Except from doing a little research on, uh, work you've done.
- JHJeremy Howard
Okay, so not at all surprising you're not familiar with it, 'cause it's not well known, but it's actually one of the main, uh, families of programming languages going back to the late '50s, early '60s. So there was a couple of major directions. One was the kind of lambda calculus Alonzo Church direction, which I guess kind of LISP and Scheme and whatever, um, which has a history going back to the early days of computing. The second was the kind of imperative/OO, you know, ALGOL, Simula, going on to C, C++, so forth. There was a third, which are called, uh, array-oriented languages, um, which started with a, um, paper by a guy called Ken Iverson, uh, which was actually a math theory paper, not a programming paper. Uh, it was, uh, called Notation as a Tool for Thought.
- LFLex Fridman
(laughs)
- JHJeremy Howard
And it was the development of a new way, a new type of math notation. And the idea is that this math notation would be, was, was much more flexible, expressive, and also well-defined than traditional math notation, which is none (laughs) of those things. Math notation is awful. Um, and so he actually turned that into a programming language, and 'cause this was the early '50s, all the... uh, sorry, late '50s, all the names were available, so he called his language A Programming Language or APL.
- LFLex Fridman
APL, wow. (laughs)
- JHJeremy Howard
So APL is a implementation of Notation as a Tool for Thought, by which he means math notation. And Ken and his son went on to do many things, but eventually they actually produced a, a, you know, a new language that was built on top of all the learnings of APL and that was called J.
- LFLex Fridman
Hmm.
- JHJeremy Howard
Um, and J is the most, uh, expressive, composable, um, language of, you know, beautifully designed language I've ever seen.
- LFLex Fridman
Does it have object-oriented components? Does it have that kind of thing or
- NANarrator
more like-
- JHJeremy Howard
Not really. It's an array-oriented language. It's a new, it's a, it's a, it's, it's the third path, so-
- LFLex Fridman
Can you... Are you saying array?
- JHJeremy Howard
Array-oriented, yeah, so it's-
- LFLex Fridman
What does it mean to be array-oriented?
- JHJeremy Howard
So array-oriented means that you generally don't use any loops, but the whole thing is done with kind of a, a, um, extreme version of broadcasting, if you're familiar with that.
- LFLex Fridman
Mm-hmm.
- JHJeremy Howard
NUM- NUMPy/Python concept. So, uh, you, you do a lot with one line of code. It, it looks a lot like math notation.
- LFLex Fridman
So it's basically a-
- JHJeremy Howard
Highly compact.
- LFLex Fridman
Mm-hmm.
- JHJeremy Howard
And the idea is that you can kind of, because you can do so much with one line of code, a single screen of code is very unlikely to... you very rarely need more than that to express your program.
- LFLex Fridman
Yeah.
- JHJeremy Howard
And so you can kind of keep it all in your head, and you can kind of clearly communicate it. It's interesting, there, there, uh, APL created two main branches, uh, K and J. Um, J is this kind of, like, open source niche community of, of crazy enthusiasts like me. And then the other path, K, was fascinating. It's an astonishingly expensive programming language which many of the world's most ludicrously rich hedge funds use.
- LFLex Fridman
Ah.
- JHJeremy Howard
So the entire K, uh, machine is so small it sits inside level 3 cache on your CPU, and it, and it s- easily wins every benchmark.... I've ever seen in terms of data processing speed. But you don't come across it very much, 'cause it's like $100,000 per CPU to, to run it.
- 12:59 – 15:01
Pragmatism wins: Perl, Python, and why libraries dictate reality
- LFLex Fridman
At the same time, you've done a lot of stuff with Perl and Python.
- JHJeremy Howard
Yeah.
- LFLex Fridman
So where does that fit into the picture of, uh, J and K and APL and Delphi?
- JHJeremy Howard
Well, you know, it's just much more pragmatic. Like, in the end, you kind of have to end up where the, where the libraries are, you know? Like, 'cause for me, my, my focus is on productivity, I just want to get stuff done and solve problems, so... Um, Perl was great for... I created an email company called Fastmail and Perl was great 'cause back in the late '90s, early 2000s, um, it just ha- had a lot of stuff it could do. I still had to write my own monitoring system and my own web framework, my own whatever, 'cause like none of that stuff existed. But it was a super flexible language to do that in.
- LFLex Fridman
And you used Perl... For Fastmail, you used it as a backend? Like, so everything was written in Perl?
- JHJeremy Howard
Yeah. Yeah, everything.
- LFLex Fridman
Wow.
- JHJeremy Howard
Everything was Perl.
- LFLex Fridman
Why do you think Perl hasn't succeeded or hasn't dominated the market where Python really takes over a lot of the same tasks as Perl?
- JHJeremy Howard
Yeah. Well, I mean, it, Perl did dominate. It was-
- LFLex Fridman
For a time, yeah.
- JHJeremy Howard
... everything everywhere. But then the, the guy that ran Perl, Larry Wall, kind of just didn't put the time in anymore.
- LFLex Fridman
Yeah.
- JHJeremy Howard
And no project can be successful if there isn't, you know, a s- uh, particularly one that started with a strong leader, that, that loses that strong leadership. So then Python has kind of replaced it, you know. Python is, um, a lot less elegant language in nearly every way but it has the data science libraries, and a lot of them are pretty great. So I kind of use it 'cause it's the best we have.
- LFLex Fridman
(laughs)
- JHJeremy Howard
But it's definitely not good enough.
- 15:01 – 21:42
Making deep learning truly hackable: Swift, MLIR, and DSLs for GPU kernels
- LFLex Fridman
But what do you think the future of programming looks like? What do you hope the future of programming looks like if we zoom in on the computational fields, on data science, on machine learning?
- JHJeremy Howard
I, I hope Swift is successful because the, the goal of Swift, the way Chris Lattner describes it, is to be infinitely hackable, and that's what I want. I want something where, um, me and the people I do research with, and my students can look at and change everything from top to bottom. There's nothing mysterious and magical and inaccessible.
- LFLex Fridman
Mm-hmm.
- JHJeremy Howard
Unfortunately with Python, it's the opposite of that because Python's so slow. It's, um, extremely unhackable. You get to a point where it's like, okay, from here on down, it's C. So your debugger doesn't work in the same way, your profiler doesn't work in the same way, your build system doesn't work in the same way, it's really not very hackable at all.
- LFLex Fridman
What's the, what's the part you would like to be hackable? Is it for the objective of optimizing training of neural networks, inference of neural networks? Is it performance of the system? Or is there some non-performance related just creative idea?
- JHJeremy Howard
It's, it's, it's everything. I mean, in the end, I want to be productive as a practitioner. So that means that, uh, so like at the moment, our understanding of deep learning is incredibly primitive. There's very little we understand, most things don't work very well even though it works better than anything else out there.
- LFLex Fridman
Right.
- JHJeremy Howard
There's so many opportunities to make it better. So you look at any domain area, like, I don't know, speech recognition with deep learning or natural language processing classification with deep learning or whatever. Every time I look at an area with deep learning, I always see like, oh, it's, it's terrible. There's lots and lots of obviously stupid ways to do things that need to be fixed. So then I want to be able to jump in there and quickly-
- LFLex Fridman
And, and-
- JHJeremy Howard
... experiment and make them better.
- LFLex Fridman
... you think the programming language is, um, has a role in that?
- JHJeremy Howard
Huge role. Yeah. So currently, Python, um, has a big, uh, gap in terms of our ability to, um, innovate particularly around recurrent neural networks and, um, natural language processing because, uh, 'cause it's so slow. The, the, the actual loop where we actually loop through words, we have to do that whole thing in CUDA C.
- LFLex Fridman
Mm-hmm.
- JHJeremy Howard
So we actually can't innovate with the, the kernel, the heart of that most important algorithm.
- LFLex Fridman
Mm-hmm.
- JHJeremy Howard
Um, and it's just a huge problem. And this happens all over the place. So we hit, you know, research limitations. Another example, convolutional neural networks which are actually the most popular architecture for lots of things, maybe most things in deep learning, we almost certainly should be using sparse convolutional neural networks.
- LFLex Fridman
Mm-hmm.
- JHJeremy Howard
Um, but only like two people are, because to do it, you have to rewrite all of that CUDA C level stuff. And yeah, just researchers and practitioners don't. So like there's just big gaps in like s- what people actually research on, what people actually implement, because of the programming language problem.
- LFLex Fridman
So you think, uh...... do you think it's, it's just too difficult to write in CUDA C, uh, that a programming lang- a higher level programming language like Swift should enable the, the easier imp- th- fooling around creative stuff with RNNs or with sparse convolutional networks?
- JHJeremy Howard
Kind of.
- LFLex Fridman
Who, who, who's, uh, who's at fault? Who's e- who's in charge of making it easy for a researcher to play around?
- JHJeremy Howard
I mean, no one's at fault.
- LFLex Fridman
(laughs)
- JHJeremy Howard
It's just nobody's got around to it yet, or-
- LFLex Fridman
Yeah.
- JHJeremy Howard
... it's just, it's hard, right? And, I mean, part, part of the fault is that we ignored that whole APL kind of direction most... well, nearly everybody did for 60 years, 50 years. But, uh, recently, people have been starting to reinvent pieces of that and kind of create some interesting new directions in the compiler technology. So the place, um, where that's particularly happening right now is, uh, something called MLIR, which is something that, again, Chris Lattner, the Swift guy, is leading. And, uh, yeah, 'cause it's actually not going to be Swift on its own that solves this problem.
- LFLex Fridman
Mm-hmm.
- JHJeremy Howard
Because the problem is that currently writing a acceptably fast, you know, GPU program is too complicated regardless of what language you use.
- LFLex Fridman
Right.
- JHJeremy Howard
And that's just because if you have to deal with the fact that I've got, you know, 10,000 threads and I have to synchronize between them all, and I have to put my thing into grid blocks and think about warps and all this stuff, it's just, it's just so much boilerplate that to do that well, you have to be a specialist at that, and it's going to be a year's work to, you know, optimize that algorithm in that way. But, uh, with things like tensor comprehensions and Tile and MLIR and TVM, there's all these various projects which are all about saying, "Let's let people create, like, domain-specific languages for tensor computations." These are the kinds of things we do, uh, generally in, on the GPU for deep learning, and then have a compiler which can optimize that tensor computation. Um, a lot of this work is actually sitting on top of, uh, a project called Halide, which, uh, was, uh, is a mind-blowing project where they came up with such a domain-specific language. In fact, two. One domain-specific language for expressing "this is what my tensor computation is," and another domain-specific language for expressing "this is the, kind of the way I want you to structure the compilation of that," like, do it block by block and do these bits in parallel.
- 21:42 – 23:34
Hardware reality: NVIDIA lock-in, TPUs, and the need for competition
- LFLex Fridman
Now, does it all eventually boil down to CUDA and NVIDIA GPUs?
- JHJeremy Howard
Unfortunately, at the moment, it does. But one of the nice things about MLIR, if AMD ever gets their act together, which they probably won't, is that they or others could write MLIR backends for other GPUs or other, or other tensor computation devices, of which today there are increasing number like Graphcore or Vertex AI or whatever.
- LFLex Fridman
Mm-hmm.
- JHJeremy Howard
So, yeah, being able to target lots of backends would be another benefit of this. And the market really needs competition because, at the moment, NVIDIA is massively overcharging for their-
- LFLex Fridman
Right.
- JHJeremy Howard
... kind of enterprise class cards because there is no serious competition because nobody else is doing the software properly.
- LFLex Fridman
In the cloud, there is some competition, right? But, um-
- JHJeremy Howard
Not really, other than TPUs perhaps.
- LFLex Fridman
TPUs, yeah.
- JHJeremy Howard
But TPUs are almost unprogrammable at the moment. Um-
- LFLex Fridman
So you can't... The TPUs has the same problem that you can't-
- JHJeremy Howard
It's even worse. So TPUs, they, Google actually made an explicit decision to make them almost entirely unprogrammable because they felt that there was too much IP in there, and if they gave people direct access to program them, people would learn their secrets.
- LFLex Fridman
Yeah.
- JHJeremy Howard
So you can't actually directly program the memory-
- LFLex Fridman
Okay.
- JHJeremy Howard
... in a TPU. You can't even directly, like, create code that runs on and that you look at-
- LFLex Fridman
Mm-hmm.
- JHJeremy Howard
... on the machine that has the GPU. It all goes through a virtual machine. So all you can really do is this kind of cookie-cutter thing of, like, plug in high-level stuff together, which is just super tedious and annoying-
- LFLex Fridman
(laughs)
- JHJeremy Howard
... and totally unnecessary.
- 23:34 – 28:16
From Enlitic to fast.ai: deep learning for medicine and the doctor shortage
- LFLex Fridman
So what was the... Tell me, if you could, the origin story of fast.ai. What is the motivation, its mission, its dream?
- JHJeremy Howard
So I guess the founding story is ti- heavily tied to my previous startup, which is a company called Enlitic, which was the first company to focus on deep learning for medicine.
- LFLex Fridman
Mm-hmm.
- JHJeremy Howard
And I created that because I saw there was a huge opportunity to, uh... there's a, there's a, about a 10X shortage of the number of doctors in the world, in the developing world, that we need. Um, expected it would take about 300 years to train enough doctors to meet that gap. But, um, I guessed that maybe if we used-... um, deep learning for some of the analytics. We could maybe make it so you don't need as highly trained doctors-
- LFLex Fridman
For diagnosis?
- JHJeremy Howard
... for diagnosis and treatment planning.
- LFLex Fridman
Where's the biggest benefit? Just to, before we get to fast.ai-
- JHJeremy Howard
Sure.
- LFLex Fridman
... wh- where's the biggest benefit of AI in medicine that you see today? Uh, and maybe in the next 10 years?
- JHJeremy Howard
Not much, not much happening today, in terms of, like, stuff that's actually out there. It's very early. But in terms of the opportunity, it's to take, uh, uh, markets like, uh, India and China and Indonesia, which have big populations, uh, um, uh, Africa, small numbers of doctors... And provide diagnostic, particularly treatment planning and triage kind of on device. So that if you do a, you know, test for malaria or tuberculosis or whatever, you immediately get something that, that even a healthcare worker that's had a month of training-
- LFLex Fridman
Mm-hmm.
- JHJeremy Howard
... can get a very high-quality assessment of whether the patient might be at risk and tell, you know, "Okay, we'll send them off to a hospital." Uh, so for example, in Africa, outside of South Africa, there's only five pediatric radiologists for the entire continent.
- LFLex Fridman
Wow.
- JHJeremy Howard
So, most countries don't h- have any. So, if your kid is sick and they need something diagnosed through medical imaging, the person... E- even if you're able to get medical imaging done, the person that looks at it will be, you know, a nurse at best.
- LFLex Fridman
Yeah.
- JHJeremy Howard
Uh, but actually in, in India, for example, and, and China, almost no X-rays are read by anybody, by any trained professional because they don't have enough.
- LFLex Fridman
Mm-hmm.
- JHJeremy Howard
So, if instead we had a algorithm that could take the most likely high-risk 5%-
- LFLex Fridman
Mm-hmm.
- JHJeremy Howard
... and say... Triage, basically. Say, "Okay, someone needs to look at this." Um, it would massively change the kind of way that... What's possible with medicine in the developing world. And remember, they have... Increasingly, they have money. They're the developing world.
- LFLex Fridman
Mm-hmm.
- JHJeremy Howard
They're not the poor world, the developing world. So, they have the money, so they're, they're building the hospitals, they're getting the, uh, diagnostic equipment, but they just... There's no way for a very long time will they be able to have the expertise.
- LFLex Fridman
Shortage of expertise? Okay. And that's where the deep learning systems can step in and-
- JHJeremy Howard
Exactly.
- LFLex Fridman
... magnify the expertise they do have-
- JHJeremy Howard
Exactly.
- LFLex Fridman
... essentially.
- JHJeremy Howard
Yeah.
- LFLex Fridman
Yeah. So, you do see... Just to linger on, a little bit longer.
- JHJeremy Howard
Sure.
- 28:16 – 32:26
Why medical AI adoption is slow: regulation, hospital lawyers, and data portability
- LFLex Fridman
Why do you think we haven't quite made progress on that yet, in terms of the h- the scale of, uh, how much, uh, AI is applied in the medical field?
- JHJeremy Howard
Oh, there's a lot of reasons. I mean, one is it's pretty new. I only started Enlitic in like 2014. And before that, like, it, it's hard to express to what degree the medical world was not aware of the o- o- opportunities here. So, I went to RSNA, which is the world's largest radiology conference.
- LFLex Fridman
Mm-hmm.
- JHJeremy Howard
And I told everybody I could, you know, like, "I'm doing this thing with deep learning. Please come and check it out." (laughs) And no one had any idea what I was talking about and no one had any interest in it. Um, so like we, we, we, we've come from absolute zero, which is hard. And then the whole regulatory framework, education system, everything is just set up to think of doctoring in a very different way.
- LFLex Fridman
Mm-hmm.
- JHJeremy Howard
So, today there is a small number of people who are deep learning practitioners and doctors at the same time.
- LFLex Fridman
Wow.
- JHJeremy Howard
And that we're starting to see the first ones come out of their PhD programs, so... Um, uh, Zach Kahane, uh, uh, uh, over in Boston, Cambridge has a number of students now who are data, uh, data science experts, deep learning experts and, and, uh, actual medical doctors.
- LFLex Fridman
Mm-hmm.
- JHJeremy Howard
Uh, quite a few doctors have completed our fast.ai course now and are publishing papers and creating journal reading groups in the American Council of Radiology. And, like it's just starting to happen.
- LFLex Fridman
Go ahead.
- JHJeremy Howard
But it's gonna be a long process. The regulators have to learn how to regulate this. They have to build, you know, um, guidelines. And then, um, the lawyers at hospitals have to develop a new way of understanding that sometimes it makes sense for data to be-... you know, looked at in raw form in large quantities in order to create world-changing results.
- LFLex Fridman
Yeah, so the regulation around data, all that, um, it sounds, uh, probably the hardest problem, but sounds reminiscent of autonomous vehicles as well, many of the same regulatory challenges, many of the same data challenges.
- JHJeremy Howard
Yeah, I mean, funnily enough, the problem is less the regulation and more the interpretation of that regulation by l- by lawyers in hospitals. So HIPAA is actually, was designed to, it's, it... The P in HIPAA is not standing, does not stand for privacy. It stands for portability. It's actually meant to be a way that data can be used. Uh, and it was created with lots of gray areas because the idea is that would be more practical and it would help people to use this, this legislation to actually share data in a more thoughtful way. Unfortunately, it's done the opposite because when a lawyer sees a gray area, they see, "Oh, if we don't know we won't get sued, then we can't do it."
- LFLex Fridman
Right.
- JHJeremy Howard
So HIPAA is not exactly the problem. The problem is more that there's... Hospital lawyers are not incented to make bold decisions (laughs) about data portability.
- LFLex Fridman
Or even to, uh, embrace technology that saves lives.
- JHJeremy Howard
Right.
- LFLex Fridman
They more want to not get in trouble for embracing that technology.
- JHJeremy Howard
Right. Also, it em- it, it has also saves lives in a very abstract way, which is like, "Oh, we've been to release these 100,000 anonymized records."
- LFLex Fridman
Right.
- JHJeremy Howard
I can't point at the specific person whose life that saved. I can say like, "Oh, we ended up with this paper which found this result, which, you know, diagnosed 1,000 more people than we would have otherwise." But it's like which ones were helped?
- LFLex Fridman
Yeah.
- JHJeremy Howard
It's, it's very abstract.
- LFLex Fridman
Yeah. And on the counter side of that, you may be able to point to, uh, a life that was taken because of something that was-
- JHJeremy Howard
Yeah. Or, or, or a person whose privacy was violated.
- LFLex Fridman
Violated.
- JHJeremy Howard
It's like, "Oh, this specific person-"
- LFLex Fridman
Right.
- JHJeremy Howard
... "you know, was de-identified."
- 32:26 – 37:59
Privacy and “do more with less”: transfer learning as a practical antidote
- LFLex Fridman
S- just a fascinating topic. We're jumping around. We'll get back to fast.ai, but on the question of privacy, data is the fuel for so much innovation in deep learning. What's your sense in privacy, whether we're talking about Twitter, Facebook, YouTube, uh, just the technologies like in the medical field that rely on people's data in order to create impact? How do we get that right, uh, respecting people's privacy and yet creating technology that-
- JHJeremy Howard
Well-
- LFLex Fridman
... uh, is learned from data?
- JHJeremy Howard
... one of my areas of focus is on doing more with less data. Um, which, so most vendors, unfortunately, are strongly incented to find ways to require more data and more computation. So Google and IBM being the most obvious, uh-
- LFLex Fridman
IBM?
- JHJeremy Howard
Yeah.
- LFLex Fridman
Sorry.
- JHJeremy Howard
So Watson.
- LFLex Fridman
Watson, yeah.
- JHJeremy Howard
You know, so Google and IBM both strongly push the idea that you have to be, you know, that they have more data and more computation and more intelligent people than anybody else.
- LFLex Fridman
Right.
- JHJeremy Howard
And so you have to trust them to do things because nobody else can do it. Um, and Google's very, uh, upfront about this. Like, Jeff Dean has g- gone out there and given talks and said, "Our goal is to require 1,000 times more computation-"
- LFLex Fridman
Mm-hmm.
- JHJeremy Howard
"... but less people." Um, our goal is to use the people that you have better and the data you have better and the computation you have better. So one of the things that we've discovered is, or, or at least highlighted, is that you very, very, very often don't need much data at all. And so the data you already have in your organization will be enough to get state-of-the-art results.
- LFLex Fridman
Mm-hmm.
- JHJeremy Howard
So like, my starting point would be just kind of, say, around privacy is a lot of people are looking for ways to sh- share data and aggregate data, but I think often that's unnecessary. They assume that they need more data than they do 'cause they're not familiar with the basics of transfer learning, which is this, uh, critical technique for needing orders of magnitude less data.
- LFLex Fridman
Is your sense one reason you might want to collect data from everyone is, uh, like in a recommender system context where your individual, Jeremy Howard's, individual data is the most useful for figu- for providing a product that's impactful for you? So f- for giving you advertisements, for recommending to you movies, for doing medical diagnosis. Uh, i- is your sense we can build with a small amount of data general models that will have a huge impact for most people that-
- JHJeremy Howard
Mm-hmm.
- LFLex Fridman
... don't need to sp- have data from each individual?
- JHJeremy Howard
On, on, on the whole, I'd say yes. I mean, there are things like, you know, recommender systems have this, uh, cold start problem where, you know, Jeremy is a new customer. We haven't seen him before so we can't recommend him things based on what else he's bought and liked with us.
- LFLex Fridman
Right.
- JHJeremy Howard
Um, and there's various workarounds to that. Like in a lot of music programs will start out by saying, "Which of these artists do you like? Which of these albums do you like? Which of these songs do you like?" Netflix used to do that. Nowadays they, they tend not to. Uh, people kind of don't like that 'cause they think, "Oh, we don't want to bother the user." So you could work around that by having some kind of data sharing where you get my marketing record from Acxiom or whatever-
- LFLex Fridman
Mm-hmm.
- JHJeremy Howard
... and try to guess from that. To, to me the, the benefit to me and to society of saving me five minutes on answering some questions versus the negative externalities of, of the privacy issue-... doesn't add up. So I think, like, a lot of the time, the places where people are, uh, invading our privacy in order to provide convenience is really about just trying to make them more money. And, and they move these negative externalities into places that they don't have to pay for them.
- LFLex Fridman
Mm-hmm.
- JHJeremy Howard
So when you actually see regulations appear that actually cause the companies that create these negative externalities to have to pay for it themselves, they say, "Well, we can't do it anymore." So the cost is actually too high.
- LFLex Fridman
Right.
- JHJeremy Howard
But for something like medicine, yeah, I mean the hospital has my, you know, medical imaging, my pathology studies, my medical records. Um, and also I own my medical data. So you can, um... So I, I help a startup called, uh, Doc.ai. One of the things Doc.ai does is that there- it has an app that you can connect to, you know, Sutter Health and Labcorp and-
- LFLex Fridman
Mm-hmm.
- JHJeremy Howard
... Walgreens, and download your medical data to your phone, and then up- upload it, again, at your discretion, to share it as you wish. Um, so with that kind of approach, we can share our medical information with the people we want to.
- 37:59 – 40:56
fast.ai’s mission: empower domain experts (who may not code) to use deep learning
- JHJeremy Howard
Right. So, so before I started fast.ai, I spent a year researching, where are the biggest opportunities for deep learning? 'Cause I knew from my time at Kaggle, in particular, that deep learning had kind of hit this threshold point where it was rapidly becoming the state-of-the-art approach in every area that looked at it. And I'd been working with neural nets for over 20 years. I knew that, from a theoretical point of view, once it hit that point, it would do that in kind of just about every domain.
- LFLex Fridman
Mm-hmm.
- JHJeremy Howard
And so I kind of spent a year researching, what are the domains it's gonna have the biggest low-hanging fruit in the shortest time period? And I picked medicine, but there were so many I could have picked. And so there was a kind of level of frustration for me of like, okay, I'm really glad we've opened up the medical deep learning world, and today it's huge, as you know, um, but we can't do, (laughs) you know, I can't do everything. I don't even know, like... It took, like, in medicine, it took me a really long time to even get a sense of like, what kind of problems do medical practitioners-
- LFLex Fridman
Mm-hmm.
- JHJeremy Howard
... solve? What kind of data do they have? Who has that data? So, I kind of felt like I need to approach this differently if I want to maximize the positive impact of deep learning. Um, rather than m- me picking an area and trying to become good at it and building something, I should let people who are already domain experts in those areas and who already have the data do it themselves.
- LFLex Fridman
Mm-hmm.
- JHJeremy Howard
So that was the reason for fast.ai-
- LFLex Fridman
Mm-hmm.
- JHJeremy Howard
... is to basically try and figure out how to get deep learning into the hands of people who could benefit from it-
- LFLex Fridman
Mm-hmm.
- JHJeremy Howard
... and help them to do so in as quick and easy and effective a way as possible.
- LFLex Fridman
Got it. So sort of empower the, the, the domain experts.
- JHJeremy Howard
Yeah. And, like, partly it's 'cause, like, unlike most people in this field, my background is very applied and industrial. Like, my first job at m- was at McKinsey & Company. I spent 10 years in management consulting. I, I spent a lot of time with domain experts.
- LFLex Fridman
Mm-hmm. Right.
- JHJeremy Howard
You know, so I kind of respect them and appreciate them, and know, I know that's where their value generation in society is. And so I also know how m- most of them can't code.
- LFLex Fridman
Mm-hmm.
- JHJeremy Howard
And most of them don't have the time to invest, you know, three years in a g- graduate degree or whatever. So it's like, how do I upskill those domain experts? I think that would be a super powerful thing, you know, bi- biggest societal impact I could have. So that, yeah, that was the thinking.
- LFLex Fridman
So, so much of fast.ai students and researchers and the things you teach are, uh, pragmatically minded.
- JHJeremy Howard
Right.
- LFLex Fridman
Practically minded. Figuring, figuring out ways how to solve real problems and fast.
- JHJeremy Howard
Right.
- 40:56 – 45:41
Theory vs practice: why much deep learning research misses what matters
- LFLex Fridman
So from your experience, what's the difference between theory and practice of deep learning?
- JHJeremy Howard
Hmm. Well, most of the research in the deep learning world is a total waste of time, uh, becau-
- LFLex Fridman
Right, that's what I was getting at. (laughs)
- JHJeremy Howard
Yeah. It's a s- it's a problem of, in science in general. Um, scientists need to be published, which means they need to work on things that their peers are extremely familiar with and can recognize and advance in that area. So that means that they all need to work on the same thing.
- LFLex Fridman
Yeah.
- JHJeremy Howard
And so it really enc- and, and the thing they work on, there's nothing to encourage them to work on things that are practically useful. So you get just a whole lot of research which is minor advances in stuff that's been very highly studied and has no significant practical impact.
- LFLex Fridman
Mm-hmm.
- JHJeremy Howard
Whereas the things that really make a difference, like I mentioned transfer learning. Like, if we can do better at transfer learning, then it's this like world-changing thing, where suddenly like, lots more people can do world-class work with less resources and less data, and... But almost nobody works on that. Or another example, active learning, which is the study of like, how do we get more out of the human beings-
- LFLex Fridman
Yeah.
- JHJeremy Howard
... in the loop?
- LFLex Fridman
That's my favorite topic.
- JHJeremy Howard
Yeah. So active learning's great, but it's, almost nobody working on it, um, because it's just not a trendy thing right now. You know, it's-
- LFLex Fridman
You know what somebody s- sorry to interrupt, uh-
- JHJeremy Howard
Yeah.
- LFLex Fridman
You were saying that nobody's publishing on active learning.
- JHJeremy Howard
Right.
- LFLex Fridman
But there's people inside companies, anybody who actually has to solve a problem, they're going to innovate on active learning.
- JHJeremy Howard
Yeah, everybody kind of reinvents active learning-
- LFLex Fridman
Right.
- JHJeremy Howard
... when they actually have to work in practice, because they start labeling things and they think, "Gosh, this is taking a long time and it's very expensive."
- LFLex Fridman
Yeah.
- JHJeremy Howard
And then they start thinking, "Well, why am I labeling everything? I'm only... The machine's only making mistakes on those two classes. They're the hard ones. Maybe I'll just start labeling those two classes." And then you start thinking, "Well, why did I do that manually? Why can't I just get the system to tell me which things are gonna be hardest?"
- LFLex Fridman
Yeah. Yep.
- JHJeremy Howard
It's, it's an obvious thing to do, but, um, yeah, it's, it's just like, like transfer learning, it's, it's under-studied and the academic world just has no reason to care about practical results. The funny thing is, like I've only really ever written one paper. I hate writing papers.
- LFLex Fridman
(laughs)
- JHJeremy Howard
Um, and I didn't even write it. It was my colleague, Sebastian Ruda, who actually wrote it. I just-
- LFLex Fridman
Yeah.
- JHJeremy Howard
... did the research for it. Uh, but it was basically introducing transfer learning, successful transfer learning, to NLP-
- LFLex Fridman
Mm-hmm.
- JHJeremy Howard
... for the first time. Uh, and the algorithm is called ULMFiT. And it, it actually, I actually wrote it for the course, for the fast.ai course. Uh, I wanted to teach people NLP and I thought, "I only want to teach people practical stuff." And I think the only practical stuff is transfer learning, and I couldn't find any examples of transfer learning in NLP, so I just did it. And I was shocked to find that as soon as I did it, which, you know, the basic prototype took a couple of days, smashed the state of the art on one of the most important datasets in a field that I knew nothing about. And I just thought, "Well, this is ridiculous." And so, um, I spoke to Sebastian about it, and he kindly offered to write it up-
- 45:41 – 52:02
DAWNBench: fast.ai’s speed-and-cost wins and the ImageNet resolution trick
- LFLex Fridman
Once again, just as you described. Can you tell the story of, uh, the Stanford competition, DAWNBench, and fast.ai's achievement on it?
- JHJeremy Howard
Sure. So, something which I really enjoy is that I, I basically teach two courses a year. Um, the practical Deep Learning for Coders, which is kind of the introductory course, and then Cutting-Edge Deep Learning for Coders, which is the kind of research-level course. And when I teach those courses, um, I have a, a, I, I basically have a big office, uh, at, at University of San Francisco, big enough for like 30 people, and I invite anybody, any student who wants to come and hang out with me while I build the course.
- LFLex Fridman
Mm-hmm.
- JHJeremy Howard
And so generally it's full. And so we have 20 or 30 people in a big office with nothing to do but study deep learning.
- LFLex Fridman
Mm-hmm.
- JHJeremy Howard
Uh, so it was during one of these times that somebody in the group said, "Oh, there's a thing called DAWNBench that looks interesting." And I was like, "What the hell is that?" And they said, "Oh, it's some competition to see how quickly you can train a model. Seems kind of, not exactly relevant to what we're doing, but it sounds like the kind of thing which you might be interested in." And I checked it out and I was like, "Oh, crap, there's only 10 days till it's over."
- LFLex Fridman
(laughs)
- JHJeremy Howard
"It's pretty much too late. And we're kind of busy trying to teach this course." (laughs)
- LFLex Fridman
Yeah.
- JHJeremy Howard
But we're like, "Uh, it would make an interesting case study for the course. Like, it's all the stuff we're already doing. Why don't we just put together our current best practices and ideas?"
- LFLex Fridman
Mm-hmm.
- JHJeremy Howard
So me and, I guess, about four students just decided to give it a go. And we focused on this small one called CIFAR-10, which is little 32 by 32 pixel images.
- LFLex Fridman
Can you say what DAWNBench is that-
- JHJeremy Howard
Yeah, so it's a, it's competition to, to train a model as fast as possible. It was run by Stanford. Uh-
- LFLex Fridman
And as cheap as possible too.
- JHJeremy Howard
Uh, that's also another one, for as cheap as possible. And there was a couple of categories, uh, ImageNet and CIFAR-10. So ImageNet's this big 1.3 million, uh, image thing that took couple of days to train. Remember a friend of mine, uh, Pete Warden, who's now at Google, um, I remember he told me how he trained ImageNet a few years ago when he basically, like, had this, uh, uh, little granny flat out the back that he turned into his ImageNet training center.
- LFLex Fridman
(laughs)
- JHJeremy Howard
And he figu- you know, after like a year of work, he figured out how to train it in like a 10 days or something. It was like, that was a big job. Whereas CIFAR-10, at that time, you could train in a few hours. You know, it's much smaller and easier. So we thought we'd try CIFAR-10. And yeah, I'd really never done that before. Like I'd never really... Like things like using more than one GP- uh, GPU at a time was something I...... try to avoid, 'cause to me it's like very against the whole idea of accessibility is should be able to do things with one GPU.
- LFLex Fridman
Uh, I mean, have you asked in the past before, after having accomplished something, "How do I do this faster, much faster?"
- JHJeremy Howard
Oh, always. But it's always, for me, it's always, "How do I make it much faster on a single GPU-"
- LFLex Fridman
On a single GPU.
- JHJeremy Howard
"... that a normal person could afford in their day-to-day life?"
- LFLex Fridman
Yeah.
- JHJeremy Howard
It's not, "How could I do it faster by, you know, having a huge data center?" 'Cause, uh, to me, it's all about, like, as many people should be able to use something as possible without fussing around with infrastructure. So, anyways, so in this case, it's like, well, we can use, uh, eight GPUs just by renting a AWS machine.
- LFLex Fridman
Mm-hmm.
- JHJeremy Howard
So we thought we'd try that. And, um, yeah, basically using the stuff we were already doing, we were able to get, you know, the speed, you know, within a few days, we had the speed down to, I don't know, a, a s- a very small number of minutes. I can't remember exactly how many minutes it was, but it might have been like 10 minutes or something.
- LFLex Fridman
Mm-hmm.
- JHJeremy Howard
And so, yeah, we found ourselves at the top of the leaderboard easily for both time and money.
- LFLex Fridman
(laughs)
- JHJeremy Howard
Which really shocked me 'cause the other people competing in this were like Google and Intel and stuff, who are like, know a lot more about this stuff than I think we do. So then we were emboldened. We thought, "Let's try the ImageNet one too." (laughs) I mean, it seemed way out of our league. But, uh, our goal was to get under 12 hours.
- 52:02 – 58:57
Single-GPU creativity: smaller benchmarks, DeOldify, and an underused audio frontier
- LFLex Fridman
So what's your view on multi-GPU or multiple-machine training in general as, as a way to speed code up?
- JHJeremy Howard
I think it's largely a waste of time.
- LFLex Fridman
Both multi-GPU on a single machine and...
- JHJeremy Howard
Yeah. Particularly multi-machines 'cause it's just clunky. Uh, multi-GPUs is less clunky than it used to be. But to me, anything that slows down your iteration speed is a waste of time. So, you could maybe do your very last, you know, perfecting of the model on multi-GPUs if you need to. But, so, for example, I think doing stuff on ImageNet is generally a waste of time. Why test things on 1.3 million images? Most of us don't use 1.3 million images. And we've also done research that shows that doing things on a smaller subset of images gives you the same relative answers anyway. So from a research point of view, why waste that time? So, actually, I released a couple of new datasets recently. One is called Imagenet.
- LFLex Fridman
(laughs)
- JHJeremy Howard
Uh, the French ImageNet, uh, which is a small subset of ImageNet, which is designed to be easy to classify.
- LFLex Fridman
Uh, what's, uh, how do you spell Imagenet?
- JHJeremy Howard
It's got an extra T and E at the end 'cause it's very French.
- LFLex Fridman
Imagem- okay. (laughs)
- JHJeremy Howard
Yeah, and then I cr- and then another one called, um, Imageworf, which is a subset of ImageNet that only contains dog breeds. And both-
- LFLex Fridman
That's a hard one, right?
- JHJeremy Howard
That's a hard one.
- LFLex Fridman
Yeah.
- JHJeremy Howard
And I've discovered that if you just look at these two subsets, you can train things on a single GPU in 10 minutes. And the results you get are directly transferrable to ImageNet nearly all the time. And so now, I'm starting to see some researchers start to use these-
- LFLex Fridman
Oh, man, I-
- JHJeremy Howard
... much smaller datasets.
- LFLex Fridman
... I so deeply love the way you think, because, um, I think you might have written a blog post saying, um, that sort of going to these big datasets is, um, encouraging people to, to not think creatively.
- JHJeremy Howard
Absolutely.
- LFLex Fridman
So you're too... It, it sort of, um, constrains you to train on large resources, and because you have these resources, you think more resources will be b- better, and then you start... So like, for, some- somehow you kill the creativity.
- JHJeremy Howard
Yeah. And even worse than that, Lex, I, I keep hearing from people who say, "I decided not to get into deep learning-"
- LFLex Fridman
Mm-hmm.
- JHJeremy Howard
"... because I don't believe it's accessible to people outside of Google to do useful work." So, like, I see a lot of people make an explicit decision to not learn this incredibly valuable tool because they've...... they've drunk the Google Kool-Aid, which is that only Google's big enough and smart enough-
- LFLex Fridman
Yeah.
- JHJeremy Howard
... to, to do it. And I just find that so disappointing and it's so wrong.
- LFLex Fridman
And I think all the major breakthroughs in AI in the next 20 years will be doable on a single GPU. Like I would, I would say, my sense is all the big sort of, uh, even-
- JHJeremy Howard
Well, let's put it this way, none of the big breakthroughs of the last 20 years have required multiple GPUs.
- LFLex Fridman
Right.
- JHJeremy Howard
So like batchnorm, ReLU, Dropout, conv-
- LFLex Fridman
To, to demonstrate that there's something to them.
- JHJeremy Howard
... ConvNets in general, every one of them-
- 58:57 – 1:06:16
Learning-rate breakthroughs: superconvergence and the limits of current DL science norms
- LFLex Fridman
Okay. Uh, also learning rate in terms of DAWNBench. Um, there's some magic on learning rate that you played around with-
- JHJeremy Howard
Yeah.
- LFLex Fridman
... that's kind of interesting.
- JHJeremy Howard
Yeah. So this is all work that came from a guy called Leslie Smith. Uh, Leslie's a researcher who, like us, cares a lot about just the practicalities of training neural networks quickly and accurately. Which you would think is what everybody should care about but almost nobody does. (laughs) Um, and, uh, he discovered something very interesting which he calls super convergence, which is there are certain networks that with certain settings of hyper parameters could suddenly be trained ten times faster-
- LFLex Fridman
Mm-hmm.
- JHJeremy Howard
... by using a ten times higher learning rate. Now, um, no one published that paper because it's not an area of kind of active research in the academic world, no academics recognize this is important. And also deep learning in academia is not considered a experimental science. So unlike in physics where you could say like, "I just saw a s- a sub-atomic particle do something which the theory doesn't explain."
- LFLex Fridman
Mm-hmm.
- JHJeremy Howard
You could publish that without an explanation.
- LFLex Fridman
Right.
- JHJeremy Howard
And then in the next 60 years people can try to work out how to explain it. We don't allow this in the deep learning world so it's, it's literally impossible for Leslie to publish a paper that says, "I've just seen something amazing happen, this thing trained ten times faster than it should have, I don't know why." And so the reviewers were like, "Well, you can't publish that 'cause you don't why." So anyway.
- LFLex Fridman
That's important to pause on because there's so many discoveries that would need to start like that.
- JHJeremy Howard
Every, every other scientific field I know of works of that way. I don't know why ours is uniquely disinterested in-... publishing unexplained experimental results. But there it is. So it wasn't published. Having said that, uh, I read a lot more unpublished papers than published papers, 'cause that's where you find the interesting insights.
- LFLex Fridman
Mm-hmm.
- JHJeremy Howard
So I absolutely read this paper. And I was just like, "This is astonishingly mind-blowing and weird and awesome." And like, "Why isn't everybody only talking about this?" Because like, if you can train these things 10 times faster... They also generalize better because you're, you're doing less epochs, which means you look at the data less, so you get better accuracy. So I've been kind of studying that ever since. And, uh, eventually, Leslie kind of figured out a lot of how to get this done, and we added minor tweaks. And a big part of the trick is starting at a very low learning rate, very gradually increasing it. So as you're training your model, you take very small steps at the start, and you gradually make them bigger and bigger until eventually you're taking much bigger steps than anybody thought was possible.
- LFLex Fridman
Right.
- JHJeremy Howard
There's a few other little tricks to make it work, but e- ev- basically, we can reliably get super convergence. And so for the DAWNBench thing, we were using just much higher learning rates than people expected to work.
- LFLex Fridman
What do you think the future of... I mean, it makes so much sense for that to be a critical hyper-parameter, learning rate that you vary. What do you think the future of learning rate magic looks like?
- JHJeremy Howard
Well, there's been a lot of great work in the last 12 months in this area. It's... And people are increasingly realizing that optimize... Like, we just have no idea really how optimizers work. And, uh, the combination of weight decay, which is how we regularize optimizers, and the learning rate, and then other things like the epsilon we use in, in the Adam optimizer, they all work together in weird ways. And different parts of the model... This is another thing we've done a lot of work on, is research into how different parts of the model should be trained at different rates in different ways.
- LFLex Fridman
Mm-hmm.
- JHJeremy Howard
Um, so we do something we call discriminative learning rates, which is really important, particularly for transfer learning. Um, so really, I think in the last 12 months, a lot of people have realized that this, all this stuff is important. There's been a lot of great work coming out. And we're starting to see algorithms appear which have very, very few dials, if any, that you have to touch. So like the... I think what's going to happen is the idea of a learning rate will... It almost already has disappeared-
- LFLex Fridman
Mm-hmm.
- JHJeremy Howard
... in the latest research. And instead, it's just like, you know, we, we know enough about how to interpret the gradients and the change of gradients we see to know how to set every parameter off to
- LFLex Fridman
That you can automate it. So you see the future of, uh, of, uh, deep learning where really... Where is the input of a human expert needed in the future?
- JHJeremy Howard
Well, hopefully, the input of the human expert will be almost entirely unneeded from the deep learning point of view. So, um, again, like Google's approach to this is to try and use thousands of times more compute to run lots and lots of models at the same time and hope that one of them is good.
- LFLex Fridman
AutoML kind of thing?
- JHJeremy Howard
Yeah, AutoML kind of stuff, which I think is insane.
- LFLex Fridman
(laughs)
- JHJeremy Howard
Um, when you better understand the mechanics of how models learn, you don't have to try a thousand different models to find which one happens to work the best.
- LFLex Fridman
Mm-hmm.
- JHJeremy Howard
You can just jump straight to the best one, uh, which means that it's more accessible in terms of compute, cheaper, and also with less hyper-parameters to set. It means you don't need deep learning experts to train your deep learning model for you, which means that domain experts can do more of the work, which means that now you can focus the human time on the kind of interpretation, the data gathering, identifying what all errors, and stuff like that.
- 1:06:16 – 1:17:51
Cloud, tooling, and frameworks: why fast.ai chose PyTorch (and where Swift fits)
- LFLex Fridman
... (sighs) what are the different cloud options for training your networks? Last question related to DAWNBench. Well, it's part of a lot of the work you do, but from a perspective of performance, I think you've written this in a blog post. There's AWS. There's T- TPU from Google. What's your sense what the future holds? What would you recommend now-
- JHJeremy Howard
Right.
- LFLex Fridman
... in terms of, uh-
- JHJeremy Howard
So from a hardware point of view-
- LFLex Fridman
... training in the cloud?
- JHJeremy Howard
... Google's TPUs and the best NVIDIA GPUs are similar. I mean, maybe the TPUs are like 30% faster, but they're also much harder to program.... where there isn't a clear leader in terms of hardware right now. Although, much more importantly, the GPU, NVIDIA's GPUs are much more programmable, they've got much more written for them. So like that's the clear leader for me and where I would spend my time as a researcher and practitioner. But then in terms of the platform, I mean, we're, we're super lucky now with stuff like Google, uh, GCP, Google Cloud, and, um, AWS that you can access a GPU pretty quickly and easily. But, I mean, for, for AWS it's still too hard. Like you have to find an AMI and get the instance running and then install the software you want and blah, blah, blah. GCP is still, is currently the, the best way to get started on a full-
- LFLex Fridman
Easiest?
- JHJeremy Howard
... server environment because, um, they have a fantastic fast.ai and PyTorch ready to go instance, which has all the courses pre-installed, it has Jupyter Notebook pre-running. Jupyter Notebook is this, uh, wonderful interactive computing system which everybody basically should (laughs) be using-
- LFLex Fridman
Mm-hmm.
- JHJeremy Howard
... for any kind of data-driven research. But then even better than that, uh, uh, there are platforms like, um, Salamander, which we own, and, uh, Paperspace where literally you click a single button and it pops up a Jupyter Notebook straightaway without any kind of, uh-
- LFLex Fridman
(laughs)
- JHJeremy Howard
... um, installation or anything.
- LFLex Fridman
Yeah.
- JHJeremy Howard
And all the course notebooks are all preinstalled. So like for me, we, um, this is one of the things we spent a lot of time, uh, kind of curating and working on, 'cause when we first started our courses, the biggest problem was people dropped out of lesson one 'cause they couldn't get an AWS instance running.
- LFLex Fridman
Right.
- JHJeremy Howard
So, um, things are so much better now. And like we actually have, if you go to course.fast.ai, the first thing it says is, "Here's how to get started with your GPU." And it's like you just click on the link and you click start and-
- LFLex Fridman
And it's-
- JHJeremy Howard
... you're going.
- LFLex Fridman
... it w- you will go GCP. I have to confess, I've never used the Google GCP.
- JHJeremy Howard
Yeah. GCP gives you $300 of compute for free, which is really nice. But as I say, uh, Salamander and Paperspace are even, even easier still.
- LFLex Fridman
Okay. So, uh, the, from the perspective of deep learning frameworks, you work with fast.ai, particular to this framework, and, uh, PyTorch and TensorFlow. What are the strengths of each platform-
- JHJeremy Howard
Sure.
- LFLex Fridman
... in your perspective?
- JHJeremy Howard
So in terms of what we've done our research on and taught in our course, we started with Theano and Keras, and then we switched to TensorFlow and Keras. And then we switched to PyTorch, and then we switched to PyTorch and fast.ai. Um, and that, th- that kind of reflects a growth and development of the, the ecosystem of deep learning libraries. Th- Theano and TensorFlow were great, but were much harder to teach and to do research and development on because they define what's called a computational graph upfront, a static graph, where you basically have to say, "Here are all the things that I'm going to eventually do in my model."
- LFLex Fridman
Mm-hmm.
- JHJeremy Howard
And then later on you say, "Okay, do those things with this data." And you can't, like, debug them. You can't do them step by step. You can't program them interactively in a Jupyter Notebook and so forth. PyTorch was not the first, but PyTorch was certainly the, the s- the strongest entrant to come along and say, "Let's not do it that way. Let's just use normal Python. And everything you know about in Python is just gonna work, and we'll figure out how to make that run on the GPU as and when n- necessary."
- LFLex Fridman
Mm-hmm.
- JHJeremy Howard
That turned out to be a huge, a huge leap in terms of what we could do with our research and what we could do with our teaching, and-
- LFLex Fridman
'Cause it wasn't limiting.
- JHJeremy Howard
Yeah, I mean, it was critical for us for something like DAWNBench to be able to rapidly try things. It's just so much harder to be a researcher and practitioner when you have to do everything up front and you can't inspect it. Um, problem with PyTorch is it's not at all accessible to newcomers because you have to, like, write your own training loop and manage the-
- 1:17:51 – 1:26:59
How to get started (and become an expert): train models, fine-tune, and pick a domain
- LFLex Fridman
How long does it take, um... So th- if you look at the two- two fast.ai courses, how long does it take to get from point zero to completing both courses?
- JHJeremy Howard
Um, it varies a lot. Somewhere between two months and two years, generally.
- LFLex Fridman
So for two months, how many hours a day on average?
- JHJeremy Howard
So like, so like a- a- a- somebody who is a very competent coder can- can do 70 hours per course and pick up-
- LFLex Fridman
70? Seven zero?
- JHJeremy Howard
70, yeah.
- LFLex Fridman
That's it? Okay.
- JHJeremy Howard
Yeah. But a lot of people I know who take a year off to study fast.ai full time-
- LFLex Fridman
Yeah.
- JHJeremy Howard
... and say at the end of the year, they feel pretty competent 'cause generally there's a lot of other things you do. Like they're, generally they'll be entering Kaggle competitions, they-
- LFLex Fridman
Yeah, exactly.
- JHJeremy Howard
... might be reading Ian Goodfellow's book. They might, you know, they'll be doing a bunch of stuff.
- LFLex Fridman
Yeah.
- JHJeremy Howard
And often, you know, particularly if they are a domain expert, their coding skills might be-... a little on the pedestrian side. So part of it's just, like, doing a lot more writing.
- LFLex Fridman
What do you find is the bottleneck for people usually? Uh, except getting started and setting stuff up.
- JHJeremy Howard
I would say coding.
- LFLex Fridman
Just-
- JHJeremy Howard
Yeah, I would say the best... The people who are strong coders pick it up the best. Although another bottleneck is people who have a lot of experience of classic statistics can really struggle because it... The intuition is so the opposite of what they're used to. They're very used to-
- LFLex Fridman
Hmm.
- JHJeremy Howard
... like, trying to reduce the number of parameters in their model and looking at individual coefficients and stuff like that. So I, I find people who have a lot of coding background and know nothing about statistics are generally going to be the best off.
- LFLex Fridman
So, uh, you taught several courses on deep learning. And as Feynman says, "Best way to understand something is to teach it." What have you learned about deep learning from teaching it?
- JHJeremy Howard
A lot. That's a key reason for me to, to teach the courses. I mean, obviously it's going to be necessary to achieve our goal of getting domain experts to be familiar with deep learning, but that's also necessary for me to achieve my goal (laughs) of being really familiar with deep learning.
- LFLex Fridman
Yeah.
- JHJeremy Howard
I, I mean (sighs) to see so many domain experts from so many different backgrounds, it's definitely, I wouldn't say taught me, but convinced me something that I liked to believe was true, which was anyone can do it. So there's a lot of, kind of snobbishness out there about only certain people can learn to code, only certain people are going to be smart enough to, like, do AI. That's definitely bullshit, you know?
- LFLex Fridman
Mm-hmm.
- JHJeremy Howard
I've seen so many people from so many different backgrounds get state-of-the-art results in their domain areas now. Uh, the, uh, it's definitely taught me that the key differentiator between people that succeed and people that fail is tenacity. That seems to be basically the only thing that matters. Um, the people... A lot of people give up, and... But of the ones who don't give up, pretty much everybody succeeds, you know, even if at first I'm just kind of, like, thinking like, "Wow, they really aren't quite getting it yet, are they?" But, uh, eventually people get it and they succeed. So I think that's been... I think they're both things I'd liked to believe was true, but I don't feel like I really had strong evidence (laughs) for them to be true. But now I can say I've seen it again and again.
- LFLex Fridman
So what advice do you have for someone, uh, who wants to get started in deep learning?
- JHJeremy Howard
Train lots of models. That's, that's how you, that's how you learn it. So, like... So I would... You know, uh, I think... Uh, uh, it's not just me. I think, I think our course is very good, but also lots of people independently have said it's very good. It recently won the CogX award for AI courses as being the best in the world. So I, I'd say come to our course, course.fast.ai.
- LFLex Fridman
Of course.
- JHJeremy Howard
And the thing I keep on harping on in my lessons is train models, print out the inputs to the models, print out the outputs to the models. Like study, you know, change, change the inputs a bit, look at how the outputs vary. Just run lots of experiments to get a, you know, an intuitive understanding of what's going on.
Episode duration: 1:44:10
Install uListen for AI-powered chat & search across the full episode — Get Full Transcript
Transcript of episode J6XcP4JOHmk