Skip to content
Lex Fridman PodcastLex Fridman Podcast

Rajat Monga: TensorFlow | Lex Fridman Podcast #22

Lex Fridman and Rajat Monga on rajat Monga on TensorFlow’s evolution, ecosystem, and open-source impact.

Lex FridmanhostRajat Mongaguest
Jun 3, 20191h 10mWatch on YouTube ↗

EVERY SPOKEN WORD

  1. 0:002:40

    Google Brain’s early mission: scaling deep learning with Google’s compute and data

    1. LF

      The following is a conversation with Rajat Monga. He's an engineering director at Google, leading the TensorFlow team. TensorFlow is an open source library at the center of much of the work going on in the world in deep learning, both the cutting edge research and the large-scale application of learning-based approaches. But it's quickly becoming much more than a software library. It's now an ecosystem of tools for the deployment of machine learning in the cloud, on the phone, in the browser, on both generic and specialized hardware, TPU, GPU, and so on. Plus, there's a big emphasis on growing a passionate community of developers. Rajat, Jeff Dean, and a large team of engineers at Google Brain are working to define the future of machine learning with TensorFlow 2.0, which is now in alpha. I think the decision to open source TensorFlow was a definitive moment in the tech industry. It showed that open innovation could be successful, and inspired many companies to open source their code, to publish, and in general, engage in the open exchange of ideas. This conversation is part of the Artificial Intelligence podcast. If you enjoy it, subscribe on YouTube, iTunes, or simply connect with me on Twitter, @LexFridman, spelled F-R-I-D. And now, here's my conversation with Rajat Monga. You were involved with Google Brain since its start in 2011 with, uh, Jeff Dean. It started with this belief the proprietary machine learning library and turned into TensorFlow in 2014, the open source library. So, what were the early days of Google Brain like? What were the goals, the missions? How do you even proceed forward once there's so much possibilities before you?

    2. RM

      It was interesting back then, you know, when I started out, when you, you were even just talking about it. The idea of deep learning was interesting and intriguing in some ways. It hadn't yet taken off, but it held some promise and it showed some very promising and early results. I think the, the idea where Andrew and Jeff had started was, what if we can take this, what people are doing in research, and scale it to what Google has in terms of the compute power? And, uh, also put that kind of data together, what does it mean? And so far, the results have been if you scale the compute, scale the data, it does better, and would that work. And so that, that was the first year or two, "Can we prove that out right?" And with this belief, when we started the first year, we got some early wins, which, which is always great.

  2. 2:403:10

    First proof points: speech recognition and the “cat paper” image breakthrough

    1. LF

      What were the wins like? What was the wins where you were, "There's some promise to this, this is gonna be good"?

    2. RM

      I think the two early wins were, one was speech that we collaborated very closely with the speech research team who was also getting interested in this, and the other one was on images where we, you know, the cat paper as we call it-

    3. LF

      Mm-hmm.

    4. RM

      ... that was covered by-

    5. LF

      Yeah.

    6. RM

      ... (laughs) uh, a lot of folks.

    7. LF

      And, uh, the birth of Google Brain was a- around neural networks. That was ... So, it was deep learning from the very beginning.

    8. RM

      That's right.

    9. LF

      That was the whole mission.

    10. RM

      Yeah.

  3. 3:104:46

    From experiments to massive scale: thousands to 10,000-machine training runs

    1. LF

      So, what, what, uh, in terms of scale, what was the sort of, uh, dream of what this could become? Like, what, were there echoes of this open source TensorFlow community that might be brought in? Was there a sense of TPUs? Was there a sense of like, machine learning is now gonna be at the core of the entire company, is g- going to grow into that direction?

    2. RM

      Yeah, I, I think ... So, so that was interesting, and like, if I think back to 2012 or 2011-

    3. LF

      Right.

    4. RM

      ... and first was, can we scale it? And in the year or so, we had started scaling it to hundreds and thousands of machines. In fact, we had some runs even going to 10,000 machines, and all of those shows great promise. Uh, in terms of machine learning at Google, the good thing was Google's been doing machine learning for a long time. Deep learning was new, but as we scaled this up, we showed that, yes, that was possible, and it was gonna impact lots of things, like we started seeing real products wanting to use this. Again, speech was the first. There were image things that photos came out of, and, and then many other products as well. So, so that was exciting. Um, as we went into that a couple of years, externally also, academia started to, you know, there was lots of push on, "Okay, deep learning's interesting. We should be doing more," and so on. And so, by 2014, we were looking at, "Okay, this is a big thing. It's gonna grow," and, uh, not just internally, externally as well. Yes, maybe Google's ahead of where everybody is, but there's a lot to do, so a lot of this start to make sense and come together.

  4. 4:467:18

    Why open source TensorFlow: research sharing, better standards, and avoiding “Hadoop repeats”

    1. LF

      So, the decision to open source ... I was just chatting with, uh, with Chris Lattner about this. Uh, the decision to go open source for TensorFlow, I w- I would say is that for me personally seems to be one of the big seminal moments in all of software engineering ever.

    2. RM

      (laughs) .

    3. LF

      I think that's a ... When a large company like Google decides to take a large project that many lawyers might argue has a lot of IP, just decide to go open source with it, and in so doing, lead the entire world in saying, "You know what? Open innovation is, is, is a pretty powerful thing, and it's okay to do." (laughs) Uh, that, that was ... I mean, that's an, uh, that's an incredible, credible moment in time. So, do you remember those discussions happening?

    4. RM

      Yeah.

    5. LF

      Whether open source should be happening? What was that like?

    6. RM

      I would say, I think I ... So the, the initial idea came from Jeff, who was a big proponent of this. I think it came off of two big things. Uh, one was, research-wise, we were a research group. We were putting all our research out there, if you wanted to ... We were building on others' research, and we wanted to push the state of the art forward, and part of that was to share the research. That's how I think deep learning and machine learning has really grown so fast.So the next step was, okay, now would software help with that? And it seemed like there were existing a few libraries out there, Theano being one, Torch being another, and a few others, but they were all done by academia, and so the level was, was significantly different. The other one was, from a software perspective, Google had done lots of software or that we used internally, you know, and we published papers. Often, there was an open source project that came out of that, that somebody else picked up that paper and implemented, and they were very successful. Back then, it was like, "Okay, there's Hadoop, which has come off of tech that we built." We know the tech we've built is way better for a number of different reasons. We've, you know, invested a lot of effort in that. And turns out, we have Google Cloud and we are now not really providing our tech, but we are saying, "Okay, we have Bigtable, which is the original thing. We are gonna now provide HBase APIs on top of that, which isn't as good, but that's what everybody's used to." So there's, there's like, can we make something that is better and really just provide... Helps the community in lots of ways, but also helps push the right... a good standard forward.

  5. 7:187:47

    TensorFlow + Cloud: open everywhere, optimized integrations on Google Cloud

    1. LF

      So how does cloud fit into that? There's a TensorFlow open source-

    2. RM

      Right.

    3. LF

      ... library. And how does the fact that you can, uh, use so many of the resources that Google provides in the cloud fit into that strategy?

    4. RM

      So, so TensorFlow itself is open and you can use it anywhere, right? And we wanna make sure that continues to be the case. On Google Cloud, we do make sure that there's lots of integrations with everything else, and we wanna make sure that it works really, really well there, so...

  6. 7:4711:47

    TensorFlow’s early design decisions (2014–2015): production, hardware diversity, mobile, customization

    1. LF

      You're leading the TensorFlow effort. Can you tell me the history and the timeline of TensorFlow project in terms of major design decisions? So like the open source decision, but really, uh, you know, what to include and not. There's this incredible ecosystem that I'd like to talk about.

    2. RM

      Yeah.

    3. LF

      There's all these parts, but what, uh, if you just... Some sample moments that, uh, de- defined what TensorFlow eventually became through its... I don't know if you're allowed to say history when it's just...

    4. RM

      (laughs)

    5. LF

      But in deep learning, everything moves so fast-

    6. RM

      Yeah.

    7. LF

      ... and just a few years is, is already history.

    8. RM

      Yes, yes. So looking back, we were building TensorFlow, I guess we open sourced it in 2015, November 2015. We started on it in summer of 2014, I guess, and somewhere like 3 to 6... late 2014, by then we had decided that, okay, there's a high likelihood we'll open source it. So we started thinking about that and making sure we're, we're heading down that path. At that point, by that point, we had seen a few, you know, lots of different use cases at Google. So there were things like, okay, yes, we want to run in... at large scale in the data center. Yes, we need to support different kind of hardware. We had GPUs at that point. We had our first TPU at that point, or it was about to come out, you know, roughly around that time. So the design sort of included those. We had started to push on mobile, so we were running models on mobile. At that point, people were customizing code, so we wanted to make sure TensorFlow could support that as well. So that, that sort of, uh, became part of that overall design.

    9. LF

      W- when you say mobile, you mean like pretty complicated algorithms running on the phone?

    10. RM

      That's correct. So, so when you have a model that you deployed on the phone and run it there, right?

    11. LF

      So already at that time, there was ideas of running machine learning on the phone?

    12. RM

      That's correct. We already had a couple of products that were doing that by then.

    13. LF

      Right.

    14. RM

      And, and in those cases we had basically customized handcrafted code or, or some internal libraries that we're using.

    15. LF

      So I was actually at Google during this time in a parallel, I guess, universe.

    16. RM

      (laughs)

    17. LF

      But we were using Theano and Caffe.

    18. RM

      Yeah.

    19. LF

      W- d- w- was there some degree to which you were bouncing ide- like trying to see what Caffe was offering people, trying to see what Theano was offering, that you want to make sure you're delivering on whatever that is?

    20. RM

      Yeah.

    21. LF

      Perhaps the Python part of thing, maybe. Uh, did that influence any design decisions?

    22. RM

      Um, totally. So when we built this belief, and, and some of that was in parallel with some of these libraries coming up. I mean, Theano itself is older, but we were buil- building this belief focused on our internal thing, because our systems were very different. By the time we got to this, we looked at, um, a number of libraries that were out there. Theano, there were folks in the group who had experience with Torch, with Lua. There were folks here who had seen Caffe. I mean, actually Yangqing was here as well. There's, uh, what other libraries? I think we looked at a number of things, might even have looked at Chainer back then. I'm trying to remember if it was there. In fact, yeah, we did discuss ideas around, okay, should we have a graph or not?

    23. LF

      Mm-hmm.

    24. RM

      And, uh, they were... So, so putting all these together was definitely, you know, they were key decisions that we wanted. We, we had seen limitations in our prior disbelief things. A few of them were just in terms of research was moving so fast, we wanted the flexibility. Uh, we want... The hardware was changing fast, we expected to change that. So that, those probably were two things. And yeah, I think the flexibility in terms of being able to express all kinds of crazy things was definitely a big one then.

  7. 11:4714:07

    Graph vs eager: why TensorFlow started graph-first and how TF 2.0 changes the default experience

    1. LF

      So what... The, the graph decision.

    2. RM

      Mm-hmm.

    3. LF

      So that was... Moving towards TensorFlow 2.0-

    4. RM

      Yeah.

    5. LF

      ... there's, uh, more... By default it will be eager, eager execution. So sort of hiding the graph a little bit-

    6. RM

      Yeah.

    7. LF

      ... because it's less intuitive in terms of, uh, the way people develop and so on. What was that discussion like with in terms of, uh, using graphs? It seemed- It's kind of the Theano way, uh, did it seem the obvious choice?

    8. RM

      So, I think where it came from was... or, like, DisBelief had a graph-like-

    9. LF

      Okay.

    10. RM

      ... thing as well. A much more sim- ... It wasn't a general graph, it was more like a straight line, you know, thing. More like what you might think of Caffe, I guess, in that sense? But the graph was... And we always wa- cared about the production stuff. Like, even with DisBelief, we were deploying a whole bunch of stuff in production. So, the graph did come from that when we thought of, "Okay, should we do that in Python?" And, and we experimented with some ideas where it looked a lot simpler to use. But not having a graph meant, "Okay, how do you deploy now?" So that was probably what filtered the balance for us, and eventually we ended up with the graph (laughs) .

    11. LF

      And I guess the question there is, did you (laughs) ... I mean, uh, so production seems to be the... really good thing to f- to focus on. But did you even anticipate the other side of it where there could be, uh... What is it, what are the numbers? Something crazy. Uh, 41 million downloads?

    12. RM

      Yep (laughs) .

    13. LF

      TensorFlow (laughs) , uh, d- did... I mean, was that even, like, a possibility in your mind, that- that it would be as popular as it became?

    14. RM

      So, I- I think we- we did see a need for this a lot from the research perspective, and, like, early days of deep learning in some ways. 41 million, no, I don't think I imagined this number then (laughs) . Uh, there- it- it seemed like there's a potential future where lots more people would be doing this, and how do we en- enable that? I- I would say this kind of growth, uh, probably started seeing somewhat after the open-sourcing, where it was like, "Okay, you know, deep learning is actually growing way faster for a lot of different reasons, and we are in just the right place to push on that, and- and leverage that, and- and deliver on lots of things that people want."

  8. 14:0718:07

    After open sourcing: explosive adoption, documentation as a catalyst, and the road to 1.0 stability

    1. LF

      So what changed once you open-sourced? Like, how... You know, this incredible amount of attention from a global population of developers, what... How did the project start changing? I don't think you actually remember, uh, uh, during those times. I- I know looking now, there's really good documentation, there's an ecosystem of tools.

    2. RM

      Yep.

    3. LF

      There's a community blog (laughs) , there's a YouTube channel now, right?

    4. RM

      (laughs) Yeah.

    5. LF

      (laughs) It's very, uh, very community-driven. Uh, back then, uh, I guess 0.1 version.

    6. RM

      (laughs) .

    7. LF

      Is that the version when it opens-

    8. RM

      I- I think we called it 0.6 or 5, something like that, I forget-

    9. LF

      Something like that.

    10. RM

      ... but, uh-

    11. LF

      What changed, leading into 1.0?

    12. RM

      It- it's interesting. You know, I think we've gone through a few things there. When we started out, when we first came out, people loved the documentation we have.

    13. LF

      Mm-hmm.

    14. RM

      Because it was just a huge step up from everything else-

    15. LF

      Yeah.

    16. RM

      ... because all of those were academic projects, people doing... You know, we don't think about documentation. Um, I think what that changed was instead of deep learning being a research thing-

    17. LF

      Mm-hmm.

    18. RM

      ... some people who were just developers could now suddenly take this out and do some interesting things with it-

    19. LF

      Mm-hmm.

    20. RM

      ... like, who had no clue what machine learning was before then. Um, and that, I think, really changed how things started to scale up in some ways, and- and, uh, pushed on it. Over the next few months, as we looked at, you know, "How do we stabilize things?" As we look at not just researchers now, we want stability, people who want to deploy things. That's how we started planning for 1.0. And there are certain needs for that perspective. And so, again, documentation comes up, um, designs, more kinds of things to put that together. And so, that was exciting to get that to a stage where more and more enterprises wanted to buy in and really get behind that. And I think post-1.0 and, you know, with the next few re- releases, that enterprise adoption also started to take off. I would say between the initial release and 1.0, it was... Okay, researchers, of course, uh, then a lot of hobbyists and early interest- people excited about this who started to get on board, and then over the 1.X thing, lots of enterprises.

    21. LF

      I imagine anything that's, you know, below 1.0-

    22. RM

      (laughs) .

    23. LF

      ... gets pressured to be, uh... Yeah, everybody probably wants something that's stable.

    24. RM

      Exactly.

    25. LF

      And, uh, do you have a sense now that TensorFlow is sta-

    26. RM

      (laughs) .

    27. LF

      Like, it feels like deep learning in general is extremely dynamic field, uh, so much is changing. Do you ever... And TensorFlow has been growing incredibly. Do you have a sense of stability-

    28. RM

      (laughs) .

    29. LF

      ... at the helm of this? I mean, I know you're in the midst of it, but...

    30. RM

      Yeah. It- it's... Yeah, it's... I think in the midst of it, it's often easy to forget what, uh, an enterprise wants and what some of the people, uh, uh, on that side want. There are still people running models that are three years old, four years old.

  9. 18:0722:00

    What real users need: transfer learning for hobbyists vs structured-data pipelines for enterprises (TFX)

    1. LF

      So I imagine, maybe you can correct me if I'm wrong, one of the biggest use cases is essentially taking something like ResNet-50 and doing some kind of, uh, transfer learning on a very particular problem that you have. I- it's basically probably what majority of the world does.... and you want to make that as easy as possible, say?

    2. RM

      Yes. So, so I would say for the hobbyist perspective, that's the most common case, right? In fact, the apps on phones and stuff that you will see, the early ones, that's the most common case. I, I would say there are a couple of reasons for that. One is that everybody talks about that.

    3. LF

      Mm-hmm.

    4. RM

      It looks great on slides.

    5. LF

      Yeah.

    6. RM

      That's a part of a great presentation.

    7. LF

      Visual. Yeah.

    8. RM

      Yeah, exactly. What enterprises want is ... That is part of it, but that's not the big thing. Enterprises really have data that they, they want to make predictions on. This is often what they used to do with the people who were doing ML was just regression models, linear regression, logistic regression, linear models, or, uh, maybe gradient booster trees and so on. Some of them still benefit from deep learning, but they want that, that ... That's the bread and butter, like the structured data and so on. So depending-

    9. LF

      Right.

    10. RM

      ... on the audience you look at, they're a little bit different there.

    11. LF

      And they just have ... I mean, the best of enterprise probably just has a very large dataset where deep learning can probably shine.

    12. RM

      That's correct.

    13. LF

      Right.

    14. RM

      That's right. The- and then they ... I think the other pieces that they want, like again at the 2.0 or the, the developer summit we put together, is the, the whole TensorFlow Extended piece-

    15. LF

      Mm-hmm.

    16. RM

      Which is the entire pipeline. They care about stability across doing their entire thing. They want simplicity across the entire thing. I don't need to just train a model. I need to do that every day again, over and over again.

    17. LF

      I wonder to which degree you have a role in, uh ... I don't know. So I teach a course on deep learning. Uh, the ... I have people, like lawyers come up to me and say, uh, you know, say, "When is machine learning gonna enter legal, the legal realm?"

    18. RM

      (laughs)

    19. LF

      The same thing in, um, all, all, all kinds of disciplines, uh, eh, immigration, insurance. Often when I see what it boils down to is, these companies are often a little bit old school in the way they organize the data. So the data is just not ready yet. It's not digitized.

    20. RM

      Yeah.

    21. LF

      Do you also find yourself being in the role of an, an, an evangelist for like, uh, let's get ... Organize your data, folks, and then you'll get the big benefit of TensorFlow? Do you get those, have those conversations?

    22. RM

      So yeah, yeah. I, I ... You know, I get all kinds of questions there from, "Okay, what can I ... What do I need to make this work," right?

    23. LF

      Right.

    24. RM

      "Do, do we really need deep learning? I mean, there are all these things ... I already use this linear model. Why would this help? I don't have enough data," let's say. You know, or, "Uh, I want to use machine learning, but I have no clue where to start." So, so it varies. Start to all the way to the experts who ask for very specific things. So it, it's interesting.

    25. LF

      Is there a good answer? Is ... It boils down to oftentimes digitizing data. So whatever you want automated, the ... Whatever data you're wanting to make prediction based on, you have to make sure that it's in an organized form. You've d- ... Like with the ... Within the TensorFlow, uh, ecosystem, there's now ... You're providing more and more datasets and more and more pretrained models. Are you finding yourself also the organizer of datasets?

    26. RM

      Yes. I think with TensorFlow datasets that we just released-

    27. LF

      Right.

    28. RM

      ... that's definitely come up where people want these datasets. Can we organize them and can we make that easier? And so, so that's, that's definitely one important thing. The other related thing I would say is I often tell people, "You know what? Don't think of the most fanciest thing that ... the newest model that you see."

    29. LF

      Mm-hmm.

    30. RM

      "Make something very basic work, and then you can improve it. There's just lots of things you can do with it."

  10. 22:0026:23

    Keras becomes the front door: how it joined TensorFlow and why TF 2.0 standardizes on it

    1. LF

      One of the big things that makes it s- makes TensorFlow even more accessible was the appearance, whenever that happened, of Keras.

    2. RM

      Mm-hmm.

    3. LF

      The Keras Standard, sort of, uh, outside of TensorFlow.

    4. RM

      Yeah.

    5. LF

      I, I think it was, uh, Keras on top of, um, Theano at first only, and then, uh, m- Keras became on top of TensorFlow. Do you know when, uh, Keras chose to also, uh, add TensorFlow as a backend? Who was the ... Was it just the community that drove that initially? Do you know if there was, uh, discussions, conversations?

    6. RM

      (laughs) Yeah. So Francois started the Keras project, uh, before he was at Google, and the first thing was Theano. I don't remember if that was after TensorFlow was created or way before. And then, at some point, when TensorFlow started becoming popular, there were enough similarities that he decided to, okay, create this interface and, and put TensorFlow as a backend. Um, I believe that might still have been before he joined Google, so I ... You know-

    7. LF

      (laughs)

    8. RM

      We weren't really talking about that. He decided on his own and, uh, thought that was interesting and relevant to the community. In fact, I didn't find out about him being at Google until a few months after he was here. He was working on some research ideas and doing Keras on his nights and weekends project and stuff.

    9. LF

      Oh, interesting.

    10. RM

      Yeah.

    11. LF

      So he wasn't like part of the TensorFlow ... He didn't join initially.

    12. RM

      He joined research and he was doing-

    13. LF

      Okay. Interesting.

    14. RM

      ... some amazing research. He has some papers on that, on research that he, he's done.

    15. LF

      Yeah.

    16. RM

      He's a great researcher as well.

    17. LF

      Yeah.

    18. RM

      And at some point we realized, oh, he's, he's doing this good stuff. People seem to like the API and he's right here. So we talked to him and he said, "Okay, why don't I come over to your team and work with you for a quarter-"

    19. LF

      Mm-hmm.

    20. RM

      "... and let's make that integration happen." And we talked to his manager and he said, "Sure, my quarter's fine" (laughs) . And that quarter's been something like two years now.

    21. LF

      (laughs)

    22. RM

      (laughs) And so he, he's fully on this.

    23. LF

      So Keras got integrated into TensorFlow like in, in a deep way.

    24. RM

      Yeah.

    25. LF

      And now with 2.0, with TensorFlow 2.0, sort of Keras is kinda the recommended way for a beginner to interact with TensorFlow.

    26. RM

      That's right.

    27. LF

      Which makes that initial sort of transfer learning or the basic use cases, even for a enterprise, uh, s- super simple, right?

    28. RM

      That's correct. That's right.

    29. LF

      So what, what was that decision like? That seems like a...... a, it's kind of a bold decision, uh, (laughs) as well.

    30. RM

      Yeah. (laughs) We did spend a lot of time thinking about that one. We had a bunch of APIs, some built by us. Uh, there was a parallel layers API that we were building, and when we decided to do Keras in parallel, so they were like, "Okay, two things that we are looking at." And the first thing we was trying to do is just have them look similar-

  11. 26:2328:03

    Open-source governance at scale: no single ‘BDFL’, more transparency via RFCs and SIGs

    1. LF

      So, uh, Python has, uh, Guido van Rossum, who, uh, until recently held the position of Benevolent Dictator-

    2. RM

      (laughs)

    3. LF

      ... for Life. Right? So, but does a huge, successful open source project like TensorFlow need one person who makes a final decision? So you've d- did a pretty successful, uh, TensorFlow dev summit just now, last couple of days. There's clearly a lot of different new features, uh, being incorporated, an amazing ecosystem, so on. Who's, uh... How are those design decisions made? Is there, is there a BDFL in TensorFlow, and, uh, or is it more distributed and, uh, organic?

    4. RM

      I think it's- it's somewhat different, I would say. I've always been involved in the key design directions, but there are lots of things that are distributed where, uh, there are a number of people, Martin Wick being one who has really driven a lot of our open source stuff, a lot of the APIs. Um, and there, there are a number of other people who have been, you know, pushed and been responsible for different parts of it. We do have regular design reviews. Over the last year, we've really spent a lot of time opening up to the community and adding transparency. We're setting more processes in place, so RFCs, special interest groups, to really grow that community and, and scale that. I think the kind of scale that ecosystem is in, I don't think we could scale with having me as the-

    5. LF

      (laughs)

    6. RM

      ... sole point of decision-making, so...

    7. LF

      I got it. So-

    8. RM

      Yeah.

  12. 28:0332:34

    The ecosystem vision: ML on every device + tooling cohesion (SavedModel, Hub, Lite, JS, TFX)

    1. LF

      So, yeah, the growth of that ecosystem, may- maybe you can talk about it a little bit. First of all, when I, uh... It started with Andrej Karpathy when he first did convnet.js. The fact that you can, uh, train in your own network in the browser was, in, uh, JavaScript was incredible.

    2. RM

      Yup.

    3. LF

      So now TensorFlow JS is really making that a serious, like, uh, a legit thing, uh, a way to operate, whether it's in the back end or the front end. Then there's the TensorFlow Extended, like you've mentioned. There's the TensorFlow Lite for mobile.

    4. RM

      Yup.

    5. LF

      And all of it, as far as I can tell, it's really converging towards being able to, uh, you know, save models in the same kind of way. You can move around, you can train on the desktop and then move it to mobile and so on. Like-

    6. RM

      That's right.

    7. LF

      So there's that cohesiveness. So, uh, can you maybe give me, whatever I missed-

    8. RM

      (laughs)

    9. LF

      ... a bigger overview of the mission of the ecosystem that's trying to be built, and where is it moving forward?

    10. RM

      Yeah. So in short, the way I like to think of this is, our goal is to enable machine learning and in couple of ways. You know, one is, we have lots of exciting things going on in ML today. Uh, we started with deep learning, but we now support a bunch of other algorithms too. So, so one is to, on the research side, keep pushing on the state of the art. Can we... You know, how do we enable researchers to build the next amazing thing? So, BERT came out recently. You know, it's great that people are able to do new kinds of research. And there are lots of, you know, amazing research that happens across the world. Uh, so that's one direction. The other is, how do you take that across all the people outside who want to take that research and do some great things with it and integrate it to build real products, to- to have a real impact on people? Um, and so that's the other axis in some ways. Um, you know, at a high level, one way I think about it is, there are a crazy number of compute devices across the world. And we often used to think of ML and training and all of this as, okay, something you do either in the workstation or the data center or cloud, but we see things running on the phones, we see things running on really tiny chips. I mean, we had some demos at the developer summit. And so the way I've... think about this ecosystem is, how do we help get machine learning on every device that has a compute capability?

    11. LF

      Right.

    12. RM

      And that continues to grow. And, and so, uh, in some ways, this ecosystem has looked at, you know, various aspects of that and grown over time to cover more of those, and we continue to push the boundaries. In some areas, we've built, um, more tooling and things around that to help you. I mean, the first tool we started was TensorBoard, if you wanted to learn just the training piece. Um, TFX or TensorFlow Extended to really do your entire ML pipelines if you're, you know, care about all that production stuff. Uh, but then going to the edge, going to different kinds of things. And it's not just us now. Um, we are at a place where, you know, there are lots of libraries being built on top. So there are some for research, maybe things like TensorFlow Agents or TensorFlow Probability that started as research things or for researchers for focusing on certain kinds of algorithms, but they're also being deployed or used by, you know, production folks. And, uh, some have come from within Google, just teams across Google who wanted to, to build these things. Others have come from just the community, because there are different pieces that different parts of the community care about, and I, I see our goal as enabling even that, right? It's not we, we cannot and won't build every single thing. That just doesn't make sense. But if we can enable others to build the things that they care about, and there's n- a broader community that cares about that, and we can help encourage that, and, um, that, that's great. That really helps the entire ecosystem, not just those. Uh, one of the big things about 2.0 that we're pushing on is, okay, we have these so many different pieces, right?

    13. LF

      Mm-hmm.

    14. RM

      How do we help make all of them work well together? So there are few key pieces there that we're pushing on, uh, one being the core format in there and how we share the models themselves through SavedModel and Mo- TensorFlow Hub and so on, um, and, you know, a few of the pieces that we really put this together.

  13. 32:3437:07

    Hard engineering problems: integrating new hardware, breaking up a monolith, and backward compatibility trade-offs

    1. LF

      I was very skeptical that that's... You know, when TensorFlow.js came out, it didn't seem... Or deeplearning.js, as it was earlier called.

    2. RM

      Yeah, that was the first.

    3. LF

      It seems like technically very difficult project. As a standalone, it's not as difficult, but as a thing that integrates into the ecosystem, it seems very difficult. So y- I mean, there's a lot of aspects of this you make it look easy, but, uh-

    4. RM

      (laughs)

    5. LF

      ... on the technical side, how many challenges have to be overcome here?

    6. RM

      A lot. (laughs)

    7. LF

      And still have to be overcome.

    8. RM

      Yes, yes.

    9. LF

      That's the question here too.

    10. RM

      There, there are lots of steps to it, right? And we reiterated over the last few years that there's a lot we've learned. I, I, yeah, n- Often when things come together well, things look easy, and that's exactly the point. It should be easy for the end user, but there are lots of things that go behind that. If I think about still, um, challenges ahead, there are... You know, we have a lot more devices coming on board, for example, from the hardware perspective. How do we make it really easy for these vendors to integrate with something like TensorFlow, right? Uh, so there's a lot of compiler stuff that others are working on. There are, uh, things we can do in terms of our APIs and so on that we can do. As we... You know, TensorFlow started as a very monolithic system, and to some extent it still is. There are less, lots of tools around it, but the core is still pretty large and monolithic. One of the key challenges for us to scale that out is how do we break that apart with clearer interfaces. It's, um... You know, in some ways, it's Software Engineering 101, but for a system that's now four years old, I guess, or more, and that's still rapidly evolving, and that we're not slowing down with, it's hard to, you know, change and modify and really break apart.

    11. LF

      Mm-hmm.

    12. RM

      It, it's sort of like, as people say, right, it's like, uh, changing the engine with the car running or trying to fix that.

    13. LF

      That's right.

    14. RM

      That's exactly what we're trying to do.

    15. LF

      So there, there's a challenge here because the downside of so many people, uh, being excited about TensorFlow and be- coming to rely on it in many of their applications is that you're kind of responsible. L- it's the technical debt. You're responsible for previous versions to some degree still working. So when you're trying to innovate, I mean, uh, it's probably easier to just start from scratch every few months. (laughs)

    16. RM

      (laughs) Absolutely. (laughs)

    17. LF

      You know? So do you feel the pain of that? Uh, 2.0 does break some back compatibility, but not too much. It seems like the conversion is pretty straightforward. Uh, do, do you think that's still important given how quickly deep learning is changing? Can you just... (laughs) The things that d- you've learned, can you just start over, or is there pressure to not?

    18. RM

      It, it's a, it's a tricky balance. So i- if it was, um, just a researcher writing a paper who a year later will not look at that code again, sure, it doesn't matter. Uh, there are a lot of production systems that rely on TensorFlow, both at Google and across the world, and people worry about this. I mean, they're- these systems run for a long time.

    19. LF

      Right.

    20. RM

      Uh, so it is important to keep that compatibility and so on. And yes, it does come with a huge cost. There's, uh... We have to think about a lot of things as we do new things and make new changes. Uh, I think the... It, it's a trade-off, right? You can, you might slow certain kinds of things down, but the overall value you're bringing because of that is, is much bigger because it's not just about breaking the person yesterday, it's also about ge- telling the person tomorrow that, "You know what? This is how we do things. We're not going to break you when you come on board-"

    21. LF

      Right.

    22. RM

      "... because there are lots of new people who are also going to come on board."

    23. LF

      Right.

    24. RM

      Um, a... You know, one way I, I like to think about this, and I always push the team to think about it as well, when you want to do new things, you want to start with a clean slate.... design with a clean slate in mind, and then we'll figure out how to make sure all the other things work. And, yes, we do make compromises occasionally. But unless you design with the clean slate and not worry about that, you'll never get to a good place.

    25. LF

      Oh, that's brilliant. So even if you're do- you are responsible when in the idea stage, when you're thinking of new-

    26. RM

      Right.

    27. LF

      ... just put all that behind you.

    28. RM

      Yes.

  14. 37:0739:42

    TensorFlow vs PyTorch: learning from competition and accelerating eager execution

    1. LF

      That's, okay, that's really, really well put. So I have to ask this because a lot of students, developers ask me how I feel about PyTorch versus TensorFlow.

    2. RM

      (laughs)

    3. LF

      So, uh, I've recently completely switched my, uh, research group to TensorFlow. I wish everybody would just use the same thing, and TensorFlow is as close to that, I believe, as we have. But, uh, do you enjoy competition?

    4. RM

      (laughs)

    5. LF

      Uh, so TensorFlow is leading in many ways, on, uh, many dimensions, in terms of ecosystem, in terms of the number of users, uh, momentum, power, production levels, so on. But, you know, a lot of researchers are now also using PyTorch.

    6. RM

      Mm-hmm.

    7. LF

      Do you enjoy that kind of competition, or do you just ignore it and focus on making TensorFlow the best that it can be?

    8. RM

      So, just like research or anything people are doing, right, it's great to get different kinds of ideas. And when we started with TensorFlow, like I was saying earlier, one, i- it was very important for us to also have production in mind. We didn't want just research, right? And that's why we chose certain things. Now, PyTorch came along and said, "You know what? I only care about research. This is what I'm trying to do. What's the best thing I can do for this?" And it started iterating and said, "Okay, I don't need to worry about graphs. Let me just run things. Um, I don't care if it's not as fast as it can be, but let me just, you know, make this part easy." And there are things you can learn from that, right? They, they again had the benefit of seeing what had come before, uh, but also, uh, exploring certain different kinds of spaces and, and they, uh, had some good things there, you know, building on, say, things like Chainer and so on before that. So, uh, competition is definitely interesting. It made us, you know, this is an area that we had thought about, like I said, you know, way early on. Over time, we had, uh, revisited this a couple of times. "Should we add this again?" Uh, at some point we said, "You know what? Here's, it seems like this can be done well, so let's try it again." And we, that's how, you know, we started pushing on eager execution and how do we combine those two together, which has finally come very well together in 2.0, but it took us a while to get all the things together and so on.

    9. LF

      So let me, I mean, ask, uh, put another way, I think eager execution is a really powerful thing that was added. Do you think it wouldn't have been, you know, Muhammad Ali versus Frazier, right?

    10. RM

      (laughs)

    11. LF

      Do you think, uh, it wouldn't have been added as quickly if PyTorch wasn't there?

    12. RM

      It, it, it might have taken longer.

    13. LF

      Longer.

    14. RM

      Yeah, yeah. It was, I mean, we had tried some variants of that before, so I'm sure it would have happened, but it might have taken longer.

  15. 39:4251:12

    Looking ahead: performance-by-default, modularity, Swift for TensorFlow, and the unpredictability of ‘TF 3.0’

    1. LF

      I'm grateful that TensorFlow has responded in the way they did. It's g- doing some incredible work last couple years. What other things that we didn't talk about are you looking forward in 2.0 that comes to mind? So we talked about some of the ecosystem stuff, making it, uh, easily accessible through Keras, eager execution. Is there other things that we missed, you think?

    2. RM

      Yeah. So I would say one is just where 2.0 is and, you know, with all the things that we've talked about. I think as we think beyond that, there are lots of other things that it enables us to do and that we're excited about. So what it's setting us up for, okay, here are these really clean APIs. We've cleaned up the surface for what the users want. What it a- also allows us to do a whole bunch of stuff behind the scenes once we've, we are ready with 2.0. So, uh, for example, in TensorFlow with graphs and all the things you could do, you could always get a lot of good performance if you spent the time to tune it, right? And we've clearly shown that, lots of people do that. With 2.0, with these APIs, where we are, we can give you a lot of performance just with whatever you do. You know, if you're, because we see these, it's much cleaner. We know most people are going to do things this way. We can really optimize for that and, and get a lot of those things out of the box. Uh, and it really allows us, you know, both for single machine and distributed and so on, to really explore other spaces behind the scenes after, you know, 2.0 and the future versions as well. So, uh, right now the team's really excited about that, that over time I think we'll see that. The, the other piece that I was talking about in terms of just restructuring the monolithic thing into more pieces and making it more modular, I think that's b- gonna be really important for, uh, a lot of the other people in the ecosystem, other organizations and so on that wanted to build things.

    3. LF

      Can you elaborate a little bit what you mean by making TensorFlow more, uh, ecosystem more modular?

    4. RM

      So the way it's organized today is there's one, there are lots of repositories in the TensorFlow organization at GitHub. The core one where we have TensorFlow, it has the execution engine, it has, you know, the key backends for CPUs and GPUs. It has, uh, the work to do distributed stuff, and all of these just work together in a single library or binary. Uh, there's no way to split them apart easily. I mean, there are some interfaces, but they're not very clean. In a perfect world, you would have clean interfaces where, "Okay, I want to run it on my fancy cluster with some custom networking, just implement this and do that." I mean, we kind of support that, but it's hard for people today. Uh, I think as we are starting to see more interesting things in some of these spaces, having that clean separation will really start to help.Um, and, and again, going to the, the large size of the ecosystem and the different groups involved there, enabling people to evolve and push on things more independently just allows it to scale better.

    5. LF

      And by people you mean individual developers and ...

    6. RM

      And organizations.

    7. LF

      And organizations?

    8. RM

      That's right.

    9. LF

      So the hope is that everybody, sort of major ... I don't know, Pepsi or something, uses ... Like, major corporations go to TensorFlow to this kind of thing?

    10. RM

      Yeah. If you look at enterprises like Pepsi or these... I mean a lot of them are already using TensorFlow. They, they are not the ones that do the development or changes in the core. Some of them do, but a lot of them don't. I mean, they touch small pieces. There are lots of these, um, some of them being, let's say, hardware vendors who are building their custom hardware and they want their own pieces.

    11. LF

      Ah, got it.

    12. RM

      Or some of them being bigger companies, say IBM. I mean, they are involved in some of our special interest groups, and they see a lot of users who want certain things and they want to optimize for that. It's folks like that, often.

    13. LF

      Autonomous vehicle companies, perhaps. (laughs)

    14. RM

      Exactly, yes.

    15. LF

      So, uh, yeah, like I mentioned, TensorFlow has been downloaded 41 million times. 50,000 commits, almost 10,000 pull requests, and 1800 contributors. So, uh, I'm not sure if you can explain it, but-

    16. RM

      (laughs) .

    17. LF

      W- wha- what does it take to build a community like that? What ... If, in retrospect, what do you think ... What, what is the critical thing that allowed for this growth to happen, and how does that growth continue?

    18. RM

      Yeah. Uh, yeah, that's a interesting question. I wish I had all the answers there-

    19. LF

      Yes.

    20. RM

      ... I guess, so we could replicate it. I, I think there's, um ... There are a number of things that need to come together, right? Um, one, m- you know, just like any new thing, it is about ... There's a sweet spot of timing, what's needed, you know, does it grow with what's needed? So in this case, for example, uh, TensorFlow has not just grown because it was a good tool, it's also grown with the growth of deep learning itself. Uh, so tho- those factors come into play. Other than that, though, I think just hearing, listening to the community, what they'll do, what they need. Being open to ... Like, in terms of external contributions, we've spent a lot of time in making sure we can accept those contributions well, we can help the contributors in, in adding those, putting the right process in place, getting the right kind of community, welcoming them, and so on. Like over the last year, we've really pushed on transparency. That, that's important for an open source project. Uh, people want to know where things are going, and we're like, "Okay, here's a process where you can do that. Here are our fees and so on." Uh, so thinking th- through ... There are lots of community asp- aspects that come into that you can really work on. As a small project, it's maybe easy to do, because there's like two developers and, and you can do those. A- as you grow, putting more of these processes in place, thinking about the documentation, thinking about, "What do developers care about? What kind of tools would they want to use?" All of these come into play, I think.

    21. LF

      So one, one of the big things, I think, that feeds the TensorFlow fire is, uh, people building something on TensorFlow. And, uh, you know, some, uh, implement a particular architecture that does something cool and useful, and then put it, that on GitHub. And so it just feeds this, uh, this growth. Do you s- have a sense that with 2.0 and 1.0 that there may be a little bit of a partitioning like there is with Python 2 and, and 3? That there will be a codebase in, in the older versions of TensorFlow that will not be as compatible easily? Or d- are you pretty confident that this kind of, uh, conversion is pretty natural and easy to do?

    22. RM

      So, we're definitely working, uh, hard to make that very easy to do. There is lots of tooling that we talked about at the developer summit this week, and we'll continue to invest in that tooling. It's ... Um, you know, when you think of these significant version changes, that's always a risk.

    23. LF

      Yeah.

    24. RM

      And we, we are really pushing hard to make that transition very, very smooth. I, I think ... So, so at some level people want to move and they see the value in the new thing. They don't want to move just because it's a new thing. I mean some people do, but most people want a, a really good thing. And I think over the next few months, as people start to see the value, we'll definitely see that shift happening. So I'm, I'm pretty excited and confident that we, we'll see people moving. Um, as you said earlier, this field is also moving rapidly, so that'll help because we can do more things, and, you know, all the new things will clearly happen in 2.X, so people will have lots of good reasons to move.

    25. LF

      So what do you think, uh, TensorFlow 3.0 looks like?

    26. RM

      (laughs) .

    27. LF

      Is th- is there a ... Are things happening so crazily that even at th- the end of this year seems impossible to plan for? Or is it possible to plan for the next five years?

    28. RM

      I, I think it's tricky. There are some things that we can expect in terms of, okay, change. Yes, change is gonna happen.

    29. LF

      (laughs) .

    30. RM

      (laughs) . Uh, are, are, are there some going, things going to stick around and some things not gonna stick around? I, I would say the, the basics of deep learning, the, you know, say, convolutional models or the, the basic kind of things, they'll probably be around in some form still in five years. Uh, will RL and GANs stay? Very likely, based on where they are. Will we have new things? Probably, but those are hard to predict and ... Some ... Directionally some things that we can see as ... You know, in, in things that we're starting to do, right, with some of our projects right now, is, uh, just 2.0 combining Eager Execution and, and Graphs where we're starting to make it more like just your natural programming language. You're not trying to program something else. Uh, similarly with Swift for TensorFlow, we are taking that approach. Can you do something round-up right? So, so some of those ideas seem like, okay, that's the right direction. In five years, we expect to see more in that area.Um, other things we don't know is, will hardware accelerators be the same? Will we be able to train with, uh, four bits instead of 32 bits? (laughs) Uh-

  16. 51:121:10:42

    Leading the project: team culture, hiring for motivation, and balancing speed vs quality and community input

    1. LF

      ... uh, in, in a sense that when they grow up, um, it's, uh, some incredible ideas will be coming from them. So there's certainly a technical aspect to your work, but you also have a management aspect to your role with Tensor Flow, leading the project, uh, large number of developers and people. So what do you look for in a good team? What do you think? You know, Google has been at the forefront of exploring what it takes to build a good team, and Tensor Flow is one of the most cutting-edge technologies in the world. So in this context, what do you think makes for a good team?

    2. RM

      It's definitely something I think a fair bit about. I think the, i- in terms of, you know, the team being able to deliver something well, one of the things that's important is, uh, a cohesion across the team, so being able to execute together in doing things. It's not an en-... Like at this scale, an individual engineer can only do so much. There's a lot more that they, they can do together, even though we have some amazing superstars across Google and in the team. But there's e- you know, often the way I see it is the product of what the team generates is way larger than, um, the whole or, you know, the, uh, e- uh, each individual put together. And so how do we have all of them work together, the, the culture of the team itself? Um, hiring good people i- is important, uh, but part of that is it's not just that, okay, we hire a bunch of smart people and throw them together and let them do things. It's also people have to care about what they're building, people have to be motivated f- for the right kind of things. Uh, that's often an important factor. Um, and, you know, finally, how do you put that together with a somewhat unified vision of where we want to go? So are we all looking in the same direction or each of us going all over? A- and sometimes it's a mix. Uh, Google's a very bottom-up organization in some sense, um, also research even more so, uh, and that's how we started. But as we've become this larger product and ecosystem, I think it's also important to combine that well with a mix of, okay, here's the direction we want to go in. There is exploration we'll do around that, but let's keep staying in that direction, not just all over the place.

    3. LF

      And is there a way you monitor the health of the team, sort of like, is, is there a way you know you did a good job?

    4. RM

      (laughs)

    5. LF

      (laughs) The team is good. Like, uh, I mean, you're sort of, uh, you're saying nice things, but it's sometimes difficult to determine-

    6. RM

      Yes.

    7. LF

      ... how, how aligned-

    8. RM

      Yes.

    9. LF

      ... 'cause it's not binary.

    10. RM

      No.

    11. LF

      It's not like ev- it's, it's- there's tensions and complexities and so on. And the other element of it is, you mention superstars, you know, there's so much, even at Google, such a large percentage of work is done by individual superstars too. So there's a-

    12. RM

      Yeah.

    13. LF

      ... uh, and sometimes those superstars can be against the dynamic of a team in those, those tensions. Have, was that, has that, I mean, I'm sure in Tensor Flow it might be a little bit easier because the mission of the project is so, sort of beautiful. You're, you're at the cutting edge, so it's exciting.

    14. RM

      Yeah.

    15. LF

      Uh, but have you had struggle with that? Has there been challenges?

    16. RM

      There are always people challenges in different kinds of ways.

    17. LF

      (laughs) I'm into this.

    18. RM

      That, that said, I think we've been... What's good about getting people who care and are, you know, uh, have the same kind of culture, and that's Google in general to a large extent, but also like you said, given that the project has had so many exciting things to do, there's been room for lots of people to do different kinds of things and grow, which, which does make the problem a bit easier, I guess.

    19. LF

      Yeah.

    20. RM

      And it allows people, depending on what they're doing, if there's room around them, then that's fine. But, yes, we do, we do care about whether a superstar or not, that they need to work well with the team across Google. It's not something that-

    21. LF

      That's interesting to hear. So it's like, um, superstar or not, the productivity broadly is about the team.

    22. RM

      Yeah.

    23. LF

      Yeah.

    24. RM

      Yeah. I mean, they, they might add a lot of value, but if they're hurting the team, then that's a problem.

    25. LF

      So in hiring engineers, it's so interesting, right? The, the hiring process. Wh- what do you look for? How do you determine a good developer or a good member of a team from just a few minutes or hours together?

    26. RM

      (laughs) So, yeah, I guess-

    27. LF

      Again, no magic answers, I'm sure, but-

    28. RM

      Yeah, yeah. I mean, Google has a, a hiring process that we've refined over the last 20 years, I guess, and that you've probably heard and seen a lot about. So we, we do work with the same hiring process, and that, that's really helped. For me in particular, I would say, in addition to the, the core technical skills, what does matter is their motivation in what they want to do. Because if that doesn't align well with where we want to go, that's not going to lead to long-term success, for either them or their team.

    29. LF

      Mm-hmm.

    30. RM

      Um, and I think that becomes more important the more senior the person is, but it's important at every level. Like even the junior-most engineer, if they're not motivated to do well at what they're trying to do, however smart they are, it's going to be hard for them to succeed.

Episode duration: 1:10:57

Install uListen for AI-powered chat & search across the full episode — Get Full Transcript

Transcript of episode NERNE4UThHU

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.