Skip to content
Rohit Prasad: Amazon Alexa and Conversational AI | Lex Fridman Podcast #57
This video isn’t embeddableWatch on YouTube →
Lex Fridman PodcastLex Fridman Podcast

Rohit Prasad: Amazon Alexa and Conversational AI | Lex Fridman Podcast #57

Rohit Prasad is the vice president and head scientist of Amazon Alexa and one of its original creators. Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep57-sb See below for timestamps, and to give feedback, submit questions, contact Lex, etc. *CONTACT LEX:* *Feedback* - give feedback to Lex: https://lexfridman.com/survey *AMA* - submit questions, videos or call-in: https://lexfridman.com/ama *Hiring* - join our team: https://lexfridman.com/hiring *Other* - other ways to get in touch: https://lexfridman.com/contact *OUTLINE:* 0:00 - Introduction 4:34 - Her 6:31 - Human-like aspects of smart assistants 8:39 - Test of intelligence 13:04 - Alexa prize 21:35 - What does it take to win the Alexa prize? 27:24 - Embodiment and the essence of Alexa 34:35 - Personality 36:23 - Personalization 38:49 - Alexa's backstory from her perspective 40:35 - Trust in Human-AI relations 44:00 - Privacy 47:45 - Is Alexa listening? 53:51 - How Alexa started 54:51 - Solving far-field speech recognition and intent understanding 1:11:51 - Alexa main categories of skills 1:13:19 - Conversation intent modeling 1:17:47 - Alexa memory and long-term learning 1:22:50 - Making Alexa sound more natural 1:27:16 - Open problems for Alexa and conversational AI 1:29:26 - Emotion recognition from audio and video 1:30:53 - Deep learning and reasoning 1:36:26 - Future of Alexa 1:41:47 - The big picture of conversational AI *PODCAST LINKS:* - Podcast Website: https://lexfridman.com/podcast - Apple Podcasts: https://apple.co/2lwqZIr - Spotify: https://spoti.fi/2nEwCF8 - RSS: https://lexfridman.com/feed/podcast/ - Podcast Playlist: https://www.youtube.com/playlist?list=PLrAXtmErZgOdP_8GztsuKi9nrraNbKKp4 - Clips Channel: https://www.youtube.com/lexclips *SOCIAL LINKS:* - X: https://x.com/lexfridman - Instagram: https://instagram.com/lexfridman - TikTok: https://tiktok.com/@lexfridman - LinkedIn: https://linkedin.com/in/lexfridman - Facebook: https://facebook.com/lexfridman - Patreon: https://patreon.com/lexfridman - Telegram: https://t.me/lexfridman - Reddit: https://reddit.com/r/lexfridman

Lex FridmanhostRohit Prasadguest
Dec 14, 20191h 45mWatch on YouTube ↗

EVERY SPOKEN WORD

  1. 0:00 – 4:34

    Introduction

    1. LF

      The following is a conversation with Rohit Prasad. He's the vice president and head scientist of Amazon Alexa, and one of its original creators. The Alexa team embodies some of the most challenging, incredible, impactful, and inspiring work that is done in AI today. The team has to both solve problems at the cutting edge of natural language processing, and provide a trustworthy, secure, and enjoyable experience to millions of people. This is where state of the art methods in computer science meet the challenges of real-world engineering. In many ways, Alexa and the other voice assistants are the voices of artificial intelligence to millions of people, and an introduction to AI for people who have only encountered it in science fiction. This is an important and exciting opportunity. So the work that Rohit and the Alexa team are doing is an inspiration to me, and to many researchers and engineers in the AI community. This is the Artificial Intelligence Podcast. If you enjoy it, subscribe on YouTube, give it five stars on Apple Podcast, support it on Patreon, or simply connect with me on Twitter @lexfridman, spelled F-R-I-D-M-A-N. If you leave a review on Apple Podcast especially, but also Castbox, or comment on YouTube, consider mentioning topics, people, ideas, questions, quotes, in science, tech, or philosophy that you find interesting, and I'll read them on this podcast. I won't call out names, but I love comments with kindness and thoughtfulness in them, so I thought I'd share them. Someone on YouTube highlighted a quote from the conversation with Ray Dalio, where he said that you have to appreciate all the different ways that people can be A players. This connected with me, too. Uh, on teams of engineers, it's easy to think that raw productivity is the measure of excellence, but there are others. I've worked with people who brought a smile to my face every time I got to work in the morning. Their contribution to the team is immeasurable. I recently started doing podcast ads at the end of the introduction. I'll do one or two minutes after introducing the episode, and never any ads in the middle that break the flow of the conversation. I hope that works for you and doesn't hurt the listening experience. This show is presented by Cash App, the number one finance app in the App Store. I personally use Cash App to send money to friends, but you can also use it to buy, sell, and deposit bitcoin in just seconds. Cash App also has a new investing feature. You can buy fractions of a stock, say $1 worth, no matter what the stock price is. Brokerage services are provided by Cash App Investing, a subsidiary of Square and member SIPC. I'm excited to be working with Cash App to support one of my favorite organizations called FIRST, best known for their FIRST Robotics and LEGO competitions. They educate and inspire hundreds of thousands of students in over 110 countries, and have a perfect rating on Charity Navigator, which means the donated money is used to maximum effectiveness. When you get Cash App from the App Store or Google Play and use code LEXPODCAST, you'll get $10, and Cash App will also donate $10 to FIRST, which again, is an organization that I've personally seen inspire girls and boys to dream of engineering a better world. This podcast is also supported by ZipRecruiter. Hiring great people is hard, and to me, is one of the most important elements of a successful mission-driven team. I've been fortunate to be a part of and lead several great engineering teams. The hiring I've done in the past was mostly through tools we built ourselves, but reinventing the wheel was painful. ZipRecruiter is a tool that's already available for you. It seeks to make hiring simple, fast, and smart. For example, Codable co-founder Gretchen Heubner used ZipRecruiter to find a new game artist to join her education tech company. By using ZipRecruiter's screening questions to filter candidates, Gretchen found it easier to focus on the best candidates, and finally hiring the perfect person for the role, in less than two weeks from start to finish. ZipRecruiter, the smartest way to hire. See why ZipRecruiter is effective for businesses of all sizes by signing up, as I did, for free at ziprecruiter.com/lexpod. That's ziprecruiter.com/lexpod. And now, here's my conversation with Rohit Prasad.

  2. 4:34 – 6:31

    Her

    1. LF

      In the movie Her, I'm not sure if you've ever seen.

    2. RP

      Yeah.

    3. LF

      A human falls in love with the voice of an AI system. Let's start at the highest philosophical level before we get to deep learning and some of the fun things. Do you think this, what the movie Her shows, is within our reach?

    4. RP

      I think, uh, not specifically about Her, but I think what we are seeing is a massive increase in adoption of AI assistance or AI in all parts of our social fabric. And I think it's... What I do believe is that the utility these AIs provide, some of the functionalities that sh- are shown are absolutely within reach.

    5. LF

      So the, some of the functionalities in terms of the interactive elements, but in terms of the deep connection that's purely voice-based, do you think such a close connection is possible with voice alone?

    6. RP

      It's been a while since I saw Her, but I would say in terms of the, uh, in terms of interactions which are both humanlike and, in these AI assistants, you have to value what is also superhuman. We, as humans, can be in only one place. AI assistants can be in multiple places at the same time. One with you on your mobile device, one at your home, one at work. So you have to respect these superhuman capabilities too.Plus, as humans, we have certain attributes we are very good at. Very good at reasoning. AI assistance, not yet there. Uh, but in the realm of AI assistance, what they're great at is computation, memory, it's infinite and pure. These are the attributes you have to start respecting. So, I think the comparison with human-like versus the other, other aspect, which is also super human, has to be taken into consideration. So, I think we need to elevate the discussion to not just human-like.

    7. LF

      So, there's certainly elements where you just mentioned,

  3. 6:31 – 8:39

    Human-like aspects of smart assistants

    1. LF

      Alexa is everywhere, uh, computationally speaking. So, this is a much bigger infrastructure than just the thing that sits there in the room with you. But it certainly feels, to us mere humans, that there's just another little creature there when you're interacting with it. You're not interacting with the entirety of the infrastructure, you're interacting with the device. The feeling is, okay, sure, we anthropomorphize things, but, uh, that feeling is still there. So, what do you think we, as humans, the purity of the interaction with a smart assistant, what do you think we look for in, in that interaction?

    2. RP

      I think in the certain interactions, I think will be very much where it does feel like a human, uh, because it has a persona of its own. And in certain ones, it wouldn't be. So, I think a simple example to think of it is if you're walking through the house and you just wanna turn on your lights on and off, and you're issuing a command, that's not very much like a human-like interaction, and that's where the AI shouldn't come back and have a conversation with you. Just, it should simply complete that command. Uh, so those, I think the blend of we have to think about this as not human-human alone.

    3. LF

      Mm-hmm.

    4. RP

      It is a human-machine interaction, and certain aspects of humans are needed, and certain aspects or in situations demand it to be like a machine.

    5. LF

      So, I told you, it's gonna be philosophical-

    6. RP

      (laughs) .

    7. LF

      ... in parts. Uh, what, what's the difference between human and machine in that interaction? When we interact two humans, especially those that are friends and loved ones, versus you and a machine that you also are close with?

    8. RP

      I think the, uh, you have to think about the roles the AI plays, right? So, and it differs from different customer to customer, different situation to situation. Uh, especially I can speak from Alexa's perspective, it is a companion, a friend at times, an assistant, and an advisor down the line. So, I think most AIs will have this kind of attributes, and it will be very situational in nature. So, where is the boundary? I think the boundary depends on exact context in which you're interacting with the AI.

  4. 8:39 – 13:04

    Test of intelligence

    1. RP

    2. LF

      So, the depth and the richness of natural language conversation has been, by Alan Turing, been used to try to define what it means to be intelligent.

    3. RP

      Mm-hmm.

    4. LF

      You know, there's a lot of criticism of that kind of test, but w- what do you think is a good test of intelligence, in your view, in the context of the Turing test, and Alexa, with the Alexa Prize, this whole realm? Do you think about this human intelligence, what it means to define it, what it means to reach that level?

    5. RP

      I do think the ability to converse is an, uh, sign of an ultimate intelligence. I think that is no question about it. So, if you think about all aspects of humans, there are sensors we have, and, uh, those are basically a data collection mechanism. And based on that, we make some decisions with our sensory brains, right? And from that perspective, I think that there are elements we have to talk about, how we sense the world, and then how we act based on what we sense. Those elements, clearly machines have. But then, there's the other aspects of computation that is way better. I also mentioned about memory again, in terms of being near infinite, depending on the storage capacity you have. And the retrieval can be extremely fast and pure, uh, in terms of like, there's no ambiguity of, "Who did I see when?" (laughs) , right? I mean, if you, machines can remember that quite well. So, it, again, on a philosophical level, I do subscribe to the fact that to con-, be able to converse, and as part of that, to be able to reason based on the world knowledge you've acquired and the sensory knowledge that is there, is definitely very much the essence of intelligence. But intelligence can go beyond human level intelligence, based on what machines are getting capable of.

    6. LF

      So, what do you think, maybe stepping outside of Alexa, broadly as an AI field, what do you think is a good test of intelligence? Put it another way, outside of Alexa, because so much of Alexa is a product, is an experience for the customer. On the research side, what would impress the heck out of you if you saw? You know, what is the test where you said, "Wow, this thing is now starting to encroach into the realm of what we loosely think of as human intelligence"?

    7. RP

      So, well, we think of it as AGI and human intelligence-

    8. LF

      AGI.

    9. RP

      ... all together, right? So, in some sense (laughs) . And I think we are quite far from that. Uh, I think, uh, an unbiased view I have is that the Alexa's ca-, intelligence capability is a great test. I think of it as there are many other proof points, like self-driving cars, game playing, like Go or chess. Let's take those two for, as an example.

    10. LF

      Sure.

    11. RP

      Uh, clearly requires a lot of data-driven learning and intelligence, but it's not as hard a problem as conversing with, as an AI is with humans to accomplish certain tasks, or open domain chat, as you mentioned-

    12. LF

      Mm-hmm.

    13. RP

      ... Alexa Prize.In those settings, the key difference is that the end goal is not defined, unlike game playing. You also do not know exactly what state you are in, in a particular goal completion scenario. (laughs) And so sometimes... sometimes you can if it's a simple goal, but if you're... Even certain examples like planning a weekend or... You- you can imagine how many th- things change along the way. Uh, you look for weather, you may change your mind and you, uh, you change the destination, or you want to catch a particular event and then you decide, no, I want this other event I want to go to. So these dimensions of how many different steps are possible when you're conversing as a human with a machine makes it an extremely daunting problem, and I think it is the ultimate test for intelligence.

    14. LF

      And don't you think that natural language is enough to prove that conversation?

    15. RP

      Uh...

    16. LF

      Just pure conversation?

    17. RP

      From a scientific standpoint, natural language, uh, is a great test. Uh, but I would go beyond... Uh, I don't wanna l- limit it to as natural language or simply understanding an intent or parsing for entities and so forth. We are really talking about dialogue.

    18. LF

      Dialogue. Yeah.

    19. RP

      Uh, so- so I would say human machine dialogue is definitely one of the best tests of intelligence.

    20. LF

      So

  5. 13:04 – 21:35

    Alexa prize

    1. LF

      can you briefly speak to the Alexa Prize for people who are not familiar with it?

    2. RP

      Uh-huh.

    3. LF

      And- and also just maybe where things stand, and what have you learned and what's surprising? What have you seen that's surprising-

    4. RP

      (laughs)

    5. LF

      ... from this incredible competition?

    6. RP

      Absolutely. It's a very exciting competition. Uh, Alexa Prize is essentially a grand challenge in conversational artificial intelligence, where we threw the gauntlet to the universities, uh, who do, uh, active research in the field, to say, "Can you build what we call a social bot that can converse with you coherently and engagingly for 20 minutes?" Uh, that is an extremely hard challenge, uh, talking to someone in a- uh, who you're meeting for the first time or even if you're... (laughs) You've met them quite often, uh, to speak at 20 minutes on any topic, an evolving nature of topics is super hard. Uh, we have completed two successful years of the competition. The first was won with the University of Washington, second with University of California. We are in our third instance. We have a extremely strong team of ten cohorts and the third instance of the, uh, of the Alexa Prize is underway now. And we are seeing a constant evolution. First year was definitely a learning. It was a lot of things to be put together. We had to build, uh, a lot of infrastructure to enable these universities to be able to build magical experiences and c- uh, and do h- high-quality research.

    7. LF

      Just a few quick questions.

    8. RP

      Yeah.

    9. LF

      Sorry for the interruption. What does failure look like af- uh, in the 20-minute session? So what does it mean to fail not to reach the 20-minute mark?

    10. RP

      Oh, awesome question. Uh, so there are... One, first of all, uh, I forgot to mention one more detail. It's not just 20 minutes, but the quality of the conversation too that matters, and, uh, the beauty of this competition, before I answer that question on what failure means, is first that you actually converse with millions and millions of customers as these social bots. Uh, so during the judging phases, uh, there are multiple phases, uh, before we get to the finals, which is a very controlled judging in a situation where we have, uh, we bring in judges and we have interactors who interact with these social bots, that is a much more controlled setting. But to- till the point we get to the finals, the... all the judging is essentially by the customers of Alexa, and there you basically rate, uh, on a simple question how good your experience was.

    11. LF

      Mm.

    12. RP

      Uh, so that's where we are not testing for a 20-minute boundary being cla- uh, crossed, because you do want it to be very much like a clear-cut winner we chosen and- and it's an absolute bar. (laughs) So did you really, uh, break that 20-minute barrier is why we have to test it in a more controlled setting with actors, essentially interactors, and see how the conversation goes. So this is why it's a subtle difference between how it's being tested in the field with real customers versus in the lab to award the prize. So on the latter one, what it means is that essentially the, uh, the j- there are three judges and two of them have to say this conversation has stalled, essentially.

    13. LF

      Mm-hmm. Got it, and the judges are human experts that are-

    14. RP

      Judges are human experts.

    15. LF

      Okay, great. So th- th- this in the third year, so what's been the evolution? How far? So the DARPA challenge in the first year n- the autonomous vehicles-

    16. RP

      (laughs)

    17. LF

      ... nobody finished. In the second year, a few more finished in the desert. Uh, so how far along with- in this, I would say, much harder challenge are we?

    18. RP

      This challenge has come a long way, to the extent that, uh, we're definitely not close to the 20-minute barrier being with coherence and engaging, uh, conversation. I think we are still five to 10 years away in that horizon to complete that. Uh, but the progress is immense, like, uh, you... what you're finding is the accuracy and what kind of responses these social bots generate is getting better and better. Uh, what's even amazing to see that, uh, now there's humor coming in. The bots are quite, uh-

    19. LF

      Awesome. (laughs)

    20. RP

      (laughs) You know, y- you're talking about ultimate signs of intelli- uh, signs of intelligence, I think humor is a very high bar-

    21. LF

      Mm-hmm.

    22. RP

      ... in terms of what it takes to create humor, uh, and I don't mean just being goofy, I really mean good sense of humor-

    23. LF

      Yeah.

    24. RP

      ... is also a sign of intelligence in my mind, and something very hard to do. So these social bots are now exploring not only what we think of natural language abilities, but also personality attributes and, uh, aspects of when to inject an appropriate joke, when to, uh... when you don't know the ques- uh, the domain, how you come back with something more intelligible so that you can continue the conversation if- if you and I are talking about AI and we are domain experts, we can speak to it, but if you suddenly switch a topic to that I don't know of, how do I change the conversation? So you're starting to notice these elements as well.And that's coming from... partly by (laughs) by the a- nature of the 20 minute t- challenge.

    25. LF

      Mm-hmm.

    26. RP

      That people are getting quite clever on how to, uh, really converse and, m- uh, essentially mask some of the understanding defects if they exist.

    27. LF

      So some of this, this is not Alexa the product, this is f- somewhat for fun, for research, for innovation, and so on. Uh, I have a question sort of in this modern era, there's a lot of... if you look at, uh, Twitter and Facebook and so on, there's- there's discourse, public discourse going on, and some things that are a little bit too edgy, people get blocked and so on. I- um, just outta curiosity, are people in this context pushing the limits? Is anyone using the F word? Is anyone s- uh, sort of pushing back, uh, sort of, uh, you know, arguing, I- I guess I should say, in- as part of the dialogue to really draw people in?

    28. RP

      First of all, let me just back up a bit-

    29. LF

      Yeah.

    30. RP

      ... in terms of why we are doing this, right? So, uh, you said it's fun. Uh, I think fun is, uh, more part of the, uh, engaging part for customers. It is one of the most, uh, used skills as well in our skill store. But up- that apart, the real goal was essentially what was happening is with lot of AI research moving to industry, uh, we felt that academia has the risk of not being able to have the same resources at disposal that we have, which is lots of data, massive computing power, uh, and, uh, clear, uh, ways to test these AI advances with, uh, real customer benefits. So we brought all these three together in the Alexa Prize, that's why it's one of my favorite projects in Amazon. And, uh, with that, uh, eh- the secondary effect is, yes, it has become engaging for our customers as well. Uh, we're not there in terms of where we want to- it to be (laughs) , right? But it's a huge progress. But coming back to your question on-

  6. 21:35 – 27:24

    What does it take to win the Alexa prize?

    1. RP

    2. LF

      So what do you think it takes to build a system that wins the Alexa Prize?

    3. RP

      I think you have to start focusing on aspects of reasoning, uh, that it is... there are still more lookups of what intents customers asking for and responding to those, uh, rather than, uh, really reasoning about (laughs) the elements of the, uh, of the conversation. For instance, if you have... if you're playing... if the conversation is about games and it's about a recent sports event-

    4. LF

      Mm-hmm.

    5. RP

      ... there's so much context involved and you have to understand the entities that are being mentioned so that the conversation is coherent rather than you suddenly just switch to knowing some fact about, uh, a sports entity and you're just relaying that rather than understanding the true context of the game. Like, you, uh... if you just said, "I learned this fun fact about Tom Brady," rather than really say how he played the game, eh, the previous night, then the conversation is not really that intelligent. Uh, so you have to go to more reasoning elements of understanding the context of the dialogue and giving more appropriate responses, which tells you that we are still quite far. Because a lot of times it's, uh, more facts being looked up (laughs) and something that's close enough as an answer, but not really the answer. Uh, so that is where the research needs to go more, in actual true understanding and reasoning. And that's why I feel it's a great way to do it, because you have an engaged set of, uh, users working to make... help, uh, these AI advances happen in this case.

    6. LF

      Y- right. So you mention customers, uh, uh, there quite a bit t- and there's a skill. Uh, what is the-

    7. RP

      Mm-hmm.

    8. LF

      ... experience for- for the, uh, for the user that it's helping? So just to clarify-

    9. RP

      Yeah.

    10. LF

      ... this isn't, as far as I understand, the Alexa... so this skill is a standalone for the Alexa Prize. I mean, it's focused on the Alexa Prize.

    11. RP

      Yes.

    12. LF

      It's not you o- ordering certain things on amazon.com or tr- checking the weather or-

    13. RP

      Yeah.

    14. LF

      ... playing Spotify, right? This is a s- separate skill.

    15. RP

      Exactly.

    16. LF

      And so you're focused on helping not... uh, h- I don't know, how- how do people, how do customers think of it? Are they having fun? Are they helping teach the system? What's the experience like for them?

    17. RP

      I think it's both actually. Um, uh, and let me tell you how the, uh, how you invoke this skill. So you, all you have to say, "Alexa, let's chat."

    18. LF

      Right.

    19. RP

      And then, the first time you, uh, say, "Alexa, let's chat," it comes back with a clear message that you're interacting with one of those university social bots, uh, and there's a clear... so you know exactly-

    20. LF

      That's so awesome.

    21. RP

      ... how we interact, right?

    22. LF

      Yeah.

    23. RP

      And that is why it's very transparent. Uh, you are being asked to help, right? And, uh, and we have lot of mechanisms where as the, uh, we are in the f- uh, first phase or feedback phase then, you send a lot of emails to our customers. And then they s- uh, they know that this, uh, the team needs a lot of interactions to improve the accuracy of the system. So, we know we have a lot of customers who really want to help these university bots, and they're conversing with that, and some are just having fun with just saying, "Alexa, let's chat." And, uh, also some adversarial behavior to see whether, how much do you understand (laughs) as a social bot. So, I think we have a good healthy mix of all three situations.

    24. LF

      So, what is the, if we talk about solving the Alexa challenge, the Alexa Prize, uh, wh- what's the data set of really engaging pleasant conversations look like? 'Cause if we think of this as a supervised learning problem-

    25. RP

      Mm-hmm.

    26. LF

      ... I don't know if it has to be, but if it does, maybe you can comment on that.

    27. RP

      Mm-hmm.

    28. LF

      Do you think there needs to be a data set of, of, uh, what it means to be an engaging, successful, fulfilling conversation?

    29. RP

      I think that's part of the research question here. This was, I think, uh, ex- uh, we at least got the first part right, which is have a way for universities to, uh, build and test in a real world setting.

    30. LF

      Right.

  7. 27:24 – 34:35

    Embodiment and the essence of Alexa

    1. LF

      You said the only way to make a smart assistant really smart is to give it eyes and let it explore the world. Uh, I'm not sure, it might have been taken out of context, but can you, uh, comment on that? Can you elaborate on that idea? 'Cause I personally also find that idea super exciting from a social robotics, personal robotics perspective.

    2. RP

      Yeah. A lot of things do get taken out of context. My, this particular one was just as philosophical discussion we were having on terms of what does intelligence look like. And the context was, in terms of learning, I think, uh, just we said, we as humans are empowered with many different sensory abilities.

    3. LF

      Yeah.

    4. RP

      Uh, I do believe that, uh, eyes are an important aspect of it in terms of, uh, if you think about how we as humans learn, it is quite complex and it's also not unimodal that you are fed a ton of text-

    5. LF

      Mm-hmm.

    6. RP

      ... or audio and you just learn that way. No. You are, you learn by experience, you learn by seeing, you're taught by, uh, humans. And we are very efficient in how we learn. Uh, machines, on the contrary, are very inefficient on how they learn, especially these AIs. I think the next wave of research is going to be with less data, not just less human, uh, not, uh, just with less label data, but also with a lot of weak supervision. And where you can increase the, uh, learning rate, I don't mean less data in terms of not having a lot of data to learn from, that we are generating so much data, but it is more about from an aspect of how fast can you learn?

    7. LF

      Mm-hmm. So, improving the, the, the quality of the data that's, uh, the quality of the data and the learning process, but-

    8. RP

      I think more on the learning process. I think we have to, we as humans learn with a lot of noisy data, right? (laughs)

    9. LF

      Yeah.

    10. RP

      And, uh, and I think that's the part that, uh, I don't think should change. What should change is how we learn, right? So if you look at, you mentioned supervised learning, we have making transformative shifts from moving to more unsupervised, more weak supervision. Those are the key aspects of how to learn. Uh, and I think in that setting, you, I hope you agree with me that, uh, having other senses is very crucial in terms of how you learn.

    11. LF

      So, absolutely. And from a, from a machine learning perspective, which I hope we get a chance to talk to a few aspects that are fascinating there. But to stick on the point of sort of, um, a body, you know, uh, embodiment. So Alexa has a body, has a very s- minimalistic-

    12. RP

      Mm-hmm.

    13. LF

      ... beautiful interface, or there's a ring and so on. I mean, I'm not sure of all the flavors of, uh, the devices that Alexa lives on, but there's a minimalistic basic interface, uh-And nevertheless, we're humans so I have a Roomba, I have all kinds of robots all over everywhere. So, uh, what do you think the, uh, Alexa of the future looks like if it begins to shift what its body looks like? What, uh, what maybe beyond Alexa, what do you think are the different devices in the home as they start to embody their intelligence more and more, what do you think that looks like? Philosophically-

    14. RP

      Yeah.

    15. LF

      ... a fu- a future, what do you think that looks like?

    16. RP

      I think, uh, let's look at what's happening today. You mentioned, I think, our devices as in Amazon devices.

    17. LF

      Yeah.

    18. RP

      But I also wanted to point out, Alexa is already integrated on a lot of third-party devices, which also come in lots of forms and shapes. Some in robots, right? Some in, uh, microwaves. (laughs) Some in appliances of, uh, that you use in e- everyday life. So, I think it is... it's not just the shape Alexa takes in terms of form factors, but it's also where all it's available. Uh, it's getting in cars, it's getting in different appliances in homes, even toothbrushes. (laughs) Right?

    19. LF

      Yeah.

    20. RP

      So, I think you have to think about it as not, uh, a physical assistant. It will be in some embodiment. As you said, we already have these nice devices, uh, but I think it's also important to think of it, uh, it is a virtual assistant. It is superhuman in the sense that it is in multiple places at the same time. Uh, so I think the, uh, the actual embodiment to, in some sense, to me, doesn't matter.

    21. LF

      Right.

    22. RP

      I think you have to think of it as not as human-like, and more of what its capabilities are that derive a lot of benefit for customers.

    23. LF

      Right.

    24. RP

      And how there are different ways to delight it... and, uh, delight customers in different experiences. And I think I'm a big fan of it not being ins- just human-like. It should be human-like in certain situations. Alexa, prior social bot in terms of conversation is a great way to look at it. But there are other, uh, uh, scenarios where human-like I think is underselling the abilities of this AI.

    25. LF

      So, if I could, uh, trivialize what we're talking about. So, if you look at the way Steve Jobs thought about the interaction with the device that the- that Apple produced, there was a- a extreme focus on controlling the experience by making sure there's only these sp- Apple-produced devices. You see the voice of Alexa being... taking all kinds of forms depending on what the customers want, and that means, uh, that means it could be anywhere from the microwave, to a vacuum cleaner, to the home, and so on. The voice is the essential element-

    26. RP

      Correct.

    27. LF

      ... of the interaction.

    28. RP

      I think voice is an essence. It's not all, but it's a key aspect. I think, uh, to your question, in terms of, uh, you should be able to recognize Alexa.

    29. LF

      Yeah.

    30. RP

      And that's a huge problem, I think, in terms of... a huge scientific problem, I should say. Like, what are the traits? What makes it look like Alexa? Especially in different settings, and especially if it's primarily voice, what it is. But Alexa's not just voice either, right? I mean, we have devices with a screen. Uh, now you're seeing just other behaviors of Alexa. So, I think we are in very early stages of what that means, and this will be an important topic for the following years. Uh, but I do believe that being able to recognize and tell when it's Alexa versus it's not is going to be important from an Alexa perspective. I'm not speaking for the entire AI (laughs) -

  8. 34:35 – 36:23

    Personality

    1. LF

      w-... there's a, there's a team that works on personality.

    2. RP

      Yeah.

    3. LF

      So if we talk about those different flavors of what it means, culturally speaking, India, UK, US, what does it mean to add... so, so the problem that we just stated, which is fascinating, how do we make it purely recognizable that it's Alexa, assuming that the qualities of the voice are not sufficient? Uh, it, uh, it's also the content of what is being said.

    4. RP

      Yeah.

    5. LF

      W- how do we, how do we do that? How does the personality c-

    6. RP

      (laughs)

    7. LF

      ... come into play? What's, uh, what, what's that research even look like? I mean, it's such a fascinating speaker

    8. RP

      We have some very fascinating, uh, uh, folks who, from both a UX background and human factors, are looking at these aspects and these exact questions.

    9. LF

      Yeah.

    10. RP

      Uh, but I'll definitely say it's not just how it sounds. The choice of words, the tone, not just, I mean the voice identity of it, (laughs) but-

    11. LF

      Yeah.

    12. RP

      ... the tone matters, the speed matters. Uh, how you speak, how you enunciate words, how, uh... what choice of words are you using? Uh, how terse are you or how, uh, lengthy in your explanations you are. All of these are factors. Uh, and y- also, you mentioned something crucial that it's... may have... you may have personalized it, uh, Alexa, to some extent-

    13. LF

      Yeah.

    14. RP

      ... in your homes or in the devices you are interacting with. So, you as... your individual, how you prefer Alexa sounds can be different than how I prefer. (laughs)

    15. LF

      Yeah.

    16. RP

      And we may... uh, and the amount of customizability you want to give is also a key debate we always have.

    17. LF

      Right.

    18. RP

      Uh, but I do want to point out it's more than the voice actor that recorded, and it sounds like that actor. (laughs) It is more about the choices of words, the attributes of, uh, tonality, the volume in terms of how you raise your pitch and so forth,-... all of that

  9. 36:23 – 38:49

    Personalization

    1. RP

      matters.

    2. LF

      This is such a fascinating problem from a product perspective.

    3. RP

      Uh-huh.

    4. LF

      I could see those debates just happening inside of the Alexa team-

    5. RP

      (laughs)

    6. LF

      ... of how much personalization do you do for, for the specific customer, 'cause you're taking a risk if you over personalize, uh, because you don't... I, I if you create a personality for a million people, you can test that better, you can create a-

    7. RP

      Yeah.

    8. LF

      ... rich, fulfilling experience that will do well. But if you... the more you personalize it, the less you can test it, the less you can know that it's a, it's a great-

    9. RP

      Mm-hmm.

    10. LF

      ... experience. So how much personalization? What's the right balance?

    11. RP

      I think the right balance depends on the customer. Give them the control.

    12. LF

      Yeah.

    13. RP

      So I'll say, uh, I think the, uh, more control you give customers, the better it is for everyone. And, uh, I'll give you some key personalization features.

    14. LF

      Mm-hmm.

    15. RP

      I think we have a feature called Remember This, which is where you can tell Alexa to remember something. Uh, there, you have an explicit sort of control in customer's hand because they have to say, "Alexa, remember X, Y, Z."

    16. LF

      What kind of things would that be used for?

    17. RP

      So you can like- (laughs)

    18. LF

      For song title or something or...

    19. RP

      I, I have stored my tire specs for my car-

    20. LF

      Nice.

    21. RP

      ... because it's so hard to go and find and see what it is, (laughs) right? When you're having some issues. So, uh, I, I store my mileage plan, uh, numbers for all the frequent flyer ones where I'm sometimes just looking at it and it's not handy.

    22. LF

      Mm-hmm.

    23. RP

      Uh, so, and, so those are my own personal choices I've made for Alexa to remember something on my behalf, right? So again, I think the choice was be explicit about how you provide that to a customer as a control. So I think these are the aspects of, uh, what you do. Like think about where we can use speaker recognition capabilities that it's if you taught Alexa that you are Lex and this person in your household is Person 2-

    24. LF

      Yeah.

    25. RP

      ... then you can personalize the experiences. Again, these are very in the C- uh, in the CX, customer experience, patterns are very clear about and transparent when a personalization action is happening. And then you have other ways like you go through explicit control right now through your app that, uh, you have multiple service providers, let's say for music. Which one is your preferred one? So when you say, "Play Sting," depend on your- whether you have preferred Spotify or Amazon Music or Apple Music that the decision is made where to play it from.

  10. 38:49 – 40:35

    Alexa's backstory from her perspective

    1. LF

      So what's Alexa's backstory from her perspective?

    2. RP

      (laughs)

    3. LF

      Does... is there, um... I, I remember just asking, as probably a lot of us are, just the basic questions about love and so on-

    4. RP

      Yeah.

    5. LF

      ... of Alexa just to see what the answer would be, just the... Uh, I... It feels like there's a little bit of a back... Like there's a, it feels like there's a little bit of personality, but not too much. Is, is Alexa have a metaphysical presence in this human universe we live in or is it something more ambiguous? Is there a past? Is there a birth? Is there a family kind of-

    6. RP

      (laughs)

    7. LF

      ... idea even for joking purposes and so on?

    8. RP

      I think, uh, well, it does tell you if I think you... uh, I should double-check this, but if you said, "When were you born?"

    9. LF

      When were you born, yeah.

    10. RP

      I think we do respond. I need to double-check that, but I'm pretty positive about it.

    11. LF

      I think you, you do actually-

    12. RP

      Yes.

    13. LF

      ... 'cause I think I've tested that.

    14. RP

      So, uh-

    15. LF

      But that's like, uh, that's like how-

    16. RP

      (laughs)

    17. LF

      Like I was, I was born in Urbana-Champaign and whatever the year-

    18. RP

      Yeah.

    19. LF

      ... kind of thing, yeah.

    20. RP

      So in terms of the metaphysical, I think it's early. Uh, does it have the historic knowledge about herself to be able to do that? Maybe. Have we crossed that boundary? Not yet, right? In terms of being... thinking of... Uh, have we thought about it? Quite a bit, uh, but I wouldn't say that we have come to a clear decision in terms of what it should look like. But, uh, you can imagine though, and I bring this back to the Alexa Prize social bot one, there you will start seeing some of that. Like you-

    21. LF

      Yeah.

    22. RP

      ... these bots have their identity and in terms of that you may find, uh, you know, this is such a great research topic that some academia team may think of these problems and start solving them too.

  11. 40:35 – 44:00

    Trust in Human-AI relations

    1. RP

    2. LF

      So let me ask, uh, a question. It's kinda difficult I think. Um, but it feels, and fascinating to me 'cause I'm fascinated with psychology. It feels that the more personality you have, the more dangerous it is in terms of a customer perspec- a product, if you wanna create a product that's useful. Da- by dangerous I mean creating an experience that upsets me. (laughs) And so, uh, uh, what... how do you, how do you get that right? Because, uh, if you look at the, the relationships, may- maybe I'm just a screwed up Russian-

    3. RP

      (laughs)

    4. LF

      ... but, uh, if you look at the re- human to human relationship, some of our deepest relationships have fights, have tension, have the push and pull, have a little flavor in them. Uh, do you want to have such flavor in an, an interaction with Alexa? How do you think about that?

    5. RP

      So there's one other common thing that you didn't say but is, we think of it as paramount-

    6. LF

      Mm-hmm.

    7. RP

      ... for any deep relationship. That's trust.

    8. LF

      Trust, yeah.

    9. RP

      So I think if you trust every attribute you said-

    10. LF

      Mm-hmm.

    11. RP

      ... a fight, some tension-

    12. LF

      Yeah.

    13. RP

      ... is all healthy. But the, what is, uh, sort of unnegotiable in this instance is trust. And I think the bar to earn customer trust for AI is very high.

    14. LF

      Yeah.

    15. RP

      In some sense more than a human.

    16. LF

      Yes.

    17. RP

      It's, uh, it's not just about personal information or your data, it's also about your actions on a daily basis. How trustworthy are you in terms of consistency, in terms of, uh, how accurate are you in understanding me? Like if, if you're talking to a person on the phone, if you have a problem with your, let's say, your internet or something, if the person is not understanding, you lose trust right away. You don't want to talk to that person.

    18. LF

      Yeah.

    19. RP

      S- that whole example gets amplified by a factor of 10 because as... when you're a human interacting with an AI-... you have a certain expectation. Either you expect it to be very intelligent, and then you get upset, why is it behaving this way-

    20. LF

      Mm-hmm.

    21. RP

      ... or you expect it to be, uh, not so intelligent, and when it surprises you, you're like, "Really?" (laughs) "You were trying to be too smart?" So, I think we grapple with these hard questions as well, but I think the key is, actions need to be trustworthy from these AIs. Not just about data protection, your personal information protection, uh, but also from how accurately it accomplishes all commands or all interactions.

    22. LF

      Well, it's tough to hear because trust, you're absolutely right, but trust is such a high bar with AI systems-

    23. RP

      Yep.

    24. LF

      ... because people, and I see this 'cause I work with autonomous vehicles, I mean-

    25. RP

      (laughs)

    26. LF

      ... the bar that's placed on AI system is unreasonably high.

    27. RP

      Yeah, that is going to be a s- uh, I agree with you. And, uh, I think of it as, it's, uh-

    28. LF

      A challenge? (laughs)

    29. RP

      It's a challenge, and it's also keeps my job. (laughs)

    30. LF

      (laughs)

  12. 44:00 – 47:45

    Privacy

    1. RP

      I think, for us.

    2. LF

      Well, one of the questions that we grapple as a society now, that I think about a lot, that I think a lot of people in the AI think about a lot, and Alexa's taking on head on, is privacy.

    3. RP

      Mm-hmm.

    4. LF

      Is, th- the reality is, us giving over data to any AI system can be used to enrich our lives in h- in, in, in profound ways.

    5. RP

      Mm-hmm.

    6. LF

      So if it, may, basically any product that does anything awesome for you, would, m- the more data it has, the more awesome things it can do. And yet, uh, on the other side, people imagine the worst case possible scenario of what can you possibly do with that data. People h- it's, it go- boils down to trust, as you said before.

    7. RP

      Yeah.

    8. LF

      There's a fundamental distrust of, uh, in certain groups of governments and so on. Depending on the government, depending on who's in power, depending on all these kinds of factors. And so here's Alexa in the middle of all of it-

    9. RP

      (laughs)

    10. LF

      ... uh, in the home, trying to do good things for the customers, so how do you think about privacy in this context, the smart assistants in the home? How do you maintain, how do you earn trust?

    11. RP

      Absolutely. So, as you said, trust is the key here.

    12. LF

      Yeah.

    13. RP

      So you start with trust, and then privacy is a key aspect of it. It's, has to be designed from very beginning about that. And we believe in two, uh, fundamental principles. One is transparency, and second is control. So if you, uh, by transparency I mean, uh, when we built, uh, what is now called Smart Speaker, or the first Echo, we were quite judicious about making these right trade-offs on customers' behalf that it is pretty clear when, when the audio's being sent to cloud, the light ring comes on when it has heard you say the word, wake word, and then the streaming happens, right? So and the light ring comes up. We also had, we put a physical mute button on it, just so you're, if you didn't want it to, uh, be listening even for the wake word, then you turn the, uh, power button, uh, the mute button on, and that, uh, disables the microphones. That's just the first decision on essentially transparency and control. O- then, even when we launched, we gave the control in the hands of the customers that you can go and look at any of your individual utterances that is recorded and delete them anytime, and, uh, we have cut to, true to that promise, right? So and that is super, again, a great instance of showing how you have the control. Then we made it even easier. You can say, "Alexa, delete what I said today." So that is now making it even just, (laughs) just more control in your hands-

    14. LF

      Mm-hmm.

    15. RP

      ... but what's most convenient about this technology is voice. You delete it with your voice now.

    16. LF

      Yeah.

    17. RP

      Uh, so these are the types of decisions we continually make. Uh, we just recently launched this feature called, uh, what we think of it as if you wanted humans not to review your tr- uh, data, uh, because sm- you've mentioned supervised learning, right?

    18. LF

      Yeah.

    19. RP

      So you, in supervised learning, humans, uh, have to give some annotation. And that also is now a feature where you can, uh, essentially, if you've selected that flag, your data will not be reviewed by a human. So, these are the types of controls that we have to constantly offer with customers.

    20. LF

      So, why do you think it bothers people so much that, so th- uh, so everything you just said, uh, is really powerful, so the control, the ability to delete, 'cause we collect, we have studies here running at MIT that collects huge amounts of data-

    21. RP

      Mm-hmm.

    22. LF

      ... and people consent and so on. Uh, the b- ability to delete that data is really empowering, and m- almost nobody ever asks to delete it, but the ability to have that control is, is really powerful. But still,

  13. 47:45 – 53:51

    Is Alexa listening?

    1. LF

      you know, there's these popular anecdote, anecdotal evidence that people say, they like to tell that, uh, them and a friend were talking about something, I don't know, uh, sweaters for cats, and all of a sudden they'll have advertisements for cat sweaters-

    2. RP

      Mm-hmm.

    3. LF

      ... on Amazon. There's that, that's a popular anecdote, as if something is always listening. What, can you explain that anecdote, that experience that people have? What's the psychology of that?

    4. RP

      Mm-hmm.

    5. LF

      What's that experience? Uh, and can you ... You've answered it, but let me just ask, is Alexa listening? (laughs)

    6. RP

      No, Alexa listens only for the wake word on the device, right? Uh-

    7. LF

      And a wake word is ...

    8. RP

      The words like, "Alexa," "Amazon," "Echo," and you, uh, but you only choose one at a time.

    9. LF

      Yeah.

    10. RP

      So you choose one and it listens only for that on our devices.

    11. LF

      Okay.

    12. RP

      So that's first. From a listening perspective, we have to be very clear that it's just a wake word. So you said, "Why is there, uh, this anxiety?" If you may.

    13. LF

      Yeah, exactly.

    14. RP

      It's because there's lot of confusion what (laughs) it really listens to, right? And you, uh, and I think it's partly on us to keep educating, uh, our, our customers and the general media more in terms of like how, what really happens, and we have done, uh, a lot of it. Uh, and w- uh, our pages on information are clear, but still people have to have, uh, more... There's always a hunger for information (laughs) -

    15. LF

      Yeah.

    16. RP

      ... and clarity. And we'll constantly look at how best to communicate. If you go back and read everything, yes, it states exactly that, uh, and then people could still question it. And I think that's absolutely okay to question. Uh, what we have to make sure is that we are, uh, because our fundamental philosophy is customer first, customer obsession is our leadership principle, if you put... As researchers, I put myself in the f- uh, shoes of the customer, and all decisions in Amazon are made with that in line, so. And trust has to be earned and we have to keep earning (laughs) the trust of our customers in this setting. Uh, and to your other point on like, uh, is there something showing up based on your conversations? No.

    17. LF

      Mm-hmm.

    18. RP

      I think the answer is like you, uh, a lot of times when those experiences happen, you have to also be, know that, okay, it may be a winter season. (laughs) People are looking for sweaters, right?

    19. LF

      Yeah.

    20. RP

      And it shows up on your amazon.com because it is popular. So-

    21. LF

      Okay.

    22. RP

      ... there are many of these, uh, uh, you mentioned that, uh, personality or personalization. Turns out we are not that unique either. (laughs)

    23. LF

      Yeah.

    24. RP

      Right? So those things we, we as humans start thinking, "Oh, must be because something was heard and that's why this other thing showed up," the answer is no. Probably it is just the season (laughs) for sweaters.

    25. LF

      I'm not gonna ask you this question 'cause it's just, 'cause you're also, 'cause people have so much paranoia. But from my, let me just say from my perspective, I hope there's a day when customer can ask Alexa to listen all the time, uh, to improve the experience, to improve... Because I, I personally don't see the negative of, because if you have the control and if you have the trust, there's no reason why you shouldn't be listening all the time to the conversations to learn more about you. Because ultimately, as long as you have control and trust, every data you provide to the device that the device wants (laughs) is going to be useful.

    26. RP

      Mm-hmm.

    27. LF

      And that, so that to me, I, as, as a machine learning person, I think it worries me how sensitive people are about their data relative to how, uh, empowering it could be for the devices around them, how m- enriching it could be for their own life to ex- improve the product. So I just, (laughs) it's something I think about sort of a lot, how do we make that in devices. Obviously Alexa thinks about it a, a lot as well. I don't know if you wanna comment on that-

    28. RP

      (laughs)

    29. LF

      ... sort of. Okay, have you seen, let me ask it-

    30. RP

      So you, do-

  14. 53:51 – 54:51

    How Alexa started

    1. RP

    2. LF

      (laughs) So I know we, we kind of started with a lot of philosophy and a lot of interesting topics-

    3. RP

      Yeah.

    4. LF

      ... and we're just jumping all over the place, but really some of the fascinating things that, uh, the Alexa team and Amazon is doing is in the, the algorithm side, the data side, the technology, the deep learning and machine learning and, and so on. So can you give a brief history of Alexa-

    5. RP

      Uh-huh.

    6. LF

      ... from the perspective of just innovation, the algorithms, the data of-

    7. RP

      Mm-hmm.

    8. LF

      ... how, how it was born, how it came to be, how it's grown-

    9. RP

      Yeah.

    10. LF

      ... where it is today?

    11. RP

      Yeah. It start with the... In Amazon, everything starts with the customer, and we have a process called working backwards. Alexa, and more specifically than the product Echo, there was a working backwards document essentially that reflected what it would be. Started with a very simple...... a vision statement, for instance, (laughs) that, uh, morphed into a full-fledged document along the way changed into what all it can do, uh, right?

  15. 54:51 – 1:11:51

    Solving far-field speech recognition and intent understanding

    1. RP

      You can, uh... But the inspiration was the Star Trek computer. So when you think of it that way, you know, e- everything is possible, but when you launch a product, you have to start with some place. And when I joined, we... the product was already in conception, and we started working on the far-field speech recognition because that was the first thing to solve. By that, we mean that you should be able to speak to the, uh, device from a distance, and in those days, that wasn't a common practice. Uh, and even in the previous research world I was in, was considered to, uh, unsolvable problem then-

    2. LF

      Mm-hmm.

    3. RP

      ... in terms of whether you can converse from a length.

    4. LF

      Yeah.

    5. RP

      And here, I'm still talking about the first part of the problem where you say... get the attention of the device, as in by saying what we call the wake word, uh, which means the word Alexa has to be detected with a very high accuracy because it is a very common word. It has sound units that map with words like I like you or Alek, Alex. (laughs)

    6. LF

      Yeah.

    7. RP

      Right? So it's a, uh, undoubtably hard problem to detect the right mentions of Alexa's address to the device versus I like Alexa.

    8. LF

      So you have to pick up that signal when there's a lot of noise and all kinds of background-

    9. RP

      Not only noise, but a lot of conversation-

    10. LF

      Yeah.

    11. RP

      ... are in the house, right?

    12. LF

      Oh-

    13. RP

      You remember on the device, you're s- simply listening for the wake word, Alexa. And there's a lot of words being spoken in the house. How do you know it's Alexa and directed at Alexa? Because I could say, "I love my Alexa." "I hate my Alexa." Uh, "I want Alexa to do this." And in all these three sentences, I said, "Alexa," I didn't want it to wake up.

    14. LF

      Yeah.

    15. RP

      Uh, so, um-

    16. LF

      Can I just pause on that second?

    17. RP

      Yeah.

    18. LF

      What would be your device that I should probably in the introduction of this conversation give to people in terms of with them turning off their Alexa device if they're listening to this podcast conversation out loud? Like w- what's the probability that an Alexa device will go off because we mentioned Alexa like a million times?

    19. RP

      (laughs) So it will... Uh, we have done lot of different things where we can, uh, figure out that there is the device, the speech is coming from a human versus, uh, over the air. Also, I mean, in terms of like, also it is... Think about ads or, uh, so we have also-

    20. LF

      Ads.

    21. RP

      ... launched a technology for watermarking kind of approaches in terms of-

    22. LF

      Mm-hmm.

    23. RP

      ... filtering it out. But, yes, if this kind of a podcast is happening, it's possible your device will wake up a few times.

    24. LF

      Okay.

    25. RP

      Right? It's an unsolved problem-

    26. LF

      (laughs)

    27. RP

      ... but it is, uh, definitely, uh, uh, something we care very much about, right?

    28. LF

      But the idea is you wanna detect Alexa first of all.

    29. RP

      Meant for the device.

    30. LF

      Meant for the device. I mean, f- first of all, just even hearing Alexa versus I like...

  16. 1:11:51 – 1:13:19

    Alexa main categories of skills

    1. RP

      So-

    2. LF

      So what, what are the big skills? Can you just go over them? 'Cause the only thing I use it for is music, weather, and shopping.

    3. RP

      (laughs) So, uh, we think of it as music information, uh, right? So it sh- uh, weather is a part of information, right?

    4. LF

      Weather is definitely... right.

    5. RP

      So, uh, when we launched, we didn't have smart home, but within s-... uh, by smart home I mean you connect your smart devices, you control them with voice. If you haven't done it, it's worth... it will change your life.

    6. LF

      Like turning on the lights and so on? Yeah.

    7. RP

      Yeah, turning on your light to, uh, anything that's connected, uh, and has a... it's just that-

    8. LF

      What's your favorite smart device for you that you-

    9. RP

      My light. (laughs)

    10. LF

      Light?

    11. RP

      And now you have the smart plug with... and you don't... uh, we also have this Echo Plug, uh, which is-

    12. LF

      Oh yeah, you can plug in anything-

    13. RP

      ... so you can plug in anything, and now you can turn on... that one on and off, right? So-

    14. LF

      I'll use this conversation motivation and get one at some point.

    15. RP

      The garage door, you can check, uh, your status of the garage door and things like... and we have gone make Alexa more and more proactive, where it even have a hunch- has hunches now that, uh... or looks-

    16. LF

      Hunches?

    17. RP

      Hunches like you left your light on. Uh, uh, let's say you've gone to your bed and you left the garage light on. So, uh, so it'll help-

    18. LF

      It will help you out.

    19. RP

      ... it will help you out in these settings, right? So-

    20. LF

      That's smart devices.

    21. RP

      Right.

    22. LF

      Information, smart devices, you said music.

    23. RP

      Yeah, so I don't remember everything we had, but alarms, timers were the big ones.

    24. LF

      Yeah, those were the big ones. Yeah, yeah, to have.

    25. RP

      Like, that was...... you know, the timers were very popular fr- right away. Uh, music also, like, you could play song, artist, album, everything, uh, and so that was, like, a, a clear win in terms-

    26. LF

      Yeah.

    27. RP

      ... of the customer experience.

  17. 1:13:19 – 1:17:47

    Conversation intent modeling

    1. RP

      So, that's ... again, this is language understanding. Now things have evolved, right? So, where we want Alexa definitely to be more accurate, competent, trustworthy, based on how well it does these core things, but we have evolved in many different dimensions. First is what I think of her doing more conversational for high utility, not just for chat. Right? And there we, uh, at re:MARS this year, which is our AI conference, we launched what is called Alexa Conversations. Uh, that is providing, uh, the ability for developers to author multi-turn experiences on Alexa with no code, essentially, where, uh, in terms of the dialogue code. Initially it was like, uh, you know, all these IVR systems, uh, you have to fully author if the customer says this, do that, right? So, the whole dialogue flow is hand-authored.

    2. LF

      Mm-hmm.

    3. RP

      And with Alexa Conversations, the way it is that you just provide a sample interaction data with your service order API, let's say your ATOM tickets that, uh, provides a service for buying movie tickets.

    4. LF

      Mm-hmm.

    5. RP

      You provide a few examples of how your customers will interact with your APIs, and then the dialogue flow is automatically constructed using a recurrent neural network, uh, trained on that data. So, that simplifies the developer experience. We just launched our preview for the developers to try this capability out. And then the second part of it, which shows i- even increased utility for customers is you and I, when we interact with Alexa or any customer, as I'm coming back to our initial part of the conversation, the goal is often unclear, uh, or unknown to the AI. If I say, "Alexa, what movies are playing nearby?" Am I trying to just buy movie tickets? Am I actually even ... do you think I'm looking for just movies for curiosity, whether the-

    6. LF

      Yeah.

    7. RP

      ... Avengers is still in theater or when is (laughs) maybe it's gone, and maybe it will come on my ... missed it, so I may watch it on Prime? (laughs)

    8. LF

      Yeah. Right.

    9. RP

      Which happened to me.

    10. LF

      Yeah. (laughs)

    11. RP

      Uh, so, uh, so from that perspective now, you're looking into, what is my goal? And let's say I now complete the movie ticket purchase. Maybe I would like to get dinner nearby.

    12. LF

      Mm-hmm.

    13. RP

      Uh, so what is really the goal here? Is it night out or is it movies, as in just go watch a movie?

    14. LF

      Yeah.

    15. RP

      The answer is we don't know. So, can Alexa now figuratively have the, uh, intelligence that I think this meta goal is really night out, or at least say to the customer, "When you've completed the purchase of movie tickets from ATOM Tickets or Fandango or Pick Your or any one, then the next thing is do you want to, uh, get to, uh, get a Uber to the theater?" Right?

    16. LF

      Mm-hmm.

    17. RP

      Or, uh, "Do you want to book a restaurant next to it?" And, uh, and then not ask the same information over and over again, what time? (laughs)

    18. LF

      Yeah.

    19. RP

      What, uh, uh, how many people in your party, right? So, uh, so this is where you shift the cognitive burden from the customer to the AI, where it's thinking the, of what is your, uh, it anticipates your goal and takes the next best action to complete it. Now, that's the machine learning problem. Uh, s- but essentially you're, uh, the way we solved this first instance, and we have, uh, got a long way to go to make it-

    20. LF

      Yeah.

    21. RP

      ... scale to everything possible in the world, but at least for this situation, it is from, uh, at every instance, Alexa is making the determination whether it should stick with the experience with ATOM Tickets or offer, uh-

    22. LF

      Mm-hmm.

    23. RP

      ... or you, based on what you say, whether either you have completed the interaction or you said, "No, get me an Uber now," so it will shift context into another experience or skill or another service. So, that's a dynamic decision-making. That's making Alexa, you can say, more conversational for the benefit of the customer, rather than simply complete transactions which are well thought through, you have i- uh, you as a customer has fully specified what you want to be accomplished.

    24. LF

      Yeah.

    25. RP

      It's accomplishing that.

    26. LF

      So, it's kinda as, uh, in, uh, we do this with pedestrians, right, intent modeling. It's, uh, predicting w- what your possible goals are-

    27. RP

      Exactly.

    28. LF

      ... and what's the most likely goal and then switching that depending on the things you say. So, the ... my question is there ... it seems ... maybe it's a dumb question, but

  18. 1:17:47 – 1:22:50

    Alexa memory and long-term learning

    1. LF

      it would help a lot if Alexa remembered me, what I said previously.

    2. RP

      Right. So-

    3. LF

      Is, is it, is it trying to use some memory for the customer's purpose?

    4. RP

      So, it- yeah, it is using a lot of memory within that. So, right now not so much in terms of, okay, which restaurant do you prefer-

    5. LF

      Right.

    6. RP

      ... right? That is a more long-term memory. But within the short-term memory, within the session, it is remembering how many people did you ... so if you said-

    7. LF

      Oh, with the-

    8. RP

      ... uh, buy four tickets, now it has made an implicit assumption that, uh, you are gonna have ... you need four, at least four seats at a restaurant. Right? So, these are the kind of, uh, contexts it's preserving between these skills but within that session. But you're asking the right question in terms of for it to be more and more useful-

    9. LF

      Yes.

    10. RP

      ... it has to have more long-term memory, and that's also an open question. And, uh, again, this is still early days.

    11. LF

      That's right. So, so for me, I, I mean, everybody's different, but, uh, yeah, I'm definitely not representative of the general population in the sense that I do the same thing every day.

    12. RP

      (laughs)

    13. LF

      (laughs) Like, I eat the same ... like, I do everything the same, the same thing. Um, wear the same thing clearly, uh, this or the black shirt. Uh, so it's frustrating when Alexa doesn't get what I'm saying-

    14. RP

      Okay.

    15. LF

      ... because I have to correct her every time the exact same way. This has to do with certain songs.

    16. RP

      Yeah.

    17. LF

      Like sh- she doesn't know certain weird songs I like and doesn't know-... I've complained to Spotify about this. Talked to the-

    18. RP

      (laughs)

    19. LF

      ... RD, uh, head of RD at Spotify. Stairway to Heaven, I have to correct it every time.

    20. RP

      Really?

    21. LF

      It do- doesn't play Led Zeppelin correctly.

    22. RP

      (laughs)

    23. LF

      It plays-

    24. RP

      Good-

    25. LF

      ... a cover of Led Z- of Stairway to Heaven.

    26. RP

      (laughs)

    27. LF

      So, I'm-

    28. RP

      You should figure... You should send me your, uh... Next time it fails-

    29. LF

      This-

    30. RP

      ... feel free to send it to me. We'll take care of it.

Episode duration: 1:45:57

Install uListen for AI-powered chat & search across the full episode — Get Full Transcript

Transcript of episode Ad89JYS-uZM

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.