Skip to content
a16za16z

Building the Cloud for AI Agents | AWS CEO Matt Garman

a16z’s Raghu Raghuram sits down with AWS CEO Matt Garman to discuss how AI is reshaping the cloud, from the needs of AI-native startups to infrastructure increasingly designed for agents. Matt explains how AWS is adapting as agents write code and manage infrastructure, why it’s reserving scarce GPU capacity for startups, and where custom chips like Trainium and Graviton fit into the AI stack. They also discuss Amazon's $220 billion capital investment, the shifting bottlenecks in the infrastructure buildout, what enterprises need to trust autonomous agents, and how AWS's own teams are building with agents. Timestamps: 00:00 - Intro 00:55 - $169B, growing 37% 02:51 - Why startups are the lifeblood 12:26 - Rethinking cloud for agents 16:24 - The GPU allocation problem 21:27 - Is the AI CapEx a bubble? 30:22 - Clearing up data center myths 33:07 - The Graviton and Trainium bet 40:06 - Where enterprises are stuck on agents 45:43 - AI risk, Hugging Face, and Continuum Resources: Follow Matt Garman on X: https://x.com/mattsgarman Follow Matt Garman on LinkedIn: https://www.linkedin.com/in/mattgarman Follow Raghu Raghuram on X: https://x.com/RaghuRaghuram Learn more about AWS: https://aws.amazon.com/ Stay Updated: If you enjoyed this episode, be sure to like, subscribe, and share with your friends! Find a16z on X: https://twitter.com/a16z Find a16z on LinkedIn: https://www.linkedin.com/company/a16z Listen to the a16z Show on Spotify: https://open.spotify.com/show/5bC65RDvs3oxnLyqqvkUYX Listen to the a16z Show on Apple Podcasts: https://podcasts.apple.com/us/podcast/a16z-podcast/id842818711 Follow our host: https://x.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see http://a16z.com/disclosures.

Matt GarmanguestRaghu Raghuramhost
Oct 8, 202656mWatch on YouTube ↗

EVERY SPOKEN WORD

  1. 0:00 – 0:55

    Intro

    1. MG

      The agentic workflows tend to perform better on AWS than anywhere else. Compute sandboxes, gateways, agent permissions versus people, a lot of those are things that we have built and are building and thinking actively about.

    2. RR

      The top frontier labs gobble up every available GPUs. At the same time-

    3. MG

      Yeah

    4. RR

      ... you wanna promote newer companies that are gonna become the enterprises of tomorrow. How are you thinking about balancing?

    5. MG

      From the very beginning of when we launched AWS, startups have been the lifeblood of the core of what we do. These are the innovators that are at the edge of technology, understanding what's possible. We're very intentional about keeping capacity available for the startups. We recently announced we're gonna be buying two million NVIDIA GPUs over the next couple of years. We-

    6. RR

      Your CapEx this year is what, 200 or-

    7. MG

      $220 billion dollars, uh, for '26, and we don't anticipate slowing down any time soon because the demand is just massive.

    8. RR

      With all the debate around AI extension risks and the Hugging Face attack, what are CEOs asking you about all these things?

  2. 0:55 – 2:51

    $169B, growing 37%

    1. RR

      Welcome to the pod, Matt. What a time we are living in. So I have a lot of topics to talk to you about.

    2. MG

      Awesome.

    3. RR

      So-

    4. MG

      Thanks for having me. I'm excited.

    5. RR

      Yeah, yeah, absolutely. So let's start actually right from the beginning, right? So you were the first GM for EC2-

    6. MG

      Mm-hmm

    7. RR

      ... and that was 2006, right? And today you guys are, what, one sixty, one seventy billion in revenue?

    8. MG

      Yeah, about a hundred and sixty-nine, hundred seventy billion, yep.

    9. RR

      Yeah, growing thirty-five-

    10. MG

      Thirty-seven, 37%, yeah

    11. RR

      ... thirty-- That's 37% at 169.

    12. MG

      Yeah, it's, uh-

    13. RR

      That's crazy

    14. MG

      ... there's a ton of op-- And it, it's, yeah, it's, it's, uh, it's interesting to think about it from, you know, we're there in day, day, day one when-

    15. RR

      Yeah

    16. MG

      ... we had like, you know, the first-

    17. RR

      Yeah

    18. MG

      ... dollar of revenue. But-

    19. RR

      Yeah

    20. MG

      ... um, yeah. And it's, you know, the interesting thing is we're still at the early stages of what the business can be and what the opportunity is for customers. Um, most workloads still, you know, there's a huge amount of workloads still that live on-prem today, and the amount of compute, um, that people are doing every single day is, is more than it was the day before. And so you see the tailwind from AI, you see the tailwind from migration into the cloud, and, um, you know, the, the, the business has grown really fast, and it's, it's been a super fun part to be a, um, thing to be a part of.

    21. RR

      Yeah, it is. It is. I mean, they'll be writing, uh, history books about this and business books about this for a long time to come. I wanna touch on the on-prem. That's, that sort of boggles my mind. I obviously did my best to keep them there for a long time.

    22. MG

      You did. You built a lot of stuff-

    23. RR

      [laughs]

    24. MG

      ... on-prem back in the day of-

    25. RR

      Yeah, yeah. I know, but, uh-

    26. MG

      That's true. We're trying to now get all of that to move into AWS.

    27. RR

      Yeah. So we can talk about that later. But, uh, so if you think about EC2 in the early days, and you got your start, uh, with obviously selling to startups and, and today the AWS cloud service, I mean, as you reflect on the evolution, what's, what stands out?

    28. MG

      Yeah.

    29. RR

      Part one, and then part two we'll talk about how this nature of how you serve startups has changed.

  3. 2:51 – 12:26

    Why startups are the lifeblood

    1. MG

      Sure. Well, l- like you said, actually, it's, uh, it's funny. So I actually interned for, um, AWS in 2005 when we were first-- kind of was an internal project. It was my business school internship, and my project was actually to come up with, uh, an analysis of who we thought AWS would be most interesting to, and, uh, the, the answer was startups, like probably-

    2. RR

      Yeah

    3. MG

      ... not surprisingly. And, uh, and so, you know, from the very beginning of when we launched AWS, startups have been the lifeblood of, of the core of what we do for, for a number of reasons. One is the value proposition is just so attractive-

    4. RR

      Yep

    5. MG

      ... what AWS provides to startups. Um, and, and we, we spend a lot of time and effort making sure that we're great partners to the startups, helping them not just provide infrastructure, but also advice on how to get your company up and running and how to think about your architecture so it'll scale eventually and, and a whole bunch of things that we do for the startups. But we also think that for us it's just good business, 'cause-

    6. RR

      Yeah

    7. MG

      ... the startups today are the enterprises of tomorrow. And, and, you know, it's, it's an imprecise number, but today we estimate that, you know, maybe thirty, forty percent of, of AWS revenue comes from companies that was a one time, you know, startups in, in AWS' lifetime.

    8. RR

      Yeah. Yeah, yeah.

    9. MG

      So and it's, uh, you know, and it-- So it's fun to see, for us to see the companies grow over time and, and to get bigger and become enterprises effectively. Um, and so that's why we invest so much in startups and why we pay so much attention to the, the brand new, you know, two people in a garage type startups. And, um, and, and not f- um, and, and not just because of the business outcomes, but also because that's who we learn from. These are the, these are the innovators that are at the edge of technology, understanding what's possible, uh, pushing our services to say what they'd want more of, what could help them go faster, what could help them achieve their, uh, outcomes. And a lot of times they're pushing more than the, the banks or the healthcare companies or the governments. The startups are the ones that are pushing that envelope, and it really helps us to be better and make sure that we're ahead of that wave where the banks are gonna want some tech, some capability five years down the road that startups want today.

    10. RR

      Yeah, yeah. And, uh, compared from then to now-

    11. MG

      Mm-hmm

    12. RR

      ... what have the, how have the startups changed in what they want from you? Obviously-

    13. MG

      Yeah

    14. RR

      ... they all want a lot of GPUs, and we can talk about that. But-

    15. MG

      Yeah

    16. RR

      ... besides that.

    17. MG

      Well, I will say, uh, there's a couple of things that have changed. One is, um, you know, startups started at a much smaller size than, than [chuckles] when, when we first started, right? They, they might have gotten ten million dollars of funding, and they had a app idea that they were kind of slowly iterating on. Now it's like, you know, from day one, they're valued at a billion dollars. They have two hundred million dollars of funding. They-

    18. RR

      Yeah. Yeah.

    19. MG

      [laughs] It's like it's a, you know, it's a, a, a, a team and an idea, and all of a sudden they're worth a billion dollars, which, um... So it's, you know, I think the brand-

    20. RR

      They just go from our offices to your offices.

    21. MG

      [laughs] That's right, that's right. Um, uh, and you know, like the, the size and scale and ambition of the ideas, I think requires more capital. They're, they're bigger to start with. Um, they're obviously more expensive too, right? To, you know, to, um, uh, to go after kind of a training a model or doing something that, that a lot of the folks are doing today. Um, so that's number one, which is just the size that they're, that, that they start from is, is really big. Um, but I would say number two is, uh- There's some things haven't changed, right? They, they actually are still thinking about, okay, once I scale, how do I think about an architecture? Um, how do I think about, um, security? How do I think about performance? How do I think about having all the capabilities that I need? How do I think about setting up my IM setup so that when I have more than three employees, this thing-

    22. RR

      Yeah

    23. MG

      ... is gonna work and scale? Um, which is a lot of times why they like AWS as opposed to just, you know, going to a, a Neocloud or something like that. They actually need the, all of those other security and, and, um, and capabilities that, that come around from, um, from training. Um, and so I think that's, that's, you know, that's something that hasn't changed and, and is exactly the same. I think the, the scale at which the ramp-up, um, is definitely different today. I also think increasingly we're seeing, um, teams that want a cloud that is great to work with agents and not just with people.

    24. RR

      Mm-hmm.

    25. MG

      Um, and I think that's where we've spent a lot of time is-

    26. RR

      Yeah, that plays to your strengths. Yeah

    27. MG

      ... how do you think about, you know, e-exactly, how do you think about broad scale? How do you think about performance? How do you think about, um, you know, an interface that is, you know, a, a, a well-defined API interface that agents can actually-

    28. RR

      Yeah

    29. MG

      ... easily traverse and work across? Um, and so that's also something that we both, I think we're naturally set up to do well, but that we've actually doubled down on to ensure things like, you know, how can you start a database in three seconds and, and, and, and really get to, um, some capabilities that agents are excited about.

    30. RR

      Um, I should have looked, but I haven't lately. But have you introduced, uh, any specific new services that are explicitly targeted at agents or people building agents?

  4. 12:26 – 16:24

    Rethinking cloud for agents

    1. RR

      what has been some of the hardest things to accommodate as agents have taken over? I mean, you made a name-

    2. MG

      Mm-hmm

    3. RR

      ... just the way you guys became the phenomenon that you became was serving developers-

    4. MG

      Yeah

    5. RR

      ... right? Um, and then, uh, infra teams. Um, now developers are substituted by agents, and pretty soon infra teams are being substituted by agents.

    6. MG

      Yeah.

    7. RR

      Are there-- What has been some of the hardest things for you as the largest service provider in the world to, to handle that transition?

    8. MG

      Well, you know, look, I think, um, like I said, much of our infrastructure was, is-

    9. RR

      Yeah, the whole building blocks are broad

    10. MG

      ... is, is actually pretty well set up-

    11. RR

      Yeah

    12. MG

      ... to, to handle the scale, um, which is good. I, I think there's been a couple of things which are interesting, which are, I think, interesting paradigm shifts where, um, you could argue that some of our systems are, um, I don't, wouldn't say overengineered, but, um, as an example, many agents wanna create a database, uh, do a little bit of work, and then have the database go away. You don't really need that database to have five nines of durability.

    13. RR

      Exactly.

    14. MG

      Right. And so there's some things that we rethink there, where when you're creating an Aurora database, like does your production database, like you do want five or-

    15. RR

      Yeah

    16. MG

      ... n nines of data, you know, you want-

    17. RR

      Yeah, yeah

    18. MG

      ... you, you, you need durability, you want availability, you want all of those things. And for the agent use case, that's, that's arguably or maybe not arguably like overengineered for, for what we need. And so, [clears throat] you know, it's hard. We don't, um-- And we kind of have this, this, uh, um, like a belief that we don't really wanna have like a non-durable option that's-

    19. RR

      Yeah

    20. MG

      ... that's gonna cause problems either, 'cause you never really quite know if the database that's created wants to stay around for a long time or a short amount of time. And so we're, we're, um, trying to do the hard work to think about how do you accomplish both of those things, where you can create it quickly and throw it away, and you don't really waste a lot of resources. But, um, but, um, but if you do want it, and it can be durable and stay for a long time, it can actually grow into a big production database. So those are some trade-offs that we think about actively as we think about how does the, the, the more traditional kind of, "This is gonna be my production system," kind of capabilities, um, match with, uh, with, with some of the more transient nature of the infrastructure that agents, uh, want to use. Um, so that, that's one, I think. But there's, there's a, there's a bunch. I think the, um, scale and speed and latency of creation of, of durable resources is another one that's interesting. Um, you know, I think the other one that we actively think about too is just are there new building blocks that agents are gonna want that we just didn't really need before?

    21. RR

      Exactly.

    22. MG

      And so compute sandboxes, um, uh, gateways, um, agent permissions versus people or, or, or service role permissions. Um, a lot of those are things that we have built and are building and thinking actively about 'cause they are just brand-new building blocks. It's not using the existing building blocks differently, but it's brand-new ones, where pretty, pretty clearly you want a different permissions for agents. You don't just wanna give it Raghu's permissions-

    23. RR

      Yeah, exactly

    24. MG

      ... and let it go do whatever you can do. You actually want, um, you know, very time boxed, short-term permissions to just go do a task. You actually may not wanna give permissions to a whole tool in at all. You may wanna get very fine-grained permissions-

    25. RR

      Exactly

    26. MG

      ... of what an agent's able to do from a sandbox, right? You actually want a sandbox-

    27. RR

      Yeah

    28. MG

      ... that, um, uh, you wanna, um-

    29. RR

      Virtualization is having its act too-

    30. MG

      That's right

  5. 16:24 – 21:27

    The GPU allocation problem

    1. MG

      And yeah.

    2. RR

      So I've rat-holed enough into agents. I'll come back to it later, but-

    3. MG

      Yeah

    4. RR

      ... such a fascinating topic.

    5. MG

      Super cool.

    6. RR

      But, uh, let me t- uh, switch gears a little bit and ask about something that, uh, every one of our startups face-

    7. MG

      Yeah

    8. RR

      ... and we get asked most frequently, which is: How can we get GPUs, right?

    9. MG

      [chuckles] Yeah.

    10. RR

      I mean, obviously you guys are running massive GPU farms-

    11. MG

      Yeah

    12. RR

      ... and increasing it every day. But, uh, the top frontier labs gobble up every available GPUs, more power to them. How do you deal with the internal-- At the same time you wanna-

    13. MG

      Yeah

    14. RR

      ... promote newer companies that are gonna become the enterprises of tomorrow, like you said. Um, how are you thinking about balancing-

    15. MG

      Yeah, it's, it's a great question

    16. RR

      ... the capacity needs and these small companies that don't have a lot of credit and whatnot, so.

    17. MG

      Yeah. Well, yeah, it, it's a great question. Um, and, and a couple of these things make it more challenging too. One is, um, the, the CapEx expense needed to go deploy kind of all of the compute that everyone needs right now, um, is massive, right? And so we are, uh, I think-

    18. RR

      I think your CapEx this year is what, two hundred or something?

    19. MG

      Two hundred and twenty billion dollars, uh, for '26. It's, uh, I, I-- By, by-- At, at one point I saw that it's, you know, it's, it's a pretty-- I mean, that is a larger expense than we've ever had, maybe any company has ever had in a, in a single year. Um, and, um, and we don't anticipate slowing down anytime soon because the demand is just massive. And so at some point, you're limited by how fast can you build data centers, how fast can you deploy capital, how, how fast can you get memory and chips and all of those kind of things. Um, and so all of those are different points, constraints that we're, you know, whether it's- Power data centers, capital memory chips, they're all kind of constraints at various times. Construction people that build buildings like-

    20. RR

      Yeah

    21. MG

      ... you know, construction people are, are, are at a premium today. Um, and so all of those we, you know, we, we work really hard to make sure happen. Um, and then with the capacity that we are able to deploy, which is still a huge amount and, and, and not enough, we know, um, we really think intentionally about what is that allocation strategy. And so, um, we're great partners with the, the large frontier labs, the Anthropics and OpenAIs and Meta and, and other large customers. And so those folks are, are really good customers of ours.

    22. RR

      Yep.

    23. MG

      And, uh, and we wanna make sure that we invest in them. We have large enterprise customers, whether it's, uh, Salesforces or, or JPMCs or, you know-

    24. RR

      Yep

    25. MG

      ... other large companies that, that have demand for fewer numbers of GPUs, but or, or accelerators. Sometimes they want Trainium, sometimes they want, um, uh, NVIDIA GPUs. Um, but we wanna make sure that we can support them as well. And we're very intentional about how we make sure that we have capacity for startups. And, and so what we do is we actually do allocate, and we basically say, okay, we're gonna keep-- We, we could... You're right, we could sell every single, um, uh, ex- you know, GPU or, or AI accelerator we had to probably just to the la- the, the big frontier labs and call-

    26. RR

      Yeah

    27. MG

      ... it a day. Um, we choose not to do that 'cause we actually wanna keep growing the, the full ecosystem. They get a large number, but, but, um, we wanna keep supporting, um, a broader set of customers 'cause we actually think that both the whole ecosystem will be more healthy for us. There's some diversification, but it's also just we know these are gonna be big companies over time. Um, and so we try to support them. Um, I saw recently that we say, you know, yes, in some way, shape, or form to something like 60% of the, the requests we eve-eventually get. Sometimes it's a little bit later, sometimes it's in a different region, sometimes it's a slightly different configuration than the customers are looking for. Um, but we really try to lean in and, and, and try to allocate as much as we can. And every single startup you have wants more. [laughs]

    28. RR

      Yeah, exactly.

    29. MG

      You know, and so-

    30. RR

      There's no question about it

  6. 21:27 – 30:22

    Is the AI CapEx a bubble?

    1. RR

      Yeah. Yeah, of course.

    2. MG

      And, um, uh, and you know, I think we were for the longest time actually, um, we, you know, we, we were investing ahead of where the demand was and, um-

    3. RR

      Yeah. No, I think that's-

    4. MG

      ... you know, one of the, one of the most painful things is that with the real ramp of, of GPUs, like a lot of the elasticity has unfortunately, you know, kind of gone away.

    5. RR

      Just gone. Yeah.

    6. MG

      Um, and so hopefully we'll get back to... And in our core compute and, and storage and things, that elasticity is still there and that, but, um, but you know, the, the, but, but so we've been spending for, for quite a bit of time and, and we feel really good about the spend that we're making now. Um, which is, you know, I think I get lots of questions sometimes about how you feel about that spend and like, are you nervous about-

    7. RR

      Yeah

    8. MG

      ... the bubble and other things like that.

    9. RR

      Yep.

    10. MG

      And I will say, you know, we, um, because of the position we have, like one, we do take this diversified approach, and so not all of our capacity is bundled up in one customer. And I think you go to a, some of these, whether they're Neoclouds or some of the, the other providers out there and, you know, you'll see sometimes concentrations of 30, 40, 50, 60%-

    11. RR

      Yep. Yep

    12. MG

      ... with one or two customers. We're nowhere near that obviously. We're, you know, single digit percentages at the highest and, and, and usually it's less than that. For one, I think we have a lot re- less risk on, on one particular customer. But also because we have that rich set of services, like AWS is where people really are coming to launch their production workloads. And so the majority of our usage today actually is either core compute and storage and inference, which is part of that application. And so those are the workloads that I think just aren't gonna go away.

    13. RR

      Yep.

    14. MG

      Just 'cause we see enterprises getting positive ROI. You go talk to the customers, and you say, "At the capability today and the cost today, are you seeing positive returns to your business?" And almost to a person, they'll say like, "Oh, yeah."

    15. RR

      Yeah.

    16. MG

      And so you're like, well, that's not gonna go away. There's no bubble in which they, they stop spending-

    17. RR

      Yep

    18. MG

      ... on that, you know.

    19. RR

      Yep.

    20. MG

      And so, you know, and you know this, like the VC model, like is every billion-dollar startup gonna make it? No, they won't. But you know, that's, that's kind of the game.

    21. RR

      That is the game.

    22. MG

      And it's been, that's been true for 50 years. [laughs]

    23. RR

      Yeah. Yep.

    24. MG

      It is. They haven't always been billion-dollar startups, but-

    25. RR

      The numbers have changed, but the principles haven't changed

    26. MG

      ... but, but the principles are the same, right?

    27. RR

      Yeah.

    28. MG

      You, you bet on 10, and one makes it-

    29. RR

      Yeah

    30. MG

      ... and pays for the other 10 or whatever the, whatever the percentage is. Hopefully it's higher than one out of 10.

  7. 30:22 – 33:07

    Clearing up data center myths

    1. RR

      Yeah. So, um- Obviously, there's a lot of wide-ranging debate about data centers.

    2. MG

      Mm-hmm.

    3. RR

      Right? And, uh, it's clear that folks like us where we stand, but do you think as an industry we have not done a good job-

    4. MG

      Yeah

    5. RR

      ... of explaining, uh, why data centers are good for America-

    6. MG

      Mm-hmm

    7. RR

      ... and generally the world, but, uh-

    8. MG

      Yeah

    9. RR

      ... and, uh, how do you-- what's the internal talk-

    10. MG

      Yeah

    11. RR

      ... amongst Andy's team on, uh, how do we deal with this?

    12. MG

      Yeah. Well, look, I, I think, um, and, and I think you'll hear more from us over this, and I agree, I think we need to be more vocal and be more fr-upfront, um, 'cause we actually do a ton that's, that's really beneficial, both for the communities we operate in, um, for the... You know, we, we think a ton about how do we bring renewable energy to these data centers?

    13. RR

      Yeah.

    14. MG

      How do we think about being water positive? Um, actually, our, our, our data centers use a, a really, really small amount of water. We mostly use free air cooling. How do we think about being great participants in the communities where we are and, um, and how do we bring high-paying jobs to the communities we operate in? Um, and not all data center operators do that.

    15. RR

      Yeah.

    16. MG

      And I think there are some well-

    17. RR

      And you have a 20 year-

    18. MG

      There are some well-chronicled examples of others out there that are not great, um, at that, and they just don't really pay attention to regulations. They think that the rules don't apply or they just, you know, launch really quickly without thinking about those. And I think that it causes a problem for the whole industry 'cause everybody kinda gets lumped into that.

    19. RR

      Yeah.

    20. MG

      And so, um, so look, I think we'll be-- we, we, I think you're right. We, um, you know, vocally self-critical, we need to be more vocal about the benefits that we do bring, and think about additional ways that we can help communities understand the benefits that, that we bring to them, both for the, the services they use, right? If, if you usually go to a community and say, "Well, do you not wanna use Netflix?" And they'll be like, "No, no. Like I, I still want Netflix." [laughs] Like it's, uh, um... And um, you know, like it's important for us to think about and, and highlight the benefits that we bring, where I, I recently saw a report where one of the communities that we operate in, the-- everybody in that county pays $5,000 a year less in taxes because of the taxes that we bring to that.

    21. RR

      Yep.

    22. MG

      And we don't, we don't tell them. They don't even know it, right?

    23. RR

      Yep. Exactly.

    24. MG

      It's just invisible to them.

    25. RR

      Yeah.

    26. MG

      And so I think we just need to be more, um, um, clear about those benefits that we bring. 'Cause I think if you told the communities, "By the way, your tax bill is $5,000 less than it would otherwise be if we weren't here," they might have a little bit of different thought about-

    27. RR

      Yeah

    28. MG

      ... the building that's over there.

    29. RR

      Yeah, yeah.

    30. MG

      Um, so, and, and but, but not everyone does that. Not everyone kinda, um, um-

  8. 33:07 – 40:06

    The Graviton and Trainium bet

    1. RR

      Yeah. And, uh, before we leave the hardware topic, uh, um, I wanna touch upon Trainium.

    2. MG

      Mm-hmm.

    3. RR

      And your whole history with, uh, building your own, uh, chips, right?

    4. MG

      Yeah.

    5. RR

      Um, we were one of your first partners using Nitro a long time ago.

    6. MG

      Mm-hmm.

    7. RR

      And since then, uh, Graviton and you made tremendous progress. So what was the thinking that led to saying, "Look, we're gonna do our own thing"?

    8. MG

      Yeah.

    9. RR

      And then how has that progress been and where are y-you-

    10. MG

      Yeah

    11. RR

      ... with respect to Trainium?

    12. MG

      It's, it's, it's actually a fascinating story, and I think it's a great example of where, um, Amazon AWS will innovate and will iterate over time and continue to think bigger about what we can do, but, but, um, but kind of prove our way there as opposed to, you know... And so take this as an example. Um, this is probably now, I don't know, ten, eh, it was probably about 13, 14 years ago, we were seeing that there was a pretty significant, um, virtualization tax on the overall number of resources. And, um, and we were kinda thinking about how do we-- Our customers were telling us, "I want bare metal performance," and they, you know, they're comparing to-

    13. RR

      Yeah

    14. MG

      ... having all the resources of a, of a server. And so the first thing that we did is that, um, we took a, a network offload card and virtualized all of our network virtualization and pulled it off onto a, an offload card so that network virtualization got closer to, to bare metal performance.

    15. RR

      Yeah.

    16. MG

      And back then it was not quite bare metal, but it was close- it was closer. And then we got really excited about that, and we said, "Okay, what if we could move storage virtualization off as well?" Right? And, and, um, and none of the network offload cards could do that. And then we s- we found this one company who had some Arm cores on an offload card, and they were doing it for other reasons. I can't remember their original purpose, but we're like, "Could you use those to do storage virtualization and some of these other functions?" And they're like, "Maybe." And so we really iterated with them. This was the Annapurna team.

    17. RR

      Yeah, this is Annapurna team.

    18. MG

      And, um, and just loved that team. Like really innovative, really mission driven-

    19. RR

      Fantastic team

    20. MG

      ... really wanting to solve problems. Um, and so we acquired them, and we said, "Look, could you build us a slightly bigger card that actually could take all the network virtualization off and basically give us a bare metal server that has no virtualization on it, no VM virtualization, everything is through APIs on the card?" And 'cause we had this view that, one, performance would be much better, resource re- uh, resource utilization would be better. It'd be-- We'd, we'd get-

    21. RR

      The security isolation.

    22. MG

      Se- and security isolation would be much, much better. And we tell people, you know, could then legitimately tell people, we have no access to any of your VMs that are running there. And, and, um, and this has been a huge benefit for us for the last decade, where frankly, like we've been leading and, and others-

    23. RR

      Yeah

    24. MG

      ... have been kinda slow to do this because they... This is not a generalized thing that people can do. But, um, so we got to that, and we basically said, "Look, we're, we're making a lot of progress here. Um, what if we take..." You know, and there's a bunch of Arm cores that were-

    25. RR

      Yeah

    26. MG

      ... on this offload card, and we said, "What if we turn that into a server?" Uh, and we did that first with Graviton. It was a very underpowered, um, very small server that we launched. Um, and, and customers were excited. They're like, "Ah, I'd love to have an Arm server. This is super interesting." And so we went down the path and, um, and Graviton, um, you know... And part of what we did is we looked and saw that there's the, the slope of, you know, Arm cores were getting faster.

    27. RR

      Yep.

    28. MG

      And where the-- You saw the power utilization and, and the graph where the... And you, you knew the intercept was gonna happen for where this architecture was gonna be really good for, for parts, and they just needed somebody to drive the ecosystem and, and get some of the pieces in place. So we did that with Graviton, um, and Graviton's been a, a runaway hit at this point. Um, you know, it's-

    29. RR

      Have you been public about what percent of your fleet is Graviton?

    30. MG

      We land more Graviton ships, uh, uh, every year than, um- Uh, than, than any other-

  9. 40:06 – 45:43

    Where enterprises are stuck on agents

    1. MG

      it's both.

    2. RR

      Yeah. Now, uh, let's get back to, uh, uh, talking about agents, but from a perspective of, uh, large enterprises-

    3. MG

      Mm-hmm

    4. RR

      ... or large medium enterprises.

    5. MG

      Yeah.

    6. RR

      Where are they in their adoption? Um-

    7. MG

      Mm-hmm

    8. RR

      ... and, uh, have they-- What sort of benefits are you seeing them reap-

    9. MG

      Yeah

    10. RR

      ... already, and, uh, what is the roadmap for them as far as you can tell from your vantage point?

    11. MG

      Yeah. It's, it's a really good question. I think it's one that, um, that we've spent a lot of time thinking about. And when I talk to customers all out there, um, today they view... You know, they're getting a lot of value out of what they've done today, and I would say the agents that most enterprises have built are, are relatively simple and straightforward. Um, and they're starting to think about, um... And they're, um, mostly non-autonomous, right?

    12. RR

      Yeah.

    13. MG

      There are still kind of people in the loop, if you will. And so I think we're... And, and by the way, there's, customers are still getting lots of value out of that today, and so they're really thinking about, "How do I have these be autonomous but in a safe way?" And I think there's two things that, that I think hold customers back today from just continuing to, to scale. And it's, it's already a pretty big business today, but I think it's, it has a massive opportunity to really change every single customer out there and every single workflow and, and really thinking about it. And so number one is just how to think about it. I think what we vers- originally saw was that enterprises had a workflow, and they're saying, "Great, I would have the enterpri-- I would have an agent go do the same workflow."

    14. RR

      Yeah.

    15. MG

      And what we encourage them to really do is think, not just replicate, you know, Bob does step one, two, three, four, five, so agent is gonna do one, step one, two, three, four, five, and then Bob's gonna check it at the end. That's not really the model you want. You want to actually step back and say, "If I wanna accomplish something, how can an agent do it differently?"

    16. RR

      Yeah.

    17. MG

      It can do it in a massively parallelized way. It can try 50 different things and get to that. And, and how do you help it get to that right outcome and rethink how a computer would solve a problem versus a human solving a problem? And so one of the things is us just helping customers understand how to think about that and really kind of have that blank slate, because that's where you really get value, is not just replicating what you're doing today, but, but thinking from a greenfield approach about how you go solve a problem completely differently. And I'm sure that's how many of your startups are thinking about this too.

    18. RR

      Yeah, yeah.

    19. MG

      Is how do you help customers greenfield solve a problem, not replicate the thing that happens today.

    20. RR

      Yeah.

    21. MG

      So that, that is number one.

    22. RR

      People running fleets of agents and swarms or whatever you wanna call them-

    23. MG

      Yeah

    24. RR

      ... very common these days.

    25. MG

      And, and, and you just wanna think about it, you know. And, and so enterprises are not as, as... Again, this is one where you learn from the startups and you try to how do you apply that to an enterprise world where they're, you know, an insurance company is not necessarily as forward-leaning, but they would love to figure out how they can have a better, um, you know, approval workflow or something like that. Um, so that's number one. But then the second one is this, this how do you turn those into fully autonomous workflows, and how do you actually trust the agents?

    26. RR

      Yeah.

    27. MG

      And so we're spending a lot of time thinking about how do we build services to- Help enterprises feel like their systems are secure, and that they can trust an agent to make a decision, that they can have the right guardrails, they can have the right permissions on their data, that it's not gonna delete production systems, that it's not gonna make tragic mistakes. And right now, I think that nervousness is probably holding people back-

    28. RR

      Yeah, that's huge

    29. MG

      ... that maybe appropriately, by the way, is holding enterprises back from just saying, "Okay, go nuts." Like, you don't actually want an agent to just go crazy and accidentally delete a production database. That's gonna be-

    30. RR

      Yeah

  10. 45:43 – 56:03

    AI risk, Hugging Face, and Continuum

    1. RR

      going into a big frontier model. What enterprises should really do is to take an open source model and then post train on your own data and workflows and traces and whatnot.

    2. MG

      Mm-hmm.

    3. RR

      Where do you stand on that? Are you seeing customers actually trying to do that-

    4. MG

      Yeah

    5. RR

      ... or do you guys-

    6. MG

      Well, yeah

    7. RR

      How do, how do you think about that?

    8. MG

      It's, it's a great-- The first point, I, I wholeheartedly agree on that first point, like the customer's en-- like enterprise data is their most valuable asset.

    9. RR

      Yeah.

    10. MG

      And so from the very beginning, it's why we built Bedrock like we did. Um, we have a guarantee that your data never leaves your VPC. And so you're-- if you're running inside of Bedrock, um, your data doesn't go back to the model provider.

    11. RR

      Yeah.

    12. MG

      They never see your prompts. Um, that stays inside of your own trusted environment. Um, and so that, that is why, um, enterprises kind of run, they, they prefer to run on top of Bedrock, and it's why you see that business growing massively, like hundreds and hundreds. I mean, it's e-every, um, every month we see that just business, just every week we see that business exploding, and then, and it's why, um, you see OpenAI workloads migrating to Bedrock. It's why you see Anthropic really growing really rapidly. I mean, so whether you're using open models or, um, uh, or, or, um, closed frontier models, um, I think Bedrock is a great solution that our en-- our, our customers told us, by the way, like, you know we-- if you remember three years ago, I got a lot of heat-

    13. RR

      Speaking of bad names, that's a good name, though.

    14. MG

      Yeah, Bedrock is a good name.

    15. RR

      Yeah.

    16. MG

      That's good. But we, we got a lot of heat actually for being slow to the AI world-

    17. RR

      Yes

    18. MG

      ... because we actually built the foundations of this where we said, "Look, we're not just gonna rush out a service. We really wanna think about how do we make sure that we protect our customers' data and build a service that we think is gonna be durable for the use cases that we knew about." And if you remember, we got a lot of heat, and we said, "Look, we're gonna go build the right thing." And now as people move from proof of concepts to production, vast majority of them are landing in AWS on Bedrock for, for muchly-- One of the reasons is because of this. It's also because of the, the set of services that we have. We also offer open models. We offer proprietary models. Um, we offer, um, a whole set of capabilities around those Agent Core. We build these building blocks, so it's easier to build agents with any of the models that you want, um, whether they're in Bedrock or out of Bedrock-

    19. RR

      Yeah, yeah

    20. MG

      ... for that matter. You can use Gemini or other things for it. But, um, but I think that's a, it's a, it's a differentiating piece for us, and it's a super important thing to think about because having that data go back into the, the model provider, I think is a, is a dangerous thing. You talk about open-weights models, though. I do think that there's a, a scenario that I'm excited about. Um, uh, we're, we're really ramping up our support of open-weights models and, and trying to build a good, um, environment. And, and frankly, this is where today, I, I think, and I think it's true, a lot of customers believe that they have meaningful proprietary data that if they could-

    21. RR

      Yeah

    22. MG

      ... mix in, you know, do some post-training, do some fine-tuning, um, um, to an open-weights model that they could distill down, they can actually get a better performing model at a lower price. Um, there's a bunch of pieces in here that have to work out well. They actually have a good eval to actually prove that that's true.

    23. RR

      Yeah.

    24. MG

      Uh, most people are doing that in SageMaker today. I think there's more that we can do to make that easier.

    25. RR

      Mm-hmm.

    26. MG

      But actually, like, if you go look at where are people doing that, they actually do it in SageMaker on AWS today.

    27. RR

      Oh, really? Okay.

    28. MG

      And they actually host the inference via SageMaker.

    29. RR

      So SageMaker, SageMaker is getting a new lease of life as a-

    30. MG

      Uh, I mean, it is. It's, it's-

Episode duration: 56:17

Install uListen for AI-powered chat & search across the full episode — Get Full Transcript

Transcript of episode rn_afJaPldg

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.