Skip to content
a16za16z

Building the Cloud for AI Agents | AWS CEO Matt Garman

a16z’s Raghu Raghuram sits down with AWS CEO Matt Garman to discuss how AI is reshaping the cloud, from the needs of AI-native startups to infrastructure increasingly designed for agents. Matt explains how AWS is adapting as agents write code and manage infrastructure, why it’s reserving scarce GPU capacity for startups, and where custom chips like Trainium and Graviton fit into the AI stack. They also discuss Amazon's $220 billion capital investment, the shifting bottlenecks in the infrastructure buildout, what enterprises need to trust autonomous agents, and how AWS's own teams are building with agents. Timestamps: 00:00 - Intro 00:55 - $169B, growing 37% 02:51 - Why startups are the lifeblood 12:26 - Rethinking cloud for agents 16:24 - The GPU allocation problem 21:27 - Is the AI CapEx a bubble? 30:22 - Clearing up data center myths 33:07 - The Graviton and Trainium bet 40:06 - Where enterprises are stuck on agents 45:43 - AI risk, Hugging Face, and Continuum Resources: Follow Matt Garman on X: https://x.com/mattsgarman Follow Matt Garman on LinkedIn: https://www.linkedin.com/in/mattgarman Follow Raghu Raghuram on X: https://x.com/RaghuRaghuram Learn more about AWS: https://aws.amazon.com/ Stay Updated: If you enjoyed this episode, be sure to like, subscribe, and share with your friends! Find a16z on X: https://twitter.com/a16z Find a16z on LinkedIn: https://www.linkedin.com/company/a16z Listen to the a16z Show on Spotify: https://open.spotify.com/show/5bC65RDvs3oxnLyqqvkUYX Listen to the a16z Show on Apple Podcasts: https://podcasts.apple.com/us/podcast/a16z-podcast/id842818711 Follow our host: https://x.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see http://a16z.com/disclosures.

Matt GarmanguestRaghu Raghuramhost
Oct 8, 202656mWatch on YouTube ↗

CHAPTERS

  1. 0:00 – 1:03

    AWS is tuning the cloud for agentic workflows

    Matt Garman opens by arguing that agentic workflows perform especially well on AWS because of performance work and emerging agent-specific building blocks. He frames the shift as moving from cloud interfaces built for humans to interfaces and primitives built for machines.

    • •Agentic workflows depend on fast, well-defined APIs and low tail latency
    • •AWS is building agent-oriented primitives like compute sandboxes, gateways, and agent-specific permissions
    • •Performance characteristics (e.g., P999 latency) matter more for agents than humans
  2. 1:03 – 2:51

    From EC2’s early days to $169B+ at 37% growth: why AWS still feels “early”

    Raghu and Matt reflect on AWS’s growth from the first dollar of revenue to roughly $169–170B and 37% growth. Matt emphasizes that migration from on-prem remains a huge runway, now amplified by AI-driven compute demand.

    • •AWS scale today: ~$169–170B revenue and ~37% growth
    • •Large share of workloads still remain on-prem, leaving significant migration tailwinds
    • •AI adds a new compute tailwind on top of ongoing cloud adoption
  3. 2:51 – 5:04

    Startups as the lifeblood: learning, future enterprise revenue, and ecosystem strategy

    Matt explains why startups have been central since AWS’s inception, both as a business engine and as a source of product insight. He notes that a substantial portion of AWS revenue comes from companies that started as AWS-native startups.

    • •AWS originally identified startups as the primary early adopter segment
    • •Startups push the tech envelope sooner than regulated enterprises, shaping AWS’s roadmap
    • •Estimated ~30–40% of AWS revenue comes from companies that were once startups
    • •AWS aims to be more than infrastructure: architecture guidance, security, scaling help
  4. 5:04 – 7:25

    What startups want now: bigger from day one, and increasingly “cloud for agents”

    The discussion shifts to how startups have changed: larger funding rounds, higher valuations, and more compute-intensive ambitions. Matt also highlights a new requirement—cloud platforms that work well for agents, not just human operators.

    • •Startups start bigger (capital, ambition, burn) than in the mid-2000s
    • •They still need scaling architecture, IAM, security, and operational depth
    • •Agents introduce new expectations: fast provisioning, machine-friendly interfaces, and predictable latency
  5. 7:25 – 9:25

    Agent-focused AWS services and performance: Bedrock, Agent Core, AWS Context, tail latency

    Matt describes how AWS is adapting services and introducing new layers so agents can discover and use data across systems. He points to AWS Context (preview) as a context layer across data stores, and explains why tail latency becomes a key bottleneck for agent workflows.

    • •Agent Core and Bedrock as explicit building blocks for agent development
    • •AWS Context aims to help agents find relevant data across S3, Aurora, and other stores
    • •Agents are sensitive to tail latency (e.g., P999 S3 latency) that humans often ignore
    • •AWS claims agentic workflows tend to perform better on AWS due to these optimizations
  6. 9:25 – 12:21

    Re-architecting onboarding and usability for agents: instant AWS accounts and no forced migration

    Raghu presses on whether “legacy” AWS services must be rethought for an agent-driven world. Matt says core primitives are strong, but AWS is simplifying the initial user experience so new users (and agents) can start quickly without losing the ability to scale into full AWS controls later.

    • •Agents can deploy to AWS today if instructed (“build on AWS; use these credentials”)
    • •AWS is simplifying account setup: fast signup, defaults behind the scenes, no credit card required (rolling out)
    • •Goal: an easy on-ramp that still becomes a full-featured AWS account when needed
    • •Avoids a ‘simple now, migrate later’ trap by making the initial account real AWS
  7. 12:21 – 16:30

    New cloud primitives for agents: transient infrastructure, sandboxes, and time-boxed permissions

    Matt outlines harder technical shifts driven by agents: ephemeral resource creation and deletion, and the need for agent-specific security models. He highlights microVM-based sandboxing and Firecracker as a foundation for fast, secure isolation.

    • •Agent workflows often need temporary databases/resources, conflicting with “always durable, always production-grade” defaults
    • •AWS is balancing fast creation/teardown with the option to grow into durable production systems
    • •New primitives: compute sandboxes, gateways, agent permissions distinct from human roles
    • •Time-boxed, fine-grained permissions are critical to prevent destructive actions
    • •Firecracker microVMs enable fast startup with strong security boundaries
  8. 16:30 – 20:40

    The GPU allocation problem: balancing frontier labs, enterprises, and startups

    The conversation turns to GPU scarcity and how AWS allocates limited accelerator capacity across customer segments. Matt stresses intentional allocation for startups despite pressure to sell everything to the largest frontier labs.

    • •GPU capacity is constrained by power, construction, chips, memory, and supply chain limits
    • •AWS supports frontier labs and large enterprises but reserves capacity intentionally for startups
    • •AWS aims to say “yes” to most requests eventually, sometimes via different regions/configs/timing
    • •AWS announced plans to buy ~2 million NVIDIA GPUs over the next couple of years
  9. 20:40 – 24:43

    Is AI CapEx a bubble? AWS’s risk posture and why ROI-driven workloads persist

    Raghu asks about bubble risk amid surging CapEx, and Matt explains why AWS is comfortable investing aggressively. He contrasts AWS’s diversified customer base and production workload focus against more concentrated providers, and points to customer ROI as a stabilizing force.

    • •AWS CapEx cited at ~$220B for ’26; demand-driven and not expected to slow soon
    • •Diversification reduces concentration risk versus providers dependent on 1–2 large customers
    • •Most usage is production workloads: compute, storage, inference—less likely to vanish in a downturn
    • •Enterprises report positive ROI from AI at current capability/cost levels
    • •VC power-law failures are expected; durable companies and infrastructure remain valuable
  10. 24:43 – 30:22

    Planning myths and real constraints: power, renewables, multi-year supply chains, and “The Goal”

    Matt explains how scaling cloud infrastructure now requires long-horizon planning for power, grid interconnects, and component supply chains. He argues there is never a single bottleneck—constraints shift over time and vary by geography.

    • •AWS now invests in power projects (renewables and nuclear), both grid and behind-the-meter
    • •Supply planning has moved from quarters to multi-year horizons (’26–’28 and beyond)
    • •Constraints rotate: power, HBM/memory, foundry capacity, networking parts, SSDs, connectors, etc.
    • •Geographic non-fungibility matters (capacity and power differ by region)
    • •AWS tracks deep supply-chain tiers to prevent shocks like prior disk drive shortages
  11. 30:22 – 33:07

    Clearing up data center myths: water use, community benefits, and transparency

    Raghu raises public skepticism around data centers, and Matt argues the industry hasn’t communicated benefits well enough. He highlights AWS efforts on renewables, water efficiency, and local economic impact, while noting that a few bad actors can tarnish the industry’s reputation.

    • •AWS claims minimal water usage and heavy reliance on free-air cooling
    • •Data centers can bring high-paying jobs and meaningful local tax base benefits
    • •Example cited: communities paying ~$5,000 less per year in taxes due to data center taxes
    • •Need for better transparency and proactive communication to communities
    • •Industry reputational risk driven by a small number of poor operators
  12. 33:07 – 38:27

    The in-house silicon strategy: Nitro → Graviton → Trainium (and why it worked)

    Matt walks through AWS’s chip journey, beginning with offload cards for virtualization (Nitro) and evolving into full CPUs (Graviton) and AI accelerators (Trainium). He frames the strategy as iterative problem-solving: improve performance, security isolation, and cost efficiency at scale.

    • •Nitro began as network virtualization offload, expanded to storage and full virtualization separation
    • •Annapurna acquisition enabled AWS to build bespoke hardware for performance and security isolation
    • •Graviton started modestly but became a major cost/performance lever for customers
    • •Claim: Graviton offers ~20% better performance at ~20% lower cost; broad adoption among top customers
    • •Trainium built in anticipation of AI growth; now on Trainium3 with strong demand
  13. 38:27 – 40:06

    Trainium’s role today: sold-out capacity, inference strength, and Bedrock’s backbone

    Matt clarifies that despite the name, Trainium is heavily used for inference—especially for large-model serving—due to architecture and cost-performance advantages. He notes Bedrock traffic largely runs on Trainium and that demand is strong enough to sell out capacity through much of next year.

    • •Trainium is used for both training and inference; naming split with Inferentia is outdated for modern large models
    • •Claim: Trainium may be among the best inference options on cost/performance for large workloads
    • •Most of Bedrock inference runs on Trainium
    • •Partnerships include major labs (Anthropic, OpenAI) plus startups building on Trainium
    • •Trainium4 announced (not yet launched in transcript) as future training-cluster direction
  14. 40:06 – 45:34

    Where enterprises are stuck on agents: rethinking workflows, autonomy, evals, and guardrails

    Matt describes enterprise agent adoption as valuable but mostly non-autonomous, with humans still in the loop. The biggest blockers are conceptual (not just automating existing human workflows) and operational/safety (trust, permissions, guardrails, and continuous evaluation).

    • •Enterprises mostly deploy simpler, non-autonomous agents today but still see ROI
    • •Big unlock is redesigning workflows for agent-native parallelism rather than copying human steps
    • •Trust and safety concerns slow autonomy: permissions, guardrails, sandboxing, preventing destructive actions
    • •Enterprises need help with evals, drift monitoring, data labeling, and production measurement
    • •AWS’s FDE approach aims to train customers in ~45 days rather than create long-term consultant dependence
  15. 45:34 – 56:17

    Data privacy, open-weights models, and security: Bedrock isolation, SageMaker post-training, Continuum

    The closing section covers data governance (enterprise data as crown jewels), model choices, and security concerns highlighted by recent attacks. Matt positions Bedrock as privacy-preserving (data stays in customer VPC) and points to SageMaker for fine-tuning open-weights models, then pivots to Continuum as AI-powered security scanning and prioritization.

    • •Bedrock design goal: enterprise prompts/data stay inside the customer’s VPC; model providers don’t see them
    • •Customers increasingly move from PoCs to production on Bedrock; multi-model support (open and proprietary)
    • •Enterprises fine-tune/post-train open-weights models primarily in SageMaker and host inference there
    • •CEO concerns: safe agent deployment, attack surfaces, and software supply-chain threats (e.g., Hugging Face incident)
    • •Continuum uses AI to find and prioritize vulnerabilities using environmental context; positioned as exposing AWS internal practices

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.