Skip to content
a16za16z

Building the Cloud for AI Agents | AWS CEO Matt Garman

a16z’s Raghu Raghuram sits down with AWS CEO Matt Garman to discuss how AI is reshaping the cloud, from the needs of AI-native startups to infrastructure increasingly designed for agents. Matt explains how AWS is adapting as agents write code and manage infrastructure, why it’s reserving scarce GPU capacity for startups, and where custom chips like Trainium and Graviton fit into the AI stack. They also discuss Amazon's $220 billion capital investment, the shifting bottlenecks in the infrastructure buildout, what enterprises need to trust autonomous agents, and how AWS's own teams are building with agents. Timestamps: 00:00 - Intro 00:55 - $169B, growing 37% 02:51 - Why startups are the lifeblood 12:26 - Rethinking cloud for agents 16:24 - The GPU allocation problem 21:27 - Is the AI CapEx a bubble? 30:22 - Clearing up data center myths 33:07 - The Graviton and Trainium bet 40:06 - Where enterprises are stuck on agents 45:43 - AI risk, Hugging Face, and Continuum Resources: Follow Matt Garman on X: https://x.com/mattsgarman Follow Matt Garman on LinkedIn: https://www.linkedin.com/in/mattgarman Follow Raghu Raghuram on X: https://x.com/RaghuRaghuram Learn more about AWS: https://aws.amazon.com/ Stay Updated: If you enjoyed this episode, be sure to like, subscribe, and share with your friends! Find a16z on X: https://twitter.com/a16z Find a16z on LinkedIn: https://www.linkedin.com/company/a16z Listen to the a16z Show on Spotify: https://open.spotify.com/show/5bC65RDvs3oxnLyqqvkUYX Listen to the a16z Show on Apple Podcasts: https://podcasts.apple.com/us/podcast/a16z-podcast/id842818711 Follow our host: https://x.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see http://a16z.com/disclosures.

Matt GarmanguestRaghu Raghuramhost
Oct 8, 202656mWatch on YouTube ↗

At a glance

WHAT IT’S REALLY ABOUT

AWS reshapes cloud for agents, GPUs, and custom AI silicon

  1. Matt Garman describes how AWS is rethinking cloud infrastructure for AI agents, emphasizing API-driven interfaces, fast ephemeral resource creation, and low tail latency that matters more to agents than humans.
  2. He explains AWS’s GPU allocation approach amid extreme scarcity, including intentional capacity reservations for startups and plans to acquire roughly two million NVIDIA GPUs over the next few years.
  3. Garman argues AWS’s soaring AI-driven CapEx is justified by diversified demand and enterprise ROI, while noting that real-world constraints increasingly shift to power, construction, and multi-tier supply chain limits.
  4. He outlines AWS’s custom silicon journey from Nitro and Graviton to Trainium, positioning Trainium as cost-effective for both training and inference and central to Bedrock’s underlying infrastructure.
  5. On enterprise agent adoption, he says most deployments remain human-in-the-loop, with major blockers being safe autonomy (permissions/sandboxing/guardrails) and operational maturity (evals, labeling, drift monitoring), areas where AWS is investing in products and field enablement.

IDEAS WORTH REMEMBERING

5 ideas

Cloud UX is shifting from “developer-first” to “agent-first,” with performance and tail latency as differentiators.

Garman argues that agents stress different parts of cloud infrastructure than humans do—especially API consistency, fast resource creation, and very low tail latency (e.g., P999). AWS is tuning foundational services (like S3) and adding new layers to make agent orchestration more reliable and “machine-friendly.”

AWS is trying to remove account and security setup friction without sacrificing enterprise-grade depth later.

AWS is introducing a faster onboarding path (e.g., Gmail signup, no credit card, defaults for VPC/IAM) to reduce friction for experimentation and agent-driven deployment. The goal is an easy on-ramp that still becomes a “real” AWS account without later migration.

GPU scarcity is managed as an ecosystem strategy: large labs matter, but startups are “lifeblood,” so capacity is intentionally partitioned.

Garman describes a core tension: frontier labs can absorb nearly all GPU supply, but AWS deliberately reserves capacity for startups because they become future enterprise revenue and push product innovation. He cites planned purchases of ~2M NVIDIA GPUs over a couple of years and an allocation approach that tries to say “yes” (even if later/elsewhere/config-changed) to a majority of requests.

AWS’s confidence in massive AI CapEx rests on diversified customers and production ROI, not single-customer bets.

He rejects the “AI CapEx bubble” framing by pointing to diversified demand across many customers and to production inference plus core compute/storage workloads that are delivering ROI. AWS believes enterprise spending persists because customers report positive returns at today’s capability and cost.

The bottleneck isn’t “GPUs” alone; it’s a moving target across power, supply chain, and geography.

At hyperscale, constraints rotate: power, construction labor, memory/HBM, foundry capacity, networking parts, region-specific limits, etc. AWS now plans power (including renewables and some behind-the-meter), supply chain, and server components multiple years out—work most customers can’t do themselves.

WORDS WORTH SAVING

5 quotes

From the very beginning of when we launched AWS, startups have been the lifeblood of, of the core of what we do.

— Matt Garman

Two hundred and twenty billion dollars, uh, for '26... and we don't anticipate slowing down anytime soon because the demand is just massive.

— Matt Garman

We could sell every single, um, uh, ex- you know, GPU or, or AI accelerator we had to probably just to the la- the, the big frontier labs and call it a day. Um, we choose not to do that 'cause we actually wanna keep growing the, the full ecosystem.

— Matt Garman

You actually want, um, you know, very time boxed, short-term permissions to just go do a task.

— Matt Garman

You want to actually step back and say, "If I wanna accomplish something, how can an agent do it differently?"

— Matt Garman

AWS growth and cloud migration tailwindsStartups as AWS “lifeblood” and future enterprisesAgent-first cloud requirements (APIs, latency, ephemeral resources)Compute sandboxes and agent-specific permissionsGPU scarcity, allocation strategy, and CapEx scalingData center power, renewables, and supply-chain constraintsCustom silicon: Nitro, Graviton, Trainium and Bedrock economics

High quality AI-generated summary created from speaker-labeled transcript.

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.