CHAPTERS
- 0:00 – 1:03
AWS is tuning the cloud for agentic workflows
Matt Garman opens by arguing that agentic workflows perform especially well on AWS because of performance work and emerging agent-specific building blocks. He frames the shift as moving from cloud interfaces built for humans to interfaces and primitives built for machines.
- •Agentic workflows depend on fast, well-defined APIs and low tail latency
- •AWS is building agent-oriented primitives like compute sandboxes, gateways, and agent-specific permissions
- •Performance characteristics (e.g., P999 latency) matter more for agents than humans
- 1:03 – 2:51
From EC2’s early days to $169B+ at 37% growth: why AWS still feels “early”
Raghu and Matt reflect on AWS’s growth from the first dollar of revenue to roughly $169–170B and 37% growth. Matt emphasizes that migration from on-prem remains a huge runway, now amplified by AI-driven compute demand.
- •AWS scale today: ~$169–170B revenue and ~37% growth
- •Large share of workloads still remain on-prem, leaving significant migration tailwinds
- •AI adds a new compute tailwind on top of ongoing cloud adoption
- 2:51 – 5:04
Startups as the lifeblood: learning, future enterprise revenue, and ecosystem strategy
Matt explains why startups have been central since AWS’s inception, both as a business engine and as a source of product insight. He notes that a substantial portion of AWS revenue comes from companies that started as AWS-native startups.
- •AWS originally identified startups as the primary early adopter segment
- •Startups push the tech envelope sooner than regulated enterprises, shaping AWS’s roadmap
- •Estimated ~30–40% of AWS revenue comes from companies that were once startups
- •AWS aims to be more than infrastructure: architecture guidance, security, scaling help
- 5:04 – 7:25
What startups want now: bigger from day one, and increasingly “cloud for agents”
The discussion shifts to how startups have changed: larger funding rounds, higher valuations, and more compute-intensive ambitions. Matt also highlights a new requirement—cloud platforms that work well for agents, not just human operators.
- •Startups start bigger (capital, ambition, burn) than in the mid-2000s
- •They still need scaling architecture, IAM, security, and operational depth
- •Agents introduce new expectations: fast provisioning, machine-friendly interfaces, and predictable latency
- 7:25 – 9:25
Agent-focused AWS services and performance: Bedrock, Agent Core, AWS Context, tail latency
Matt describes how AWS is adapting services and introducing new layers so agents can discover and use data across systems. He points to AWS Context (preview) as a context layer across data stores, and explains why tail latency becomes a key bottleneck for agent workflows.
- •Agent Core and Bedrock as explicit building blocks for agent development
- •AWS Context aims to help agents find relevant data across S3, Aurora, and other stores
- •Agents are sensitive to tail latency (e.g., P999 S3 latency) that humans often ignore
- •AWS claims agentic workflows tend to perform better on AWS due to these optimizations
- 9:25 – 12:21
Re-architecting onboarding and usability for agents: instant AWS accounts and no forced migration
Raghu presses on whether “legacy” AWS services must be rethought for an agent-driven world. Matt says core primitives are strong, but AWS is simplifying the initial user experience so new users (and agents) can start quickly without losing the ability to scale into full AWS controls later.
- •Agents can deploy to AWS today if instructed (“build on AWS; use these credentials”)
- •AWS is simplifying account setup: fast signup, defaults behind the scenes, no credit card required (rolling out)
- •Goal: an easy on-ramp that still becomes a full-featured AWS account when needed
- •Avoids a ‘simple now, migrate later’ trap by making the initial account real AWS
- 12:21 – 16:30
New cloud primitives for agents: transient infrastructure, sandboxes, and time-boxed permissions
Matt outlines harder technical shifts driven by agents: ephemeral resource creation and deletion, and the need for agent-specific security models. He highlights microVM-based sandboxing and Firecracker as a foundation for fast, secure isolation.
- •Agent workflows often need temporary databases/resources, conflicting with “always durable, always production-grade” defaults
- •AWS is balancing fast creation/teardown with the option to grow into durable production systems
- •New primitives: compute sandboxes, gateways, agent permissions distinct from human roles
- •Time-boxed, fine-grained permissions are critical to prevent destructive actions
- •Firecracker microVMs enable fast startup with strong security boundaries
- 16:30 – 20:40
The GPU allocation problem: balancing frontier labs, enterprises, and startups
The conversation turns to GPU scarcity and how AWS allocates limited accelerator capacity across customer segments. Matt stresses intentional allocation for startups despite pressure to sell everything to the largest frontier labs.
- •GPU capacity is constrained by power, construction, chips, memory, and supply chain limits
- •AWS supports frontier labs and large enterprises but reserves capacity intentionally for startups
- •AWS aims to say “yes” to most requests eventually, sometimes via different regions/configs/timing
- •AWS announced plans to buy ~2 million NVIDIA GPUs over the next couple of years
- 20:40 – 24:43
Is AI CapEx a bubble? AWS’s risk posture and why ROI-driven workloads persist
Raghu asks about bubble risk amid surging CapEx, and Matt explains why AWS is comfortable investing aggressively. He contrasts AWS’s diversified customer base and production workload focus against more concentrated providers, and points to customer ROI as a stabilizing force.
- •AWS CapEx cited at ~$220B for ’26; demand-driven and not expected to slow soon
- •Diversification reduces concentration risk versus providers dependent on 1–2 large customers
- •Most usage is production workloads: compute, storage, inference—less likely to vanish in a downturn
- •Enterprises report positive ROI from AI at current capability/cost levels
- •VC power-law failures are expected; durable companies and infrastructure remain valuable
- 24:43 – 30:22
Planning myths and real constraints: power, renewables, multi-year supply chains, and “The Goal”
Matt explains how scaling cloud infrastructure now requires long-horizon planning for power, grid interconnects, and component supply chains. He argues there is never a single bottleneck—constraints shift over time and vary by geography.
- •AWS now invests in power projects (renewables and nuclear), both grid and behind-the-meter
- •Supply planning has moved from quarters to multi-year horizons (’26–’28 and beyond)
- •Constraints rotate: power, HBM/memory, foundry capacity, networking parts, SSDs, connectors, etc.
- •Geographic non-fungibility matters (capacity and power differ by region)
- •AWS tracks deep supply-chain tiers to prevent shocks like prior disk drive shortages
- 30:22 – 33:07
Clearing up data center myths: water use, community benefits, and transparency
Raghu raises public skepticism around data centers, and Matt argues the industry hasn’t communicated benefits well enough. He highlights AWS efforts on renewables, water efficiency, and local economic impact, while noting that a few bad actors can tarnish the industry’s reputation.
- •AWS claims minimal water usage and heavy reliance on free-air cooling
- •Data centers can bring high-paying jobs and meaningful local tax base benefits
- •Example cited: communities paying ~$5,000 less per year in taxes due to data center taxes
- •Need for better transparency and proactive communication to communities
- •Industry reputational risk driven by a small number of poor operators
- 33:07 – 38:27
The in-house silicon strategy: Nitro → Graviton → Trainium (and why it worked)
Matt walks through AWS’s chip journey, beginning with offload cards for virtualization (Nitro) and evolving into full CPUs (Graviton) and AI accelerators (Trainium). He frames the strategy as iterative problem-solving: improve performance, security isolation, and cost efficiency at scale.
- •Nitro began as network virtualization offload, expanded to storage and full virtualization separation
- •Annapurna acquisition enabled AWS to build bespoke hardware for performance and security isolation
- •Graviton started modestly but became a major cost/performance lever for customers
- •Claim: Graviton offers ~20% better performance at ~20% lower cost; broad adoption among top customers
- •Trainium built in anticipation of AI growth; now on Trainium3 with strong demand
- 38:27 – 40:06
Trainium’s role today: sold-out capacity, inference strength, and Bedrock’s backbone
Matt clarifies that despite the name, Trainium is heavily used for inference—especially for large-model serving—due to architecture and cost-performance advantages. He notes Bedrock traffic largely runs on Trainium and that demand is strong enough to sell out capacity through much of next year.
- •Trainium is used for both training and inference; naming split with Inferentia is outdated for modern large models
- •Claim: Trainium may be among the best inference options on cost/performance for large workloads
- •Most of Bedrock inference runs on Trainium
- •Partnerships include major labs (Anthropic, OpenAI) plus startups building on Trainium
- •Trainium4 announced (not yet launched in transcript) as future training-cluster direction
- 40:06 – 45:34
Where enterprises are stuck on agents: rethinking workflows, autonomy, evals, and guardrails
Matt describes enterprise agent adoption as valuable but mostly non-autonomous, with humans still in the loop. The biggest blockers are conceptual (not just automating existing human workflows) and operational/safety (trust, permissions, guardrails, and continuous evaluation).
- •Enterprises mostly deploy simpler, non-autonomous agents today but still see ROI
- •Big unlock is redesigning workflows for agent-native parallelism rather than copying human steps
- •Trust and safety concerns slow autonomy: permissions, guardrails, sandboxing, preventing destructive actions
- •Enterprises need help with evals, drift monitoring, data labeling, and production measurement
- •AWS’s FDE approach aims to train customers in ~45 days rather than create long-term consultant dependence
- 45:34 – 56:17
Data privacy, open-weights models, and security: Bedrock isolation, SageMaker post-training, Continuum
The closing section covers data governance (enterprise data as crown jewels), model choices, and security concerns highlighted by recent attacks. Matt positions Bedrock as privacy-preserving (data stays in customer VPC) and points to SageMaker for fine-tuning open-weights models, then pivots to Continuum as AI-powered security scanning and prioritization.
- •Bedrock design goal: enterprise prompts/data stay inside the customer’s VPC; model providers don’t see them
- •Customers increasingly move from PoCs to production on Bedrock; multi-model support (open and proprietary)
- •Enterprises fine-tune/post-train open-weights models primarily in SageMaker and host inference there
- •CEO concerns: safe agent deployment, attack surfaces, and software supply-chain threats (e.g., Hugging Face incident)
- •Continuum uses AI to find and prioritize vulnerabilities using environmental context; positioned as exposing AWS internal practices
