Skip to content
a16za16z

Building the Real-World Infrastructure for AI, with Google, Cisco & a16z

AI isn’t just changing software, it’s causing the biggest buildout of physical infrastructure in modern history. In this episode, live from Runtime, a16z's Raghu Raghuram speaks with Amin Vahdat, VP and GM of AI and Infrastructure at Google, and Jeetu Patel, President and Chief Product Officer at Cisco, about the unprecedented scale of what’s being built, from chips to power grids to global data centers. They discuss the new “AI industrial revolution,” where power, compute, and network are the new scarce resources; how geopolitical competition is shaping chip design and data center placement; and why the next generation of AI infrastructure will demand co-design across hardware, software, and networking. The conversation also covers how enterprises will adapt, why we’re still in the earliest phase of this CapEx supercycle, and how AI inference, reinforcement learning, and multi-site computing will transform how systems are built and run. 00:00 Intro 01:16 The Scale of the AI Buildout 03:00 CapEx, Demand Signals, and the Power Bottleneck 05:56 Data Centers, Scarcity, and Global Power Constraints 08:18 Rethinking Systems and Networking 10:08 Scale-Out vs. Mainframe Architectures 12:18 The Next Wave in Processor Innovation 14:36 Specialized Chips, Power Efficiency, and Geopolitics 16:14 Networking Evolution and Scale Challenges 18:52 Building Networks for AI: Power, Bursts, and Bottlenecks 21:00 Inference Architecture and Cost Reduction 24:00 AI Inside the Enterprise: Code Migration and Productivity 27:30 Rewiring Culture Around Rapid AI Adoption 29:40 Startups, Models, and Intelligent Routing Layers 31:55 The Future of AI Models, Agents, and Media 33:10 Closing Thoughts Resources Full Transcript: https://a16z.substack.com/p/surviving-the-ai-sprint-up-close Follow Raghu on X: https://x.com/RaghuRaghuram Follow Jeetu on X: https://x.com/jpatel41 Follow Amin on LinkedIn: https://www.linkedin.com/in/vahdat/ Find a16z on X: https://x.com/a16z Find a16z on LinkedIn: https://www.linkedin.com/company/a16z Listen to the a16z Podcast on Spotify: https://open.spotify.com/show/5bC65RDvs3oxnLyqqvkUYX Listen to the a16z Podcast on Apple Podcasts: https://podcasts.apple.com/us/podcast/a16z-podcast/id842818711 Follow our host: https://x.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures.

Jeetu PatelguestAmin Vahdatguest
Oct 29, 202532mWatch on YouTube ↗

CHAPTERS

  1. 0:00 – 3:20

    AI infrastructure is back—and the buildout is unprecedented

    Jeetu Patel and Amin Vahdat frame the current AI infrastructure moment as unlike any prior cycle, exceeding even the late-’90s internet buildout. They emphasize the mix of economic, geopolitical, national security, and speed implications driving urgency.

    • Scale and pace are described as 10–100x the internet era buildout
    • Infrastructure has become strategically critical again (beyond just tech)
    • No meaningful historical “priors” for this speed/scale
    • Panelists argue demand will likely exceed today’s projections (not a bubble)
  2. 3:20 – 5:30

    CapEx planning signals: utilization, turned-away demand, and hard physical limits

    The discussion shifts to where the industry sits in the spend cycle and what internal signals guide long-range planning. Amin highlights extreme utilization—even older TPU generations are fully booked—while warning that power, land, permitting, and supply chain constrain delivery.

    • Cycle is still early relative to demand pressure
    • Older TPUs (7–8 years old) at 100% utilization signals severe scarcity
    • Opportunity cost: real use cases are being deprioritized due to limited supply
    • Bottlenecks include power availability, land transformation, permitting, and supply chain
    • Constraint may persist 3–5 years; money isn’t the limiting factor—execution is
  3. 5:30 – 5:52

    Depreciation reality: hardware vs. power-and-space lifetimes

    A quick but important tangent clarifies that depreciation curves differ across infrastructure layers. Compute hardware can be procured more “just-in-time,” while the underlying facility investments are long-lived assets.

    • Hardware can be purchased closer to need (JIT)
    • Space and power infrastructure depreciate over 25–40 years
    • Mismatch between demand growth and facility timelines shapes planning risk
  4. 5:52 – 8:08

    Data center scarcity goes global: build where power is, then connect it

    Jeetu explains how enterprise adoption is still nascent while hyperscalers/neo-clouds are already confronting constraints. Because power can’t be concentrated easily, data centers are increasingly sited around available energy, driving demand for new network architectures including ‘scale-across’ between distant sites.

    • Enterprise data centers lag: major re-racking and per-rack power upgrades still ahead
    • Hyperscalers/neo-clouds face immediate scarcity of power, compute, and networking
    • Data centers are being built near power sources rather than importing power
    • Rising need for scale-up (within rack), scale-out (across racks), and scale-across (across sites) networking
    • Concept of two data centers behaving like one logical DC across ~800–900 km
  5. 8:08 – 10:40

    Scale-out isn’t dead: mainframe comparisons and stack reinvention

    The panel debates whether AI superclusters imply a return to mainframe-like architectures. Amin argues scale-out resource pooling remains dominant, but expects the full hardware-to-software stack to be reinvented through tight co-design—similar to how Google’s earlier systems matched its cluster model.

    • AI systems still operate as pooled, schedulable fleets (not fixed ‘supercomputers’)
    • Historical parallel: commodity scale-out was once radical but became default
    • Next 5 years likely produce an unrecognizable compute stack
    • Co-design lesson: scale-out software (e.g., storage/scheduling) enabled scale-out hardware
    • Future ‘mainframe-like’ systems will look fundamentally different
  6. 10:40 – 12:10

    Integrated systems and open ecosystems: designing like “one company”

    Jeetu emphasizes that integration across silicon-to-app will determine efficiency and lossiness, making partnerships essential. He argues the industry must collaborate across company boundaries while still preserving openness to avoid walled gardens.

    • Tight integration across the stack becomes a primary performance constraint (after power)
    • Multi-company ecosystems must operate with deep, months-long design partnerships
    • Pressure to move fast increases after deals—execution matters
    • Open ecosystem capabilities become important at every layer
  7. 12:10 – 14:11

    The golden age of specialized processors—and shrinking the chip innovation loop

    Amin predicts more specialization beyond GPUs/TPUs as power efficiency dominates economics. He highlights the long hardware turnaround (concept-to-production ~2.5 years) as a core constraint, and argues specialization will expand to serving and agentic workloads.

    • Specialized accelerators can be 10–100x more efficient per watt than CPUs
    • Next specialization targets: serving and agentic workloads, not just training
    • Major bottleneck: hardware iteration cycle is too slow for today’s pace
    • Need to shorten concept-to-production timelines dramatically
    • Power/space/cost savings will force specialization even if complexity increases
  8. 14:11 – 15:42

    Geopolitics shapes architectures: power abundance vs. advanced nodes

    The conversation connects processor choices to geopolitical realities. Amin contrasts China’s constraints on leading-edge nodes with potential advantages in power and engineering scale, suggesting regional optimization paths could diverge and evolve with regulatory and geopolitical shifts.

    • Different regions may optimize differently (e.g., older nodes + more engineering vs. leading-edge nodes)
    • Power availability becomes a strategic differentiator, not just a cost line item
    • Thermal and efficiency tradeoffs complicate ‘smaller node = better’ assumptions
    • Regulatory and geopolitical expansions could drive distinct architecture ecosystems
    • Emerging metric idea: engineering effort as a resource alongside watts
  9. 15:42 – 18:37

    Networking becomes the bottleneck: bandwidth, predictability, and burstiness

    Amin argues AI training/inference traffic pushes intra-building bandwidth demands to extreme levels and turns networks into primary bottlenecks. Because AI communication patterns are more knowable than general-purpose workloads, there’s room for optimization, but bursty utilization and shifting training sites make right-sizing difficult.

    • Intra-DC bandwidth demands are ‘astounding’ and rising
    • Networking is increasingly the limiting factor for performance
    • Known communication patterns enable potential switch/fabric optimizations
    • Workloads are extremely bursty—networks must handle short 100% spikes then idle
    • Wide-area ‘scale-across’ is intermittent; hard to justify permanent peak capacity
  10. 18:37 – 20:32

    Network as force multiplier: scale-up/scale-out/scale-across and silicon diversity

    Jeetu frames networking as the lever that converts scarce power and expensive compute into delivered performance. He also argues that inference vs. training will drive distinct infrastructure designs and stresses the strategic importance of silicon choice rather than a single-vendor ‘wrapper’ ecosystem.

    • If power is constrained and compute is the asset, the network is the multiplier
    • Efficiency gains in networking translate into more power budget for GPUs/accelerators
    • Inference and training optimize for different constraints (latency vs. memory, etc.)
    • Prediction: inference-native infrastructure will emerge (not just reused training stacks)
    • Emphasis on avoiding monoculture: silicon diversity matters for high-volume deployments
  11. 20:32 – 23:37

    Inference architecture: specialization, RL on the serving path, and cost-vs-quality dynamics

    Amin details how inference is being deployed with specialized configurations and software, with reinforcement learning and latency increasingly on the critical path. He explains prefill vs. decode as fundamentally different phases and notes that efficiency gains are real, but users continuously trade them away for higher model quality.

    • Inference is deployed with specialized configurations (hardware + software)
    • Reinforcement learning and latency constraints increasingly shape serving systems
    • Prefill and decode have different balance points; potential for different hardware
    • Inference cost is dropping 10–100x, but demand shifts toward higher quality models
    • Cycle repeats: efficiency improvements enable—and are consumed by—bigger models
  12. 23:37 – 28:09

    Enterprise wins: code migration, debugging, and the ‘culture reset’ for rapid adoption

    The panel moves to internal use cases at Google and Cisco, with coding as the standout. Amin describes AI-assisted instruction-set migration at massive scale, while Jeetu stresses that organizational behavior must adapt—tools evolve too quickly to dismiss after one bad trial.

    • Google: AI applied to large-scale instruction-set migration (x86→ARM; future agnosticism)
    • Large migrations historically cost ‘staff millennia,’ making AI leverage essential
    • Cisco: code migrations, debugging, and greenfield front-end work show strong gains
    • Older infrastructure code is harder; productivity gains vary by stack depth
    • Key cultural shift: retest tools frequently (weeks), not on 6–9 month cycles
  13. 28:09 – 28:46

    Non-engineering productivity: sales prep, legal review, and marketing workflows

    Jeetu outlines additional internal wins outside pure software engineering. He notes strong results in account research, contract review, and marketing—especially using LLMs to avoid starting from a blank page.

    • Sales/account call preparation improved materially
    • Legal contract review quality exceeded expectations
    • Marketing content benefits from ‘LLM-first draft’ workflows
    • Principle: never start from a blank slate—iterate from model output
  14. 28:46 – 32:47

    Advice to founders and what’s next: agents, durable moats, routing layers, and media

    In closing, the panel gives forward-looking guidance: agents and frameworks will extend how long tasks can run correctly, and startups should avoid thin wrappers around foundation models. They also predict major gains in image/video inputs and outputs as practical productivity tools, not just novelty generation.

    • Agents and agent frameworks will be transformative in the next 12 months
    • Startup warning: thin wrappers around others’ models are not durable
    • Moat idea: intelligent routing layers that choose among proprietary vs. foundation models
    • Cisco outlook: innovation across silicon, networking, security, observability, and apps
    • Next wave: image/video model capabilities become practical for education and productivity

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.