Skip to content
a16za16z

Why AI Demand Is Outrunning Compute Supply

a16z’s David George sits down with Gavin Baker to unpack the state of the AI boom, why demand for intelligence may still be dramatically underestimated, and why the outcome doesn't necessarily have to be winner-take-all. David and Gavin explore the possibility that frontier labs, open-source models, applications, clouds, and NVIDIA can all capture significant value as AI adoption expands. They dig into the economics of the infrastructure buildout, why compute investments can have unusually fast payback periods, and what happens when today's relatively small group of heavy AI users expands to hundreds of millions of people. They also debate the risk of an AI bubble versus an AI shortage, the backlash against data centers, orbital compute, the rise of multi-model architectures, and NVIDIA's position at the center of the AI supply chain. Gavin makes the case that the AI buildout could help reindustrialize America, while David explores whether the bigger near-term risk is not overbuilding, but failing to build enough. Timestamps: 00:00 - Intro 01:06 - Finding the Bear Case: Why Gavin Can't Find One 08:06 - How's This All Gonna Go Wrong? An "And" Thing, Not an "Or" 09:09 - Will Labs Reinvest All Their Profits Into Training Forever? 14:44 - The Demand Side: 30 Million Heavy Users & the Diffusion Question 19:16 - From Reactive Coding to Fully Autonomous Agents 30:18 - What Happens If There's a Massive Supply Shortage? 34:21 - Orbital Data Centers: The SpaceX Compute Play 40:00 - Starlink's $2 Trillion Market & the Heads-You-Win Compute Bet 44:04 - The Most Futuristic SpaceX Idea: Asteroid Mining 54:01 - Who Becomes the Abstraction Layer of Intelligence? 57:57 - Harvey, Cursor & Vertical AI Winners 01:00:26 - Jensen, Nvidia & the Central Bank of AI 01:12:16 - How Chip Deal Structures Reveal True Customer Preference Resources: Follow Gavin Baker on X: https://x.com/GavinSBaker Follow David George on X: https://x.com/DavidGeorge83 Stay Updated: If you enjoyed this episode, be sure to like, subscribe, and share with your friends! Find a16z on X: https://twitter.com/a16z Find a16z on LinkedIn: https://www.linkedin.com/company/a16z Listen to the a16z Show on Spotify: https://open.spotify.com/show/5bC65RDvs3oxnLyqqvkUYX Listen to the a16z Show on Apple Podcasts: https://podcasts.apple.com/us/podcast/a16z-podcast/id842818711 Follow our host: https://x.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see http://a16z.com/disclosures.

Gavin BakerguestDavid Georgehost
Aug 31, 20261h 14mWatch on YouTube ↗

CHAPTERS

  1. 0:00 – 5:50

    AI momentum feels relentless: no quantitative bear case (yet)

    Gavin says he’s actively tried to find a credible, quantitative “something is getting worse” datapoint across AI businesses—and hasn’t found one. Despite meaningful drawdowns in some AI public stocks, he argues underlying usage, revenue, and activity are still accelerating across major labs and open-source ecosystems.

    • Gavin’s standing question: “one quantitative datapoint getting worse” and he can’t find it
    • AI revenue/usage acceleration (OpenAI, open source, Groq) vs. weak stock performance
    • Market skepticism vs. business reality: the river-average-depth analogy
    • ‘Maybe everyone wins’ framing across labs, infra, and apps
    • Anthropic IPO dynamics: quiet period, rebasing revenue/accounting, gamesmanship
  2. 5:50 – 8:01

    Training vs. inference: the public-market tension for AI labs

    They explore how frontier labs can swing revenue dramatically by reallocating compute between inference (monetized) and training (capability). This creates a business model unlike prior internet companies: a lab can choose to sacrifice near-term revenue to pursue a step-function research gain.

    • Compute allocation can cut revenue quickly if shifted from inference to training
    • Public markets must adapt to revenue being partially ‘management-controlled’
    • Pricing and checkpoint release strategies materially affect monetization
    • Labs likely generate operating cash flow but little/no free cash flow soon
    • Strategic contrasts: OpenAI more commercial; Anthropic more mission-driven but still compute-constrained
  3. 8:01 – 11:51

    Why the ‘AI goes wrong’ story is an “and,” not an “or”

    David argues most bearish narratives assume a zero-sum outcome (frontier vs. open source vs. apps vs. clouds). Instead, they expect multiple layers to succeed simultaneously, with NVIDIA central across them all.

    • LP fear framing: “what crashes?” and why that misses the multi-winner structure
    • Frontier + N-1 models + open source + apps + clouds can all thrive together
    • NVIDIA positioned at the center regardless of which model paradigm wins
    • Fragmentation of models/approaches increases total compute demand
    • Zero-sum thinking is misleading for a general-purpose technology wave
  4. 11:51 – 14:32

    ROI math: sub-1-year paybacks drive massive compute buildouts

    They dig into the surprisingly fast payback periods for GPU clusters and “neocloud” economics, including financing structures. The conversation emphasizes that compute investments can scale into tens or hundreds of billions with unusually quick paybacks, attracting sophisticated capital.

    • Nebius/CoreWeave-style disclosures suggest ~9–10 month paybacks (or faster in spot)
    • Upfront customer payments reduce required equity capital
    • SpaceX clusters come online fast, improving paybacks further
    • Shift from per-GPU to per-megawatt thinking for pricing and planning
    • Financing is becoming standardized; useful lives extend, improving economics
  5. 14:32 – 19:16

    Demand-side reality: heavy users are few, diffusion is just starting

    Even today’s large AI revenues may be driven by a surprisingly small base of power users (likely under 10 million). With ~1.5B knowledge workers globally, they argue adoption is early—and supply constraints may be the real limiter.

    • Revenue concentration: a power law of token spend among engineers and companies
    • Many firms spend ~1% of comp on tokens; AI-native can reach 10%+
    • With 1.5B knowledge workers, diffusion runway remains enormous
    • Token consumption inside firms is growing extremely fast (100x in months)
    • Agentic workflows could create “endless token consumption” once automation clicks
  6. 19:16 – 20:36

    From reactive coding to autonomous agents: why usage explodes

    They contrast today’s “reactive” copilots/summarizers with emerging agentic systems that recommend actions and execute tasks. The leap from assistance to automation is framed as the next driver of demand and compute scarcity.

    • Claude/Codex boosted coding output; next is action-taking agents
    • ‘Recommended actions’ bots compound what other bots learned
    • Racing different agent stacks (GrokBot/Codex/Town) to automate workflows
    • Once users trust ‘Yes, automate,’ token demand scales dramatically
    • Coding as the leading wedge: highly verifiable, well-documented domain
  7. 20:36 – 23:53

    Bubbles, constraints, and the politics of data centers

    Gavin acknowledges tech bubbles are normal—overvaluation can cause overbuild—yet argues AI faces hard physical and political constraints that may prevent oversupply. He emphasizes copper/power bottlenecks and warns regulation and local opposition could worsen shortages.

    • Historical pattern: transformational tech → bubble → overbuild (often debt-fueled)
    • AI buildout is constrained by materials, power, and supply chain realities
    • Rising real rates and heavier regulation slow capacity expansion
    • Copper, power gear, labor, and permitting become limiting factors
    • Net: higher risk is undersupply through 2028 rather than glut
  8. 23:53 – 29:57

    The ‘tell the truth’ narrative: data centers as working-class revitalization

    They argue the AI industry is losing the public narrative and needs to communicate tangible benefits. Gavin claims data centers can transform towns via tax revenue and jobs, and many common objections (like water use) are overstated or debunked.

    • ‘Stay ahead of China’ is correct but too abstract to persuade the public
    • Data centers can 10x local tax revenue and revive declining towns
    • Trades may become more attractive economically than college in many cases
    • Water consumption critique is characterized as “debunked” in this discussion
    • Call for companies to publish concrete, human stories of AI benefits
  9. 29:57 – 34:21

    If supply stays tight: token prices rise and compute inequality emerges

    They outline a scenario where constrained buildouts lead to higher prices for intelligence—contrary to the assumption that AI always gets cheaper. The risk is a period where only large companies and wealthy individuals can afford frontier compute before advertising-like business models mature.

    • Possible 10x token price increase under severe shortages
    • Massive user surplus implies willingness to pay could absorb price hikes
    • Compute inequality becomes a political and social risk
    • Open source helps access, but tokens still require real compute
    • Conclusion: society should build far more data centers to avoid this outcome
  10. 34:21 – 39:39

    Orbital data centers: SpaceX’s swing-capacity bet on compute

    They describe orbital compute as small rack-like platforms with solar arrays and radiators—not ‘buildings in space.’ The key variable is Starship reusability: if launch costs fall enough, orbit can become a meaningful share of global compute, especially as Earth-based power/cooling costs inflate.

    • Orbital compute architecture: solar wings + radiators in sun-synchronous orbit
    • Physics feasibility vs. engineering execution: SpaceX scale as the advantage
    • Economics hinge on launch cost; Starship reusability flips the equation
    • Training likely stays on Earth (latency/speed-of-light constraints)
    • Timeline claim: co-designed “Rubin rack” targeting 2027–2028 window
  11. 39:39 – 44:04

    Starlink’s $2T market and SpaceX’s ‘heads-you-win’ compute optionality

    Even if orbital compute is debated, they argue SpaceX has multiple gigantic markets: broadband + mobile (approaching $2T) plus a rapidly growing AI ARR base. Heavy compute investment is framed as win-win: it powers first-party AI, and any excess capacity can be monetized as infrastructure with fast paybacks.

    • Broadband + mobile connectivity expands SpaceX’s addressable market dramatically
    • Bundling possibilities: Starlink + AI assistant + ads/enterprise offers
    • ‘Heads-you-win, tails-you-win’: first-party AI growth or infra monetization
    • Starbase Louisiana and high launch cadence support long-term capacity expansion
    • SpaceX as a durable infrastructure business even under app-level stumbles
  12. 44:04 – 54:01

    The most futuristic SpaceX idea: asteroid mining and off-world industry

    Gavin argues asteroid mining could become economically viable with Starship-scale logistics and robotics, enabling extraction of massive precious-metal resources. He links this to Bezos’s notion of Earth becoming ‘zoned residential’ while heavy industry moves to space, potentially reducing pollution on Earth.

    • Asteroid Psyche as an example of enormous precious-metal supply
    • Robotics (e.g., Optimus-like systems) enabling space-based industrial work
    • Concept: relocate heavy industry to space to reduce Earth-based pollution
    • Speculation: controlled capture/orbit and safe processing zones
    • Nearer-term futurism: Mars landings with robots, solar, batteries, and compute
  13. 54:01 – 55:30

    Who becomes the abstraction layer of intelligence? The platform battle

    They frame the next giant prize as owning the interface/router layer that brokers models, data, and workflows for enterprises and consumers. This is described as extraordinarily hard to execute, with competition from hyperscalers, data platforms, inference providers, and vertical apps.

    • Abstraction layer = ‘arbiter of intelligence’ for organizations and users
    • Execution is harder than it sounds: trust, routing, continuous upgrades, UX simplicity
    • Competition spans Microsoft, Databricks, Snowflake, Palantir, inference clouds, apps
    • Vertical integration and compute ownership drive long-term cost advantage
    • Analogy: building large-scale retail ops is simple to describe but hard to run
  14. 55:30 – 1:00:25

    Vertical AI winners: Harvey, Cursor, and the multi-model enterprise stack

    They highlight product-led vertical AI companies and platforms that already implement multi-model routing and fine-tuning. They argue enterprises will increasingly run their own tuned models on proprietary data while using frontier models selectively for planning or verification.

    • Cursor’s product-first path toward autonomy; coding is uniquely verifiable
    • Harvey’s success in legal as a documented, semi-verifiable domain
    • Fireworks Nexus as a concrete multi-model + fine-tuning + router implementation
    • Enterprise trend: own your tuned model to protect proprietary context/data
    • ‘Specialized intelligence’ vision is plausible but operationally complex
  15. 1:00:25 – 1:12:06

    Jensen, NVIDIA, and the ‘Central Bank of AI’ advantage

    Gavin argues NVIDIA’s strategy—vertically integrated yet broadly enabling—creates a powerful ecosystem moat across chips, systems, financing, and supply chain. He advises would-be accelerator competitors to find niches and integrate rather than confront NVIDIA head-on.

    • NVIDIA as ‘Central Bank/Fed of AI’: ecosystem control and standard-setting
    • Competitors should avoid direct confrontation; small share can still be huge
    • Cost of capital advantage: NVIDIA systems are the most financeable
    • Supply chain lock-up (fabs, memory, optics, components) amplifies dominance
    • Open source is framed as bullish for NVIDIA: lower margins → more consumption
  16. 1:12:06 – 1:14:25

    Deal structures as preference signals: how customers reveal what they want

    In a supply-constrained world, product demand is hard to measure because everything sells out. Gavin suggests reading true customer preference through deal structures—equity investments, residual value guarantees (RVGs), warrants tied to token pricing—and what these imply about risk sharing and confidence.

    • Scarcity distorts ‘preference’—customers may take anything available
    • Hierarchy of deals: chipmaker invests in customer vs. RVGs vs. warrants
    • RVGs can be positive NPV if covered by chip gross profit and structured well
    • Google/Amazon investments helped TPU/Trainium overcome cold-start adoption
    • Analyzing deal terms can reveal real confidence in performance and economics

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.