Skip to content
a16za16z

How Fei-Fei Li Is Rebuilding AI for the Real World

What if the next leap in artificial intelligence isn’t about better language—but better understanding of space? In this episode, a16z General Partner Erik Torenberg moderates chats with Fei-Fei Li, cofounder and CEO of World Labs, and Martin Casado, a16z General Partner and early investor in the company. Together, they explore the concept of world models—AI systems that understand and reason about the physical, 3D world—not just text. Fei-Fei, often called the “godmother of AI,” explains why spatial intelligence is a critical (and missing) component of today's AI systems, and why her new company is going all-in on solving this challenge. Martin shares the story of how he and Fei-Fei aligned on this vision long before it was trendy - and why it may define the future of robotics, creativity, and computation itself. From the limitations of LLMs to the promise of embodied AI, from personal anecdotes to deep technical insights, this is a discussion on what it truly means to build intelligence for the real (and virtual) world. Timecodes: 00:00 Spatial Intelligence 00:39 Fei-Fei Li’s Background 01:17 Building a World Model 05:14 Reflecting on AI's Evolution 08:07 The Importance of 3D Understanding 10:20 Unrolling Evolution: Why 3D Intelligence Is Harder Than Language 12:19 From Single Reality to Infinite Virtual Universes 16:52 3D vs 2D: Why 2D Isn’t Enough for Machines 17:57 Fei-Fei’s Personal Story of Losing Stereo Vision 19:24 Research and Development at World Labs Resources: Find Fei-Fei on X: https://x.com/drfeifei Find Martin on X: https://x.com/martin_casado Learn more about World Labs: https://www.worldlabs.ai/ Stay Updated: Let us know what you think: https://ratethispodcast.com/a16z Find a16z on Twitter: https://twitter.com/a16z Find a16z on LinkedIn: https://www.linkedin.com/company/a16z Subscribe on your favorite podcast app: https://a16z.simplecast.com/ Follow our host: https://x.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures.

Fei-Fei LiguestErik TorenberghostMartin Casadoguest
Jun 4, 202522mWatch on YouTube ↗

CHAPTERS

  1. 0:00 – 0:39

    Why spatial intelligence matters: AI needs a 3D “world model”

    Fei-Fei Li frames spatial (3D) intelligence as a core component of real-world intelligence, distinct from language. She hints at a future where accurate 3D understanding enables “infinite universes” for robots, creativity, and social experiences.

    • Spatial intelligence as a critical pillar of intelligence beyond language
    • The concept of “world models” (3D understanding + interaction)
    • A future of many simulated/virtual worlds built from 3D AI
    • Why she never needed LLM success to believe in world models
  2. 0:39 – 2:57

    Fei-Fei’s AI impact and the “data” revolution (and why World Labs needs a true partner)

    Erik and Martin contextualize Fei-Fei’s contributions, emphasizing her role in making data central to modern AI. Fei-Fei explains why she chose Martin as an early investor: not just capital, but an intellectual partner for deep-tech execution.

    • Fei-Fei’s reputation and track record across academia and industry
    • The shift from model-centric to data-centric AI progress
    • What she wanted in a “unicorn investor”: technical + go-to-market + constant collaboration
    • World Labs positioned as deep tech aiming at something fundamentally new
  3. 2:57 – 5:12

    The lunch-table moment: defining a “world model” as 3D structure and compositionality

    Martin recounts the origin story: amid LLM excitement, Fei-Fei points out what’s missing—an AI that models the world itself. Fei-Fei describes how she tested whether people truly understood the idea, and Martin was the first to clearly define it in 3D terms.

    • LLM hype created a contrast that clarified the missing piece
    • “World model” defined as understanding 3D structure, shape, and compositionality
    • Most people nodded at the term without substance
    • A key alignment that catalyzed World Labs’ direction
  4. 5:12 – 6:01

    Looking back on AI’s evolution: surprise at emergent behavior from data-hungry models

    Asked what would surprise her younger self, Fei-Fei notes an emotional surprise: how far data-driven methods have gone, producing unexpected emergent capabilities. This sets up the next question—why push beyond LLMs if scaling works so well?

    • Fei-Fei’s continued surprise despite championing data early
    • Emergence as a real phenomenon in large-scale models
    • Recognition of how far “data hungry” approaches can go
    • Transition toward questioning LLM sufficiency
  5. 6:01 – 8:07

    Why LLMs aren’t enough: language is lossy and the physical world isn’t made of words

    Fei-Fei explains her North Star: intelligence grounded in the 3D physical world, not just language. Language is powerful but incomplete—especially for perception, action, and building in the real world—so concentrated, industry-scale effort is needed to build world models.

    • Her motivation is the “North Star problem,” not company-building
    • Language as a lossy encoding of a richly structured 3D world
    • Physical perception and embodied intelligence are foundational in evolution and civilization
    • Need for focused compute, data, and talent to bring world models to life
  6. 8:07 – 8:58

    A simple thought experiment: blindfolded instructions vs seeing and reconstructing 3D

    Martin offers an intuition pump: describing a room via language is insufficient for reliable action, but seeing enables the brain to reconstruct 3D and manipulate objects. The chapter draws a sharp line between communicating concepts and navigating reality.

    • Language descriptions fail at conveying precise spatial reality
    • Vision supports internal 3D reconstruction for action
    • Navigation and manipulation require exact geometry and context
    • Why a model of the world is fundamentally different from a model of text
  7. 8:58 – 10:51

    “Unrolling evolution”: why spatial intelligence is older—and harder—than language

    Martin argues it’s surprising language came first in AI given the massive investment in robotics and autonomy. They connect difficulty to biology: spatial cognition is ancient (hundreds of millions of years), while language is recent—so replicating spatial intelligence may be the deeper challenge.

    • Robotics/AV as a costly, slow-moving proving ground for spatial difficulty
    • Language competence may be easier to emulate than embodied navigation
    • Spatial cognition predates humans by vast evolutionary timescales
    • Generative AI offers clues for tackling world-model learning
  8. 10:51 – 13:35

    What 3D intelligence unlocks: science, creativity, and embodied machines

    Fei-Fei broadens the stakes: spatial reasoning underpins breakthroughs from DNA structure to molecular geometry, and fuels human creativity. She outlines high-impact domains—from design and architecture to robotics—where 3D understanding is essential.

    • 3D reasoning as central to scientific discovery (DNA, molecular structures)
    • Creativity and design as inherently visual-spatial activities
    • Robotics defined broadly as embodied machines beyond humanoids/cars
    • Spatial intelligence as a horizontal capability across industries
  9. 13:35 – 14:28

    From one reality to infinite virtual universes: generation + reconstruction as the engine

    Fei-Fei describes a shift from humanity living in a single physical 3D world to creating countless digital worlds. The key is combining reconstruction (building 3D from observations) and generation (inventing new worlds), enabling a “multiverse” of experiences and training environments.

    • Human civilization historically constrained to one physical 3D world
    • Digital 3D worlds expand possibilities for travel, storytelling, socialization, robotics
    • Reconstruction + generation as the foundational technical combination
    • “Multiverse” framing: infinite, purpose-built universes
  10. 14:28 – 16:51

    Concrete capabilities: infer a full 3D scene from 2D views, then measure and manipulate

    Martin makes the abstract tangible: a system that can build a complete 3D representation from partial 2D inputs (including occluded surfaces). Once you have that representation, you can perform spatial operations—and also generate entirely new environments for games, media, and design.

    • 3D completion: modeling unseen parts (e.g., the back of a table)
    • Actionability: moving, stacking, measuring, and interacting with objects
    • Creativity: turning a single image into a navigable 360° world
    • Why “horizontal” platforms can power many applications like LLMs do
  11. 16:51 – 17:53

    Why 2D isn’t enough: physics and interaction demand the missing Z dimension

    They address a key objection—why not stay in 2D? Fei-Fei emphasizes physics and interaction happen in 3D, and Martin explains computers/robots can’t “mentally reconstruct” depth the way humans do; the Z-axis must be explicit for reliable action.

    • Physics, navigation, and composition occur in 3D space
    • Humans can infer 3D from 2D; robots generally can’t without explicit depth
    • Spatial tasks require distance and geometry for grasping and movement
    • 3D representation as a prerequisite for embodied autonomy
  12. 17:53 – 19:24

    Fei-Fei’s stereo-vision loss: a personal demonstration of depth’s role in safety

    Fei-Fei shares losing stereo vision temporarily after a corneal injury and how it immediately affected driving. Even with strong prior knowledge of the environment, the lack of depth cues made distance estimation unreliable—illustrating why machines need true 3D understanding.

    • Temporary loss of stereo vision created real-world functional limitations
    • Driving became frightening due to unreliable distance estimates
    • Prior knowledge couldn’t substitute for real-time depth perception
    • A vivid analogy for why 3D perception is indispensable in machines
  13. 19:24 – 22:25

    World Labs R&D: the technical lineage (NeRF, Gaussian splats) and building the right team

    Fei-Fei situates world-model research as newer than LLMs but built on years of progress in computer vision and 3D representation. She highlights key innovations (NeRF, Gaussian splats, early deep generative vision work) and explains World Labs’ bet: assemble elite cross-disciplinary talent to solve and productize this singular problem.

    • 3D world-modeling draws from prior vision research rather than starting from scratch
    • Key building blocks: NeRF (3D reconstruction), Gaussian splats (3D representation), early generative vision techniques
    • World Labs’ strategy: concentrated effort on one North Star problem
    • Need for a hybrid team spanning AI/data/modeling + graphics/representation + optimization/productization

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.