Skip to content
a16za16z

“The Future of AI is Here” — Fei-Fei Li Unveils the Next Frontier of AI

Fei-Fei Li and Justin Johnson are pioneers in AI. While the world has only recently witnessed a surge in consumer AI, our guests have long been laying the groundwork for innovations that are transforming industries today. In this episode, a16z General Partner Martin Casado joins Fei-Fei and Justin to explore the journey from early AI winters to the rise of deep learning and the rapid expansion of multimodal AI. From foundational advancements like ImageNet to the cutting-edge realm of spatial intelligence, Fei-Fei and Justin share the breakthroughs that have shaped the AI landscape and reveal what's next for innovation at World Labs. If you're curious about how AI is evolving beyond language models and into a new realm of 3D, generative worlds, this episode is a must-listen. Timestamps: 00:00 - Spatial Intelligence: A New Frontier 01:38 - Scaling AI: The Impact of ImageNet on Computer Vision 06:56 - The Role of Compute 09:16 - Data as the Key Driver 17:01 - Defining AI’s Ultimate Goal 18:58 - What is Spatial Intelligence? Unlocking 3D Understanding in AI 26:35 - Comparing Models: Spatial Intelligence vs. Language-Based AI 29:41 - 1D vs. 3D 32:39 - Building Immersive Worlds with Spatial Intelligence 35:11 - From Static Scenes to Dynamic Worlds 37:42 - The Future of VR and AR 40:42 - Creating Deep Tech Platforms 44:26 - Building a World-Class Team 45:54 - Measuring Success: Milestones in Spatial Intelligence Resources: Learn more about World Labs: https://www.worldlabs.ai Find Fei-Fei on Twitter: https://x.com/drfeifei Find Justin on Twitter: https://x.com/jcjohnss Find Martin on Twitter: https://x.com/martin_casado Stay Updated: Let us know what you think: https://ratethispodcast.com/a16z Find a16z on Twitter: https://twitter.com/a16z Find a16z on LinkedIn: https://www.linkedin.com/company/a16z Subscribe on your favorite podcast app: https://a16z.simplecast.com/ Follow our host: https://twitter.com/stephsmithio Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures.

Fei-Fei LiguestJustin JohnsonguestMartin Casadohost
Sep 20, 202448mWatch on YouTube ↗

CHAPTERS

  1. 0:00 – 0:29

    Why spatial intelligence is the next major AI bet

    Fei-Fei Li frames visual-spatial intelligence as a peer to language—possibly even more foundational for embodied agents and real-world interaction. She argues the timing is right due to a convergence of compute, better data understanding, and algorithmic advances, motivating World Labs’ focus.

    • Spatial/visual intelligence is positioned as a core pillar of intelligence alongside language
    • A ‘right moment’ thesis: compute + data understanding + algorithms have converged
    • North Star framing: focus the field/company around a fundamental capability
    • Sets up World Labs as the vehicle to pursue spatial intelligence
  2. 0:29 – 1:48

    How we got here: AI winters, deep learning takeoff, and multimodal expansion

    Martin Casado asks for a historical arc of modern AI. Fei-Fei describes moving from AI winter to the birth of modern AI, deep learning breakthroughs, and today’s “Cambrian explosion” across text, pixels, video, and audio.

    • From AI winter to deep learning’s breakout successes
    • Industry adoption accelerates first in language models
    • AI expands beyond text into vision, video, and audio
    • Framing today as a rapid proliferation of new model/application possibilities
  3. 1:48 – 3:43

    Justin Johnson’s path: the deep learning ‘recipe’ and the research boom

    Justin recounts getting hooked on deep learning after early landmark work (the ‘Cat paper’) and learning the now-classic formula: algorithms + lots of compute + lots of data. He describes the intense pace of discovery during his PhD years as core generative and vision building blocks emerged from academia.

    • Deep learning as a general recipe that scales with data and compute
    • Computer vision as an early proving ground for deep nets
    • Early generative modeling foundations developed during the PhD-era paper boom
    • Cultural shift: the broader world recently discovered what researchers felt for years
  4. 3:43 – 7:25

    Fei-Fei Li’s path: physics to AI, and the overlooked power of data

    Fei-Fei explains her route into AI via physics and computational neuroscience during a period when AI looked stalled publicly but was active scientifically. She highlights a key insight from her lab: scaling data, not just model cleverness, was crucial for generalization—leading to ImageNet and internet-scale vision datasets.

    • Physics training encouraged ‘audacious questions’ like intelligence
    • Machine learning era preceded deep learning, experimenting with many model families
    • Data was an underappreciated driver of progress and generalization
    • ImageNet as a deliberate bet to push vision datasets to internet scale
  5. 7:25 – 9:18

    Compute as the underrated unlock: AlexNet to modern GPUs

    Justin argues compute is the biggest unlock and still underestimated. He contrasts AlexNet’s 2012 training run (days on two consumer GPUs) with modern NVIDIA hardware, where an equivalent run could complete in minutes—illustrating how rapidly the feasible model/search space has expanded.

    • Compute growth over the last decade is ‘in the thousands’ in raw factor terms
    • AlexNet: ~60M params, trained for days on two GTX 580s
    • Modern hardware (e.g., GB200) collapses training time dramatically
    • Hardware scaling reshaped what research and products are feasible
  6. 9:18 – 12:10

    Data vs. compute—and the shift from supervised to self-supervised learning

    The conversation weighs two narratives: compute scaling vs. new data sources. Justin distinguishes the ImageNet era of supervised learning (human-labeled ontologies) from later epochs where models learn from less explicitly labeled data and broader internet structure, enabling more general capabilities.

    • ‘Bitter lesson’ perspective: prioritize compute-friendly approaches
    • ImageNet era: heavy human labeling, constrained tasks and ontologies
    • Later era: learning from data without explicit per-example labels
    • Language data carries implicit human structure differently than pixels
  7. 12:10 – 16:58

    Generative AI as a continuum: from image-text matching to text-to-image

    Fei-Fei and Justin map generative AI’s rise as a gradual progression rather than a sudden break. They trace steps through image-text alignment, captioning (pixels to words), style transfer, and early structured text-to-image generation using scene graphs and GANs.

    • Generative modeling existed theoretically earlier, but results weren’t compelling
    • Progression: alignment → captioning → style transfer → structured generation
    • 2015 neural style transfer as a visceral ‘gen AI’ moment for researchers
    • Scene graphs as an early bridge from language-like structure to image synthesis
  8. 16:58 – 19:58

    From research North Stars to World Labs: why focus on spatial intelligence now

    Martin transitions from Fei-Fei’s research journey to the founding of World Labs. Fei-Fei describes spatial intelligence as her next North Star after major milestones in visual storytelling, arguing today’s compute, data maturity, and new algorithms (including NeRF-related advances) make the bet timely.

    • North Stars guide long-term research direction and field advancement
    • Visual-spatial reasoning is essential for seeing, interacting, and building in the world
    • World Labs’ mission: unlock spatial intelligence as the next frontier
    • Key enablers: compute, improved data practices, and cutting-edge 3D methods (e.g., NeRF)
  9. 19:58 – 21:20

    Defining spatial intelligence: perceiving, reasoning, and acting in 3D/4D

    Justin offers a crisp definition: machines understanding space and time—objects, events, and their interactions across 3D plus dynamics. They clarify spatial intelligence applies to both the physical world and generated/simulated worlds.

    • Spatial intelligence = perception + reasoning + action in 3D and time
    • 4D framing: positions and interactions evolve over space-time
    • Applies to real environments and synthetic/generated environments
    • Goal is to move AI from data centers into rich world contexts
  10. 21:20 – 26:33

    Why now: NeRF and the merge of reconstruction with generation

    Justin explains the pivot toward 3D understanding from 2D observations and highlights NeRF as a breakthrough that catalyzed the field. Fei-Fei adds that computer vision’s long tradition in 3D reconstruction is now converging with generative modeling, making the reconstruction-vs-generation distinction increasingly blurred.

    • 3D data is hard to collect; leverage 2D projections to infer 3D structure
    • NeRF as a simple, powerful method to back out 3D from 2D views
    • Academia could still innovate here even as LLM compute needs outpaced labs
    • Reconstruction (seeing) and generation (imagining) are rapidly converging in vision
  11. 26:33 – 29:42

    Spatial intelligence vs. LLMs: 1D token sequences versus 3D-native representations

    Martin probes how spatial approaches differ from multimodal LLMs that also ‘see.’ Justin and Fei-Fei argue LLMs are fundamentally built on 1D sequence representations, whereas spatial intelligence puts 3D structure at the core—better matching physical reality and enabling richer interactions.

    • LLMs/multimodal models operate on 1D token sequences under the hood
    • Other modalities often get ‘shoehorned’ into sequence form
    • Spatial intelligence treats 3D structure as first-class in representation
    • Physical world constraints (geometry/physics) make 3D understanding qualitatively different from language modeling
  12. 29:42 – 32:36

    2D video vs. 3D worlds: affordances, interaction, and an ‘arc of intelligence’

    They unpack why 3D representations matter even if humans ultimately view 2D renders. A 3D-native model better supports user affordances like moving cameras/objects and interacting with environments, aligning with intelligence as the capacity to navigate and manipulate the world.

    • Separate ‘representation’ from ‘user-facing output’ (often still 2D)
    • 2D perception can imply 3D, but explicit 3D representations fit tasks better
    • Affordances: camera motion, object manipulation, consistent geometry
    • Spatial intelligence as a prerequisite for broad real-world and creative applications
  13. 32:36 – 37:42

    Use cases: world generation, new media, and dynamic interactive environments

    Justin outlines an evolution from generating images/clips to generating full interactive 3D worlds. They position this as a new media layer that could dramatically reduce the cost and labor of producing AAA-quality virtual experiences, unlocking many non-gaming applications.

    • From text-to-image/video to generating full 3D worlds
    • A ‘new media’ thesis: interactive experiences become far cheaper to create
    • Economic unlock: beyond $70 AAA games into niche/personalized worlds
    • Progression expected: static scenes first, then dynamic physics and interaction
  14. 37:42 – 40:42

    AR/VR and robotics: spatial intelligence as the operating system for mixed reality

    Fei-Fei connects spatial intelligence to spatial computing (e.g., Vision Pro) and the blending of real and virtual worlds. They extend the case to robotics, where a robot’s digital ‘brain’ must act in a 3D physical environment—making spatial intelligence the key bridge.

    • AR/MR needs real-time 3D understanding to blend digital content with reality
    • Hardware form factors may evolve (goggles, glasses, contact lenses)
    • AR could collapse the need for many separate screens via contextual overlays
    • Robotics: spatial intelligence connects digital compute to physical-world action
  15. 40:42 – 48:09

    Company strategy, team-building, and what success looks like for World Labs

    Martin asks how World Labs balances deep tech with multiple application areas. Fei-Fei positions the company as a platform model provider, then both discuss the multidisciplinary talent needed across ML, data, systems, and graphics; they close with milestones defined by real-world deployment and expanding possibilities.

    • World Labs as a deep tech platform company enabling many downstream use cases
    • Near-term pragmatism: some markets/devices aren’t ready for mass adoption
    • Team composition spans ML, infra, data, 3D vision, and computer graphics
    • Success metrics: widespread deployment/impact; long journey with expanding horizons

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.