a16z“The Future of AI is Here” — Fei-Fei Li Unveils the Next Frontier of AI
CHAPTERS
- 0:00 – 0:29
Why spatial intelligence is the next major AI bet
Fei-Fei Li frames visual-spatial intelligence as a peer to language—possibly even more foundational for embodied agents and real-world interaction. She argues the timing is right due to a convergence of compute, better data understanding, and algorithmic advances, motivating World Labs’ focus.
- •Spatial/visual intelligence is positioned as a core pillar of intelligence alongside language
- •A ‘right moment’ thesis: compute + data understanding + algorithms have converged
- •North Star framing: focus the field/company around a fundamental capability
- •Sets up World Labs as the vehicle to pursue spatial intelligence
- 0:29 – 1:48
How we got here: AI winters, deep learning takeoff, and multimodal expansion
Martin Casado asks for a historical arc of modern AI. Fei-Fei describes moving from AI winter to the birth of modern AI, deep learning breakthroughs, and today’s “Cambrian explosion” across text, pixels, video, and audio.
- •From AI winter to deep learning’s breakout successes
- •Industry adoption accelerates first in language models
- •AI expands beyond text into vision, video, and audio
- •Framing today as a rapid proliferation of new model/application possibilities
- 1:48 – 3:43
Justin Johnson’s path: the deep learning ‘recipe’ and the research boom
Justin recounts getting hooked on deep learning after early landmark work (the ‘Cat paper’) and learning the now-classic formula: algorithms + lots of compute + lots of data. He describes the intense pace of discovery during his PhD years as core generative and vision building blocks emerged from academia.
- •Deep learning as a general recipe that scales with data and compute
- •Computer vision as an early proving ground for deep nets
- •Early generative modeling foundations developed during the PhD-era paper boom
- •Cultural shift: the broader world recently discovered what researchers felt for years
- 3:43 – 7:25
Fei-Fei Li’s path: physics to AI, and the overlooked power of data
Fei-Fei explains her route into AI via physics and computational neuroscience during a period when AI looked stalled publicly but was active scientifically. She highlights a key insight from her lab: scaling data, not just model cleverness, was crucial for generalization—leading to ImageNet and internet-scale vision datasets.
- •Physics training encouraged ‘audacious questions’ like intelligence
- •Machine learning era preceded deep learning, experimenting with many model families
- •Data was an underappreciated driver of progress and generalization
- •ImageNet as a deliberate bet to push vision datasets to internet scale
- 7:25 – 9:18
Compute as the underrated unlock: AlexNet to modern GPUs
Justin argues compute is the biggest unlock and still underestimated. He contrasts AlexNet’s 2012 training run (days on two consumer GPUs) with modern NVIDIA hardware, where an equivalent run could complete in minutes—illustrating how rapidly the feasible model/search space has expanded.
- •Compute growth over the last decade is ‘in the thousands’ in raw factor terms
- •AlexNet: ~60M params, trained for days on two GTX 580s
- •Modern hardware (e.g., GB200) collapses training time dramatically
- •Hardware scaling reshaped what research and products are feasible
- 9:18 – 12:10
Data vs. compute—and the shift from supervised to self-supervised learning
The conversation weighs two narratives: compute scaling vs. new data sources. Justin distinguishes the ImageNet era of supervised learning (human-labeled ontologies) from later epochs where models learn from less explicitly labeled data and broader internet structure, enabling more general capabilities.
- •‘Bitter lesson’ perspective: prioritize compute-friendly approaches
- •ImageNet era: heavy human labeling, constrained tasks and ontologies
- •Later era: learning from data without explicit per-example labels
- •Language data carries implicit human structure differently than pixels
- 12:10 – 16:58
Generative AI as a continuum: from image-text matching to text-to-image
Fei-Fei and Justin map generative AI’s rise as a gradual progression rather than a sudden break. They trace steps through image-text alignment, captioning (pixels to words), style transfer, and early structured text-to-image generation using scene graphs and GANs.
- •Generative modeling existed theoretically earlier, but results weren’t compelling
- •Progression: alignment → captioning → style transfer → structured generation
- •2015 neural style transfer as a visceral ‘gen AI’ moment for researchers
- •Scene graphs as an early bridge from language-like structure to image synthesis
- 16:58 – 19:58
From research North Stars to World Labs: why focus on spatial intelligence now
Martin transitions from Fei-Fei’s research journey to the founding of World Labs. Fei-Fei describes spatial intelligence as her next North Star after major milestones in visual storytelling, arguing today’s compute, data maturity, and new algorithms (including NeRF-related advances) make the bet timely.
- •North Stars guide long-term research direction and field advancement
- •Visual-spatial reasoning is essential for seeing, interacting, and building in the world
- •World Labs’ mission: unlock spatial intelligence as the next frontier
- •Key enablers: compute, improved data practices, and cutting-edge 3D methods (e.g., NeRF)
- 19:58 – 21:20
Defining spatial intelligence: perceiving, reasoning, and acting in 3D/4D
Justin offers a crisp definition: machines understanding space and time—objects, events, and their interactions across 3D plus dynamics. They clarify spatial intelligence applies to both the physical world and generated/simulated worlds.
- •Spatial intelligence = perception + reasoning + action in 3D and time
- •4D framing: positions and interactions evolve over space-time
- •Applies to real environments and synthetic/generated environments
- •Goal is to move AI from data centers into rich world contexts
- 21:20 – 26:33
Why now: NeRF and the merge of reconstruction with generation
Justin explains the pivot toward 3D understanding from 2D observations and highlights NeRF as a breakthrough that catalyzed the field. Fei-Fei adds that computer vision’s long tradition in 3D reconstruction is now converging with generative modeling, making the reconstruction-vs-generation distinction increasingly blurred.
- •3D data is hard to collect; leverage 2D projections to infer 3D structure
- •NeRF as a simple, powerful method to back out 3D from 2D views
- •Academia could still innovate here even as LLM compute needs outpaced labs
- •Reconstruction (seeing) and generation (imagining) are rapidly converging in vision
- 26:33 – 29:42
Spatial intelligence vs. LLMs: 1D token sequences versus 3D-native representations
Martin probes how spatial approaches differ from multimodal LLMs that also ‘see.’ Justin and Fei-Fei argue LLMs are fundamentally built on 1D sequence representations, whereas spatial intelligence puts 3D structure at the core—better matching physical reality and enabling richer interactions.
- •LLMs/multimodal models operate on 1D token sequences under the hood
- •Other modalities often get ‘shoehorned’ into sequence form
- •Spatial intelligence treats 3D structure as first-class in representation
- •Physical world constraints (geometry/physics) make 3D understanding qualitatively different from language modeling
- 29:42 – 32:36
2D video vs. 3D worlds: affordances, interaction, and an ‘arc of intelligence’
They unpack why 3D representations matter even if humans ultimately view 2D renders. A 3D-native model better supports user affordances like moving cameras/objects and interacting with environments, aligning with intelligence as the capacity to navigate and manipulate the world.
- •Separate ‘representation’ from ‘user-facing output’ (often still 2D)
- •2D perception can imply 3D, but explicit 3D representations fit tasks better
- •Affordances: camera motion, object manipulation, consistent geometry
- •Spatial intelligence as a prerequisite for broad real-world and creative applications
- 32:36 – 37:42
Use cases: world generation, new media, and dynamic interactive environments
Justin outlines an evolution from generating images/clips to generating full interactive 3D worlds. They position this as a new media layer that could dramatically reduce the cost and labor of producing AAA-quality virtual experiences, unlocking many non-gaming applications.
- •From text-to-image/video to generating full 3D worlds
- •A ‘new media’ thesis: interactive experiences become far cheaper to create
- •Economic unlock: beyond $70 AAA games into niche/personalized worlds
- •Progression expected: static scenes first, then dynamic physics and interaction
- 37:42 – 40:42
AR/VR and robotics: spatial intelligence as the operating system for mixed reality
Fei-Fei connects spatial intelligence to spatial computing (e.g., Vision Pro) and the blending of real and virtual worlds. They extend the case to robotics, where a robot’s digital ‘brain’ must act in a 3D physical environment—making spatial intelligence the key bridge.
- •AR/MR needs real-time 3D understanding to blend digital content with reality
- •Hardware form factors may evolve (goggles, glasses, contact lenses)
- •AR could collapse the need for many separate screens via contextual overlays
- •Robotics: spatial intelligence connects digital compute to physical-world action
- 40:42 – 48:09
Company strategy, team-building, and what success looks like for World Labs
Martin asks how World Labs balances deep tech with multiple application areas. Fei-Fei positions the company as a platform model provider, then both discuss the multidisciplinary talent needed across ML, data, systems, and graphics; they close with milestones defined by real-world deployment and expanding possibilities.
- •World Labs as a deep tech platform company enabling many downstream use cases
- •Near-term pragmatism: some markets/devices aren’t ready for mass adoption
- •Team composition spans ML, infra, data, 3D vision, and computer graphics
- •Success metrics: widespread deployment/impact; long journey with expanding horizons