Skip to content
a16za16z

Why World Models Could Change Robotics, 3D, and Creativity

World Labs co-founders Fei-Fei Li, Justin Johnson, and Ben Mildenhall join a16z General Partner Martin Casado to discuss Atlas, their latest world model, and what it reveals about the pursuit of spatial intelligence. At the center of Atlas is what the team calls “new view prediction”: given images or views of a scene, the model predicts what that environment should look like from a different position in space and time. This brings generation and 3D reconstruction into the same model, and raises a broader question about whether predicting views could become a useful primitive for understanding the physical world. They discuss the technical bets behind the model, what it can and can’t yet capture, and the importance of dynamics, editability, and simulation as world models develop. The conversation also explores applications in creative work, architecture, and robotics, where Fei-Fei argues that one of today’s biggest constraints is access to real-world training data. Timestamps: 00:00 - Intro 00:51 - What Atlas Is & Why It Matters 05:15 - Is This a Scaled-Up Video Model or a New Architecture? 08:15 - Spatial Intelligence & Why New View Prediction Matters 21:27 - Did You Know It Was Going to Work? 24:42 - Use Cases: Creatives, Games & Robotics 35:21 - The Elephant in the Room: Video Models vs World Models 37:55 - Will We Get 4D Video You Can Walk Around In? 42:22 - Why New View Prediction Is the Next Token Prediction Resources: Follow Fei-Fei Li on X: https://x.com/drfeifei Follow Justin Johnson on X: https://x.com/jcjohnss Follow Ben Mildenhall on X: https://x.com/BenMildenhall Follow Martin Casado on X: https://x.com/martin_casado Learn more about Atlas: https://www.worldlabs.ai/blog/atlas Stay Updated: If you enjoyed this episode, be sure to like, subscribe, and share with your friends! Find a16z on X: https://twitter.com/a16z Find a16z on LinkedIn: https://www.linkedin.com/company/a16z Listen to the a16z Show on Spotify: https://open.spotify.com/show/5bC65RDvs3oxnLyqqvkUYX Listen to the a16z Show on Apple Podcasts: https://podcasts.apple.com/us/podcast/a16z-podcast/id842818711 Follow our host: https://x.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see http://a16z.com/disclosures.

Fei-Fei LiguestJustin JohnsonguestMartin Casadohost
Sep 4, 202643mWatch on YouTube ↗

CHAPTERS

  1. 0:00 – 0:51

    Atlas’s big leap: spatially grounded pixel generation

    The conversation opens by framing Atlas as a difficult but crucial step toward “spatial intelligence”: not just generating pretty pixels, but generating pixels that remain consistent with 3D geometry and viewpoint. The hosts tease the practical impact—major reductions in capture requirements—and set up Atlas as a new primitive beyond typical video generation.

    • Spatial intelligence requires pixels that are grounded in space, not just visually plausible
    • Atlas is positioned as a hard-won step change vs. “many video models”
    • Promises large workflow efficiency gains (orders-of-magnitude reductions)
    • Teaser example: Matrix-style “bullet time” with far fewer cameras
  2. 0:51 – 1:49

    What Atlas is: a world model for generation, reconstruction, and simulation

    Justin Johnson defines Atlas as a next-generation world model with three pillars: generating scenes, reconstructing 3D from sparse views, and simulating outcomes. The chapter clarifies inputs/outputs and highlights camera-conditioned generation as a central capability.

    • Atlas supports generation, reconstruction, and simulation in one model
    • Camera-conditioned generation: steer output with an image + camera trajectory
    • Sparse 3D reconstruction from as few as 1–100 frames
    • Outputs can be novel-view video fly-throughs or explicit 3D reconstructions
  3. 1:49 – 2:49

    Bullet time demo: from hundreds of cameras to three phones

    The team explains “bullet time” (Matrix-style camera orbit around frozen action) and why Atlas makes it dramatically easier. They describe capturing an event with just a few iPhones and then synthesizing new camera paths through time-frozen moments.

    • Classic bullet time required a ring of hundreds of calibrated cameras
    • Atlas can recreate similar shots with as few as three cameras
    • No studio capture, green screen, or expensive calibration required
    • Reframing and virtual camera fly-ins enable dramatic creative shots
  4. 2:49 – 3:34

    The core primitive: new view prediction (not next-frame prediction)

    Atlas’s simplest description is introduced: it performs new view prediction by building a “spatial context” from posed inputs and rendering what a virtual camera would see at any point in space and time. This reframes world modeling as viewpoint-grounded generation rather than sequential frame extrapolation.

    • LLMs: next token; video models: next frame; Atlas: new view prediction
    • Inputs become a spatial context representing the scene
    • A virtual camera can be placed anywhere in space/time to render consistent views
    • Emphasis on geometry-aware outputs rather than purely temporal continuation
  5. 3:34 – 5:23

    How Atlas differs from typical video models: pose-grounded precision over “slot machine” prompting

    The discussion contrasts Atlas with common video models that rely heavily on ambiguous conditioning and repeated retries. Atlas ties each input frame to an explicit 3D camera pose, enabling controlled reconstruction and directed fly-throughs with high precision and less guesswork.

    • Many models accept images/text but lack per-frame spatial grounding
    • Atlas associates each image with a 3D camera pose for consistent reconstruction
    • Enables exact replication of known areas vs. guessing relationships
    • Supports intentional staging and camera choreography for creative pipelines
  6. 5:23 – 8:15

    New architecture: unifying generation + reconstruction with native multimodality

    Justin and Fei-Fei argue Atlas is architecturally new because it jointly solves tasks that historically lived in separate CV subfields. The model is multimodal from pretraining—handling text, images, video, camera poses, and depth—so reconstruction and generation can reinforce each other.

    • First unification of pixel generation and pixel reconstruction in one model
    • Historically separate tracks: generation vs. 3D reconstruction vs. recognition
    • Native inputs include camera pose and depth maps (3D as a modality)
    • Viewpoint anchoring is presented as the key to the unification
  7. 8:15 – 10:54

    Spatial intelligence roadmap: geometry first, then dynamics and interaction

    Fei-Fei outlines spatial intelligence as the ability to generate, reason, edit, and interact in 3D/4D space. Atlas is framed as progress on geometry and viewpoint estimation, with time/dynamics and higher-fidelity simulation as the next ladder rungs.

    • Spatial intelligence includes generation, reasoning, editing, and interaction
    • 4D = 3D plus time; dynamics are essential long-term
    • Camera pose is described as critical geometric information
    • Atlas enables emergent downstream behaviors by grounding pixels in space
  8. 10:54 – 13:16

    Why Atlas now: from Marble’s Gaussian splat bottleneck to a scalable formulation

    The team explains why they didn’t jump straight to Atlas: compute, iteration, and choosing the right representation. Marble focused on outputting Gaussian splats; Atlas removes that bottleneck by centering on new view prediction and only producing splats when useful.

    • Compute constraints required incremental “scaling ladder” progress
    • Marble: strong product, but bottlenecked by Gaussian splat outputs
    • Atlas: bifurcates modalities earlier and unifies them in-model
    • New view prediction becomes the fundamental primitive; splats are optional outputs
  9. 13:16 – 17:30

    Dense vs. sparse reconstruction: collapsing capture effort by 50–100×

    Ben explains why traditional dense reconstruction is laborious: you need many views of every surface to avoid holes. Atlas aims to make sparse capture viable—using far fewer images—unlocking reconstruction from casual footage and older captures that previously failed.

    • Dense reconstruction requires exhaustive coverage (100–300+ photos per room)
    • Sparse capture targets “three-ish” views, changing feasibility dramatically
    • Enables using existing imagery, internet photos, and casual videos
    • Practical impact: old incomplete captures can reconstruct convincingly with Atlas
  10. 17:30 – 21:27

    Why generation must complement reconstruction: filling holes and scaling context

    They argue that even expert dense capture misses occlusions, so generative ability is necessary to fill gaps. The team connects this to an LLM-like “context length” analogy: reconstruction becomes generation with a long, structured context of posed views.

    • Classic triangulation fails where pixels were never observed (holes)
    • Generative modeling fills inevitable occlusions and missing regions
    • Analogy: reconstruction is generation with a long context window
    • Atlas supports many inputs (e.g., dozens of images) to improve grounded outputs
  11. 21:27 – 24:41

    Conviction, scaling laws, and the moment it ‘clicked’ (the under-the-table flythrough)

    The hosts probe whether the team knew Atlas would work; they describe confidence in scaling and viewpoint prediction, but uncertainty about details and speed. A breakthrough demo—flying the camera under a classic NeRF table scene—cemented commitment to the approach.

    • Team had conviction in scaling laws + viewpoint prediction, not exact details
    • Quality improved consistently with larger models and longer training
    • Compute, not ideas, is portrayed as the main current limiter
    • A notable internal milestone: under-the-table camera flythrough impressed everyone
  12. 24:41 – 29:10

    Use cases in creative pipelines: persistent 3D state, controllability, and faster design iteration

    Ben connects Atlas to real creative workflows: creators want stable, persistent 3D scenes and controllable viewpoints, not one-off generations. He extends the value proposition to architecture, construction, fabrication, and any domain where 3D revisions are slow and costly today.

    • Creators used Marble mainly to get consistent multi-view images
    • Atlas directly serves view synthesis without degrading through intermediate renders
    • Persistent 3D “state” matches how artists/designers build worlds over time
    • Large opportunity: reduce manual 3D revision cycles in industrial/design workflows
  13. 29:10 – 35:22

    Robotics: real-to-sim pipelines, data scarcity, and simulation as the bridge to policy learning

    Fei-Fei explains how Atlas fits robotics via accelerating real-to-sim environment creation, historically blocked by dense reconstruction effort. They emphasize robotics’ core bottleneck is data and variation (domain randomization), making scalable simulation and learned world models strategically important.

    • Acquired robotics tech focuses on real-to-sim and sim-to-real loops
    • Dense reconstruction slows sim creation; Atlas can speed environment capture
    • Robotics bottleneck today: data collection + necessary randomization
    • Learned simulators could train policies and eventually inform planning
  14. 35:22 – 43:42

    Dynamics, 4D worlds, and editability: where Atlas goes next

    The group addresses the “elephant in the room”: richer dynamics are needed for true world modeling, especially for robotics. They argue Atlas already contains “baby dynamics,” and the next frontier is stronger multimodal control/editability without sacrificing output quality—culminating in the claim that new view prediction could be an AI-complete primitive.

    • Atlas supports dynamics in architecture/data, even if release emphasized static reconstruction
    • Key idea: expose the model to dynamics to help it factor motion out when needed
    • Future focus: editability and multimodal control (layout, identity, time) at high quality
    • Thesis: generative new view prediction may be ‘AI complete’ like next token prediction

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.