Skip to content
a16za16z

Why World Models Could Change Robotics, 3D, and Creativity

World Labs co-founders Fei-Fei Li, Justin Johnson, and Ben Mildenhall join a16z General Partner Martin Casado to discuss Atlas, their latest world model, and what it reveals about the pursuit of spatial intelligence. At the center of Atlas is what the team calls “new view prediction”: given images or views of a scene, the model predicts what that environment should look like from a different position in space and time. This brings generation and 3D reconstruction into the same model, and raises a broader question about whether predicting views could become a useful primitive for understanding the physical world. They discuss the technical bets behind the model, what it can and can’t yet capture, and the importance of dynamics, editability, and simulation as world models develop. The conversation also explores applications in creative work, architecture, and robotics, where Fei-Fei argues that one of today’s biggest constraints is access to real-world training data. Timestamps: 00:00 - Intro 00:51 - What Atlas Is & Why It Matters 05:15 - Is This a Scaled-Up Video Model or a New Architecture? 08:15 - Spatial Intelligence & Why New View Prediction Matters 21:27 - Did You Know It Was Going to Work? 24:42 - Use Cases: Creatives, Games & Robotics 35:21 - The Elephant in the Room: Video Models vs World Models 37:55 - Will We Get 4D Video You Can Walk Around In? 42:22 - Why New View Prediction Is the Next Token Prediction Resources: Follow Fei-Fei Li on X: https://x.com/drfeifei Follow Justin Johnson on X: https://x.com/jcjohnss Follow Ben Mildenhall on X: https://x.com/BenMildenhall Follow Martin Casado on X: https://x.com/martin_casado Learn more about Atlas: https://www.worldlabs.ai/blog/atlas Stay Updated: If you enjoyed this episode, be sure to like, subscribe, and share with your friends! Find a16z on X: https://twitter.com/a16z Find a16z on LinkedIn: https://www.linkedin.com/company/a16z Listen to the a16z Show on Spotify: https://open.spotify.com/show/5bC65RDvs3oxnLyqqvkUYX Listen to the a16z Show on Apple Podcasts: https://podcasts.apple.com/us/podcast/a16z-podcast/id842818711 Follow our host: https://x.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see http://a16z.com/disclosures.

Fei-Fei LiguestJustin JohnsonguestMartin Casadohost
Sep 4, 202643mWatch on YouTube ↗

At a glance

WHAT IT’S REALLY ABOUT

Atlas reframes world models as controllable, pose-grounded new-view prediction

  1. Atlas is presented as a next-generation world model whose central capability is controllable “new view prediction,” letting users render a scene from arbitrary camera trajectories using a spatially grounded context.
  2. The model’s novelty is the architectural unification of precise 3D reconstruction and generative completion, with camera pose and depth treated as first-class modalities during pretraining.
  3. Compared with prior video and 3D systems, Atlas aims to replace dense, tedious capture with sparse inputs (often just a few views), enabling consistent fly-throughs and reframing without repeated prompt retries.
  4. The conversation highlights practical creative workflows (film, games, design, architecture) where persistent 3D state and editability matter more than one-off generations, citing large reductions in production friction.
  5. For robotics, Atlas is positioned as a key building block for accelerating real-to-sim pipelines today and enabling learned simulation and planning tomorrow, with dynamics and richer controls as near-term roadmap items.

IDEAS WORTH REMEMBERING

5 ideas

Atlas’s core primitive is “new view prediction,” not next-frame video generation.

Unlike next-token (LLMs) or next-frame (video diffusion) objectives, Atlas conditions on a spatial context built from posed images and can render what the scene looks like from any queried camera position (and time). This reframes “world modeling” around geometry-anchored controllability rather than prompt-driven guesswork.

Atlas unifies pixel reconstruction and pixel generation in one architecture.

The team emphasizes that reconstruction (precise, geometry-consistent replication) and generation (imagination to fill gaps) historically lived in separate CV subfields and models. Atlas merges them so it can both faithfully reproduce observed structure and plausibly complete unobserved regions.

Camera pose as a native input is the key to precision and consistency.

Each input frame is paired with an explicit 3D camera pose (and depth as a 3D signal), making every pixel spatially grounded. This enables precise multi-view consistency and “directed” fly-throughs without the slot-machine rerolling common in text-only control.

Atlas targets a step-change reduction in capture effort for usable 3D.

Traditional dense reconstruction demands exhaustive capture (hundreds/thousands of images) to avoid holes; Atlas aims to drop that to a handful of views by combining triangulation with learned priors that fill missing regions. The discussion cites ~50–100× reductions that can make 3D capture practical for everyday workflows.

Atlas turns sparse real-world capture into studio-grade reframing (e.g., bullet time) with minimal hardware.

The “bullet time” demo illustrates high-value reframing: where films once needed rings of hundreds of calibrated cameras, Atlas can synthesize similar shots from as few as three iPhones on tripods. This showcases both reconstruction fidelity and controllable virtual cinematography.

WORDS WORTH SAVING

5 quotes

On the path to spatial intelligence, generating pixels that are truly spatially contextualized and grounded is absolutely another major step, and that is the very hard step that Atlas has taken.

Fei-Fei Li

We know LLMs are built on next token prediction. We've seen video models as being built on next frame prediction. Atlas is really new view prediction.

Justin Johnson

But now with Atlas, we can do this with just as few as through three cameras. So, like, no studio capture, no green screen, no expensive calibration.

Justin Johnson

That morning, the three of us looked at each other in the eyes and say, "That's it, this is-- We're gonna build this."

Fei-Fei Li

We do believe very strongly that next viewpoint prediction is, is the equivalent of next token prediction.

Fei-Fei Li

Atlas world model capabilities: generate, reconstruct, simulateNew view prediction vs next-frame video modelsCamera pose conditioning and spatial contextUnifying reconstruction and generationSparse vs dense 3D capture; context scalingBullet time reframing from few camerasRobotics real-to-sim, simulation, and data bottlenecks

High quality AI-generated summary created from speaker-labeled transcript.

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.