Skip to content
a16za16z

Why World Models Could Change Robotics, 3D, and Creativity

World Labs co-founders Fei-Fei Li, Justin Johnson, and Ben Mildenhall join a16z General Partner Martin Casado to discuss Atlas, their latest world model, and what it reveals about the pursuit of spatial intelligence. At the center of Atlas is what the team calls “new view prediction”: given images or views of a scene, the model predicts what that environment should look like from a different position in space and time. This brings generation and 3D reconstruction into the same model, and raises a broader question about whether predicting views could become a useful primitive for understanding the physical world. They discuss the technical bets behind the model, what it can and can’t yet capture, and the importance of dynamics, editability, and simulation as world models develop. The conversation also explores applications in creative work, architecture, and robotics, where Fei-Fei argues that one of today’s biggest constraints is access to real-world training data. Timestamps: 00:00 - Intro 00:51 - What Atlas Is & Why It Matters 05:15 - Is This a Scaled-Up Video Model or a New Architecture? 08:15 - Spatial Intelligence & Why New View Prediction Matters 21:27 - Did You Know It Was Going to Work? 24:42 - Use Cases: Creatives, Games & Robotics 35:21 - The Elephant in the Room: Video Models vs World Models 37:55 - Will We Get 4D Video You Can Walk Around In? 42:22 - Why New View Prediction Is the Next Token Prediction Resources: Follow Fei-Fei Li on X: https://x.com/drfeifei Follow Justin Johnson on X: https://x.com/jcjohnss Follow Ben Mildenhall on X: https://x.com/BenMildenhall Follow Martin Casado on X: https://x.com/martin_casado Learn more about Atlas: https://www.worldlabs.ai/blog/atlas Stay Updated: If you enjoyed this episode, be sure to like, subscribe, and share with your friends! Find a16z on X: https://twitter.com/a16z Find a16z on LinkedIn: https://www.linkedin.com/company/a16z Listen to the a16z Show on Spotify: https://open.spotify.com/show/5bC65RDvs3oxnLyqqvkUYX Listen to the a16z Show on Apple Podcasts: https://podcasts.apple.com/us/podcast/a16z-podcast/id842818711 Follow our host: https://x.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see http://a16z.com/disclosures.

Fei-Fei LiguestJustin JohnsonguestMartin Casadohost
Sep 4, 202643mWatch on YouTube ↗

Episode Details

EPISODE INFO

Released
September 4, 2026
Duration
43m
Channel
a16z
Watch on YouTube
▶ Open ↗

EPISODE DESCRIPTION

World Labs co-founders Fei-Fei Li, Justin Johnson, and Ben Mildenhall join a16z General Partner Martin Casado to discuss Atlas, their latest world model, and what it reveals about the pursuit of spatial intelligence. At the center of Atlas is what the team calls “new view prediction”: given images or views of a scene, the model predicts what that environment should look like from a different position in space and time. This brings generation and 3D reconstruction into the same model, and raises a broader question about whether predicting views could become a useful primitive for understanding the physical world. They discuss the technical bets behind the model, what it can and can’t yet capture, and the importance of dynamics, editability, and simulation as world models develop. The conversation also explores applications in creative work, architecture, and robotics, where Fei-Fei argues that one of today’s biggest constraints is access to real-world training data. Timestamps: 00:00 - Intro 00:51 - What Atlas Is & Why It Matters 05:15 - Is This a Scaled-Up Video Model or a New Architecture? 08:15 - Spatial Intelligence & Why New View Prediction Matters 21:27 - Did You Know It Was Going to Work? 24:42 - Use Cases: Creatives, Games & Robotics 35:21 - The Elephant in the Room: Video Models vs World Models 37:55 - Will We Get 4D Video You Can Walk Around In? 42:22 - Why New View Prediction Is the Next Token Prediction Resources: Follow Fei-Fei Li on X: https://x.com/drfeifei Follow Justin Johnson on X: https://x.com/jcjohnss Follow Ben Mildenhall on X: https://x.com/BenMildenhall Follow Martin Casado on X: https://x.com/martin_casado Learn more about Atlas: https://www.worldlabs.ai/blog/atlas Stay Updated: If you enjoyed this episode, be sure to like, subscribe, and share with your friends! Find a16z on X: https://twitter.com/a16z Find a16z on LinkedIn: https://www.linkedin.com/company/a16z Listen to the a16z Show on Spotify: https://open.spotify.com/show/5bC65RDvs3oxnLyqqvkUYX Listen to the a16z Show on Apple Podcasts: https://podcasts.apple.com/us/podcast/a16z-podcast/id842818711 Follow our host: https://x.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see http://a16z.com/disclosures.

SPEAKERS

  • Fei-Fei Li

    guest

    Stanford computer vision professor and AI leader focused on “spatial intelligence” and embodied AI.

  • Justin Johnson

    guest

    Computer vision researcher describing the technical design and roadmap of the Atlas world model.

  • Martin Casado

    host

    General Partner at Andreessen Horowitz (a16z) who moderates and interviews guests on AI and infrastructure.

EPISODE SUMMARY

In this episode of a16z, featuring Fei-Fei Li and Justin Johnson, Why World Models Could Change Robotics, 3D, and Creativity explores atlas reframes world models as controllable, pose-grounded new-view prediction Atlas is presented as a next-generation world model whose central capability is controllable “new view prediction,” letting users render a scene from arbitrary camera trajectories using a spatially grounded context.

RELATED EPISODES

Why AI Agents Could Finally Reinvent the Credit Card

Why AI Agents Could Finally Reinvent the Credit Card

What Today’s Best Models Still Can’t Do in Math

What Today’s Best Models Still Can’t Do in Math

How Whatnot's Live Shopping Beats Traditional E-Commerce

How Whatnot's Live Shopping Beats Traditional E-Commerce

Tokens Are the New Dollars | Stripe's Will Gaybrick & David George

Tokens Are the New Dollars | Stripe's Will Gaybrick & David George

Why Attention is Becoming a Competitive Advantage | a16z

Why Attention is Becoming a Competitive Advantage | a16z

The Media Game Has Changed

The Media Game Has Changed

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.