a16zSolving the Hardest Problem in Robotics | Fei-Fei Li with a16z
At a glance
WHAT IT’S REALLY ABOUT
Fei-Fei Li explains spatial intelligence powering real-to-sim robotics foundation models
- World Labs frames “spatial intelligence” as the next AI frontier: models that understand, generate, and reason about 3D spaces and interactions across physical and virtual worlds.
- SceniX tackles robotics’ core bottleneck—scarce, slow, costly real-world data—by mapping real environments into aligned digital twins for scalable training and evaluation (real-to-sim-to-real).
- Their synergy centers on Marble, World Labs’ base model that generates geometrically consistent 3D worlds from prompts, enabling more efficient environment reconstruction and simulation infrastructure.
- The discussion argues simulation is essential (not optional) because it enables counterfactual reasoning, systematic coverage of edge cases via randomization, and faster iteration than real-world robotics testing.
- They outline a pragmatic go-to-market: serve near-deployment customers in semi-structured environments (warehouses, assembly, hospitality) and measure success via lighthouse customers and proven automation value within two years.
IDEAS WORTH REMEMBERING
5 ideasRobotics needs “world understanding + action,” not just perception.
World Labs positions spatial intelligence as the capability to generate and reason about 3D space, while a robotics foundation model likely must also model actions—both as inputs (forward simulation) and outputs (policy).
Marble’s value is geometrically consistent 3D worlds from sparse input.
Marble converts text/images into consistent 3D representations (e.g., Gaussian splats/meshes), which SceniX can leverage to make environment reconstruction and simulation setup more efficient and scalable.
Simulation is a scaling unlock because real-world robotics data is fundamentally constrained.
Unlike internet-scale language data, robotics requires moving atoms under physics, making collection slow, dangerous, and expensive; simulation provides a lever to approximate scaling laws via synthetic generation and faster iteration.
Consistency across time, viewpoints, and interactions is where video-only approaches often fail.
Yunzhu contrasts their approach with video prediction models that can violate object permanence or interaction logic; a robot learning world must remain stable under actions to produce trustworthy training signals.
The right goal is not perfect fidelity, but “essential structure” plus systematic randomization.
They argue effective sim-to-real can work without modeling every bush or surface precisely, as long as the simulator captures key dynamics and uses controlled variation (lighting, friction, geometry) to build robustness.
WORDS WORTH SAVING
5 quotesWe are building the next frontier of AI, which is what we call spatial intelligence.
— Fei-Fei Li
Think about human intelligence. We do a lot of simulation in our head. You know why? There's a very important role simulation plays that real world data doesn't play, which is counterfactual reasoning.
— Fei-Fei Li
What we are building is a consistent world. Consistent both over space, over time, over different viewpoints, and over different type of interactions. My North Star is I want the robot to work.
— Yunzhu Li
I want to add to this and be s- uh, slightly philosophical here, is there isn't a, a, uh, binary choice between simulation or no simulation. All this come, um, in together, um, to, to make robotics work.
— Fei-Fei Li
Martin, the hardest thing in today's AI is to have the right measured optimism.
— Fei-Fei Li
High quality AI-generated summary created from speaker-labeled transcript.