Skip to content
a16za16z

Why World Models Could Change Robotics, 3D, and Creativity

World Labs co-founders Fei-Fei Li, Justin Johnson, and Ben Mildenhall join a16z General Partner Martin Casado to discuss Atlas, their latest world model, and what it reveals about the pursuit of spatial intelligence. At the center of Atlas is what the team calls “new view prediction”: given images or views of a scene, the model predicts what that environment should look like from a different position in space and time. This brings generation and 3D reconstruction into the same model, and raises a broader question about whether predicting views could become a useful primitive for understanding the physical world. They discuss the technical bets behind the model, what it can and can’t yet capture, and the importance of dynamics, editability, and simulation as world models develop. The conversation also explores applications in creative work, architecture, and robotics, where Fei-Fei argues that one of today’s biggest constraints is access to real-world training data. Timestamps: 00:00 - Intro 00:51 - What Atlas Is & Why It Matters 05:15 - Is This a Scaled-Up Video Model or a New Architecture? 08:15 - Spatial Intelligence & Why New View Prediction Matters 21:27 - Did You Know It Was Going to Work? 24:42 - Use Cases: Creatives, Games & Robotics 35:21 - The Elephant in the Room: Video Models vs World Models 37:55 - Will We Get 4D Video You Can Walk Around In? 42:22 - Why New View Prediction Is the Next Token Prediction Resources: Follow Fei-Fei Li on X: https://x.com/drfeifei Follow Justin Johnson on X: https://x.com/jcjohnss Follow Ben Mildenhall on X: https://x.com/BenMildenhall Follow Martin Casado on X: https://x.com/martin_casado Learn more about Atlas: https://www.worldlabs.ai/blog/atlas Stay Updated: If you enjoyed this episode, be sure to like, subscribe, and share with your friends! Find a16z on X: https://twitter.com/a16z Find a16z on LinkedIn: https://www.linkedin.com/company/a16z Listen to the a16z Show on Spotify: https://open.spotify.com/show/5bC65RDvs3oxnLyqqvkUYX Listen to the a16z Show on Apple Podcasts: https://podcasts.apple.com/us/podcast/a16z-podcast/id842818711 Follow our host: https://x.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see http://a16z.com/disclosures.

Fei-Fei LiguestJustin JohnsonguestMartin Casadohost
Sep 4, 202643mWatch on YouTube ↗

EVERY SPOKEN WORD

  1. 0:000:51

    Intro

    1. FL

      On the path to spatial intelligence, generating pixels that are truly spatially contextualized and grounded, that is the very hard step that Atlas has taken.

    2. JJ

      We know LLMs are built on next token prediction. We've seen video models as being built on next frame prediction. Atlas is really new view prediction.

    3. SP

      This is the real place where AI can actually unlock a ton of value for people and their process. We're seeing like 50, 100 X reduction.

    4. JJ

      There was a famous shot in the first Matrix movie where Neo is, like, falling down.

    5. SP

      [laughs] Yeah.

    6. JJ

      Exactly. They had hundreds of cameras viewing that angle on a green screen. On Atlas, we can do this with just three cameras. No studio capture, no green screen, no expensive calibration.

    7. FL

      No one has ever seen this result.

    8. JJ

      When you set out to do this, did you know it was gonna work? I was pretty sure. Each time we made the model bigger, and each time we trained it for longer, it got significantly better.

    9. SP

      Does that mean we're gonna get 4D video? Can I go walk around?

    10. FL

      Uh-

  2. 0:515:15

    What Atlas Is & Why It Matters

    1. SP

      So big day yesterday, you launched, uh, a new frontier model, which has got an amazing reception, which is still coming in. I think maybe a good way to structure this conversation, let's just talk about exactly what that was, and then we'll go back to history and work our way back up. So maybe, Justin, do you wanna talk about what was launched yesterday, why it's significant?

    2. JJ

      Yeah. So Atlas is our new next gen- generation world model. Um, it has three basic things. It can, uh, generate, reconstruct, and simulate the world. Um, so within that, there's a couple different major capabilities. It has really good camera condition generation, so you can input an image together with a camera traject- with a camera trajectory and steer the model and have it generate, you know, video frames along any per- any perspective you want. It's really good at sparse 3D reconstruction. You can input m- one or mul- multiple, up to 100 frames, um, that are views of the real world and use those to reconstruct the real world, and that reconstruction can take the case either of, of a, of novel, of video flying through the space or an explicit 3D reconstruction of the space. Um, then finally, it can be used for simulation. Um, and for this, we show off, um, you know, these awesome bullet time videos, which got a lot of attention online, and then also robotics simulation. We can-

    3. SP

      What's a, what's a bullet time video?

    4. JJ

      A bullet time video, this comes from, uh, The Matrix. You know, there, there was a famous shot in the first Matrix movie where Neo is, like, falling down.

    5. SP

      Oh, that like... [laughs] Yes.

    6. JJ

      Exactly.

    7. SP

      Yeah.

    8. JJ

      So and then remember-

    9. SP

      Yeah

    10. JJ

      ... in that famous shot, he's, like, falling down, it's in slow motion, and the camera flies-

    11. SP

      Yeah

    12. JJ

      ... all the way around. Um, so that's... The, the, the way that they did that shot is they had a ring of, like, hundreds of cameras. So then, like, he fell over in the studio. They had hundreds of cameras viewing that angle on a green screen, and then they used those hundreds and hundreds of cameras to make that, that famous shot in, in The Matrix. But now with Atlas, we can do this with just as few as through three cameras. So, like, no studio capture, no green screen, no expensive calibration. We can literally stick, like, three cameras on tri- three iPhones on tripods-

    13. SP

      Mm-hmm

    14. JJ

      ... um, use these to take sort of v- a video of something happening, um, like someone shooting a basket, someone dropping a strawberry into a bowl of milk, and then from those three, like, iPhone videos, we can then reframe the shot and imagine, like, a fr- like freeze time, have the camera fly in, like, as the milk is splashing up and get these amazing frozen time views.

    15. SP

      Yeah.

    16. JJ

      Um, and we can do this with just, uh, just a couple cameras.

    17. SP

      Can, can you, um, just maybe, what is the simplest description of what Atlas does? Like, what goes in and what comes out?

    18. JJ

      Yeah. So one of the, one of the really core principles of Atlas, like, the most fundamental thing is it does new view prediction. Um, and this is, uh, a really fundamental primitive that we think is super exciting, uh, a super new primitive for, for base models that no one's ever done before. Right? So we know LLMs are built on next token prediction. We've seen video models as being built on next frame prediction. Atlas is really new view prediction, right? That given some number of views of a scene or a description of a scene, um, those go into what we call a spatial context that describes implicitly what is the world that we wanna talk about. Then you can point a virtual camera at, at an arbitrary point in space and time, and Atlas will understand what that world is supposed to look like from that position in space and time.

    19. SP

      You know, you know, Ben, that, you know, with, with a, with a bajillion video models out there all claiming to be world models and all claiming to have novel views and... Can you maybe tease apart kind of more concretely how this is different from, like, the myriad models that have come before? Mm-hmm. Yeah. I think what Justin was saying about the spatial context aspect is super important here. Um, so there's many video models, a lot of video models actually got their claim to fame from their single image input or their start to last frame interpolation. Now we're starting to see models that can do this kind of omni-reference thing with, you know, 20, 30, 50 images. But what's key with Atlas is that it actually has a kind of, like, spatially grounded meaning to every frame you put into it. So it's not just an image that the model's gonna interpret whatever way it wants, or you can kind of try to argue with it in the, the text prompting and get it to do something specific. With Atlas, every image actually has an associated three-dimensional camera pose, and that means that you can perform this task of reconstruction with an extremely high degree of precision, right? So if we had four views of this room, one at each corner, you can put those into the model and then get an exact replication of everything you see in this room, and it's not going to guess what's in the other corner, like the relationship between things. It's just gonna reproduce exactly what you give it. Um, and you can also do that in a kind of creative or imaginative sense too. If you take two photos from different, you know, AI generations or real world locations, you can actually position and stage those to build these kind of intentionally directed fly-throughs, um, that are really governed by exactly the precise place that you put the content you want and where the camera's gonna look a- and travel. Uh, which is very different, I think, than the kind of like more slot machine effect you get of having to retry generations over and over with just that kind of higher level of text control you get with video models.

  3. 5:158:15

    Is This a Scaled-Up Video Model or a New Architecture?

    1. SP

      Is this just kind of an obvious, you know, scaled-up version of a traditional video model, or is it a new architecture?

    2. JJ

      I think it's, it's a pretty new thing for a couple different reasons.

    3. SP

      Okay.

    4. JJ

      Um, one that we talk about is it does both generation and reconstruction jointly in the same model. Um, like Ben was saying, this thing can take a couple views of this room and then reconstruct everything in this room exactly as you see it. And historically, reconstruction has been its own subfield in computer vision with its own specialized task, its own specialized models. Um, and generation is what all, all the text, all the text-to-video models are really good at, like what all the, all the big diffusion models we've seen the last couple years. And those are great for creative applications. I wanna imagine something that's never been there before. Um, but now with Atlas, for the first time, we're putting these two different parts of visual intelligence together in one model. So it can do both 3D reconstruction and generation together in one architecture. So to do that, we had to make a couple changes. Um, one is we had to make it multimodal from the start. So this thing natively works on text, it works on images, it works on videos. It also works on camera poses, um, as a native input to the model, which I don't think anyone's ever done at the pre-training phase before.

    5. MC

      Yeah.

    6. JJ

      Um, and it uses, um, it uses 3D as a native modality that it works on. So this thing from the beginning was designed to be natively multimodal in a way that no one else I think-

    7. MC

      Sorry, I just have, um... I, I don't know the space super well. By 3D, is this like depth or-

    8. JJ

      Yeah, so the-

    9. MC

      ... models or like what does that n- what does that mean?

    10. JJ

      Yeah, so the formulation we used so far is depth maps.

    11. MC

      Okay.

    12. JJ

      Right? So, um, right now you can have a... When you have a frame that has a, a pos- a virtual camera telling its position in 3D space-

    13. MC

      Yeah

    14. JJ

      ... that camera position and, and camera parameters are a native input to the model. And then what, attached to that camera position, you can have both, um, like RGB telling you what does that, what does that position in space look like, and you can have a depth map that tells you what is the spatial structure of that position in 3D space. So then, you know, text, image, video, 3D cameras are these modalities that this thing all does jointly in, in a multimodal way.

    15. FL

      I wanna add something-

    16. MC

      Yeah, yeah, yeah

    17. FL

      ... 'cause I think what Justin just said is actually so important, and also what Ben said, that it's underappreciated. It's the first time we have a unification of pixel generation and pixel reconstruction. In the world of computer vision, this field has been around for more than half a century. Um, sitting here, having been in this field for decades, I cannot tell you how many PhD thesis have been written on the problem of reconstruction or novel view things, uh, synthesis. And also, our field traditionally, uh, have multiple tracks. You go to a computer vision conference, you have the pixel generation track, you have s-some recognition ch-track, and you have 3D reconstruction track. This is a elegant model that combines, uh, or unifies the problem of reconstruction and generation by anchoring on, um, viewpoints and, uh, the, the viewpoint, uh, estimation and th-that's just in-incredibly powerful.

  4. 8:1521:27

    Spatial Intelligence & Why New View Prediction Matters

    1. MC

      Can, can you, can you maybe... Well, can we take a step back and then maybe you just fill something out? So wh-when you started the company, I remember you saying, uh, you know, you wanna tackle, um, spatial intelligence, right? And, uh, you know, now we have this new model, and so it feel-- I mean, like as a layperson, it feels very general to me. You've got next token prediction, and this is next-

    2. JJ

      New view prediction.

    3. FL

      New view.

    4. MC

      New view prediction, right?

    5. FL

      Yeah.

    6. MC

      And so, like you can get one viewer, set of views, and you get a new view. Can maybe you pencil out like, like how, how this is a significant step to this general problem of spatial intelligence? May-maybe by like starting to describe what spatial intelligence is.

    7. FL

      Well, spatial intelligence e-eventually must enable us to both generate what the space is, reason within it-

    8. MC

      Yeah

    9. FL

      ... and being able to edit and interact within it.

    10. MC

      Yeah.

    11. FL

      Now, we can argue is it 3D or 4D? Ultimately, it's 4D with the time dimension, but even just 3D, these are the fundamental tasks that one has to do, or spatial intelligence has to enable. And then we talk about with that you can render, you can simulate, and you can plan actions. But to do that, a fundamental problem to solve is to understand the geometry and structure and the physics of the space.

    12. MC

      Yeah.

    13. FL

      And I do believe Atlas is a significant step forward because now with every single frame, you have a... You, you can generate and estimate an important piece of information, which is the, the view, viewpoint, the camera pose. And that is the most critical information one needs about the geometry of the, of the space, and that can lead to all the emergent, uh, behaviors we see in our, um, in, in the downstream of the model, which we showed in the blog. So, so in the, on the path to spatial intelligence, generating pixels is definitely a, a early step-

    14. MC

      Yeah

    15. FL

      ... which we have seen with what you call it, gazillions of, uh, models. But generating pixels that are truly spatially contextualized and grounded is absolutely another major step, and that is the very hard step that Atlas has taken.

    16. MC

      Yeah.

    17. FL

      I, we definitely have, you know, we can just keep going here, right? Like, there is the fourth dimension of time which will bring in dynamics, and there is more, a higher fidelity simulation and delineation of the space. So this is part of the, the roadmap of spatial intelligence.

    18. MC

      Great, yeah. I mean, and I definitely wanna like dig into like where this is going, but first maybe let's talk about getting here.

    19. FL

      Mm-hmm.

    20. MC

      Uh, how long has World Labs been in existence?

    21. FL

      Two, uh-

    22. JJ

      Two and a half, right?

    23. FL

      Two and a half, yeah.

    24. MC

      And so you, you've actually released models before, so why, why didn't you just jump right to Atlas? [laughing]

    25. FL

      Great question.

    26. MC

      It's so magic, right? Yeah.

    27. FL

      Justin's team needs a lot of chips.

    28. JJ

      Yeah, you need a lot of GPUs to actually scale this thing up. Um-

    29. MC

      Yeah

    30. JJ

      ... so one, one of the th-- Well, like last year we, we released our Marble world model, and that was the first kind of big major world model that we put out, um, that, that powers our current Marble product. And Mar- and Marble's really cool. Marble can take images, it can take videos, it can take text prompts, and use these to generate 3D worlds. Um, but one of the biggest differences between Marble and Atlas is exactly what is that output modality.

  5. 21:2724:42

    Did You Know It Was Going to Work?

    1. MC

      Did you-- By the way, but I have to ask, when you set out to do this, did you know it was gonna work?

    2. JJ

      I was pretty sure. [laughs]

    3. MC

      [laughs]

    4. SP

      That's, uh, so, so okay.

    5. MC

      [laughs] W-were you sure?

    6. SP

      I think three of us have total conviction about the scaling law.

    7. JJ

      Right.

    8. SP

      That, that I think we do. I do think the exact architecture choices and m- data mixtures is where the, the devils are in the details. I, you know, have watched Justin and his team going from, "We really don't know how long this is gonna take," to, "Oh, maybe sign of life," to, "Wow, this is gonna work." So it, it, no one, what, no one has done it. But I think the hypothesis, two hypothesis, one is scaling law hypothesis, another one is, uh, next, uh, viewpoint prediction. We had, uh, conviction of these two things-

    9. JJ

      Yeah

    10. SP

      ... from early on.

    11. JJ

      So, so I think I was very convicted that it was going to work. I was not sure it was gonna work this well at this fast, right? Like, I thought there's a chance that we do this, maybe, like, it's not clear that, like, the first cycle of pre-training a new model with a new architecture and a new paradigm, like-

    12. MC

      Yeah

    13. JJ

      ... the first cycle of that working is insane. So I thought there was a chance in which we had to, we might have had to do a couple more turns of that, of that pre-training cycle before we got to the level of quality we were ex- we, we wanted.

    14. MC

      Is there, is it, are, are we kind of like at the end of, like, the scaling for this architectural approach? We need another breakthrough or is there-

    15. SP

      No, no, we're at the beginning.

    16. MC

      Really?

    17. SP

      Yes.

    18. JJ

      Yeah.

    19. MC

      Without changing, without changing the architecture?

    20. SP

      Yeah, we're at the-

    21. JJ

      Yeah, I think we're basically at the beginning.

    22. SP

      Yep.

    23. JJ

      I think we're basically at the beginning and we're basically limited by compute at this point.

    24. MC

      Wow.

    25. JJ

      Right? Like, data is very important, as Fei-Fei likes to point out, but, like, everything has a bottleneck, and I think the main bottleneck on continuing to scale this thing is actually training compute, right? Like, during development, we trained a sequence of models. We wrote about this in the blog post a little bit, but we trained a couple models that, like, uh, the first couple rungs of the scaling ladder. Um, and each time we made the model bigger, and each time we trained it for longer, each time we put it on more chips, like, it got significantly better.

    26. MC

      Right.

    27. JJ

      And the model size that we, like, the model that we showed in the blog post is obviously the biggest and best one that we trained, but the thing that was limiting it was not the scale or the data or anything like that. It was literally like we had a deadline of when we wanted to release this thing, and therefore we backed up what we could afford to train and after that deadline. [laughs]

    28. SP

      But here's a, a little bit of a insider story, right? Like, Justin and team are, are training the, from the smaller and slightly bigger, you know, are train- uh, having these roadmaps. And then there was one day in summer, early summer, that it's not even this, the current Atlas, uh, model size, it's a smaller model, and then Ben, Justin, Ben feed it into, you know, the viewpoint generation. And remember that famous table, the garden table for NeRF paper-

    29. JJ

      Yeah

    30. SP

      ... and many papers that the

  6. 24:4235:21

    Use Cases: Creatives, Games & Robotics

    1. MC

      can you talk through maybe more specifically the use cases? So, so, uh, World Labs has historically had a lot of users that were creatives.

    2. SP

      Mm-hmm.

    3. MC

      And they use it for, like, consistency and, you know, like, whatever, 2D images and for movies and for 3D and for games, et cetera. And so maybe can you talk about how this extends use cases or cater to the existing ones, and then we'll-- I would like to talk about robotics, actually.

    4. SP

      Mm-hmm. Yeah, sure. Um, yeah, I mean, it, it's kind of funny actually. One of the kind of main ways we even saw people using Marble plays exactly into this new view prediction case. Like, a lot of our, um-

    5. MC

      Marble being the previous-

    6. SP

      Sorry, yeah, Marble, our previous product

    7. MC

      ... the, the previous model.

    8. SP

      Um, like, people would take that product, put an image in, get a full 3D scene as a Gaussian splat, take a couple screenshots of it from different points of view and leave.

    9. MC

      [chuckles]

    10. SP

      Right? And we're like, "We can just make those images, and that's a data, data control," right? So, so I think, like-- And, you know, there's a lot of degradation there. They're like, "Oh, this splat could look better." And it's like, "Okay, what if we just generatively model those viewpoints with that exact modality of control?"

    11. MC

      That's a good point.

    12. SP

      So I think, like, even that core capability of just, like, view synthesis, um, it's sort of been this academic problem for a long time, but in the sense of, oh, you're gonna do this really dense capture. Like, like, generative view synthesis is a relatively quite a new problem. And we just see so many people who, uh, i-in this creative pipeline, right, people have a multi-stage workflow, right? I don't think there's a single person out there using one monolithic model, not even CDance or whatever, for, for their entire, uh, task. Uh, people will have this, like, you know, kind of bunch of storyboards and mood boards of images they pull out from, like, their favorite collection of image models, and then they'll go to different video tools and, like, build those together as key frames. And then they'll go and, like, clip and edit those later, right? So we were seeing this, like, sort of, you know, niche but very specific use case for Marble as just providing that, like, sanity that you can ground your generations in some kind of 3D consistent world, right? People, you know, I know I've, I've fought with, with various image models to ask them to, like, give me different viewpoints of a room. And every time you can just look and see, oh, things kinda moved around, like, it's not stable. And, like, even that one seed of a use case I think kinda signals that there's this value and there's, hiding under the surface there, like, there's just decades of people being used to persistent 3D state, like, virtually modeling what they would be doing in the real world and having, you know, a stage and props and, like, elements there, whether it is for, uh, a movie or a show or a marketing shot or, like, building out game environments. Like, this, this statefulness and persistence is so key in how people think about spatial reasoning and, like, developing an environment over time. Like, people don't think in this ephemeral, like, generate a thing, generate a thing, like, just throw it away, keep my text prompts. Like, people wanna build this, like, collection of assets and, like, model a world in that way. So we're, we're trying to provide, like, again, with, with this spatial context mechanism and other things, like, we're trying to provide that level of control and precision, uh, and the ability to ingest different modalities of input, starting with the pose images. But, you know, we wanna give people more control over the elements of the things in the scenes they're looking at, and editing and interaction and all that as we go forward.

    13. MC

      Yeah.

    14. SP

      Um, and I think that that, it unlocks, like, further use cases in those areas we're already seeing, but, um, also expanding out into kind of any place people want to create a virtual replication or, or, like, you know, a pre-imagination of a real-world space they need to build, right-

    15. MC

      Yeah

    16. SP

      ... for architecture and construction. Like, I talked to a guy at some point building booths for conferences, right? There's just so many things in the world you don't think about need to be fabricated.

    17. MC

      Yeah.

    18. SP

      And every single one of those basically goes through this, like, pretty painstaking virtual design phase. Uh, and of that process, like, the part where you go into 3D software is kind of one of the most, like, arduous and, like, labor-intensive parts right now. Like, like, taking feedback on a 3D design from kind of, like, verbal commentary or sketches or really, really quick, um, stuff you got from, like, a creative director or, like, a design director or an architect or whatever. Like, mapping that back into the 3D representation is, like, 95% of the work, right? You can have a meeting, get feedback, and then you go back and do a week of revisions.

    19. MC

      Yeah.

    20. SP

      And that's just because, like, our software is kind of decades old at this point, and it, it just never became as intuitive as, you know, playing with Legos or, like, pottery or doing this stuff with your hands or sketching with a pencil.

    21. MC

      Yeah.

    22. SP

      Um, and this is the real place where AI can actually unlock a ton of value for people in their process, whether it's a creative application or something more industrial or design or, or whatever. Uh, and that, like, really motivates me to kind of build different flavors of, of our model to cater to those kind of people.

    23. MC

      Yeah. I, I, I can understand, uh, how it helps with the creatives 'cause, like, Marble did that, and also how that extends to things like design or architecture.

    24. SP

      Mm-hmm.

    25. MC

      Uh, but Fei-Fei, you acquired a robotics company, and so it's less, it's less-

    26. FL

      We just talked about it two weeks ago.

    27. MC

      I know, I know. [chuckles] But it's less obvious to me, especially in the context of Atlas, like, how that maps to robotics. So if you wouldn't mind just penciling that out.

    28. FL

      Yeah. Actually, Atlas is a, a, a key part of the puzzle. So, um, we acquired this company that was formerly known as SceneX, and what is their key technology? Right now, their key technology is a system that goes from real to sim and s- and then sim to real.

    29. MC

      Mm-hmm.

    30. FL

      And what does that mean in robotics situation? You want to train a robotic arm to, you know, uh, figure out how to, um, do cabling, let's say, in a, in a industrial setting. Well, you need a whole bunch of data to first train the robotic policy to do these cable-cables, cabling activity, and then you wanna evaluate if the robotic policy is doing a good job, and then you deploy the robot into the cabling- Environment

  7. 35:2137:55

    The Elephant in the Room: Video Models vs World Models

    1. JJ

      way.

    2. FL

      Yep.

    3. MC

      What, what, um, one piece of feedback that I got... I mean, by the way, congrats on the launch. It was overwhelmingly positive. I think it was probably the most significant model launch this year. And, you know, everybody said glowing things. But one person who's an expert in the space who I texted was like, "What'd you think?" And it's great presence. It's fantastic. It's amazing, but there needs to be more dynamics. [laughing] And so it seemed to be, at least in the robotics case, but generally it was kind of ideal to actually have a world that moves. And so maybe talk a little bit about that and then any other future directions that, A, you're comfortable sharing, but you think are worth talking through.

    4. JJ

      Yeah, I mean, like dynamics is clearly gonna happen. Like actually, um-

    5. FL

      We have baby dynamics.

    6. JJ

      We actually do have baby dynamics already, and this is something I think people didn't quite appreciate-

    7. FL

      Yeah

    8. JJ

      ... we didn't really highlight in the blog post. But like the previous Marble world model, it was like fundamentally static.

    9. MC

      Yeah.

    10. JJ

      Like the, the, the model just like could not handle any dynamics at all, and that was just like baked into the model architecture, baked into the training, like the whole thing was fundamentally static.

    11. MC

      Yeah.

    12. JJ

      Um, we already knew that that was a big problem post-Marble, and we already fixed it in Atlas, right?

    13. FL

      Yeah.

    14. JJ

      Like the Atlas architecture is already fundamentally supports dynamics. Um, the Atlas training data fundamentally has dynamics. And if you look carefully in some of the videos that we even posted-

    15. FL

      Yes

    16. JJ

      ... there actually is-

    17. MC

      I saw, I saw you steering the thing a little bit.

    18. FL

      The waves. Yeah, the waves, the water waves.

    19. JJ

      Yeah. So like some of the examples, there's like waves in the water, like in some of the like air-generated aerial views, there's like little cars moving around. So like dynamics is actually already in this model.

    20. MC

      But is it-- By the way, dynamics seems very problematic to me if you're trying to reconstruct 3D from multiple views, right?

    21. JJ

      It is. It is.

    22. MC

      And so like, are these things like at odds or-

    23. JJ

      No. So, so actually one of our theses here is that like, you know, if you're gonna do fundamental 3D reconstruction, you actually wanna have no dynamics. Like you wanna be able to model like exact views of the scene with exact frozen time.

    24. MC

      Right.

    25. JJ

      Um, but then like this is actually kind of a problem with our previous Marble approach, right? Like there, like you can try to find data that's fully static, but that's really hard to scale and really hard to get more of. And the thing we realized is that even in the 3D case where I want static output in the end, the best way to get it is actually expose the model to dynamics.

    26. MC

      Right.

    27. JJ

      Right? Like expose the model to as much dynamic stuff as you got, as much static stuff as you got, and let the model figure out how to factor out the dynamic stuff.

    28. MC

      Interesting.

    29. JJ

      So in, in, especially in like the-

    30. MC

      Oh, interesting. Yeah, yeah

  8. 37:5542:22

    Will We Get 4D Video You Can Walk Around In?

    1. MC

      So, so Ben, does that mean we're gonna get 4D video? Just gonna like go walk around. [chuckles]

    2. JJ

      I mean, I think-

    3. FL

      You can see the smile on their face, so. [laughing]

    4. MC

      So I, I actually can see if like you just stopped now and you only did kind of bigger, you know, faster, better, you could build almost an entire industry, right? It f- it feels like a very horizontal primitive and then, and if you did nothing else. But are there other things that are not just kind of bigger, faster that you're excited about for the applications you're focused on, which tend to be kind of more on the-

    5. FL

      Yeah

    6. MC

      ... kind of content creative 3D side?

    7. JJ

      Yeah, I'm really excited about pushing that kind of multimodal aspect. I think different modes of control is so critical here. I think like it's super underappreciated, especially in the academic community, how critical it is to add control conditioning to these models to kind of get out what's inside. I mean, honestly, this dynamic versus static is-

    8. MC

      I'm not, I, I'm not sure I even understand what those words mean.

    9. FL

      To, to translate in layman's language is-

    10. JJ

      Different inputs

    11. FL

      ... just editability.

    12. JJ

      Yeah. Yeah.

    13. FL

      I think editability is, is the key here.

    14. JJ

      Yeah. So I mean, this is something we've seen in like sort of single image models, uh, and starting this year in video models is starting to be unlocked in terms of, oh, like getting that flavor of like multi-turn or really like intuitively interpreting like I want like this person and this object and this thing to happen, and kind of combining those all together in like one pastiche without having to do a lot of like manual work with the system. Like it just interprets it, like kind of frontier image models are kind of there, right? For in terms of editing. But we haven't seen that propagate out as strongly into video yet, and then into world models, right? We've seen some really kind of toy examples of, oh, I can like put in a sentence and like, you know, a dinosaur appears or something with these like sort of real-time models. Um, but I wanna like turn that up to really industrial strength and make that... 'Cause like the, the trick here is you gotta add control but not compromise the quality of the model, or it just becomes a, a party trick basically. Like it's like no one is going to seriously think about swapping their like cutting-edge frontier video model usage for your model if you give them extra knobs, but the quality degrades. So I think it's really that game of like, how can we maintain like the high bar we've set with the outputs we're able to get in the current model, and then add all kinds of interesting stuff that people will ask us for in terms of like, I want to interact with the scene, or control the layout, or control like the identity of the objects and the things that we're seeing within there, or control time, right? Um, and I think that's like an axis where it opens up like a ton of really interesting product and interface work, the more complexity you add there and richness in terms of kind of like enabling you to really think about like redesigning almost from scratch the way people interact with sort of like stateful, you know, 3D worlds in the computer. Like that's, that's really the end goal here, is getting like all the capabilities you need to build that kind of system.

    15. MC

      Awesome. Anything you'd add to that as far as new functionality that you'd be excited about that's not just bigger, better?

    16. FL

      [chuckles] I think for me, let's go back to the first principle of intelligence. Intelligence is not sitting there stack and just seeing something or interpreting something when it comes to space and physical space, right? It's really this, uh, closing the loop between seeing and experiencing an interaction. So just thinking about going up that ladder is exactly what, uh, Ben said.

    17. MC

      Cool.

    18. JJ

      I think one interesting notion there is this notion of AI completeness. Have you heard this before?

    19. MC

      Yeah, yeah, I have. Yeah.

    20. JJ

      So like everyone-

    21. MC

      I, I hear about AI complete, by the way, in terms of LLMs-

    22. JJ

      Exactly

    23. MC

      ... which is like you have to be basically, you know, like the smartest LLM to answer the question what the smartest LLM will need to answer, or you have to solve general intelligence.

    24. JJ

      No, no. It's basically, it's a, it's a connection to Turing completeness, right?

    25. FL

      Mm-hmm.

    26. JJ

      Like the idea being that like i- a, a task is Turing complete, like in classical complexity theory, if like I can take any class and any problem in this category, reduce to that one problem, right?

    27. MC

      Yeah, yeah, yeah, yeah, yeah, yeah.

    28. JJ

      Like 3SAT's the classic example, right?

    29. MC

      Right, yeah.

    30. JJ

      So you can take any NP-hard problem, then reduce it to 3SAT, therefore, therefore you can use-

  9. 42:2243:32

    Why New View Prediction Is the Next Token Prediction

    1. JJ

      like-

    2. MC

      You could, you could have the, the movie and you do all of the frames of the movie, and then like the killer walks out, and then you predict exactly who walks out.

    3. JJ

      Exactly. [laughing] Not just that, we could say like, I want, I wanna have a world where like Martine is like writing a proof of the Riemann hypothesis on a blackboard.

    4. MC

      [laughing]

    5. JJ

      And then the camera pans over to the next whiteboard and like-

    6. FL

      And he solves the, uh...

    7. MC

      That's fucking-

    8. FL

      So, okay, so to take a evolutionary view, right? That new viewpoint, uh, prediction is exactly evolution had to solve by making animals move.

    9. MC

      Yeah.

    10. FL

      You, you, nature give animals eyes.

    11. MC

      Yeah.

    12. FL

      But nature didn't give trees eye.

    13. MC

      Yeah.

    14. FL

      Eyes. Why? Because when you move, you see a new viewpoint.

    15. MC

      Yeah.

    16. FL

      And that is the, the, the, whether you call it AI complete or intelligence complete. So, so we do believe very strongly that next viewpoint prediction is, is the equivalent of next token prediction.

    17. MC

      Amazing. Well, with that, congratulations all of you on a phenomenal model launch. We're very excited for future model launches, and thanks for coming.

    18. FL

      Thank yous.

    19. JJ

      Yeah, thanks so much

Episode duration: 43:42

Install uListen for AI-powered chat & search across the full episode — Get Full Transcript

Transcript of episode qn1QDDBnTA0

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.