Skip to content
a16za16z

Google DeepMind Lead Researchers on Genie 3 & the Future of World-Building

Genie 3 can generate fully interactive, persistent worlds from just text, in real time. In this episode, Google DeepMind’s Jack Parker-Holder (Research Scientist) and Shlomi Fruchter (Research Director) join Anjney Midha, Marco Mascorro, and Justine Moore of a16z, with host Erik Torenberg, to discuss how they built it, the breakthrough “special memory” feature, and the future of AI-powered gaming, robotics, and world models. They share: - How Genie 3 generates interactive environments in real time - Why its “special memory” feature is such a breakthrough - The evolution of generative models and emergent behaviors - Instruction following, text adherence, and model comparisons - Potential applications in gaming, robotics, simulation, and more - What’s next: Genie 4, Genie 5, and the future of world models This conversation offers a first-hand look at one of the most advanced world models ever created. Timecodes: 0:00 Introduction 0:29 The Evolution of Generative Models 1:10 Real-Time Interactivity & User Experience 4:35 Applications and Use Cases 8:15 The Importance of Special Memory 13:12 Emergent Behaviors & Model Capabilities 19:45 Instruction Following & Text Adherence 20:48 Comparing Genie 3 and Other Models 21:56 The Future of World Models & Modalities 32:23 Robotics, Simulation, and Real-World Impact 37:58 Looking Ahead: Genie 4, 5, and Future World Models 40:41 Are We Living in a Simulation? Resources: Find Shlomi on X: https://x.com/shlomifruchter Find Jack on X: https://x.com/jparkerholder Find Anjney on X: https://x.com/anjneymidha Find Justine on X: https://x.com/venturetwins Find Marco on X: https://x.com/Mascobot Stay Updated: Let us know what you think: https://ratethispodcast.com/a16z Find a16z on Twitter: https://twitter.com/a16z Find a16z on LinkedIn: https://www.linkedin.com/company/a16z Subscribe on your favorite podcast app: https://a16z.simplecast.com/ Follow our host: https://x.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details, please see a16z.com/disclosures.

Shlomi FruchterguestJack Parker-HolderguestErik TorenberghostMarco MascorrohostJustine MoorehostAnjney Midhahost
Aug 16, 202542mWatch on YouTube ↗

CHAPTERS

  1. 0:00 – 1:47

    Why Genie 3 feels like a breakthrough: real-time worlds from a few words

    The conversation opens with the core “wow” of Genie 3: prompting a world into existence and moving through it interactively. The hosts and researchers reflect on why the internet reaction was so strong and what felt game-changing internally.

    • World generation from short text prompts feels instantly tangible
    • Non-experts perceive outputs as “real,” signaling a quality threshold crossing
    • The team aimed to push the boundary of what real-time environment generation could be
    • Early reflections on why the release resonated so broadly
  2. 1:47 – 3:30

    From Genie 1/2 and GameNGen to Genie 3: combining projects into an ambitious model

    Jack explains how Genie 3 emerged from multiple internal lines of work: Genie 2’s world generation, Veo’s video quality progress, and Shlomi’s GameNGen (Doom) work. The key was realizing these threads could be unified into a single, more ambitious effort.

    • Genie 2: 3D-ish environments, lower fidelity; Veo 2: higher-quality video generation
    • GameNGen/Doom highlighted real-time simulation potential
    • Internal cross-pollination accelerated progress and raised ambition
    • The timeline and how strongly it resonated exceeded expectations
  3. 3:30 – 4:35

    The magic of latency: why real-time interactivity changes user experience

    Shlomi describes the qualitative shift when the model responds immediately—turning video generation into something that feels like “walking around.” The group emphasizes that real-time responsiveness is not just a speed metric; it fundamentally changes how people relate to generated worlds.

    • Real-time response creates a “presence” effect unlike offline video generation
    • Early “wow moment” came when GameNGen became fast enough to navigate
    • Release artifacts (overlays, keyboard controls) helped convey interactivity
    • The team intentionally pushed to the edge of feasibility
  4. 4:35 – 8:16

    Applications: entertainment, games, education, and agent training—enabled by world generation

    Justine asks about use cases, and both researchers argue the downstream applications flow from the same core capability: generating interactive worlds on demand. Jack ties this back to reinforcement learning’s need for scalable, diverse environments.

    • Core unlock: unlimited environments generated from text
    • Potential uses: gaming, controllable video, education, agent training, robotics
    • RL motivation: after Go/StarCraft, the limiting factor became the environment
    • Foundation-model pattern: build capability first, then discover emergent uses
  5. 8:16 – 11:33

    Spatial memory and persistence: the “paint stays on the wall” moment

    The discussion turns to Genie 3’s most surprising capability: persistent spatial memory across navigation. Jack and Shlomi explain it was an explicit goal—yet still shocking to see working reliably—and describe the current practical limits.

    • Spatial persistence example: painting, leaving, returning, and seeing changes remain
    • Planned capability, but stronger-than-expected results felt unbelievable even to the team
    • Genie 2 showed early signs of memory, but Genie 3 made it a headline target
    • Current design supports ~minute-scale memory; longer horizons are a future goal
  6. 11:33 – 13:13

    How they got consistency without explicit 3D reconstruction

    Shlomi clarifies what Genie 3 is not doing: it doesn’t rely on explicit 3D representations like NeRFs or Gaussian splatting for persistence. Instead, it generates frame-by-frame while maintaining consistency, which they believe helps generalization.

    • Avoided explicit 3D world reconstruction to prevent limiting assumptions
    • Frame-by-frame generation still achieves strong consistency
    • Generalization is prioritized over rigid/static-world priors
    • Reliability of “look away and look back” remains a striking demo
  7. 13:13 – 17:12

    Emergent behaviors from scale: physics, realism, and terrain-aware interactions

    The hosts ask about surprising behaviors emerging from scaling data/compute. The researchers point to improved physics and realism—especially water, lighting, storms—and to behavior that changes appropriately with terrain (skiing, swimming, puddles).

    • Scaling improves “world sense” rather than LLM-style reasoning
    • Large jump in photorealism from Genie 2 to Genie 3
    • Emergent physics cues: water behavior, lighting, storms, motion dynamics
    • Terrain-conditioned interaction: downhill speed, swimming in water, etc.
  8. 17:12 – 18:26

    Instruction following vs world priors: pushing into low-probability prompts

    They discuss the tension between generating what’s statistically likely and obeying specific instructions that may be unusual. Genie 3’s strong text adherence allows more controllable, imaginative worlds, though edge cases remain challenging.

    • Trade-off: realism/consistency vs strict adherence to unlikely instructions
    • Model often follows ‘silly’ or specific prompts surprisingly well
    • Text adherence enables creativity beyond mundane real-world scenes
    • Some low-probability scenarios remain harder for video/world models
  9. 18:26 – 20:48

    Why Genie 3’s text control improved: leveraging DeepMind’s internal learnings

    Genie 1/2 leaned on image prompting; Genie 3 moved to direct text control and gained major instruction-following improvements. Jack credits internal knowledge-sharing and expertise from other projects (including Veo) for accelerating this leap.

    • Shift from image prompting to direct text-to-world improved controllability
    • Image-prompt transfer issues: good images aren’t always good ‘world starts’
    • DeepMind’s internal expertise and collaboration “turbocharged” progress
    • Text control can recreate highly specific descriptions (e.g., a dog) convincingly
  10. 20:48 – 28:19

    Genie vs Veo: why interactive world models aren’t just “real-time video”

    Anjney probes whether video generation and interactive world generation will converge. Shlomi and Jack argue they’re currently different products with different trade-offs—Genie emphasizes navigation/control; Veo emphasizes cinematic quality and other features like audio.

    • Genie: navigation/action and interactive control; Veo: higher cinematic quality and audio
    • Modality vs axes like speed, controllability, and product goals
    • Combining everything into one model is technically and product-wise nontrivial
    • Different user needs: agent training vs filmmaking demand different capabilities
  11. 28:19 – 42:21

    What’s next: Genie 4/5 directions, access plans, and the long arc to embodied agents

    They discuss future directions—multiplayer/shared worlds, more capability, better realism—and emphasize they’re still early in accurate world simulation. The segment also covers robotics/agent composability (Genie as environment, not agent), plus public access and the closing simulation question.

    • Near-term: collect feedback; long-term: build more capable, more realistic world models
    • Embodied AI vision: learning from experience in rich simulated environments
    • Robotics: bridge data-driven realism with simulation-scale learning; acknowledge remaining gaps beyond vision
    • Access: desire to broaden availability, but no firm timeline yet; ends with a philosophical ‘are we in a simulation?’ wrap-up

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.