Skip to content
a16za16z

Google DeepMind Developers: How Nano Banana Was Made

Google DeepMind’s new image model Nano Banana took the internet by storm. In this episode, we sit down with Principal Scientist Oliver Wang and Group Product Manager Nicole Brichtova to discuss how Nano Banana was created, why it’s so viral, and the future of image and video editing. Timestamps: 00:00 Intro 02:00 The Origin of Nano Banana and How It Got Its Name 04:15 The “Wow” Moments and Viral Launch 06:20 Seeing Yourself in AI 08:40 How AI Is Changing Art and Creative Work 11:00 Control, Customization & Character Consistency 14:00 Building Interfaces for Artists and Everyday Users 17:10 AI in Education and Visual Learning 20:25 Multimodal AI and the Future of Creativity 24:10 2D vs 3D: The Debate Over World Models 27:20 The Challenge of Taste, Preference & Artistic Style 31:10 The Japan Phenomenon & Creative Communities 35:00 From Images to Video: The Next Frontier 41:00 Working With Artists and Designing With Intent 47:30 The Next Era of Image Models 53:50 Closing Thoughts Follow Oliver on X: https://x.com/oliver_wang2 Follow Nicole on X: https://x.com/nbrichtova Follow Guido on X: https://x.com/appenz Follow Yoko on X: https://x.com/stuffyokodraws Follow Justine on X: https://x.com/venturetwins Stay Updated: If you enjoyed this episode, be sure to like, subscribe, and share with your friends! Follow a16z on X: https://x.com/a16z Subscribe to a16z on Substack: https://a16z.substack.com/ Follow a16z on LinkedIn: https://www.linkedin.com/company/a16z Listen to the a16z Podcast on Spotify: https://open.spotify.com/show/5bC65RDvs3oxnLyqqvkUYX Listen to the a16z Podcast on Apple Podcasts: https://podcasts.apple.com/us/podcast/a16z-podcast/id842818711 Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures.

Nicole Brichtovaguest
Oct 28, 202554mWatch on YouTube ↗

CHAPTERS

  1. 0:00 – 0:29

    Creators get time back: AI as a “tedious work” eliminator

    The conversation opens on how modern image/editing models shift creative work away from repetitive manual tasks. The guests frame AI as a new medium that can amplify artists rather than replace them.

    • AI reduces time spent on tedious editing and manual operations
    • Creative professionals can spend more time on ideation and composition
    • AI tools are compared to new artistic mediums (e.g., “watercolors for Michelangelo”)
    • The underlying thesis: tools expand what artists can do
  2. 0:29 – 2:07

    From Imagine + Gemini to “Nano Banana”: the model’s origin story

    Oliver and Nicole describe the lineage from DeepMind/Google’s Imagine image models to Gemini’s multimodal, conversational use cases. “Nano Banana” emerges as a best-of-both-worlds effort: Gemini’s multimodal intelligence plus Imagine’s visual quality—and the nickname that stuck.

    • Team background: Imagine family of image models and earlier Gemini image generation
    • Shift toward Gemini use cases: interactive, conversational, editing-focused workflows
    • Nano Banana combines Gemini multimodal “smartness” with high visual quality
    • The name “Nano Banana” becomes the memorable, sticky brand
  3. 2:07 – 3:18

    The viral inflection: LLM Arena demand and the first “this is big” signal

    Oliver recounts that the team didn’t anticipate virality until the public release dynamics made it obvious. Usage on LLM Arena surged so hard they had to increase capacity repeatedly, revealing unexpected mainstream pull for conversational image editing.

    • Virality wasn’t assumed during development
    • LLM Arena usage exceeded provisioned queries-per-second
    • Users were willing to “wait their turn” for access due to perceived value
    • Early proof that conversational image editing had broad appeal
  4. 3:18 – 7:01

    “It finally looked like me”: personalization and emotional resonance

    Nicole describes the breakthrough moment when the model could generate a convincing likeness from a single image—without fine-tuning. The discussion highlights why personal identity (self, family, pets) turns a neat demo into a deeply engaging product experience.

    • Zero-shot self-likeness from one image felt like a qualitative leap
    • Prior approaches required fine-tuning (e.g., LoRA) and multiple images
    • Emotional resonance is strongest when users try it on themselves/family
    • Internal adoption accelerated via playful trends (e.g., ’80s makeovers)
  5. 7:01 – 8:15

    What art becomes: intent, authorship, and human creativity in an AI era

    The hosts and guests debate whether “out-of-distribution” is a useful definition of art and land on “intent” as the core. They argue great creators will continue to differentiate themselves, using AI as a tool while contributing taste, vision, and meaning.

    • “Out-of-distribution” is considered too restrictive for defining art
    • Intent is emphasized as the defining element of meaningful art
    • Pros will keep using state-of-the-art tools; AI becomes another tool in the belt
    • AI doesn’t erase the need for skilled creatives with ideas and craft
  6. 8:15 – 9:59

    Control, customization, and character consistency as product-defining capabilities

    They dig into why many artists previously rejected AI tools: insufficient control and inconsistent characters. Nano Banana’s focus on customizability, consistency, and iterative conversational editing is positioned as key to enabling storytelling and professional workflows.

    • Artists demand control: consistent characters/objects across iterations
    • Style transfer and multi-image referencing unlock new editing workflows
    • Interactive, multi-turn iteration matches how art is actually made
    • Challenges remain: long conversations can degrade instruction-following
  7. 9:59 – 15:36

    Interfaces for everyone: from chat to pro “knobs,” nodes, and ComfyUI workflows

    The group explores the UI spectrum: simple chat-based creation for everyday users, sophisticated node-based pipelines for power users, and an emerging middle. They discuss future interfaces that suggest next steps automatically—without forcing users to learn dozens of controls.

    • Professional tools historically offer many knobs/dials; consumers want simplicity
    • Node-based UIs (e.g., ComfyUI) enable robust, composable workflows
    • A “middle UI” opportunity: more control than chat, less complexity than pro tools
    • Future UIs may proactively suggest edits based on context and intent
  8. 15:36 – 17:51

    Education and visual learning: using images for tutoring, diagrams, and understanding

    The conversation shifts to education, arguing that many people learn visually and current text-only tutoring is limiting. They envision AI that generates diagrams, figures, and step-by-step visuals to teach—not merely to beautify outputs.

    • AI tutoring today is overly text-based; visuals can improve comprehension
    • Models could pair explanations with generated figures/diagrams
    • Goal is guidance and critique (image “autocomplete”), not perfecting every child’s art
    • Generating “childlike” drawings is surprisingly hard and used in evals
  9. 17:51 – 20:04

    Multimodal + agentic futures: visual deep research, long-context design, and self-critique

    They forecast multimodal systems where image generation is tightly coupled with reasoning and long-context instructions. Examples include “visual deep research” for tasks like home redesign, and models that iterate, critique, and refine outputs—similar to inference-time scaling in text.

    • Multimodal models need both reasoning and visual generation to be truly useful
    • “Visual deep research” concept: the model explores options and returns structured proposals
    • Long context enables adhering to dense constraints (e.g., brand guidelines)
    • Self-critique loops and iterative refinement are framed as a major next step
  10. 20:04 – 22:25

    2D vs 3D world models: projections, consistency, data constraints, and robotics needs

    Oliver outlines the tradeoffs between explicit 3D representations and learning latent 3D from 2D projections. While 3D helps enforce consistency, training data is overwhelmingly 2D; humans also naturally work in 2D interfaces—though robotics ultimately needs 3D for movement.

    • Explicit 3D world models provide consistency advantages
    • Training data availability heavily favors 2D projections
    • Video models already show strong latent 3D understanding (reconstruction works well)
    • Robotics likely requires 3D for locomotion even if planning can be 2D-ish
  11. 22:25 – 34:55

    Evaluating “taste”: uncanny valley, subjective tradeoffs, and the lemon-picking era

    They unpack why character consistency is uniquely hard: people are hypersensitive to familiar faces, making benchmarks insufficient. The discussion broadens to how labs encode preferences into models, and why the frontier is improving worst-case outputs rather than cherry-picked best cases.

    • Uncanny valley is strongest for faces you know personally; evals must reflect that
    • Benchmarks struggle to compress multi-dimensional quality into one score
    • Model “taste” and tradeoffs reflect lab preferences and target users
    • Shift from “cherry-picking” to “lemon-picking”: raising the floor is the new goal
  12. 34:55 – 54:11

    Community acceleration and what’s next: Japan’s workflows, video, artist collaboration, and interleave

    The closing stretch highlights how communities (notably in Japan) build extensions and workflows that amplify the base model. They discuss images as frames on a continuum toward video, collaboration with artists/designers, an underused “interleave” feature, and near-term technical priorities like factuality and guideline compliance.

    • Japan’s creator community built tools/extensions for manga/anime workflows
    • Images and video converge: chaining frames and “what happens next” interactions
    • Working with artists (e.g., fine-tuning on sketches) to design real products
    • Underrated capability: interleave generation (multi-image story sequences)
    • Next priorities: factuality for education, stronger constraint-following, higher worst-case quality

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.