Skip to content
Lex Fridman PodcastLex Fridman Podcast

Gavin Miller: Adobe Research | Lex Fridman Podcast #23

Gavin Miller is the Head of Adobe Research. Adobe have empowered artists, designers, and creative minds from all professions working in the digital medium for over 30 years with software such as Photoshop, Illustrator, Premiere, After Effects, InDesign, Audition that work with images, video, and audio. Adobe Research is working to define the future evolution of these products in a way that makes the life of creatives easier, automates the tedious tasks, and gives more & more time to operate in the idea space instead of pixel space. This is where the cutting-edge deep learning methods of the past decade can shine more than perhaps any other application. Gavin is the embodiment of combing tech and creativity. Outside of Adobe Research, he writes poetry & builds robots. Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep23-sb See below for timestamps, and to give feedback, submit questions, contact Lex, etc. *CONTACT LEX:* *Feedback* - give feedback to Lex: https://lexfridman.com/survey *AMA* - submit questions, videos or call-in: https://lexfridman.com/ama *Hiring* - join our team: https://lexfridman.com/hiring *Other* - other ways to get in touch: https://lexfridman.com/contact *OUTLINE:* 0:00 - Introduction 1:11 - Poetry & crossover to creative work 6:35 - Turning one medium into another 7:45 - Creative process in both the space pixels and ideas 10:00 - Improving workflow in Adobe tools with AI 14:31 - Taking ideas from prototype to product 16:22 - Learning how to use Adobe tools 21:13 - Applications of deep learning 28:46 - Improving user experience from data 34:30 - Augmented reality and virtual reality 39:57 - Resistance to change 43:40 - Poem - Today I Left My Phone at Home 44:17 - Illusion of beauty in digital space 49:17 - Secret to a thriving research lab 55:27 - Future ideas in Adobe Research 58:13 - Robotics and animation in the physical world 1:08:01 - Poem - Cast My Ashes Wide and Far *PODCAST LINKS:* - Podcast Website: https://lexfridman.com/podcast - Apple Podcasts: https://apple.co/2lwqZIr - Spotify: https://spoti.fi/2nEwCF8 - RSS: https://lexfridman.com/feed/podcast/ - Podcast Playlist: https://www.youtube.com/playlist?list=PLrAXtmErZgOdP_8GztsuKi9nrraNbKKp4 - Clips Channel: https://www.youtube.com/lexclips *SOCIAL LINKS:* - X: https://x.com/lexfridman - Instagram: https://instagram.com/lexfridman - TikTok: https://tiktok.com/@lexfridman - LinkedIn: https://linkedin.com/in/lexfridman - Facebook: https://facebook.com/lexfridman - Patreon: https://patreon.com/lexfridman - Telegram: https://t.me/lexfridman - Reddit: https://reddit.com/r/lexfridman

Lex FridmanhostGavin Millerguest
Jun 10, 20191h 9mWatch on YouTube ↗

CHAPTERS

  1. 0:00 – 3:22

    Poetry as a window into Gavin Miller’s creative mind

    Lex opens with a humorous poem excerpt to introduce Gavin’s artistic side alongside his role leading Adobe Research. Gavin explains how poetry and technology have run as parallel threads throughout his life, occasionally cross-pollinating in surprising ways.

    • Lex frames Gavin as a hybrid of technologist, artist, and builder
    • A playful poem prompts discussion of Gavin’s creative origins
    • Gavin describes “parallel strands” of personal writing and technical work
    • Creativity as a recurring theme across career and hobbies
  2. 3:22 – 5:44

    From voice synthesis to smart homes: early experiments in “magical realism”

    Gavin connects writing to early AI/agent ideas, including composing a poem for a 1990s voice synthesizer. He describes building a proto–smart home and interactive photo albums that combine sensors, audio, and context-aware behavior.

    • Writing a poem iteratively by testing lines in an early speech synthesizer
    • Imagining an “intelligent agent” personality long before today’s AI
    • Early 1990s smart home reminders and context-triggered behaviors
    • Interactive photo albums that detect page turns and play matching sounds
    • Technology as a form of literary/experiential “magical realism”
  3. 5:44 – 7:43

    Uncanny valley in conversation: what it means for AI to “understand”

    The conversation shifts to dialogue systems and the uncanny valley—when AI sounds convincing but lacks grounded understanding. Gavin argues that flexible, multi-perspective explanations and varied phrasing can make systems feel less canned and more competent.

    • Uncanny valley arises when fluent output masks shallow understanding
    • Varied ways of describing the same concept boosts perceived intelligence
    • “Understanding” is partly in the eye of the beholder
    • Links to image captioning: multiple valid descriptions for one scene
  4. 7:43 – 10:00

    Creativity from pixels to ideas: Adobe’s spectrum from low-level craft to automation

    Lex and Gavin discuss how AI can shift creative work from tedious pixel labor toward higher-level ideation—without abandoning hands-on artistry. Gavin describes Adobe’s goal of supporting both highly controllable “analog-like” simulation tools and AI-driven acceleration for production workflows.

    • Support both expressive, low-level tools (e.g., oil paint/watercolor simulation) and high-level automation
    • AI frees creators from “perspiration” tasks like resizing and reformatting across devices
    • “Content velocity” challenges: reflowing layouts for aspect ratios and languages
    • Evolving roles: artisan vs. art director/conceptual designer with AI as partner
  5. 10:00 – 14:56

    AI inside Photoshop/Premiere: smart defaults, faster selection, and background removal

    Lex asks how AI can improve day-to-day creative workflows, especially for users who work manually. Gavin highlights practical entry points: smart “auto” settings, dramatically improved selection tools, and reliable background removal—often as a strong starting point with human refinement.

    • AI as an “auto button” that picks context-aware default parameters
    • Selection as a core primitive: from color boundaries to semantic object understanding
    • One-click or no-click selection for known categories (people/animals)
    • Background removal and matting: quick results plus touch-up for perfection
    • Robustness matters: good enough often beats waiting for perfect automation
  6. 14:56 – 16:32

    From research demo to shipped feature: robustness, intervention, and product reality

    Gavin contrasts academic novelty with the demands of shipping tools used by professionals. He describes how product success depends on high success rates plus UI mechanisms for recovery when models fail, enabling creators to move from 99% to 100% quickly.

    • Industrial research must pass “shipping review,” not just peer review
    • Tools don’t need perfection if users can intervene effectively
    • Designing workflows that recover gracefully from errors
    • Professional users tolerate small fixes if automation saves major time
  7. 16:32 – 21:12

    Teaching users in the moment: AI-guided learning from tutorials and context

    Lex raises a key adoption issue: powerful tools are hard to learn. Gavin describes research that mines thousands of tutorial hours and uses recent user actions to recommend next steps, relevant learning content, and context-aware guidance—moving toward an “assistant + teacher” experience inside the app.

    • Two goals: help with the current task and support long-term skill growth
    • Mining online tutorials to understand workflows and instructional structure
    • Predicting “what users do next” based on recent actions (CHI work)
    • Reducing the need for keyword search by using in-app context
    • Assistants as additive capability without increasing GUI complexity
  8. 21:12 – 25:33

    Compound AI workflows: Sky Replace and spatial search that feels like design

    Gavin highlights AI projects that compress multi-step workflows into a single action. He explains Sky Replace as a compound operation (selection, stock search, compositing, relighting) and introduces Concept Canvas for spatially constrained image search where layout intent becomes part of retrieval.

    • Sky Replace combines selection, geometry-aware search, compositing, and foreground recoloring
    • Relighting prevents surreal mismatches (Magritte reference as cautionary example)
    • Exploring many variations quickly supports “I’ll know it when I see it” ideation
    • Concept Canvas: assign keywords to regions (person center, dog right, etc.)
    • Spatial constraints increase user ownership and make search feel like creation
  9. 25:33 – 28:42

    Deep Fill and generative models: structure inference, resolution limits, and ensembles

    The discussion turns to removing objects and filling missing regions, contrasting classic patch-based methods with neural generative approaches. Gavin explains why global structure understanding is hard, why high-resolution generation remains challenging, and how future systems may route tasks to specialized “experts” with confidence estimation.

    • Content-aware fill works well for textures but struggles with hidden structure
    • GANs add learned “common sense” but can be brittle outside training distribution
    • High-resolution stability remains a key research barrier
    • Need diverse training data and competence estimation/guardrails
    • Potential future: ensembles of specialized models with dispatch/voting
  10. 28:42 – 34:28

    Data, trust, and privacy: learning from users without crossing the line

    Lex asks about leveraging Adobe’s massive user base to learn real workflows. Gavin emphasizes explicit permission, clear user benefit, and privacy-preserving approaches, describing a spectrum from high-level aggregate signals to opt-in detailed studies and careful governance.

    • Users can “teach” Adobe what matters—if trust and value exchange are clear
    • Privacy requires explicit consent and careful product framing
    • Learning without permanently storing sensitive user histories
    • Different tiers: broad coarse telemetry vs. opt-in fine-grained participation
    • Trust as a product asset protected by dedicated privacy leadership
  11. 34:28 – 43:37

    AR/VR and 3D creation: immersive design, spontaneity, and UI challenges

    Gavin explains how VR and AR differ: VR transports you to new worlds; AR brings digital assets into real context and demands adaptation to physical environments. They explore the promise of immersive tools for 3D layout, and the open question of whether precise CAD-like tasks can become practical in AR/VR through better interfaces.

    • VR: immersion, evolving hardware, professional use cases (architecture, scale, spatial reasoning)
    • AR: placing assets into real environments requires context-aware adaptation
    • Fidelity vs storytelling depends on the media and expectations (celebrity vs character)
    • Immersive 3D layout is more natural than juggling multiple 2D views
    • Key frontier: enabling fine-grained precision work with new UI paradigms
  12. 43:37 – 49:14

    Deepfakes, idealized selves, and imagined realities: benefits, harms, and media literacy

    Prompted by another poem, Lex asks about living in increasingly artificial digital worlds. Gavin notes the long history of flattering representation (portraits) while acknowledging modern risks: impossible standards, manipulation, and the need for public literacy about what images can and cannot prove.

    • Self-presentation and “improved” portrayals predate social media
    • Harm: people comparing themselves to impossible, edited standards
    • Benefit: visualizing alternate realities can inspire real-world creation
    • Concern: pre-visualization may dull the thrill of exploration (moon/Mars example)
    • Need for skepticism and multiple forms of evidence as synthetic media improves
  13. 49:14 – 55:26

    How Adobe Research stays inventive: interns, freedom, and the funnel from novelty to impact

    Gavin lays out his philosophy that interns are essential to a thriving lab: they bring fresh ideas and enable lightweight exploration of risky concepts. He explains Adobe’s culture of researcher autonomy balanced with strategic priorities, and how projects evolve into papers, features, or longer-term bets.

    • Interns inject new ideas and counterbalance conservatism in industrial research
    • Intern projects form a funnel: explore broadly, then invest to ship robustly
    • Flexible early phase: projects can pivot after initial investigation
    • Culture: freedom of project choice paired with reward for impact
    • Strategic priorities influence direction without heavy-handed mandates
  14. 55:26 – 58:11

    2019 and beyond: assistants, high-res GANs, and the Sensei platform for scaling impact

    Looking forward, Gavin highlights creative/analytics assistants that infer intent and offer helpful suggestions as a major direction. He also describes Adobe Sensei as a shared platform to standardize and deploy AI models across products, shortening the path from research idea to real user impact.

    • Assistants that understand intent and guide actions are a key future bet
    • GANs are promising but must become practical at high resolution and high quality
    • Sensei platform centralizes AI capabilities for reuse across product teams
    • Standardized deployment reduces time from invention to shipped feature
    • Convergence with graphics advances (e.g., real-time ray tracing) for “dancing with light”
  15. 58:11 – 1:09:11

    Snake robots as a lifelong muse: from animation physics to autonomy and personality

    The conversation closes on Gavin’s robotics passion—especially snake robots—originating from 1980s animation and soft-body simulation research. He recounts iterative hardware builds, the constraints of early onboard compute, modern autonomy enabled by embedded AI accelerators, and the long-term vision of robots that can explain their reasoning and feel meaningfully alive.

    • Robotics interest grew from simulation/animation: modeling muscles and motion
    • A multi-decade series of snake robots (and later a spider), with real-world engineering lessons
    • Early compute limits: discrete logic, 8-bit processors, minimal RAM, radio control
    • Modern embedded compute + neural acceleration enables real autonomy and perception
    • Goal: “illusion of life” plus explainable reasoning to support trust and debugging

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.