Skip to content
No PriorsNo Priors

No Priors Ep. 60 | With Playground AI Founder Suhail Doshi

Multimodal models are making it possible to create AI art and augment creativity across artistic mediums. This week on No Priors, Sarah and Elad talk with Suhail Doshi, the founder of Playground AI, an image generator and editor. Playground AI has been open-sourcing foundation diffusion models, most recently releasing Playground V2.5. In this episode, Suhail talks with Sarah and Elad about how the integration of language and vision models enhances the multimodal capabilities, how the Playground team thought about creating a user-friendly interface to make AI-generated content more accessible, and the future of AI-powered image generation and editing. Sign up for new podcasts every week. Email feedback to show@no-priors.com Follow us on Twitter: @NoPriorsPod | @Saranormous | @EladGil | @Suhail Show Notes: 0:00 Introduction 0:52 Focusing on image generation 3:01 Differentiating from other AI creative tools 5:58 Training a Stable Diffusion model 8:31 Long term vision for Playground AI 15:00 Evolution of AI architecture 17:21 Capabilities of multimodal models 22:30 Parallels between audio AI tools and image-generation

Sarah GuohostSuhail DoshiguestElad Gilhost
Apr 18, 202424mWatch on YouTube ↗

CHAPTERS

  1. 0:00 – 1:34

    Why Suhail Doshi started Playground AI after DALL·E 2 and early Stable Diffusion

    Sarah and Suhail set the context for Playground AI as his third company and trace the spark back to early 2022 releases like DALL·E 2 and Stable Diffusion. Suhail explains why the moment felt like a step-change for creativity, and why the tooling ecosystem (e.g., Colab notebooks) needed a more accessible UI.

    • DALL·E 2 and Stable Diffusion as the inflection point for generative images
    • Early exposure to SD 1.4 reshaped his conviction about the space
    • Motivation: moving creation from technical notebooks into an approachable product UI
    • Founding framing: build for creators and sharing/distribution of images
  2. 1:34 – 4:03

    Choosing images over language or music: market dynamics and personal fit

    Suhail discusses why he focused on image generation rather than language or music, despite having a deep personal interest in music. He contrasts the competitive intensity and funding in language with the relative openness in images at the time, and emphasizes the importance of working where commitment and focus could win.

    • Music was tempting personally but unclear what product utility would be
    • Language felt crowded with many highly funded, highly capable teams
    • Images had fewer long-term, singularly focused efforts early on
    • Personal fit: creativity + tooling + natural sharing distribution of images
  3. 4:03 – 5:13

    From 'text-to-art' to practical utility: editing, blending, and real workflows

    The conversation shifts from pure text-to-image generation toward the broader promise of image manipulation. Suhail argues most models are effectively 'text-to-art' today and explains what’s missing for utility—editing existing images, compositing, lighting consistency, and stylization.

    • Distinction: current systems excel at art, not broad image utility
    • Key missing capability: robust editing of existing images with consistent lighting/realism
    • Opportunity: blending real and synthetic imagery in one coherent output
    • Utility framing as the path beyond novelty art generation
  4. 5:13 – 6:44

    Why train foundation models (Playground 2.5) instead of just fine-tuning

    Elad asks about Playground’s decision to train models rather than relying on Stable Diffusion fine-tunes. Suhail explains that training a top-tier image model is far more complex than architecture + data + compute, and describes their goal: push the SDXL recipe to deliver the best open-source model they could release.

    • Training from scratch is significantly more complex than it appears
    • Strategy: start with a proven recipe (SDXL-style U-Net + CLIP + VAE) and push it
    • Explicit objective: surpass SDXL as state-of-the-art open source
    • Team-building and research iteration as a core competency
  5. 6:44 – 8:02

    Aesthetics breakthroughs: fixing 'average brightness' with EDM sampling

    Suhail shares a concrete example of model quality improvements: SDXL-like outputs had an 'average brightness' look and weaker contrast. He explains how adopting an EDM formulation (noise/sampling changes) produced more vibrant colors and better contrast, highlighting how small mathematical choices can drive big perceptual gains.

    • Observed issue in baseline models: flat/average brightness and weaker contrast
    • EDM formulation changes noise sampling to improve color and contrast
    • Perceptual evaluation can initially look like an 'eval bug' when quality shifts
    • Small, targeted changes can yield outsized aesthetic improvements
  6. 8:02 – 10:01

    The real recipe: tricks, fast convergence, and meticulous curated data

    Pressed on how much is hand-tuning vs emergent improvement, Suhail describes a landscape of rapidly evolving techniques. He lists examples (power EMA, offset noise, DPO) but emphasizes the biggest gains often come from painstaking supervised fine-tuning on carefully curated datasets—where taste and judgment matter.

    • The field is nascent; new techniques emerge monthly
    • Technique examples: power EMA for faster convergence, offset noise, EDM, DPO-like methods
    • Large gains often come from late-stage supervised fine-tuning on curated data
    • Images require stronger human taste/judgment than typical language workflows
  7. 10:01 – 11:54

    Evaluating image models is flawed: coverage gaps, realism failures, and better taste signals

    Suhail critiques industry evaluation practices, noting that benchmarks can be misaligned with what users actually care about. He describes how gaps (e.g., photorealistic faces) show up after launch and why true coverage across styles and tasks requires extensive qualitative review of thousands of generations.

    • Many evals optimize for marketing-friendly metrics, not user satisfaction
    • Common post-eval discovery: unexpected capability gaps (e.g., photorealism/face quality)
    • Need broader coverage: paintings, logos, 3D, illustrations, world knowledge, celebrities
    • Current practice: manual inspection of thousands of images across checkpoints/grids
  8. 11:54 – 13:05

    Community feedback loops and in-product preference voting (with some secret sauce)

    Sarah asks about Playground’s use of user preference signals inside the product. Suhail explains they keep user interactions lightweight while running a sophisticated behind-the-scenes curation and ranking pipeline, though he notes the details are proprietary.

    • In-product studies/voting provide scalable preference signals
    • Design principle: users come to create, not to label—keep prompts simple
    • Backend: complex ranking/curation process built on lightweight user actions
    • Community becomes part of improving and selecting release candidates
  9. 13:05 – 14:40

    Where Playground aims to win: editing-first utility, consistency, and text rendering

    Suhail characterizes Playground’s current standing in text-to-art and outlines differentiation. He emphasizes shifting focus toward editing, character/likeness consistency, reducing the 'loot box' feel of prompting, and expanding into higher-utility tasks like text synthesis in images.

    • Positioning: strong text-to-art performance and closing the gap quickly
    • Differentiator: prioritize editing over one-shot generation
    • Problems to solve: character consistency, controllability, iterative tweaks
    • Roadmap utility: text synthesis and practical creative workflows (logos, compositions)
  10. 14:40 – 16:41

    Long-term vision: 'scaling pixels' into a large vision model that creates, edits, and understands

    Elad asks about the multi-year product and company vision. Suhail describes a focus on 'scaling pixels' starting with images (more efficient than video, more commercially direct than 3D tools) and a broader ambition: a multitask large vision model that can generate, edit, and understand visual content.

    • Thesis: complement text scaling with pixel scaling
    • Why images first: better compute economics than video; clearer business than 3D tooling
    • North star: a multitask large vision model
    • Three pillars: create, edit, understand (including VLM-like understanding)
  11. 16:41 – 18:26

    Architecture evolution: beyond U-Nets toward transformers—and merging with language knowledge

    The discussion turns to where generative vision architectures are heading, including diffusion transformers (DiT) and variants like MM-DiT. Suhail’s view is that transformers are the direction, but pure caption-to-image systems may lack interpretable knowledge; the big challenge is marrying language-model knowledge with image generation architectures.

    • DiT/MM-DiT as emerging architectures in the ecosystem
    • Transformers likely win long-term, but current forms may be insufficient
    • Caption-to-image alone may not encode enough interpretable/usable knowledge
    • Key research problem: integrate LLM-like knowledge/control with pixel generation
  12. 18:26 – 21:32

    Multimodal 'one model' debate: language control vs pixel density, plus data scaling constraints

    Sarah asks whether the future is one general multimodal model. Suhail agrees multimodality is inevitable and contrasts language’s low-dimensional controllability with vision’s high-density information about physics and spatial relationships. He also questions whether internet text is sufficient long-term and notes vision data can be collected endlessly—though filtering gets harder with more synthetic content online.

    • Agreement: multimodal systems are the destination
    • Language: compressed, controllable interface; may face a scaling ceiling
    • Vision: information-dense representation capturing spatial/physical details
    • Data view: vision can be collected at scale; internet data may be limited and noisier with synthetic content
  13. 21:32 – 24:31

    Parallels with audio AI: vocals/lyrics as the scarce resource and Suhail’s Suno workflow

    Elad pivots to music and audio generation tools. Suhail explains why audio is exciting and highlights that instrumentals are relatively easy compared to the scarcity of strong vocals and lyrics; he describes how tools like Suno unlock convincing flow and emotional delivery, and shares a personal workflow of generating songs, splitting stems, and recombining for higher quality production.

    • Audio is a major interest; would be his next focus absent Playground
    • Key scarcity in music: vocals and lyrics, not instrumentals
    • Suno as a breakthrough for flow, lyric quality, and vocal realism
    • Practical workflow: generate with AI, split stems, keep vocals/lyrics, rebuild instrumentals

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.