Skip to content
No PriorsNo Priors

Why Diffusion Will Win AI Inference with Inception Co-Founder and CEO Stefano Ermon

As generative AI hits hardware and latency bottlenecks, Stanford professor, diffusion pioneer, and Inception co-founder and CEO Stefano Ermon is betting on a radical new architecture. Stefano joins Sarah Guo to talk about Inception, and how his team is applying diffusion architecture beyond images and video into discrete text and code generation. Stefano explains the limitations of autoregressive LLMs, as well as why parallel token generation in diffusion models offers superior inference scaling and hardware utilization on standard GPUs. He also shares details about Inception’s Mercury models, real-world voice agent applications, the software stack required to serve diffusion-based models at scale, academia’s role at the frontier of AI innovations, and why the next era of AI competition will be defined by efficiency. Sign up for new podcasts every week. Email feedback to show@no-priors.com Follow us on Twitter: @NoPriorsPod | @Saranormous | @EladGil | @StefanoErmon | @_inception_ai Chapters: 00:00 – Stefano Ermon Introduction 00:35 – Research Background 02:54 – Starting Inception 05:59 – Why Diffusion Beats Autoregressive 11:10 – Discrete vs. Continuous Modalities 13:19 – Inception Today 16:45 – Where Speed Wins 17:31 – Inception Customer Base 18:49 – Interaction with Hardware Landscape 19:34 – Inception and the Broader Industry 21:41 – Data Compression and Structure 24:45 – Controllability of Diffusion Modeles 27:25 – Emergent Capabilities at Scale 29:02 – Future Workload Split Between Diffusion vs. Traditional 30:03 – Adoption Challenges 31:44 – Hiring and Team Organization 32:50 – Recursive Self Improvement 34:02 – Resource Allocation 35:10 – Impact of Academia 38:13 – Conclusion

Sarah GuohostStefano Ermonguest
Sep 18, 202638mWatch on YouTube ↗

Episode Details

EPISODE INFO

Released
September 18, 2026
Duration
38m
Channel
No Priors
Watch on YouTube
▶ Open ↗

EPISODE DESCRIPTION

As generative AI hits hardware and latency bottlenecks, Stanford professor, diffusion pioneer, and Inception co-founder and CEO Stefano Ermon is betting on a radical new architecture. Stefano joins Sarah Guo to talk about Inception, and how his team is applying diffusion architecture beyond images and video into discrete text and code generation. Stefano explains the limitations of autoregressive LLMs, as well as why parallel token generation in diffusion models offers superior inference scaling and hardware utilization on standard GPUs. He also shares details about Inception’s Mercury models, real-world voice agent applications, the software stack required to serve diffusion-based models at scale, academia’s role at the frontier of AI innovations, and why the next era of AI competition will be defined by efficiency. Sign up for new podcasts every week. Email feedback to show@no-priors.com Follow us on Twitter: @NoPriorsPod | @Saranormous | @EladGil | @StefanoErmon | @_inception_ai Chapters: 00:00 – Stefano Ermon Introduction 00:35 – Research Background 02:54 – Starting Inception 05:59 – Why Diffusion Beats Autoregressive 11:10 – Discrete vs. Continuous Modalities 13:19 – Inception Today 16:45 – Where Speed Wins 17:31 – Inception Customer Base 18:49 – Interaction with Hardware Landscape 19:34 – Inception and the Broader Industry 21:41 – Data Compression and Structure 24:45 – Controllability of Diffusion Modeles 27:25 – Emergent Capabilities at Scale 29:02 – Future Workload Split Between Diffusion vs. Traditional 30:03 – Adoption Challenges 31:44 – Hiring and Team Organization 32:50 – Recursive Self Improvement 34:02 – Resource Allocation 35:10 – Impact of Academia 38:13 – Conclusion

SPEAKERS

  • Sarah Guo

    host

    Venture capitalist and co-host of No Priors.

  • Stefano Ermon

    guest

    Stanford professor and co-founder/CEO of Inception working on diffusion-based generative models.

EPISODE SUMMARY

In this episode of No Priors, featuring Sarah Guo and Stefano Ermon, Why Diffusion Will Win AI Inference with Inception Co-Founder and CEO Stefano Ermon explores why diffusion could outscale autoregressive LLMs on inference speed Stefano Ermon traces his path from early Stanford generative modeling work (VAEs/GANs) to score-based methods that became modern diffusion, and explains why he believes diffusion is the next major paradigm shift for language inference.

RELATED EPISODES

Coinbase’s Everything Exchange: Agentic Finance, Stablecoins & Tokenization with CEO Brian Armstrong

Coinbase’s Everything Exchange: Agentic Finance, Stablecoins & Tokenization with CEO Brian Armstrong

From Restoring Sight to Reimagining the Brain, with Max Hodak

From Restoring Sight to Reimagining the Brain, with Max Hodak

Rethinking Legacy Data Infrastructure with Eon Co-Founders Ofir Ehrlich and Gonen Stein

Rethinking Legacy Data Infrastructure with Eon Co-Founders Ofir Ehrlich and Gonen Stein

Redefining Chip Architecture with Arm CEO Rene Haas

Redefining Chip Architecture with Arm CEO Rene Haas

How Nuclear Will Unlock Energy Abundance with Valar Atomics Founder Isaiah Taylor

How Nuclear Will Unlock Energy Abundance with Valar Atomics Founder Isaiah Taylor

How Chess.com Became the World’s Biggest Chess Community with CEO Erik Allebest

How Chess.com Became the World’s Biggest Chess Community with CEO Erik Allebest

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.