As generative AI hits hardware and latency bottlenecks, Stanford professor, diffusion pioneer, and Inception co-founder and CEO Stefano Ermon is betting on a radical new architecture. Stefano joins Sarah Guo to talk about Inception, and how his team is applying diffusion architecture beyond images and video into discrete text and code generation. Stefano explains the limitations of autoregressive LLMs, as well as why parallel token generation in diffusion models offers superior inference scaling and hardware utilization on standard GPUs. He also shares details about Inception’s Mercury models, real-world voice agent applications, the software stack required to serve diffusion-based models at scale, academia’s role at the frontier of AI innovations, and why the next era of AI competition will be defined by efficiency. 

Sign up for new podcasts every week. Email feedback to show@no-priors.com

Follow us on Twitter: @NoPriorsPod | @Saranormous | @EladGil | @StefanoErmon | @_inception_ai

Chapters:

00:00 – Stefano Ermon Introduction

00:35 – Research Background

02:54 – Starting Inception

05:59 – Why Diffusion Beats Autoregressive

11:10 – Discrete vs. Continuous Modalities

13:19 – Inception Today

16:45 – Where Speed Wins

17:31 – Inception Customer Base

18:49 – Interaction with Hardware Landscape

19:34 – Inception and the Broader Industry

21:41 – Data Compression and Structure

24:45 – Controllability of Diffusion Modeles

27:25 – Emergent Capabilities at Scale

29:02 – Future Workload Split Between Diffusion vs. Traditional

30:03 – Adoption Challenges

31:44 – Hiring and Team Organization

32:50 – Recursive Self Improvement

34:02 – Resource Allocation

35:10 – Impact of Academia

38:13 – Conclusion

Podden och tillhörande omslagsbild på den här sidan tillhör Conviction. Innehållet i podden är skapat av Conviction och inte av, eller tillsammans med, Poddtoppen.