← Browse

Europe 2026

Building Generative Image & Video models at Scale - Sander Dieleman, Google DeepMind

Sander Dieleman

Overview

This talk provides a behind-the-scenes look at training generative image and video diffusion models at scale. It covers essential aspects from data curation and representation to modeling, architecture, training, sampling, and control mechanisms. The core thesis emphasizes that while diffusion models are powerful, their effectiveness relies heavily on careful data handling, efficient latent representations, and sophisticated sampling techniques like guidance.

Who should watch

Key takeaways

Notable quotes

*Time spent on improving the data is sometimes a better investment of that time than actually trying to tweak the model.*
*Diffusion is basically spectral auto-regression.*
*Guidance is a no-brainer these days. So everyone just always leaves it on.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.