Building Generative Image & Video models at Scale - Sander Dieleman, Google DeepMind
This talk provides a behind-the-scenes look at training generative image and video diffusion models at scale. It covers essential aspects from data curation and representation to modeling, architecture, training, sampling, and control mechanisms. The core thesis emphasizes that while diffusion models are powerful, their effectiveness relies heavily on careful data handling, efficient latent representations, and sophisticated sampling techniques like guidance.
Europe 2026 41 min