← Browse

Europe 2026

From Transcription to Live Music: Gemini's Audio Stack — Thor Schaeff, Google DeepMind

Thor Schaeff

Overview

This talk explores advancements in AI audio processing, focusing on Google DeepMind's Gemini models. The core thesis is that Gemini's sophisticated audio understanding capabilities enable richer transcription, robust reasoning, and more nuanced speech generation, moving beyond simple speech-to-text to a more comprehensive audio comprehension and synthesis system.

Who should watch

Key takeaways

Notable quotes

*Gemini 3 is incredibly good at understanding audio. And that's not just transcribing it but really understanding all the nuances that are in there.*
*Our goal is to build models that deeply comprehend, richly transcribe and robustly reason through audio.*
*We can go from a small set of base voices to a very specific kind of voice that we're looking for for our speech generation.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.