Europe 2026
From Transcription to Live Music: Gemini's Audio Stack — Thor Schaeff, Google DeepMind
Overview
This talk explores advancements in AI audio processing, focusing on Google DeepMind's Gemini models. The core thesis is that Gemini's sophisticated audio understanding capabilities enable richer transcription, robust reasoning, and more nuanced speech generation, moving beyond simple speech-to-text to a more comprehensive audio comprehension and synthesis system.
Who should watch
- AI engineers working with audio data.
- Product Managers seeking to integrate advanced audio features into applications.
- Builders interested in real-time conversational AI and speech synthesis.
- Developers exploring multimodal AI capabilities.
Key takeaways
- Gemini 3 models offer deep audio comprehension, going beyond transcription to understand nuances like emotion, pacing, and overlapping speech across multiple languages.
- The Gemini 3.1 Flash preview is a full-duplex, real-time conversational model capable of ingesting text, voice, and vision, and providing real-time audio and text responses.
- Speech generation is enhanced by audio understanding, allowing for the modification of a smaller set of base voices to achieve specific accents, emotions, and performance styles, rather than relying on large voice libraries.
- Tools like Echo Script, available in Google AI Studio, demonstrate the ability to extract structured information, including speaker labels, timestamps, language identification, emotion, and summaries, from a single audio request.
- The Voice Library application showcases how detailed prompts can direct speech generation to adopt specific accents and personas, such as an Irish accent or a Singaporean colloquial style.
- Gemini 3.1 Flashlight is a real-time speech-to-speech multimodal model that can be integrated via web sockets, offering immediate audio and text feedback.
- Lyra 3 is a music generation model with capabilities for generating jingles (Lyra 3 Clip) and full-length songs with lyrics (Lyra 3 Pro).
- A demo integrated Gemini Live with Lyra to create a German techno Schlager song about the UK startup scene, showcasing the combination of real-time interaction and music generation.
Notable quotes
*Gemini 3 is incredibly good at understanding audio. And that's not just transcribing it but really understanding all the nuances that are in there.*
*Our goal is to build models that deeply comprehend, richly transcribe and robustly reason through audio.*
*We can go from a small set of base voices to a very specific kind of voice that we're looking for for our speech generation.*
Unofficial community note. Prefer the recording for nuance.