Speaker
Thor Schaeff
3 sessions in this library.
- From Transcription to Live Music: Gemini's Audio Stack — Thor Schaeff, Google DeepMind
This talk explores advancements in AI audio processing, focusing on Google DeepMind's Gemini models. The core thesis is that Gemini's sophisticated audio understanding capabilities enable richer transcription, robust reasoning, and more nuanced speech generation, moving beyond simple speech-to-text to a more comprehensive audio comprehension and synthesis system.
- Building Conversational Agents — Thor Schaeff and Philipp Schmid, Google DeepMind
This talk introduces the Gemini Interactions API, a new unified API designed to simplify building with large language models and agents. It emphasizes a more developer-friendly interface, akin to industry standards, and introduces server-side state management for improved agent development. The session also showcases the Gemini Live API for real-time conversational AI applications, including audio and video processing, and demonstrates how to build a coding agent with file I/O and bash command execution capabilities.
- [Full Workshop] Building Conversational AI Agents - Thor Schaeff, ElevenLabs
This workshop focuses on building multilingual conversational AI agents, detailing the pipeline from speech-to-text to text-to-speech. It highlights the integration of large language models as the agent's "brain" and showcases ElevenLabs' tools for creating dynamic, responsive AI interactions across numerous languages. The session emphasizes practical application and developer experience, offering insights into configuring and deploying these agents.