World's Fair 2025
Milliseconds to Magic: Real‑Time Workflows using the Gemini Live API and Pipecat
Overview
This talk explores the development of real-time voice AI workflows, emphasizing voice as a natural and universal interface for the next generation of AI applications. It highlights the complexities involved in creating seamless voice interactions, from foundational LLMs and real-time APIs like Gemini Live API to orchestration frameworks such as Pipecat and application code. The presentation suggests that while significant progress has been made, many aspects of voice AI are still in early stages of development, with capabilities progressively moving down the technology stack.
Who should watch
- AI Engineers
- Product Managers
- Developers building voice-enabled applications
- Those interested in real-time AI interactions
- Builders exploring multimodal AI
Key takeaways
- Voice is considered a critical and universal building block for next-generation AI, offering a more natural interface than typing.
- Building effective voice AI involves addressing numerous challenges, including real-time responsiveness and dynamic UI generation.
- The voice AI stack comprises large language models, real-time APIs (e.g., Gemini Live API), orchestration frameworks (e.g., Pipecat), and application code.
- Capabilities in voice AI tend to mature and move down the stack over time, from application code to frameworks and eventually into APIs.
- Developing voice applications requires a shift in thinking from traditional programming to understanding how models drive application cycles, often leading to unexpected but potentially beneficial outcomes.
- A demo showcased a voice-driven application for managing lists (groceries, reading, work tasks), illustrating both successes and areas for improvement in model understanding and execution.
- The evolution of technology allows for new ways to manage information and tasks, drawing parallels to historical methods like tying knots or strings to remember things.
Notable quotes
*Voice is the most natural of interfaces.*
*We believe that voice is the most natural of interfaces and there will come a most of the interaction with language models will happen via voice.*
Unofficial community note. Prefer the recording for nuance.