← Browse

World's Fair 2025

Building Effective Voice Agents — Toki Sherbakov + Anoop Kotha, OpenAI

Overview

This talk explores the advancements and practical applications of building effective voice agents, moving beyond traditional text-based AI. It highlights the emergence of speech-to-speech technology as a key component of the multimodal AI era, emphasizing that current models are now fast, expressive, and accurate enough for scalable production use. The discussion covers architectural patterns, key trade-offs in development, and best practices for creating robust and engaging audio experiences.

Who should watch

Key takeaways

Notable quotes

*The models are incredibly slow. They're very robotic in how they sound. And they're quite brittle.*
*The models are really, we do believe, at that good enough tipping point where you can start to build much more reliable experiences around it.*
*So, you can see it's much faster. It's incredibly expressive, emotional. You can steer it, you can get it excited, you can slow it down, and it's accurate.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.