← Browse

Session brief

Voice Agents: the good, the bad, and the ugly

19 min

Overview

This talk explores the complexities and challenges of building AI voice agents, using a case study of an automated interview system. It highlights that while LLMs offer powerful capabilities, developing robust voice applications requires overcoming issues like transcription errors, conversational flow, and agent behavior. The presentation emphasizes the need for sophisticated architectural patterns beyond simple prompt engineering to achieve reliable and effective voice AI.

Who should watch

Key takeaways

Notable quotes

*Voice models are just generally tough to wrangle.*
*When we're dealing with audio and these voice agents everything's on hard mode.*
*The basic kind of call the API do some prompt engineering and get it into a good place is very helpful it gets you a very far in the development process but it's not enough to build your robust app.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.