World's Fair 2025
[Full Workshop] Building Conversational AI Agents - Thor Schaeff, ElevenLabs
Overview
This workshop focuses on building multilingual conversational AI agents, detailing the pipeline from speech-to-text to text-to-speech. It highlights the integration of large language models as the agent's "brain" and showcases ElevenLabs' tools for creating dynamic, responsive AI interactions across numerous languages. The session emphasizes practical application and developer experience, offering insights into configuring and deploying these agents.
Who should watch
- AI Engineers
- Product Managers
- Developers building AI applications
- Anyone interested in multilingual conversational AI
- Those looking to integrate advanced speech and language models into their products
Key takeaways
- Conversational AI agents typically follow a pipeline: speech-to-text, LLM processing, and text-to-speech.
- ElevenLabs offers a benchmark-leading automatic speech recognition (ASR) model supporting 99 languages, with features like speaker diarization and word-level timestamps.
- The platform provides a vast library of over 5,000 voices and allows for voice cloning, with a marketplace that has paid out over $5 million to voice actors.
- Agents can be configured through a dashboard or API, supporting custom LLMs and integrating tools for enhanced functionality like scheduling via webhooks.
- System tools, such as language detection, are built-in to facilitate seamless multilingual conversations and language switching.
- Latency is a key consideration, with strategies like using flash models for speech generation and configuring tool timeouts to maintain a natural conversational flow.
- Safety features include voice cloning verification, live moderation, and watermarking of generated speech to trace its origin and mitigate misuse.
- The platform supports agent-to-agent transfers, allowing for complex workflows and task routing to specialized agents.
Notable quotes
*The actual sounds that you're hearing are generated by our sound effects model.*
*We deploy kind of all these models very close to each other to kind of bring down the latency as as much as possible.*
*We have more than 5,000 different voices that you can choose from.*
Unofficial community note. Prefer the recording for nuance.