Europe 2026
Voice AI: when is the \"Her\" moment? — Neil Zeghidour, CEO, Gradium AI
Overview
The talk explores the current state of voice AI and the challenges in achieving a "Her moment," where AI voices are indistinguishable from humans and seamlessly integrated into conversations. While significant progress has been made in speech-to-text and text-to-speech, true conversational fluency, low latency, and the ability to handle complex tool calls remain major hurdles. The presentation contrasts cascaded systems with end-to-end speech-to-speech models, highlighting the limitations of current half-duplex speech-to-speech models in replicating human conversational nuances like overlapping speech and backchanneling.
Who should watch
- AI Engineers
- Product Managers
- Builders of voice applications
- Those interested in the future of human-AI interaction
- Developers facing challenges with voice AI latency and naturalness
Key takeaways
- Current voice AI, while advanced, still struggles with human-like conversational latency, often exceeding 200 milliseconds for a full response, compared to human conversational speed.
- The primary bottleneck in voice AI is shifting from TTS latency to the unpredictable latency of tool calls, which can range from 500 milliseconds to several seconds.
- Full-duplex speech-to-speech models, which can handle simultaneous speaking and listening, are crucial for natural human conversation, unlike current half-duplex models that break with overlapping speech or backchanneling.
- The Moshi model is presented as a robust, full-duplex conversational experience, capable of handling overlapping speech and paralinguistic cues, though it may lack the intelligence and observability of cascaded systems.
- Scalability and cost are significant challenges; voice AI, particularly TTS, is expensive and can hinder the profitability of consumer applications.
- On-device TTS solutions, like Gradion Phonon, offer a path to lower costs, enhanced privacy, and consistent performance by running models directly on smartphone CPUs.
- Achieving the "Her moment" requires not just natural-sounding voices but also reliability, intelligence, personalization, and cost-effectiveness, which are still areas of active research and engineering.
Notable quotes
*The last mile is going to be the most difficult to solve, and for us, it's really about science and engineering.*
*Voice is very challenging. The last mile is going to be the most difficult to solve.*
Unofficial community note. Prefer the recording for nuance.