← Browse

World's Fair 2025

Serving Voice AI at Scale — Arjun Desai (Cartesia) & Rohit Talluri (AWS)

Overview

This talk addresses the challenges and advancements in serving voice AI at scale, focusing on the critical need for low latency and high quality in real-time interactive applications. It introduces state space models (SSMs) as a more efficient alternative to transformers for handling long sequences, enabling faster and more natural voice interactions across various devices. The discussion highlights how these advancements are crucial for enterprise voice AI, customer support, gaming, and content creation.

Who should watch

Key takeaways

Notable quotes

*In voice, you don't have seconds to actually give your response back. You have milliseconds.*
*With state based models or SSM generation at inference time is O of one. We maintain a state that you can generate from.*
*Running our models on edge are about five times faster than if you were to roundtrip.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.