← Browse

World's Fair 2025

Optimizing inference for voice models in production - Philip Kiely, Baseten

Philip Kiely , Baseten

Overview

This talk focuses on optimizing inference for voice models in production, emphasizing runtime performance and infrastructure considerations. It highlights how the architectural similarity between Text-to-Speech (TTS) models and Large Language Models (LLMs) allows for the application of LLM optimization techniques. The core thesis is that while runtime optimizations are crucial, non-runtime factors like infrastructure and client code implementation can significantly impact overall latency and cost-efficiency.

Who should watch

Key takeaways

Notable quotes

*The most important thing here is again, while you can have great runtimes, the infrastructure to connect these three together is really what's going to determine your latency.*
*As much fun as it is to talk about the runtime stuff and as much work as we do there, the the infrastructure and the client implementation is equally important if not more so.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.