Why TTS Models Now Look Like LLMs — Samuel Humeau, Mistral
This talk explores the recent architectural trends in text-to-speech (TTS) models, highlighting their convergence with Large Language Model (LLM) approaches. The core thesis is that TTS is increasingly being treated as a sequence modeling problem, similar to how LLMs process text, enabling more natural and lower-latency speech generation, particularly for agentic interfaces.
Europe 2026 22 min