← Browse

Europe 2026

Why TTS Models Now Look Like LLMs — Samuel Humeau, Mistral

Samuel Humeau

Overview

This talk explores the recent architectural trends in text-to-speech (TTS) models, highlighting their convergence with Large Language Model (LLM) approaches. The core thesis is that TTS is increasingly being treated as a sequence modeling problem, similar to how LLMs process text, enabling more natural and lower-latency speech generation, particularly for agentic interfaces.

Who should watch

Key takeaways

Notable quotes

*The persistent anxiety that fills the rest of my life is calmed for as long as I have the flavor of something good in my mouth.*
*The dominant trend is to generate each sample one after the other. Then another era was generating the whole audio at once, but as we can as we saw, it's very interesting to have the beginning of the audio generated first so that we can start to play it out.*
*My take on this is that you we can go very, very far by just using speech as an interface.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.