← Browse

Europe 2026

Road to 5 Million Tokens: Breaking Barriers in Long Context Training — Max Ryabinin, Together AI

Max Ryabinin

Overview

This talk details Together AI's research into training large context models, specifically aiming to break the 5 million token barrier. The core challenge lies in overcoming memory and computational bottlenecks inherent in standard transformer architectures when dealing with extremely long sequences. The research explores and combines various techniques to enable efficient training at unprecedented context lengths.

Who should watch

Key takeaways

Notable quotes

*The primary reasons for that are twofold, I would say. First of all, with the explosion in popularity of agents, you can see a lot of different applications where you might want to put as many tokens as you want in your context.*
*The problem here is that if you are taking a standard transformer-based language model and trying to extend its context, you can run into two bottlenecks. Bottleneck number one is that you are faced with quadratic computation.*
*The second problem is more insidious, one might say. As you continue scaling your context, your memory keeps growing linearly, which is not as bad, but still pretty difficult to deal with, unless you apply a range of specific techniques.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.