← Browse

World's Fair 2026

SWE-Marathon: Evaluating Coding Agents at Billion-Token Scale - Rishi Desai, Abundant AI

Rishi Desai , Abundant AI

Overview

This talk introduces SWE Marathon, a benchmark designed to evaluate the capabilities of coding agents on large-scale, project-level tasks. It addresses the growing need to assess if agents can maintain coherence and successfully complete complex engineering projects over extended operational periods, such as a billion token budget, moving beyond simple bug fixes to end-to-end project ownership.

Who should watch

Key takeaways

Notable quotes

*Can coding agents stay coherent over a billion token budget?*
*The important thing is that these aren't shallow failures. The average trial used 31 million tokens, and the longest rollout consumed 877 million tokens.*
*If you remember one thing from this video, it's that the future of SWE evals is not just harder unit tests.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.