SWE-Marathon: Evaluating Coding Agents at Billion-Token Scale - Rishi Desai, Abundant AI
This talk introduces SWE Marathon, a benchmark designed to evaluate the capabilities of coding agents on large-scale, project-level tasks. It addresses the growing need to assess if agents can maintain coherence and successfully complete complex engineering projects over extended operational periods, such as a billion token budget, moving beyond simple bug fixes to end-to-end project ownership.
World's Fair 2026 13 min