← Browse

Europe 2026

SWE-rebench: Lessons from Evaluating Coding Agents — Ibragim Badertdinov, Nebius

Ibragim Badertdinov , Nebius

Overview

This talk details the practical lessons learned from evaluating coding agents on real-world software engineering tasks using the rebench leaderboard. It emphasizes the critical need for robust evaluation beyond gut feelings or limited testing, especially as models are deployed to production. The presentation highlights the challenges and methodologies involved in creating and maintaining a "fresh" and "real-world" benchmark, focusing on the complexities of software engineering tasks that require understanding repository structures, writing and running tests, and handling multi-turn interactions and tool use.

Who should watch

Key takeaways

Notable quotes

*I believe that for the AI domain, we also could say that the cost of each mistake is higher than traditional software engineering.*
*I think that we need to evaluate everything.*
*I believe that it is better to have some minimalistic agent with strong infrastructure than having over-engineering agent with weak infrastructure.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.