← Browse

World's Fair 2025

Agents reported thousands of bugs, how many were real? - Ian Butler and Nick Gregory

Ian Butler , Nick Gregory

Overview

This talk investigates the effectiveness of AI agents in identifying and fixing bugs within software development, moving beyond traditional feature development benchmarks. The presenters introduce a new benchmark designed to evaluate agents on maintenance tasks, highlighting that while agents can often patch simple bugs, their ability to comprehensively detect them, especially complex ones, is still in its early stages. The research suggests current agents struggle with holistic code evaluation and deep reasoning, leading to missed bugs and high false positive rates.

Who should watch

Key takeaways

Notable quotes

*Agents struggle with holistic evaluation of files and systems, only finding subsets of bugs per run.*
*The most used agents in the world are the worst at finding and fixing complex bugs.*
*While that's still useful to know they're possible, it's more useful to know in everyday usage how are they capable on just everyday tasks.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.