Agents reported thousands of bugs, how many were real? - Ian Butler and Nick Gregory
This talk investigates the effectiveness of AI agents in identifying and fixing bugs within software development, moving beyond traditional feature development benchmarks. The presenters introduce a new benchmark designed to evaluate agents on maintenance tasks, highlighting that while agents can often patch simple bugs, their ability to comprehensively detect them, especially complex ones, is still in its early stages. The research suggests current agents struggle with holistic code evaluation and deep reasoning, leading to missed bugs and high false positive rates.
World's Fair 2025 19 min