← Browse

Code 2025

Why Agent Hype can fall short of reality – Joel Becker, METR

Joel Becker

Overview

This talk addresses the discrepancy between AI capabilities suggested by benchmarks and real-world performance, particularly in developer productivity. It introduces two distinct methods of evaluation: benchmark-style assessments measuring AI performance on diverse tasks against human baselines, and field experiments examining AI's impact on experienced developers in complex, real-world coding environments. The core thesis is that while benchmarks show rapid AI advancement, practical application, especially in messy, high-context scenarios, reveals a more nuanced and sometimes even negative impact on productivity.

Who should watch

Key takeaways

Notable quotes

*Benchmarks seem to have less and less time between coming online and being fully saturated.*
*We find that developers are slowed down by 19%. They take 19% more time when AI is allowed relative to when AI is not allowed.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.