Speaker
Joel Becker
2 sessions in this library.
- How METR measures Long Tasks and Experienced Open Source Dev Productivity - Joel Becker, METR
This talk explores the challenges and potential of measuring AI capabilities, particularly focusing on developer productivity and long-term task completion. It questions the extrapolation of current AI progress trends, suggesting that physical and economic constraints, alongside potential technological breakthroughs, could alter the trajectory of AI development. The discussion also delves into the complexities of evaluating AI in real-world scenarios beyond controlled benchmarks, highlighting the gap between AI capabilities and practical application in fields like software engineering and data science.
- Why Agent Hype can fall short of reality – Joel Becker, METR
This talk addresses the discrepancy between AI capabilities suggested by benchmarks and real-world performance, particularly in developer productivity. It introduces two distinct methods of evaluation: benchmark-style assessments measuring AI performance on diverse tasks against human baselines, and field experiments examining AI's impact on experienced developers in complex, real-world coding environments. The core thesis is that while benchmarks show rapid AI advancement, practical application, especially in messy, high-context scenarios, reveals a more nuanced and sometimes even negative impact on productivity.