← Browse

Code 2025

How METR measures Long Tasks and Experienced Open Source Dev Productivity - Joel Becker, METR

Joel Becker

Overview

This talk explores the challenges and potential of measuring AI capabilities, particularly focusing on developer productivity and long-term task completion. It questions the extrapolation of current AI progress trends, suggesting that physical and economic constraints, alongside potential technological breakthroughs, could alter the trajectory of AI development. The discussion also delves into the complexities of evaluating AI in real-world scenarios beyond controlled benchmarks, highlighting the gap between AI capabilities and practical application in fields like software engineering and data science.

Who should watch

Key takeaways

Notable quotes

*The log linear lines continue through approximately the same number of orders of magnitude except maybe if there's some significant break in the inputs.*
*I think the lesson I learn over and over again at this data specs really matter. Really really matter.*
*It's not that surprising. My my only feeling about AI abilities is like, well, today is the 200th day my car didn't rocket off the Earth and escape velocity and fly to the moon.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.