Code 2025
How METR measures Long Tasks and Experienced Open Source Dev Productivity - Joel Becker, METR
Overview
This talk explores the challenges and potential of measuring AI capabilities, particularly focusing on developer productivity and long-term task completion. It questions the extrapolation of current AI progress trends, suggesting that physical and economic constraints, alongside potential technological breakthroughs, could alter the trajectory of AI development. The discussion also delves into the complexities of evaluating AI in real-world scenarios beyond controlled benchmarks, highlighting the gap between AI capabilities and practical application in fields like software engineering and data science.
Who should watch
- AI engineers and researchers interested in the future trajectory of AI capabilities.
- Product Managers and builders evaluating the impact of AI on developer productivity.
- Open-source contributors and maintainers seeking to understand AI's role in software development.
- Anyone curious about the limitations and potential of current AI models in complex, real-world tasks.
Key takeaways
- Extrapolating AI progress based on historical log-linear trends may be unreliable due to potential physical constraints, economic limitations, and unpredictable technological advancements.
- Measuring AI productivity in software development is complex, with initial studies showing a "J-curve" effect where familiarity with tools can initially decrease productivity before improving it.
- Current AI models struggle with complex, real-world tasks, particularly in domains like data science where data quality and implicit knowledge are critical, often failing to provide value beyond basic code generation.
- The effectiveness of AI tools in open-source projects is influenced by the high quality bar and maintenance focus inherent in such environments, often requiring human oversight for verification and cleanup.
- Evaluating AI capabilities requires moving beyond controlled benchmarks to real-world "in the wild" data and complex, fuzzy goals, though these methods also present significant methodological challenges.
- The development of AI for R&D automation faces hurdles, including the need for AI to understand and adapt to complex, often messy, real-world problems rather than just benchmark-style tasks.
- Advancements in robotics and chip manufacturing are also discussed, with skepticism regarding the speed at which AI might automate these complex, hardware-intensive fields.
Notable quotes
*The log linear lines continue through approximately the same number of orders of magnitude except maybe if there's some significant break in the inputs.*
*I think the lesson I learn over and over again at this data specs really matter. Really really matter.*
*It's not that surprising. My my only feeling about AI abilities is like, well, today is the 200th day my car didn't rocket off the Earth and escape velocity and fly to the moon.*
Unofficial community note. Prefer the recording for nuance.