Europe 2026
The Art & Science of Benchmarking Agents — Vincent Chen, Snorkel AI
Overview
This talk addresses the growing gap between the capabilities of AI agents and the methods used to evaluate them. It argues that while progress is evident in areas like coding agents, current evaluation benchmarks are falling behind, creating hesitation in deploying these agents in high-stakes environments. The presentation outlines key principles for building effective benchmarks, distinguishing between the scientific rigor required for measurement and the artistic vision needed to shape the future of AI development.
Who should watch
- AI Engineers
- Product Managers
- Builders working with AI agents
- Researchers focused on AI evaluation
- Anyone concerned with the practical deployment and trustworthiness of AI agents
Key takeaways
- Effective benchmarks should not only measure progress but also define it, setting goals for future AI capabilities.
- The science of benchmarking involves ensuring individual task quality through rigorous validation and expert review, maintaining distributional diversity to represent real-world scenarios, and designing tasks with sufficient difficulty to reveal model limitations and headroom.
- Robust evaluation methodologies should extend beyond simple accuracy to capture critical real-world dimensions like cost, latency, and adherence to policy constraints.
- The art of benchmarking lies in developing a clear thesis about the future direction of AI, creating roadmaps that inspire new research, and prioritizing researcher user experience for broader adoption.
- Future benchmarks should focus on increasing environment complexity, extending autonomy horizons to reflect real-world operational lengths, and capturing a wider range of output complexity and nuanced signals.
- The Open Benchmarks Grant is funding the development of new benchmarks to accelerate progress in these areas.
Notable quotes
*Our ability to actually measure these agents in practice is falling behind of where the capabilities actually are.*
*The best open benchmarks aren't just about taking a snapshot of progress looking backwards. They're actually about defining progress and shaping the field.*
*Benchmarks should have a thesis on where the field is going that inspire new road maps.*
Unofficial community note. Prefer the recording for nuance.