← Browse

Europe 2026

The Art & Science of Benchmarking Agents — Vincent Chen, Snorkel AI

Vincent Chen , Snorkel AI

Overview

This talk addresses the growing gap between the capabilities of AI agents and the methods used to evaluate them. It argues that while progress is evident in areas like coding agents, current evaluation benchmarks are falling behind, creating hesitation in deploying these agents in high-stakes environments. The presentation outlines key principles for building effective benchmarks, distinguishing between the scientific rigor required for measurement and the artistic vision needed to shape the future of AI development.

Who should watch

Key takeaways

Notable quotes

*Our ability to actually measure these agents in practice is falling behind of where the capabilities actually are.*
*The best open benchmarks aren't just about taking a snapshot of progress looking backwards. They're actually about defining progress and shaping the field.*
*Benchmarks should have a thesis on where the field is going that inspire new road maps.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.