The Art & Science of Benchmarking Agents — Vincent Chen, Snorkel AI
This talk addresses the growing gap between the capabilities of AI agents and the methods used to evaluate them. It argues that while progress is evident in areas like coding agents, current evaluation benchmarks are falling behind, creating hesitation in deploying these agents in high-stakes environments. The presentation outlines key principles for building effective benchmarks, distinguishing between the scientific rigor required for measurement and the artistic vision needed to shape the future of AI development.
Europe 2026 23 min