← Browse

Europe 2026

Agentic Evaluations at Scale, For Everybody — Nicholas Kang & Michael Aaron, Google DeepMind

Nicholas Kang , Michael Aaron

Overview

This talk addresses the challenges of AI evaluations, which are currently scattered, quickly become outdated, and often lack transparency and verifiability. The speakers propose solutions to democratize the evaluation process, enabling a broader community to contribute to and benefit from robust AI benchmarking. Their work aims to foster more equitable AI development by allowing diverse expertise to shape evaluation standards.

Who should watch

Key takeaways

Notable quotes

*Evals are scattered, decentralized, and get stale fast.*
*We expect AI to help most of humanity, but then a very small percentage of people are creating all these evals.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.