Europe 2026
Agentic Evaluations at Scale, For Everybody — Nicholas Kang & Michael Aaron, Google DeepMind
Overview
This talk addresses the challenges of AI evaluations, which are currently scattered, quickly become outdated, and often lack transparency and verifiability. The speakers propose solutions to democratize the evaluation process, enabling a broader community to contribute to and benefit from robust AI benchmarking. Their work aims to foster more equitable AI development by allowing diverse expertise to shape evaluation standards.
Who should watch
- AI Engineers and researchers looking to improve model evaluation practices.
- Product Managers and builders seeking to understand the state of AI model performance.
- Developers working with coding agents and open-source AI models.
- Anyone interested in contributing to the open-source AI evaluation ecosystem.
- Individuals concerned about the equitable development and assessment of AI capabilities.
Key takeaways
- Current AI evaluations are fragmented, with numerous benchmarks appearing daily and quickly becoming obsolete, making it difficult to track progress.
- Many benchmarks lack transparency, making it hard to verify results or understand the specific configurations and testing methodologies used.
- The creation of AI evaluations is concentrated among a small group of researchers, potentially leading to biased assessments and overlooking critical real-world applications.
- Kaggle is developing platforms like hackathons, standardized agent exams, Game Arena, and a general benchmark platform to address these issues.
- Hackathons can channel community energy towards solving specific evaluation problems, with open-source results benefiting everyone.
- Standardized agent exams aim to provide accessible, baseline evaluations for consumer-facing agents before deployment.
- Game Arena uses PvP models in games like Werewolf, Poker, and Chess to create continuously evolving, unsaturated benchmarks.
- The Kaggle benchmark platform allows anyone to build, run, and share evaluations in an open and verifiable manner.
Notable quotes
*Evals are scattered, decentralized, and get stale fast.*
*We expect AI to help most of humanity, but then a very small percentage of people are creating all these evals.*
Unofficial community note. Prefer the recording for nuance.