← Browse

World's Fair 2024

E-Values Evaluating the Values of AI: Sheila Gulati and Nischal Nadhamuni

Overview

This talk addresses the critical and evolving landscape of AI evaluations, particularly as systems become more agentic and automated. It argues that current evaluation methods are often simplistic and prone to "benchmark hacking," failing to capture the true performance and values embedded in AI systems. The discussion emphasizes the need for more robust, multifaceted evaluation strategies that consider real-world scenarios, user experience, and the underlying values driving AI development to ensure these systems align with human goals.

Who should watch

Key takeaways

Notable quotes

*Evals may not be the most sexy talk at this conference but it might be one of the most important.*
*Can we evaluate the performance of these systems in relationship to our goals for those systems?*
*The best solvers for X are at the top of those leaderboards and that's a problem.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.