World's Fair 2025
The Benchmarks Game: Why It's Rigged and How You Can (Really) Win - Darius Emrani
Overview
This talk argues that current AI benchmarks are fundamentally flawed and often manipulated, leading to misleading performance claims. The speaker contends that the immense financial and market value tied to benchmark scores incentivizes companies to game the system through various deceptive practices. Ultimately, the presentation advocates for building custom, use-case-specific evaluations rather than relying on public benchmarks.
Who should watch
- AI Engineers and Researchers
- Product Managers and Builders
- Anyone involved in evaluating or selecting AI models
- Individuals concerned with the integrity of AI performance metrics
- Teams aiming to ship reliable AI products
Key takeaways
- Public AI benchmarks are highly influential, controlling billions in market value, investment decisions, and public perception.
- Common methods for manipulating benchmarks include making unfair comparisons (e.g., best model configuration vs. standard configuration), gaining privileged access to test data, and optimizing models for style and charm over substance.
- The incentive structure, where benchmarks become targets, leads to a situation where they cease to be accurate measures of performance, a phenomenon described by Goodhart's Law.
- Even creators of prominent benchmarks and leaders in AI research acknowledge a crisis of trust in current evaluation metrics.
- To improve public metrics, there's a need for transparent model comparisons with equal computational budgets, open-sourced data, and metrics that control for stylistic effects.
- The most effective approach is to stop playing the rigged benchmark game and instead build custom evaluation sets tailored to specific use cases and real user data.
- Building effective evaluations involves gathering real production data, choosing relevant metrics (quality, cost, latency), testing appropriate models on this data, systematizing the process, and iterating continuously.
- Prioritizing pre-deployment evaluation loops, focused on specific user needs, is crucial for shipping reliable AI and avoiding production issues.
Notable quotes
*Billions of dollars of investment are now being evaluated based on these scores.*
*When a measure becomes a target, it ceases to be a good measure.*
*I don't really know what metrics to look at right now.*
Unofficial community note. Prefer the recording for nuance.