Europe 2026
20 days of compute vs 7 hours: rethinking what state-of-the-art means — Bertrand Charpentier, Pruna
Overview
This talk challenges the conventional understanding of "state-of-the-art" AI models, arguing that relying solely on public leaderboards or internal manual evaluations can lead to suboptimal choices. It emphasizes that true state-of-the-art is context-dependent and that efficiency, not just raw quality, is a critical factor in model selection for practical applications. The core thesis is that a more nuanced approach to benchmarking, considering specific use cases and efficiency metrics, reveals a landscape of multiple specialized, high-performing models rather than a single dominant one.
Who should watch
- AI engineers and researchers evaluating model performance.
- Product Managers and builders selecting models for applications.
- Anyone seeking to understand the trade-offs between model quality and computational cost.
- Teams looking to optimize AI model deployment for efficiency and effectiveness.
Key takeaways
- Public leaderboards often present conflicting rankings and may not reflect performance on specific use cases.
- Internal manual evaluations are prone to personal bias and limited sample size, leading to unreliable conclusions.
- Automated metrics can be inconsistent and require a deep understanding of what they actually measure.
- Efficiency, measured by compute time or cost, is as crucial as quality and should be considered alongside it.
- Pareto plots are a useful tool for visualizing the trade-off between model quality and efficiency, highlighting multiple optimal choices.
- For specific tasks, evaluating models against use-case-tailored metrics yields more relevant results than general benchmarks.
- The pursuit of state-of-the-art should focus on identifying specialized, efficient models rather than solely large foundational models.
- Techniques like quantization, pruning, and optimizing generation steps (e.g., denoising) can significantly improve model efficiency.
Notable quotes
*The problem with these methods is like in most cases if you apply them naively, you will always find like a kind of lazy solution which is just to to use a large foundation model.*
*So, the idea is that each leaderboard has a different perspective, and sometimes there are also some models that have duplicate entries.*
*So, in general, the idea is like you should never only trust the the the the manual inspection. It's good to get a feeling, but it's not enough.*
Unofficial community note. Prefer the recording for nuance.