← Browse

Europe 2026

20 days of compute vs 7 hours: rethinking what state-of-the-art means — Bertrand Charpentier, Pruna

Bertrand Charpentier , Pruna

Overview

This talk challenges the conventional understanding of "state-of-the-art" AI models, arguing that relying solely on public leaderboards or internal manual evaluations can lead to suboptimal choices. It emphasizes that true state-of-the-art is context-dependent and that efficiency, not just raw quality, is a critical factor in model selection for practical applications. The core thesis is that a more nuanced approach to benchmarking, considering specific use cases and efficiency metrics, reveals a landscape of multiple specialized, high-performing models rather than a single dominant one.

Who should watch

Key takeaways

Notable quotes

*The problem with these methods is like in most cases if you apply them naively, you will always find like a kind of lazy solution which is just to to use a large foundation model.*
*So, the idea is that each leaderboard has a different perspective, and sometimes there are also some models that have duplicate entries.*
*So, in general, the idea is like you should never only trust the the the the manual inspection. It's good to get a feeling, but it's not enough.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.