20 days of compute vs 7 hours: rethinking what state-of-the-art means — Bertrand Charpentier, Pruna
This talk challenges the conventional understanding of "state-of-the-art" AI models, arguing that relying solely on public leaderboards or internal manual evaluations can lead to suboptimal choices. It emphasizes that true state-of-the-art is context-dependent and that efficiency, not just raw quality, is a critical factor in model selection for practical applications. The core thesis is that a more nuanced approach to benchmarking, considering specific use cases and efficiency metrics, reveals a landscape of multiple specialized, high-performing models rather than a single dominant one.
Europe 2026 20 min