← Browse

Europe 2026

What Do Models Still Suck At? - Peter Gostev, Arena.ai, BullshitBench

Peter Gostev

Overview

This talk challenges the perception that AI models are rapidly approaching general intelligence, as suggested by steadily increasing benchmark scores. It argues that despite impressive progress, models still exhibit significant weaknesses, particularly in handling nonsensical or complex, real-world tasks. The presentation introduces a benchmark focused on nonsensical questions and analyzes user dissatisfaction data to reveal areas where models continue to struggle.

Who should watch

Key takeaways

Notable quotes

*I think we could be deceiving ourselves a little bit.*
*It's really surprising me how easy it was for the models to just go along with a complete nonsense questions.*
*There's something that this kind of fuzziness that we all have in our hearts in our experience about the judgment that we have that doesn't necessarily match all of these super narrow very well defined very well specified tasks.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.