Europe 2026
What Do Models Still Suck At? - Peter Gostev, Arena.ai, BullshitBench
Overview
This talk challenges the perception that AI models are rapidly approaching general intelligence, as suggested by steadily increasing benchmark scores. It argues that despite impressive progress, models still exhibit significant weaknesses, particularly in handling nonsensical or complex, real-world tasks. The presentation introduces a benchmark focused on nonsensical questions and analyzes user dissatisfaction data to reveal areas where models continue to struggle.
Who should watch
- AI engineers and researchers evaluating model capabilities.
- Product managers assessing the readiness of AI for complex applications.
- Builders seeking to understand the limitations of current LLMs.
- Anyone questioning the narrative of rapid AGI advancement based solely on benchmark performance.
Key takeaways
- Standard benchmarks may not reflect true model capabilities, as models can be trained to perform well on specific, narrow tasks.
- A benchmark designed with nonsensical questions revealed that many models, including prominent ones, readily accept and attempt to answer illogical prompts.
- Anthropic's Claude models, particularly newer versions like Sonnet 4.5, showed stronger performance in resisting nonsensical questions compared to GPT and Gemini models.
- The "reasoning" capability in models does not always improve performance and can sometimes lead to worse outcomes when dealing with flawed premises.
- Analysis of user dissatisfaction data from Arena.ai indicates that while overall model performance has improved, a significant percentage of users still encounter unsatisfactory responses, especially in expert-level tasks.
- Quantitative and mathematical tasks have seen substantial improvement, but areas like creative writing, finance, law, and game design show less dramatic progress.
- Models may be overly trained to "solve the task at any cost," lacking the ability to recognize and refuse to engage with nonsensical or inappropriate requests.
- There is a gap between the performance shown on well-defined benchmarks and the nuanced judgment required for complex, real-world white-collar work.
Notable quotes
*I think we could be deceiving ourselves a little bit.*
*It's really surprising me how easy it was for the models to just go along with a complete nonsense questions.*
*There's something that this kind of fuzziness that we all have in our hearts in our experience about the judgment that we have that doesn't necessarily match all of these super narrow very well defined very well specified tasks.*
Unofficial community note. Prefer the recording for nuance.