What Do Models Still Suck At? - Peter Gostev, Arena.ai, BullshitBench
This talk challenges the perception that AI models are rapidly approaching general intelligence, as suggested by steadily increasing benchmark scores. It argues that despite impressive progress, models still exhibit significant weaknesses, particularly in handling nonsensical or complex, real-world tasks. The presentation introduces a benchmark focused on nonsensical questions and analyzes user dissatisfaction data to reveal areas where models continue to struggle.
Europe 2026 20 min