World's Fair 2026
Stop Evaluating Models Like It's the 50s - Alejandro Vidal, Mindmakers
Overview
This talk proposes a shift from traditional, simplistic model evaluation methods to more sophisticated techniques borrowed from psychology and psychometrics. The core argument is that simply counting correct answers, a method akin to classical test theory, is insufficient. Instead, the presentation advocates for Item Response Theory (IRT) to provide a more nuanced understanding of model capabilities by calibrating individual questions and estimating model intelligence more accurately.
Who should watch
- AI Engineers
- Product Managers
- Builders evaluating LLM performance
- Anyone seeking to improve LLM benchmarking methodologies
- Those interested in applying psychometric principles to AI
Key takeaways
- Traditional LLM evaluation often relies on counting correct answers, a method called classical test theory, which is outdated and insufficient.
- Item Response Theory (IRT), adapted from psychometrics, offers a more robust approach by modeling item difficulty (B parameter) and model intelligence (theta parameter).
- IRT allows for the calibration of each question (item), recognizing that not all questions are equally important or difficult, leading to more accurate intelligence estimations.
- The A parameter in IRT quantifies item discrimination, helping to identify high-quality, informative questions and flag problematic or mislabeled items.
- This methodology enables auditing benchmarks to remove or improve weak items, and can significantly reduce benchmark size without losing significant ranking correlation.
- IRT can be used to detect potential issues like data leakage, overfitting, or inconsistencies in model inference platforms by analyzing residuals.
- Advanced applications include adaptive testing for benchmark protection, identifying model bias across different groups, and creating model "DNA" to understand relationships and detect distillations.
- Future research directions include multidimensionality, merging benchmarks, incorporating other signals like latency, and using psychometric models for alignment and interpretability.
Notable quotes
*At this moment the state in the industry is counting the number of right answers.*
*We have better questions, more complicated questions that may maybe we should pay more attention to.*
*Counting the number of right answers is not a good approach because I can create benchmarks that are not calibrated.*
Unofficial community note. Prefer the recording for nuance.