← Browse

Session brief

Your Evals Are Meaningless (And Here’s How to Fix Them)

Overview

This talk argues that traditional evaluation methods for AI systems are often insufficient and can lead to meaningless results. The core thesis is that evaluations must be dynamic and continuously aligned with real-world usage and user expectations, rather than relying on static, generalized criteria.

Who should watch

Key takeaways

Notable quotes

*Your LM evals are really only as good as the alignment with real world usage.*
*Don't fall into the trap of static evaluation; don't treat tests like static tests in traditional software.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.