← Browse

World's Fair 2026

Build Evals That Actually Matter - Nick Ung, Lyft

Nick Ung , Lyft

Overview

This talk addresses the common problem of AI evaluations that fail to predict real-world performance. The core thesis is that offline evaluations often use simplistic "customers" and test sets that don't reflect the complexity and adversarial nature of actual user interactions, leading to shipped models that fail in production. A solution involves building more realistic, adversarial user simulators trained on real data.

Who should watch

Key takeaways

Notable quotes

*Your agent passes offline evals at 90%. You ship. Production immediately finds failure modes your eval never saw.*
*The culprit is almost always the same: the customer in your offline eval is an off-the-shelf LLM that sounds nothing like your real users.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.