Speaker
Nick Ung
2 sessions in this library.
- Build Evals That Actually Matter - Nick Ung & Akshay Sharma, Lyft
This talk focuses on the critical importance of building effective evaluation systems for AI models, moving beyond superficial metrics to create benchmarks that genuinely reflect real-world performance and user needs. The presenters emphasize that robust evals are essential for shipping reliable AI products and driving meaningful improvements in model quality.
- Build Evals That Actually Matter - Nick Ung, Lyft
This talk addresses the common problem of AI evaluations that fail to predict real-world performance. The core thesis is that offline evaluations often use simplistic "customers" and test sets that don't reflect the complexity and adversarial nature of actual user interactions, leading to shipped models that fail in production. A solution involves building more realistic, adversarial user simulators trained on real data.