World's Fair 2026
Build Evals That Actually Matter - Nick Ung & Akshay Sharma, Lyft
Nick Ung , Akshay Sharma , Lyft
Overview
This talk focuses on the critical importance of building effective evaluation systems for AI models, moving beyond superficial metrics to create benchmarks that genuinely reflect real-world performance and user needs. The presenters emphasize that robust evals are essential for shipping reliable AI products and driving meaningful improvements in model quality.
Who should watch
- AI Engineers
- Product Managers
- Machine Learning Engineers
- Anyone building or evaluating AI systems
- Teams struggling with AI model quality and performance
Key takeaways
- Traditional AI evaluation metrics often fail to capture nuanced performance or user experience.
- Developing effective evals requires a deep understanding of the specific product use case and desired outcomes.
- Consider user-centric metrics and real-world scenarios when designing evaluation frameworks.
- Iterative evaluation is crucial; build, test, and refine your models and their corresponding evals continuously.
- The process of building good evals can itself reveal insights into model limitations and areas for improvement.
- Focus on creating evals that provide actionable feedback for development teams.
Notable quotes
The presenters stressed the need for evals that *actually matter* in reflecting real-world utility.
Building good evals is presented as a core part of the product development lifecycle for AI.
Unofficial community note. Prefer the recording for nuance.