World's Fair 2025
[Full Workshop] Building Metrics that actually work — David Karam, Pi Labs (fmr Google Search)
Overview
This workshop focuses on building effective evaluation metrics for AI systems, moving beyond basic testing to create robust scoring systems. It emphasizes that evaluations are not just for testing but are the primary place where domain knowledge resides, enabling significant improvements in AI development. The session introduces a methodology for creating nuanced, calibrated metrics that correlate with desired outcomes, ultimately simplifying the AI development stack.
Who should watch
- AI engineers struggling to define and implement effective evaluation metrics.
- Product Managers and builders seeking to improve the quality and reliability of AI features.
- Teams finding traditional testing methods insufficient for complex AI applications.
- Developers looking to create more nuanced and domain-specific evaluations.
Key takeaways
- Evaluation is a critical component of AI development, serving as the repository for domain knowledge.
- Effective evaluation metrics are calibrated and correlate with user feedback or desired outcomes, rather than being abstract measures.
- A scoring system breaks down complex evaluation into multiple, inspectable signals, leading to more precise and reliable assessments.
- Starting with simple, correlated signals and iteratively adding complexity is a practical approach to building evaluation systems.
- Techniques like generating multiple responses and scoring them online can significantly improve output quality without model changes.
- The process of building and refining evaluation metrics can be integrated into development workflows, including using tools like co-pilots and spreadsheets.
- Sophisticated evaluation systems can be built using specialized models designed for high precision and low variance, enabling online deployment.
- *Metrics that work are not necessarily good metrics or bad metrics. They're either calibrated metrics or uncalibrated metrics.*
Notable quotes
The speaker noted that *evals are actually the only place you're going to spend most of your time because that's where domain knowledge is going to live.*
Regarding metrics, it was stated that *evals is not like oh there's one way to do it and then you're done it's it's just part of how you do development.*
Unofficial community note. Prefer the recording for nuance.