← Browse

World's Fair 2025

[Full Workshop] Building Metrics that actually work — David Karam, Pi Labs (fmr Google Search)

David Karam

Overview

This workshop focuses on building effective evaluation metrics for AI systems, moving beyond basic testing to create robust scoring systems. It emphasizes that evaluations are not just for testing but are the primary place where domain knowledge resides, enabling significant improvements in AI development. The session introduces a methodology for creating nuanced, calibrated metrics that correlate with desired outcomes, ultimately simplifying the AI development stack.

Who should watch

Key takeaways

Notable quotes

The speaker noted that *evals are actually the only place you're going to spend most of your time because that's where domain knowledge is going to live.*
Regarding metrics, it was stated that *evals is not like oh there's one way to do it and then you're done it's it's just part of how you do development.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.