← Browse

World's Fair 2025

Engineering Better Evals: Scalable LLM Evaluation Pipelines That Work — Dat Ngo, Aman Khan, Arize

Dat Ngo , Aman Khan

Overview

This talk focuses on building scalable LLM evaluation pipelines. It emphasizes that effective evaluation is crucial for developing high-quality AI products and goes beyond simple "LLM as a judge" approaches. The core thesis is that a comprehensive evaluation strategy involves multiple methods, continuous tuning, and integration into the development lifecycle to accelerate iteration and improve AI system performance.

Who should watch

Key takeaways

Notable quotes

*Evals are really important in this space because the reality of the fact is if you've ever seen a trace or something like that you're not going to inspect every single trace manually.*
*When we talk about architectures and things like that, um when the industry first started, this was state-of-the-art. Routers, right?*
*My hot take is that don't use out of the box evals. If you get out of the box, if you use out of the box evals, you'll get out of the box results.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.