← Browse

World's Fair 2024

Lessons from the Trenches: Building LLM Evals That Work IRL: Aparna Dhinkaran

Overview

This talk focuses on the practical challenges and solutions for building effective LLM evaluation systems in real-world applications. It distinguishes between model evals, which rank models against benchmarks, and task evals, which assess whether an LLM application is functioning correctly for its intended purpose. The core thesis is that robust task evals, especially those providing explanations for failures, are crucial for iterating and improving deployed LLM applications.

Who should watch

Key takeaways

Notable quotes

*Evals with explanations are by far what we see real people deploying applications finding the most useful in production.*
*A single incorrect not incorrect is just really hard to know what to go fix but when you have something like an explanation like we were looking at it makes it easier for teams to go okay here's what I go fix here's what I go dig into.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.