← Browse

World's Fair 2025

Five hard earned lessons about Evals — Ankur Goyal, Braintrust

Ankur Goyal

Overview

This talk emphasizes the critical role of robust evaluation (evals) in the development and deployment of AI systems. It argues that evals should be actively engineered, not passively accepted, and that they are essential for both playing offense by identifying new use cases and defense by ensuring quality. The core thesis is that by treating evals as a continuous engineering process, organizations can better adapt to rapid model advancements and ship higher-quality AI products.

Who should watch

Key takeaways

Notable quotes

*The best data sets are those that you can continuously reconcile as you actually experience what happens in reality.*
*You have to think about tools in terms of what the LLM wants to see and how you can use exactly what you present to the LLM to make it work really well.*
*You want to create evals so that if there's something ambitious that you want to do in the future, you are very well prepared when a new model comes out to just push a button and find that out.*

Watch on YouTube →

Up next · Evals first

Watch next

Where the practice is headed once measurement is table stakes.

The Future of Evals - Ankur Goyal, Braintrust

Full trail →

Unofficial community note. Prefer the recording for nuance.