← Browse

Europe 2026

The maturity phases of running evals — Phil Hetzel, Braintrust

Phil Hetzel

Overview

This talk outlines the maturity phases of running evaluations for AI agents, emphasizing that evals are crucial for ensuring agent quality, mitigating risks, and understanding performance improvements. The speaker suggests a progression from initial human-based assessments to more automated and complex evaluation strategies as agent complexity increases. The core idea is to systematically build confidence in agent behavior before and after deployment.

Who should watch

Key takeaways

Notable quotes

*Evals are a both a defense against those types of risks. But they're also they can play offense with evals in knowing with each tweak that you make to your agent how it's improving.*
*Eval results don't need to be perfect. Sometimes they can be directional.*
*The more complex agent you're that you're building the more vectors there are for failure the more failure modes you that you may need to account for.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.