← Browse

Europe 2026

Why building eval platforms is hard — Phil Hetzel, Braintrust

Phil Hetzel

Overview

Building effective evaluation (eval) platforms for AI agents is a complex systems problem, not just a UI challenge. While starting with simple tools like spreadsheets is a valid first step, maturing requires moving towards more robust solutions that facilitate experimentation and integrate with production data. The core difficulty lies in managing the unique characteristics of AI agent traces, which are often large, unstructured, and high-velocity, demanding specialized data infrastructure.

Who should watch

Key takeaways

Notable quotes

*LLMs have extreme variability. Agents are becoming the norm in how customers are interacting with companies.*
*Evals are important because LLMs have extreme variability.*
*The best way to perform evals is to really think about the failure modes that your agent can fall into and build scoring functions around those failure modes.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.