← Browse

Europe 2026

Ship Real Agents: Hands-On Evals for Agentic Applications — Laurie Voss, Arize

Laurie Voss

Overview

This talk focuses on the practical aspects of evaluating and improving AI agents, moving beyond simple "vibe checks" to implement robust testing frameworks. It emphasizes that effective evaluation is crucial for shipping reliable AI applications, especially agents, which are prone to cascading failures. The session introduces methods for capturing agent behavior through tracing, categorizing failures, and implementing various types of evaluations (code-based, LLM-as-judge, human) to ensure agents perform as expected and to drive iterative improvements.

Who should watch

Key takeaways

Notable quotes

*Pretty legit is not enough to ship something to production. That is the whole point of evals is I is that the vibe is good, but we want to be doing better than vibes.*
*An eval that you haven't validated is just a fancy way of being wrong at scale.*
*The first time a regression shows up before it reaches your users instead of after, you will have justified the cost of building your evals.*

Watch on YouTube →

Up next · Evals first

Watch next

The category error — why unit-test instincts fail for stochastic systems.

Evals Are Not Unit Tests — Ido Pesok, Vercel v0

Full trail →

Unofficial community note. Prefer the recording for nuance.