← Browse

Session brief

Agent Evals: Finally, With The Map

Overview

This talk introduces a framework for evaluating AI agents, dividing agent evaluation into semantic and behavioral aspects. Semantic evaluation focuses on how an agent's internal representations align with reality, while behavioral evaluation assesses how an agent's actions and tool usage contribute to achieving its goals. Both aspects are further categorized into single-turn and multi-turn scenarios, providing a comprehensive map for understanding and measuring agent performance.

Who should watch

Key takeaways

Notable quotes

*Agent evaluation is rather art and science and ultimately it is nonetheless required to actually ensure that your agents do what you expect them to do especially when you're launching them in production.*
*The representations are in a sense a special case of tools special case of behaviors.*
*Let's go out there and make our agents measurable controllable and let's make sure they are actually doing our biding and not not rebelling against our ultimate intentions.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.