← Browse

World's Fair 2025

Taming Rogue AI Agents with Observability-Driven Evaluation — Jim Bennett, Galileo

Jim Bennett , Galileo

Overview

This talk addresses the challenge of ensuring AI agents function reliably by introducing observability-driven evaluation. It highlights that AI's non-deterministic nature makes traditional testing methods insufficient. The core thesis is that by using AI itself to evaluate AI outputs, developers can gain crucial insights into agent performance, identify failures at granular levels, and implement targeted improvements.

Who should watch

Key takeaways

Notable quotes

*Detecting problems with AI is hard. It is a nondeterministic problem.*
*We can use AI to evaluate is our AI application actually working.*
*The best time to put evaluations in is as you're doing prompt engineering model selection. The second best time is now.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.