Taming Rogue AI Agents with Observability-Driven Evaluation — Jim Bennett, Galileo
This talk addresses the challenge of ensuring AI agents function reliably by introducing observability-driven evaluation. It highlights that AI's non-deterministic nature makes traditional testing methods insufficient. The core thesis is that by using AI itself to evaluate AI outputs, developers can gain crucial insights into agent performance, identify failures at granular levels, and implement targeted improvements.
World's Fair 2025 16 min