← Browse

World's Fair 2025

Agentic Excellence: Mastering AI Agent Evals w/ Azure AI Evaluation SDK — Cedric Vidal, Microsoft

Cedric Vidal

Overview

This talk focuses on the critical process of evaluating AI agents, emphasizing a shift from ad-hoc testing to a more methodical approach. It highlights that effective evaluation should begin at the earliest stages of AI development, not as an afterthought. The presentation introduces tools and frameworks, particularly the Azure AI Evaluation SDK, to ensure AI agents behave correctly and safely as they gain more independence.

Who should watch

Key takeaways

Notable quotes

*Real safety comes from layering smart mitigations at the application layer.*
*Before evaluating at scale you need first to cherry pick and look at specific examples.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.