← Browse

Session brief

Ensure AI Agents Work: Evaluation Frameworks for Scaling Success — Aparna Dhinkaran, CEO Arize

Aparna Dhinkaran , CEO Arize

Overview

This talk addresses the critical need for robust evaluation frameworks for AI agents, moving beyond initial development to production readiness. It emphasizes that as AI agents become more sophisticated, particularly with multimodal and voice capabilities, ensuring their reliability and effectiveness in real-world scenarios requires a structured approach to testing and validation. The core thesis is that comprehensive evaluation at every stage of an agent's operation is essential for scaling success and building trust.

Who should watch

Key takeaways

Notable quotes

Aparna Dhinkaran, CEO Arize, stated that *evals aren't just at one layer of your entire trace*.
She also emphasized that *the goal here is really how do you make sure that you have evals throughout your application*.

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.