← Browse

World's Fair 2025

Break It 'Til You Make It: Building the Self-Improving Stack for AI Agents - Aparna Dhinakaran

Aparna Dhinakaran

Overview

This talk addresses the challenges of building and iterating on AI agents, particularly in evaluating their performance and identifying bottlenecks. It emphasizes the need for systematic evaluation beyond manual inspection of a few examples. The core thesis is that a self-improving stack for AI agents requires not only improving the agent's prompts and models but also continuously refining the evaluation methods themselves.

Who should watch

Key takeaways

Notable quotes

*Building agents is incredibly hard. There is a lot of iteration that goes on at the prompt level, at the model level, at iterating on the different tool call definitions.*
*The eval prompts that you're actually using, which is kind of the crux of how you identify those failure cases, end up becoming crucial to calling out what you need to improve.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.