← Browse

Session brief

Building and evaluating AI Agents — Sayash Kapoor, AI Snake Oil

Sayash Kapoor , AI Snake Oil

Overview

This talk addresses the current limitations and challenges in building and evaluating AI agents. While agents are increasingly integrated into products, ambitious visions of their capabilities are far from realized. The core thesis is that AI engineering must prioritize reliability and rigorous evaluation, treating these as first-class concerns to overcome the inherent stochasticity of language models and move beyond misleading benchmarks.

Who should watch

Key takeaways

Notable quotes

*Evaluating agents is genuinely a very hard problem; it needs to be treated as a first-class citizen in the AI engineering toolkit.*
*The challenge for AI Engineers is to figure out what sorts of software optimizations and abstractions are needed for working with inherently stochastic components like LLMs.*
*AI Engineers need a reliability shift in your mindset to think of yourselves as the people who are ensuring that this next wave of computing is as reliable for end users as possible.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.