Building and evaluating AI Agents — Sayash Kapoor, AI Snake Oil
This talk addresses the current limitations and challenges in building and evaluating AI agents. While agents are increasingly integrated into products, ambitious visions of their capabilities are far from realized. The core thesis is that AI engineering must prioritize reliability and rigorous evaluation, treating these as first-class concerns to overcome the inherent stochasticity of language models and move beyond misleading benchmarks.
20 min