← Browse

World's Fair 2025

Turning Fails into Features: Zapier’s Hard-Won Eval Lessons — Rafal Willinski, Vitor Balocco, Zapier

Rafal Willinski , Vitor Balocco , Zapier

Overview

Building effective AI agents and the platforms that enable non-technical users to create them is a complex challenge due to the inherent nondeterminism of AI and unpredictable user behavior. The process requires a shift from traditional software development to building a data flywheel that continuously collects feedback, understands usage patterns, and identifies failures to improve the product iteratively. This involves robust instrumentation, strategic feedback collection, and a tiered approach to evaluation.

Who should watch

Key takeaways

Notable quotes

*Building good AI agents is hard and building good platform to enable nontechnical people to build AI agents is even harder.*
*The ultimate judge are your users. You shouldn't be optimizing for the biggest scores for the evos and completely disregard the vibes.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.