Turning Fails into Features: Zapier’s Hard-Won Eval Lessons — Rafal Willinski, Vitor Balocco, Zapier
Building effective AI agents and the platforms that enable non-technical users to create them is a complex challenge due to the inherent nondeterminism of AI and unpredictable user behavior. The process requires a shift from traditional software development to building a data flywheel that continuously collects feedback, understands usage patterns, and identifies failures to improve the product iteratively. This involves robust instrumentation, strategic feedback collection, and a tiered approach to evaluation.
World's Fair 2025 16 min