World's Fair 2024
How Zapier Builds AI Products and Features with the Help of Braintrust: Ankur Goyal & Olmo Maldonado
Overview
This talk details Zapier's approach to building AI-powered products, emphasizing an iterative development process heavily reliant on robust evaluation and observability. The speakers share their journey of integrating AI features, highlighting the challenges and successes encountered when developing and refining tools like the AI Zap Builder and Zapier Copilot. Their strategy involves close collaboration between product and engineering teams, continuous testing, and leveraging platforms like Braintrust to ensure product quality and performance.
Who should watch
- AI Engineers and Builders
- Product Managers working on AI features
- Teams looking to improve AI product development cycles
- Those interested in AI evaluation and observability strategies
- Developers building with large language models
Key takeaways
- Zapier employs an iterative approach to AI product development, prioritizing rapid prototyping and user feedback.
- Involving product managers early in the AI development process is crucial for aligning engineering efforts with product goals.
- A comprehensive evaluation suite, moving beyond basic unit tests, is essential for measuring and improving AI feature accuracy.
- Zapier built an in-house framework with Braintrust's assistance, utilizing synthetic data for seeding evaluations and running tests on a CI basis.
- Custom graders, both logic-based and LLM-based, are used to assess AI outputs against defined criteria.
- Observability tools, like those provided by Braintrust, are vital for understanding AI agent behavior, especially in complex, multi-tool frameworks.
- Switching to newer models like GPT-4o required prompt engineering adjustments to maintain performance and accuracy, demonstrating the need for careful model migration.
- The implementation of a robust evaluation system led to a nearly 300% improvement in accuracy for their AI Zap Builder.
Notable quotes
*We make adjustments through our evals and evals are the way that we make decisions.*
*The long tail of integrations is vast so how well are we doing it.*
*This is an engineering problem as well as a product problem.*
Unofficial community note. Prefer the recording for nuance.