World's Fair 2025
How to build world-class AI products — Sarah Sachs (AI lead @ Notion) & Carlos Esteban (Braintrust)
Overview
This talk, presented by Sarah Sachs of Notion AI and Carlos Esteban of Braintrust, focuses on the critical role of observability and rigorous evaluation in building world-class AI products. The core thesis is that the quality and scalability of AI products stem directly from robust evaluation processes, which allow teams to iterate effectively and ensure consistent performance beyond simple demos.
Who should watch
- AI Engineers
- Product Managers
- Builders of AI applications
- Teams struggling with AI product quality and iteration
- Those looking to scale AI development effectively
Key takeaways
- Building high-quality AI products requires a significant investment in evaluation and observability, with a suggested balance of 10% prompting and 90% evaluation and iteration.
- Notion AI prioritizes polish and exceptional customer experience in its AI products, even when rapidly integrating new foundation models, which necessitates a strong evaluation framework.
- The development of Notion AI features, from early AI writers to complex agentic capabilities like deep research, has been guided by iterative evaluation and user feedback.
- Effective AI product development involves curating targeted datasets, defining clear scoring functions (often LLM-as-a-judge or heuristic-based), and integrating these into a continuous iteration cycle.
- Tools like Braintrust are essential for managing the complexity of AI evaluation, enabling teams to track performance, identify regressions, and ensure product quality across diverse user bases, including multilingual ones.
- The process involves defining tasks, curating datasets, and implementing scoring mechanisms, which can be done both offline for iteration and online for production monitoring.
- Human input remains crucial for establishing ground truth, auditing edge cases, and training LLM judges, complementing automated evaluation processes.
- Remote evals extend the capabilities of platforms like Braintrust, allowing complex local code and custom tooling to be integrated into the evaluation workflow, bridging the gap between technical and non-technical teams.
Notable quotes
*All of the rigor and excellence that comes from building great AI products comes from observability and good evals and that's how you scale an engineering team.*
*I believe that's the right balance of work in order to know that you're not just shipping something that worked well in a demo... but actually worked consistently and for the for the users you were curious about.*
*Quality is much more important than quantity in terms of the insights and things that you extract.*
Unofficial community note. Prefer the recording for nuance.