World's Fair 2025
The Future of Evals - Ankur Goyal, Braintrust
Overview
This talk discusses the evolution of AI model evaluation (evals), highlighting the shift from manual processes to automated optimization. It introduces Loop, an agent designed to enhance prompts, datasets, and scorers, thereby revolutionizing the eval process. The core thesis is that frontier models, particularly recent advancements, are now capable of significantly improving AI development workflows.
Who should watch
- AI Engineers
- Product Managers
- Builders working on AI products
- Those involved in AI model testing and quality assurance
- Individuals seeking to automate and improve their AI evaluation pipelines
Key takeaways
- The average organization using Braintrust runs nearly 13 evals daily, with some advanced users exceeding 3,000 evals per day.
- Historically, AI evals have been a manual process, requiring significant human effort to analyze dashboards and make code or prompt adjustments.
- Recent breakthroughs in frontier models, such as Claude 4, have dramatically improved their ability to refine prompts, data, and scoring mechanisms.
- Loop, an agent integrated into Braintrust, automates the optimization of prompts, datasets, and scorers, making the eval process more efficient.
- Loop can be configured to use various models, including those from OpenAI, Gemini, or custom-built LLMs.
- Users can review Loop's suggested edits side-by-side with their existing data and prompts, or enable an automatic optimization mode.
- The combination of optimized prompts, data, and scorers is crucial for achieving high-quality evals.
- The future of evals is expected to be heavily influenced by advancements in frontier models, leading to a more automated and effective development cycle.
Notable quotes
*The average org that signs up for Brain Trust runs almost 13 evals a day.*
*I actually think that is all going to change.*
*Eval themselves are going to be completely revolutionized by the latest and greatest that's coming out.*
Unofficial community note. Prefer the recording for nuance.