← Browse

World's Fair 2025

7 Habits of Highly Effective Generative AI Evaluations - Justin Muller

Justin Muller

Overview

This talk emphasizes that robust evaluations are the most critical, yet often overlooked, component for successfully scaling generative AI workloads. The speaker argues that evaluations are not merely for measuring quality but are primarily a tool for discovering problems within AI systems. Implementing a strong evaluation framework is presented as the key differentiator between a project that remains a science experiment and one that achieves production-level success and scalability.

Who should watch

Key takeaways

Notable quotes

*The number one thing that I see across all workloads is a lack of evaluations.*
*Evals are so important and we recognize that and this project is so important that we're going to invest the time.*
*The main goal with any evaluation framework should be to discover problems.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.