World's Fair 2025
2025 is the Year of Evals! Just like 2024, and 2023, and … — John Dickerson, CEO Mozilla AI
John Dickerson , CEO Mozilla AI
Overview
The core thesis of this talk is that 2025 will finally be the year for robust AI evaluation, driven by the convergence of several key factors. The increasing understanding of AI by business leaders, coupled with budget shifts towards generative AI projects, has set the stage. Crucially, the rise of agentic systems, which make decisions and take actions, introduces significant complexity and risk, making rigorous evaluation a necessity rather than an option.
Who should watch
- AI Engineers
- Product Managers
- Builders working with AI systems
- Those concerned with AI quality and risk
- Anyone deploying or considering agentic AI
Key takeaways
- AI monitoring and evaluation are two sides of the same coin, with measurement being core to evaluation.
- The widespread understanding of AI by CEOs and CFOs, spurred by ChatGPT, and a budget freeze that prioritized GenAI pet projects, shifted focus.
- Agentic systems, which act autonomously or semi-autonomously, introduce complexity and risk, making evaluation critical.
- Historically, ML model outputs were often obscured within larger systems, reducing the perceived need for direct evaluation by non-technical leadership.
- Enterprise budgets are now increasingly earmarked for AI applications, moving beyond science projects to production systems.
- Business leaders (CEOs, CFOs, CISOs, CIOs, CTOs) are now aligning on the need for AI evaluation due to concerns about ROI, governance, risk, compliance, and brand optics.
- The "LLM as a judge" paradigm is emerging as a practical, though imperfect, solution for data set creation and evaluation, requiring careful validation.
- Domain expertise remains crucial for evaluating complex AI outputs, with human experts being hired to validate agentic system performance in high-stakes scenarios.
Notable quotes
*AI monitoring and evaluation as two sides of the same sword or ruler.*
*2025 we all hear it. It's year of the agent. I no longer is a question mark needed here. It's clearly the the year of the agent.*
*The LLM as a judge paradigm is and I think this was talked about in in the previous talk as well.*
Unofficial community note. Prefer the recording for nuance.