← Browse

Europe 2026

Judge the Judge: Building LLM Evaluators That Actually Work with GEPA — Mahmoud Mabrouk, Agenta AI

Mahmoud Mabrouk , Agenta AI

Overview

This talk introduces GEPA, a method for building calibrated LLM evaluators that align with human annotations. The core idea is to move beyond generic LLM judges, which often fail in production, by optimizing prompts using algorithms like GEPA. This approach aims to accelerate development cycles by providing reliable signals for both offline evaluations and online monitoring, ultimately contributing to the creation of a data flywheel for continuous AI improvement.

Who should watch

Key takeaways

Notable quotes

*The bottleneck in this loop is actually the evaluation. How fast can you evaluate?*
*Having calibrated LLM as a judge with a similar quality let's say as human annotator will make your development much faster.*
*The metrics need to come from the use case itself. It does not make sense to have general metrics like hallucination when you're evaluating your AI agent.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.