Judge the Judge: Building LLM Evaluators That Actually Work with GEPA — Mahmoud Mabrouk, Agenta AI
This talk introduces GEPA, a method for building calibrated LLM evaluators that align with human annotations. The core idea is to move beyond generic LLM judges, which often fail in production, by optimizing prompts using algorithms like GEPA. This approach aims to accelerate development cycles by providing reliable signals for both offline evaluations and online monitoring, ultimately contributing to the creation of a data flywheel for continuous AI improvement.
Europe 2026 41 min