RAG Evaluation Is Broken! Here's Why (And How to Fix It) - Yuval Belfer and Niv Granot
This talk addresses the shortcomings of current Retrieval Augmented Generation (RAG) evaluation methods, arguing that they are fundamentally broken. The presenters contend that most benchmarks rely on simple, local questions with easily identifiable answers within specific text chunks, which does not reflect real-world data complexity. This leads to a cycle of optimizing for flawed benchmarks, resulting in RAG systems that perform poorly when deployed with actual user data.
World's Fair 2025 11 min