← Browse

World's Fair 2025

RAG Evaluation Is Broken! Here's Why (And How to Fix It) - Yuval Belfer and Niv Granot

Yuval Belfer , Niv Granot

Overview

This talk addresses the shortcomings of current Retrieval Augmented Generation (RAG) evaluation methods, arguing that they are fundamentally broken. The presenters contend that most benchmarks rely on simple, local questions with easily identifiable answers within specific text chunks, which does not reflect real-world data complexity. This leads to a cycle of optimizing for flawed benchmarks, resulting in RAG systems that perform poorly when deployed with actual user data.

Who should watch

Key takeaways

Notable quotes

*Most benchmarks comprise from local questions that has local answers.*
*Existing benchmarks fail to capture these use cases; they are very limited.*
*We should note that the regular pipeline of chunking embedding retrieving reranking is not good enough for many questions.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.