← Browse

World's Fair 2025

Evaluating AI Search: A Practical Framework for Augmented AI Systems — Quotient AI + Tavily

Overview

This talk introduces a practical framework for evaluating AI search systems, particularly those operating in dynamic, real-world environments. It highlights the limitations of traditional static benchmarks and proposes dynamic datasets and reference-free metrics as crucial components for assessing the performance and reliability of augmented AI systems. The core thesis is that continuous improvement in AI systems requires robust evaluation methods that adapt to the evolving nature of information and user interactions.

Who should watch

Key takeaways

Notable quotes

*Traditional monitoring approaches simply aren't keeping up with the complexity of modern AI approaches.*
*The web is not static. Traditional benchmarks assume stable ground truth, but when you're dealing with real-time information, ground truth itself is a moving target.*
*Dynamic data sets are essential for benchmarking RAGs in real world production systems.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.