← Browse

World's Fair 2026

Production Evals For Agentic AI Systems - Nishant Gupta, Meta Superintelligence Labs

Nishant Gupta

Overview

This talk argues that traditional benchmarks are insufficient for evaluating agentic AI systems in production. Agentic systems, which plan, call tools, and execute workflows, require a shift in evaluation focus from model capability to system behavior. The core thesis is that production telemetry and reliability metrics are paramount for dependable outcomes, moving evaluation from a pre-deployment phase to a continuous operational capability integrated into the system's control plane.

Who should watch

Key takeaways

Notable quotes

The question is no longer did the model generate the right answer? The question is did the system behave correctly?
*We are moving from evaluating answers to evaluating workflows. And that requires fundamentally different evaluation architectures.*
*Production telemetry is a most important evaluation signal.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.