← Browse

World's Fair 2025

Evals Are Not Unit Tests — Ido Pesok, Vercel v0

Ido Pesok

Overview

This talk introduces the concept of evals at the application layer, distinguishing them from traditional unit tests. It emphasizes that Large Language Models (LLMs) can be unreliable, leading to unexpected failures in AI applications even when basic functionality appears to work. The core thesis is that robust evals are crucial for building reliable AI products by systematically testing and measuring performance across a spectrum of user-driven scenarios.

Who should watch

Key takeaways

Notable quotes

*Improvement without measurement is limited and imprecise.*
*Evals give you the clarity you need to systematically improve your app.*

Watch on YouTube →

Up next · Evals first

Watch next

Field notes from running evals in production: what actually breaks.

Five hard earned lessons about Evals — Ankur Goyal, Braintrust

Full trail →

Unofficial community note. Prefer the recording for nuance.