← Browse

World's Fair 2026

Don't Ship Skills Without Evals — Philipp Schmid, Google DeepMind

Philipp Schmid

Overview

This talk emphasizes the critical need for rigorous evaluation of AI skills before deployment, arguing that shipping skills without proper testing can lead to unpredictable behavior and degraded performance. The speaker highlights the distinction between agents developers use for personal productivity and agents built for consumers, noting that end-users lack the context to troubleshoot skill invocation issues. The core thesis is that comprehensive evaluations are essential for ensuring skill reliability, managing costs, and determining when skills can be retired as models improve.

Who should watch

Key takeaways

Notable quotes

*Human-written skills are the best we can provide. AI-generated skills can impact performance negatively.*
*Don't ship skills without evals.*

Watch on YouTube →

Up next · Evals first

Watch next

Hands-on: what evals look like when the unit of work is an agent, not a completion.

Ship Real Agents: Hands-On Evals for Agentic Applications — Laurie Voss, Arize

Full trail →

Unofficial community note. Prefer the recording for nuance.