← Browse

Europe 2026

Spec-Driven Testing for Agents With A Brain the Size of A Planet — Steven Willmott, SafeIntelligence

Steven Willmott

Overview

This talk introduces the concept of spec-driven testing for AI agents, arguing that simply using larger models or extensive datasets is insufficient for ensuring agent safety and reliability. It proposes a more comprehensive approach to defining and testing agent behavior by specifying not just expected outputs but also the context, rules, domain knowledge, and robustness requirements relevant to the agent's intended task.

Who should watch

Key takeaways

Notable quotes

*A smarter agent is a better agent, right? So if I have a smarter agent, I'm using a bigger model, uh it's going to be better at doing the job that it's supposed to do.*
*It's not obvious that bigger is safer and it's not obvious that bigger is better.*
*So what you're really seeking is like a model uh an agent that's built on a model that's good enough to perform but it's not capable of doing arbitrary harm.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.