Spec-Driven Testing for Agents With A Brain the Size of A Planet — Steven Willmott, SafeIntelligence
This talk introduces the concept of spec-driven testing for AI agents, arguing that simply using larger models or extensive datasets is insufficient for ensuring agent safety and reliability. It proposes a more comprehensive approach to defining and testing agent behavior by specifying not just expected outputs but also the context, rules, domain knowledge, and robustness requirements relevant to the agent's intended task.
Europe 2026 13 min