← Browse

Europe 2026

Malleable Evals: Why Are We Evaluating Adaptive Systems with Static Tests? — Vincent Koc, OpenClaw

Vincent Koc

Overview

This talk argues that traditional static testing methods are insufficient for evaluating adaptive AI systems. As AI applications become more dynamic and intent-driven, evaluation strategies must evolve to become equally malleable. The core thesis is that static benchmarks fail to capture the emergent behaviors and changing user interactions characteristic of modern AI, necessitating a shift towards more adaptive and continuous evaluation approaches.

Who should watch

Key takeaways

Notable quotes

*Our AI applications are not static, but we're treating them like they're static software.*
*Why are stat benchmarks static? Like why don't we test in a more sort of adaptive manner?*
*People need to start looking at the evals not as this like static data set thing, but actually as like code as like software or as like a a living agent.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.