← Browse

World's Fair 2025

Fuzzing in the GenAI Era — Leonard Tang, Haize Labs

Leonard Tang

Overview

This talk introduces hazing, a form of fuzz testing adapted for Generative AI systems. It addresses the critical challenge of validating and verifying AI outputs, which are inherently subjective and unstructured. Traditional evaluation methods relying on static datasets are insufficient due to the brittle and non-deterministic nature of GenAI, where minor input variations can lead to drastically different outputs. Hazing aims to pressure-test AI systems through large-scale simulation and optimization before deployment to ensure robustness and reliability.

Who should watch

Key takeaways

Notable quotes

*We know that AI systems are extremely unreliable. They're hard to trust in practice, and you sort of need to pressure test them before you put them out into the wild.*
*GenAI apps are incredibly brittle. And I think this is the actual core property that makes building with AI so difficult.*
*Standard evals don't have sufficient coverage.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.