← Browse

World's Fair 2026

The Miranda Hypothesis: How Hamilton Poisoned Persona Evals - Jacob E. Thomas, Results Gen

Jacob E. Thomas , Results Gen

Overview

This talk argues that current evaluation methods for role-playing language agents (RPLAs) are fundamentally flawed. These methods, which prioritize fluency and personality consistency, fail to detect a dominant failure mode: anachronistic compositing. This occurs when models generate personas that sound like historical figures but reason using knowledge and moral frameworks that postdate them, often due to the overwhelming influence of culturally dominant, later representations in training data. The proposed solution is a shift towards epistemic simulation, emphasizing corpus-boundedness, temporal anchoring, and expert-loop evaluation.

Who should watch

Key takeaways

Notable quotes

*If a dominant failure mode is anachronistic compositing, and your evals measure fluency and personality consistency, then your evals cannot detect the dominant failure.*
*Convincingness and fidelity are independent properties. A system can score perfectly on personality consistency and still produce a figure reasoning from knowledge his historical counterpart never possessed.*
*The persona is the configuration, not the checkpoint. No more located in the weights than Hamlet is located in Laurence Olivier's body.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.