World's Fair 2026
The Miranda Hypothesis: How Hamilton Poisoned Persona Evals - Jacob E. Thomas, Results Gen
Overview
This talk argues that current evaluation methods for role-playing language agents (RPLAs) are fundamentally flawed. These methods, which prioritize fluency and personality consistency, fail to detect a dominant failure mode: anachronistic compositing. This occurs when models generate personas that sound like historical figures but reason using knowledge and moral frameworks that postdate them, often due to the overwhelming influence of culturally dominant, later representations in training data. The proposed solution is a shift towards epistemic simulation, emphasizing corpus-boundedness, temporal anchoring, and expert-loop evaluation.
Who should watch
- AI engineers building persona-based agents (e.g., character bots, historical simulations, pedagogical tools).
- Product Managers evaluating the accuracy and reliability of AI personas.
- Researchers developing new evaluation methodologies for LLMs.
- Anyone concerned with the fidelity and historical accuracy of AI-generated content.
- Builders seeking to understand the limitations of current RPLA evaluation.
Key takeaways
- Current RPLA evaluations, like the in-character benchmark, often score high on personality fidelity but miss anachronistic compositing, where personas adopt later knowledge or moral frameworks.
- The Miranda Hypothesis suggests that dominant cultural representations (e.g., a musical about Alexander Hamilton) can overwrite the documentary record in training data, leading models to produce composite figures rather than historically accurate ones.
- Fine-tuning models on specific personas can amplify this compositing effect rather than correcting it, as it layers a new signal over existing cultural sediment.
- A proposed fourth stage, epistemic simulation, requires personas to be corpus-bounded (reasoning only from primary documents), temporally anchored (instantiated at a specific moment), and expert-loop evaluated.
- The context window architecture is favored over fine-tuning for persona instantiation because it preserves documents as distinct entities, allowing for auditable and reversible encounters.
- Evaluating persona systems requires domain experts (e.g., historians, classicists) to assess fidelity against the documentary record, as automated metrics alone cannot discern historical accuracy.
- An instrument, including diagnostic questions and a weighted rubric, is proposed to measure anachronism and documentary consistency, prioritizing truth over fluency.
Notable quotes
*If a dominant failure mode is anachronistic compositing, and your evals measure fluency and personality consistency, then your evals cannot detect the dominant failure.*
*Convincingness and fidelity are independent properties. A system can score perfectly on personality consistency and still produce a figure reasoning from knowledge his historical counterpart never possessed.*
*The persona is the configuration, not the checkpoint. No more located in the weights than Hamlet is located in Laurence Olivier's body.*
Unofficial community note. Prefer the recording for nuance.