← Browse

World's Fair 2026

Your LLM Deception Monitor Is Broken. The Fix Is in the Training Data - Sachin Kumar, LexisNexis

Sachin Kumar

Overview

This talk addresses a critical vulnerability in Large Language Models (LLMs): "sleeper agents" or backdoors that remain undetected by standard evaluations. The core thesis is that current defenses, which focus on model behavior or joint feature analysis, are insufficient. The proposed solution lies in analyzing the difference between a base model's activations and a fine-tuned model's activations, a method that reveals these hidden backdoors with high precision.

Who should watch

Key takeaways

Notable quotes

*Your LLM deception monitor's broken, the fix is in training data.*
*A model can pass every eval you have and every behavioral monitor you run, it still be carrying a backdoor that flips it into malicious on a trigger you never tested.*
*Backdoors are directions, and the difference is where they live.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.