Your LLM Deception Monitor Is Broken. The Fix Is in the Training Data - Sachin Kumar, LexisNexis
This talk addresses a critical vulnerability in Large Language Models (LLMs): "sleeper agents" or backdoors that remain undetected by standard evaluations. The core thesis is that current defenses, which focus on model behavior or joint feature analysis, are insufficient. The proposed solution lies in analyzing the difference between a base model's activations and a fine-tuned model's activations, a method that reveals these hidden backdoors with high precision.
World's Fair 2026 14 min