"Did you lie?" Evaluating Lie Detectors across Model Scale and Belief-Verified Model Organisms
Current language models lack a reliable evaluation framework for deceptive behavior due to the inaccessibility of their true beliefs. This work introduces the first set of reasoning-based model organisms (13 in total) with verifiable internal beliefs and constructs Varied Deception, a prompt-based lying benchmark. Coupled with a novel probe-training method, Did-You-Lie (DYL), the study systematically evaluates four deception detection approaches across models ranging from 2B to 1T parameters. Experiments reveal that while all detectors improve with model capability on prompt-induced lying tasks, only chain-of-thought–based discriminators maintain high accuracy (balanced accuracy of 0.82) on trained model organisms; other methods degrade significantly. This work establishes the first reliable methodology for assessing inconsistencies between a model’s internal beliefs and its generated outputs.