Pressure-Testing Deception Probes in LLMs: Scaling, Robustness, and the Geometry of Deceptive Representations
This study addresses the sharp performance degradation of existing linear probes in detecting deception by large language models under distribution shift, a phenomenon whose root cause remains unclear. Through a systematic analysis of the geometric structure of deceptive representations in the Gemma 3 model family, the work proposes a cross-domain transfer matrix, a permutation-null-baseline multidimensional probe, an entropy-residualization test, and an evaluation framework incorporating eight stylistic perturbations. The research demonstrates for the first time that probe fragility stems from narrow training distributions rather than inherent architectural limitations, thereby refuting prevailing hypotheses such as single-direction encoding, entropy-based proxies, and linear subspace assumptions. The proposed style-augmented probes achieve average AUROCs of 0.979–0.983 on unseen styles, while multidimensional probes (k≥5) successfully recover distributed weak signals, overturning the inverse scaling conjecture.