Pressure-Testing Deception Probes in LLMs: Scaling, Robustness, and the Geometry of Deceptive Representations

📅 2026-05-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the sharp performance degradation of existing linear probes in detecting deception by large language models under distribution shift, a phenomenon whose root cause remains unclear. Through a systematic analysis of the geometric structure of deceptive representations in the Gemma 3 model family, the work proposes a cross-domain transfer matrix, a permutation-null-baseline multidimensional probe, an entropy-residualization test, and an evaluation framework incorporating eight stylistic perturbations. The research demonstrates for the first time that probe fragility stems from narrow training distributions rather than inherent architectural limitations, thereby refuting prevailing hypotheses such as single-direction encoding, entropy-based proxies, and linear subspace assumptions. The proposed style-augmented probes achieve average AUROCs of 0.979–0.983 on unseen styles, while multidimensional probes (k≥5) successfully recover distributed weak signals, overturning the inverse scaling conjecture.
📝 Abstract
Linear probes trained on LLM activations are increasingly proposed as deception-detection metrics, yet report AUROC exceeding 0.96 on clean benchmarks while collapsing under distributional shift. This paper systematically pressure-tests probe-based metrics across the Gemma 3 model family (1B-27B parameters), diagnosing why they fail rather than merely documenting that they fail. We test four hypotheses about deception encoding: (1) single linear direction, (2) multi-dimensional subspace, (3) convex conic hull, (4) entropy proxy. Our design includes cross-domain transfer matrices, multi-dimensional probe analysis with permutation null baselines, entropy-residualization tests, and distractor evaluations across 8 stylistic shifts. We find that: (a) probes achieve near-perfect AUROC (>=0.998) on clean data but collapse under stylistic shifts; style-augmented probes recover near-perfect detection (mean AUROC 0.979-0.983) on unseen styles; (b) the single-direction hypothesis is rejected (k=1 captures only 0.61-0.80 AUROC), with cross-domain transfer failure confirmed as geometric rather than layer-mismatch-driven; (c) the entropy-proxy hypothesis is rejected (max |rho|=0.454, max Delta-AUROC after residualization=0.004); and (d) deception does not form a significant linear subspace (per-domain k*=0), yet multi-dimensional probes (k>=5) recover the signal through distributed sub-threshold features. Probe fragility reflects distributional narrowness rather than an architectural limitation: style-augmented probes recover near-perfect detection at both 4B and 27B, establishing that the inverse scaling pattern is a training-distribution artifact rather than a genuine scale-dependent phenomenon.
Problem

Research questions and friction points this paper is trying to address.

deception detection
distributional shift
linear probes
representation geometry
robustness
Innovation

Methods, ideas, or system contributions that make the work stand out.

deception detection
linear probes
distributional shift
representation geometry
style augmentation
🔎 Similar Papers
No similar papers found.
S
Sachin Kumar
LexisNexis, USA