Building Better Deception Probes Using Targeted Instruction Pairs
This work addresses the vulnerability of existing linear probes to spurious correlations when detecting AI deception, which often leads to false positives on non-deceptive responses. To mitigate this, the authors propose a targeted instruction-pair design method grounded in a taxonomy of deceptive behaviors. By constructing interpretable instruction pairs that isolate specific deception types, the approach trains linear probes to focus on deceptive intent rather than superficial content patterns. Experimental results demonstrate that instruction selection is the dominant factor in probe performance, accounting for 70.6% of variance. Probes tailored to specific threat models significantly outperform general-purpose detectors, achieving higher detection accuracy and lower false-positive rates on evaluation datasets. This study underscores the importance of aligning probing mechanisms with concrete deception categories and establishes a new paradigm for interpretable AI safety evaluation.