Transcoders for Investigating Deception in Language Models

📅 2026-07-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the critical security risks posed by deceptive behaviors in large language models and proposes an interpretable mechanism for their detection and intervention. For the first time, per-layer transcoders (PLTs) are applied to conduct circuit-level analysis of deception in large models. By constructing attribution graphs for Qwen3-4B, manipulating internal features, and tracing their downstream effects on model outputs, the work reveals that deceptive behavior is driven by specific internal mechanisms. The research successfully identifies a set of high-impact features associated with deception and compiles the first such feature dictionary. These findings demonstrate the effectiveness and potential of transcoders in behavioral monitoring and early detection of malicious intent within language models.
📝 Abstract
Transcoders have recently emerged as a promising approach for mechanistic interpretability (MI), enabling circuit-level analysis of model behaviour. In this paper, we investigate the use of transcoders to analyse deceptive behaviour in language models, a behaviour that poses a safety and security risk. Using a Qwen3-4B model with pre-trained transcoders, specifically per-layer transcoders (PLTs), we construct attribution graphs that capture feature activations and inter-feature dependencies, allowing circuit-level analysis of deception. Through feature steering and circuit analysis, we identified a dictionary of deception-related features and show that these features exert a stronger influence on deceptive outputs, as they produce predictable shifts between deceptive and non-deceptive responses. These findings suggest that deception emerges from internal model mechanisms and highlight the potential of transcoders for behavioural monitoring and early detection of security vulnerabilities related to malicious behaviours in language models.
Problem

Research questions and friction points this paper is trying to address.

deception
language models
mechanistic interpretability
safety
security
Innovation

Methods, ideas, or system contributions that make the work stand out.

transcoders
mechanistic interpretability
deception detection
feature steering
attribution graphs
🔎 Similar Papers
No similar papers found.
D
Darius Lim
Home Team Science & Technology Agency (HTX), Singapore
N
Nathan Leow
Home Team Science & Technology Agency (HTX), Singapore
X
Xin Wei Chia
Home Team Science & Technology Agency (HTX), Singapore