Why Are LLM Backdoor Defenses Fragmented? A Feature-Level Explanation with Sparse Autoencoders

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究使用稀疏自动编码器分析大型语言模型后门攻击的防御碎片化问题,揭示了不同后门特征的作用机制,并通过特征夹紧方法验证了其有效性。
📝 Abstract
Backdoor attacks pose a serious threat to large language models (LLMs), but existing defenses remain fragmented, failing to pro?vide unified defense against both dirty-label and clean-label attacks. To investigate why such fragmentation arises, we present the first systematic feature-level mechanistic analysis of LLM backdoors using sparse autoencoders (SAEs). Starting from a 2 x 2 comparison of clean and poisoned models on clean and triggered inputs, we trace backdoor-induced logit shifts to high-contributing SAE features and categorize them into four roles: interac?tion, suppressed, mixed, and weight-modified features. This taxonomy reveals system?atic encoding differences: dirty-label back?doors are dominated by isolated interaction features, whereas clean-label backdoors rely more on heterogeneous mixtures of mixed and weight-modified features. These differ?ences explain why existing defenses remain fragmented across attack paradigms. We val?idate this hypothesis through inference-time feature clamping, which reduces ASR to at most 10.8% in most dirty-label settings and at most 15.4% in the majority of clean-label settings, while preserving benign-task perfor?mance. These results show that SAE-based analysis can explain defense fragmentation and guide interpretable backdoor mitigation.
Problem

Research questions and friction points this paper is trying to address.

backdoor attacks
large language models
defense fragmentation
clean-label attacks
dirty-label attacks
Innovation

Methods, ideas, or system contributions that make the work stand out.

sparse autoencoders
feature-level analysis
backdoor defense
interaction features
weight-modified features
Y
Yizhe Zeng
Institute of Information Engineering, Chinese Academy of Sciences; School of Cyber Security, University of Chinese Academy of Sciences
C
Chenxu Niu
Institute of Information Engineering, Chinese Academy of Sciences; School of Cyber Security, University of Chinese Academy of Sciences
W
Wei Zhang
Beijing University of Posts and Telecommunications
H
Hao Huang
Institute of Information Engineering, Chinese Academy of Sciences; School of Cyber Security, University of Chinese Academy of Sciences
Yunpeng Li
Yunpeng Li
Institute of Information Engineering,Chinese Academy of Sciences
Large Language ModelsCyber Security
D
Dongxu Han
Institute of Information Engineering, Chinese Academy of Sciences; School of Cyber Security, University of Chinese Academy of Sciences
D
Dan Du
Institute of Information Engineering, Chinese Academy of Sciences
C
Cheng Hong
Ant Group
H
Hequn Xian
College of Computer Science and Technology, Qingdao University; Institute of Cryptography and Cyber Security (Whampoa)
Y
Yuling Liu
Institute of Information Engineering, Chinese Academy of Sciences; School of Cyber Security, University of Chinese Academy of Sciences