When Pruning Meets Interpretability: Preserving Sparse Autoencoder Robustness in LLMs

📅 2026-08-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究探讨了剪枝对稀疏自动编码器在大型语言模型中表现的影响,提出了一种基于激活感知的方法和分层稀疏分配策略以保持其功能稳定性。
📝 Abstract
Sparse autoencoders (SAEs) are widely used to interpret the internal representations of large language models (LLMs), yet their reliability under post-hoc model compression remains poorly understood. We present a systematic study of how pruning affects SAE behavior and theoretically show that, for a fixed SAE, its impact is governed by perturbation energy, a covariance-weighted norm. This perspective exposes a key limitation of magnitude pruning: by ignoring activation geometry, it distorts the learned representation space and degrades SAE functionality. Activation-aware methods such as Wanda and SparseGPT, in contrast, implicitly control perturbation energy and are therefore substantially more robust at preserving SAE behavior. We further reveal a consistent structural vulnerability across all pruning methods: middle layers are significantly more sensitive to pruning than early or late layers. Guided by this insight, we propose a layer-wise sparsity allocation strategy, achieving lower perplexity under the same average pruning sparsity. Experiments across four model architectures validate our theoretical findings. Code is publicly available at https://github.com/osu-srml/sae-robustness-under-pruning/tree/main.
Problem

Research questions and friction points this paper is trying to address.

pruning
sparse autoencoders
large language models
robustness
perturbation energy
Innovation

Methods, ideas, or system contributions that make the work stand out.

activation-aware pruning
perturbation energy
layer-wise sparsity allocation
sparse autoencoders (SAEs)
large language models (LLMs)
🔎 Similar Papers