Measuring Semantic Abstractness of SAE Features via Nonlocality

📅 2026-08-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work proposes a novel metric, Feature Non-Locality (FNL), to distinguish shallow lexical features from high-level semantic features in sparse autoencoders (SAEs). FNL quantifies the degree of semantic abstraction by computing the normalized entropy of feature activations across sequence positions. Notably, FNL requires no labels and does not rely on large language models, offering the first unsupervised criterion for hierarchically ordering interpretable features in mechanistic interpretability research. Experiments demonstrate that FNL correctly identifies context-reasoning features in 73%–84% of tested feature pairs. Furthermore, interventions targeting high-FNL features improve performance on MATH-500 by 4.6 accuracy points, while analysis reveals that existing jailbreak mitigation strategies predominantly exploit low-FNL, position-dependent features.
📝 Abstract
Sparse autoencoders (SAEs) have helped uncover mechanistic explanations for LLM behaviours such as reasoning, jailbreaking etc., via understanding the corresponding task-relevant and causally effective features. To evaluate such mechanistic explanations, downstream studies must distinguish surface lexical features from genuinely high-level ones. However, neither an autointerp-based semantic description nor causal steering utility fully resolves the abstraction level of a feature. To this end, we introduce \emph{Feature Nonlocality} (FNL), defined as the entropy of the normalized per-position influence on an SAE feature's activation. We report that FNL correlates with existing LLM-based proxy metrics of feature semantic abstractness, and successfully distinguishes context-dependent reasoning features from token-driven ones, correctly assigning the higher FNL to the contextual feature in $73$--$84\%$ of randomly drawn pairs that consist of one contextual and one token-level feature. We demonstrate two downstream applications. We audit SAE-based features used for jailbreak mitigation and find surprisingly that most effective features are positional features with low FNL rather than genuinely recognizing harmful intents. We report that steering high-FNL features in DeepSeek-R1-Distill-Llama-8B improves MATH-500 accuracy by $4.6$ points over the unsteered model and outperforms steering low-FNL features, though the gains are model-specific. We conclude that FNL provides an LLM-independent, label-free, correlational witness of the abstraction level of an SAE feature, with applications in evaluating mechanistic explanations as well as selecting features for downstream interventions.
Problem

Research questions and friction points this paper is trying to address.

Semantic Abstractness
Sparse Autoencoders
Feature Abstraction
Mechanistic Interpretability
Nonlocality
Innovation

Methods, ideas, or system contributions that make the work stand out.

Feature Nonlocality
Sparse Autoencoders
Semantic Abstractness
Mechanistic Interpretability
Causal Steering
Chuqiao Lin
Chuqiao Lin
PhD student, University of Oxford
condensed matter physicsquantum informationmechanistic interpretability
S
Shivaji Sondhi
Rudolf Peierls Centre for Theoretical Physics, University of Oxford
X
Xiao-Liang Qi
Leinweber Institute for Theoretical Physics, Stanford University