Intrinsic Structure: Spectral Identifiability for Mechanistic Interpretability

📅 2026-08-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge in mechanistic interpretability that existing methods struggle to disentangle intrinsic model structure from methodological artifacts, lacking theoretical identifiability guarantees. The authors model neural network forward propagation as a depth-indexed controlled dynamical system and lift it—via the Koopman operator—to a finite-dimensional linear system, whose spectrum serves as a coordinate-invariant intrinsic property for analysis. They establish, for the first time, an identifiability theory for fundamental units of mechanistic interpretability, introducing spectral convergence rates, a median-of-means variant tailored to heavy-tailed activations, and a theorem decoupling information and variance directions in non-normal systems. Experiments on GPT-2 Small, Gemma-2-2B, and Qwen3-8B-Base confirm spectral convergence (empirical exponent 0.506 ± 0.031 for Qwen3), demonstrate that Koopman modes significantly outperform random directions, and reveal a 4.1× depth-dependent decay in the gap between Koopman modes and principal components.
📝 Abstract
Mechanistic interpretability explains models by identifying circuits inside them, but has no way to tell whether a circuit is a property of the model or an artifact of the method that found it. Sparse autoencoders illustrate the problem: different seeds and widths recover materially different features from the same activations, and no theory says whether that variability is incidental or structural. We put dictionary learning for interpretability on an identifiability footing. Treating the forward pass as a controlled dynamical system with depth as time and lifting it with the Koopman operator yields a finite linear realisation whose \emph{spectrum} is a coordinate-free property of the model. We prove the spectrum is recoverable from $M$ calibration samples at rate $M^{-1/2}$ up to permutation - to our knowledge the first identifiability theorem for a mechanistic-interpretability primitive, with a matching minimax lower bound, a median-of-means variant for heavy-tailed activations, and a dissociation theorem: whenever the realisation is non-normal, the directions carrying activation variance and the directions carrying information across depth cannot coincide. The identifiable object and the legible object are not the same object. On GPT-2 small, Gemma-2-2B and Qwen3-8B-Base the spectrum converges everywhere and attains the predicted exponent on Qwen3-8B-Base ($0.506 \pm 0.031$); shortfalls collapse onto one curve against each cell's sample threshold. Koopman modes beat random directions but lose to principal components on indirect-object identification, with the gap decaying $4.1\times$ in depth-distance, as the theorem predicts. The Koopman spectrum is an identifiable, model-intrinsic fingerprint with a stated error bar, not a legible decomposition.
Problem

Research questions and friction points this paper is trying to address.

mechanistic interpretability
identifiability
dictionary learning
model-intrinsic structure
spectral properties
Innovation

Methods, ideas, or system contributions that make the work stand out.

Koopman spectrum
mechanistic interpretability
identifiability
dictionary learning
dynamical systems