π€ AI Summary
Linear probes are widely used for interpreting and evaluating neural representations, yet their reliability lacks theoretical grounding, often exhibiting abrupt performance transitions or outright failure. To address this, we propose the Spectral Identifiability Principle (SIP), the first theoretical framework linking probe stability to the spectral geometry of representations. SIP quantitatively relates spectral gaps in the representation covariance to the Fisher estimation error, thereby characterizing the phase-transition mechanism governing probe performance. It yields verifiable stability criteria that enable early detection of probe failureβgoing beyond conventional generalization bounds. Our method integrates finite-sample analysis, spectral graph theory, and the Fisher information framework. On synthetic data, we analytically compute and empirically validate SIP-predicted phase transitions. Experiments demonstrate that spectral analysis reliably identifies unstable probes, substantially enhancing the trustworthiness and interpretability of neural representation evaluation.
π Abstract
Linear probes are widely used to interpret and evaluate neural representations, yet their reliability remains unclear, as probes may appear accurate in some regimes but collapse unpredictably in others. We uncover a spectral mechanism behind this phenomenon and formalize it as the Spectral Identifiability Principle (SIP), a verifiable Fisher-inspired condition for probe stability. When the eigengap separating task-relevant directions is larger than the Fisher estimation error, the estimated subspace concentrates and accuracy remains consistent, whereas closing this gap induces instability in a phase-transition manner. Our analysis connects eigengap geometry, sample size, and misclassification risk through finite-sample reasoning, providing an interpretable diagnostic rather than a loose generalization bound. Controlled synthetic studies, where Fisher quantities are computed exactly, confirm these predictions and show how spectral inspection can anticipate unreliable probes before they distort downstream evaluation.