Feature Superposition in Neural Networks: From Theory to Practice

📅 2026-09-06
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
该论文探讨了神经网络中的特征叠加问题,通过理论和实践方法分析了特征叠加的几何、学习和计算,并评估了从训练网络中恢复和分析特征的方法。
📝 Abstract
Superposition refers to neural networks representing more features than they have dimensions. It offers a possible explanation for polysemantic neurons and motivates methods for recovering interpretable features from neural activations. Theoretical models typically start with a given set of input features and assumptions about how their values vary across inputs, then study how a network encodes those values in a lower-dimensional hidden representation. Empirical work, by contrast, seeks to identify the features encoded in trained networks and determine their role in computation. In this survey, we review the geometry, learning, and computation of superposed representations, explaining how feature statistics and decoder choice affect the conclusions. To connect these theoretical accounts with evidence from trained networks, we compare practical methods for recovering and analyzing features and examine what their evaluations establish. Since accurate activation reconstruction alone does not establish feature identity or causal use, we discuss the methods'documented failures and applications in light of the evidence available for these different claims. Finally, we assess previously stated open problems and identify remaining theoretical and empirical questions about superposition in trained networks. We hope our work can pave the way for a deeper understanding of superposition and more reliable methods for interpreting neural networks.
Problem

Research questions and friction points this paper is trying to address.

Superposition
Neural Networks
Polysemantic Neurons
Feature Recovery
Interpretable Features
Innovation

Methods, ideas, or system contributions that make the work stand out.

Feature Superposition
Neural Networks
Interpretable Features
Decoder Choice
Activation Reconstruction
🔎 Similar Papers
No similar papers found.