Deep Learning Models Also Recall Features

📅 2026-08-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文探讨了深度学习模型如何通过特征回忆来检索存储信息,定义了特征回忆,并将其与特征组合进行了对比。
📝 Abstract
Recent work in mechanistic interpretability has studied how large language models recall facts stored in their weights. This paper argues that factual recall points to something broader: a general kind of operation in deep learning models, which I call feature recall. The core observation is that a linear projection can be read as retrieving stored information scaled by input activations. I define feature recall, show it applies across architectures, and contrast it with the established paradigm of feature combination. I also consider how cases of feature recall might be mechanistically identified. The account gives philosophers a new conceptual tool for understanding deep learning, and points to empirical directions for mechanistic interpretability research.
Problem

Research questions and friction points this paper is trying to address.

feature recall
deep learning models
mechanistic interpretability
Innovation

Methods, ideas, or system contributions that make the work stand out.

feature recall
linear projection
mechanistic interpretability