A Unifying Perspective on Language Model Representations: From Filler-Role Structure to Mechanistic Interpretability

📅 2026-08-29
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文使用张量积表示法(TPRs)统一解释语言模型的内部结构,通过数学和实证方法整合多种现有解释技术,旨在提供神经网络性质的一致性理解。
📝 Abstract
A wide range of methods have been proposed for interpreting language models, delivering important insights into their inner workings. However, different methods and their resulting insights stand in relative isolation: what could the underlying structure of language models be, such that they give rise to all our interpretations? In this work, we propose using Tensor Product Representations (TPRs) as a unifying hypothesis. TPRs give a concrete proposal for how compositional structure could be represented in vector space --- as filler-role bindings. We show, both mathematically and empirically, that TPRs can unify several prior interpretability methods: additive analogies, linear probing, sparse autoencoders, and activation patching. Mathematically, we show that these methods can all be derived from TPRs. Empirically, we apply the derivations to a range of different models --- from small toy models to LLMs --- to construct instances of each of the above interpretability methods; these constructed variants perform comparably to their standard variants. We view this work as a step toward what interpretability will ideally provide: a unified account of the nature of neural networks, corroborated not just by individual observations but also by an explanation of the connections between them.
Problem

Research questions and friction points this paper is trying to address.

language models
interpretability methods
Tensor Product Representations
compositional structure
Innovation

Methods, ideas, or system contributions that make the work stand out.

Tensor Product Representations
mechanistic interpretability
filler-role bindings
🔎 Similar Papers