Representation of syntax in LLMs through the lens of linear distance and similarity-aware entropy

📅 2026-08-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过线性距离和相似性感知熵分析了大型语言模型中的句法表示问题,使用结构探针方法评估不同句法关系的重建准确性。
📝 Abstract
Structural probes were introduced by Hewitt and Manning to reconstruct syntactic trees from a neural language model's latent representations. They are evaluated by calculating the proportion of syntactic tree edges correctly reconstructed over an annotated corpus (as measured by undirected unlabeled attachment score). Here, we disaggregate this measure, considering undirected attachment score by label (UASL), which assesses the reconstruction accuracy of each syntactic relation separately, establishing important differences among relations that overlap linguistic distinctions. Moreover, we identify two factors that predict most of UASL's variability across relations: (i) the mean and dispersion of the linear distance (on a log scale) between the related words, and (ii) the diversity (similarity-aware entropy) of the syntactic relation's head. These results, which hold across a range of model sizes and architectures, shed light on the degree of abstraction of the representation of syntax in language models and the dependence of such representation on geometric properties of the embedding space.
Problem

Research questions and friction points this paper is trying to address.

syntax
linear distance
similarity-aware entropy
structural probes
UASL
Innovation

Methods, ideas, or system contributions that make the work stand out.

linear distance
similarity-aware entropy
undirected attachment score by label (UASL)
syntactic representation
🔎 Similar Papers
No similar papers found.