The Approximation Rank of Softmax Attention: Sharp Geometric Laws and Robust Interaction Dimension

📅 2026-08-28
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过分析几何结构对softmax注意力的近似秩复杂度的影响,提出两种最坏情况下的法则,并通过实验验证了有效维度减少与秩上界之间的正相关关系。
📝 Abstract
Which geometry controls the rank complexity of normalized softmax attention? We study maximum-row-$\ell_1$ approximation rank, exactly the least unrestricted rank preserving every bounded vector-valued output. Two sharp worst-case laws isolate support geometry: for fixed $d$ and error $\varepsilon$, spherical self-attention has rank $Θ_{d,\varepsilon}(\min\{n,(1+β)^{(d-1)/2}\})$, while full-ball geometry adds one radial degree and, for $β\geβ_0(d,\varepsilon)$ and $n\ge C_d e^{β/8}$, gives $Θ_{d,\varepsilon}(β^{d/2})$. For a fixed head, row-softmax quotients out row-scalar logit directions: the remaining visible query--key interaction dimension $r$ yields an $r/2$ per-instance upper law, and bounded constructions show this exponent is minimax sharp. Approximate interaction subspaces incur an explicit residual output error and yield a tolerance-indexed SVD dimension. On an 84-head BERT-base calibration set, we observe modest effective-dimension reductions across many head--temperature settings, together with positive associations with finite constructive rank upper certificates. Together, these results separate support geometry, which sets worst-case temperature scaling, from softmax-visible interaction geometry, which controls per-head approximation complexity.
Problem

Research questions and friction points this paper is trying to address.

softmax attention
approximation rank
geometry
rank complexity
support geometry
Innovation

Methods, ideas, or system contributions that make the work stand out.

approximation rank
softmax attention
geometric laws
interaction dimension
temperature scaling
🔎 Similar Papers
No similar papers found.