Three Tokens Force Exponential Feature Rank in Nonnegative Kernel Attention
This work establishes fundamental limitations of non-negative kernel attention mechanisms, showing that their feature dimension must grow exponentially even to handle contexts as short as three tokens. Focusing on the exact modeling of Boolean inputs under the Min-IP task, the authors employ rank analysis, information-theoretic lower bounds, and constructive counterexamples within a framework of causal queries and position-dependent mappings. Their key contribution is the first proof that standard Softmax attention solves this task with only linear feature dimensionality, whereas non-negative kernel attention heads require at least $2^{\Omega(m)}$ dimensions. Furthermore, they derive an information transmission lower bound for multi-head, multi-layer models operating over finite alphabets.