Influence Malleability in Linearized Attention: Dual Implications of Non-Convergent NTK Dynamics

📅 2026-03-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work uncovers the fundamental reason why linearized attention mechanisms fail to converge within the Neural Tangent Kernel (NTK) framework and elucidates their dual impact on model performance and robustness. By constructing a linearized attention operator that exactly corresponds to the data-dependent Gram kernel, and integrating NTK theory with spectral analysis, the study establishes—for the first time—a theoretical link between the cubic amplification of the Gram matrix condition number and the required network width (m = Ω(κ⁶)). It introduces the notion of “influence plasticity” to unify the explanation of attention’s expressive power and its vulnerability to adversarial perturbations. Empirical results show that practical training widths fall far below the theoretical threshold, yet exhibit 6–9 times higher influence plasticity than ReLU networks, conferring superior task adaptability at the cost of heightened sensitivity to data poisoning.

Technology Category

Application Category

📝 Abstract
Understanding the theoretical foundations of attention mechanisms remains challenging due to their complex, non-linear dynamics. This work reveals a fundamental trade-off in the learning dynamics of linearized attention. Using a linearized attention mechanism with exact correspondence to a data-dependent Gram-induced kernel, both empirical and theoretical analysis through the Neural Tangent Kernel (NTK) framework shows that linearized attention does not converge to its infinite-width NTK limit, even at large widths. A spectral amplification result establishes this formally: the attention transformation cubes the Gram matrix's condition number, requiring width $m = Ω(κ^6)$ for convergence, a threshold that exceeds any practical width for natural image datasets. This non-convergence is characterized through influence malleability, the capacity to dynamically alter reliance on training examples. Attention exhibits 6--9$\times$ higher malleability than ReLU networks, with dual implications: its data-dependent kernel can reduce approximation error by aligning with task structure, but this same sensitivity increases susceptibility to adversarial manipulation of training data. These findings suggest that attention's power and vulnerability share a common origin in its departure from the kernel regime.
Problem

Research questions and friction points this paper is trying to address.

linearized attention
Neural Tangent Kernel
influence malleability
non-convergence
Gram matrix
Innovation

Methods, ideas, or system contributions that make the work stand out.

linearized attention
Neural Tangent Kernel
influence malleability
non-convergent dynamics
Gram matrix condition number
🔎 Similar Papers
No similar papers found.