🤖 AI Summary
研究通过将Softmax注意力分解为几何、势能和流通三个部分,揭示其与扩散映射的联系,并在预训练模型上验证了该方法的有效性。
📝 Abstract
Transformers, diffusion-maps, and magnetic Laplacians are usually treated as separate tools; we show they are all different regimes of a single Markov geometry built from pre-softmax query-scores. We define a QK"bidivergence"whose exponentiated and normalized forms yield attention, diffusion-maps, and magnetic diffusion. And use product of experts and Schr\"odinger-bridges to connect and organize them into equilibrium, nonequilibrium steady-state, and driven dynamics.