Measure-to-measure interpolation using Transformers
This paper investigates the expressive power of Transformers as arbitrary input-to-output measure mappings. Method: We reformulate Transformers from a measure-theoretic perspective, modeling them as differentiable maps on continuous measure spaces—departing from conventional discrete token-based interpretations. Leveraging continuity equations to describe particle dynamics, we design an attention mechanism incorporating spherical geometry constraints and optimal transport theory. Contribution/Results: We propose the first Transformer architecture provably capable of exact matching between arbitrary input–target measure pairs. Under the minimal assumption that a transport map exists between each pair, a single model achieves precise matching for N arbitrary measure pairs. We establish theoretical completeness by proving that Transformers serve as universal interpolators between measures and provide explicit parameter constructions. This work fundamentally characterizes the expressive capacity of Transformers for measure transformation tasks.