Trading Complexity for Expressivity Through Structured Generalized Linear Token Mixing
This work addresses the challenge in autoregressive generation of simultaneously achieving fast decoding, low memory overhead, and effective long-range dependency modeling. To this end, the authors propose a unified structured generalized linear recurrence framework that decouples the direct single-step input–output influence from the multi-step state propagation mechanism. By incorporating recurrence designs that depend on multiple historical states, the framework substantially enhances model expressivity while maintaining controllable computational complexity. This formulation generalizes both state space models and attention mechanisms, establishing a unified token mixing paradigm. Empirical validation on synthetic tasks and language modeling benchmarks demonstrates the framework’s efficiency and strong representational capacity, offering a novel toolkit for designing high-performance token mixers.