Length Generalization Bounds for Transformers
This study investigates the length generalization capability of Transformers—specifically, their ability to generalize to arbitrarily long inputs when trained on sequences of bounded length—and the computability of associated generalization bounds. Leveraging formal language theory and computational complexity analysis, and employing C-RASP, a mathematically rigorous abstraction of the Transformer architecture, the work establishes for the first time that C-RASP models with two or more layers admit no computable length generalization bound. In contrast, both the positive fragment of C-RASP and fixed-precision Transformers possess tight, optimal, and computable exponential generalization bounds. These results uncover fundamental limitations inherent in deep Transformer models regarding length generalization while providing an exact characterization of the conditions under which such bounds remain computable.