🤖 AI Summary
研究探讨了基于Transformer的语言模型学习k-反局部语言的能力,发现随着k增大,模型收敛速度变慢,表明非局部依赖更难学习。
📝 Abstract
We consider the ability of transformer-based language models (LLMs) to learn what we call k-antilocal languages, i.e., languages that have no mutual information across any span of $k$ contiguous symbols. We construct such languages with increasing $k$, finding that LLMs trained on them achieve comparable cross-entropy loss regardless of antilocality, but converge more slowly on more antilocal languages. Our findings support the idea that non-local dependencies are more difficult to learn, but the evidence for this bias comes from learning speed rather than learning success.