🤖 AI Summary
This study investigates the origins of word order preferences in large language models and their implications for linguistic diversity. Leveraging decoder-based models across 192 languages, we systematically compare synthetic and natural language experiments. Results indicate that word order biases stem from data distribution rather than inherent architectural constraints, with high-resource SVO language dominance correlating positively with data scale. Crucially, the model exhibits opposing preferences for synthetic versus natural languages, revealing a risk of word order homogenization driven by training data imbalance. These findings warn that the proliferation of large language models may erode global word order diversity, providing critical empirical evidence for advancing fairness research in multilingual modeling.
📝 Abstract
We systematically compare word order preferences in decoder-only language models across 192 artificial languages and typologically diverse natural languages. On artificial languages, models exhibit a left-branching preference that aligns with neither natural language universals nor human word order learning biases. On natural languages, monolingual models show no clear base word order bias at small scales, but as data grows, a preference for right-branching subject-verb-object (SVO) languages emerges while SOV falls behind despite being the most frequent order cross-linguistically. This SVO advantage extends to multilingual models and correlates with language resource level and data quality rather than word order. Thus, the same architecture exhibits opposite preferences on artificial and natural languages, establishing that word order biases observed in practice are data-driven. Since highly-resourced languages are overwhelmingly SVO, these biases risk gradually reducing word order diversity, particularly in languages that productively use multiple word orders, with the widespread adoption of LLMs.