MultiSynt/MT: Trillion-Token Multi-Parallel Pre-Training Data Translated Across 36 Languages
This work addresses the scarcity of non-English open-source pretraining corpora, which hinders the development of multilingual large language models. To overcome this limitation, the authors propose a machine translation–based synthesis approach that leverages the high-quality Nemotron-CC corpus and employs both Tower+ and OPUS-MT/HPLT-MT systems to generate a sentence-aligned parallel corpus spanning 36 European languages and approximately 4.8 trillion tokens—the first large-scale, open, multi-system-fused multilingual pretraining dataset of its kind. Experimental results demonstrate that, under a fixed budget of 100 billion tokens, models trained on this synthetic data achieve a ~15% performance gain over the native-data baseline HPLT 2.0; moreover, they reach equivalent final performance using only 72% of the token budget, substantially reducing reliance on scarce native multilingual data and exposing limitations in current evaluation benchmarks.