One Form to Transfer Them All: Pretraining Multilingual Language Models Beyond Native Orthography

📅 2026-08-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文探讨了通过罗马化预训练方法解决不同书写系统间多语言模型知识迁移问题,实验表明该方法优于正字法和IPA。
📝 Abstract
Multilingual language models transfer knowledge across languages through shared subword vocabulary, a mechanism that breaks down when related languages use different writing systems. Prior work addresses this via script equalization (romanization or IPA transcription), but direct comparisons are rare; the focus has been on encoder-only models, with most work adapting existing pretrained models. We systematically compare different input representations in autoregressive multilingual pretraining, comparing orthographic text, IPA, and romanization in a controlled setup across three scales (467M, 709M, and 1.03B) on eight languages in four typologically motivated pairs. Across a wide range of downstream tasks on seen and unseen languages, romanized pretraining yields the strongest cross-lingual transfer, and the advantage over text widens with scale. IPA improves over text in most settings but trails romanization. Surprisingly, finetuning a text-pretrained model on romanized data hurts performance on languages already covered by the base model, only marginally helping when the model lacks script coverage. Our results indicate that for multilingual models spanning typologically diverse scripts, to obtain maximum benefits, romanization should be treated as a core design choice applied at pretraining rather than a post hoc fix.
Problem

Research questions and friction points this paper is trying to address.

Multilingual Language Models
Cross-lingual Transfer
Orthography
Script Differences
Pretraining
Innovation

Methods, ideas, or system contributions that make the work stand out.

romanization
cross-lingual transfer
multilingual pretraining
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Muge Zhang
Department of Computer Science and Engineering, Ohio State University
A
Aaron Jencks
Department of Computer Science and Engineering, Ohio State University
K
Krishna Badikela
Department of Computer Science and Engineering, Ohio State University
Yulia Tsvetkov
Yulia Tsvetkov
University of Washington
Natural Language Processing
Sachin Kumar
Sachin Kumar
The Ohio State University
Natural Language ProcessingMachine Learning