Introducing OmniGEC: A Silver Multilingual Dataset for Grammatical Error Correction
Existing grammatical error correction (GEC) datasets suffer from severe multilingual undercoverage, hindering the development of non-English GEC models. To address this, we introduce OmniGEC—the first large-scale, silver-standard multilingual GEC dataset covering 11 languages. It integrates Wikipedia edit histories, Reddit user posts, and UberText 2.0, with high-quality automatic corrections generated by GPT-4o-mini. Unlike conventional sentence-level annotations, OmniGEC supports paragraph-level error correction modeling. Leveraging this dataset, we fine-tune Aya-Expanse and Gemma-3, achieving state-of-the-art performance on multilingual GEC benchmarks. All data and best-performing models are publicly released on Hugging Face, substantially alleviating the scarcity of non-English GEC resources and advancing multilingual grammatical error correction research.