GRPO Beyond English: A Large-Scale Study of GRPO in Non-English and Multilingual Settings

📅 2026-08-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the overreliance on English and the scarcity of multilingual reasoning optimization research in Group Relative Policy Optimization (GRPO). By integrating verifiable reward reinforcement learning with GRPO, we conduct large-scale empirical evaluations across non-English and multilingual settings. This work provides the first systematic validation of non-English GRPO effectiveness, demonstrating significant cross-lingual transfer gains and native-language training performance comparable to English. Concurrently, it reveals potential out-of-domain capability regression risks under specific model-language configurations. Ultimately, this research bridges the gap in non-English GRPO literature, offering critical empirical evidence and safety considerations for optimizing multilingual reasoning capabilities.
📝 Abstract
Reinforcement Learning with Verifiable Rewards (RLVR), often optimized with Group Relative Policy Optimization (GRPO), has become a central recipe for improving the reasoning capabilities of pretrained language models but current studies remain heavily English-centric. We conduct a large-scale empirical study of multilingual and non-English GRPO across a wide range of base models, training languages, and different reasoning language rewards. We find that training to reason in the native language often leaves only a small gap to training for English reasoning. We further observe strong crosslingual transfer: training in one language often improves performance in many others. However, specific trends are highly model- and language-dependent. In some cases, training in a particular language induces severe regressions on out-of-domain capabilities in other languages. Our analysis shows that RLVR beyond English can provide broad crosslingual gains, but also requires broad evaluation to detect language-specific regressions.
Problem

Research questions and friction points this paper is trying to address.

GRPO
Multilingual Reasoning
Reinforcement Learning with Verifiable Rewards
Cross-lingual Transfer
Non-English LLMs
Innovation

Methods, ideas, or system contributions that make the work stand out.

GRPO
Multilingual RLVR
Cross-lingual Transfer
Non-English Reasoning
Language-specific Regression