🤖 AI Summary
This study addresses the overreliance on English and the scarcity of multilingual reasoning optimization research in Group Relative Policy Optimization (GRPO). By integrating verifiable reward reinforcement learning with GRPO, we conduct large-scale empirical evaluations across non-English and multilingual settings. This work provides the first systematic validation of non-English GRPO effectiveness, demonstrating significant cross-lingual transfer gains and native-language training performance comparable to English. Concurrently, it reveals potential out-of-domain capability regression risks under specific model-language configurations. Ultimately, this research bridges the gap in non-English GRPO literature, offering critical empirical evidence and safety considerations for optimizing multilingual reasoning capabilities.
📝 Abstract
Reinforcement Learning with Verifiable Rewards (RLVR), often optimized with Group Relative Policy Optimization (GRPO), has become a central recipe for improving the reasoning capabilities of pretrained language models but current studies remain heavily English-centric. We conduct a large-scale empirical study of multilingual and non-English GRPO across a wide range of base models, training languages, and different reasoning language rewards. We find that training to reason in the native language often leaves only a small gap to training for English reasoning. We further observe strong crosslingual transfer: training in one language often improves performance in many others. However, specific trends are highly model- and language-dependent. In some cases, training in a particular language induces severe regressions on out-of-domain capabilities in other languages. Our analysis shows that RLVR beyond English can provide broad crosslingual gains, but also requires broad evaluation to detect language-specific regressions.