Cross-lingual Self-Consistency for Multilingual Reasoning with Language Models
This work addresses the limited multilingual reasoning capabilities of large language models on low-resource and unseen languages, as well as the absence of effective methods that operate without labeled or parallel data. The authors propose an unsupervised reinforcement learning framework that enhances reasoning performance by enforcing cross-lingual self-consistency—requiring the model to produce consistent answers to semantically equivalent questions across languages—without relying on gold labels or multilingual alignment data. This approach substantially improves generalization to both unseen languages and out-of-distribution tasks, achieving an average accuracy gain of 21.7% across the ten languages in the MGSM benchmark, with an 18.2% improvement specifically on unseen languages, and up to a 6.2% increase on three out-of-distribution evaluation benchmarks.