UrduBench: An Urdu Reasoning Benchmark using Contextually Ensembled Translations with Human-in-the-Loop
This work addresses the absence of standardized reasoning benchmarks for low-resource languages like Urdu and the inability of existing machine translation approaches to preserve contextual and structural integrity in reasoning tasks. The authors propose a context-integrated translation framework that combines outputs from multiple translation systems with human validation to construct UrduBench—the first high-quality Urdu reasoning benchmark spanning multiple difficulty levels and task types, including MGSM and MATH-500. Using this benchmark, they systematically evaluate various large language models under diverse prompting strategies, revealing significant performance degradation in multi-step and symbolic reasoning. Their findings underscore the critical role of linguistic consistency in enabling robust cross-lingual reasoning and establish a scalable paradigm for evaluating reasoning capabilities in low-resource languages.