π€ AI Summary
This study addresses the limitations of existing cross-lingual reasoning benchmarks, which rely on translated English datasets and thus introduce linguistic bias while failing to assess modelsβ analogical reasoning within authentic cultural contexts. To overcome this, the authors propose the first language-agnostic, culturally grounded evaluation framework for analogical reasoning. They construct native, high-difficulty analogy datasets for Arabic, Amharic, and Japanese through a collaborative process involving native speakers and large language models, entirely bypassing translation. Experiments across 14 open-source models reveal a stark performance gap: despite strong results on English proverb-based tasks, model accuracy drops by 12β52 percentage points on the localized benchmarks, exposing significant deficiencies in cultural reasoning. The complete pipeline, datasets, and evaluation suite are publicly released.
π Abstract
Multilingual reasoning evaluation overwhelmingly relies on translating English benchmarks, a practice that introduces linguistic artifacts and fails to test culturally-grounded reasoning. We introduce ADAGE (Analogical Difficulty-by-design Assessment for Grounded Evaluation), a language-agnostic pipeline that combines native-speaker curation with LLM-assisted generation to construct challenging, translation-free benchmarks for abstract analogical reasoning. We validate ADAGE by constructing benchmarks for Arabic, Amharic, and Japanese. Evaluating 14 open-weight models, we find a consistent cultural reasoning gap: models that perform well on English proverb reasoning struggle substantially on all three native benchmarks, with accuracy dropping by 12--52 percentage points relative to English. We release the pipeline, all three benchmarks, and the full evaluation suite.