Measuring the Cross-Lingual Comprehension Gap: How the language of the evidence shapes what language models understand

📅 2026-08-06
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of isolating linguistic variables in evaluating large language models, which has obscured whether non-English comprehension genuinely lags behind English performance. The authors introduce and quantify the “Cross-Linguistic Comprehension Gap” (CLCG) through a rigorously controlled parallel evaluation framework where content is held constant across languages. Leveraging ParallelQA-18—a human-translated dataset spanning 18 languages—and combining token-level F1 micro-averaging, passage clustering with bootstrapping, and blind human preference trials, they find an overall CLCG of 0.078, corresponding to an approximate 17% performance drop. Crucially, CLCG exhibits a significant negative correlation with language resource availability: responses in high-resource languages are consistently preferred by human evaluators, revealing a systematic overestimation of model capabilities for low-resource language users under current English-centric evaluation paradigms.
📝 Abstract
Language models are often evaluated as though capabilities demonstrated in English remain equally available when the same content is presented in other languages. Traditional multilingual benchmarks rarely isolate language while holding content, question, reference answer, model, and evaluation unit constant. We define the Cross-Lingual Comprehension Gap (CLCG) as the reduction in response quality when the same content and question are presented in a target language rather than in English. Using ParallelQA-18, a professionally human-translated parallel corpus, we evaluate five models from five laboratories on a stratified sample of 150 articles across 18 languages (English reference; Portuguese high-resource baseline; 16 targets spanning Joshi et al. 2020 classes 0-4). A within-item design varies only passage language. The primary estimator contrasts English versus pooled target-language Token-F1 micro-means on higher-complexity open-ended questions, with article-cluster bootstrap intervals. The primary pooled CLCG is 0.078 (95% CI 0.072-0.084), about a 17% reduction relative to the English score; the equal-language macro summary is 0.077. Net of Portuguese, the macro gap is 0.016 (95% CI 0.013-0.020). Language-level CLCG is negatively associated with Joshi resource class (rho = -0.594, p = 0.015, n = 16). In blinded paired human evaluations, higher-resource responses are preferred in 61.6% of decisive judgments (estimated preference probability 0.655, 95% CI 0.558-0.741). Capabilities shown in English should not be assumed to transfer equally to other languages; English-centered evaluations may overestimate quality for users of low-resource languages.
Problem

Research questions and friction points this paper is trying to address.

Cross-Lingual Comprehension Gap
multilingual evaluation
language models
low-resource languages
language bias
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cross-Lingual Comprehension Gap
ParallelQA-18
multilingual evaluation
language resource bias
within-item design
💼 Related Jobs
No related jobs found.
R
Rafael da Silva
PhD in Applied Data Science, Eastern University
J
Jeff Eicher
PhD in Applied Data Science, Eastern University