MultiWikiQA: A Reading Comprehension Benchmark in 300+ Languages

📅 2025-09-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work introduces XRC, the first large-scale multilingual reading comprehension benchmark covering 306 languages, designed to systematically evaluate cross-lingual understanding capabilities of language models. Methodologically, it automatically generates question-answer pairs from Wikipedia text, leveraging large language models for question generation and employing rigorous human crowd-sourced validation—assessing question fluency and answer accuracy—across 30 representative languages. Key contributions include: (1) establishing the first reading comprehension benchmark spanning over 300 languages; (2) revealing a substantial performance gap—up to 70 percentage points—between high-resource and low-resource languages in mainstream multilingual models; and (3) open-sourcing the complete dataset, annotation protocols, and evaluation code, thereby significantly advancing NLP research for low-resource languages.

Technology Category

Application Category

📝 Abstract
We introduce a new reading comprehension dataset, dubbed MultiWikiQA, which covers 306 languages. The context data comes from Wikipedia articles, with questions generated by an LLM and the answers appearing verbatim in the Wikipedia articles. We conduct a crowdsourced human evaluation of the fluency of the generated questions across 30 of the languages, providing evidence that the questions are of good quality. We evaluate 6 different language models, both decoder and encoder models of varying sizes, showing that the benchmark is sufficiently difficult and that there is a large performance discrepancy amongst the languages. The dataset and survey evaluations are freely available.
Problem

Research questions and friction points this paper is trying to address.

Creating a multilingual reading comprehension dataset covering 306 languages
Evaluating question quality through human assessment across 30 languages
Assessing performance disparities among language models on diverse languages
Innovation

Methods, ideas, or system contributions that make the work stand out.

Uses Wikipedia articles as context data
Generates questions via an LLM model
Evaluates multiple language model types
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
D
Dan Saattrup Smart
Alexandra Institute