A multilingual hallucination benchmark: MultiWikiQHalluA

📅 2026-05-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the critical gap in hallucination evaluation, which has been predominantly confined to English and thus fails to capture model fidelity issues in low-resource languages. Leveraging the multilingual MultiWikiQA dataset and the LettuceDetect framework, the authors construct the first large-scale synthetic hallucination benchmark spanning 306 languages and train token-level hallucination classifiers for 30 European languages to systematically assess cross-lingual hallucination behavior in mainstream large language models. Experimental results reveal that smaller models, such as Qwen3-0.6B, exhibit hallucination rates as high as 60% in low-resource languages like Icelandic, while larger models generally perform better. Crucially, hallucination rates are consistently higher across low-resource languages, underscoring both the necessity and the challenges of multilingual hallucination evaluation.
📝 Abstract
Most hallucination evaluations focus on English, leaving it unclear whether findings transfer to lower-resource languages. We investigate faithfulness hallucinations, defined as model-generated content that is fluent and plausible but diverges from the provided input or is internally inconsistent. Leveraging the multilingual MultiWikiQA dataset, we utilize the LettuceDetect framework to create synthetic hallucination datasets for 306 languages, from which we train token-level hallucination classifiers for 30 European languages. In this work, we present evaluations of model hallucinations on a selection of languages: English, Danish, German, and Icelandic. Using these classifiers, we evaluate the hallucination rates for Qwen3-0.6B, Qwen3-14B, Gemma-3-12B-IT, cogito-v1-preview-qwen-32B, and cogito-v1-preview-llama-70B. Our classifiers reveal notably higher hallucination rates for Qwen3-0.6B (up to 60\% of answers containing at least one hallucination, peaking in Icelandic) and generally lower rates for larger models, with cogito-v1-preview-qwen-32B and cogito-v1-preview-llama-70B performing best on most languages. Hallucination rates are consistently higher for lower-resource languages, particularly Icelandic.
Problem

Research questions and friction points this paper is trying to address.

hallucination
multilingual
low-resource languages
faithfulness
language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

multilingual hallucination benchmark
faithfulness hallucination
token-level classifier
low-resource languages
LettuceDetect
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
F
Freja Thoresen
Alexandra Institute
D
Dan Saattrup Smart
Alexandra Institute