A multilingual hallucination benchmark: MultiWikiQHalluA
This study addresses the critical gap in hallucination evaluation, which has been predominantly confined to English and thus fails to capture model fidelity issues in low-resource languages. Leveraging the multilingual MultiWikiQA dataset and the LettuceDetect framework, the authors construct the first large-scale synthetic hallucination benchmark spanning 306 languages and train token-level hallucination classifiers for 30 European languages to systematically assess cross-lingual hallucination behavior in mainstream large language models. Experimental results reveal that smaller models, such as Qwen3-0.6B, exhibit hallucination rates as high as 60% in low-resource languages like Icelandic, while larger models generally perform better. Crucially, hallucination rates are consistently higher across low-resource languages, underscoring both the necessity and the challenges of multilingual hallucination evaluation.