Institution profile

Alexandra Institute

Academic institutioneurope · dk
Official website
Research library5linked papers
Opportunities0open roles
Selected work

Representative Papers

A multilingual hallucination benchmark: MultiWikiQHalluA

May 04, 2026

This study addresses the critical gap in hallucination evaluation, which has been predominantly confined to English and thus fails to capture model fidelity issues in low-resource languages. Leveraging the multilingual MultiWikiQA dataset and the LettuceDetect framework, the authors construct the first large-scale synthetic hallucination benchmark spanning 306 languages and train token-level hallucination classifiers for 30 European languages to systematically assess cross-lingual hallucination behavior in mainstream large language models. Experimental results reveal that smaller models, such as Qwen3-0.6B, exhibit hallucination rates as high as 60% in low-resource languages like Icelandic, while larger models generally perform better. Crucially, hallucination rates are consistently higher across low-resource languages, underscoring both the necessity and the challenges of multilingual hallucination evaluation.

0 citationsRead paper

MultiZebraLogic: A Multilingual Logical Reasoning Benchmark

Nov 05, 2025

Existing logical reasoning benchmarks lack multilingual coverage and controllable difficulty. Method: We propose MultiZebraLogic—the first multilingual logical reasoning benchmark targeting the Germanic language family—employing 14 clue types and 8 distractor categories, integrated with rule-based automatic generation, multilingual text synthesis, tunable difficulty control, and red herring injection across nine Germanic languages. Contribution/Results: We construct a high-quality dataset comprising 128 (core) + 1,024 (extended) problems per language; enable cross-lingual fair evaluation and scalable topic/difficulty customization; and empirically demonstrate that 4×5 logic puzzles pose significant challenges to state-of-the-art LLMs, with distractor clues reducing average accuracy by 23.6%, thereby validating the benchmark’s sensitivity and effectiveness.

0 citationsRead paper

MultiWikiQA: A Reading Comprehension Benchmark in 300+ Languages

Sep 04, 2025

This work introduces XRC, the first large-scale multilingual reading comprehension benchmark covering 306 languages, designed to systematically evaluate cross-lingual understanding capabilities of language models. Methodologically, it automatically generates question-answer pairs from Wikipedia text, leveraging large language models for question generation and employing rigorous human crowd-sourced validation—assessing question fluency and answer accuracy—across 30 representative languages. Key contributions include: (1) establishing the first reading comprehension benchmark spanning over 300 languages; (2) revealing a substantial performance gap—up to 70 percentage points—between high-resource and low-resource languages in mainstream multilingual models; and (3) open-sourcing the complete dataset, annotation protocols, and evaluation code, thereby significantly advancing NLP research for low-resource languages.

0 citationsRead paper

Hotter and Colder: A New Approach to Annotating Sentiment, Emotions, and Bias in Icelandic Blog Comments

Feb 24, 2025

This work addresses the critical lack of high-quality annotated resources for detecting harmful online behaviors—including sentiment, emotion, bias, hate speech, and group generalizations—in Icelandic-language online reviews. We introduce the first large-scale, multitask Icelandic blog comment dataset, comprising 12,232 unique comments and 19,301 high-quality annotations across 25 fine-grained behavioral categories. Our methodology employs a two-stage paradigm: (1) high-confidence automated pre-screening using GPT-4o mini, integrated with probability-driven sampling and quantification via a 5-point Likert scale; and (2) rigorous human verification and refinement through crowdsourcing. This approach enables the first reproducible, multidimensional, and fine-grained analysis of content safety in Icelandic. It substantially improves benchmark performance across related tasks and fills a longstanding gap in both data and methodology for harmful content detection in Nordic low-resource languages.

0 citationsRead paper

FoQA: A Faroese Question-Answering Dataset

Feb 11, 2025

Faroean lacks high-quality question answering (QA) datasets, hindering NLP advancement for this low-resource language. To address this, we introduce FoQA—the first extractive QA dataset for Faroese—constructed from Faroese Wikipedia. We employ GPT-4-turbo to generate initial QA pairs, augment question difficulty via paraphrasing, and conduct multi-round native-speaker validation, yielding 2,000 high-quality instances. Our methodology establishes a semi-automated “LLM-assisted + human refinement” paradigm for QA dataset construction, pioneering the first Faroese QA benchmark. We publicly release the development set, full dataset, and an error analysis subset. Comprehensive baseline evaluations on BERT-based models and state-of-the-art LLMs confirm FoQA’s validity and challenge level. All resources are fully open-sourced, filling a critical gap in QA research for under-resourced languages.

0 citationsRead paper
Recent publications

Latest Papers

A multilingual hallucination benchmark: MultiWikiQHalluA

May 04, 2026

This study addresses the critical gap in hallucination evaluation, which has been predominantly confined to English and thus fails to capture model fidelity issues in low-resource languages. Leveraging the multilingual MultiWikiQA dataset and the LettuceDetect framework, the authors construct the first large-scale synthetic hallucination benchmark spanning 306 languages and train token-level hallucination classifiers for 30 European languages to systematically assess cross-lingual hallucination behavior in mainstream large language models. Experimental results reveal that smaller models, such as Qwen3-0.6B, exhibit hallucination rates as high as 60% in low-resource languages like Icelandic, while larger models generally perform better. Crucially, hallucination rates are consistently higher across low-resource languages, underscoring both the necessity and the challenges of multilingual hallucination evaluation.

0 citationsRead paper

MultiZebraLogic: A Multilingual Logical Reasoning Benchmark

Nov 05, 2025

Existing logical reasoning benchmarks lack multilingual coverage and controllable difficulty. Method: We propose MultiZebraLogic—the first multilingual logical reasoning benchmark targeting the Germanic language family—employing 14 clue types and 8 distractor categories, integrated with rule-based automatic generation, multilingual text synthesis, tunable difficulty control, and red herring injection across nine Germanic languages. Contribution/Results: We construct a high-quality dataset comprising 128 (core) + 1,024 (extended) problems per language; enable cross-lingual fair evaluation and scalable topic/difficulty customization; and empirically demonstrate that 4×5 logic puzzles pose significant challenges to state-of-the-art LLMs, with distractor clues reducing average accuracy by 23.6%, thereby validating the benchmark’s sensitivity and effectiveness.

0 citationsRead paper

MultiWikiQA: A Reading Comprehension Benchmark in 300+ Languages

Sep 04, 2025

This work introduces XRC, the first large-scale multilingual reading comprehension benchmark covering 306 languages, designed to systematically evaluate cross-lingual understanding capabilities of language models. Methodologically, it automatically generates question-answer pairs from Wikipedia text, leveraging large language models for question generation and employing rigorous human crowd-sourced validation—assessing question fluency and answer accuracy—across 30 representative languages. Key contributions include: (1) establishing the first reading comprehension benchmark spanning over 300 languages; (2) revealing a substantial performance gap—up to 70 percentage points—between high-resource and low-resource languages in mainstream multilingual models; and (3) open-sourcing the complete dataset, annotation protocols, and evaluation code, thereby significantly advancing NLP research for low-resource languages.

0 citationsRead paper

Hotter and Colder: A New Approach to Annotating Sentiment, Emotions, and Bias in Icelandic Blog Comments

Feb 24, 2025

This work addresses the critical lack of high-quality annotated resources for detecting harmful online behaviors—including sentiment, emotion, bias, hate speech, and group generalizations—in Icelandic-language online reviews. We introduce the first large-scale, multitask Icelandic blog comment dataset, comprising 12,232 unique comments and 19,301 high-quality annotations across 25 fine-grained behavioral categories. Our methodology employs a two-stage paradigm: (1) high-confidence automated pre-screening using GPT-4o mini, integrated with probability-driven sampling and quantification via a 5-point Likert scale; and (2) rigorous human verification and refinement through crowdsourcing. This approach enables the first reproducible, multidimensional, and fine-grained analysis of content safety in Icelandic. It substantially improves benchmark performance across related tasks and fills a longstanding gap in both data and methodology for harmful content detection in Nordic low-resource languages.

0 citationsRead paper

FoQA: A Faroese Question-Answering Dataset

Feb 11, 2025

Faroean lacks high-quality question answering (QA) datasets, hindering NLP advancement for this low-resource language. To address this, we introduce FoQA—the first extractive QA dataset for Faroese—constructed from Faroese Wikipedia. We employ GPT-4-turbo to generate initial QA pairs, augment question difficulty via paraphrasing, and conduct multi-round native-speaker validation, yielding 2,000 high-quality instances. Our methodology establishes a semi-automated “LLM-assisted + human refinement” paradigm for QA dataset construction, pioneering the first Faroese QA benchmark. We publicly release the development set, full dataset, and an error analysis subset. Comprehensive baseline evaluations on BERT-based models and state-of-the-art LLMs confirm FoQA’s validity and challenge level. All resources are fully open-sourced, filling a critical gap in QA research for under-resourced languages.

0 citationsRead paper