KatotohananQA: Evaluating Truthfulness of Large Language Models in Filipino

📅 2025-09-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
The lack of factual evaluation benchmarks for low-resource languages—particularly Filipino—hinders reliable assessment of large language models’ (LLMs) truthfulness in such linguistic contexts. Method: We introduce KatotohananQA, the first Filipino factual evaluation benchmark, constructed via expert-translated and verified adaptations of TruthfulQA, and evaluated using a binary discrimination framework across mainstream LLMs. This work pioneers the first systematic cross-lingual factual transfer study from English to Filipino, exposing critical gaps in multilingual factuality evaluation. Contributions/Results: (1) We release the first open-source Filipino truthfulness benchmark; (2) Empirical results reveal significantly lower factual consistency of mainstream LLMs in Filipino compared to English, indicating exacerbated cross-lingual hallucination; (3) GPT-5 series models demonstrate superior multilingual robustness, and question-type sensitivity to hallucination varies across languages. These findings underscore the urgent need to expand multilingual factual evaluation to ensure model fairness, reliability, and equitable performance across linguistic communities.

Technology Category

Application Category

📝 Abstract
Large Language Models (LLMs) achieve remarkable performance across various tasks, but their tendency to produce hallucinations limits reliable adoption. Benchmarks such as TruthfulQA have been developed to measure truthfulness, yet they are primarily available in English, leaving a gap in evaluating LLMs in low-resource languages. To address this, we present KatotohananQA, a Filipino translation of the TruthfulQA benchmark. Seven free-tier proprietary models were assessed using a binary-choice framework. Findings show a significant performance gap between English and Filipino truthfulness, with newer OpenAI models (GPT-5 and GPT-5 mini) demonstrating strong multilingual robustness. Results also reveal disparities across question characteristics, suggesting that some question types, categories, and topics are less robust to multilingual transfer which highlight the need for broader multilingual evaluation to ensure fairness and reliability in LLM usage.
Problem

Research questions and friction points this paper is trying to address.

Evaluating truthfulness of LLMs in Filipino language
Addressing multilingual evaluation gap for low-resource languages
Assessing performance disparities across question characteristics
Innovation

Methods, ideas, or system contributions that make the work stand out.

Filipino translation of TruthfulQA benchmark
Binary-choice framework for model assessment
Evaluating multilingual robustness across question types
💼 Related Jobs
No related jobs found.
L
Lorenzo Alfred Nery
De La Salle University, Manila, Philippines
R
Ronald Dawson Catignas
De La Salle University, Manila, Philippines
Thomas James Tiam-Lee
Thomas James Tiam-Lee
Assistant Professor, De La Salle University
gamificationaffective computingeducationdata science