Language models should be subject to repeatable, open, domain-contextualized hallucination benchmarking

📅 2025-05-22
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Current evaluation frameworks for language model hallucinations lack systematicity and reproducibility, often脱离 real-world domain contexts and suffering from low validity. Method: We propose the first domain-contextualized, open, and reproducible hallucination evaluation framework, featuring: (1) an original hallucination taxonomy; (2) an expert-collaborative annotation protocol, empirically demonstrating that early expert involvement critically enhances evaluation validity; and (3) context-sensitive benchmark tasks with rigorous validity verification procedures. Results: Experiments reveal that non-expert–driven evaluations frequently distort metric outcomes. Our framework substantially improves reliability, construct validity, and cross-domain generalizability of hallucination assessment. It provides both theoretical foundations and a practical paradigm for developing high-trust hallucination benchmarks.

Technology Category

Application Category

📝 Abstract
Plausible, but inaccurate, tokens in model-generated text are widely believed to be pervasive and problematic for the responsible adoption of language models. Despite this concern, there is little scientific work that attempts to measure the prevalence of language model hallucination in a comprehensive way. In this paper, we argue that language models should be evaluated using repeatable, open, and domain-contextualized hallucination benchmarking. We present a taxonomy of hallucinations alongside a case study that demonstrates that when experts are absent from the early stages of data creation, the resulting hallucination metrics lack validity and practical utility.
Problem

Research questions and friction points this paper is trying to address.

Measure prevalence of language model hallucination comprehensively
Evaluate models with repeatable open contextualized benchmarking
Address validity issues in hallucination metrics without expert input
Innovation

Methods, ideas, or system contributions that make the work stand out.

Repeatable hallucination benchmarking for language models
Domain-contextualized evaluation to measure hallucination prevalence
Expert-involved data creation for valid hallucination metrics
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Justin D. Norman
University of California, Berkeley School of Information
M
Michael U. Rivera
University of California, Berkeley School of Information