Can We Trust the Judges? Validation of Factuality Evaluation Methods via Answer Perturbation

📅 2026-09-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过引入一种元评估框架,使用控制下的答案扰动测试事实性评估方法的敏感性和可靠性,揭示了基于流水线的方法比LLM作为评判者的方法更能准确追踪退化。
📝 Abstract
Evaluating the factual correctness of large language models (LLMs) is vital for many applications. But are our evaluation tools themselves trustworthy? Despite the rise of factuality-based metrics, their sensitivity and reliability remain underexplored. This paper introduces a meta-evaluation framework that systematically tests these metrics using controlled corruptions of gold standard answers. Our method generates ranked outputs with known degrees of degradation to probe how metrics capture nuanced changes in truthfulness. Our experiments reveal that pipeline-based methods, such as the RAGAS's factual correctness metric, better track degradation than LLM-as-judge approaches. We also propose a new variant of the factual correctness metric that provides a competitive and cost-efficient.
Problem

Research questions and friction points this paper is trying to address.

factuality
evaluation metrics
large language models
reliability
Innovation

Methods, ideas, or system contributions that make the work stand out.

meta-evaluation framework
controlled corruptions
factual correctness metric
pipeline-based methods
cost-efficient
💼 Related Jobs
No related jobs found.
S
Sarra Gharsallah
Data Science Department, EURECOM, Campus SophiaTech, Biot, France
A
Adele Robaldo
Data Science Department, EURECOM, Campus SophiaTech, Biot, France
M
Mariia Tokareva
Data Science Department, EURECOM, Campus SophiaTech, Biot, France
G
Giovanni Gatti Pinheiro
Data Science Department, EURECOM, Campus SophiaTech, Biot, France
I
Ilyana Guendouz
Data Science Department, EURECOM, Campus SophiaTech, Biot, France
Raphaël Troncy
Raphaël Troncy
Associate Professor of Computer Science, EURECOM
Artificial IntelligenceSemantic WebNatural Language ProcessingKnowledge GraphRecSys
Paolo Papotti
Paolo Papotti
Professor at EURECOM
Data ManagementInformation QualityLLMs
Pietro Michiardi
Pietro Michiardi
EURECOM
Bayesian InferenceRepresentation LearningGenerative Modeling