Human-Anchored Factuality Evaluation with Strategic Annotation

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过策略性标注解决大模型事实判断与人类判断偏差问题,提出基于失效空间分析的标注策略以提高评估效率。
📝 Abstract
LLM-based factuality judges provide scalable evaluation signals, but their metrics are often systematically biased relative to human judgments. We study human-anchored factuality evaluation under limited annotation budgets, where judge predictions on the full dataset are combined with human labels on a small selectively sampled subset to obtain statistically valid estimates. The efficiency of this approach depends critically on which examples receive human annotation: in factuality evaluation, judge-human misalignment is not driven solely by low confidence, but also by structured failure modes such as incomplete evidence, temporal mismatch, unverifiable claims, and rubric misalignment. To exploit this structure, we introduce a factuality-specific annotation policy design pipeline that uses failure-space analysis (FSA) to derive diverse predictive signals for modeling human-judge misalignment. On an internal reference-based factuality evaluation system (AutoFA) and RAGTruth, where judge-predicted estimates substantially underestimate human-annotated factual accuracy, our FSA-guided policy improves annotation efficiency over uniform sampling and uncertainty-driven baselines, achieving effective-sample-size gains of 40.3% on AutoFA and 27.1% on RAGTruth.
Problem

Research questions and friction points this paper is trying to address.

factuality evaluation
human-anchored
systematic bias
limited annotation budgets
statistically valid estimates
Innovation

Methods, ideas, or system contributions that make the work stand out.

Human-anchored Evaluation
Failure Space Analysis (FSA)
Annotation Efficiency
Misalignment Modeling
🔎 Similar Papers
No similar papers found.