Judging by the Cover: Cleaning LLM Truthfulness Benchmarks to Avoid Surface-Level Feature Leakage

📅 2026-09-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对LLM真实性基准测试中的表面特征泄露问题,提出了一种名为Audit-Prune的方法来清理这些数据集,从而减少模型利用非预期线索的机会。
📝 Abstract
Binary-choice truth benchmarks ask models to choose between a correct and an incorrect answer, but if the two answers differ systematically in surface-level features, models can exceed chance without performing the intended reasoning. We show that this failure mode is detectable and can be exploited by downstream classifiers. In TruthfulQA, a simple six-feature logistic classifier achieves substantial accuracy in separating correct from incorrect answers. We further show that similar surface-level artifacts are present in additional benchmarks. To counteract this, we developed a general mechanism to clean them by removing the most leakage-reinforcing pairs. We release a version of TruthfulQA with surface-feature leakage reduced close to chance and provide a mechanism, Audit-Prune, so that the datasets can be cleaned before release.
Problem

Research questions and friction points this paper is trying to address.

surface-level features
truth benchmarks
model reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

surface-level feature leakage
binary-choice truth benchmarks
Audit-Prune
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
F
Foad Namjoo
University of Utah
R
Remy Ogasawara
University of Utah
Amirali Abdullah
Amirali Abdullah
Thoughtworks
Neural ReasoningMech interpDeep LearningCS Theory
C
Cullen Anderson
University of Massachusetts Amherst
N
Narmeen Fatimah Oozeer
Martian AI
J
Jeff M. Phillips
University of Utah