Debiasing as a Measurement Intervention: Calibrated Ties and Resolution Loss in LLM-as-a-Judge Evaluation

📅 2026-09-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过TraceJudgeBench基准测试评估LLM作为评判时去除引用偏见的方法,发现强去偏指令虽然能减少偏见但会损害分辨能力,并提出了一种分离技术以恢复分辨率。
📝 Abstract
LLM-as-a-judge protocols are commonly debiased by instructing judges to ignore presentation cues such as citation formatting, source labels, and evidence-display style. We show that this intervention can suppress bias while damaging the resolution of the measurement instrument. We introduce TraceJudgeBench, a diagnostic benchmark for auditing citation-like artifacts in RAG and agent-workflow evaluation, covering content-equivalent pairs, citation ablations, correctness conflicts, human-validated soft and moderate quality gaps, prompt-strength ladders, decoupled judging, and a controlled workflow-ranking probe. Across GPT-5.5, Claude Sonnet 4.6, and DeepSeek V4-Flash, stronger anti-citation prompts reduce worse-cited wins from up to 50.5% to 0%; yet some operating points already convert validated moderate-gap decisions into Tie before the strict stress-test endpoint, while correctness-conflict accuracy remains at or above 93.0%. A second, 50-pair FinQA moderate-gap construction reproduces the qualitative frontier, and open-weight Qwen2.5-14B-Instruct-AWQ and Gemma-3-12B-IT runs reproduce the central HotpotQA frontier. TRACE-style decoupling recovers 96.5-100.0% better-plain resolution across the reported settings. Human validation separates three meanings of Tie: correct equivalence Tie, calibrated soft-boundary Tie, and resolution-destroying Tie on validated quality gaps. We frame debiasing as a measurement intervention whose bias suppression, resolution retention, Tie cost, and protocol cost must be reported jointly. The supplementary artifact contains benchmark splits, prompts, raw judge outputs, validation summaries, and analysis.
Problem

Research questions and friction points this paper is trying to address.

LLM-as-a-judge
debiased
resolution loss
citation formatting
bias suppression
Innovation

Methods, ideas, or system contributions that make the work stand out.

TraceJudgeBench
anti-citation prompts
resolution retention
decoupling
measurement intervention
🔎 Similar Papers
No similar papers found.
L
Liang Zhao
Shanghai University of International Business and Economics
Yong Wang
Yong Wang
Shanghai University of International Business and Economics
J
Jiangzhe Chen
Shanghai University of International Business and Economics