Bias and Uncertainty in LLM-as-a-Judge Estimation

📅 2026-05-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses systematic scoring biases in LLM-as-a-Judge evaluations, where existing correction methods suffer from cross-model calibration instability, often leading to directional misjudgments in model comparisons. Through theoretical analysis, Monte Carlo simulations, and experiments on MMLU-Pro, the work uncovers the hidden risks of shared calibration strategies and proposes two diagnostic metrics—judge quality (J) and cross-model calibration instability (ΔJ)—to assess the reliability of bias correction outcomes. The research successfully reproduces the sign-flipping phenomenon observed in prior work, empirically validating the effectiveness of the proposed indicators. Building on these insights, the authors establish a reporting protocol for LLM-as-a-Judge evaluations, offering both theoretical grounding and practical guidance to enhance the trustworthiness of comparative model assessments.
📝 Abstract
LLM-as-a-Judge evaluation has become a standard tool for assessing base model performance. However, characterizing performance via the naive estimator, i.e., raw judge outputs, is systematically biased. Recent work has proposed estimators to correct this bias, but their reliability depends critically on judge quality and, for model comparisons, on calibration stability. Sharing calibration across compared models is practically attractive but can introduce severe bias, including cases where the comparison estimate points in the wrong direction with high apparent confidence. We study these failure modes through analytical results, simulations over judge quality ($J$) and cross-model calibration instability ($ΔJ$), and a real-data MMLU-Pro case study with sign reversal. We propose $J$ and $ΔJ$ as diagnostics for when corrected estimates, especially shared-calibration comparisons, are likely unreliable, and provide reporting guidance for LaaJ evaluation.
Problem

Research questions and friction points this paper is trying to address.

LLM-as-a-Judge
bias
calibration instability
model comparison
uncertainty
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM-as-a-Judge
bias correction
calibration stability
uncertainty quantification
model comparison
💼 Related Jobs
No related jobs found.
J
James Fiedler
Indeed Inc.