CARE: Confounder-Aware Aggregation for Reliable LLM Evaluation

📅 2026-02-09
🏛️ arXiv.org
📈 Citations: 9
Influential: 1
📄 PDF
📝 Abstract
LLM-as-a-judge ensembles are the standard paradigm for scalable evaluation, but their aggregation mechanisms suffer from a fundamental flaw: they implicitly assume that judges provide independent estimates of true quality. However, in practice, LLM judges exhibit correlated errors caused by shared latent confounders -- such as verbosity, stylistic preferences, or training artifacts -- causing standard aggregation rules like majority vote or averaging to provide little gain or even amplify systematic mistakes. To address this, we introduce CARE, a confounder-aware aggregation framework that explicitly models LLM judge scores as arising from both a latent true-quality signal and shared confounding factors. Rather than heuristically re-weighting judges, CARE separates quality from confounders without access to ground-truth labels. We provide theoretical guarantees for identifiability and finite-sample recovery under shared confounders, and we quantify the systematic bias incurred when aggregation models omit confounding latent factors. Across 12 public benchmarks spanning continuous scoring, binary classification, and pairwise preference settings, CARE improves aggregation accuracy, reducing error by up to 26.8\%. Code is released in \href{https://github.com/SprocketLab/CARE}{https://github.com/SprocketLab/CARE}.
Problem

Research questions and friction points this paper is trying to address.

LLM-as-a-judge
confounders
aggregation mechanisms
correlated errors
latent factors
Innovation

Methods, ideas, or system contributions that make the work stand out.

confounder-aware aggregation
latent confounders
systematic bias reduction
LLM evaluation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Jitian Zhao
Department of Statistics, University of Wisconsin–Madison, Madison, WI, USA
Changho Shin
Changho Shin
University of Wisconsin-Madison
Machine learningdata science
Tzu-Heng Huang
Tzu-Heng Huang
Ph.D. student, University of Wisconsin-Madison
Data-centric AIData CurationMultimodal ModelsLLM Evaluation
S
Satya Sai Srinath Namburi GNVV
Department of Computer Sciences, University of Wisconsin–Madison, Madison, WI, USA
Frederic Sala
Frederic Sala
Assistant Professor, University of Wisconsin
Data-centric AIMachine learningInformation theory