Judging LLM-as-a-Judge: Concerning Rubric Artifacts in LLM-based Automated Text Generation Evaluation

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究揭示基于规则的大规模语言模型评价方法存在缺陷,通过仅训练分类器识别规则文本即可预测评分,表明需要改进自动评价方法。
📝 Abstract
LLM-as-a-Judge pipelines are increasingly used to evaluate AI-generated text, based on the assumption that judgments arise from reasoning over candidate responses with respect to a rubric. We show that this assumption warrants further scrutiny. Classifiers trained only on rubric text, without access to any evaluated response, achieve nontrivial predictive performance on judge outputs. This suggests that rubric formulations encode recoverable evaluative signals, allowing scores to be partially anticipated independently of model outputs. Finally, counterfactual perturbations reveal that judges often fail to reliably update their decisions when either the candidate response or the rubric criterion is reversed. Our findings raise concerns about the reliability of rubric-based LLM evaluation and highlight the need for further methodological study of automated evaluation via LLMs.
Problem

Research questions and friction points this paper is trying to address.

LLM-as-a-Judge
Automated Text Generation Evaluation
Rubric Artifacts
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM-as-a-Judge
rubric artifacts
automated text generation evaluation
predictive performance
counterfactual perturbations
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Anshul Bagaria
Centre for Responsible AI (CeRAI), Wadhwani School of Data Science and AI (WSAI), IIT Madras, Chennai, India
S
Sowmya S Sundaram
Centre for Responsible AI (CeRAI), Wadhwani School of Data Science and AI (WSAI), IIT Madras, Chennai, India
Gokul S Krishnan
Gokul S Krishnan
Senior Research Scientist, CeRAI, IIT Madras
Natural Language ProcessingMachine LearningData ScienceHealthcare Informatics
Balaraman Ravindran
Balaraman Ravindran
Professor of Data Science and AI, Wadhwani School of Data Science and AI, IIT Madras
Reinforcement LearningData MiningNetwork AnalysisResponsible AI