Judge Reliability Harness: Stress Testing the Reliability of LLM Judges

📅 2026-03-05
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the growing reliance on large language models (LLMs) as automated judges in AI evaluation, despite a lack of systematic validation of their reliability. We propose the first open-source stress-testing framework specifically designed for LLM-based judges, which automatically generates perturbations in text formatting, phrasing, and level of detail to assess accuracy and robustness in both binary classification and ordinal scoring tasks. The framework incorporates multidimensional metrics and spans diverse scenarios, including free-form responses and agent-based tasks. Experiments with four state-of-the-art LLM judges across four established benchmarks reveal that none consistently maintain reliable performance across all settings, exposing significant robustness deficiencies in current judge systems.

Technology Category

Application Category

📝 Abstract
We present the Judge Reliability Harness, an open source library for constructing validation suites that test the reliability of LLM judges. As LLM based scoring is widely deployed in AI benchmarks, more tooling is needed to efficiently assess the reliability of these methods. Given a benchmark dataset and an LLM judge configuration, the harness generates reliability tests that evaluate both binary judgment accuracy and ordinal grading performance for free-response and agentic task formats. We evaluate four state-of-the-art judges across four benchmarks spanning safety, persuasion, misuse, and agentic behavior, and find meaningful variation in performance across models and perturbation types, highlighting opportunities to improve the robustness of LLM judges. No judge that we evaluated is uniformly reliable across benchmarks using our harness. For example, our preliminary experiments on judges revealed consistency issues as measured by accuracy in judging another LLM's ability to complete a task due to simple text formatting changes, paraphrasing, changes in verbosity, and flipping the ground truth label in LLM-produced responses. The code for this tool is available at: https://github.com/RANDCorporation/judge-reliability-harness
Problem

Research questions and friction points this paper is trying to address.

LLM judges
reliability
stress testing
benchmark evaluation
judgment consistency
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM judges
reliability testing
stress testing
benchmark robustness
judgment consistency
💼 Related Jobs
No related jobs found.
S
Sunishchal Dev
RAND Corporation, Santa Monica, CA, USA
A
Andrew Sloan
RAND Corporation, Santa Monica, CA, USA
J
Joshua Kavner
RAND Corporation, Santa Monica, CA, USA
N
Nicholas Kong
RAND Corporation, Santa Monica, CA, USA
M
Morgan Sandler
RAND Corporation, Santa Monica, CA, USA