Reliable to Expressive: A Curriculum for Rubric-Following Safety Judges

📅 2026-06-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the fragility of existing safety evaluation models when confronted with variations in scoring criteria and prompts, which undermines their ability to consistently adhere to diverse judgment standards. The authors frame safety assessment as a criterion-following problem and propose a curriculum learning framework that progresses from “reliable” to “expressive” behaviors. By integrating dynamically generated instance-conditional scoring rubrics with supervised fine-tuning, they train a 12B-parameter language model to robustly align with shifting evaluation guidelines. Their approach is the first to systematically resolve judgment instability under varying criteria, achieving accuracies of 94.12%–94.88% across three distinct scoring standards with a remarkably low cross-criterion performance variance of only 0.76—significantly outperforming general-purpose large language models, specialized safety classifiers, and reasoning-based evaluators with fewer than 30B parameters.
📝 Abstract
Safety judges are increasingly deployed to evaluate model outputs against evolving criteria, yet recent meta-evaluation work shows they remain brittle under prompt and rubric variation, with false negative-rate swings of up to 0.24 reported for stylistic perturbations alone. We argue that safety judgment is fundamentally a rubric-following problem: a robust judge must apply the given evaluation criteria consistently across rubric formulations rather than memorize one specific template. We propose a training strategy that combines (i) instance-conditioned dynamic rubrics generated from prompt-response-label triples to expose the judge to the variability of evaluation criteria, and (ii) a reliable-to-expressive curriculum that begins with clean fixed-rubric supervision and progressively introduces noisier dynamic-rubric data. We evaluate on a single human-labeled set under three contrasting rubric prompts (HarmBench-style, ShieldGemma-style, and a domain-specific rubric). Our 12B curriculum judge achieves 94.12-94.88% accuracy across the three rubrics with a cross-rubric range of only 0.76, outperforming general-purpose LLMs, dedicated safety classifiers, and reasoning-oriented judges up to 30B in both peak accuracy and stability. An ablation shows that naively mixing dynamic rubrics into SFT increases cross rubric variance (1.44 -> 3.60); only the curriculum schedule recovers and improves on the fixed rubric baseline (variance 0.76).
Problem

Research questions and friction points this paper is trying to address.

safety judges
rubric-following
prompt variation
evaluation robustness
false negative rate
Innovation

Methods, ideas, or system contributions that make the work stand out.

curriculum learning
dynamic rubrics
safety judges
rubric-following
cross-rubric robustness
💼 Related Jobs
No related jobs found.
Y
Yongtaek Lim
AI Safety Team, DATUMO.INC, Seoul, South Korea
H
Hyeji Choi
AI Safety Team, DATUMO.INC, Seoul, South Korea
M
Minwoo Kim
AI Safety Team, DATUMO.INC, Seoul, South Korea