Evaluator Ensembles Under Reward Hacking: Covariance Geometry and Finite-Search Guarantees

πŸ“… 2026-08-08
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the vulnerability of reward models in language model alignment to β€œreward hacking,” which can yield superficially high scores without genuine improvements in response quality. The study introduces a novel covariance-geometric perspective to analyze ensemble evaluators, revealing an intrinsic link between common-mode errors and evaluator disagreement, and formally demonstrates that internal scoring alone cannot detect such errors. To mitigate this, the authors propose a dual-anchor Bernstein error bound, integrating sub-Gaussian joint modeling, covariance decomposition, and adaptive search under conditional calibration, yielding a search error bound proportional to √log K. The theoretical claims are validated through 120 Gaussian stress tests and real-model audits, which also expose the limitations of disagreement-based metrics under strong search pressure.
πŸ“ Abstract
Language-model judges and reward models enable scalable supervision, but finite optimization can exploit evaluator errors rather than improve response quality. We characterize this failure through the covariance geometry of evaluator ensembles. For calibrated judges, the ensemble mean retains common-mode error along the all-ones direction, whereas cross-judge disagreement captures only orthogonal error. Consequently, disagreement can be high despite robust aggregation, or low while shared response-dependent errors persist. We prove that common-mode error is not identifiable from internal judge scores alone. Under a joint sub-Gaussian model, we bound best-of-K selection overstatement and target-quality regret, extending the guarantees to predictably adaptive search under conditional calibration. The resulting search terms scale as the square root of log K and are asymptotically tight for Gaussian projected errors. We further show that noisy quality proxies introduce artificial rank-one covariance without changing disagreement, and propose a bounded two-anchor Bernstein certificate for finite-search error and regret. Fixed-seed Gaussian stress tests over 120 (J, rho, K) configurations and real-model audits validate the theory while revealing the limits of disagreement-based diagnostics under increasing search pressure.
Problem

Research questions and friction points this paper is trying to address.

reward hacking
evaluator ensembles
common-mode error
response quality
finite-search guarantees
Innovation

Methods, ideas, or system contributions that make the work stand out.

covariance geometry
evaluator ensembles
reward hacking
common-mode error
finite-search guarantees
πŸ’Ό Related Jobs
No related jobs found.
F
Fariya Afrin
Department of Computer Science, Kalinga Institute of Industrial Technology
Ibne Farabi Shihab
Ibne Farabi Shihab
Iowa State University
Deep LearningroboticsLarge Language Model