π€ AI Summary
This work addresses the vulnerability of reward models in language model alignment to βreward hacking,β which can yield superficially high scores without genuine improvements in response quality. The study introduces a novel covariance-geometric perspective to analyze ensemble evaluators, revealing an intrinsic link between common-mode errors and evaluator disagreement, and formally demonstrates that internal scoring alone cannot detect such errors. To mitigate this, the authors propose a dual-anchor Bernstein error bound, integrating sub-Gaussian joint modeling, covariance decomposition, and adaptive search under conditional calibration, yielding a search error bound proportional to βlog K. The theoretical claims are validated through 120 Gaussian stress tests and real-model audits, which also expose the limitations of disagreement-based metrics under strong search pressure.
π Abstract
Language-model judges and reward models enable scalable supervision, but finite optimization can exploit evaluator errors rather than improve response quality. We characterize this failure through the covariance geometry of evaluator ensembles. For calibrated judges, the ensemble mean retains common-mode error along the all-ones direction, whereas cross-judge disagreement captures only orthogonal error. Consequently, disagreement can be high despite robust aggregation, or low while shared response-dependent errors persist. We prove that common-mode error is not identifiable from internal judge scores alone. Under a joint sub-Gaussian model, we bound best-of-K selection overstatement and target-quality regret, extending the guarantees to predictably adaptive search under conditional calibration. The resulting search terms scale as the square root of log K and are asymptotically tight for Gaussian projected errors. We further show that noisy quality proxies introduce artificial rank-one covariance without changing disagreement, and propose a bounded two-anchor Bernstein certificate for finite-search error and regret. Fixed-seed Gaussian stress tests over 120 (J, rho, K) configurations and real-model audits validate the theory while revealing the limits of disagreement-based diagnostics under increasing search pressure.