Harness-agnostic detection and immunization of reward hacking in self-evolving language models

📅 2026-09-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出HackProbe方法,通过黑盒监测自进化语言模型中的奖励黑客行为,并使用风险意识免疫层重新选择候选更新以解决该问题。
📝 Abstract
Self-evolving language models improve by proposing candidate updates and keeping whatever raises a visible score. When that score is an imperfect proxy for the capability one actually wants, sustained selection widens the gap between the two. This is reward hacking. We introduce HackProbe, a monitor that attaches to an arbitrary self-evolving loop through two black-box hooks, with no access to weights or activations. It keeps a secret, distribution-fixed comparison core, whose frozen distribution makes its capability proxy comparable across generations, alongside a rotated fresh layer that hardens the bank against co-adaptation. Four tests built on that proxy cover the level gap, a scale-aligned divergence with online change-point detection, capability stagnation, and a conditional confidently-wrong rate; a Sidak correction turns them into a calibrated family-wise p-value. Diagnosis alone recovers nothing, so a risk-aware immunization layer reselects an honest candidate from the proposal pool using the core together with a purely structural gaming footprint, disclosing at most log2 Pi bits per generation to the host. We prove a detectability bound that converts a target error rate into an explicit probe-size budget, and we delimit what probe rotation does and does not buy. On a controlled prompt-level host with four injected hacking channels and ground-truth labels, HackProbe reaches 0.763 AUROC against 0.663 for the strongest baseline and cuts the false-positive rate from 0.706 to 0.434. Its bandwidth-limited reselection is the only immunization level that returns more true capability under hacking, 5.2 points on average, than it forfeits on clean runs, 4.7; per-channel effects are mostly not individually significant.
Problem

Research questions and friction points this paper is trying to address.

reward hacking
self-evolving language models
imperfect proxy
Innovation

Methods, ideas, or system contributions that make the work stand out.

HackProbe
Reward Hacking
Self-Evolving Language Models
Black-Box Monitoring
Immunization Layer
🔎 Similar Papers
No similar papers found.
R
Rongxin Yang
Fullive-AI
Y
Yang Liu
Supply Chain Tech Team Y, JD.com
S
Shang Luo
Peking University
H
Haoxuan Jia
Nanyang Technological University
Chongyang Zhang
Chongyang Zhang
Professor, Shanghai Jiao Tong University
Computer VisionMachine Learningand Artificial Intelligence
H
Hao Zheng
Fullive-AI
Yingguang Yang
Yingguang Yang
University of Science and Technology of China
Y
Yulin Huang
Supply Chain Tech Team Y, JD.com
J
Jianshen Zhang
Supply Chain Tech Team Y, JD.com
Y
Yongzhi Qi
Supply Chain Tech Team Y, JD.com
K
Kefu Xu
Peking University
C
Congjing Ran
Wuhan University
B
Bin Chong
Peking University