The BatchNorm Illusion: Diagnosing Normalization Artifacts in Machine Unlearning Evaluation
研究发现并诊断了BatchNorm架构中机器遗忘评估的未记录混淆问题,通过权重保留固定点操作方法区分测量失败与编码器失败。
研究发现并诊断了BatchNorm架构中机器遗忘评估的未记录混淆问题,通过权重保留固定点操作方法区分测量失败与编码器失败。
研究通过引入GeoAgent,利用具身导航改进视觉语言模型在地理位置定位中的表现,解决了静态图像分析的局限性。
研究通过文本扰动和归因分析,探讨了多模态时间序列预测中模型是否对文本语义敏感,发现文本内容并非提升预测性能的关键因素。
This work addresses the vulnerability of reward models in language model alignment to “reward hacking,” which can yield superficially high scores without genuine improvements in response quality. The study introduces a novel covariance-geometric perspective to analyze ensemble evaluators, revealing an intrinsic link between common-mode errors and evaluator disagreement, and formally demonstrates that internal scoring alone cannot detect such errors. To mitigate this, the authors propose a dual-anchor Bernstein error bound, integrating sub-Gaussian joint modeling, covariance decomposition, and adaptive search under conditional calibration, yielding a search error bound proportional to √log K. The theoretical claims are validated through 120 Gaussian stress tests and real-model audits, which also expose the limitations of disagreement-based metrics under strong search pressure.
This work addresses the vulnerability of Process Reward Models (PRMs) to exploitation, where adversarial inputs can inflate reasoning scores while yielding incorrect answers—posing a correctness reversal risk. Framing PRM stress testing as a quality-diversity search problem, the study employs MAP-Elites to identify the most severe reversal instances in behavior space and introduces a verifiable safety framework based on archive coverage. It demonstrates that coverage ratio alone cannot guarantee worst-case safety and establishes, for the first time under Lipschitz continuity, a theoretical upper bound on residual error. Applied to Qwen2.5-Math-PRM-7B, the analysis reveals a flaw in its aggregation mechanism, which is mitigated via LoRA-based adversarial fine-tuning. This yields 44 verified attack samples (max score gain of 0.294 under mean pooling); post-repair, attack success rates drop from 0.148 to 0.037–0.074, worst-case gains fall from 0.333 to 0.177–0.212, and ranking AUROC improves without compromising accuracy.
研究发现并诊断了BatchNorm架构中机器遗忘评估的未记录混淆问题,通过权重保留固定点操作方法区分测量失败与编码器失败。
研究通过引入GeoAgent,利用具身导航改进视觉语言模型在地理位置定位中的表现,解决了静态图像分析的局限性。
研究通过文本扰动和归因分析,探讨了多模态时间序列预测中模型是否对文本语义敏感,发现文本内容并非提升预测性能的关键因素。
This work addresses the vulnerability of reward models in language model alignment to “reward hacking,” which can yield superficially high scores without genuine improvements in response quality. The study introduces a novel covariance-geometric perspective to analyze ensemble evaluators, revealing an intrinsic link between common-mode errors and evaluator disagreement, and formally demonstrates that internal scoring alone cannot detect such errors. To mitigate this, the authors propose a dual-anchor Bernstein error bound, integrating sub-Gaussian joint modeling, covariance decomposition, and adaptive search under conditional calibration, yielding a search error bound proportional to √log K. The theoretical claims are validated through 120 Gaussian stress tests and real-model audits, which also expose the limitations of disagreement-based metrics under strong search pressure.
This work addresses the vulnerability of Process Reward Models (PRMs) to exploitation, where adversarial inputs can inflate reasoning scores while yielding incorrect answers—posing a correctness reversal risk. Framing PRM stress testing as a quality-diversity search problem, the study employs MAP-Elites to identify the most severe reversal instances in behavior space and introduces a verifiable safety framework based on archive coverage. It demonstrates that coverage ratio alone cannot guarantee worst-case safety and establishes, for the first time under Lipschitz continuity, a theoretical upper bound on residual error. Applied to Qwen2.5-Math-PRM-7B, the analysis reveals a flaw in its aggregation mechanism, which is mitigated via LoRA-based adversarial fine-tuning. This yields 44 verified attack samples (max score gain of 0.294 under mean pooling); post-repair, attack success rates drop from 0.148 to 0.037–0.074, worst-case gains fall from 0.333 to 0.177–0.212, and ranking AUROC improves without compromising accuracy.