Institution profile

Kalinga Institute of Industrial Technology

Academic institutionasia · in
Official website
Research library62linked papers
Opportunities0open roles
Selected work

Representative Papers

Evaluator Ensembles Under Reward Hacking: Covariance Geometry and Finite-Search Guarantees

Aug 08, 2026

This work addresses the vulnerability of reward models in language model alignment to “reward hacking,” which can yield superficially high scores without genuine improvements in response quality. The study introduces a novel covariance-geometric perspective to analyze ensemble evaluators, revealing an intrinsic link between common-mode errors and evaluator disagreement, and formally demonstrates that internal scoring alone cannot detect such errors. To mitigate this, the authors propose a dual-anchor Bernstein error bound, integrating sub-Gaussian joint modeling, covariance decomposition, and adaptive search under conditional calibration, yielding a search error bound proportional to √log K. The theoretical claims are validated through 120 Gaussian stress tests and real-model audits, which also expose the limitations of disagreement-based metrics under strong search pressure.

0 citationsRead paper

Quality-Diversity Stress Tests for Process Reward Models:What Archive Coverage Can and Cannot Certify

Aug 08, 2026

This work addresses the vulnerability of Process Reward Models (PRMs) to exploitation, where adversarial inputs can inflate reasoning scores while yielding incorrect answers—posing a correctness reversal risk. Framing PRM stress testing as a quality-diversity search problem, the study employs MAP-Elites to identify the most severe reversal instances in behavior space and introduces a verifiable safety framework based on archive coverage. It demonstrates that coverage ratio alone cannot guarantee worst-case safety and establishes, for the first time under Lipschitz continuity, a theoretical upper bound on residual error. Applied to Qwen2.5-Math-PRM-7B, the analysis reveals a flaw in its aggregation mechanism, which is mitigated via LoRA-based adversarial fine-tuning. This yields 44 verified attack samples (max score gain of 0.294 under mean pooling); post-repair, attack success rates drop from 0.148 to 0.037–0.074, worst-case gains fall from 0.333 to 0.177–0.212, and ranking AUROC improves without compromising accuracy.

0 citationsRead paper
Recent publications

Latest Papers

Evaluator Ensembles Under Reward Hacking: Covariance Geometry and Finite-Search Guarantees

Aug 08, 2026

This work addresses the vulnerability of reward models in language model alignment to “reward hacking,” which can yield superficially high scores without genuine improvements in response quality. The study introduces a novel covariance-geometric perspective to analyze ensemble evaluators, revealing an intrinsic link between common-mode errors and evaluator disagreement, and formally demonstrates that internal scoring alone cannot detect such errors. To mitigate this, the authors propose a dual-anchor Bernstein error bound, integrating sub-Gaussian joint modeling, covariance decomposition, and adaptive search under conditional calibration, yielding a search error bound proportional to √log K. The theoretical claims are validated through 120 Gaussian stress tests and real-model audits, which also expose the limitations of disagreement-based metrics under strong search pressure.

0 citationsRead paper

Quality-Diversity Stress Tests for Process Reward Models:What Archive Coverage Can and Cannot Certify

Aug 08, 2026

This work addresses the vulnerability of Process Reward Models (PRMs) to exploitation, where adversarial inputs can inflate reasoning scores while yielding incorrect answers—posing a correctness reversal risk. Framing PRM stress testing as a quality-diversity search problem, the study employs MAP-Elites to identify the most severe reversal instances in behavior space and introduces a verifiable safety framework based on archive coverage. It demonstrates that coverage ratio alone cannot guarantee worst-case safety and establishes, for the first time under Lipschitz continuity, a theoretical upper bound on residual error. Applied to Qwen2.5-Math-PRM-7B, the analysis reveals a flaw in its aggregation mechanism, which is mitigated via LoRA-based adversarial fine-tuning. This yields 44 verified attack samples (max score gain of 0.294 under mean pooling); post-repair, attack success rates drop from 0.148 to 0.037–0.074, worst-case gains fall from 0.333 to 0.177–0.212, and ranking AUROC improves without compromising accuracy.

0 citationsRead paper