IF-RewardBench: Benchmarking Judge Models for Instruction-Following Evaluation
Existing instruction-following meta-evaluation benchmarks suffer from insufficient data coverage and oversimplified evaluation paradigms, limiting their ability to accurately reflect the performance of discriminative models in real-world alignment scenarios. To address this, this work proposes IF-RewardBench, a comprehensive benchmark encompassing diverse instruction types and constraints, which introduces—for the first time—a listwise ranking evaluation paradigm based on multi-response preference graphs. This approach better aligns with practical alignment requirements and significantly enhances the correlation between evaluation outcomes and downstream task performance. Experimental results reveal substantial deficiencies in current discriminative models’ instruction-following capabilities, while demonstrating that IF-RewardBench achieves stronger positive correlation and greater evaluative validity compared to existing benchmarks.