Process Reward Models Meet Planning: Generating Precise and Scalable Datasets for Step-Level Rewards
This work addresses the limitations of existing process reward models (PRMs), which rely on costly and error-prone human annotations and are largely confined to mathematical domains, thereby lacking the fine-grained feedback required for general reasoning tasks. To overcome these challenges, the study introduces a novel paradigm that integrates automated planning with PRM training. Specifically, logical problems are formalized using the Planning Domain Definition Language (PDDL), and large-scale datasets comprising millions of reasoning steps are generated via automated planning algorithms. This approach yields a cross-domain, scalable framework for PRM training. Experimental results demonstrate significant performance gains across multiple mathematical and non-mathematical reasoning benchmarks, confirming the effectiveness and generalizability of planning-generated data in enhancing model reasoning capabilities.