Institution profile

Sea Group

Industry researchasia · sg
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

SCOPE: Compress Mathematical Reasoning Steps for Efficient Automated Process Annotation

May 20, 2025

To address the prohibitively high annotation cost of Process Reward Models (PRMs) for mathematical reasoning, this paper proposes a compression-driven automated annotation framework. It converts natural-language reasoning steps into AST-normalized code, identifies and merges semantically equivalent steps, and constructs an equivalence-step prefix tree to enable compression-aware high-quality sample generation. By replacing inefficient Monte Carlo sampling, the method reduces annotation complexity from *O(NMK)* to *O(N)*, achieving a 20× speedup—generating a 196K-sample high-quality dataset using only 5% of the computational budget. On both Best-of-N and ProcessBench benchmarks, it outperforms all existing automated annotation methods. Crucially, it is the first approach to jointly improve PRM training data quality and annotation efficiency, establishing a new paradigm for low-cost, high-fidelity alignment of mathematical reasoning.

0 citationsRead paper

Full-Step-DPO: Self-Supervised Preference Optimization with Step-wise Rewards for Mathematical Reasoning

Feb 20, 2025

Existing direct preference optimization (DPO) methods—e.g., Step-DPO—are limited in long-chain mathematical reasoning: they focus only on the first erroneous step, rely on manual or GPT-4–generated annotations for error localization, and lack fine-grained process-level supervision. Method: We propose Full-step Dynamic DPO, the first DPO framework incorporating self-supervised process reward modeling, which automatically generates learnable, step-wise rewards for every reasoning step. It introduces a step-weighted DPO loss to enable end-to-end optimization of reasoning trajectory quality. Contribution/Results: Our method requires no external annotations. On domain-specific (MathQA, AMC23) and cross-domain (AIME) mathematical benchmarks, it consistently surpasses state-of-the-art methods. It significantly improves both reasoning accuracy and robustness of foundational models—including LLaMA-3 and Qwen2—demonstrating superior generalization across diverse mathematical problem types and difficulty levels.

0 citationsRead paper
Recent publications

Latest Papers

SCOPE: Compress Mathematical Reasoning Steps for Efficient Automated Process Annotation

May 20, 2025

To address the prohibitively high annotation cost of Process Reward Models (PRMs) for mathematical reasoning, this paper proposes a compression-driven automated annotation framework. It converts natural-language reasoning steps into AST-normalized code, identifies and merges semantically equivalent steps, and constructs an equivalence-step prefix tree to enable compression-aware high-quality sample generation. By replacing inefficient Monte Carlo sampling, the method reduces annotation complexity from *O(NMK)* to *O(N)*, achieving a 20× speedup—generating a 196K-sample high-quality dataset using only 5% of the computational budget. On both Best-of-N and ProcessBench benchmarks, it outperforms all existing automated annotation methods. Crucially, it is the first approach to jointly improve PRM training data quality and annotation efficiency, establishing a new paradigm for low-cost, high-fidelity alignment of mathematical reasoning.

0 citationsRead paper

Full-Step-DPO: Self-Supervised Preference Optimization with Step-wise Rewards for Mathematical Reasoning

Feb 20, 2025

Existing direct preference optimization (DPO) methods—e.g., Step-DPO—are limited in long-chain mathematical reasoning: they focus only on the first erroneous step, rely on manual or GPT-4–generated annotations for error localization, and lack fine-grained process-level supervision. Method: We propose Full-step Dynamic DPO, the first DPO framework incorporating self-supervised process reward modeling, which automatically generates learnable, step-wise rewards for every reasoning step. It introduces a step-weighted DPO loss to enable end-to-end optimization of reasoning trajectory quality. Contribution/Results: Our method requires no external annotations. On domain-specific (MathQA, AMC23) and cross-domain (AIME) mathematical benchmarks, it consistently surpasses state-of-the-art methods. It significantly improves both reasoning accuracy and robustness of foundational models—including LLaMA-3 and Qwen2—demonstrating superior generalization across diverse mathematical problem types and difficulty levels.

0 citationsRead paper