Learning from Hard Prompts: Difficulty-aware Advantage Amplification in Dynamic Sampling

📅 2026-08-28
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出Direct Advantage Amplification方法,解决Dynamic Sampling在处理难样本时训练效率低的问题,通过放大难样本中正确响应的优势,提高训练效率。
📝 Abstract
Decoupled Clip and Dynamic Sampling Policy Optimization (DAPO) is a prominent variant of Group Relative Policy Optimization (GRPO). DAPO introduces several improvements over GRPO. Among these, Dynamic Sampling contributes the most to DAPO's accuracy gains relative to GRPO. To improve accuracy, Dynamic Sampling enhances training stability by eliminating zero policy gradients from zero advantages. Specifically, it avoids such zero gradients by filtering out prompts where sampled responses are either entirely correct or incorrect. However, our theoretical analysis shows that Dynamic Sampling decrease training efficiency as it cannot effectively utilize hard-to-sample correct responses on hard prompts. Formally, it asymmetrically amplifies the advantages of distinct responses to the same prompts. On hard prompts, incorrect responses undergo greater amplification than correct ones. This leads the model to avoid generating the observed incorrect responses rather than capitalizing on the hard-to-sample correct ones on hard prompts, resulting in low training efficiency. To improve training efficiency, we propose Direct Advantage Amplification (DAA), which amplifies the advantages of hard-to-sample correct responses on hard prompts, as obtained by Dynamic Sampling. This ensures that, when Dynamic Sampling is used, these hard-to-sample responses can be effectively capitalized on, implying higher training efficiency. By integrating DAA into DAPO, we obtain Difficulty-aware Advantage Amplification Policy Optimization (DA3PO), which is implemented with fewer than 30 lines of code from DAPO. Experiments show that DA3PO significantly outperforms GRPO and other classical GRPO variants.
Problem

Research questions and friction points this paper is trying to address.

Dynamic Sampling
Training Efficiency
Hard Prompts
Innovation

Methods, ideas, or system contributions that make the work stand out.

Direct Advantage Amplification
Dynamic Sampling
Hard Prompts
Training Efficiency
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Siyuan Gan
State Key Laboratory of Novel Software Technology, Nanjing University, Nanjing, China
Y
Yuhan Li
State Key Laboratory of Novel Software Technology, Nanjing University, Nanjing, China; Shanghai Artificial Intelligence Laboratory, Shanghai, China
X
Xiran Wang
State Key Laboratory of Novel Software Technology, Nanjing University, Nanjing, China; Shanghai Artificial Intelligence Laboratory, Shanghai, China
L
Linjian Meng
Shanghai Artificial Intelligence Laboratory, Shanghai, China
B
Boyan Wang
State Key Laboratory of Novel Software Technology, Nanjing University, Nanjing, China
Zhen Zhao
Zhen Zhao
Researcher @ Shanghai AI Lab
Machine LearningAI4ScienceComputer VisionMedical Image Analysis
Jing Huo
Jing Huo
Nanjing University
Machine LearningComputer Vision
Lei Bai
Lei Bai
Shanghai AI Laboratory
Foundation ModelScience IntelligenceMulti-Agent SystemAutonomous Discovery
Y
Yang Gao
State Key Laboratory of Novel Software Technology, Nanjing University, Nanjing, China