GMTS: Gradient Magnitude-based Token Selection Improves RLVR Training for LLM Reasoning

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过引入基于梯度幅度的令牌选择方法(GMTS)改进了RLVR训练,以提高大语言模型的推理能力,解决了高熵令牌在不同答案中重要性不一致的问题。
📝 Abstract
Reinforcement learning (RL), particularly RL with Verifiable Rewards (RLVR), has recently emerged as a central paradigm for enhancing large language models' (LLMs) reasoning abilities, demonstrating remarkable effectiveness across reasoning tasks. Recent studies suggest that high-entropy tokens play an exceptionally important role in model training, since training with only the highest 20% entropy tokens yields significant performance gains. However, why such high-entropy tokens are beneficial remains insufficiently understood. In this work, we find that although high-entropy tokens within one answer tend to correlate with large gradient magnitude, entropy alone fails to consistently reflect token importance across different answers, considering the variations in the answer-level reward signals. Based on this observation, we introduce the Gradient Magnitude-based Token Selection (GMTS) method to quantify token importance, which leverages the entropy-gradient connection to approximate gradient-magnitude rankings for token selection. We find that training on the top 20% tokens ranked by GMTS consistently outperforms entropy-based token selection across three reasoning domains and various model sizes, suggesting that GMTS provides a more fine-grained estimate of token contribution for RLVR training.
Problem

Research questions and friction points this paper is trying to address.

Reinforcement Learning
Large Language Models
Token Selection
Gradient Magnitude
Reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Gradient Magnitude-based Token Selection (GMTS)
entropy-gradient connection
fine-grained estimate of token contribution
RLVR training
reasoning abilities
O
Outongyi Lv
School of Mathematical Sciences, Shanghai Jiao Tong University
Y
Yuanwei Zhang
School of Mathematical Sciences, Shanghai Jiao Tong University
Xiaoqun Zhang
Xiaoqun Zhang
Shanghai Jiao Tong University
Imaging sciencesinverse problemsoptimizationcompressive sensing