ERR+: Sequential Entropy Resolution for Efficient and Decisive LLM Reasoning

📅 2026-08-28
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究解决了大型推理模型内部推理结构优化问题,通过提出ERR+框架,采用两阶段强化学习方法,提高推理准确性和简洁性。
📝 Abstract
Large reasoning models achieve strong performance on complex tasks by generating extended chain-of-thought (CoT) traces via reinforcement learning with verifiable rewards (RLVR). While current RLVR methods have achieved strong results with correctness-based reward signals, they provide limited guidance on the quality of the reasoning process itself, leaving the internal reasoning structure largely unoptimized. Through empirical analysis across multiple model families, we identify a consistent pattern: correct reasoning trac es exhibit more frequent and larger token-level entropy drops within the thinking phase than incorrect ones. We propose ERR+, a two-phase RLVR framework grounded in this observation. The first phase trains with the Entropy Relief Reward (ERR), a bonus proportional to cumulative token-level entropy drops in the thinking phase, log-normalized by response length. Unlike prior methods that suppress entropy, ERR rewards the resolution of uncertainty while leaving exploratory high-entropy states unconstrained. The second phase introduces the Robust Relative Efficiency Reward, which scores each response's length against co-generated peers via a $\tanh$-transformed within-group $z$-score. We provide a formal analysis showing that joint optimization of the two objectives induces gradient conflict in early training, motivating the sequential design . Experiments on five datasets demonstrate consistent improvements in both accuracy and response conciseness across model backbones. Our code is available at https://github.com/XrkArul/err_response
Problem

Research questions and friction points this paper is trying to address.

reinforcement learning
verifiable rewards
reasoning process
entropy drops
internal reasoning structure
Innovation

Methods, ideas, or system contributions that make the work stand out.

Entropy Relief Reward
Sequential RLVR Framework
Reasoning Optimization
Token-level Entropy Drops
Robust Relative Efficiency Reward
Xin Jiang
Xin Jiang
Nanjing University of Science and Technology; Hidream.ai
Computer VisionAIGC
Minhao Wang
Minhao Wang
Data Scientist, MindRank AI
deep learning
Wen Wu
Wen Wu
Associate Researcher, Pengcheng Laboratory, China, IEEE Senior Member
Wireless networkingnetwork AInetwork slicingdigital twin
Z
Zhentao Xie
School of Computer Science and Technology, East China Normal University, Shanghai 200241, China
S
Shangheng Du
School of Computer Science and Technology, East China Normal University, Shanghai 200241, China
Jinxin Shi
Jinxin Shi
East China Normal Unversity
J
Jiabao Zhao
School of Computer Science and Technology, East China Normal University, Shanghai 200241, China