Granularity-Adaptive Credit Assignment for Long-Horizon LLM Agent Reinforcement Learning

📅 2026-09-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究提出GACA方法,通过基于不确定性的动态权重调整,解决长时序任务中信用分配精度问题,提升大型语言模型在强化学习中的表现。
📝 Abstract
Reinforcement learning is now the standard way to train large language model agents on long-horizon tasks, where dozens of interdependent actions precede a single sparse reward. Critic-free, group-relative methods such as GRPO suit this regime, but they broadcast one trajectory-level scalar to every step and cannot say which decision drove the outcome. GiGPO recovers a step-level signal by grouping time steps that share an anchor state, yet it merges the step- and episode-level estimates under one fixed weight, spending the same resolution on a pivotal branching decision as on a routine, near-deterministic transition. We argue that the right resolution is state-dependent, and propose GACA, a critic-free estimator whose granularity follows an uncertainty-based criticality proxy. GACA scores every step by the negative log-likelihood its own rollout already records, then blends the two advantages with a per-step weight that grows with that score, so the gradient places more weight on the fine-grained signal at above-average NLL and on the episode-level signal below it. We derive an exact risk decomposition for the implemented mixture and show that sufficiently small modulation improves on fixed mixing under positive directional alignment. A separate conditional result bounds local action-value variation using expected NLL, while an error-projection analysis characterizes when mixing adds value beyond scalar uncertainty reweighting. On ALFWorld and WebShop, GACA improves task success over GRPO and GiGPO at both 1.5B and 7B scales.
Problem

Research questions and friction points this paper is trying to address.

Reinforcement Learning
Credit Assignment
Long-Horizon Tasks
Large Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Granularity-Adaptive
Credit Assignment
Uncertainty-based
Reinforcement Learning
Large Language Models
🔎 Similar Papers
No similar papers found.
T
Taoran Liang
Nankai University
Y
Yang Liu
Supply Chain Tech Team Y, JD.com
S
Shang Luo
Peking University
Yingguang Yang
Yingguang Yang
University of Science and Technology of China
Rongrong Zhang
Rongrong Zhang
Georgia Southern University
Empirical corporate financeCorporate governanceInnovationand Institutional investors
Y
Yingzong Min
Shanghai Waybot Technology Co., Ltd.
Y
Yulin Huang
Supply Chain Tech Team Y, JD.com
J
Jianshen Zhang
Supply Chain Tech Team Y, JD.com
Y
Yongzhi Qi
Supply Chain Tech Team Y, JD.com
K
Kefu Xu
Peking University
C
Congjing Ran
Wuhan University
B
Bin Chong
Peking University