PlanPO: Group Planning-Aware Policy Optimization for Multi-Turn Agentic LLMs

📅 2026-08-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决多轮交互任务中成功轨迹间效率差异问题,提出PlanPO方法,通过引入粗到细的优势信号优化策略,提高模型在多轮基准测试中的表现。
📝 Abstract
Group-relative policy optimization has emerged as a key paradigm for training agentic large language models (LLMs) on multi-turn interactive tasks. However, most existing variants fail to distinguish advantages among successful trajectories even when these trajectories differ substantially in their interaction efficiency. For instance, circuitous successes are often assigned the identical outcome reward, causing advantage collapse and severe performance bottlenecks. To this end, we propose Group Planning-aware Policy Optimization (PlanPO), a simple yet effective RL method for learning generalizable planning abilities beyond task-specific high-quality behavior patterns. Specifically, PlanPO introduces coarse-to-fine advantage signals, which capture the relative differences in trajectory-level lengths and turn-level response lengths conditioned on successful trajectories sampled for the same task. Within the group-relative optimization structure, this enables agents to actively learn generalizable and deliberate behaviors spanning interaction planning and textual generation from high-quality rollouts, without degenerating into vanilla length minimization. Experimentally, PlanPO improves over GRPO by 27.2\% on average across the challenging multi-turn benchmarks ALFWorld, WebShop, and SciWorld, outperforming recent powerful baselines while incurring negligible additional training cost.
Problem

Research questions and friction points this paper is trying to address.

Group-relative policy optimization
multi-turn interactive tasks
advantage collapse
interaction efficiency
trajectory
Innovation

Methods, ideas, or system contributions that make the work stand out.

Group Planning-aware Policy Optimization
coarse-to-fine advantage signals
multi-turn interactive tasks
interaction efficiency
D
Dayang Liang
Department of Automation, Xiamen University, Xiamen, China
Liyuan He
Liyuan He
School of Artificial Intelligence, Shanghai Jiaotong University, Shanghai, China
X
Xuan Feng
College of Cyberspace Security, Jinan University, Guangzhou, China
S
Shuxin Li
College of Computing and Data Science, Nanyang Technological University, Singapore
Bo An
Bo An
Nanyang Technological University
Artificial intelligencemulti-agent systemsgame theoryreinforcement learningoptimization
Y
Yunlong Liu
Department of Automation, Xiamen University, Xiamen, China