PACE: Phase-Progress-Aware Credit for Long-Horizon Embodied Manipulation

📅 2026-08-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of policy optimization in long-horizon manipulation tasks caused by the absence of step-level credit assignment. We propose a Stage-Progress-Aware Credit framework that employs a global-local collaborative critic to achieve precise step-level credit allocation. Furthermore, this method integrates progressive policy distillation to translate assigned credits into positive and negative conditioning signals for guiding action generation. Experimental results demonstrate that the proposed framework significantly outperforms state-of-the-art baselines in both simulated and real-world robotic environments. By effectively enhancing post-training performance for long-horizon embodied manipulation, this work establishes a novel paradigm for learning complex sequential tasks.
📝 Abstract
Post-training of vision-language-action (VLA) models typically relies on expert demonstrations and policy interaction trajectories. However, in long-horizon manipulation, a single episode often spans hundreds of control steps and multiple phases, while success or failure is only revealed at episode termination. Policy improvement therefore requires step-level credit signals to distinguish behaviors that advance the task from those that stall or regress. We present PACE, a credit-assignment framework for post-training on long-horizon manipulation, centered on a phase-progress-aware critic. PACE consists of two key modules: (1) the Global-Local Cooperative Value-Correction Critic (GLC-Critic) aggregates visual and motion-difference features within local temporal windows to infer the phase and intra-phase progress of each step, and applies residual correction to a discretized remaining-cost distribution accordingly, enabling step-level credit assignment; (2) Progressive Policy Distillation (PPD) converts credit into positive and negative conditions via task-wise thresholds and trains a credit-conditioned action generation policy: it first protects the pretrained policy with high-credit positive samples, then incorporates all positive and negative credits to learn the quality boundary, and at inference amplifies high-credit behaviors through the difference between conditional outputs. Extensive simulation experiments and diverse real-world robotic-arm experiments demonstrate that PACE consistently achieves significant improvements over the strongest baseline.
Problem

Research questions and friction points this paper is trying to address.

Long-Horizon Embodied Manipulation
Credit Assignment
Vision-Language-Action Models
Sparse Reward
Post-training
Innovation

Methods, ideas, or system contributions that make the work stand out.

Phase-Progress-Aware Critic
Global-Local Cooperative Value-Correction Critic
Progressive Policy Distillation
Step-Level Credit Assignment
Long-Horizon Embodied Manipulation
🔎 Similar Papers
2024-07-16Neural Information Processing SystemsCitations: 16
C
Chengye Song
Intelligence Science and Technology, Dalian University of Technology, No. 2 Linggong Road, Ganjingzi District, Dalian 116024, China
J
Jiawei Zhang
Jilin University, No. 2699 Qianjin Street, Changchun 130012, China
R
Rui Song
Intelligence Science and Technology, Dalian University of Technology, No. 2 Linggong Road, Ganjingzi District, Dalian 116024, China
S
Shengqi Wang
Intelligence Science and Technology, Dalian University of Technology, No. 2 Linggong Road, Ganjingzi District, Dalian 116024, China
Xiangrong Zhang
Xiangrong Zhang
Professor, Xidian University
Image processing and understandingpattern recognitionmachine learning
Ziyi Wang
Ziyi Wang
University of Electronic Science and Technology of China
CVMLLMLLM
H
Huanbin Zhou
Jilin University, No. 2699 Qianjin Street, Changchun 130012, China
H
Hongzhou Wang
Jilin University, No. 2699 Qianjin Street, Changchun 130012, China