Beyond Noise Steering: Dual-Latent Space Reinforcement Learning for Generative Robot Policy

📅 2026-09-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对生成式机器人策略中仅能控制噪声空间而无法调节动作表示的问题,提出了一种双潜在空间强化学习框架,通过控制初始噪声和中间动作表示来提高性能。
📝 Abstract
Pretrained generative robot policies learn expressive action priors from demonstrations. However, existing reinforcement learning methods only steer the noisy space but fail to modulate intermediate action representations during the generation process, resulting in performance degradation and inefficiency. To address this limitation, we propose a novel Dual-Latent Space Reinforcement Learning (DLSRL) framework, which complements initial-noise steering with representation-level control inside the frozen generator. Specifically, our actor network predicts two distinct latent variables: an initial-noise latent variable that steers behavior generation, and an action-representation latent variable for intermediate feature modulation. Moreover, this representation latent variable is mapped to adapter features and ingeniously injected into the hidden states of intermediate action tokens via residual connections. Our dual-control design enables direct adjustment of action representations without updating the base policy. Experiments across generative policy architectures and robotic manipulation tasks show that DLSRL effectively accelerates online robot policy adaptation and achieves competitive performance. Our code is available at \href{https://github.com/xianchaoxiu/DLSRL}{https://github.com/xianchaoxiu/DLSRL}.
Problem

Research questions and friction points this paper is trying to address.

Reinforcement Learning
Generative Robot Policy
Action Representation
Noise Steering
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dual-Latent Space Reinforcement Learning
representation-level control
initial-noise steering
adapter features
P
Pengfei Zhang
School of Mechatronic Engineering and Automation, Shanghai University, Shanghai 200444, China
Teng Sun
Teng Sun
Shandong University
Multimedia computinginformation retrievalcausal inference
X
Xianchao Xiu
School of Mechatronic Engineering and Automation, Shanghai University, Shanghai 200444, China