Implicit Action Chunking for Smooth Continuous Control

📅 2026-05-19
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of deploying reinforcement learning in physical systems, where high-frequency oscillatory control signals often compromise safety. To this end, the authors propose the Dual-Window Smoothing (DWS) framework, which achieves smooth continuous control through implicit action chunking without expanding the action space, thereby balancing temporal abstraction with per-frame responsiveness. DWS introduces a novel dual-window mechanism: an execution window ensures action smoothness, while a value window corrects critic bias, complemented by a lightweight first-order action-difference regularizer to enhance global continuity. Experiments demonstrate that DWS consistently outperforms existing methods across diverse domains—including the DeepMind Control Suite, industrial energy management, and vision-based autonomous driving—delivering jitter-free, highly safe, and stable control, with a 100% success rate in the autonomous driving task.
📝 Abstract
Reinforcement learning often produces high-frequency oscillatory control signals that undermine the safety and stability required for physical deployment. Explicit action chunking addresses this by predicting fixed-horizon trajectories but scales the policy output dimension proportionally with the horizon length, leading to optimization difficulties and incompatibility with standard step-wise interaction. To overcome these challenges, this paper proposes Dual-Window Smoothing (DWS), an implicit action chunking framework for smooth continuous control. Unlike explicit methods, DWS enforces temporal coherence without expanding the action space. It uses a dual-window design: an execution window that ensures physical smoothness through deterministic modulation, and a value window that aligns temporal-difference targets over the horizon to correct critic bias caused by open-loop execution. DWS also includes a lightweight actor-side temporal regularizer based on first-order action differences to promote global continuity. This design effectively bridges the gap between temporal abstraction and reactive step-wise control. Experiments on benchmarks including the DeepMind Control Suite and industrial energy management tasks show that DWS outperforms state-of-the-art (SOTA) baselines. In complex vision-based autonomous driving tasks, DWS achieves smoother control, safer behavior with reduced jitter, and attains a 100% success rate.
Problem

Research questions and friction points this paper is trying to address.

smooth continuous control
high-frequency oscillation
action chunking
reinforcement learning
temporal coherence
Innovation

Methods, ideas, or system contributions that make the work stand out.

Implicit Action Chunking
Dual-Window Smoothing
Temporal Coherence
Continuous Control
Reinforcement Learning
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
B
Bosun Liang
Department of Data and Systems Engineering, The University of Hong Kong, Hong Kong SAR, China
S
Shuo Pei
Department of Data and Systems Engineering, The University of Hong Kong, Hong Kong SAR, China
Z
Zirui Chen
Beijing Institute of Technology, Zhuhai, China
C
Chuanzhi Fan
Beijing Institute of Technology, Zhuhai, China
Chen Sun
Chen Sun
The University of Hong Kong, University of Waterloo
autonomous drivingscene understandingintelligent system's safety
Y
Yuankai Wu
College of Computer Science, Sichuan University, Chengdu, China
Huachun Tan
Huachun Tan
Beijing Institute of Technology
Image ProcessingPattern Recognitioncomputer visionIntelligent Transportation Systems
Yong Wang
Yong Wang
The University of Hong Kong
Large Language ModelNatural Language ProcessingMachine Learning