High-Fidelity One-Step Generative Visuomotor Policy via Recursive Correction, Frequency Consistency, and Contrastive Flow Matching

📅 2026-07-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing single-step generative visuomotor policies often suffer from spatial bias, loss of high-frequency details, and ambiguous action modes due to oversimplified inference. To address these limitations, this work proposes a high-fidelity single-step framework that integrates Recursive Consistent Action Flow (RCAF) to correct spatial truncation errors, Dual-Timestep Frequency Consistency (DTFC) to preserve high-frequency manipulation details, and Contrastive Flow Matching (CFM) to disentangle multimodal action distributions. Requiring only a single forward pass (1 NFE), the method achieves or surpasses the control performance of multi-step baselines—such as 10-step policies—on RoboTwin, Adroit, DexArt, and real robotic platforms, significantly reducing latency while enhancing both precision and action diversity.
📝 Abstract
Generative models such as diffusion and flow matching have advanced robotic visuomotor policies by modeling multimodal action distributions, but their multi-step sampling or ODE solving introduces inference latency. Existing one-step acceleration methods often compress the whole generation process into a single large update, leading to spatial deviation, frequency distortion, and mode averaging. This paper proposes a high-fidelity one-step generative visuomotor policy framework that addresses these issues with three complementary mechanisms. Recursive Consistent Action Flow (RCAF) uses recursive correction to compensate for spatial truncation errors and align one-step predictions with refined flow trajectories. Dual-Timestep Frequency Consistency (DTFC) preserves high-frequency manipulation details through adaptive spectral consistency across flow timesteps. Contrastive Flow Matching (CFM) separates entangled action flows with a margin-based repulsive objective, reducing ambiguous actions in multimodal manipulation. Experiments on RoboTwin, RoboTwin 2.0, Adroit, DexArt, and real-world robot platforms show that the proposed method achieves competitive or superior performance compared with strong 10-step generative policy baselines while requiring only one forward pass (1 NFE), enabling low-latency visuomotor control.
Problem

Research questions and friction points this paper is trying to address.

one-step generative policy
spatial deviation
frequency distortion
mode averaging
visuomotor control
Innovation

Methods, ideas, or system contributions that make the work stand out.

Recursive Correction
Frequency Consistency
Contrastive Flow Matching
One-Step Generation
Visuomotor Policy
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yuran Chen
School of Safety Science and Engineering, Anhui University of Science and Technology, Huainan 232001, China; State Key Laboratory of Digital Intelligent Technology for Unmanned Coal Mining, Anhui University of Science and Technology, Huainan 232001, China
X
Xinye Cai
State Key Laboratory of Digital Intelligent Technology for Unmanned Coal Mining, Anhui University of Science and Technology, Huainan 232001, China; School of Artificial Intelligence, Anhui University of Science and Technology, Hefei 231131, China
Z
Zhonglin Gong
School of Safety Science and Engineering, Anhui University of Science and Technology, Huainan 232001, China; State Key Laboratory of Digital Intelligent Technology for Unmanned Coal Mining, Anhui University of Science and Technology, Huainan 232001, China
Yang Huang
Yang Huang
Nanjing University of Aeronautics and Astronautics
Reinforcement LearningSignal ProcessingOptimizationsMachine LearningWireless Communications