ForceDelta-VLA: Distilling Force-Conditioned ActionCorrections for Contact-Rich Manipulation

📅 2026-09-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为了解决接触丰富操作中任务级运动和接触调整结合的问题,提出ForceDelta-VLA方法,通过从力条件模式和无力量模式的预测中提取校正目标来改进动作预测。
📝 Abstract
Force-aware Vision-Language-Action (VLA) policies improve contact-rich manipulation, but typically combine task-level motion and contact-dependent adjustment in a single action prediction. Demonstrations provide no explicit labels for decomposing that prediction into a reusable reference action and a correction. We present ForceDelta-VLA, a correction-distillation framework that constructs an explicit force-correction target using paired predictions from a frozen teacher's force-conditioned and learned force-agnostic modes. A separate delay-correction target accounts for reference-action mismatch and the change in reference state. Training uses asynchronous schedule replay with the cached task context available during execution. The resulting lightweight policy adjusts the reference actions using recent force history and robot state, responding to contact changes between reference-action updates without regenerating complete action chunks. Across nine single-arm and bimanual contact-rich tasks, ForceDelta-VLA achieves an 82.2% mean success rate, compared with 54.4% for the original ForceVLA baseline. Direct execution of our Stage-1 Temporal Teacher achieves 70.6%. Relative to ForceVLA, the complete system reduces mean peak contact force over successful trials by approximately 26% on both platforms.
Problem

Research questions and friction points this paper is trying to address.

contact-rich manipulation
force-aware VLA policies
action prediction
Innovation

Methods, ideas, or system contributions that make the work stand out.

force-correction target
delay-correction target
asynchronous schedule replay
contact-rich manipulation
reference-action update
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Ju Dong
TAMS (Technical Aspects of Multimodal Systems), Department of Informatics, University of Hamburg, Hamburg, Germany.
Y
Yu Fu
TAMS (Technical Aspects of Multimodal Systems), Department of Informatics, University of Hamburg, Hamburg, Germany.
J
Jian Chen
University of Science and Technology of China, Hefei, China.
Yimeng Liu
Yimeng Liu
University of California, Santa Barbara
Human-Computer InteractionHuman-AI InteractionHuman-Centered AI
Haocheng Zhao
Haocheng Zhao
Xi'an Jiaotong Liverpool University
Neural NetworksNeural Network PruningRadar-Camera Fusion
L
Lei Zhang
TAMS (Technical Aspects of Multimodal Systems), Department of Informatics, University of Hamburg, Hamburg, Germany.
K
Kaixin Bai
TAMS (Technical Aspects of Multimodal Systems), Department of Informatics, University of Hamburg, Hamburg, Germany.
L
Liding Zhang
Technical University of Munich, Germany.
D
Diwen Zheng
Technical University of Munich, Germany.
A
Alois Christian Knoll
Technical University of Munich, Germany.
A
Angela P. Schoellig
Technical University of Munich, Germany.
J
Jianwei Zhang
TAMS (Technical Aspects of Multimodal Systems), Department of Informatics, University of Hamburg, Hamburg, Germany.