TAO-Force: Unifying Force-Aware Perception and Fast-Slow Control for Contact-Rich Manipulation

📅 2026-09-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
针对接触丰富操作中视觉感知和位置控制不足的问题,提出TAO-Force框架,通过力感知学习与接触调节执行相结合的方法来解决。
📝 Abstract
Vision-Language-Action (VLA) models have demonstrated strong performance across diverse robotic manipulation tasks, yet their predominantly vision-centric perception and position-controlled execution remain insufficient for contact-rich manipulation. Visual observations alone often provide limited evidence of contact onset and interaction magnitude, while position-control policies cannot respond compliantly to rapidly changing contact dynamics. To bridge both the perception and control gaps, we propose TAO-Force, a force-conditioned VLA framework that combines force-aware policy learning with contact-regulated execution. For force-aware perception, TAO-Force introduces Force-conditioned Feature-wise Linear Modulation (F-FiLM) to inject encoded force feedback into the representations of a frozen pretrained visual-language backbone while preserving its semantic priors. For responsive control, it employs a contact-gated fast-slow architecture, with a slow position-control branch tracking nominal trajectories during non-contact phases and a fast admittance-control branch regulating physical interaction during contact phases. Detailed analyses on a force-perception task and real-world evaluations across four contact-rich manipulation tasks validate the effectiveness and robustness of TAO-Force.
Problem

Research questions and friction points this paper is trying to address.

contact-rich manipulation
vision-centric perception
position-controlled execution
Innovation

Methods, ideas, or system contributions that make the work stand out.

Force-conditioned Feature-wise Linear Modulation (F-FiLM)
contact-gated fast-slow architecture
force-aware perception
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
B
Bohan Gan
China Mobile (Hangzhou) Information Technology Co., Ltd., Hangzhou 310023, China
X
Xuanzhang Wen
China Mobile (Hangzhou) Information Technology Co., Ltd., Hangzhou 310023, China
Y
Yongsheng Zhao
China Mobile (Hangzhou) Information Technology Co., Ltd., Hangzhou 310023, China
B
Baoping Cheng
China Mobile (Hangzhou) Information Technology Co., Ltd., Hangzhou 310023, China
Wenhe Jia
Wenhe Jia
China Mobile (Hangzhou) Information Technology Co., Ltd., Hangzhou 310023, China
Y
Ye Wang
China Mobile (Hangzhou) Information Technology Co., Ltd., Hangzhou 310023, China
G
Gongxin Yao
China Mobile (Hangzhou) Information Technology Co., Ltd., Hangzhou 310023, China
Han Gao
Han Gao
Alibaba Group, SenseTime Research, University of Electronic Science and Technology of China
Neural video/image compression
Jingyao Tang
Jingyao Tang
Dalian University of Technology
Natural language processingCausal inferenceLow-resourse issueInterpretability
L
Lei Zhao
China Mobile (Hangzhou) Information Technology Co., Ltd., Hangzhou 310023, China
J
Ji Ge
China Mobile (Hangzhou) Information Technology Co., Ltd., Hangzhou 310023, China