๐ค AI Summary
This work addresses the challenge that vision alone struggles to capture fine-grained contact interactions in robotic insertion tasks, and existing visionโtactile fusion methods often exhibit unstable performance. To this end, the authors propose a Cross-Modal Transformer (CMT) that effectively integrates wrist-mounted visual and tactile signals through structured self-attention and cross-attention mechanisms. Notably, they introduce, for the first time, a physics-based bilateral force equilibrium prior as a regularization term to stabilize tactile embeddings and enhance symmetry perception. Evaluated on the TacSL benchmark, the proposed method achieves a 96.59% insertion success rate, substantially outperforming existing baselines and approaching the performance of an oracle system with access to ideal contact force information (96.09%), thereby demonstrating the efficacy of integrating multimodal perception with physical priors.
๐ Abstract
Insertion tasks in robotic manipulation demand precise, contact-rich interactions that vision alone cannot resolve. While tactile feedback is intuitively valuable, existing studies have shown that na\"ive visuo-tactile fusion often fails to deliver consistent improvements. In this work, we propose a Cross-Modal Transformer (CMT) for visuo-tactile fusion that integrates wrist-camera observations with tactile signals through structured self- and cross-attention. To stabilize tactile embeddings, we further introduce a physics-informed regularization that encourages bilateral force balance, reflecting principles of human motor control. Experiments on the TacSL benchmark show that CMT with symmetry regularization achieves a 96.59% insertion success rate, surpassing na\"ive and gated fusion baselines and closely matching the privileged"wrist + contact force"configuration (96.09%). These results highlight two central insights: (i) tactile sensing is indispensable for precise alignment, and (ii) principled multimodal fusion, further strengthened by physics-informed regularization, unlocks complementary strengths of vision and touch, approaching privileged performance under realistic sensing.