TacUMI: A Multi-Modal Universal Manipulation Interface for Contact-Rich Tasks
This work addresses the challenge of accurately identifying semantic event boundaries in long-horizon manipulation tasks rich in physical contact, where reliance solely on visual and proprioceptive cues proves insufficient for effective task segmentation. To overcome this limitation, the authors propose TacUMI—a compact, multimodal data acquisition system that integrates ViTac visuo-tactile sensing with force-torque and pose perception—and, for the first time, embed it within a general-purpose manipulation interface to enable highly synchronized multimodal recording. Building upon this hardware foundation, they further introduce a temporal modeling–based multimodal fusion framework to automatically extract event boundaries from human demonstrations. Evaluated on a cable assembly task, the method achieves over 90% segmentation accuracy, significantly outperforming unimodal baselines and demonstrating the critical role of multimodal perception in enhancing task decomposition performance.