VT-MUSE: Multimodal Unified Sequential Visuotactile Representation Learning for Manipulation

📅 2026-08-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
VT-MUSE通过两阶段学习框架解决视觉和触觉模态独立编码及忽略接触时间演变的问题,提高操作任务表现。
📝 Abstract
We propose VT-MUSE, a Multimodal Unified SEquential representation learning framework for visuotactilemanipulation. Existing approaches often encode visual and tactile observations independently before fusion, limiting their ability to capture fine-grained cross-modal dependencies. Moreover, most methods focus on observations at the current time step and overlook the temporal evolution of contact. VT-MUSE addresses both limitations through a two-stage representation learning framework. In Stage I, modality specific encoders are jointly adapted via cross-modal temporal alignment and masked-view consistency. In Stage II, a conditional variational latent model processes masked visual sequences together with full tactile histories. Auxiliary decoders reconstruct the masked recent visual observations and predict tactile depth changes, encouraging the latent representation to retain both global visual context and local contact dynamics. The learned representation is subsequently integrated into a lightweight Transformer policy through gated cross-attention. On the simulation benchmark, VT-MUSE outperforms the strongest baseline evaluated on all tasks by 11 percentage points and also achieves substantial improvements in real-world experiments.
Problem

Research questions and friction points this paper is trying to address.

multimodal
sequential representation learning
visuotactile manipulation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal Unified SEquential representation learning
cross-modal temporal alignment
masked-view consistency
conditional variational latent model
gated cross-attention
💼 Related Jobs
No related jobs found.
Congsheng Xu
Congsheng Xu
Undergraduate, SEIEE, Shanghai Jiao Tong University
Human Motion Generation Digital Twin Embodied AI
Q
Qiaochu Yang
Shanghai Jiao Tong University
F
Fangyuan Shi
Xense Robotics
Y
Yifan Han
Shanghai Jiao Tong University
Baijun Chen
Baijun Chen
Nanjing University
Artificial Intelligence
Yiming Wang
Yiming Wang
Shanghai Jiao Tong University
Large Language ModelsComplex ReasoningAI Interpretability
H
Haonan Zhao
Shanghai Jiao Tong University
Daolin Ma
Daolin Ma
Department of Engineering Mechanics, SJTU
Tactile SensingRobotic ManipulationContact MechanicsDynamics
X
Xiaokang Yang
Shanghai Jiao Tong University
H
Hesheng Wang
Shanghai Jiao Tong University