DeCAL: Towards Physically-Grounded Dexterous Vision-Language-Action Models via Contact-Aware Latent Co-Imagination

📅 2026-09-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决灵巧操作中的视觉遮挡和复杂接触动力学问题,DeCAL通过自适应触觉融合与视觉-触觉潜在协同想象方法提高模型性能。
📝 Abstract
Dexterous manipulation involves contact-rich and fine-grained interactions with the physical world, posing significant challenges for existing vision-language-action (VLA) models due to severe visual occlusions and complex contact dynamics. While recent works have incorporated tactile sensing into robotic manipulation, most approaches still rely on homogeneous multimodal fusion, lacking adaptive tactile integration and explicit modeling of physical dynamics. In this work, we present DeCAL, a physically-grounded dexterous vision-language-action model that unifies understanding, imagination and action generation for contact-rich dexterous manipulation. Built upon a Mixture-of-Transformers (MoT) architecture, DeCAL leverages specialized experts for each capability while enabling efficient information flow among them. To effectively leverage tactile information, we introduce Adaptive Visuo-Tactile Fusion that dynamically regulates tactile interactions via a contact-aware gating strategy. Furthermore, we propose Visuo-Tactile Latent Co-Imagination to jointly model visual and tactile dynamics, equipping the policy with implicit physical world knowledge. Experimental results show that DeCAL consistently achieves state-of-the-art performance across all tasks, attaining a 71% average success rate and an 83.4% progress success rate, while also demonstrating strong generalization to unseen scenarios. The website is available at https://aureleopku.github.io/DeCAL.
Problem

Research questions and friction points this paper is trying to address.

Dexterous Manipulation
Vision-Language-Action Models
Tactile Sensing
Contact Dynamics
Adaptive Tactile Integration
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adaptive Visuo-Tactile Fusion
Visuo-Tactile Latent Co-Imagination
Mixture-of-Transformers
💼 Related Jobs
No related jobs found.
Y
Yankai Fu
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University; Beijing Academy of Artificial Intelligence
Ning Chen
Ning Chen
Peking University
satellite image understandingdeep learningrecommendation systemlarge-scale sparse learning
J
Junkai Zhao
Beijing Academy of Artificial Intelligence
H
Heng Zhang
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University
G
Guocai Yao
Beijing Academy of Artificial Intelligence
Pengwei Wang
Pengwei Wang
University of Calgary
Computer Science Security
Z
Zhongyuan Wang
Beijing Academy of Artificial Intelligence
Shanghang Zhang
Shanghang Zhang
Peking University
Embodied AIFoundation Models