STAR: Sparse Tactile Representation Learning in Vision-Tactile-Language-Action Models for Dexterous Manipulation

📅 2026-09-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决灵巧操作中稀疏触觉信号表示学习难题,通过构建大规模数据集并提出STAR模型进行视觉-触觉联合预训练等方法。
📝 Abstract
Dexterous manipulation requires coordinated multi-finger control and effective tactile feedback, yet learning these capabilities remains challenging due to the lack of large-scale real-world data and the difficulty of extracting effective representations from sparse tactile signals. We build a robot platform and teleoperation system to collect a 200-hour bimanual dexterous manipulation dataset with synchronized visual, tactile, and language annotations, comprising 10,576 trajectories across 65 tasks, 69.5% of which involve dexterous multi-finger manipulation. We further propose STAR, an integrated training recipe for vision-tactile-language-action (VTLA) models that addresses the spatial, temporal, and informational sparsity of tactile signals through visual-tactile joint pre-training, sparse-global tactile token representation, and sparse future tactile prediction. Trained on this dataset, STAR achieves a 61% average success rate across four real-world tasks with 100 post-training trajectories per task, demonstrating dexterous performance under task-specific post-training.
Problem

Research questions and friction points this paper is trying to address.

Dexterous Manipulation
Tactile Feedback
Sparse Tactile Signals
Innovation

Methods, ideas, or system contributions that make the work stand out.

Sparse Tactile Representation
Vision-Tactile-Language-Action Models
Dexterous Manipulation
Bimanual Dataset
Joint Pre-training