T-Rex: Tactile-Reactive Dexterous Manipulation
This work addresses the limitation of existing vision–language–action (VLA) models in robotic dexterous manipulation, which typically neglect tactile feedback or rely solely on static tactile representations, thereby failing to support dynamic tactile responses. To overcome this, the study introduces dynamic tactile perception into the VLA framework for the first time, proposing three core innovations: a large-scale, motion-primitive-based dataset enriched with high-frequency tactile signals, a temporal tactile VQ-VAE encoder that captures time-varying tactile features, and a variable-rate Mixture-of-Transformers architecture. The proposed method effectively leverages rich tactile dynamics, achieving an average success rate improvement of over 30% compared to the strongest baseline across twelve fine-grained force-control and deformable object manipulation tasks.