ViTaR: Visuo-Tactile Residual Adaptation for Foundation VLA Manipulation

📅 2026-08-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the failure of foundational Vision-Language-Action (VLA) models in contact-rich manipulation tasks caused by the absence of tactile feedback. We propose ViTaR, a framework that reconceptualizes tactile signals as execution modulators rather than generative inputs. By employing effect-guided modeling and bounded residual action modulation, ViTaR applies corrective overlays atop frozen VLA backbones, thereby preserving pretrained generalization capabilities while preventing catastrophic forgetting. Experimental evaluations on seven UniVTAC tasks demonstrate that ViTaR achieves an average success rate of 61.3%, outperforming baselines by 30.6 percentage points. Furthermore, real-world robotic validation confirms its effectiveness, successfully unifying tactile adaptability with policy robustness in complex physical interactions.
📝 Abstract
As Vision-Language-Action (VLA) models scale toward real-world deployment, contact-rich manipulation exposes a critical blind spot: these policies encode broad visual-semantic priors yet remain unaware of local contact events, producing identical actions whether contact is established, lost, or destabilized. Existing remedies either modify VLA internals, risking catastrophic forgetting, or demand online reinforcement under near-failure contact conditions. Both grant tactile unbounded influence over action generation, conflicting with the priors that make VLAs generalizable. We introduce ViTaR, which reframes tactile feedback from an action-generating perceptual input to an execution modulator that selects and scales bounded residual corrections atop a frozen VLA, preserving pretrained capabilities by construction. ViTaR decomposes adaptation into two stages: Effect-Guided Modeling determines whether and which correction is locally justified via outcome-grounded preference evidence, and Residual Action Modulation converts this evidence into a residual choice with continuously scaled gain from real-time visuotactile observations. On the UniVTAC benchmark spanning seven contact-rich tasks, ViTaR achieves 61.3% average success, a 30.6 percentage-point improvement over its frozen VLA base that also surpasses purpose-built tactile baselines. Physical-robot experiments confirm that bounded tactile modulation transfers to real sensor noise and dynamics.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action
Contact-rich manipulation
Tactile feedback
Catastrophic forgetting
Generalization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Visuo-Tactile Residual Adaptation
Bounded Residual Correction
Effect-Guided Modeling
Frozen VLA
Execution Modulator
💼 Related Jobs
No related jobs found.
Y
Yi Wang
Beijing Institute of Technology
R
Renjun Wu
Beijing Institute of Technology
J
Jinyan Liu
Beijing Institute of Technology
Xuesong Li
Xuesong Li
Beijing Institute of Technology