A Causality-aware Infer-diagnose-refine Framework for Test-time Modality Adaptation in VLA Models
This work addresses the challenge of multimodal fusion in vision-language-action (VLA) models, where the importance of visual inputs dynamically shifts across different stages of robotic manipulation. To tackle this, the authors propose an Inference-Diagnosis-Refinement (IDR) framework that operates at test time without requiring retraining. IDR infers actions using both factual and counterfactual visual inputs, where counterfactual predictions are generated via zero-filling interventions. By quantifying the causal effect of visual observations through norm-based metrics, the method dynamically estimates visual modality importance and adaptively refines action outputs via a gated residual fusion module. As the first approach to incorporate causal awareness for test-time, adaptive modality weighting, IDR significantly enhances the performance of diverse VLA backbone models on both simulated and real-world robotic tasks.