Don't Blind Your VLA: Aligning Visual Representations for OOD Generalization
This work addresses the degradation of vision-language (VL) representations during action fine-tuning of Vision-Language-Action (VLA) models, which impairs out-of-distribution (OOD) generalization. We systematically characterize the trade-off between action adaptation and visual representation collapse. To mitigate this, we propose a lightweight hidden-layer representation alignment strategy that explicitly preserves pre-trained VL knowledge via cross-task feature constraints and attention-guided regularization—without incurring additional inference overhead. Through representation probing, attention visualization, and ablation on contrastive tasks, we demonstrate that our method significantly alleviates visual representation degradation. Empirically, it improves OOD generalization across multiple robotic manipulation benchmarks. The implementation is publicly available.