π€ AI Summary
This study investigates whether visual language models (VLMs) genuinely rely on visual information localized to target regions when generating answers. Employing causal tracing, the authors systematically analyze which token positions in the visual encoder exert causal influence on model outputs and evaluate the modelsβ capacity to understand visual structure under conditions where appearance cues are absent. The findings reveal that tokens with high causal impact are often distributed outside target regions, challenging the assumption that strong multimodal performance implies spatially localized causal representations. Furthermore, the work demonstrates that prevailing VLMs predominantly depend on superficial appearance cues to infer structural relationships, exposing a significant gap between their abilities to perceive, utilize, and reason about visual structure. This research establishes a causal framework for analyzing how visual information is transformed, preserved, and leveraged in VLMs.
π Abstract
Vision-language models can describe an image with remarkable accuracy, yet a more fundamental question remains unanswered: what visual information actually drives their answers? In this work, we investigate this question through causal tracing, and we observe that highly causal vision tokens often lie outside the target region. Extending the analysis to larger vision-language models reveals a similar pattern across models and corruption settings, suggesting that strong multimodal performance does not necessarily imply spatially localized causal representations. We further investigate: can these models preserve visual structure when appearance cues are removed? and find that visual cues are exploited to understand visual structures. Together, our experiments expose a gap between seeing, using, and reasoning over visual structure, and provide a causal framework for studying how visual information is transformed, preserved, and ultimately used by modern vision-language models.