Where To Look? : Causal Tracing of Vision Encoders in VLM

πŸ“… 2026-08-11
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study investigates whether visual language models (VLMs) genuinely rely on visual information localized to target regions when generating answers. Employing causal tracing, the authors systematically analyze which token positions in the visual encoder exert causal influence on model outputs and evaluate the models’ capacity to understand visual structure under conditions where appearance cues are absent. The findings reveal that tokens with high causal impact are often distributed outside target regions, challenging the assumption that strong multimodal performance implies spatially localized causal representations. Furthermore, the work demonstrates that prevailing VLMs predominantly depend on superficial appearance cues to infer structural relationships, exposing a significant gap between their abilities to perceive, utilize, and reason about visual structure. This research establishes a causal framework for analyzing how visual information is transformed, preserved, and leveraged in VLMs.
πŸ“ Abstract
Vision-language models can describe an image with remarkable accuracy, yet a more fundamental question remains unanswered: what visual information actually drives their answers? In this work, we investigate this question through causal tracing, and we observe that highly causal vision tokens often lie outside the target region. Extending the analysis to larger vision-language models reveals a similar pattern across models and corruption settings, suggesting that strong multimodal performance does not necessarily imply spatially localized causal representations. We further investigate: can these models preserve visual structure when appearance cues are removed? and find that visual cues are exploited to understand visual structures. Together, our experiments expose a gap between seeing, using, and reasoning over visual structure, and provide a causal framework for studying how visual information is transformed, preserved, and ultimately used by modern vision-language models.
Problem

Research questions and friction points this paper is trying to address.

vision-language models
causal tracing
visual structure
multimodal reasoning
vision encoders
Innovation

Methods, ideas, or system contributions that make the work stand out.

causal tracing
vision-language models
visual tokens
multimodal representation
visual structure reasoning
N
Naren Kumar S
LINGO Research Group, Indian Institute of Technology Gandhinagar, India
T
Tirth Bhatt
LINGO Research Group, Indian Institute of Technology Gandhinagar, India
Mayank Singh
Mayank Singh
Assistant Professor, Computer Science and Engineering, IIT Gandhinagar
LLMsInterpretabilityExplainabilityCode-mixingNLP/ML/AI