Mechanisms of Object Localization in Vision-Language Models
This study addresses the unclear internal mechanisms underlying object localization in vision-language models (VLMs), which hinders their interpretability and performance improvement. It reveals, for the first time at the layer and attention head granularity, that object localization in LLaVA-1.5 and InternVL-3.5 relies on narrow computational pathways formed by a small subset of specialized attention heads—rather than internal semantic rearrangements—exhibiting a “containerized” mechanism. Through token ablation, attention knockout, and causal mediation analysis, the work demonstrates that localization and classification tasks share early visual processing but are driven by distinct sets of attention heads: in LLaVA, critical heads concentrate in early-to-mid layers, whereas in InternVL, they are distributed across mid-to-late layers.