Visual Attention Faithfulness in Vision-Language Models is Heterogeneous

📅 2026-09-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
该研究通过因果扰动分析评估视觉-语言模型中视觉注意力的忠实性,揭示了三种处理模式,并指出模型视觉依赖与人类直觉存在系统性差异。
📝 Abstract
Whether attention weights faithfully reflect model reasoning has been actively debated in NLP, yet this question remains largely unexplored for the visual modality in Vision-Language Models (VLMs). We address this gap through causal perturbation analysis on current VLMs, evaluating both the comprehensiveness and sufficiency gap of attention-ranked visual tokens. Our analysis reveals that visual attention faithfulness is heterogeneous, manifesting in three distinct processing modes: Faithful-Sufficient, where top-$k$ attention tokens are both necessary and sufficient for prediction; Faithful-Distributed, where they are necessary but broader visual context remains required; and Non-Focal, where no localized attention region is individually necessary while visual information remains an essential trigger for prediction. Furthermore, human-annotated ground-truth regions satisfy comprehensiveness in only $\sim 60$% of cases compared with model attention rankings, revealing systematic divergence between model visual reliance and human intuition. We demonstrate these patterns across both general VQA on VQAv2 and document tasks on VRDU and ChartQA, showing that visual attention faithfulness varies systematically with processing demands and model architectures rather than being uniformly faithful or unfaithful.
Problem

Research questions and friction points this paper is trying to address.

Visual Attention Faithfulness
Vision-Language Models
Causal Perturbation Analysis
Innovation

Methods, ideas, or system contributions that make the work stand out.

causal perturbation analysis
visual attention faithfulness
processing modes
VLMs
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
X
Xurui Song
SAP; College of Computing and Data Science, Nanyang Technological University, Singapore
Weishi Wang
Weishi Wang
SAP
Document AILLM+RLNLP/ML/AILM for Code
Z
Zhongqi Yue
Microsoft
Kuluhan Binici
Kuluhan Binici
SAP
T
Tao Bai
SAP
H
Hongxin Shao
SAP
Daniel Dahlmeier
Daniel Dahlmeier
SAP SE
Machine LearningNatural Language Processing
J
Jun Luo
College of Computing and Data Science, Nanyang Technological University, Singapore