From Vision to Language: Investigating Causal Information Flow in Multimodal Decision-Making

📅 2026-09-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过在视频-文本注意路径上进行层级因果干预,探讨视觉信息如何影响基于语言的决策,揭示了名词和动词在多模态推理中的不同作用及模型处理时间关系时的挑战。
📝 Abstract
Vision-Language Models are commonly evaluated through their final predictions, but understanding whether these decisions are grounded in visual evidence requires tracing how visual information contributes to language-based decisions. With this purpose in mind, we investigate cross-modal information flow in a video-based generative multiple-choice-like setting by applying a layer-wise causal intervention on video-text attention pathways. We target spatial, causal, and temporal visual reasoning. Our results show that visual information is mainly integrated while the model processes the candidate answer options, which serve as the primary textual grounding sites for the final decision. We further show that nouns play an important role as semantic anchors during multimodal enrichment, while verbs are more relevant when temporal relations are processed. Finally, we identify a distinct pattern in temporal reasoning, suggesting that VLMs struggle to reconstruct sequential information across video frames, but we remark that such fragility may also reflect linguistic biases associated with specific temporal expressions used for defining the relation between events within a scene.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Models
multimodal decision-making
visual information flow
cross-modal information
Innovation

Methods, ideas, or system contributions that make the work stand out.

layer-wise causal intervention
visual-textual information flow
multimodal decision-making
temporal reasoning
💼 Related Jobs
No related jobs found.
D
Davide Testa
Fondazione Bruno Kessler (FBK) - Trento, Italy; Università di Roma La Sapienza - Rome, Italy
H
Hugh Mee Wong
Utrecht University - Utrecht, The Netherlands
A
Alessandro Lenci
University of Pisa - Pisa, Italy
Bernardo Magnini
Bernardo Magnini
Researcher, Fondazone Bruno Kessler - FBK, Trento, Italy
Intelligenza ArtificialeComputational Linguistics
Albert Gatt
Albert Gatt
Professor of Natural Language Generation, Utrecht University
Computational LinguisticsNatural Language GenerationVision and LanguageLanguage Production