Visual Input and Its Framing Affect Attribute-based Descriptions Produced by Large Vision-Language Models

📅 2026-09-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究探讨了图像及其构图对大型视觉-语言模型基于属性描述的影响,揭示了视觉线索存在时即使文本提示不具体指向图像中的实例也会导致响应变化。
📝 Abstract
Large vision-language models (LVLMs) are commonly used with only a single text prompt as the input, or plus an image. In this paper, we demonstrate that when the image exists, even if the text prompt is not about the specific instance (but only the concept it belongs to) in that image, the response would still be affected. For example, when the text prompt only asks for the attribute descriptions of a dog breed, an image depicting a specific dog from that breed would shift the response. Further, how the specific instance is framed in that image would determine towards which the response shifts. Detailed analyses also reveal that in the response, physical terms increase from 18% for text-only to 45% (40%) for subject-focused (subject-in-situation) framings. Overall, the unexpected effects of visual cues on LVLMs highlight the need to understand the presence of an image and its framing when evaluating the robustness of LVLMs.
Problem

Research questions and friction points this paper is trying to address.

Large Vision-Language Models
Visual Cues
Attribute Descriptions
Framing Effects
Innovation

Methods, ideas, or system contributions that make the work stand out.

Large Vision-Language Models
Visual Cues Impact
Attribute Descriptions
Image Framing
Robustness Evaluation
🔎 Similar Papers
2024-02-26Conference on Empirical Methods in Natural Language ProcessingCitations: 1