Slow to See, Slow to Suppress: Understanding the Effects of Modality in Context-Memory Conflicts

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究视觉-语言模型处理上下文记忆冲突的问题,通过增加视觉信息量来缓解模型在图像实体上偏好参数化信息的倾向。
📝 Abstract
We investigate how vision-language models (VLMs) handle context-memory conflicts; that is, situations in which the model is given information in context that differs from what was stored parametrically during training. We document asymmetric biases: models tend to prefer in-context information about entities which appear in text, but prefer parametric information about entities which appear in images. We relate this asymmetry to the late representational alignment across modalities, showing that the longer processing time associated with resolving visual entities prevents the suppression of the model's usual factual recall mechanism, thus resulting in more parametric answers. Chain-of-thought reasoning does not appear to resolve the gap, but increasing the amount of visual information in the context does show an effect. These results illustrate the complexity of ensuring consistent behavior as models become increasingly multimodal and retrieval-augmented.
Problem

Research questions and friction points this paper is trying to address.

context-memory conflicts
vision-language models
asymmetric biases
Innovation

Methods, ideas, or system contributions that make the work stand out.

context-memory conflicts
asymmetric biases
multimodal processing
visual information
🔎 Similar Papers
No similar papers found.
A
Athulith Paraselli
Brown University
E
Etha Tianze Hua
Brown University
Ellie Pavlick
Ellie Pavlick
Brown University
Natural Language Processing