Same Semantics, Different Outcome: On the Modality Robustness of Multimodal LLMs under Knowledge Conflict

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究了多模态大语言模型在知识冲突下的模态鲁棒性问题,发现模型处理不同形式证据时存在不一致,并通过监督微调等方法尝试缓解这一脆弱性。
📝 Abstract
Multimodal large language models (MLLMs) are increasingly provided with contextual evidence in heterogeneous forms: as a text passage, as a rendered image of the same passage, or as both together. However, it remains unclear how consistently these surface forms are processed, especially when the evidence conflicts with the model's parametric knowledge. We study modality robustness under knowledge conflict across 13 MLLMs and two datasets, and find them far from robust. (1) Contrary to common belief, models favor a context that contradicts parametric knowledge more readily in image form than in text form; (2) when a contradicting text and image are presented together, the preferred modality is essentially arbitrary, varying with input order, model, and dataset. We further demonstrate that this instability has practical consequences: it degrades performance in multimodal RAG and can be exploited by adversarial attacks. To alleviate this brittleness, we examine several simple techniques---prompting, steering, supervised fine-tuning (SFT), and direct preference optimization; the majority prove ineffective, whereas SFT achieves moderate success. We therefore call for greater awareness of this inconsistency and argue that it is fundamental, demanding attention at multiple training stages.
Problem

Research questions and friction points this paper is trying to address.

modality robustness
knowledge conflict
multimodal LLMs
contextual evidence
Innovation

Methods, ideas, or system contributions that make the work stand out.

Modality Robustness
Knowledge Conflict
Supervised Fine-Tuning (SFT)
Multimodal LLMs
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Jungyeon Lee
Hanyang University, Seoul, Republic of Korea
Y
Yejin Yoon
Hanyang University, Seoul, Republic of Korea
Taeuk Kim
Taeuk Kim
Assistant Professor, Hanyang University.
Natural Language ProcessingLarge Language ModelsMachine Learning