Do Multimodal LLMs See Before They Read? Diagnosing Contextual Sycophancy

📅 2026-08-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过998个案例诊断多模态大语言模型中外部文本覆盖图像证据的问题,提出System-2 Visual Arbitration方法以改善模型表现。
📝 Abstract
External text can override conflicting image evidence in multimodal large language models, a failure we call multimodal contextual sycophancy. We introduce a 998-case diagnostic that independently varies visual evidence, commonsense priors, and external text, and probe when this failure arises by moving the information boundary around a context-blind visual witness. On abnormal images paired with Gemini-generated false text, GPT-5.1 scores 7.9% under joint conditioning, 49.7% when the context-blind witness report is scored directly, 63.7% under a matched two-call witness-arbiter pipeline that exposes the witness to the text, and 84.2% under System-2 Visual Arbitration (S2VA), which withholds the text from the witness. Across six models, S2VA improves over the direct witness report by 19.7 to 44.1 points, with all paired 95% confidence intervals excluding zero. The best information boundary is not uniform: textual context scaffolds some models, and a GPT-4o-regenerated subset changes the relative ordering of joint conditioning, Witness-Only, and S2VA. Contextual sycophancy is therefore sensitive to when text is introduced, as well as to the model and context source.
Problem

Research questions and friction points this paper is trying to address.

multimodal large language models
contextual sycophancy
external text
visual evidence
information boundary
Innovation

Methods, ideas, or system contributions that make the work stand out.

multimodal contextual sycophancy
diagnostic tool
System-2 Visual Arbitration (S2VA)
information boundary
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.