Beyond What Meets the Eye: Unveiling Situational Illusions for Multimodal Large Language Models

📅 2026-08-23
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文探讨了情境幻觉对多模态大语言模型的挑战,并通过开发评估基准和改进方法(如提示工程和监督微调)来提高模型在复杂环境中的可靠性。
📝 Abstract
Real-world situation appearances can deviate from their underlying physical states, challenging the reliability of multimodal large language models (MLLMs) in practical applications. In this paper, we term this phenomenon situational illusions and investigate: (1) how MLLMs perform under such illusions, and (2) how to mitigate the limitations. We first develop a comprehensive where-what-how taxonomy that characterizes where situational illusions occur, what targets they take, and how they arise. Building on this taxonomy, we introduce MSIBench, a benchmark designed to assess the discrimination, understanding, and reasoning capabilities of MLLMs under situational illusions. Evaluations of 27 model configurations reveal that current MLLMs are highly vulnerable to these illusions and exhibit 6 typical failure modes related to visual observation, grounding, and reasoning. To mitigate the limitations, we build on the core idea of systematically inspecting and reasoning over visual evidence for contextual understanding, developing prompting for closed-source models and supervised fine-tuning for open-source models, respectively. These two simple yet effective methods improve model performances by 20% at most, suggesting a practical path toward more reliable multimodal perception and reasoning in complex real-world environments.
Problem

Research questions and friction points this paper is trying to address.

Situational Illusions
Multimodal Large Language Models
Reliability
Innovation

Methods, ideas, or system contributions that make the work stand out.

situational illusions
MSIBench
multimodal large language models
supervised fine-tuning
prompting
Z
Zhiming Yang
School of Artificial Intelligence, Beijing University of Posts and Telecommunications, China
Z
Zhuoxi Xiong
School of Artificial Intelligence, Beijing University of Posts and Telecommunications, China
D
Donglin Zhou
School of Artificial Intelligence, Beijing University of Posts and Telecommunications, China
W
Wenjun Wei
School of Artificial Intelligence, Beijing University of Posts and Telecommunications, China
Shiyao Cui
Shiyao Cui
Tsinghua University
J
Jinqiao Shi
Beijing University of Posts and Telecommunications, China