Multimodal Language Models Cannot Spot Spatial Inconsistencies
Current multimodal large language models struggle to recognize three-dimensional spatial inconsistencies across different viewpoints of the same scene. This work introduces a novel task: given a pair of images depicting the same scene from two distinct viewpoints, detect objects that violate 3D motion consistency. To facilitate research on this task, we develop a scalable synthetic framework capable of generating multiview image pairs with controllable spatial inconsistencies, and establish an evaluation protocol integrating human comparative experiments with model assessments. This study presents the first systematic evaluation of multimodal large language models’ ability to reason about 3D spatial consistency, revealing significant limitations in their understanding of physical world dynamics—state-of-the-art models perform substantially worse than humans and exhibit unstable performance across varying scene attributes.