MV-STRIDE: Enabling MLLMs to Master Multi-View Spatial Reasoning via Hierarchical Capability Modeling

📅 2026-09-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过构建MV-STRIDE数据集,解决多视角空间推理问题,采用层次化能力建模方法,提升MLLMs在3D一致的空间推理上的性能。
📝 Abstract
Despite the rapid progress of Multimodal Large Language Models (MLLMs) in 2D vision-language tasks, robust multi-view spatial reasoning remains a fundamental bottleneck due to the lack of structured 3D cognitive pathways in existing datasets. To address this, we introduce MV-STRIDE, a Multi-View hierarchical SpaTial Reasoning dataset with Interdependent and DEcomposed capabilitiEs. Moving beyond flat data structures, MV-STRIDE explicitly models the dependency relationships between foundational perception, scene understanding, and complex contextual reasoning, providing a coherent learning pathway aligned with human spatial cognition. We develop a systematic QA generation pipeline leveraging diverse 3D scene sources that enforces cross-view dependency constraints to prevent single-view solvability, generating multi-level spatial reasoning tasks supported by cognitively grounded chain-of-thought supervision for complex inference. Extensive evaluations demonstrate that our multi-stage training framework based on our hierarchical dataset achieves state-of-the-art performance across multiple spatial reasoning benchmarks, notably the multi-view oriented MMSI-Bench. Our approach enables MLLMs to maintain robust, 3D-consistent spatial reasoning across diverse viewpoints. The code and dataset are available at https://co1dspring.github.io/MV-STRIDE/.
Problem

Research questions and friction points this paper is trying to address.

Multimodal Large Language Models
Multi-View Spatial Reasoning
Structured 3D Cognitive Pathways
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-View Spatial Reasoning
Hierarchical Capability Modeling
Cross-view Dependency
Chain-of-thought Supervision
🔎 Similar Papers
No similar papers found.
J
Jin Xu
MoE Key Laboratory of Brain-inspired Intelligent Perception and Cognition, University of Science and Technology of China, China
X
Xiaojian Huang
MoE Key Laboratory of Brain-inspired Intelligent Perception and Cognition, University of Science and Technology of China, China
Z
Zhuodong Luo
MoE Key Laboratory of Brain-inspired Intelligent Perception and Cognition, University of Science and Technology of China, China
Z
Zhihong Zhang
MoE Key Laboratory of Brain-inspired Intelligent Perception and Cognition, University of Science and Technology of China, China
X
Xin Liu
Huawei Noah’s Ark Lab
J
Jiansheng Wei
Huawei Noah’s Ark Lab
X
Xinzhi Wang
Huawei Noah’s Ark Lab
Jie Zhao
Jie Zhao
Dalian University of Technology
computer visionvisual object tracking
Xuejin Chen
Xuejin Chen
University of Science and Technology of China
3D modelingScene understandingImage Generation