Graph-MambaNav: Spatial-Temporal Graph Mamba Leveraging Object-Relation Knowledge for Object-Goal Navigation

📅 2026-08-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of graph-based methods in object goal navigation, specifically their lack of node prioritization and insufficient long-range dependency modeling. To overcome these challenges, we propose Graph-MambaNav, a framework that leverages large language model commonsense to construct target-aware spatiotemporal graphs. By integrating heuristic node ordering with selective scanning mechanisms, the approach effectively combines local message passing and global sequential reasoning for efficient structured decision-making. Experimental results demonstrate that Graph-MambaNav significantly enhances navigation performance and generalization capabilities on the AI2-THOR and RoboTHOR benchmarks. Furthermore, real-world robot deployment validates its practical effectiveness, establishing a novel modeling paradigm for embodied navigation.
📝 Abstract
Object-goal navigation requires an agent to reason over object relationships and prioritize target-relevant objects for efficient decision making in unseen environments. While existing graph-based methods incorporate target-awareness at the feature or attention level, they remain permutation-invariant and lack an explicit mechanism to control information propagation order, limiting their ability to model target-dependent importance and long-range dependencies. In contrast, Graph-Mamba highlights that node prioritization through sequence ordering is critical for effective global reasoning. In this work, we investigate the node prioritization mechanism in Graph-Mamba and study its role in object navigation. We propose Graph-MambaNav, a target-aware spatial-temporal graph encoding framework that introduces a heuristic ordering over objects based on their relevance to the target, allowing more informative objects to be processed later to aggregate richer context. Both node ordering and edge weights are initialized from LLM-derived commonsense object relationships, providing a unified prior for structured reasoning. A spatial module integrates local message passing with global GraphMamba-based selective scanning, while a temporal module applies Mamba-based sequence modeling over object-wise temporal orders, allowing selective aggregation of historical context for long-range temporal reasoning. Experiments on AI2-THOR and RoboTHOR demonstrate improved navigation performance with generalization, and additional real-world robot deployment further validates the effectiveness of our proposed approach.
Problem

Research questions and friction points this paper is trying to address.

Object-Goal Navigation
Graph-based Methods
Permutation Invariance
Long-range Dependencies
Node Prioritization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Graph-Mamba
Node Prioritization
Object-Goal Navigation
LLM Commonsense Knowledge
Spatio-Temporal Reasoning
Leyuan Sun
Leyuan Sun
Wuxi University
Embodied AI navigation3D vision
G
Genxin Chen
School of Internet of Things Engineering, Wuxi University, Wuxi, Jiangsu 214105, China
Linwei Ye
Linwei Ye
School of Internet of Things Engineering, Wuxi University, Wuxi, Jiangsu 214105, China
Y
Yan Zhang
School of Internet of Things Engineering, Wuxi University, Wuxi, Jiangsu 214105, China; School of Communications and Information Engineering, Nanjing University of Posts and Telecommunications, Nanjing, Jiangsu 210003, China
X
Xi Kan
School of Internet of Things Engineering, Wuxi University, Wuxi, Jiangsu 214105, China
Y
Yanfei Sun
Wuxi Key Laboratory of Artificial Intelligence and Security, Wuxi, Jiangsu 214105, China; Wuxi University, Wuxi, Jiangsu 214105, China