Hypergraph-based Multimodal Retrieval-Augmented Generation with Incremental Refinement

πŸ“… 2026-08-17
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the challenges of missing higher-order correlations and cross-page semantic fragmentation in multimodal Retrieval-Augmented Generation (RAG) by proposing the Hyper-M2RAG framework. This approach introduces a novel multimodal hypergraph as a unified semantic container to model higher-order dependencies and employs an anchor-driven incremental refinement mechanism to replace global reconstruction, effectively bridging cross-page knowledge gaps through one-hop neighborhood context restructuring. Experimental results demonstrate that Hyper-M2RAG significantly outperforms state-of-the-art methods on multimodal benchmarks in both retrieval accuracy and generation coherence. Consequently, this work establishes a new paradigm for mitigating computational redundancy and noise in complex multimodal retrieval tasks.
πŸ“ Abstract
Modern Multimodal Retrieval-Augmented Generation (M-RAG) systems are fundamentally limited by the binary connectivity paradigm of traditional simple graphs, which fails to capture the intricate, high-order correlations among heterogeneous entities, such as the N-ary relationships between a visual chart, its scattered textual descriptions, and underlying numerical data. Furthermore, existing refinement strategies often rely on exhaustive, full-page reconstruction to align cross-modal information, leading to prohibitive computational redundancy and the introduction of contextual noise in long-form document processing. In this paper, we propose Hyper-M2RAG, a novel framework that redefines multimodal document retrieval through High-order Hypergraph Representation Learning. We first formalize the document structure as a Multimodal Hypergraph, utilizing hyperedges as unified semantic containers to encapsulate multi-way associations across text, images, and tables, thereby transcending point-to-point modeling. To mitigate semantic fragmentation caused by physical pagination, we introduce an Anchor-driven Incremental Refinement mechanism. Rather than performing a global sweep, our approach identifies boundary-crossing anchor nodes and reconstructs their local hyper-topology using one-hop neighborhood contexts. This targeted refinement effectively bridges cross-page knowledge gaps with minimal computational footprints. Extensive evaluations on multimodal benchmarking datasets demonstrate that Hyper-M2RAG significantly outperforms state-of-the-art methods in both retrieval precision and generation coherence. Our code is available at https://github.com/ShenAoChen2001/MMHRAG.
Problem

Research questions and friction points this paper is trying to address.

Multimodal Retrieval-Augmented Generation
High-order Correlations
Computational Redundancy
Contextual Noise
Semantic Fragmentation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hypergraph Representation Learning
Multimodal Retrieval-Augmented Generation
Incremental Refinement
High-order Correlations
Anchor-driven Mechanism
πŸ’Ό Related Jobs
No related jobs found.
S
Shenao Chen
Hangzhou Dianzi University
Y
Yidan Xu
Hangzhou Dianzi University
Xiangmin Han
Xiangmin Han
Postdoctoral, Tsinghua University
hypergraphmedical image analysisbrain network analysispathology analysis
R
Rundong Xue
Xi’an Jiaotong University
D
Duanpo Wu
Hangzhou Dianzi University
Y
Yuhan Gao
Hangzhou Dianzi University
Chenggang Yan
Chenggang Yan
Hangzhou Dianzi University
Yue Gao
Yue Gao
Tsinghua University
Artificial IntelligenceComputer VisionHypergraph Computation