MCite-RL: Towards Reliable Multimodal RAG via Citation-enhanced Agentic Reinforcement Learning

📅 2026-08-22
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决多模态RAG中视觉引用不准确及与生成答案脱节问题,提出MCite-RL框架,通过代理精炼模块和增强奖励机制提高引用精确度和答案质量。
📝 Abstract
Multimodal Retrieval-Augmented Generation (RAG) with visual citation is crucial for ensuring the traceability and verifiability of MLLMs. However, current RAG and SFT-based methods struggle to achieve robust cross-modal reasoning, causing imprecise visual citations or decoupling between the citation and the generated answers. To address these limitations, we propose MCite-RL, a citation-enhanced agentic reinforcement learning framework designed for reliable multimodal RAG. MCite-RL introduces an Agentic Refinement module for visual citation that employs iterative retrieval, reasoning, and recursive cropping to progressively narrow the search space, transforming citation into a dynamic, evidence-driven reasoning process rather than a static step. Furthermore, we incorporate a Citation-enhanced Reward mechanism that integrates both process-level and outcome-level feedback within a reinforcement learning paradigm to jointly optimize answer accuracy and source traceability. Extensive experiments on benchmarks such as Wiki-VISA, FinRAGBench-V, and MMLongBench-Doc demonstrate that MCite-RL effectively achieves the joint optimization of citation precision and answer quality.
Problem

Research questions and friction points this paper is trying to address.

Multimodal RAG
cross-modal reasoning
visual citation
answer accuracy
source traceability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Agentic Refinement
Citation-enhanced Reward
Multimodal RAG
S
Suifeng Zhao
Key Laboratory of High Confidence Software Technologies, CS, Peking University, China
Zida Liu
Zida Liu
Key Laboratory of High Confidence Software Technologies, CS, Peking University, China
Xinyu Lei
Xinyu Lei
Michigan Technological University
Mobile ComputingData SciencePrivacy-Preserving Protocols
L
Lei Sun
Panasonic Connect Co., Ltd. Tokyo, Japan
Jun Gao
Jun Gao
Peking University
S
Sujian Li
Key Laboratory of Computational Linguistics(MOE), CS, Peking University, China