SAFE-G: Structure-aware Faithful Evidence-guided Generation for Knowledge-based Visual Question Answering

📅 2026-08-22
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决知识型视觉问答中结构关联捕捉与忠实推理问题,提出SAFE-G框架,通过结构感知细粒度图检索和基于证据的强化学习策略来精准定位证据并确保推理过程忠实于所获取证据。
📝 Abstract
Knowledge-based Visual Question Answering (KB-VQA) aims to answer queries that necessitate reasoning over external knowledge sources beyond the visual content. Typically, current methods fuse multimodal features to retrieve external information, subsequently leveraging Multimodal Large Language Models (MLLMs) to derive answers from the retrieved evidence. However, these methods often struggle to capture structural associations within complex contexts to effectively filter noise. Furthermore, they frequently fail to ensure that the reasoning process remains strictly faithful to the retrieved evidence. To address these challenges, we propose SAFE-G, a Structure-Aware Faithful Evidence-guided Generation framework, which enables precise evidence localization and trustworthy reasoning. Specifically, we first employ a coarse-grained hybrid search fusing visual and textual modalities to recall candidate documents, and subsequently implement a structure-aware fine-grained graph retrieval that captures structural dependencies to filter noise and pinpoint precise evidence. Moreover, we introduce a reinforcement learning (RL) strategy with an evidence-grounded reward that assigns credit to correct answers only when the selected evidence is correct. This strict alignment constraint compels the model to anchor its response in the retrieved context, effectively enhancing its capability to locate evidence via multimodal features and perform faithful reasoning. Extensive experiments on the Encyclopedic-VQA and InfoSeek benchmarks demonstrate that SAFE-G outperforms prior methods by a margin of 8.9% and 3.5%, substantially enhancing the overall reasoning accuracy. Our source code is publicly available at: https://github.com/MINE-USTC/SAFE-G.
Problem

Research questions and friction points this paper is trying to address.

Knowledge-based Visual Question Answering
structural associations
faithful reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Structure-Aware
Faithful Evidence-guided Generation
Reinforcement Learning
Evidence-grounded Reward
Graph Retrieval
🔎 Similar Papers
No similar papers found.
L
Long Shu
State Key Laboratory of Cognitive Intelligence, University of Science and Technology of China
Shuochen Liu
Shuochen Liu
University of Science and Technology of China
Large Language Model
W
Wei Chen
State Key Laboratory of Cognitive Intelligence, University of Science and Technology of China
J
Junda Lin
State Key Laboratory of Cognitive Intelligence, University of Science and Technology of China
Z
Zhi Zheng
State Key Laboratory of Cognitive Intelligence, University of Science and Technology of China
H
Huijun Hou
NIO
Tong Xu
Tong Xu
Professor, University of Science and Technology of China
Data Mining