Bridging the Semantic-Utility Gap in Multimodal RAG via Generator-in-the-Loop Alignment

📅 2026-09-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决视觉-语言模型中语义与实用性差距问题,提出一种两阶段生成器循环对齐框架,通过生成器引导信号优化重排序器,提高答案准确性。
📝 Abstract
Vision-language models (VLMs) augmented with retrieval-augmented generation (RAG) benefit from access to external evidence. However, standard retrievers and rerankers optimize for semantic similarity rather than answer utility, creating a preference gap: documents that appear relevant may not help the generator produce a correct answer. Motivated by this, we propose a two-stage generator-in-the-loop alignment framework that closes this gap without human document-level relevance annotations. Our framework consists of two stages: in Stage 1, a VLM generates a hypothetical text passage from the image-query pair, which is used as the retrieval query for dense text search, bridging the image-to-text modality gap. In Stage 2, a cross-encoder reranker adapted with low-rank adaptation (LoRA) is fine-tuned using answer-supervised preference pairs mined from the frozen VLM: given the dataset answer label, a candidate document is labeled positive if the VLM produces the correct answer when given that document as context, and negative otherwise. This generator-guided signal is compatible with multiple alignment loss functions, including contrastive (triplet) loss, pairwise direct preference optimization (DPO), and supervised fine-tuning (SFT), and supports periodic re-mining to refresh preference pairs as the reranker improves. Experiments on VQA-X and A-OKVQA with Qwen3.5-2B and Qwen3-VL-4B-Instruct show that our proposed framework consistently outperforms rank-order, random, and REPLUG-style likelihood baselines under various alignment losses and pool size settings, suggesting that answer-level generator feedback is an effective supervision signal for preference alignment.
Problem

Research questions and friction points this paper is trying to address.

Semantic-Utility Gap
Multimodal RAG
Generator-in-the-Loop
Answer Utility
Retrieval-Augmented Generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

generator-in-the-loop
alignment framework
cross-encoder reranker
low-rank adaptation (LoRA)
answer-level feedback
💼 Related Jobs
No related jobs found.
Z
Zhan-Lun Chang
Elmore Family School of Electrical and Computer Engineering, Purdue University, West Lafayette, IN 47907, USA
Dong-Jun Han
Dong-Jun Han
Assistant Professor, Yonsei University
Edge AIOn-Device AIFederated LearningDistributed Machine LearningWireless Network
S
Seyyedali Hosseinalipour
Department of Electrical Engineering, University at Buffalo–SUNY, Buffalo, NY 14260, USA
M
Mung Chiang
Elmore Family School of Electrical and Computer Engineering, Purdue University, West Lafayette, IN 47907, USA
Christopher G. Brinton
Christopher G. Brinton
Elmore Associate Professor of ECE, Purdue University
NetworkingMachine LearningCommunicationsEdge ComputingNextG Wireless