See First, Answer Later: Visual Evidence Pre-Alignment via Sufficiency-Driven RL

πŸ“… 2026-06-16
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the frequent inconsistency in multimodal large language models’ responses due to their inadequate utilization of fine-grained visual evidence from images. To mitigate this issue, the authors introduce a Visual Evidence Pre-Alignment (VEPA) stage positioned between pretraining and post-training. VEPA employs a sufficiency-driven objective conditioned on questions to refine visual caption generation and integrates Group Relative Policy Optimization (GRPO), a reinforcement learning algorithm, to enhance the model’s perception and exploitation of critical visual evidence. The proposed approach yields substantial performance gains across multiple vision-intensive benchmarks, with improvements attributed to transferable visual grounding capabilities rather than task-specific overfitting.
πŸ“ Abstract
Multimodal large language models (MLLMs) integrate strong text reasoning with visual inputs, yet their responses can be inconsistent with the underlying images, indicating ineffective utilization of visual evidence during inference. The prevailing training paradigm relies on large-scale caption-based pretraining for general alignment, followed by supervised fine-tuning and reinforcement learning to enable instruction following and complex reasoning. However, such pretraining provides only weak visual grounding: short, coarse captions bias models toward salient objects while neglecting fine-grained visual evidence. In this paper, we introduce Visual Evidence Pre-Alignment (VEPA), an intermediate stage between pretraining and post-training that explores a novel sufficiency-driven objective with Group Relative Policy Optimization (GRPO) to optimize question-conditioned visual evidence descriptions. Extensive experiments across diverse benchmarks show that our VEPA consistently enhances performance on visually demanding evaluations and complements standard supervised post-training. Further analyses show that the income stems from strengthened, transferable visual grounding, rather than from additional task-specific training.
Problem

Research questions and friction points this paper is trying to address.

multimodal large language models
visual grounding
visual evidence
inconsistent responses
fine-grained visual information
Innovation

Methods, ideas, or system contributions that make the work stand out.

Visual Evidence Pre-Alignment
Sufficiency-Driven RL
Group Relative Policy Optimization
Multimodal LLMs
Visual Grounding
πŸ”Ž Similar Papers
No similar papers found.
Y
Yilian Liu
Beijing University of Posts and Telecommunications, China
Sicong Leng
Sicong Leng
Nanyang Technological University
Multi-modal Learning
Guoshun Nan
Guoshun Nan
Professor of Beijing University of Posts and Telecommunications
Multimodal LearningVideo LLM6G SecuritySemantic Communications
J
Junyi Zhu
Beijing University of Posts and Telecommunications, China
J
Jiayu Huang
Beijing University of Posts and Telecommunications, China
M
Minghao Sun
Beijing University of Posts and Telecommunications, China
X
Xuancheng Zhu
Beijing University of Posts and Telecommunications, China
Yisong Chen
Yisong Chen
Associate Professor of Computer Science, Peking University
computer vision
Z
Zexian Wei
Beijing University of Posts and Telecommunications, China
Xiaofeng Tao
Xiaofeng Tao
Beijing University of Posts and Telecommunications
wireless communication