Anchoring What Matters: A Dual-Level Learning Framework for Visually-Grounded Multimodal Reasoning

📅 2026-09-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决视觉基础推理行为强化问题,提出PIVOT框架,通过自校准经验回放和视觉引导的优势分配机制来优化策略。
📝 Abstract
Reinforcement learning with verifiable rewards (RLVR) has significantly improved the reasoning capabilities of large vision-language models (LVLMs). However, standard on-policy RLVR algorithms face a critical optimization bottleneck in preserving and reinforcing visually grounded reasoning behaviors: valuable visually-grounded reasoning trajectories are discarded after a single update, while uniform token advantage allocation prevents the model from reinforcing critical perception or reasoning steps. To bridge this gap, we propose PIVOT, a dual-level learning framework that anchors policy optimization around informative visual reasoning signals. Specifically, PIVOT introduces a self-calibrated experience replay mechanism, which selectively collects and replays visually-grounded historical experiences as stable reference anchors for policy optimization. Building upon this, we further design a vision-guided advantage allocation mechanism to allocate additional vision-aware advantages to tokens based on their local visual support and impact on downstream reasoning. Extensive experiments across diverse benchmarks demonstrate that PIVOT achieves highly competitive performance in enhancing the multimodal reasoning capabilities of LVLMs.
Problem

Research questions and friction points this paper is trying to address.

Reinforcement Learning
Visually-Grounded Reasoning
Large Vision-Language Models
Policy Optimization
Experience Replay
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dual-Level Learning Framework
Self-Calibrated Experience Replay
Vision-Guided Advantage Allocation
🔎 Similar Papers
No similar papers found.
X
Xinxin Song
Department of Automation, Tsinghua University
S
Siyuan Li
Department of Automation, Tsinghua University
Tingxiong Xiao
Tingxiong Xiao
Tsinghua University
XAIDeep LearningAI4Science
Jinli Suo
Jinli Suo
Tsinghua University
Computer VisionComputational PhotographyComputational Imaging