Long-to-Short Video Evidence Reasoning for Grounded Question Answering

📅 2026-09-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
该研究提出LOVER模型,通过长到短视频证据课程学习、GQA奖励和自适应时间戳渲染方法,解决了基于强化学习的视频推理问题。
📝 Abstract
We present LOVER, a \underline{L}ong to sh\underline{O}rt \underline{V}ideo \underline{E}vidence \underline{R}einforced model for grounded question answering (GQA). LOVER highlights three innovations over existing reinforcement-learning (RL) based video reasoning models: (1) \textbf{Long-to-short Video Evidence Curriculum Learning}, which organizes RL training according to evidence duration and progressively adapts the model from long-range grounding to short-term reasoning; (2) \textbf{GQA Rewards}, which underscore the benefit of IoP reward over IoU for evidence spotting rather than strict temporal span overlap; (3) \textbf{Adaptive Timestamp Rendering}, which adaptively renders timestamps onto video frames using background-aware position and color selection to enhance temporal observability. The three designs are model-agnostic and reciprocal. They effectively improve QA, grounding, and grounded QA performance over different backbones. Notably, LOVER built on Time-R1 achieves new state-of-the-art (SOTA) results among open-source models on popular GQA benchmarks: NExT-GQA and ReXTime. Comprehensive ablation studies further validate the effectiveness of our three innovative components.
Problem

Research questions and friction points this paper is trying to address.

Long-to-Short Video Evidence
Grounded Question Answering
Reinforcement Learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Long-to-short Video Evidence Curriculum Learning
GQA Rewards
Adaptive Timestamp Rendering
K
Kaiyan Chen
University of Science and Technology of China
Junbin Xiao
Junbin Xiao
National University of Singapore
Video and LanguageEmbodied InteractionTrustworthy Multimodality
X
Xun Yang
University of Science and Technology of China