REVEAL: A Rubric-Guided Agent for Explicit Evidence Sufficiency Verificationin Long-Video Question Answering

📅 2026-08-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing long-form video question answering methods rely on fixed-duration chunking and static memory banks, which fragment continuous events and answer questions based solely on semantic relevance while neglecting evidential sufficiency. This often leads to the loss of critical temporal, causal, or fine-grained action information. To address these limitations, this work proposes REVEAL, a novel framework that adaptively constructs dynamic event units based on visual similarity, employs a hybrid offline-online memory architecture, and introduces—for the first time—an explicit evidential sufficiency verification mechanism. This mechanism leverages an automatically constructed rule-based scoring system to evaluate evidence completeness, identify missing cues, and guide targeted re-retrieval. Without requiring additional training, REVEAL consistently outperforms both open-source and closed-source state-of-the-art methods across multiple benchmarks, significantly enhancing the accuracy and reliability of long video question answering.
📝 Abstract
Recently, retrieval-augmented and memory-augmented methods have emerged as two promising paradigms for long-video question answering. However, existing methods typically rely on rigid, fixed-length temporal chunking (e.g., 10s) and static offline memory banks, which not only fragment coherent continuous events but also fail to adapt during real-time reasoning. Moreover, whether using multi-scale summaries or multimodal knowledge graphs, current approaches prioritize retrieval relevance while overlooking evidence sufficiency, often stopping to answer once only semantically relevant clues are retrieved, even when key temporal, causal, or fine-grained action evidence is still missing. To tackle these challenges, we propose REVEAL, a rubric-guided agent framework. As a foundation, we introduce an adaptive visual-similarity-based preprocessing pipeline that groups visually coherent adjacent frames into natural event units to construct an offline-online video memory---capturing global video context offline while dynamically maintaining question-conditioned memory online. Built upon this structured memory, REVEAL uses an automatically constructed rubric library to explicitly verify whether retrieved evidence satisfies sufficiency criteria, pinpoints missing clues upon verification failure, and directs targeted re-retrieval for complementary information. Without any extra training, REVEAL consistently outperforms both closed-source and open-source state-of-the-art methods across extensive experiments. These results show that explicitly verifying evidence sufficiency, rather than stopping at semantic relevance, retrieves the decisive clues that prior methods miss and yields more reliable long-video reasoning.
Problem

Research questions and friction points this paper is trying to address.

long-video question answering
evidence sufficiency
retrieval relevance
adaptive memory
temporal chunking
Innovation

Methods, ideas, or system contributions that make the work stand out.

evidence sufficiency verification
adaptive event chunking
rubric-guided reasoning
offline-online video memory
targeted re-retrieval
🔎 Similar Papers
C
Caijun Yan
Zhejiang University
Y
Yang Zhou
Zhejiang University
M
Meixing Shi
Zhejiang University
H
Haoran Sun
Shanghai AI Laboratory
Y
Yichen Li
Shanghai AI Laboratory
Y
Yuxiang Cai
Zhejiang University
Yankai Jiang
Yankai Jiang
Shanghai AI Laboratory
Multimodal LLMVision-Language PretrainingAI for Science