RIDGE: Region-Informed Derivative-Guided Evidence Selection for Long Video Understanding

📅 2026-08-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为了解决长视频理解中帧选择问题,提出RIDGE框架,通过分析帧-查询相似度曲线的局部变化和曲率,将时间线划分为结构区域,并在固定预算下进行区域特定选择。
📝 Abstract
Long videos contain far more visual content than Large Vision-Language Models (LVLMs) can process under a fixed visual-token budget, making frame selection essential. Existing query-aware selectors usually estimate frame-query relevance and build a compact subset from high-scoring frames. Although their mechanisms differ, the similarity sequence is still often treated primarily as values to rank or sample from, rather than as an ordered signal whose shape reflects how query-relevant evidence emerges, peaks, and fades over time. This can obscure frames that explain, contextualize, or follow an event, because such evidence may lie on the rising or falling sides of a nearby relevance peak and receive lower absolute scores. We propose RIDGE, a frame selection framework that reads the frame-query similarity curve as a temporal signal. By using local changes and curvature, RIDGE partitions the timeline into structural regions and applies region-specific selection to preserve event cores, transitions, buildup, aftermath, and contextual frames under a fixed budget. It is a lightweight post-processing step on precomputed frame-query scores and requires neither training nor iterative LVLM calls. Across four long-video benchmarks and three backbones, RIDGE achieves the best performance in most settings and remains competitive in the others.
Problem

Research questions and friction points this paper is trying to address.

frame selection
long video understanding
query-aware selectors
visual-token budget
temporal signal
Innovation

Methods, ideas, or system contributions that make the work stand out.

Region-Informed Selection
Derivative-Guided Evidence
Temporal Signal Processing
Frame-Query Similarity
S
Shanqing Xu
Huazhong University of Science and Technology
Meng Luo
Meng Luo
National University of Singapore
Human-Centered AIMultimodal UnderstandingMultimodal Reasoning
M
Mengchen Qian
Huazhong University of Science and Technology
Y
Yuhui Gao
Huazhong University of Science and Technology
S
Siyue Peng
Huazhong University of Science and Technology
X
Xiaohan Zhong
Huazhong University of Science and Technology
X
Xiaojin Zhang
Huazhong University of Science and Technology
Z
Zhongyu Wei
Fudan University
W
Wei Chen
Huazhong University of Science and Technology
Xiang Bai
Xiang Bai
Huazhong University of Science and Technology (HUST)
Computer VisionOCR