MEDR: Query-Independent Frame Selection via Multi-Signal Event Modeling and Dynamic Rescoring

📅 2026-08-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of poor reusability in query-dependent frame selection and information loss from uniform sampling in long video processing with multimodal large language models. We propose MEDR, a training-free and query-agnostic method that constructs a reusable set of fixed keyframes by integrating multi-signal event modeling across visual, motion, and textual modalities with a dynamic rescoring strategy to enhance information coverage. Experiments demonstrate that MEDR improves accuracy by 0.89% on Video-MME and 1.23% on LongVideoBench. By enabling efficient cross-query reuse of the same frame set, the proposed approach effectively balances inference efficiency with model performance, offering a robust solution for long video understanding.
📝 Abstract
Frame selection is a fundamental component of multimodal large language models, enabling long videos to be processed under limited visual-token and computational budgets. Uniform sampling preserves temporal coverage but may miss informative content that appears only briefly. To alleviate this limitation, query-dependent methods can retrieve question-relevant frames. However, because the selected frames depend on the current question, the same visual input cannot be directly shared across different questions, and frame selection must be repeated in multi-turn video dialogue. This motivates us to seek a query-independent frame selection method that preserves the reusability of a fixed visual input while improving the coverage of informative events beyond uniform sampling. We propose Multi-Signal Event Modeling and Dynamic Rescoring (MEDR), a training-free and query-independent frame selection method. Multi-Signal Event Modeling organizes complementary visual, motion, and text signals into signal-specific temporal events. Dynamic Rescoring then iteratively reevaluates each candidate relative to the current selected set, updating its score according to frame-level signal strength, additional event coverage, and temporal proximity. The resulting fixed frame set is constructed without observing the query and can be reused across different questions. On the standard benchmark evaluations, MEDR improves model accuracy by 0.63%-0.89% on Video-MME. On the long-video subset of LongVideoBench, it improves accuracy by up to 1.23% with Qwen3-VL-8B. MEDR further improves overall accuracy by 0.53%, while reusing exactly the same frame set for every question about a video.
Problem

Research questions and friction points this paper is trying to address.

Frame Selection
Query-Independent
Multimodal Large Language Models
Long Video Understanding
Visual Token Budget
Innovation

Methods, ideas, or system contributions that make the work stand out.

Query-Independent Frame Selection
Multi-Signal Event Modeling
Dynamic Rescoring
Training-Free
Long Video Understanding
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
X
Xinlei Pu
School of Computer Science and Technology, Soochow University
Weijie Shi
Weijie Shi
Hong Kong University of Science and Technology
Wen Yang
Wen Yang
Professor, School of Electronic Information, Wuhan University
Image ProcessingPattern RecognitionMachine Learning
Y
Yi Cao
School of Computer Science and Technology, Soochow University
H
Hao Chen
FiT, Tencent
Yuanjun Liu
Yuanjun Liu
Soochow University
Trajectory data
W
Wenwei Ding
Suzhou Rural Commercial Bank
Jia Zhu
Jia Zhu
Zhejiang Normal University
Artificial IntelligenceKnowledge GraphData QualityComputational Pedagogy
J
Jiajie Xu
School of Computer Science and Technology, Soochow University