MarKey: Marginal Utility Guided Greedy Keyframe Selection for Long Video Understanding

📅 2026-09-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为了解决长视频理解中计算成本高和证据覆盖不全的问题,提出MarKey方法,通过边际效用引导的贪婪关键帧选择来提高效率和性能。
📝 Abstract
Long-video understanding remains challenging for multimodal large language models (MLLMs) because densely encoding long frame sequences is computationally expensive, while uniform sampling under a limited visual budget can miss sparse yet decisive evidence. Recent training-free keyframe selection methods have enabled more efficient inference and yielded promising performance gains. However, many existing methods score frames largely in isolation without explicitly considering how each candidate complements the currently selected subset, potentially resulting in redundant selections and incomplete evidence coverage. To address this limitation, we propose MarKey, a training-free framework that formulates keyframe selection as subset-aware greedy optimization. At each iteration, MarKey scores each candidate using a tractable surrogate that jointly accounts for query relevance, marginal coverage gain, and context-dependent redundancy, and selects the frame with the highest utility. To make this iterative subset-aware evaluation efficient, MarKey uses a compact set of representative anchors to approximate full-video coverage and a bounded window of previously selected frames to limit context-dependent comparisons. Experiments on six benchmarks spanning holistic video understanding, human-centric video understanding, and open-ended video understanding demonstrate that MarKey consistently outperforms existing methods. Further analyses show robust gains across different MLLM backbones, model scales, and frame budgets.
Problem

Research questions and friction points this paper is trying to address.

Long-video understanding
Multimodal large language models
Keyframe selection
Evidence coverage
Redundant selections
Innovation

Methods, ideas, or system contributions that make the work stand out.

subset-aware greedy optimization
marginal coverage gain
context-dependent redundancy
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
H
Hongchang Shi
Hefei University of Technology, Hefei, China
Jinpeng Hu
Jinpeng Hu
Hefei University of Technology
natural language processingnamed entity recognitionsummarization
A
Ao Wang
Hefei University of Technology, Hefei, China
W
Wenzheng Zhou
Hefei University of Technology, Hefei, China
H
Hui Ma
Hefei University of Technology, Hefei, China
Feng Li
Feng Li
Hefei University of Technology
Video and Image ProcessingLow-level Computer-vision
Zenglin Shi
Zenglin Shi
Professor of Artificial Intelligence, Hefei University of Technology
Deep LearningComputer VisionMachine LearningMultimedia