Too Many Frames, not all Useful: Efficient Strategies for Long-Form Video QA
Long-form video question answering (LVQA) suffers from high visual redundancy and sparse salient information, while existing methods inefficiently process uniformly sampled frames via independent vision-language model (VLM) descriptions, leading to poor semantic utilization. To address this, we propose the Hierarchical Keyframe Selector (HKFS), the first framework to jointly perform question-guided dynamic temporal segment localization and semantic keyframe selection. HKFS integrates multi-granularity temporal modeling, question-driven visual attention, and a lightweight VLM adaptation architecture—LVNet—to substantially reduce visual-language modeling overhead. Our approach achieves state-of-the-art performance on three major LVQA benchmarks—EgoSchema, NExT-QA, and IntentQA—and demonstrates strong generalization on VideoMME. Notably, it supports LVQA over videos up to one hour in length, enabling scalable, efficient, and semantically grounded long-video understanding.