Vista: Scene-Aware Optimization for Streaming Video Question Answering under Post-Hoc Queries
This work proposes a scene-aware streaming video question answering framework to address context loss and memory overflow caused by the continuous arrival of video frames and arbitrary-time queries. By integrating dynamic scene segmentation, GPU/CPU-coordinated compressed storage, and a query-driven selective recall mechanism, the approach enables efficient, low-latency comprehension of long videos. Key innovations include dynamic video segmentation based on scene clustering, heterogeneous memory management, index-based retrieval, and a model-agnostic architecture. Evaluated on StreamingBench, the method achieves state-of-the-art performance, significantly enhancing system scalability and inference completeness.