Unified Interactive Multimodal Moment Retrieval via Cascaded Embedding-Reranking and Temporal-Aware Score Fusion
To address three key challenges in video moment retrieval—poor cross-modal noise robustness, weak modeling of event temporal coherence, and heavy reliance on manual modality selection—this paper proposes an end-to-end interactive multimodal moment retrieval framework. Methodologically: (1) a cascaded dual-embedding re-ranking architecture enhances cross-modal alignment; (2) a temporal-aware exponentially decaying scoring mechanism explicitly models event continuity while suppressing implausible time intervals; (3) GPT-4o enables automatic query decomposition and adaptive score fusion, eliminating manual modality specification. The system integrates BEIT-3, SigLIP, BLIP-2, and ASR/OCR features, with beam search for temporally constrained optimization. Experiments demonstrate substantial improvements in recall-precision trade-off under ambiguous queries, generation of highly coherent event sequences, and enhanced practicality of video retrieval systems.