Beyond Similarity: Foundation Models as an Efficient Backbone for Training-Free Composed Video Retrieval

๐Ÿ“… 2026-09-09
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
ๆœฌๆ–‡ๆๅ‡บไธ€็งๆ— ้œ€่ฎญ็ปƒ็š„ๆ–นๆณ•๏ผŒ้€š่ฟ‡่‡ช้€‚ๅบ”่ฐƒๅบฆๅŸบ็ก€ๆจกๅž‹่ƒฝๅŠ›่งฃๅ†ณๅคง่ง„ๆจก่ง†้ข‘ๆฃ€็ดขไธญ็š„ๆ•ˆ็އไธŽ็ป†็ฒ’ๅบฆๆŽจ็†ไน‹้—ด็š„็Ÿ›็›พใ€‚
๐Ÿ“ Abstract
Composed video retrieval (CoVR) searches a gallery for the target video that realizes a natural-language modification of a source clip. However, at gallery scale, this creates a fundamental tension: compact embeddings enable efficient, reusable search but can miss the transient actions, state changes, and subtle constraints that demand fine-grained video reasoning, whereas applying large multimodal models uniformly sacrifices scalability. To address these limitations, we propose that frozen foundation models should instead occupy complementary roles, with inference depth adapted to query difficulty. Based on this premise, we introduce \methodname{}, a framework for training-free \methodexpansion{}. Specifically, a composed-query embedding first searches reusable video-only gallery representations; uncertain queries undergo bounded reranking and candidate expansion; ambiguous edits trigger target-description generation; and only close leading candidates reach multimodal verification. To support these roles, frame selection, spatial resolution, and time cues are adapted to each stage. Across complete target-gallery evaluations, our method reaches state-of-the-art performance among training-free approaches, with 89.55 and 93.43 R@1 on Dense-WebVid-CoVR and CoVR-R, respectively (with more than +35\% and +25\% absolute margins to the closest counterpart). These results show that adaptively orchestrating foundation-model capabilities can combine scalable retrieval with fine-grained reasoning without task-specific training. The source code and all relevant guidelines are available on https://github.com/demidovd98/CoVRAGE.
Problem

Research questions and friction points this paper is trying to address.

Composed Video Retrieval
Efficiency
Scalability
Fine-grained Reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

frozen foundation models
training-free retrieval
adaptive inference depth
fine-grained video reasoning
multimodal verification
๐Ÿ”Ž Similar Papers
2024-02-20International Conference on Machine LearningCitations: 30
Dmitry Demidov
Dmitry Demidov
PhD Student, Mohamed bin Zayed University of Artificial Intelligence
Deep LearningComputer VisionMachine LearningArtificial Intelligence
M
Muhammad Zaigham Zaheer
Mohamed bin Zayed University of Artificial Intelligence, UAE
Omkar Thawakar
Omkar Thawakar
MBZUAI,UAE
Computer VisionMachine LearningGenerative AILLMFoundation Models
A
Abdelrahman Mohamed Shaker
Mohamed bin Zayed University of Artificial Intelligence, UAE
R
Rao Anwer
Mohamed bin Zayed University of Artificial Intelligence, UAE