MELON: A Large-Scale Dataset for Multi-Event Text-to-Long-Video Retrieval

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决长视频多事件文本检索问题,构建了MELON数据集,并提出了一种多事件感知损失函数以提高检索精度。
📝 Abstract
Existing text-video retrieval datasets primarily consist of short-form clips containing a single dominant event. While suitable for measuring basic vision-language alignment, they are limited in capturing real-world retrieval scenarios, where long-form videos naturally contain multiple semantically distinct events and a single text query may correspond to several non-contiguous temporal segments. To bridge this gap, we introduce MELON, the first large-scale dataset designed to extend text-video retrieval to long-form videos featuring complex, multi-event structures. MELON explicitly annotates multiple event intervals per video along with their corresponding textual descriptions, enabling both training and evaluation of multi-event understanding in long, untrimmed videos. In addition, we propose a multi-event aware loss that encourages models to differentiate between full-event and partial-event matches, yielding substantial improvements in retrieval accuracy. Together, the MELON dataset and our proposed loss establish a robust foundation for expanding text-to-video retrieval to complex long-form scenarios and provide a more realistic evaluation setting for future research in the field.
Problem

Research questions and friction points this paper is trying to address.

text-video retrieval
long-form videos
multi-event structures
real-world scenarios
Innovation

Methods, ideas, or system contributions that make the work stand out.

multi-event text-to-long-video retrieval
MELON dataset
multiple event intervals
multi-event aware loss
full-event and partial-event matches
🔎 Similar Papers