MeViS: A Multi-Modal Dataset for Referring Motion Expression Video Segmentation
Existing video segmentation datasets emphasize static attribute descriptions, neglecting the critical role of motion in video understanding. To address this, we introduce MeViS—the first multimodal video segmentation dataset explicitly guided by motion expression—comprising 33K human-annotated text and audio motion descriptions across 2,006 complex scenes and 8,171 objects, supporting four tasks: Referring Video Object Segmentation (RVOS), Audio-Visual Object Segmentation (AVOS), Referring Multi-Object Tracking (RMOT), and Referring Motion Expression Grounding (RMEG). MeViS pioneers motion semantics as the core referential cue, breaking the static-dominant paradigm. We further propose LMPM++, a model integrating multimodal aligned annotation, motion-aware modeling, and joint audio-visual-linguistic representation, achieving new state-of-the-art performance on RVOS, AVOS, and RMOT. Comprehensive evaluation of 15 mainstream methods reveals systematic motion reasoning bottlenecks; leveraging MeViS significantly improves segmentation and tracking accuracy, advancing motion-centric video understanding.