Motion-Aware Reasoning from Speech to Mask Tracks: Runner-up Solution for the MeViS-Audio Track of the 8th LSVOS Challenge 2026

📅 2026-08-23
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出Speech2MaskTrack方法,通过语音识别、运动中心时间定位、掩模跟踪等步骤,解决语音引导的视频对象分割问题。
📝 Abstract
Speech-guided referring video object segmentation aims to recover the mask tracks of objects specified by a spoken motion description. Here, speech carries a linguistic instruction rather than acoustic evidence from a sounding object, so a solution must connect speech recognition, motion-centric temporal grounding, mask tracking, and explicit no-target handling. We introduce Speech2MaskTrack, our approach for the MeViS-Audio track of the 8th LSVOS Challenge. Speech2MaskTrack transcribes the spoken query and compiles it into structured constraints over category, count, direction, interaction role, and temporal phase. SAM3.1 enumerates multiple instance tracks, which TRACE ranks using complete-trajectory motion and relation evidence. A frozen lexical presence gate may suppress the ranked SAM3.1 base prediction. When the gate predicts that a target is present, an available full-expression-conditioned SaSaSa2VA track replaces the SAM3.1 mask. Only outputs that remain empty enter GPT-assisted recovery, which invokes SaSaSa2VA again under query- and mask-level verification. Speech2MaskTrack achieved second place in the official challenge ranking.
Problem

Research questions and friction points this paper is trying to address.

speech-guided
referring video object segmentation
mask tracks
motion description
temporal grounding
Innovation

Methods, ideas, or system contributions that make the work stand out.

Speech2MaskTrack
Structured Constraints
TRACE Ranking
SaSaSa2VA
GPT-assisted Recovery
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Jinxing Zhou
Mohamed bin Zayed University of Artificial Intelligence
S
Suiyi Zhao
Anhui University of Science and Technology
Y
Yanghao Zhou
National University of Singapore
Ruohao Guo
Ruohao Guo
Peking University
Multi-Modal LearningComputer VisionVideo Generation