Unlocking Motion in Expressions: Temporal Calibration for Referring Video Object Segmentation

📅 2026-08-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of unmodeled motion-semantic dependencies and adaptive adjustment in referring video segmentation by proposing an Expression-Driven Motion Calibration framework. The method explicitly decouples motion semantics from textual expressions and employs motion signal processing alongside influence calibration modules to achieve adaptive weighting of motion cues. Furthermore, it constructs a compact semantic-temporal candidate space to optimize matching precision. Extensive experiments demonstrate that the proposed framework achieves superior performance across six mainstream benchmarks, including Ref-YouTubeVOS. These results confirm its effectiveness in significantly enhancing segmentation accuracy and robustness within complex scenarios, validating the efficacy of integrating explicit motion calibration with linguistic understanding for precise video object segmentation.
📝 Abstract
Referring Video Object Segmentation (RVOS) aims to segment referred objects at the pixel level in video sequences based on natural language descriptions. Existing methods typically introduce motion information within a unified cross-modal temporal modeling framework, where language cues are used for target localization and segmentation. However, the dependency of expressions on motion semantics is not explicitly modeled, making it difficult to adaptively adjust the use of motion information according to different semantic requirements. To address these issues, we propose an Expression-driven Motion Calibration (EMC) framework for RVOS that explicitly unlocks and leverages the motion semantics within expressions. The proposed method extracts interpretable motion control signals from expressions via a Motion Signal Processing (MSP) module, and employs a Motion Influence Calibration (MIC) module to adjust the contribution of motion cues during temporal decision making. In addition, a Semantic Temporal Stage Construction (STSC) module is introduced to build expression-relevant temporal stages, providing a compact temporal candidate space for motion calibration. Through extensive evaluation on six standard benchmarks, including Ref-YouTubeVOS, Ref-DAVIS17, MeViS (valid/valid$^u$), A2D-Sentences, and JHMDB-Sentences, the superiority of our method is validated. We will release the code on https://github.com/Jeven7/EMC.
Problem

Research questions and friction points this paper is trying to address.

Referring Video Object Segmentation
Motion Semantics
Temporal Calibration
Cross-modal Modeling
Innovation

Methods, ideas, or system contributions that make the work stand out.

Expression-driven Motion Calibration
Motion Signal Processing
Motion Influence Calibration
Semantic Temporal Stage Construction
Referring Video Object Segmentation
🔎 Similar Papers
No similar papers found.
Y
Yiwen Jiang
School of Computer Science & Technology, Soochow University, Suzhou, China
Z
Zhengtong Zhu
School of Computer Science & Technology, Soochow University, Suzhou, China
Ruixin Zhang
Ruixin Zhang
tencent
computer vision
J
Jiaqing Fan
School of Computer Science & Technology, Soochow University, Suzhou, China