Report of the 8th LSVOS Challenge: Complex and Multimodal Video Object Segmentation

📅 2026-09-05
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
该报告总结了第8届大规模视频对象分割挑战赛,通过结合基础分割模型与多模态推理等方法解决了复杂和多模态视频对象分割问题。
📝 Abstract
This report summarizes the 8th Large-scale Video Object Segmentation (LSVOS) Challenge, held in conjunction with ECCV 2026. The challenge evaluates video segmentation in three complementary settings: complex semi-supervised video object segmentation on MOSEv2, text-guided referring video object segmentation on MeViSv2-Text, and audio-guided referring video object segmentation on MeViSv2-Audio. We describe the tasks and evaluation protocols and review the methods of the top three teams in each track. Across the nine leading solutions, foundation segmentation models are combined with target-aware memory, multimodal reasoning, explicit target-existence verification, agentic interaction, and corrective tracking. These systems illustrate a broader transition from single-model mask propagation toward modular pipelines that reason about object identity, query validity, and temporal reliability.
Problem

Research questions and friction points this paper is trying to address.

Video Object Segmentation
Semi-supervised
Text-guided
Audio-guided
Innovation

Methods, ideas, or system contributions that make the work stand out.

foundation segmentation models
multimodal reasoning
target-aware memory
🔎 Similar Papers
No similar papers found.