SAM3Dual: A 3rd Place Solution to the MOSEv2 Track, 8th LSVOS Challenge

📅 2026-08-22
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出了SAM3Dual,通过重组时间记忆在推理阶段提高长期视频对象分割性能,无需额外训练,实现了64.37的J&F得分,在MOSEv2赛道中排名第三。
📝 Abstract
We present SAM3Dual, our third-place solution to the MOSEv2 track of the 8th Large-scale Video Object Segmentation (LSVOS) Challenge at ECCV 2026. SAM3Dual is a training-free inference extension of pretrained SAM 3 that explicitly separates temporal memory into a short-term branch for recent observations and a long-term branch for interval-sampled historical representations. The two memory responses are combined using a deterministic sequence-relative fusion schedule and conservatively modulated by the previous-frame object confidence. All pretrained SAM 3 parameters remain frozen, requiring no task-specific training, fine-tuning, test-time training, or online parameter optimization. The complete system achieved an official J&F score of 64.37 and ranked third in the MOSEv2 track. This result highlights the potential of reorganizing temporal memory entirely at inference time to obtain competitive long-term VOS performance while preserving the pretrained model.
Problem

Research questions and friction points this paper is trying to address.

Large-scale Video Object Segmentation
Temporal Memory
Pretrained Model
Inference Time
Innovation

Methods, ideas, or system contributions that make the work stand out.

training-free inference
temporal memory separation
sequence-relative fusion schedule
long-term VOS performance
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
J
JeongRae Kim
Chung-Ang University, Seoul, Republic of Korea
Chaehyun Kim
Chaehyun Kim
M.S. / Ph.D Student, KAIST
Deep LearningComputer Vision
C
Changwon Lim
Chung-Ang University, Seoul, Republic of Korea