Segment Anything with Motion, Geometry, and Semantic Adaptation for Complex Nonlinear Visual Object Tracking

📅 2026-05-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limited generalization of existing visual object tracking methods, which rely heavily on task-specific supervision and struggle in complex scenarios involving occlusion, distractors, or nonlinear motion. To overcome these challenges, we propose SAMOSA, a novel framework built upon SAM 2 that introduces, for the first time, a lightweight nonlinear motion predictor, a semantic consistency verification mechanism, and a geometric structure constraint module. These components explicitly integrate motion, semantic, and geometric cues to enhance cross-frame consistency and dynamic modeling. By effectively combining implicit video priors with explicit tracking mechanisms, SAMOSA optimizes memory filtering and mask selection. Experimental results demonstrate that our method outperforms current SAM 2–based trackers on standard benchmarks, exhibits superior generalization compared to supervised visual object tracking models, and achieves significant performance gains on anti-drone datasets.
📝 Abstract
Traditional visual object tracking (VOT) methods typically rely on task-specific supervised training, limiting their generalization to unseen objects and challenging scenarios with distractors, occlusion, and nonlinear motion. Recent vision foundation models, exemplified by SAM 2, learn strong video understanding priors from large-scale pretraining and offer a promising foundation for building more robust and generalizable trackers. However, directly applying SAM 2 to VOT remains suboptimal, as it does not explicitly model target motion dynamics or enforce geometric and semantic consistency across frames, both of which are essential for reliable tracking. To address this issue, we propose SAMOSA, a new tracking framework that adapts SAM 2 to complex VOT scenarios by explicitly leveraging motion, geometry, and semantic cues. Specifically, we introduce a lightweight nonlinear motion predictor to model target dynamics and guide mask selection as well as memory filtering. We further exploit semantic cues to detect target shifts and recover from tracking failures, while geometric cues are incorporated as structural constraints to improve tracking stability. In this way, SAMOSA bridges the gap between the implicit video understanding prior of SAM 2 and explicit tracking-oriented modeling. Extensive experiments show that SAMOSA consistently outperforms state-of-the-art SAM 2--based approaches on general benchmarks, demonstrates stronger generalization than supervised VOT methods, and achieves substantial gains on anti-UAV datasets, which typify complex nonlinear motion scenarios. Our code is available at https://github.com/DurYi/SAMOSA.
Problem

Research questions and friction points this paper is trying to address.

visual object tracking
nonlinear motion
occlusion
generalization
foundation models
Innovation

Methods, ideas, or system contributions that make the work stand out.

motion modeling
geometric constraints
semantic adaptation
visual object tracking
foundation model adaptation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
D
Deyi Zhu
Shenzhen International Graduate School, Tsinghua University, Shenzhen 518055, China
Yuji Wang
Yuji Wang
Tsinghua University
CVMultimodalSegmentationMLLM
Yong Liu
Yong Liu
Tsinghua University
Video SegmentationMultimodal SegmentationComputer Vision
Y
Yansong Tang
Shenzhen International Graduate School, Tsinghua University, Shenzhen 518055, China
B
Bingyao Yu
Department of Automation, Tsinghua University, Beijing 100084, China
J
Jiwen Lu
Department of Automation, Tsinghua University, Beijing 100084, China
Jie Zhou
Jie Zhou
Tsinghua University
Graph Neural NetworksNatural Language Processing