Efficient Language-to-Vision Feature Injection for Referring Single-Object Tracking

📅 2026-08-29
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出LVTrack,通过模式条件门控特征注入器和预训练模型来解决语言到视觉特征的有效注入问题,减少训练成本并提高单目标跟踪性能。
📝 Abstract
Referring single-object tracking enables language-grounded target initialization and subsequent tracking by jointly leveraging semantic cues and visual templates. The core difficulty is to use language differently across stages: it is indispensable for grounding but can induce semantic drift during tracking when overemphasized. Meanwhile, current methods often require costly vision-language alignment training. We present LVTrack, a pure transformer framework that introduces a mode-conditioned Gated Feature Injector to adaptively regulate textual guidance and alleviate semantic drift. Together with targeted adaptations, it directly harnesses a frozen vision-language pretrained model, greatly reducing training cost and preserving strong language understanding. To further improve temporal localization, LVTrack integrates hybrid relative-absolute positional encodings with a lightweight memory mechanism and optimizes autoregressive box prediction using a Gaussian-smoothed KL loss. Extensive experiments on standard benchmarks demonstrate that LVTrack achieves strong performance.
Problem

Research questions and friction points this paper is trying to address.

Referring single-object tracking
language-grounded target initialization
semantic drift
vision-language alignment training
Innovation

Methods, ideas, or system contributions that make the work stand out.

Gated Feature Injector
vision-language pretrained model
hybrid positional encodings
lightweight memory mechanism
Gaussian-smoothed KL loss
🔎 Similar Papers
H
Han Wang
Technology and Engineering Center for Space Utilization, Chinese Academy of Sciences; University of Chinese Academy of Sciences
Y
Yuxuan Liu
Technology and Engineering Center for Space Utilization, Chinese Academy of Sciences; University of Chinese Academy of Sciences
Yuhan Sun
Yuhan Sun
Ph.D. student of Computer Science, Arizona State Unviersity
GeoSpatial Graphdatabase
J
Jian Yang
Technology and Engineering Center for Space Utilization, Chinese Academy of Sciences; University of Chinese Academy of Sciences
X
Xiaotong Xu
Technology and Engineering Center for Space Utilization, Chinese Academy of Sciences; University of Chinese Academy of Sciences
Y
Yixuan Lv
Technology and Engineering Center for Space Utilization, Chinese Academy of Sciences
Z
Zhuang Zhou
Technology and Engineering Center for Space Utilization, Chinese Academy of Sciences
S
Shengyang Li
Technology and Engineering Center for Space Utilization, Chinese Academy of Sciences; University of Chinese Academy of Sciences