COST: Contrastive One-Stage Transformer for Vision-Language Small Object Tracking

📅 2025-04-02
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
To address redundancy in multi-stage fusion, cross-modal representation inconsistency, and the difficulty of tracking small targets (<32×32) due to weak visual appearance in vision-language (VL) tracking, this paper proposes COST: a Contrastive One-stage Transformer fusion framework. COST employs an end-to-end single-stage architecture integrated with contrastive alignment and mutual information maximization to achieve semantically consistent modeling between video and language descriptions. We introduce VL-SOT500—the first dedicated benchmark for small-object VL tracking—comprising VL-SOT230 and VL-SOT270 subsets. Extensive experiments demonstrate that linguistic cues significantly enhance weak visual feature representation. COST achieves state-of-the-art performance across five mainstream VL tracking benchmarks and VL-SOT500, with particularly notable improvements in small-target tracking accuracy.

Technology Category

Application Category

📝 Abstract
Transformer has recently demonstrated great potential in improving vision-language (VL) tracking algorithms. However, most of the existing VL trackers rely on carefully designed mechanisms to perform the multi-stage multi-modal fusion. Additionally, direct multi-modal fusion without alignment ignores distribution discrepancy between modalities in feature space, potentially leading to suboptimal representations. In this work, we propose COST, a contrastive one-stage transformer fusion framework for VL tracking, aiming to learn semantically consistent and unified VL representations. Specifically, we introduce a contrastive alignment strategy that maximizes mutual information (MI) between a video and its corresponding language description. This enables effective cross-modal alignment, yielding semantically consistent features in the representation space. By leveraging a visual-linguistic transformer, we establish an efficient multi-modal fusion and reasoning mechanism, empirically demonstrating that a simple stack of transformer encoders effectively enables unified VL representations. Moreover, we contribute a newly collected VL tracking benchmark dataset for small object tracking, named VL-SOT500, with bounding boxes and language descriptions. Our dataset comprises two challenging subsets, VL-SOT230 and VL-SOT270, dedicated to evaluating generic and high-speed small object tracking, respectively. Small object tracking is notoriously challenging due to weak appearance and limited features, and this dataset is, to the best of our knowledge, the first to explore the usage of language cues to enhance visual representation for small object tracking. Extensive experiments demonstrate that COST achieves state-of-the-art performance on five existing VL tracking datasets, as well as on our proposed VL-SOT500 dataset. Source codes and dataset will be made publicly available.
Problem

Research questions and friction points this paper is trying to address.

Improves vision-language tracking with one-stage transformer fusion
Addresses cross-modal alignment for consistent feature representation
Introduces a new dataset for small object tracking
Innovation

Methods, ideas, or system contributions that make the work stand out.

Contrastive alignment strategy for cross-modal fusion
One-stage transformer for unified VL representations
VL-SOT500 dataset for small object tracking
💼 Related Jobs
No related jobs found.
C
Chunhui Zhang
Cooperative Medianet Innovation Center, Shanghai Jiao Tong University, Shanghai, 200240, China; The Hong Kong University of Science and Technology (Guangzhou), Guangzhou, 511458, China; CloudWalk Technology Co., Ltd, Shanghai, 201203, China
L
Li Liu
The Hong Kong University of Science and Technology (Guangzhou), Guangzhou, 511458, China
Jialin Gao
Jialin Gao
National University of Singapore
Video Understanding Multi-modal Understanding
X
Xin Sun
Cooperative Medianet Innovation Center, Shanghai Jiao Tong University, Shanghai, 200240, China
H
Hao Wen
CloudWalk Technology Co., Ltd, Shanghai, 201203, China
X
Xi Zhou
CloudWalk Technology Co., Ltd, Shanghai, 201203, China
Shiming Ge
Shiming Ge
Institute of Information Engineering, Chinese Academy of Sciences
Computer VisionArtificial Intelligence
Yanfeng Wang
Yanfeng Wang
Shanghai Jiao Tong University