MAETrack: Unleashing the Potential of Pretrained Geometric Priors for 3D Single Object Tracking

📅 2026-09-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对3D单目标跟踪中预训练几何先验的迁移问题,提出MAETrack框架,通过分层选择性初始化和几何残差门控方法有效改善了从3D重建到3D跟踪的任务适应。
📝 Abstract
Large-scale pre-training has transformed representation learning in 2D vision, yet its transferability to 3D single object tracking (SOT) remains insufficiently understood. Directly fine-tuning self-supervised 3D encoders, such as masked autoencoders (MAE), often leads to sub-optimal adaptation because the reconstruction objective is not fully aligned with the spatial-temporal matching requirements of tracking. In this paper, we observe that this difficulty can be interpreted as a layer-wise transfer mismatch: shallow layers tend to preserve transferable geometric cues, while deeper layers become increasingly specialized to the reconstruction pretext task and are less suitable for downstream tracking. Based on this observation, we propose MAETrack, a lightweight adaptation framework for transferring pre-training MAE representations to 3D SOT. MAETrack includes Layer-Selective Initialization (LSI), which initializes only the shallow stages of the tracking backbone from pre-trained weights while re-initializing deeper stages, and Geometric Residual Gating (GRG), which reinforces structurally salient regions in the search BEV features before template-search fusion through residual spatial modulation. Extensive experiments on standard 3D SOT benchmarks show that MAETrack consistently improves upon vanilla fine-tuning baselines with limited computational overhead. More broadly, our results suggest that effective transfer from 3D reconstruction pre-training to 3D tracking is not merely a matter of partial fine-tuning, but depends on a tracking-oriented transfer principle that preserves shallow geometry while adapting deeper representations to the downstream objective.
Problem

Research questions and friction points this paper is trying to address.

3D Single Object Tracking
Pretrained Models
Masked Autoencoders
Layer-wise Transfer Mismatch
Representation Learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Layer-Selective Initialization
Geometric Residual Gating
3D Single Object Tracking
Pretrained Geometric Priors
Transfer Learning
🔎 Similar Papers
2023-09-05International Joint Conference on Artificial IntelligenceCitations: 5
💼 Related Jobs
No related jobs found.
Sifan Zhou
Sifan Zhou
Southeast University
RoboticsM/LLMsSpatial AIQuantization
Qiwei Wang
Qiwei Wang
ShanghaiTech University
computer vision
L
Linyue Tan
University of Pennsylvania, Philadelphia, PA, 19104, USA
Z
Ziyu Liu
University of Pennsylvania, Philadelphia, PA, 19104, USA
Ziyu Zhao
Ziyu Zhao
University of South Carolina
computer vision. 2D/3D segmentationGenerative 3D reconstruction
X
Xiaobo Lu
School of Automation, Southeast University, Nanjing, 210096, China; Key Laboratory of Measurement and Control of Complex Systems of Engineering, Ministry of Education, Nanjing, 210096, China