FocusTrack: One-Stage Focus-and-Suppress Framework for 3D Point Cloud Object Tracking
This work addresses the limitations of existing two-stage 3D point cloud object tracking methods, which rely on explicit foreground segmentation and consequently suffer from error accumulation and computational bottlenecks. To overcome these issues, we propose the first end-to-end single-stage tracking framework that jointly models motion and semantics without explicit segmentation, enabling both efficiency and accuracy. The core innovation lies in a focus-suppression attention mechanism, integrated with a temporal difference Siamese encoder to model inter-frame motion dynamics, thereby adaptively enhancing foreground features while suppressing background noise. Extensive experiments demonstrate that our method achieves state-of-the-art performance on major benchmarks—including KITTI, nuScenes, and Waymo—while running at an impressive inference speed of 105 FPS.