Distilling Future Temporal Knowledge with Masked Feature Reconstruction for 3D Object Detection
To address the challenge of incorporating future-frame information in online 3D object detection, this paper proposes a Future Temporal Knowledge Distillation (FTKD) framework that relaxes the strict frame-wise alignment constraint inherent in conventional knowledge distillation. Methodologically, FTKD introduces a sparse query mechanism and a future-aware feature reconstruction strategy, jointly optimized with foreground-background contextual modeling to efficiently extract and transfer future temporal knowledge from an offline teacher model to an online student. Additionally, future-guided logit distillation is incorporated to enhance the student’s modeling capability for motion dynamics and temporal consistency. Evaluated on the nuScenes dataset, FTKD achieves consistent improvements of +1.3 mAP and +1.3 NDS over strong baselines, with notable gains in velocity estimation accuracy—while incurring no additional computational overhead during online inference.