RoPETR: Improving Temporal Camera-Only 3D Detection by Integrating Enhanced Rotary Position Embedding
This work addresses the critical limitation of inaccurate velocity estimation in vision-only temporal 3D object detection, which severely constrains NuScenes detection performance—particularly the NuScenes Detection Score (NDS). To tackle this, we propose a velocity-optimized enhanced Rotary Position Encoding (Rotary PE) that explicitly incorporates motion priors and strengthens cross-frame feature alignment and temporal motion representation. We further design an end-to-end trainable temporal fusion module, tightly integrated with the StreamPETR architecture built upon a ViT-L backbone. Crucially, our method improves velocity prediction accuracy without requiring additional sensors or post-processing. Evaluated on the NuScenes test set, it achieves a new state-of-the-art NDS of 70.86%, demonstrating that refined temporal position modeling is pivotal for accurate motion estimation in vision-only 3D detection.