ResWorld: Temporal Residual World Model for End-to-End Autonomous Driving
Existing end-to-end autonomous driving world models suffer from redundant modeling of static regions and insufficient interaction between trajectories and scene dynamics, limiting planning performance. To address these issues, this work proposes the Temporal Residual World Model (TR-World), which directly extracts dynamic object information through temporal residuals without relying on explicit detection or tracking, and predicts high-fidelity future bird’s-eye-view (BEV) representations by leveraging current BEV features. Furthermore, a Future-Guided Trajectory Refinement module (FGTR) is introduced to enable bidirectional co-optimization between trajectories and future scene context, while sparse spatiotemporal supervision is employed to prevent training instability. Evaluated on nuScenes and NAVSIM, the proposed approach significantly improves planning accuracy and robustness, achieving state-of-the-art performance.