Long-Horizon Consistent and Interaction-Aware World Models for Multi-Style End-to-End Driving

📅 2026-09-02
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出StyleDrive框架,通过增强长期一致性、分离交互状态和支持多种驾驶风格优化,解决现有世界模型在长时预测、环境互动建模及适应性上的局限。
📝 Abstract
End-to-end autonomous driving has increasingly adopted world model-based reinforcement learning frameworks to improve learning efficiency through \textit{imagined rollouts}. However, existing world models suffer from three key limitations: temporal inconsistency in long-horizon imagined rollouts, inadequate modeling of ego-environment interactions, and limited adaptability to diverse driving styles. To address these challenges, we propose \textit{StyleDrive}, a world-model-based learning framework that jointly enforces long-horizon consistency, explicitly disentangles interactive traffic states, and supports multi-style policy optimization within a unified learning paradigm. First, we introduce a temporal consistency regularization that integrates historical latent states through gated cross-attention, stabilizing long-horizon imagined rollouts and mitigating error accumulation. Second, we design an explicit state disentanglement module that separates ego-relevant from ego-irrelevant interactive states, enabling more interpretable and efficient decision-making in complex traffic scenarios. Third, we enable multi-style driving behaviors through Group Relative Policy Optimization, which replaces per-step reward optimization with trajectory-wise relative advantages, reducing reward variance and supporting diverse driving styles without retraining. We evaluate StyleDrive on the Bench2Drive closed-loop driving benchmark, achieving a driving score of 88.44 (+17.08 over the previous best world model-based method) and a success rate of 66.82 (+16.58). Furthermore, we deploy StyleDrive on a real automated guided vehicle platform and demonstrate promising sim-to-real transfer capability in dynamic driving scenarios.
Problem

Research questions and friction points this paper is trying to address.

long-horizon consistency
ego-environment interactions
multi-style driving
Innovation

Methods, ideas, or system contributions that make the work stand out.

long-horizon consistency
interaction-aware
multi-style policy optimization
temporal consistency regularization
state disentanglement
💼 Related Jobs
No related jobs found.
Yuxuan Han
Yuxuan Han
Tsinghua University
computer visioncomputer graphics
K
Kunyuan Wu
School of Intelligence Science and Engineering, Harbin Institute of Technology, Shenzhen, China
L
Liyunong Yang
School of Intelligence Science and Engineering, Harbin Institute of Technology, Shenzhen, China
Z
Zilu Wang
School of Intelligence Science and Engineering, Harbin Institute of Technology, Shenzhen, China
Cansen Jiang
Cansen Jiang
PhD, Le2i, CNRS, Université de Bourgogne
Computer VisionSLAMCamera Calibration
Yi Xiao
Yi Xiao
Professor of Physics, Huazhong University of Science and Technology
BiophysicsNonlinear PhysicsPhysics
Liang Hu
Liang Hu
Professor, Harbin Institute of Technology, Shenzhen
State Estimation and SLAMNavigation and ControlAutonomous Systems