What Emerges and What Breaks in Self-Play Driving

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过自博弈训练自动驾驶策略,使用Transformer模型在真实城市高清地图上进行训练,分析了交通规则的自涌现情况及与人类驾驶行为的匹配度。
📝 Abstract
Training autonomous driving policies through pure self-play has recently shown promising results. Following Gigaflow and Puffer- Drive, we train driving policies in a similar self-play fashion, but extend the models from MLPs to Transformers and train on the high-definition map of a real city, where we ultimately aim to deploy them. On the CARLA and Waymax benchmarks, our policies fall short of Gigaflow, and we trace the gap to specific failure modes, including reward hacking at traffic lights and a missing incentive to stop at stop signs. We further analyze which traffic rules emerge from self-play and how closely they match human driving, and we confirm that reward conditioning yields the intended diversity of driving behaviors. A demonstration of a trained policy is available at https://laursisask-ut.github.io/eccvdemo.
Problem

Research questions and friction points this paper is trying to address.

Self-Play
Autonomous Driving
Traffic Rules
Innovation

Methods, ideas, or system contributions that make the work stand out.

Transformers
Self-Play
High-Definition Map
Reward Hacking
Diversity of Driving Behaviors
🔎 Similar Papers
2024-04-122024 IEEE Intelligent Vehicles Symposium (IV)Citations: 8