Pictura: Perspective-View Self-Play at Scale for Driving
This work addresses the representational gap that arises when driving policies trained with privileged state information are deployed using only first-person visual inputs. To bridge this gap, the authors propose the first purely vision-based self-play training framework that learns driving policies end-to-end directly from agent-centric images. Leveraging the GPU-accelerated multi-agent simulator Pictura and the PPO algorithm, the method achieves highly efficient training—processing 500,000 agent steps (equivalent to 2 million images) per second on a single H100 GPU. The resulting policy, Alberti, trained over 50 billion agent steps (approximately 35 million kilometers), closely matches the performance of privileged-observation baselines and demonstrates zero-shot superiority on re-rendered Waymo scenarios, effectively closing the perception gap between simulation and real-world deployment.