Beyond the Beep: Scalable Collision Anticipation and Real-Time Explainability with BADAS-2.0
This work addresses the limited capability of advanced driver-assistance systems to anticipate collisions in long-tail, high-risk driving scenarios by proposing an efficient and interpretable real-time prediction framework. Leveraging the Nexar Atlas platform, the authors curate a large-scale dataset comprising 178,500 annotated video clips. The approach integrates V-JEPA2 self-supervised pretraining, active learning for targeted annotation of hazardous scenarios, edge-device-oriented knowledge distillation, and a vision-language model (BADAS-Reason) that generates object-level attention heatmaps alongside natural language rationales. The resulting models are compressed to 22–86 million parameters, achieving 7–12× inference speedup while preserving accuracy, thereby significantly enhancing both predictive performance and interpretability in long-tail safety-critical situations.