Explainable deep learning improves human mental models of self-driving cars
The opacity of black-box motion planners in autonomous driving severely undermines human–machine collaborative safety, while existing eXplainable AI (XAI) methods remain largely confined to simulation or simplified scenarios, lacking real-road validation. To address this, we propose the Concept-Wrapping Network (CW-Net), the first approach enabling causally faithful, performance-preserving, and human-interpretable decision explanations within production-grade autonomous driving systems. CW-Net maps neural network outputs onto a semantically grounded driving concept space—e.g., “yielding” or “emergency evasive maneuver”—by jointly integrating causal reasoning and human cognitive modeling. Real-world deployment evaluations demonstrate that CW-Net significantly improves drivers’ prediction accuracy of vehicle behavior (+28.6%) and enhances response adaptability. This work establishes the first practically deployable, explanation-aware motion planning paradigm for trustworthy human–autonomy collaboration.