Generating Causal Explanations of Vehicular Agent Behavioural Interactions with Learnt Reward Profiles
To address the lack of transparency and interpretability in causal explanations for autonomous driving human–machine interaction, this paper proposes a causal inference framework based on implicit reward modeling. Methodologically, it introduces a learnable reward profile as the core mediator for generating causal explanations, unifying inverse reinforcement learning with structural causal models and integrating multi-task optimization and differentiable causal discovery to enable counterfactual reasoning and semantically interpretable inference over multi-vehicle interactions. Experiments on three real-world driving datasets demonstrate that the method achieves state-of-the-art performance across key evaluation metrics—including explanation fidelity, consistency, and human comprehensibility—significantly outperforming existing baselines.