From Prompts to Pavement Through Time: Temporal Grounding in Agentic Scene-to-Plan Reasoning
Current large model–based approaches to autonomous driving scene understanding and planning lack effective temporal modeling, leading to inconsistent reasoning over sequential actions and compromising both safety and interpretability. To address this, this work proposes three multi-agent planner architectures incorporating varying degrees of temporal conditioning constraints. The authors establish the first empirical benchmark for temporally aware scene-to-planning reasoning on a subset of BDD-X and introduce evaluation metrics assessing semantic, syntactic, and logical consistency. Experimental results show that while explicit temporal constraints do not significantly improve standard NLP metrics, qualitative analysis reveals their capacity to elicit forward-looking risk assessment, stabilize corrective behaviors, and enhance strategic diversity. The study also highlights limitations in current prompt engineering practices regarding temporal grounding.