Cognitive Dual-Process Planning for Autonomous Driving with Structured Scene Knowledge and Verifiable Reasoning-Action Consistency

📅 2026-07-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses key limitations in high-level autonomous driving planning—namely inaccurate scene understanding, redundant reasoning, and misalignment between reasoning and action—particularly when relying on vision-language models that lack efficient, verifiable structured decision-making mechanisms. The authors propose a cognitive dual-path planning framework that represents scene knowledge via Structured Chain-of-Thought (S-CoT), leverages an automated data engine to generate unlabeled supervision signals, and dynamically routes decisions based on scene complexity to either a fast meta-action prediction path or a slow structured reasoning path. A deterministic rule-based verifier ensures logical consistency between reasoning and actions. Integrating dual-process cognition with verifiable reasoning-action constraints, the method achieves 91.8% S-CoT accuracy and 98.5% consistency across 195 human-evaluated scenarios, and 80.14% planning accuracy with 97.20% consistency on 574 NAVSim samples, while reducing average latency by 17.39%.
📝 Abstract
High-level planning for autonomous driving is a knowledge-intensive engineering decision task that requires accurate scene understanding, timely inference, and internally consistent action selection. Vision-language models (VLMs) can make intermediate reasoning explicit, but their use in deployed planners is constrained by costly structured supervision, unnecessary reasoning in routine scenes, and possible inconsistencies between generated rationales and driving actions. We present a cognitive dual-process planning framework that represents planning-relevant scene knowledge in a machine-parsable structured chain-of-thought (S-CoT) schema. An automated data engine integrates perception foundation models, critical-path filtering, and an expert VLM to generate S-CoT supervision without manual annotation of individual rationales. A lightweight visual Arbiter estimates scene complexity from multilevel vision-encoder features before language decoding and routes each input to either fast meta-action prediction or slow structured reasoning. For slow-path outputs, a deterministic rule-based validator checks whether the parsed S-CoT fields are consistent with the final meta-action and provides verifiable rewards for Group Relative Policy Optimization (GRPO). In a 195-scene manual audit, the generated annotations achieve 91.8\% CoT accuracy and a 98.5\% Logical Consistency Score (LCS). On 574 manually verified NAVSIM test samples, the planner achieves 80.14\% planning accuracy and 97.20\% LCS while reducing average latency by 17.39\% relative to applying slow reasoning to every scene. Evaluation on external long-tail subsets further identifies conditions under which routing and planning performance degrade. Together, these results show how explicit scene knowledge can be operationalized through adaptive reasoning and rule-based verification to support high-level VLM planning decisions.
Problem

Research questions and friction points this paper is trying to address.

autonomous driving
high-level planning
reasoning-action consistency
structured supervision
vision-language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

structured chain-of-thought
cognitive dual-process planning
verifiable reasoning-action consistency
automated data engine
visual arbiter
🔎 Similar Papers
No similar papers found.
Z
Zhongyao Yang
aSchool of Mechanical Engineering, Beijing Institute of Technology, Beijing, 100081, Beijing, China; bNational Engineering Research Center of Electric Vehicles, Beijing Institute of Technology, Beijing, 100081, Beijing, China
Haoyu Li
Haoyu Li
National Institute of Informatics
speech processing
Y
Yu Yan
aSchool of Mechanical Engineering, Beijing Institute of Technology, Beijing, 100081, Beijing, China; bNational Engineering Research Center of Electric Vehicles, Beijing Institute of Technology, Beijing, 100081, Beijing, China
Z
Zhuangxuan Yu
aSchool of Mechanical Engineering, Beijing Institute of Technology, Beijing, 100081, Beijing, China; bNational Engineering Research Center of Electric Vehicles, Beijing Institute of Technology, Beijing, 100081, Beijing, China
J
Jiangfeng Nan
cSchool of Mechanical Engineering, Southeast University, Nanjing, 211189, Jiangsu, China
J
Jinrui Nan
aSchool of Mechanical Engineering, Beijing Institute of Technology, Beijing, 100081, Beijing, China; bNational Engineering Research Center of Electric Vehicles, Beijing Institute of Technology, Beijing, 100081, Beijing, China