DH-VLM: Dual-Horizon Cooperative Latent Reasoning for Autonomous Driving

📅 2026-08-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of单车智能 in long-horizon decision-making—namely restricted perception range, occlusion, and high computational overhead—and the lack of deep semantic reasoning in existing cooperative driving approaches. To overcome these challenges, the authors propose a dual-temporal cooperative implicit reasoning framework that enables asymmetric semantic collaboration between infrastructure and vehicles, integrating global reasoning guidance with local planning. Key innovations include an infrastructure-driven implicit state evolution mechanism, multi-layer implicit state aggregation, and the first question-answering dataset tailored for cooperative driving, which supports counterfactual and safety-aware reasoning. Experiments demonstrate that the proposed method reduces L2 error by 14.6%, decreases collision rate by 26.9%, lowers communication overhead by 57.3%, reduces GPU memory usage by 25.5%, and exhibits strong robustness to erroneous infrastructure guidance.
📝 Abstract
Large-scale language models for autonomous driving enable enhanced global understanding and long-horizon planning. However, when deployed in isolated vehicles, limited sensing range and occlusions restrict reliable decision-making, and the substantial computational and latency overhead makes on-board deployment impractical. Cooperative driving provides a potential solution by leveraging external agents for information exchange, but existing methods remain limited in semantic reasoning capability under practical constraints. To address these challenges, we propose DH-VLM, a dual-horizon cooperative latent reasoning framework that enables asymmetric semantic cooperation between the infrastructure and ego vehicle. The infrastructure aggregates multi-layer hidden states to form a global-reasoning horizon latent guidance, which is integrated into the ego model through an Infrastructure-Driven Latent Evolution mechanism for conditional latent refinement. This enables the ego vehicle to leverage long-range contextual understanding while preserving autonomous decision-making within its local planning horizon. Furthermore, we construct a cooperation-oriented question-answer (QA) dataset covering fundamental scene understanding and ego-personalized comprehension to support counterfactual and safety-aware reasoning. Extensive experiments demonstrate that DH-VLM achieves state-of-the-art planning performance, outperforming the previous state of the art by 14.6% in L2 error and 26.9% in collision rate. Compared with query-based end-to-end cooperative driving methods, our approach reduces the communication cost by 57.3% and GPU memory usage by 25.5%, while maintaining strong robustness against infrastructure guidance errors, providing a practical and robust paradigm for cooperative autonomous driving.
Problem

Research questions and friction points this paper is trying to address.

autonomous driving
cooperative perception
semantic reasoning
long-horizon planning
occlusion
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dual-Horizon Reasoning
Latent Guidance
Infrastructure-Vehicle Cooperation
Conditional Latent Refinement
Cooperative Autonomous Driving
Z
Ziyi Song
Department of Electronic Engineering, Tsinghua University
C
Chen Xia
Department of Electronic Engineering, Tsinghua University
H
Hang Yu
Department of Electronic Engineering, Tsinghua University
S
Sheng Zhou
Department of Electronic Engineering, Tsinghua University; State Key Laboratory of Intelligent Green Vehicle and Mobility, Tsinghua University
Zhisheng Niu
Zhisheng Niu
Professor of Electronic Engineering, Tsinghua University
Green CommunicationRadio Resource ManagementQueueing Theory