CRUISE: Vision-Language Model-Guided Uncertainty-Aware Cross-Modal Sensor Fusion for Robust Autonomous Driving

📅 2026-08-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of robust cross-modal fusion in autonomous driving under complex and out-of-distribution scenarios, where reliability discrepancies among multi-sensor modalities degrade perception performance. To this end, the paper proposes CRUISE, a novel framework that, for the first time, leverages vision-language models (VLMs) to guide pixel-level uncertainty quantification and incorporates a dynamic adaptive mechanism to explicitly model cross-modal complementarity. This enables fine-grained, reliability-aware fusion that adapts to varying environmental conditions. Experimental results demonstrate that CRUISE significantly outperforms existing uncertainty-aware fusion methods under adverse weather and low-visibility conditions, achieving enhanced perception robustness and generalization capability.
📝 Abstract
Modern autonomous vehicles are equipped with multiple sensors, such as cameras, LiDAR, and radar, for comprehensive environmental perception. However, robust cross-modal feature fusion remains a critical challenge, as the reliability of each sensor varies significantly across diverse real-world driving conditions, including poor visibility and adverse weather. While uncertainty quantification (UQ) mitigates this issue by allowing models to prioritize reliable signals, existing uncertainty-aware fusion methods typically rely on simple feature-level uncertainty estimates and thus often fail to generalize effectively in complex, out-of-distribution scenarios. To address this limitation, we propose CRUISE, a novel uncertainty-aware cross-modal sensor fusion framework. CRUISE integrates a vision-language model (VLM)-guided UQ module that generates fine-grained, pixel-level uncertainty estimates. By leveraging the VLM's rich prior knowledge and superior contextual reasoning, our approach provides a highly informative guide for the fusion process. Furthermore, we introduce a dynamic adaptive mechanism that explicitly models and captures cross-modal dependencies, ensuring the framework fully exploits the inherent complementary nature of multi-sensor inputs.
Problem

Research questions and friction points this paper is trying to address.

cross-modal sensor fusion
uncertainty quantification
autonomous driving
robust perception
out-of-distribution generalization
Innovation

Methods, ideas, or system contributions that make the work stand out.

vision-language model
uncertainty quantification
cross-modal fusion
autonomous driving
sensor fusion
🔎 Similar Papers
No similar papers found.