M2Depth: Unifying Monocular Depth Foundation Priors with Multi-View Stereo
本文针对多视图立体视觉在未见过场景中的泛化问题,提出了一种结合单目深度基础模型和级联MVS的双向互精炼框架,提高了深度估计的完整性和细节。
本文针对多视图立体视觉在未见过场景中的泛化问题,提出了一种结合单目深度基础模型和级联MVS的双向互精炼框架,提高了深度估计的完整性和细节。
This work addresses the challenge that existing hybrid reasoning models struggle to dynamically allocate reasoning budgets, often leading to over-reasoning on simple problems and under-reasoning on complex ones. The authors propose a training-free adaptive routing mechanism that samples two zero-thought drafts and uses their consistency to decide whether to answer directly; if inconsistent, it predicts the required reasoning budget based on draft entropy. This approach achieves dynamic budget allocation for the first time without labeled data or gradient updates, leveraging the model’s own generated signals for decision-making and demonstrating compatibility across diverse model scales and architectures. Evaluated on mathematical and code reasoning tasks, the method improves accuracy by up to 9.0 and 22.5 percentage points while reducing reasoning tokens by 15–69% and 51–63%, respectively, confirming its efficiency and generality.
To address the degradation of end-to-end autonomous driving robustness caused by inter-camera viewpoint discrepancies, this paper proposes VR-Drive—a multi-view robust end-to-end framework. Methodologically, it introduces feedforward 3D Gaussian splatting for unsupervised novel-view synthesis, jointly optimizing 3D scene reconstruction and trajectory planning; it further designs a cross-view hybrid memory bank and a consistency distillation mechanism to enable online augmentation under sparse-view conditions and cross-view temporal modeling. Experiments on a custom-built multi-view benchmark demonstrate that VR-Drive significantly suppresses synthesis artifacts, markedly improving planning generalization and stability under unseen viewpoints. The framework establishes a novel paradigm for scalable, real-world deployment of end-to-end autonomous driving systems.
To address the degradation of camera-radar fusion performance in back-projection-based BEV transformation caused by image depth ambiguity, this paper proposes CRAB: a novel framework that (1) explicitly constrains image depth distributions using high-precision sparse depth priors from radar, thereby mitigating depth ambiguity during inverse projection; and (2) introduces a radar-context-enhanced cross-attention mechanism to achieve fine-grained alignment and fusion of image features with radar occupancy information directly in BEV space. CRAB jointly integrates inverse projection, view-specific feature aggregation, and spatially adaptive radar fusion into a single end-to-end trainable architecture for high-fidelity BEV representation learning. Evaluated on nuScenes, CRAB achieves 62.4% NDS and 54.0% mAP—setting the new state of the art among back-projection-based camera-radar fusion methods for 3D detection and semantic segmentation.
Existing online 3D Gaussian Splatting (3DGS) methods rely solely on sparse keyframes, resulting in incomplete scene coverage, geometric incompleteness, and a fundamental trade-off between modeling fidelity and scalability under real-time constraints. Method: We propose an online, end-to-end 3DGS framework for pure RGB video streams. It introduces a reconstruction-quality-driven adaptive view selection mechanism that jointly optimizes keyframes and dynamically selected non-keyframes. Integrating online SLAM, multi-view stereo matching, and incremental 3DGS training, the framework enables real-time, dynamic identification and efficient incorporation of informative non-keyframes. Contribution/Results: Evaluated on complex dynamic outdoor scenes, our method significantly improves reconstruction completeness and geometric accuracy. It outperforms state-of-the-art approaches in robustness and efficiency, achieving—for the first time—high-fidelity, full-coverage online 3DGS reconstruction under strict real-time constraints.
本文针对多视图立体视觉在未见过场景中的泛化问题,提出了一种结合单目深度基础模型和级联MVS的双向互精炼框架,提高了深度估计的完整性和细节。
This work addresses the challenge that existing hybrid reasoning models struggle to dynamically allocate reasoning budgets, often leading to over-reasoning on simple problems and under-reasoning on complex ones. The authors propose a training-free adaptive routing mechanism that samples two zero-thought drafts and uses their consistency to decide whether to answer directly; if inconsistent, it predicts the required reasoning budget based on draft entropy. This approach achieves dynamic budget allocation for the first time without labeled data or gradient updates, leveraging the model’s own generated signals for decision-making and demonstrating compatibility across diverse model scales and architectures. Evaluated on mathematical and code reasoning tasks, the method improves accuracy by up to 9.0 and 22.5 percentage points while reducing reasoning tokens by 15–69% and 51–63%, respectively, confirming its efficiency and generality.
To address the degradation of end-to-end autonomous driving robustness caused by inter-camera viewpoint discrepancies, this paper proposes VR-Drive—a multi-view robust end-to-end framework. Methodologically, it introduces feedforward 3D Gaussian splatting for unsupervised novel-view synthesis, jointly optimizing 3D scene reconstruction and trajectory planning; it further designs a cross-view hybrid memory bank and a consistency distillation mechanism to enable online augmentation under sparse-view conditions and cross-view temporal modeling. Experiments on a custom-built multi-view benchmark demonstrate that VR-Drive significantly suppresses synthesis artifacts, markedly improving planning generalization and stability under unseen viewpoints. The framework establishes a novel paradigm for scalable, real-world deployment of end-to-end autonomous driving systems.
To address the degradation of camera-radar fusion performance in back-projection-based BEV transformation caused by image depth ambiguity, this paper proposes CRAB: a novel framework that (1) explicitly constrains image depth distributions using high-precision sparse depth priors from radar, thereby mitigating depth ambiguity during inverse projection; and (2) introduces a radar-context-enhanced cross-attention mechanism to achieve fine-grained alignment and fusion of image features with radar occupancy information directly in BEV space. CRAB jointly integrates inverse projection, view-specific feature aggregation, and spatially adaptive radar fusion into a single end-to-end trainable architecture for high-fidelity BEV representation learning. Evaluated on nuScenes, CRAB achieves 62.4% NDS and 54.0% mAP—setting the new state of the art among back-projection-based camera-radar fusion methods for 3D detection and semantic segmentation.
Existing online 3D Gaussian Splatting (3DGS) methods rely solely on sparse keyframes, resulting in incomplete scene coverage, geometric incompleteness, and a fundamental trade-off between modeling fidelity and scalability under real-time constraints. Method: We propose an online, end-to-end 3DGS framework for pure RGB video streams. It introduces a reconstruction-quality-driven adaptive view selection mechanism that jointly optimizes keyframes and dynamically selected non-keyframes. Integrating online SLAM, multi-view stereo matching, and incremental 3DGS training, the framework enables real-time, dynamic identification and efficient incorporation of informative non-keyframes. Contribution/Results: Evaluated on complex dynamic outdoor scenes, our method significantly improves reconstruction completeness and geometric accuracy. It outperforms state-of-the-art approaches in robustness and efficiency, achieving—for the first time—high-fidelity, full-coverage online 3DGS reconstruction under strict real-time constraints.