BEVCon: Advancing Bird's Eye View Perception with Contrastive Learning

📅 2025-08-06
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
To address insufficient representation learning in BEV perception for autonomous driving, this paper proposes a dual-path contrastive learning framework that enforces semantic consistency at both the instance level (BEV features) and the view level (image backbone outputs). Specifically, it introduces an instance-feature contrastive module and a perspective-to-BEV contrastive module, integrated with dense pixel- and instance-level contrastive strategies. The approach enhances the discriminability and geometric robustness of BEV representations without increasing inference overhead. It seamlessly integrates into existing BEV detection pipelines by adding only a lightweight contrastive regularization term to the standard detection loss. Evaluated on the nuScenes dataset, the method achieves significant improvements across 3D detection, segmentation, and trajectory prediction—yielding up to +2.4% mAP gain over state-of-the-art baselines with negligible computational cost. This work provides the first systematic empirical validation that contrastive learning substantially improves BEV representation quality.

Technology Category

Application Category

📝 Abstract
We present BEVCon, a simple yet effective contrastive learning framework designed to improve Bird's Eye View (BEV) perception in autonomous driving. BEV perception offers a top-down-view representation of the surrounding environment, making it crucial for 3D object detection, segmentation, and trajectory prediction tasks. While prior work has primarily focused on enhancing BEV encoders and task-specific heads, we address the underexplored potential of representation learning in BEV models. BEVCon introduces two contrastive learning modules: an instance feature contrast module for refining BEV features and a perspective view contrast module that enhances the image backbone. The dense contrastive learning designed on top of detection losses leads to improved feature representations across both the BEV encoder and the backbone. Extensive experiments on the nuScenes dataset demonstrate that BEVCon achieves consistent performance gains, achieving up to +2.4% mAP improvement over state-of-the-art baselines. Our results highlight the critical role of representation learning in BEV perception and offer a complementary avenue to conventional task-specific optimizations.
Problem

Research questions and friction points this paper is trying to address.

Improving Bird's Eye View perception in autonomous driving
Enhancing representation learning in BEV models
Boosting feature quality via contrastive learning modules
Innovation

Methods, ideas, or system contributions that make the work stand out.

Contrastive learning for BEV perception
Instance feature contrast module
Perspective view contrast module
💼 Related Jobs
No related jobs found.
Z
Ziyang Leng
University of California, Los Angeles
J
Jiawei Yang
University of Southern California
Z
Zhicheng Ren
Aurora Innovation
Bolei Zhou
Bolei Zhou
Associate Professor at UCLA
Computer VisionRoboticsArtificial Intelligence