Improving 3D Labeling in Self-Driving by Inferring Vehicle Information using Vision Language Models
This work addresses the challenges of error-prone and inefficient manual 3D vehicle annotation in complex autonomous driving scenarios, particularly under occlusion. To overcome these limitations, the study introduces vision-language models (VLMs) into the 3D annotation pipeline for the first time, leveraging zero-shot inference to predict vehicle make, model, and generation from cropped image regions and generate accurate initial 3D bounding box dimensions. By integrating iterative prompt engineering with Vehicle Make and Model Recognition (VMMR) techniques, the proposed method demonstrates strong generalization across both public and proprietary datasets. It significantly outperforms conventional LiDAR-assisted annotation approaches, markedly improving annotation accuracy while substantially reducing human effort, especially in mitigating annotation failures caused by occlusion.