Evaluating 2D and 3D-Aware Vision Foundation Models for Vehicle Attribute Recognition

📅 2026-08-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过对比14种2D和3D感知视觉基础模型在车辆属性识别中的表现,发现2D自监督模型如DINOv3在细粒度任务上优于3D感知模型。
📝 Abstract
Vehicle attribute recognition is an important task in intelligent transportation systems, particularly when Automatic License Plate Recognition (ALPR) is unavailable or unreliable. Although vision foundation models have shown strong transferability across domains, their effectiveness for fine-grained vehicle classification remains underexplored. Moreover, given the inherently three-dimensional structure of vehicles, it is unclear whether emerging 3D-aware foundation models offer advantages over standard 2D architectures. This paper presents an empirical benchmark of 14 state-of-the-art 2D and 3D-aware vision foundation models. Using the challenging real-world UFPR-VeSV dataset, we evaluate these models as frozen feature extractors via linear probing for vehicle type, make, and model recognition. We further stress-test the best-performing models under few-shot learning and Out-of-Distribution (OOD) domain shifts. Our results show that standard 2D self-supervised models, particularly DINOv3, substantially outperform 3D-aware models in fine-grained tasks, achieving over 93% Macro-Accuracy for make and model recognition. However, the 3D-aware Depth Anything v2 exhibits stronger invariance to viewing angles in vehicle type classification. These findings motivate hybrid approaches that combine 2D and 3D priors for robust vehicle recognition. Our code is publicly available at https://github.com/UFPR-IPASPPR/3D-Vision-Benchmark/.
Problem

Research questions and friction points this paper is trying to address.

Vehicle Attribute Recognition
2D and 3D-Aware Vision Foundation Models
Fine-Grained Vehicle Classification
Intelligent Transportation Systems
Automatic License Plate Recognition
Innovation

Methods, ideas, or system contributions that make the work stand out.

2D and 3D-aware vision foundation models
vehicle attribute recognition
DINOv3
viewing angle invariance
hybrid approaches
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Alexandre V. Delazeri
Department of Informatics, Federal University of Paraná, Curitiba, Brazil
G
Gabriel E. Lima
Department of Informatics, Federal University of Paraná, Curitiba, Brazil
E
Eduil Nascimento Jr
Department of Technological Development and Quality, Paraná Military Police, Curitiba, Brazil
Rayson Laroca
Rayson Laroca
Pontifical Catholic University of Paraná (PUCPR)
Computer VisionDeep LearningPattern Recognition
David Menotti
David Menotti
Department of Informatics, Universidade Federal do Paraná
Computer VisionImage ProcessingPattern RecognitionMachine Learning