Evaluating Vision-Language Models as a Zero-Shot Learning Alternative to You Only Look Once and Optical Character Recognition for Nigerian License Plate Recognition

📅 2026-07-02
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of traditional license plate recognition systems, which rely on multi-stage pipelines combining YOLO and OCR and suffer from performance degradation, high computational costs, and dependence on annotated data in unstructured environments. For the first time, this work introduces vision-language models (VLMs) to license plate recognition in Nigeria’s challenging real-world conditions, proposing a zero-shot, end-to-end unified framework that enables direct recognition without task-specific training. The authors evaluate leading VLMs—including Gemini 2.0 Flash, Qwen2.5-VL, GPT-4o, Claude 4 Sonnet, and Llama 3.2 Vision—on a dataset of 88 real-world images. Results demonstrate that Gemini and Qwen significantly outperform other models, achieving lower character error rates and confirming the feasibility and robustness of VLMs as a viable alternative to conventional YOLO+OCR pipelines.
📝 Abstract
License Plate Recognition (LPR) systems are critical tools in traffic monitoring, security enforcement, and urban mobility management. Traditional LPR systems often rely on a multi-stage pipeline involving object detection using You Only Look Once (YOLO) and Optical Character Recognition (OCR), which suffer from limitations such as high resource demands, poor performance in unstructured environments, and the need for large annotated datasets. This study explores the potential of Vision-Language Models (VLMs) as a unified, zeroshot learning solution for Nigerian license plate recognition. Using a curated dataset of 88 challenging real-world images collected in Nigeria, we evaluate five selected VLMs: Gemini 2.0 Flash Exp (Google DeepMind), Qwen2.5-VL-7B-Instruct (Alibaba), GPT-4o (OpenAI), Claude 4 Sonnet (Anthropic), and Llama 3.2 Vision 90b (Meta). Results based on Character Error Rate (CER) reveal that Gemini and Qwen significantly outperform other models in both accuracy and robustness, on the challenging image scenarios. This work highlights the practical advantages of VLMs over YOLO+OCR, questions the claims by model providers, and compares the performances of the VLMs.
Problem

Research questions and friction points this paper is trying to address.

License Plate Recognition
Zero-Shot Learning
Vision-Language Models
YOLO
OCR
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language Models
Zero-Shot Learning
License Plate Recognition
YOLO
OCR
I
Ismail Ismail Tijjani
department of Mechatronics Engineering, Bayero University Kano(BUK), Kano, Nigeria
A
Ahmad Abubakar Mustapaha
School of Computer Science and Engineering, VIT-AP University, Andhra Pradesh, India
S
Sunusi Ibrahim Muhammad
department of Petroleum Engineering, Bayero University Kano(BUK), Kano, Nigeria
M
Muhammad Bashir Aliyu
department of Mechatronics Engineering, Aliko Dangote University of Science and Technology, Wudil, Kano, Nigeria