A Comparative Evaluation of Deep Learning Object Detection Models on a Real-World Multi-Plant Dataset from Africa
This study addresses the limited real-world applicability of existing agricultural vision models, which are predominantly trained on idealized datasets and struggle in the complex, dynamic field conditions prevalent in under-resourced regions such as Africa. To bridge this gap, the authors present a systematic evaluation of six state-of-the-art object detection models—YOLOv5, YOLOv8, YOLO11, YOLO26, Faster R-CNN, and RT-DETR—on AgriAISeg, the first multi-crop, multi-challenge dataset collected directly from African farmlands. Experimental results demonstrate that RT-DETR achieves the highest performance with an mAP@0.5:0.95 of 0.624, while YOLO-family models exhibit consistently strong accuracy and training efficiency. In contrast, Faster R-CNN suffers significant performance degradation in complex scenarios. This work provides the first empirical evidence of the varying suitability of modern detectors in authentic African agricultural settings, offering critical guidance for deploying AI solutions in resource-constrained environments.