Fusing Perceptual Vision Experts with Multimodal Large Language Models for Explainable Plant Disease Diagnosis: From Benchmark Imagery to Real-World Robotic Field Validation

📅 2026-08-23
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
"This study addresses the challenge of enhancing the accuracy of field plant disease diagnosis by proposing a Hybrid Hierarchical Multi-Agent Framework (H²MAF). This framework integrates the decision outputs from EfficientNet-B3 and ConvNeXt-Tiny visual models, leveraging multimodal large language models such as Gemma 4 E4B and Qwen3.5 4B for semantic arbitration. The method generates explainable diagnoses, risk levels, and treatment urgency assessments based on JSON-formatted evidence. Experimental results demonstrate that H²MAF significantly improves diagnostic accuracy, particularly in conflict data subsets, with an increase from 63.9% to 68.5% on the PlantDoc dataset, and a 7.6% improvement in CNN prediction conflicts. Furthermore, the framework achieves high accuracy rates of 99.3% and 98.9% on real-world field datasets from Cornell University."
📝 Abstract
Accurate field plant disease diagnosis requires reliable fusion of uncertain and conflicting perceptual evidence. We present the Hybrid Hierarchical Multi-Agent Framework (H$^{2}$MAF), combining decision-level fusion of EfficientNet-B3 and ConvNeXt-Tiny with semantic arbitration by open-weight multimodal large language models (MLLMs), Gemma 4 E4B and Qwen3.5 4B, using structured JSON evidence to generate explainable diagnoses, risk levels, treatment urgency, and financial exposure. (H$^{2}$MAF) is evaluated on 14,364 images (1,370 test images) across PlantDoc (2,922 images, 27 classes) and two non-public, continuously captured Cornell robot-acquired field datasets: Stage 2 (20 GB; 4,215 images) and Stage 4 (40 GB; 7,227 images), covering Early Blight, Late Blight, and Septoria Leaf Spot under uncontrolled field conditions. On PlantDoc, Gemma improves accuracy from 63.9% to 68.5%, achieving +7.6 points on the 41.7% CNN-conflict subset. Cornell accuracies reach 99.3% and 98.9%, with only 1.7-4.1% disagreement, demonstrating conflict-dependent MLLM utility. The critical-risk error of gemma is 0.14-0.5 points, whereas Qwen overflags by 3.5-14.4 points. These results establish MLLM arbitration as a promising, yet calibration-dependent, approach for explainable agricultural AI and robotic field decision support. Github Link: https://github.com/Applied-AI-Research-Lab/Explainable-AI-Plant-Disease-Detection
Problem

Research questions and friction points this paper is trying to address.

plant disease diagnosis
perceptual evidence
multimodal large language models
explainable diagnoses
field conditions
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hybrid Hierarchical Multi-Agent Framework (H²MAF)
multimodal large language models
explainable diagnoses
structured JSON evidence
robotic field validation
🔎 Similar Papers
No similar papers found.