CoAtNet-DeepMoE: A Convolution-Attention Hybrid with DeepSeek Mixture-of-Experts for Parameter-Efficient Tomato Disease Classification

📅 2026-09-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决番茄疾病分类中模型参数过多的问题,提出了一种结合卷积-注意力机制与深度混合专家系统的CoAtNet-DeepMoE方法,在保持高精度的同时显著减少了参数量。
📝 Abstract
The world population is growing rapidly, and technology is improving in parallel. Meeting the huge demand for food for these 7 billion people not only depends on increasing food production but also on reducing food loss. Crop losses due to disease affect both the food supply and the financial and economic stability of a country. Tomatoes are among the top food-producing crops globally, and a significant portion of this production is lost due to disease. People have used Machine Learning techniques for feature extraction and early diagnosis of tomato diseases, and nowadays, Deep Learning-based models are widely used for disease recognition. However, most existing models are highly parameter-intensive, which increases the time required for training and inference. As a result, while lightweight models are more suitable for user-friendly applications, they often show a reduction in performance. To balance performance and model size, we propose CoAtNet-DeepMoE, a Convolution-Attention hybrid architecture for rich feature extraction, further enhanced with a DeepSeek Mixture of Experts to substantially reduce the number of parameters without sacrificing accuracy. We evaluate our model on both balanced and imbalanced datasets from Kaggle and PlantVillage, demonstrating robustness and achieving 99.80% accuracy, 99.80% precision, 99.80% recall, and 99.80% F1-score on Kaggle, and 99.83% accuracy, 99.85% precision, 99.76% recall, and 99.80% F1-score on PlantVillage, representing state-of-the-art performance with only 2.47M parameters. The source code will be available at https://github.com/nadimbrur/CoAt-MoE.
Problem

Research questions and friction points this paper is trying to address.

Tomato Disease Classification
Parameter-Efficient
Deep Learning Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Convolution-Attention Hybrid
DeepSeek Mixture-of-Experts
Parameter-Efficient
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Md Nadim Mahamood
Department of Computer Science and Engineering, Begum Rokeya University, Rangpur, Bangladesh
M
Md Arif Shahriar
Department of Computer Science, Texas State University, San Marcos, Texas, USA
M
Md Shafi Ud Doula
Department of Information and Communication Technologies, Asian Institute of Technology, Pathum Thani, Thailand
Kamrul Hasan
Kamrul Hasan
Assistant Professor of Computer Engineering, Tennessee State University, Nashville, TN
Digital Security & PrivacyCyber-Physical SystemsAI/ML