A Lightweight Multimodal Vision-Language Framework for Early-Stage Anatomical Green Fruit Classification in Commercial Orchards

📅 2026-08-23
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究提出了一种轻量级的多模态视觉-语言框架,使用TinyCLIP进行苹果幼果解剖结构分类,以支持果园精准操作。
📝 Abstract
Accurate identification of early-stage apple fruitlet anatomical structures, including the calyx, fruitlet body, and peduncle, is essential for robotic thinning, crop-load management, and other precision orchard operations. This study presents a lightweight multimodal vision-language framework that adapts TinyCLIP for fine-grained fruitlet anatomy classification in complex orchard environments. A dataset of 600 high-resolution RGB images collected from Scilate and Scifresh apple orchards was converted into 224 x 224 image patches and annotated for three anatomical classes. Domain-specific language prompts, such as ``a photo of a class,'' were used to guide multimodal alignment between orchard imagery and horticultural structures. A sliding-window inference strategy with a stride of 112 pixels aggregates patch-level predictions into spatial heatmaps, enabling interpretable whole-image localization of fruitlet components relevant to robotic thinning. Patch-level evaluation on an NVIDIA T4 GPU achieved F1-scores of 0.95 for calyx, 0.98 for fruitlet, and 0.85 for peduncle, with a macro-F1 score of 0.93. Deployment-oriented optimization using ONNX and TensorRT enabled efficient inference on NVIDIA Jetson hardware, preserved accuracy under INT8 quantization, and supported model sizes of approximately 127-137 MB with millisecond-level patch inference. These results demonstrate that lightweight vision-language models can provide interpretable and edge-deployable perception for automated fruitlet analysis and future robotic thinning systems. The source code and implementation details are publicly available at https://github.com/WilliamBu1/A-Lightweight-Vision-Language-Model-for-Early-Stage-Fruitlet-Classification-in-Apple-Orchards.
Problem

Research questions and friction points this paper is trying to address.

Early-Stage Anatomical Green Fruit Classification
Multimodal Vision-Language Framework
Orchard Environments
Robotic Thinning
Precision Agriculture
Innovation

Methods, ideas, or system contributions that make the work stand out.

Lightweight Multimodal Vision-Language Framework
TinyCLIP
Sliding-Window Inference Strategy
Domain-Specific Language Prompts
Edge-Deployable
🔎 Similar Papers
No similar papers found.
Ranjan Sapkota
Ranjan Sapkota
Cornell University
Artificial IntelligenceAgentic AIAgricultural AutomationAgricultural Robotics
W
William Bu
Department of Computer Science, University of Central Florida, USA
C
Chen Chen
Institute of Artificial Intelligence (IAI) & Department of Computer Science, University of Central Florida, USA
Y
Yunjun Xu
UCF Department of Mechanical and Aerospace Engineering, University of Central Florida, USA
Manoj Karkee
Manoj Karkee
Cornell University
Agricultural AutomationAgricultural RoboticsSmart FarmingDigital AgriculturePrecision Ag