How Merge-Tolerant Are Vision Transformers for Wheat Phenotyping?

📅 2026-08-24
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过对比不同视觉变换器在小麦表型任务中的性能,探讨了无训练令牌合并方法对提高处理速度和减少资源消耗的效果。
📝 Abstract
Vision-based wheat phenotyping requires repeated measurements under deployment constraints, from growth-stage recognition to wheat-head counting and organ segmentation. Plain Vision Transformers (ViTs) provide a common architecture for these tasks, but quadratic attention limits high-throughput and edge inference. Training-free token merging is attractive because it can be inserted into trained models without retraining. We provide a systematic benchmark of ToMe and Mutual Pair Merging across growth-stage classification, wheat-head detection, and wheat-organ segmentation, measuring task quality, throughput, token count, and peak GPU memory, with additional Raspberry Pi 5 measurements. The benchmark reveals a clear hierarchy: classification is highly merge-tolerant, while detection and segmentation are constrained by repeated instances, thin organs, dense boundaries, reconstruction, and runtime overhead. Optimized attention backends can erase apparent speedups, so deployment value must be profiled on the target runtime rather than inferred from token count.
Problem

Research questions and friction points this paper is trying to address.

Vision-based wheat phenotyping
Vision Transformers (ViTs)
quadratic attention
high-throughput
edge inference
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision Transformers
Token Merging
Wheat Phenotyping
Throughput Optimization
Edge Inference
🔎 Similar Papers
No similar papers found.
S
Simon Ravé
LARIS, University of Angers, Angers, France
P
Pejman Rasti
LARIS, University of Angers, Angers, France; UMR INRAE-IRHS, Angers, France
D
David Rousseau
LARIS, University of Angers, Angers, France; UMR INRAE-IRHS, Angers, France