Can Edge-Deployable Vision-Language Models Identify Species?

📅 2026-09-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究测试了2-8B参数范围内的视觉-语言模型在边缘设备上识别物种的能力,与专业模型BioCLIP对比,发现所有模型在野外图像上的表现均下降,表明数据训练的重要性。
📝 Abstract
Camera traps often run in the field on edge hardware with limited or no connectivity, making small, locally-deployable vision-language models (VLMs) -- not frontier-scale ones -- the practically relevant class to evaluate for species identification. We test whether models in this deployment-relevant 2--8B range carry genuine taxonomic knowledge, evaluating four such VLMs (Qwen3-VL 2B/4B/8B, Gemma3 4B) against the domain-specific specialist BioCLIP (300M parameters) on a 96-species task, comparing clean iNaturalist photographs against camera-trap imagery from 6 LILA.science collections, on two independently-sampled evaluation sets. All models identify species far above chance, but every model -- general-purpose or specialist -- degrades sharply on field imagery (domain gaps of 9.6--26.6 percentage points, consistent across taxonomic levels and both evaluation sets), indicating the degradation reflects general image legibility rather than fine-grained discrimination failure. BioCLIP substantially outperforms every VLM tested (by 33.2--59.2 percentage points across an expanded 200-image sample for every model) despite its far smaller size, suggesting the gap reflects specialized training data rather than model scale; yet BioCLIP's own domain gap (18.0 points) is statistically indistinguishable from the best VLM's (22.3 points), suggesting the clean-to-field degradation itself is a property of the image-quality shift rather than a general-purpose-model weakness. Under open-set prompting, 5.9--9.6% of responses are syntactically valid but taxonomically nonexistent species names; the relative fabrication-rate ranking across models replicates exactly across both evaluation sets, a more robust finding than any single point estimate.
Problem

Research questions and friction points this paper is trying to address.

species identification
edge-deployable vision-language models
camera trap imagery
Innovation

Methods, ideas, or system contributions that make the work stand out.

Edge-Deployable Vision-Language Models
Species Identification
Domain Gaps
Specialized Training Data
Open-Set Prompting
💼 Related Jobs
No related jobs found.
W
William Zhou
Plano West Senior High School
M
Mayukha Siripuram
Centennial High School
Xiao Yan
Xiao Yan
The University of Texas at Dallas
Z
Ziqi Liu
The University of Texas at Dallas
Yi Ding
Yi Ding
University of Texas at Dallas
Cyber-Physical SystemsMobile ComputingMachine Learning