What Matters for Grocery Product Retrieval with Open Source Vision Language Models

๐Ÿ“… 2026-05-18
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the challenge of distinguishing highly similar products in fine-grained, zero-shot multimodal SKU retrieval for grocery items. The authors systematically evaluate 190 open-source vision-language models on the GroceryVision benchmark, analyzing the impact of pretraining data quality, model architecture, and input resolution within a contrastive learning embedding and multimodal retrieval framework. Their findings reveal that data quality outweighs scale: lightweight models such as MobileCLIP-B (150M parameters) trained on high-quality data outperform larger models trained on noisier datasets. The work introduces โ€œsemantic power densityโ€ as a novel efficiency metric. The best-performing model achieves 94.5% Recall@5, yet a 17.5% performance gap remains at Recall@1; notably, data filtering improves accuracy by up to 16.6%.
๐Ÿ“ Abstract
Multimodal product retrieval (MPR) underpins checkout-free retail and automated inventory systems, yet it demands fine-grained SKU discrimination that standard vision-language benchmarks fail to capture. We present the first systematic zero-shot evaluation of 190 open-source VLMs on the MPR task of the GroceryVision Challenge, isolating pre-training data, architecture, and input resolution. Our analysis yields three actionable findings. \textbf{(1) Data quality trumps scale.} Switching from raw web-scrapes to filtered datasets delivers up to 16.6\% accuracy gains, exceeding the benefit of doubling model parameters. \textbf{(2) Efficient models can win.} MobileCLIP-B (150M parameters) outperforms 351M counterparts trained on noisy data. We introduce \textit{semantic power density} ($ฯ†$), an efficiency metric that penalizes sub-threshold accuracy. \textbf{(3) A precision gap persists.} State-of-the-art models achieve 94.5\% Recall@5 but suffer a 17.5\% drop at Recall@1, revealing that contrastive embeddings cluster categories effectively but fail to rank visually similar SKUs. Code and evaluation scripts are available at \url{https://github.com/upeee/openmpr}.
Problem

Research questions and friction points this paper is trying to address.

multimodal product retrieval
SKU discrimination
vision-language models
zero-shot evaluation
fine-grained retrieval
Innovation

Methods, ideas, or system contributions that make the work stand out.

vision-language models
multimodal product retrieval
data quality
semantic power density
zero-shot evaluation
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
E
Emmanuel G. Maminta
AI Graduate Program, University of the Philippines, Diliman, Quezon City
R
Rowel O. Atienza
1 AI Graduate Program, University of the Philippines, Diliman, Quezon City; 2 EEEI, University of the Philippines, Diliman, Quezon City