Edge Reliability Gap in Vision-Language Models: Quantifying Failure Modes of Compressed VLMs Under Visual Corruption

πŸ“… 2026-03-24
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study investigates whether compressed vision-language models (VLMs) exhibit systematic failure modes beyond mere increases in error rates when deployed on edge devices. Through categorizing errors into object blindness, semantic drift, and prior bias, and leveraging confidence calibration (ECE), structured negation reasoning probes, controlled ambiguity experiments, and a GPT-4o-based discriminator, the work reveals for the first time that compact VLMs undergo significant and non-uniform qualitative degradation on benchmarks like COCOβ€”most notably a collapse in negation reasoning capabilities. For instance, SmolVLM2-500M achieves a 100% error rate on the false_yn template, underperforming Qwen2.5-VL-7B by 12.5 percentage points. The study further introduces a reproducible safety auditing pipeline, establishing a new paradigm for evaluating the reliability of compressed VLMs.

Technology Category

Application Category

πŸ“ Abstract
The rapid compression of large vision-language models (VLMs) for edge deployment raises an underexplored question: do compact models fail differently, not merely more often? This study compares a 7-billion-parameter quantised VLM (Qwen2.5-VL-7B, 4-bit NF4) against a 500-million-parameter FP16 model (SmolVLM2-500M) across 4,000 samples from VQAv2 and COCO Captions. A three-category error taxonomy (Object Blindness, Semantic Drift, Prior Bias) is applied as a diagnostic framework. A text-only GPT-4o judge reveals Semantic Drift (B) as the dominant failure mode on VQAv2 and on COCO for Qwen, with a mixed Object Blindness / Semantic Drift profile for SmolVLM2 on COCO; Prior Bias (C) is present on VQAv2 but absent on COCO for both models. Confidence calibration is measured via Expected Calibration Error (ECE) using geometric mean token probability, compositional reasoning is probed with structured negation probes across four templates, and a blur robustness experiment completes the evaluation. For this model pair, the compact model exhibits a qualitatively distinct failure signature: a 12.5pp larger negation collapse (-33.2pp vs. -20.8pp, Wald 95% CI [8.2, 16.8]pp, p < 10^-8), driven almost entirely by COCO while the VQAv2 gap is not statistically significant (4.5pp, p=0.19). The most discriminating template is false_yn: SMOLVLM2-500M responds "Yes" (incorrectly claiming a depicted object is absent) on 100% of COCO trials vs. 14% for Q WEN 2.5-VL-7B. Asymmetric dataset-dependent miscalibration and a blur experiment with two controlled ablations complete the analysis. The fully reproducible pipeline is released for systematic safety auditing of compressed VLMs prior to edge deployment.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Models
Model Compression
Failure Modes
Edge Deployment
Visual Corruption
Innovation

Methods, ideas, or system contributions that make the work stand out.

failure mode analysis
compressed vision-language models
semantic drift
negation reasoning
confidence calibration
πŸ’Ό Related Jobs
No related jobs found.
M
Mehmet Kaan Erol
Marmara University, Institute of Pure and Applied Sciences