InspectVLM: Unified in Theory, Unreliable in Practice
This work investigates the feasibility and robustness of unified vision-language models (VLMs) as replacements for task-specific models in industrial inspection. To address diverse requirements—including image classification, object detection, and keypoint localization—it proposes a language-interface-based unified paradigm, introduces InspectMM—the first large-scale multimodal industrial inspection dataset—and performs instruction tuning on Florence-2. Experiments show strong performance on classification and structured keypoint tasks, but limited robustness on fine-grained detection, high sensitivity to prompt engineering, and weaker visual grounding compared to specialized architectures like ResNet. The core contribution lies in empirically exposing the fundamental tension between *linguistic unification* and *visual reliability* in VLMs for industrial applications, thereby establishing an empirical benchmark and identifying concrete directions for advancing VLMs toward high-precision industrial deployment.