VDiff-Bench: A Challenging Benchmark for Fine-Grained Image Difference Identification

📅 2026-09-05
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对多模态大语言模型在细粒度图像差异识别上的不足,提出了VDiff-Bench基准测试,通过1,756个四选一问题来评估模型在10种变化类别上的表现。
📝 Abstract
Multimodal Large Language Models (MLLMs) perform strongly on general visual understanding tasks such as visual question answering, yet they often struggle with a basic comparative skill: identifying what has changed between two similar images. We introduce VDiff-Bench, a challenging multiple-choice benchmark for fine-grained Image Difference Identification. VDiff-Bench contains 1,756 four-way questions over image pairs and covers 10 change categories: position, motion, regional image color, overall image color, appearance/disappearance, noise/resolution, texture, substitution/size, OCR/text, and illumination. Each question corresponds to two image inputs with 4 choices: the true difference, two hard negative descriptions, and a"no difference"distractor. To make the task challenging, we specifically curate ground-truth-conditioned negatives that require models to distinguish the actual change from nearby semantic alternatives. Experiments with 11 state-of-the-art open- and closed-source MLLMs show that fine-grained visual comparison remains brittle: models exhibit uneven performance across sources and change categories, with persistent failures on subtle low-level changes like noises and textures. For instance, three 7-8B-scale open-source MLLMs score 52.5-70.6% on semantic changes but only 8.7-33.3% on low-level changes like noise and texture, falsely assuming no changes between two image inputs. Surprisingly, despite strong performance of other closed-source commercial models, Grok 4.3 demonstrate remarkable performance drop on identifying noise and texture differences between images, falling significantly behind large open-source models like Kimi K2.5 and K3. Overall, VDiff-Bench provides a targeted diagnostic for evaluating comparative visual understanding in MLLMs, exposing failures that are not captured by standard single-image vision-language tasks.
Problem

Research questions and friction points this paper is trying to address.

Multimodal Large Language Models
Image Difference Identification
Fine-Grained Visual Comparison
Innovation

Methods, ideas, or system contributions that make the work stand out.

VDiff-Bench
fine-grained Image Difference Identification
multimodal large language models
challenging multiple-choice benchmark
ground-truth-conditioned negatives
🔎 Similar Papers
2024-07-08Asian Conference on Computer VisionCitations: 2
Yixin Wan
Yixin Wan
PhD student in Computer Science, University of California, Los Angeles
MultimodalLLMNatural Language ProcessingFairnessTrustworthiness
T
Tianle Zheng
Department of Computer Science, University of California, Los Angeles
K
Kai-Wei Chang
Department of Computer Science, University of California, Los Angeles