PragMatch: Separating Pragmatic Incongruity from Cross-Modal Mismatch in Large Vision-Language Models

📅 2026-08-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Current large vision-language models struggle to distinguish pragmatic incongruity from superficial image-text mismatches in multimodal sarcasm detection, often relying on surface-level cues rather than deep reasoning. To address this limitation, this work introduces PragMatch, a novel benchmark that explicitly disentangles pragmatic incongruity from cross-modal mismatch by constructing 3,000 controlled image-text pairs. The benchmark evaluates models’ reasoning capabilities through systematic masking, injection of surface cues, analysis of OCR and stylistic features, and human-curated literal versus hard negative examples. Experimental results reveal that state-of-the-art models are significantly biased by lexical content, OCR-derived text, and visual style—exhibiting substantial prediction shifts even when semantic relationships remain unchanged—thereby exposing their lack of genuine multimodal pragmatic reasoning ability.
📝 Abstract
Large Vision-Language Models (LVLMs) have demonstrated strong performance on multimodal benchmarks, yet it remains unclear whether they genuinely reason about relationships between images and text or rely on superficial correlations, known as shortcut learning. This question is particularly important for multimodal sarcasm detection, where successful prediction depends on recognizing pragmatic incongruity rather than treating sarcasm as simple image-text mismatch. We introduce PragMatch, a controlled benchmark of 3,000 image-text pairs derived from MMSD2.0, including original sarcastic examples and constructed literal and hard-negative pairs. We identify influential shortcut cues through systematic masking and evaluate their impact through targeted injection experiments. Our results show that LVLM predictions are sensitive to lexical, OCR-derived and stylistic cues, with injected surface signals causing substantial changes in model predictions despite unchanged underlying image-text relationships. Our findings reveal limitations in current LVLMs while PragMatch provides a systematic testbed for evaluating multimodal pragmatic reasoning beyond surface-level image-text alignment.
Problem

Research questions and friction points this paper is trying to address.

pragmatic incongruity
cross-modal mismatch
shortcut learning
multimodal sarcasm detection
vision-language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

pragmatic incongruity
shortcut learning
multimodal sarcasm detection
controlled benchmark
surface cues
🔎 Similar Papers