EDCT-Bench: Uncovering Faithfulness Gaps in VLMs via Explanation-Driven Counterfactual Testing

📅 2026-09-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过Explanation-Driven Counterfactual Testing(EDCT)方法,揭示了视觉-语言模型(VLMs)在生成自然语言解释时存在的忠实度问题,并创建了EDCT-Bench基准来测试模型一致性。
📝 Abstract
Vision-Language Models (VLMs) can produce Natural Language Explanations (NLEs) that sound plausible yet remain inconsistent with the visual evidence they cite. We present Explanation-Driven Counterfactual Testing (EDCT), an intervention-based protocol that extracts visual concepts cited in a model's explanation, applies verified minimal edits to them, and tests whether the resulting answer and explanation remain consistent with the edited image. Using this protocol, we create EDCT-Bench, a comprehensive benchmark spanning three complementary domains: knowledge-intensive visual question answering (OK-VQA), safety-critical driving (DriveLM), and 3D spatial reasoning (3DSRBench). Across the evaluated VLMs, EDCT reveals substantial faithfulness gaps, with models frequently producing responses inconsistent with verified visual changes. Finally, our fine-tuning study suggests that EDCT-generated counterfactuals provide high-impact training signals.
Problem

Research questions and friction points this paper is trying to address.

Visual-Language Models
Natural Language Explanations
Faithfulness Gaps
Explanation-Driven Counterfactual Testing
Visual Evidence
Innovation

Methods, ideas, or system contributions that make the work stand out.

Explanation-Driven Counterfactual Testing
Visual-Language Models
Natural Language Explanations