EviScope: Paired Counterfactual Evidence Diagnostics for Faithful and Efficient Grounded Language Models

📅 2026-09-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
针对基于证据的语言模型评价问题,提出EviScope方法,通过固定问题并改变证据来诊断模型的推理过程,从而更准确地评估模型性能。
📝 Abstract
Grounded language-model systems are often evaluated by final answer accuracy, yet a correct answer can be unsupported, drawn from the wrong source, or produced when evidence is insufficient or contradictory. We introduce EviScope, a paired counterfactual benchmark that holds the question fixed while adding, removing, distracting, or contradicting its evidence. EviScope-v1.1 contains 40 four-condition quartets with repaired counterfactual claims and span-level support labels for automatic evaluation. Across 960 gold-blind generations from Qwen2.5-7B, Llama 3.1 8B, and Gemini 3.5 Flash, paired metrics expose model-dependent grounding behavior that answer accuracy hides. On two local open models, an explicit evidence-action gate underperforms vanilla RAG on QCS: 0.15 vs. 0.50 for Qwen and 0.10 vs. 0.375 for Llama. Gemini reaches 0.944 joint success under both prompts, yet still answers 5% of conflict cases after contradiction insertion. EviScope therefore distinguishes unsupported answering, conflict blindness, and wrong non-answer actions rather than scoring answers alone.
Problem

Research questions and friction points this paper is trying to address.

Grounded Language Models
Evidence Diagnostics
Counterfactual Benchmark
Answer Accuracy
Evidence Sufficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

EviScope
counterfactual benchmark
evidence diagnostics
grounded language models
conflict blindness
🔎 Similar Papers
2024-08-21BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLPCitations: 1
💼 Related Jobs
No related jobs found.
S
Suryadeep Singh Deswal
Indian Institute of Technology Roorkee