When Models Defer to Wrong Answers: A Robustness Audit of Source-Attributed Cues in Multiple-Choice QA

📅 2026-09-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过引入NC-MCAR指标,审计语言模型在多选题中受到来源声明误导时答案的不稳定性问题。
📝 Abstract
Language models often receive a question together with a claim about what another source answered. We audit whether such claims destabilize answers in multiple-choice question answering. For each item, we hold one wrong option fixed across misleading conditions and vary the cue template attached to it. We introduce \emph{neutral-conditioned misleading cue adoption rate} (NC-MCAR), which measures switches to that option only on valid cued trials where the same model first selected the gold answer under a neutral prompt. This is a measure of answer instability, not proof that the model knew the answer or that all deference is irrational. We evaluate four instruction-following models on MMLU-Pro and IndicMMLU-Pro in English, Hindi, Bengali, Tamil, and Telugu. Across 220{,}000 outputs, the expert template yields 41.1\% aggregate NC-MCAR, compared with 12.5\% for the majority template. These two conditions use the same wrong option and final instruction. Filler accuracy remains well above expert-wrong accuracy, while correct-cue prompts have high valid-response accuracy. The audit documents answer instability relevant to grounding under the tested forced-choice prompts: a bare, unverified source claim can outweigh an answer that was previously consistent with the task evidence.
Problem

Research questions and friction points this paper is trying to address.

language models
multiple-choice question answering
answer instability
source-claimed answers
misleading cues
Innovation

Methods, ideas, or system contributions that make the work stand out.

Neutral-Conditioned Misleading Cue Adoption Rate
Answer Instability
Source-Attributed Cues
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Manikandan Ravikiran
Manikandan Ravikiran
Thoughtworks AI Research
Machine learningComputer visionNatural Language Processing
S
Siddharth Vohra
Carnegie Mellon University; Amazon Web Services AI Native, Pittsburgh, PA, USA