Untangling the Mechanisms of Misleading Context in Medical Question Answering

📅 2026-09-02
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究探讨误导性上下文如何影响医学问答模型的判断,通过在MedMisBench上测试三种模型对两种误导线索的反应,发现开放推理路径有助于监测错误决策。
📝 Abstract
Large language models now answer medical questions with expert-level performance. However, the context these systems act on can be misleading, and misleading context can corrupt a model's medical judgment. To understand how misleading context corrupts this judgment, we examine the model's susceptibility to the context, disclosure of it, mechanism of corrupted reasoning, and monitorability of the decision. On the medical reasoning subset of MedMisBench, a clinician-reviewed question-answering benchmark of 8,627 questions, we inject two types of misleading context cues, fabricated evidence and a bare assertion. We test three reasoning models, two that expose their full reasoning trace and one frontier model that exposes only its response. All three are more susceptible to the assertion than to the fabricated evidence, adopting the asserted answer 10 to 27 points more often. The misleading cues are disclosed in 81 to 98% of traces but only 7 to 90% of responses, and the assertion is disclosed less often than evidence based cues. Resampling from reasoning traces without disclosure shows the two cues corrupt reasoning differently, evidence entering early and accumulating while the assertion redirects the conclusion near its end. An LLM monitor catches 78% of corrupted decisions at 5% false positives when reading an open model's trace with guidance, against at most 32% from any response. The misleading context that models are most susceptible to is disclosed least, and was caught reliably only from an open reasoning trace, which frontier providers withhold.
Problem

Research questions and friction points this paper is trying to address.

misleading context
medical judgment
reasoning models
fabricated evidence
bare assertion
Innovation

Methods, ideas, or system contributions that make the work stand out.

misleading context
medical question answering
reasoning trace
LLM monitor
disclosure of misleading cues
💼 Related Jobs
No related jobs found.