Beyond Majority Vote: Multi-Perspective Adjudication for Medical Hallucination Detection

📅 2026-09-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过多视角标注、LLM作为裁判的候选发现及医学专家和基于证据的事实核查来改进医疗聊天机器人生成文本中事实错误的检测,揭示单次标注可能低估错误。
📝 Abstract
Understanding the frequency of factual errors in chatbot-generated text and evaluating systems that detect these errors is critical for determining chatbot safety. Yet factual-error detection is often treated as a single-pass, single-annotator labeling problem. In long-form chatbot responses, factual errors can be subtle and embedded within mostly correct text. We develop a multi-perspective annotation study of medically relevant chatbot responses, combining first-pass annotation, LLM-as-a-Judge (LaJ) candidate discovery, and two forms of adjudication: medical-expert and evidence-based fact-checking. First-pass annotators frequently miss factual errors later validated by adjudicators. LaJ improves candidate discovery, but is insufficient on its own: It misses factual errors that annotators catch. We also find disagreement among adjudicators, suggesting that adjudication over multiple candidate sources can improve benchmark completeness, but does not eliminate the need to apply judgment and expertise. Applied to an existing benchmark, this technique reveals a similar pattern of missing annotations. Together, these results suggest that in the settings examined here, single-pass hallucination benchmarks may achieve scale at the cost of undercounting factual errors. Multi-pass adjudication can improve coverage, but inferences drawn from the benchmarks are still sensitive to the judgment, expertise, and evidence used to determine error presence.
Problem

Research questions and friction points this paper is trying to address.

factual errors
chatbot responses
medical relevance
single-annotator labeling
multi-perspective annotation
Innovation

Methods, ideas, or system contributions that make the work stand out.

multi-perspective adjudication
LLM-as-a-Judge (LaJ)
evidence-based fact-checking
medical-expert review
💼 Related Jobs
No related jobs found.
J
Joe Cecil
Information Sciences Institute, University of Southern California
Marjorie Freedman
Marjorie Freedman
Research Team Lead, USC Information Sciences Institute
NLP