MedRoundsQA: A Persona and Difficulty Aware Evaluation for Multi-Turn Medical Consultations

📅 2026-09-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
针对单轮医疗咨询无法反映真实情况的问题,通过构建多轮对话基准MedRoundsQA并分类难度来评估模型性能。
📝 Abstract
Medical benchmarks are dominated by single-turn, multiple-choice clinical cases that poorly reflect real consultations. Practically, clinicians elicit evidence interactively and patient communication varies widely. We introduce MedRoundsQA, a multi-turn diagnostic benchmark derived from 1,387 board-exam cases across 17 specialties. Each case is converted into a structured 24-slot clinical record, and then instantiated as controlled doctor-patient dual-agent dialogues under varying patient personas, with the underlying clinical content held fixed. We further classify cases by difficulty using model-based uncertainty to enable easy-to-hard analysis. Evaluations of fifteen LLM doctor agents show that (i) moving from a single-turn diagnosis on the standardized records to multi-turn consultations causes large degradations of roughly 13-39 points; (ii) more turns reliably improves question relevance, but diagnostic accuracy exhibits diminishing returns and typically plateaus after 6-12 turns; and (iii) patient persona differences can shift diagnosis accuracy by about 7-8 points (lowest to highest education), highlighting equity risks that single-turn benchmarks miss.
Problem

Research questions and friction points this paper is trying to address.

multi-turn consultations
medical benchmarks
patient communication
Innovation

Methods, ideas, or system contributions that make the work stand out.

multi-turn diagnostic benchmark
persona and difficulty aware
doctor-patient dual-agent dialogues
model-based uncertainty
🔎 Similar Papers
No similar papers found.