MTDiag: A Multi-Turn Diagnostic Dataset Towards Clinically Meaningful LLM Evaluation

📅 2026-08-25
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决LLM在多轮临床诊断对话中准确性和可靠性下降的问题,构建了MTDiag数据集,并提出了基于临床知识的评估指标。
📝 Abstract
Clinical diagnosis is fundamentally interactive and incremental, yet the dominant paradigm for evaluating Large Language Models (LLMs) in medicine remains static QA benchmarks or template-based dialogues. These benchmarks say little about whether a model can serve as a diagnostic agent in a dynamic clinical encounter, with LLMs showing significant accuracy and reliability degradation in multi-turn settings. To address this issue, we present MTDiag, a large multi-turn diagnostic dialogue dataset constructed from three heterogeneous sources: DDXPlus, MIMIC-IV, and published case reports (AJCR), covering common ED presentations as well as long-tail rare and atypical conditions. All cases are normalized into a canonical schema anchored in the most comprehensive and widely-adopted medical knowledge bases (UMLS concept identifiers, with ICD-10 diagnosis codes). We release the schema, a UserLM-8B-based utterance-generation pipeline, and the physician-validated dataset that converts structured clinical evidence into natural-language utterances. Importantly, we introduce and motivate clinical knowledge-grounded metrics for evaluating LLMs as diagnostic agents, beyond diagnostic accuracy, for the task of multi-turn differential diagnosis.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Clinical Diagnosis
Multi-turn Dialogue
Evaluation Metrics
Diagnostic Accuracy
Innovation

Methods, ideas, or system contributions that make the work stand out.

multi-turn diagnostic dialogue
dynamic clinical encounter
UMLS concept identifiers
clinical knowledge-grounded metrics
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
P
Pia Chouayfati
Technical University of Munich
A
Alexander M. Fichtl
Technical University of Munich
Miriam Anschütz
Miriam Anschütz
PhD Student of Computer Science, Technical University of Munich
Natural language processingeasy-to-readtext simplification
G
George Doumat
Department of Internal Medicine, UT Southwestern
Georg Groh
Georg Groh
Adjunct Professor
Social ComputingNatural Language Processing