Evaluating Scaffolding-Oriented Multi-Agent Large Language Model System for Clinical Interview Training
为解决临床访谈培训资源密集和难以扩展的问题,研究开发了一个基于多代理大型语言模型的AI标准化患者培训平台,并通过随机对照实验验证了其在提高医学生沟通、同理心及病史采集行为方面的有效性。
为解决临床访谈培训资源密集和难以扩展的问题,研究开发了一个基于多代理大型语言模型的AI标准化患者培训平台,并通过随机对照实验验证了其在提高医学生沟通、同理心及病史采集行为方面的有效性。
Rare diseases suffer from prolonged diagnostic timelines and fragmented clinical evidence; moreover, general-purpose large language models exhibit limited clinical reasoning capabilities due to scarce real-world electronic health records (EHRs), outdated medical knowledge, and hallucination. To address these challenges, we propose a domain-specific clinical reasoning paradigm centered on “narrative-first, knowledge-enhanced” inference. Our approach comprises: (1) constructing a physician-validated rare-disease reasoning dataset and domain-specific corpus; (2) designing a knowledge graph–anchored retrieval mechanism and phased chain-of-thought training to integrate non-phenotypic evidence (e.g., imaging, functional tests); and (3) enhancing robustness under noisy conditions and phenotypic overlap via instruction fine-tuning, knowledge graph fusion, and structured reasoning. Evaluated on multicenter real-world EHRs and public benchmarks, our method achieves state-of-the-art performance—matching the diagnostic accuracy of senior clinicians while significantly shortening diagnostic pathways and enabling transparent, auditable clinical decision-making.
为解决临床访谈培训资源密集和难以扩展的问题,研究开发了一个基于多代理大型语言模型的AI标准化患者培训平台,并通过随机对照实验验证了其在提高医学生沟通、同理心及病史采集行为方面的有效性。
Rare diseases suffer from prolonged diagnostic timelines and fragmented clinical evidence; moreover, general-purpose large language models exhibit limited clinical reasoning capabilities due to scarce real-world electronic health records (EHRs), outdated medical knowledge, and hallucination. To address these challenges, we propose a domain-specific clinical reasoning paradigm centered on “narrative-first, knowledge-enhanced” inference. Our approach comprises: (1) constructing a physician-validated rare-disease reasoning dataset and domain-specific corpus; (2) designing a knowledge graph–anchored retrieval mechanism and phased chain-of-thought training to integrate non-phenotypic evidence (e.g., imaging, functional tests); and (3) enhancing robustness under noisy conditions and phenotypic overlap via instruction fine-tuning, knowledge graph fusion, and structured reasoning. Evaluated on multicenter real-world EHRs and public benchmarks, our method achieves state-of-the-art performance—matching the diagnostic accuracy of senior clinicians while significantly shortening diagnostic pathways and enabling transparent, auditable clinical decision-making.