HealMed: Multilingual Evaluation of Large Language Models in Medicine

📅 2026-08-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过构建HealMed基准,采用多语言医学数据集评估大型语言模型性能,揭示了资源少的语言中模型表现下降等问题。
📝 Abstract
We present HealMed, an expert-reviewed benchmark for multilingual evaluation of large language models in medicine. HealMed contains 1,000 examples in each of nine languages, drawn from nine datasets and covering three task formats: MCQA, NLI and open-ended QA. The benchmark was developed over two years by 23 physicians and medical experts based across nine countries and regions. Each translation was evaluated and revised by two experts fluent in English and the corresponding target language. On HealMed, performance declined most in low-resource languages, although the size of the gap varied markedly across languages and models. The strongest proprietary models were the most stable across languages, whereas many open-source and medically specialized models showed larger and less consistent gaps. Medical specialization alone did not ensure multilingual robustness. Furthermore, expert revision could either raise or lower measured performance, indicating that translation quality materially affects cross-language evaluation results.
Problem

Research questions and friction points this paper is trying to address.

multilingual evaluation
large language models
medicine
low-resource languages
performance consistency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multilingual Evaluation
Large Language Models in Medicine
Expert-Reviewed Benchmark
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yingjian Chen
Fan Gao
Fan Gao
Caltech; MIT
NGS BioinformaticsImage data processingAI/MLNeurodegenerationProtein Bioinformatics
S
Sherry T. Tong
H
Haoyu Zhang
A
Aosong Feng
K
Kevin W. Jin
X
Xing Wu
Jinghui Lu
Jinghui Lu
ByteDance Inc., School of Computer Science, University College Dublin
Natural Language ProcessingMulti-ModalityLLMHuman-in-the-loop Learning
A
Abdul Samad
A
Akbar Faruqi
C
Cesar Caraballo
C
Cibele Brandão
D
Dhruva
G
Gupta
E
Eunji Jeon
G
Gabriel Madera-Santiago
G
Geon Lee
H
Hugo Toshio Itikawa
Insook Cho
Insook Cho
I
Isabelli Martins
I
Isarar Siddique
I
Israr Ahmed
J
Jihyo Kwak
K
Kanyakorn Veerakanjana
L
Luis Guilherme Cardoso