En-ViMedNER: An English-Vietnamese Parallel Biomedical Corpus with UMLS Semantic Type Annotations

📅 2026-08-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文构建了En-ViMedNER,首个英越平行生物医学命名实体识别语料库,通过自动翻译、专家校对等方法解决越南语生物医学命名实体识别资源缺乏的问题。
📝 Abstract
Biomedical Named Entity Recognition (NER) is fundamental to healthcare AI applications, including clinical decision support and medical information extraction. While corpora with Unified Medical Language System (UMLS) annotations, such as MedMentions, have driven progress in English biomedical NER, no comparable resource exists for Vietnamese. This paper presents En-ViMedNER, the first English-Vietnamese parallel biomedical NER corpus annotated with UMLS semantic types, which are language-neutral codes providing a shared cross-lingual label space and ensuring direct comparability with existing UMLS-based resources. The corpus contains 4,392 PubMed abstract pairs, 44,892 English-Vietnamese sentence pairs, and 202,949 aligned entity-mention pairs across 21 semantic types adapted from the MedMentions ST21pv dataset. To balance quality and scalability, we have constructed the corpus through automatic translation, expert post-editing, LLM-assisted label projection, and human verification and adjudication. We characterize En-ViMedNER as a large-scale silver-standard corpus with a human-audited and consensus-corrected mini-test subset. We evaluate En-ViMedNER in two settings: (i) Vietnamese-input/Vietnamese-output biomedical NER and (ii) English-input/Vietnamese-output cross-lingual NER. For Vietnamese NER, we benchmark Vietnamese-supervised encoder models, English-supervised multilingual encoder models, and prompt-based LLMs. The best model achieves an F1 score of 52.70 on the test set and 53.78 on the mini-test set. For cross-lingual NER, we benchmark encoder-decoder models and prompt-based LLMs. The best model achieves an F1 score of 45.44 on the mini-test set. We publicly release our corpus, corpus construction pipeline, and baseline models to facilitate future Vietnamese biomedical NLP research.
Problem

Research questions and friction points this paper is trying to address.

Biomedical NER
UMLS Semantic Types
English-Vietnamese Parallel Corpus
Cross-lingual NER
Innovation

Methods, ideas, or system contributions that make the work stand out.

En-ViMedNER
UMLS Semantic Types
Cross-lingual NER
LLM-assisted label projection
Biomedical NER
🔎 Similar Papers
2024-06-19arXiv.orgCitations: 0
💼 Related Jobs
No related jobs found.
N
Nhu Vo
College of Engineering and Computer Science, VinUniversity, Vietnam
Phuong Nguyen
Phuong Nguyen
Associate Professor of Computer Science, DISIM, University of L'Aquila, Italy
Recommender SystemsMachine LearningAI4SEMining Software RepositoriesLinked Data
N
Nu Uyen Phuong Le
College of Engineering and Computer Science, VinUniversity, Vietnam
I
Inigo Jauregi Unanue
Faculty of Engineering and IT, University of Technology Sydney, Australia
D
Dung D. Le
College of Engineering and Computer Science, VinUniversity, Vietnam; Center for AI Research, VinUniversity, Vietnam
Massimo Piccardi
Massimo Piccardi
Professor, University of Technology Sydney
natural language processingcomputer visionpattern recognition
Wray Buntine
Wray Buntine
Professor, VinUniversity
Machine Learning