ViQA-COVID: COVID-19 Machine Reading Comprehension Dataset for Vietnamese

πŸ“… 2025-04-21
πŸ›οΈ arXiv.org
πŸ“ˆ Citations: 1
✨ Influential: 1
πŸ“„ PDF
πŸ€– AI Summary
To address the scarcity of Vietnamese-language machine reading comprehension (MRC) resources and dedicated NLP benchmarks for pandemic-related text in low-resource languages, this work introduces ViQA-COVIDβ€”the first Vietnamese multi-span, multi-paragraph MRC dataset focused on COVID-19. Constructed manually from authentic pandemic documents, it comprises thousands of high-quality question-answer pairs annotated under a fine-grained multi-span schema, enabling robust fine-tuning and evaluation of mainstream models (e.g., PhoBERT). Its key contributions are twofold: (1) it pioneers multi-span answer extraction for Vietnamese MRC, and (2) it fills a critical gap by providing the first domain-specific MRC benchmark for pandemic-related text in a low-resource language. Empirical results demonstrate that ViQA-COVID substantially improves model performance on Vietnamese pandemic text understanding and has already enabled multiple follow-up studies in health-focused NLP.

Technology Category

Application Category

πŸ“ Abstract
After two years of appearance, COVID-19 has negatively affected people and normal life around the world. As in May 2022, there are more than 522 million cases and six million deaths worldwide (including nearly ten million cases and over forty-three thousand deaths in Vietnam). Economy and society are both severely affected. The variant of COVID-19, Omicron, has broken disease prevention measures of countries and rapidly increased number of infections. Resources overloading in treatment and epidemics prevention is happening all over the world. It can be seen that, application of artificial intelligence (AI) to support people at this time is extremely necessary. There have been many studies applying AI to prevent COVID-19 which are extremely useful, and studies on machine reading comprehension (MRC) are also in it. Realizing that, we created the first MRC dataset about COVID-19 for Vietnamese: ViQA-COVID and can be used to build models and systems, contributing to disease prevention. Besides, ViQA-COVID is also the first multi-span extraction MRC dataset for Vietnamese, we hope that it can contribute to promoting MRC studies in Vietnamese and multilingual.
Problem

Research questions and friction points this paper is trying to address.

Creating first Vietnamese COVID-19 MRC dataset
Addressing resource overload in pandemic prevention
Advancing multilingual machine reading comprehension
Innovation

Methods, ideas, or system contributions that make the work stand out.

First Vietnamese COVID-19 MRC dataset
Multi-span extraction MRC approach
AI application for epidemic prevention
πŸ”Ž Similar Papers
πŸ’Ό Related Jobs
No related jobs found.
H
Hai-Chung Nguyen
Hanoi University of Science and Technology
P
Phung Ngoc C. LΓͺ
Hanoi University of Science and Technology
V
Van-Chien Nguyen
Hanoi University of Science and Technology
H
Hang Thi Nguyen
Hanoi University of Science and Technology
T
Thuy Phuong Thi Nguyen
Hanoi University of Science and Technology