AutoViVQA: A Large-Scale Automatically Constructed Dataset for Vietnamese Visual Question Answering

📅 2026-03-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the scarcity of large-scale, high-quality, and visually grounded datasets for Vietnamese Visual Question Answering (VQA) by introducing AutoViVQA, the first large-scale automatically constructed Vietnamese VQA dataset. The authors propose a multimodal modeling approach that integrates PhoBERT and Vision Transformer to effectively leverage both linguistic and visual information. Furthermore, they conduct a systematic evaluation of automatic metrics—including BLEU, METEOR, CIDEr, and F1—in the context of multilingual VQA, analyzing their alignment with human judgments. This study not only establishes a new benchmark for low-resource language VQA but also provides critical insights into the reliability and limitations of commonly used automatic evaluation metrics across languages, offering valuable guidance for future research in multilingual multimodal assessment.

Technology Category

Application Category

📝 Abstract
Visual Question Answering (VQA) is a fundamental multimodal task that requires models to jointly understand visual and textual information. Early VQA systems relied heavily on language biases, motivating subsequent work to emphasize visual grounding and balanced datasets. With the success of large-scale pre-trained transformers for both text and vision domains -- such as PhoBERT for Vietnamese language understanding and Vision Transformers (ViT) for image representation learning -- multimodal fusion has achieved remarkable progress. For Vietnamese VQA, several datasets have been introduced to promote research in low-resource multimodal learning, including ViVQA, OpenViVQA, and the recently proposed ViTextVQA. These resources enable benchmarking of models that integrate linguistic and visual features in the Vietnamese context. Evaluation of VQA systems often employs automatic metrics originally designed for image captioning or machine translation, such as BLEU, METEOR, CIDEr, Recall, Precision, and F1-score. However, recent research suggests that large language models can further improve the alignment between automatic evaluation and human judgment in VQA tasks. In this work, we explore Vietnamese Visual Question Answering using transformer-based architectures, leveraging both textual and visual pre-training while systematically comparing automatic evaluation metrics under multilingual settings.
Problem

Research questions and friction points this paper is trying to address.

Vietnamese Visual Question Answering
low-resource multimodal learning
automatic evaluation metrics
dataset construction
multimodal alignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

AutoViVQA
Vietnamese Visual Question Answering
Transformer-based multimodal fusion
Automatic dataset construction
Multilingual VQA evaluation
💼 Related Jobs
No related jobs found.
N
Nguyen Anh Tuong
Faculty of Information Technology, University of Science, VNU-HCM, Vietnam
P
Phan Ba Duc
Faculty of Information Technology, University of Science, VNU-HCM, Vietnam
N
Nguyen Trung Quoc
Faculty of Information Technology, University of Science, VNU-HCM, Vietnam
T
Tran Dac Thinh
Faculty of Information Technology, University of Science, VNU-HCM, Vietnam
D
Dang Duy Lan
Faculty of Information Technology, University of Science, VNU-HCM, Vietnam
N
Nguyen Quoc Thinh
Faculty of Information Technology, University of Science, VNU-HCM, Vietnam
Tung Le
Tung Le
Dr, , Lecturer.
Natural Language ProcessingVisual Question Answering