Nepali Passport Question Answering: A Low-Resource Dataset for Public Service Applications

📅 2026-03-04
🏛️ arXiv.org
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决尼泊尔语信息检索系统缺乏标注数据的问题,构建了护照服务相关的问答数据集,并使用微调的基于转换器的嵌入模型和混合检索方法提高检索性能。
📝 Abstract
Nepali, a low-resource language, faces significant challenges in building an effective information retrieval system due to the unavailability of annotated data and computational linguistic resources. In this study, we attempt to address this gap by preparing a pair-structured Nepali Question-Answer dataset. We focus on Frequently Asked Questions (FAQs) for passport-related services, building a data set for training and evaluation of IR models. In our study, we have fine-tuned transformer-based embedding models for semantic similarity in question-answer retrieval. The fine-tuned models were compared with the baseline BM25. In addition, we implement a hybrid retrieval approach, integrating fine-tuned models with BM25, and evaluate the performance of the hybrid retrieval. Our results show that the fine-tuned SBERT-based models outperform BM25, whereas multilingual E5 embedding-based models achieve the highest retrieval performance among all evaluated models.
Problem

Research questions and friction points this paper is trying to address.

Nepali
low-resource language
information retrieval
annotated data
FAQs
Innovation

Methods, ideas, or system contributions that make the work stand out.

fine-tuned transformer-based embedding models
hybrid retrieval approach
low-resource language dataset
💼 Related Jobs
No related jobs found.
F
Funghang Limbu Begha
ILPRL, Department of Computer Science & Engineering, Kathmandu University, Dhulikhel, Nepal
P
Praveen Acharya
School of Computing, Dublin City University, Dublin, Ireland
Bal Krishna Bal
Bal Krishna Bal
Professor of Computer Engineering, Kathmandu University
Natural Language ProcessingSentiment AnalysisSoftware Localization