Bangla Sentence Function Classification: Corpus Development, Model Benchmarking, and Interpretability

📅 2026-09-12
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过开发包含10,000句孟加拉语句子的语料库,并使用多种特征表示方法及异构集成模型,解决了孟加拉语文本句子功能分类缺乏基准资源的问题。
📝 Abstract
Automatic sentence function identification is important for many downstream natural language processing (NLP) applications such as dialogue systems, text-to-speech synthesis, and machine translation. However, benchmark resources for Bangla sentence function classification remain limited. To mitigate this gap, this paper introduces a corpus of 10,000 Bangla sentences, manually annotated into four functional categories, namely declarative, interrogative, imperative, and exclamatory. The corpus is nearly balanced across the four classes, with high annotation reliability reflected by a Fleiss\' Kappa of 0.82. Furthermore, we evaluate multiple feature representations, including Bag-of-Words (BoW), TF-IDF, and Word2Vec, with several classical machine learning classifiers. In addition, two heterogeneous ensemble models, namely Single-Level Ensemble (SLE) and Double-Level Ensemble (DLE), are utilized to improve classification performance. Experimental results show that TF-IDF consistently outperforms Word2Vec, likely due to its ability to emphasize discriminative lexical cues associated with sentence functions, particularly given the relatively small corpus used to train Word2Vec. The DLE model with TF-IDF features achieves the best performance with accuracy and macro-F1 of 0.95, demonstrating the effectiveness of sparse lexical representations and heterogeneous ensemble learning for this task. Further cross-validation confirms the robustness of the approach, while LIME-based interpretability provides insights into model predictions. The developed corpus and model benchmarking establish strong baselines for Bangla sentence function classification.
Problem

Research questions and friction points this paper is trying to address.

Bangla
sentence function classification
corpus development
benchmarking
Innovation

Methods, ideas, or system contributions that make the work stand out.

Bangla Sentence Function
TF-IDF
Heterogeneous Ensemble Models
Corpus Development
Model Interpretability
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Swapnil Kundu Argha
Department of CSE, Khulna University of Engineering & Technology, Bangladesh
A
Abdullah Al Shafi
Department of CSE, Khulna University of Engineering & Technology, Bangladesh; Institute of ICT, Khulna University of Engineering & Technology, Bangladesh
R
Rowzatul Zannat
Department of CSE, Daffodil International University, Bangladesh
S
Shoumik Barman Polok
Department of CSE, Khulna University of Engineering & Technology, Bangladesh
A
Abdul Muntakim
Department of Computer Science, Kennesaw State University, United States
J
Jannatul Ferdousi
Department of Computer Science, Kennesaw State University, United States
M
M. A. Moyeen
Department of CSE, Khulna University of Engineering & Technology, Bangladesh