BeliN: A Novel Corpus for Bengali Religious News Headline Generation using Contextual Feature Fusion

📅 2025-01-02
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the low-resource, multi-dimensional modeling challenge of Bangla religious news headline generation. We propose MultiGen, the first model to jointly encode article content with multi-faceted contextual features—including news category, sentiment polarity, and core topic—to enable fine-grained semantic guidance. To support this research, we introduce and publicly release BeliN, the first benchmark corpus for Bangla religious news headline generation, featuring human-annotated headlines, categories, sentiment labels, and thematic tags. MultiGen builds upon multilingual pretrained language models (e.g., BanglaT5, mBART, mT5, mT0) and incorporates a lightweight multi-source feature fusion architecture. On BeliN, MultiGen significantly outperforms content-only baselines: BLEU-4 improves by +2.53 to 18.61, and ROUGE-L by +1.11 to 24.19. Both code and dataset are openly available.

Technology Category

Application Category

📝 Abstract
Automatic text summarization, particularly headline generation, remains a critical yet underexplored area for Bengali religious news. Existing approaches to headline generation typically rely solely on the article content, overlooking crucial contextual features such as sentiment, category, and aspect. This limitation significantly hinders their effectiveness and overall performance. This study addresses this limitation by introducing a novel corpus, BeliN (Bengali Religious News) - comprising religious news articles from prominent Bangladeshi online newspapers, and MultiGen - a contextual multi-input feature fusion headline generation approach. Leveraging transformer-based pre-trained language models such as BanglaT5, mBART, mT5, and mT0, MultiGen integrates additional contextual features - including category, aspect, and sentiment - with the news content. This fusion enables the model to capture critical contextual information often overlooked by traditional methods. Experimental results demonstrate the superiority of MultiGen over the baseline approach that uses only news content, achieving a BLEU score of 18.61 and ROUGE-L score of 24.19, compared to baseline approach scores of 16.08 and 23.08, respectively. These findings underscore the importance of incorporating contextual features in headline generation for low-resource languages. By bridging linguistic and cultural gaps, this research advances natural language processing for Bengali and other underrepresented languages. To promote reproducibility and further exploration, the dataset and implementation code are publicly accessible at https://github.com/akabircs/BeliN.
Problem

Research questions and friction points this paper is trying to address.

Automatic Title Generation
Bangla Religious News
Multi-faceted Information Incorporation
Innovation

Methods, ideas, or system contributions that make the work stand out.

BeliN Database
MultiGen Method
Contextual Title Generation
🔎 Similar Papers
No similar papers found.
Md Osama
Md Osama
Chittagong University of Engineering and Technology (CUET)
Bangla Language ProcessingNatural Language Processing
A
Ashim Dey
Department of Computer Science and Engineering, Chittagong University of Engineering and Technology, Raozan, Chittagong, 4349, Bangladesh
Kawsar Ahmed
Kawsar Ahmed
Research Assistant, Dept. of CSE, Chittagong University of Engineering and Technology
Natural Language ProcessingMachine LearningDeep Learning
M
Muhammad Ashad Kabir
School of Computing, Mathematics and Engineering, Charles Sturt University, Bathurst, NSW, 2795, Australia