๐ค AI Summary
This study addresses the longstanding scarcity of localized disease prediction resources for Bengali-speaking populations, which has hindered equitable access to reliable health information. To bridge this gap, the authors present the first large-scale Bengali dataset comprising 85 diseases and 758 symptomโdisease associations, which is publicly released. Leveraging this dataset, they design an ensemble learning model that integrates multiple machine learning algorithms with both soft and hard voting strategies to enable high-accuracy disease classification from Bengali symptom descriptions. Experimental results demonstrate that the proposed model achieves 98% accuracy on the newly curated dataset, significantly advancing the accessibility and equity of localized healthcare information for Bengali speakers.
๐ Abstract
Increased access to reliable health information is essential for non-English-speaking populations, yet resources in Bangla for disease prediction remain limited. This study addresses this gap by developing a comprehensive Bangla symptoms-disease dataset containing 758 unique symptom-disease relationships spanning 85 diseases. The dataset enables the prediction of diseases based on Bangla symptom inputs, supporting healthcare accessibility for Bengali-speaking populations. Using this dataset, we evaluated multiple machine learning models to predict diseases based on symptoms provided in Bangla and analyzed their performance on our dataset. Both soft and hard voting ensemble approaches combining top-performing models achieved 98% accuracy, demonstrating superior robustness and generalization. Our work establishes a foundational resource for disease prediction in Bangla, paving the way for future advancements in localized health informatics and diagnostic tools. This contribution aims to enhance equitable access to health information for Bangla-speaking communities, particularly for early disease detection and healthcare interventions.