Bridging the Gap in Bangla Healthcare: Machine Learning Based Disease Prediction Using a Symptoms-Disease Dataset
This study addresses the longstanding scarcity of localized disease prediction resources for Bengali-speaking populations, which has hindered equitable access to reliable health information. To bridge this gap, the authors present the first large-scale Bengali dataset comprising 85 diseases and 758 symptom–disease associations, which is publicly released. Leveraging this dataset, they design an ensemble learning model that integrates multiple machine learning algorithms with both soft and hard voting strategies to enable high-accuracy disease classification from Bengali symptom descriptions. Experimental results demonstrate that the proposed model achieves 98% accuracy on the newly curated dataset, significantly advancing the accessibility and equity of localized healthcare information for Bengali speakers.