Predicting Anemia Among Under-Five Children in Nepal Using Machine Learning and Deep Learning

📅 2026-02-01
📈 Citations: 1
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high prevalence of anemia among children under five in Nepal by leveraging the 2022 Nepal Demographic and Health Survey data. To balance model interpretability with the identification of key risk factors, a consensus feature set was constructed through the integration of four feature selection methods: chi-square test, mutual information, point-biserial correlation, and Boruta. The binary classification performance of logistic regression, XGBoost, support vector machines (SVM), deep neural networks (DNN), and TabNet was systematically evaluated on imbalanced data. Results indicate that logistic regression achieved the highest F1-score (0.649) and recall (0.701), SVM attained the best AUC (0.736), and DNN yielded the highest accuracy (0.709), collectively demonstrating the feasibility and potential of machine learning approaches for early screening of childhood anemia.

Technology Category

Application Category

📝 Abstract
Childhood anemia remains a major public health challenge in Nepal and is associated with impaired growth, cognition, and increased morbidity. Using World Health Organization hemoglobin thresholds, we defined anemia status for children aged 6-59 months and formulated a binary classification task by grouping all anemia severities as \emph{anemic} versus \emph{not anemic}. We analyzed Nepal Demographic and Health Survey (NDHS 2022) microdata comprising 1,855 children and initially considered 48 candidate features spanning demographic, socioeconomic, maternal, and child health characteristics. To obtain a stable and substantiated feature set, we applied four features selection techniques (Chi-square, mutual information, point-biserial correlation, and Boruta) and prioritized features supported by multi-method consensus. Five features: child age, recent fever, household size, maternal anemia, and parasite deworming were consistently selected by all methods, while amenorrhea, ethnicity indicators, and provinces were frequently retained. We then compared eight traditional machine learning classifiers (LR, KNN, DT, RF, XGBoost, SVM, NB, LDA) with two deep learning models (DNN and TabNet) using standard evaluation metrics, emphasizing F1-score and recall due to class imbalance. Among all models, logistic regression attained the best recall (0.701) and the highest F1-score (0.649), while DNN achieved the highest accuracy (0.709), and SVM yielded the strongest discrimination with the highest AUC (0.736). Overall, the results indicate that both machine learning and deep learning models can provide competitive anemia prediction and the interpretable features such as child age, infection proxy, maternal anemia, and deworming history are central for risk stratification and public health screening in Nepal.
Problem

Research questions and friction points this paper is trying to address.

anemia
under-five children
Nepal
public health
prediction
Innovation

Methods, ideas, or system contributions that make the work stand out.

feature selection consensus
anemia prediction
machine learning
deep learning
class imbalance
D
Deepak Bastola
Florida Atlantic University, Department of Mathematics and Statistics
P
Pitambar Acharya
University of Alabama at Birmingham, Department of Applied Mathematics
D
Dipak Dulal
Eastern New Mexico University, Department of Mathematical Science
R
Rabina Dhakal
Apara Innovations
Y
Yang Li
Florida Atlantic University, Department of Mathematics and Statistics