Explainable Machine Learning for Current Stunting Status Classification in Toddlers: A Comparative Evaluation of SMOTEENN and SHAP-Based Interpretation
DOI:
https://doi.org/10.37034/medinftech.v4i3.170Keywords:
Explainable Machine Learning, Machine Learning, SHAP, SMOTEENN, Stunting ClassificationAbstract
Stunting remains a major nutritional problem that requires accurate and reliable assessment of children’s current growth status. Machine learning can support stunting classification; however, class imbalance and limited model interpretability remain important challenges. This study aims to evaluate the effect of SMOTEENN across classification algorithms with different learning characteristics and interpret the best-performing model using SHAP. Two secondary datasets were integrated, resulting in 41,176 records, of which 41,173 were retained after data cleaning and numerical conversion. Logistic Regression, Decision Tree, Random Forest, and XGBoost were evaluated using original and SMOTEENN-balanced training data. Hyperparameter optimization was performed using RandomizedSearchCV with stratified five-fold cross-validation and ROC-AUC as the optimization metric. The optimized configurations were fixed and applied to both experimental scenarios. Model performance was evaluated using Accuracy, Precision, Recall, F1-Score, and ROC-AUC on an independent test set. Random Forest trained on the original data achieved the best overall performance, with an Accuracy of 0.9600, Precision of 0.9584, F1-Score of 0.9505, and ROC-AUC of 0.9927. SMOTEENN consistently improved Recall across all models but generally reduced the overall performance of tree-based models. The largest improvement occurred in Logistic Regression, with Recall increasing from 0.7058 to 0.8657. SHAP analysis identified Height and Age (Month) as the most influential features. These findings show that the effect of SMOTEENN depends on dataset and algorithm characteristics, while SHAP provides transparent interpretation of current stunting-status classification within the analyzed datasets.
Downloads
References
C. P. W. Kase and S. Y. J. Prasetyo, “Analisis Faktor Risiko Stunting pada Balita di Desa Kesetnana Menggunakan Metode Random Forest,” J. Indones. Manaj. Inform. dan Komun., vol. 6, no. 3, pp. 1556–1566, 2025, doi: 10.63447/jimik.v6i3.1449.
I. Indriani, M. Mujahadatuljannah, and R. Rabiatunnisa, “Factors Affecting Incidence of Stunting in Infants and Toddlers,” J. Surya Med., vol. 9, no. 3, pp. 131–136, 2023, [Online]. Available: https://akper-sandikarsa.e-journal.id/JIKSH/article/view/1019/584
K. K. R. Indonesia, “Survei Status Gizi Indonesia (SSGI) 2024 dalam Angka,” Jakarta, 2025. [Online]. Available: https://stunting.go.id/ssgi-2024-prevalensi-stunting-nasional-turun-jadi-198-capai-angka-di-bawah-proyeksi-bappenas/
World Health Organization, WHO Child Growth Standards: Length/Height-for-Age, Weight-for-Age, Weight-for-Length, Weight-for-Height and Body Mass Index-for-Age: Methods and Development. Geneva, Switzerland: World Health Organization, 2006. [Online]. Available: https://www.who.int/publications/i/item/924154693X
R. Ratnasari, A. J. Wahidin, and T. H. Andika, “Deteksi Dini Stunting Pada Anak Berdasarkan Indikator Antropometri dengan Menggunakan Algoritma Machine Learning,” J. Algoritm., vol. 21, no. 2, pp. 378–387, 2024, doi: 10.33364/algoritma/v.21-2.2122.
T. Hidayat, I. Sembiring, H. D. Purnomo, and A. Iriani, “Prediksi Prevalensi Stunting Balita dengan Pendekatan Algoritma Support Vector Machine dan Synthetic Minority Oversampling Technique (SMOTE),” J. Pekommas, vol. 10, no. 1, pp. 9–16, 2025, doi: 10.56873/jpkm.v9i1.5389.
N. I. Syahfitri, A. P. Juledi, and R. Muti’ah, “Comparative Analysis of Machine Learning Algorithm Performance in Predicting Stunting in Toddlers,” Sinkron, vol. 8, no. 3, pp. 1452–1462, 2024, doi: 10.33395/sinkron.v8i3.13698.
E. N. Candra, I. Cholissodin, and R. C. Wihandika, “Klasifikasi Status Gizi Balita Menggunakan Metode Optimasi Random Forest Dengan Algoritme Genetika (Studi Kasus: Puskesmas Cakru),” J. Pengemb. Teknol. Inf. dan Ilmu Komput., vol. 6, no. 5, pp. 2188–2197, 2022, [Online]. Available: http://j-ptiik.ub.ac.id
M. K. Ayele, G. A. Baye, S. H. Yesuf, A. A. Engda, and E. T. Mitiku, “Predicting stunting status among under five children in ethiopia using ensemblemachine learning algorithms,” Sci. Rep., vol. 15, no. 1, pp. 1–11, 2025, doi: 10.1038/s41598-025-03206-1.
T. Sugihartono, B. Wijaya, Marini, A. F. Alkayes, and H. A. Anugrah, “Optimizing Stunting Detection through SMOTE and Machine Learning: a Comparative Study of XGBoost, Random Forest, SVM, and k-NN,” J. Appl. Data Sci., vol. 6, no. 1, pp. 667–682, 2025, doi: 10.47738/jads.v6i1.494.
W. R. Mgomezulu, P. Thangata, B. Mkandawire, and N. Amoah, “Advancing predictive analytics in child malnutrition: Machine, ensemble and deep learning models with balanced class distribution for early detection of stunting and wasting,” Hum. Nutr. Metab., vol. 42, no. August, p. 200340, 2025, doi: 10.1016/j.hnm.2025.200340.
N. El Furqany, M. Subianto, and A. Rusyana, “Hybrid Ensemble Learning with SMOTEENN and Soft Voting for Stunting Risk Prediction: A SHAP-Based Explainable Approach,” J. Appl. Data Sci., vol. 6, no. 4, pp. 2989–3004, 2025, doi: 10.47738/jads.v6i4.829.
N. Novalina, I. A. A. Tarigan, F. K. Kameela, and M. Rizkinia, “Benchmarking machine learning algorithm for stunting risk prediction in Indonesia,” Bull. Electr. Eng. Informatics, vol. 14, no. 3, pp. 2252–2263, 2025, doi: 10.11591/eei.v14i3.8997.
Y. Li, Y. Yang, P. Song, L. Duan, and R. Ren, “An improved SMOTE algorithm for enhanced imbalanced data classification by expanding sample generation space,” Sci. Rep., vol. 15, no. 1, pp. 1–21, 2025, doi: 10.1038/s41598-025-09506-w.
H. W. Loh, C. P. Ooi, S. Seoni, P. D. Barua, F. Molinari, and U. R. Acharya, “Application of explainable artificial intelligence for healthcare: A systematic review of the last decade (2011–2022),” Comput. Methods Programs Biomed., vol. 226, p. 107161, 2022, doi: 10.1016/j.cmpb.2022.107161.
N. Hettikankanamage, N. Shafiabady, F. Chatteur, R. M. X. Wu, F. Ud Din, and J. Zhou, “eXplainable Artificial Intelligence (XAI): A Systematic Review for Unveiling the Black Box Models and Their Relevance to Biomedical Imaging and Sensing,” Sensors, vol. 25, no. 21, pp. 1–31, 2025, doi: 10.3390/s25216649.
A. H. RS et al., “Dataset Stunting and Nutritional Status of Toddler from Jeneponto Regency, South Sulawesi, Indonesia,” 2026, Mendeley Data. doi: 10.17632/WZWPC9J5BX.
J. Peng, J. Hahn, and K. W. Huang, “Handling Missing Values in Information Systems Research: A Review of Methods and Assumptions,” Inf. Syst. Res., vol. 34, no. 1, pp. 5–26, 2023, doi: 10.1287/isre.2022.1104.
V. Çetin and O. Yıldız, “A comprehensive review on data preprocessing techniques in data analysis,” Pamukkale Univ. J. Eng. Sci., vol. 28, no. 2, pp. 299–312, 2022, doi: 10.5505/pajes.2021.62687.
I. Souiden, M. N. Omri, and Z. Brahmi, “A survey of outlier detection in high dimensional data streams,” Comput. Sci. Rev., vol. 44, p. 100463, 2022, doi: 10.1016/j.cosrev.2022.100463.
F. Pargent, F. Pfisterer, J. Thomas, and B. Bischl, “Regularized target encoding outperforms traditional methods in supervised machine learning with high cardinality features,” Comput. Stat., vol. 37, no. 5, pp. 2671–2692, 2022, doi: 10.1007/s00180-022-01207-6.
M. R. Islam, A. A. Lima, S. C. Das, M. F. Mridha, A. R. Prodeep, and Y. Watanobe, “A Comprehensive Survey on the Process, Methods, Evaluation, and Challenges of Feature Selection,” IEEE Access, vol. 10, no. July, pp. 99595–99632, 2022, doi: 10.1109/ACCESS.2022.3205618.
A. Géron, Hands-on machine learning with Scikit-Learn, Keras, and TensorFlow. “ O’Reilly Media, Inc.,” 2022.
L. B. V De Amorim, G. D. C. Cavalcanti, and R. M. O. Cruz, “The choice of scaling technique matters for classification performance,” Appl. Soft Comput., vol. 133, p. 109924, 2023, doi: 10.1016/j.asoc.2022.109924.
G. E. A. P. A. Batista, R. C. Prati, and M. C. Monard, “A study of the behavior of several methods for balancing machine learning training data,” ACM SIGKDD Explor. Newsl., vol. 6, no. 1, pp. 20–29, 2004, doi: 10.1145/1007730.1007735.
E. Dumitrescu, S. Hué, C. Hurlin, and S. Tokpavi, “Machine learning for credit scoring: Improving logistic regression with non-linear decision-tree effects,” Eur. J. Oper. Res., vol. 297, no. 3, pp. 1178–1192, 2022, doi: 10.1016/j.ejor.2021.06.053.
C. Kwenda, M. Gwetu, and J. V. F. Dombeu, “Machine Learning Methods for Forest Image Analysis and Classification : A Survey of the State of the Art,” IEEE Access, vol. 10, pp. 45290–45316, 2022, doi: 10.1109/ACCESS.2022.3170049.
I. D. Mienye, Y. Sun, and S. Member, “A Survey of Ensemble Learning : Concepts , Algorithms , Applications , and Prospects,” IEEE Access, vol. 10, no. September, pp. 99129–99149, 2022, doi: 10.1109/ACCESS.2022.3207287.
G. Naidu, T. Zuva, and E. M. Sibanda, “A Review of Evaluation Metrics in Machine Learning Algorithms BT - Artificial Intelligence Application in Networks and Systems,” R. Silhavy and P. Silhavy, Eds., Cham: Springer International Publishing, 2023, pp. 15–25.
H. Chen, I. C. Covert, S. M. Lundberg, and S.-I. Lee, “Algorithms to estimate Shapley value feature attributions,” Nat. Mach. Intell., vol. 5, no. 6, pp. 590–601, 2023, doi: 10.1038/s42256-023-00657-x.
M. Altalhan, A. Algarni, and M. T. Alouane, “Imbalanced Data Problem in Machine Learning : A Review,” IEEE Access, vol. 13, no. January, pp. 13686–13699, 2025, doi: 10.1109/ACCESS.2025.3531662.
T. Sugihartono, D. Soetarno, R. Sulaiman, Sarwindah, Marini, and Fitriyani, “Ensemble learning for pediatric stunting detection: A comparative study of XGBoost, Random Forest, and LightGBM with oversampling techniques,” J. Inf. Syst. Inform., vol. 8, no. 2, pp. 1672–1692, 2026, doi: 10.63158/journalisi.v8i2.1568.
M. Carvalho, A. J. Pinho, and S. Brás, “Resampling approaches to handle class imbalance : a review from a data perspective,” J. Big Data, 2025, doi: 10.1186/s40537-025-01119-4.







