Perbandingan Kinerja XGBoost dan Random Forest Menggunakan SMOTE pada Klasifikasi Diabetes Multi-Kelas


  • Java Sika Maulana Universitas Budi Luhur, Jakarta, Indonesia
  • Safitri Juanita * Mail Universitas Budi Luhur, Jakarta, Indonesia
  • (*) Corresponding Author
Keywords: Diabetes Mellitus; Multiclass Classification; SMOTE; XGBoost; Random Forest

Abstract

Type 2 diabetes mellitus is a metabolic disorder that requires an early detection system to support accurate diagnosis identification. One of the major challenges in developing classification models based on clinical medical records is class imbalance, which may cause models to be biased toward the majority class, particularly in multiclass classification involving the Prediabetes class, which represents only 5.3% of the total data. Failure to accurately identify the Prediabetes class may have serious clinical consequences, as this stage still provides an opportunity for early intervention to prevent progression to Diabetes. This study compares the performance of Extreme Gradient Boosting (XGBoost) and Random Forest for multiclass diabetes classification (Normal, Prediabetes, and Diabetes) using clinical data obtained from Medical City Hospital and Al-Kindy Teaching Hospital, Iraq. A Split-First, Resample-Later procedure was employed to prevent data leakage, while Synthetic Minority Over-sampling Technique (SMOTE) was applied to balance the training data, and Grid Search was used for hyperparameter optimization across different train-test split ratios (60:40, 70:30, 80:20, and 90:10). This study provides a comprehensive evaluation of XGBoost and Random Forest on imbalanced multiclass clinical data by comparing their performance before and after SMOTE application across different train–test split ratios using the Split-First, Resample-Later procedure to prevent information leakage between the training and test set. The experimental results demonstrate that Random Forest achieved more stable performance than XGBoost across all evaluation scenarios, both before and after SMOTE application. Both models achieved their best performance with an 80:20 train–test split ratio, whereas SMOTE significantly improved the performance of XGBoost only under the 60:40 split ratio. Furthermore, feature importance analysis identified HbA1c, BMI, and AGE as the most influential clinical attributes for diabetes classification.

Downloads

Download data is not yet available.

References

Adhi Pratama, N., & Wahyu Utomo, D. (2026). Deteksi diabetes mellitus dengan menggunakan teknik ensemble XGBoost dan LightGBM. Jurnal Informatika Sunan Kalijaga, 11(1), 1–12. https://doi.org/10.14421/jiska.4908

Aghware, F. O., Akazue, M. I., Okpor, M. D., Malasowe, B. O., Aghaunor, T. C., Ugbotu, E. V., Ojugo, A. A., Ako, R. E., Geteloma, V. O., Odiakaose, C. C., Eboka, A. O., & Onyemenem, S. I. (2025). Effects of data balancing in diabetes mellitus detection: A comparative XGBoost and Random Forest learning approach. NIPES - Journal of Science and Technology Research, 7(1), 1–11. https://doi.org/10.37933/nipes/7.1.2025.1

Agyemang, E. F., Mensah, J. A., Nyarko, E., Arku, D., Mbeah-Baiden, B., Opoku, E., & Noye Nortey, E. N. (2025). Addressing class imbalance problem in health data classification: Practical application from an oversampling viewpoint. Applied Computational Intelligence and Soft Computing, 2025(1), 1–20. https://doi.org/10.1155/acis/1013769

Alfarisi, M. A., Sumanto, B., Putu, I., & Rachmawan, F. (2025). Pengembangan model machine learning untuk deteksi penyakit diabetes menggunakan analisis gini importance. Journal of Internet and Software Engineering, 6(2), 127–134. https://doi.org/10.22146/jise.v6i2.16551

Amri, Z., Rodi, M., Wathani, N., Bagja, A., & Zulkipli. (2025). Prediksi diabetes menggunakan algoritma K-Nearest (KNN) teknik SMOTE-ENN. Infotek: Jurnal Informatika Dan Teknologi, 8(1), 193–204. https://doi.org/10.29408/jit.v8i1.27975

Astofa, A., Rosyani, P., & Apandi, S. (2025). Evaluasi komparatif algoritma machine learning untuk prediksi dini diabetes. Bulletin of Computer Science and Research, 6(1), 558–565. https://doi.org/10.47065/bulletincsr.v6i1.859

Bagus, G., Sidi, A., Arsana, M., & Gunawan, R. (2022). Peningkatan akurasi algoritma C4.5 menggunakan particle swarm optimization untuk mendeteksi penyakit diabetes. Bit, 19(2), 90–97. https://doi.org/10.36080/bit.v19i2.2044

Goldney, J., Sargeant, J. A., & Davies, M. J. (2023). Incretins and microvascular complications of diabetes: Neuropathy, nephropathy, retinopathy and microangiopathy. Diabetologia, 66(10), 1832–1845. https://doi.org/10.1007/s00125-023-05988-3

Haris Yunianto, A., & Rosi Subhiyakto, E. (2025). Perbandingan kinerja algoritma k-nearest neighbors dan decision tree untuk klasifikasi diabetes. Technology and Science (BITS), 6(4), 2601–2611. https://doi.org/10.47065/bits.v6i4.6550

Iftikhar, W., Yaseen, M., Rahman, G., Nauman, M. A., & Khattak, U. F. (2026). Improving diabetes prediction accuracy and interpretability with SMOTE and SHAP. Engineering, Technology & Applied Science Research, 16(2), 34276–34282. https://doi.org/10.48084/etasr.16247

International Diabetes Federation. (2025). IDF Diabetes Atlas (D. J. Magliano, E. J. Boyko, I. Genitsaridi, L. Piemonte, P. Riley, & P. Salpea, Eds.; 11th ed.). International Diabetes Federation. https://diabetesatlas.org/resources/idf-diabetes-atlas-2025/

Jang, Y. (2025). Feature-based ensemble modeling for addressing diabetes data imbalance using the SMOTE, RUS, and random forest methods: A prediction study. Ewha Medical Journal, 48(2), 1–8. https://doi.org/10.12771/emj.2025.00353

Munshi, R. M., Munshi, L. R., Himdi, H., Qashlan, A., Munshi, R., Alyahyawy, O. Y., & Khayyat, M. M. (2025). Optimising hyperparameters with a tree structured parzen estimator to improve diabetes prediction. Scientific Reports, 15(1), 1–10. https://doi.org/10.1038/s41598-025-19295-x

Pratama, A., Nurcahyo, A. C., & Firgia, L. (2023). Penerapan machine learning dengan algoritma logistik regresi untuk memprediksi diabetes. Prosiding CORISINDO, 116–121. https://ojs.stmikpontianak.ac.id/corisindo/article/view/30

Prayoga, N., & Ridla, M. A. (2026). Analisis prediksi diabetes menggunakan decision tree pada dataset diabetes pima indians. Journal of Information System and Application Development, 4(1), 146–153. https://doi.org/10.26905/jisad.v4i1.16494

Rashid, A. (2026, April 30). Diabetes Dataset [Data set]. Mendeley Data. https://doi.org/10.17632/wj9rwkp9c2.1

Sembiring Depari, A. D., Tania, K. D., & Sevtiyuni, P. E. (2025). Penerapan metode machine learning dan teknik SMOTE untuk prediksi diabetes. Jurnal Sistem Komputer Dan Informatika (JSON), 7(2), 436–447. https://doi.org/10.30865/json.v7i2.9032

Susanto, E. R., & Cahyana, A. (2025). Penerapan algoritma XGBoost untuk prediksi diabetes: Analisis confusion matrix dan ROC curve. Fountain of Informatics Journal, 10(1), 40–50. https://doi.org/10.21111/fij.v10i1.14311

Syahla, H., Izzudin, H., Pratama, F. A., Rahmatullah, B., Wahidin, A. J., & Kurniawati, I. (2025). Klasifikasi indikator kesehatan diabetes menggunakan algoritma random forest. Jurnal Teknik Informatika Dan Teknologi Informasi, 5(3), 401–416. https://doi.org/10.55606/jutiti.v5i3.6338

Utomo, S., Sigit, S., Sulistyowati, N., Arman Prasetya, D., & Maulana Fahrudin, T. (2025). Perbandingan kinerja algoritma XGBoost dan CatBoost dalam klasifikasi risiko penyakit diabetes. Alinier, 6(2), 84–96. https://doi.org/10.36040/alinier.v6i2.14772

Vlachas, C., Damianos, L., Gousetis, N., Mouratidis, I., Kelepouris, D., Kollias, K.-F., Asimopoulos, N., & Fragulis, G. F. (2022). Random forest classification algorithm for medical industry data. SHS Web of Conferences, 1–6. https://doi.org/10.1051/shsconf/202213903008

Wongso Prawiro, A., Rofilah, A. K., Putri, I. P. H., Praditna, L. M. A., Hasanah, M., & Husodo, D. P. (2025). Diabetic ketoacidosis and hyperosmolar hyperglycemic state: Diagnosis and management in emergency condition : A literature review. Jurnal Biologi Tropis, 25(4a), 620–626. https://doi.org/10.29303/jbt.v25i4a.11083

Zhou, F., Hu, S., Du, X., & Lu, Z. (2025). Anston attentional network for structured data based stroke risk prediction in smart aging. Scientific Reports, 15(1), 1–20. https://doi.org/10.1038/s41598-025-18758-5


Bila bermanfaat silahkan share artikel ini

Berikan Komentar Anda terhadap artikel Perbandingan Kinerja XGBoost dan Random Forest Menggunakan SMOTE pada Klasifikasi Diabetes Multi-Kelas

Dimensions Badge
Article History
Published: 2026-06-30
Abstract View: 0 times
PDF Download: 0 times
Issue
Section
Articles