Identifikasi Faktor Dominan Kegagalan Akademik pada Data Tidak Seimbang Menggunaan Ensemble Learning dan Hybrid SMOTE-ENN
Abstract
Academic failure and student dropout remain significant challenges in higher education because they affect educational quality and institutional performance. One of the major challenges in developing an effective dropout prediction model is class imbalance, which causes classification algorithms to be biased toward the majority class and reduces their ability to identify minority-class instances. This study aims to identify the dominant factors influencing academic failure by integrating the Hybrid Synthetic Minority Oversampling Technique–Edited Nearest Neighbour (SMOTE-ENN) with Ensemble Learning algorithms, including Decision Tree, Random Forest, XGBoost, and Voting Ensemble. Hybrid SMOTE-ENN was selected because it combines minority-class oversampling with noise and overlapping data removal, resulting in a more balanced training dataset and improving classification performance. The experiment was conducted using the Predict Students' Dropout and Academic Success dataset containing 4,424 student records. The research procedure consisted of data preprocessing, train–test splitting, Hybrid SMOTE-ENN resampling, model training, performance evaluation using accuracy, precision, recall, and F1-score, followed by feature importance analysis. Experimental results demonstrate that XGBoost with Hybrid SMOTE-ENN achieved the best performance, obtaining an accuracy of 87.12%, precision of 87.20%, recall of 87.12%, and F1-score of 87.15%. More importantly, the proposed model achieved a dropout-class recall of 80.99%, indicating its effectiveness in identifying students at risk of academic failure. Feature importance analysis revealed that Curricular Units 2nd Semester (Approved), Curricular Units 1st Semester (Approved), and Curricular Units 2nd Semester (Grade) are the three most influential factors affecting student dropout risk. These findings contribute to the development of an early warning system based on machine learning to support academic decision-making for imbalanced educational datasets.
Downloads
References
D. Anggraini, P. Korespondensi, and P. Mesin, “Data Mining Pendidikan: Predeksi Gaya Belajar Mahasiswa Teknik Menggunakan Machine Learning,” J. Teknol. Inf. dan Ilmu Komput., vol. 11, no. 3, pp. 563–572, 2025, doi: https://doi.org/10.25126/jtiik.2025129190.
Z. Parinzka, S. Surono, and A. Thobirin, “Algortima Dupport Vector Regression dan Analisis Long Short-Term Memory Sebagai Penaganan Missing Data,” J. Teknol. Inf. dan Ilmu Komput., vol. 13, no. 1, pp. 73–82, 2026, doi: https://doi.org/10.25126/jtiik.2026131.
M. A. Sadat, A. Pambudi, S. Ibad, U. D. Nuswantoro, I. Systems, and S. Program, “Comparison of Algoritm Between Classfication & Regression Trees and Support Vector Machine in Determining Student Accptance in State Universities,” J. Tek. Inform., vol. 4, no. 6, pp. 1589–1604, 2024, doi: https://doi.org/10.52436/1.jutif.2023.4.6.1565.
N. Mubarak, S. Susandri, and A. Zamsuri, “Geodetically-Enhanced Hybrid GRU with Adaptive Dropout and Dynamic L2 Regularization for Earthquake Parameter Prediction in Indonesia,” J. Tek. Inform., vol. 7, no. 2, pp. 1483–1499, 2026, doi: https://doi.org/10.52436/1.jutif.2026.7.2.5478.
N. Fadilah et al., “Perbandingan Kinerja Algortima Klasifikasi untuk Pemilihan Tempat Promosi Upaya Meningkatkan Jumlah Mahasiswa Baru,” J. Teknol. Inf. dan Ilmu Komput., vol. 13, no. 1, pp. 1–10, 2026, doi: https://doi.org/10.25126/jtiik.2026131.
F. Maulana and E. B. Setiawan, “Performance of Deep Feed-Forward Neural Network Algorithm Based on Content-Based Filtering Approach,” INTENSIF, vol. 8, no. 2, pp. 278–294, 2024, doi: https://doi.org/10.29407/intensif.v8i2.22904.
J. J. Purnama et al., “Klasifikasi Mahasiswa Her Berbasis Algoritma SVM dan Decision Tree,” J. Teknol. Inf. dan Ilmu Komput., vol. 7, no. 6, pp. 1253–1260, 2020, doi: 10.25126/jtiik.202073080.
T. Suprapti, B. Nurhakim, B. Warni, A. Hermina, and V. A. Syahputra, “A Decision Tree Model with Grid Search Optimization for Scholarship Recipient Classification,” J. Tek. Inform., vol. 6, no. 5, pp. 3800–3813, 2025, doi: https://doi.org/10.52436/1.jutif.2025.6.5.5235.
E. J. Kusuma, R. Nurmandhani, I. Pantiawati, and Y. M. Manglapy, “A Random Forest and SMOTE-Based Machine Learning Model for Predicting Recurrence in Papillary Thyroid Carcinoma,” J. Tek. Inform., vol. 6, no. 4, pp. 2019–2034, 2025, doi: https://doi.org/10.52436/1.jutif.2025.6.4.4854.
M. Rizki, A. Hermawan, and D. Avianto, “Learning Accuracy with Particle Swarm Optimization for Music Genre Classification Using Recurrent Neural Networks,” Matrik J. Manajemen, Tek. Inform. dan Rekayasa Komput., vol. 23, no. 2, pp. 297–308, 2024, doi: 10.30812/matrik.v23i2.3037.
Y. Desnelita, N. Nasution, L. Suryati, and F. Zoromi, “Dampak SMOTE terhadap Kinerja Random Forest Classifier berdasarkan Data Tidak seimbang Impact of SMOTE on Random Forest Classifier Performance based on Imbalanced Data,” Matrik J. Manajemen, Tek. Inform. dan Rekayasa Komput., vol. 21, no. 3, 2022, doi: 10.30812/matrik.v21i3.1726.
I. G. T. Suryawan, I. P. Agus, and E. Darma, “Optimasi Convulutional Neural Netwoek untuk Deteksi Covid-19 pada X-RAY Thorax Berbasis Dropout,” J. Teknol. Inf. dan Ilmu Komput., vol. 9, no. 3, pp. 551–558, 2022, doi: 10.25126/jtiik.202295143.
M. Lalu Ganda Rady Putra, Didik Dwi Prasetya, “Student Dropout Prediction Using Random Forest and XGBoost Method,” INTENSIF, vol. 9, no. 1, pp. 147–157, 2025, doi: https://doi.org/10.29407/intensif.v9i1.21191.
M. Ricky, P. Putra, and E. Utami, “Comparative Analysis of Hybrid Model Performance Using Stacking and Blending Techniques for Student Drop-Out Prediction in MOOC,” J. RESTI (Rekayasa Sist. dan Teknol. Informasi), vol. 5, no. 158, pp. 1–6, 2026, doi: https://doi.org/10.29207/resti.v8i3.5760.
S. H. Hasanah, E. Julianti, U. Terbuka, S. Informasi, and U. Terbuka, “Analysis of CART and Random Forest on Statistics Student Status at Universitas Terbuka,” INTENSIF, vol. 6, no. 1, pp. 56–65, 2022, doi: https://doi.org/10.29407/intensif.v6i1.16156.
A. A. Nababan, M. Jannah, M. Aulina, and D. Andrian, “Prediksi Kualitas Udara menggunakan XGBoost dengan Synthetic Minority Overampling Technique (SMOTE) Beradasarkan Indeks Standar Pencemaran Udara (ISPU),” J. Tek. Inform. Kaputama, vol. 7, no. 1, pp. 214–219, 2023, doi: https://doi.org/10.59697/jtik.v7i1.66.
A. Anggrawan, H. Hairani, and C. Satria, “Improving SVM Classification Performance on Unbalanced Student Graduation Time Data Using SMOTE,” Int. J. Inf. Educ. Technol., vol. 13, no. 2, 2023, doi: 10.18178/ijiet.2023.13.2.1806.
V. Realinho, M. Vieira Martins, J. Machado, and L. Baptista, “Predict Students Dropout and Academic Success,” 2021, UCI Machine Learning Repository, Irvine, California. doi: 10.24432/C5MC89.
T. Gori et al., “Preprocessing Data dan Klasifikasi untuk Prediksi Akademik Siswa,” J. Teknol. Inf. dan Ilmu Komput., vol. 11, no. 1, pp. 215–224, 2024, doi: 10.25126/jtiik.20241118074.
S. Sidiq, N. S. Mabrur, S. T. Informatika, F. Teknik, and U. M. Tangerang, “Pengembangan Model Prediksi Risiko Diabetes Menggunakan Pendekatan AdaBoost dan Teknik Oversampling,” J. Ilm. Inform. dan Ilmu Komput., vol. 4, pp. 13–23, 2025, doi: https://doi.org/10.58602/jima-ilkom.v4i1.41.
D. Kurniadi, F. Nuraeni, M. Firmansyah, K. Garut, and P. Korespondensi, “Klasifikasi Masyarakat Penerima Bantuan Langsung Tunai Dana Desa Menggunakan Naive Bayes dan SMOTE,” J. Teknol. Inf. dan Ilmu Komput., vol. 10, no. 2, pp. 309–320, 2023, doi: 10.25126/jtiik.2023106453.
H. Ali, N. Hendrastuty, C. Science, and U. T. Indonesia, “Comparison of Naive Bayes Classdifier, Support Vector Machine, Rando Forest Algorithms Public Sentiment Analysis od KIP-K Program on Twitter,” J. Tek. Inform., vol. 5, no. 6, pp. 1701–1712, 2024, doi: https://doi.org/10.52436/1.jutif.2024.5.6.4030.
R. Siringoringo, D. Arisandi, E. Kurniawan, and E. B. Nababan, “Model Klasifikasi dengan Logistic Regression dan Recursive Feature Elimation pada Data Tidak Seimbang,” J. Teknol. Inf. dan Ilmu Komput., vol. 11, no. 4, pp. 735–742, 2024, doi: 10.25126/jtiik.1148198.
D. Kurnia et al., “Seleksi Fitur Dengan Particle Swarm Optimization pada Klasifikasi Penyakit Parkinson Menggunakan XBboost,” J. Teknol. Inf. dan Ilmu Komput., vol. 10, no. 5, pp. 1083–1094, 2023, doi: 10.25126/jtiik.2023107252.
M. D. Simbolon, D. D. Bukit, R. E. Purba, and F. H. Ketaren, “Perbandingan Algoritma K-Nearest Neighboor dan Naive Bayes Dalam Prediksi Penyakit Ginjal Kronis Pada Lansia,” J. Inf. Syst. Res., vol. 6, no. 2, pp. 1519–1525, 2025, doi: 10.47065/josh.v6i2.5938.
E. Junianto and S. Nurkhodijah, “Explainable Ensemble Learning for Depression Risk Classification Using Multidomain Behavioral Features,” J. Tek. Inform., vol. 7, no. 2, pp. 778–792, 2026.
L. Efrizoni, S. Defit, M. Tajuddin, and A. Anggrawan, “Komparasi Ekstraksi Fitur dalam Klasifikasi Teks Multilabel Menggunakan Algoritma Machine Learning Comparison of Feature Extraction in Multilabel Text Classification Using Machine Learning Algorithm,” Matrik J. Manajemen, Tek. Inform. dan Rekayasa Komput., vol. 21, no. 3, pp. 653–66, 2022, doi: 10.30812/matrik.v21i3.1851.
W. S. Pratama, D. D. Prasetya, T. Widyaningtyas, M. Z. Wiryawan, and L. G. Rady, “Performance Evaluation of Artificial Intelligence Models for Classification in Concept Map Quality Assessment,” Matrik J. Manajemen, Tek. Inform. dan Rekayasa Komput., vol. 24, no. 3, pp. 407–422, 2025, doi: 10.30812/matrik.v24i3.4729.
A. H. M. Suparyati, Emma Utami, “Applying Different Resampling Strategies In Random Forest Algorithm To Predict Lumpy Skin Disease,” J. RESTI (Rekayasa Sist. dan Teknol. Informasi), vol. 5, no. 158, pp. 555–562, 2026, doi: https://doi.org/10.29207/resti.v6i4.4147.
I. Raga et al., “Analisis Sentimen Pasien Terhadap Layanan Anatrmedika DentalCare Menggunakan Metode XGBoost,” J. Teknol. Inf. dan Ilmu Komput., vol. 12, no. 6, pp. 1377–1384, 2025, doi: https://doi.org/10.25126/jtiik.2025126.
N. Putri et al., “Penerapan Feature Enngineering dan Hyperparameter Tuning untuk Meningkatkan Akurasi Model Random Forest Pada Klasifikasi Risiko Kredit,” J. Teknol. Inf. dan Ilmu Komput., vol. 12, no. 2, pp. 251–262, 2025, doi: 10.25126/jtiik.2025128472.
D. Kurniadi, A. I. Zulkarnaen, A. Mulyani, K. Garut, and P. Korespondensi, “Perancangan Model Convolutional Neural Network pada Aplikasi Pengenalan Askara Sunda Berbasis Mobile,” J. Teknol. Inf. dan Ilmu Komput., vol. 12, no. 5, 2025, doi: https://doi.org/10.25126/jtiik.2025125.
D. W. W. Haryono Setiadi , Indah Paksi Larasati, Esti Suryani, A. W. Hasan Dwi Cahyono, and A. Doewes, “Comparing Correlation-Based Feature Selection and Symmetrical,” J. RESTI (Rekayasa Sist. dan Teknol. Informasi), vol. 5, no. 158, pp. 542–554, 2026, doi: https://doi.org/10.29207/resti.v8i4.5911.
Bila bermanfaat silahkan share artikel ini
Berikan Komentar Anda terhadap artikel Identifikasi Faktor Dominan Kegagalan Akademik pada Data Tidak Seimbang Menggunaan Ensemble Learning dan Hybrid SMOTE-ENN
Pages: 798-807
Copyright (c) 2026 Rizky Nurhasanah, Solikhun Solikhun

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under Creative Commons Attribution 4.0 International License that allows others to share the work with an acknowledgment of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgment of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (Refer to The Effect of Open Access).






















