Komparasi Decision Tree, Random Forest, dan XGBoost untuk Deteksi Dini Risiko Akademik, Kehadiran, dan Karakter Siswa
Abstract
Schools routinely collect academic, attendance, and character data, but their use remains limited to administrative recording. As a result, changes in student conditions are not always transformed into early risk signals, such as academic decline, repeated absenteeism, or accumulated character-related violations. The problem addressed in this study is how to evaluate an early detection model that can capture these three risk dimensions simultaneously using imbalanced and temporally ordered monthly longitudinal data. This study compares Decision Tree, Random Forest, and XGBoost for detecting Academic Risk, Attendance Risk, and Character Risk in the following month. The dataset consists of longitudinal student records from SMKS YUM Pesantren Teknologi Riau in 2025, covering 206 students across 12 months. The raw dataset contained 2,472 rows and was reduced to 2,266 valid rows after feature engineering and time-shifting target construction. Model evaluation used rolling-origin temporal evaluation, majority baseline, temporal persistence baseline, and imbalance handling strategies: no balancing, class_weight, and partial SMOTE on training data. Macro F1 was used as the primary metric because the class distribution was imbalanced. The best models were Random Forest with class_weight for Academic Risk, with Macro F1 of 0.5005 and High-Risk recall of 0.6185; Decision Tree with class_weight for Attendance Risk, with Macro F1 of 0.4586 and High-Risk recall of 0.5339; and Random Forest with class_weight for Character Risk, with Macro F1 of 0.4761 and High-Risk recall of 0.5698. These findings indicate that machine learning models can provide early High-Risk signals, but predictions should be verified by teachers or homeroom teachers.
Downloads
References
World Bank, UNESCO, “Indonesia - Learning Poverty Brief 2024,” World Bank, Washington, DC, 2024. [Online]. Available: https://documents.worldbank.org/curated/en/099082924151529593
OECD, PISA 2022 Results (Volume I): The State of Learning and Equity in Education. Paris: OECD Publishing, 2023. doi: 10.1787/53f23881-en.
UNESCO, Global Education Monitoring Report 2023: Technology in Education: A Tool on Whose Terms?, First. Paris: UNESCO, 2023. doi: 10.54676/UZQV8501.
S. Batool, J. Rashid, M. W. Nisar, J. Kim, H.-Y. Kwon, and A. Hussain, “Educational data mining to predict students’ academic performance: A survey study,” Educ. Inf. Technol., vol. 28, no. 1, pp. 905–971, 2023, doi: 10.1007/s10639-022-11152-y.
H. T.-H. Duong, L. T.-M. Tran, H. Q. To, and K. Van Nguyen, “Academic performance warning system based on data driven for higher education,” Neural Comput. Appl., vol. 35, no. 8, pp. 5819–5837, 2023, doi: 10.1007/s00521-022-07997-6.
M. Skittou, S. Raghay, and A. Zellou, “Development of an Early Warning System to Support Educational Planning Process by Identifying At-Risk Students,” IEEE Access, vol. 12, p. 2260, 2024, doi: 10.1109/ACCESS.2023.3348091.
A. N. H. Zaied, E. M. Hamza, and R. W. Ismael, “An Advisory Student Achievement Model Based on Data Mining Techniques,” in 2022 International Conference on Computing and Information, 2022. doi: 10.1109/ICCI54321.2022.9756121.
Y. Jang, S. Choi, H. Jung, and H. Kim, “Practical early prediction of students’ performance using machine learning and eXplainable AI,” Educ. Inf. Technol., vol. 27, no. 9, pp. 12855–12889, 2022, doi: 10.1007/s10639-022-11120-6.
A. M. Kord, A. Aboelfetouh, and S. M. Shohieb, “Academic Course Planning Recommendation and Students’ Performance Prediction Multi-Modal Based on Educational Data Mining Techniques,” J. Comput. High. Educ., 2025, doi: 10.1007/s12528-024-09426-0.
A. Gupta, D. Garg, and P. Kumar, “Mining Sequential Learning Trajectories with Hidden Markov Models for Early Prediction of At-Risk Students in E-Learning Environments,” IEEE Trans. Learn. Technol., vol. 15, no. 6, pp. 783–799, 2022, doi: 10.1109/TLT.2022.3202277.
M. Kumar, N. Singh, J. Wadhwa, P. Singh, G. Kumar, and A. Qtaishat, “Utilizing Random Forest and XGBoost DataMining Algorithms for Anticipating Students’ Academic Performance,” Int. J. Mod. Educ. Comput. Sci., vol. 16, no. 2, 2024, doi: 10.5815/ijmecs.2024.02.03.
C. Aurelio and J. Setiawan, “In-Depth Examination of Student Attrition in Higher Education Through Decision Tree, Random Forest, and XGBoost Algorithms,” 2024. doi: 10.1109/EECSI63442.2024.10776324.
J. Niyogisubizo, L. Liao, E. Nziyumva, E. Murwanashyaka, and P. C. Nshimyumukiza, “Predicting Student’s Dropout in University Classes Using Two-Layer Ensemble Machine learning Approach: A Novel Stacked Generalization,” Comput. Educ. Artif. Intell., vol. 3, p. 100066, 2022, doi: 10.1016/j.caeai.2022.100066.
B. Carballo Mendívil, A. Arellano González, N. J. Ríos-Vázquez, and M. del P. Lizardi-Duarte, “Predicting Student Dropout from Day One: XGBoost-Based Early Warning System Using Pre-Enrollment Data,” Appl. Sci., vol. 15, no. 16, p. 9202, 2025, doi: 10.3390/app15169202.
S. Almaqbali, M.-C. Leow, B. Shannaq, O. Farouk, A. H. Marhoubi, and L.-Y. Ong, “Predicting Learner Disengagement in E-Learning Platforms Using Interpretable Machine learning: A Comparative Study of Tree-Based Classifiers,” 2025. doi: 10.1109/ISCI65687.2025.11167693.
Andri, R. Yunis, and Tanti, “Optimizing Random Forest Classification Using Chi-Square and SMOTE-ENN on Student Drop-Out Data,” 2023. doi: 10.1109/ICIC60109.2023.10382055.
M. Alhammouri, Z. A. A. Hammouri, I. T. Almalkawi, and A. Lafee, “Optimizing Multi-Class Classification in Educational Data with Ensemble Learning and Data Balancing Techniques,” 2024. doi: 10.1109/IDSTA62194.2024.10746987.
V. Cerqueira, L. Torgo, and I. Mozetič, “Evaluating Time Series Forecasting Models: An Empirical Study on Performance Estimation Methods,” Mach. Learn., vol. 109, pp. 1997–2028, 2020, doi: 10.1007/s10994-020-05910-7.
T. A. Marzuqi, E. Kristiani, and M. Marcel, “Prediksi Mahasiswa Drop-Out di Universitas XYZ,” J. Teknol. Inf. dan Ilmu Komput., 2024, doi: 10.25126/jtiik.2024118689.
P.-N. Tan, M. Steinbach, A. Karpatne, and V. Kumar, Introduction to Data Mining, 2nd ed. Boston, MA: Pearson, 2019.
A. Géron, Hands-On Machine learning with Scikit-Learn, Keras, and TensorFlow, 3rd ed. Sebastopol, CA: O’Reilly Media, 2022.
Y. Zoralioğlu, M. Gul, F. Azizoğlu, G. Azizoğlu, and A. N. Toprak, “Predicting Academic Performance of Students Using Machine learning Techniques,” in 2023 Innovations in Intelligent Systems and Applications Conference, 2023. doi: 10.1109/ASYU58738.2023.10296648.
Bila bermanfaat silahkan share artikel ini
Berikan Komentar Anda terhadap artikel Komparasi Decision Tree, Random Forest, dan XGBoost untuk Deteksi Dini Risiko Akademik, Kehadiran, dan Karakter Siswa
Pages: 497-509
Copyright (c) 2026 M. Hafidhatul Fathoni, Junadhi Junadhi, Susanti Susanti, Hadi Asnal

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under Creative Commons Attribution 4.0 International License that allows others to share the work with an acknowledgment of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgment of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (Refer to The Effect of Open Access).





















