Komparasi Decision Tree, Random Forest, dan XGBoost untuk Deteksi Dini Risiko Akademik, Kehadiran, dan Karakter Siswa


  • M. Hafidhatul Fathoni * Mail Universitas Sains dan Teknologi Indonesia, Pekanbaru, Indonesia
  • Junadhi Junadhi Universitas Sains dan Teknologi Indonesia, Pekanbaru, Indonesia
  • Susanti Susanti Universitas Sains dan Teknologi Indonesia, Pekanbaru, Indonesia
  • Hadi Asnal Universitas Sains dan Teknologi Indonesia, Pekanbaru, Indonesia
  • (*) Corresponding Author
Keywords: Early Detection; Machine Learning; Student Risk; Temporal Evaluation; Time-Shifting Target

Abstract

Schools routinely collect academic, attendance, and character data, but their use remains limited to administrative recording. As a result, changes in student conditions are not always transformed into early risk signals, such as academic decline, repeated absenteeism, or accumulated character-related violations. The problem addressed in this study is how to evaluate an early detection model that can capture these three risk dimensions simultaneously using imbalanced and temporally ordered monthly longitudinal data. This study compares Decision Tree, Random Forest, and XGBoost for detecting Academic Risk, Attendance Risk, and Character Risk in the following month. The dataset consists of longitudinal student records from SMKS YUM Pesantren Teknologi Riau in 2025, covering 206 students across 12 months. The raw dataset contained 2,472 rows and was reduced to 2,266 valid rows after feature engineering and time-shifting target construction. Model evaluation used rolling-origin temporal evaluation, majority baseline, temporal persistence baseline, and imbalance handling strategies: no balancing, class_weight, and partial SMOTE on training data. Macro F1 was used as the primary metric because the class distribution was imbalanced. The best models were Random Forest with class_weight for Academic Risk, with Macro F1 of 0.5005 and High-Risk recall of 0.6185; Decision Tree with class_weight for Attendance Risk, with Macro F1 of 0.4586 and High-Risk recall of 0.5339; and Random Forest with class_weight for Character Risk, with Macro F1 of 0.4761 and High-Risk recall of 0.5698. These findings indicate that machine learning models can provide early High-Risk signals, but predictions should be verified by teachers or homeroom teachers.

Downloads

Download data is not yet available.

References

World Bank, UNESCO, “Indonesia - Learning Poverty Brief 2024,” World Bank, Washington, DC, 2024. [Online]. Available: https://documents.worldbank.org/curated/en/099082924151529593

OECD, PISA 2022 Results (Volume I): The State of Learning and Equity in Education. Paris: OECD Publishing, 2023. doi: 10.1787/53f23881-en.

UNESCO, Global Education Monitoring Report 2023: Technology in Education: A Tool on Whose Terms?, First. Paris: UNESCO, 2023. doi: 10.54676/UZQV8501.

S. Batool, J. Rashid, M. W. Nisar, J. Kim, H.-Y. Kwon, and A. Hussain, “Educational data mining to predict students’ academic performance: A survey study,” Educ. Inf. Technol., vol. 28, no. 1, pp. 905–971, 2023, doi: 10.1007/s10639-022-11152-y.

H. T.-H. Duong, L. T.-M. Tran, H. Q. To, and K. Van Nguyen, “Academic performance warning system based on data driven for higher education,” Neural Comput. Appl., vol. 35, no. 8, pp. 5819–5837, 2023, doi: 10.1007/s00521-022-07997-6.

M. Skittou, S. Raghay, and A. Zellou, “Development of an Early Warning System to Support Educational Planning Process by Identifying At-Risk Students,” IEEE Access, vol. 12, p. 2260, 2024, doi: 10.1109/ACCESS.2023.3348091.

A. N. H. Zaied, E. M. Hamza, and R. W. Ismael, “An Advisory Student Achievement Model Based on Data Mining Techniques,” in 2022 International Conference on Computing and Information, 2022. doi: 10.1109/ICCI54321.2022.9756121.

Y. Jang, S. Choi, H. Jung, and H. Kim, “Practical early prediction of students’ performance using machine learning and eXplainable AI,” Educ. Inf. Technol., vol. 27, no. 9, pp. 12855–12889, 2022, doi: 10.1007/s10639-022-11120-6.

A. M. Kord, A. Aboelfetouh, and S. M. Shohieb, “Academic Course Planning Recommendation and Students’ Performance Prediction Multi-Modal Based on Educational Data Mining Techniques,” J. Comput. High. Educ., 2025, doi: 10.1007/s12528-024-09426-0.

A. Gupta, D. Garg, and P. Kumar, “Mining Sequential Learning Trajectories with Hidden Markov Models for Early Prediction of At-Risk Students in E-Learning Environments,” IEEE Trans. Learn. Technol., vol. 15, no. 6, pp. 783–799, 2022, doi: 10.1109/TLT.2022.3202277.

M. Kumar, N. Singh, J. Wadhwa, P. Singh, G. Kumar, and A. Qtaishat, “Utilizing Random Forest and XGBoost DataMining Algorithms for Anticipating Students’ Academic Performance,” Int. J. Mod. Educ. Comput. Sci., vol. 16, no. 2, 2024, doi: 10.5815/ijmecs.2024.02.03.

C. Aurelio and J. Setiawan, “In-Depth Examination of Student Attrition in Higher Education Through Decision Tree, Random Forest, and XGBoost Algorithms,” 2024. doi: 10.1109/EECSI63442.2024.10776324.

J. Niyogisubizo, L. Liao, E. Nziyumva, E. Murwanashyaka, and P. C. Nshimyumukiza, “Predicting Student’s Dropout in University Classes Using Two-Layer Ensemble Machine learning Approach: A Novel Stacked Generalization,” Comput. Educ. Artif. Intell., vol. 3, p. 100066, 2022, doi: 10.1016/j.caeai.2022.100066.

B. Carballo Mendívil, A. Arellano González, N. J. Ríos-Vázquez, and M. del P. Lizardi-Duarte, “Predicting Student Dropout from Day One: XGBoost-Based Early Warning System Using Pre-Enrollment Data,” Appl. Sci., vol. 15, no. 16, p. 9202, 2025, doi: 10.3390/app15169202.

S. Almaqbali, M.-C. Leow, B. Shannaq, O. Farouk, A. H. Marhoubi, and L.-Y. Ong, “Predicting Learner Disengagement in E-Learning Platforms Using Interpretable Machine learning: A Comparative Study of Tree-Based Classifiers,” 2025. doi: 10.1109/ISCI65687.2025.11167693.

Andri, R. Yunis, and Tanti, “Optimizing Random Forest Classification Using Chi-Square and SMOTE-ENN on Student Drop-Out Data,” 2023. doi: 10.1109/ICIC60109.2023.10382055.

M. Alhammouri, Z. A. A. Hammouri, I. T. Almalkawi, and A. Lafee, “Optimizing Multi-Class Classification in Educational Data with Ensemble Learning and Data Balancing Techniques,” 2024. doi: 10.1109/IDSTA62194.2024.10746987.

V. Cerqueira, L. Torgo, and I. Mozetič, “Evaluating Time Series Forecasting Models: An Empirical Study on Performance Estimation Methods,” Mach. Learn., vol. 109, pp. 1997–2028, 2020, doi: 10.1007/s10994-020-05910-7.

T. A. Marzuqi, E. Kristiani, and M. Marcel, “Prediksi Mahasiswa Drop-Out di Universitas XYZ,” J. Teknol. Inf. dan Ilmu Komput., 2024, doi: 10.25126/jtiik.2024118689.

P.-N. Tan, M. Steinbach, A. Karpatne, and V. Kumar, Introduction to Data Mining, 2nd ed. Boston, MA: Pearson, 2019.

A. Géron, Hands-On Machine learning with Scikit-Learn, Keras, and TensorFlow, 3rd ed. Sebastopol, CA: O’Reilly Media, 2022.

Y. Zoralioğlu, M. Gul, F. Azizoğlu, G. Azizoğlu, and A. N. Toprak, “Predicting Academic Performance of Students Using Machine learning Techniques,” in 2023 Innovations in Intelligent Systems and Applications Conference, 2023. doi: 10.1109/ASYU58738.2023.10296648.


Bila bermanfaat silahkan share artikel ini

Berikan Komentar Anda terhadap artikel Komparasi Decision Tree, Random Forest, dan XGBoost untuk Deteksi Dini Risiko Akademik, Kehadiran, dan Karakter Siswa

Dimensions Badge
Article History
Submitted: 2026-02-25
Published: 2026-06-30
Abstract View: 0 times
PDF Download: 0 times
How to Cite
Fathoni, M. H., Junadhi, J., Susanti, S., & Asnal, H. (2026). Komparasi Decision Tree, Random Forest, dan XGBoost untuk Deteksi Dini Risiko Akademik, Kehadiran, dan Karakter Siswa. Building of Informatics, Technology and Science (BITS), 8(1), 497-509. https://doi.org/10.47065/bits.v8i1.9452
Issue
Section
Articles