Evaluasi Komparatif Algoritma Decision Tree, Random Forest, dan XGBoost untuk Software Defect Prediction Menggunakan Dataset NASA Software Metrics
Abstract
Software Defect Prediction (SDP) has become an important approach for identifying software modules that are likely to contain defects during the early stages of software development. However, the performance of prediction algorithms remains highly dependent on dataset characteristics, and no single algorithm has consistently demonstrated superior performance across different datasets. In addition, the use of synthetic datasets in SDP research still requires empirical validation to ensure that their characteristics remain representative of benchmark datasets such as the NASA Metric Data Program (NASA MDP). Therefore, this study aims to compare the performance of Decision Tree, Random Forest, and XGBoost using the Playground Series Season 3 Episode 23 dataset, a synthetic dataset developed based on the characteristics of the NASA MDP dataset. Prior to model training, the dataset underwent preprocessing, including missing value imputation, label encoding, and feature standardization. Model performance was evaluated using 10-fold stratified cross-validation with Accuracy, F1-Score, and AUC-ROC as the primary evaluation metrics. The experimental results indicate that ensemble learning methods achieved competitive performance compared with the single-classifier approach. Random Forest achieved the highest Accuracy of 0.8144, while XGBoost obtained the highest AUC-ROC score of 0.7929, indicating strong capability in distinguishing between defective and non-defective software modules on the evaluated dataset. Furthermore, feature importance analysis identified Lines of Code (LOC), Cyclomatic Complexity, Halstead Volume, IOCode, and branchCount as the most influential factors affecting software defect prediction. Based on the experimental results obtained from the selected dataset, the ensemble learning approach demonstrated competitive predictive performance and may be considered a promising alternative for developing Software Defect Prediction models. Nevertheless, the selection of the most appropriate algorithm should remain dependent on dataset characteristics and specific implementation requirements
Downloads
References
S. R. Goyal, “Effective software defect prediction with deep neural networks,” Results in Engineering, vol. 29, no. November 2025, p. 108378, 2026, doi: 10.1016/j.rineng.2025.108378.
Y. Ding et al., “Metric information mining with metric attention to boost software defect prediction performance,” Science of Computer Programming, vol. 248, no. August 2025, 2026, doi: 10.1016/j.scico.2025.103381.
L. Shen and H. Bai, “Engineering software-based evaluation of pelvic floor muscle defects in anterior vaginal wall prolapse,” Journal of Radiation Research and Applied Sciences, vol. 19, no. 1, p. 102119, 2026, doi: 10.1016/j.jrras.2025.102119.
X. Wei, “Research on Preprocessing Techniques for Software Defect Prediction Dataset Based on Hybrid Category Balance and Synthetic Sampling Algorithm,” Procedia Computer Science, vol. 262, pp. 840–848, 2025, doi: 10.1016/j.procs.2025.05.117.
L. Lavazza, S. Morasca, and G. Rotoloni, “Software Defect Prediction evaluation: New metrics based on the ROC curve,” Information and Software Technology, vol. 187, no. August, p. 107865, 2025, doi: 10.1016/j.infsof.2025.107865.
T. Bayramova, “Software Defect Prediction Using the Machine Learning Methods,” Problems of Information Technology, vol. 14, no. 2, pp. 23–31, 2023, doi: 10.25045/jpit.v14.i2.03.
Akhlas Tariq Hasan and Shayma Mustafa Mohi-Aldeen, “Software Defect Prediction Based On Deep Learning Algorithms : A Systematic Literature Review,” AL-Rafidain Journal of Computer Sciences and Mathematics, vol. 19, no. 1, pp. 67–79, 2025, doi: 10.33899/csmj.2025.156086.1160.
E. A. Kusnanti et al., “Prediksi cacat perangkat lunak menggunakan recurrent neural network berbasis pca,” Jurnal Ilmiah Teknologi Informasi, pp. 23–31, 2024.
Emma Andini, M. R. Faisal, Rudy Herteno, R. A. Nugroho, Friska Abadi, and Muliadi, “Peningkatan Kinerja Prediksi Cacat Software Dengan Hyperparameter Tuning Pada Algoritma Klasifikasi Deep Forest,” Jurnal Mnemonic, vol. 5, no. 2, pp. 119–127, 2022, doi: 10.36040/mnemonic.v5i2.4793.
Emma Andini, M. R. Faisal, Rudy Herteno, R. A. Nugroho, Friska Abadi, and Muliadi, “Peningkatan Kinerja Prediksi Cacat Software Dengan Hyperparameter Tuning Pada Algoritma Klasifikasi Deep Forest,” Jurnal Mnemonic, vol. 5, no. 2, pp. 119–127, 2022, doi: 10.36040/mnemonic.v5i2.4793.
Y. B. L. Kintomonho, M. N. Atchadé, and D. Daddah, “Decision tree-based statistical learning and quantile regression adjustment: Insights from pregnant women in Benin,” Scientific African, vol. 29, pp. 1–10, Sep. 2025, doi: 10.1016/j.sciaf.2025.e02832.
P. Rosyani, A. M. Lutfi, E. Purwadi, Kamaluddin, Y. A. Hanaan, and I. H. Ikasari, “Application of Random Forest for Rice Plant Disease Classification,” International Journal of Integrative Sciences, vol. 4, no. 1, pp. 141–150, Feb. 2025, doi: 10.55927/ijis.v4i1.13477.
M. Abdullahi et al., “Detecting Cybersecurity Attacks in Internet of Things Using Artificial Intelligence Methods: A Systematic Literature Review,” 2022. doi: 10.3390/electronics11020198.
N. * Risky, D. Setiyawan, D. Hermawan, and O. Herdiyanto, “Prediksi Kelulusan Mahasiswa Menggunakan Algoritma Decision Tree C4.5 Berbasis Data Akademik dengan Validasi 10-Fold,” TIN: Terapan Informatika Nusantara, vol. 6, no. 6, pp. 670–678, Nov. 2025, doi: 10.47065/tin.v6i6.8662.
N. Istiqomah and M. Murinto, “Klasifikasi Penyakit Tanaman Padi Berbasis Citra Daun Menggunakan Convolutional Neural Network (CNN),” JSTIE (Jurnal Sarjana Teknik Informatika) (E-Journal), vol. 12, no. 1, p. 18, Feb. 2024, doi: 10.12928/jstie.v12i1.27314.
N. Paul, G. C. Sunil, D. Horvath, and X. Sun, “Deep learning for plant stress detection: A comprehensive review of technologies, challenges, and future directions,” Feb. 01, 2025, Elsevier B.V. doi: 10.1016/j.compag.2024.109734.
R. Rashkovits and I. Lavy, “Mapping Common Errors in Entity Relationship Diagram Design of Novice Designers,” International Journal of Database Management Systems, vol. 13, no. 1, pp. 1–19, 2021, doi: 10.5121/ijdms.2021.13101.
D. Pradhan and D. Muduli, “Improved prediction of software defect: A particle swarm optimization based ELM approach,” e-Prime - Advances in Electrical Engineering, Electronics and Energy, vol. 13, no. July, p. 101081, 2025, doi: 10.1016/j.prime.2025.101081.
N. Camelia-Petrina and V. Andreea, “Software Defect Prediction Models. A Replication and Extension Study,” Procedia Computer Science, vol. 270, pp. 1936–1945, 2025, doi: 10.1016/j.procs.2025.09.314.
W. Ustyannie, E. Setyaningsih, and C. Iswahyudi, “Optimization of software defects prediction in imbalanced class using a combination of resampling methods with support vector machine and logistic regression,” Jurnal Infotel, vol. 13, no. 4, pp. 176–184, 2021, doi: 10.20895/infotel.v13i4.726.
Bila bermanfaat silahkan share artikel ini
Berikan Komentar Anda terhadap artikel Evaluasi Komparatif Algoritma Decision Tree, Random Forest, dan XGBoost untuk Software Defect Prediction Menggunakan Dataset NASA Software Metrics
Pages: 615-628
Copyright (c) 2026 Okky Prasetia, Syaeful Machfud, Syaeful Machfud

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under Creative Commons Attribution 4.0 International License that allows others to share the work with an acknowledgment of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgment of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (Refer to The Effect of Open Access).





















