Comparison of XGBoost and Random Forest for Prediction of Male Fertility Status Based on Semen Analysis Parameters
Abstract
Male infertility is a major reproductive health problem that contributes to approximately half of infertility cases among couples of reproductive age. Accurate evaluation of male fertility status commonly relies on semen analysis, including parameters such as semen volume, sperm concentration, motility, morphology, and vitality. However, manual interpretation of these parameters remains time-consuming and is highly dependent on clinical expertise. This study aims to compare the performance of the XGBoost and Random Forest algorithms in predicting male fertility status based on semen analysis parameters. The study employed a secondary dataset consisting of 1,000 semen analysis records with 11 predictor variables and one target variable representing fertility status. Data preprocessing included categorical encoding, data cleaning, and an 80:20 train–test split before model development and evaluation using accuracy, precision, recall, F1-score, and ROC-AUC. Experimental results showed that XGBoost outperformed Random Forest, achieving an accuracy of 99.00%, precision of 100.00%, recall of 96.88%, F1-score of 98.41%, and ROC-AUC of 99.86%, while Random Forest achieved an accuracy of 97.50%. Feature importance analysis identified Total Motility, Vitality, and Progressive Motility as the most influential predictors of male fertility status. The main contribution of this study is the direct empirical comparison of Random Forest and XGBoost under identical experimental settings using comprehensive semen analysis parameters, providing evidence on the relative effectiveness of ensemble learning algorithms for male fertility status prediction. Although the proposed model demonstrates excellent predictive performance, it was developed using secondary data and is intended to support, rather than replace, clinical decision-making. Future studies should validate the model using larger multicenter clinical datasets to improve its generalizability and practical applicability.
Downloads
References
World Health Organization, “WHO Laboratory Manual for the Examination and Processing of Human Semen,” 6th ed., H. R. P. World Health Organization, Ed., Geneva: World Health Organization, 2021. [Online]. Available: https://www.who.int/publications/i/item/9789240030787
S. Dehghan, R. Rabiei, H. Choobineh, K. Maghooli, M. Nazari, and M. Vahidi-Asl, “Comparative study of machine learning approaches integrated with genetic algorithm for IVF success prediction,” PLoS One, vol. 19, no. 10 October, pp. 1–19, 2024, doi: 10.1371/journal.pone.0310829.
A. Aykaç, C. Kaya, Ö. Çelik, M. E. Aydın, and M. Sungur, “The prediction of semen quality based on lifestyle behaviours by the machine learning based models,” Reprod. Biol. Endocrinol., vol. 22, p. 112, 2024, doi: 10.1186/s12958-024-01268-w.
S. S. Cao, X. M. Liu, B. T. Song, and Y. Y. Hu, “Interpretable machine learning models for predicting clinical pregnancies associated with surgical sperm retrieval from testes of different etiologies: a retrospective study,” BMC Urol., vol. 24, no. 1, pp. 1–12, 2024, doi: 10.1186/s12894-024-01537-1.
Gyorgy J. Simon & Constantin Aliferis, “Artificial Intelligence and Machine Learning in Health Care and Medical Sciences,” Springer, 2024, p. 810. [Online]. Available: https://link.springer.com/book/10.1007/978-3-031-39355-6
A. Alanazi, “Using machine learning for healthcare challenges and opportunities,” Informatics Med. Unlocked, vol. 30, p. 100924, 2022, doi: 10.1016/j.imu.2022.100924.
L. Lu, Y. Qian, Y. Dong, H. Su, Y. Deng, and L. Lu, “A systematic study of the performance of machine learning models on analyzing the association between semen quality and environmental pollutants,” Front. Phys., vol. 11, p. 1259273, 2023, doi: 10.3389/fphy.2023.1259273.
I. D. Mienye, Y. Sun, and S. Member, “A Survey of Ensemble Learning: Concepts, Algorithms, Applications, and Prospects,” IEEE Access, vol. 10, pp. 99129–99149, 2022, doi: 10.1109/ACCESS.2022.3207287.
H. A. Salman, A. Kalakech, and A. Steiti, “Random Forest Algorithm Overview,” Babylonian J. Mach. Learn., vol. 2024, pp. 69–79, 2024, doi: 10.58496/BJML/2024/007.
M. Niazkar et al., “Applications of XGBoost in water resources engineering : A systematic literature review ( Dec 2018 – May 2023 ),” Environ. Model. Softw., vol. 174, p. 105971, 2024, doi: 10.1016/j.envsoft.2024.105971.
A. Mehrjerd, T. Dehghani, M. Jajroudi, S. Eslami, H. Rezaei, and N. K. Ghaebi, “Ensemble machine learning models for sperm quality evaluation concerning success rate of clinical pregnancy in assisted reproductive techniques,” Sci. Rep., vol. 14, no. 1, pp. 1–10, 2024, doi: 10.1038/s41598-024-73326-7.
H. Huang, S. Hsieh, M. Chen, M. Jhou, and T. Liu, “Machine Learning Predictive Models for Evaluating Risk Factors Affecting Sperm Count : Predictions Based on Health Screening Indicators,” J. Clin. Med., vol. 12, no. 3, p. 1220, 2023, doi: 10.3390/jcm12031220.
D. Ghoshroy, P. A. Alvi, and K. C. Santosh, “Explainable AI to Predict Male Fertility Using Extreme Gradient Boosting Algorithm with SMOTE,” Electronics, vol. 12, no. 1, p. 15, 2023, doi: 10.3390/electronics12010015.
H. Mohammadi, S. Khoddam, and F. Golbabaei, “Analyzing the impact of occupational exposures on male fertility indicators : A machine learning approach,” Reprod. Toxicol., vol. 136, p. 108959, 2025, doi: 10.1016/j.reprotox.2025.108959.
Z. Chen, D. Zhang, J. Zhen, Z. Sun, Q. Yu, and Y. Yin, “Predicting cumulative live birth rate for patients undergoing in vitro fertilization (IVF)/intracytoplasmic sperm injection (ICSI) for tubal and male infertility: a machine learning approach using XGBoost,” Chin. Med. J. (Engl)., vol. 135, no. 8, pp. 997–999, 2022, doi: 10.1097/CM9.0000000000001874.
J. Kim, “The research gap in evaluating community-based mental health interventions in Korea : A comparative analysis with the United Kingdom,” Asian J. Psychiatr., vol. 103, p. 104348, 2025, doi: 10.1016/j.ajp.2024.104348.
I. Chouvarda et al., “Differences in technical and clinical perspectives on AI validation in cancer imaging : mind the gap !,” Eur. Radiol. Exp., vol. 9, no. 7, 2025, doi: 10.1186/s41747-024-00543-0.
A. Tawakuli, B. Havers, V. Gulisano, D. Kaiser, and T. Engel, “Survey : Time-series data preprocessing : A survey and an empirical analysis,” J. Eng. Res., vol. 13, no. 2, pp. 674–711, 2025, doi: 10.1016/j.jer.2024.02.018.
L. A. Yates, Z. Aandahl, S. A. Richards, and B. W. Brook, “Cross validation for model selection : A review with examples from ecology,” Ecol. Monogr., vol. 93, no. 1, pp. 1–24, 2023, doi: 10.1002/ecm.1557.
W. Hong, X. Zhou, S. Jin, Y. Lu, J. Pan, and Q. Lin, “A Comparison of XGBoost , Random Forest , and Nomograph for the Prediction of Disease Severity in Patients With COVID-19 Pneumonia : Implications of Cytokine and Immune Cell Pro fi le,” Front. Cell. Infect. Microbiol., vol. 12, no. April, pp. 1–13, 2022, doi: 10.3389/fcimb.2022.819267.
Bila bermanfaat silahkan share artikel ini
Berikan Komentar Anda terhadap artikel Comparison of XGBoost and Random Forest for Prediction of Male Fertility Status Based on Semen Analysis Parameters
Pages: 538-549
Copyright (c) 2026 Kecitaan Harefa, Joko Priambodo

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under Creative Commons Attribution 4.0 International License that allows others to share the work with an acknowledgment of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgment of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (Refer to The Effect of Open Access).





















