Optimasi Hyperparameter Gaussian Naive Bayes Untuk Prediksi Risiko Stroke Pada Data Tidak Seimbang
Abstract
Stroke is a serious disease with global impact that requires high-accuracy early detection. Significant difficulties in designing machine learning-based predictive models arise due to disproportionate data conditions (imbalanced datasets). This occurs because the number of stroke cases (minority class) is very small compared to non-stroke cases. This imbalanced data situation often causes models to become biased and potentially produce high false negative rates, which is very risky in a clinical setting. This study focuses on improving the sensitivity of the Gaussian Naive Bayes (GNB) model through hyperparameter optimization and classification threshold adjustment. The research process included data preprocessing, stratified dataset division (70% training and 30% testing), feature scaling, var_smoothing parameter optimization using GridSearchCV, and threshold adjustment to maximize the Recall value. The results showed that the standard GNB model only achieved a Recall value of 0.4400. However, after var_smoothing optimization (1.00×10⁻¹⁰) and threshold adjustment to 0.0100, the Recall value increased significantly to 0.8000. This increase was accompanied by a decrease in Accuracy (0.5988) and Precision (0.0909). This improvement was accompanied by a decrease in Accuracy (0.5988) and Precision (0.0909). The high Recall (0.8000) indicates that the model is better for mass screening (early detection phase), although it must be balanced with further diagnostic processes due to low precision. This high Recall value confirms the model's success in minimizing False Negatives, which is a top priority in stroke risk prediction cases.
Downloads
References
Y. Y. Saifullah, Mochammad Erwin Rachman, Ramlian, Lilian Triana Limoa, and Nurussyariah Hamado, “Literature Review: Hubungan Hipertensi dengan Kejadian Stroke Iskemik dan Stroke Hemoragik,” Fakumi Med. J. J. Mhs. Kedokt., vol. 4, no. 10, pp. 695–708, Oct. 2024, doi: 10.33096/fmj.v4i10.477.
A. T. P. Ilham Darmawan, Indhit Tri Utami, “Penerapan Range Of Motion (ROM) Exercise Bola Karet Terhadap Kekuatan Otot Pasien Stroke Non Hemoragik,” J. Cendikia Muda, vol. 4, pp. 246–254, 2024.
E. A. Handoko and Z. Muslim, “Profil Pasien Stroke Di Rs Bhayangkara Bondowoso Tahun 2024,” Syntax Idea, vol. 7, no. 3, pp. 475–486, Mar. 2025, doi: 10.46799/syntaxidea.v7i3.12722.
A. Adi Bhirawa and U. Pradema Sanjaya, “From Data Imbalance to Precision: SMOTE-Driven Machine Learning for Early Detection of Kidney Disease,” INOVTEK Polbeng - Seri Inform., vol. 10, no. 1, pp. 514–525, Mar. 2025, doi: 10.35314/7jgjmg64.
D. M. Makarov and A. M. Kolker, “Viscosity of deep eutectic solvents: Predictive modeling with experimental validation,” Fluid Phase Equilib., vol. 587, p. 114217, Jan. 2025, doi: 10.1016/j.fluid.2024.114217.
F. Y. A’la, “Optimasi Klasifikasi Sentimen Ulasan Game Berbahasa Indonesia: IndoBERT dan SMOTE untuk Menangani Ketidakseimbangan Kelas,” Edumatic J. Pendidik. Inform., vol. 9, no. 1, pp. 256–265, Apr. 2025, doi: 10.29408/edumatic.v9i1.29666.
S. Ernawati and I. Maulana, “Meningkatkan Klasifikasi Penyakit Diabetes Menggunakan Metode Ensemble Softvoting Dengan SMOTE-ENN dan Optimasi Bayesian,” Evolusi J. Sains dan Manaj., vol. 13, no. 1, pp. 71–86, Mar. 2025, doi: 10.31294/evolusi.v13i1.8267.
M. Sulistiyono, Y. Pristyanto, S. Adi, and G. Gumelar, “Implementasi Algoritma Synthetic Minority Over-Sampling Technique untuk Menangani Ketidakseimbangan Kelas pada Dataset Klasifikasi,” Sistemasi, vol. 10, no. 2, p. 445, 2021, doi: 10.32520/stmsi.v10i2.1303.
Joshua Agung Nurcahyo and Theopilus Bayu Sasongko, “Hyperparameter Tuning Algoritma Supervised Learning untuk Klasifikasi Keluarga Penerima Bantuan Pangan Beras,” Indones. J. Comput. Sci., vol. 12, no. 3, pp. 1351–1365, 2023, doi: 10.33022/ijcs.v12i3.3254.
M. Hasan et al., “Enhancing stroke disease classification through machine learning models via a novel voting system by feature selection techniques,” PLoS One, vol. 20, no. 1, p. e0312914, Jan. 2025, doi: 10.1371/journal.pone.0312914.
A. J. Appukutty, L. E. Skolarus, M. V. Springer, W. J. Meurer, and J. F. Burke, “Increasing false positive diagnoses may lead to overestimation of stroke incidence, particularly in the young: a cross-sectional study,” BMC Neurol., vol. 21, no. 1, pp. 1–10, 2021, doi: 10.1186/s12883-021-02172-1.
K. Apostolidis et al., “Innovative Visualization Approach for Biomechanical Time Series in Stroke Diagnosis Using Explainable Machine Learning Methods: A Proof-of-Concept Study,” Information, vol. 14, no. 10, p. 559, Oct. 2023, doi: 10.3390/info14100559.
S. S. Rambe, A. Asriyanik, and P. Prajoko, “PENERAPAN MODEL CONVOLUTIONAL NEURAL NETWORK (CNN) BERBASIS MOBILENETV2 UNTUK KLASIFIKASI TINGKAT KESEGARAN IKAN NILA,” J. Inform. dan Tek. Elektro Terap., vol. 13, no. 3, Jul. 2025, doi: 10.23960/jitet.v13i3.6744.
T. Saito and M. Rehmsmeier, “The Precision-Recall Plot Is More Informative than the ROC Plot When Evaluating Binary Classifiers on Imbalanced Datasets,” PLoS One, vol. 10, no. 3, p. e0118432, Mar. 2015, doi: 10.1371/journal.pone.0118432.
Md. Khalilur Rahman, Md. Ashikur Rahman Khan, Ishtiaq Ahammad, and Joysri Rani Das, “Comparative Evaluation of Machine Learning Models for Stroke Prediction in Clinical Settings,” Cloud Comput. Data Sci., pp. 196–216, May 2025, doi: 10.37256/ccds.6220256976.
P. L. Reddy and M. Amanullah, “Prediction of wireless sensor network attack by using the gradient boosting classifier algorithm compared with Gaussian Naive Bayes with improved accuracy,” 2025, p. 020146. doi: 10.1063/5.0258920.
K. Phung, E. Ogunshile, and M. E. Aydin, Domain-specific implications of error-type metrics in risk-based software fault prediction, vol. 33, no. 1. 2025. doi: 10.1007/s11219-024-09704-1.
S. K. Singhi and H. Liu, “Feature subset selection bias for classification learning,” in Proceedings of the 23rd international conference on Machine learning - ICML ’06, New York, New York, USA: ACM Press, 2006, pp. 849–856. doi: 10.1145/1143844.1143951.
J. J. Chen, C.-A. Tsai, H. Moon, H. Ahn, J. J. Young, and C.-H. Chen, “Decision threshold adjustment in class prediction,” SAR QSAR Environ. Res., vol. 17, no. 3, pp. 337–352, Jun. 2006, doi: 10.1080/10659360600787700.
R. S. Rohman, R. A. Saputra, and D. A. Firmansaha, “Komparasi Algoritma C4.5 Berbasis PSO Dan GA Untuk Diagnosa Penyakit Stroke,” CESS (Journal Comput. Eng. Syst. Sci., vol. 5, no. 1, p. 155, 2020, doi: 10.24114/cess.v5i1.15225.
W. García, “Detecting CSV file dialects by table uniformity measurement and data type inference,” Data Sci., vol. 7, no. 2, pp. 55–72, Nov. 2024, doi: 10.3233/DS-240062.
T. Li et al., “Comprehensive bioinformatics analysis identifies LAPTM5 as a potential blood biomarker for hypertensive patients with left ventricular hypertrophy,” Aging (Albany. NY)., vol. 14, no. 3, pp. 1508–1528, Feb. 2022, doi: 10.18632/aging.203894.
S. Sidiq, Alfian, and N. S. Mabrur, “Pengembangan Model Prediksi Risiko Diabetes Menggunakan Pendekatan AdaBoost dan Teknik Oversampling SMOTE,” J. Ilm. Inform. dan Ilmu Komput., vol. 4, no. 1, pp. 13–23, 2025.
D. McMahon, C. Micallef, and T. J. Quinn, “Review of clinical practice guidelines relating to cognitive assessment in stroke,” Disabil. Rehabil., vol. 44, no. 24, pp. 7632–7640, Nov. 2022, doi: 10.1080/09638288.2021.1980122.
T. Taslim, S. Handayani, and F. Fajrizal, “Kinerja Komparatif Optimasi Algoritma Naive Bayes dalam Klasifikasi Teks untuk Uji Klinis Kanker,” J. Eksplora Inform., vol. 13, no. 1, pp. 113–123, Sep. 2023, doi: 10.30864/eksplora.v13i1.994.
Bila bermanfaat silahkan share artikel ini
Berikan Komentar Anda terhadap artikel Optimasi Hyperparameter Gaussian Naive Bayes Untuk Prediksi Risiko Stroke Pada Data Tidak Seimbang
Pages: 1490-1499
Copyright (c) 2025 Khoirun Nida, Ridwan Mahenra, Erliyan Redi Susanto

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under Creative Commons Attribution 4.0 International License that allows others to share the work with an acknowledgment of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgment of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (Refer to The Effect of Open Access).





















