Evaluasi Seleksi Fitur Algoritma Genetika pada Klasifikasi Penyakit Stroke Menggunakan K-Nearest Neighbor pada Dataset Overlapping
Abstract
Stroke is a disease that disrupts brain function due to obstructed blood flow and requires early detection to reduce the risk of disability and death. One of the classification methods widely used in stroke prediction is K-Nearest Neighbor (KNN), but this algorithm is sensitive to irrelevant features and complex data distribution. This study aims to evaluate the effectiveness of feature selection using Genetic Algorithm (GA) on the performance of KNN classification on stroke prediction datasets characterized by high overlap between classes. The dataset used consists of 15,000 patient data with 21 initial attributes, which after the preprocessing stage changed to 38 attributes. The evaluation model was carried out using 10-Fold Stratified Cross Validation at three data sharing ratios, namely 90:10, 80:20, and 70:30, with parameter K = 13. The results showed that GA was effective in reducing the data dimensionality significantly to 15–17 selected features and increasing the average internal validation fitness value by around 3%. However, evaluation of the test data shows that the performance of both the KNN and GA-KNN models remains stagnant at an accuracy range of 48–52%, with ROC-AUC values approaching 0.500. This condition indicates that GA tends to find local optimal solutions that are sensitive to the division of the training data, making it difficult for the model to generalize when processing new data. The extreme level of overlap in this dataset is further demonstrated through comparative experiments using feature engineering and the Random Forest algorithm, which both stalled at around 50% accuracy. These results indicate that GA feature selection successfully addresses data complexity, but is not automatically capable of solving class criteria problems if overlap is an inherent characteristic of the data itself.
Downloads
References
K. Wirastuti, N. S. Riasari, D. Djannah, and M. Silviana, “Upaya Pencegahan Stroke melalui Skrining Skor Risiko Stroke dengan Intervensi Penyuluhan dan Pemeriksaan Faktor Risiko Stroke di Kelurahan Bojong Salaman Kecamatan Pusponjolo Selatan Semarang Barat,” J. ABDIMAS-KU J. Pengabdi. Masy. Kedokt., vol. 2, no. 1, pp. 23–29, Jan. 2023, doi: 10.30659/abdimasku.2.1.23-29.
S. Rahayu and Y. Yamasari, “Klasifikasi Penyakit Stroke dengan Metode Support Vector Machine (SVM),” J. Informatics Comput. Sci., vol. 5, no. 03, pp. 440–446, Feb. 2024, doi: 10.26740/jinacs.v5n03.p440-446.
K. Sari et al., “Deteksi Dini Stroke Menggunakan Machine Learning,” INSOLOGI J. Sains dan Teknol., vol. 4, no. 4, pp. 706–720, Aug. 2025, doi: 10.55123/INSOLOGI.V4I4.5590.
Z. Zuriati and N. Qomariyah, “Klasifikasi Penyakit Stroke Menggunakan Algoritma K-Nearest Neighbor (KNN),” ROUTERS J. Sist. dan Teknol. Inf., vol. 1, no. 1, pp. 1–8, Nov. 2022, doi: 10.25181/RT.V1I1.2665.
T. G. Rahayu, “Analisis Faktor Risiko Terjadinya Stroke Serta Tipe Stroke,” Faletehan Heal. J., vol. 10, no. 01, pp. 48–53, Mar. 2023, doi: 10.33746/fhj.v10i01.410.
A. Fitri, I. Afrianty, E. Budianita, and S. K. Gusti, “Implementation of Feature Selection Information Gain in Support Vector Machine Method for Stroke Disease Classification,” Bull. Informatics Data Sci., vol. 4, no. 1, pp. 22–33, May 2025, doi: 10.61944/bids.v4i1.116.
K. Wirastuti, N. Sofi, and I. Rosdiana, “Pendampingan Kader Kesehatan dengan Intervensi Penyuluhan Peduli Cegah Kenali dan Atasi Stroke (Cekatan Stroke) Kelurahan Banjardowo Kecamatan Genuk Kota Semarang,” J. ABDIMAS-KU J. Pengabdi. Masy. Kedokt., vol. 3, no. 1, pp. 6–12, Jan. 2024, doi: 10.30659/abdimasku.3.1.6-12.
I. O. Herlistiono and S. Violina, “Model Prediksi Risiko Stroke Menggunakan Machine Learning,” INTECOMS J. Inf. Technol. Comput. Sci., vol. 7, no. 4, pp. 1230–1238, Jul. 2024, doi: 10.31539/INTECOMS.V7I4.10942.
C. Paramita, C. S. Simbolon, A. S. Pamungkas, J. M. Triono, E. P. Widi Utomo, and E. R. Subhiyakto, “Analisis Pengaruh SMOTE terhadap Kinerja Model KNN untuk Prediksi Risiko Stroke,” J. Inform. J. Pengemb. IT, vol. 10, no. 4, pp. 978–988, Sep. 2025, doi: 10.30591/JPIT.V10I4.8809.
I. Bagus, A. Indra Iswara, G. Anandita, and M. Dahul, “Comparative Analysis of Naïve Bayes and K-Nearest Neighbor (KNN) Algorithms in Stroke Classification,” J. Comput. Networks, Archit. High Perform. Comput., vol. 6, no. 3, pp. 1415–1424, Jul. 2024, doi: 10.47709/CNAHPC.V6I3.4395.
F. Putra, H. F. Tahiyat, R. M. Ihsan, R. Rahmaddeni, and L. Efrizoni, “Penerapan Algoritma K-Nearest Neighbor Menggunakan Wrapper Sebagai Preprocessing untuk Penentuan Keterangan Berat Badan Manusia,” MALCOM Indones. J. Mach. Learn. Comput. Sci., vol. 4, no. 1, pp. 273–281, Jan. 2024, doi: 10.57152/MALCOM.V4I1.1085.
M. A. V. Darmawan, M. M. Al Haromainy, and A. Junaidi, “OPTIMASI ALGORITMA K-NEAREST NEIGHBOR DENGAN ALGORITMA GENETIKA PADA DETEKSI PENYAKIT DIABETES MELLITUS,” JATISI (Jurnal Tek. Inform. dan Sist. Informasi), vol. 12, no. 2, Jun. 2025, doi: 10.35957/JATISI.V12I2.11353.
F. Novianti and N. Ulinnuha, “SELEKSI FITUR ALGORITMA GENETIKA DALAM KLASIFIKASI DATA REKAM MEDIS PCOS MENGGUNAKAN SVM,” NERO (Networking Eng. Res. Oper., vol. 9, no. 1, pp. 9–20, Apr. 2024, doi: 10.21107/NERO.V9I1.25399.
A. Alassaf et al., “Genetic Algorithms and Feature Selection for Improving the Classification Performance in Healthcare,” Int. J. Adv. Comput. Sci. Appl., vol. 15, no. 3, pp. 737–744, Mar. 2024, doi: 10.14569/IJACSA.2024.0150375.
S. Wulandari, Y. I. Mukti, and T. Susanti, “Optimization of the Artificial Neural Network Algorithm with Genetic Algorithm in Stroke Prediction,” Sink. J. dan Penelit. Tek. Inform., vol. 8, no. 2, pp. 1056–1063, Apr. 2024, doi: 10.33395/SINKRON.V8I2.13609.
G. Feng, “Feature selection algorithm based on optimized genetic algorithm and the application in high-dimensional data processing,” PLoS One, vol. 19, no. 5, p. e0303088, May 2024, doi: 10.1371/JOURNAL.PONE.0303088.
P. Koukaras and C. Tjortjis, “Data Preprocessing and Feature Engineering for Data Mining: Techniques, Tools, and Best Practices,” MDPI, vol. 6, no. 10, p. 257, Oct. 2025, doi: 10.3390/AI6100257.
M. R. Firmansyah and Y. P. Astuti, “Stroke Classification Comparison with KNN through Standardization and Normalization Techniques,” Adv. Sustain. Sci. Eng. Technol., vol. 6, no. 1, p. 02401012, Jan. 2024, doi: 10.26877/ASSET.V6I1.17685.
C. Herdian, A. Kamila, and I. G. A. M. Budidarma, “Studi Kasus Feature Engineering Untuk Data Teks: Perbandingan Label Encoding dan One-Hot Encoding Pada Metode Linear Regresi,” Technol. J. Ilm., vol. 15, no. 1, pp. 93–108, Jan. 2024, doi: 10.31602/TJI.V15I1.13457.
T. Abedin, H. Xu, and S. Uddin, “The impact of K selection in K‑fold cross-validation on bias and variance in supervised learning models,” Sci. Reports 2026 161, vol. 16, no. 1, pp. 6084-, Jan. 2026, doi: 10.1038/s41598-026-37247-x.
T. H. Raza, A. Ahmad, M. Z. Hussain, M. Z. Hasan, H. R. Basra, and A. Bilal, “Enhancing Multivariate Data Classification Using Graph Convolutional Networks: A Comparative Evaluation with PCA and t-SNE,” VAWKUM Trans. Comput. Sci., vol. 13, no. 2, pp. 149–165, Oct. 2025, doi: 10.21015/VTCS.V13I2.2052.
Bila bermanfaat silahkan share artikel ini
Berikan Komentar Anda terhadap artikel Evaluasi Seleksi Fitur Algoritma Genetika pada Klasifikasi Penyakit Stroke Menggunakan K-Nearest Neighbor pada Dataset Overlapping
Pages: 233-243
Copyright (c) 2026 Aqmal Syarif Fadilah, Siska Kurnia Gusti, Iwan Iskandar, Iis Afrianty

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under Creative Commons Attribution 4.0 International License that allows others to share the work with an acknowledgment of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgment of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (Refer to The Effect of Open Access).


