Hybrid Weighted K-Nearest Neighbor dan Naive Bayes untuk Klasifikasi Diabetes Melitus


  • Aklima Laduna Ramadya * Mail STIKOM Tunas Bangsa, Pematang Siantar, Indonesia
  • Solikhun Solikhun STIKOM Tunas Bangsa, Pematang Siantar, Indonesia
  • Timbo Faritcan P. Siallagan Universitas Mandiri, Jawa Barat, Indonesia
  • (*) Corresponding Author
Keywords: Diabetes Melitus; Klasifikasi Hybrid; Weighted K-Nearest Neighbor; Naive Bayes

Abstract

AbstractDiabetes mellitus is a chronic metabolic disease with a continuously increasing prevalence and the potential to cause various serious complications if not detected early. The utilization of machine learning has become one of the effective approaches to support a faster and more accurate diagnostic process. This study proposes a Hybrid Weighted K-Nearest Neighbor (WKNN)–Naive Bayes model that combines a distance-based approach with a probabilistic approach for the classification of diabetes mellitus. The dataset used is the Pima Indians Diabetes Dataset, consisting of 768 patient records with eight predictor attributes and one target variable. The preprocessing stage included missing value handling using imputation, data normalization using RobustScaler, feature engineering, and class balancing using the Synthetic Minority Over-sampling Technique (SMOTE). Model evaluation was performed using Stratified 10-Fold Cross Validation and compared against K-Nearest Neighbor (KNN), Weighted K-Nearest Neighbor (WKNN), and Naive Bayes models. The results show that the Hybrid WKNN–Naive Bayes model achieved an Accuracy of 88.50%, Precision of 85.32%, Recall of 93.00%, F1-Score of 88.99%, and AUC-ROC of 91.09%, demonstrating good classification performance for diabetes mellitus diagnosis. The Weighted K-Nearest Neighbor (WKNN) model also achieved an Accuracy of 88.50%, with a Recall of 94.00%, F1-Score of 89.10%, and AUC-ROC of 91.73%, while the Naive Bayes model produced an Accuracy of 71.00%. These results indicate that the Hybrid WKNN–Naive Bayes approach is able to deliver competitive classification performance through the combination of distance-based and probabilistic methods, making it a promising candidate for use as a decision support system in the early detection of diabetes mellitus.

Downloads

Download data is not yet available.

References

A. Bangun and E. Rachmat, “Analisis Perbandingan Algoritma KNN dan Naïve Bayes dalam Mendiagnosis Penyakit Diabetes Mellitus,” Ilm. KOMPUTASI, vol. 23, no. September, pp. 387–396, 2024.

A. Mulyani, S. Khoerunisa, and D. Kurniadi, “Comparison of KNN and SVM Algorithms Performance Using SMOTE to Classify Diabetes,” J. Nas. Tek. Elektro dan Teknol. Inf., vol. 14, no. 1, pp. 25–34, 2025.

A. F. N. M. Selly Rahmawati, Arief Wibowo, “Improving Diabetes Prediction Accuracy in Indonesia : A Comparative Analysis of SVM , Logistic Regression , and Naive Bayes with SMOTE and,” RESTI (Rekayasa Sist. dan Teknol. Informasi), vol. 5, no. 158, pp. 7–10, 2026.

R. Pratama, “Implementation of Diabetes Prediction Model Using Machine Learning,” JUTIF, vol. 5, no. 3, pp. 455–466, 2024, doi: 10.20884/1.jutif.2024.5.3.2593.

A. Rifai, “Analysis of Classification Algorithm in Unbalanced Diabetes Dataset,” J. Ris. Inform., vol. 7, no. 1, pp. 33–42, 2025.

A. K. E. Pily, “Komparasi Algoritma KNN dan Naive Bayes pada Diabetes Gestasional,” Indones. J. Comput. Sci., vol. 13, no. 2, pp. 210–220, 2024.

M. C. Ibrahim, “Comparison of Diabetes Prediction Data Using Machine Learning,” MALCOM, vol. 5, no. 4, pp. 550–560, 2025, doi: 10.57152/malcom.v5i4.2301.

N. K. Sowabi and others, “Optimasi Algoritma K-Nearest Neighbors Menggunakan Teknik Bayesian Optimization untuk Klasifikasi Diabetes,” JOSH, vol. 6, no. 1, pp. 294–301, 2024, doi: 10.47065/josh.v6i1.5975.

H. A. D. Fasnuari and others, “Application of K-Nearest Neighbor Algorithm for Classification of Diabetes Mellitus,” Antivirus, vol. 16, no. 2, pp. 133–142, 2022, doi: 10.35457/antivirus.v16i2.2445.

F. Ruziq, M. R. Wayahdi, and S. H. N. Ginting, “Diabetes Prediction Based on Medical Records Using K-NN,” J. Sci. Soc. Res., vol. 8, no. 3, pp. 3776–3782, 2025, doi: 10.54314/JSSR.V8I3.2981.

N. C. Sari, “Komparasi Algoritma Naive Bayes dan Gradient Boosting untuk Prediksi Diabetes,” TEKNOSI, vol. 10, no. 2, pp. 101–110, 2024.

V. Wulandari and others, “Implementasi Algoritma Naive Bayes dan KNN untuk Klasifikasi Penyakit Ginjal Kronik,” MALCOM, vol. 4, no. 2, pp. 420–428, 2024, doi: 10.57152/malcom.v4i2.1229.

Q. R. Cahyani, “Prediksi Risiko Penyakit Diabetes Menggunakan Algoritma Regresi Logistik,” JOMLAI, vol. 1, no. 2, pp. 120–129, 2022, doi: 10.55123/jomlai.v1i2.598.

M. Ardiansyah, “Analisis Perbandingan Akurasi Algoritma Naive Bayes dan C4.5 untuk Klasifikasi Diabetes,” Edumatic, vol. 5, no. 2, pp. 145–154, 2021, doi: 10.29408/edumatic.v5i2.3424.

R. Maulana and Eliyani, “Diabetes Classification Algorithm Optimization Using Particle Swarm Optimization on Naive Bayes, C4.5 and Random Forest,” J. Sisfokom, vol. 14, no. 4, pp. 1–12, 2025, doi: 10.32736/sisfokom.v14i4.2431.

F. Sholekhah and others, “Perbandingan Algoritma Naive Bayes dan K-Nearest Neighbor untuk Klasifikasi Metabolik Sindrom,” MALCOM, vol. 4, no. 2, pp. 507–514, 2024, doi: 10.57152/malcom.v4i2.1249.

A. Ridwan, “Penerapan Algoritma Naive Bayes Untuk Klasifikasi Penyakit Diabetes Mellitus,” SISKOM-KB, vol. 4, no. 1, pp. 1–9, 2021, doi: 10.47970/siskom-kb.v4i1.169.

G. Sulaeman, Y. Setiya, R. Nur, A. W. Paramadini, and D. Aldo, “Optimization Of Extreme Learning Machine Models Using Metaheuristic Approaches For Diabetes Classification,” J. Tek. Inform., vol. 6, no. 3, pp. 1503–1515, 2025.

D. Nasien and others, “Perbandingan Implementasi Machine Learning Menggunakan Metode KNN, Naive Bayes, dan Logistic Regression Untuk Mengklasifikasi Penyakit Diabetes,” JEKIN, vol. 4, no. 1, pp. 10–17, 2024, doi: 10.58794/jekin.v4i1.640.

N. L. Anggreini, A. Yuliana, D. S. Ramdan, and W. Al-dayyeni, “Improving Diabetes Prediction Performance Using Random Forest Classifier with Hyperparameter Tuning,” J. Tek. Inform., vol. 6, no. 4, pp. 1847–1860, 2025.

I. Riadi, A. Yudhana, and G. C. Kurniawan, “Evaluating Synthetic Minority Oversampling Technique Strategies for Diabetes Mellitus Classification using K-Nearest Neighbors Algorithm,” J. Tek. Inform., vol. 6, no. 5, pp. 3958–3970, 2025.

N. Charibaldi, N. H. Cahyana, M. S. Wicaksono, and C. Author, “Optimization of Foodstuffs for Patients with Hypertension Using the Improved Particle Swarm Optimization Method,” Int. J. Artif. Intell. Robot., vol. 4, no. 1, pp. 31–38, 2022.

A. Wantoro, A. F. Yuliana, D. Yana, A. Andini, and I. Awaliyani, “Optimizing Type 2 Diabetes Classification with Feature Selection and Class Balancing in Machine Learning,” J. Tek. Inform., vol. 6, no. 4, pp. 2625–2637, 2025.

N. A. Nanda, Y. Farida, and W. Dianita, “Implementation of SMOTE to Improve the Performance of Random Forest Classification in Credit Risk Assessment in Banking,” INTENSIF, vol. 9, no. 2, pp. 158–177, 2025.

E. P. H, S. Assegaff, and J. Jasmir, “Comparison of AdaBoost and Random Forest Methods in Osteoporosis Risk Prediction Based on Machine Learning,” J. Tek. Inform., vol. 7, no. 1, pp. 201–209, 2026.

S. Gesti, A. Utami, H. Setiadi, and A. Rohmadi, “Comparative Analysis of Machine Learning Algorithms with RFE-CV for Student Dropout Prediction,” J. Tek. Inform., vol. 6, no. 3, pp. 1319–1338, 2025.

F. Ruziq and M. R. Wayahdi, “Web-Based Diabetes Risk Prediction System Using K-NN on Kaggle Early Stage Diabetes Dataset,” J. Tek. Inform., vol. 6, no. 5, pp. 3217–3229, 2025.

M. H. Putri, U. Zaky, and B. A. Prabawa, “Optimizing Data Augmentation Parameters in YOLOv11 for Enhanced Rip Current Detection on Small Datasets from Depok-Parangtritis Coastline,” J. Tek. Inform., vol. 6, no. 5, pp. 3938–3957, 2025.

A. Suris, A. Thobirin, S. Surono, N. A. Abdulnazar, and U. A. Dahlan, “Enhancing Multi-Class Classification of Non- Functional Requirements Using a BERT-DBN Hybrid Model,” INTENSIF, vol. 9, no. 2, pp. 195–210, 2025.

A. S. Sunge, H. Muhammad, and M. Putra, “Interpretable Machine Learning for Employee Recruitment Prediction Using Boruta , CatBoost , Lasso , Logistic Regression , NLP , and RFE Feature Selection,” J. Tek. Inform., vol. 6, no. 4, pp. 2153–2170, 2025.

S. Salsabila, Y. Sibaroni, and S. S. Prasetiyowati, “Geo-Sentiment Analysis of Public Opinion of X Users towards the Documentary Film Dirty Vote using the Bidirectional Long Short-Term Memory Method,” J. Tek. Inform., vol. 6, no. 2, pp. 539–556, 2025.

B. V. Indriyono, M. S. Hidajat, T. E. Rahayuningtyas, Z. Pratama, I. Irdinawati, and E. Citra, “Expert System for Detecting Diseases of Potatoes of Granola Varieties Using Certainty Factor Method,” Int. J. Artif. Intell. Robot., vol. 4, no. 2, pp. 70–77, 2022.

A. M. Argina, “Penerapan Metode Klasifikasi K-Nearest Neighbor pada Dataset Penderita Penyakit Diabetes,” Indones. J. Data Sci., vol. 1, no. 2, pp. 29–33, 2020, doi: 10.33096/ijodas.v1i2.11.

M. Saputra and others, “Analisis Metode Algoritma K-Nearest Neighbor dan Naive Bayes untuk Klasifikasi Diabetes Mellitus,” Tekinkom, vol. 6, no. 2, pp. 723–729, 2023, doi: 10.37600/tekinkom.v6i2.942.


Bila bermanfaat silahkan share artikel ini

Berikan Komentar Anda terhadap artikel Hybrid Weighted K-Nearest Neighbor dan Naive Bayes untuk Klasifikasi Diabetes Melitus

Dimensions Badge
Article History
Published: 2026-06-30
Abstract View: 0 times
PDF Download: 0 times
How to Cite
Ramadya, A., Solikhun, S., & P. Siallagan, T. (2026). Hybrid Weighted K-Nearest Neighbor dan Naive Bayes untuk Klasifikasi Diabetes Melitus. Bulletin of Data Science, 5(3), 357-368. https://doi.org/10.47065/bulletinds.v5i3.10391
Issue
Section
Articles