Revisiting SMOTE in Balanced Medical Data: A Comparative Evaluation of SVM, Random Forest, and KNN


  • Kelvin Leonardi Kohsasih * Mail STMIK TIME, Medan, Indonesia
  • Joni Joni STMIK TIME, Medan, Indonesia
  • Herman Herman STMIK TIME, Medan, Indonesia
  • Pieter Octaviandy STMIK TIME, Medan, Indonesia
  • (*) Corresponding Author
Keywords: Heart Disease; Classification; SMOTE; Support Vector Machine; Random Forest; K-Nearest Neighbor

Abstract

Heart disease is one of the leading causes of death worldwide, making data-driven early detection crucial for supporting medical decision-making systems. A major challenge in developing heart disease prediction models is dataset quality, including the often imbalanced class distribution, which can impact the performance of classification algorithms. This study aims to analyze the effect of the Synthetic Minority Oversampling Technique (SMOTE) on the performance of three classification algorithms: Support Vector Classifier (SVC), Random Forest (RF), and K-Nearest Neighbor (KNN). The dataset used is heart_disease50.csv with 4,001 patient data consisting of 21 predictor attributes and one target variable (heart disease status: “Yes” or “No”) with a relatively balanced class distribution. The research process includes data preprocessing (data cleaning, normalization, and encoding), data partitioning using Stratified K-Fold Cross Validation (k=5), applying SMOTE to training data, building a classification model, and evaluation using accuracy, precision, recall, F1-score, and AUC-ROC metrics. The results showed that applying SMOTE did not always improve performance. The SVC model with SMOTE experienced a decrease in accuracy (0.4819) compared to the one without SMOTE (0.5106), while Random Forest remained relatively stable with insignificant differences (0.4669 without SMOTE and 0.4644 with SMOTE). KNN with SMOTE emerged as the best model with an accuracy of 0.5268 and a precision of 0.5271, although the AUC-ROC remained the same as KNN without SMOTE (0.5135). Overall, these results confirm that the effectiveness of SMOTE is highly dependent on dataset conditions, and in cases with relatively balanced data, SMOTE does not provide significant benefits. Therefore, improving the performance of heart disease prediction classification is recommended through hyperparameter optimization strategies, relevant feature selection, or the use of more sophisticated algorithms such as Gradient Boosting or Neural Networks.

Downloads

Download data is not yet available.

References

M. Said, Y. Omar, S. Safwat, and A. Salem, “Explainable Artificial Intelligence Powered Model for Explainable Detection of Stroke Disease BT - Proceedings of the 8th International Conference on Advanced Intelligent Systems and Informatics 2022,” A. E. Hassanien, V. Snášel, M. Tang, T.-W. Sung, and K.-C. Chang, Eds., Cham: Springer International Publishing, 2023, pp. 211–223.

P. R. Kumar, “Identification of noteworthy features and data mining techniques for heart disease prediction,” Int. J. Model. Simulation, Sci. Comput., 2024, doi: 10.1142/S1793962325500102.

N. A. M. Zaini and M. K. Awang, “Performance Comparison between Meta-classifier Algorithms for Heart Disease Classification,” Int. J. Adv. Comput. Sci. Appl., vol. 13, no. 10, pp. 323–328, 2022, doi: 10.14569/IJACSA.2022.0131039.

S. Sowmya and D. Jose, “Contemplate on ECG signals and classification of arrhythmia signals using CNN-LSTM deep learning model,” Meas. Sensors, vol. 24, no. October, p. 100558, 2022, doi: 10.1016/j.measen.2022.100558.

P. Alkhairi, A. P. Windarto, and M. M. Efendi, “Optimasi LSTM Mengurangi Overfitting untuk Klasifikasi Teks Menggunakan Kumpulan Data Ulasan Film Kaggle IMDB,” vol. 6, no. 2, pp. 1142–1150, 2024, doi: 10.47065/bits.v6i2.5850.

J. Lee et al., “Unsupervised machine learning for identifying important visual features through bag-of-words using histopathology data from chronic kidney disease,” Sci. Rep., vol. 12, no. 1, pp. 1–13, 2022, doi: 10.1038/s41598-022-08974-8.

S. A. Alex, “Classification of Imbalanced Data Using SMOTE and AutoEncoder Based Deep Convolutional Neural Network,” Int. J. Uncertain. Fuzziness Knowl. Based Syst., vol. 31, no. 3, pp. 437–469, 2023, doi: 10.1142/S0218488523500228.

A. Desiani, “Handling the imbalanced data with missing value elimination smote in the classification of the relevance education background with graduates employment,” Iaes Int. J. Artif. Intell., vol. 10, no. 2, pp. 346–354, 2021, doi: 10.11591/ijai.v10.i2.pp346-354.

J. A. Adisa, “The Effect of Imbalanced Data and Parameter Selection via Genetic Algorithm Long Short-Term Memory (LSTM) for Financial Distress Prediction,” IAENG Int. J. Appl. Math., vol. 53, no. 3, 2023.

Asniar, “SMOTE-LOF for noise identification in imbalanced data classification,” J. King Saud Univ. Comput. Inf. Sci., vol. 34, no. 6, pp. 3413–3423, 2022, doi: 10.1016/j.jksuci.2021.01.014.

J. Park, “A study on improving turnover intention forecasting by solving imbalanced data problems: focusing on SMOTE and generative adversarial networks,” J. Big Data, vol. 10, no. 1, 2023, doi: 10.1186/s40537-023-00715-6.

K. V. Rani, “Lung Lesion Classification Scheme Using Optimization Techniques and Hybrid (KNN-SVM) Classifier,” IETE J. Res., vol. 68, no. 2, pp. 1485–1499, 2022, doi: 10.1080/03772063.2019.1654935.

K. M. Sunnetci, “Face Mask Detection Using GoogLeNet CNN-Based SVM Classifiers,” Gazi Univ. J. Sci., vol. 36, no. 2, pp. 645–658, 2023, doi: 10.35378/gujs.1009359.

S. DEMİR and E. K. ŞAHİN, “Evaluation of Oversampling Methods (OVER, SMOTE, and ROSE) in Classifying Soil Liquefaction Dataset based on SVM, RF, and Naïve Bayes,” Eur. J. Sci. Technol., no. 34, pp. 142–147, 2022, doi: 10.31590/ejosat.1077867.

L. Liu, “Average AoI Minimization in UAV-Assisted Data Collection With RF Wireless Power Transfer: A Deep Reinforcement Learning Scheme,” IEEE Internet Things J., vol. 9, no. 7, pp. 5216–5228, 2022, doi: 10.1109/JIOT.2021.3110138.

J. Huang, “Deciphering decision-making mechanisms for the susceptibility of different slope geohazards: A case study on a SMOTE-RF-SHAP hybrid model,” J. Rock Mech. Geotech. Eng., vol. 17, no. 3, pp. 1612–1630, 2025, doi: 10.1016/j.jrmge.2024.03.008.

Z. A. Sejuti and M. S. Islam, “A hybrid CNN–KNN approach for identification of COVID-19 with 5-fold cross validation,” Sensors Int., vol. 4, no. November 2022, p. 100229, 2023, doi: 10.1016/j.sintl.2023.100229.

W. Sinhashthita, “Improving KNN Algorithm Based on Weighted Attributes by Pearson Correlation Coefficient and PSO Fine Tuning,” 2020.

T. Benil, “Efficient data pruning using optimal KNN for weather forecasting in cloud computing,” Int. J. Glob. Warm., vol. 30, no. 2, pp. 137–151, 2023, doi: 10.1504/IJGW.2023.130984.

İ. Atik, “Pneumonia detection on chest x-ray images using residual convolutional neural network,” J. Fac. Eng. Archit. Gazi Univ., vol. 39, no. 3, pp. 1719–1731, 2024, doi: 10.17341/gazimmfd.1271385.

D. I. Swasono, “Classification of Air-Cured Tobacco Leaf Pests Using Pruning Convolutional Neural Networks and Transfer Learning,” Int. J. Adv. Sci. Eng. Inf. Technol., vol. 12, no. 3, pp. 1229–1235, 2022, doi: 10.18517/ijaseit.12.3.15950.

J. Chen, H. Huang, A. G. Cohn, D. Zhang, and M. Zhou, “Machine learning-based classification of rock discontinuity trace: SMOTE oversampling integrated with GBT ensemble learning,” Int. J. Min. Sci. Technol., vol. 32, no. 2, pp. 309–322, 2022, doi: 10.1016/j.ijmst.2021.08.004.

H. Hairani, “Improvement Performance of the Random Forest Method on Unbalanced Diabetes Data Classification Using Smote-Tomek Link,” Int. J. Informatics Vis., vol. 7, no. 1, pp. 258–264, 2023, doi: 10.30630/joiv.7.1.1069.

T. Wu, “Intrusion detection system combined enhanced random forest with SMOTE algorithm,” EURASIP J. Adv. Signal Process., vol. 2022, no. 1, 2022, doi: 10.1186/s13634-022-00871-6.

N. G. Ramadhan, “MODIFIED SMOTE AND ENSEMBLE LEARNING BASED ON EXPERT JUDGMENT FOR CHRONIC DISEASES PREDICTION,” Int. J. Innov. Comput. Inf. Control, vol. 21, no. 4, pp. 1003–1023, 2025, doi: 10.24507/ijicic.21.04.1003.

M. Mishra, “DTCDWT-SMOTE-XGBoost-Based Islanding Detection for Distributed Generation Systems: An Approach of Class-Imbalanced Issue,” IEEE Syst. J., vol. 16, no. 2, pp. 2008–2019, 2022, doi: 10.1109/JSYST.2021.3086298.

H. Naz, “SMOTE-SMO-based expert system for type II diabetes detection using PIMA dataset,” Int. J. Diabetes Dev. Ctries., vol. 42, no. 2, pp. 245–253, 2022, doi: 10.1007/s13410-021-00969-x.

A. R. Vaka, B. Soni, and S. R. K., “Breast cancer detection by leveraging Machine Learning,” ICT Express, vol. 6, no. 4, pp. 320–324, 2020, doi: 10.1016/j.icte.2020.04.009.

A. Rahmat, H. Hardi, F. A. Syam, Z. Zamzami, B. Febriadi, and A. P. Windarto, “Utilization of the field of data mining in mapping the area of the Human Development Index (HDI) in Indonesia,” J. Phys. Conf. Ser., vol. 1783, no. 1, 2021, doi: 10.1088/1742-6596/1783/1/012035.

H. Wang, “Research on the Application of Random Forest-based Feature Selection Algorithm in Data Mining Experiments,” Int. J. Adv. Comput. Sci. Appl., vol. 14, no. 10, pp. 505–518, 2023, doi: 10.14569/IJACSA.2023.0141054.

M. Irshada, “SMOTE and ExtraTreesRegressor based random forest technique for predicting Australian rainfall,” Int. J. Inf. Technol. Singapore, vol. 15, no. 3, pp. 1679–1687, 2023, doi: 10.1007/s41870-023-01185-y.

A. E. K. Gunawan, “Stock Price Movement Classification Using Ensembled Model of Long Short-Term Memory (LSTM) and Random Forest (RF),” Int. J. Informatics Vis., vol. 7, no. 4, pp. 2255–2262, 2023, doi: 10.30630/joiv.7.4.1640.

C. Y. Lee, K. Y. Huang, Y. X. Shen, and Y. C. Lee, “Improved weighted k-nearest neighbor based on PSO for wind power system state recognition,” Energies, vol. 13, no. 20, 2020, doi: 10.3390/en13205520.

S. Pande, A. Khamparia, and D. Gupta, “Feature selection and comparison of classification algorithms for wireless sensor networks,” J. Ambient Intell. Humaniz. Comput., vol. 14, no. 3, pp. 1977–1989, 2023, doi: 10.1007/s12652-021-03411-6.

A. P. Ayudhitama and U. Pujianto, “Analisa 4 Algoritma Dalam Klasifikasi Liver Menggunakan Rapidminer,” J. Inform. Polinema, vol. 6, no. 2, pp. 1–9, 2020, doi: 10.33795/jip.v6i2.274.

H. Sun, “Multi-variety monitoring of potato late blight severity using UAV data with improved SMOTE-CS for small sample modeling and deep feature learning,” Eur. J. Agron., vol. 169, 2025, doi: 10.1016/j.eja.2025.127702.

A. R. Doni, “Artificial Intelligence Driven Skin Cancer Detection Using R-FCN Enhanced Deep Convolutional Neural Networks with SMOTE Balancing,” Int. J. Eng. Sci. Inf. Technol., vol. 5, no. 4, pp. 28–36, 2025, doi: 10.52088/ijesty.v5i4.1196.

L. Azoulay-Younes, “Zinc electrode manufacturing quality control with machine learning: Using SMOTE & image augmentation to prevent overfitting,” J. Eng. Res. Kuwait, 2025, doi: 10.1016/j.jer.2025.01.002.


Bila bermanfaat silahkan share artikel ini

Berikan Komentar Anda terhadap artikel Revisiting SMOTE in Balanced Medical Data: A Comparative Evaluation of SVM, Random Forest, and KNN

Dimensions Badge
Article History
Submitted: 2025-10-06
Published: 2025-12-16
Abstract View: 54 times
PDF Download: 40 times
How to Cite
Kohsasih, K., Joni, J., Herman, H., & Octaviandy, P. (2025). Revisiting SMOTE in Balanced Medical Data: A Comparative Evaluation of SVM, Random Forest, and KNN. Building of Informatics, Technology and Science (BITS), 7(3), 1753-1760. https://doi.org/10.47065/bits.v7i3.8471
Issue
Section
Articles