Credit Card Fraud Detection Using Support Vector Machine: A Study on Data Balancing Strategies
Abstract
The rise in credit card transactions has been accompanied by an increase in fraudulent activities. One of the key challenges in detecting fraud is the distribution of the dataset, where fraudulent transactions are significantly outnumbered by normal ones. Despite their low occurrence, fraudulent transactions have a significant impact on the banking sector. Therefore, an effective model is needed to identify and estimate fraudulent transactions. This study aims to generate optimal training dataset from an imbalanced one using Adaptive Synthetic Sampling (ADASYN) to enhance the training process of Support Vector Machine (SVM) model. The dataset used consists of anonymized credit card transactions and labeled as either fraudulent or normal, sourced from the Kaggle dataset. It contains transactions made by European cardholders in September 2013, covering a two-day period with 492 fraud cases out of 284,807 transactions. Three datasets were derived from the original: raw, balanced, and support vector-based balanced. The SVM model training on these datasets resulted in sensitivities of 0.39, 0.64, and 0.70, respectively, while the precision values were 0.92, 0.72, and 0.01. The corresponding f-measure values were 0.55, 0.68, and 0.02. The best performance based on the f-measure was achieved using the balanced version of the raw dataset.
Downloads
References
Asosiasi Kartu Kredit Indonesia, “Credit Card Growth.” [Online]. Available: https://www.akki.or.id/index.php/credit-card-growth. [Accessed: 31-Jan-2024].
UK Finance, “Annual Fraud Report,” 2022.
R. K. L. Kennedy, Z. Salekshahrezaee, F. Villanustre, and T. M. Khoshgoftaar, “Iterative Cleaning and Learning of Big Highly-Imbalanced Fraud Data Using Unsupervised Learning,” J. Big Data, vol. 10, no. 1, 2023. doi: 10.1186/s40537-023-00750-3.
M. S. Kraiem, F. Sánchez-Hernández, and M. N. Moreno-García, “Selecting The Suitable Resampling Strategy for Imbalanced Data Classification Regarding Dataset Properties. An Approach Based On Association Models,” Appl. Sci., vol. 11, no. 18, 2021. doi: 10.3390/app11188546.
Y. H. Liu and Y. T. Chen, “Total Margin Based Adaptive Fuzzy Support Vector Machines for Multiview Face Recognition,” Conf. Proc. - IEEE Int. Conf. Syst. Man Cybern., vol. 2, pp. 1704–1711, 2005. doi: 10.1109/icsmc.2005.1571394.
A. S. Desuky and S. Hussain, “An Improved Hybrid Approach for Handling Class Imbalance Problem,” Arab. J. Sci. Eng., vol. 46, no. 4, pp. 3853–3864, 2021. doi: 10.1007/s13369-021-05347-7.
N. V Chawla, K. W. Bowyer, L. O. Hall, and W. P. Kegelmeyer, “SMOTE: Synthetic Minority Over-sampling Technique,” Journal of Artificial Intelligence Research., 2002. doi: https://doi.org/10.1613/jair.953.
H. He, Y. Bai, E. A. Garcia, and S. Li, “ADASYN: Adaptive Synthetic Sampling Approach for Imbalanced Learning,” Proc. Int. Jt. Conf. Neural Networks, no. 3, pp. 1322–1328, 2008. doi: 10.1109/IJCNN.2008.4633969.
C. Li, N. Ding, Y. Zhai, and H. Dong, “Comparative Study on Credit Card Fraud Detection Based On Different Support Vector Machines,” Intell. Data Anal., vol. 25, no. 1, pp. 105–119, 2021. doi: 10.3233/IDA-195011.
V. Chang, L. M. T. Doan, A. Di Stefano, Z. Sun, and G. Fortino, “Digital Payment Fraud Detection Methods In Digital Ages and Industry 4.0,” Comput. Electr. Eng., vol. 100, p. 107734, May 2022. doi: 10.1016/J.COMPELECENG.2022.107734.
J. Chung and K. Lee, “Credit Card Fraud Detection: An Improved Strategy for High Recall Using KNN, LDA, and Linear Regression,” Sensors, vol. 23, no. 18, 2023. doi: 10.3390/s23187788.
L. S. Hasibuan and F. A. Jannah, “Deteksi Penipuan Kartu Kredit Menggunakan Support Vector Machine Dengan Optimasi Grid Search Dan Genetic Algorithm,” Building of Informatics, Technology and Science, vol. 6, no. 1, pp. 344–353, 2023. doi: 10.47065/bits.v6i1.5355.
N. G. Ramadhan, “Comparative Analysis of ADASYN-SVM and SMOTE-SVM Methods on The Detection of Type 2 Diabetes Mellitus,” Sci. J. Informatics, vol. 8, no. 2, pp. 276–282, 2021. doi: 10.15294/sji.v8i2.32484.
K. Shing Lim, L. Hong Lee, and Y.-W. Sim, “A Review of Machine Learning Algorithms for Fraud Detection in Credit Card Transaction,” Int. J. Comput. Sci. Netw. Secur., vol. 21, no. 9, pp. 31–40, 2021. doi: 10.22937/IJCSNS.2021.21.9.4.
“Credit Card Fraud Detection.” [Online]. Available: https://www.kaggle.com/datasets/mlg-ulb/creditcardfraud/data. [Accessed: 16-Jun-2024].
C. T. Chen, C. Lee, S. H. Huang, and W. C. Peng, “Credit Card Fraud Detection via Intelligent Sampling and Self-supervised Learning,” ACM Trans. Intell. Syst. Technol., vol. 15, no. 2, pp. 1–29, 2024. doi: 10.1145/3641283.
D. Singh and B. Singh, “Investigating The Impact Of Data Normalization On Classification Performance,” Appl. Soft Comput., vol. 97, p. 105524, Dec. 2020. doi: 10.1016/J.ASOC.2019.105524.
K. Hornik, A. Weingessel, F. Leisch, and M. D. M. Davidmeyerr-projectorg, Package ‘e1071.’ 2021.
B. Juba and H. S. Le, “Precision-Recall Versus Accuracy And The Role Of Large Data Sets,” 33rd AAAI Conf. Artif. Intell. AAAI 2019, 31st Innov. Appl. Artif. Intell. Conf. IAAI 2019 9th AAAI Symp. Educ. Adv. Artif. Intell. EAAI 2019, pp. 4039–4048, 2019. doi: 10.1609/aaai.v33i01.33014039.
A. Balla, M. H. Habaebi, E. A. A. Elsheikh, M. R. Islam, and F. M. Suliman, “The Effect of Dataset Imbalance on the Performance of SCADA Intrusion Detection Systems,” Sensors, vol. 23, no. 2, 2023. doi: 10.3390/s23020758.
S. J. Yen and Y. S. Lee, “Cluster-based Under-sampling Approaches for Imbalanced Data Distributions,” Expert Syst. Appl., vol. 36, no. 3 PART 1, pp. 5718–5727, 2009. doi: 10.1016/j.eswa.2008.06.108.
S. Bagui and K. Li, “Resampling imbalanced data for network intrusion detection datasets,” J. Big Data, vol. 8, no. 1, 2021. doi: 10.1186/s40537-020-00390-x
Bila bermanfaat silahkan share artikel ini
Berikan Komentar Anda terhadap artikel Credit Card Fraud Detection Using Support Vector Machine: A Study on Data Balancing Strategies
Pages: 1900-1909
Copyright (c) 2025 Lailan Sahrina Hasibuan

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under Creative Commons Attribution 4.0 International License that allows others to share the work with an acknowledgment of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgment of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (Refer to The Effect of Open Access).





















