Credit Card Fraud Detection Using Support Vector Machine: A Study on Data Balancing Strategies


  • Lailan Sahrina Hasibuan * Mail IPB University, Bogor, Indonesia
  • (*) Corresponding Author
Keywords: Adaptive Synthetic Sampling (ADASYN); Support Vector-Based Dataset, Imbalanced Dataset; Credit Card Fraud Detection; Support Vector Machine (SVM)

Abstract

The rise in credit card transactions has been accompanied by an increase in fraudulent activities. One of the key challenges in detecting fraud is the distribution of the dataset, where fraudulent transactions are significantly outnumbered by normal ones. Despite their low occurrence, fraudulent transactions have a significant impact on the banking sector. Therefore, an effective model is needed to identify and estimate fraudulent transactions. This study aims to generate optimal training dataset from an imbalanced one using Adaptive Synthetic Sampling (ADASYN) to enhance the training process of Support Vector Machine (SVM) model. The dataset used consists of anonymized credit card transactions and labeled as either fraudulent or normal, sourced from the Kaggle dataset. It contains transactions made by European cardholders in September 2013, covering a two-day period with 492 fraud cases out of 284,807 transactions. Three datasets were derived from the original: raw, balanced, and support vector-based balanced. The SVM model training on these datasets resulted in sensitivities of 0.39, 0.64, and 0.70, respectively, while the precision values were 0.92, 0.72, and 0.01. The corresponding f-measure values were 0.55, 0.68, and 0.02. The best performance based on the f-measure was achieved using the balanced version of the raw dataset.

Downloads

Download data is not yet available.

References

Asosiasi Kartu Kredit Indonesia, “Credit Card Growth.” [Online]. Available: https://www.akki.or.id/index.php/credit-card-growth. [Accessed: 31-Jan-2024].

UK Finance, “Annual Fraud Report,” 2022.

R. K. L. Kennedy, Z. Salekshahrezaee, F. Villanustre, and T. M. Khoshgoftaar, “Iterative Cleaning and Learning of Big Highly-Imbalanced Fraud Data Using Unsupervised Learning,” J. Big Data, vol. 10, no. 1, 2023. doi: 10.1186/s40537-023-00750-3.

M. S. Kraiem, F. Sánchez-Hernández, and M. N. Moreno-García, “Selecting The Suitable Resampling Strategy for Imbalanced Data Classification Regarding Dataset Properties. An Approach Based On Association Models,” Appl. Sci., vol. 11, no. 18, 2021. doi: 10.3390/app11188546.

Y. H. Liu and Y. T. Chen, “Total Margin Based Adaptive Fuzzy Support Vector Machines for Multiview Face Recognition,” Conf. Proc. - IEEE Int. Conf. Syst. Man Cybern., vol. 2, pp. 1704–1711, 2005. doi: 10.1109/icsmc.2005.1571394.

A. S. Desuky and S. Hussain, “An Improved Hybrid Approach for Handling Class Imbalance Problem,” Arab. J. Sci. Eng., vol. 46, no. 4, pp. 3853–3864, 2021. doi: 10.1007/s13369-021-05347-7.

N. V Chawla, K. W. Bowyer, L. O. Hall, and W. P. Kegelmeyer, “SMOTE: Synthetic Minority Over-sampling Technique,” Journal of Artificial Intelligence Research., 2002. doi: https://doi.org/10.1613/jair.953.

H. He, Y. Bai, E. A. Garcia, and S. Li, “ADASYN: Adaptive Synthetic Sampling Approach for Imbalanced Learning,” Proc. Int. Jt. Conf. Neural Networks, no. 3, pp. 1322–1328, 2008. doi: 10.1109/IJCNN.2008.4633969.

C. Li, N. Ding, Y. Zhai, and H. Dong, “Comparative Study on Credit Card Fraud Detection Based On Different Support Vector Machines,” Intell. Data Anal., vol. 25, no. 1, pp. 105–119, 2021. doi: 10.3233/IDA-195011.

V. Chang, L. M. T. Doan, A. Di Stefano, Z. Sun, and G. Fortino, “Digital Payment Fraud Detection Methods In Digital Ages and Industry 4.0,” Comput. Electr. Eng., vol. 100, p. 107734, May 2022. doi: 10.1016/J.COMPELECENG.2022.107734.

J. Chung and K. Lee, “Credit Card Fraud Detection: An Improved Strategy for High Recall Using KNN, LDA, and Linear Regression,” Sensors, vol. 23, no. 18, 2023. doi: 10.3390/s23187788.

L. S. Hasibuan and F. A. Jannah, “Deteksi Penipuan Kartu Kredit Menggunakan Support Vector Machine Dengan Optimasi Grid Search Dan Genetic Algorithm,” Building of Informatics, Technology and Science, vol. 6, no. 1, pp. 344–353, 2023. doi: 10.47065/bits.v6i1.5355.

N. G. Ramadhan, “Comparative Analysis of ADASYN-SVM and SMOTE-SVM Methods on The Detection of Type 2 Diabetes Mellitus,” Sci. J. Informatics, vol. 8, no. 2, pp. 276–282, 2021. doi: 10.15294/sji.v8i2.32484.

K. Shing Lim, L. Hong Lee, and Y.-W. Sim, “A Review of Machine Learning Algorithms for Fraud Detection in Credit Card Transaction,” Int. J. Comput. Sci. Netw. Secur., vol. 21, no. 9, pp. 31–40, 2021. doi: 10.22937/IJCSNS.2021.21.9.4.

“Credit Card Fraud Detection.” [Online]. Available: https://www.kaggle.com/datasets/mlg-ulb/creditcardfraud/data. [Accessed: 16-Jun-2024].

C. T. Chen, C. Lee, S. H. Huang, and W. C. Peng, “Credit Card Fraud Detection via Intelligent Sampling and Self-supervised Learning,” ACM Trans. Intell. Syst. Technol., vol. 15, no. 2, pp. 1–29, 2024. doi: 10.1145/3641283.

D. Singh and B. Singh, “Investigating The Impact Of Data Normalization On Classification Performance,” Appl. Soft Comput., vol. 97, p. 105524, Dec. 2020. doi: 10.1016/J.ASOC.2019.105524.

K. Hornik, A. Weingessel, F. Leisch, and M. D. M. Davidmeyerr-projectorg, Package ‘e1071.’ 2021.

B. Juba and H. S. Le, “Precision-Recall Versus Accuracy And The Role Of Large Data Sets,” 33rd AAAI Conf. Artif. Intell. AAAI 2019, 31st Innov. Appl. Artif. Intell. Conf. IAAI 2019 9th AAAI Symp. Educ. Adv. Artif. Intell. EAAI 2019, pp. 4039–4048, 2019. doi: 10.1609/aaai.v33i01.33014039.

A. Balla, M. H. Habaebi, E. A. A. Elsheikh, M. R. Islam, and F. M. Suliman, “The Effect of Dataset Imbalance on the Performance of SCADA Intrusion Detection Systems,” Sensors, vol. 23, no. 2, 2023. doi: 10.3390/s23020758.

S. J. Yen and Y. S. Lee, “Cluster-based Under-sampling Approaches for Imbalanced Data Distributions,” Expert Syst. Appl., vol. 36, no. 3 PART 1, pp. 5718–5727, 2009. doi: 10.1016/j.eswa.2008.06.108.

S. Bagui and K. Li, “Resampling imbalanced data for network intrusion detection datasets,” J. Big Data, vol. 8, no. 1, 2021. doi: 10.1186/s40537-020-00390-x


Bila bermanfaat silahkan share artikel ini

Berikan Komentar Anda terhadap artikel Credit Card Fraud Detection Using Support Vector Machine: A Study on Data Balancing Strategies

Dimensions Badge
Article History
Submitted: 2025-10-12
Published: 2025-12-26
Abstract View: 37 times
PDF Download: 30 times
How to Cite
Hasibuan, L. (2025). Credit Card Fraud Detection Using Support Vector Machine: A Study on Data Balancing Strategies. Building of Informatics, Technology and Science (BITS), 7(3), 1900-1909. https://doi.org/10.47065/bits.v7i3.8514
Issue
Section
Articles