Credit Card Fraud Detection under Extreme Class Imbalance: A Comparison of KNN and Logistic Regression

Authors

  • Felicia Sword Fakultas Teknologi Informasi, Universitas Ciputra, Surabaya, Indonesia
  • Christopher Andreas Informatics, School of Information Technology, Universitas Ciputra

DOI:

https://doi.org/10.24036/ujsds/vol4-iss3/539

Keywords:

Credit Card Fraud Detection, Class Imbalance, K-Nearest Neighbor, Logistic Regression, SMOTE

Abstract

Credit card fraud is a serious threat in the digital financial ecosystem and is characterised by extreme class imbalance, with fraudulent transactions typically below 1%. This study compares two standard classification algorithms, K-Nearest Neighbor (KNN) and Logistic Regression (LR), for detecting fraudulent transactions on the Sparkov dataset (1.85 million transactions; a stratified subsample of 100,000 rows; 0.52% fraud rate), and analyses the effect of the Synthetic Minority Over-sampling Technique (SMOTE). Preprocessing includes temporal feature engineering, haversine distance, leak-free per-card behavioural features, one-hot and label encoding, and z-score standardisation. Models are evaluated on a stratified 80:20 split using the confusion matrix, accuracy, precision, recall, F1-score, ROC-AUC, and PR-AUC, complemented by decision-threshold tuning, confidence intervals over five repetitions, and the McNemar test. No single model dominates across all metrics. At the default threshold, KNN baseline achieves the highest F1 (0.407) and precision (0.540), while LR baseline achieves the highest PR-AUC (0.246); LR+SMOTE leads on recall (0.712) and ROC-AUC (0.861) but with very low precision (0.025). Threshold tuning lets LR baseline reach the best F1 (0.422 at a 0.071 cut-off). McNemar shows the KNN–LR difference is not significant at baseline (p = 0.282) but significant under SMOTE (p < 0.001). The main finding is that under severe imbalance ROC-AUC can be misleading and PR-AUC is more informative; KNN baseline is a balanced detector without tuning, threshold-tuned LR baseline gives the best single operating point, and LR+SMOTE suits cases where recall is the priority.

Published

2026-08-31

How to Cite

Sword, F., & Andreas, C. (2026). Credit Card Fraud Detection under Extreme Class Imbalance: A Comparison of KNN and Logistic Regression. UNP Journal of Statistics and Data Science, 4(3), 417–424. https://doi.org/10.24036/ujsds/vol4-iss3/539

Similar Articles

1 2 3 4 5 6 7 8 9 10 > >> 

You may also start an advanced similarity search for this article.