1. Project Overview

The problem

Digital payments make it easy to move money in seconds, and just as easy for fraudsters to move stolen money through “phantom” (mule) accounts: freshly opened accounts that receive a large transfer and pass it on immediately. By the time a bank notices, the money is gone.

Fraud is also rare. Only a small share of transactions are fraudulent, so a model that simply predicts “legit” every time can look 85-90% accurate while catching nothing. This project is built around that tension.

Objective: build a classifier that flags fraudulent transactions with high recall (catch as many real frauds as possible) while keeping precision reasonable, and show how handling class imbalance with SMOTE changes the result.

Project at a glance

Domain Financial fraud detection (transaction monitoring)
Task Binary classification: Legit vs Fraud
Models Decision Tree (rpart) and Random Forest (500 trees)
Imbalance handling SMOTE (Synthetic Minority Over-sampling Technique), applied to the training set only
Key engineered features is_emptied, dest_is_new, balance_error, the “Phantom Signature”
Evaluation Confusion matrix, Accuracy, Precision, Recall, F1, ROC curve and AUC
Tools R, tidyverse (dplyr, ggplot2), caret, rpart, randomForest, smotefamily, pROC

Workflow

Load data → Clean → Explore → Encode and scale → Engineer “Phantom Signature” features → Train/test split (80/20) → Baseline models → SMOTE → Retrain models → Evaluate (ROC/AUC, feature importance) → Compare and conclude


Packages needed (install once, then load as shown in the setup chunk):

install.packages(c("readxl", "dplyr", "rpart", "rpart.plot", "randomForest",
                   "caret", "ggplot2", "smotefamily", "pROC", "knitr"))

2. Data

df <- as.data.frame(read_excel("phantom_fraud_dataset.xlsx"))

cat("Shape:", nrow(df), "rows x", ncol(df), "columns\n")
## Shape: 50000 rows x 9 columns

The dataset contains 50,000 transactions and 9 columns.

kable(head(df, 5), caption = "First five rows of the raw data")
First five rows of the raw data
step type amount oldbalanceOrg newbalanceOrig oldbalanceDest isFraud Transaction_ID User_ID
5420 PAYMENT 39.79 93213.17 93173.38 39029.23 0 TXN_33553 USER_1834
3773 TRANSFER 1.19 75725.25 75724.06 10096.53 1 TXN_9427 USER_7875
4096 DEBIT 28.96 1588.96 1560.00 72311.72 1 TXN_199 USER_2734
8161 CASH_OUT 254.32 76807.20 76552.88 46887.76 1 TXN_12447 USER_2617
7560 PAYMENT 31.28 92354.66 92323.38 0.00 1 TXN_39489 USER_2014

Fields used in the analysis

Column Meaning
step Time step of the transaction
type Transaction type (CASH_IN, CASH_OUT, DEBIT, TRANSFER, PAYMENT)
amount Transaction amount
oldbalanceOrg / newbalanceOrig Sender’s balance before / after the transaction
oldbalanceDest Receiver’s balance before the transaction
isFraud Target label: 1 = fraud, 0 = legitimate

3. Data Cleaning

missing_tbl <- data.frame(Column = names(df), Missing = colSums(is.na(df)), row.names = NULL)
kable(missing_tbl, caption = "Missing values per column")
Missing values per column
Column Missing
step 0
type 0
amount 0
oldbalanceOrg 0
newbalanceOrig 0
oldbalanceDest 0
isFraud 0
Transaction_ID 0
User_ID 0

Total missing values: 0. The data is complete, so no imputation or row removal is needed.

4. Exploratory Analysis

4.1 The class imbalance

class_tbl <- df %>%
  count(isFraud) %>%
  mutate(Label = ifelse(isFraud == 1, "Fraud", "Legit"),
         Pct   = round(100 * n / sum(n), 1))

ggplot(class_tbl, aes(x = Label, y = n, fill = Label)) +
  geom_col(width = 0.55, show.legend = FALSE) +
  geom_text(aes(label = paste0(format(n, big.mark = ","), " (", Pct, "%)")),
            vjust = -0.5, size = 5, fontface = "bold") +
  scale_fill_manual(values = c(Legit = "steelblue", Fraud = "firebrick")) +
  scale_y_continuous(expand = expansion(mult = c(0, 0.15))) +
  labs(title = "Fraud is the minority class",
       x = NULL, y = "Number of transactions") +
  theme_minimal(base_size = 13)

Fraud makes up only 32.1% of transactions. A model that always answered “Legit” would score about 67.9% accuracy and detect zero fraud, which is why this project looks at Recall, Precision, F1 and AUC, not accuracy alone.

4.2 Where does fraud happen?

type_tbl <- df %>%
  group_by(type) %>%
  summarise(Transactions = n(),
            Frauds       = sum(isFraud),
            Fraud_Rate_Pct = round(100 * mean(isFraud), 1),
            .groups = "drop") %>%
  arrange(desc(Fraud_Rate_Pct))

kable(type_tbl, caption = "Fraud rate by transaction type")
Fraud rate by transaction type
type Transactions Frauds Fraud_Rate_Pct
CASH_OUT 12453 4046 32.5
DEBIT 12546 4031 32.1
TRANSFER 12452 3995 32.1
PAYMENT 12549 3995 31.8
ggplot(type_tbl, aes(x = reorder(type, Fraud_Rate_Pct), y = Fraud_Rate_Pct)) +
  geom_col(fill = "firebrick", width = 0.6) +
  geom_text(aes(label = paste0(Fraud_Rate_Pct, "%")), hjust = -0.15, size = 4.5) +
  coord_flip() +
  scale_y_continuous(expand = expansion(mult = c(0, 0.15))) +
  labs(title = "Fraud rate by transaction type", x = NULL, y = "Fraud rate (%)") +
  theme_minimal(base_size = 13)

The table shows which transaction types carry the most fraud. These are the types the model should focus on, and the reason type_TRANSFER and type_CASH_OUT are kept as model inputs.

4.3 Transaction amounts

ggplot(df, aes(x = factor(isFraud, levels = c(0, 1), labels = c("Legit", "Fraud")),
               y = amount + 1,
               fill = factor(isFraud, levels = c(0, 1), labels = c("Legit", "Fraud")))) +
  geom_boxplot(show.legend = FALSE, outlier.alpha = 0.3) +
  scale_y_log10(labels = scales::comma) +
  scale_fill_manual(values = c(Legit = "steelblue", Fraud = "firebrick")) +
  labs(title = "Transaction amount: Legit vs Fraud (log scale)",
       x = NULL, y = "Amount (log scale)") +
  theme_minimal(base_size = 13)

5. Preprocessing

5.1 One-hot encoding

Machine-learning models need numbers, so the categorical type column is converted into 0/1 indicator columns.

df$type_CASH_IN  <- ifelse(df$type == "CASH_IN",  1, 0)
df$type_CASH_OUT <- ifelse(df$type == "CASH_OUT", 1, 0)
df$type_DEBIT    <- ifelse(df$type == "DEBIT",    1, 0)
df$type_TRANSFER <- ifelse(df$type == "TRANSFER", 1, 0)
df$type_PAYMENT  <- ifelse(df$type == "PAYMENT",  1, 0)

kable(head(df[, c("type", "type_CASH_IN", "type_CASH_OUT", "type_TRANSFER", "type_PAYMENT")], 5),
      caption = "One-hot encoding preview")
One-hot encoding preview
type type_CASH_IN type_CASH_OUT type_TRANSFER type_PAYMENT
PAYMENT 0 0 0 1
TRANSFER 0 0 1 0
DEBIT 0 0 0 0
CASH_OUT 0 1 0 0
PAYMENT 0 0 0 1

5.2 Min-max scaling

Amounts and balances can differ by orders of magnitude. Min-max scaling brings them into the [0, 1] range so no variable dominates just because of its units.

min_max_scale <- function(x) (x - min(x)) / (max(x) - min(x))

df$amount_scaled         <- min_max_scale(df$amount)
df$oldbalanceOrg_scaled  <- min_max_scale(df$oldbalanceOrg)
df$newbalanceOrig_scaled <- min_max_scale(df$newbalanceOrig)
df$step_scaled           <- min_max_scale(df$step)

kable(data.frame(Original = head(df$amount, 5),
                 Scaled   = round(head(df$amount_scaled, 5), 4)),
      caption = "`amount` before and after scaling (first 5 rows)")
amount before and after scaling (first 5 rows)
Original Scaled
39.79 0.0339
1.19 0.0010
28.96 0.0247
254.32 0.2166
31.28 0.0266

6. Feature Engineering: The “Phantom Signature”

Raw columns do not tell the full story. Three engineered features capture how a mule-account fraud actually looks:

Feature Logic Fraud intuition
is_emptied Sender’s new balance is 0 The account is drained completely
dest_is_new Receiver’s old balance was 0 Money lands in a fresh / phantom account
balance_error Old balance - amount - new balance The books don’t reconcile, a hidden discrepancy
df$is_emptied       <- ifelse(df$newbalanceOrig == 0, 1, 0)
df$dest_is_new      <- ifelse(df$oldbalanceDest == 0, 1, 0)
df$expected_balance <- df$oldbalanceOrg - df$amount
df$balance_error    <- df$oldbalanceOrg - df$amount - df$newbalanceOrig

fraud_summary <- df %>%
  group_by(isFraud) %>%
  summarise(
    Count        = n(),
    Pct_Emptied  = round(mean(is_emptied)  * 100, 1),
    Pct_DestNew  = round(mean(dest_is_new) * 100, 1),
    Avg_BalError = round(mean(abs(balance_error)), 2),
    .groups = "drop"
  ) %>%
  mutate(Class = ifelse(isFraud == 1, "Fraud", "Legit")) %>%
  select(Class, Count, Pct_Emptied, Pct_DestNew, Avg_BalError)

kable(fraud_summary, caption = "Do the engineered features separate fraud from legit?")
Do the engineered features separate fraud from legit?
Class Count Pct_Emptied Pct_DestNew Avg_BalError
Legit 33933 0 35.0 0
Fraud 16067 0 34.8 0
sig_long <- rbind(
  data.frame(Class = fraud_summary$Class, Feature = "Account emptied (%)",
             Value = fraud_summary$Pct_Emptied),
  data.frame(Class = fraud_summary$Class, Feature = "Destination was new (%)",
             Value = fraud_summary$Pct_DestNew)
)

ggplot(sig_long, aes(x = Feature, y = Value, fill = Class)) +
  geom_col(position = position_dodge(width = 0.7), width = 0.6) +
  geom_text(aes(label = paste0(Value, "%")),
            position = position_dodge(width = 0.7), vjust = -0.5, size = 4.5) +
  scale_fill_manual(values = c(Legit = "steelblue", Fraud = "firebrick")) +
  scale_y_continuous(expand = expansion(mult = c(0, 0.15))) +
  labs(title = "The Phantom Signature: Fraud vs Legit",
       x = NULL, y = "Share of transactions (%)", fill = NULL) +
  theme_minimal(base_size = 13) +
  theme(legend.position = "top")

Reading the chart: if the Fraud bars are clearly taller than the Legit bars, the engineered features carry real signal. These are the same patterns an investigator would look for by hand, and the models below learn them automatically.

7. Model Setup

7.1 Feature set and train/test split

features <- c("type_TRANSFER", "type_CASH_OUT",
              "amount_scaled", "oldbalanceOrg_scaled", "newbalanceOrig_scaled",
              "is_emptied", "dest_is_new", "balance_error",
              "isFraud")

model_df <- df[, features]
model_df$isFraud <- factor(model_df$isFraud, levels = c(0, 1), labels = c("Legit", "Fraud"))

set.seed(42)
train_idx  <- createDataPartition(model_df$isFraud, p = 0.80, list = FALSE)
train_data <- model_df[ train_idx, ]
test_data  <- model_df[-train_idx, ]

split_tbl <- data.frame(
  Set   = c("Training (80%)", "Testing (20%)"),
  Rows  = c(nrow(train_data), nrow(test_data)),
  Legit = c(sum(train_data$isFraud == "Legit"), sum(test_data$isFraud == "Legit")),
  Fraud = c(sum(train_data$isFraud == "Fraud"), sum(test_data$isFraud == "Fraud"))
)
kable(split_tbl, caption = "Stratified split: the fraud ratio is preserved in both sets")
Stratified split: the fraud ratio is preserved in both sets
Set Rows Legit Fraud
Training (80%) 40001 27147 12854
Testing (20%) 9999 6786 3213

7.2 Helpers

Small helper functions keep the evaluation consistent across all four models.

metrics_row <- function(cm, name) {
  data.frame(
    Model     = name,
    Accuracy  = round(unname(cm$overall["Accuracy"])  * 100, 2),
    Precision = round(unname(cm$byClass["Precision"]) * 100, 2),
    Recall    = round(unname(cm$byClass["Recall"])    * 100, 2),
    F1_Score  = round(unname(cm$byClass["F1"])        * 100, 2),
    stringsAsFactors = FALSE
  )
}

plot_cm <- function(cm, title) {
  d <- as.data.frame(cm$table)
  ggplot(d, aes(x = Reference, y = Prediction, fill = Freq)) +
    geom_tile(colour = "white", linewidth = 1.2) +
    geom_text(aes(label = Freq), size = 7, fontface = "bold") +
    scale_fill_gradient(low = "#eef3fb", high = "#5a9bd5") +
    labs(title = title, x = "Actual", y = "Predicted") +
    theme_minimal(base_size = 13) +
    theme(legend.position = "none")
}

How to read a confusion matrix here: the bottom-right cell is frauds correctly caught, the top-right cell is frauds missed (the costly mistake), and the bottom-left cell is legitimate transactions wrongly flagged (the annoying mistake).

8. Baseline Models (No SMOTE)

8.1 Decision Tree

A decision tree learns simple if/then rules, which makes it easy to explain to a non-technical audience.

dt_model <- rpart(
  isFraud ~ .,
  data    = train_data,
  method  = "class",
  control = rpart.control(minsplit = 10, cp = 0.001)
)

rpart.plot(dt_model, type = 4, extra = 104,
           main = "Decision Tree: Baseline (No SMOTE)", cex = 0.7)

dt_preds <- predict(dt_model, test_data, type = "class")
dt_cm    <- confusionMatrix(dt_preds, test_data$isFraud, positive = "Fraud")
plot_cm(dt_cm, "Decision Tree (Baseline): Confusion Matrix")

kable(metrics_row(dt_cm, "Decision Tree (No SMOTE)"), caption = "Decision Tree baseline metrics (%)")
Decision Tree baseline metrics (%)
Model Accuracy Precision Recall F1_Score
Decision Tree (No SMOTE) 67.87 NA 0 NA

8.2 Random Forest

A random forest averages 500 trees, each trained on a random sample of rows and features, which usually makes it more accurate and stable than a single tree.

set.seed(42)
rf_model <- randomForest(
  isFraud ~ .,
  data       = train_data,
  ntree      = 500,
  mtry       = 3,
  importance = TRUE
)
print(rf_model)
## 
## Call:
##  randomForest(formula = isFraud ~ ., data = train_data, ntree = 500,      mtry = 3, importance = TRUE) 
##                Type of random forest: classification
##                      Number of trees: 500
## No. of variables tried at each split: 3
## 
##         OOB estimate of  error rate: 32.3%
## Confusion matrix:
##       Legit Fraud class.error
## Legit 27018   129 0.004751906
## Fraud 12792    62 0.995176599
rf_preds <- predict(rf_model, test_data)
rf_cm    <- confusionMatrix(rf_preds, test_data$isFraud, positive = "Fraud")
plot_cm(rf_cm, "Random Forest (Baseline): Confusion Matrix")

kable(metrics_row(rf_cm, "Random Forest (No SMOTE)"), caption = "Random Forest baseline metrics (%)")
Random Forest baseline metrics (%)
Model Accuracy Precision Recall F1_Score
Random Forest (No SMOTE) 67.6 28.57 0.56 1.1

9. Handling Class Imbalance with SMOTE

Because fraud is rare, the models see far more legitimate examples than fraudulent ones and tend to under-detect fraud. SMOTE fixes this by creating synthetic fraud examples: for each real fraud case it picks one of its K = 5 nearest fraud neighbours and generates a new point somewhere between them.

Important: SMOTE is applied only to the training data. The test set is left untouched so evaluation reflects the real-world imbalance and nothing leaks from synthetic data into the score.

cat("Class distribution BEFORE SMOTE:\n")
## Class distribution BEFORE SMOTE:
print(table(train_data$isFraud))
## 
## Legit Fraud 
## 27147 12854
# SMOTE needs a numeric target, so convert temporarily to 0/1
train_smote         <- train_data
train_smote$isFraud <- ifelse(train_data$isFraud == "Fraud", 1, 0)

set.seed(42)
smote_result <- SMOTE(
  X        = train_smote[, !names(train_smote) %in% "isFraud"],
  target   = train_smote$isFraud,
  K        = 5,
  dup_size = 0     # 0 = generate enough synthetic cases to balance the classes
)

train_balanced         <- smote_result$data
names(train_balanced)[names(train_balanced) == "class"] <- "isFraud"

feature_cols <- setdiff(names(train_balanced), "isFraud")
train_balanced[feature_cols] <- lapply(train_balanced[feature_cols], as.numeric)

train_balanced$isFraud <- factor(
  as.numeric(as.character(train_balanced$isFraud)),
  levels = c(0, 1),
  labels = c("Legit", "Fraud")
)

cat("\nClass distribution AFTER SMOTE:\n")
## 
## Class distribution AFTER SMOTE:
print(table(train_balanced$isFraud))
## 
## Legit Fraud 
## 27147 25708
before <- as.data.frame(table(train_data$isFraud));     before$Stage <- "Before SMOTE"
after  <- as.data.frame(table(train_balanced$isFraud)); after$Stage  <- "After SMOTE"
ba     <- rbind(before, after)
names(ba)[1:2] <- c("Class", "Count")
ba$Stage <- factor(ba$Stage, levels = c("Before SMOTE", "After SMOTE"))

ggplot(ba, aes(x = Class, y = Count, fill = Class)) +
  geom_col(width = 0.55, show.legend = FALSE) +
  geom_text(aes(label = format(Count, big.mark = ",")), vjust = -0.5, size = 4.5) +
  facet_wrap(~Stage) +
  scale_fill_manual(values = c(Legit = "steelblue", Fraud = "firebrick")) +
  scale_y_continuous(expand = expansion(mult = c(0, 0.15))) +
  labs(title = "Training set before and after SMOTE", x = NULL, y = "Rows") +
  theme_minimal(base_size = 13)

10. Models Retrained on Balanced Data

test_features <- test_data[, feature_cols]
test_features[feature_cols] <- lapply(test_features[feature_cols], as.numeric)

10.1 Decision Tree (SMOTE)

dt_model_bal <- rpart(
  isFraud ~ .,
  data    = train_balanced,
  method  = "class",
  control = rpart.control(minsplit = 10, cp = 0.001)
)

rpart.plot(dt_model_bal, type = 4, extra = 104,
           main = "Decision Tree: After SMOTE", cex = 0.7)

dt_preds_bal <- predict(dt_model_bal, test_features, type = "class")
dt_cm_bal    <- confusionMatrix(dt_preds_bal, test_data$isFraud, positive = "Fraud")
plot_cm(dt_cm_bal, "Decision Tree (SMOTE): Confusion Matrix")

kable(metrics_row(dt_cm_bal, "Decision Tree (SMOTE)"), caption = "Decision Tree (SMOTE) metrics (%)")
Decision Tree (SMOTE) metrics (%)
Model Accuracy Precision Recall F1_Score
Decision Tree (SMOTE) 67.58 36.19 1.18 2.29

10.2 Random Forest (SMOTE)

set.seed(42)
rf_model_bal <- randomForest(
  isFraud ~ .,
  data       = train_balanced,
  ntree      = 500,
  mtry       = 3,
  importance = TRUE
)
print(rf_model_bal)
## 
## Call:
##  randomForest(formula = isFraud ~ ., data = train_balanced, ntree = 500,      mtry = 3, importance = TRUE) 
##                Type of random forest: classification
##                      Number of trees: 500
## No. of variables tried at each split: 3
## 
##         OOB estimate of  error rate: 37.64%
## Confusion matrix:
##       Legit Fraud class.error
## Legit 25848  1299  0.04785059
## Fraud 18595  7113  0.72331570
rf_preds_bal <- predict(rf_model_bal, test_features)
rf_probs_bal <- predict(rf_model_bal, test_features, type = "prob")[, "Fraud"]
rf_cm_bal    <- confusionMatrix(rf_preds_bal, test_data$isFraud, positive = "Fraud")
plot_cm(rf_cm_bal, "Random Forest (SMOTE): Confusion Matrix")

kable(metrics_row(rf_cm_bal, "Random Forest (SMOTE)"), caption = "Random Forest (SMOTE) metrics (%)")
Random Forest (SMOTE) metrics (%)
Model Accuracy Precision Recall F1_Score
Random Forest (SMOTE) 66.12 31.42 4.61 8.03

The forest also outputs a fraud probability for every transaction, which is more useful in practice than a hard yes/no, because a bank can send high-probability cases to investigators first.

prob_preview <- data.frame(
  Actual            = head(test_data$isFraud, 10),
  Fraud_Probability = round(head(rf_probs_bal, 10), 4)
)
kable(prob_preview, caption = "Fraud probabilities for the first 10 test transactions")
Fraud probabilities for the first 10 test transactions
Actual Fraud_Probability
2 Fraud 0.386
14 Legit 0.162
27 Legit 0.178
35 Legit 0.100
45 Fraud 0.076
54 Legit 0.288
55 Legit 0.222
57 Fraud 0.196
63 Legit 0.268
86 Legit 0.024

11. ROC Curve and AUC

The ROC curve shows the trade-off between catching fraud (true positive rate) and false alarms (false positive rate) at every possible threshold. AUC summarises it in one number: 1.0 is perfect, 0.9+ is excellent, 0.5 is no better than a coin flip.

actual_bin <- ifelse(test_data$isFraud == "Fraud", 1, 0)

dt_probs_base <- predict(dt_model,     test_features, type = "prob")[, "Fraud"]
dt_probs_bal  <- predict(dt_model_bal, test_features, type = "prob")[, "Fraud"]
rf_probs_base <- predict(rf_model,     test_features, type = "prob")[, "Fraud"]

roc_list <- list(
  "Decision Tree (No SMOTE)" = roc(actual_bin, dt_probs_base,  quiet = TRUE),
  "Decision Tree (SMOTE)"    = roc(actual_bin, dt_probs_bal,   quiet = TRUE),
  "Random Forest (No SMOTE)" = roc(actual_bin, rf_probs_base,  quiet = TRUE),
  "Random Forest (SMOTE)"    = roc(actual_bin, rf_probs_bal,   quiet = TRUE)
)

auc_vals <- sapply(roc_list, function(r) round(as.numeric(auc(r)), 4))

ggroc(roc_list, linewidth = 1.1) +
  geom_abline(intercept = 1, slope = 1, linetype = "dashed", colour = "grey50") +
  scale_colour_manual(values = c("#7fb3d5", "#2b6cb0", "#f1948a", "#c0392b")) +
  labs(title = "ROC curves: all four models",
       subtitle = paste0("AUC: Random Forest (SMOTE) = ", auc_vals["Random Forest (SMOTE)"]),
       x = "Specificity", y = "Sensitivity (Recall)", colour = NULL) +
  theme_minimal(base_size = 13) +
  theme(legend.position = "bottom")

kable(data.frame(Model = names(auc_vals), AUC = unname(auc_vals)),
      caption = "Area under the ROC curve")
Area under the ROC curve
Model AUC
Decision Tree (No SMOTE) 0.5000
Decision Tree (SMOTE) 0.5080
Random Forest (No SMOTE) 0.5112
Random Forest (SMOTE) 0.5096

12. Feature Importance

Which signals does the final model rely on most?

imp_df         <- as.data.frame(importance(rf_model_bal))
imp_df$Feature <- rownames(imp_df)
imp_df         <- imp_df[order(-imp_df$MeanDecreaseGini), ]

ggplot(imp_df,
       aes(x = reorder(Feature, MeanDecreaseGini),
           y = MeanDecreaseGini, fill = MeanDecreaseGini)) +
  geom_col(show.legend = FALSE) +
  coord_flip() +
  scale_fill_gradient(low = "steelblue", high = "firebrick") +
  labs(title = "Random Forest: Feature Importance (SMOTE)",
       subtitle = "Higher = more important for detecting phantom fraud",
       x = NULL, y = "Mean Decrease in Gini") +
  theme_minimal(base_size = 13)

The most important feature is oldbalanceOrg_scaled, followed by newbalanceOrig_scaled and amount_scaled. Compare these with the engineered features from Section 6 to see how much the “Phantom Signature” contributes.

13. Final Model Comparison

final_comparison <- rbind(
  metrics_row(dt_cm,     "Decision Tree (No SMOTE)"),
  metrics_row(dt_cm_bal, "Decision Tree (SMOTE)"),
  metrics_row(rf_cm,     "Random Forest (No SMOTE)"),
  metrics_row(rf_cm_bal, "Random Forest (SMOTE)")
)
final_comparison$AUC <- unname(auc_vals[final_comparison$Model])

kable(final_comparison, caption = "Test-set performance (Accuracy, Precision, Recall, F1 in %)")
Test-set performance (Accuracy, Precision, Recall, F1 in %)
Model Accuracy Precision Recall F1_Score AUC
Decision Tree (No SMOTE) 67.87 NA 0.00 NA 0.5000
Decision Tree (SMOTE) 67.58 36.19 1.18 2.29 0.5080
Random Forest (No SMOTE) 67.60 28.57 0.56 1.10 0.5112
Random Forest (SMOTE) 66.12 31.42 4.61 8.03 0.5096
comp_long <- rbind(
  data.frame(Model = final_comparison$Model, Metric = "Recall",    Value = final_comparison$Recall),
  data.frame(Model = final_comparison$Model, Metric = "Precision", Value = final_comparison$Precision),
  data.frame(Model = final_comparison$Model, Metric = "F1 Score",  Value = final_comparison$F1_Score)
)
comp_long$Model <- factor(comp_long$Model, levels = final_comparison$Model)

ggplot(comp_long, aes(x = Model, y = Value, fill = Metric)) +
  geom_col(position = position_dodge(width = 0.8), width = 0.7) +
  geom_text(aes(label = round(Value, 1)),
            position = position_dodge(width = 0.8), vjust = -0.4, size = 3.5) +
  scale_fill_manual(values = c(Recall = "#c0392b", Precision = "#2b6cb0", `F1 Score` = "#7f8c8d")) +
  scale_y_continuous(limits = c(0, 110), expand = c(0, 0)) +
  labs(title = "Fraud-class performance across models", x = NULL, y = "Score (%)", fill = NULL) +
  theme_minimal(base_size = 12) +
  theme(axis.text.x = element_text(angle = 15, hjust = 1), legend.position = "top")

14. Key Findings

  • Best overall model (highest F1): Random Forest (SMOTE), with Recall = 4.61%, Precision = 31.42%, F1 = 8.03% and AUC = 0.5096.
  • Effect of SMOTE on the Random Forest: Recall moved from 0.56% to 4.61% (+4.05 points) and Precision from 28.57% to 31.42% (+2.85 points).
  • Effect of SMOTE on the Decision Tree: Recall moved from 0% to 1.18% (+1.18 points) and Precision from NA% to 36.19% (NA points).
  • Accuracy alone is misleading: with only 32.1% fraud, even a useless model scores high accuracy. Recall, F1 and AUC tell the real story.
  • Fraud has a recognisable signature: accounts that are emptied, paying into brand-new destination accounts, with balances that do not reconcile.

Business takeaway

For a fraud team, a missed fraud (false negative) usually costs far more than a false alarm (false positive), which only costs an analyst a few minutes. That is why Recall is prioritised here. In deployment, the probability score from the Random Forest would let the bank rank alerts and tune a threshold to match its investigation capacity.

15. Limitations and Next Steps

  • Single train/test split. Results come from one 80/20 split. Repeated k-fold cross-validation would give a more reliable estimate of performance.
  • Small, likely simulated data. Conclusions should not be assumed to carry over to a live production system with millions of rows and evolving fraud tactics.
  • Possible information leakage in balance_error and is_emptied. These use the sender’s balance after the transaction. If a real-time system must decide before the transaction completes, these inputs would not be available, so a pre-transaction version of the model should be tested.
  • SMOTE is not the only option. Class weights, cost-sensitive learning or threshold tuning on the probability output are worth comparing.
  • More algorithms. Gradient boosting (XGBoost / LightGBM) and logistic regression would make a stronger benchmark.
  • Explainability. SHAP values or partial-dependence plots could show how each feature moves an individual prediction.

16. Save Outputs and Reproducibility

test_out <- test_data
test_out$DT_Baseline_Pred     <- dt_preds
test_out$RF_Baseline_Pred     <- rf_preds
test_out$DT_SMOTE_Pred        <- dt_preds_bal
test_out$RF_SMOTE_Pred        <- rf_preds_bal
test_out$RF_Fraud_Probability <- round(rf_probs_bal, 4)

write.csv(df,               "fraud_data_cleaned.csv",       row.names = FALSE)
write.csv(test_out,         "fraud_predictions_final.csv",  row.names = FALSE)
write.csv(final_comparison, "model_comparison_summary.csv", row.names = FALSE)

Files written next to this document: fraud_data_cleaned.csv (full engineered dataset), fraud_predictions_final.csv (test set with all predictions) and model_comparison_summary.csv (metrics table).

sessionInfo()
## R version 4.5.1 (2025-06-13 ucrt)
## Platform: x86_64-w64-mingw32/x64
## Running under: Windows 11 x64 (build 26200)
## 
## Matrix products: default
##   LAPACK version 3.12.1
## 
## locale:
## [1] LC_COLLATE=English_India.utf8  LC_CTYPE=English_India.utf8   
## [3] LC_MONETARY=English_India.utf8 LC_NUMERIC=C                  
## [5] LC_TIME=English_India.utf8    
## 
## time zone: Asia/Calcutta
## tzcode source: internal
## 
## attached base packages:
## [1] stats     graphics  grDevices utils     datasets  methods   base     
## 
## other attached packages:
##  [1] knitr_1.51           pROC_1.19.0.1        smotefamily_1.4.0   
##  [4] caret_7.0-1          lattice_0.22-7       ggplot2_4.0.1       
##  [7] randomForest_4.7-1.2 rpart.plot_3.1.4     rpart_4.1.24        
## [10] dplyr_1.1.4          readxl_1.5.0.1      
## 
## loaded via a namespace (and not attached):
##  [1] tidyselect_1.2.1     timeDate_4051.111    farver_2.1.2        
##  [4] S7_0.2.1             fastmap_1.2.0        digest_0.6.39       
##  [7] timechange_0.3.0     lifecycle_1.0.5      survival_3.8-3      
## [10] magrittr_2.0.4       dbscan_1.2.4         compiler_4.5.1      
## [13] rlang_1.1.7          sass_0.4.10          tools_4.5.1         
## [16] igraph_2.2.2         yaml_2.3.12          data.table_1.18.2.1 
## [19] FNN_1.1.4.1          labeling_0.4.3       plyr_1.8.9          
## [22] RColorBrewer_1.1-3   withr_3.0.2          purrr_1.2.1         
## [25] nnet_7.3-20          grid_4.5.1           stats4_4.5.1        
## [28] e1071_1.7-17         future_1.70.0        globals_0.19.1      
## [31] scales_1.4.0         iterators_1.0.14     MASS_7.3-65         
## [34] cli_3.6.5            rmarkdown_2.32       generics_0.1.4      
## [37] otel_0.2.0           rstudioapi_0.18.0    future.apply_1.20.2 
## [40] reshape2_1.4.5       proxy_0.4-29         cachem_1.1.0        
## [43] stringr_1.6.0        splines_4.5.1        parallel_4.5.1      
## [46] cellranger_1.1.0     vctrs_0.7.0          hardhat_1.4.2       
## [49] Matrix_1.7-3         jsonlite_2.0.0       listenv_0.10.1      
## [52] foreach_1.5.2        gower_1.0.2          jquerylib_0.1.4     
## [55] recipes_1.3.2        glue_1.8.0           parallelly_1.47.0   
## [58] codetools_0.2-20     lubridate_1.9.4      stringi_1.8.7       
## [61] gtable_0.3.6         tibble_3.3.1         pillar_1.11.1       
## [64] htmltools_0.5.9      ipred_0.9-15         lava_1.9.0          
## [67] R6_2.6.1             evaluate_1.0.5       bslib_0.10.0        
## [70] class_7.3-23         Rcpp_1.1.1           nlme_3.1-168        
## [73] prodlim_2026.03.11   xfun_0.56            pkgconfig_2.0.3     
## [76] ModelMetrics_1.2.2.2

Author: Shreeya. Built in R with R Markdown. Random seed fixed at 42 for reproducibility.

install.packages(“rmarkdown”) install.packages(“readxl”) library(readxl)

phantom_fraud_dataset <- read_excel(“phantom_fraud_dataset.xlsx”)

View(phantom_fraud_dataset)