Analytics & predictive modelling on food products sales and marketing
For this study, an analysis of a dataset from a company in the retail food sector that sells products from 5 major categories: wines, rare meat products, exotic fruits, specially prepared fish and sweet products will be conducted. The customers can order and acquire products through 3 sales channels: physical stores, catalogs and company’s website. Then a predictive model will be built to support direct marketing initiatives and improve the company’s revenue.
The primary objective of this study is to leverage this data to build predictive models using both classification and regression techniques. Specifically, this study aim to:
Classification Modeling: Determine if a customer would accept the next marketing campaign. This binary classification task will help us understand the factors that influence the effectiveness of marketing campaigns.
Regression Modeling: Predict number of purchases for each customer at physical stores. By identifying customers who are likely to make purchases at physical stores, more effective marketing campaigns can be generated by targeting the correct customer interest.
This will be done by employing three algortihms for each of the classification and regression modelling in which the best model will be chosen from the algorithms which have the best performance metric. With this, best results can be achieved for each of the objectives.
The dataset utilized for this study is a IFood Marketing dataset taken from https://www.kaggle.com/datasets/jackdaoud/marketing-data/data. The dataset contains 2206 observations (customers) with 39 variables related to marketing data. The dataset contains both numeric and categorical variables. The numeric variables are Year_Birth, Income, Kidhome, Teenhome, Recency, MntWines, MntFruits, MntMeatProducts, MntFishproducts, MntSweetProducts, MntGoldProds, NumDealsPurchases, NumWebPurchases, NumCatalogPurchases, NumStorePurchase and NumWebVisitsMont. The categorical variabls are ID, Education, Marital_Status, Dt_Customer, AcceptedCmp1, AcceptedCmp2, AcceptedCmp3, AcceptedCmp4, AcceptedCmp5, Response, Complain and Country.
# Set a CRAN mirror
options(repos = c(CRAN = "https://cran.r-project.org"))
library(dplyr)
library(knitr)
library(summarytools)
library(tidyr)
library(VIM)
library(ggplot2)
install.packages("caret")
## package 'caret' successfully unpacked and MD5 sums checked
##
## The downloaded binary packages are in
## C:\Users\izzul\AppData\Local\Temp\RtmpgzAIZA\downloaded_packages
install.packages("glmnet")
## package 'glmnet' successfully unpacked and MD5 sums checked
##
## The downloaded binary packages are in
## C:\Users\izzul\AppData\Local\Temp\RtmpgzAIZA\downloaded_packages
install.packages("randomForest")
## package 'randomForest' successfully unpacked and MD5 sums checked
##
## The downloaded binary packages are in
## C:\Users\izzul\AppData\Local\Temp\RtmpgzAIZA\downloaded_packages
install.packages("xgboost")
## package 'xgboost' successfully unpacked and MD5 sums checked
##
## The downloaded binary packages are in
## C:\Users\izzul\AppData\Local\Temp\RtmpgzAIZA\downloaded_packages
install.packages("GGally")
## package 'GGally' successfully unpacked and MD5 sums checked
##
## The downloaded binary packages are in
## C:\Users\izzul\AppData\Local\Temp\RtmpgzAIZA\downloaded_packages
install.packages("corrplot")
## package 'corrplot' successfully unpacked and MD5 sums checked
##
## The downloaded binary packages are in
## C:\Users\izzul\AppData\Local\Temp\RtmpgzAIZA\downloaded_packages
install.packages("ggrepel")
## package 'ggrepel' successfully unpacked and MD5 sums checked
##
## The downloaded binary packages are in
## C:\Users\izzul\AppData\Local\Temp\RtmpgzAIZA\downloaded_packages
library(xgboost)
library(randomForest)
library(glmnet)
library(caret)
library(GGally)
library(corrplot)
library(ggrepel)
desc_df <- read.csv("Description.csv")
desc_df[1:29,1:2] %>%knitr::kable()
| Feature | Description | |
|---|---|---|
| 1 | ID | Customer’s Unique Identifier |
| 2 | Year_Birth | Customer’s Birth Year |
| 3 | Education | Customer’s education level |
| 4 | Marital_Status | Customer’s marital status |
| 5 | Income | Customer’s yearly household income |
| 6 | Kidhome | Number of children in customer’s household |
| 7 | Teenhome | Number of teenagers in customer’s household |
| 8 | Dt_Customer | Date of customer’s enrollment with the company |
| 9 | Recency | Number of days since customer’s last purchase |
| 10 | MntWines | Amount spent on wine in the last 2 years |
| 11 | MntFruits | Amount spent on fruits in the last 2 years |
| 12 | MntMeatProducts | Amount spent on meat in the last 2 years |
| 13 | MntFishProducts | Amount spent on fish in the last 2 years |
| 14 | MntSweetProducts | Amount spent on sweets in the last 2 years |
| 15 | MntGoldProds | Amount spent on gold in the last 2 years |
| 16 | NumDealsPurchases | Number of purchases made with a discount |
| 17 | NumWebPurchases | Number of purchases made through the company’s web site |
| 18 | NumCatalogPurchases | Number of purchases made using a catalogue |
| 19 | NumStorePurchases | Number of purchases made directly in stores |
| 20 | NumWebVisitsMonth | Number of visits to company’s web site in the last month |
| 21 | AcceptedCmp1 | 1 if customer accepted the offer in the 1st campaign, 0 otherwise (Target variable) |
| 22 | AcceptedCmp2 | 1 if customer accepted the offer in the 2nd campaign, 0 otherwise (Target variable) |
| 23 | AcceptedCmp3 | 1 if customer accepted the offer in the 3rd campaign, 0 otherwise (Target variable) |
| 24 | AcceptedCmp4 | 1 if customer accepted the offer in the 4th campaign, 0 otherwise (Target variable) |
| 25 | AcceptedCmp5 | 1 if customer accepted the offer in the 5th campaign, 0 otherwise (Target variable) |
| 26 | Response | 1 if customer accepted the offer in the last campaign, 0 otherwise (Target variable) |
| 27 | Complain | 1 if customer complained in the last 2 years, 0 otherwise |
| 28 | Country | Customer’s location |
| NA | NA | NA |
df <- read.csv("Izzul_recommend_ml_project1_data.csv")
glimpse(df)
## Rows: 2,240
## Columns: 29
## $ ID <int> 5524, 2174, 4141, 6182, 5324, 7446, 965, 6177, 485…
## $ Year_Birth <int> 1957, 1954, 1965, 1984, 1981, 1967, 1971, 1985, 19…
## $ Education <chr> "Graduation", "Graduation", "Graduation", "Graduat…
## $ Marital_Status <chr> "Single", "Single", "Together", "Together", "Marri…
## $ Income <int> 58138, 46344, 71613, 26646, 58293, 62513, 55635, 3…
## $ Kidhome <int> 0, 1, 0, 1, 1, 0, 0, 1, 1, 1, 1, 0, 0, 1, 0, 0, 1,…
## $ Teenhome <int> 0, 1, 0, 0, 0, 1, 1, 0, 0, 1, 0, 0, 0, 1, 0, 0, 1,…
## $ Dt_Customer <chr> "2012-09-04", "2014-03-08", "2013-08-21", "2014-02…
## $ Recency <int> 58, 38, 26, 26, 94, 16, 34, 32, 19, 68, 11, 59, 82…
## $ MntWines <int> 635, 11, 426, 11, 173, 520, 235, 76, 14, 28, 5, 6,…
## $ MntFruits <int> 88, 1, 49, 4, 43, 42, 65, 10, 0, 0, 5, 16, 61, 2, …
## $ MntMeatProducts <int> 546, 6, 127, 20, 118, 98, 164, 56, 24, 6, 6, 11, 4…
## $ MntFishProducts <int> 172, 2, 111, 10, 46, 0, 50, 3, 3, 1, 0, 11, 225, 3…
## $ MntSweetProducts <int> 88, 1, 21, 3, 27, 42, 49, 1, 3, 1, 2, 1, 112, 5, 1…
## $ MntGoldProds <int> 88, 6, 42, 5, 15, 14, 27, 23, 2, 13, 1, 16, 30, 14…
## $ NumDealsPurchases <int> 3, 2, 1, 2, 5, 2, 4, 2, 1, 1, 1, 1, 1, 3, 1, 1, 3,…
## $ NumWebPurchases <int> 8, 1, 8, 2, 5, 6, 7, 4, 3, 1, 1, 2, 3, 6, 1, 7, 3,…
## $ NumCatalogPurchases <int> 10, 1, 2, 0, 3, 4, 3, 0, 0, 0, 0, 0, 4, 1, 0, 6, 0…
## $ NumStorePurchases <int> 4, 2, 10, 4, 6, 10, 7, 4, 2, 0, 2, 3, 8, 5, 3, 12,…
## $ NumWebVisitsMonth <int> 7, 5, 4, 6, 5, 6, 6, 8, 9, 20, 7, 8, 2, 6, 8, 3, 8…
## $ AcceptedCmp3 <int> 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0,…
## $ AcceptedCmp4 <int> 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,…
## $ AcceptedCmp5 <int> 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0,…
## $ AcceptedCmp1 <int> 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0,…
## $ AcceptedCmp2 <int> 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,…
## $ Complain <int> 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,…
## $ Z_CostContact <int> 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3,…
## $ Z_Revenue <int> 11, 11, 11, 11, 11, 11, 11, 11, 11, 11, 11, 11, 11…
## $ Response <int> 1, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 1, 0,…
summary(df)
## ID Year_Birth Education Marital_Status
## Min. : 0 Min. :1893 Length:2240 Length:2240
## 1st Qu.: 2828 1st Qu.:1959 Class :character Class :character
## Median : 5458 Median :1970 Mode :character Mode :character
## Mean : 5592 Mean :1969
## 3rd Qu.: 8428 3rd Qu.:1977
## Max. :11191 Max. :1996
##
## Income Kidhome Teenhome Dt_Customer
## Min. : 1730 Min. :0.0000 Min. :0.0000 Length:2240
## 1st Qu.: 35303 1st Qu.:0.0000 1st Qu.:0.0000 Class :character
## Median : 51382 Median :0.0000 Median :0.0000 Mode :character
## Mean : 52247 Mean :0.4442 Mean :0.5062
## 3rd Qu.: 68522 3rd Qu.:1.0000 3rd Qu.:1.0000
## Max. :666666 Max. :2.0000 Max. :2.0000
## NA's :24
## Recency MntWines MntFruits MntMeatProducts
## Min. : 0.00 Min. : 0.00 Min. : 0.0 Min. : 0.0
## 1st Qu.:24.00 1st Qu.: 23.75 1st Qu.: 1.0 1st Qu.: 16.0
## Median :49.00 Median : 173.50 Median : 8.0 Median : 67.0
## Mean :49.11 Mean : 303.94 Mean : 26.3 Mean : 166.9
## 3rd Qu.:74.00 3rd Qu.: 504.25 3rd Qu.: 33.0 3rd Qu.: 232.0
## Max. :99.00 Max. :1493.00 Max. :199.0 Max. :1725.0
##
## MntFishProducts MntSweetProducts MntGoldProds NumDealsPurchases
## Min. : 0.00 Min. : 0.00 Min. : 0.00 Min. : 0.000
## 1st Qu.: 3.00 1st Qu.: 1.00 1st Qu.: 9.00 1st Qu.: 1.000
## Median : 12.00 Median : 8.00 Median : 24.00 Median : 2.000
## Mean : 37.53 Mean : 27.06 Mean : 44.02 Mean : 2.325
## 3rd Qu.: 50.00 3rd Qu.: 33.00 3rd Qu.: 56.00 3rd Qu.: 3.000
## Max. :259.00 Max. :263.00 Max. :362.00 Max. :15.000
##
## NumWebPurchases NumCatalogPurchases NumStorePurchases NumWebVisitsMonth
## Min. : 0.000 Min. : 0.000 Min. : 0.00 Min. : 0.000
## 1st Qu.: 2.000 1st Qu.: 0.000 1st Qu.: 3.00 1st Qu.: 3.000
## Median : 4.000 Median : 2.000 Median : 5.00 Median : 6.000
## Mean : 4.085 Mean : 2.662 Mean : 5.79 Mean : 5.317
## 3rd Qu.: 6.000 3rd Qu.: 4.000 3rd Qu.: 8.00 3rd Qu.: 7.000
## Max. :27.000 Max. :28.000 Max. :13.00 Max. :20.000
##
## AcceptedCmp3 AcceptedCmp4 AcceptedCmp5 AcceptedCmp1
## Min. :0.00000 Min. :0.00000 Min. :0.00000 Min. :0.00000
## 1st Qu.:0.00000 1st Qu.:0.00000 1st Qu.:0.00000 1st Qu.:0.00000
## Median :0.00000 Median :0.00000 Median :0.00000 Median :0.00000
## Mean :0.07277 Mean :0.07455 Mean :0.07277 Mean :0.06429
## 3rd Qu.:0.00000 3rd Qu.:0.00000 3rd Qu.:0.00000 3rd Qu.:0.00000
## Max. :1.00000 Max. :1.00000 Max. :1.00000 Max. :1.00000
##
## AcceptedCmp2 Complain Z_CostContact Z_Revenue
## Min. :0.00000 Min. :0.000000 Min. :3 Min. :11
## 1st Qu.:0.00000 1st Qu.:0.000000 1st Qu.:3 1st Qu.:11
## Median :0.00000 Median :0.000000 Median :3 Median :11
## Mean :0.01339 Mean :0.009375 Mean :3 Mean :11
## 3rd Qu.:0.00000 3rd Qu.:0.000000 3rd Qu.:3 3rd Qu.:11
## Max. :1.00000 Max. :1.000000 Max. :3 Max. :11
##
## Response
## Min. :0.0000
## 1st Qu.:0.0000
## Median :0.0000
## Mean :0.1491
## 3rd Qu.:0.0000
## Max. :1.0000
##
print(dfSummary(df,
varnumbers = FALSE,
valid.col = FALSE,
graph.magnif = 0.76),
method = 'render')
| Variable | Stats / Values | Freqs (% of Valid) | Graph | Missing | |||||||||||||||||||||||||||||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ID [integer] |
|
2240 distinct values | 0 (0.0%) | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Year_Birth [integer] |
|
59 distinct values | 0 (0.0%) | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Education [character] |
|
|
0 (0.0%) | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Marital_Status [character] |
|
|
0 (0.0%) | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Income [integer] |
|
1974 distinct values | 24 (1.1%) | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Kidhome [integer] |
|
|
0 (0.0%) | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Teenhome [integer] |
|
|
0 (0.0%) | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Dt_Customer [character] |
|
|
0 (0.0%) | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Recency [integer] |
|
100 distinct values | 0 (0.0%) | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| MntWines [integer] |
|
776 distinct values | 0 (0.0%) | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| MntFruits [integer] |
|
158 distinct values | 0 (0.0%) | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| MntMeatProducts [integer] |
|
558 distinct values | 0 (0.0%) | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| MntFishProducts [integer] |
|
182 distinct values | 0 (0.0%) | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| MntSweetProducts [integer] |
|
177 distinct values | 0 (0.0%) | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| MntGoldProds [integer] |
|
213 distinct values | 0 (0.0%) | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| NumDealsPurchases [integer] |
|
15 distinct values | 0 (0.0%) | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| NumWebPurchases [integer] |
|
15 distinct values | 0 (0.0%) | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| NumCatalogPurchases [integer] |
|
14 distinct values | 0 (0.0%) | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| NumStorePurchases [integer] |
|
14 distinct values | 0 (0.0%) | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| NumWebVisitsMonth [integer] |
|
16 distinct values | 0 (0.0%) | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| AcceptedCmp3 [integer] |
|
|
0 (0.0%) | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| AcceptedCmp4 [integer] |
|
|
0 (0.0%) | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| AcceptedCmp5 [integer] |
|
|
0 (0.0%) | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| AcceptedCmp1 [integer] |
|
|
0 (0.0%) | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| AcceptedCmp2 [integer] |
|
|
0 (0.0%) | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Complain [integer] |
|
|
0 (0.0%) | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Z_CostContact [integer] | 1 distinct value |
|
0 (0.0%) | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Z_Revenue [integer] | 1 distinct value |
|
0 (0.0%) | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Response [integer] |
|
|
0 (0.0%) |
Generated by summarytools 1.0.1 (R version 4.3.1)
2024-06-06
# Display the total missing values for each variable
data.frame("variable"=c(colnames(df)),
"missing values count"=sapply(df, function(x) sum(is.na(x))),
row.names=NULL)
## variable missing.values.count
## 1 ID 0
## 2 Year_Birth 0
## 3 Education 0
## 4 Marital_Status 0
## 5 Income 24
## 6 Kidhome 0
## 7 Teenhome 0
## 8 Dt_Customer 0
## 9 Recency 0
## 10 MntWines 0
## 11 MntFruits 0
## 12 MntMeatProducts 0
## 13 MntFishProducts 0
## 14 MntSweetProducts 0
## 15 MntGoldProds 0
## 16 NumDealsPurchases 0
## 17 NumWebPurchases 0
## 18 NumCatalogPurchases 0
## 19 NumStorePurchases 0
## 20 NumWebVisitsMonth 0
## 21 AcceptedCmp3 0
## 22 AcceptedCmp4 0
## 23 AcceptedCmp5 0
## 24 AcceptedCmp1 0
## 25 AcceptedCmp2 0
## 26 Complain 0
## 27 Z_CostContact 0
## 28 Z_Revenue 0
## 29 Response 0
df <- kNN(df, variable = "Income", k = 5, imp_var = FALSE)
# Display the total missing values for each variable
data.frame("variable"=c(colnames(df)),
"missing values count"=sapply(df, function(x) sum(is.na(x))),
row.names=NULL)
## variable missing.values.count
## 1 ID 0
## 2 Year_Birth 0
## 3 Education 0
## 4 Marital_Status 0
## 5 Income 0
## 6 Kidhome 0
## 7 Teenhome 0
## 8 Dt_Customer 0
## 9 Recency 0
## 10 MntWines 0
## 11 MntFruits 0
## 12 MntMeatProducts 0
## 13 MntFishProducts 0
## 14 MntSweetProducts 0
## 15 MntGoldProds 0
## 16 NumDealsPurchases 0
## 17 NumWebPurchases 0
## 18 NumCatalogPurchases 0
## 19 NumStorePurchases 0
## 20 NumWebVisitsMonth 0
## 21 AcceptedCmp3 0
## 22 AcceptedCmp4 0
## 23 AcceptedCmp5 0
## 24 AcceptedCmp1 0
## 25 AcceptedCmp2 0
## 26 Complain 0
## 27 Z_CostContact 0
## 28 Z_Revenue 0
## 29 Response 0
df <- df %>%select(-ID)
df <- df %>%mutate(Dt_Customer = as.Date(Dt_Customer, format = "%Y-%m-%d"))
df <- df %>%mutate(Income = as.numeric(Income))
custom_encoding <- c( "Divorced" = 1, "Single" = 1, "Married" = 2, "Together" = 2, "Widow" = 1, "YOLO" = 1, "Alone" = 1, "Absurd" = 1)
df <- df %>%mutate(Marital_Status = custom_encoding[Marital_Status])
# Create a function to plot boxplots for a specified number of numeric variables excluding specified variables
plot_boxplots_exclude <- function(data, exclude_vars = c(), num_per_page = 4) {
# Get numeric variables excluding the ones to be excluded
numeric_vars <- names(data)[sapply(data, is.numeric) & !names(data) %in% exclude_vars]
# Calculate the number of pages needed
num_plots <- length(numeric_vars)
num_pages <- ceiling(num_plots / num_per_page)
# Loop through each page
for (page in 1:num_pages) {
# Determine the variables to plot on this page
start_index <- (page - 1) * num_per_page + 1
end_index <- min(page * num_per_page, num_plots)
vars_to_plot <- numeric_vars[start_index:end_index]
# Set up the plotting layout with adjusted margins
par(mfrow = c(ceiling(length(vars_to_plot) / 2), 2), mar = c(4, 4, 2, 1))
# Loop through each numeric variable on this page and create a boxplot
for (col in vars_to_plot) {
boxplot(data[[col]], main = col)
}
# Reset plotting layout
par(mfrow = c(1, 1))
}
}
# Usage example excluding specific variables, with 4 variables per page
exclude_vars <- c('AcceptedCmp1', 'AcceptedCmp2', 'AcceptedCmp3', 'AcceptedCmp4',
'AcceptedCmp5', 'Response', 'Complain', 'Education', 'Country')
plot_boxplots_exclude(df, exclude_vars, num_per_page = 4)
# Create a function to clean specific variables using IQR method
clean_specific_variables <- function(data, vars_to_clean) {
cleaned_data <- data # Create a copy of the dataset
# Loop through each variable to clean
for (col in vars_to_clean) {
# Calculate interquartile range (IQR)
Q1 <- quantile(data[[col]], 0.25, na.rm = TRUE)
Q3 <- quantile(data[[col]], 0.75, na.rm = TRUE)
IQR <- Q3 - Q1
# Define outlier thresholds
lower_threshold <- Q1 - 1.5 * IQR
upper_threshold <- Q3 + 1.5 * IQR
# Identify outliers
outliers <- data[[col]] < lower_threshold | data[[col]] > upper_threshold
# Replace outliers with NA
cleaned_data[[col]][outliers] <- NA
# Replace NA values with median of the variable
cleaned_data[[col]] <- ifelse(is.na(cleaned_data[[col]]), median(data[[col]], na.rm = TRUE), cleaned_data[[col]])
}
return(cleaned_data)
}
# Target variables to clean using IQR
vars_to_clean <- c("Year_Birth")
# Clean Outliner variables using IQR
clean_data <- clean_specific_variables(df, vars_to_clean)
# 1. Age variable instead of year of birth
# 2. Customer_Days variable instead of year of birth
# 3. FamilyMember variable the total number of family members
# 4. TotalPurchases variable Total places of purchase
# 5. Spending variable the sum of the amount spent on product categories
# 6. AcceptedCmpOverall variable Total amount spent on all the products
#Age of customer
clean_data$Age <- as.integer(format(Sys.Date(), "%Y")) - clean_data$Year_Birth
#number of days since registration as a customer
clean_data$Customer_Days <- as.integer(Sys.Date() - clean_data$Dt_Customer)
# Create the new variable Number of Family Member
clean_data$FamilyMember <- clean_data$Marital_Status + clean_data$Kidhome + clean_data$Teenhome
# Create the new variable TotalPurchases
clean_data$TotalPurchases <- clean_data$NumWebPurchases + clean_data$NumCatalogPurchases + clean_data$NumStorePurchases
#Total amount spent on all the products
clean_data$Spending <- rowSums(clean_data[, grep("^Mnt", names(clean_data))])
#Total overall number of accepted campaigns
clean_data$AcceptedCmpOverall <- rowSums(clean_data[, grep("^Accept", names(clean_data))])
# remove columns with minimal usability in the machine learning modelling
clean_data <- subset(clean_data, select = -c(Year_Birth, Dt_Customer, Z_CostContact, Z_Revenue))
# 6 new variables; removed 3 variables
glimpse(clean_data)
## Rows: 2,240
## Columns: 30
## $ Education <chr> "Graduation", "Graduation", "Graduation", "Graduat…
## $ Marital_Status <dbl> 1, 1, 2, 2, 2, 2, 1, 2, 2, 2, 2, 2, 1, 1, 2, 1, 2,…
## $ Income <dbl> 58138, 46344, 71613, 26646, 58293, 62513, 55635, 3…
## $ Kidhome <int> 0, 1, 0, 1, 1, 0, 0, 1, 1, 1, 1, 0, 0, 1, 0, 0, 1,…
## $ Teenhome <int> 0, 1, 0, 0, 0, 1, 1, 0, 0, 1, 0, 0, 0, 1, 0, 0, 1,…
## $ Recency <int> 58, 38, 26, 26, 94, 16, 34, 32, 19, 68, 11, 59, 82…
## $ MntWines <int> 635, 11, 426, 11, 173, 520, 235, 76, 14, 28, 5, 6,…
## $ MntFruits <int> 88, 1, 49, 4, 43, 42, 65, 10, 0, 0, 5, 16, 61, 2, …
## $ MntMeatProducts <int> 546, 6, 127, 20, 118, 98, 164, 56, 24, 6, 6, 11, 4…
## $ MntFishProducts <int> 172, 2, 111, 10, 46, 0, 50, 3, 3, 1, 0, 11, 225, 3…
## $ MntSweetProducts <int> 88, 1, 21, 3, 27, 42, 49, 1, 3, 1, 2, 1, 112, 5, 1…
## $ MntGoldProds <int> 88, 6, 42, 5, 15, 14, 27, 23, 2, 13, 1, 16, 30, 14…
## $ NumDealsPurchases <int> 3, 2, 1, 2, 5, 2, 4, 2, 1, 1, 1, 1, 1, 3, 1, 1, 3,…
## $ NumWebPurchases <int> 8, 1, 8, 2, 5, 6, 7, 4, 3, 1, 1, 2, 3, 6, 1, 7, 3,…
## $ NumCatalogPurchases <int> 10, 1, 2, 0, 3, 4, 3, 0, 0, 0, 0, 0, 4, 1, 0, 6, 0…
## $ NumStorePurchases <int> 4, 2, 10, 4, 6, 10, 7, 4, 2, 0, 2, 3, 8, 5, 3, 12,…
## $ NumWebVisitsMonth <int> 7, 5, 4, 6, 5, 6, 6, 8, 9, 20, 7, 8, 2, 6, 8, 3, 8…
## $ AcceptedCmp3 <int> 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0,…
## $ AcceptedCmp4 <int> 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,…
## $ AcceptedCmp5 <int> 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0,…
## $ AcceptedCmp1 <int> 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0,…
## $ AcceptedCmp2 <int> 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,…
## $ Complain <int> 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,…
## $ Response <int> 1, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 1, 0,…
## $ Age <dbl> 67, 70, 59, 40, 43, 57, 53, 39, 50, 74, 41, 48, 65…
## $ Customer_Days <int> 4293, 3743, 3942, 3769, 3791, 3923, 4223, 4047, 40…
## $ FamilyMember <dbl> 1, 3, 2, 3, 3, 3, 2, 3, 3, 4, 3, 2, 1, 3, 2, 1, 4,…
## $ TotalPurchases <int> 22, 4, 20, 6, 14, 20, 17, 8, 5, 1, 3, 5, 15, 12, 4…
## $ Spending <dbl> 1617, 27, 776, 53, 422, 716, 590, 169, 46, 49, 19,…
## $ AcceptedCmpOverall <dbl> 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 2, 0,…
# cast Response column as factor for classification modelling
clean_data$Response <- as.factor(clean_data$Response)
clean_data$Education <- as.factor(clean_data$Education)
clean_data$Marital_Status <- as.factor(clean_data$Marital_Status)
print(dfSummary(clean_data,
varnumbers = FALSE,
valid.col = FALSE,
graph.magnif = 0.76),
method = 'render')
| Variable | Stats / Values | Freqs (% of Valid) | Graph | Missing | ||||||||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Education [factor] |
|
|
0 (0.0%) | |||||||||||||||||||||||||||||||||||
| Marital_Status [factor] |
|
|
0 (0.0%) | |||||||||||||||||||||||||||||||||||
| Income [numeric] |
|
1974 distinct values | 0 (0.0%) | |||||||||||||||||||||||||||||||||||
| Kidhome [integer] |
|
|
0 (0.0%) | |||||||||||||||||||||||||||||||||||
| Teenhome [integer] |
|
|
0 (0.0%) | |||||||||||||||||||||||||||||||||||
| Recency [integer] |
|
100 distinct values | 0 (0.0%) | |||||||||||||||||||||||||||||||||||
| MntWines [integer] |
|
776 distinct values | 0 (0.0%) | |||||||||||||||||||||||||||||||||||
| MntFruits [integer] |
|
158 distinct values | 0 (0.0%) | |||||||||||||||||||||||||||||||||||
| MntMeatProducts [integer] |
|
558 distinct values | 0 (0.0%) | |||||||||||||||||||||||||||||||||||
| MntFishProducts [integer] |
|
182 distinct values | 0 (0.0%) | |||||||||||||||||||||||||||||||||||
| MntSweetProducts [integer] |
|
177 distinct values | 0 (0.0%) | |||||||||||||||||||||||||||||||||||
| MntGoldProds [integer] |
|
213 distinct values | 0 (0.0%) | |||||||||||||||||||||||||||||||||||
| NumDealsPurchases [integer] |
|
15 distinct values | 0 (0.0%) | |||||||||||||||||||||||||||||||||||
| NumWebPurchases [integer] |
|
15 distinct values | 0 (0.0%) | |||||||||||||||||||||||||||||||||||
| NumCatalogPurchases [integer] |
|
14 distinct values | 0 (0.0%) | |||||||||||||||||||||||||||||||||||
| NumStorePurchases [integer] |
|
14 distinct values | 0 (0.0%) | |||||||||||||||||||||||||||||||||||
| NumWebVisitsMonth [integer] |
|
16 distinct values | 0 (0.0%) | |||||||||||||||||||||||||||||||||||
| AcceptedCmp3 [integer] |
|
|
0 (0.0%) | |||||||||||||||||||||||||||||||||||
| AcceptedCmp4 [integer] |
|
|
0 (0.0%) | |||||||||||||||||||||||||||||||||||
| AcceptedCmp5 [integer] |
|
|
0 (0.0%) | |||||||||||||||||||||||||||||||||||
| AcceptedCmp1 [integer] |
|
|
0 (0.0%) | |||||||||||||||||||||||||||||||||||
| AcceptedCmp2 [integer] |
|
|
0 (0.0%) | |||||||||||||||||||||||||||||||||||
| Complain [integer] |
|
|
0 (0.0%) | |||||||||||||||||||||||||||||||||||
| Response [factor] |
|
|
0 (0.0%) | |||||||||||||||||||||||||||||||||||
| Age [numeric] |
|
56 distinct values | 0 (0.0%) | |||||||||||||||||||||||||||||||||||
| Customer_Days [integer] |
|
663 distinct values | 0 (0.0%) | |||||||||||||||||||||||||||||||||||
| FamilyMember [numeric] |
|
|
0 (0.0%) | |||||||||||||||||||||||||||||||||||
| TotalPurchases [integer] |
|
33 distinct values | 0 (0.0%) | |||||||||||||||||||||||||||||||||||
| Spending [numeric] |
|
1054 distinct values | 0 (0.0%) | |||||||||||||||||||||||||||||||||||
| AcceptedCmpOverall [numeric] |
|
|
0 (0.0%) |
Generated by summarytools 1.0.1 (R version 4.3.1)
2024-06-06
For univariate analysis, we normally look at the distribution of each variable. This can be done using summary statistics and visualizations like histograms and boxplots.Univariate analysis is primarily used to summarize and find patterns in the data for a single variable. It can reveal insights such as the distribution, central tendency, variability, and presence of outliers in the data.
# Plot histograms for all numeric variables
numeric_cols <- sapply(clean_data, is.numeric)
data_numeric <- clean_data[, numeric_cols]
# Histograms
for (col in names(data_numeric)) {
print(ggplot(clean_data, aes_string(x = col)) +
geom_histogram(bins = 30, fill = 'blue', alpha = 0.7) +
labs(title = paste('Histogram of', col), x = col, y = 'Count') +
theme_minimal())
}
# Boxplots for numeric variables
for (col in names(data_numeric)) {
print(ggplot(clean_data, aes_string(y = col)) +
geom_boxplot(fill = 'blue', alpha = 0.7) +
labs(title = paste('Boxplot of', col), y = col) +
theme_minimal())
}
For bivariate analysis, we explore the correlation between two variables using scatter plots, correlation matrices, and box plots for categorical vs. numerical data. Bivariate analysis extends univariate analysis by examining the correlation, association, or causal relationships between two variables, rather than focusing on a single variable.
# 'Response' is a factor or a binary numeric variable
numeric_vars <- sapply(clean_data, is.numeric)
data_numeric <- clean_data[, numeric_vars]
# Automatically generate plots for each numeric variable
plots_list <- lapply(names(data_numeric), function(var) {
ggplot(clean_data, aes_string(x = "Response", y = var)) +
geom_boxplot(alpha = 0.5) +
labs(title = paste("Response vs", var),
x = "Response",
y = var) +
theme_minimal()
})
for (i in 1:length(data_numeric)) {
print(plots_list[[i]])
}
# Correlation matrix
cor_matrix <- cor(data_numeric, use = "complete.obs")
# Visualize the correlation matrix
library(corrplot)
corrplot(cor_matrix, method = "circle")
Multivariate analysis is a statistical technique used to explore patterns and understand relationships among multiple variables simultaneously. It is crucial for analyzing datasets with multiple outcomes or predictors, providing insights that cannot be obtained through univariate or bivariate analysis alone. Techniques such as pair plots, Principal Component Analysis (PCA), and others are employed in multivariate analysis to investigate the complex interactions between variables.
# Principal Component Analysis (PCA)
pca_result <- prcomp(data_numeric, scale. = TRUE)
# Coefficients for each principal componentAccessing the loadings
loadings <- pca_result$rotation
# Explained variance
scores <- as.data.frame(pca_result$x)
explained_variance <- pca_result$sdev^2
total_variance <- sum(explained_variance)
explained_variance_percent <- explained_variance / total_variance * 100
variance_df <- data.frame(Principal_Component = 1:length(explained_variance),
Variance = explained_variance,
Percent_Variance = explained_variance_percent)
# Convert these to data frames for easier viewing/manipulation
scores_df <- as.data.frame(scores)
loadings_df <- as.data.frame(loadings)
variance_df <- as.data.frame(variance_df)
knitr::kable(loadings_df, caption = "Loadings from PCA")
| PC1 | PC2 | PC3 | PC4 | PC5 | PC6 | PC7 | PC8 | PC9 | PC10 | PC11 | PC12 | PC13 | PC14 | PC15 | PC16 | PC17 | PC18 | PC19 | PC20 | PC21 | PC22 | PC23 | PC24 | PC25 | PC26 | PC27 | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Income | -0.2474238 | 0.0510403 | -0.0033790 | 0.2117790 | -0.0962527 | 0.1207169 | -0.0593574 | -0.0632547 | 0.0611844 | 0.0971675 | -0.1854461 | -0.0861370 | -0.0231676 | -0.0859476 | -0.1247436 | 0.2430403 | 0.3826866 | -0.0551579 | -0.6093311 | -0.3391945 | 0.2163488 | 0.1295045 | -0.1471943 | 0.0364914 | 0.0000000 | 0.0000000 | 0.0000000 |
| Kidhome | 0.2234915 | -0.0423642 | -0.0826165 | -0.1915886 | -0.2363109 | 0.4230030 | -0.1601178 | -0.0736641 | 0.0431119 | -0.0284936 | -0.0870678 | 0.0400124 | -0.0156810 | 0.0069876 | 0.3226397 | 0.3909151 | -0.0027144 | 0.2085642 | -0.0519891 | 0.1250196 | -0.0241753 | 0.0228068 | 0.0798953 | -0.5481803 | 0.0000000 | 0.0000000 | 0.0000000 |
| Teenhome | 0.0538909 | 0.3880974 | -0.2836427 | 0.3085219 | -0.1915492 | -0.0198456 | 0.0378686 | 0.0252252 | -0.0538301 | -0.1297632 | 0.0178129 | -0.2257228 | 0.0479937 | -0.0407911 | -0.3797323 | -0.2106846 | 0.0058301 | -0.2213995 | 0.1662052 | -0.0455061 | 0.0274238 | -0.0453872 | -0.0477707 | -0.5331839 | 0.0000000 | 0.0000000 | 0.0000000 |
| Recency | -0.0036527 | 0.0198565 | 0.0031635 | 0.0292237 | 0.1166838 | 0.4461439 | 0.4849463 | 0.7166148 | 0.1022060 | 0.0723574 | 0.1070610 | 0.0425847 | 0.0264951 | -0.0551839 | -0.0046644 | -0.0060868 | -0.0264051 | -0.0216052 | -0.0401302 | -0.0406435 | -0.0023489 | -0.0003103 | -0.0086872 | 0.0094117 | 0.0000000 | 0.0000000 | 0.0000000 |
| MntWines | -0.2768113 | 0.0475687 | -0.1997572 | -0.0002879 | 0.1133021 | -0.0082339 | -0.0024376 | -0.0069221 | 0.0145631 | 0.1853988 | -0.1088318 | -0.0771569 | -0.0099701 | -0.0990420 | 0.0668597 | 0.1758664 | -0.1495899 | -0.1294096 | 0.2141323 | 0.0359911 | 0.4490774 | 0.1471007 | 0.4989278 | 0.0255073 | 0.3574596 | 0.0422282 | -0.2860754 |
| MntFruits | -0.2242979 | 0.0328905 | 0.2042765 | -0.0926419 | -0.1175893 | 0.1141482 | -0.0388986 | -0.0408341 | -0.0384957 | -0.2854561 | 0.1166721 | -0.1491460 | 0.2967797 | 0.0893744 | 0.1063613 | -0.0306865 | -0.3976360 | -0.4895254 | -0.3947937 | 0.2290737 | -0.1418625 | 0.0824998 | 0.0715942 | 0.0045158 | 0.0422386 | 0.0049898 | -0.0338036 |
| MntMeatProducts | -0.2704691 | -0.0418625 | 0.1272996 | -0.0367045 | -0.0386954 | 0.1689676 | -0.0286397 | -0.0170271 | -0.0804237 | 0.1069907 | -0.3393575 | 0.1435522 | -0.0160640 | 0.1064598 | 0.0301504 | -0.0050955 | 0.0963505 | -0.2181125 | 0.1787384 | 0.1292784 | 0.0080031 | -0.6031031 | -0.3798367 | -0.0389868 | 0.2397051 | 0.0283174 | -0.1918363 |
| MntFishProducts | -0.2333203 | 0.0224587 | 0.2050504 | -0.0871227 | -0.1082586 | 0.1094124 | -0.0433039 | -0.0110680 | -0.1176603 | -0.3097960 | 0.1115610 | 0.0990328 | 0.1623428 | 0.1533030 | -0.0157039 | 0.1428898 | -0.1024382 | 0.0933840 | 0.3304513 | -0.7217191 | -0.0949799 | 0.0476926 | 0.0272983 | -0.0155074 | 0.0580149 | 0.0068535 | -0.0464294 |
| MntSweetProducts | -0.2264169 | 0.0262186 | 0.1672398 | -0.0869353 | -0.1052916 | 0.1510988 | -0.0513058 | -0.0034338 | -0.0798910 | -0.2769654 | 0.2066254 | -0.1171019 | 0.2730590 | -0.1636144 | 0.1363912 | -0.3717832 | 0.4760808 | 0.2870570 | 0.0990230 | 0.2588484 | 0.2904431 | 0.0694753 | 0.0046255 | 0.0075336 | 0.0438390 | 0.0051789 | -0.0350845 |
| MntGoldProds | -0.1903065 | 0.1280807 | 0.0191564 | -0.2100637 | -0.1467101 | -0.1595949 | 0.0633211 | 0.1391735 | 0.0499653 | -0.2354976 | 0.2565309 | -0.2127333 | -0.6327050 | 0.4121770 | -0.0512170 | 0.1594286 | 0.0260871 | 0.0883096 | -0.0529456 | 0.1361586 | 0.1448304 | -0.0235420 | -0.0938911 | 0.0012293 | 0.0554008 | 0.0065447 | -0.0443373 |
| NumDealsPurchases | 0.0426387 | 0.3769631 | -0.2617530 | -0.2148019 | -0.0692239 | 0.1779661 | -0.1330049 | -0.0381924 | 0.0996303 | 0.0742314 | -0.1041213 | 0.3445281 | -0.0659482 | 0.2987329 | 0.1368582 | -0.5151350 | -0.0745514 | 0.0675630 | -0.2449732 | -0.1822997 | 0.0790348 | -0.0899422 | 0.1849939 | 0.1059469 | 0.0000000 | 0.0000000 | 0.0000000 |
| NumWebPurchases | -0.1909665 | 0.2883206 | -0.1600416 | -0.1523009 | 0.0461628 | -0.1428420 | -0.0220804 | -0.0128536 | 0.1385948 | 0.1524247 | 0.3777742 | 0.0115522 | -0.0064218 | -0.2735923 | 0.2457810 | 0.1548419 | 0.3077669 | -0.1871995 | 0.0110317 | -0.0653714 | -0.3999869 | -0.2268123 | 0.1418735 | -0.0089651 | 0.0121873 | -0.3113610 | -0.0307323 |
| NumCatalogPurchases | -0.2780998 | 0.0455758 | 0.0288867 | -0.0076440 | -0.0959634 | 0.0252347 | 0.0362362 | 0.0301886 | 0.0237436 | 0.1043032 | -0.3069157 | 0.2344031 | -0.0693895 | 0.1700393 | -0.0644675 | -0.0598746 | 0.1348800 | -0.0543128 | 0.1953056 | 0.2043220 | -0.3090514 | 0.6213083 | -0.1066718 | -0.0558955 | 0.0128206 | -0.3275398 | -0.0323292 |
| NumStorePurchases | -0.2507971 | 0.2006405 | -0.0161438 | 0.0587103 | 0.1113048 | -0.0003747 | -0.0495889 | -0.0695334 | 0.2435670 | 0.0597094 | 0.0630121 | -0.0247002 | 0.1146149 | -0.1957180 | -0.0814262 | 0.0394501 | -0.4961738 | 0.4755377 | -0.0755546 | 0.0635467 | 0.1177376 | -0.0629106 | -0.3334400 | -0.0329746 | 0.0142585 | -0.3642770 | -0.0359552 |
| NumWebVisitsMonth | 0.2036132 | 0.0931622 | -0.2372912 | -0.3824040 | 0.1477998 | -0.0769292 | -0.0016237 | 0.0079741 | -0.1207554 | -0.0049792 | 0.1317588 | 0.0872254 | 0.1038293 | -0.0881924 | 0.1750260 | 0.0675043 | -0.0005463 | -0.3252006 | 0.0894338 | -0.1000730 | 0.3511411 | 0.2710048 | -0.5528394 | 0.0049959 | 0.0000000 | 0.0000000 | 0.0000000 |
| AcceptedCmp3 | -0.0236862 | -0.1442482 | -0.1432878 | -0.2682694 | -0.5818633 | -0.4083800 | 0.2373163 | 0.1887612 | 0.1120028 | 0.1429383 | -0.1543051 | -0.0617243 | 0.3313161 | 0.0364380 | -0.0159329 | 0.0278554 | -0.0173936 | 0.1087654 | -0.0487479 | -0.0364580 | 0.0359893 | -0.0833048 | 0.0167482 | 0.0158872 | -0.1898965 | 0.0157579 | -0.2349552 |
| AcceptedCmp4 | -0.0943480 | -0.1523833 | -0.3807425 | 0.1350660 | 0.3587343 | 0.0900348 | -0.0698468 | -0.0506669 | 0.0373027 | 0.0357090 | 0.1733508 | -0.1350709 | 0.3555736 | 0.5736249 | -0.0374135 | 0.1248157 | 0.1424796 | 0.0867930 | 0.0014549 | 0.0785563 | -0.0788801 | -0.0305300 | -0.0476358 | -0.0012981 | -0.1920272 | 0.0159347 | -0.2375915 |
| AcceptedCmp5 | -0.1706010 | -0.2946844 | -0.2050694 | 0.0193514 | -0.0281177 | 0.1656063 | -0.0214075 | -0.0503540 | -0.1789474 | 0.1216350 | -0.0532742 | -0.4450925 | -0.3108967 | -0.1849214 | 0.3169922 | -0.3513413 | -0.1463420 | 0.0082562 | 0.0166538 | -0.2080379 | -0.1729566 | 0.0771509 | -0.1220543 | -0.0236794 | -0.1898965 | 0.0157579 | -0.2349552 |
| AcceptedCmp1 | -0.1512244 | -0.2571029 | -0.2032774 | -0.0079016 | -0.1650511 | 0.1406652 | -0.0859987 | -0.0449785 | -0.3321207 | -0.0047433 | 0.3575113 | 0.5060082 | -0.1428392 | -0.2113210 | -0.3731005 | 0.0379299 | -0.0736383 | -0.0144407 | -0.0954028 | 0.1157220 | 0.0341983 | -0.0267502 | 0.0414106 | -0.0006799 | -0.1793006 | 0.0148786 | -0.2218451 |
| AcceptedCmp2 | -0.0571148 | -0.1896916 | -0.2956346 | 0.0328896 | 0.1670118 | -0.0955575 | 0.0548912 | 0.0247733 | 0.4235079 | -0.6704752 | -0.2796780 | 0.1900572 | -0.0980933 | -0.2124280 | 0.0677366 | -0.0008377 | 0.0476798 | -0.0691934 | -0.0066742 | -0.0039986 | -0.0260142 | -0.0540206 | 0.0202522 | 0.0001696 | -0.0840353 | 0.0069734 | -0.1039752 |
| Complain | 0.0131909 | 0.0096340 | 0.0072370 | -0.0429894 | -0.0324495 | 0.1779010 | 0.7391715 | -0.6287884 | 0.0961444 | 0.0181117 | 0.0865489 | 0.0263419 | -0.0323748 | 0.0437998 | -0.0150963 | -0.0108617 | 0.0202436 | -0.0151135 | 0.0186802 | -0.0067636 | 0.0341260 | -0.0031672 | 0.0062002 | 0.0081768 | 0.0000000 | 0.0000000 | 0.0000000 |
| Age | -0.0463659 | 0.2194763 | -0.1172498 | 0.4355307 | -0.0360791 | -0.2420475 | 0.2306301 | 0.0721888 | -0.5118382 | -0.1620272 | -0.0714697 | 0.2064057 | 0.0162425 | 0.0523011 | 0.4806713 | 0.1259111 | -0.0554754 | 0.1491605 | -0.1002449 | 0.0661328 | 0.0084083 | -0.0206541 | -0.0360179 | 0.0239185 | 0.0000000 | 0.0000000 | 0.0000000 |
| Customer_Days | -0.0371158 | 0.1841063 | -0.0921633 | -0.4853823 | 0.3211784 | -0.0035616 | 0.1222621 | 0.0107145 | -0.4525948 | -0.1406240 | -0.3317354 | -0.1883660 | 0.0470718 | -0.1240848 | -0.2982994 | 0.0819031 | 0.0395570 | 0.1969330 | -0.1537487 | 0.0336068 | -0.1908819 | -0.0552311 | 0.1190810 | -0.0131660 | 0.0000000 | 0.0000000 | 0.0000000 |
| FamilyMember | 0.1732379 | 0.2391520 | -0.2591193 | 0.0983714 | -0.3408401 | 0.3423015 | -0.1147247 | -0.0484444 | -0.0098295 | -0.1300427 | -0.0708658 | -0.1961264 | 0.0204616 | -0.0765589 | -0.0786423 | 0.2212764 | -0.0154572 | 0.0046046 | 0.2040647 | 0.1241961 | -0.1038013 | 0.0051553 | -0.0680792 | 0.6264487 | 0.0000000 | 0.0000000 | 0.0000000 |
| TotalPurchases | -0.2996064 | 0.2201935 | -0.0572813 | -0.0353440 | 0.0290893 | -0.0450157 | -0.0161877 | -0.0240811 | 0.1729658 | 0.1280293 | 0.0496036 | 0.0883996 | 0.0210848 | -0.1248260 | 0.0318908 | 0.0532204 | -0.0504565 | 0.1203232 | 0.0493949 | 0.0863468 | -0.2264969 | 0.1361942 | -0.1389983 | -0.0410088 | -0.0316040 | 0.8074191 | 0.0796947 |
| Spending | -0.3240594 | 0.0279976 | -0.0187210 | -0.0520931 | 0.0113111 | 0.0727207 | -0.0166249 | -0.0021310 | -0.0363655 | 0.0573819 | -0.1338044 | -0.0166420 | -0.0133563 | 0.0288419 | 0.0591800 | 0.0956432 | -0.0481553 | -0.1506057 | 0.1927706 | 0.0477665 | 0.2684582 | -0.1313233 | 0.1358815 | -0.0008413 | -0.6395765 | -0.0755559 | 0.5118539 |
| AcceptedCmpOverall | -0.1753157 | -0.3522464 | -0.4044890 | -0.0403069 | -0.1260631 | -0.0234389 | 0.0338446 | 0.0213195 | -0.0595153 | -0.0001924 | 0.0695176 | -0.0312225 | 0.0772507 | 0.0528687 | -0.0226273 | -0.0619781 | -0.0260766 | 0.0614826 | -0.0473590 | -0.0220460 | -0.0750486 | -0.0330099 | -0.0403726 | -0.0037042 | 0.4958262 | -0.0411443 | 0.6134760 |
knitr::kable(variance_df, caption = "Variance Explained by Each Principal Component")
| Principal_Component | Variance | Percent_Variance |
|---|---|---|
| 1 | 8.6152397 | 31.9082953 |
| 2 | 2.4830288 | 9.1964031 |
| 3 | 2.3465471 | 8.6909151 |
| 4 | 1.4928754 | 5.5291680 |
| 5 | 1.1631892 | 4.3081081 |
| 6 | 1.0597817 | 3.9251173 |
| 7 | 1.0153490 | 3.7605520 |
| 8 | 0.9840128 | 3.6444919 |
| 9 | 0.8262506 | 3.0601873 |
| 10 | 0.8020947 | 2.9707210 |
| 11 | 0.7591514 | 2.8116717 |
| 12 | 0.6340123 | 2.3481937 |
| 13 | 0.6163340 | 2.2827186 |
| 14 | 0.5847525 | 2.1657500 |
| 15 | 0.5581573 | 2.0672491 |
| 16 | 0.4822097 | 1.7859619 |
| 17 | 0.4489330 | 1.6627150 |
| 18 | 0.4231662 | 1.5672824 |
| 19 | 0.3997372 | 1.4805081 |
| 20 | 0.3679080 | 1.3626223 |
| 21 | 0.3193783 | 1.1828826 |
| 22 | 0.2588496 | 0.9587022 |
| 23 | 0.2241497 | 0.8301840 |
| 24 | 0.1348918 | 0.4995994 |
| 25 | 0.0000000 | 0.0000000 |
| 26 | 0.0000000 | 0.0000000 |
| 27 | 0.0000000 | 0.0000000 |
# Scree plot for PCA
pca_var <- pca_result$sdev^2
pca_var_percent <- pca_var / sum(pca_var) * 100
barplot(pca_var_percent,
main = "Scree Plot",
xlab = "Principal Component",
ylab = "Percentage of Variance Explained",
col = "blue")
# Loadings
loadings <- as.data.frame(pca_result$rotation)
# Prepare loadings for plotting
loadings$variable <- rownames(loadings)
loadings_long <- reshape2::melt(loadings, id.vars = "variable")
# Create base plot with scores
biplot_gg <- ggplot(data = scores, aes(x = PC1, y = PC2)) +
geom_point(alpha = 0.5) + # Adjust alpha for point transparency if needed
theme_minimal()
# Add loadings as vectors
biplot_gg <- biplot_gg + geom_segment(data = loadings, aes(x = 0, y = 0, xend = PC1, yend = PC2), arrow = arrow(length = unit(0.02, "npc")), color = "red")
# Add labels using ggrepel library to avoid overlap
biplot_gg <- biplot_gg + geom_text_repel(data = loadings,
aes(x = PC1, y = PC2, label = variable),
size = 3, # Adjust text size as needed
box.padding = unit(0.35, "lines"),
point.padding = unit(0.3, "lines"))
# Display the plot
print(biplot_gg)
In this section we would develop two types of machine learning models as per specified in the Objectives section
# Import library for model training and evaluation
library(dplyr)
library(caret)
# reassigned clean_data into data for ease of model training
data <- clean_data
For classification model training we would be using the following:
Target : Response (1 if customer accept the last campaign offer, 0 otherwise) Features : All other columns in the ‘data’ dataframe except for column ID
We would perform the model training using the following three classification algorithms 1) Logistic Regression 2) Random Forest 3) XGBoost (Etreme Gradient Boosting)
Before modelling we need to split the dataset into training and testing set
set.seed(123) # Set a random seed for reproducibility
train_class_index <- createDataPartition(data$Response, p = 0.7, list = FALSE)
train_class_data <- data[train_class_index, ]
test_class_data <- data[-train_class_index, ]
#preprocess data before training to manage the diffrent data scaling and normalise data
preProcValue_class <- preProcess(train_class_data, method = c('center','scale','YeoJohnson'))
train_class_data <- predict(preProcValue_class, newdata = train_class_data)
test_class_data <- predict(preProcValue_class, newdata = test_class_data)
As there some columns such as Marital Status which are categorical we will to perform encoding to translate the information into inouts that is acceptable by the machine learning algorithms. For this project, we choose One Hot Encoding as it was a common method to encode categorical columns
dummies_class_model <- dummyVars(Response ~. , data = train_class_data )
train_class_data_encoded <- predict(dummies_class_model, newdata = train_class_data) %>% as.data.frame()
test_class_data_encoded <- predict(dummies_class_model, newdata = test_class_data) %>% as.data.frame()
train_class_data_encoded$Response <- train_class_data$Response
test_class_data_encoded$Response <- test_class_data$Response
To achieve a better generalisation on the model, we would use the cross validation technique when training the model. We would set the control variable for 3 fold cross validation
fitControl_class <- trainControl(method = 'cv', number = 5)
tune_num <- 5
Firstly we would train a model using logistic regression
set.seed(123)
model_class_lr <- train(Response ~., train_class_data_encoded , trControl = fitControl_class, method = 'glmnet', tuneLength = tune_num)
print(model_class_lr)
## glmnet
##
## 1569 samples
## 34 predictor
## 2 classes: '0', '1'
##
## No pre-processing
## Resampling: Cross-Validated (5 fold)
## Summary of sample sizes: 1255, 1255, 1256, 1255, 1255
## Resampling results across tuning parameters:
##
## alpha lambda Accuracy Kappa
## 0.100 0.0001403266 0.8999430 0.5578265
## 0.100 0.0006513384 0.8999430 0.5540880
## 0.100 0.0030232451 0.8986691 0.5393692
## 0.100 0.0140326608 0.9018518 0.5215486
## 0.100 0.0651338415 0.8878248 0.3888577
## 0.325 0.0001403266 0.8999430 0.5578265
## 0.325 0.0006513384 0.8999430 0.5558961
## 0.325 0.0030232451 0.8986691 0.5369850
## 0.325 0.0140326608 0.8999430 0.5068941
## 0.325 0.0651338415 0.8814452 0.3224139
## 0.550 0.0001403266 0.8999430 0.5578265
## 0.550 0.0006513384 0.8993061 0.5540029
## 0.550 0.0030232451 0.8973953 0.5289237
## 0.550 0.0140326608 0.8993000 0.4988733
## 0.550 0.0651338415 0.8744409 0.2629674
## 0.775 0.0001403266 0.8999430 0.5578265
## 0.775 0.0006513384 0.8999430 0.5559040
## 0.775 0.0030232451 0.8967583 0.5233209
## 0.775 0.0140326608 0.8935675 0.4573158
## 0.775 0.0651338415 0.8680654 0.1914112
## 1.000 0.0001403266 0.8999430 0.5578265
## 1.000 0.0006513384 0.8999430 0.5559040
## 1.000 0.0030232451 0.8986691 0.5342438
## 1.000 0.0140326608 0.8903828 0.4319175
## 1.000 0.0651338415 0.8680654 0.1914112
##
## Accuracy was used to select the optimal model using the largest value.
## The final values used for the model were alpha = 0.1 and lambda = 0.01403266.
plot(model_class_lr)
We we will calculate the confusion matrix to evaluate the performance of the models
predicted_model_class_lr <- predict(model_class_lr, test_class_data_encoded)
cm_model_class_lr <- confusionMatrix(reference = test_class_data_encoded$Response, data = predicted_model_class_lr, mode='everything',positive = '1')
print(cm_model_class_lr)
## Confusion Matrix and Statistics
##
## Reference
## Prediction 0 1
## 0 555 56
## 1 16 44
##
## Accuracy : 0.8927
## 95% CI : (0.8668, 0.9151)
## No Information Rate : 0.851
## P-Value [Acc > NIR] : 0.0009786
##
## Kappa : 0.4934
##
## Mcnemar's Test P-Value : 4.303e-06
##
## Sensitivity : 0.44000
## Specificity : 0.97198
## Pos Pred Value : 0.73333
## Neg Pred Value : 0.90835
## Precision : 0.73333
## Recall : 0.44000
## F1 : 0.55000
## Prevalence : 0.14903
## Detection Rate : 0.06557
## Detection Prevalence : 0.08942
## Balanced Accuracy : 0.70599
##
## 'Positive' Class : 1
##
library(pROC)
test_target <- unclass(test_class_data_encoded$Response)
pred_target <- unclass(predicted_model_class_lr)
auc_model <- roc(test_target,pred_target)
print(auc_model)
##
## Call:
## roc.default(response = test_target, predictor = pred_target)
##
## Data: pred_target in 571 controls (test_target 1) < 100 cases (test_target 2).
## Area under the curve: 0.706
plot(auc_model, ylim=c(0,1),xlim=c(1,0), print.thres=TRUE, main=paste('AUC:',round(auc_model$auc[[1]],2)))
abline(h=1,col='blue',lwd=2)
abline(h=0,col='red',lwd=2)
Secondly we would train a model using random forest
set.seed(123)
model_class_rf <- train(Response ~., train_class_data_encoded , trControl = fitControl_class, method = 'rf', tuneLength = tune_num)
print(model_class_rf)
## Random Forest
##
## 1569 samples
## 34 predictor
## 2 classes: '0', '1'
##
## No pre-processing
## Resampling: Cross-Validated (5 fold)
## Summary of sample sizes: 1255, 1255, 1256, 1255, 1255
## Resampling results across tuning parameters:
##
## mtry Accuracy Kappa
## 2 0.8801795 0.3499323
## 10 0.8859160 0.4383262
## 18 0.8852852 0.4558195
## 26 0.8852872 0.4642413
## 34 0.8859262 0.4725744
##
## Accuracy was used to select the optimal model using the largest value.
## The final value used for the model was mtry = 34.
plot(model_class_rf)
The top 10 performing variables are shown, since 10 randomly selected
predictors give almost as good an accuracy score as all 34 predictors,
with 34 selected as the final due to it marginally higher accuracy and
higher Kappa score.
# Extract variable importance
importance <- varImp(model_class_rf, scale = FALSE)
importance_df <- as.data.frame(importance$importance)
importance_df$Variable <- rownames(importance_df)
importance_df <- importance_df[order(-importance_df$Overall), ]
top_10_predictors <- head(importance_df, 10)
print(top_10_predictors)
## Overall Variable
## AcceptedCmpOverall 59.44783 AcceptedCmpOverall
## Recency 40.23875 Recency
## Customer_Days 39.38606 Customer_Days
## MntMeatProducts 21.18782 MntMeatProducts
## Income 17.66550 Income
## MntGoldProds 17.14471 MntGoldProds
## Age 15.74375 Age
## MntWines 15.56271 MntWines
## NumWebVisitsMonth 14.78894 NumWebVisitsMonth
## FamilyMember 13.88956 FamilyMember
We we will calculate the confusion matrix to evaluate the performance of the models
predicted_model_class_rf <- predict(model_class_rf, test_class_data_encoded)
cm_model_class_rf <- confusionMatrix(reference = test_class_data_encoded$Response, data = predicted_model_class_rf, mode='everything',positive ='1')
print(cm_model_class_rf)
## Confusion Matrix and Statistics
##
## Reference
## Prediction 0 1
## 0 551 54
## 1 20 46
##
## Accuracy : 0.8897
## 95% CI : (0.8635, 0.9124)
## No Information Rate : 0.851
## P-Value [Acc > NIR] : 0.002114
##
## Kappa : 0.4943
##
## Mcnemar's Test P-Value : 0.000125
##
## Sensitivity : 0.46000
## Specificity : 0.96497
## Pos Pred Value : 0.69697
## Neg Pred Value : 0.91074
## Precision : 0.69697
## Recall : 0.46000
## F1 : 0.55422
## Prevalence : 0.14903
## Detection Rate : 0.06855
## Detection Prevalence : 0.09836
## Balanced Accuracy : 0.71249
##
## 'Positive' Class : 1
##
library(pROC)
test_target_rf <- unclass(test_class_data_encoded$Response)
pred_target_rf <- unclass(predicted_model_class_rf)
auc_model_rf <- roc(test_target_rf,pred_target_rf)
print(auc_model_rf)
##
## Call:
## roc.default(response = test_target_rf, predictor = pred_target_rf)
##
## Data: pred_target_rf in 571 controls (test_target_rf 1) < 100 cases (test_target_rf 2).
## Area under the curve: 0.7125
plot(auc_model_rf, ylim=c(0,1),xlim=c(1,0), print.thres=TRUE, main=paste('AUC:',round(auc_model_rf$auc[[1]],2)))
abline(h=1,col='blue',lwd=2)
abline(h=0,col='red',lwd=2)
Lastly we would train a model using extreme gradient boosting
set.seed(123)
model_class_xg <- train(Response ~., train_class_data_encoded , trControl = fitControl_class, method = 'xgbTree', tuneLength = tune_num)
print(model_class_xg)
## eXtreme Gradient Boosting
##
## 1569 samples
## 34 predictor
## 2 classes: '0', '1'
##
## No pre-processing
## Resampling: Cross-Validated (5 fold)
## Summary of sample sizes: 1255, 1255, 1256, 1255, 1255
## Resampling results across tuning parameters:
##
## eta max_depth colsample_bytree subsample nrounds Accuracy Kappa
## 0.3 1 0.6 0.500 50 0.8897418 0.4501731
## 0.3 1 0.6 0.500 100 0.8891130 0.4851387
## 0.3 1 0.6 0.500 150 0.8910136 0.4986783
## 0.3 1 0.6 0.500 200 0.8897377 0.5076492
## 0.3 1 0.6 0.500 250 0.8884638 0.5041478
## 0.3 1 0.6 0.625 50 0.8859201 0.4385504
## 0.3 1 0.6 0.625 100 0.8942085 0.5019622
## 0.3 1 0.6 0.625 150 0.8935675 0.4923017
## 0.3 1 0.6 0.625 200 0.8910197 0.4998586
## 0.3 1 0.6 0.625 250 0.8935756 0.5153310
## 0.3 1 0.6 0.750 50 0.8871920 0.4286475
## 0.3 1 0.6 0.750 100 0.8942085 0.4967151
## 0.3 1 0.6 0.750 150 0.8929346 0.5009365
## 0.3 1 0.6 0.750 200 0.8891048 0.4899555
## 0.3 1 0.6 0.750 250 0.8897458 0.5104621
## 0.3 1 0.6 0.875 50 0.8897458 0.4521982
## 0.3 1 0.6 0.875 100 0.8954742 0.4959770
## 0.3 1 0.6 0.875 150 0.8961112 0.5104826
## 0.3 1 0.6 0.875 200 0.8897438 0.4918651
## 0.3 1 0.6 0.875 250 0.8916546 0.5024662
## 0.3 1 0.6 1.000 50 0.8910116 0.4367461
## 0.3 1 0.6 1.000 100 0.8935634 0.4712649
## 0.3 1 0.6 1.000 150 0.8961112 0.5003373
## 0.3 1 0.6 1.000 200 0.8948393 0.5046629
## 0.3 1 0.6 1.000 250 0.8948393 0.5108036
## 0.3 1 0.8 0.500 50 0.8935655 0.4878948
## 0.3 1 0.8 0.500 100 0.8935634 0.5052486
## 0.3 1 0.8 0.500 150 0.8903807 0.5051362
## 0.3 1 0.8 0.500 200 0.8871920 0.4847349
## 0.3 1 0.8 0.500 250 0.8903868 0.5086101
## 0.3 1 0.8 0.625 50 0.8910156 0.4558072
## 0.3 1 0.8 0.625 100 0.8891028 0.4704589
## 0.3 1 0.8 0.625 150 0.8910116 0.4968637
## 0.3 1 0.8 0.625 200 0.8865571 0.4918523
## 0.3 1 0.8 0.625 250 0.8884719 0.5009598
## 0.3 1 0.8 0.750 50 0.8929326 0.4681610
## 0.3 1 0.8 0.750 100 0.8916546 0.4886093
## 0.3 1 0.8 0.750 150 0.8922956 0.4978692
## 0.3 1 0.8 0.750 200 0.8897458 0.4978666
## 0.3 1 0.8 0.750 250 0.8935695 0.5237914
## 0.3 1 0.8 0.875 50 0.8916506 0.4572300
## 0.3 1 0.8 0.875 100 0.8910136 0.4782912
## 0.3 1 0.8 0.875 150 0.8897397 0.4798847
## 0.3 1 0.8 0.875 200 0.8884699 0.4899331
## 0.3 1 0.8 0.875 250 0.8897479 0.4973454
## 0.3 1 0.8 1.000 50 0.8884638 0.4325485
## 0.3 1 0.8 1.000 100 0.8916526 0.4703323
## 0.3 1 0.8 1.000 150 0.8942044 0.4927625
## 0.3 1 0.8 1.000 200 0.8967522 0.5109089
## 0.3 1 0.8 1.000 250 0.8980281 0.5208046
## 0.3 2 0.6 0.500 50 0.8852791 0.4767340
## 0.3 2 0.6 0.500 100 0.8865652 0.4783355
## 0.3 2 0.6 0.500 150 0.8859201 0.4915353
## 0.3 2 0.6 0.500 200 0.8859120 0.4933389
## 0.3 2 0.6 0.500 250 0.8789015 0.4674777
## 0.3 2 0.6 0.625 50 0.8897499 0.4825680
## 0.3 2 0.6 0.625 100 0.8859282 0.4785155
## 0.3 2 0.6 0.625 150 0.8872042 0.4991469
## 0.3 2 0.6 0.625 200 0.8814554 0.4745419
## 0.3 2 0.6 0.625 250 0.8859262 0.4982142
## 0.3 2 0.6 0.750 50 0.8840052 0.4592690
## 0.3 2 0.6 0.750 100 0.8903767 0.5010621
## 0.3 2 0.6 0.750 150 0.8891048 0.4966118
## 0.3 2 0.6 0.750 200 0.8884638 0.5034521
## 0.3 2 0.6 0.750 250 0.8865530 0.4991468
## 0.3 2 0.6 0.875 50 0.8891028 0.4669051
## 0.3 2 0.6 0.875 100 0.8878309 0.4759269
## 0.3 2 0.6 0.875 150 0.8865550 0.4866814
## 0.3 2 0.6 0.875 200 0.8910177 0.5118197
## 0.3 2 0.6 0.875 250 0.8916546 0.5188643
## 0.3 2 0.6 1.000 50 0.8910218 0.4795279
## 0.3 2 0.6 1.000 100 0.8859201 0.4765891
## 0.3 2 0.6 1.000 150 0.8897438 0.4960913
## 0.3 2 0.6 1.000 200 0.8846442 0.4825711
## 0.3 2 0.6 1.000 250 0.8897418 0.5054875
## 0.3 2 0.8 0.500 50 0.8897438 0.4921774
## 0.3 2 0.8 0.500 100 0.8865571 0.4927610
## 0.3 2 0.8 0.500 150 0.8846442 0.4983884
## 0.3 2 0.8 0.500 200 0.8846462 0.5026798
## 0.3 2 0.8 0.500 250 0.8789076 0.4724174
## 0.3 2 0.8 0.625 50 0.8846503 0.4718662
## 0.3 2 0.8 0.625 100 0.8871920 0.4983521
## 0.3 2 0.8 0.625 150 0.8833622 0.4864065
## 0.3 2 0.8 0.625 200 0.8884638 0.5069367
## 0.3 2 0.8 0.625 250 0.8865530 0.5036860
## 0.3 2 0.8 0.750 50 0.8846442 0.4574420
## 0.3 2 0.8 0.750 100 0.8865550 0.4789298
## 0.3 2 0.8 0.750 150 0.8929326 0.5181332
## 0.3 2 0.8 0.750 200 0.8922936 0.5168331
## 0.3 2 0.8 0.750 250 0.8922895 0.5233382
## 0.3 2 0.8 0.875 50 0.8840113 0.4394101
## 0.3 2 0.8 0.875 100 0.8948495 0.5150217
## 0.3 2 0.8 0.875 150 0.8929326 0.5181830
## 0.3 2 0.8 0.875 200 0.8916567 0.5106411
## 0.3 2 0.8 0.875 250 0.8922956 0.5195807
## 0.3 2 0.8 1.000 50 0.8891089 0.4621115
## 0.3 2 0.8 1.000 100 0.8871940 0.4742488
## 0.3 2 0.8 1.000 150 0.8884699 0.4941551
## 0.3 2 0.8 1.000 200 0.8903787 0.5146909
## 0.3 2 0.8 1.000 250 0.8929305 0.5266725
## 0.3 3 0.6 0.500 50 0.8910156 0.5061423
## 0.3 3 0.6 0.500 100 0.8897397 0.5165739
## 0.3 3 0.6 0.500 150 0.8840072 0.4964908
## 0.3 3 0.6 0.500 200 0.8820964 0.4854413
## 0.3 3 0.6 0.500 250 0.8808103 0.4840408
## 0.3 3 0.6 0.625 50 0.8916648 0.4981944
## 0.3 3 0.6 0.625 100 0.8935634 0.5321672
## 0.3 3 0.6 0.625 150 0.8948373 0.5306381
## 0.3 3 0.6 0.625 200 0.8916526 0.5121277
## 0.3 3 0.6 0.625 250 0.8948393 0.5301234
## 0.3 3 0.6 0.750 50 0.8820923 0.4542442
## 0.3 3 0.6 0.750 100 0.8827313 0.4641349
## 0.3 3 0.6 0.750 150 0.8820964 0.4737390
## 0.3 3 0.6 0.750 200 0.8789097 0.4582159
## 0.3 3 0.6 0.750 250 0.8821005 0.4737782
## 0.3 3 0.6 0.875 50 0.8891069 0.4769394
## 0.3 3 0.6 0.875 100 0.8820923 0.4659213
## 0.3 3 0.6 0.875 150 0.8808225 0.4709135
## 0.3 3 0.6 0.875 200 0.8827374 0.4858815
## 0.3 3 0.6 0.875 250 0.8776378 0.4583886
## 0.3 3 0.6 1.000 50 0.8884719 0.4904275
## 0.3 3 0.6 1.000 100 0.8865530 0.4807136
## 0.3 3 0.6 1.000 150 0.8865509 0.4909173
## 0.3 3 0.6 1.000 200 0.8859140 0.4851201
## 0.3 3 0.6 1.000 250 0.8820883 0.4673490
## 0.3 3 0.8 0.500 50 0.8801815 0.4683445
## 0.3 3 0.8 0.500 100 0.8769866 0.4552086
## 0.3 3 0.8 0.500 150 0.8833683 0.4896464
## 0.3 3 0.8 0.500 200 0.8776256 0.4522729
## 0.3 3 0.8 0.500 250 0.8814513 0.4830644
## 0.3 3 0.8 0.625 50 0.8840072 0.4722397
## 0.3 3 0.8 0.625 100 0.8884699 0.4973485
## 0.3 3 0.8 0.625 150 0.8878370 0.4984396
## 0.3 3 0.8 0.625 200 0.8833723 0.4787472
## 0.3 3 0.8 0.625 250 0.8865591 0.4975426
## 0.3 3 0.8 0.750 50 0.8891150 0.4786158
## 0.3 3 0.8 0.750 100 0.8852852 0.4879364
## 0.3 3 0.8 0.750 150 0.8808225 0.4607804
## 0.3 3 0.8 0.750 200 0.8846401 0.4724511
## 0.3 3 0.8 0.750 250 0.8814574 0.4689260
## 0.3 3 0.8 0.875 50 0.8821005 0.4551528
## 0.3 3 0.8 0.875 100 0.8852832 0.4828043
## 0.3 3 0.8 0.875 150 0.8801774 0.4653196
## 0.3 3 0.8 0.875 200 0.8801774 0.4652434
## 0.3 3 0.8 0.875 250 0.8833683 0.4790410
## 0.3 3 0.8 1.000 50 0.8859221 0.4624276
## 0.3 3 0.8 1.000 100 0.8865571 0.4794460
## 0.3 3 0.8 1.000 150 0.8846503 0.4821156
## 0.3 3 0.8 1.000 200 0.8821005 0.4710670
## 0.3 3 0.8 1.000 250 0.8833744 0.4784156
## 0.3 4 0.6 0.500 50 0.8769988 0.4363817
## 0.3 4 0.6 0.500 100 0.8846483 0.4744830
## 0.3 4 0.6 0.500 150 0.8840154 0.4770269
## 0.3 4 0.6 0.500 200 0.8859160 0.4942915
## 0.3 4 0.6 0.500 250 0.8859201 0.4928492
## 0.3 4 0.6 0.625 50 0.8808266 0.4564525
## 0.3 4 0.6 0.625 100 0.8865611 0.4940356
## 0.3 4 0.6 0.625 150 0.8833683 0.4785639
## 0.3 4 0.6 0.625 200 0.8884658 0.4903480
## 0.3 4 0.6 0.625 250 0.8859120 0.4901645
## 0.3 4 0.6 0.750 50 0.8846442 0.4723500
## 0.3 4 0.6 0.750 100 0.8808205 0.4620963
## 0.3 4 0.6 0.750 150 0.8808185 0.4590159
## 0.3 4 0.6 0.750 200 0.8820923 0.4701355
## 0.3 4 0.6 0.750 250 0.8814554 0.4656110
## 0.3 4 0.6 0.875 50 0.8840011 0.4626501
## 0.3 4 0.6 0.875 100 0.8827313 0.4707950
## 0.3 4 0.6 0.875 150 0.8820923 0.4692096
## 0.3 4 0.6 0.875 200 0.8782707 0.4587346
## 0.3 4 0.6 0.875 250 0.8776337 0.4578959
## 0.3 4 0.6 1.000 50 0.8852872 0.4710697
## 0.3 4 0.6 1.000 100 0.8865591 0.4760889
## 0.3 4 0.6 1.000 150 0.8827334 0.4681652
## 0.3 4 0.6 1.000 200 0.8827273 0.4693231
## 0.3 4 0.6 1.000 250 0.8820923 0.4707374
## 0.3 4 0.8 0.500 50 0.8814635 0.4752814
## 0.3 4 0.8 0.500 100 0.8865591 0.5117289
## 0.3 4 0.8 0.500 150 0.8891109 0.5102340
## 0.3 4 0.8 0.500 200 0.8891069 0.4999257
## 0.3 4 0.8 0.500 250 0.8821005 0.4843797
## 0.3 4 0.8 0.625 50 0.8795486 0.4581119
## 0.3 4 0.8 0.625 100 0.8833642 0.4762670
## 0.3 4 0.8 0.625 150 0.8846401 0.4882879
## 0.3 4 0.8 0.625 200 0.8808144 0.4715985
## 0.3 4 0.8 0.625 250 0.8833662 0.4873938
## 0.3 4 0.8 0.750 50 0.8833703 0.4685182
## 0.3 4 0.8 0.750 100 0.8789117 0.4581926
## 0.3 4 0.8 0.750 150 0.8795446 0.4557248
## 0.3 4 0.8 0.750 200 0.8820964 0.4736929
## 0.3 4 0.8 0.750 250 0.8782687 0.4591820
## 0.3 4 0.8 0.875 50 0.8865611 0.4797866
## 0.3 4 0.8 0.875 100 0.8865550 0.4877692
## 0.3 4 0.8 0.875 150 0.8865530 0.4816771
## 0.3 4 0.8 0.875 200 0.8859120 0.4830791
## 0.3 4 0.8 0.875 250 0.8833703 0.4724094
## 0.3 4 0.8 1.000 50 0.8865571 0.4883996
## 0.3 4 0.8 1.000 100 0.8795425 0.4517548
## 0.3 4 0.8 1.000 150 0.8820903 0.4640475
## 0.3 4 0.8 1.000 200 0.8789015 0.4551226
## 0.3 4 0.8 1.000 250 0.8788975 0.4584234
## 0.3 5 0.6 0.500 50 0.8827435 0.4698319
## 0.3 5 0.6 0.500 100 0.8795568 0.4654709
## 0.3 5 0.6 0.500 150 0.8808164 0.4677435
## 0.3 5 0.6 0.500 200 0.8827293 0.4825383
## 0.3 5 0.6 0.500 250 0.8808185 0.4712533
## 0.3 5 0.6 0.625 50 0.8795527 0.4546192
## 0.3 5 0.6 0.625 100 0.8808225 0.4633466
## 0.3 5 0.6 0.625 150 0.8827313 0.4719869
## 0.3 5 0.6 0.625 200 0.8808185 0.4715168
## 0.3 5 0.6 0.625 250 0.8827334 0.4735184
## 0.3 5 0.6 0.750 50 0.8801836 0.4580151
## 0.3 5 0.6 0.750 100 0.8846523 0.4840933
## 0.3 5 0.6 0.750 150 0.8878330 0.5002090
## 0.3 5 0.6 0.750 200 0.8852872 0.4874882
## 0.3 5 0.6 0.750 250 0.8846483 0.4880802
## 0.3 5 0.6 0.875 50 0.8884679 0.4891121
## 0.3 5 0.6 0.875 100 0.8846483 0.4733757
## 0.3 5 0.6 0.875 150 0.8852872 0.4819864
## 0.3 5 0.6 0.875 200 0.8852852 0.4845564
## 0.3 5 0.6 0.875 250 0.8840093 0.4806004
## 0.3 5 0.6 1.000 50 0.8846483 0.4619816
## 0.3 5 0.6 1.000 100 0.8789056 0.4480411
## 0.3 5 0.6 1.000 150 0.8808164 0.4564025
## 0.3 5 0.6 1.000 200 0.8782687 0.4475672
## 0.3 5 0.6 1.000 250 0.8769948 0.4409518
## 0.3 5 0.8 0.500 50 0.8852791 0.4908475
## 0.3 5 0.8 0.500 100 0.8839991 0.4737506
## 0.3 5 0.8 0.500 150 0.8871879 0.4944062
## 0.3 5 0.8 0.500 200 0.8859099 0.4938328
## 0.3 5 0.8 0.500 250 0.8852750 0.4860967
## 0.3 5 0.8 0.625 50 0.8821025 0.4718937
## 0.3 5 0.8 0.625 100 0.8840072 0.4818342
## 0.3 5 0.8 0.625 150 0.8820964 0.4836657
## 0.3 5 0.8 0.625 200 0.8846462 0.4959368
## 0.3 5 0.8 0.625 250 0.8840011 0.4909613
## 0.3 5 0.8 0.750 50 0.8821005 0.4591192
## 0.3 5 0.8 0.750 100 0.8833723 0.4747919
## 0.3 5 0.8 0.750 150 0.8820944 0.4721017
## 0.3 5 0.8 0.750 200 0.8820903 0.4675382
## 0.3 5 0.8 0.750 250 0.8833662 0.4699626
## 0.3 5 0.8 0.875 50 0.8878350 0.4835257
## 0.3 5 0.8 0.875 100 0.8859181 0.4778555
## 0.3 5 0.8 0.875 150 0.8871920 0.4921527
## 0.3 5 0.8 0.875 200 0.8865530 0.4880826
## 0.3 5 0.8 0.875 250 0.8814554 0.4669461
## 0.3 5 0.8 1.000 50 0.8878330 0.4828387
## 0.3 5 0.8 1.000 100 0.8871940 0.4877724
## 0.3 5 0.8 1.000 150 0.8891069 0.5016576
## 0.3 5 0.8 1.000 200 0.8871940 0.4928198
## 0.3 5 0.8 1.000 250 0.8884699 0.4968876
## 0.4 1 0.6 0.500 50 0.8897540 0.4656471
## 0.4 1 0.6 0.500 100 0.8865632 0.4864245
## 0.4 1 0.6 0.500 150 0.8903848 0.5110276
## 0.4 1 0.6 0.500 200 0.8878228 0.5014712
## 0.4 1 0.6 0.500 250 0.8890987 0.5095266
## 0.4 1 0.6 0.625 50 0.8891069 0.4640523
## 0.4 1 0.6 0.625 100 0.8935614 0.5144692
## 0.4 1 0.6 0.625 150 0.8916526 0.5125268
## 0.4 1 0.6 0.625 200 0.8922936 0.5158455
## 0.4 1 0.6 0.625 250 0.8795466 0.4806179
## 0.4 1 0.6 0.750 50 0.8891048 0.4572448
## 0.4 1 0.6 0.750 100 0.8929326 0.5042712
## 0.4 1 0.6 0.750 150 0.8903807 0.4998426
## 0.4 1 0.6 0.750 200 0.8891130 0.5130891
## 0.4 1 0.6 0.750 250 0.8884760 0.5164418
## 0.4 1 0.6 0.875 50 0.8922916 0.4687275
## 0.4 1 0.6 0.875 100 0.8935655 0.5017034
## 0.4 1 0.6 0.875 150 0.8929285 0.5113189
## 0.4 1 0.6 0.875 200 0.8929305 0.5118669
## 0.4 1 0.6 0.875 250 0.8884699 0.4996746
## 0.4 1 0.6 1.000 50 0.8916526 0.4636074
## 0.4 1 0.6 1.000 100 0.8973871 0.5019091
## 0.4 1 0.6 1.000 150 0.8980241 0.5139523
## 0.4 1 0.6 1.000 200 0.8967542 0.5171781
## 0.4 1 0.6 1.000 250 0.8948434 0.5166475
## 0.4 1 0.8 0.500 50 0.8954865 0.4916241
## 0.4 1 0.8 0.500 100 0.8929305 0.5062620
## 0.4 1 0.8 0.500 150 0.8897418 0.5136517
## 0.4 1 0.8 0.500 200 0.8871960 0.4985260
## 0.4 1 0.8 0.500 250 0.8922875 0.5257200
## 0.4 1 0.8 0.625 50 0.8948393 0.4972656
## 0.4 1 0.8 0.625 100 0.8980302 0.5270807
## 0.4 1 0.8 0.625 150 0.8922916 0.5062835
## 0.4 1 0.8 0.625 200 0.8891150 0.5118286
## 0.4 1 0.8 0.625 250 0.8840093 0.4869609
## 0.4 1 0.8 0.750 50 0.8922997 0.4777555
## 0.4 1 0.8 0.750 100 0.8923017 0.5005721
## 0.4 1 0.8 0.750 150 0.8903828 0.4993005
## 0.4 1 0.8 0.750 200 0.8884740 0.4989195
## 0.4 1 0.8 0.750 250 0.8865632 0.5001667
## 0.4 1 0.8 0.875 50 0.8922956 0.4694639
## 0.4 1 0.8 0.875 100 0.8986590 0.5216626
## 0.4 1 0.8 0.875 150 0.8929305 0.5086597
## 0.4 1 0.8 0.875 200 0.8891089 0.5044414
## 0.4 1 0.8 0.875 250 0.8865611 0.5031779
## 0.4 1 0.8 1.000 50 0.8910197 0.4549497
## 0.4 1 0.8 1.000 100 0.8942044 0.4885365
## 0.4 1 0.8 1.000 150 0.8929285 0.4930973
## 0.4 1 0.8 1.000 200 0.8973932 0.5184701
## 0.4 1 0.8 1.000 250 0.8948434 0.5116537
## 0.4 2 0.6 0.500 50 0.8840052 0.4465662
## 0.4 2 0.6 0.500 100 0.8814595 0.4670190
## 0.4 2 0.6 0.500 150 0.8776236 0.4662151
## 0.4 2 0.6 0.500 200 0.8776297 0.4634836
## 0.4 2 0.6 0.500 250 0.8782646 0.4746354
## 0.4 2 0.6 0.625 50 0.8891048 0.4834822
## 0.4 2 0.6 0.625 100 0.8839930 0.4859340
## 0.4 2 0.6 0.625 150 0.8865509 0.5091132
## 0.4 2 0.6 0.625 200 0.8833662 0.4881404
## 0.4 2 0.6 0.625 250 0.8846442 0.4873026
## 0.4 2 0.6 0.750 50 0.8878330 0.4848085
## 0.4 2 0.6 0.750 100 0.8954763 0.5393536
## 0.4 2 0.6 0.750 150 0.8910156 0.5342457
## 0.4 2 0.6 0.750 200 0.8922936 0.5338572
## 0.4 2 0.6 0.750 250 0.8897438 0.5266528
## 0.4 2 0.6 0.875 50 0.8871899 0.4731488
## 0.4 2 0.6 0.875 100 0.8852791 0.4838698
## 0.4 2 0.6 0.875 150 0.8859160 0.5030910
## 0.4 2 0.6 0.875 200 0.8865509 0.4918282
## 0.4 2 0.6 0.875 250 0.8878309 0.5089728
## 0.4 2 0.6 1.000 50 0.8910136 0.4846384
## 0.4 2 0.6 1.000 100 0.8884679 0.4891221
## 0.4 2 0.6 1.000 150 0.8871940 0.4917762
## 0.4 2 0.6 1.000 200 0.8840093 0.4906225
## 0.4 2 0.6 1.000 250 0.8852771 0.4915306
## 0.4 2 0.8 0.500 50 0.8916587 0.5109445
## 0.4 2 0.8 0.500 100 0.8910156 0.5181769
## 0.4 2 0.8 0.500 150 0.8916506 0.5268326
## 0.4 2 0.8 0.500 200 0.8903787 0.5233980
## 0.4 2 0.8 0.500 250 0.8827395 0.4979880
## 0.4 2 0.8 0.625 50 0.8833662 0.4638966
## 0.4 2 0.8 0.625 100 0.8852791 0.4948892
## 0.4 2 0.8 0.625 150 0.8859181 0.5006028
## 0.4 2 0.8 0.625 200 0.8808246 0.4858520
## 0.4 2 0.8 0.625 250 0.8865652 0.5102571
## 0.4 2 0.8 0.750 50 0.8884699 0.4854830
## 0.4 2 0.8 0.750 100 0.8954783 0.5181806
## 0.4 2 0.8 0.750 150 0.8942085 0.5314046
## 0.4 2 0.8 0.750 200 0.8916485 0.5184262
## 0.4 2 0.8 0.750 250 0.8871920 0.5054930
## 0.4 2 0.8 0.875 50 0.8884699 0.4816760
## 0.4 2 0.8 0.875 100 0.8884658 0.4880504
## 0.4 2 0.8 0.875 150 0.8903726 0.5074869
## 0.4 2 0.8 0.875 200 0.8929224 0.5199940
## 0.4 2 0.8 0.875 250 0.8916526 0.5208145
## 0.4 2 0.8 1.000 50 0.8846483 0.4552875
## 0.4 2 0.8 1.000 100 0.8878350 0.4817496
## 0.4 2 0.8 1.000 150 0.8878370 0.4888407
## 0.4 2 0.8 1.000 200 0.8910116 0.5106068
## 0.4 2 0.8 1.000 250 0.8897397 0.5174440
## 0.4 3 0.6 0.500 50 0.8789036 0.4666215
## 0.4 3 0.6 0.500 100 0.8846483 0.4921267
## 0.4 3 0.6 0.500 150 0.8852811 0.4890953
## 0.4 3 0.6 0.500 200 0.8846422 0.4897169
## 0.4 3 0.6 0.500 250 0.8871899 0.5058836
## 0.4 3 0.6 0.625 50 0.8820985 0.4581436
## 0.4 3 0.6 0.625 100 0.8865611 0.4825040
## 0.4 3 0.6 0.625 150 0.8808185 0.4622916
## 0.4 3 0.6 0.625 200 0.8801795 0.4677106
## 0.4 3 0.6 0.625 250 0.8782626 0.4571415
## 0.4 3 0.6 0.750 50 0.8763619 0.4394070
## 0.4 3 0.6 0.750 100 0.8820985 0.4747621
## 0.4 3 0.6 0.750 150 0.8801876 0.4702776
## 0.4 3 0.6 0.750 200 0.8782748 0.4562484
## 0.4 3 0.6 0.750 250 0.8763578 0.4625099
## 0.4 3 0.6 0.875 50 0.8884760 0.4792909
## 0.4 3 0.6 0.875 100 0.8852933 0.4913365
## 0.4 3 0.6 0.875 150 0.8770009 0.4581046
## 0.4 3 0.6 0.875 200 0.8770009 0.4549690
## 0.4 3 0.6 0.875 250 0.8789097 0.4642187
## 0.4 3 0.6 1.000 50 0.8840174 0.4695941
## 0.4 3 0.6 1.000 100 0.8763680 0.4422011
## 0.4 3 0.6 1.000 150 0.8808266 0.4613033
## 0.4 3 0.6 1.000 200 0.8808246 0.4653336
## 0.4 3 0.6 1.000 250 0.8801897 0.4669457
## 0.4 3 0.8 0.500 50 0.8865693 0.5043372
## 0.4 3 0.8 0.500 100 0.8871981 0.4984338
## 0.4 3 0.8 0.500 150 0.8859262 0.5050299
## 0.4 3 0.8 0.500 200 0.8814635 0.4842173
## 0.4 3 0.8 0.500 250 0.8808185 0.4789274
## 0.4 3 0.8 0.625 50 0.8872001 0.5015598
## 0.4 3 0.8 0.625 100 0.8827313 0.4732149
## 0.4 3 0.8 0.625 150 0.8795425 0.4660224
## 0.4 3 0.8 0.625 200 0.8763538 0.4475097
## 0.4 3 0.8 0.625 250 0.8782666 0.4628714
## 0.4 3 0.8 0.750 50 0.8846483 0.4760307
## 0.4 3 0.8 0.750 100 0.8808246 0.4733572
## 0.4 3 0.8 0.750 150 0.8789117 0.4659185
## 0.4 3 0.8 0.750 200 0.8814615 0.4846650
## 0.4 3 0.8 0.750 250 0.8808246 0.4741255
## 0.4 3 0.8 0.875 50 0.8852791 0.4637003
## 0.4 3 0.8 0.875 100 0.8827354 0.4756585
## 0.4 3 0.8 0.875 150 0.8865632 0.4854230
## 0.4 3 0.8 0.875 200 0.8827354 0.4750055
## 0.4 3 0.8 0.875 250 0.8814615 0.4770945
## 0.4 3 0.8 1.000 50 0.8859262 0.4676501
## 0.4 3 0.8 1.000 100 0.8763660 0.4352504
## 0.4 3 0.8 1.000 150 0.8789117 0.4504320
## 0.4 3 0.8 1.000 200 0.8757270 0.4398793
## 0.4 3 0.8 1.000 250 0.8744490 0.4376969
## 0.4 4 0.6 0.500 50 0.8795568 0.4583321
## 0.4 4 0.6 0.500 100 0.8814656 0.4726724
## 0.4 4 0.6 0.500 150 0.8801856 0.4707612
## 0.4 4 0.6 0.500 200 0.8801815 0.4723852
## 0.4 4 0.6 0.500 250 0.8808185 0.4777287
## 0.4 4 0.6 0.625 50 0.8859262 0.4857153
## 0.4 4 0.6 0.625 100 0.8827313 0.4743800
## 0.4 4 0.6 0.625 150 0.8859120 0.4951494
## 0.4 4 0.6 0.625 200 0.8865469 0.5005396
## 0.4 4 0.6 0.625 250 0.8859160 0.5068756
## 0.4 4 0.6 0.750 50 0.8821005 0.4736556
## 0.4 4 0.6 0.750 100 0.8840113 0.4827349
## 0.4 4 0.6 0.750 150 0.8840113 0.4908930
## 0.4 4 0.6 0.750 200 0.8846523 0.4871538
## 0.4 4 0.6 0.750 250 0.8884821 0.5080462
## 0.4 4 0.6 0.875 50 0.8801836 0.4579866
## 0.4 4 0.6 0.875 100 0.8827252 0.4722261
## 0.4 4 0.6 0.875 150 0.8808164 0.4684492
## 0.4 4 0.6 0.875 200 0.8833662 0.4861884
## 0.4 4 0.6 0.875 250 0.8865550 0.4975602
## 0.4 4 0.6 1.000 50 0.8840174 0.4774107
## 0.4 4 0.6 1.000 100 0.8757209 0.4422266
## 0.4 4 0.6 1.000 150 0.8795446 0.4657512
## 0.4 4 0.6 1.000 200 0.8776317 0.4633426
## 0.4 4 0.6 1.000 250 0.8789076 0.4708957
## 0.4 4 0.8 0.500 50 0.8782809 0.4644463
## 0.4 4 0.8 0.500 100 0.8820964 0.4734484
## 0.4 4 0.8 0.500 150 0.8757250 0.4549266
## 0.4 4 0.8 0.500 200 0.8731772 0.4457879
## 0.4 4 0.8 0.500 250 0.8795446 0.4679524
## 0.4 4 0.8 0.625 50 0.8903767 0.5090175
## 0.4 4 0.8 0.625 100 0.8903746 0.5130431
## 0.4 4 0.8 0.625 150 0.8859160 0.4964367
## 0.4 4 0.8 0.625 200 0.8827252 0.4807698
## 0.4 4 0.8 0.625 250 0.8820903 0.4839194
## 0.4 4 0.8 0.750 50 0.8827354 0.4538944
## 0.4 4 0.8 0.750 100 0.8871960 0.4929340
## 0.4 4 0.8 0.750 150 0.8852852 0.4770518
## 0.4 4 0.8 0.750 200 0.8840093 0.4816522
## 0.4 4 0.8 0.750 250 0.8846483 0.4817025
## 0.4 4 0.8 0.875 50 0.8846503 0.4770389
## 0.4 4 0.8 0.875 100 0.8846462 0.4771880
## 0.4 4 0.8 0.875 150 0.8820985 0.4690039
## 0.4 4 0.8 0.875 200 0.8840032 0.4855191
## 0.4 4 0.8 0.875 250 0.8865509 0.4936460
## 0.4 4 0.8 1.000 50 0.8833703 0.4661730
## 0.4 4 0.8 1.000 100 0.8852771 0.4783444
## 0.4 4 0.8 1.000 150 0.8846401 0.4761695
## 0.4 4 0.8 1.000 200 0.8814554 0.4672024
## 0.4 4 0.8 1.000 250 0.8846422 0.4810287
## 0.4 5 0.6 0.500 50 0.8763741 0.4528014
## 0.4 5 0.6 0.500 100 0.8814656 0.4753982
## 0.4 5 0.6 0.500 150 0.8795548 0.4585240
## 0.4 5 0.6 0.500 200 0.8789117 0.4620105
## 0.4 5 0.6 0.500 250 0.8814635 0.4751659
## 0.4 5 0.6 0.625 50 0.8789097 0.4592301
## 0.4 5 0.6 0.625 100 0.8801815 0.4554747
## 0.4 5 0.6 0.625 150 0.8808144 0.4730219
## 0.4 5 0.6 0.625 200 0.8840032 0.4832253
## 0.4 5 0.6 0.625 250 0.8801795 0.4708314
## 0.4 5 0.6 0.750 50 0.8827211 0.4660527
## 0.4 5 0.6 0.750 100 0.8865550 0.4899385
## 0.4 5 0.6 0.750 150 0.8891028 0.5018053
## 0.4 5 0.6 0.750 200 0.8890987 0.4991246
## 0.4 5 0.6 0.750 250 0.8859120 0.4897526
## 0.4 5 0.6 0.875 50 0.8789056 0.4437569
## 0.4 5 0.6 0.875 100 0.8782626 0.4355743
## 0.4 5 0.6 0.875 150 0.8801774 0.4481918
## 0.4 5 0.6 0.875 200 0.8833601 0.4653886
## 0.4 5 0.6 0.875 250 0.8814534 0.4586813
## 0.4 5 0.6 1.000 50 0.8789056 0.4463260
## 0.4 5 0.6 1.000 100 0.8808103 0.4537012
## 0.4 5 0.6 1.000 150 0.8820801 0.4615110
## 0.4 5 0.6 1.000 200 0.8820822 0.4584398
## 0.4 5 0.6 1.000 250 0.8808103 0.4514820
## 0.4 5 0.8 0.500 50 0.8903807 0.5084045
## 0.4 5 0.8 0.500 100 0.8859201 0.4896653
## 0.4 5 0.8 0.500 150 0.8840072 0.4898299
## 0.4 5 0.8 0.500 200 0.8820964 0.4812064
## 0.4 5 0.8 0.500 250 0.8840072 0.4913145
## 0.4 5 0.8 0.625 50 0.8846523 0.4770083
## 0.4 5 0.8 0.625 100 0.8827374 0.4636750
## 0.4 5 0.8 0.625 150 0.8859262 0.4914480
## 0.4 5 0.8 0.625 200 0.8871940 0.4904907
## 0.4 5 0.8 0.625 250 0.8827334 0.4772702
## 0.4 5 0.8 0.750 50 0.8840093 0.4658642
## 0.4 5 0.8 0.750 100 0.8827374 0.4660067
## 0.4 5 0.8 0.750 150 0.8897418 0.4942577
## 0.4 5 0.8 0.750 200 0.8922956 0.5132319
## 0.4 5 0.8 0.750 250 0.8916567 0.5169039
## 0.4 5 0.8 0.875 50 0.8776276 0.4463162
## 0.4 5 0.8 0.875 100 0.8776317 0.4522553
## 0.4 5 0.8 0.875 150 0.8782666 0.4531335
## 0.4 5 0.8 0.875 200 0.8769907 0.4491802
## 0.4 5 0.8 0.875 250 0.8769907 0.4515099
## 0.4 5 0.8 1.000 50 0.8916546 0.5037535
## 0.4 5 0.8 1.000 100 0.8871940 0.4935714
## 0.4 5 0.8 1.000 150 0.8884618 0.4998231
## 0.4 5 0.8 1.000 200 0.8865509 0.4930364
## 0.4 5 0.8 1.000 250 0.8846381 0.4873438
##
## Tuning parameter 'gamma' was held constant at a value of 0
## Tuning
## parameter 'min_child_weight' was held constant at a value of 1
## Accuracy was used to select the optimal model using the largest value.
## The final values used for the model were nrounds = 100, max_depth = 1, eta
## = 0.4, gamma = 0, colsample_bytree = 0.8, min_child_weight = 1 and subsample
## = 0.875.
plot(model_class_xg)
We we will calculate the confusion matrix to evaluate the performance of the models
predicted_model_class_xg <- predict(model_class_xg, test_class_data_encoded)
cm_model_class_xg <- confusionMatrix(reference = test_class_data_encoded$Response, data = predicted_model_class_xg, mode='everything', positive= '1')
print(cm_model_class_xg)
## Confusion Matrix and Statistics
##
## Reference
## Prediction 0 1
## 0 551 50
## 1 20 50
##
## Accuracy : 0.8957
## 95% CI : (0.87, 0.9178)
## No Information Rate : 0.851
## P-Value [Acc > NIR] : 0.0004272
##
## Kappa : 0.5306
##
## Mcnemar's Test P-Value : 0.0005279
##
## Sensitivity : 0.50000
## Specificity : 0.96497
## Pos Pred Value : 0.71429
## Neg Pred Value : 0.91681
## Precision : 0.71429
## Recall : 0.50000
## F1 : 0.58824
## Prevalence : 0.14903
## Detection Rate : 0.07452
## Detection Prevalence : 0.10432
## Balanced Accuracy : 0.73249
##
## 'Positive' Class : 1
##
library(pROC)
test_target_xg <- unclass(test_class_data_encoded$Response)
pred_target_xg <- unclass(predicted_model_class_xg)
auc_model_xg <- roc(test_target_xg,pred_target_xg)
print(auc_model_xg)
##
## Call:
## roc.default(response = test_target_xg, predictor = pred_target_xg)
##
## Data: pred_target_xg in 571 controls (test_target_xg 1) < 100 cases (test_target_xg 2).
## Area under the curve: 0.7325
plot(auc_model_xg, ylim=c(0,1),xlim=c(1,0), print.thres=TRUE, main=paste('AUC:',round(auc_model_xg$auc[[1]],2)))
abline(h=1,col='blue',lwd=2)
abline(h=0,col='red',lwd=2)
For regression model training we would be using the following:
Target : NumStorePurchases (number of purchases bought directly through stores) Features : All other columns in the ‘data’ dataframe except for column ID
We would perform the model training using the following three regression algorithms 1) Generalised Linear Models 2) Random Forest 3) XGBoost (Extreme Gradient Boosting)
# import library for evaluation
library(Metrics)
library(ggplot2)
library(caret)
Before modelling we need to split the dataset into training and testing set
set.seed(123) # Set a random seed for reproducibility
train_reg_index <- createDataPartition(data$NumStorePurchases, p = 0.8, list = FALSE)
# we remove the classification model target and remove Total Purchase to prevent target data leakage
train_reg_data <- data[train_reg_index, ] %>% select(-Response,-TotalPurchases)
test_reg_data <- data[-train_reg_index, ] %>% select(-Response,-TotalPurchases)
#preprocess data before training to manage the diffrent data scaling and normalise data
preProcValue_reg <- preProcess(train_reg_data, method = c('center','scale','YeoJohnson'))
train_reg_data <- predict(preProcValue_reg, newdata = train_reg_data)
test_reg_data <- predict(preProcValue_reg, newdata = test_reg_data)
As there some columns such as Marital Status which are categorical we will to perform encoding to translate the information into inouts that is acceptable by the machine learning algorithms. For this project, we choose One Hot Encoding as it was a common method to encode categorical columns
dummies_reg_model <- dummyVars(NumStorePurchases ~. , data = train_reg_data )
train_reg_data_encoded <- predict(dummies_reg_model, newdata = train_reg_data) %>% as.data.frame()
test_reg_data_encoded <- predict(dummies_reg_model, newdata = test_reg_data) %>% as.data.frame()
ytrain_reg <- train_reg_data$NumStorePurchases %>% as.numeric()
ytest_reg <- test_reg_data$NumStorePurchases %>% as.numeric()
To achieve a better generalisation on the model, we would use the cross validation technique when training the model. We would set the control variable for 3 fold cross validation
fitControl_reg <- trainControl(method = 'cv', number = 5)
Firstly we would train a model using linear regression
set.seed(123)
model_reg_lr <- train(train_reg_data_encoded, ytrain_reg, trControl = fitControl_reg, method = 'glmnet', tuneLength = tune_num)
print(model_reg_lr)
## glmnet
##
## 1792 samples
## 32 predictor
##
## No pre-processing
## Resampling: Cross-Validated (5 fold)
## Summary of sample sizes: 1433, 1434, 1433, 1434, 1434
## Resampling results across tuning parameters:
##
## alpha lambda RMSE Rsquared MAE
## 0.100 0.0007397605 0.5520033 0.6959649 0.4000305
## 0.100 0.0034336640 0.5517122 0.6962337 0.3990185
## 0.100 0.0159376564 0.5522664 0.6956047 0.3988961
## 0.100 0.0739760479 0.5575874 0.6903214 0.4062899
## 0.100 0.3433663980 0.5744052 0.6790707 0.4270654
## 0.325 0.0007397605 0.5517175 0.6962486 0.4000120
## 0.325 0.0034336640 0.5512936 0.6966676 0.3984394
## 0.325 0.0159376564 0.5514851 0.6965210 0.4000902
## 0.325 0.0739760479 0.5606251 0.6882617 0.4129546
## 0.325 0.3433663980 0.6064166 0.6591300 0.4584588
## 0.550 0.0007397605 0.5517333 0.6962233 0.3997760
## 0.550 0.0034336640 0.5507964 0.6972015 0.3981929
## 0.550 0.0159376564 0.5514984 0.6965973 0.4017035
## 0.550 0.0739760479 0.5642450 0.6859263 0.4172899
## 0.550 0.3433663980 0.6324571 0.6525721 0.4870368
## 0.775 0.0007397605 0.5516304 0.6963299 0.3995677
## 0.775 0.0034336640 0.5505391 0.6974786 0.3982908
## 0.775 0.0159376564 0.5522985 0.6958255 0.4041945
## 0.775 0.0739760479 0.5704840 0.6806846 0.4238431
## 0.775 0.3433663980 0.6570954 0.6526833 0.5142494
## 1.000 0.0007397605 0.5515181 0.6964475 0.3993660
## 1.000 0.0034336640 0.5503211 0.6977125 0.3984237
## 1.000 0.0159376564 0.5542517 0.6937861 0.4071645
## 1.000 0.0739760479 0.5792857 0.6721214 0.4325911
## 1.000 0.3433663980 0.6872251 0.6491934 0.5452336
##
## RMSE was used to select the optimal model using the smallest value.
## The final values used for the model were alpha = 1 and lambda = 0.003433664.
plot(model_reg_lr)
We will calculate the evaluation metrics such as RMSE and MAE to evaluate the performance of the models
predicted_model_reg_lr <- predict(model_reg_lr, test_reg_data_encoded) %>% as.vector()
MAE_model_reg_lr <- mae(predicted_model_reg_lr, ytest_reg)
RMSE_model_reg_lr <- rmse(predicted_model_reg_lr, ytest_reg)
pred_df_model_reg_lr <- data.frame(Predicted = predicted_model_reg_lr, Observed = ytest_reg)
pred_df_plot_model_reg_lr <- ggplot(pred_df_model_reg_lr, aes(x = Predicted, y = Observed)) + geom_point() + geom_abline(intercept = 0, slope = 1, color = "blue", size = 0.5)
cat("RMSE value for the model is",RMSE_model_reg_lr,"and MAE value for the model is",MAE_model_reg_lr)
## RMSE value for the model is 0.5310193 and MAE value for the model is 0.3911684
print(pred_df_plot_model_reg_lr)
Secondly we would train a model using Random Forest
set.seed(123)
model_reg_rf <- train(train_reg_data_encoded, ytrain_reg , trControl = fitControl_reg, method = 'rf', tuneLength = tune_num)
print(model_reg_rf)
## Random Forest
##
## 1792 samples
## 32 predictor
##
## No pre-processing
## Resampling: Cross-Validated (5 fold)
## Summary of sample sizes: 1433, 1434, 1433, 1434, 1434
## Resampling results across tuning parameters:
##
## mtry RMSE Rsquared MAE
## 2 0.5044062 0.7531703 0.3763324
## 9 0.4643397 0.7872469 0.3328227
## 17 0.4608097 0.7900660 0.3258071
## 24 0.4602050 0.7904606 0.3224093
## 32 0.4607650 0.7897579 0.3206299
##
## RMSE was used to select the optimal model using the smallest value.
## The final value used for the model was mtry = 24.
plot(model_reg_rf)
We will calculate the evaluation metrics such as RMSE and MAE to evaluate the performance of the models
predicted_model_reg_rf <- predict(model_reg_rf, test_reg_data_encoded) %>% as.vector()
MAE_model_reg_rf <- mae(predicted_model_reg_rf, ytest_reg)
RMSE_model_reg_rf <- rmse(predicted_model_reg_rf, ytest_reg)
pred_df_model_reg_rf <- data.frame(Predicted = predicted_model_reg_rf, Observed = ytest_reg)
pred_df_plot_model_reg_rf <- ggplot(pred_df_model_reg_rf, aes(x = Predicted, y = Observed)) + geom_point() + geom_abline(intercept = 0, slope = 1, color = "blue", size = 0.5)
cat("RMSE value for the model is",RMSE_model_reg_rf,"and MAE value for the model is",MAE_model_reg_rf)
## RMSE value for the model is 0.4218924 and MAE value for the model is 0.2927821
print(pred_df_plot_model_reg_rf)
Thirdly we would train a model using XGBoost
set.seed(123)
model_reg_xg <- train(train_reg_data_encoded , ytrain_reg, trControl = fitControl_reg, method = 'xgbLinear', tuneLength = tune_num)
print(model_reg_xg)
## eXtreme Gradient Boosting
##
## 1792 samples
## 32 predictor
##
## No pre-processing
## Resampling: Cross-Validated (5 fold)
## Summary of sample sizes: 1433, 1434, 1433, 1434, 1434
## Resampling results across tuning parameters:
##
## lambda alpha nrounds RMSE Rsquared MAE
## 0e+00 0e+00 50 0.4758148 0.7750077 0.3144964
## 0e+00 0e+00 100 0.4760087 0.7750112 0.3095776
## 0e+00 0e+00 150 0.4758406 0.7751895 0.3081221
## 0e+00 0e+00 200 0.4759093 0.7751400 0.3077437
## 0e+00 0e+00 250 0.4759030 0.7751474 0.3076687
## 0e+00 1e-04 50 0.4715026 0.7787929 0.3146065
## 0e+00 1e-04 100 0.4724218 0.7779893 0.3109026
## 0e+00 1e-04 150 0.4722834 0.7781801 0.3095036
## 0e+00 1e-04 200 0.4723532 0.7781226 0.3091803
## 0e+00 1e-04 250 0.4723686 0.7781103 0.3091545
## 0e+00 1e-03 50 0.4758353 0.7753730 0.3162185
## 0e+00 1e-03 100 0.4753438 0.7760733 0.3104419
## 0e+00 1e-03 150 0.4756369 0.7758454 0.3092253
## 0e+00 1e-03 200 0.4756076 0.7758818 0.3087813
## 0e+00 1e-03 250 0.4755985 0.7758907 0.3087625
## 0e+00 1e-02 50 0.4827162 0.7686028 0.3180740
## 0e+00 1e-02 100 0.4824779 0.7690614 0.3121215
## 0e+00 1e-02 150 0.4824033 0.7691744 0.3107047
## 0e+00 1e-02 200 0.4824203 0.7691674 0.3104509
## 0e+00 1e-02 250 0.4824203 0.7691674 0.3104509
## 0e+00 1e-01 50 0.4713617 0.7781720 0.3124329
## 0e+00 1e-01 100 0.4709738 0.7787461 0.3061882
## 0e+00 1e-01 150 0.4708227 0.7789182 0.3050950
## 0e+00 1e-01 200 0.4708227 0.7789182 0.3050950
## 0e+00 1e-01 250 0.4708227 0.7789182 0.3050950
## 1e-04 0e+00 50 0.4784612 0.7725980 0.3168511
## 1e-04 0e+00 100 0.4792317 0.7721495 0.3115519
## 1e-04 0e+00 150 0.4790089 0.7724081 0.3100046
## 1e-04 0e+00 200 0.4790271 0.7724070 0.3095872
## 1e-04 0e+00 250 0.4790279 0.7724067 0.3095399
## 1e-04 1e-04 50 0.4703850 0.7798370 0.3140872
## 1e-04 1e-04 100 0.4717423 0.7786216 0.3091998
## 1e-04 1e-04 150 0.4718875 0.7785094 0.3080324
## 1e-04 1e-04 200 0.4719176 0.7785009 0.3077161
## 1e-04 1e-04 250 0.4719024 0.7785159 0.3076695
## 1e-04 1e-03 50 0.4764438 0.7750276 0.3164445
## 1e-04 1e-03 100 0.4777324 0.7741524 0.3124333
## 1e-04 1e-03 150 0.4776174 0.7743049 0.3107200
## 1e-04 1e-03 200 0.4776335 0.7742921 0.3103377
## 1e-04 1e-03 250 0.4776255 0.7742991 0.3103088
## 1e-04 1e-02 50 0.4819374 0.7693732 0.3177501
## 1e-04 1e-02 100 0.4806159 0.7707887 0.3119756
## 1e-04 1e-02 150 0.4808028 0.7706859 0.3105421
## 1e-04 1e-02 200 0.4807797 0.7707075 0.3102560
## 1e-04 1e-02 250 0.4807797 0.7707075 0.3102560
## 1e-04 1e-01 50 0.4693588 0.7800977 0.3104813
## 1e-04 1e-01 100 0.4691340 0.7805222 0.3038821
## 1e-04 1e-01 150 0.4689777 0.7806929 0.3027121
## 1e-04 1e-01 200 0.4689765 0.7806970 0.3027070
## 1e-04 1e-01 250 0.4689765 0.7806970 0.3027070
## 1e-03 0e+00 50 0.4742083 0.7765168 0.3156252
## 1e-03 0e+00 100 0.4748572 0.7760890 0.3099200
## 1e-03 0e+00 150 0.4749940 0.7760072 0.3085381
## 1e-03 0e+00 200 0.4749687 0.7760424 0.3080999
## 1e-03 0e+00 250 0.4749645 0.7760462 0.3080350
## 1e-03 1e-04 50 0.4795206 0.7717281 0.3180303
## 1e-03 1e-04 100 0.4798691 0.7717254 0.3126553
## 1e-03 1e-04 150 0.4798387 0.7718275 0.3113165
## 1e-03 1e-04 200 0.4798516 0.7718253 0.3109401
## 1e-03 1e-04 250 0.4798474 0.7718313 0.3109112
## 1e-03 1e-03 50 0.4754384 0.7754055 0.3172529
## 1e-03 1e-03 100 0.4761789 0.7749396 0.3118034
## 1e-03 1e-03 150 0.4761859 0.7749438 0.3102127
## 1e-03 1e-03 200 0.4762768 0.7748652 0.3098922
## 1e-03 1e-03 250 0.4762735 0.7748686 0.3098811
## 1e-03 1e-02 50 0.4800423 0.7704593 0.3158503
## 1e-03 1e-02 100 0.4803509 0.7703571 0.3101656
## 1e-03 1e-02 150 0.4801553 0.7705715 0.3087141
## 1e-03 1e-02 200 0.4801261 0.7706058 0.3084537
## 1e-03 1e-02 250 0.4801261 0.7706058 0.3084537
## 1e-03 1e-01 50 0.4693674 0.7801512 0.3098154
## 1e-03 1e-01 100 0.4688522 0.7807636 0.3034559
## 1e-03 1e-01 150 0.4687328 0.7808872 0.3024416
## 1e-03 1e-01 200 0.4687328 0.7808872 0.3024416
## 1e-03 1e-01 250 0.4687328 0.7808872 0.3024416
## 1e-02 0e+00 50 0.4781164 0.7724604 0.3154184
## 1e-02 0e+00 100 0.4782573 0.7726215 0.3091114
## 1e-02 0e+00 150 0.4780783 0.7728806 0.3078032
## 1e-02 0e+00 200 0.4780173 0.7729542 0.3074114
## 1e-02 0e+00 250 0.4780162 0.7729564 0.3073505
## 1e-02 1e-04 50 0.4773410 0.7735084 0.3144942
## 1e-02 1e-04 100 0.4776839 0.7736237 0.3090814
## 1e-02 1e-04 150 0.4780312 0.7733838 0.3080931
## 1e-02 1e-04 200 0.4781359 0.7732962 0.3077234
## 1e-02 1e-04 250 0.4781197 0.7733105 0.3076554
## 1e-02 1e-03 50 0.4783425 0.7728832 0.3177717
## 1e-02 1e-03 100 0.4763503 0.7749471 0.3112502
## 1e-02 1e-03 150 0.4764368 0.7749169 0.3101474
## 1e-02 1e-03 200 0.4763984 0.7749614 0.3098543
## 1e-02 1e-03 250 0.4764022 0.7749583 0.3098349
## 1e-02 1e-02 50 0.4852040 0.7659554 0.3192278
## 1e-02 1e-02 100 0.4851050 0.7661552 0.3125495
## 1e-02 1e-02 150 0.4851815 0.7661145 0.3113439
## 1e-02 1e-02 200 0.4851641 0.7661383 0.3111627
## 1e-02 1e-02 250 0.4851641 0.7661383 0.3111627
## 1e-02 1e-01 50 0.4697725 0.7798006 0.3099871
## 1e-02 1e-01 100 0.4693396 0.7803831 0.3036628
## 1e-02 1e-01 150 0.4694546 0.7803109 0.3028700
## 1e-02 1e-01 200 0.4694546 0.7803109 0.3028700
## 1e-02 1e-01 250 0.4694546 0.7803109 0.3028700
## 1e-01 0e+00 50 0.4732530 0.7768813 0.3180303
## 1e-01 0e+00 100 0.4735564 0.7768471 0.3107496
## 1e-01 0e+00 150 0.4734706 0.7769621 0.3093117
## 1e-01 0e+00 200 0.4733858 0.7770550 0.3088033
## 1e-01 0e+00 250 0.4733876 0.7770542 0.3087757
## 1e-01 1e-04 50 0.4789986 0.7720186 0.3207332
## 1e-01 1e-04 100 0.4783983 0.7727268 0.3148178
## 1e-01 1e-04 150 0.4781111 0.7730526 0.3131981
## 1e-01 1e-04 200 0.4781016 0.7730793 0.3128099
## 1e-01 1e-04 250 0.4781139 0.7730734 0.3127781
## 1e-01 1e-03 50 0.4802048 0.7702477 0.3192761
## 1e-01 1e-03 100 0.4819928 0.7688056 0.3151464
## 1e-01 1e-03 150 0.4818692 0.7689848 0.3137789
## 1e-01 1e-03 200 0.4818297 0.7690339 0.3132709
## 1e-01 1e-03 250 0.4818301 0.7690355 0.3132399
## 1e-01 1e-02 50 0.4724710 0.7773864 0.3151450
## 1e-01 1e-02 100 0.4722422 0.7777912 0.3080061
## 1e-01 1e-02 150 0.4721785 0.7779024 0.3063013
## 1e-01 1e-02 200 0.4722603 0.7778398 0.3059700
## 1e-01 1e-02 250 0.4722603 0.7778398 0.3059700
## 1e-01 1e-01 50 0.4614027 0.7879385 0.3044169
## 1e-01 1e-01 100 0.4608002 0.7886797 0.2991013
## 1e-01 1e-01 150 0.4610048 0.7885135 0.2985066
## 1e-01 1e-01 200 0.4610057 0.7885126 0.2985075
## 1e-01 1e-01 250 0.4610057 0.7885126 0.2985075
##
## Tuning parameter 'eta' was held constant at a value of 0.3
## RMSE was used to select the optimal model using the smallest value.
## The final values used for the model were nrounds = 100, lambda = 0.1, alpha
## = 0.1 and eta = 0.3.
plot(model_reg_xg)
We will calculate the evaluation metrics such as RMSE and MAE to evaluate the performance of the models
predicted_model_reg_xg <- predict(model_reg_xg, train_reg_data_encoded) %>% as.vector()
MAE_model_reg_xg <- mae(predicted_model_reg_xg, ytest_reg)
RMSE_model_reg_xg <- rmse(predicted_model_reg_xg, ytest_reg)
pred_df_model_reg_xg <- data.frame(Predicted = predicted_model_reg_xg, Observed = ytest_reg)
pred_df_plot_model_reg_xg <- ggplot(pred_df_model_reg_xg, aes(x = Predicted, y = Observed)) + geom_point() + geom_abline(intercept = 0, slope = 1, color = "blue", size = 0.5)
cat("RMSE value for the model is",RMSE_model_reg_xg,"and MAE value for the model is",MAE_model_reg_xg)
## RMSE value for the model is 1.44003 and MAE value for the model is 1.157122
print(pred_df_plot_model_reg_xg)
# Compare model performances using resample()
models_class_compare <- resamples(list(Class_model_LR=model_class_lr, Class_model_RF=model_class_rf, Class_model_XG=model_class_xg))
# Summary of the models performances
summary(models_class_compare)
##
## Call:
## summary.resamples(object = models_class_compare)
##
## Models: Class_model_LR, Class_model_RF, Class_model_XG
## Number of resamples: 5
##
## Accuracy
## Min. 1st Qu. Median Mean 3rd Qu. Max. NA's
## Class_model_LR 0.8917197 0.8949045 0.9044586 0.9018518 0.9073482 0.9108280 0
## Class_model_RF 0.8566879 0.8885350 0.8885350 0.8859262 0.8917197 0.9041534 0
## Class_model_XG 0.8757962 0.8945687 0.8949045 0.8986590 0.9140127 0.9140127 0
##
## Kappa
## Min. 1st Qu. Median Mean 3rd Qu. Max. NA's
## Class_model_LR 0.4087939 0.5138817 0.5246353 0.5215486 0.5443959 0.6160363 0
## Class_model_RF 0.2948398 0.4854387 0.5052368 0.4725744 0.5334918 0.5438648 0
## Class_model_XG 0.3607892 0.4815540 0.5682140 0.5216626 0.5947031 0.6030527 0
# Draw box plots to compare models
scales <- list(x=list(relation="free"), y=list(relation="free"))
bwplot(models_class_compare, scales=scales)
# Compare model performances using resample()
models_reg_compare <- resamples(list(Reg_model_LR=model_reg_lr, Reg_model_RF=model_reg_rf, Reg_model_XG=model_reg_xg))
# Summary of the models performances
summary(models_reg_compare)
##
## Call:
## summary.resamples(object = models_reg_compare)
##
## Models: Reg_model_LR, Reg_model_RF, Reg_model_XG
## Number of resamples: 5
##
## MAE
## Min. 1st Qu. Median Mean 3rd Qu. Max. NA's
## Reg_model_LR 0.3815123 0.3957786 0.3985945 0.3984237 0.4035077 0.4127253 0
## Reg_model_RF 0.3009488 0.3048611 0.3343564 0.3224093 0.3344527 0.3374277 0
## Reg_model_XG 0.2878518 0.2887132 0.3001687 0.2991013 0.3093112 0.3094616 0
##
## RMSE
## Min. 1st Qu. Median Mean 3rd Qu. Max. NA's
## Reg_model_LR 0.5242501 0.5311280 0.5466092 0.5503211 0.5714119 0.5782062 0
## Reg_model_RF 0.4258321 0.4376468 0.4669112 0.4602050 0.4816229 0.4890118 0
## Reg_model_XG 0.4254316 0.4502910 0.4528905 0.4608002 0.4836984 0.4916895 0
##
## Rsquared
## Min. 1st Qu. Median Mean 3rd Qu. Max. NA's
## Reg_model_LR 0.6802060 0.6884836 0.6918154 0.6977125 0.7078001 0.7202571 0
## Reg_model_RF 0.7667545 0.7761271 0.7890160 0.7904606 0.8006387 0.8197669 0
## Reg_model_XG 0.7633222 0.7816011 0.7883903 0.7886797 0.7937976 0.8162870 0
# Draw box plots to compare models
scales <- list(x=list(relation="free"), y=list(relation="free"))
bwplot(models_reg_compare, scales=scales)
Two objectives were set for this project:
Classification Modeling: Determine if a customer would accept the next marketing campaign. This binary classification task will help us understand the factors that influence the effectiveness of marketing campaigns.
Regression Modeling: Predict number of purchases for each customer at physical stores. By identifying customers who are likely to make purchases at physical stores, more effective marketing campaigns can be generated by targeting the correct customer interest.
The predicted outcome for Logistic Regression showed the highest accuracy score, at 0.902 mean accuracy for 5-fold cross validation. Similarly, the Kappa score for the Logistic Regression model was an extremely close second (0.5215) to XGBoost (0.5217). As accuracy is our main performance metric while Kappa serves as the measurement of non-random occurrence due to the imbalanced distribution (0.1491:0.8509) of our target variable “response”, Logistic Regression is selected as the top performing classification model.
The Random Forest model appears to be the lowest-performing, yet showed valuable information regarding the feature importance of predictor variables during its tuning phase. While tuning the hyperparameter of mtry, where the number of predictor variables used are checked against the accuracy score, the performance of 10 random predictors showed an accuracy of 88.592%, compared with the 34-predictor score of 88.593%. However, 34 predictors are selected as the best hyperparameter due to the higher Kappa score, suggesting better non-random prediction.
Hence, Logistic Regression is the best model for determining if a customer would accept the next marketing campaign while all 34 predictors are important, with emphasis given to the top 10:
For predicting the number of customer purchases at physical stores, Random Forest and XGBoost performed neck-to-neck in terms of RMSE and Rsquared metrics, with Random Forest performing marginally better for each. However, the MAE became the tie-breaker as XGBoost has a mean MAE score of 0.299 compared with Random Forest’s 0.322.
This shows that XGBoost is the overall most accurate model in terms of prediction number of customer purchases, but is affected by outliers or extreme values marginally more than Random Forest. However, the difference is insignificant as MAE is deemed the most important metrics due to the its equal treatment of error margins.
Hence, XGBoost is deemed the top-performing models and should be deployed to predict and identify customers who are most likely to make physical store purchases; and marketing strategies should be adjusted accordingly to these customers.