Introduction

Introduction to Linear Regression

When attempting to understand the relationship between two variables, specifically how an outcome or dependent variable is affected by predictor or independent variables, we can utilize the method of linear regression. A linear regression model is commonly used to ascertain the value of Y based on the observed value of X. The regression line represents the average value of Y conditional on each level of X.

Linear regression models utilize this mathematical equation: \(\hat{y}\) = b0 + b1X + e
The values of b0 and b1 represent the model’s intercept and slope, respectively, while error is represented by e. \(\hat{y}\) is used to signify that it is an estimate.

The most common practice is to choose the line that minimizes the sum of squared residuals, called the least squares line. Statistical software is commonly used to compute the least squares line. Linear regression can be performed with a single predictor or with multiple predictors, called multiple regression. A multiple regression model with many predictors can me written as: \(\hat{y}\) = b0 + b1X + b2X + b3X … + e

The quality of a linear regression model can be determined by the strength of the relationship between X and Y of the data. If the correlation is weak, then the linear regression model built will not operate any better than just guessing the average of Y. However, if the correlation is strong, the linear regression model will provide a lot more accurate predictions using the conditional mean of Y.

Correlation can take on values between -1 and 1, is unit less, and can quantify the strength of a linear trend. Correlation describes the strength and direction of a linear relationship between two variables. Correlation is commonly denoted by r.

Methods to Select Predictor Variables

To develop a strong model, one must develop selection strategies to choose which predictor variables to include in the model. The model that includes all of the variables in the dataset is referred to as the full model. There are two common strategies for adding or removing predictor variables from a multiple regression model - backward elimination and forward selection.

  • Backward elimination starts with the full model and you remove predictor variables one at a time until the model cannot be improved anymore.

  • Forward selection is the opposite of backward elimination and requires predictor variables to be added to the model one at a time until the model cannot be improved anymore.

Another way to explore the relationship between variables is the use of a pairwise scatter plot which creates a grid of graphs that show the relationship between different variables. It includes a scatter plot, histogram or kernel density estimate (KDE), and the correlation value of each relationship. This method provides a quick snapshot of predictor variables of interest that you may be considering to add or eliminate from the model. To create pairwise scatter plots, you can use the package “GGally” or the Base R function, pairs()

#Code for pairwise scatter plots 

# install.packages("GGally"))
#library(GGally) 
#ggpairs(data, columns = c("X1", "X2", "X3"))

#X1, X2, and X3 represents the predictor variables you are interested in checking. 

While the use of pairwise scatter plot to explore the relationships between predictor variables is commonly used, it can become cumbersome and time consuming when checking many variables, as the plot become large and difficult to interpret. In addition, the use of pairs() with Base R can lead to errors for non-numeric columns.

Figure 1: Pariwise Scatter Plot with Multiple Predictor Variables Pairwise scatter plot for mtcars dataset

Introduction to correlationfunnel

A method to speed up the exploring of correlation among predictor and outcome variables for a model is the use of the package, correlationfunnel. This package allows the process of determining relationships between variables to become more concise, interactive, and less time consuming. The correlationfunnel allows you to determine the relationship of the predictor variables to the outcome of focus. Correlationfunnel speeds up exploratory data analysis and improves predictor variable selection.

Learning Objectives

  • First, we will install and load the package, correlationfunnel and dplyr, and then bring up the dataset, mtcars.
  • Second, we will learn to perform a binary correlation analysis of the dataset to prepare it for the correlation analysis.
  • Third, we will perform a correlation analysis and visualize the correlation funnel.
  • And, lastly, we will learn to interpret the results of the correlation funnel to make insights into model selection.



Step 1: Install & Load Required Libraries & Data

# Install the required libraries
# install.packages(pkgs = c("correlationfunnel", "dplyr"))

# Import packages

library(correlationfunnel)
library(dplyr)

We will utilize the mtcars dataset. You will not need to upload an external file as it is a built-in dataset that is included with R.

#Load the built-in dataset

data("mtcars")

#View the dataset with glimpse()

glimpse(mtcars)
## Rows: 32
## Columns: 11
## $ mpg  <dbl> 21.0, 21.0, 22.8, 21.4, 18.7, 18.1, 14.3, 24.4, 22.8, 19.2, 17.8,…
## $ cyl  <dbl> 6, 6, 4, 6, 8, 6, 8, 4, 4, 6, 6, 8, 8, 8, 8, 8, 8, 4, 4, 4, 4, 8,…
## $ disp <dbl> 160.0, 160.0, 108.0, 258.0, 360.0, 225.0, 360.0, 146.7, 140.8, 16…
## $ hp   <dbl> 110, 110, 93, 110, 175, 105, 245, 62, 95, 123, 123, 180, 180, 180…
## $ drat <dbl> 3.90, 3.90, 3.85, 3.08, 3.15, 2.76, 3.21, 3.69, 3.92, 3.92, 3.92,…
## $ wt   <dbl> 2.620, 2.875, 2.320, 3.215, 3.440, 3.460, 3.570, 3.190, 3.150, 3.…
## $ qsec <dbl> 16.46, 17.02, 18.61, 19.44, 17.02, 20.22, 15.84, 20.00, 22.90, 18…
## $ vs   <dbl> 0, 0, 1, 1, 0, 1, 0, 1, 1, 1, 1, 0, 0, 0, 0, 0, 0, 1, 1, 1, 1, 0,…
## $ am   <dbl> 1, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 1, 1, 0, 0,…
## $ gear <dbl> 4, 4, 4, 3, 3, 3, 3, 4, 4, 4, 4, 3, 3, 3, 3, 3, 3, 4, 4, 4, 3, 3,…
## $ carb <dbl> 4, 4, 1, 1, 2, 1, 4, 2, 2, 4, 4, 3, 3, 3, 4, 4, 4, 1, 2, 1, 1, 2,…


Step 2: Perform a Binary Correlation Analysis


Here, we’ll show how to perform a Binary Correlation Analysis, which is the process of converting continuous (numeric) and categorical (character/factor) data to a binary (0/1) format. Creating a regression model often requires an outcome variable with many predictors, and we must determine the predictors that are related to the outcome. For this example, we will us mpg as the outcome variable and determine what predictor variables are related to mpg.

With the use of binarize() numeric variables are binned into ranges or are binned by their value, and categorical variables are converted to binary numerical format. Make sure to de-select any variables that are non-predictive.

#Create a high and low bin for mpg variable with cut-off at the median of mpg values
#De-select original mpg column 
#Convert numerical and categorical mtcars data into binary format

mtcars_binarized <- mtcars %>%
  mutate(mpg_bin = ifelse(mpg > median(mpg, na.rm = TRUE), "high", "low")) %>%
  select(-mpg) %>%
  binarize(n_bins = 4, thresh_infreq = 0.01)
  
#View the new data frame output with glimpse()

glimpse(mtcars_binarized)
## Rows: 32
## Columns: 38
## $ cyl__4               <dbl> 0, 0, 1, 0, 0, 0, 0, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0…
## $ cyl__6               <dbl> 1, 1, 0, 1, 0, 1, 0, 0, 0, 1, 1, 0, 0, 0, 0, 0, 0…
## $ cyl__8               <dbl> 0, 0, 0, 0, 1, 0, 1, 0, 0, 0, 0, 1, 1, 1, 1, 1, 1…
## $ `disp__-Inf_120.825` <dbl> 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0…
## $ disp__120.825_196.3  <dbl> 1, 1, 0, 0, 0, 0, 0, 1, 1, 1, 1, 0, 0, 0, 0, 0, 0…
## $ disp__196.3_326      <dbl> 0, 0, 0, 1, 0, 1, 0, 0, 0, 0, 0, 1, 1, 1, 0, 0, 0…
## $ disp__326_Inf        <dbl> 0, 0, 0, 0, 1, 0, 1, 0, 0, 0, 0, 0, 0, 0, 1, 1, 1…
## $ `hp__-Inf_96.5`      <dbl> 0, 0, 1, 0, 0, 0, 0, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0…
## $ hp__96.5_123         <dbl> 1, 1, 0, 1, 0, 1, 0, 0, 0, 1, 1, 0, 0, 0, 0, 0, 0…
## $ hp__123_180          <dbl> 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 1, 1, 1, 0, 0, 0…
## $ hp__180_Inf          <dbl> 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 1, 1, 1…
## $ `drat__-Inf_3.08`    <dbl> 0, 0, 0, 1, 0, 1, 0, 0, 0, 0, 0, 1, 1, 1, 1, 1, 0…
## $ drat__3.08_3.695     <dbl> 0, 0, 0, 0, 1, 0, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0, 1…
## $ drat__3.695_3.92     <dbl> 1, 1, 1, 0, 0, 0, 0, 0, 1, 1, 1, 0, 0, 0, 0, 0, 0…
## $ drat__3.92_Inf       <dbl> 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0…
## $ `wt__-Inf_2.58125`   <dbl> 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0…
## $ wt__2.58125_3.325    <dbl> 1, 1, 0, 1, 0, 0, 0, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0…
## $ wt__3.325_3.61       <dbl> 0, 0, 0, 0, 1, 1, 1, 0, 0, 1, 1, 0, 0, 0, 0, 0, 0…
## $ wt__3.61_Inf         <dbl> 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 1, 1, 1, 1, 1…
## $ `qsec__-Inf_16.8925` <dbl> 1, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0…
## $ qsec__16.8925_17.71  <dbl> 0, 1, 0, 0, 1, 0, 0, 0, 0, 0, 0, 1, 1, 0, 0, 0, 1…
## $ qsec__17.71_18.9     <dbl> 0, 0, 1, 0, 0, 0, 0, 0, 0, 1, 1, 0, 0, 1, 1, 1, 0…
## $ qsec__18.9_Inf       <dbl> 0, 0, 0, 1, 0, 1, 0, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0…
## $ vs__0                <dbl> 1, 1, 0, 0, 1, 0, 1, 0, 0, 0, 0, 1, 1, 1, 1, 1, 1…
## $ vs__1                <dbl> 0, 0, 1, 1, 0, 1, 0, 1, 1, 1, 1, 0, 0, 0, 0, 0, 0…
## $ am__0                <dbl> 0, 0, 0, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1…
## $ am__1                <dbl> 1, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0…
## $ gear__3              <dbl> 0, 0, 0, 1, 1, 1, 1, 0, 0, 0, 0, 1, 1, 1, 1, 1, 1…
## $ gear__4              <dbl> 1, 1, 1, 0, 0, 0, 0, 1, 1, 1, 1, 0, 0, 0, 0, 0, 0…
## $ gear__5              <dbl> 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0…
## $ carb__1              <dbl> 0, 0, 1, 1, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0…
## $ carb__2              <dbl> 0, 0, 0, 0, 1, 0, 0, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0…
## $ carb__3              <dbl> 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 1, 1, 0, 0, 0…
## $ carb__4              <dbl> 1, 1, 0, 0, 0, 0, 1, 0, 0, 1, 1, 0, 0, 0, 1, 1, 1…
## $ carb__6              <dbl> 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0…
## $ carb__8              <dbl> 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0…
## $ mpg_bin__high        <dbl> 1, 1, 1, 1, 0, 0, 0, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0…
## $ mpg_bin__low         <dbl> 0, 0, 0, 0, 1, 1, 1, 0, 0, 1, 1, 1, 1, 1, 1, 1, 1…

The result is a data frame that has only binary data with columns representing the bins that the observations fall into. In the original dataset, there were 32 rows and 11 columns, in the binary format of the dataset, there are 32 rows, but now there are 38 columns.


Step 3: Perform & Visualize a Correlation Funnel


Next, we’ll perform and visualize a Correlation Analysis between the outcome variable, mpg, and the other predictor variables. We will select the bin with the high mpg to determine if high mpg is related to any of the predictor variables.

#Perform correlation analysis

mtcars_correlated <- mtcars_binarized %>%
  correlate(target = mpg_bin__high)
            
mtcars_correlated

This output shows the target variable (high mpg) and its correlation to all the other predictor variables. Next, we will visualize the Correlation Funnel. To zoom into the plot and make it interactive, you can set interactive set to TRUE.

#Perform correlation visualization

mtcars_correlated %>% 
  plot_correlation_funnel(interactive = FALSE)

A correlation funnel looks like a tornado plot with the highest correlation predictor variables at the top, and the lowers correlation predictor variables at the bottom.


Step 4: Interpret the Correlation Funnel


Lastly, we will interpret the results of the Correlation Funnel. We can focus on the predictor variables toward the top half as they have the most correlation with the predictor variable, high mpg.

#Filter the desired variables with highest correlation and visualize the correlation funnel  

mtcars_correlated %>%
  filter(feature %in% c("cyl", "gear", "am", "disp", "hp")) %>%
  plot_correlation_funnel(interactive = FALSE)

With this visualization, we can see that certain bins within the predictor variables are a lot more correlated with the outcome variable, high mpg:

  • cyl: Higher correlation between a 4 cylinder vehicle than when it is 8 cylinder.
  • gear: Higher correlation with 4 gear vehicle than 3 gear vehicle.
  • am: Higher correlation for manual vehicles.
  • disp: High correlation for displacement (in cu.in.) of 0 to 120.825.
  • hp: High correlation for horsepower of 0 to 96.5.



Conclusion

To conclude, we can now use the findings from the correlation analysis to determine which predictor variables to include in a model, dependent on the outcome variable of focus. Correlation analysis and visualization speeds up the process of choosing which predictor variables to include in the model versus removing or adding a predictor variable one by one, or using the pairwise scatterplot that may become too large to interpret adequately. The correlation funnel would especially work for large datasets with many variables of interest.

Resources

These are additional resources for the application of correlationfunnel

  • Tutorial of how to create a correlation funnel. Hyperlink Text

  • How to explore and analyze a dataset with correlation funnel. Hyperlink Text

  • Application of correlation funnel for business science. Hyperlink Text

Works Cited

This code through references and cites the following sources: