When attempting to understand the relationship between two variables, specifically how an outcome or dependent variable is affected by predictor or independent variables, we can utilize the method of linear regression. A linear regression model is commonly used to ascertain the value of Y based on the observed value of X. The regression line represents the average value of Y conditional on each level of X.
Linear regression models utilize this mathematical equation: \(\hat{y}\) = b0 + b1X
+ e
The values of b0 and b1 represent the model’s
intercept and slope, respectively, while error is represented by e.
\(\hat{y}\) is used to signify that it
is an estimate.
The most common practice is to choose the line that minimizes the sum of squared residuals, called the least squares line. Statistical software is commonly used to compute the least squares line. Linear regression can be performed with a single predictor or with multiple predictors, called multiple regression. A multiple regression model with many predictors can me written as: \(\hat{y}\) = b0 + b1X + b2X + b3X … + e
The quality of a linear regression model can be determined by the strength of the relationship between X and Y of the data. If the correlation is weak, then the linear regression model built will not operate any better than just guessing the average of Y. However, if the correlation is strong, the linear regression model will provide a lot more accurate predictions using the conditional mean of Y.
To develop a strong model, one must develop selection strategies to choose which predictor variables to include in the model. The model that includes all of the variables in the dataset is referred to as the full model. There are two common strategies for adding or removing predictor variables from a multiple regression model - backward elimination and forward selection.
Backward elimination starts with the full model and you remove predictor variables one at a time until the model cannot be improved anymore.
Forward selection is the opposite of backward elimination and requires predictor variables to be added to the model one at a time until the model cannot be improved anymore.
Another way to explore the relationship between variables is the use of a pairwise scatter plot which creates a grid of graphs that show the relationship between different variables. It includes a scatter plot, histogram or kernel density estimate (KDE), and the correlation value of each relationship. This method provides a quick snapshot of predictor variables of interest that you may be considering to add or eliminate from the model. To create pairwise scatter plots, you can use the package “GGally” or the Base R function, pairs()
#Code for pairwise scatter plots
# install.packages("GGally"))
#library(GGally)
#ggpairs(data, columns = c("X1", "X2", "X3"))
#X1, X2, and X3 represents the predictor variables you are interested in checking. While the use of pairwise scatter plot to explore the relationships between predictor variables is commonly used, it can become cumbersome and time consuming when checking many variables, as the plot become large and difficult to interpret. In addition, the use of pairs() with Base R can lead to errors for non-numeric columns.
Figure 1: Pariwise Scatter Plot with Multiple Predictor Variables
A method to speed up the exploring of correlation among predictor and outcome variables for a model is the use of the package, correlationfunnel. This package allows the process of determining relationships between variables to become more concise, interactive, and less time consuming. The correlationfunnel allows you to determine the relationship of the predictor variables to the outcome of focus. Correlationfunnel speeds up exploratory data analysis and improves predictor variable selection.
# Install the required libraries
# install.packages(pkgs = c("correlationfunnel", "dplyr"))
# Import packages
library(correlationfunnel)
library(dplyr)We will utilize the mtcars dataset. You will not need to upload an external file as it is a built-in dataset that is included with R.
## Rows: 32
## Columns: 11
## $ mpg <dbl> 21.0, 21.0, 22.8, 21.4, 18.7, 18.1, 14.3, 24.4, 22.8, 19.2, 17.8,…
## $ cyl <dbl> 6, 6, 4, 6, 8, 6, 8, 4, 4, 6, 6, 8, 8, 8, 8, 8, 8, 4, 4, 4, 4, 8,…
## $ disp <dbl> 160.0, 160.0, 108.0, 258.0, 360.0, 225.0, 360.0, 146.7, 140.8, 16…
## $ hp <dbl> 110, 110, 93, 110, 175, 105, 245, 62, 95, 123, 123, 180, 180, 180…
## $ drat <dbl> 3.90, 3.90, 3.85, 3.08, 3.15, 2.76, 3.21, 3.69, 3.92, 3.92, 3.92,…
## $ wt <dbl> 2.620, 2.875, 2.320, 3.215, 3.440, 3.460, 3.570, 3.190, 3.150, 3.…
## $ qsec <dbl> 16.46, 17.02, 18.61, 19.44, 17.02, 20.22, 15.84, 20.00, 22.90, 18…
## $ vs <dbl> 0, 0, 1, 1, 0, 1, 0, 1, 1, 1, 1, 0, 0, 0, 0, 0, 0, 1, 1, 1, 1, 0,…
## $ am <dbl> 1, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 1, 1, 0, 0,…
## $ gear <dbl> 4, 4, 4, 3, 3, 3, 3, 4, 4, 4, 4, 3, 3, 3, 3, 3, 3, 4, 4, 4, 3, 3,…
## $ carb <dbl> 4, 4, 1, 1, 2, 1, 4, 2, 2, 4, 4, 3, 3, 3, 4, 4, 4, 1, 2, 1, 1, 2,…
Here, we’ll show how to perform a Binary Correlation Analysis, which is the process of converting continuous (numeric) and categorical (character/factor) data to a binary (0/1) format. Creating a regression model often requires an outcome variable with many predictors, and we must determine the predictors that are related to the outcome. For this example, we will us mpg as the outcome variable and determine what predictor variables are related to mpg.
With the use of binarize() numeric variables are binned into ranges or are binned by their value, and categorical variables are converted to binary numerical format. Make sure to de-select any variables that are non-predictive.
#Create a high and low bin for mpg variable with cut-off at the median of mpg values
#De-select original mpg column
#Convert numerical and categorical mtcars data into binary format
mtcars_binarized <- mtcars %>%
mutate(mpg_bin = ifelse(mpg > median(mpg, na.rm = TRUE), "high", "low")) %>%
select(-mpg) %>%
binarize(n_bins = 4, thresh_infreq = 0.01)
#View the new data frame output with glimpse()
glimpse(mtcars_binarized)## Rows: 32
## Columns: 38
## $ cyl__4 <dbl> 0, 0, 1, 0, 0, 0, 0, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0…
## $ cyl__6 <dbl> 1, 1, 0, 1, 0, 1, 0, 0, 0, 1, 1, 0, 0, 0, 0, 0, 0…
## $ cyl__8 <dbl> 0, 0, 0, 0, 1, 0, 1, 0, 0, 0, 0, 1, 1, 1, 1, 1, 1…
## $ `disp__-Inf_120.825` <dbl> 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0…
## $ disp__120.825_196.3 <dbl> 1, 1, 0, 0, 0, 0, 0, 1, 1, 1, 1, 0, 0, 0, 0, 0, 0…
## $ disp__196.3_326 <dbl> 0, 0, 0, 1, 0, 1, 0, 0, 0, 0, 0, 1, 1, 1, 0, 0, 0…
## $ disp__326_Inf <dbl> 0, 0, 0, 0, 1, 0, 1, 0, 0, 0, 0, 0, 0, 0, 1, 1, 1…
## $ `hp__-Inf_96.5` <dbl> 0, 0, 1, 0, 0, 0, 0, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0…
## $ hp__96.5_123 <dbl> 1, 1, 0, 1, 0, 1, 0, 0, 0, 1, 1, 0, 0, 0, 0, 0, 0…
## $ hp__123_180 <dbl> 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 1, 1, 1, 0, 0, 0…
## $ hp__180_Inf <dbl> 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 1, 1, 1…
## $ `drat__-Inf_3.08` <dbl> 0, 0, 0, 1, 0, 1, 0, 0, 0, 0, 0, 1, 1, 1, 1, 1, 0…
## $ drat__3.08_3.695 <dbl> 0, 0, 0, 0, 1, 0, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0, 1…
## $ drat__3.695_3.92 <dbl> 1, 1, 1, 0, 0, 0, 0, 0, 1, 1, 1, 0, 0, 0, 0, 0, 0…
## $ drat__3.92_Inf <dbl> 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0…
## $ `wt__-Inf_2.58125` <dbl> 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0…
## $ wt__2.58125_3.325 <dbl> 1, 1, 0, 1, 0, 0, 0, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0…
## $ wt__3.325_3.61 <dbl> 0, 0, 0, 0, 1, 1, 1, 0, 0, 1, 1, 0, 0, 0, 0, 0, 0…
## $ wt__3.61_Inf <dbl> 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 1, 1, 1, 1, 1…
## $ `qsec__-Inf_16.8925` <dbl> 1, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0…
## $ qsec__16.8925_17.71 <dbl> 0, 1, 0, 0, 1, 0, 0, 0, 0, 0, 0, 1, 1, 0, 0, 0, 1…
## $ qsec__17.71_18.9 <dbl> 0, 0, 1, 0, 0, 0, 0, 0, 0, 1, 1, 0, 0, 1, 1, 1, 0…
## $ qsec__18.9_Inf <dbl> 0, 0, 0, 1, 0, 1, 0, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0…
## $ vs__0 <dbl> 1, 1, 0, 0, 1, 0, 1, 0, 0, 0, 0, 1, 1, 1, 1, 1, 1…
## $ vs__1 <dbl> 0, 0, 1, 1, 0, 1, 0, 1, 1, 1, 1, 0, 0, 0, 0, 0, 0…
## $ am__0 <dbl> 0, 0, 0, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1…
## $ am__1 <dbl> 1, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0…
## $ gear__3 <dbl> 0, 0, 0, 1, 1, 1, 1, 0, 0, 0, 0, 1, 1, 1, 1, 1, 1…
## $ gear__4 <dbl> 1, 1, 1, 0, 0, 0, 0, 1, 1, 1, 1, 0, 0, 0, 0, 0, 0…
## $ gear__5 <dbl> 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0…
## $ carb__1 <dbl> 0, 0, 1, 1, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0…
## $ carb__2 <dbl> 0, 0, 0, 0, 1, 0, 0, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0…
## $ carb__3 <dbl> 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 1, 1, 0, 0, 0…
## $ carb__4 <dbl> 1, 1, 0, 0, 0, 0, 1, 0, 0, 1, 1, 0, 0, 0, 1, 1, 1…
## $ carb__6 <dbl> 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0…
## $ carb__8 <dbl> 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0…
## $ mpg_bin__high <dbl> 1, 1, 1, 1, 0, 0, 0, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0…
## $ mpg_bin__low <dbl> 0, 0, 0, 0, 1, 1, 1, 0, 0, 1, 1, 1, 1, 1, 1, 1, 1…
The result is a data frame that has only binary data with columns representing the bins that the observations fall into. In the original dataset, there were 32 rows and 11 columns, in the binary format of the dataset, there are 32 rows, but now there are 38 columns.
Next, we’ll perform and visualize a Correlation Analysis between the outcome variable, mpg, and the other predictor variables. We will select the bin with the high mpg to determine if high mpg is related to any of the predictor variables.
#Perform correlation analysis
mtcars_correlated <- mtcars_binarized %>%
correlate(target = mpg_bin__high)
mtcars_correlatedThis output shows the target variable (high mpg) and its correlation to all the other predictor variables. Next, we will visualize the Correlation Funnel. To zoom into the plot and make it interactive, you can set interactive set to TRUE.
#Perform correlation visualization
mtcars_correlated %>%
plot_correlation_funnel(interactive = FALSE)
A correlation funnel looks like a tornado plot with the highest
correlation predictor variables at the top, and the lowers correlation
predictor variables at the bottom.
Lastly, we will interpret the results of the Correlation Funnel. We can focus on the predictor variables toward the top half as they have the most correlation with the predictor variable, high mpg.
#Filter the desired variables with highest correlation and visualize the correlation funnel
mtcars_correlated %>%
filter(feature %in% c("cyl", "gear", "am", "disp", "hp")) %>%
plot_correlation_funnel(interactive = FALSE)
With this visualization, we can see that certain bins within the
predictor variables are a lot more correlated with the outcome variable,
high mpg:
To conclude, we can now use the findings from the correlation analysis to determine which predictor variables to include in a model, dependent on the outcome variable of focus. Correlation analysis and visualization speeds up the process of choosing which predictor variables to include in the model versus removing or adding a predictor variable one by one, or using the pairwise scatterplot that may become too large to interpret adequately. The correlation funnel would especially work for large datasets with many variables of interest.
These are additional resources for the application of correlationfunnel
Tutorial of how to create a correlation funnel. Hyperlink Text
How to explore and analyze a dataset with correlation funnel. Hyperlink Text
Application of correlation funnel for business science. Hyperlink Text
This code through references and cites the following sources:
Çetinkaya-Rundel, M. & Hardin, J. (2024). Source I. Hyperlink Text
Çetinkaya-Rundel, M. & Hardin, J. (2024). Source II. Hyperlink Text
Data Science for Public Service (n.d.). Source III. Hyperlink Text
Datanovia (2026). Source IV. Hyperlink Text