Getting Started with FinalProject13

Braxton Stacey, Joshua Fahlgren, Thomas Jackson

Installation

Our package can be installed from our public github repository using the function install_github from the package devtools.

library(devtools)
#> Loading required package: usethis
install_github("AU-R-Programming/FinalProject13")
#> Skipping install of 'FinalProject13' from a github remote, the SHA1 (e8aa8692) has not changed since last install.
#>   Use `force = TRUE` to force installation
library(FinalProject13)

Introduction

Welcome to the FinalProject13 package! This package implements supervised binary classification using numerical optimization techniques, with a particular focus on logistic regression. It is designed to help users efficiently estimate the coefficients of logistic regression models and generate confidence intervals using bootstrapping. Additionally, the package includes tools for model evaluation, including confusion matrix calculations and other diagnostic metrics.

Whether you are exploring relationships in your data or building robust predictive models, FinalProject13 provides a set of functions for model fitting, evaluation, and interpretation, helping you easily integrate these tools into your analytical workflow.

In this vignette, we will guide you through the primary functionalities of the package, showing you how to:

Let’s get started!

Data Preprocessing

Before we begin using the functions in the FinalProject13 package, let’s first preprocess the data. For this example, we will be using a dataset called expenses.csv, which contains various information, including smoking status, BMI, and the number of children for a set of individuals. Our goal is to process this data and prepare it for binary classification analysis.

Step 1: Loading the Data

We will begin by loading the expenses.csv file. This dataset includes multiple columns, but we are particularly interested in three columns:

For our purposes, we will create new adjusted columns based off of the previously mentioned columns.

# Load the dataset from the 'inst/extdata' directory in the package
expenses <- read.csv(system.file("extdata", "expenses.csv", package = "FinalProjectGroup13"))
head(expenses)
#>   age    sex    bmi children smoker    region   charges
#> 1  19 female 27.900        0    yes southwest 16884.924
#> 2  18   male 33.770        1     no southeast  1725.552
#> 3  28   male 33.000        3     no southeast  4449.462
#> 4  33   male 22.705        0     no northwest 21984.471
#> 5  32   male 28.880        0     no northwest  3866.855
#> 6  31 female 25.740        0     no southeast  3756.622

# Step 2: Converting Data for Use
# Set the median value for charges
median_charges <- 9382

# Create a new column indicating whether charges are greater than the median
expenses$charges_above_median <- ifelse(expenses$charges > median_charges, 1, 0)

# Create a new column for age in years older than 18
expenses$years_older_than_min <- expenses$age - 18

# Create a new column for bmi greater than min of 15.96
expenses$bmi_greater_than_min <- expenses$bmi - 15.96

# Step 3: Inspecting the Data
# Inspect the 'years_older_than_mi' and 'bmi_greater_than_min' columns for any issues
summary(expenses$years_older_than_min)
#>    Min. 1st Qu.  Median    Mean 3rd Qu.    Max. 
#>    0.00    9.00   21.00   21.21   33.00   46.00
summary(expenses$bmi_greater_than_min)
#>    Min. 1st Qu.  Median    Mean 3rd Qu.    Max. 
#>    0.00   10.34   14.44   14.70   18.73   37.17

# Step 4: Data Summary
# Summary of the processed data
summary(expenses)
#>       age            sex                 bmi           children    
#>  Min.   :18.00   Length:1338        Min.   :15.96   Min.   :0.000  
#>  1st Qu.:27.00   Class :character   1st Qu.:26.30   1st Qu.:0.000  
#>  Median :39.00   Mode  :character   Median :30.40   Median :1.000  
#>  Mean   :39.21                      Mean   :30.66   Mean   :1.095  
#>  3rd Qu.:51.00                      3rd Qu.:34.69   3rd Qu.:2.000  
#>  Max.   :64.00                      Max.   :53.13   Max.   :5.000  
#>     smoker             region             charges      charges_above_median
#>  Length:1338        Length:1338        Min.   : 1122   Min.   :0.0         
#>  Class :character   Class :character   1st Qu.: 4740   1st Qu.:0.0         
#>  Mode  :character   Mode  :character   Median : 9382   Median :0.5         
#>                                        Mean   :13270   Mean   :0.5         
#>                                        3rd Qu.:16640   3rd Qu.:1.0         
#>                                        Max.   :63770   Max.   :1.0         
#>  years_older_than_min bmi_greater_than_min
#>  Min.   : 0.00        Min.   : 0.00       
#>  1st Qu.: 9.00        1st Qu.:10.34       
#>  Median :21.00        Median :14.44       
#>  Mean   :21.21        Mean   :14.70       
#>  3rd Qu.:33.00        3rd Qu.:18.73       
#>  Max.   :46.00        Max.   :37.17

Applying Package Functions to the Data

Now that we have preprocessed the data, let’s begin applying the functions from the FinalProject13 package to perform logistic regression analysis on the data.

Step 1: Estimating Beta Coefficients Using Logistic Regression

The first step in logistic regression is to estimate the coefficients for the predictor variables. We will use the estimate_beta() function from our package to perform this estimation. The function takes a matrix of predictor variables (in our case, the years_older_than_min column and the bmi_greater_than_min column) and a binary response variable (the charges_above_median column), and returns the estimated beta coefficients.

# Ensure that both 'bmi' and 'children' columns are numeric
expenses$years_older_than_min <- as.numeric(expenses$years_older_than_min)
expenses$bmi_greater_than_min <- as.numeric(expenses$bmi_greater_than_min)

# Create the predictor matrix (bmi and children columns) as a numeric matrix
X <- as.matrix(expenses[, c("years_older_than_min", "bmi_greater_than_min")])

# Response vector (smoker column)
y <- expenses$charges_above_median

# Apply the estimate_beta function to estimate the beta coefficients
model <- estimate_beta(X, y)

# View the estimated coefficients
model$coefficients
#>                             [,1]
#>                      -2.07264785
#> years_older_than_min  0.08744715
#> bmi_greater_than_min  0.01529154

Explaining the Results

Our estimate_beta() function returned us three values:

1. The Intercept (-2.073).

This value indicates the log-odds of an individuals charges being above the median when both years_older_than_min and bmi_greater_than_min are equal to 0. In other words, it’s the base. It being large and negative indicates that an individual with the base values for the predictor variables is relatively unlikely to have their charges be above the median.

2. The Coefficient for years_older_than_min (0.0874)

This value indicates the change in the log-odds of an individuals charges being above the median for each additional year above 18. The number being positive and relatively large indicates relatively significantly that the older an individual is they are more likely to have their charges be above the median.

3. The Coefficient for years_older_than_min (0.0153)

This value indicates the change in the log-odds of an individuals charges being above the median for each additional additonal unit increase in BMI above 15.96. The number being positive and relatively large indicates relatively significantly (though not as significantly as for years_older_than_min) that the higher bmi an individual has they are more likely to have their charges be above the median.

Step 2: Bootstrapping Confidence Intervals for Beta Coefficients

Now that we have estimated the beta coefficients using logistic regression, we can use bootstrapping to estimate confidence intervals for these coefficients. Bootstrapping is a resampling technique that allows us to estimate the sampling distribution of a statistic by resampling with replacement from the original data. We can do this using the bootstrap_ci() function from our package:


# Perform bootstrap to calculate confidence intervals for the beta coefficients
bootstrap_result <- bootstrap_ci(X, y, alpha = 0.05, num_bootstrap = 20)

# Display the confidence intervals
print(bootstrap_result)
#> $lower_bound
#> [1] -2.4852720086  0.0810256769 -0.0003556741
#> 
#> $upper_bound
#> [1] -1.75035424  0.10059669  0.03740458

Explaining the Results:

Our bootstrap_ci() function yielded us with two vectors consisting of three numbers. The first vector represents the lower bound of the confident interval (in our case 95% given alpha is 0.05) for each of the estimated coefficients, while the second vector represents the upper bound.

Intercept

Our intercept confidence interval ranges from -2.3432 to -1.6474, indicating that on both the high and low end an individual with the minimum age and bmi is unlikely ot have charges above the median.

years_older_than_min

Our years_older_than_min confidence interval ranges from 0.07998 to 0.09822, indicating that this predictor is very likely to have a relatively strong positive correlation with charges_above_median as thought before.

bmi_greater_than_min

Our bmi_greater_than_min confidence interval ranges from -0.01034 to 0.02901, this indicates that our determination that bmi has a positive correlation with charges_above_median may actually be incorrect, although we cannot know for sure either way as the true coefficient could be anywhere within this range, which spans from negative to positive.

Step 3: Predicting Probabilities for Logistic Regression

After estimating the beta coefficients using the logistic regression model, the next step is to predict the probabilities of the outcome variable (whether an individual has charges above the median) for each observation in the dataset. This is done using the predict_prob() function from our package.

The predict_prob() function takes the predictor matrix \(X\) and the estimated beta coefficients, both gathered earlier, and it returns the predicted probabilities for each observation in the dataset. These predicted probabilities indicate the likelihood that each individual has charges above the median based on their age and bmi.

The formula used for prediction is the logistic function applied to the linear predictor:

\[ p_i = \frac{1}{1 + e^{-(\beta_0 + \beta_1 \cdot \text{BMI})}} \]

where: - \(p_i\) is the predicted probability of being a smoker for the \(i\)-th observation, - \(\beta_0\) is the intercept, - \(\beta_1\) is the coefficient for BMI, - and \(\text{BMI}\) is the predictor variable.

Using the estimated coefficients, the predict_prob() function computes the probability for each observation in the dataset, which can then be used for classification or further model evaluation.

# Apply the predict_prob function to generate predicted probabilities
predictions <- predict_prob(X, model$coefficients)

# View the predicted probabilities for the first few observations
head(predictions)
#>           [,1]
#> [1,] 0.1415325
#> [2,] 0.1418139
#> [3,] 0.2813837
#> [4,] 0.3412342
#> [5,] 0.3428045
#> [6,] 0.3129672

Explaining the Results:

Our predict_prob() function yielded us with predicted probabilities for each individual in the dataset. Looking at the first few, we can see that the prediction function is working as expected. For example, the first individual is predicted to be very unlikely to have charges above the median as they have a predicted probability of 14%. This lines up with our coefficients, as, looking at our dataset, the first individual is young with a low bmi. In the other direction, we can look at the highest probability of having charges above the median in the group, which is the fifth individual. This lines up with our coefficients, as, looking at our dataset, the individual has the second highest age of our group of 6 and the third highest bmi.

Explaining the Results:

Confusion Matrix

In our confusion matrix we can see that our model is generating more correct outcomes than incorrect outomes, indicating positive things about our model.

Prevalence

Our value here is 0.5. This indicates that half of the observations in your dataset have charges greater than the median, which makes sense considering the nature of our variable.

Accuracy

Our value here is 71.23%. This means that the model is correctly classifying around 71% of the cases, which indicates positive things about our model

Sensitivity

Our value here is 71.60%. This means the model is identifying about 71.6% of the true positives, which indicates positive things about our model.

Specificity

Our value here is 70.85%, meaning the model is correctly identifying around 70.85% of the true negatives, which indicates positive things about our model.

False Discovery Rate

Our value here is 28.93%, indicating that only about 29% of the predicted positive cases were actually false positives, which indicates positive things about our model.

Diagnostic Odds Ratio

Our value here is 6.13, suggesting that the model has some discriminatory power between the classes (charges above and below the median), which indicates positive things about our model.

Summary

All metrics considered, it appears our model is strong and that the coefficients of age and bmi can be used together effectively to predict whether an individual has charges above the median.

In Closing

Through this process, the merits of the FinalProject13 package become clear. We were able to estimate coefficients from a dataset using estimate_beta, determine the confidence interval for the coefficients using bootstrap_ci, create and use a prediction model on the dataset using the coefficients using predict_prob, and finally we were able to determine the quality of that model using confusion_metrics.

References