Our package can be installed from our public github repository using
the function install_github from the package
devtools.
Welcome to the FinalProject13 package! This package implements supervised binary classification using numerical optimization techniques, with a particular focus on logistic regression. It is designed to help users efficiently estimate the coefficients of logistic regression models and generate confidence intervals using bootstrapping. Additionally, the package includes tools for model evaluation, including confusion matrix calculations and other diagnostic metrics.
Whether you are exploring relationships in your data or building robust predictive models, FinalProject13 provides a set of functions for model fitting, evaluation, and interpretation, helping you easily integrate these tools into your analytical workflow.
In this vignette, we will guide you through the primary functionalities of the package, showing you how to:
Estimate logistic regression coefficients using numerical optimization
Generate bootstrap confidence intervals for model parameters
Make predictions and evaluate model performance using various metrics
Let’s get started!
Before we begin using the functions in the FinalProject13 package, let’s first preprocess the data. For this example, we will be using a dataset called expenses.csv, which contains various information, including smoking status, BMI, and the number of children for a set of individuals. Our goal is to process this data and prepare it for binary classification analysis.
We will begin by loading the expenses.csv file. This dataset includes multiple columns, but we are particularly interested in three columns:
The charges column, which indicates an individuals medical charges.
The age column, which indicates how old an individual is.
The bmi colum, which represents the Body Mass Index (BMI) of each individual.
For our purposes, we will create new adjusted columns based off of the previously mentioned columns.
charges_above_median will consist of 1’s for if the individuals charges are above the median of the dataset or a 0 if it is less than or equal to the median of the dataset, the median of the dataset being 9382.
years_older_than_min will consist of a numeric value that indicates how many years older an individual is than the minimum age value of the dataset, which is 18.
bmi_greater_than_min will consist of a numeric value that indicates how much greater an individual’s bmi is than the minimum bmi value of the dataset, which is 15.96.
# Load the dataset from the 'inst/extdata' directory in the package
expenses <- read.csv(system.file("extdata", "expenses.csv", package = "FinalProjectGroup13"))
head(expenses)
#> age sex bmi children smoker region charges
#> 1 19 female 27.900 0 yes southwest 16884.924
#> 2 18 male 33.770 1 no southeast 1725.552
#> 3 28 male 33.000 3 no southeast 4449.462
#> 4 33 male 22.705 0 no northwest 21984.471
#> 5 32 male 28.880 0 no northwest 3866.855
#> 6 31 female 25.740 0 no southeast 3756.622
# Step 2: Converting Data for Use
# Set the median value for charges
median_charges <- 9382
# Create a new column indicating whether charges are greater than the median
expenses$charges_above_median <- ifelse(expenses$charges > median_charges, 1, 0)
# Create a new column for age in years older than 18
expenses$years_older_than_min <- expenses$age - 18
# Create a new column for bmi greater than min of 15.96
expenses$bmi_greater_than_min <- expenses$bmi - 15.96
# Step 3: Inspecting the Data
# Inspect the 'years_older_than_mi' and 'bmi_greater_than_min' columns for any issues
summary(expenses$years_older_than_min)
#> Min. 1st Qu. Median Mean 3rd Qu. Max.
#> 0.00 9.00 21.00 21.21 33.00 46.00
summary(expenses$bmi_greater_than_min)
#> Min. 1st Qu. Median Mean 3rd Qu. Max.
#> 0.00 10.34 14.44 14.70 18.73 37.17
# Step 4: Data Summary
# Summary of the processed data
summary(expenses)
#> age sex bmi children
#> Min. :18.00 Length:1338 Min. :15.96 Min. :0.000
#> 1st Qu.:27.00 Class :character 1st Qu.:26.30 1st Qu.:0.000
#> Median :39.00 Mode :character Median :30.40 Median :1.000
#> Mean :39.21 Mean :30.66 Mean :1.095
#> 3rd Qu.:51.00 3rd Qu.:34.69 3rd Qu.:2.000
#> Max. :64.00 Max. :53.13 Max. :5.000
#> smoker region charges charges_above_median
#> Length:1338 Length:1338 Min. : 1122 Min. :0.0
#> Class :character Class :character 1st Qu.: 4740 1st Qu.:0.0
#> Mode :character Mode :character Median : 9382 Median :0.5
#> Mean :13270 Mean :0.5
#> 3rd Qu.:16640 3rd Qu.:1.0
#> Max. :63770 Max. :1.0
#> years_older_than_min bmi_greater_than_min
#> Min. : 0.00 Min. : 0.00
#> 1st Qu.: 9.00 1st Qu.:10.34
#> Median :21.00 Median :14.44
#> Mean :21.21 Mean :14.70
#> 3rd Qu.:33.00 3rd Qu.:18.73
#> Max. :46.00 Max. :37.17Now that we have preprocessed the data, let’s begin applying the functions from the FinalProject13 package to perform logistic regression analysis on the data.
The first step in logistic regression is to estimate the coefficients
for the predictor variables. We will use the
estimate_beta() function from our package to perform this
estimation. The function takes a matrix of predictor variables (in our
case, the years_older_than_min column and the
bmi_greater_than_min column) and a binary response
variable (the charges_above_median column), and returns
the estimated beta coefficients.
# Ensure that both 'bmi' and 'children' columns are numeric
expenses$years_older_than_min <- as.numeric(expenses$years_older_than_min)
expenses$bmi_greater_than_min <- as.numeric(expenses$bmi_greater_than_min)
# Create the predictor matrix (bmi and children columns) as a numeric matrix
X <- as.matrix(expenses[, c("years_older_than_min", "bmi_greater_than_min")])
# Response vector (smoker column)
y <- expenses$charges_above_median
# Apply the estimate_beta function to estimate the beta coefficients
model <- estimate_beta(X, y)
# View the estimated coefficients
model$coefficients
#> [,1]
#> -2.07264785
#> years_older_than_min 0.08744715
#> bmi_greater_than_min 0.01529154Our estimate_beta() function returned us three
values:
This value indicates the log-odds of an individuals charges being above the median when both years_older_than_min and bmi_greater_than_min are equal to 0. In other words, it’s the base. It being large and negative indicates that an individual with the base values for the predictor variables is relatively unlikely to have their charges be above the median.
This value indicates the change in the log-odds of an individuals charges being above the median for each additional year above 18. The number being positive and relatively large indicates relatively significantly that the older an individual is they are more likely to have their charges be above the median.
This value indicates the change in the log-odds of an individuals charges being above the median for each additional additonal unit increase in BMI above 15.96. The number being positive and relatively large indicates relatively significantly (though not as significantly as for years_older_than_min) that the higher bmi an individual has they are more likely to have their charges be above the median.
Now that we have estimated the beta coefficients using logistic
regression, we can use bootstrapping to estimate confidence intervals
for these coefficients. Bootstrapping is a resampling technique that
allows us to estimate the sampling distribution of a statistic by
resampling with replacement from the original data. We can do this using
the bootstrap_ci() function from our package:
# Perform bootstrap to calculate confidence intervals for the beta coefficients
bootstrap_result <- bootstrap_ci(X, y, alpha = 0.05, num_bootstrap = 20)
# Display the confidence intervals
print(bootstrap_result)
#> $lower_bound
#> [1] -2.4852720086 0.0810256769 -0.0003556741
#>
#> $upper_bound
#> [1] -1.75035424 0.10059669 0.03740458Our bootstrap_ci() function yielded us with two vectors
consisting of three numbers. The first vector represents the lower bound
of the confident interval (in our case 95% given alpha is 0.05) for each
of the estimated coefficients, while the second vector represents the
upper bound.
Our intercept confidence interval ranges from -2.3432 to -1.6474, indicating that on both the high and low end an individual with the minimum age and bmi is unlikely ot have charges above the median.
Our years_older_than_min confidence interval ranges from 0.07998 to 0.09822, indicating that this predictor is very likely to have a relatively strong positive correlation with charges_above_median as thought before.
Our bmi_greater_than_min confidence interval ranges from -0.01034 to 0.02901, this indicates that our determination that bmi has a positive correlation with charges_above_median may actually be incorrect, although we cannot know for sure either way as the true coefficient could be anywhere within this range, which spans from negative to positive.
After estimating the beta coefficients using the logistic regression
model, the next step is to predict the probabilities of the outcome
variable (whether an individual has charges above the median) for each
observation in the dataset. This is done using the
predict_prob() function from our package.
The predict_prob() function takes the predictor matrix
\(X\) and the estimated beta
coefficients, both gathered earlier, and it returns the predicted
probabilities for each observation in the dataset. These predicted
probabilities indicate the likelihood that each individual has charges
above the median based on their age and bmi.
The formula used for prediction is the logistic function applied to the linear predictor:
\[ p_i = \frac{1}{1 + e^{-(\beta_0 + \beta_1 \cdot \text{BMI})}} \]
where: - \(p_i\) is the predicted probability of being a smoker for the \(i\)-th observation, - \(\beta_0\) is the intercept, - \(\beta_1\) is the coefficient for BMI, - and \(\text{BMI}\) is the predictor variable.
Using the estimated coefficients, the predict_prob()
function computes the probability for each observation in the dataset,
which can then be used for classification or further model
evaluation.
# Apply the predict_prob function to generate predicted probabilities
predictions <- predict_prob(X, model$coefficients)
# View the predicted probabilities for the first few observations
head(predictions)
#> [,1]
#> [1,] 0.1415325
#> [2,] 0.1418139
#> [3,] 0.2813837
#> [4,] 0.3412342
#> [5,] 0.3428045
#> [6,] 0.3129672Our predict_prob() function yielded us with predicted
probabilities for each individual in the dataset. Looking at the first
few, we can see that the prediction function is working as expected. For
example, the first individual is predicted to be very unlikely to have
charges above the median as they have a predicted probability of 14%.
This lines up with our coefficients, as, looking at our dataset, the
first individual is young with a low bmi. In the other direction, we can
look at the highest probability of having charges above the median in
the group, which is the fifth individual. This lines up with our
coefficients, as, looking at our dataset, the individual has the second
highest age of our group of 6 and the third highest bmi.
In our confusion matrix we can see that our model is generating more correct outcomes than incorrect outomes, indicating positive things about our model.
Our value here is 0.5. This indicates that half of the observations in your dataset have charges greater than the median, which makes sense considering the nature of our variable.
Our value here is 71.23%. This means that the model is correctly classifying around 71% of the cases, which indicates positive things about our model
Our value here is 71.60%. This means the model is identifying about 71.6% of the true positives, which indicates positive things about our model.
Our value here is 70.85%, meaning the model is correctly identifying around 70.85% of the true negatives, which indicates positive things about our model.
Our value here is 28.93%, indicating that only about 29% of the predicted positive cases were actually false positives, which indicates positive things about our model.
Our value here is 6.13, suggesting that the model has some discriminatory power between the classes (charges above and below the median), which indicates positive things about our model.
All metrics considered, it appears our model is strong and that the coefficients of age and bmi can be used together effectively to predict whether an individual has charges above the median.
Through this process, the merits of the
FinalProject13 package become clear. We were able to
estimate coefficients from a dataset using estimate_beta,
determine the confidence interval for the coefficients using
bootstrap_ci, create and use a prediction model on the
dataset using the coefficients using predict_prob, and
finally we were able to determine the quality of that model using
confusion_metrics.