This dataset is from Professor Catlin’s public repository on GitHub called “data” located here
Tackling the problem: My plan is to review some of the concepts such as null error rates and harmonic mean that tripped me up in the reading Performance Metrics for Classification Problems in Machine Learning. Then, I will ingest the dataset “penguin_predictions” and go through the calculation exercises. Since I am familiar with confusion matrices (2x2 tables) I likely will use SAS knowledge to re create them in R.
Anticipated data challenges: The biggest hurdle may be calculating in R rather than manually. Second, is making sure the performance metrics are correctly done.
Citations:
Anthropic. (2025). Claude Opus 4.5 [Large language model]. (https://claude.ai/) Accessed January 8, 2026. Link to chat here.
Analysis
#Not shown: Location on computer where I am grabbing the data and creating a saved location
#Install packages needed pacman::p_load(# Package Install and Management pacman, # package install/load janitor, #clean up data names # Project and File Management readr, # import data httr, #github passkey # General Data Management dplyr, # data management tidyr, # data management lubridate, # work with dates zoo, # work with dates tidyverse, # work with dates scales, # ggplot gmodels # freq tables SAS )#Load data#desktoppenguins <-read.csv(paste0(csv_location, "penguin_predictions.csv"), header=TRUE, stringsAsFactors =FALSE)#Review and prepare data for #looking at dataframe ls(penguins)
#change notation to easier format options(scipen =100) # take out sci notation clean$pred_female <-round(clean$pred_female, 2) #round #sex = actual#pred_female = model predicted probability that sex is female #pred_class = #1 = if the model probability is greater than .5#0 = if not #change pred_class in 1/0 binary clean <- clean%>%mutate( pred_class_50edit =case_when( pred_female > .5~1, pred_female <= .5~2), true_class_edit =case_when( sex =="female"~1, TRUE~2 ))
Task #1: Null Error Rate
A. Calculate the null error rate (majority-class error rate).
#what is majority class (higher count of actual data) majority <-table(clean$true_class_edit)addmargins(majority)
1 2 Sum
39 54 93
prop.table(majority)
1 2
0.4193548 0.5806452
#calc null error rate = 1 - (majority class proportion) = lower count of actual data # 1- (54/93)
[1] 0.4193548
B. Create a plot showing the distribution of the actual class (sex).
#sum up true counts (sex) and percents clean_graph<- clean%>%count(sex)%>%mutate( percent = n /sum(n))ggplot(clean_graph, aes(x = sex, y = percent)) +geom_col(fill ="darkblue") +scale_y_continuous(labels = percent) +labs(title ="True Class Percent",x ="Class",y ="Percent" ) +theme_minimal()
C. Explain why knowing the null error rate is important when evaluating models.
It’s important to understand the proportion of the actual distribution of positive and negative values (in this case (sex of penguin) before evaluating how the model does. This of course will be helpful once we examine the confusion matrix and other performance metrics.
Task #2: Confusion Matrices at Multiple Thresholds
################### Importants note:##In order for the confusion matrix to work with out flipping the table, I needed to change the formatting, so now: 1 = TRUE 2 = FALSE ###################adding in thresholds clean <- clean%>%mutate( #first threshold : .2 pred_class_20edit =case_when( pred_female > .2~1, pred_female <= .2~2), #Second threshold: .5 (already done)#Third threshold: .8 pred_class_80edit =case_when( pred_female > .8~1, pred_female <= .8~2 ))#confusion matrix #looking at it in SAS style library(gmodels)#20% threshold CrossTable(clean$pred_class_20edit, clean$true_class_edit, expected =FALSE, prop.r =FALSE, prop.c =FALSE, prop.t =FALSE)
##Precision## #meaning: Estimated outcome or Positive Predictive Value (PPV) #“Among those the model called positive, how many were actually positive?”#equation: Precision = TP / (TP + FP) <- aka row totaladdmargins(table(clean$pred_class_20edit, clean$true_class_edit))
##Recall## #meaning: Sensitivity or True Positive Rate #“Among those who were actually positive, how many did the model identify?”#equation: Recall = TP / (TP + FN) <- aka column total#20% threshold CrossTable(clean$pred_class_20edit, clean$true_class_edit, expected =FALSE, prop.r =FALSE, prop.c =TRUE, prop.t =FALSE)
A .2 threshold would be preferable when we want to maximize the predicted positives regardless of they are actually positive. An example of this is 911 call screenings. Regardless of whether there is actually an emergency, it’s important that filtered like it is one.
A .8 threshold would be preferable when we want to minimize false positives. An example of this could be making sure a donor’s blood type matches with a patient’s blood type. We would want to make sure there are no patients receiving the wrong blood match.
Conclusion
I did a mixture of manual calculations, and I would like to understand how to code these calculations better. It would be interesting to use modeling techniques to play out which thresholds work better.