The National Football League (NFL), the premier American Football competition globally, is one of the most data-rich sports, having a large number of metrics to not only evaluate individual players, but also teams and their performance as a unit and then against specific schemes. One highly used metric is the Quaterback Rating (QBR), a complex metric used to evaluate the performance of the Quarterback each game, and then across the season. Many believe, with the quarterbacks importance, that high performance in this position is required to ensure the team can win and challenge to win the Super Bowl yearly.
The QBR of a quarterback is calculated using the following methods, which involve scaling for each component based on the significance of the metric to a quarterback’s performance on the field.
The passer rating is built from four components, each capped between 0 and 2.375:
\[A = \left(\frac{\text{Completions}}{\text{Attempts}} - 0.3\right) \times 5\]
\[B = \left(\frac{\text{Passing Yards}}{\text{Attempts}} - 3\right) \times 0.25\]
\[C = \left(\frac{\text{Touchdowns}}{\text{Attempts}}\right) \times 20\]
\[D = 2.375 - \left(\frac{\text{Interceptions}}{\text{Attempts}} \times 25\right)\]
The final rating is then:
\[\text{Passer Rating} = \left(\frac{A + B + C + D}{6}\right) \times 100\]
The analyses in this portfolio were conducted in R and R Studio. This analysis will utilise a Backward Model Selection process against a Linear Regression model to determine the most important metrics to a Quarterback’s rating. Ultimately, the output from the linear regression model will be compared with the current, accepted QBR formula to determine how accurate the regression model was at identifying important metrics.
Data for the analysis was sourced in two different ways - the testing data was sourced from Kaggle, which was a cleaned NFL passing statistics dataset for all observations between 2001 and 2023 inclusive. The second dataset, which was used to test the linear regression model predictions, was scraped from Pro Football Reference’s 2025 NFL Passing data tab.
The aim of this analytical portfolio is to create a linear regression model which predicts a quarterback’s passer rating on their seasonal metric, and comparing this against the observed value to garner how successful this model is. Furthermore, we will also consider the variables the model selects and how many of these selected parameters are currently incorporated in the quarterback passer rating formula currently, albeit in different ways.
#Analysis
#There are a number of required packages for this analysis which will need to be installed
library(dplyr)
##
## Attaching package: 'dplyr'
## The following objects are masked from 'package:stats':
##
## filter, lag
## The following objects are masked from 'package:base':
##
## intersect, setdiff, setequal, union
library(tidyr)
library(tidyverse)
## ── Attaching core tidyverse packages ──────────────────────── tidyverse 2.0.0 ──
## ✔ forcats 1.0.1 ✔ readr 2.2.0
## ✔ ggplot2 4.0.3 ✔ stringr 1.6.0
## ✔ lubridate 1.9.5 ✔ tibble 3.3.1
## ✔ purrr 1.2.2
## ── Conflicts ────────────────────────────────────────── tidyverse_conflicts() ──
## ✖ dplyr::filter() masks stats::filter()
## ✖ dplyr::lag() masks stats::lag()
## ℹ Use the conflicted package (<http://conflicted.r-lib.org/>) to force all conflicts to become errors
library(ggplot2)
library(nlme)
## Warning: package 'nlme' was built under R version 4.6.1
##
## Attaching package: 'nlme'
##
## The following object is masked from 'package:dplyr':
##
## collapse
library(readr)
qb_pass_data <- readr::read_csv("passing_cleaned.csv") |>
glimpse()
## New names:
## Rows: 2350 Columns: 27
## ── Column specification
## ──────────────────────────────────────────────────────── Delimiter: "," chr
## (2): Player, Tm dbl (25): ...1, Age, G, GS, Cmp, Att, Cmp%, Yds, TD, TD%, Int,
## Int%, 1D, Lng...
## ℹ Use `spec()` to retrieve the full column specification for this data. ℹ
## Specify the column types or set `show_col_types = FALSE` to quiet this message.
## • `` -> `...1`
## Rows: 2,350
## Columns: 27
## $ ...1 <dbl> 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, …
## $ Player <chr> "Kurt Warner", "Peyton Manning", "Brett Favre", "Aaron Brooks"…
## $ Tm <chr> "STL", "IND", "GNB", "NOR", "OAK", "KAN", "NYG", "ARI", "SFO",…
## $ Age <dbl> 30, 25, 32, 25, 36, 31, 29, 27, 31, 39, 33, 28, 31, 30, 25, 29…
## $ G <dbl> 16, 16, 16, 16, 16, 16, 16, 16, 16, 16, 16, 15, 15, 16, 16, 16…
## $ GS <dbl> 16, 16, 16, 16, 16, 16, 16, 16, 16, 16, 16, 15, 15, 16, 16, 15…
## $ Cmp <dbl> 375, 343, 314, 312, 361, 296, 327, 304, 316, 294, 340, 264, 28…
## $ Att <dbl> 546, 547, 510, 558, 549, 523, 568, 525, 504, 521, 559, 431, 47…
## $ `Cmp%` <dbl> 68.7, 62.7, 61.6, 55.9, 65.8, 56.6, 57.6, 57.9, 62.7, 56.4, 60…
## $ Yds <dbl> 4830, 4131, 3921, 3832, 3828, 3783, 3764, 3653, 3538, 3464, 34…
## $ TD <dbl> 36, 26, 32, 26, 27, 17, 19, 18, 32, 15, 13, 21, 19, 20, 25, 12…
## $ `TD%` <dbl> 6.6, 4.8, 6.3, 4.7, 4.9, 3.3, 3.3, 3.4, 6.3, 2.9, 2.3, 4.9, 4.…
## $ Int <dbl> 22, 23, 15, 22, 9, 24, 16, 14, 12, 18, 11, 12, 13, 19, 12, 22,…
## $ `Int%` <dbl> 4.0, 4.2, 2.9, 3.9, 1.6, 4.6, 2.8, 2.7, 2.4, 3.5, 2.0, 2.8, 2.…
## $ `1D` <dbl> 233, 201, 187, 182, 195, 176, 189, 177, 184, 168, 181, 161, 16…
## $ Lng <dbl> 65, 86, 67, 63, 49, 67, 76, 68, 61, 78, 47, 71, 44, 74, 64, 49…
## $ `Y/A` <dbl> 8.8, 7.6, 7.7, 6.9, 7.0, 7.2, 6.6, 7.0, 7.0, 6.6, 6.1, 7.8, 7.…
## $ `AY/A` <dbl> 8.4, 6.6, 7.6, 6.0, 7.2, 5.8, 6.0, 6.4, 7.2, 5.7, 5.7, 7.5, 6.…
## $ `Y/C` <dbl> 12.9, 12.0, 12.5, 12.3, 10.6, 12.8, 11.5, 12.0, 11.2, 11.8, 10…
## $ `Y/G` <dbl> 301.9, 258.2, 245.1, 239.5, 239.3, 236.4, 235.3, 228.3, 221.1,…
## $ Rate <dbl> 101.4, 84.1, 94.1, 76.4, 95.5, 71.1, 77.1, 79.6, 94.8, 72.0, 7…
## $ Sk <dbl> 38, 29, 22, 50, 27, 39, 36, 29, 26, 25, 44, 37, 57, 27, 39, 25…
## $ `Yds-s` <dbl> 233, 232, 151, 330, 155, 198, 206, 204, 114, 168, 269, 251, 38…
## $ `Sk%` <dbl> 6.5, 5.0, 4.1, 8.2, 4.7, 6.9, 6.0, 5.2, 4.9, 4.6, 7.3, 7.9, 10…
## $ `NY/A` <dbl> 7.87, 6.77, 7.09, 5.76, 6.38, 6.38, 5.89, 6.23, 6.46, 6.04, 5.…
## $ `ANY/A` <dbl> 7.41, 5.88, 7.02, 4.99, 6.61, 5.06, 5.33, 5.74, 6.65, 5.10, 4.…
## $ Year <dbl> 2001, 2001, 2001, 2001, 2001, 2001, 2001, 2001, 2001, 2001, 20…
#A large number of variables have initialised abbreviations
#Despite their cleanliness, it is difficult to deduce some from others so some renaming of variables will occur
##The full metric name is listed above each name change line, with names pulled in line with the data from Pro Football Reference
qb_pass_data <- qb_pass_data |>
dplyr::rename(
#Completion Percentage - Completed Passes / Attempted Passes
comp_pct = "Cmp%",
#TD Percentage (Percentage of throws which result in a Touchdown)
TD_pct = "TD%",
#Interception Percentage (Percentage of throws which result in an Interception)
int_pct = "Int%",
#Sack Percentage (Percentage of pass plays which resulted in the quarterback getting sacked)
sack_pct = "Sk%",
#Yards Per Attempt
yards_per_att = "Y/A",
#Adjusted Yards Gained per Attempt
adj_yards_gain_per_att = "AY/A",
#Yards Gained per Completion
yards_per_comp = "Y/C",
#Yards Gained from Throws per Game (seasonal data)
yards_per_game = "Y/G",
#Total Sack Yards cinceded
tot_sack_yds = "Yds-s",
#Net Yards Gained per Attempt
net_yds_per_att = "NY/A",
#Adjusted Net Yards Gained per Attempt
adj_net_yds_per_att = "ANY/A"
)
We will also remove all seasonal Quarterback occurences where the quarterback did not start more than 1 game, removing observations where the quarterback may have had minimal pass plays due to filling in for injury, a spot-start due to illness etc.
qb_pass_data <- qb_pass_data |>
subset(
GS > 1
)
#Removed approximately 800 observations of QBs with minimal playing time
To ensure a strong understanding of the dataset and the relationships involved within it, Exploratory Data Analysis is critical. This will visually show relationships between key metrics, both showing positive data interaction, any applicable outliers as well as further data investigation which may warrant more cleaning if applicable.
ggplot(data = qb_pass_data) +
geom_histogram(aes(x = Att), bins = 20)
#Largely biased to QBs who have had minimal passes
Although this plot is bimodal and could be accepted, the large occurence of quarterbacks still remain with minimal pass attempts and may be due to quarterbacks who filled in late in blowout games and accrued a number of attempts in meaningless situations. Therefore, further data cleaning will be run to clean out quarterbacks who may had started multiple games but did not so lots of playtime in them nor made a large number of throws to influence their quarterback rating.
#Understanding a QB may attempt, on avg, 30 passes per game
#Subset out any QBs with less than 50 passes (2 games essentially)
qb_pass_data <- qb_pass_data |>
subset(
Att > 50
)
#Removed a further 400 observations
ggplot(data = qb_pass_data) +
geom_histogram(aes(x = Att), bins = 20)
#Far greater model, displaying a bimodel distribution of back-ups who are in the game when QB injured
#And permanent starting QBs
#Such distribution should be kept to display the fact players get injured and backups see gametime throughout season
ggplot(data = qb_pass_data) +
geom_histogram(aes(x = Rate), bins = 20)
#Normally distirbuted dataset with a slight left skew, demonstrating far greater sparsity of QBs putting up poor ratings
ggplot(aes(x = Att, y = Yds), data = qb_pass_data) +
geom_point(alpha = 0.2) +
geom_smooth(method = "lm")
## `geom_smooth()` using formula = 'y ~ x'
#More attempts = more pass yards (pbvious)
#However, really high pass yards (4500+) often do it under the lm projected pass attempt threshold (efficient passing offenses)
ggplot(aes(x = comp_pct, y = TD_pct), data = qb_pass_data) +
geom_point(alpha = 0.2) +
geom_smooth(method = "loess")
## `geom_smooth()` using formula = 'y ~ x'
#Geom_smooth layer shows no obvious relationship between completion percentage and TD percentage
#A large number of observations occur above line, and loess shows a curve through the data
ggplot(aes(x = Cmp, y = Att), data = qb_pass_data) +
geom_point(alpha = 0.2) +
geom_smooth(method = "loess")
## `geom_smooth()` using formula = 'y ~ x'
#Loess wrapper shows no discernable window either side of the line of best fit
#The line of best fit does curve at the end, showing a point where more completions does not come at the cost of more attempts
ggplot(aes(x = TD, y = Att), data = qb_pass_data) +
geom_point(alpha = 0.2) +
geom_smooth(method = "loess")
## `geom_smooth()` using formula = 'y ~ x'
#Once again, a certain pass attempt threshold does not discern any gain in benefit for a metric
As the previous two EDA plots are now demonstrating, a certain metric does not gain any discernable benefit from exceeding a certain pass attempts threshold. This threshold is occurring around 600 Pass Attempts for the season, which still averages to approximately 37.5 pass attempts in a 16 game season. This indicates that, at some threshold, quarterback (and team talent) play a greater role in the performance of a metric over just the raw volume of pass attempts they have.
ggplot() +
geom_histogram(aes(x = Att), data = qb_pass_data) +
facet_wrap(~Year)
## `stat_bin()` using `bins = 30`. Pick better value `binwidth`.
#More of the observations of over 600 pass attempts coming in recent years
#Ties in to trends of the game which say the NFL is more pass heavy
Ultimately, the question at-hand investigates which statistics influence Quarterback Rating the most, without the linear regression model having prior knowledge of what currently formulates this calculation. Now the EDA has shown us strong data integrity and distribution, and with no further cleaning or removal of data required, the next portion of EDA will cover the relationship of metrics with QBR, allowing for some personal thoughts of important metrics to be formed based on relationships.
#Does more sacks leads to a lower QBR?
ggplot(aes(x = Rate, y = Sk), data = qb_pass_data) +
geom_point()
#And, if not the number of sacks, does the total yards lost to sacks lower your QBR?
ggplot(aes(x = Rate, y = tot_sack_yds), data = qb_pass_data) +
geom_point()
#Does more TDs leads to a greater QBR
ggplot(aes(x = Rate, y = TD), data = qb_pass_data) +
geom_point() +
geom_smooth(method = "lm")
## `geom_smooth()` using formula = 'y ~ x'
#Pairs plot to gain each variable against eachother
numeric_qb_data <- qb_pass_data |>
dplyr::select(-c(Player, Tm, Year, ...1))
From the Pairs plots, a number of distributions are exhibited. For counting variables including Games, Games Started, a Bi-Modal distirbution is still evident. However, this time, the largest peak occurs towards the right hand side, indicating the starters who are fit and play the majority, if not all, of their games. Beyond that, all scaled metrics show a normal distribution for their graph, with the skew dependent on the metric in question and how well the dataset of quarterbacks performed in these metrics.
For this analysis, we will be using a Linear Regression model. When utilising linear regression models, there are a number of ways that the models can be constructed. For the purpose of this analysis, a BACKWARD MODEL SELECTION PROCESS will be used. This means, we will commence with every variable in the dataset informing the Quarterback Rating and, from there, remove variables from the model based on how insignificant they are in contributing to the Quarterback Rating calculation.
For the purpose of the analyses, the p-value for the regression model will be set at XXX. Traditionally, the p-value exists at 0.05. However, the large number of variables informing the Quarterback Rating to begin courtesy of the large dataset means a scaling of the p-value should be considered to ensure effective selection of parameters is made.
qbr1 <- lm(Rate ~ ., data = numeric_qb_data)
summary(qbr1)
##
## Call:
## lm(formula = Rate ~ ., data = numeric_qb_data)
##
## Residuals:
## Min 1Q Median 3Q Max
## -1.3129 -0.1252 -0.0029 0.1147 12.1973
##
## Coefficients:
## Estimate Std. Error t value Pr(>|t|)
## (Intercept) -0.6906423 1.1540215 -0.598 0.549653
## Age 0.0050528 0.0029303 1.724 0.084930 .
## G -0.0271991 0.0167041 -1.628 0.103752
## GS -0.0011848 0.0121832 -0.097 0.922549
## Cmp -0.0020213 0.0018284 -1.106 0.269163
## Att 0.0040445 0.0011861 3.410 0.000673 ***
## comp_pct 0.8680053 0.0203658 42.621 < 2e-16 ***
## Yds -0.0001505 0.0001377 -1.093 0.274579
## TD 0.0037755 0.0066860 0.565 0.572405
## TD_pct 2.1247809 0.1128190 18.834 < 2e-16 ***
## Int -0.0337380 0.0064527 -5.229 2.05e-07 ***
## int_pct -1.3666992 0.2461606 -5.552 3.54e-08 ***
## `1D` -0.0009815 0.0014811 -0.663 0.507659
## Lng 0.0012086 0.0010393 1.163 0.245138
## yards_per_att 1.8550752 0.3460273 5.361 1.01e-07 ***
## adj_yards_gain_per_att 2.4327503 0.3245841 7.495 1.36e-13 ***
## yards_per_comp 0.1538026 0.1010353 1.522 0.128231
## yards_per_game -0.0018981 0.0007779 -2.440 0.014846 *
## Sk 0.0117998 0.0059448 1.985 0.047406 *
## tot_sack_yds -0.0010555 0.0008608 -1.226 0.220396
## sack_pct -0.0738531 0.0317991 -2.322 0.020389 *
## net_yds_per_att -4.0097705 0.5842412 -6.863 1.12e-11 ***
## adj_net_yds_per_att 3.7480665 0.5265180 7.119 1.97e-12 ***
## ---
## Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
##
## Residual standard error: 0.419 on 1096 degrees of freedom
## Multiple R-squared: 0.9992, Adjusted R-squared: 0.9992
## F-statistic: 6.138e+04 on 22 and 1096 DF, p-value: < 2.2e-16
The three variables with the most insignificant p-values (the largest p-value number) will be removed from the model. Therefore, GS, TD and 1D will be removed
qbr2 <- lm(Rate ~ . - GS - TD - `1D`, data = numeric_qb_data)
summary(qbr2)
##
## Call:
## lm(formula = Rate ~ . - GS - TD - `1D`, data = numeric_qb_data)
##
## Residuals:
## Min 1Q Median 3Q Max
## -1.3225 -0.1259 -0.0015 0.1125 12.2023
##
## Coefficients:
## Estimate Std. Error t value Pr(>|t|)
## (Intercept) -0.6640560 1.1515058 -0.577 0.564270
## Age 0.0051101 0.0029232 1.748 0.080717 .
## G -0.0275423 0.0150161 -1.834 0.066897 .
## Cmp -0.0020899 0.0018087 -1.156 0.248131
## Att 0.0038929 0.0011510 3.382 0.000745 ***
## comp_pct 0.8679514 0.0203441 42.664 < 2e-16 ***
## Yds -0.0001503 0.0001057 -1.421 0.155500
## TD_pct 2.1376676 0.1113811 19.192 < 2e-16 ***
## Int -0.0341089 0.0064164 -5.316 1.28e-07 ***
## int_pct -1.3745147 0.2448448 -5.614 2.51e-08 ***
## Lng 0.0012958 0.0010251 1.264 0.206478
## yards_per_att 1.8681514 0.3450999 5.413 7.59e-08 ***
## adj_yards_gain_per_att 2.4161031 0.3235406 7.468 1.66e-13 ***
## yards_per_comp 0.1536845 0.1009195 1.523 0.128086
## yards_per_game -0.0018814 0.0007713 -2.439 0.014878 *
## Sk 0.0116298 0.0059008 1.971 0.048986 *
## tot_sack_yds -0.0010572 0.0008582 -1.232 0.218247
## sack_pct -0.0739455 0.0317309 -2.330 0.019965 *
## net_yds_per_att -4.0140914 0.5812149 -6.906 8.40e-12 ***
## adj_net_yds_per_att 3.7458529 0.5244708 7.142 1.67e-12 ***
## ---
## Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
##
## Residual standard error: 0.4185 on 1099 degrees of freedom
## Multiple R-squared: 0.9992, Adjusted R-squared: 0.9992
## F-statistic: 7.122e+04 on 19 and 1099 DF, p-value: < 2.2e-16
qbr3 <- lm(Rate ~ . - GS - TD - `1D` - Cmp - Lng, data = numeric_qb_data)
summary(qbr3)
##
## Call:
## lm(formula = Rate ~ . - GS - TD - `1D` - Cmp - Lng, data = numeric_qb_data)
##
## Residuals:
## Min 1Q Median 3Q Max
## -1.3023 -0.1244 -0.0004 0.1081 12.2371
##
## Coefficients:
## Estimate Std. Error t value Pr(>|t|)
## (Intercept) -2.353e-01 1.101e+00 -0.214 0.830781
## Age 4.936e-03 2.922e-03 1.689 0.091500 .
## G -2.408e-02 1.487e-02 -1.619 0.105655
## Att 2.971e-03 7.676e-04 3.871 0.000115 ***
## comp_pct 8.592e-01 1.911e-02 44.951 < 2e-16 ***
## Yds -2.231e-04 8.951e-05 -2.492 0.012853 *
## TD_pct 2.129e+00 1.113e-01 19.127 < 2e-16 ***
## Int -3.155e-02 6.109e-03 -5.165 2.86e-07 ***
## int_pct -1.362e+00 2.448e-01 -5.565 3.29e-08 ***
## yards_per_att 1.924e+00 3.433e-01 5.604 2.65e-08 ***
## adj_yards_gain_per_att 2.421e+00 3.236e-01 7.479 1.52e-13 ***
## yards_per_comp 1.398e-01 1.002e-01 1.396 0.163141
## yards_per_game -1.725e-03 7.656e-04 -2.253 0.024449 *
## Sk 1.210e-02 5.895e-03 2.053 0.040325 *
## tot_sack_yds -1.131e-03 8.572e-04 -1.319 0.187423
## sack_pct -7.586e-02 3.172e-02 -2.391 0.016956 *
## net_yds_per_att -4.066e+00 5.806e-01 -7.003 4.36e-12 ***
## adj_net_yds_per_att 3.784e+00 5.242e-01 7.218 9.83e-13 ***
## ---
## Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
##
## Residual standard error: 0.4187 on 1101 degrees of freedom
## Multiple R-squared: 0.9992, Adjusted R-squared: 0.9992
## F-statistic: 7.952e+04 on 17 and 1101 DF, p-value: < 2.2e-16
qbr4 <- lm(Rate ~ . - GS - TD - `1D` - Cmp - Lng - G - yards_per_comp, data = numeric_qb_data)
summary(qbr4)
##
## Call:
## lm(formula = Rate ~ . - GS - TD - `1D` - Cmp - Lng - G - yards_per_comp,
## data = numeric_qb_data)
##
## Residuals:
## Min 1Q Median 3Q Max
## -1.2831 -0.1245 -0.0008 0.1133 12.2727
##
## Coefficients:
## Estimate Std. Error t value Pr(>|t|)
## (Intercept) 1.124e+00 2.837e-01 3.962 7.91e-05 ***
## Age 5.141e-03 2.923e-03 1.759 0.078848 .
## Att 2.802e-03 6.820e-04 4.109 4.26e-05 ***
## comp_pct 8.332e-01 3.260e-03 255.590 < 2e-16 ***
## Yds -2.854e-04 8.419e-05 -3.390 0.000724 ***
## TD_pct 2.128e+00 1.105e-01 19.257 < 2e-16 ***
## Int -3.320e-02 6.043e-03 -5.494 4.89e-08 ***
## int_pct -1.356e+00 2.432e-01 -5.577 3.08e-08 ***
## yards_per_att 2.044e+00 3.328e-01 6.143 1.13e-09 ***
## adj_yards_gain_per_att 2.486e+00 3.168e-01 7.848 9.98e-15 ***
## yards_per_game -6.591e-04 4.098e-04 -1.608 0.108071
## Sk 9.315e-03 5.756e-03 1.618 0.105896
## tot_sack_yds -9.879e-04 8.553e-04 -1.155 0.248326
## sack_pct -6.490e-02 3.132e-02 -2.072 0.038474 *
## net_yds_per_att -3.963e+00 5.793e-01 -6.841 1.30e-11 ***
## adj_net_yds_per_att 3.718e+00 5.239e-01 7.096 2.29e-12 ***
## ---
## Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
##
## Residual standard error: 0.4192 on 1103 degrees of freedom
## Multiple R-squared: 0.9992, Adjusted R-squared: 0.9992
## F-statistic: 8.991e+04 on 15 and 1103 DF, p-value: < 2.2e-16
After the removal of a large number of variables, a model comparison is a useful way to track how the model is shaping.
sjPlot::tab_model(qbr1, qbr2, qbr3, qbr4, show.aic = T, show.aicc = T)
| Rate | Rate | Rate | Rate | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Predictors | Estimates | CI | p | Estimates | CI | p | Estimates | CI | p | Estimates | CI | p |
| (Intercept) | -0.69 | -2.95 – 1.57 | 0.550 | -0.66 | -2.92 – 1.60 | 0.564 | -0.24 | -2.39 – 1.92 | 0.831 | 1.12 | 0.57 – 1.68 | <0.001 |
| Age | 0.01 | -0.00 – 0.01 | 0.085 | 0.01 | -0.00 – 0.01 | 0.081 | 0.00 | -0.00 – 0.01 | 0.091 | 0.01 | -0.00 – 0.01 | 0.079 |
| G | -0.03 | -0.06 – 0.01 | 0.104 | -0.03 | -0.06 – 0.00 | 0.067 | -0.02 | -0.05 – 0.01 | 0.106 | |||
| GS | -0.00 | -0.03 – 0.02 | 0.923 | |||||||||
| Cmp | -0.00 | -0.01 – 0.00 | 0.269 | -0.00 | -0.01 – 0.00 | 0.248 | ||||||
| Att | 0.00 | 0.00 – 0.01 | 0.001 | 0.00 | 0.00 – 0.01 | 0.001 | 0.00 | 0.00 – 0.00 | <0.001 | 0.00 | 0.00 – 0.00 | <0.001 |
| comp pct | 0.87 | 0.83 – 0.91 | <0.001 | 0.87 | 0.83 – 0.91 | <0.001 | 0.86 | 0.82 – 0.90 | <0.001 | 0.83 | 0.83 – 0.84 | <0.001 |
| Yds | -0.00 | -0.00 – 0.00 | 0.275 | -0.00 | -0.00 – 0.00 | 0.156 | -0.00 | -0.00 – -0.00 | 0.013 | -0.00 | -0.00 – -0.00 | 0.001 |
| TD | 0.00 | -0.01 – 0.02 | 0.572 | |||||||||
| TD pct | 2.12 | 1.90 – 2.35 | <0.001 | 2.14 | 1.92 – 2.36 | <0.001 | 2.13 | 1.91 – 2.35 | <0.001 | 2.13 | 1.91 – 2.34 | <0.001 |
| Int | -0.03 | -0.05 – -0.02 | <0.001 | -0.03 | -0.05 – -0.02 | <0.001 | -0.03 | -0.04 – -0.02 | <0.001 | -0.03 | -0.05 – -0.02 | <0.001 |
| int pct | -1.37 | -1.85 – -0.88 | <0.001 | -1.37 | -1.85 – -0.89 | <0.001 | -1.36 | -1.84 – -0.88 | <0.001 | -1.36 | -1.83 – -0.88 | <0.001 |
| 1D | -0.00 | -0.00 – 0.00 | 0.508 | |||||||||
| Lng | 0.00 | -0.00 – 0.00 | 0.245 | 0.00 | -0.00 – 0.00 | 0.206 | ||||||
| yards per att | 1.86 | 1.18 – 2.53 | <0.001 | 1.87 | 1.19 – 2.55 | <0.001 | 1.92 | 1.25 – 2.60 | <0.001 | 2.04 | 1.39 – 2.70 | <0.001 |
| adj yards gain per att | 2.43 | 1.80 – 3.07 | <0.001 | 2.42 | 1.78 – 3.05 | <0.001 | 2.42 | 1.79 – 3.06 | <0.001 | 2.49 | 1.86 – 3.11 | <0.001 |
| yards per comp | 0.15 | -0.04 – 0.35 | 0.128 | 0.15 | -0.04 – 0.35 | 0.128 | 0.14 | -0.06 – 0.34 | 0.163 | |||
| yards per game | -0.00 | -0.00 – -0.00 | 0.015 | -0.00 | -0.00 – -0.00 | 0.015 | -0.00 | -0.00 – -0.00 | 0.024 | -0.00 | -0.00 – 0.00 | 0.108 |
| Sk | 0.01 | 0.00 – 0.02 | 0.047 | 0.01 | 0.00 – 0.02 | 0.049 | 0.01 | 0.00 – 0.02 | 0.040 | 0.01 | -0.00 – 0.02 | 0.106 |
| tot sack yds | -0.00 | -0.00 – 0.00 | 0.220 | -0.00 | -0.00 – 0.00 | 0.218 | -0.00 | -0.00 – 0.00 | 0.187 | -0.00 | -0.00 – 0.00 | 0.248 |
| sack pct | -0.07 | -0.14 – -0.01 | 0.020 | -0.07 | -0.14 – -0.01 | 0.020 | -0.08 | -0.14 – -0.01 | 0.017 | -0.06 | -0.13 – -0.00 | 0.038 |
| net yds per att | -4.01 | -5.16 – -2.86 | <0.001 | -4.01 | -5.15 – -2.87 | <0.001 | -4.07 | -5.20 – -2.93 | <0.001 | -3.96 | -5.10 – -2.83 | <0.001 |
| adj net yds per att | 3.75 | 2.71 – 4.78 | <0.001 | 3.75 | 2.72 – 4.77 | <0.001 | 3.78 | 2.76 – 4.81 | <0.001 | 3.72 | 2.69 – 4.75 | <0.001 |
| Observations | 1119 | 1119 | 1119 | 1119 | ||||||||
| R2 / R2 adjusted | 0.999 / 0.999 | 0.999 / 0.999 | 0.999 / 0.999 | 0.999 / 0.999 | ||||||||
| AIC | 1253.374 | 1248.086 | 1247.129 | 1247.890 | ||||||||
| AICc | 1254.471 | 1248.928 | 1247.820 | 1248.445 | ||||||||
As evidenced, the AIC and AICc metrics have lowered from the original, full model. This is a positive progression as we progress through the linear regression model selection, as a low AIC/AICc value indicates a model which loses less information when predicting as well as penalising for the balance between model fit and complexity.
Despite the AIC and AICc continuing to decrease with the QBR3 model showing the lowest AIC and AICc figures, there are a number of variables which can still be considered insignificant at the p-value < 0.05 level. As we are backwards selecting the model, we will continue to proceed to remove these variables until we get as set of fully significant variables, continuing to compare models built along the way with QBR3 and QBR4 linear regression models which have the better AIC/AICc.
QBR model 5 proceeds to remove Percentage of dropbacks resulting in sacks, sacks total, player age and yards per game from the linear regression model. Age, for example, is a feasible removal and should have been considered earlier. Although Age may impact Quarterback performance, it isn’t an on-field performance indicator that should be baked into the regression model if we are only considering their performance as a player.
qbr5 <- lm(Rate ~ Att + comp_pct + Yds + TD_pct + Int + int_pct + yards_per_att + adj_yards_gain_per_att +
net_yds_per_att + adj_net_yds_per_att, data = numeric_qb_data)
summary(qbr5)
##
## Call:
## lm(formula = Rate ~ Att + comp_pct + Yds + TD_pct + Int + int_pct +
## yards_per_att + adj_yards_gain_per_att + net_yds_per_att +
## adj_net_yds_per_att, data = numeric_qb_data)
##
## Residuals:
## Min 1Q Median 3Q Max
## -1.2806 -0.1257 0.0008 0.1127 12.4053
##
## Coefficients:
## Estimate Std. Error t value Pr(>|t|)
## (Intercept) 8.850e-01 2.378e-01 3.722 0.000207 ***
## Att 3.266e-03 6.277e-04 5.204 2.33e-07 ***
## comp_pct 8.333e-01 3.199e-03 260.478 < 2e-16 ***
## Yds -3.333e-04 8.156e-05 -4.087 4.69e-05 ***
## TD_pct 2.204e+00 1.024e-01 21.512 < 2e-16 ***
## Int -3.514e-02 6.007e-03 -5.849 6.50e-09 ***
## int_pct -1.517e+00 2.254e-01 -6.730 2.71e-11 ***
## yards_per_att 1.748e+00 3.034e-01 5.761 1.08e-08 ***
## adj_yards_gain_per_att 2.411e+00 3.058e-01 7.884 7.57e-15 ***
## net_yds_per_att -3.249e+00 4.607e-01 -7.052 3.10e-12 ***
## adj_net_yds_per_att 3.401e+00 4.856e-01 7.003 4.33e-12 ***
## ---
## Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
##
## Residual standard error: 0.4201 on 1108 degrees of freedom
## Multiple R-squared: 0.9992, Adjusted R-squared: 0.9992
## F-statistic: 1.343e+05 on 10 and 1108 DF, p-value: < 2.2e-16
As evidenced, QBR6’s linear regression model has achieved all 10 variables, along with the intercept, having statistical significance at the level of p < 0.05. This is a strong sign that the model is become much more refined to the point wherein all performance metrics are being considered as significant to impacting a Quarterback’s rating value.
sjPlot::tab_model(qbr3, qbr4, qbr5, show.aic = T, show.aicc = T)
| Rate | Rate | Rate | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Predictors | Estimates | CI | p | Estimates | CI | p | Estimates | CI | p |
| (Intercept) | -0.24 | -2.39 – 1.92 | 0.831 | 1.12 | 0.57 – 1.68 | <0.001 | 0.88 | 0.42 – 1.35 | <0.001 |
| Age | 0.00 | -0.00 – 0.01 | 0.091 | 0.01 | -0.00 – 0.01 | 0.079 | |||
| G | -0.02 | -0.05 – 0.01 | 0.106 | ||||||
| Att | 0.00 | 0.00 – 0.00 | <0.001 | 0.00 | 0.00 – 0.00 | <0.001 | 0.00 | 0.00 – 0.00 | <0.001 |
| comp pct | 0.86 | 0.82 – 0.90 | <0.001 | 0.83 | 0.83 – 0.84 | <0.001 | 0.83 | 0.83 – 0.84 | <0.001 |
| Yds | -0.00 | -0.00 – -0.00 | 0.013 | -0.00 | -0.00 – -0.00 | 0.001 | -0.00 | -0.00 – -0.00 | <0.001 |
| TD pct | 2.13 | 1.91 – 2.35 | <0.001 | 2.13 | 1.91 – 2.34 | <0.001 | 2.20 | 2.00 – 2.40 | <0.001 |
| Int | -0.03 | -0.04 – -0.02 | <0.001 | -0.03 | -0.05 – -0.02 | <0.001 | -0.04 | -0.05 – -0.02 | <0.001 |
| int pct | -1.36 | -1.84 – -0.88 | <0.001 | -1.36 | -1.83 – -0.88 | <0.001 | -1.52 | -1.96 – -1.07 | <0.001 |
| yards per att | 1.92 | 1.25 – 2.60 | <0.001 | 2.04 | 1.39 – 2.70 | <0.001 | 1.75 | 1.15 – 2.34 | <0.001 |
| adj yards gain per att | 2.42 | 1.79 – 3.06 | <0.001 | 2.49 | 1.86 – 3.11 | <0.001 | 2.41 | 1.81 – 3.01 | <0.001 |
| yards per comp | 0.14 | -0.06 – 0.34 | 0.163 | ||||||
| yards per game | -0.00 | -0.00 – -0.00 | 0.024 | -0.00 | -0.00 – 0.00 | 0.108 | |||
| Sk | 0.01 | 0.00 – 0.02 | 0.040 | 0.01 | -0.00 – 0.02 | 0.106 | |||
| tot sack yds | -0.00 | -0.00 – 0.00 | 0.187 | -0.00 | -0.00 – 0.00 | 0.248 | |||
| sack pct | -0.08 | -0.14 – -0.01 | 0.017 | -0.06 | -0.13 – -0.00 | 0.038 | |||
| net yds per att | -4.07 | -5.20 – -2.93 | <0.001 | -3.96 | -5.10 – -2.83 | <0.001 | -3.25 | -4.15 – -2.34 | <0.001 |
| adj net yds per att | 3.78 | 2.76 – 4.81 | <0.001 | 3.72 | 2.69 – 4.75 | <0.001 | 3.40 | 2.45 – 4.35 | <0.001 |
| Observations | 1119 | 1119 | 1119 | ||||||
| R2 / R2 adjusted | 0.999 / 0.999 | 0.999 / 0.999 | 0.999 / 0.999 | ||||||
| AIC | 1247.129 | 1247.890 | 1247.558 | ||||||
| AICc | 1247.820 | 1248.445 | 1247.840 | ||||||
Despite all variables becoming significant in the QBR5 linear regression model, the AIC and AICc values have seen a small increase against the prior models. This indicates that, although all variables are significant to the model, some of the ones removed as a result of not being statistically significant at our set threshold are actually influential, to a degree, on the outcome of a quarterback’s rating calculation.
When examining the linear model summary for the QBR5 model, it is interesting to note that yards and interceptions variable have such low estimate co-effecients that they would be 0.00 when rounded to 2 decimal points. Considering this fact, we will further experiment with a more refined model which removes such variables where, although significant, there low numerical count for a QBR calculation could very well result in the value generated from the co-efficient multiplied with the observed value being mathematically minute.
qbr6 <- lm(Rate ~ comp_pct + TD_pct + int_pct + yards_per_att + adj_yards_gain_per_att +
net_yds_per_att + adj_net_yds_per_att, data = numeric_qb_data)
summary(qbr6)
##
## Call:
## lm(formula = Rate ~ comp_pct + TD_pct + int_pct + yards_per_att +
## adj_yards_gain_per_att + net_yds_per_att + adj_net_yds_per_att,
## data = numeric_qb_data)
##
## Residuals:
## Min 1Q Median 3Q Max
## -1.1713 -0.1174 -0.0029 0.1069 12.9160
##
## Coefficients:
## Estimate Std. Error t value Pr(>|t|)
## (Intercept) 1.588767 0.183188 8.673 < 2e-16 ***
## comp_pct 0.835188 0.003188 261.983 < 2e-16 ***
## TD_pct 2.162783 0.101867 21.231 < 2e-16 ***
## int_pct -1.495482 0.224896 -6.650 4.60e-11 ***
## yards_per_att 1.762938 0.307636 5.731 1.29e-08 ***
## adj_yards_gain_per_att 2.294655 0.310165 7.398 2.72e-13 ***
## net_yds_per_att -3.585178 0.447287 -8.015 2.76e-15 ***
## adj_net_yds_per_att 3.750596 0.470547 7.971 3.89e-15 ***
## ---
## Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
##
## Residual standard error: 0.4275 on 1111 degrees of freedom
## Multiple R-squared: 0.9991, Adjusted R-squared: 0.9991
## F-statistic: 1.853e+05 on 7 and 1111 DF, p-value: < 2.2e-16
sjPlot::tab_model(qbr3, qbr4, qbr5, qbr6, show.aic = T, show.aicc = T)
| Rate | Rate | Rate | Rate | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Predictors | Estimates | CI | p | Estimates | CI | p | Estimates | CI | p | Estimates | CI | p |
| (Intercept) | -0.24 | -2.39 – 1.92 | 0.831 | 1.12 | 0.57 – 1.68 | <0.001 | 0.88 | 0.42 – 1.35 | <0.001 | 1.59 | 1.23 – 1.95 | <0.001 |
| Age | 0.00 | -0.00 – 0.01 | 0.091 | 0.01 | -0.00 – 0.01 | 0.079 | ||||||
| G | -0.02 | -0.05 – 0.01 | 0.106 | |||||||||
| Att | 0.00 | 0.00 – 0.00 | <0.001 | 0.00 | 0.00 – 0.00 | <0.001 | 0.00 | 0.00 – 0.00 | <0.001 | |||
| comp pct | 0.86 | 0.82 – 0.90 | <0.001 | 0.83 | 0.83 – 0.84 | <0.001 | 0.83 | 0.83 – 0.84 | <0.001 | 0.84 | 0.83 – 0.84 | <0.001 |
| Yds | -0.00 | -0.00 – -0.00 | 0.013 | -0.00 | -0.00 – -0.00 | 0.001 | -0.00 | -0.00 – -0.00 | <0.001 | |||
| TD pct | 2.13 | 1.91 – 2.35 | <0.001 | 2.13 | 1.91 – 2.34 | <0.001 | 2.20 | 2.00 – 2.40 | <0.001 | 2.16 | 1.96 – 2.36 | <0.001 |
| Int | -0.03 | -0.04 – -0.02 | <0.001 | -0.03 | -0.05 – -0.02 | <0.001 | -0.04 | -0.05 – -0.02 | <0.001 | |||
| int pct | -1.36 | -1.84 – -0.88 | <0.001 | -1.36 | -1.83 – -0.88 | <0.001 | -1.52 | -1.96 – -1.07 | <0.001 | -1.50 | -1.94 – -1.05 | <0.001 |
| yards per att | 1.92 | 1.25 – 2.60 | <0.001 | 2.04 | 1.39 – 2.70 | <0.001 | 1.75 | 1.15 – 2.34 | <0.001 | 1.76 | 1.16 – 2.37 | <0.001 |
| adj yards gain per att | 2.42 | 1.79 – 3.06 | <0.001 | 2.49 | 1.86 – 3.11 | <0.001 | 2.41 | 1.81 – 3.01 | <0.001 | 2.29 | 1.69 – 2.90 | <0.001 |
| yards per comp | 0.14 | -0.06 – 0.34 | 0.163 | |||||||||
| yards per game | -0.00 | -0.00 – -0.00 | 0.024 | -0.00 | -0.00 – 0.00 | 0.108 | ||||||
| Sk | 0.01 | 0.00 – 0.02 | 0.040 | 0.01 | -0.00 – 0.02 | 0.106 | ||||||
| tot sack yds | -0.00 | -0.00 – 0.00 | 0.187 | -0.00 | -0.00 – 0.00 | 0.248 | ||||||
| sack pct | -0.08 | -0.14 – -0.01 | 0.017 | -0.06 | -0.13 – -0.00 | 0.038 | ||||||
| net yds per att | -4.07 | -5.20 – -2.93 | <0.001 | -3.96 | -5.10 – -2.83 | <0.001 | -3.25 | -4.15 – -2.34 | <0.001 | -3.59 | -4.46 – -2.71 | <0.001 |
| adj net yds per att | 3.78 | 2.76 – 4.81 | <0.001 | 3.72 | 2.69 – 4.75 | <0.001 | 3.40 | 2.45 – 4.35 | <0.001 | 3.75 | 2.83 – 4.67 | <0.001 |
| Observations | 1119 | 1119 | 1119 | 1119 | ||||||||
| R2 / R2 adjusted | 0.999 / 0.999 | 0.999 / 0.999 | 0.999 / 0.999 | 0.999 / 0.999 | ||||||||
| AIC | 1247.129 | 1247.890 | 1247.558 | 1283.590 | ||||||||
| AICc | 1247.820 | 1248.445 | 1247.840 | 1283.752 | ||||||||
Although the variable co-efficients all become far more feasible and their interaction with observed values will bring much more feasible numbers into the model, the model information criteria suffers severely as the AIC and AICc values climb above the values observed for the null model which included all variables. Understanding this and examining the tab_model plot produced comparing the models and their coeffecients, some variables will be selected to re-integrate into the model. Understanding Interceptions was significant with a miniscule coeffecient, it is settled that this will remain out of the model as the variable is still covered throught the interception percentage of a quarterback, which would look to penalise quarterbacks who pass less but make poor decisions resulting in turnovers. Attempts was removed, however, will be elected to be reintegrated in. This metric will have a high observed value and should ensure the small coeffecient doesn’t impact a reasonable value building into the model. Furthermore, without it influencing the decision too heavily, we note that the current Quarterback Ratings formulas incorporate attempts into the calculations (which has been done already through the percentages of everything)
qbr7 <- lm(Rate ~ Att + comp_pct + TD_pct + int_pct + yards_per_att + adj_yards_gain_per_att +
net_yds_per_att + adj_net_yds_per_att, data = numeric_qb_data)
summary(qbr7)
##
## Call:
## lm(formula = Rate ~ Att + comp_pct + TD_pct + int_pct + yards_per_att +
## adj_yards_gain_per_att + net_yds_per_att + adj_net_yds_per_att,
## data = numeric_qb_data)
##
## Residuals:
## Min 1Q Median 3Q Max
## -1.1711 -0.1179 -0.0024 0.1068 12.9161
##
## Coefficients:
## Estimate Std. Error t value Pr(>|t|)
## (Intercept) 1.589e+00 1.833e-01 8.669 < 2e-16 ***
## Att -3.932e-06 8.430e-05 -0.047 0.963
## comp_pct 8.352e-01 3.219e-03 259.465 < 2e-16 ***
## TD_pct 2.163e+00 1.020e-01 21.204 < 2e-16 ***
## int_pct -1.495e+00 2.253e-01 -6.637 5.00e-11 ***
## yards_per_att 1.763e+00 3.078e-01 5.726 1.32e-08 ***
## adj_yards_gain_per_att 2.294e+00 3.103e-01 7.393 2.82e-13 ***
## net_yds_per_att -3.586e+00 4.481e-01 -8.003 3.04e-15 ***
## adj_net_yds_per_att 3.752e+00 4.721e-01 7.947 4.65e-15 ***
## ---
## Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
##
## Residual standard error: 0.4277 on 1110 degrees of freedom
## Multiple R-squared: 0.9991, Adjusted R-squared: 0.9991
## F-statistic: 1.62e+05 on 8 and 1110 DF, p-value: < 2.2e-16
sjPlot::tab_model(qbr3, qbr4, qbr5, qbr6, qbr7, show.aic = T, show.aicc = T)
| Rate | Rate | Rate | Rate | Rate | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Predictors | Estimates | CI | p | Estimates | CI | p | Estimates | CI | p | Estimates | CI | p | Estimates | CI | p |
| (Intercept) | -0.24 | -2.39 – 1.92 | 0.831 | 1.12 | 0.57 – 1.68 | <0.001 | 0.88 | 0.42 – 1.35 | <0.001 | 1.59 | 1.23 – 1.95 | <0.001 | 1.59 | 1.23 – 1.95 | <0.001 |
| Age | 0.00 | -0.00 – 0.01 | 0.091 | 0.01 | -0.00 – 0.01 | 0.079 | |||||||||
| G | -0.02 | -0.05 – 0.01 | 0.106 | ||||||||||||
| Att | 0.00 | 0.00 – 0.00 | <0.001 | 0.00 | 0.00 – 0.00 | <0.001 | 0.00 | 0.00 – 0.00 | <0.001 | -0.00 | -0.00 – 0.00 | 0.963 | |||
| comp pct | 0.86 | 0.82 – 0.90 | <0.001 | 0.83 | 0.83 – 0.84 | <0.001 | 0.83 | 0.83 – 0.84 | <0.001 | 0.84 | 0.83 – 0.84 | <0.001 | 0.84 | 0.83 – 0.84 | <0.001 |
| Yds | -0.00 | -0.00 – -0.00 | 0.013 | -0.00 | -0.00 – -0.00 | 0.001 | -0.00 | -0.00 – -0.00 | <0.001 | ||||||
| TD pct | 2.13 | 1.91 – 2.35 | <0.001 | 2.13 | 1.91 – 2.34 | <0.001 | 2.20 | 2.00 – 2.40 | <0.001 | 2.16 | 1.96 – 2.36 | <0.001 | 2.16 | 1.96 – 2.36 | <0.001 |
| Int | -0.03 | -0.04 – -0.02 | <0.001 | -0.03 | -0.05 – -0.02 | <0.001 | -0.04 | -0.05 – -0.02 | <0.001 | ||||||
| int pct | -1.36 | -1.84 – -0.88 | <0.001 | -1.36 | -1.83 – -0.88 | <0.001 | -1.52 | -1.96 – -1.07 | <0.001 | -1.50 | -1.94 – -1.05 | <0.001 | -1.49 | -1.94 – -1.05 | <0.001 |
| yards per att | 1.92 | 1.25 – 2.60 | <0.001 | 2.04 | 1.39 – 2.70 | <0.001 | 1.75 | 1.15 – 2.34 | <0.001 | 1.76 | 1.16 – 2.37 | <0.001 | 1.76 | 1.16 – 2.37 | <0.001 |
| adj yards gain per att | 2.42 | 1.79 – 3.06 | <0.001 | 2.49 | 1.86 – 3.11 | <0.001 | 2.41 | 1.81 – 3.01 | <0.001 | 2.29 | 1.69 – 2.90 | <0.001 | 2.29 | 1.69 – 2.90 | <0.001 |
| yards per comp | 0.14 | -0.06 – 0.34 | 0.163 | ||||||||||||
| yards per game | -0.00 | -0.00 – -0.00 | 0.024 | -0.00 | -0.00 – 0.00 | 0.108 | |||||||||
| Sk | 0.01 | 0.00 – 0.02 | 0.040 | 0.01 | -0.00 – 0.02 | 0.106 | |||||||||
| tot sack yds | -0.00 | -0.00 – 0.00 | 0.187 | -0.00 | -0.00 – 0.00 | 0.248 | |||||||||
| sack pct | -0.08 | -0.14 – -0.01 | 0.017 | -0.06 | -0.13 – -0.00 | 0.038 | |||||||||
| net yds per att | -4.07 | -5.20 – -2.93 | <0.001 | -3.96 | -5.10 – -2.83 | <0.001 | -3.25 | -4.15 – -2.34 | <0.001 | -3.59 | -4.46 – -2.71 | <0.001 | -3.59 | -4.47 – -2.71 | <0.001 |
| adj net yds per att | 3.78 | 2.76 – 4.81 | <0.001 | 3.72 | 2.69 – 4.75 | <0.001 | 3.40 | 2.45 – 4.35 | <0.001 | 3.75 | 2.83 – 4.67 | <0.001 | 3.75 | 2.83 – 4.68 | <0.001 |
| Observations | 1119 | 1119 | 1119 | 1119 | 1119 | ||||||||||
| R2 / R2 adjusted | 0.999 / 0.999 | 0.999 / 0.999 | 0.999 / 0.999 | 0.999 / 0.999 | 0.999 / 0.999 | ||||||||||
| AIC | 1247.129 | 1247.890 | 1247.558 | 1283.590 | 1285.588 | ||||||||||
| AICc | 1247.820 | 1248.445 | 1247.840 | 1283.752 | 1285.786 | ||||||||||
Examining the output, the integration of just Attempts back into the model negatively impacts the model, as the model information criteria sees a rise in score once again, while the Attempts variable is strongly insignificant in the model with a p-value of 0.963. This is more than likely a result of all variables in the model being a result of percentages (observation of the metric divided by attempts) or the statistical value per attempt. Ultimately, this means the attempt variable is currently baked into all metrics observed into the model at this stage.
Before accepting the model chosen to estimate the QBR of the quarterbacks, we will integrate sack percentage back into a model as a result of penalise quarterbacks who fail to throw the ball prior to being tackled in a situation where their offense loses yardage and they have to play “behind the sticks”.
qbr8 <- lm(Rate ~ sack_pct + comp_pct + TD_pct + int_pct + yards_per_att + adj_yards_gain_per_att +
net_yds_per_att + adj_net_yds_per_att, data = numeric_qb_data)
summary(qbr8)
##
## Call:
## lm(formula = Rate ~ sack_pct + comp_pct + TD_pct + int_pct +
## yards_per_att + adj_yards_gain_per_att + net_yds_per_att +
## adj_net_yds_per_att, data = numeric_qb_data)
##
## Residuals:
## Min 1Q Median 3Q Max
## -1.2016 -0.1147 -0.0022 0.1067 12.8543
##
## Coefficients:
## Estimate Std. Error t value Pr(>|t|)
## (Intercept) 1.756945 0.202789 8.664 < 2e-16 ***
## sack_pct -0.040319 0.020964 -1.923 0.0547 .
## comp_pct 0.834947 0.003187 262.021 < 2e-16 ***
## TD_pct 2.102851 0.106408 19.762 < 2e-16 ***
## int_pct -1.363613 0.234855 -5.806 8.34e-09 ***
## yards_per_att 1.923635 0.318421 6.041 2.09e-09 ***
## adj_yards_gain_per_att 2.398748 0.314482 7.628 5.12e-14 ***
## net_yds_per_att -4.075163 0.514285 -7.924 5.56e-15 ***
## adj_net_yds_per_att 3.956870 0.482059 8.208 6.17e-16 ***
## ---
## Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
##
## Residual standard error: 0.427 on 1110 degrees of freedom
## Multiple R-squared: 0.9991, Adjusted R-squared: 0.9991
## F-statistic: 1.625e+05 on 8 and 1110 DF, p-value: < 2.2e-16
sjPlot::tab_model(qbr3, qbr4, qbr5, qbr6, qbr8, show.aic = T, show.aicc = T)
| Rate | Rate | Rate | Rate | Rate | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Predictors | Estimates | CI | p | Estimates | CI | p | Estimates | CI | p | Estimates | CI | p | Estimates | CI | p |
| (Intercept) | -0.24 | -2.39 – 1.92 | 0.831 | 1.12 | 0.57 – 1.68 | <0.001 | 0.88 | 0.42 – 1.35 | <0.001 | 1.59 | 1.23 – 1.95 | <0.001 | 1.76 | 1.36 – 2.15 | <0.001 |
| Age | 0.00 | -0.00 – 0.01 | 0.091 | 0.01 | -0.00 – 0.01 | 0.079 | |||||||||
| G | -0.02 | -0.05 – 0.01 | 0.106 | ||||||||||||
| Att | 0.00 | 0.00 – 0.00 | <0.001 | 0.00 | 0.00 – 0.00 | <0.001 | 0.00 | 0.00 – 0.00 | <0.001 | ||||||
| comp pct | 0.86 | 0.82 – 0.90 | <0.001 | 0.83 | 0.83 – 0.84 | <0.001 | 0.83 | 0.83 – 0.84 | <0.001 | 0.84 | 0.83 – 0.84 | <0.001 | 0.83 | 0.83 – 0.84 | <0.001 |
| Yds | -0.00 | -0.00 – -0.00 | 0.013 | -0.00 | -0.00 – -0.00 | 0.001 | -0.00 | -0.00 – -0.00 | <0.001 | ||||||
| TD pct | 2.13 | 1.91 – 2.35 | <0.001 | 2.13 | 1.91 – 2.34 | <0.001 | 2.20 | 2.00 – 2.40 | <0.001 | 2.16 | 1.96 – 2.36 | <0.001 | 2.10 | 1.89 – 2.31 | <0.001 |
| Int | -0.03 | -0.04 – -0.02 | <0.001 | -0.03 | -0.05 – -0.02 | <0.001 | -0.04 | -0.05 – -0.02 | <0.001 | ||||||
| int pct | -1.36 | -1.84 – -0.88 | <0.001 | -1.36 | -1.83 – -0.88 | <0.001 | -1.52 | -1.96 – -1.07 | <0.001 | -1.50 | -1.94 – -1.05 | <0.001 | -1.36 | -1.82 – -0.90 | <0.001 |
| yards per att | 1.92 | 1.25 – 2.60 | <0.001 | 2.04 | 1.39 – 2.70 | <0.001 | 1.75 | 1.15 – 2.34 | <0.001 | 1.76 | 1.16 – 2.37 | <0.001 | 1.92 | 1.30 – 2.55 | <0.001 |
| adj yards gain per att | 2.42 | 1.79 – 3.06 | <0.001 | 2.49 | 1.86 – 3.11 | <0.001 | 2.41 | 1.81 – 3.01 | <0.001 | 2.29 | 1.69 – 2.90 | <0.001 | 2.40 | 1.78 – 3.02 | <0.001 |
| yards per comp | 0.14 | -0.06 – 0.34 | 0.163 | ||||||||||||
| yards per game | -0.00 | -0.00 – -0.00 | 0.024 | -0.00 | -0.00 – 0.00 | 0.108 | |||||||||
| Sk | 0.01 | 0.00 – 0.02 | 0.040 | 0.01 | -0.00 – 0.02 | 0.106 | |||||||||
| tot sack yds | -0.00 | -0.00 – 0.00 | 0.187 | -0.00 | -0.00 – 0.00 | 0.248 | |||||||||
| sack pct | -0.08 | -0.14 – -0.01 | 0.017 | -0.06 | -0.13 – -0.00 | 0.038 | -0.04 | -0.08 – 0.00 | 0.055 | ||||||
| net yds per att | -4.07 | -5.20 – -2.93 | <0.001 | -3.96 | -5.10 – -2.83 | <0.001 | -3.25 | -4.15 – -2.34 | <0.001 | -3.59 | -4.46 – -2.71 | <0.001 | -4.08 | -5.08 – -3.07 | <0.001 |
| adj net yds per att | 3.78 | 2.76 – 4.81 | <0.001 | 3.72 | 2.69 – 4.75 | <0.001 | 3.40 | 2.45 – 4.35 | <0.001 | 3.75 | 2.83 – 4.67 | <0.001 | 3.96 | 3.01 – 4.90 | <0.001 |
| Observations | 1119 | 1119 | 1119 | 1119 | 1119 | ||||||||||
| R2 / R2 adjusted | 0.999 / 0.999 | 0.999 / 0.999 | 0.999 / 0.999 | 0.999 / 0.999 | 0.999 / 0.999 | ||||||||||
| AIC | 1247.129 | 1247.890 | 1247.558 | 1283.590 | 1281.867 | ||||||||||
| AICc | 1247.820 | 1248.445 | 1247.840 | 1283.752 | 1282.066 | ||||||||||
Ultimately, reintegrating the sack metric back into the model proves no advantageous gain for the linear regression model looking to predict a quarterback’s passer rating. This is conceptually interesting as a portion of sacks are as a result of the quarterback moving into trouble, rather than solely the offensive line failing to protect the passer adequately. This could be a consideration if further tuning was performed to see if a penalisation factor for sacks, along with the current penalisation for interceptions, could be incorporated.
Ultimately, the QBR6 linear model shows both the strongest AIC and AICc information metrics, indicating that this model both has the best model fit of all built linear models as well as the best balance between model fit without becoming too complex in its formulation. It must be noted that there is a “double up” of variable observations when consider attempts is baked into a number of variables, while Interceptions and yards are also included in a number of variable formats in this model. However, for the simplistic sake of this investigation, the QBR6 model will be selected as the best model.
When considering the data used to fit, all Quarterback data was sourced from Pro Football Reference, a reputable open source statistics site for NFL and a number of other sports. The dataset used utilised quarterback passing data from 2001-2023. Therefore, using PFR to scrape data, the accuracy of the model was determined by calculating the QBR of quarterback’s in 2025 using their end-of-season data,l with comparisons to the actual calculated score made.
qb_2025 <- read.csv("2025 QB data.csv") |>
dplyr::glimpse()
## Rows: 102
## Columns: 34
## $ Rk <int> 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 1…
## $ Player <chr> "Matthew Stafford", "Jared Goff", "Dak Prescott", "D…
## $ Age <int> 37, 31, 32, 23, 28, 26, 24, 25, 27, 30, 29, 30, 27, …
## $ Team <chr> "LAR", "DET", "DAL", "NWE", "SEA", "JAX", "CHI", "DE…
## $ Pos <chr> "QB", "QB", "QB", "QB", "QB", "QB", "QB", "QB", "QB"…
## $ G <int> 17, 17, 17, 17, 17, 17, 17, 17, 16, 17, 17, 14, 15, …
## $ GS <int> 17, 17, 17, 17, 17, 17, 17, 17, 16, 17, 17, 14, 15, …
## $ QBrec <chr> "12/5/2000", "9/8/2000", "7/9/2001", "14/3/2000", "1…
## $ Cmp <int> 388, 393, 404, 354, 323, 341, 330, 388, 340, 343, 31…
## $ Att <int> 597, 578, 600, 492, 477, 560, 568, 612, 512, 543, 46…
## $ Cmp. <dbl> 65.0, 68.0, 67.3, 72.0, 67.7, 60.9, 58.1, 63.4, 66.4…
## $ Yds <int> 4707, 4564, 4552, 4394, 4048, 4007, 3942, 3931, 3727…
## $ TD <int> 46, 34, 30, 31, 25, 29, 27, 25, 26, 26, 25, 22, 23, …
## $ TD. <dbl> 7.7, 5.9, 5.0, 6.3, 5.2, 5.2, 4.8, 4.1, 5.1, 4.8, 5.…
## $ Int <int> 8, 8, 10, 8, 14, 12, 7, 11, 13, 11, 10, 11, 6, 8, 7,…
## $ Int. <dbl> 1.3, 1.4, 1.7, 1.6, 2.9, 2.1, 1.2, 1.8, 2.5, 2.0, 2.…
## $ X1D <int> 236, 223, 220, 203, 178, 194, 187, 196, 175, 172, 17…
## $ Succ. <dbl> 54.4, 48.7, 48.2, 54.7, 50.0, 45.8, 42.9, 45.6, 47.2…
## $ Lng <int> 88, 64, 86, 72, 67, 63, 65, 52, 60, 77, 54, 61, 59, …
## $ Y.A <dbl> 7.9, 7.9, 7.6, 8.9, 8.5, 7.2, 6.9, 6.4, 7.3, 6.8, 8.…
## $ AY.A <dbl> 8.8, 8.4, 7.8, 9.5, 8.2, 7.2, 7.3, 6.4, 7.2, 6.8, 8.…
## $ Y.C <dbl> 12.1, 11.6, 11.3, 12.4, 12.5, 11.8, 11.9, 10.1, 11.0…
## $ Y.G <dbl> 276.9, 268.5, 267.8, 258.5, 238.1, 235.7, 231.9, 231…
## $ Rate <dbl> 109.2, 105.5, 99.5, 113.5, 99.1, 91.0, 90.1, 87.8, 9…
## $ QBR <dbl> 71.2, 57.3, 70.2, 77.1, 55.6, 58.3, 58.2, 58.3, 60.6…
## $ Sk <int> 23, 38, 31, 47, 27, 41, 24, 22, 54, 36, 40, 34, 21, …
## $ Yds.1 <int> 150, 259, 208, 201, 186, 247, 165, 119, 301, 234, 29…
## $ Sk. <dbl> 3.71, 6.17, 4.91, 8.72, 5.36, 6.82, 4.05, 3.47, 9.54…
## $ NY.A <dbl> 7.4, 7.0, 6.9, 7.8, 7.7, 6.3, 6.4, 6.0, 6.1, 6.0, 6.…
## $ ANY.A <dbl> 8.3, 7.5, 7.1, 8.3, 7.4, 6.3, 6.8, 6.0, 5.9, 6.0, 6.…
## $ X4QC <int> 1, 1, 4, 1, 2, 3, 6, 5, 3, 4, 4, 1, 4, 0, 2, 1, 1, 3…
## $ GWD <int> 1, 3, 3, 2, 4, 5, 6, 7, 3, 4, 4, 1, 4, 0, 3, 1, 1, 3…
## $ Awards <chr> "PBAP-1AP MVP-1", "PB", "PBAP CPoY-3", "PBAP-2AP MVP…
## $ Player.additional <chr> "StafMa00", "GoffJa00", "PresDa01", "MayeDr00", "Dar…
All data cleaning and manipulation will follow the same stages as the prior set, renaming respective variables while removing any observations based off the same criteria used previously.
qb_2025 <- qb_2025 |>
dplyr::rename(
#Completion Percentage - Completed Passes / Attempted Passes
comp_pct = "Cmp.",
#TD Percentage (Percentage of throws which result in a Touchdown)
TD_pct = "TD.",
#Interception Percentage (Percentage of throws which result in an Interception)
int_pct = "Int.",
#Sack Percentage (Percentage of pass plays which resulted in the quarterback getting sacked)
sack_pct = "Sk.",
#Yards Per Attempt
yards_per_att = "Y.A",
#Adjusted Yards Gained per Attempt
adj_yards_gain_per_att = "AY.A",
#Yards Gained per Completion
yards_per_comp = "Y.C",
#Yards Gained from Throws per Game (seasonal data)
yards_per_game = "Y.G",
#Total Sack Yards cinceded
tot_sack_yds = "Yds.1",
#Net Yards Gained per Attempt
net_yds_per_att = "NY.A",
#Adjusted Net Yards Gained per Attempt
adj_net_yds_per_att = "ANY.A"
)
qb_2025 <- qb_2025 |>
subset(
GS > 1
)
qb_2025 <- qb_2025 |>
subset(
Att > 50
)
#Simple EDA
ggplot(data = qb_2025) +
geom_histogram(aes(x = Att), bins = 10)
#Low number of observations - 10 bins, highest observations in a bin is 8.
#Expected though as data only encompasses 1 season - 55 quarterback observation considered in the end
qb_2025$lm_qbr <- predict(qbr6, newdata = qb_2025)
glimpse(qb_2025)
## Rows: 55
## Columns: 35
## $ Rk <int> 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, …
## $ Player <chr> "Matthew Stafford", "Jared Goff", "Dak Prescott…
## $ Age <int> 37, 31, 32, 23, 28, 26, 24, 25, 27, 30, 29, 30,…
## $ Team <chr> "LAR", "DET", "DAL", "NWE", "SEA", "JAX", "CHI"…
## $ Pos <chr> "QB", "QB", "QB", "QB", "QB", "QB", "QB", "QB",…
## $ G <int> 17, 17, 17, 17, 17, 17, 17, 17, 16, 17, 17, 14,…
## $ GS <int> 17, 17, 17, 17, 17, 17, 17, 17, 16, 17, 17, 14,…
## $ QBrec <chr> "12/5/2000", "9/8/2000", "7/9/2001", "14/3/2000…
## $ Cmp <int> 388, 393, 404, 354, 323, 341, 330, 388, 340, 34…
## $ Att <int> 597, 578, 600, 492, 477, 560, 568, 612, 512, 54…
## $ comp_pct <dbl> 65.0, 68.0, 67.3, 72.0, 67.7, 60.9, 58.1, 63.4,…
## $ Yds <int> 4707, 4564, 4552, 4394, 4048, 4007, 3942, 3931,…
## $ TD <int> 46, 34, 30, 31, 25, 29, 27, 25, 26, 26, 25, 22,…
## $ TD_pct <dbl> 7.7, 5.9, 5.0, 6.3, 5.2, 5.2, 4.8, 4.1, 5.1, 4.…
## $ Int <int> 8, 8, 10, 8, 14, 12, 7, 11, 13, 11, 10, 11, 6, …
## $ int_pct <dbl> 1.3, 1.4, 1.7, 1.6, 2.9, 2.1, 1.2, 1.8, 2.5, 2.…
## $ X1D <int> 236, 223, 220, 203, 178, 194, 187, 196, 175, 17…
## $ Succ. <dbl> 54.4, 48.7, 48.2, 54.7, 50.0, 45.8, 42.9, 45.6,…
## $ Lng <int> 88, 64, 86, 72, 67, 63, 65, 52, 60, 77, 54, 61,…
## $ yards_per_att <dbl> 7.9, 7.9, 7.6, 8.9, 8.5, 7.2, 6.9, 6.4, 7.3, 6.…
## $ adj_yards_gain_per_att <dbl> 8.8, 8.4, 7.8, 9.5, 8.2, 7.2, 7.3, 6.4, 7.2, 6.…
## $ yards_per_comp <dbl> 12.1, 11.6, 11.3, 12.4, 12.5, 11.8, 11.9, 10.1,…
## $ yards_per_game <dbl> 276.9, 268.5, 267.8, 258.5, 238.1, 235.7, 231.9…
## $ Rate <dbl> 109.2, 105.5, 99.5, 113.5, 99.1, 91.0, 90.1, 87…
## $ QBR <dbl> 71.2, 57.3, 70.2, 77.1, 55.6, 58.3, 58.2, 58.3,…
## $ Sk <int> 23, 38, 31, 47, 27, 41, 24, 22, 54, 36, 40, 34,…
## $ tot_sack_yds <int> 150, 259, 208, 201, 186, 247, 165, 119, 301, 23…
## $ sack_pct <dbl> 3.71, 6.17, 4.91, 8.72, 5.36, 6.82, 4.05, 3.47,…
## $ net_yds_per_att <dbl> 7.4, 7.0, 6.9, 7.8, 7.7, 6.3, 6.4, 6.0, 6.1, 6.…
## $ adj_net_yds_per_att <dbl> 8.3, 7.5, 7.1, 8.3, 7.4, 6.3, 6.8, 6.0, 5.9, 6.…
## $ X4QC <int> 1, 1, 4, 1, 2, 3, 6, 5, 3, 4, 4, 1, 4, 0, 2, 1,…
## $ GWD <int> 1, 3, 3, 2, 4, 5, 6, 7, 3, 4, 4, 1, 4, 0, 3, 1,…
## $ Awards <chr> "PBAP-1AP MVP-1", "PB", "PBAP CPoY-3", "PBAP-2A…
## $ Player.additional <chr> "StafMa00", "GoffJa00", "PresDa01", "MayeDr00",…
## $ lm_qbr <dbl> 109.30511, 105.28385, 99.25668, 113.61002, 98.9…
qbr_comparisons <- qb_2025 |>
dplyr::select(Player, Rate, lm_qbr) |>
glimpse()
## Rows: 55
## Columns: 3
## $ Player <chr> "Matthew Stafford", "Jared Goff", "Dak Prescott", "Drake Maye",…
## $ Rate <dbl> 109.2, 105.5, 99.5, 113.5, 99.1, 91.0, 90.1, 87.8, 94.1, 90.6, …
## $ lm_qbr <dbl> 109.30511, 105.28385, 99.25668, 113.61002, 98.99027, 90.81450, …
Using the QBR6’s model summary, we can determine that the residual standard error is 0.4275. This will be used as the comparison metric to see whether the predicted QBR and the true QBR fall within such a window from eachother.
qbr6_rse <- summary(qbr6)$sigma # residual standard error from training
qbr_comparisons <- qbr_comparisons %>%
dplyr::mutate(within_2sd = factor(abs(Rate - lm_qbr) <= 2 * qbr6_rse,
levels = c(TRUE, FALSE)))
good_predictions <- sum(qbr_comparisons$within_2sd == TRUE, na.rm = TRUE)
good_predictions_pct <- sum(qbr_comparisons$within_2sd == TRUE, na.rm = TRUE) / nrow(qbr_comparisons)
good_predictions
## [1] 54
good_predictions_pct
## [1] 0.9818182
As evidenced above, in a 2025 Quarterback Passing dataset of 55 rows, 54 of those player-seasonal observations calculated a quarterback passer rating that was considered a strong predicted value, falling within 2 residual standard error’s from the true observed value. Although a small sample size, this is a strong indicator that the linear regression model chosen successfully model’s a player-seasonal quarterback passer rating.
Further extensions of these analytics and research would be to consider subsetting the data into player-game observations, where the accumulation of metrics would not be as great and seeing how the linear regression model performs and any adjustments that may be required as a result of adjusting the data it fits against. Furthermore, using this data could be used to refine the quarterback passer rating system if desired. In saying that, the rating system comprises of four scaled metrics in a total formula - Completion Percentage, Touchdown Percentage, Interception Percentage and a Yards per Attempt component. All of these factors are baked into our model along with additional adjusted yards per attempt metrics, scaling the observed value by certain preset formulas (with the net yards formulas including sack yards in their as a penalising factor). Therefore, it is also a strong suggestion that, when using a backwards selection method for a simple linear regression on quarterback passer rating, the analytical methods in R are able to decipher the key metrics for a QBR calculation are, in fact, those that are in the current calculation method. This is a strong reinforcement that the linear regression method worked well on this data and the final linear regression model was the correct choice of all options.
Kaggle Passing Data - Click here to visit Google
Pro Football Reference 2025 Quarterback Passing Data - Click here to visit Google