Logistic regression is an important statistical methodology to determine the probability of an event occurring based on a binary response variable. This binary variable takes on two responses: 1 and 0. This run through looks at conducting logistic regression on NBA data from the 2024 season. Each row of data contains one team’s performance in a game, using W_L as the binary variable for win (1) and loss (0), while the rest of the remaining statistics are candidate predictors.
## Load in packages
library(hoopR)
library(dplyr)
library(tidyr)
library(skimr)
library(ggplot2)
library(GGally)
library(FactoMineR)
library(factoextra)
library(sjPlot)
library(performance)
Use the hoopR package to import the 2024 game-by-game team data. Clean the data frame to allow for analysis.
## Load data from hoopR package
nba_data <- load_nba_team_box(seasons = 2024)
## Clean, select and update column names
nba <- nba_data %>% filter(season_type == 2) %>%
transmute(TEAM = team_abbreviation, DATE = game_date, W_L = if_else(team_winner, "W", "L"),
PTS = team_score, FGM = field_goals_made, FGA = field_goals_attempted, FG_P = field_goal_pct,
`3PM` = three_point_field_goals_made, `3PA` = three_point_field_goals_attempted,
`3P_P` = three_point_field_goal_pct, FTM = free_throws_made, FTA = free_throws_attempted,
FT_P = free_throw_pct, OREB = offensive_rebounds, DREB = defensive_rebounds, REB = total_rebounds,
AST = assists, STL = steals, BLK = blocks, TOV = turnovers, PF = fouls,
POINT_DFF = team_score - opponent_team_score)
nba <- nba %>% group_by(TEAM) %>%
mutate(WIN_P = (sum(W_L == "W", na.rm = TRUE)/sum(W_L %in% c("W","L")))*100) %>%
ungroup()
## Review descriptive statistics
skim(nba)
| Name | nba |
| Number of rows | 2464 |
| Number of columns | 23 |
| _______________________ | |
| Column type frequency: | |
| character | 2 |
| Date | 1 |
| numeric | 20 |
| ________________________ | |
| Group variables | None |
Variable type: character
| skim_variable | n_missing | complete_rate | min | max | empty | n_unique | whitespace |
|---|---|---|---|---|---|---|---|
| TEAM | 0 | 1 | 2 | 4 | 0 | 32 | 0 |
| W_L | 0 | 1 | 1 | 1 | 0 | 2 | 0 |
Variable type: Date
| skim_variable | n_missing | complete_rate | min | max | median | n_unique |
|---|---|---|---|---|---|---|
| DATE | 0 | 1 | 2023-10-24 | 2024-04-14 | 2024-01-19 | 162 |
Variable type: numeric
| skim_variable | n_missing | complete_rate | mean | sd | p0 | p25 | p50 | p75 | p100 | hist |
|---|---|---|---|---|---|---|---|---|---|---|
| PTS | 0 | 1 | 114.28 | 13.06 | 73.0 | 105.0 | 114.00 | 123.00 | 211.0 | ▂▇▂▁▁ |
| FGM | 0 | 1 | 42.20 | 5.46 | 26.0 | 38.0 | 42.00 | 46.00 | 83.0 | ▂▇▂▁▁ |
| FGA | 0 | 1 | 88.95 | 7.18 | 67.0 | 84.0 | 89.00 | 93.00 | 146.0 | ▂▇▁▁▁ |
| FG_P | 0 | 1 | 47.53 | 5.51 | 27.7 | 43.8 | 47.50 | 51.20 | 67.1 | ▁▃▇▃▁ |
| 3PM | 0 | 1 | 12.85 | 3.89 | 2.0 | 10.0 | 13.00 | 15.00 | 42.0 | ▃▇▁▁▁ |
| 3PA | 0 | 1 | 35.14 | 6.71 | 12.0 | 30.0 | 35.00 | 39.00 | 97.0 | ▂▇▁▁▁ |
| 3P_P | 0 | 1 | 36.48 | 8.35 | 6.9 | 31.0 | 36.50 | 41.70 | 64.5 | ▁▃▇▃▁ |
| FTM | 0 | 1 | 17.03 | 5.91 | 0.0 | 13.0 | 17.00 | 21.00 | 44.0 | ▁▇▆▁▁ |
| FTA | 0 | 1 | 21.71 | 7.02 | 0.0 | 17.0 | 21.00 | 26.00 | 52.0 | ▁▇▇▂▁ |
| FT_P | 0 | 1 | 78.31 | 10.28 | 0.0 | 72.0 | 78.95 | 85.20 | 100.0 | ▁▁▁▇▇ |
| OREB | 0 | 1 | 10.56 | 3.83 | 0.0 | 8.0 | 10.00 | 13.00 | 28.0 | ▁▇▅▁▁ |
| DREB | 0 | 1 | 32.99 | 5.41 | 16.0 | 29.0 | 33.00 | 36.00 | 55.0 | ▁▆▇▂▁ |
| REB | 0 | 1 | 43.55 | 6.59 | 25.0 | 39.0 | 43.00 | 48.00 | 74.0 | ▁▇▆▁▁ |
| AST | 0 | 1 | 26.69 | 5.16 | 11.0 | 23.0 | 27.00 | 30.00 | 60.0 | ▁▇▂▁▁ |
| STL | 0 | 1 | 7.47 | 2.82 | 0.0 | 6.0 | 7.00 | 9.00 | 20.0 | ▂▇▅▁▁ |
| BLK | 0 | 1 | 5.14 | 2.60 | 0.0 | 3.0 | 5.00 | 7.00 | 17.0 | ▆▇▅▁▁ |
| TOV | 0 | 1 | 12.90 | 3.76 | 2.0 | 10.0 | 13.00 | 15.00 | 28.0 | ▁▇▇▂▁ |
| PF | 0 | 1 | 18.72 | 4.19 | 1.0 | 16.0 | 19.00 | 21.00 | 34.0 | ▁▂▇▅▁ |
| POINT_DFF | 0 | 1 | 0.00 | 15.79 | -62.0 | -10.0 | 0.00 | 10.00 | 62.0 | ▁▂▇▂▁ |
| WIN_P | 0 | 1 | 50.00 | 16.13 | 0.0 | 37.8 | 56.63 | 59.76 | 100.0 | ▁▃▇▃▁ |
## Field goal percentage vs win percentage plot
ggplot(nba, aes(x = FG_P, y = WIN_P)) +
geom_point() +
geom_smooth(method = "lm") +
xlab("Field Goals (%)") + ylab("Wins (%)")
## `geom_smooth()` using formula = 'y ~ x'
## Free throw percentage vs win percentage plot
ggplot(nba, aes(x = FT_P, y = WIN_P)) +
geom_point() +
geom_smooth(method = "lm") +
xlab("FT%") + ylab("Win Percentage")
## `geom_smooth()` using formula = 'y ~ x'
## Three point percentage vs win percentage plot
ggplot(nba, aes(x = `3P_P`, y = WIN_P)) +
geom_point() +
geom_smooth(method = "lm") +
xlab("3 point pct") + ylab("Win Percentage")
## `geom_smooth()` using formula = 'y ~ x'
## Three point attempts vs win percentage plot
ggplot(nba, aes(x = `3PA`, y = WIN_P)) +
geom_point() +
geom_smooth(method = "lm") +
xlab("3P Attempts") + ylab("Win Percentage")
## `geom_smooth()` using formula = 'y ~ x'
## Average assists vs win percentage plot
ggplot(nba, aes(x = AST, y = WIN_P)) +
geom_point() +
geom_smooth(method = "lm") +
xlab("Average assists") + ylab("Win Percentage")
## `geom_smooth()` using formula = 'y ~ x'
# Remove categorical predictors
nba_cor <- nba |> dplyr::select(-c(TEAM:W_L))
# Check correlations between variables
ggpairs(nba_cor)
Principal component analysis, or PCA, is an exploratory technique used to simplify a dataset with many correlated variables.
PCA is a useful step before modelling because it helps identify redundancy and potential multicollinearity among predictors. In other words, it can show when several variables are measuring nearly the same underlying idea, which is important when building a stable and interpretable model.
Including multiple highly related predictors may make individual regression coefficients more difficult to interpret because the model can struggle to distinguish their separate contributions.
## Remove any meta data variables
nba_pca_data <- nba |> dplyr::select(-TEAM, -DATE, -W_L, -POINT_DFF,-PTS, -WIN_P)
## Keep numeric variables only and remove missing rows
nba_pca_numeric <- nba_pca_data |> dplyr::select(where(is.numeric)) |> drop_na()
## Run PCA using FactoMineR package
nba_pca_fit <- PCA(nba_pca_numeric, scale.unit = TRUE, graph = TRUE)
## Check Eigenvalues
nba_pca_fit$eig
## eigenvalue percentage of variance cumulative percentage of variance
## comp 1 3.484232e+00 2.049548e+01 20.49548
## comp 2 2.586347e+00 1.521381e+01 35.70929
## comp 3 2.142486e+00 1.260286e+01 48.31215
## comp 4 1.572948e+00 9.252638e+00 57.56478
## comp 5 1.355378e+00 7.972811e+00 65.53759
## comp 6 1.096665e+00 6.450970e+00 71.98856
## comp 7 9.966475e-01 5.862632e+00 77.85120
## comp 8 9.470554e-01 5.570914e+00 83.42211
## comp 9 8.463786e-01 4.978698e+00 88.40081
## comp 10 7.744839e-01 4.555788e+00 92.95660
## comp 11 5.688270e-01 3.346041e+00 96.30264
## comp 12 3.519958e-01 2.070564e+00 98.37320
## comp 13 2.551864e-01 1.501096e+00 99.87430
## comp 14 1.006953e-02 5.923253e-02 99.93353
## comp 15 8.836880e-03 5.198165e-02 99.98551
## comp 16 2.463021e-03 1.448836e-02 100.00000
## comp 17 3.390429e-28 1.994370e-27 100.00000
## Check variable contributions and correlations
nba_pca_fit$var$contrib
## Dim.1 Dim.2 Dim.3 Dim.4 Dim.5
## FGM 19.005616115 0.050337419 2.654707937 0.363238975 14.62548960
## FGA 4.101018713 20.078361307 0.079975037 8.083203551 6.38184364
## FG_P 11.664839936 11.405264856 4.132936086 1.732808106 6.27042759
## 3PM 16.558283883 0.225401303 0.003894039 3.952706856 23.34654531
## 3PA 5.322995961 6.645474627 1.358475653 6.725240787 22.57932738
## 3P_P 12.077821638 7.854291076 1.067095207 0.215113839 6.41769112
## FTM 5.751856593 2.549198315 24.083427674 10.307275959 0.21488030
## FTA 5.774554199 2.063835438 21.346021030 9.213659051 0.07171156
## FT_P 0.256156474 0.554061958 3.440240185 1.397832946 0.45903289
## OREB 0.054020961 20.610903134 0.731398980 4.392537566 0.69329513
## DREB 0.787301649 5.845037698 15.957481034 17.248919333 1.02194976
## REB 0.352615571 21.338007620 14.262996655 4.816691601 0.12047925
## AST 17.079472434 0.484578236 1.193028821 0.005339695 2.32175513
## STL 0.008955865 0.009319926 0.370363993 5.483201453 10.49176885
## BLK 0.177440222 0.141028944 4.114177675 4.319564973 0.01168627
## TOV 0.871299106 0.103653671 1.262078276 13.458880305 4.51135472
## PF 0.155750679 0.041244472 3.941701717 8.283785005 0.46076150
nba_pca_fit$var$cos2
## Dim.1 Dim.2 Dim.3 Dim.4 Dim.5
## FGM 0.6621997281 0.0013019004 5.687674e-02 5.713562e-03 0.1982306588
## FGA 0.1428889997 0.5192961455 1.713454e-03 1.271446e-01 0.0864981006
## FG_P 0.4064300671 0.2949797540 8.854756e-02 2.725618e-02 0.0849879920
## 3PM 0.5769289992 0.0058296604 8.342924e-05 6.217404e-02 0.3164339236
## 3PA 0.1854655201 0.1718750503 2.910515e-02 1.057846e-01 0.3060352210
## 3P_P 0.4208193071 0.2031392413 2.286236e-02 3.383630e-03 0.0869839694
## FTM 0.2004080188 0.0659311205 5.159840e-01 1.621281e-01 0.0029124402
## FTA 0.2011988560 0.0533779511 4.573354e-01 1.449261e-01 0.0009719627
## FT_P 0.0089250854 0.0143299662 7.370665e-02 2.198719e-02 0.0062216305
## OREB 0.0018822155 0.5330695264 1.567012e-02 6.909235e-02 0.0093967692
## DREB 0.0274314147 0.1511729718 3.418867e-01 2.713166e-01 0.0138512815
## REB 0.0122859440 0.5518749733 3.055827e-01 7.576408e-02 0.0016329491
## AST 0.5950884166 0.0125328759 2.556047e-02 8.399066e-05 0.0314685567
## STL 0.0003120431 0.0002410457 7.934995e-03 8.624793e-02 0.1422031199
## BLK 0.0061824287 0.0036474982 8.814567e-02 6.794453e-02 0.0001583932
## TOV 0.0303580809 0.0026808439 2.703985e-02 2.117012e-01 0.0611459064
## PF 0.0054267147 0.0010667253 8.445039e-02 1.302997e-01 0.0062450598
## Scree plot
fviz_eig(nba_pca_fit)
We can see with the variable correlation circle specific groupings of variables that may show redundency. From the above, we can see the following groupings: - Field goals: FGM, FGA, FG_P - Three-point shooting: 3PM, 3PA, 3P_P - Free throws: FTM, FTA, FT_P - Rebounding: OREB, DREB, REB
A sensible reduction would include: - Keep FG_P and drop FGM and FGA - Keep 3P_P and drop 3PM and 3PA - Keep FT_P and drop FTM and FTA - Keep REB and drop OREB and DREB
We will also coerce W_L to a factor using as.factor() so the matrix coding is 0 and 1 for modelling.
## Create reduced dataset and coerce W_L to factor
nba_red <- nba |> dplyr::select(W_L, FG_P, `3P_P`, FT_P, REB, AST, STL, BLK, TOV, PF) |>
dplyr::mutate(W_L = as.factor(W_L))
This is a practical reduction based on the PCA visualisation and the definitions of the variables. The correlation matrix can be used to confirm the strength of the relationships, and the reduced model can then be assessed through the logistic regression results.
Perform a logistic regression to determine the variables from our reduced dataset that best contribute towards wins vs. losses. The response/target variable is W_L.
The glm() function fits a generalised linear model, while family = binomial specifies logistic regression for a binary outcome. The formula W_L ~ . indicates that all remaining variables in nba_red, other than the response variable, are included as predictors.
The model summary provides estimated coefficients, standard errors, test statistics and p-values. Logistic regression coefficients are expressed in log-odds, so exponentiating a coefficient converts it into an odds ratio. An odds ratio greater than one indicates higher odds of winning for a one-unit increase in the predictor, while an odds ratio below one indicates lower odds, holding the other predictors constant.
## Logistic regression using glm
nba_1_glm <- glm(W_L ~ ., family = binomial, data = nba_red)
## Model summary
summary(nba_1_glm)
##
## Call:
## glm(formula = W_L ~ ., family = binomial, data = nba_red)
##
## Coefficients:
## Estimate Std. Error z value Pr(>|z|)
## (Intercept) -34.223346 1.537404 -22.260 < 2e-16 ***
## FG_P 0.391262 0.020179 19.389 < 2e-16 ***
## `3P_P` 0.089749 0.009531 9.416 < 2e-16 ***
## FT_P 0.039277 0.006179 6.357 2.06e-10 ***
## REB 0.283230 0.014050 20.158 < 2e-16 ***
## AST -0.063336 0.015095 -4.196 2.72e-05 ***
## STL 0.309375 0.024882 12.434 < 2e-16 ***
## BLK 0.155353 0.024069 6.454 1.09e-10 ***
## TOV -0.210338 0.018018 -11.674 < 2e-16 ***
## PF -0.093090 0.015181 -6.132 8.68e-10 ***
## ---
## Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
##
## (Dispersion parameter for binomial family taken to be 1)
##
## Null deviance: 3415.8 on 2463 degrees of freedom
## Residual deviance: 1733.4 on 2454 degrees of freedom
## AIC: 1753.4
##
## Number of Fisher Scoring iterations: 6
## Check model with sjPlot package
tab_model(nba_1_glm,show.aicc = TRUE)
| W L | |||
|---|---|---|---|
| Predictors | Odds Ratios | CI | p |
| (Intercept) | 0.00 | 0.00 – 0.00 | <0.001 |
| FG P | 1.48 | 1.42 – 1.54 | <0.001 |
| 3P P | 1.09 | 1.07 – 1.11 | <0.001 |
| FT P | 1.04 | 1.03 – 1.05 | <0.001 |
| REB | 1.33 | 1.29 – 1.37 | <0.001 |
| AST | 0.94 | 0.91 – 0.97 | <0.001 |
| STL | 1.36 | 1.30 – 1.43 | <0.001 |
| BLK | 1.17 | 1.11 – 1.23 | <0.001 |
| TOV | 0.81 | 0.78 – 0.84 | <0.001 |
| PF | 0.91 | 0.88 – 0.94 | <0.001 |
| Observations | 2464 | ||
| R2 Tjur | 0.551 | ||
| AICc | 1753.508 | ||
## Overdispersion check
check_overdispersion(nba_1_glm) # No overdispersion
## Registered S3 method overwritten by 'lme4':
## method from
## na.action.merMod car
## # Overdispersion test
##
## dispersion ratio = 1.006
## p-value = 0.84
## No overdispersion detected.
## Model diagnostics check
check_model(nba_1_glm)
The fitted coefficients describe the association between each predictor and the odds of winning, after accounting for the other predictors in the model. The odds ratios provide a more intuitive way to interpret these relationships than the raw log-odds coefficients.
The p-values provide evidence about whether individual coefficients differ from zero under the model assumptions. However, statistical significance alone does not establish that a predictor is practically important or that the model will accurately predict future results.
The diagnostic checks provide additional information about model assumptions and potential problems. The results should be interpreted alongside the data structure, the relationships between predictors and the purpose of the analysis.