Linear Regression

Linear regression is a statistical method used to examine the relationship between a continuous outcome variable and one or more predictor variables. This resource demonstrates how linear regression can be applied to NBA team statistics to investigate which performance measures are associated with season winning percentage.

## Load in packages
library(hoopR)
library(dplyr)
library(ggplot2)
library(performance)
library(sjPlot)

Data Import and Cleaning

The analysis uses data from the 2023–24 NBA regular season. The data contains one row per team per game. To investigate season winning percentage, the game-level statistics are aggregated to produce one observation per team.

You can also embed plots, for example:

## Load data from hoopR package
nba_data <- load_nba_team_box(seasons = 2024)

## Clean, select and update column names
nba <- nba_data %>% filter(season_type == 2) %>% 
  transmute(TEAM = team_abbreviation, DATE = game_date, W_L = if_else(team_winner, "W", "L"),
    PTS = team_score, FGM = field_goals_made, FGA = field_goals_attempted, FG_P = field_goal_pct,
    `3PM` = three_point_field_goals_made, `3PA` = three_point_field_goals_attempted,
    `3P_P` = three_point_field_goal_pct, FTM = free_throws_made, FTA = free_throws_attempted,
    FT_P = free_throw_pct, OREB = offensive_rebounds, DREB = defensive_rebounds, REB = total_rebounds,
    AST = assists, STL = steals, BLK = blocks, TOV = turnovers, PF = fouls, 
    POINT_DFF = team_score - opponent_team_score)

## Aggregate statistics to team level
nba_team <- nba %>% group_by(TEAM) %>%summarise(WIN_P = mean(W_L == "W", na.rm = TRUE) * 100,
                                                FG_P = mean(FG_P, na.rm = TRUE),
                                                `3P_P` = mean(`3P_P`, na.rm = TRUE),
                                                FT_P = mean(FT_P, na.rm = TRUE),
                                                REB = mean(REB, na.rm = TRUE),
                                                AST = mean(AST, na.rm = TRUE),
                                                TOV = mean(TOV, na.rm = TRUE), .groups = "drop")

head(nba_team)
## # A tibble: 6 × 8
##   TEAM  WIN_P  FG_P `3P_P`  FT_P   REB   AST   TOV
##   <chr> <dbl> <dbl>  <dbl> <dbl> <dbl> <dbl> <dbl>
## 1 ATL    43.9  46.6   36.4  79.4  44.7  26.6  12.8
## 2 BKN    39.0  45.7   36.2  76.2  44.1  25.6  12.3
## 3 BOS    78.0  48.8   38.8  79.6  46.3  26.9  11.3
## 4 CHA    25.6  46.1   35.3  78.9  40.3  24.8  13.0
## 5 CHI    47.6  47.1   35.7  79.3  43.8  25.0  11.7
## 6 CLE    58.5  48.0   36.6  76.2  43.3  28.0  12.7

Exploratory Data Analysis

ggplot(nba_team, aes(x = FG_P, y = WIN_P)) + geom_point() + geom_smooth(method = "lm", se = TRUE) +
  labs(title = "Field-Goal Percentage and Winning Percentage", x = "Average Field-Goal Percentage",
       y = "Season Winning Percentage (%)") + theme_minimal()
## `geom_smooth()` using formula = 'y ~ x'

The plot illustrates the relationship between average field-goal percentage and season winning percentage. The fitted line summarises the overall linear trend, while the shaded region represents its confidence interval.

Fit the Linear Regression Model

Linear regression estimates the relationship between the outcome and the selected predictors. Here, season winning percentage is the response variable, and shooting percentages, rebounds, assists and turnovers are the predictors.

## Fit linear regression model
nba_lm <- lm(WIN_P ~ FG_P + `3P_P` + FT_P + REB + AST + TOV, data = nba_team)

# Examine model results
summary(nba_lm)
## 
## Call:
## lm(formula = WIN_P ~ FG_P + `3P_P` + FT_P + REB + AST + TOV, 
##     data = nba_team)
## 
## Residuals:
##      Min       1Q   Median       3Q      Max 
## -14.7243  -3.5699   0.9421   3.7551  12.9138 
## 
## Coefficients:
##              Estimate Std. Error t value Pr(>|t|)   
## (Intercept) -222.1121    62.2156  -3.570  0.00148 **
## FG_P           3.4828     1.5879   2.193  0.03780 * 
## `3P_P`         5.8510     1.7416   3.360  0.00251 **
## FT_P          -0.6667     0.4182  -1.594  0.12347   
## REB            1.4125     0.7341   1.924  0.06579 . 
## AST           -2.3595     0.7151  -3.299  0.00291 **
## TOV           -4.1312     1.3291  -3.108  0.00465 **
## ---
## Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
## 
## Residual standard error: 7.002 on 25 degrees of freedom
## Multiple R-squared:  0.9039, Adjusted R-squared:  0.8808 
## F-statistic: 39.19 on 6 and 25 DF,  p-value: 1.563e-11
# Display formatted regression table
sjPlot::tab_model(nba_lm)
  WIN P
Predictors Estimates CI p
(Intercept) -222.11 -350.25 – -93.98 0.001
FG P 3.48 0.21 – 6.75 0.038
3P P 5.85 2.26 – 9.44 0.003
FT P -0.67 -1.53 – 0.19 0.123
REB 1.41 -0.10 – 2.92 0.066
AST -2.36 -3.83 – -0.89 0.003
TOV -4.13 -6.87 – -1.39 0.005
Observations 32
R2 / R2 adjusted 0.904 / 0.881

The model coefficients estimate the expected change in winning percentage associated with a one-unit increase in a predictor, holding the other predictors constant. The coefficient signs indicate the direction of the estimated relationship, while the p-values provide evidence about whether individual coefficients differ from zero under the model assumptions.

Model Diagnostics

Model diagnostics help assess whether the linear regression assumptions are reasonable. These include linearity, constant variance of residuals, approximate normality of residuals and the presence of influential observations.

check_model(nba_lm)

The diagnostic plots should be examined for systematic patterns, unequal residual variance, departures from normality and potentially influential observations. Diagnostic results should be interpreted together rather than relying on a single plot.