1 Introduction

A question that can be asked in basketball is, how valuable is ball movement to a team’s offence? An assist is a way to measure ball movement in basketball as it is a pass that directly leads to a teammate scoring. Teams that share the ball well tend to record more assists. A simple way to explore the relationship between assists and a team’s overall offence is by creating a simple linear regression model.

2 Load the packages

library(tidyverse)
library(hoopR)

3 Import the data

The hoopR function load_nba_team_box() downloads team statistics for every game in a season.

nba <- load_nba_team_box(seasons = 2024)

Each row of this data is one team in one game. Let’s see which columns are available:

names(nba)
##  [1] "game_id"                           "season"                           
##  [3] "season_type"                       "game_date"                        
##  [5] "game_date_time"                    "team_id"                          
##  [7] "team_uid"                          "team_slug"                        
##  [9] "team_location"                     "team_name"                        
## [11] "team_abbreviation"                 "team_display_name"                
## [13] "team_short_display_name"           "team_color"                       
## [15] "team_alternate_color"              "team_logo"                        
## [17] "team_home_away"                    "team_score"                       
## [19] "team_winner"                       "assists"                          
## [21] "blocks"                            "defensive_rebounds"               
## [23] "fast_break_points"                 "field_goal_pct"                   
## [25] "field_goals_made"                  "field_goals_attempted"            
## [27] "flagrant_fouls"                    "fouls"                            
## [29] "free_throw_pct"                    "free_throws_made"                 
## [31] "free_throws_attempted"             "largest_lead"                     
## [33] "offensive_rebounds"                "points_in_paint"                  
## [35] "steals"                            "team_turnovers"                   
## [37] "technical_fouls"                   "three_point_field_goal_pct"       
## [39] "three_point_field_goals_made"      "three_point_field_goals_attempted"
## [41] "total_rebounds"                    "total_technical_fouls"            
## [43] "total_turnovers"                   "turnover_points"                  
## [45] "turnovers"                         "opponent_team_id"                 
## [47] "opponent_team_uid"                 "opponent_team_slug"               
## [49] "opponent_team_location"            "opponent_team_name"               
## [51] "opponent_team_abbreviation"        "opponent_team_display_name"       
## [53] "opponent_team_short_display_name"  "opponent_team_color"              
## [55] "opponent_team_alternate_color"     "opponent_team_logo"               
## [57] "opponent_team_score"

We only need a few columns, so we keep those and remove any rows with missing values.

games <- nba |> 
  select(game_id, team_display_name, team_score, assists) |> 
  drop_na()
head(games)
## # A tibble: 6 × 4
##     game_id team_display_name team_score assists
##       <int> <chr>                  <int>   <int>
## 1 401656363 Dallas Mavericks          88      18
## 2 401656363 Boston Celtics           106      25
## 3 401656362 Boston Celtics            84      18
## 4 401656362 Dallas Mavericks         122      21
## 5 401656361 Boston Celtics           106      26
## 6 401656361 Dallas Mavericks          99      15

Every row now shows one team’s total points and assists in a single game.

4 Explore the data

A quick scatterplot to visualise the relationship between team points and assists.

ggplot(games, aes(x = assists, y = team_score)) +
  geom_point(alpha = 0.3, colour = "blue") +
  labs(title = "NBA team points vs assists (2023-24 season)",
       x = "Assists in a game",
       y = "Points scored in a game")

The dots drift from the bottom left to the top right. Games with more assists tend to be games with more points, showing a positive relationship. We can also get the correlation coefficient of this relationship.

cor(games$assists, games$team_score)
## [1] 0.6177835

A value of 0.62 tells us the relationship is positive. The closer to 1, the stronger the straight-line relationship.

5 Fit the linear regression model

As the two variables share a linear relationship we can now fit a linear regression model using the lm() function.

model <- lm(team_score ~ assists, data = games)

Now let’s look at the results:

summary(model)
## 
## Call:
## lm(formula = team_score ~ assists, data = games)
## 
## Residuals:
##     Min      1Q  Median      3Q     Max 
## -30.473  -7.036  -0.417   6.583  66.749 
## 
## Coefficients:
##             Estimate Std. Error t value Pr(>|t|)    
## (Intercept) 72.32939    1.04515   69.20   <2e-16 ***
## assists      1.56351    0.03875   40.35   <2e-16 ***
## ---
## Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
## 
## Residual standard error: 10.35 on 2638 degrees of freedom
## Multiple R-squared:  0.3817, Adjusted R-squared:  0.3814 
## F-statistic:  1628 on 1 and 2638 DF,  p-value: < 2.2e-16

6 Interpret the model output

The summary() output contains a table of coefficients. Here it is on its own:

coefs <- summary(model)$coefficients
coefs
##              Estimate Std. Error  t value     Pr(>|t|)
## (Intercept) 72.329391 1.04514544 69.20510  0.00000e+00
## assists      1.563512 0.03874739 40.35141 1.07729e-277

The Estimate column holds the intercept and slope, and the Pr(>|t|) column holds the p-values.

6.1 The intercept

The intercept is 72.3. This is the predicted number of points when a team has zero assists. It is where the line crosses the vertical axis. In this case it is not very meaningful by itself, because teams almost never record zero assists. It simply anchors the line.

6.2 The slope

The slope for assists is 1.56. This is the key result:

For every one extra assist, the model predicts a team scores about 1.56 more points.

So if a team recorded 10 more assists than another, we would predict about 15.6 more points.

6.3 The p-value

The p-value tests whether the relationship could just be down to chance. Here the p-value is <2e-16. As this value is below 0.05 it means that the relationship between assists and total team points is statistically significant.

6.4 R-squared

R-squared is 0.382. It tells us the proportion of the variation in points that is explained by assists. That means assists explain about 38.2% of the differences in team scores between games. The remaining 61.8% is due to everything else (shooting percentage, opponent or free throws).

7 Add the line of best fit

We can draw the model’s line on top of our scatterplot:

ggplot(games, aes(x = assists, y = team_score)) +
  geom_point(alpha = 0.3, colour = "blue") +
  geom_smooth(method = "lm", se = TRUE, colour = "red") +
  labs(title = "Line of best fit: points vs assists",
       x = "Assists in a game",
       y = "Points scored in a game")

The red line is the model’s prediction. The grey band around it shows the uncertainty in where the true line sits (the confidence interval).

8 Check the model’s assumptions

Linear regression works best when some conditions hold. We check them with four diagnostic plots:

par(mfrow = c(2, 2))
plot(model)

par(mfrow = c(1, 1))

What the plots show: the residuals are scattered evenly around zero with a flat trend line (top left), the points follow the diagonal in the Q-Q plot apart from a slightly heavy upper end (top right), and the spread of the errors stays similar for low and high predictions, with no funnel shape (bottom left). Almost every game has low leverage, and no game is influential enough to change the model much (bottom right). The exceptions are rows 997 and 998. Row 998 is an extreme game that stands out in every plot, and row 997 has unusually high leverage but sits close to the line. Overall the assumptions hold well and a straight line is a reasonable model for this data.

A residual is the gap between what the model predicted and what actually happened. Small, random residuals mean a good fit.

9 Make predictions using the dataset

Let’s now use our simple model to make predictions on the dataset. We will predict the score of every game in the dataset from its assists, then compare each prediction with what actually happened.

9.1 Add predictions to the data

The predict() function takes the model and the data and returns one predicted score per row. We store these in a new column called predicted_score, and work out the residual (actual score minus predicted score):

games <- games |> 
  mutate(predicted_score = round(predict(model, newdata = games), 0),
         residual        = team_score - predicted_score)

games |>
  select(team_display_name, assists, team_score, predicted_score, residual) |>
  head(10)
## # A tibble: 10 × 5
##    team_display_name assists team_score predicted_score residual
##    <chr>               <int>      <int>           <dbl>    <dbl>
##  1 Dallas Mavericks       18         88             100      -12
##  2 Boston Celtics         25        106             111       -5
##  3 Boston Celtics         18         84             100      -16
##  4 Dallas Mavericks       21        122             105       17
##  5 Boston Celtics         26        106             113       -7
##  6 Dallas Mavericks       15         99              96        3
##  7 Dallas Mavericks       21         98             105       -7
##  8 Boston Celtics         29        105             118      -13
##  9 Dallas Mavericks        9         89              86        3
## 10 Boston Celtics         23        107             108       -1

For each game, team_score is what the team really scored and predicted_score is what our model expected given the team’s assists. A positive residual means the team scored more than predicted. A negative residual means they scored less.

9.2 How accurate are the predictions?

Two common measures of prediction error:

  • Mean absolute error (MAE): the average size of the miss, ignoring direction.
  • Root mean squared error (RMSE): similar, but punishes big misses more.
mae  <- mean(abs(games$residual))
rmse <- sqrt(mean(games$residual^2))

c(MAE = mae, RMSE = rmse)
##       MAE      RMSE 
##  8.173106 10.352444

On average, the model’s prediction is off by about 8.2 points. Whether that is good enough depends on the purpose. In basketball, a typical miss of 8.2 points is sizeable.

This matches the R-squared we saw earlier: assists only explain 38.2% of the differences in team scores between games, and the remaining 61.8% is due to everything else.

10 Conclusion

Using every team-game from the 2023-24 season, the simple linear regression model estimates that each extra assist is associated with about 1.56 more points. This gives a direct answer to the question we started with: how valuable is ball movement to a team’s offence? Assists explain about 38.2% of the differences in team scores, and the model’s predictions are off by about 8.2 points on average, so assists are a useful predictor of scoring but far from the whole story.

The diagnostic plots suggested that a straight line is a reasonable fit, apart from a couple of extreme games. Remember that this shows an association, not cause and effect: an assist only happens when a teammate scores. A natural next step to improve model performance would be to add more predictors.