A question that can be asked in basketball is, how valuable is ball movement to a team’s offence? An assist is a way to measure ball movement in basketball as it is a pass that directly leads to a teammate scoring. Teams that share the ball well tend to record more assists. A simple way to explore the relationship between assists and a team’s overall offence is by creating a simple linear regression model.
library(tidyverse)
library(hoopR)
The hoopR function load_nba_team_box()
downloads team statistics for every game in a season.
nba <- load_nba_team_box(seasons = 2024)
Each row of this data is one team in one game. Let’s see which columns are available:
names(nba)
## [1] "game_id" "season"
## [3] "season_type" "game_date"
## [5] "game_date_time" "team_id"
## [7] "team_uid" "team_slug"
## [9] "team_location" "team_name"
## [11] "team_abbreviation" "team_display_name"
## [13] "team_short_display_name" "team_color"
## [15] "team_alternate_color" "team_logo"
## [17] "team_home_away" "team_score"
## [19] "team_winner" "assists"
## [21] "blocks" "defensive_rebounds"
## [23] "fast_break_points" "field_goal_pct"
## [25] "field_goals_made" "field_goals_attempted"
## [27] "flagrant_fouls" "fouls"
## [29] "free_throw_pct" "free_throws_made"
## [31] "free_throws_attempted" "largest_lead"
## [33] "offensive_rebounds" "points_in_paint"
## [35] "steals" "team_turnovers"
## [37] "technical_fouls" "three_point_field_goal_pct"
## [39] "three_point_field_goals_made" "three_point_field_goals_attempted"
## [41] "total_rebounds" "total_technical_fouls"
## [43] "total_turnovers" "turnover_points"
## [45] "turnovers" "opponent_team_id"
## [47] "opponent_team_uid" "opponent_team_slug"
## [49] "opponent_team_location" "opponent_team_name"
## [51] "opponent_team_abbreviation" "opponent_team_display_name"
## [53] "opponent_team_short_display_name" "opponent_team_color"
## [55] "opponent_team_alternate_color" "opponent_team_logo"
## [57] "opponent_team_score"
We only need a few columns, so we keep those and remove any rows with missing values.
games <- nba |>
select(game_id, team_display_name, team_score, assists) |>
drop_na()
head(games)
## # A tibble: 6 × 4
## game_id team_display_name team_score assists
## <int> <chr> <int> <int>
## 1 401656363 Dallas Mavericks 88 18
## 2 401656363 Boston Celtics 106 25
## 3 401656362 Boston Celtics 84 18
## 4 401656362 Dallas Mavericks 122 21
## 5 401656361 Boston Celtics 106 26
## 6 401656361 Dallas Mavericks 99 15
Every row now shows one team’s total points and assists in a single game.
A quick scatterplot to visualise the relationship between team points and assists.
ggplot(games, aes(x = assists, y = team_score)) +
geom_point(alpha = 0.3, colour = "blue") +
labs(title = "NBA team points vs assists (2023-24 season)",
x = "Assists in a game",
y = "Points scored in a game")
The dots drift from the bottom left to the top right. Games with more assists tend to be games with more points, showing a positive relationship. We can also get the correlation coefficient of this relationship.
cor(games$assists, games$team_score)
## [1] 0.6177835
A value of 0.62 tells us the relationship is positive. The closer to 1, the stronger the straight-line relationship.
As the two variables share a linear relationship we can now fit a
linear regression model using the lm() function.
model <- lm(team_score ~ assists, data = games)
Now let’s look at the results:
summary(model)
##
## Call:
## lm(formula = team_score ~ assists, data = games)
##
## Residuals:
## Min 1Q Median 3Q Max
## -30.473 -7.036 -0.417 6.583 66.749
##
## Coefficients:
## Estimate Std. Error t value Pr(>|t|)
## (Intercept) 72.32939 1.04515 69.20 <2e-16 ***
## assists 1.56351 0.03875 40.35 <2e-16 ***
## ---
## Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
##
## Residual standard error: 10.35 on 2638 degrees of freedom
## Multiple R-squared: 0.3817, Adjusted R-squared: 0.3814
## F-statistic: 1628 on 1 and 2638 DF, p-value: < 2.2e-16
The summary() output contains a table of coefficients.
Here it is on its own:
coefs <- summary(model)$coefficients
coefs
## Estimate Std. Error t value Pr(>|t|)
## (Intercept) 72.329391 1.04514544 69.20510 0.00000e+00
## assists 1.563512 0.03874739 40.35141 1.07729e-277
The Estimate column holds the intercept and slope, and the Pr(>|t|) column holds the p-values.
The intercept is 72.3. This is the predicted number of points when a team has zero assists. It is where the line crosses the vertical axis. In this case it is not very meaningful by itself, because teams almost never record zero assists. It simply anchors the line.
The slope for assists is 1.56. This is the key result:
For every one extra assist, the model predicts a team scores about 1.56 more points.
So if a team recorded 10 more assists than another, we would predict about 15.6 more points.
The p-value tests whether the relationship could just be down to chance. Here the p-value is <2e-16. As this value is below 0.05 it means that the relationship between assists and total team points is statistically significant.
R-squared is 0.382. It tells us the proportion of the variation in points that is explained by assists. That means assists explain about 38.2% of the differences in team scores between games. The remaining 61.8% is due to everything else (shooting percentage, opponent or free throws).
We can draw the model’s line on top of our scatterplot:
ggplot(games, aes(x = assists, y = team_score)) +
geom_point(alpha = 0.3, colour = "blue") +
geom_smooth(method = "lm", se = TRUE, colour = "red") +
labs(title = "Line of best fit: points vs assists",
x = "Assists in a game",
y = "Points scored in a game")
The red line is the model’s prediction. The grey band around it shows the uncertainty in where the true line sits (the confidence interval).
Linear regression works best when some conditions hold. We check them with four diagnostic plots:
par(mfrow = c(2, 2))
plot(model)
par(mfrow = c(1, 1))
What the plots show: the residuals are scattered evenly around zero with a flat trend line (top left), the points follow the diagonal in the Q-Q plot apart from a slightly heavy upper end (top right), and the spread of the errors stays similar for low and high predictions, with no funnel shape (bottom left). Almost every game has low leverage, and no game is influential enough to change the model much (bottom right). The exceptions are rows 997 and 998. Row 998 is an extreme game that stands out in every plot, and row 997 has unusually high leverage but sits close to the line. Overall the assumptions hold well and a straight line is a reasonable model for this data.
A residual is the gap between what the model predicted and what actually happened. Small, random residuals mean a good fit.
Let’s now use our simple model to make predictions on the dataset. We will predict the score of every game in the dataset from its assists, then compare each prediction with what actually happened.
The predict() function takes the model and the data and
returns one predicted score per row. We store these in a new column
called predicted_score, and work out the residual (actual
score minus predicted score):
games <- games |>
mutate(predicted_score = round(predict(model, newdata = games), 0),
residual = team_score - predicted_score)
games |>
select(team_display_name, assists, team_score, predicted_score, residual) |>
head(10)
## # A tibble: 10 × 5
## team_display_name assists team_score predicted_score residual
## <chr> <int> <int> <dbl> <dbl>
## 1 Dallas Mavericks 18 88 100 -12
## 2 Boston Celtics 25 106 111 -5
## 3 Boston Celtics 18 84 100 -16
## 4 Dallas Mavericks 21 122 105 17
## 5 Boston Celtics 26 106 113 -7
## 6 Dallas Mavericks 15 99 96 3
## 7 Dallas Mavericks 21 98 105 -7
## 8 Boston Celtics 29 105 118 -13
## 9 Dallas Mavericks 9 89 86 3
## 10 Boston Celtics 23 107 108 -1
For each game, team_score is what the team really scored
and predicted_score is what our model expected given the
team’s assists. A positive residual means the team scored
more than predicted. A negative residual means they
scored less.
Two common measures of prediction error:
mae <- mean(abs(games$residual))
rmse <- sqrt(mean(games$residual^2))
c(MAE = mae, RMSE = rmse)
## MAE RMSE
## 8.173106 10.352444
On average, the model’s prediction is off by about 8.2 points. Whether that is good enough depends on the purpose. In basketball, a typical miss of 8.2 points is sizeable.
This matches the R-squared we saw earlier: assists only explain 38.2% of the differences in team scores between games, and the remaining 61.8% is due to everything else.
Using every team-game from the 2023-24 season, the simple linear regression model estimates that each extra assist is associated with about 1.56 more points. This gives a direct answer to the question we started with: how valuable is ball movement to a team’s offence? Assists explain about 38.2% of the differences in team scores, and the model’s predictions are off by about 8.2 points on average, so assists are a useful predictor of scoring but far from the whole story.
The diagnostic plots suggested that a straight line is a reasonable fit, apart from a couple of extreme games. Remember that this shows an association, not cause and effect: an assist only happens when a teammate scores. A natural next step to improve model performance would be to add more predictors.