Logistic regression is a method used to predict the probability of a categorical outcome, usually a binary choice like yes or no, or 0 and 1. This method can be easily applied in a sporting context when predicting certain outcomes.
Over the last decade the three-point shot has become a much bigger part of NBA offences. But taking threes only helps if they go in. We will use logistic regression to investigate if teams that shoot a higher percentage from three have a better chance of making the playoffs.
library(tidyverse)
library(hoopR)
The hoopR function load_nba_team_box()
downloads team statistics for every game in our chosen time frame.
nba <- load_nba_team_box(seasons = 2021:2026)
Quick check of column names
names(nba)
## [1] "game_id" "season"
## [3] "season_type" "game_date"
## [5] "game_date_time" "team_id"
## [7] "team_uid" "team_slug"
## [9] "team_location" "team_name"
## [11] "team_abbreviation" "team_display_name"
## [13] "team_short_display_name" "team_color"
## [15] "team_alternate_color" "team_logo"
## [17] "team_home_away" "team_score"
## [19] "team_winner" "assists"
## [21] "blocks" "defensive_rebounds"
## [23] "fast_break_points" "field_goal_pct"
## [25] "field_goals_made" "field_goals_attempted"
## [27] "flagrant_fouls" "fouls"
## [29] "free_throw_pct" "free_throws_made"
## [31] "free_throws_attempted" "largest_lead"
## [33] "offensive_rebounds" "points_in_paint"
## [35] "steals" "team_turnovers"
## [37] "technical_fouls" "three_point_field_goal_pct"
## [39] "three_point_field_goals_made" "three_point_field_goals_attempted"
## [41] "total_rebounds" "total_technical_fouls"
## [43] "total_turnovers" "turnover_points"
## [45] "turnovers" "opponent_team_id"
## [47] "opponent_team_uid" "opponent_team_slug"
## [49] "opponent_team_location" "opponent_team_name"
## [51] "opponent_team_abbreviation" "opponent_team_display_name"
## [53] "opponent_team_short_display_name" "opponent_team_color"
## [55] "opponent_team_alternate_color" "opponent_team_logo"
## [57] "opponent_team_score" "lead_changes"
## [59] "lead_percentage"
We will use three_point_field_goals_made and
three_point_field_goals_attempted, Adding these up over a
whole season lets us work out a team’s season three-point percentage.
The season_type column tells us what kind of game it
was:
table(nba$season_type)
##
## 2 3 5
## 14488 1014 60
In this data, 2 means a regular season game and
3 means a playoff game. We will use this to work out which
teams made the playoffs.
games <- nba |>
select(season, season_type, game_id, team_display_name,
three_point_field_goals_made,
three_point_field_goals_attempted) |>
drop_na()
head(games)
## # A tibble: 6 × 6
## season season_type game_id team_display_name three_point_field_goals_made
## <int> <int> <int> <chr> <int>
## 1 2021 3 401344140 Phoenix Suns 6
## 2 2021 3 401344140 Milwaukee Bucks 6
## 3 2021 3 401344139 Milwaukee Bucks 14
## 4 2021 3 401344139 Phoenix Suns 13
## 5 2021 3 401344138 Phoenix Suns 7
## 6 2021 3 401344138 Milwaukee Bucks 7
## # ℹ 1 more variable: three_point_field_goals_attempted <int>
A team made the playoffs if it played in playoff games. We require at least 4 playoff games, because the shortest possible series is four games. This makes sure teams that only played in the play-in tournament and lost are not counted.
playoff_teams <- games |>
filter(season_type == 3) |>
count(season, team_display_name, name = "playoff_games") |>
filter(playoff_games >= 4) |>
mutate(made_playoffs = 1)
Now we take the regular season games only, and for each team in each
season we work out its season three-point percentage.
We keep teams with at least 50 games, which removes one-off special
games that can appear in this data as a “team” (such as the All-Star
Game). Finally we join on the playoff information. Teams that are not in
playoff_teams get a 0.
team_seasons <- games |>
filter(season_type == 2) |>
group_by(season, team_display_name) |>
summarise(games_played = n(),
threes_made = sum(three_point_field_goals_made),
threes_attempted = sum(three_point_field_goals_attempted),
three_pct = 100 * threes_made / threes_attempted,
.groups = "drop") |>
filter(games_played >= 50) |>
left_join(playoff_teams, by = c("season", "team_display_name")) |>
mutate(made_playoffs = replace_na(made_playoffs, 0)) |>
select(-playoff_games, -threes_made, -threes_attempted)
head(team_seasons)
## # A tibble: 6 × 5
## season team_display_name games_played three_pct made_playoffs
## <int> <chr> <int> <dbl> <dbl>
## 1 2021 Atlanta Hawks 72 37.3 1
## 2 2021 Boston Celtics 72 37.4 1
## 3 2021 Brooklyn Nets 72 39.2 1
## 4 2021 Charlotte Hornets 72 36.9 0
## 5 2021 Chicago Bulls 72 37.0 0
## 6 2021 Cleveland Cavaliers 72 33.6 0
Because we multiplied by 100, three_pct is on a 0 to 100
scale, so a change of 1 unit means a change of 1 percentage
point. A quick check of the values:
summary(team_seasons$three_pct)
## Min. 1st Qu. Median Mean 3rd Qu. Max.
## 31.85 34.90 36.13 36.07 37.07 41.13
What the columns mean:
three_pct is the season three-point percentage (total
made divided by total attempted), our predictormade_playoffs is 1 if the team made the
playoffs and 0 if it did not.Each row is one team-season: one team in one season.
A quick sanity check of the work. Each season should have 30 teams, and 16 of them should have made the playoffs:
team_seasons |>
group_by(season) |>
summarise(teams = n(), playoff_teams = sum(made_playoffs))
## # A tibble: 6 × 3
## season teams playoff_teams
## <int> <int> <dbl>
## 1 2021 30 16
## 2 2022 30 16
## 3 2023 30 16
## 4 2024 30 16
## 5 2025 30 16
## 6 2026 30 16
We have 180 team-seasons to work with, and 53.3% of them made the playoffs.
Compare the three-point percentage of teams that made the playoffs with those that did not:
ggplot(team_seasons,
aes(x = factor(made_playoffs, labels = c("No", "Yes")),
y = three_pct)) +
geom_boxplot(fill = "lightblue") +
labs(title = "Three-point percentage: playoff teams vs the rest",
x = "Made the playoffs?",
y = "Season three-point %")
This boxplot allows us to visually compare the three-point percentages of teams who made the playoffs and those who didn’t.
Now we fit a logistic regression using the glm()
function.
model <- glm(made_playoffs ~ three_pct, data = team_seasons,
family = binomial)
Here are the model results:
summary(model)
##
## Call:
## glm(formula = made_playoffs ~ three_pct, family = binomial, data = team_seasons)
##
## Coefficients:
## Estimate Std. Error z value Pr(>|z|)
## (Intercept) -28.6170 5.0154 -5.706 1.16e-08 ***
## three_pct 0.7983 0.1394 5.729 1.01e-08 ***
## ---
## Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
##
## (Dispersion parameter for binomial family taken to be 1)
##
## Null deviance: 248.73 on 179 degrees of freedom
## Residual deviance: 202.23 on 178 degrees of freedom
## AIC: 206.23
##
## Number of Fisher Scoring iterations: 4
Here is the table of coefficients on its own:
coefs <- summary(model)$coefficients
coefs
## Estimate Std. Error z value Pr(>|z|)
## (Intercept) -28.6170013 5.0154313 -5.705791 1.158042e-08
## three_pct 0.7982796 0.1393517 5.728522 1.013091e-08
The Estimate column holds the intercept and slope (in log-odds), and the Pr(>|z|) column holds the p-values.
The intercept is -28.62 (in log-odds). It is the model’s prediction for a team with a three-point percentage of 0%, which never happens, so it is not meaningful by itself. It simply anchors the curve.
The slope for three-point percentage is 0.798 (in
log-odds). Log-odds are hard to picture, so we convert them to an
odds ratio by taking the exponential,
exp(slope):
Each extra percentage point of three-point shooting multiplies a team’s odds of making the playoffs by about 2.22.
A bigger step makes this easier to picture. If a team improves its three-point percentage by 2 percentage points, its odds are multiplied by about 4.94. A value above 1 means the odds go up, and a value below 1 means they go down.
The p-value tests whether the relationship could just be down to chance. Here it is 1.01e-08. As this value is below 0.05, this relationship is statistically significant.
If the p-value is above 0.05, it does not prove there is no relationship. It means we do not have strong enough evidence for one, which is more likely when there are few teams.
Logistic regression does not have an R-squared, but a similar measure called McFadden’s pseudo R-squared compares our model with one that uses no predictors. Ours is 0.187. Values closer to 1 mean a better fit, and a value close to 0 means the predictor adds very little. Values between 0.2 and 0.4 are usually considered a strong fit.
The S-shaped curve shows how the predicted probability changes as three-point percentage increases. Each dot is one team-season.
ggplot(team_seasons, aes(x = three_pct, y = made_playoffs)) +
geom_jitter(height = 0.04, alpha = 0.5, colour = "blue") +
geom_smooth(method = "glm", method.args = list(family = "binomial"),
se = TRUE, colour = "red") +
labs(title = "Chance of making the playoffs by three-point percentage",
x = "Season three-point %",
y = "Probability of making the playoffs")
The red curve is the model’s predicted probability, and the grey band shows the uncertainty around it. A curve that climbs from left to right means teams with a higher three-point percentage have a higher predicted chance. A wide grey band means we are not very sure about the curve.
We can use the model to predict the probability for every team-season in our data and add it as a new column:
team_seasons <- team_seasons |>
mutate(predicted_prob = predict(model, type = "response"))
team_seasons |>
arrange(desc(predicted_prob)) |>
select(season, team_display_name, three_pct, made_playoffs,
predicted_prob) |>
mutate(three_pct = round(three_pct, 1),
predicted_prob = round(predicted_prob, 3)) |>
head(10)
## # A tibble: 10 × 5
## season team_display_name three_pct made_playoffs predicted_prob
## <int> <chr> <dbl> <dbl> <dbl>
## 1 2021 LA Clippers 41.1 1 0.985
## 2 2026 Denver Nuggets 39.6 1 0.951
## 3 2021 Brooklyn Nets 39.2 1 0.937
## 4 2021 New York Knicks 39.2 1 0.934
## 5 2021 Utah Jazz 38.9 1 0.919
## 6 2021 Milwaukee Bucks 38.9 1 0.919
## 7 2024 Oklahoma City Thunder 38.9 1 0.917
## 8 2024 Boston Celtics 38.8 1 0.913
## 9 2025 Milwaukee Bucks 38.7 1 0.908
## 10 2023 Philadelphia 76ers 38.7 1 0.906
These are the teams the model gave the highest chance. The
made_playoffs column shows whether they really made it (1)
or not (0).
The most interesting cases are the surprises. Here are the playoff teams the model gave the lowest chance:
team_seasons |>
filter(made_playoffs == 1) |>
arrange(predicted_prob) |>
select(season, team_display_name, three_pct, predicted_prob) |>
mutate(three_pct = round(three_pct, 1),
predicted_prob = round(predicted_prob, 3)) |>
head(5)
## # A tibble: 5 × 4
## season team_display_name three_pct predicted_prob
## <int> <chr> <dbl> <dbl>
## 1 2025 Orlando Magic 31.8 0.039
## 2 2022 New Orleans Pelicans 33.2 0.108
## 3 2026 Portland Trail Blazers 34.3 0.219
## 4 2026 Orlando Magic 34.3 0.23
## 5 2023 Miami Heat 34.4 0.234
And here are the teams with the highest predicted chance that did not make the playoffs:
team_seasons |>
filter(made_playoffs == 0) |>
arrange(desc(predicted_prob)) |>
select(season, team_display_name, three_pct, predicted_prob) |>
mutate(three_pct = round(three_pct, 1),
predicted_prob = round(predicted_prob, 3)) |>
head(5)
## # A tibble: 5 × 4
## season team_display_name three_pct predicted_prob
## <int> <chr> <dbl> <dbl>
## 1 2026 Milwaukee Bucks 38.7 0.906
## 2 2024 Golden State Warriors 38 0.843
## 3 2026 Charlotte Hornets 37.8 0.831
## 4 2025 Phoenix Suns 37.8 0.825
## 5 2021 Golden State Warriors 37.6 0.797
Three-point percentage does not tell the whole story. A team can shoot threes well and still miss out, for example because the team simply shoots less threes, because they’re poor defensively or because playoff spots are decided within each conference.
To turn probabilities into yes/no predictions, we pick a cut-off. Here we predict “makes the playoffs” whenever the probability is 0.5 or higher. A confusion matrix counts how often the predictions were right and wrong:
team_seasons <- team_seasons |>
mutate(predicted_playoffs = if_else(predicted_prob >= 0.5, 1, 0))
table(Actual = team_seasons$made_playoffs,
Predicted = team_seasons$predicted_playoffs)
## Predicted
## Actual 0 1
## 0 59 25
## 1 24 72
The rows are what really happened and the columns are what the model predicted. The numbers on the diagonal (top left and bottom right) are correct predictions.
Using NBA team data from the 2020-21 to 2025-26 seasons, we fitted a logistic regression to see whether teams with a higher three-point percentage are more likely to make the playoffs. Each extra percentage point of three-point shooting multiplies a team’s odds of making the playoffs by about 2.22, and the relationship is statistically significant. The model classifies 72.8% of team-seasons correctly, compared with 53.3% for a model that always guesses the most common outcome.