1 Introduction

Logistic regression is a method used to predict the probability of a categorical outcome, usually a binary choice like yes or no, or 0 and 1. This method can be easily applied in a sporting context when predicting certain outcomes.

Over the last decade the three-point shot has become a much bigger part of NBA offences. But taking threes only helps if they go in. We will use logistic regression to investigate if teams that shoot a higher percentage from three have a better chance of making the playoffs.

2 Load the Packages

library(tidyverse)
library(hoopR)

3 Get the data

3.1 Download the games

The hoopR function load_nba_team_box() downloads team statistics for every game in our chosen time frame.

nba <- load_nba_team_box(seasons = 2021:2026)

Quick check of column names

names(nba)
##  [1] "game_id"                           "season"                           
##  [3] "season_type"                       "game_date"                        
##  [5] "game_date_time"                    "team_id"                          
##  [7] "team_uid"                          "team_slug"                        
##  [9] "team_location"                     "team_name"                        
## [11] "team_abbreviation"                 "team_display_name"                
## [13] "team_short_display_name"           "team_color"                       
## [15] "team_alternate_color"              "team_logo"                        
## [17] "team_home_away"                    "team_score"                       
## [19] "team_winner"                       "assists"                          
## [21] "blocks"                            "defensive_rebounds"               
## [23] "fast_break_points"                 "field_goal_pct"                   
## [25] "field_goals_made"                  "field_goals_attempted"            
## [27] "flagrant_fouls"                    "fouls"                            
## [29] "free_throw_pct"                    "free_throws_made"                 
## [31] "free_throws_attempted"             "largest_lead"                     
## [33] "offensive_rebounds"                "points_in_paint"                  
## [35] "steals"                            "team_turnovers"                   
## [37] "technical_fouls"                   "three_point_field_goal_pct"       
## [39] "three_point_field_goals_made"      "three_point_field_goals_attempted"
## [41] "total_rebounds"                    "total_technical_fouls"            
## [43] "total_turnovers"                   "turnover_points"                  
## [45] "turnovers"                         "opponent_team_id"                 
## [47] "opponent_team_uid"                 "opponent_team_slug"               
## [49] "opponent_team_location"            "opponent_team_name"               
## [51] "opponent_team_abbreviation"        "opponent_team_display_name"       
## [53] "opponent_team_short_display_name"  "opponent_team_color"              
## [55] "opponent_team_alternate_color"     "opponent_team_logo"               
## [57] "opponent_team_score"               "lead_changes"                     
## [59] "lead_percentage"

We will use three_point_field_goals_made and three_point_field_goals_attempted, Adding these up over a whole season lets us work out a team’s season three-point percentage. The season_type column tells us what kind of game it was:

table(nba$season_type)
## 
##     2     3     5 
## 14488  1014    60

In this data, 2 means a regular season game and 3 means a playoff game. We will use this to work out which teams made the playoffs.

3.2 Keep the columns we need

games <- nba |> 
  select(season, season_type, game_id, team_display_name,
         three_point_field_goals_made,
         three_point_field_goals_attempted) |> 
  drop_na()

head(games)
## # A tibble: 6 × 6
##   season season_type   game_id team_display_name three_point_field_goals_made
##    <int>       <int>     <int> <chr>                                    <int>
## 1   2021           3 401344140 Phoenix Suns                                 6
## 2   2021           3 401344140 Milwaukee Bucks                              6
## 3   2021           3 401344139 Milwaukee Bucks                             14
## 4   2021           3 401344139 Phoenix Suns                                13
## 5   2021           3 401344138 Phoenix Suns                                 7
## 6   2021           3 401344138 Milwaukee Bucks                              7
## # ℹ 1 more variable: three_point_field_goals_attempted <int>

3.3 Which teams made the playoffs?

A team made the playoffs if it played in playoff games. We require at least 4 playoff games, because the shortest possible series is four games. This makes sure teams that only played in the play-in tournament and lost are not counted.

playoff_teams <- games |> 
  filter(season_type == 3) |> 
  count(season, team_display_name, name = "playoff_games") |> 
  filter(playoff_games >= 4) |> 
  mutate(made_playoffs = 1)

3.4 Build one row per team per season

Now we take the regular season games only, and for each team in each season we work out its season three-point percentage. We keep teams with at least 50 games, which removes one-off special games that can appear in this data as a “team” (such as the All-Star Game). Finally we join on the playoff information. Teams that are not in playoff_teams get a 0.

team_seasons <- games |> 
  filter(season_type == 2) |> 
  group_by(season, team_display_name) |> 
  summarise(games_played = n(),
            threes_made = sum(three_point_field_goals_made),
            threes_attempted = sum(three_point_field_goals_attempted),
            three_pct = 100 * threes_made / threes_attempted,
            .groups = "drop") |> 
  filter(games_played >= 50) |> 
  left_join(playoff_teams, by = c("season", "team_display_name")) |> 
  mutate(made_playoffs = replace_na(made_playoffs, 0)) |> 
  select(-playoff_games, -threes_made, -threes_attempted)

head(team_seasons)
## # A tibble: 6 × 5
##   season team_display_name   games_played three_pct made_playoffs
##    <int> <chr>                      <int>     <dbl>         <dbl>
## 1   2021 Atlanta Hawks                 72      37.3             1
## 2   2021 Boston Celtics                72      37.4             1
## 3   2021 Brooklyn Nets                 72      39.2             1
## 4   2021 Charlotte Hornets             72      36.9             0
## 5   2021 Chicago Bulls                 72      37.0             0
## 6   2021 Cleveland Cavaliers           72      33.6             0

3.5 Check the percentage

Because we multiplied by 100, three_pct is on a 0 to 100 scale, so a change of 1 unit means a change of 1 percentage point. A quick check of the values:

summary(team_seasons$three_pct)
##    Min. 1st Qu.  Median    Mean 3rd Qu.    Max. 
##   31.85   34.90   36.13   36.07   37.07   41.13

What the columns mean:

  • three_pct is the season three-point percentage (total made divided by total attempted), our predictor
  • made_playoffs is 1 if the team made the playoffs and 0 if it did not.

Each row is one team-season: one team in one season.

3.6 Check the data

A quick sanity check of the work. Each season should have 30 teams, and 16 of them should have made the playoffs:

team_seasons |> 
  group_by(season) |> 
  summarise(teams = n(), playoff_teams = sum(made_playoffs))
## # A tibble: 6 × 3
##   season teams playoff_teams
##    <int> <int>         <dbl>
## 1   2021    30            16
## 2   2022    30            16
## 3   2023    30            16
## 4   2024    30            16
## 5   2025    30            16
## 6   2026    30            16

We have 180 team-seasons to work with, and 53.3% of them made the playoffs.

4 Explore the data

Compare the three-point percentage of teams that made the playoffs with those that did not:

ggplot(team_seasons,
       aes(x = factor(made_playoffs, labels = c("No", "Yes")),
           y = three_pct)) +
  geom_boxplot(fill = "lightblue") +
  labs(title = "Three-point percentage: playoff teams vs the rest",
       x = "Made the playoffs?",
       y = "Season three-point %")

This boxplot allows us to visually compare the three-point percentages of teams who made the playoffs and those who didn’t.

5 Fit the logistic regression model

Now we fit a logistic regression using the glm() function.

model <- glm(made_playoffs ~ three_pct, data = team_seasons,
             family = binomial)

Here are the model results:

summary(model)
## 
## Call:
## glm(formula = made_playoffs ~ three_pct, family = binomial, data = team_seasons)
## 
## Coefficients:
##             Estimate Std. Error z value Pr(>|z|)    
## (Intercept) -28.6170     5.0154  -5.706 1.16e-08 ***
## three_pct     0.7983     0.1394   5.729 1.01e-08 ***
## ---
## Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
## 
## (Dispersion parameter for binomial family taken to be 1)
## 
##     Null deviance: 248.73  on 179  degrees of freedom
## Residual deviance: 202.23  on 178  degrees of freedom
## AIC: 206.23
## 
## Number of Fisher Scoring iterations: 4

6 Interpret the model output

Here is the table of coefficients on its own:

coefs <- summary(model)$coefficients
coefs
##                Estimate Std. Error   z value     Pr(>|z|)
## (Intercept) -28.6170013  5.0154313 -5.705791 1.158042e-08
## three_pct     0.7982796  0.1393517  5.728522 1.013091e-08

The Estimate column holds the intercept and slope (in log-odds), and the Pr(>|z|) column holds the p-values.

6.1 The intercept

The intercept is -28.62 (in log-odds). It is the model’s prediction for a team with a three-point percentage of 0%, which never happens, so it is not meaningful by itself. It simply anchors the curve.

6.2 The slope

The slope for three-point percentage is 0.798 (in log-odds). Log-odds are hard to picture, so we convert them to an odds ratio by taking the exponential, exp(slope):

Each extra percentage point of three-point shooting multiplies a team’s odds of making the playoffs by about 2.22.

A bigger step makes this easier to picture. If a team improves its three-point percentage by 2 percentage points, its odds are multiplied by about 4.94. A value above 1 means the odds go up, and a value below 1 means they go down.

6.3 The p-value

The p-value tests whether the relationship could just be down to chance. Here it is 1.01e-08. As this value is below 0.05, this relationship is statistically significant.

If the p-value is above 0.05, it does not prove there is no relationship. It means we do not have strong enough evidence for one, which is more likely when there are few teams.

6.4 How well does the model fit?

Logistic regression does not have an R-squared, but a similar measure called McFadden’s pseudo R-squared compares our model with one that uses no predictors. Ours is 0.187. Values closer to 1 mean a better fit, and a value close to 0 means the predictor adds very little. Values between 0.2 and 0.4 are usually considered a strong fit.

7 Visualise the model

The S-shaped curve shows how the predicted probability changes as three-point percentage increases. Each dot is one team-season.

ggplot(team_seasons, aes(x = three_pct, y = made_playoffs)) +
  geom_jitter(height = 0.04, alpha = 0.5, colour = "blue") +
  geom_smooth(method = "glm", method.args = list(family = "binomial"),
              se = TRUE, colour = "red") +
  labs(title = "Chance of making the playoffs by three-point percentage",
       x = "Season three-point %",
       y = "Probability of making the playoffs")

The red curve is the model’s predicted probability, and the grey band shows the uncertainty around it. A curve that climbs from left to right means teams with a higher three-point percentage have a higher predicted chance. A wide grey band means we are not very sure about the curve.

8 Make predictions

8.1 Predictions for every team

We can use the model to predict the probability for every team-season in our data and add it as a new column:

team_seasons <- team_seasons |>
  mutate(predicted_prob = predict(model, type = "response"))

team_seasons |> 
  arrange(desc(predicted_prob)) |> 
  select(season, team_display_name, three_pct, made_playoffs,
         predicted_prob) |> 
  mutate(three_pct = round(three_pct, 1),
         predicted_prob = round(predicted_prob, 3)) |> 
  head(10)
## # A tibble: 10 × 5
##    season team_display_name     three_pct made_playoffs predicted_prob
##     <int> <chr>                     <dbl>         <dbl>          <dbl>
##  1   2021 LA Clippers                41.1             1          0.985
##  2   2026 Denver Nuggets             39.6             1          0.951
##  3   2021 Brooklyn Nets              39.2             1          0.937
##  4   2021 New York Knicks            39.2             1          0.934
##  5   2021 Utah Jazz                  38.9             1          0.919
##  6   2021 Milwaukee Bucks            38.9             1          0.919
##  7   2024 Oklahoma City Thunder      38.9             1          0.917
##  8   2024 Boston Celtics             38.8             1          0.913
##  9   2025 Milwaukee Bucks            38.7             1          0.908
## 10   2023 Philadelphia 76ers         38.7             1          0.906

These are the teams the model gave the highest chance. The made_playoffs column shows whether they really made it (1) or not (0).

The most interesting cases are the surprises. Here are the playoff teams the model gave the lowest chance:

team_seasons |> 
  filter(made_playoffs == 1) |> 
  arrange(predicted_prob) |> 
  select(season, team_display_name, three_pct, predicted_prob) |> 
  mutate(three_pct = round(three_pct, 1),
         predicted_prob = round(predicted_prob, 3)) |> 
  head(5)
## # A tibble: 5 × 4
##   season team_display_name      three_pct predicted_prob
##    <int> <chr>                      <dbl>          <dbl>
## 1   2025 Orlando Magic               31.8          0.039
## 2   2022 New Orleans Pelicans        33.2          0.108
## 3   2026 Portland Trail Blazers      34.3          0.219
## 4   2026 Orlando Magic               34.3          0.23 
## 5   2023 Miami Heat                  34.4          0.234

And here are the teams with the highest predicted chance that did not make the playoffs:

team_seasons |> 
  filter(made_playoffs == 0) |> 
  arrange(desc(predicted_prob)) |> 
  select(season, team_display_name, three_pct, predicted_prob) |> 
  mutate(three_pct = round(three_pct, 1),
         predicted_prob = round(predicted_prob, 3)) |> 
  head(5)
## # A tibble: 5 × 4
##   season team_display_name     three_pct predicted_prob
##    <int> <chr>                     <dbl>          <dbl>
## 1   2026 Milwaukee Bucks            38.7          0.906
## 2   2024 Golden State Warriors      38            0.843
## 3   2026 Charlotte Hornets          37.8          0.831
## 4   2025 Phoenix Suns               37.8          0.825
## 5   2021 Golden State Warriors      37.6          0.797

Three-point percentage does not tell the whole story. A team can shoot threes well and still miss out, for example because the team simply shoots less threes, because they’re poor defensively or because playoff spots are decided within each conference.

9 How accurate is the model?

To turn probabilities into yes/no predictions, we pick a cut-off. Here we predict “makes the playoffs” whenever the probability is 0.5 or higher. A confusion matrix counts how often the predictions were right and wrong:

team_seasons <- team_seasons |> 
  mutate(predicted_playoffs = if_else(predicted_prob >= 0.5, 1, 0))

table(Actual = team_seasons$made_playoffs,
      Predicted = team_seasons$predicted_playoffs)
##       Predicted
## Actual  0  1
##      0 59 25
##      1 24 72

The rows are what really happened and the columns are what the model predicted. The numbers on the diagonal (top left and bottom right) are correct predictions.

  • Accuracy: the model is correct for 72.8% of team-seasons.
  • Comparison: a model that always predicted the more common outcome would be right 53.3% of the time. This is the score our model needs to beat, and the bigger the gap, the more useful three-point percentage is as a predictor.
  • The model correctly identified 75% of the teams that really made the playoffs.
  • The model correctly identified 70.2% of the teams that really missed the playoffs.

10 Conclusion

Using NBA team data from the 2020-21 to 2025-26 seasons, we fitted a logistic regression to see whether teams with a higher three-point percentage are more likely to make the playoffs. Each extra percentage point of three-point shooting multiplies a team’s odds of making the playoffs by about 2.22, and the relationship is statistically significant. The model classifies 72.8% of team-seasons correctly, compared with 53.3% for a model that always guesses the most common outcome.