1. Introduction

Three-point shooting is an important part of basketball, allowing teams to score three points from a successful shot beyond the three-point line. But are teams with higher three-point shooting accuracy more likely to win their matches?

This tutorial explores the relationship between three-point shooting accuracy and match outcomes using data from the 2022-2026 National Basketball Association (NBA) regular seasons. Binary logistic regression in R is used to examine this relationship through data preparation, visualisation, statistical modelling and interpretation.

The analysis focuses on two key variables:

  • Three-point shooting accuracy: The percentage of attempted three-point shots successfully made by a team in a match.
  • Match outcome: Whether the team won (1) or lost (0) the match.

Unlike simple linear regression, binary logistic regression is suitable when the outcome has two categories. It estimates the probability of winning rather than predicting a numerical match margin.

Because shooting accuracy is measured during a match, this tutorial examines its association with winning, rather than predicting results before a match begins.

2. Loading the Data

The hoopR package is used to retrieve NBA team box-score statistics, while tidyverse is used for data preparation and visualisation, and knitr for presenting tables.

If the packages are not already installed, they can be installed by running the following code once in the RStudio Console:

install.packages(c("hoopR", "tidyverse", "knitr"))
# Load required packages
library(hoopR)
library(tidyverse)
library(knitr)

The hoopR season numbers refer to the year in which each NBA season ends. For example, 2023 represents 2022-23, while 2026 represents 2025-26.

# Retrieve NBA team statistics for the 2022–23 to 2024–25 seasons
nba_stats <- load_nba_team_box(seasons = 2023:2025)

3. Preparing the Data

Each row of the downloaded dataset represents one team’s performance in a match. The dataset was restricted to regular-season matches, and only variables relevant to the analysis were retained.

# Select regular-season matches and relevant variables
nba_data <- nba_stats %>%
  filter(season_type == 2) %>%
  select(game_id, season, team_display_name,
         three_point_field_goal_pct, team_winner) %>%
  drop_na() %>%

  # Convert match outcomes to 1 for wins and 0 for losses
  mutate(win = as.integer(team_winner))

# Check the number of observations in each season
nba_data %>%
  count(season) %>%
  kable(caption = "Table 1. Team-game observations by NBA season")
Table 1. Team-game observations by NBA season
season n
2023 2462
2024 2464
2025 2468

The three_point_field_goal_pct variable is already recorded as a percentage, so a value of 40 represents 40% shooting accuracy. The new win variable is coded as 1 for a win and 0 for a loss.

# Display the first six team-game observations
nba_data %>%
  select(season, team_display_name, three_point_field_goal_pct, win) %>%
  head(6) %>%
  kable(caption = "Table 2. Example NBA team-game observations")
Table 2. Example NBA team-game observations
season team_display_name three_point_field_goal_pct win
2023 Atlanta Hawks 28.2 0
2023 Boston Celtics 46.3 1
2023 Philadelphia 76ers 43.8 1
2023 Brooklyn Nets 37.5 0
2023 Charlotte Hornets 18.8 1
2023 Cleveland Cavaliers 27.0 0

4. Exploring and Visualising the Data

4.1 Summary Statistics

Before fitting the model, the average three-point shooting accuracy of winning and losing teams was compared to identify differences in shooting performance.

# Calculate summary statistics for wins and losses
nba_summary <- nba_data %>%
  group_by(win) %>%
  summarise(
    `Mean accuracy (%)` = round(mean(three_point_field_goal_pct), 1),
    `Median accuracy (%)` = round(median(three_point_field_goal_pct), 1),
    `SD accuracy` = round(sd(three_point_field_goal_pct), 2),
    `Observations` = n(),
    .groups = "drop"
  ) %>%
  mutate(`Match outcome` = if_else(win == 1, "Win", "Loss")) %>%
  select(`Match outcome`, everything(), -win)

# Display the summary statistics in a formatted table
kable(nba_summary, caption = "Table 3. Three-point shooting accuracy by match outcome")
Table 3. Three-point shooting accuracy by match outcome
Match outcome Mean accuracy (%) Median accuracy (%) SD accuracy Observations
Loss 33.2 33.0 7.66 3697
Win 39.0 38.9 7.94 3697

The mean and median describe typical shooting accuracy, while the standard deviation describes variation between matches. Winning teams recorded an average accuracy of approximately 39.0%, compared with 33.2% for losing teams, a difference of 5.8 percentage points.

4.2 Comparing Shooting Accuracy in Wins and Losses

A boxplot compares the distribution of three-point shooting accuracy for winning and losing teams.

# Create a boxplot comparing shooting accuracy by match outcome
ggplot(nba_data,
       aes(x = factor(win, labels = c("Loss", "Win")),
           y = three_point_field_goal_pct,
           fill = factor(win))) +
  geom_boxplot(alpha = 0.85, width = 0.5, outlier.alpha = 0.3) +
  scale_fill_manual(values = c("0" = "orange", "1" = "lightgreen")) +
  labs(title = "Three-Point Shooting Accuracy in Wins and Losses",
       subtitle = "NBA regular seasons, 2022–23 to 2024–25",
       x = "Match outcome",
       y = "Three-point shooting accuracy (%)") +
  theme_minimal(base_size = 13) +
  theme(legend.position = "none")

The line inside each box represents the median. Winning teams generally recorded higher shooting accuracy, although the distributions overlap. This suggests that three-point shooting accuracy is associated with match outcomes but is not the only factor involved.

5. Binary Logistic Regression

The glm() function is used to estimate how the probability of winning changes with three-point shooting accuracy. Setting family = binomial fits a binary logistic regression model.

# Fit a binary logistic regression model using three-point shooting accuracy
model <- glm(win ~ three_point_field_goal_pct,
             data = nba_data,
             family = binomial)

# Display the regression coefficients and p-values
summary(model)
## 
## Call:
## glm(formula = win ~ three_point_field_goal_pct, family = binomial, 
##     data = nba_data)
## 
## Coefficients:
##                             Estimate Std. Error z value Pr(>|z|)    
## (Intercept)                -3.473610   0.124991  -27.79   <2e-16 ***
## three_point_field_goal_pct  0.096233   0.003398   28.32   <2e-16 ***
## ---
## Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
## 
## (Dispersion parameter for binomial family taken to be 1)
## 
##     Null deviance: 10250.3  on 7393  degrees of freedom
## Residual deviance:  9281.3  on 7392  degrees of freedom
## AIC: 9285.3
## 
## Number of Fisher Scoring iterations: 4

A positive coefficient indicates that higher shooting accuracy is associated with a greater probability of winning. The p-value indicates whether there is statistical evidence of an association. Because both teams from a match contribute an observation, these observations are related; the standard model p-value does not account for this dependence.

Logistic regression coefficients are expressed in log-odds. The exp() function converts the coefficient into an odds ratio, which is easier to interpret.

# Extract the coefficient, odds ratio and p-value
coefficient <- coef(model)[["three_point_field_goal_pct"]]
odds_ratio <- exp(coefficient)
p_value <- coef(summary(model))["three_point_field_goal_pct", "Pr(>|z|)"]

# Present the model results in a formatted table
tibble(
  Measure = c("Regression coefficient", "Odds ratio", "p-value"),
  Value = c(as.character(round(coefficient, 3)),
            as.character(round(odds_ratio, 3)),
            if_else(p_value < 0.001, "< 0.001",
                    as.character(signif(p_value, 3))))
) %>%
  kable(caption = "Table 4. Binary logistic regression results")
Table 4. Binary logistic regression results
Measure Value
Regression coefficient 0.096
Odds ratio 1.101
p-value < 0.001

The regression coefficient was 0.096, with an odds ratio of 1.101. This means that each one-percentage-point increase in three-point shooting accuracy was associated with approximately 10.1% higher odds of winning. The relationship was statistically significant (p < 0.001).

5.1 Interpreting Predicted Probabilities

Binary logistic regression estimates the probability of an outcome occurring based on a predictor variable. The fitted model can therefore be used to estimate the probability of winning at different levels of three-point shooting accuracy.

# Select different three-point shooting percentages
shooting_levels <- tibble(
  three_point_field_goal_pct = c(25, 30, 35, 40, 45, 50)
)

# Calculate predicted probabilities using the fitted model
shooting_levels <- shooting_levels %>%
  mutate(win_probability = predict(model, newdata = shooting_levels,
                                   type = "response"))

# Display predicted probabilities as percentages
probability_table <- shooting_levels %>%
  transmute(
    `Three-point shooting accuracy (%)` = three_point_field_goal_pct,
    `Predicted probability of winning (%)` = round(win_probability * 100, 1)
  )

kable(probability_table,
      caption = "Table 5. Predicted probability of winning at different shooting percentages")
Table 5. Predicted probability of winning at different shooting percentages
Three-point shooting accuracy (%) Predicted probability of winning (%)
25 25.6
30 35.7
35 47.4
40 59.3
45 70.2
50 79.2

The predicted probabilities demonstrate how the estimated likelihood of winning increases with three-point shooting accuracy. For example, a team recording 25% shooting accuracy had an estimated 25.6% probability of winning, compared with 70.2% for a team recording 45% accuracy. These estimates are based solely on shooting accuracy and do not account for other factors that may influence match outcomes.

5.2 Visualising Predicted Win Probability

A logistic regression curve shows how the model’s estimated probability of winning changes across different shooting percentages. Unlike a scatterplot, a smooth probability curve makes the overall relationship easier to see.

# Create shooting accuracy values for a smooth probability curve
prediction_data <- tibble(
  three_point_field_goal_pct = seq(10, 65, by = 0.5)
)

# Estimate win probability at each shooting percentage
prediction_data <- prediction_data %>%
  mutate(win_probability = predict(model, newdata = prediction_data,
                                   type = "response"))

# Plot the estimated probability of winning
ggplot(prediction_data,
       aes(x = three_point_field_goal_pct, y = win_probability)) +
  geom_line(colour = "navyblue", linewidth = 1.3) +
  geom_hline(yintercept = 0.5, linetype = "dashed", colour = "grey65") +
  scale_y_continuous(labels = scales::percent_format(accuracy = 1),
                     limits = c(0, 1)) +
  labs(title = "Three-Point Shooting Accuracy and NBA Win Probability",
       subtitle = "Estimated probabilities from binary logistic regression",
       x = "Three-point shooting accuracy (%)",
       y = "Estimated probability of winning") +
  theme_minimal(base_size = 13)

The upward-sloping curve indicates that teams with higher three-point shooting accuracy had a higher estimated probability of winning. However, the curve represents an association based on match statistics, not a guarantee of winning or a pre-match prediction.

6. Evaluating the Model

A confusion matrix compares predicted wins and losses with actual match outcomes. A team is classified as a predicted winner when its estimated win probability is at least 50%.

6.1 Classification Accuracy

The model was first evaluated using the same matches on which it was fitted. This is called in-sample accuracy.

# Calculate predicted probabilities and match outcomes
nba_data <- nba_data %>%
  mutate(
    win_probability = predict(model, type = "response"),
    predicted_win = if_else(win_probability >= 0.5, 1L, 0L)
  )

# Create a confusion matrix
confusion_matrix <- table(
  Actual = nba_data$win,
  Predicted = nba_data$predicted_win
)

# Display the confusion matrix
kable(confusion_matrix,
      caption = "Table 6. Confusion matrix for the development dataset")
Table 6. Confusion matrix for the development dataset
0 1
0 2401 1296
1 1358 2339
# Calculate classification accuracy
accuracy <- mean(nba_data$predicted_win == nba_data$win)
accuracy
## [1] 0.6410603

The model correctly classified 64.1% of team-game outcomes in the dataset used to fit it. Because each completed match contributes one win and one loss, a simple majority-class baseline is approximately 50%.

However, in-sample accuracy may overestimate how well a model performs on new matches.

6.2 Testing the Model on the Unseen 2025-26 NBA Season

To assess performance on unseen data, the binary logistic regression model developed using the 2022-23 to 2024-25 regular seasons was applied to the 2025-26 regular season. The 2025-26 observations were not included in model development. Separating seasons also prevents the two teams from the same match being divided between the development and test datasets.

# Download team box-score statistics for the unseen 2025–26 season
nba_2026 <- load_nba_team_box(seasons = 2026)

# Prepare the 2025–26 regular-season test dataset
test_data <- nba_2026 %>%
  filter(season_type == 2) %>%
  select(game_id, season, team_display_name,
         three_point_field_goal_pct, team_winner) %>%
  drop_na() %>%
  mutate(win = as.integer(team_winner))

# Apply the model trained on 2022–23 to 2024–25 to the 2025–26 season
test_data <- test_data %>%
  mutate(
    win_probability = predict(model, newdata = test_data,
                              type = "response"),
    predicted_win = if_else(win_probability >= 0.5, 1L, 0L)
  )

# Compare actual and predicted outcomes for the 2025–26 season
test_confusion <- table(
  Actual = factor(test_data$win, levels = c(0, 1)),
  Predicted = factor(test_data$predicted_win, levels = c(0, 1))
)

kable(test_confusion,
      caption = "Table 7. Confusion matrix for the unseen 2025–26 season")
Table 7. Confusion matrix for the unseen 2025–26 season
0 1
0 830 405
1 470 765
# Calculate test accuracy and the proportion of incorrect classifications
test_accuracy <- mean(test_data$predicted_win == test_data$win)
test_error <- 1 - test_accuracy
incorrect_predictions <- sum(test_data$predicted_win != test_data$win)

# Summarise model performance on the unseen season
tibble(
  Measure = c("Correct classifications", "Incorrect classifications",
              "Classification accuracy", "Incorrect classification rate"),
  Result = c(as.character(sum(test_data$predicted_win == test_data$win)),
             as.character(incorrect_predictions),
             paste0(round(test_accuracy * 100, 1), "%"),
             paste0(round(test_error * 100, 1), "%"))
) %>%
  kable(caption = "Table 8. Classification performance on the 2025–26 season")
Table 8. Classification performance on the 2025–26 season
Measure Result
Correct classifications 1595
Incorrect classifications 875
Classification accuracy 64.6%
Incorrect classification rate 35.4%

The model correctly classified 64.6% of the 2,470 team-game observations in the unseen 2025–26 season. This was similar to the 64.1% in-sample accuracy obtained from the 2022–23 to 2024–25 dataset, although the two values represent different evaluation settings.

The remaining 35.4% of team outcomes (875 observations) were incorrectly classified. Some teams won despite relatively low three-point shooting accuracy, while others lost despite shooting accurately. This is expected because match outcomes also depend on other aspects of performance, including two-point scoring, free throws, rebounding, turnovers and defence. The results indicate that three-point shooting accuracy is informative but insufficient on its own to classify every outcome correctly.

7. Conclusion

This tutorial demonstrated how binary logistic regression can be applied to NBA team statistics to investigate the relationship between three-point shooting accuracy and match outcomes.

Using 7,394 team-game observations from the 2022–23 to 2024–25 NBA regular seasons, a positive relationship was identified. Winning teams had higher average three-point shooting accuracy than losing teams, and each additional percentage point of accuracy was associated with approximately 10.1% higher odds of winning.

The model correctly classified 64.1% of team-game outcomes in the 2022–23 to 2024–25 development dataset and 64.6% in the unseen 2025–26 season. Approximately 35.4% of 2025–26 team outcomes were incorrectly classified. Although shooting accuracy provides useful information, it cannot fully explain match outcomes on its own.

Overall, the findings demonstrate how binary logistic regression can be used to examine relationships between sporting performance indicators and match outcomes. Including additional variables, such as rebounding, turnovers and defensive statistics, may improve classification accuracy in future analyses. However, shooting accuracy is measured during the match, so this analysis should not be interpreted as a pre-match forecasting model.

7.1 Key Takeaways

  • Shooting performance: Winning teams generally recorded higher three-point shooting accuracy than losing teams.
  • Model interpretation: Binary logistic regression identified a positive association between shooting accuracy and the odds of winning an NBA match.
  • Model evaluation: Testing the model on a separate season provides a more realistic assessment of classification performance than training accuracy alone.
  • Practical application: Binary logistic regression is useful for investigating relationships between sporting performance indicators and win/loss outcomes, although many factors contribute to match results.