When are we making the prediction? The prediction point will be before the start of the NBA Regular Season, after the previous season has concluded. Player fantasy point projections will be generated for the upcoming season using only information available at that time.
What data are we going to use? We will use NBA player and team statistics from the 2015/16 season through to the 2024/25 season to develop the model, with 2025/26 data used to evaluate its predictions against ESPN’s projections and actual fantasy point outcomes.
What model are we going to use? A Linear Mixed Model (LMM) will be developed using a manually selected set of player-level and team-level predictors. The model will account for variation between players and, where appropriate, other grouping factors such as teams or seasons. Its predictions will be compared against ESPN’s NBA Fantasy Point projections using three measures:
Correlation (r): How closely the model’s projections align with ESPN’s projections.
R Squared: How much variation in ESPN’s projections is associated with the model’s predictions.
Prediction error: How closely each set of projections estimates actual fantasy point production, measured using metrics such as RMSE and MAE.
The analysis uses the ESPN Points fantasy scoring system. Each player’s fantasy score is calculated from their box-score statistics using the following scoring weights:
| Statistic | Fantasy Points |
|---|---|
| Point | +1 |
| Rebound | +1 |
| Assist | +2 |
| Steal | +4 |
| Block | +4 |
| Field Goal Made | +2 |
| Field Goal Attempted | -1 |
| Three Pointer Made | +1 |
| Free Throw Made | +1 |
| Free Throw Attempted | -1 |
| Turnover | -2 |
library(dplyr)
## Warning: package 'dplyr' was built under R version 4.5.3
library(tidyr)
library(ggplot2)
## Warning: package 'ggplot2' was built under R version 4.5.3
library(readr)
The dataset was manually compiled from Basketball-Reference.com from the 2015–16 to 2025–26 NBA seasons. Rather than using a single dataset, three tables were retrieved from each season:
Player Per Game Statistics: Individual player statistics for each season. Rows containing partial-season statistics were excluded to avoid duplicate player entries where a player appeared for multiple teams during the same season. Team Advanced Statistics: Advanced statistical measures for each NBA team during the regular season.
Team Advanced Statistics: Team based Advanced statistics for each season. This was used to incorporate team level performance.
Each table was manually retrieved in CSV format from its respective season page. The tables were then organised into separate spreadsheets for further preparation before being imported into R.
library(dplyr)
# Remove league average rows
pergameraw <- pergameraw |>
filter(Player != "League Average")
teamadv <- teamadv |>
filter(Team != "League Average")
# Correct variable type
pergameraw <- pergameraw |>
mutate(across(
c(Year, Age, G, GS, MP, FG, FGA, FG_pct, X3P, X3PA, X3P_pct,
X2P, X2PA, X2P_pct, eFG_pct, FT, FTA, FT_pct, ORB, DRB, TRB,
AST, STL, BLK, TOV, PF, PTS),
as.numeric
))
## Warning: There were 26 warnings in `mutate()`.
## The first warning was:
## ℹ In argument: `across(...)`.
## Caused by warning:
## ! NAs introduced by coercion
## ℹ Run `dplyr::last_dplyr_warnings()` to see the 25 remaining warnings.
teamadv <- teamadv |>
rename(TeamId = Abr)
pergameraw <- pergameraw |>
rename(TeamId = Team)
pergameraw <- pergameraw |>
left_join(
pergameraw |>
transmute(
playerid,
Year = Year - 1,
Next_Stats = PTS + (AST * 2) + TRB + (STL * 4) + (BLK * 4) +
(FG * 2) - FGA - (TOV * 2) + X3P + FT - FTA
),
by = c("playerid", "Year")
) |>
left_join(
teamadv |>
select(TeamId, Year, ORtg, Pace, Age) |>
rename(tmAge = Age),
by = c("TeamId", "Year")
)
pergame_model <- pergameraw |>
filter(!is.na(Next_Stats), !is.na(Pace))
train <- pergame_model |>
filter(Year < 2025)
test <- pergame_model |>
filter(Year == 2025) |>
left_join (espn |>
select(Player, ESPN_FPTS),
by = c('Player')) |>
filter(!is.na(ESPN_FPTS))
# Simple PRA Model
model1 <- lm(Next_Stats ~ PTS + AST + TRB, data = train)
# Counting Stats Expanded Model
model2 <- lm(Next_Stats ~ PTS + AST + TRB + STL + BLK + TOV + FG + FGA + FT + FTA + X3P, data = train)
# Team Stats Inclusive
model3 <- lm(Next_Stats ~ PTS + AST + TRB + STL + BLK + TOV + FG + FGA + FT + FTA + X3P + Pace + ORtg + tmAge, data = train)
# Interaction Term Inclusive
model4 <- lm(
Next_Stats ~ PTS + AST + TRB + STL + BLK + TOV +
FG + FGA + FT + FTA + X3P + Pace + ORtg + tmAge +
PTS:FGA + AST:TOV + TRB:BLK + STL:BLK + FG:FGA + X3P:FGA +
Pace:PTS + ORtg:PTS,
data = train
)
test <- test |>
mutate(
Pred_1 = predict(model1, newdata = test),
Pred_2 = predict(model2, newdata = test),
Pred_3 = predict(model3, newdata = test),
Pred_4 = predict(model4, newdata = test)
)
library(ggplot2)
# Plot Model 1
ggplot(test, aes(x = Next_Stats, y = Pred_1)) +
geom_point(alpha = 0.5) +
geom_abline(slope = 1, intercept = 0, linetype = "dashed") +
labs(
title = "Model 1: Simple PRA - Predicted vs Actual FPTS",
x = "Actual FPTS",
y = "Predicted FPTS"
) +
theme_minimal()
# Plot Model 2
ggplot(test, aes(x = Next_Stats, y = Pred_2)) +
geom_point(alpha = 0.5) +
geom_abline(slope = 1, intercept = 0, linetype = "dashed") +
labs(
title = "Model 2: Counting Stats Model - Predicted vs Actual FPTS",
x = "Actual FPTS",
y = "Predicted FPTS"
) +
theme_minimal()
# Plot Model 3
ggplot(test, aes(x = Next_Stats, y = Pred_3)) +
geom_point(alpha = 0.5) +
geom_abline(slope = 1, intercept = 0, linetype = "dashed") +
labs(
title = "Model 3: Team Stats Inclusive - Predicted vs Actual FPTS",
x = "Actual FPTS",
y = "Predicted FPTS"
) +
theme_minimal()
# Plot Model 4
ggplot(test, aes(x = Next_Stats, y = Pred_4)) +
geom_point(alpha = 0.5) +
geom_abline(slope = 1, intercept = 0, linetype = "dashed") +
labs(
title = "Model 4: Interaction Term Inclusive - Predicted vs Actual FPTS",
x = "Actual FPTS",
y = "Predicted FPTS"
) +
theme_minimal()
ggplot(test, aes(x = Next_Stats, y = ESPN_FPTS)) +
geom_point(alpha = 0.5) +
geom_abline(slope = 1, intercept = 0, linetype = "dashed") +
labs(
title = "ESPN: Predicted vs Actual Fantasy Points",
x = "Actual Next_Stats",
y = "ESPN Projected Fantasy Points"
) +
theme_minimal()
# Model Summaries
library(sjPlot)
## Warning: package 'sjPlot' was built under R version 4.5.3
##
## Attaching package: 'sjPlot'
## The following object is masked from 'package:ggplot2':
##
## set_theme
tab_model(model1, model2, model3, model4)
| Next Stats | Next Stats | Next Stats | Next Stats | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Predictors | Estimates | CI | p | Estimates | CI | p | Estimates | CI | p | Estimates | CI | p |
| (Intercept) | 3.51 | 3.07 – 3.96 | <0.001 | 3.39 | 2.87 – 3.90 | <0.001 | -12.81 | -23.42 – -2.20 | 0.018 | -3.65 | -21.62 – 14.31 | 0.690 |
| PTS | 0.97 | 0.91 – 1.02 | <0.001 | 2.13 | -0.69 – 4.94 | 0.139 | 2.13 | -0.66 – 4.93 | 0.134 | 2.28 | -0.94 – 5.51 | 0.166 |
| AST | 1.30 | 1.14 – 1.46 | <0.001 | 1.44 | 1.20 – 1.67 | <0.001 | 1.25 | 1.01 – 1.49 | <0.001 | 1.08 | 0.72 – 1.45 | <0.001 |
| TRB | 1.14 | 1.03 – 1.25 | <0.001 | 0.50 | 0.34 – 0.66 | <0.001 | 0.50 | 0.34 – 0.66 | <0.001 | 0.54 | 0.35 – 0.73 | <0.001 |
| STL | 2.67 | 1.93 – 3.42 | <0.001 | 2.90 | 2.16 – 3.65 | <0.001 | 2.54 | 1.46 – 3.62 | <0.001 | |||
| BLK | 3.37 | 2.65 – 4.08 | <0.001 | 3.35 | 2.65 – 4.06 | <0.001 | 1.50 | 0.05 – 2.96 | 0.042 | |||
| TOV | -0.79 | -1.50 – -0.08 | 0.030 | -0.28 | -0.99 – 0.44 | 0.449 | -0.67 | -1.52 – 0.18 | 0.121 | |||
| FG | -0.77 | -6.40 – 4.87 | 0.790 | -1.43 | -7.02 – 4.16 | 0.615 | -3.45 | -9.16 – 2.26 | 0.236 | |||
| FGA | -0.77 | -1.06 – -0.47 | <0.001 | -0.42 | -0.73 – -0.10 | 0.009 | -0.52 | -0.85 – -0.18 | 0.003 | |||
| FT | -0.89 | -3.82 – 2.04 | 0.551 | -1.12 | -4.03 – 1.78 | 0.449 | -0.96 | -3.88 – 1.95 | 0.517 | |||
| FTA | -0.05 | -0.90 – 0.80 | 0.910 | 0.07 | -0.77 – 0.92 | 0.870 | -0.18 | -1.03 – 0.68 | 0.687 | |||
| X3P | -1.13 | -3.98 – 1.71 | 0.434 | -1.55 | -4.37 – 1.27 | 0.282 | -2.07 | -5.00 – 0.86 | 0.166 | |||
| Pace | -0.09 | -0.18 – 0.00 | 0.053 | -0.09 | -0.26 – 0.07 | 0.266 | ||||||
| ORtg | 0.22 | 0.16 – 0.28 | <0.001 | 0.16 | 0.06 – 0.26 | 0.001 | ||||||
| tmAge | 0.03 | -0.11 – 0.16 | 0.686 | 0.01 | -0.13 – 0.14 | 0.915 | ||||||
| PTS × FGA | -0.06 | -0.11 – -0.01 | 0.030 | |||||||||
| AST × TOV | 0.09 | -0.04 – 0.22 | 0.162 | |||||||||
| TRB × BLK | 0.12 | -0.05 – 0.28 | 0.162 | |||||||||
| STL × BLK | 1.75 | 0.23 – 3.26 | 0.024 | |||||||||
| FG × FGA | 0.19 | 0.05 – 0.33 | 0.007 | |||||||||
| FGA × X3P | 0.06 | -0.01 – 0.14 | 0.103 | |||||||||
| PTS × Pace | 0.00 | -0.01 – 0.01 | 0.970 | |||||||||
| PTS × ORtg | 0.00 | -0.01 – 0.01 | 0.419 | |||||||||
| Observations | 3357 | 3357 | 3357 | 3357 | ||||||||
| R2 / R2 adjusted | 0.710 / 0.710 | 0.727 / 0.726 | 0.732 / 0.731 | 0.736 / 0.735 | ||||||||
This output shows how the relationship between current season statistics and next season fantasy production changes as additional predictors are introduced. Across the four models, explanatory power increases steadily, although not every additional variable is statistically significant.
results <- test |>
summarise(
ESPN_R = cor(ESPN_FPTS, Next_Stats),
ESPN_R_Squared = ESPN_R^2,
ESPN_Variance = var(ESPN_FPTS - Next_Stats),
Pred1_R = cor(Pred_1, Next_Stats),
Pred1_R_Squared = Pred1_R^2,
Pred1_Variance = var(Pred_1 - Next_Stats),
Pred2_R = cor(Pred_2, Next_Stats),
Pred2_R_Squared = Pred2_R^2,
Pred2_Variance = var(Pred_2 - Next_Stats),
Pred3_R = cor(Pred_3, Next_Stats),
Pred3_R_Squared = Pred3_R^2,
Pred3_Variance = var(Pred_3 - Next_Stats),
Pred4_R = cor(Pred_4, Next_Stats),
Pred4_R_Squared = Pred4_R^2,
Pred4_Variance = var(Pred_4 - Next_Stats)
)
glimpse(results)
## Rows: 1
## Columns: 15
## $ ESPN_R <dbl> 0.8058955
## $ ESPN_R_Squared <dbl> 0.6494675
## $ ESPN_Variance <dbl> 37.21491
## $ Pred1_R <dbl> 0.7487029
## $ Pred1_R_Squared <dbl> 0.5605561
## $ Pred1_Variance <dbl> 47.07189
## $ Pred2_R <dbl> 0.7650477
## $ Pred2_R_Squared <dbl> 0.585298
## $ Pred2_Variance <dbl> 43.90803
## $ Pred3_R <dbl> 0.7760077
## $ Pred3_R_Squared <dbl> 0.602188
## $ Pred3_Variance <dbl> 42.0354
## $ Pred4_R <dbl> 0.7818114
## $ Pred4_R_Squared <dbl> 0.6112291
## $ Pred4_Variance <dbl> 42.19245
test |>
summarise(
ESPN_RMSE = sqrt(mean((ESPN_FPTS - Next_Stats)^2, na.rm = TRUE)),
ESPN_MAE = mean(abs(ESPN_FPTS - Next_Stats), na.rm = TRUE),
Model1_RMSE = sqrt(mean((Pred_1 - Next_Stats)^2, na.rm = TRUE)),
Model1_MAE = mean(abs(Pred_1 - Next_Stats), na.rm = TRUE),
Model2_RMSE = sqrt(mean((Pred_2 - Next_Stats)^2, na.rm = TRUE)),
Model2_MAE = mean(abs(Pred_2 - Next_Stats), na.rm = TRUE),
Model3_RMSE = sqrt(mean((Pred_3 - Next_Stats)^2, na.rm = TRUE)),
Model3_MAE = mean(abs(Pred_3 - Next_Stats), na.rm = TRUE),
Model4_RMSE = sqrt(mean((Pred_4 - Next_Stats)^2, na.rm = TRUE)),
Model4_MAE = mean(abs(Pred_4 - Next_Stats), na.rm = TRUE)
)
## ESPN_RMSE ESPN_MAE Model1_RMSE Model1_MAE Model2_RMSE Model2_MAE Model3_RMSE
## 1 6.488286 5.153731 6.888707 5.476782 6.618854 5.267533 6.471487
## Model3_MAE Model4_RMSE Model4_MAE
## 1 5.110689 6.482161 5.178391
The error results indicate that a manually specified regression model can produce fantasy point predictions comparable to ESPN’s projections. Model 3 marginally outperformed ESPN on both error metrics, suggesting that incorporating broader player statistics and team context can produce accurate predictions without requiring a complex model.
However, these differences are small, and without further testing across multiple seasons, we cannot conclude that Model 3 consistently outperforms ESPN.
The results suggest that the eyeballed linear regression models can approximate ESPN’s NBA fantasy point projections reasonably well, although ESPN performed best across all three reported metrics. Model 4 produced the strongest results of the four regression models, with the addition of team-level variables and interaction terms improving the correlation with actual fantasy points.
A logical next step would be to focus on improving the models predictive performance rather than continuing to add variables solely to increase R-squared. The current results suggest that most of the explanatory power comes from player level statistics, with smaller gains from team context and interaction terms. Future models could incorporate historical trends across multiple seasons rather than relying solely on a player’s most recent season, allowing changes in production and longer term performance patterns to inform predictions. More comprehensive team-level statistics could also be included to better understand a players potential role, such as their offensive responsibility, usage, competition for playing time and position within the teams rotation. These additions may help distinguish between players with similar individual statistics but different circumstances, improving predictions of how their fantasy production could change in the following season.
Note: The test set was restricted to the top 250 players projected by ESPN for the 2025/26 season. This was done because most fantasy leagues do not extend beyond this range, making the comparison more relevant to typical fantasy league settings. Consequently, the findings may not generalise to lower ranked players or the broader NBA player pool.