Research Question: Can we predict the ALL NBA teams using statistical modelling?

  1. A prediction point must be determined, when are we making this prediction? Our prediction point will be after the buzzer sounds of the final NBA Regular season game.

  2. What Data are we going to use? We are going to use the last 11 Seasons of NBA data. From and including the 2015/16 season up to and including the 2025/26 season.

  3. What model are we going to use? Logistic Regression will be used to determine two outcomes:

Data Collection and Preparation

Data Collection

The dataset was manually compiled from Basketball-Reference.com from the 2015–16 to 2025–26 NBA seasons. Rather than using a single dataset, three tables were retrieved from each season:

Player Per Game Statistics: Individual player statistics for each season. Rows containing partial-season statistics were excluded to avoid duplicate player entries where a player appeared for multiple teams during the same season. Team Advanced Statistics: Advanced statistical measures for each NBA team during the regular season.

All-NBA Voting: Voting results used to identify players selected to an All-NBA Team.

Team Advanced Statistics: Team based Advanced statistics for each season. This was used to incorporate team level performance.

Each table was manually retrieved in CSV format from its respective season page. The tables were then organised into separate spreadsheets for further preparation before being imported into R.

Data Preparation

Before importing the data, two adjustments were made to the Team Advanced Statistics table in Excel. Firstly, a new column was made to mark if a team made the playoffs that season (yes = 1, no = 0). Basketball-reference.com mark these teams with a asterisks (*). After that was complete, any asterisks (*) from team names were removed. Finally, each team was assigned its three-letter abbreviation to match the team identifiers used in the Player Per Game Statistics table.

All (%) symbols were replaced with “_pct”

These adjustments were made to improve consistency between the tables and support the merging of player-level statistics with team-level data during the analysis.

1. Load Key Packages

Use Install_packages() if you do not have them installed

library(dplyr) 
## Warning: package 'dplyr' was built under R version 4.5.3
library(tidyr)
library(ggplot2)
## Warning: package 'ggplot2' was built under R version 4.5.3
library(gamlss)
## Warning: package 'gamlss' was built under R version 4.5.3
## Warning: package 'gamlss.dist' was built under R version 4.5.3
library(readr)

2. Collect & Import Data

As this data was collated manually, I will provide the best guide possible to follow.

Dataframe Names for each csv file

  • All-NBA Voting = allnbavoting
  • Team Advanced Statistics = teamadv
  • Player Per Game Statistics = pergameraw

3. Data manipulation

View variable types

Ensure variable types are correct

glimpse(pergameraw)
## Rows: 7,595
## Columns: 33
## $ Year              <int> 2026, 2026, 2026, 2026, 2026, 2026, 2026, 2026, 2026…
## $ Rk                <int> 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 1…
## $ Player            <chr> "Luka Dončić", "Shai Gilgeous-Alexander", "Anthony E…
## $ Age               <int> 26, 27, 24, 29, 25, 34, 29, 30, 31, 31, 28, 37, 29, …
## $ Team              <chr> "LAL", "OKC", "MIN", "BOS", "PHI", "LAC", "CLE", "DE…
## $ Pos               <chr> "PG", "PG", "SG", "SF", "PG", "SF", "SG", "C", "PF",…
## $ G                 <int> 64, 68, 61, 71, 70, 65, 70, 65, 36, 38, 42, 43, 64, …
## $ GS                <int> 64, 68, 60, 71, 70, 65, 70, 65, 36, 38, 42, 41, 64, …
## $ MP                <dbl> 35.8, 33.2, 35.0, 34.4, 38.0, 32.1, 33.5, 34.8, 28.9…
## $ FG                <dbl> 10.8, 10.8, 9.9, 10.4, 9.9, 9.8, 9.7, 9.9, 10.4, 9.0…
## $ FGA               <dbl> 22.8, 19.4, 20.2, 21.7, 21.4, 19.4, 20.0, 17.4, 16.6…
## $ FG_pct            <dbl> 0.476, 0.553, 0.489, 0.477, 0.462, 0.505, 0.483, 0.5…
## $ X3P               <dbl> 4.0, 1.7, 3.4, 2.0, 3.1, 2.6, 3.2, 1.7, 0.4, 1.4, 2.…
## $ X3PA              <dbl> 10.8, 4.4, 8.4, 5.7, 8.6, 6.8, 8.8, 4.5, 1.3, 4.2, 7…
## $ X3P_pct           <dbl> 0.366, 0.386, 0.399, 0.347, 0.367, 0.387, 0.364, 0.3…
## $ X2P               <dbl> 6.9, 9.1, 6.5, 8.4, 6.8, 7.1, 6.5, 8.2, 9.9, 7.6, 6.…
## $ X2PA              <dbl> 11.9, 15.0, 11.8, 16.0, 12.9, 12.5, 11.2, 12.9, 15.4…
## $ X2P_pct           <dbl> 0.575, 0.602, 0.554, 0.523, 0.525, 0.569, 0.577, 0.6…
## $ eFG_pct           <dbl> 0.563, 0.597, 0.572, 0.522, 0.536, 0.573, 0.563, 0.6…
## $ FT                <dbl> 7.9, 7.9, 5.7, 6.0, 5.3, 5.7, 5.3, 6.1, 6.4, 7.5, 5.…
## $ FTA               <dbl> 10.1, 9.0, 7.2, 7.5, 6.0, 6.4, 6.1, 7.4, 9.9, 8.8, 6…
## $ FT_pct            <dbl> 0.780, 0.879, 0.796, 0.795, 0.892, 0.892, 0.865, 0.8…
## $ ORB               <dbl> 0.6, 0.6, 0.6, 1.1, 0.3, 1.1, 0.7, 3.0, 2.7, 2.0, 2.…
## $ DRB               <dbl> 7.1, 3.7, 4.4, 5.8, 3.8, 5.3, 3.8, 9.9, 7.1, 5.7, 4.…
## $ TRB               <dbl> 7.7, 4.3, 5.0, 6.9, 4.1, 6.4, 4.5, 12.9, 9.8, 7.7, 6…
## $ AST               <dbl> 8.3, 6.6, 3.7, 5.1, 6.6, 3.6, 5.7, 10.7, 5.4, 3.9, 2…
## $ STL               <dbl> 1.6, 1.4, 1.4, 1.0, 1.9, 1.9, 1.5, 1.4, 0.9, 0.6, 1.…
## $ BLK               <dbl> 0.5, 0.8, 0.8, 0.4, 0.8, 0.4, 0.3, 0.8, 0.7, 1.2, 0.…
## $ TOV               <dbl> 4.0, 2.2, 2.9, 3.6, 2.4, 2.0, 2.8, 3.7, 3.2, 2.9, 1.…
## $ PF                <dbl> 2.4, 2.0, 1.9, 2.7, 2.2, 1.2, 2.3, 2.7, 2.4, 2.2, 1.…
## $ PTS               <dbl> 33.5, 31.1, 28.8, 28.7, 28.3, 27.9, 27.9, 27.7, 27.6…
## $ Awards            <chr> "MVP-4CPOY-8ASNBA1", "MVP-1CPOY-1ASNBA1", "CPOY-3AS"…
## $ Player.additional <chr> "doncilu01", "gilgesh01", "edwaran01", "brownja02", …
glimpse(allnbavoting)
## Rows: 341
## Columns: 25
## $ Year     <int> 2026, 2026, 2026, 2026, 2026, 2026, 2026, 2026, 2026, 2026, 2…
## $ Team     <chr> "1T", "1T", "1T", "1T", "1T", "2T", "2T", "2T", "2T", "2T", "…
## $ Pos      <chr> "G", "C", "C", "G", "G", "F", "F", "G", "F", "G", "G", "G", "…
## $ Player   <chr> "Shai Gilgeous-Alexander", "Nikola Jokić", "Victor Wembanyama…
## $ Age      <int> 27, 30, 22, 26, 24, 29, 34, 29, 37, 29, 25, 28, 24, 22, 23, 2…
## $ Tm       <chr> "OKC", "DEN", "SAS", "LAL", "DET", "BOS", "LAC", "CLE", "HOU"…
## $ Pts.Won  <int> 500, 500, 498, 482, 414, 384, 277, 276, 241, 197, 168, 149, 1…
## $ Pts.Max  <int> 500, 500, 500, 500, 500, 500, 500, 500, 500, 500, 500, 500, 5…
## $ Share    <dbl> 1.000, 1.000, 0.996, 0.964, 0.828, 0.768, 0.554, 0.552, 0.482…
## $ X1st_Tm  <int> 100, 100, 99, 91, 60, 44, 4, 2, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,…
## $ X2nd_Tm  <int> 0, 0, 1, 9, 38, 54, 81, 85, 72, 49, 36, 27, 19, 16, 8, 1, 1, …
## $ X3rd_Tm  <int> 0, 0, 0, 0, 0, 2, 14, 11, 25, 50, 60, 68, 68, 73, 63, 23, 11,…
## $ G        <int> 68, 65, 64, 64, 64, 71, 65, 70, 78, 74, 70, 75, 72, 70, 69, 6…
## $ MP       <dbl> 33.2, 34.8, 29.2, 35.8, 33.9, 34.4, 32.1, 33.5, 36.4, 35.0, 3…
## $ PTS      <dbl> 31.1, 27.7, 25.0, 33.5, 23.9, 28.7, 27.9, 27.9, 26.0, 26.0, 2…
## $ TRB      <dbl> 4.3, 12.9, 11.5, 7.7, 5.5, 6.9, 6.4, 4.5, 5.5, 3.3, 4.1, 4.4,…
## $ AST      <dbl> 6.6, 10.7, 3.1, 8.3, 9.9, 5.1, 3.6, 5.7, 4.8, 6.8, 6.6, 7.1, …
## $ STL      <dbl> 1.4, 1.4, 1.0, 1.6, 1.4, 1.0, 1.9, 1.5, 0.8, 0.8, 1.9, 0.9, 1…
## $ BLK      <dbl> 0.8, 0.8, 3.1, 0.5, 0.8, 0.4, 0.4, 0.3, 0.9, 0.1, 0.8, 0.4, 0…
## $ FG_pct   <dbl> 0.553, 0.569, 0.512, 0.476, 0.461, 0.477, 0.505, 0.483, 0.520…
## $ X3P_pct  <dbl> 0.386, 0.380, 0.349, 0.366, 0.342, 0.347, 0.387, 0.364, 0.413…
## $ FT_pct   <dbl> 0.879, 0.831, 0.827, 0.780, 0.812, 0.795, 0.892, 0.865, 0.874…
## $ WS       <dbl> 15.2, 14.9, 10.0, 9.5, 7.9, 6.9, 9.2, 8.2, 10.7, 8.8, 8.7, 9.…
## $ WS.48    <dbl> 0.323, 0.316, 0.257, 0.199, 0.174, 0.135, 0.212, 0.169, 0.180…
## $ playerid <chr> "gilgesh01", "jokicni01", "wembavi01", "doncilu01", "cunnica0…
glimpse(teamadv)
## Rows: 341
## Columns: 34
## $ Year          <int> 2016, 2017, 2018, 2019, 2020, 2021, 2022, 2023, 2024, 20…
## $ Rk            <int> 17, 9, 15, 25, 25, 22, 23, 23, 28, 30, 30, 10, 6, 4, 4, …
## $ Team          <chr> "Washington Wizards", "Washington Wizards", "Washington …
## $ Abr           <chr> "WAS", "WAS", "WAS", "WAS", "WAS", "WAS", "WAS", "WAS", …
## $ Age           <dbl> 27.3, 26.0, 26.9, 26.5, 25.1, 26.6, 25.9, 26.2, 24.9, 23…
## $ W             <int> 41, 49, 43, 32, 25, 34, 35, 35, 15, 18, 17, 40, 51, 48, …
## $ L             <int> 41, 33, 39, 50, 47, 38, 47, 47, 67, 64, 65, 42, 31, 34, …
## $ PW            <int> 40, 46, 43, 34, 26, 32, 32, 38, 20, 15, 16, 46, 52, 53, …
## $ PL            <int> 42, 36, 39, 48, 46, 40, 50, 44, 62, 67, 66, 36, 30, 29, …
## $ MOV           <dbl> -0.50, 1.80, 0.59, -2.90, -4.67, -1.83, -3.38, -1.21, -9…
## $ SOS           <dbl> 0.00, -0.45, -0.06, -0.40, -0.57, -0.01, 0.15, 0.15, 0.0…
## $ SRS           <dbl> -0.50, 1.36, 0.53, -3.30, -5.24, -1.85, -3.23, -1.06, -9…
## $ ORtg          <dbl> 105.3, 111.2, 109.3, 111.1, 110.9, 111.2, 111.1, 114.4, …
## $ DRtg          <dbl> 105.8, 109.3, 108.7, 113.9, 115.5, 113.0, 114.5, 115.6, …
## $ NRtg          <dbl> -0.5, 1.9, 0.6, -2.8, -4.6, -1.8, -3.4, -1.2, -9.1, -12.…
## $ Pace          <dbl> 98.5, 97.4, 96.6, 101.4, 102.7, 104.1, 97.0, 98.6, 102.7…
## $ FTr           <dbl> 0.263, 0.254, 0.254, 0.266, 0.270, 0.288, 0.252, 0.258, …
## $ X3PAr         <dbl> 0.282, 0.284, 0.310, 0.370, 0.358, 0.319, 0.356, 0.365, …
## $ TS_pct        <dbl> 0.544, 0.564, 0.560, 0.567, 0.562, 0.569, 0.568, 0.585, …
## $ X             <lgl> NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, …
## $ eFG_pct       <dbl> 0.511, 0.528, 0.525, 0.531, 0.523, 0.531, 0.532, 0.550, …
## $ TOV_pct       <dbl> 13.1, 12.8, 13.3, 12.3, 12.2, 12.3, 12.1, 12.7, 12.2, 13…
## $ ORB_pct       <dbl> 20.6, 24.1, 23.5, 21.3, 22.2, 21.3, 20.9, 22.6, 20.0, 22…
## $ FT.FGA        <dbl> 0.192, 0.199, 0.196, 0.204, 0.213, 0.221, 0.197, 0.202, …
## $ X.1           <lgl> NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, …
## $ eFG_pct.1     <dbl> 0.515, 0.524, 0.522, 0.546, 0.558, 0.539, 0.529, 0.540, …
## $ TOV_pct.1     <dbl> 14.6, 13.8, 13.6, 13.5, 13.9, 12.5, 10.7, 11.0, 12.0, 11…
## $ DRB_pct       <dbl> 77.7, 75.5, 77.1, 74.1, 75.3, 77.6, 76.9, 76.1, 72.5, 71…
## $ FT.FGA.1      <dbl> 0.218, 0.213, 0.212, 0.199, 0.231, 0.217, 0.202, 0.194, …
## $ X.2           <lgl> NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, …
## $ Arena         <chr> "Verizon Center", "Verizon Center", "Capital One Arena",…
## $ Attend.       <int> 725426, 697107, 739302, 716996, 532702, 19198, 637215, 7…
## $ Attend..G     <int> 17693, 17003, 18032, 17448, 16647, 533, 15542, 17329, 16…
## $ Made_Playoffs <int> 0, 1, 1, 0, 0, 1, 0, 0, 0, 0, 0, 0, 1, 1, 1, 1, 1, 1, 0,…

Change variable types and remove useless rows

library(dplyr)

# Remove league average rows 
pergameraw <- pergameraw |>
  filter(Player != "League Average")
teamadv <- teamadv |>
  filter(Team != "League Average")

# Correct variable type
pergameraw <- pergameraw |>
  mutate(across(
    c(Year, Age, G, GS, MP, FG, FGA, FG_pct, X3P, X3PA, X3P_pct,
      X2P, X2PA, X2P_pct, eFG_pct, FT, FTA, FT_pct, ORB, DRB, TRB,
      AST, STL, BLK, TOV, PF, PTS),
    as.numeric
  ))

allnbavoting <- allnbavoting |>
  mutate(across(
    c(Year, Age, Pts.Won, Pts.Max, Share, X1st_Tm, X2nd_Tm,
      X3rd_Tm, G, MP, PTS, TRB, AST, STL, BLK, FG_pct, X3P_pct,
      FT_pct, WS, WS.48),
    as.numeric
  ))

teamadv <- teamadv |>
  mutate(across(
    c(Year, Rk, Age, W, L, PW, PL, MOV, SOS, SRS, ORtg,
      DRtg, NRtg, Pace, FTr, X3PAr, TS_pct, eFG_pct, TOV_pct, ORB_pct,
      FT.FGA, eFG_pct.1, TOV_pct.1, DRB_pct, FT.FGA.1, Attend.,
      Attend..G, Made_Playoffs),
    as.numeric
  ))

Standardizing labels between sheets

Some labels are different despite representing the same data.

teamadv <- teamadv |> 
  rename(TeamId = Abr)

pergameraw <- pergameraw |> 
  rename(TeamId = Team)

pergameraw <- pergameraw |> 
  rename(playerid = Player.additional)

allnbavoting <- allnbavoting |> 
  rename(TeamId = Tm)

Assigning all players a team

If a player played for multiple teams, the team they played the most games for will be assigned to them

pergame <- pergameraw |>
  filter(!TeamId %in% c("2TM", "3TM", "4TM")) |>
  group_by(playerid, Year) |>
  filter(row_number() == which.max(G)) |>
  ungroup()

Adding important columns

pergame <- pergame |> 
  left_join(allnbavoting |> 
              select(playerid, Year, Team, Share),
            by = c('playerid', 'Year'))

pergame <- pergame |>
  left_join(teamadv |>
              select(TeamId, Year, Made_Playoffs, W, L),
            by = c('Year', 'TeamId'))

pergame <- pergame |> 
  mutate(allnba = if_else(Team %in% c('1st', '1T', '2nd', '2T', '3rd', '3T'), 1, 0))
summary(pergame)
##       Year            Rk         Player               Age      
##  Min.   :2016   Min.   :  1   Length:5968        Min.   :19.0  
##  1st Qu.:2018   1st Qu.:136   Class :character   1st Qu.:23.0  
##  Median :2021   Median :272   Mode  :character   Median :25.0  
##  Mean   :2021   Mean   :273                      Mean   :25.9  
##  3rd Qu.:2024   3rd Qu.:407                      3rd Qu.:29.0  
##  Max.   :2026   Max.   :605                      Max.   :43.0  
##                                                                
##     TeamId              Pos                  G               GS      
##  Length:5968        Length:5968        Min.   : 1.00   Min.   : 0.0  
##  Class :character   Class :character   1st Qu.:23.00   1st Qu.: 0.0  
##  Mode  :character   Mode  :character   Median :49.00   Median : 7.0  
##                                        Mean   :45.17   Mean   :21.5  
##                                        3rd Qu.:68.00   3rd Qu.:40.0  
##                                        Max.   :82.00   Max.   :82.0  
##                                                                      
##        MP              FG              FGA             FG_pct      
##  Min.   : 0.50   Min.   : 0.000   Min.   : 0.000   Min.   :0.0000  
##  1st Qu.:12.00   1st Qu.: 1.400   1st Qu.: 3.300   1st Qu.:0.4030  
##  Median :19.30   Median : 2.600   Median : 5.800   Median :0.4460  
##  Mean   :19.48   Mean   : 3.186   Mean   : 6.958   Mean   :0.4486  
##  3rd Qu.:27.40   3rd Qu.: 4.400   3rd Qu.: 9.600   3rd Qu.:0.4980  
##  Max.   :43.50   Max.   :11.800   Max.   :24.500   Max.   :1.0000  
##                                                    NA's   :30      
##       X3P              X3PA           X3P_pct            X2P        
##  Min.   :0.0000   Min.   : 0.000   Min.   :0.0000   Min.   : 0.000  
##  1st Qu.:0.2000   1st Qu.: 0.800   1st Qu.:0.2780   1st Qu.: 0.900  
##  Median :0.7000   Median : 2.200   Median :0.3390   Median : 1.700  
##  Mean   :0.9178   Mean   : 2.612   Mean   :0.3126   Mean   : 2.269  
##  3rd Qu.:1.4000   3rd Qu.: 3.900   3rd Qu.:0.3810   3rd Qu.: 3.200  
##  Max.   :5.3000   Max.   :13.200   Max.   :1.0000   Max.   :11.600  
##                                    NA's   :379                      
##       X2PA           X2P_pct          eFG_pct             FT       
##  Min.   : 0.000   Min.   :0.0000   Min.   :0.0000   Min.   : 0.00  
##  1st Qu.: 1.800   1st Qu.:0.4620   1st Qu.:0.4760   1st Qu.: 0.50  
##  Median : 3.300   Median :0.5140   Median :0.5210   Median : 0.90  
##  Mean   : 4.348   Mean   :0.5114   Mean   :0.5115   Mean   : 1.36  
##  3rd Qu.: 6.000   3rd Qu.:0.5710   3rd Qu.:0.5640   3rd Qu.: 1.80  
##  Max.   :19.200   Max.   :1.0000   Max.   :1.5000   Max.   :10.20  
##                   NA's   :76       NA's   :30                      
##       FTA             FT_pct            ORB              DRB        
##  Min.   : 0.000   Min.   :0.0000   Min.   :0.0000   Min.   : 0.000  
##  1st Qu.: 0.600   1st Qu.:0.6820   1st Qu.:0.3000   1st Qu.: 1.400  
##  Median : 1.300   Median :0.7680   Median :0.6000   Median : 2.300  
##  Mean   : 1.771   Mean   :0.7465   Mean   :0.8572   Mean   : 2.687  
##  3rd Qu.: 2.300   3rd Qu.:0.8340   3rd Qu.:1.1000   3rd Qu.: 3.600  
##  Max.   :12.300   Max.   :1.0000   Max.   :5.4000   Max.   :11.400  
##                   NA's   :308                                       
##       TRB              AST              STL              BLK        
##  Min.   : 0.000   Min.   : 0.000   Min.   :0.0000   Min.   :0.0000  
##  1st Qu.: 1.800   1st Qu.: 0.700   1st Qu.:0.3000   1st Qu.:0.1000  
##  Median : 3.050   Median : 1.400   Median :0.6000   Median :0.3000  
##  Mean   : 3.541   Mean   : 1.958   Mean   :0.6291   Mean   :0.3936  
##  3rd Qu.: 4.700   3rd Qu.: 2.600   3rd Qu.:0.9000   3rd Qu.:0.5000  
##  Max.   :16.000   Max.   :11.700   Max.   :3.0000   Max.   :3.8000  
##                                                                     
##       TOV              PF            PTS            Awards         
##  Min.   :0.000   Min.   :0.00   Min.   : 0.000   Length:5968       
##  1st Qu.:0.500   1st Qu.:1.10   1st Qu.: 3.900   Class :character  
##  Median :0.900   Median :1.70   Median : 7.000   Mode  :character  
##  Mean   :1.085   Mean   :1.65   Mean   : 8.646                     
##  3rd Qu.:1.500   3rd Qu.:2.20   3rd Qu.:11.900                     
##  Max.   :5.700   Max.   :6.00   Max.   :36.100                     
##                                                                    
##    playerid             Team               Share        Made_Playoffs   
##  Length:5968        Length:5968        Min.   :0.0020   Min.   :0.0000  
##  Class :character   Class :character   1st Qu.:0.0100   1st Qu.:0.0000  
##  Mode  :character   Mode  :character   Median :0.1120   Median :1.0000  
##                                        Mean   :0.2903   Mean   :0.5171  
##                                        3rd Qu.:0.5160   3rd Qu.:1.0000  
##                                        Max.   :1.0000   Max.   :1.0000  
##                                        NA's   :5627                     
##        W               L             allnba       
##  Min.   :10.00   Min.   : 9.00   Min.   :0.00000  
##  1st Qu.:31.00   1st Qu.:32.00   1st Qu.:0.00000  
##  Median :42.00   Median :39.00   Median :0.00000  
##  Mean   :39.53   Mean   :40.56   Mean   :0.02765  
##  3rd Qu.:48.00   3rd Qu.:49.00   3rd Qu.:0.00000  
##  Max.   :73.00   Max.   :72.00   Max.   :1.00000  
## 

Visualise variable relationships

library(GGally)

ggpairplots <- ggpairs(data = pergame, columns = c("allnba", 
                                           "PTS", 
                                           "Share", 
                                           "Made_Playoffs",
                                           "W",
                                           "L",
                                           "TRB",
                                           "AST",
                                           "STL",
                                           "BLK",
                                           "eFG_pct"))

ggpairplots

Split data into training and test datasets

We’ll train on the first 9 seasons (15/16 - 23/24), and test on the final 2 (24/25 & 25/26)

train <- pergame |> 
  filter(Year < 2025)

test <- pergame |> 
  filter(Year >= 2025)

Fit 5 Different models with different reasoning

# Model 1: Basic PRA and wins
fit1 <- glm(
  allnba ~ PTS + AST + TRB + W,
  data = train,
  family = binomial
)

# Model 2: Scoring contribution and efficiency
fit2 <- glm(
  allnba ~ PTS + AST + FG_pct + X3P_pct + FT_pct + eFG_pct,
  data = train,
  family = binomial
)

# Model 3: Counting Stats performance
fit3 <- glm(
  allnba ~ PTS + AST + TRB + STL + BLK + TOV + FG + X2P + X3P + FT,
  data = train,
  family = binomial
)

# Model 4: Playing time and experience
fit4 <- glm(
  allnba ~ Age + G + GS + MP + PTS + AST + TRB,
  data = train,
  family = binomial
)

# Model 5: Comprehensive model (Subjectively Hand Picked)
fit5 <- glm(
  allnba ~ G + GS + MP + PTS + AST + TRB +
    STL + BLK + TOV + FT_pct + eFG_pct + 
    W + Made_Playoffs,
  data = train,
  family = binomial
)

Create Probabilities and Predictions on Test Set

test <- test |>
  mutate(
    prob1 = predict(fit1, newdata = test, type = "response"),
    prob2 = predict(fit2, newdata = test, type = "response"),
    prob3 = predict(fit3, newdata = test, type = "response"),
    prob4 = predict(fit4, newdata = test, type = "response"),
    prob5 = predict(fit5, newdata = test, type = "response")
  )

test <- test |>
  mutate(
    pred1 = if_else(prob1 >= 0.5, 1, 0),
    pred2 = if_else(prob2 >= 0.5, 1, 0),
    pred3 = if_else(prob3 >= 0.5, 1, 0),
    pred4 = if_else(prob4 >= 0.5, 1, 0),
    pred5 = if_else(prob5 >= 0.5, 1, 0)
  )

Measuring accuracy of models against test set

accuracy <- tibble(
  Model = paste("Model", 1:5),
  Accuracy = c(
    mean(test$pred1 == test$allnba, na.rm = TRUE),
    mean(test$pred2 == test$allnba, na.rm = TRUE),
    mean(test$pred3 == test$allnba, na.rm = TRUE),
    mean(test$pred4 == test$allnba, na.rm = TRUE),
    mean(test$pred5 == test$allnba, na.rm = TRUE)
  )
) |>
  mutate(Accuracy = round(Accuracy * 100, 2))

accuracy
## # A tibble: 5 × 2
##   Model   Accuracy
##   <chr>      <dbl>
## 1 Model 1     98.9
## 2 Model 2     98.2
## 3 Model 3     97.9
## 4 Model 4     98.5
## 5 Model 5     98.9

It is important to note, despite accuracy for each model being above 95%, this is significantly swayed by the fact only 15 players can make ALL NBA out of the ~500 players.

We will Compare the results of the 3 Models. The Highest scoring, Hand Picked Model (5) The Simplest Model (1) and the Lowest Scoring Model (3)

Top 30 in Probability

model1_26 <- test |> filter(Year == 2026) |> slice_max(prob1, n = 30) |> select(Player, prob1, allnba, Pos)

model3_26 <- test |> filter(Year == 2026) |> slice_max(prob3, n = 30) |> select(Player, prob3, allnba, Pos)

model5_26 <- test |> filter(Year == 2026) |> slice_max(prob5, n = 30) |> select(Player, prob5,allnba, Pos)

Predicted teams for each model

# Model 1
Model_1_2026_1st <- model1_26 |> arrange(desc(prob1)) |> slice(1:5)
Model_1_2026_2nd <- model1_26 |> arrange(desc(prob1)) |> slice(6:10)
Model_1_2026_3rd <- model1_26 |> arrange(desc(prob1)) |> slice(11:15)

# Model 3
Model_3_2026_1st <- model3_26 |> arrange(desc(prob3)) |> slice(1:5)
Model_3_2026_2nd <- model3_26 |> arrange(desc(prob3)) |> slice(6:10)
Model_3_2026_3rd <- model3_26 |> arrange(desc(prob3)) |> slice(11:15)

# Model 5
Model_5_2026_1st <- model5_26 |> arrange(desc(prob5)) |> slice(1:5)
Model_5_2026_2nd <- model5_26 |> arrange(desc(prob5)) |> slice(6:10)
Model_5_2026_3rd <- model5_26 |> arrange(desc(prob5)) |> slice(11:15)

Combine into one Dataframe

all_teams_2026 <- bind_rows(
  Model_1_2026_1st |> mutate(Model = "Model 1", Year = 2026, Team = "1st"),
  Model_1_2026_2nd |> mutate(Model = "Model 1", Year = 2026, Team = "2nd"),
  Model_1_2026_3rd |> mutate(Model = "Model 1", Year = 2026, Team = "3rd"),
  Model_3_2026_1st |> mutate(Model = "Model 3", Year = 2026, Team = "1st"),
  Model_3_2026_2nd |> mutate(Model = "Model 3", Year = 2026, Team = "2nd"),
  Model_3_2026_3rd |> mutate(Model = "Model 3", Year = 2026, Team = "3rd"),
  Model_5_2026_1st |> mutate(Model = "Model 5", Year = 2026, Team = "1st"),
  Model_5_2026_2nd |> mutate(Model = "Model 5", Year = 2026, Team = "2nd"),
  Model_5_2026_3rd |> mutate(Model = "Model 5", Year = 2026, Team = "3rd")
) |> 
  select(Model, Year, Team, Player, starts_with("prob"), allnba)

Model team predictions vs actual

knitr::kable(
  all_teams_2026 |> select(Model, Team, Player, allnba),
  caption = "Predicted All-NBA Teams for 2026"
)
Predicted All-NBA Teams for 2026
Model Team Player allnba
Model 1 1st Nikola Jokić 1
Model 1 1st Luka Dončić 1
Model 1 1st Shai Gilgeous-Alexander 1
Model 1 1st Victor Wembanyama 1
Model 1 1st Jaylen Brown 1
Model 1 2nd Cade Cunningham 1
Model 1 2nd Jayson Tatum 0
Model 1 2nd Jalen Johnson 1
Model 1 2nd Donovan Mitchell 1
Model 1 2nd Jamal Murray 1
Model 1 3rd Karl-Anthony Towns 0
Model 1 3rd Jalen Duren 1
Model 1 3rd Anthony Edwards 0
Model 1 3rd Kevin Durant 1
Model 1 3rd Joel Embiid 0
Model 3 1st Nikola Jokić 1
Model 3 1st Luka Dončić 1
Model 3 1st Shai Gilgeous-Alexander 1
Model 3 1st Tyrese Maxey 1
Model 3 1st Victor Wembanyama 1
Model 3 2nd Kawhi Leonard 1
Model 3 2nd Giannis Antetokounmpo 0
Model 3 2nd Jalen Johnson 1
Model 3 2nd Jayson Tatum 0
Model 3 2nd Donovan Mitchell 1
Model 3 3rd Lauri Markkanen 0
Model 3 3rd Cade Cunningham 1
Model 3 3rd James Harden 0
Model 3 3rd Anthony Edwards 0
Model 3 3rd Jamal Murray 1
Model 5 1st Nikola Jokić 1
Model 5 1st Luka Dončić 1
Model 5 1st Shai Gilgeous-Alexander 1
Model 5 1st Tyrese Maxey 1
Model 5 1st Donovan Mitchell 1
Model 5 2nd Jaylen Brown 1
Model 5 2nd Jamal Murray 1
Model 5 2nd Cade Cunningham 1
Model 5 2nd Jalen Johnson 1
Model 5 2nd Kevin Durant 1
Model 5 3rd Jalen Duren 1
Model 5 3rd Kawhi Leonard 1
Model 5 3rd Jalen Brunson 1
Model 5 3rd Alperen Şengün 0
Model 5 3rd Karl-Anthony Towns 0

Final Performance 2026

all_teams_2026 |>
  group_by(Model) |>
  summarise(All_NBA_Count = sum(allnba == 1, na.rm = TRUE))
## # A tibble: 3 × 2
##   Model   All_NBA_Count
##   <chr>           <int>
## 1 Model 1            11
## 2 Model 3            10
## 3 Model 5            13

Conclusions

Model 5 performed best, correctly predicting 13 of 15 All-NBA selections in 2026, compared with 11 for Model 1 and 10 for Model 3. This suggests that combining playing time, individual performance, efficiency and team success may be more effective than relying on basic statistics or a broader set of counting statistics. However, testing on a single season limits the reliability of this conclusion.

Potential methodological improvements

  • Expand testing: Evaluate multiple unseen seasons to assess consistency across different NBA seasons.
  • Refine predictors: Use cross-validation and regularisation to identify the most useful variables and reduce redundancy.

Future Applications

The model could be updated throughout the NBA season to track changing All-NBA predictions. The methodology could also be adapted to forecast other awards, compare predictions with media voting and identify potential snubs. With further testing and refinement, it could become a useful tool for sports analytics and data-driven content creation.