AFL fans often say scoring has gone up or down over the years. In this tutorial you will use real AFL results to answer that question with a time series analysis. By the end you will be able to:
fitzRoy
packagetsibble)Research question: Has the average total score in AFL games trended up or down since 2000, and how well can simple time series models forecast it?
No prior time series knowledge is needed. Every step has a short explanation.
Load these R packages. Use the install_packages()
function if you do not have them already:
library(fpp3) # time series tools (tsibble, feasts, fable) plus dplyr and ggplot2
## ── Attaching packages ──────────────────────────────────────────── fpp3 1.0.3 ──
## ✔ tibble 3.3.1 ✔ tsibble 1.2.0
## ✔ dplyr 1.2.1 ✔ tsibbledata 0.4.1
## ✔ tidyr 1.3.2 ✔ ggtime 1.0.0
## ✔ lubridate 1.9.5 ✔ feasts 0.5.0
## ✔ ggplot2 4.0.3 ✔ fable 0.5.0
## ── Conflicts ───────────────────────────────────────────────── fpp3_conflicts ──
## ✖ lubridate::date() masks base::date()
## ✖ dplyr::filter() masks stats::filter()
## ✖ tsibble::intersect() masks base::intersect()
## ✖ tsibble::interval() masks lubridate::interval()
## ✖ dplyr::lag() masks stats::lag()
## ✖ tsibble::setdiff() masks base::setdiff()
## ✖ tsibble::union() masks base::union()
library(fitzRoy) # AFL data
library(janitor)
##
## Attaching package: 'janitor'
## The following objects are masked from 'package:stats':
##
## chisq.test, fisher.test
fetch_results_afltables() downloads match results from
AFL Tables. We ask for every season from 2000 to 2025. This needs an
internet connection and can take a minute.
results_raw <- fetch_results_afltables(season = 2000:2025) |>
clean_names()
glimpse(results_raw)
## Rows: 5,111
## Columns: 16
## $ game <dbl> 11728, 11729, 11730, 11731, 11732, 11733, 11734, 11735, 1…
## $ date <date> 2000-03-08, 2000-03-09, 2000-03-10, 2000-03-11, 2000-03-…
## $ round <chr> "R1", "R1", "R1", "R1", "R1", "R1", "R1", "R1", "R2", "R2…
## $ home_team <chr> "Richmond", "Essendon", "North Melbourne", "Adelaide", "F…
## $ home_goals <int> 14, 24, 16, 15, 16, 15, 22, 13, 20, 23, 21, 12, 22, 14, 1…
## $ home_behinds <int> 10, 12, 15, 18, 11, 10, 20, 8, 10, 7, 13, 15, 22, 19, 8, …
## $ home_points <int> 94, 156, 111, 108, 107, 100, 152, 86, 130, 145, 139, 87, …
## $ away_team <chr> "Melbourne", "Port Adelaide", "West Coast", "Footscray", …
## $ away_goals <int> 13, 8, 24, 19, 19, 21, 16, 20, 12, 17, 15, 19, 18, 13, 21…
## $ away_behinds <int> 14, 14, 10, 17, 15, 8, 16, 20, 15, 18, 9, 11, 4, 14, 13, …
## $ away_points <int> 92, 62, 154, 131, 129, 134, 112, 140, 87, 120, 99, 125, 1…
## $ venue <chr> "M.C.G.", "Docklands", "M.C.G.", "Football Park", "Subiac…
## $ margin <int> 2, 94, -43, -23, -22, -34, 40, -54, 43, 25, 40, -38, 42, …
## $ season <dbl> 2000, 2000, 2000, 2000, 2000, 2000, 2000, 2000, 2000, 200…
## $ round_type <chr> "Regular", "Regular", "Regular", "Regular", "Regular", "R…
## $ round_number <int> 1, 1, 1, 1, 1, 1, 1, 1, 2, 2, 2, 2, 2, 2, 2, 2, 3, 3, 3, …
Each row is one game. The columns we need are season,
home_points and away_points.
Always check the data before analysing it. We expect roughly 180 to 230 games per season
results_raw |>
count(season) |>
summary()
## season n
## Min. :2000 Min. :162.0
## 1st Qu.:2006 1st Qu.:185.0
## Median :2012 Median :201.0
## Mean :2012 Mean :196.6
## 3rd Qu.:2019 3rd Qu.:207.0
## Max. :2025 Max. :216.0
We want one number per season: the average total score per game (home points + away points)
season_scoring <- results_raw |>
mutate(total_points = home_points + away_points) |>
group_by(season) |>
summarise(avg_total = mean(total_points), .groups = "drop") |>
mutate(season = as.integer(season)) |>
as_tsibble(index = season)
season_scoring
## # A tsibble: 26 x 2 [1Y]
## season avg_total
## <int> <dbl>
## 1 2000 207.
## 2 2001 195.
## 3 2002 190.
## 4 2003 188.
## 5 2004 186.
## 6 2005 189.
## 7 2006 185.
## 8 2007 191.
## 9 2008 195.
## 10 2009 182.
## # ℹ 16 more rows
as_tsibble(index = season) tells R that
season is the time variable. The data is
annual, so there is no within-year seasonality to
model.
season_scoring |>
autoplot(avg_total) +
labs(title = "Average total score per AFL game",
x = "Season", y = "Points per game (both teams)")
What to look for: Is the series trending up or down? Is there any sudden jump or dip? Think about what happened to quarter lengths in 2020.
What the plot shows: Average score fell from about 207 points per game in 2000 to about 160 in 2019, with small rises along the way (for example 2007 to 2008). There was a sharp one-year drop to about 121 in 2020, the season played with shortened quarters due to COVID, followed by a recovery to about 159 in 2021. Since 2022, scoring has been steady at roughly 166 to 169 points per game.
A lag plot compares each season with the previous one. Points close to a diagonal line mean this season’s scoring is similar to last season’s.
season_scoring |>
gg_lag(avg_total, geom = "point", lags = 1:4) +
labs(title = "Lag plots of average total score")
The ACF shows the correlation between the series and its own past values. Bars that cross the dashed blue lines are larger than expected from random noise.
season_scoring |>
ACF(avg_total, lag_max = 10) |>
autoplot() +
labs(title = "ACF of average total score")
How to read it: slowly decaying bars suggest a trend, which means the series is non-stationary (its average changes over time). Many forecasting methods work better after differencing (subtracting each value from the previous one). In our plot, only lags 1 and 2 are outside the dashed lines, and the bars shrink towards zero by about lag 7. THis fits a series with a trend and a fading memory, with no repeating seasonal pattern (as expected for annual data). We can ask R how many differences are needed:
season_scoring |>
features(avg_total, unitroot_ndiffs)
## # A tibble: 1 × 1
## ndiffs
## <int>
## 1 1
To check how good a model is, we hide some data. We fit models on 2000 to 2019 (training) and test them on 2020 to 2025.
train <- season_scoring |> filter(season <= 2019)
fit <- train |>
model(
Mean = MEAN(avg_total),
Naive = NAIVE(avg_total),
ETS = ETS(avg_total),
ARIMA = ARIMA(avg_total)
)
fit
## # A mable: 1 x 4
## Mean Naive ETS ARIMA
## <model> <model> <model> <model>
## 1 <MEAN> <NAIVE> <ETS(M,A,N)> <ARIMA(0,1,0) w/ drift>
Look at which ETS and ARIMA specifications were chosen:
fit |> select(ETS) |> report()
## Series: avg_total
## Model: ETS(M,A,N)
## Smoothing parameters:
## alpha = 0.0004338657
## beta = 0.0001000018
##
## Initial states:
## l[0] b[0]
## 199.9211 -1.557384
##
## sigma^2: 0.001
##
## AIC AICc BIC
## 135.8627 140.1484 140.8413
fit |> select(ARIMA) |> report()
## Series: avg_total
## Model: ARIMA(0,1,0) w/ drift
##
## Coefficients:
## constant
## -2.4606
## s.e. 1.3671
##
## sigma^2 estimated as 37.48: log likelihood=-60.87
## AIC=125.74 AICc=126.49 BIC=127.63
fc <- fit |> forecast(h = 6)
fc |>
autoplot(season_scoring |> filter(season >= 2000), level = 80) +
labs(title = "Forecasts for 2020 to 2025 compared to actual scoring",
x = "Season", y = "Points per game (both teams)")
The coloured lines are the forecasts, the shaded bands are 80% prediction intervals, and the black line is what actually happened.
accuracy(fc, season_scoring) |>
select(.model, RMSE, MAE) |>
arrange(RMSE)
## # A tibble: 4 × 3
## .model RMSE MAE
## <chr> <dbl> <dbl>
## 1 Naive 17.1 11.5
## 2 ETS 19.7 12.6
## 3 ARIMA 21.5 19.0
## 4 Mean 30.2 25.0
How to read it: RMSE and MAE are errors in points per game. Lower is better. A model that cannot beat the Mean or Naive benchmarks is not adding anything.
What the results show: The Naive benchmark had the lowest RMSE (17.1), followed by ETS (19.7), ARIMA (21.5) and Mean (30.2). So neither ETS nor ARIMA beat the Naive benchmark. In the forecast plot, ETS and ARIMA kept extending the downward trend seen from 2000 to 2019 (ARIMA falls to about 145 by 2025), but scoring leveled off at about 168 after 2020. The mean forecast (about 183) was too high because it averages the high-scoring early seasons. The 2020 season is also an outlier: quarters were shortened that year, which no model trained on earlier seasons could have predicted. With only six test seasons, the ranking should be treated with caution.
Now refit the models using all the data and forecast the next five seasons.
fit_all <- season_scoring |>
model(
ETS = ETS(avg_total),
ARIMA = ARIMA(avg_total)
)
fit_all |>
forecast(h = 5) |>
autoplot(season_scoring, level = 80) +
labs(title = "Forecast of average total score, next five seasons",
x = "Season", y = "Points per game (both teams)")
Both models forecast a flat level of about 168 to 169 points per game for 2026 to 2030. Prediction intervals widen the further ahead we look (ARIMA’s 80% interval reaches roughly 133 to 204 by 2030), which shows forecasting uncertainty honestly.
A good model leaves behind residuals (errors) that look like random noise. Here we check the ARIMA model.
fit_all |>
select(ARIMA) |>
gg_tsresiduals()
What to look for: (1) residuals scattered around zero with no pattern, (2) ACF bars inside the dashed lines, (3) a roughly bell-shaped histogram.
In our plot: most residuals are within about +- 12 points. The exceptions are 2020 (about -39) and 2021 (about +38), the disrupted season and its rebound. The residual ACF stays inside the dashed lines, so apart from those two seasons the model leaves little structure behind.
Scoring fell steadily from about 207 points per game in 2000 to about 160 in 2019, dropped to about 121 in the shortened 2020 season, and has leveled off near 168. On the 2020 to 2025 test set, the Naive benchmark had the lowest error (RMSE 17.1), ahead of ETS (19.7), ARIMA (21.5) and Mean (30.2), because the trend-following models kept projecting a decline that did not continue. Fitted to all the data, ETS and ARIMA both forecast roughly flat scoring of about 168 to 169 points per game for 2026 to 2030, with wide uncertainty. So scoring has trended down overall since 200, but with only 26 seasons and a major disruption in 2020, these simple models forecast it no better than a naive benchmark.
Limitations to keep in mind: