Which match statistics are associated with winning in the AFL?
AFL match statistics can help us explore how team performance relates to winning and losing. This tutorial uses data from the 2026 AFL season to examine whether differences in inside 50s, clearances and contested possessions are associated with match outcomes.
We will use logistic regression, a statistical method for modelling an outcome with two categories: a home-team win or loss. It allows us to examine each statistic’s association with winning while accounting for the other statistics included in the model.
The tutorial will guide you through downloading data using the fitzRoy package, calculating team totals, comparing opponents, creating exploratory graphs, and fitting and interpreting logistic regression models in R. Because the statistics are recorded during matches, we will explore their association with the result rather than forecast results before a match starts.
install.packages(c("fitzRoy",
"tidyverse",
"janitor",
"GGally",
"sjPlot",
"performance",
"report"))
library(fitzRoy)
## Warning: package 'fitzRoy' was built under R version 4.6.1
library(tidyverse)
library(janitor)
library(GGally)
library(sjPlot)
For this task, we need two datasets from AFL Tables:
player_stats: One row per player per match (kicks,
tackles, clearances, etc.)match_results: One row per match (teams, scores, round,
etc.)# Import - Player statistics for the 2026 season
player_stats <- fetch_player_stats_afltables(season = 2026) |>
clean_names()
# Import - Match results for the 2026 season
match_results <- fetch_results_afltables(season = 2026) |>
clean_names()
Use glimpse() to check the structure of each
dataset.
# Check - Data structure
match_results |>
select(date, home_team, away_team,
home_points, away_points, round_type) |>
glimpse()
## Rows: 218
## Columns: 6
## $ date <date> 2026-03-05, 2026-03-06, 2026-03-07, 2026-03-07, 2026-03-0…
## $ home_team <chr> "Sydney", "Gold Coast", "GWS", "Brisbane Lions", "St Kilda…
## $ away_team <chr> "Carlton", "Geelong", "Hawthorn", "Footscray", "Collingwoo…
## $ home_points <int> 132, 125, 122, 106, 66, 75, 83, 134, 110, 104, 79, 113, 12…
## $ away_points <int> 69, 69, 95, 111, 78, 71, 145, 53, 100, 60, 93, 67, 107, 72…
## $ round_type <chr> "Regular", "Regular", "Regular", "Regular", "Regular", "Re…
player_stats |>
select(date, playing_for, inside_50s,
clearances, contested_possessions) |>
glimpse()
## Rows: 10,028
## Columns: 5
## $ date <date> 2026-03-05, 2026-03-05, 2026-03-05, 2026-03-05,…
## $ playing_for <chr> "Sydney", "Sydney", "Sydney", "Sydney", "Sydney"…
## $ inside_50s <int> 2, 1, 5, 1, 3, 8, 5, 1, 2, 0, 1, 6, 2, 2, 1, 0, …
## $ clearances <int> 0, 0, 0, 0, 3, 5, 6, 1, 0, 1, 2, 3, 0, 0, 0, 1, …
## $ contested_possessions <int> 3, 6, 5, 4, 6, 8, 10, 5, 5, 4, 4, 8, 3, 5, 4, 9,…
We want to create a dataset with one row per match, with the following columns:
home_win: The outcome: 1 if the home team won, 0 if it
lostinside_50s_diff: Home team inside 50s minus away team
inside 50sclearances_diff: Home team clearances minus away team
clearancescontested_possessions_diff: Home team contested
possessions minus away team contested possessionsWe add up every player’s statistics to get a total for each team in each match.
# Add together player statistics for each team in each match
team_stats <- player_stats |>
group_by(season, date, home_team, away_team, playing_for) |>
summarise(inside_50s = sum(inside_50s),
clearances = sum(clearances),
contested_possessions = sum(contested_possessions),
.groups = "drop")
# Check - New dataset
head(team_stats)
## # A tibble: 6 × 8
## season date home_team away_team playing_for inside_50s clearances
## <int> <date> <chr> <chr> <chr> <int> <int>
## 1 2026 2026-03-05 Sydney Carlton Carlton 50 39
## 2 2026 2026-03-05 Sydney Carlton Sydney 67 40
## 3 2026 2026-03-06 Gold Coast Geelong Geelong 50 41
## 4 2026 2026-03-06 Gold Coast Geelong Gold Coast 71 34
## 5 2026 2026-03-07 Brisbane Lions Western Bu… Brisbane L… 63 45
## 6 2026 2026-03-07 Brisbane Lions Western Bu… Western Bu… 50 39
## # ℹ 1 more variable: contested_possessions <int>
Each match should now have two rows (one per team) and no missing values. Check to confirm this:
# Check - Missing team totals
team_stats |>
summarise(missing_inside_50s = sum(is.na(inside_50s)),
missing_clearances = sum(is.na(clearances)),
missing_contested_possessions = sum(is.na(contested_possessions)))
## # A tibble: 1 × 3
## missing_inside_50s missing_clearances missing_contested_possessions
## <int> <int> <int>
## 1 0 0 0
# Check - Each match has two teams
team_stats |>
count(season, date, home_team, away_team, name = "number_of_teams") |>
count(number_of_teams, name = "number_of_matches")
## # A tibble: 1 × 2
## number_of_teams number_of_matches
## <int> <int>
## 1 2 218
The first table should show zeros. The second should show that every
match has number_of_teams equal to 2.
This tutorial focuses on home-and-away matches from the 2026 AFL season. Finals are excluded to keep the analysis focused on the regular season. Let’s check how the match types are labelled:
# Check - Regular season and finals labels
match_results |>
count(round_type)
## # A tibble: 2 × 2
## round_type n
## <chr> <int>
## 1 Finals 9
## 2 Regular 209
The two datasets spell some club names differently (e.g
Footscray and Western Bulldogs). If we don’t
fix this, the datasets will not match up when we combine them.
case_when() replaces the names we list and leaves every
other name unchanged (TRUE ~ .x).
# Standardise club names in the team statistics
team_stats <- team_stats |>
mutate(
across(
c(home_team, away_team, playing_for),
~ case_when(.x == "GWS" ~ "Greater Western Sydney",
.x == "Footscray" ~ "Western Bulldogs",
TRUE ~ .x)))
# Standardise club names in the match results
match_results <- match_results |>
mutate(
across(
c(home_team, away_team),
~ case_when(.x == "GWS" ~ "Greater Western Sydney",
.x == "Footscray" ~ "Western Bulldogs",
TRUE ~ .x)))
# Ensure dates use the same format in both datasets
team_stats <- team_stats |>
mutate(date = as.Date(date))
match_results <- match_results |>
mutate(date = as.Date(date))
# Select home-team statistics and give them clear column names
home_stats <- team_stats |>
filter(playing_for == home_team) |>
select(season, date, home_team, away_team,
home_inside_50s = inside_50s,
home_clearances = clearances,
home_contested_possessions = contested_possessions)
# Select away-team statistics
away_stats <- team_stats |>
filter(playing_for == away_team) |>
select(season, date, home_team, away_team,
away_inside_50s = inside_50s,
away_clearances = clearances,
away_contested_possessions = contested_possessions)
# Combine results with both teams' statistics
afl_matches <- match_results |>
filter(round_type == "Regular") |>
select(season, date, round_number,
home_team, away_team, home_points, away_points) |>
left_join(home_stats,
by = c("season", "date", "home_team", "away_team")) |>
left_join(away_stats,
by = c("season", "date", "home_team", "away_team"))
Now we calculate the three differences and record whether the home team won, lost or drew.
afl_matches <- afl_matches |>
mutate(inside_50s_diff = home_inside_50s - away_inside_50s,
clearances_diff = home_clearances - away_clearances,
contested_possessions_diff = home_contested_possessions - away_contested_possessions,
home_result = case_when(
home_points > away_points ~ "Win",
home_points < away_points ~ "Loss",
home_points == away_points ~ "Draw"))
# Check - Count the outcomes before excluding draws
afl_matches |>
count(home_result)
## # A tibble: 3 × 2
## home_result n
## <chr> <int>
## 1 Draw 3
## 2 Loss 84
## 3 Win 122
A logistic regression needs exactly two outcomes. Since draws are
rare, we will remove them and code the outcome as home_win
(1 = win, 0 = loss).
# Prepare - Modelling dataset
afl_model_data <- afl_matches |>
filter(home_result != "Draw") |>
mutate(home_win = as.integer(home_result == "Win")) |>
select(season, date, round_number, home_team, away_team,
home_win, inside_50s_diff, clearances_diff,
contested_possessions_diff)
# Check - Missing values after joining the datasets
afl_model_data |>
summarise(
across(
c(home_win, inside_50s_diff, clearances_diff, contested_possessions_diff),
~ sum(is.na(.x))))
## # A tibble: 1 × 4
## home_win inside_50s_diff clearances_diff contested_possessions_diff
## <int> <int> <int> <int>
## 1 0 0 0 0
# Check - Compact preview
afl_model_data |>
select(home_team, away_team, home_win, inside_50s_diff,
clearances_diff, contested_possessions_diff) |>
head()
## # A tibble: 6 × 6
## home_team away_team home_win inside_50s_diff clearances_diff
## <chr> <chr> <int> <int> <int>
## 1 Sydney Carlton 1 17 1
## 2 Gold Coast Geelong 1 21 -7
## 3 Greater Western Sydney Hawthorn 1 5 7
## 4 Brisbane Lions Western Bulld… 0 13 6
## 5 St Kilda Collingwood 0 23 2
## 6 Carlton Richmond 1 -17 8
## # ℹ 1 more variable: contested_possessions_diff <int>
The dataset contains 209 regular-season matches. Removing the three draws leaves 206 matches: 122 home-team wins and 84 home-team losses. There are no missing values in the outcome or the three predictors.
summary() gives the minimum, quartiles, median, mean and
maximum.
# Descriptive statistics
afl_model_data |>
select(inside_50s_diff,
clearances_diff,
contested_possessions_diff) |>
summary()
## inside_50s_diff clearances_diff contested_possessions_diff
## Min. :-29.000 Min. :-22.000 Min. :-44.000
## 1st Qu.: -6.000 1st Qu.: -5.000 1st Qu.: -6.000
## Median : 4.500 Median : 1.000 Median : 4.000
## Mean : 4.049 Mean : 1.063 Mean : 3.092
## 3rd Qu.: 15.000 3rd Qu.: 8.000 3rd Qu.: 13.000
## Max. : 32.000 Max. : 24.000 Max. : 44.000
Across these matches, home teams averaged approximately 4.05 more inside 50s, 1.06 more clearances and 3.09 more contested possessions than their opponents. Positive differences mean the home team recorded more of that statistic; negative differences mean the away team recorded more.
We label home-team outcomes as ‘Win’ or ‘Loss’ to make the tables and graphs easier to read, then compare the average statistical differences. We keep home_win coded as 0/1 for the logistic regression.
# Add readable labels while keeping home_win for modelling
afl_model_data <- afl_model_data |>
mutate(
result = factor(home_win,
levels = c(0, 1),
labels = c("Loss", "Win")))
# Compare average statistical differences by outcome
afl_model_data |>
group_by(result) |>
summarise(matches = n(),
mean_inside_50s_diff = mean(inside_50s_diff),
mean_clearances_diff = mean(clearances_diff),
mean_contested_possessions_diff = mean(contested_possessions_diff),
.groups = "drop")
## # A tibble: 2 × 5
## result matches mean_inside_50s_diff mean_clearances_diff
## <fct> <int> <dbl> <dbl>
## 1 Loss 84 -3.31 -2.06
## 2 Win 122 9.11 3.21
## # ℹ 1 more variable: mean_contested_possessions_diff <dbl>
Home teams that won averaged 9.11 more inside 50s and 3.21 more clearances than their opponents. Home teams that lost averaged 3.31 fewer inside 50s and 2.06 fewer clearances. These comparisons suggest that the statistical differences are related to match outcomes, but they do not account for the other statistics.
We can check how closely the three statistics are related. Strong relationships can make it harder to identify each statistic’s separate association with winning.
ggpairs() displays graphs and correlation values to help
us explore these relationships.
# Select - Predictors only
afl_cor <- afl_model_data |>
select(inside_50s_diff, clearances_diff, contested_possessions_diff)
# Check - Correlations between predictors
ggpairs(afl_cor)
The correlations range from approximately 0.31 to 0.46. All three relationships are positive, meaning a higher difference in one statistic tends to occur alongside a higher difference in another. None of the pairs shows a very strong correlation.
Logistic regression models the odds of an outcome. Odds compare the chance of winning with the chance of losing. For example, odds of 2 correspond to a 67% probability of winning.
The model estimates its coefficients on the log-odds scale, which is hard to read. We convert them to odds ratios (by exponentiating), which tell us how the odds change for a one-unit increase in a predictor, holding the other predictors constant.
In our data, a one-unit increase means one extra inside 50 (or clearance, or contested possession) for the home team compared with its opponent.
We fit two models so we can see whether adding more statistics improves on a simple model:
# Model 1: inside-50 difference only
model_1 <- glm(home_win ~ inside_50s_diff,
family = binomial,
data = afl_model_data)
# Model 2: all three statistical differences
model_2 <- glm(home_win ~ inside_50s_diff + clearances_diff + contested_possessions_diff,
family = binomial,
data = afl_model_data)
tab_model() displays the odds ratios, their 95%
confidence intervals (CI) and p-values.
# Check - Model tables
tab_model(model_1,
model_2,
dv.labels = c("Inside 50s only", "All three statistics"),
show.aic = TRUE)
| Â | Inside 50s only | All three statistics | ||||
|---|---|---|---|---|---|---|
| Predictors | Odds Ratios | CI | p | Odds Ratios | CI | p |
| (Intercept) | 1.10 | 0.80 – 1.52 | 0.550 | 1.07 | 0.76 – 1.49 | 0.699 |
| inside 50s diff | 1.09 | 1.06 – 1.13 | <0.001 | 1.07 | 1.04 – 1.11 | <0.001 |
| clearances diff | 1.03 | 0.99 – 1.07 | 0.218 | |||
|
contested possessions diff |
1.04 | 1.02 – 1.08 | 0.003 | |||
| Observations | 206 | 206 | ||||
| R2 Tjur | 0.223 | 0.284 | ||||
| AIC | 232.537 | 221.123 | ||||
In Model 1, a one-unit increase in the inside-50 difference is associated with approximately 9% higher odds of a home-team win.
In Model 2, after accounting for the other statistics, a one-unit increase in the inside-50 difference is associated with approximately 7% higher odds of winning. A one-unit increase in the contested-possession difference is associated with approximately 4% higher odds.
The clearance difference has a p-value of 0.218, so this model does not provide clear evidence of an additional association with winning after accounting for inside 50s and contested possessions. This does not mean clearances are unimportant in football.
These percentages describe changes in the odds, rather than percentage-point changes in the probability of winning.
The AIC balances model fit against complexity. A lower AIC indicates a better model.
# Check - Model comparison
AIC(model_1, model_2)
## df AIC
## model_1 2 232.5367
## model_2 4 221.1226
Model 2 has a lower AIC than Model 1: 221.12 compared with 232.54. This favours the model containing all three statistics, because its improved fit outweighs the extra complexity.
# Check - Model diagnostics
par(mfrow = c(2, 2))
plot(model_2)
The diagnostic plots help identify matches that the model fits less
well. The labelled points show observations that may warrant closer
inspection. These labels refer to row numbers in the modelling dataset
and do not automatically mean those matches should be removed.
# Check - Collinearity
performance::check_collinearity(model_2)
## # Check for Multicollinearity
##
## Low Correlation
##
## Term VIF VIF 95% CI adj. VIF Tolerance
## inside_50s_diff 1.04 [1.00, 2.62] 1.02 0.96
## clearances_diff 1.10 [1.02, 1.49] 1.05 0.91
## contested_possessions_diff 1.13 [1.04, 1.47] 1.06 0.88
## Tolerance 95% CI
## [0.38, 1.00]
## [0.67, 0.98]
## [0.68, 0.96]
The VIF values range from 1.04 to 1.13 and are classified as low by the check. This suggests that collinearity is not a major concern for these predictors in this model.
# Check - Model report
report::report(model_2)
## We fitted a logistic model (estimated using ML) to predict home_win with
## inside_50s_diff, clearances_diff and contested_possessions_diff (formula:
## home_win ~ inside_50s_diff + clearances_diff + contested_possessions_diff). The
## model's explanatory power is substantial (Tjur's R2 = 0.28). The model's
## intercept, corresponding to inside_50s_diff = 0, clearances_diff = 0 and
## contested_possessions_diff = 0, is at 0.07 (95% CI [-0.27, 0.40], p = 0.699).
## Within this model:
##
## - The effect of inside 50s diff is statistically significant and positive (beta
## = 0.07, 95% CI [0.04, 0.10], p < .001; Std. beta = 0.93, 95% CI [0.54, 1.35])
## - The effect of clearances diff is statistically non-significant and positive
## (beta = 0.03, 95% CI [-0.01, 0.07], p = 0.218; Std. beta = 0.23, 95% CI [-0.14,
## 0.61])
## - The effect of contested possessions diff is statistically significant and
## positive (beta = 0.04, 95% CI [0.02, 0.07], p = 0.003; Std. beta = 0.65, 95% CI
## [0.23, 1.09])
##
## Standardized parameters were obtained by fitting the model on a standardized
## version of the dataset. 95% Confidence Intervals (CIs) and p-values were
## computed using a Wald z-distribution approximation.
This tutorial demonstrated how to download AFL data, calculate team statistics, compare opponents and fit logistic regression models in R.
In the 2026 regular season, larger inside-50 and contested-possession differences were associated with higher odds of a home-team win after accounting for the other statistics. The three-statistic model had a lower AIC than the inside-50-only model.
These findings describe associations within this season. They do not establish that the statistics cause wins or demonstrate how accurately the model would predict future matches. Because the statistics were collected during matches, the analysis helps us understand completed match outcomes.