1 Introduction

Which match statistics are associated with winning in the AFL?

AFL match statistics can help us explore how team performance relates to winning and losing. This tutorial uses data from the 2026 AFL season to examine whether differences in inside 50s, clearances and contested possessions are associated with match outcomes.

We will use logistic regression, a statistical method for modelling an outcome with two categories: a home-team win or loss. It allows us to examine each statistic’s association with winning while accounting for the other statistics included in the model.

The tutorial will guide you through downloading data using the fitzRoy package, calculating team totals, comparing opponents, creating exploratory graphs, and fitting and interpreting logistic regression models in R. Because the statistics are recorded during matches, we will explore their association with the result rather than forecast results before a match starts.

2 Install Packages

install.packages(c("fitzRoy",     
                   "tidyverse",  
                   "janitor",
                   "GGally",      
                   "sjPlot",      
                   "performance", 
                   "report"))     

3 Load Packages

library(fitzRoy)
## Warning: package 'fitzRoy' was built under R version 4.6.1
library(tidyverse)
library(janitor)
library(GGally)
library(sjPlot)

4 Import and Check Data

For this task, we need two datasets from AFL Tables:

  • player_stats: One row per player per match (kicks, tackles, clearances, etc.)
  • match_results: One row per match (teams, scores, round, etc.)
# Import - Player statistics for the 2026 season
player_stats <- fetch_player_stats_afltables(season = 2026) |>
  clean_names()

# Import - Match results for the 2026 season
match_results <- fetch_results_afltables(season = 2026) |>
  clean_names()

Use glimpse() to check the structure of each dataset.

# Check - Data structure
match_results |>
  select(date, home_team, away_team,
         home_points, away_points, round_type) |>
  glimpse()
## Rows: 218
## Columns: 6
## $ date        <date> 2026-03-05, 2026-03-06, 2026-03-07, 2026-03-07, 2026-03-0…
## $ home_team   <chr> "Sydney", "Gold Coast", "GWS", "Brisbane Lions", "St Kilda…
## $ away_team   <chr> "Carlton", "Geelong", "Hawthorn", "Footscray", "Collingwoo…
## $ home_points <int> 132, 125, 122, 106, 66, 75, 83, 134, 110, 104, 79, 113, 12…
## $ away_points <int> 69, 69, 95, 111, 78, 71, 145, 53, 100, 60, 93, 67, 107, 72…
## $ round_type  <chr> "Regular", "Regular", "Regular", "Regular", "Regular", "Re…
player_stats |>
  select(date, playing_for, inside_50s,
         clearances, contested_possessions) |>
  glimpse()
## Rows: 10,028
## Columns: 5
## $ date                  <date> 2026-03-05, 2026-03-05, 2026-03-05, 2026-03-05,…
## $ playing_for           <chr> "Sydney", "Sydney", "Sydney", "Sydney", "Sydney"…
## $ inside_50s            <int> 2, 1, 5, 1, 3, 8, 5, 1, 2, 0, 1, 6, 2, 2, 1, 0, …
## $ clearances            <int> 0, 0, 0, 0, 3, 5, 6, 1, 0, 1, 2, 3, 0, 0, 0, 1, …
## $ contested_possessions <int> 3, 6, 5, 4, 6, 8, 10, 5, 5, 4, 4, 8, 3, 5, 4, 9,…

5 Prepare the Data

We want to create a dataset with one row per match, with the following columns:

  • home_win: The outcome: 1 if the home team won, 0 if it lost
  • inside_50s_diff: Home team inside 50s minus away team inside 50s
  • clearances_diff: Home team clearances minus away team clearances
  • contested_possessions_diff: Home team contested possessions minus away team contested possessions

5.1 Add up player statistics for each team

We add up every player’s statistics to get a total for each team in each match.

# Add together player statistics for each team in each match
team_stats <- player_stats |>
  group_by(season, date, home_team, away_team, playing_for) |>
  summarise(inside_50s = sum(inside_50s),
            clearances = sum(clearances),
            contested_possessions = sum(contested_possessions),
            .groups = "drop")

# Check - New dataset
head(team_stats)
## # A tibble: 6 × 8
##   season date       home_team      away_team   playing_for inside_50s clearances
##    <int> <date>     <chr>          <chr>       <chr>            <int>      <int>
## 1   2026 2026-03-05 Sydney         Carlton     Carlton             50         39
## 2   2026 2026-03-05 Sydney         Carlton     Sydney              67         40
## 3   2026 2026-03-06 Gold Coast     Geelong     Geelong             50         41
## 4   2026 2026-03-06 Gold Coast     Geelong     Gold Coast          71         34
## 5   2026 2026-03-07 Brisbane Lions Western Bu… Brisbane L…         63         45
## 6   2026 2026-03-07 Brisbane Lions Western Bu… Western Bu…         50         39
## # ℹ 1 more variable: contested_possessions <int>

Each match should now have two rows (one per team) and no missing values. Check to confirm this:

# Check - Missing team totals
team_stats |>
  summarise(missing_inside_50s = sum(is.na(inside_50s)),
            missing_clearances = sum(is.na(clearances)),
            missing_contested_possessions = sum(is.na(contested_possessions)))
## # A tibble: 1 × 3
##   missing_inside_50s missing_clearances missing_contested_possessions
##                <int>              <int>                         <int>
## 1                  0                  0                             0
# Check - Each match has two teams
team_stats |>
  count(season, date, home_team, away_team, name = "number_of_teams") |>
  count(number_of_teams, name = "number_of_matches")
## # A tibble: 1 × 2
##   number_of_teams number_of_matches
##             <int>             <int>
## 1               2               218

The first table should show zeros. The second should show that every match has number_of_teams equal to 2.

5.2 Check match types

This tutorial focuses on home-and-away matches from the 2026 AFL season. Finals are excluded to keep the analysis focused on the regular season. Let’s check how the match types are labelled:

# Check - Regular season and finals labels
match_results |>
  count(round_type)
## # A tibble: 2 × 2
##   round_type     n
##   <chr>      <int>
## 1 Finals         9
## 2 Regular      209

5.3 Standardise club names

The two datasets spell some club names differently (e.g Footscray and Western Bulldogs). If we don’t fix this, the datasets will not match up when we combine them. case_when() replaces the names we list and leaves every other name unchanged (TRUE ~ .x).

# Standardise club names in the team statistics
team_stats <- team_stats |>
  mutate(
    across(
      c(home_team, away_team, playing_for),
      ~ case_when(.x == "GWS" ~ "Greater Western Sydney",
                  .x == "Footscray" ~ "Western Bulldogs",
                  TRUE ~ .x)))

# Standardise club names in the match results
match_results <- match_results |>
  mutate(
    across(
      c(home_team, away_team),
      ~ case_when(.x == "GWS" ~ "Greater Western Sydney",
                  .x == "Footscray" ~ "Western Bulldogs",
                  TRUE ~ .x)))

5.4 Combine both teams into one match per row

# Ensure dates use the same format in both datasets
team_stats <- team_stats |>
  mutate(date = as.Date(date))

match_results <- match_results |>
  mutate(date = as.Date(date))

# Select home-team statistics and give them clear column names
home_stats <- team_stats |>
  filter(playing_for == home_team) |>
  select(season, date, home_team, away_team,
         home_inside_50s = inside_50s,
         home_clearances = clearances,
         home_contested_possessions = contested_possessions)

# Select away-team statistics
away_stats <- team_stats |>
  filter(playing_for == away_team) |>
  select(season, date, home_team, away_team,
         away_inside_50s = inside_50s,
         away_clearances = clearances,
         away_contested_possessions = contested_possessions)

# Combine results with both teams' statistics
afl_matches <- match_results |>
  filter(round_type == "Regular") |>
  select(season, date, round_number,
         home_team, away_team, home_points, away_points) |>
  left_join(home_stats,
            by = c("season", "date", "home_team", "away_team")) |>
  left_join(away_stats,
            by = c("season", "date", "home_team", "away_team"))

5.5 Create the outcome and predictors

Now we calculate the three differences and record whether the home team won, lost or drew.

afl_matches <- afl_matches |>
  mutate(inside_50s_diff = home_inside_50s - away_inside_50s,
         clearances_diff = home_clearances - away_clearances,
         contested_possessions_diff = home_contested_possessions - away_contested_possessions,
         home_result = case_when(
           home_points > away_points ~ "Win",
           home_points < away_points ~ "Loss",
           home_points == away_points ~ "Draw"))

# Check - Count the outcomes before excluding draws
afl_matches |>
  count(home_result)
## # A tibble: 3 × 2
##   home_result     n
##   <chr>       <int>
## 1 Draw            3
## 2 Loss           84
## 3 Win           122

A logistic regression needs exactly two outcomes. Since draws are rare, we will remove them and code the outcome as home_win (1 = win, 0 = loss).

# Prepare - Modelling dataset
afl_model_data <- afl_matches |>
  filter(home_result != "Draw") |>
  mutate(home_win = as.integer(home_result == "Win")) |>
  select(season, date, round_number, home_team, away_team,
         home_win, inside_50s_diff, clearances_diff,
         contested_possessions_diff)

# Check - Missing values after joining the datasets
afl_model_data |>
  summarise(
    across(
      c(home_win, inside_50s_diff, clearances_diff, contested_possessions_diff),
      ~ sum(is.na(.x))))
## # A tibble: 1 × 4
##   home_win inside_50s_diff clearances_diff contested_possessions_diff
##      <int>           <int>           <int>                      <int>
## 1        0               0               0                          0
# Check - Compact preview
afl_model_data |>
  select(home_team, away_team, home_win, inside_50s_diff,
         clearances_diff, contested_possessions_diff) |>
  head()
## # A tibble: 6 × 6
##   home_team              away_team      home_win inside_50s_diff clearances_diff
##   <chr>                  <chr>             <int>           <int>           <int>
## 1 Sydney                 Carlton               1              17               1
## 2 Gold Coast             Geelong               1              21              -7
## 3 Greater Western Sydney Hawthorn              1               5               7
## 4 Brisbane Lions         Western Bulld…        0              13               6
## 5 St Kilda               Collingwood           0              23               2
## 6 Carlton                Richmond              1             -17               8
## # ℹ 1 more variable: contested_possessions_diff <int>

The dataset contains 209 regular-season matches. Removing the three draws leaves 206 matches: 122 home-team wins and 84 home-team losses. There are no missing values in the outcome or the three predictors.

6 Exploratory Data Analysis

6.1 Descriptive statistics

summary() gives the minimum, quartiles, median, mean and maximum.

# Descriptive statistics
afl_model_data |>
  select(inside_50s_diff,
         clearances_diff,
         contested_possessions_diff) |>
  summary()
##  inside_50s_diff   clearances_diff   contested_possessions_diff
##  Min.   :-29.000   Min.   :-22.000   Min.   :-44.000           
##  1st Qu.: -6.000   1st Qu.: -5.000   1st Qu.: -6.000           
##  Median :  4.500   Median :  1.000   Median :  4.000           
##  Mean   :  4.049   Mean   :  1.063   Mean   :  3.092           
##  3rd Qu.: 15.000   3rd Qu.:  8.000   3rd Qu.: 13.000           
##  Max.   : 32.000   Max.   : 24.000   Max.   : 44.000

Across these matches, home teams averaged approximately 4.05 more inside 50s, 1.06 more clearances and 3.09 more contested possessions than their opponents. Positive differences mean the home team recorded more of that statistic; negative differences mean the away team recorded more.

6.2 Statistics by match result

We label home-team outcomes as ‘Win’ or ‘Loss’ to make the tables and graphs easier to read, then compare the average statistical differences. We keep home_win coded as 0/1 for the logistic regression.

# Add readable labels while keeping home_win for modelling
afl_model_data <- afl_model_data |>
  mutate(
    result = factor(home_win,
                    levels = c(0, 1),
                    labels = c("Loss", "Win")))

# Compare average statistical differences by outcome
afl_model_data |>
  group_by(result) |>
  summarise(matches = n(),
            mean_inside_50s_diff = mean(inside_50s_diff),
            mean_clearances_diff = mean(clearances_diff),
            mean_contested_possessions_diff = mean(contested_possessions_diff),
            .groups = "drop")
## # A tibble: 2 × 5
##   result matches mean_inside_50s_diff mean_clearances_diff
##   <fct>    <int>                <dbl>                <dbl>
## 1 Loss        84                -3.31                -2.06
## 2 Win        122                 9.11                 3.21
## # ℹ 1 more variable: mean_contested_possessions_diff <dbl>

Home teams that won averaged 9.11 more inside 50s and 3.21 more clearances than their opponents. Home teams that lost averaged 3.31 fewer inside 50s and 2.06 fewer clearances. These comparisons suggest that the statistical differences are related to match outcomes, but they do not account for the other statistics.

6.3 Correlations between predictors

We can check how closely the three statistics are related. Strong relationships can make it harder to identify each statistic’s separate association with winning.

ggpairs() displays graphs and correlation values to help us explore these relationships.

# Select - Predictors only
afl_cor <- afl_model_data |>
  select(inside_50s_diff, clearances_diff, contested_possessions_diff)

# Check - Correlations between predictors
ggpairs(afl_cor)

The correlations range from approximately 0.31 to 0.46. All three relationships are positive, meaning a higher difference in one statistic tends to occur alongside a higher difference in another. None of the pairs shows a very strong correlation.

7 Understanding Odds Ratios

Logistic regression models the odds of an outcome. Odds compare the chance of winning with the chance of losing. For example, odds of 2 correspond to a 67% probability of winning.

The model estimates its coefficients on the log-odds scale, which is hard to read. We convert them to odds ratios (by exponentiating), which tell us how the odds change for a one-unit increase in a predictor, holding the other predictors constant.

  • An odds ratio of 1 means no change in the odds.
  • An odds ratio greater than 1 means the odds of a win increase. For example, 1.30 means the odds are 30% higher.
  • An odds ratio less than 1 means the odds of a win decrease. For example, 0.75 means the odds are 25% lower.

In our data, a one-unit increase means one extra inside 50 (or clearance, or contested possession) for the home team compared with its opponent.

8 Logistic Regression

We fit two models so we can see whether adding more statistics improves on a simple model:

  • Model 1: Inside-50 difference only
  • Model 2: Inside-50, clearance and contested-possession differences
# Model 1: inside-50 difference only
model_1 <- glm(home_win ~ inside_50s_diff,
               family = binomial,
               data = afl_model_data)

# Model 2: all three statistical differences
model_2 <- glm(home_win ~ inside_50s_diff + clearances_diff + contested_possessions_diff,
               family = binomial,
               data = afl_model_data)

tab_model() displays the odds ratios, their 95% confidence intervals (CI) and p-values.

# Check - Model tables
tab_model(model_1,
          model_2,
          dv.labels = c("Inside 50s only", "All three statistics"),
          show.aic = TRUE)
  Inside 50s only All three statistics
Predictors Odds Ratios CI p Odds Ratios CI p
(Intercept) 1.10 0.80 – 1.52 0.550 1.07 0.76 – 1.49 0.699
inside 50s diff 1.09 1.06 – 1.13 <0.001 1.07 1.04 – 1.11 <0.001
clearances diff 1.03 0.99 – 1.07 0.218
contested possessions
diff
1.04 1.02 – 1.08 0.003
Observations 206 206
R2 Tjur 0.223 0.284
AIC 232.537 221.123

In Model 1, a one-unit increase in the inside-50 difference is associated with approximately 9% higher odds of a home-team win.

In Model 2, after accounting for the other statistics, a one-unit increase in the inside-50 difference is associated with approximately 7% higher odds of winning. A one-unit increase in the contested-possession difference is associated with approximately 4% higher odds.

The clearance difference has a p-value of 0.218, so this model does not provide clear evidence of an additional association with winning after accounting for inside 50s and contested possessions. This does not mean clearances are unimportant in football.

These percentages describe changes in the odds, rather than percentage-point changes in the probability of winning.

9 Compare the Models

The AIC balances model fit against complexity. A lower AIC indicates a better model.

# Check - Model comparison
AIC(model_1, model_2)
##         df      AIC
## model_1  2 232.5367
## model_2  4 221.1226

Model 2 has a lower AIC than Model 1: 221.12 compared with 232.54. This favours the model containing all three statistics, because its improved fit outweighs the extra complexity.

10 Check the Model

10.1 Diagnostic plots

# Check - Model diagnostics
par(mfrow = c(2, 2))
plot(model_2)

The diagnostic plots help identify matches that the model fits less well. The labelled points show observations that may warrant closer inspection. These labels refer to row numbers in the modelling dataset and do not automatically mean those matches should be removed.

10.2 Collinearity

# Check - Collinearity
performance::check_collinearity(model_2)
## # Check for Multicollinearity
## 
## Low Correlation
## 
##                        Term  VIF   VIF 95% CI adj. VIF Tolerance
##             inside_50s_diff 1.04 [1.00, 2.62]     1.02      0.96
##             clearances_diff 1.10 [1.02, 1.49]     1.05      0.91
##  contested_possessions_diff 1.13 [1.04, 1.47]     1.06      0.88
##  Tolerance 95% CI
##      [0.38, 1.00]
##      [0.67, 0.98]
##      [0.68, 0.96]

The VIF values range from 1.04 to 1.13 and are classified as low by the check. This suggests that collinearity is not a major concern for these predictors in this model.

11 Report the Model

# Check - Model report
report::report(model_2)
## We fitted a logistic model (estimated using ML) to predict home_win with
## inside_50s_diff, clearances_diff and contested_possessions_diff (formula:
## home_win ~ inside_50s_diff + clearances_diff + contested_possessions_diff). The
## model's explanatory power is substantial (Tjur's R2 = 0.28). The model's
## intercept, corresponding to inside_50s_diff = 0, clearances_diff = 0 and
## contested_possessions_diff = 0, is at 0.07 (95% CI [-0.27, 0.40], p = 0.699).
## Within this model:
## 
##   - The effect of inside 50s diff is statistically significant and positive (beta
## = 0.07, 95% CI [0.04, 0.10], p < .001; Std. beta = 0.93, 95% CI [0.54, 1.35])
##   - The effect of clearances diff is statistically non-significant and positive
## (beta = 0.03, 95% CI [-0.01, 0.07], p = 0.218; Std. beta = 0.23, 95% CI [-0.14,
## 0.61])
##   - The effect of contested possessions diff is statistically significant and
## positive (beta = 0.04, 95% CI [0.02, 0.07], p = 0.003; Std. beta = 0.65, 95% CI
## [0.23, 1.09])
## 
## Standardized parameters were obtained by fitting the model on a standardized
## version of the dataset. 95% Confidence Intervals (CIs) and p-values were
## computed using a Wald z-distribution approximation.

12 Conclusion

This tutorial demonstrated how to download AFL data, calculate team statistics, compare opponents and fit logistic regression models in R.

In the 2026 regular season, larger inside-50 and contested-possession differences were associated with higher odds of a home-team win after accounting for the other statistics. The three-statistic model had a lower AIC than the inside-50-only model.

These findings describe associations within this season. They do not establish that the statistics cause wins or demonstrate how accurately the model would predict future matches. Because the statistics were collected during matches, the analysis helps us understand completed match outcomes.