Introduction

Soccer matches produce a range of statistics that can help us understand team performance. But which statistics are most strongly associated with actually winning a match?

We will use logistic regression to investigate whether shots and shots on target are associated with the probability of a team winning.

1. Load Packages

First, we load the packages required for the analysis.

library(tidyverse)

2. Import the Data

matches <- read_csv("soccer_matches.csv")

3. Prepare the Data

For this analysis, we will examine the home team’s performance.

We create a new variable called home_win.

A value of 1 means the home team won, while 0 means the home team drew or lost.

matches <- matches |>
  mutate(
    home_win = if_else(FTHG > FTAG, 1, 0)
  )

We will also keep the main variables required for the analysis.

match_data <- matches |>
  select(
    HomeTeam,
    AwayTeam,
    home_win,
    HS,
    HST
  ) |>
  drop_na()

The variables represent:

4. Explore the Data

Before building a model, we can compare the average statistics for wins and non-wins.

match_data |>
  group_by(home_win) |>
  summarise(
    average_shots = mean(HS),
    average_shots_on_target = mean(HST)
  )
## # A tibble: 2 × 3
##   home_win average_shots average_shots_on_target
##      <dbl>         <dbl>                   <dbl>
## 1        0          13.3                    3.74
## 2        1          14.5                    5.53

This provides an initial indication of whether winning teams tend to produce more shots and shots on target.

5. Visualise Shots on Target

The boxplot allows us to visually compare shots on target between matches where the home team won and matches where they did not.

6. Build a Logistic Regression Model

Because our outcome has two possibilities (win or not win), logistic regression is appropriate.

win_model <- glm(
  home_win ~ HS + HST,
  data = match_data,
  family = binomial
)

View the results:

summary(win_model)
## 
## Call:
## glm(formula = home_win ~ HS + HST, family = binomial, data = match_data)
## 
## Coefficients:
##             Estimate Std. Error z value Pr(>|z|)    
## (Intercept) -1.48530    0.35198  -4.220 2.44e-05 ***
## HS          -0.09438    0.03080  -3.065  0.00218 ** 
## HST          0.54379    0.07506   7.245 4.32e-13 ***
## ---
## Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
## 
## (Dispersion parameter for binomial family taken to be 1)
## 
##     Null deviance: 518.51  on 379  degrees of freedom
## Residual deviance: 445.48  on 377  degrees of freedom
## AIC: 451.48
## 
## Number of Fisher Scoring iterations: 3

Positive coefficients indicate that increasing the variable is associated with a higher probability of winning.

7. Generate Predicted Probabilities

We can use our model to estimate the probability of winning each match.

match_data <- match_data |>
  mutate(
    win_probability = predict(
      win_model,
      type = "response"
    )
  )

Look at the predictions:

head(match_data)
## # A tibble: 6 × 6
##   HomeTeam    AwayTeam    home_win    HS   HST win_probability
##   <chr>       <chr>          <dbl> <dbl> <dbl>           <dbl>
## 1 Liverpool   Bournemouth        1    19    10           0.897
## 2 Aston Villa Newcastle          0     3     3           0.466
## 3 Brighton    Fulham             0    10     4           0.437
## 4 Sunderland  West Ham           1    10     5           0.572
## 5 Tottenham   Burnley            1    16     6           0.566
## 6 Wolves      Man City           0     9     3           0.331

The win_probability variable ranges between 0 and 1.

For example:

8. Convert Probabilities Into Predictions

We can classify probabilities above 0.50 as a predicted win.

match_data <- match_data |>
  mutate(
    predicted_win = if_else(
      win_probability >= 0.5,
      1,
      0
    )
  )

9. Model Accuracy

We can compare the model predictions with the actual results.

mean(
  match_data$predicted_win ==
    match_data$home_win
)
## [1] 0.6894737

This gives the proportion of matches correctly classified by the model.

We can also create a confusion matrix.

table(
  Actual = match_data$home_win,
  Predicted = match_data$predicted_win
)
##       Predicted
## Actual   0   1
##      0 172  46
##      1  72  90

10. Conclusion

This analysis demonstrated how logistic regression can be used to investigate soccer match outcomes.

The model examined whether home-team shots and shots on target were associated with the probability of winning.

An important limitation is that shots and shots on target occur during the match. Therefore, this model explains which match statistics are associated with winning rather than providing a true pre-match prediction model.