Soccer matches produce a range of statistics that can help us understand team performance. But which statistics are most strongly associated with actually winning a match?
We will use logistic regression to investigate whether shots and shots on target are associated with the probability of a team winning.
First, we load the packages required for the analysis.
library(tidyverse)
matches <- read_csv("soccer_matches.csv")
For this analysis, we will examine the home team’s performance.
We create a new variable called home_win.
A value of 1 means the home team won, while 0 means the home team drew or lost.
matches <- matches |>
mutate(
home_win = if_else(FTHG > FTAG, 1, 0)
)
We will also keep the main variables required for the analysis.
match_data <- matches |>
select(
HomeTeam,
AwayTeam,
home_win,
HS,
HST
) |>
drop_na()
The variables represent:
HS = Home team shotsHST = Home team shots on targethome_win = Whether the home team wonBefore building a model, we can compare the average statistics for wins and non-wins.
match_data |>
group_by(home_win) |>
summarise(
average_shots = mean(HS),
average_shots_on_target = mean(HST)
)
## # A tibble: 2 × 3
## home_win average_shots average_shots_on_target
## <dbl> <dbl> <dbl>
## 1 0 13.3 3.74
## 2 1 14.5 5.53
This provides an initial indication of whether winning teams tend to produce more shots and shots on target.
The boxplot allows us to visually compare shots on target between matches where the home team won and matches where they did not.
Because our outcome has two possibilities (win or not win), logistic regression is appropriate.
win_model <- glm(
home_win ~ HS + HST,
data = match_data,
family = binomial
)
View the results:
summary(win_model)
##
## Call:
## glm(formula = home_win ~ HS + HST, family = binomial, data = match_data)
##
## Coefficients:
## Estimate Std. Error z value Pr(>|z|)
## (Intercept) -1.48530 0.35198 -4.220 2.44e-05 ***
## HS -0.09438 0.03080 -3.065 0.00218 **
## HST 0.54379 0.07506 7.245 4.32e-13 ***
## ---
## Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
##
## (Dispersion parameter for binomial family taken to be 1)
##
## Null deviance: 518.51 on 379 degrees of freedom
## Residual deviance: 445.48 on 377 degrees of freedom
## AIC: 451.48
##
## Number of Fisher Scoring iterations: 3
Positive coefficients indicate that increasing the variable is associated with a higher probability of winning.
We can use our model to estimate the probability of winning each match.
match_data <- match_data |>
mutate(
win_probability = predict(
win_model,
type = "response"
)
)
Look at the predictions:
head(match_data)
## # A tibble: 6 × 6
## HomeTeam AwayTeam home_win HS HST win_probability
## <chr> <chr> <dbl> <dbl> <dbl> <dbl>
## 1 Liverpool Bournemouth 1 19 10 0.897
## 2 Aston Villa Newcastle 0 3 3 0.466
## 3 Brighton Fulham 0 10 4 0.437
## 4 Sunderland West Ham 1 10 5 0.572
## 5 Tottenham Burnley 1 16 6 0.566
## 6 Wolves Man City 0 9 3 0.331
The win_probability variable ranges between 0 and 1.
For example:
We can classify probabilities above 0.50 as a predicted win.
match_data <- match_data |>
mutate(
predicted_win = if_else(
win_probability >= 0.5,
1,
0
)
)
We can compare the model predictions with the actual results.
mean(
match_data$predicted_win ==
match_data$home_win
)
## [1] 0.6894737
This gives the proportion of matches correctly classified by the model.
We can also create a confusion matrix.
table(
Actual = match_data$home_win,
Predicted = match_data$predicted_win
)
## Predicted
## Actual 0 1
## 0 172 46
## 1 72 90
This analysis demonstrated how logistic regression can be used to investigate soccer match outcomes.
The model examined whether home-team shots and shots on target were associated with the probability of winning.
An important limitation is that shots and shots on target occur during the match. Therefore, this model explains which match statistics are associated with winning rather than providing a true pre-match prediction model.