Expected Goals (xG) is a commonly used metric in soccer analytics that estimates the probability of a shot resulting in a goal.
An xG value ranges between 0 and 1. For example, an xG value of 0.10 represents approximately a 10% probability of scoring, while an xG value of 0.70 represents approximately a 70% probability.
We will build a simple expected goals model using logistic regression.
We will investigate how two important factors influence the probability of scoring:
First, we load the tidyverse package which provides
tools for importing, manipulating and visualising data.
library(tidyverse)
The dataset contains individual shots from soccer matches.
shots <- read_csv("XG_final_dataset.csv")
For our simple xG model, we only require three variables:
is_goal = whether the shot resulted in a goaldistance = distance of the shot from goalangle = angle of the shotxg_data <- shots |>
select(
is_goal,
distance,
angle
) |>
drop_na()
We can check how many shots remain in our dataset.
nrow(xg_data)
## [1] 64898
Before creating the model, we can calculate how many shots resulted in goals.
xg_data |>
summarise(
shots = n(),
goals = sum(is_goal),
goal_percentage = mean(is_goal) * 100
)
## # A tibble: 1 × 3
## shots goals goal_percentage
## <int> <dbl> <dbl>
## 1 64898 7144 11.0
This shows the overall percentage of shots that resulted in goals.
We can compare the average shot distance for goals and non-goals.
xg_data |>
group_by(is_goal) |>
summarise(
average_distance = mean(distance)
)
## # A tibble: 2 × 2
## is_goal average_distance
## <dbl> <dbl>
## 1 0 20.0
## 2 1 12.9
We can also visualise this relationship using a boxplot.
If goals generally occur closer to goal, we would expect the goal group to have a lower average shot distance.
A shot has two possible outcomes: goal or no goal.
Because the outcome is binary, logistic regression can be used to estimate the probability of a shot resulting in a goal.
xg_model <- glm(
is_goal ~ distance + angle,
data = xg_data,
family = binomial
)
We can view the model results.
summary(xg_model)
##
## Call:
## glm(formula = is_goal ~ distance + angle, family = binomial,
## data = xg_data)
##
## Coefficients:
## Estimate Std. Error z value Pr(>|z|)
## (Intercept) -0.0488736 0.0452816 -1.079 0.280
## distance -0.1282515 0.0020707 -61.936 <2e-16 ***
## angle 0.0002432 0.0003682 0.660 0.509
## ---
## Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
##
## (Dispersion parameter for binomial family taken to be 1)
##
## Null deviance: 44998 on 64897 degrees of freedom
## Residual deviance: 39936 on 64895 degrees of freedom
## AIC: 39942
##
## Number of Fisher Scoring iterations: 6
The coefficients show how distance and angle are associated with the probability of scoring.
We can now use our logistic regression model to calculate a scoring probability for every shot.
xg_data <- xg_data |>
mutate(
xG = predict(
xg_model,
type = "response"
)
)
View the first few predicted values.
head(xg_data)
## # A tibble: 6 × 4
## is_goal distance angle xG
## <dbl> <dbl> <dbl> <dbl>
## 1 0 20.2 76.0 0.0678
## 2 0 8.45 39.7 0.245
## 3 0 21.0 139. 0.0626
## 4 0 9.59 141. 0.224
## 5 0 30.9 94.6 0.0182
## 6 0 12.3 52.9 0.166
The new xG variable represents our model’s estimated
probability that each shot will result in a goal.
For example:
We can visualise how shot distance is related to our model’s expected goals value.
## `geom_smooth()` using method = 'gam' and formula = 'y ~ s(x, bs = "cs")'
Generally, we would expect shots taken closer to goal to have a higher probability of scoring.
We can compare the total number of actual goals with the total xG predicted by our model.
xg_data |>
summarise(
actual_goals = sum(is_goal),
expected_goals = sum(xG)
)
## # A tibble: 1 × 2
## actual_goals expected_goals
## <dbl> <dbl>
## 1 7144 7144.
This allows us to see how the total number of goals predicted by the model compares with the actual number scored.
Our original dataset also contains StatsBomb’s xG estimate.
We can add this back to our data and compare the two models.
xg_comparison <- shots |>
select(
is_goal,
distance,
angle,
shot_statsbomb_xg
) |>
drop_na()
xg_comparison$our_xG <- predict(
xg_model,
newdata = xg_comparison,
type = "response"
)
Now calculate the average xG from each model.
xg_comparison |>
summarise(
our_average_xG = mean(our_xG),
statsbomb_average_xG = mean(shot_statsbomb_xg)
)
## # A tibble: 1 × 2
## our_average_xG statsbomb_average_xG
## <dbl> <dbl>
## 1 0.110 0.106
This provides a simple comparison between our two-variable model and a professional xG model.
Our model only considers distance and angle, while professional xG models can consider many more characteristics of each shot.
This tutorial demonstrated how logistic regression can be used to build a simple Expected Goals model in R.
The model estimated the probability of a shot resulting in a goal using two variables: shot distance and shot angle.
The analysis demonstrates the basic concept behind xG: not every shot has the same probability of becoming a goal. Shots from more favourable locations generally have a greater chance of scoring.
Our model is deliberately simple and is designed to demonstrate the concept of expected goals rather than replicate a professional xG model.
More advanced models can consider additional information such as body part, defensive pressure, shot technique, whether the shot was a one-on-one and the type of attacking situation.