Introduction

Expected Goals (xG) is a commonly used metric in soccer analytics that estimates the probability of a shot resulting in a goal.

An xG value ranges between 0 and 1. For example, an xG value of 0.10 represents approximately a 10% probability of scoring, while an xG value of 0.70 represents approximately a 70% probability.

We will build a simple expected goals model using logistic regression.

We will investigate how two important factors influence the probability of scoring:

1. Load Packages

First, we load the tidyverse package which provides tools for importing, manipulating and visualising data.

library(tidyverse)

2. Import the Data

The dataset contains individual shots from soccer matches.

shots <- read_csv("XG_final_dataset.csv")

3. Prepare the Data

For our simple xG model, we only require three variables:

xg_data <- shots |>
  select(
    is_goal,
    distance,
    angle
  ) |>
  drop_na()

We can check how many shots remain in our dataset.

nrow(xg_data)
## [1] 64898

4. Explore the Data

Before creating the model, we can calculate how many shots resulted in goals.

xg_data |>
  summarise(
    shots = n(),
    goals = sum(is_goal),
    goal_percentage = mean(is_goal) * 100
  )
## # A tibble: 1 × 3
##   shots goals goal_percentage
##   <int> <dbl>           <dbl>
## 1 64898  7144            11.0

This shows the overall percentage of shots that resulted in goals.

5. Does Shot Distance Matter?

We can compare the average shot distance for goals and non-goals.

xg_data |>
  group_by(is_goal) |>
  summarise(
    average_distance = mean(distance)
  )
## # A tibble: 2 × 2
##   is_goal average_distance
##     <dbl>            <dbl>
## 1       0             20.0
## 2       1             12.9

We can also visualise this relationship using a boxplot.

If goals generally occur closer to goal, we would expect the goal group to have a lower average shot distance.

6. Build the Expected Goals Model

A shot has two possible outcomes: goal or no goal.

Because the outcome is binary, logistic regression can be used to estimate the probability of a shot resulting in a goal.

xg_model <- glm(
  is_goal ~ distance + angle,
  data = xg_data,
  family = binomial
)

We can view the model results.

summary(xg_model)
## 
## Call:
## glm(formula = is_goal ~ distance + angle, family = binomial, 
##     data = xg_data)
## 
## Coefficients:
##               Estimate Std. Error z value Pr(>|z|)    
## (Intercept) -0.0488736  0.0452816  -1.079    0.280    
## distance    -0.1282515  0.0020707 -61.936   <2e-16 ***
## angle        0.0002432  0.0003682   0.660    0.509    
## ---
## Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
## 
## (Dispersion parameter for binomial family taken to be 1)
## 
##     Null deviance: 44998  on 64897  degrees of freedom
## Residual deviance: 39936  on 64895  degrees of freedom
## AIC: 39942
## 
## Number of Fisher Scoring iterations: 6

The coefficients show how distance and angle are associated with the probability of scoring.

7. Calculate xG

We can now use our logistic regression model to calculate a scoring probability for every shot.

xg_data <- xg_data |>
  mutate(
    xG = predict(
      xg_model,
      type = "response"
    )
  )

View the first few predicted values.

head(xg_data)
## # A tibble: 6 × 4
##   is_goal distance angle     xG
##     <dbl>    <dbl> <dbl>  <dbl>
## 1       0    20.2   76.0 0.0678
## 2       0     8.45  39.7 0.245 
## 3       0    21.0  139.  0.0626
## 4       0     9.59 141.  0.224 
## 5       0    30.9   94.6 0.0182
## 6       0    12.3   52.9 0.166

The new xG variable represents our model’s estimated probability that each shot will result in a goal.

For example:

8. Visualise Distance and xG

We can visualise how shot distance is related to our model’s expected goals value.

## `geom_smooth()` using method = 'gam' and formula = 'y ~ s(x, bs = "cs")'

Generally, we would expect shots taken closer to goal to have a higher probability of scoring.

9. Compare Actual Goals With Expected Goals

We can compare the total number of actual goals with the total xG predicted by our model.

xg_data |>
  summarise(
    actual_goals = sum(is_goal),
    expected_goals = sum(xG)
  )
## # A tibble: 1 × 2
##   actual_goals expected_goals
##          <dbl>          <dbl>
## 1         7144          7144.

This allows us to see how the total number of goals predicted by the model compares with the actual number scored.

10. Compare Our Model With StatsBomb xG

Our original dataset also contains StatsBomb’s xG estimate.

We can add this back to our data and compare the two models.

xg_comparison <- shots |>
  select(
    is_goal,
    distance,
    angle,
    shot_statsbomb_xg
  ) |>
  drop_na()

xg_comparison$our_xG <- predict(
  xg_model,
  newdata = xg_comparison,
  type = "response"
)

Now calculate the average xG from each model.

xg_comparison |>
  summarise(
    our_average_xG = mean(our_xG),
    statsbomb_average_xG = mean(shot_statsbomb_xg)
  )
## # A tibble: 1 × 2
##   our_average_xG statsbomb_average_xG
##            <dbl>                <dbl>
## 1          0.110                0.106

This provides a simple comparison between our two-variable model and a professional xG model.

Our model only considers distance and angle, while professional xG models can consider many more characteristics of each shot.

Conclusion

This tutorial demonstrated how logistic regression can be used to build a simple Expected Goals model in R.

The model estimated the probability of a shot resulting in a goal using two variables: shot distance and shot angle.

The analysis demonstrates the basic concept behind xG: not every shot has the same probability of becoming a goal. Shots from more favourable locations generally have a greater chance of scoring.

Our model is deliberately simple and is designed to demonstrate the concept of expected goals rather than replicate a professional xG model.

More advanced models can consider additional information such as body part, defensive pressure, shot technique, whether the shot was a one-on-one and the type of attacking situation.