Logistic regression is a statistical method used when the outcome has two possible categories.
In this tutorial, we will use R to build a simple logistic regression model. The aim is to show how logistic regression can be fitted, interpreted, and used to calculate predicted probabilities.
This tutorial is designed for beginners. No previous experience with logistic regression is required.
Logistic regression can be used when the response variable has two possible outcomes.
For example, in sport analytics, an outcome could be:
Instead of predicting a continuous value, logistic regression estimates the probability of one of the two outcomes.
The probability is between 0 and 1, where:
# Load the required package
library(tidyverse)
## ── Attaching core tidyverse packages ──────────────────────── tidyverse 2.0.0 ──
## ✔ dplyr 1.2.1 ✔ readr 2.2.0
## ✔ forcats 1.0.1 ✔ stringr 1.6.0
## ✔ ggplot2 4.0.3 ✔ tibble 3.3.1
## ✔ lubridate 1.9.5 ✔ tidyr 1.3.2
## ✔ purrr 1.2.2
## ── Conflicts ────────────────────────────────────────── tidyverse_conflicts() ──
## ✖ dplyr::filter() masks stats::filter()
## ✖ dplyr::lag() masks stats::lag()
## ℹ Use the conflicted package (<http://conflicted.r-lib.org/>) to force all conflicts to become errors
In this example, we will examine whether fitness level is associated with squad membership.
The response variable is Squad_member:
The explanatory variable is Fitness.
# Create the data
data2 <- data.frame(
Subject_ID = 1:12,
Fitness = c(1, 3, 5, 10, 13, 15, 17, 20, 21, 14, 25, 30),
Distance = c(10, 2, 14, 8, 10, 7, 16, 10, 12, 13, 7, 17),
Squad_member = c(0, 0, 0, 0, 1, 0, 1, 0, 1, 1, 1, 1)
)
data2
## Subject_ID Fitness Distance Squad_member
## 1 1 1 10 0
## 2 2 3 2 0
## 3 3 5 14 0
## 4 4 10 8 0
## 5 5 13 10 1
## 6 6 15 7 0
## 7 7 17 16 1
## 8 8 20 10 0
## 9 9 21 12 1
## 10 10 14 13 1
## 11 11 25 7 1
## 12 12 30 17 1
Before fitting a model, it is useful to look at the data.
# Display the structure of the data
str(data2)
## 'data.frame': 12 obs. of 4 variables:
## $ Subject_ID : int 1 2 3 4 5 6 7 8 9 10 ...
## $ Fitness : num 1 3 5 10 13 15 17 20 21 14 ...
## $ Distance : num 10 2 14 8 10 7 16 10 12 13 ...
## $ Squad_member: num 0 0 0 0 1 0 1 0 1 1 ...
# Display the first few rows
head(data2)
## Subject_ID Fitness Distance Squad_member
## 1 1 1 10 0
## 2 2 3 2 0
## 3 3 5 14 0
## 4 4 10 8 0
## 5 5 13 10 1
## 6 6 15 7 0
# Plot fitness against squad membership
data2$Squad_label <- factor(
data2$Squad_member,
levels = c(0, 1),
labels = c("Non-squad member", "Squad member")
)
# Plot fitness against squad membership
ggplot(data2, aes(x = Fitness, y = Squad_label)) +
geom_jitter(width = 0, height = 0.05, size = 3) +
labs(
x = "Fitness",
y = "Squad Membership",
title = "Fitness and Squad Membership"
)
Because Squad_member has two possible outcomes (0 and
1), logistic regression can be used.
In R, logistic regression is fitted using glm() with
family = binomial.
# Fit the logistic regression model
model1 <- glm(
Squad_member ~ Fitness,
family = binomial,
data = data2
)
# Display the model results
summary(model1)
##
## Call:
## glm(formula = Squad_member ~ Fitness, family = binomial, data = data2)
##
## Coefficients:
## Estimate Std. Error z value Pr(>|z|)
## (Intercept) -3.6789 2.2948 -1.603 0.1089
## Fitness 0.2528 0.1465 1.726 0.0844 .
## ---
## Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
##
## (Dispersion parameter for binomial family taken to be 1)
##
## Null deviance: 16.636 on 11 degrees of freedom
## Residual deviance: 10.342 on 10 degrees of freedom
## AIC: 14.342
##
## Number of Fisher Scoring iterations: 5
The summary() function provides the estimated
coefficients for the logistic regression model.
The coefficient for Fitness is positive, which means the
estimated log-odds of being a squad member increase as fitness
increases.
In R, logistic regression is fitted using glm() with family = binomial.
The binomial family tells R that the response variable
has two possible outcomes, such as 0 and 1.
# Display the model coefficients
coef(summary(model1))
## Estimate Std. Error z value Pr(>|z|)
## (Intercept) -3.6788657 2.2947555 -1.603162 0.10889881
## Fitness 0.2528048 0.1464881 1.725771 0.08438861
The coefficients from a logistic regression model are expressed on the log-odds scale.
To make the result easier to interpret, we can convert the coefficient into an odds ratio by using the exponential function.
An odds ratio greater than 1 indicates higher odds, while an odds ratio less than 1 indicates lower odds.
For example, a probability of 0.80 means there is an 80% probability of being a squad member.
The corresponding odds are:
Odds = probability / (1 - probability)
For a probability of 0.80:
Odds = 0.80 / 0.20 = 4
This means the odds of being a squad member are 4 to 1.
# Calculate the odds ratio for Fitness
exp(coef(model1))
## (Intercept) Fitness
## 0.0252516 1.2876320
One useful feature of logistic regression is that the fitted model can be used to estimate the probability of an outcome.
In this example, we can calculate the predicted probability of being a squad member for each subject.
# Calculate predicted probabilities
data2$predicted_probability <- predict(
model1,
type = "response"
)
data2 %>%
select(
Subject_ID,
Fitness,
Squad_member,
predicted_probability
)
## Subject_ID Fitness Squad_member predicted_probability
## 1 1 1 0 0.03149085
## 2 2 3 0 0.05115180
## 3 3 5 0 0.08204794
## 4 4 10 0 0.24033983
## 5 5 13 1 0.40313902
## 6 6 15 0 0.52827154
## 7 7 17 1 0.64994936
## 8 8 20 0 0.79854594
## 9 9 21 1 0.83617457
## 10 10 14 1 0.46515708
## 11 11 25 1 0.93346997
## 12 12 30 1 0.98026210
The predicted probabilities can also be visualised against fitness level.
# Create a sequence of fitness values
new_data <- data.frame(
Fitness = seq(
min(data2$Fitness),
max(data2$Fitness),
length.out = 100
)
)
# Predict probabilities for the new fitness values
new_data$predicted_probability <- predict(
model1,
newdata = new_data,
type = "response"
)
# Plot the logistic regression curve
ggplot(data2, aes(x = Fitness, y = Squad_member)) +
geom_point(size = 3) +
geom_line(
data = new_data,
aes(x = Fitness, y = predicted_probability),
linewidth = 1
) +
scale_y_continuous(
limits = c(0, 1)
) +
labs(
x = "Fitness",
y = "Probability of Squad Membership",
title = "Logistic Regression: Predicted Probability"
)
The logistic regression model estimates a positive relationship between fitness and the probability of being a squad member.
The odds ratio for Fitness is approximately 1.29. This
means that a one-unit increase in fitness is associated with
approximately 1.29 times the odds of being a squad member.
The predicted probabilities also increase as fitness increases. For example, the predicted probability is approximately 0.03 when fitness is 1, compared with approximately 0.98 when fitness is 30.
The model therefore shows how logistic regression can be used to estimate the probability of a binary outcome from a continuous predictor.
In this tutorial, we used R to fit a logistic regression model using fitness level to predict squad membership.
The main steps were:
glm().The same process can be applied to other sport analytics problems where the outcome has two possible categories.
The complete R code used in this tutorial is provided below so the analysis can be replicated from start to finish.
# Load the required package
library(tidyverse)
# Create the data
data2 <- data.frame(
Subject_ID = 1:12,
Fitness = c(1, 3, 5, 10, 13, 15, 17, 20, 21, 14, 25, 30),
Distance = c(10, 2, 14, 8, 10, 7, 16, 10, 12, 13, 7, 17),
Squad_member = c(0, 0, 0, 0, 1, 0, 1, 0, 1, 1, 1, 1)
)
# Inspect the data
str(data2)
## 'data.frame': 12 obs. of 4 variables:
## $ Subject_ID : int 1 2 3 4 5 6 7 8 9 10 ...
## $ Fitness : num 1 3 5 10 13 15 17 20 21 14 ...
## $ Distance : num 10 2 14 8 10 7 16 10 12 13 ...
## $ Squad_member: num 0 0 0 0 1 0 1 0 1 1 ...
head(data2)
## Subject_ID Fitness Distance Squad_member
## 1 1 1 10 0
## 2 2 3 2 0
## 3 3 5 14 0
## 4 4 10 8 0
## 5 5 13 10 1
## 6 6 15 7 0
# Create readable labels for squad membership
data2$Squad_label <- factor(
data2$Squad_member,
levels = c(0, 1),
labels = c("Non-squad member", "Squad member")
)
# Plot fitness against squad membership
ggplot(data2, aes(x = Fitness, y = Squad_label)) +
geom_jitter(width = 0, height = 0.05, size = 3) +
labs(
x = "Fitness",
y = "Squad Membership",
title = "Fitness and Squad Membership"
)
# Fit the logistic regression model
model1 <- glm(
Squad_member ~ Fitness,
family = binomial,
data = data2
)
# Display model results
summary(model1)
##
## Call:
## glm(formula = Squad_member ~ Fitness, family = binomial, data = data2)
##
## Coefficients:
## Estimate Std. Error z value Pr(>|z|)
## (Intercept) -3.6789 2.2948 -1.603 0.1089
## Fitness 0.2528 0.1465 1.726 0.0844 .
## ---
## Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
##
## (Dispersion parameter for binomial family taken to be 1)
##
## Null deviance: 16.636 on 11 degrees of freedom
## Residual deviance: 10.342 on 10 degrees of freedom
## AIC: 14.342
##
## Number of Fisher Scoring iterations: 5
# Calculate odds ratios
exp(coef(model1))
## (Intercept) Fitness
## 0.0252516 1.2876320
# Calculate predicted probabilities
data2$predicted_probability <- predict(
model1,
type = "response"
)
# Display predicted probabilities
data2 %>%
select(
Subject_ID,
Fitness,
Squad_member,
predicted_probability
)
## Subject_ID Fitness Squad_member predicted_probability
## 1 1 1 0 0.03149085
## 2 2 3 0 0.05115180
## 3 3 5 0 0.08204794
## 4 4 10 0 0.24033983
## 5 5 13 1 0.40313902
## 6 6 15 0 0.52827154
## 7 7 17 1 0.64994936
## 8 8 20 0 0.79854594
## 9 9 21 1 0.83617457
## 10 10 14 1 0.46515708
## 11 11 25 1 0.93346997
## 12 12 30 1 0.98026210
# Create new fitness values
new_data <- data.frame(
Fitness = seq(
min(data2$Fitness),
max(data2$Fitness),
length.out = 100
)
)
# Predict probabilities for new fitness values
new_data$predicted_probability <- predict(
model1,
newdata = new_data,
type = "response"
)
# Plot the logistic regression curve
ggplot(data2, aes(x = Fitness, y = Squad_member)) +
geom_point(size = 3) +
geom_line(
data = new_data,
aes(x = Fitness, y = predicted_probability),
linewidth = 1
) +
scale_y_continuous(
limits = c(0, 1)
) +
labs(
x = "Fitness",
y = "Probability of Squad Membership",
title = "Logistic Regression: Predicted Probability"
)