A Beginner’s Guide to Logistic Regression in R

1. Introduction

Logistic regression is a statistical method used when the outcome has two possible categories.

In this tutorial, we will use R to build a simple logistic regression model. The aim is to show how logistic regression can be fitted, interpreted, and used to calculate predicted probabilities.

This tutorial is designed for beginners. No previous experience with logistic regression is required.

2. What is Logistic Regression?

Logistic regression can be used when the response variable has two possible outcomes.

For example, in sport analytics, an outcome could be:

  • Win or lose
  • Yes or no
  • Selected or not selected

Instead of predicting a continuous value, logistic regression estimates the probability of one of the two outcomes.

The probability is between 0 and 1, where:

  • 0 means the outcome is very unlikely
  • 1 means the outcome is very likely
# Load the required package
library(tidyverse)
## ── Attaching core tidyverse packages ──────────────────────── tidyverse 2.0.0 ──
## ✔ dplyr     1.2.1     ✔ readr     2.2.0
## ✔ forcats   1.0.1     ✔ stringr   1.6.0
## ✔ ggplot2   4.0.3     ✔ tibble    3.3.1
## ✔ lubridate 1.9.5     ✔ tidyr     1.3.2
## ✔ purrr     1.2.2     
## ── Conflicts ────────────────────────────────────────── tidyverse_conflicts() ──
## ✖ dplyr::filter() masks stats::filter()
## ✖ dplyr::lag()    masks stats::lag()
## ℹ Use the conflicted package (<http://conflicted.r-lib.org/>) to force all conflicts to become errors

3. Preparing the Data

In this example, we will examine whether fitness level is associated with squad membership.

The response variable is Squad_member:

  • 0 = Non squad member
  • 1 = Squad member

The explanatory variable is Fitness.

# Create the data
data2 <- data.frame(
  Subject_ID = 1:12,
  Fitness = c(1, 3, 5, 10, 13, 15, 17, 20, 21, 14, 25, 30),
  Distance = c(10, 2, 14, 8, 10, 7, 16, 10, 12, 13, 7, 17),
  Squad_member = c(0, 0, 0, 0, 1, 0, 1, 0, 1, 1, 1, 1)
)

data2
##    Subject_ID Fitness Distance Squad_member
## 1           1       1       10            0
## 2           2       3        2            0
## 3           3       5       14            0
## 4           4      10        8            0
## 5           5      13       10            1
## 6           6      15        7            0
## 7           7      17       16            1
## 8           8      20       10            0
## 9           9      21       12            1
## 10         10      14       13            1
## 11         11      25        7            1
## 12         12      30       17            1

3.1 Inspecting the Data

Before fitting a model, it is useful to look at the data.

# Display the structure of the data
str(data2)
## 'data.frame':    12 obs. of  4 variables:
##  $ Subject_ID  : int  1 2 3 4 5 6 7 8 9 10 ...
##  $ Fitness     : num  1 3 5 10 13 15 17 20 21 14 ...
##  $ Distance    : num  10 2 14 8 10 7 16 10 12 13 ...
##  $ Squad_member: num  0 0 0 0 1 0 1 0 1 1 ...
# Display the first few rows
head(data2)
##   Subject_ID Fitness Distance Squad_member
## 1          1       1       10            0
## 2          2       3        2            0
## 3          3       5       14            0
## 4          4      10        8            0
## 5          5      13       10            1
## 6          6      15        7            0
# Plot fitness against squad membership
data2$Squad_label <- factor(
  data2$Squad_member,
  levels = c(0, 1),
  labels = c("Non-squad member", "Squad member")
)

# Plot fitness against squad membership
ggplot(data2, aes(x = Fitness, y = Squad_label)) +
  geom_jitter(width = 0, height = 0.05, size = 3) +
  labs(
    x = "Fitness",
    y = "Squad Membership",
    title = "Fitness and Squad Membership"
  )

4. Fitting a Logistic Regression Model

Because Squad_member has two possible outcomes (0 and 1), logistic regression can be used.

In R, logistic regression is fitted using glm() with family = binomial.

# Fit the logistic regression model
model1 <- glm(
  Squad_member ~ Fitness,
  family = binomial,
  data = data2
)

# Display the model results
summary(model1)
## 
## Call:
## glm(formula = Squad_member ~ Fitness, family = binomial, data = data2)
## 
## Coefficients:
##             Estimate Std. Error z value Pr(>|z|)  
## (Intercept)  -3.6789     2.2948  -1.603   0.1089  
## Fitness       0.2528     0.1465   1.726   0.0844 .
## ---
## Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
## 
## (Dispersion parameter for binomial family taken to be 1)
## 
##     Null deviance: 16.636  on 11  degrees of freedom
## Residual deviance: 10.342  on 10  degrees of freedom
## AIC: 14.342
## 
## Number of Fisher Scoring iterations: 5

5. Understanding the Model Output

The summary() function provides the estimated coefficients for the logistic regression model.

The coefficient for Fitness is positive, which means the estimated log-odds of being a squad member increase as fitness increases.

In R, logistic regression is fitted using glm() with family = binomial.

The binomial family tells R that the response variable has two possible outcomes, such as 0 and 1.

# Display the model coefficients
coef(summary(model1))
##               Estimate Std. Error   z value   Pr(>|z|)
## (Intercept) -3.6788657  2.2947555 -1.603162 0.10889881
## Fitness      0.2528048  0.1464881  1.725771 0.08438861

5.1 Odds Ratio

The coefficients from a logistic regression model are expressed on the log-odds scale.

To make the result easier to interpret, we can convert the coefficient into an odds ratio by using the exponential function.

An odds ratio greater than 1 indicates higher odds, while an odds ratio less than 1 indicates lower odds.

For example, a probability of 0.80 means there is an 80% probability of being a squad member.

The corresponding odds are:

Odds = probability / (1 - probability)

For a probability of 0.80:

Odds = 0.80 / 0.20 = 4

This means the odds of being a squad member are 4 to 1.

# Calculate the odds ratio for Fitness
exp(coef(model1))
## (Intercept)     Fitness 
##   0.0252516   1.2876320

6. Predicting Probabilities

One useful feature of logistic regression is that the fitted model can be used to estimate the probability of an outcome.

In this example, we can calculate the predicted probability of being a squad member for each subject.

# Calculate predicted probabilities
data2$predicted_probability <- predict(
  model1,
  type = "response"
)

data2 %>%
  select(
    Subject_ID,
    Fitness,
    Squad_member,
    predicted_probability
)
##    Subject_ID Fitness Squad_member predicted_probability
## 1           1       1            0            0.03149085
## 2           2       3            0            0.05115180
## 3           3       5            0            0.08204794
## 4           4      10            0            0.24033983
## 5           5      13            1            0.40313902
## 6           6      15            0            0.52827154
## 7           7      17            1            0.64994936
## 8           8      20            0            0.79854594
## 9           9      21            1            0.83617457
## 10         10      14            1            0.46515708
## 11         11      25            1            0.93346997
## 12         12      30            1            0.98026210

7. Visualising the Predicted Probability

The predicted probabilities can also be visualised against fitness level.

# Create a sequence of fitness values
new_data <- data.frame(
  Fitness = seq(
    min(data2$Fitness),
    max(data2$Fitness),
    length.out = 100
  )
)

# Predict probabilities for the new fitness values
new_data$predicted_probability <- predict(
  model1,
  newdata = new_data,
  type = "response"
)

# Plot the logistic regression curve
ggplot(data2, aes(x = Fitness, y = Squad_member)) +
  geom_point(size = 3) +
  geom_line(
    data = new_data,
    aes(x = Fitness, y = predicted_probability),
    linewidth = 1
  ) +
  scale_y_continuous(
    limits = c(0, 1)
  ) +
  labs(
    x = "Fitness",
    y = "Probability of Squad Membership",
    title = "Logistic Regression: Predicted Probability"
  )

8. Interpreting the Results

The logistic regression model estimates a positive relationship between fitness and the probability of being a squad member.

The odds ratio for Fitness is approximately 1.29. This means that a one-unit increase in fitness is associated with approximately 1.29 times the odds of being a squad member.

The predicted probabilities also increase as fitness increases. For example, the predicted probability is approximately 0.03 when fitness is 1, compared with approximately 0.98 when fitness is 30.

The model therefore shows how logistic regression can be used to estimate the probability of a binary outcome from a continuous predictor.

9. Conclusion

In this tutorial, we used R to fit a logistic regression model using fitness level to predict squad membership.

The main steps were:

  1. Prepare the data.
  2. Identify a binary response variable.
  3. Fit a logistic regression model using glm().
  4. Interpret the model coefficients.
  5. Calculate an odds ratio.
  6. Calculate predicted probabilities.
  7. Visualise the predicted probabilities.

The same process can be applied to other sport analytics problems where the outcome has two possible categories.

10. Complete Code

The complete R code used in this tutorial is provided below so the analysis can be replicated from start to finish.

# Load the required package
library(tidyverse)

# Create the data
data2 <- data.frame(
  Subject_ID = 1:12,
  Fitness = c(1, 3, 5, 10, 13, 15, 17, 20, 21, 14, 25, 30),
  Distance = c(10, 2, 14, 8, 10, 7, 16, 10, 12, 13, 7, 17),
  Squad_member = c(0, 0, 0, 0, 1, 0, 1, 0, 1, 1, 1, 1)
)

# Inspect the data
str(data2)
## 'data.frame':    12 obs. of  4 variables:
##  $ Subject_ID  : int  1 2 3 4 5 6 7 8 9 10 ...
##  $ Fitness     : num  1 3 5 10 13 15 17 20 21 14 ...
##  $ Distance    : num  10 2 14 8 10 7 16 10 12 13 ...
##  $ Squad_member: num  0 0 0 0 1 0 1 0 1 1 ...
head(data2)
##   Subject_ID Fitness Distance Squad_member
## 1          1       1       10            0
## 2          2       3        2            0
## 3          3       5       14            0
## 4          4      10        8            0
## 5          5      13       10            1
## 6          6      15        7            0
# Create readable labels for squad membership
data2$Squad_label <- factor(
  data2$Squad_member,
  levels = c(0, 1),
  labels = c("Non-squad member", "Squad member")
)

# Plot fitness against squad membership
ggplot(data2, aes(x = Fitness, y = Squad_label)) +
  geom_jitter(width = 0, height = 0.05, size = 3) +
  labs(
    x = "Fitness",
    y = "Squad Membership",
    title = "Fitness and Squad Membership"
  )

# Fit the logistic regression model
model1 <- glm(
  Squad_member ~ Fitness,
  family = binomial,
  data = data2
)

# Display model results
summary(model1)
## 
## Call:
## glm(formula = Squad_member ~ Fitness, family = binomial, data = data2)
## 
## Coefficients:
##             Estimate Std. Error z value Pr(>|z|)  
## (Intercept)  -3.6789     2.2948  -1.603   0.1089  
## Fitness       0.2528     0.1465   1.726   0.0844 .
## ---
## Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
## 
## (Dispersion parameter for binomial family taken to be 1)
## 
##     Null deviance: 16.636  on 11  degrees of freedom
## Residual deviance: 10.342  on 10  degrees of freedom
## AIC: 14.342
## 
## Number of Fisher Scoring iterations: 5
# Calculate odds ratios
exp(coef(model1))
## (Intercept)     Fitness 
##   0.0252516   1.2876320
# Calculate predicted probabilities
data2$predicted_probability <- predict(
  model1,
  type = "response"
)

# Display predicted probabilities
data2 %>%
  select(
    Subject_ID,
    Fitness,
    Squad_member,
    predicted_probability
)
##    Subject_ID Fitness Squad_member predicted_probability
## 1           1       1            0            0.03149085
## 2           2       3            0            0.05115180
## 3           3       5            0            0.08204794
## 4           4      10            0            0.24033983
## 5           5      13            1            0.40313902
## 6           6      15            0            0.52827154
## 7           7      17            1            0.64994936
## 8           8      20            0            0.79854594
## 9           9      21            1            0.83617457
## 10         10      14            1            0.46515708
## 11         11      25            1            0.93346997
## 12         12      30            1            0.98026210
# Create new fitness values
new_data <- data.frame(
  Fitness = seq(
    min(data2$Fitness),
    max(data2$Fitness),
    length.out = 100
  )
)

# Predict probabilities for new fitness values
new_data$predicted_probability <- predict(
  model1,
  newdata = new_data,
  type = "response"
)

# Plot the logistic regression curve
ggplot(data2, aes(x = Fitness, y = Squad_member)) +
  geom_point(size = 3) +
  geom_line(
    data = new_data,
    aes(x = Fitness, y = predicted_probability),
    linewidth = 1
  ) +
  scale_y_continuous(
    limits = c(0, 1)
  ) +
  labs(
    x = "Fitness",
    y = "Probability of Squad Membership",
    title = "Logistic Regression: Predicted Probability"
  )