Introduction

Introduction

The goal of this project is to identify the best-performing NFL kickers using data from the 2022 NFL Big Data Bowl dataset, which includes plays from the 2020–2021 NFL season. Rather than evaluating kickers using field goal percentage alone, this analysis considers the distance of each field goal attempt. This is important because shorter kicks generally have a higher probability of success than longer kicks, making it difficult to compare kickers fairly using overall accuracy alone.

To account for differences in kick difficulty, I use binary logistic regression to estimate the probability of making a field goal based on kick distance. I then compare each kicker’s actual number of successful field goals with the number expected from the model. This produces a measure of field goals above expected, which helps identify kickers who performed better than predicted given the distances of their attempts.

The analysis uses visualizations to explore kicker performance, including a bar chart ranking the top five kickers by field goals above expected and a football field visualization showing predicted field goal success probabilities at different distances. Together, these methods provide a more detailed comparison of kicker performance and demonstrate how statistical modeling can help evaluate success beyond traditional field goal percentage.


Step 1. Load the CSV files

Here I explain briefly what files I am loading into memory.

setwd("/Users/rubynguyen/sports analytics/fall sports analytics 2026/nfl-big-data-bowl-2022")



library(data.table)
library(ggplot2)
library(plotly)
library(stringr)
library(dplyr)
library(htmltools)
library(htmlwidgets)

y1 <- fread("tracking2018.csv")
y2 <- fread("tracking2019.csv")
y3 <- fread("tracking2020.csv")



tracking_df <- rbind(y1, y2, y3)
rm(y1, y2, y3)

tracking_df$year <- year(tracking_df$time)
tracking_df$month <- month(tracking_df$time)



plays <- fread("plays.csv")
players <- fread("players.csv")
games <- fread("games.csv")


plays_df <- left_join(plays, players, by = c("kickerId" = "nflId"))
plays_df <- left_join(plays_df, games, by = c("gameId"))

Data Frame Prep

Here I explain how I built up my data frame.

source("https://raw.githubusercontent.com/mlfurman3/gg_field/main/gg_field.R")

# 1. Prepare field goal data
field_goals <- plays %>%
  filter(
    specialTeamsPlayType == "Field Goal",
    specialTeamsResult %in% c("Kick Attempt Good", "Kick Attempt No Good"),
    !is.na(kickerId),
    !is.na(kickLength)
  ) %>%
  mutate(
    made = ifelse(specialTeamsResult == "Kick Attempt Good", 1, 0),
    distance_group = case_when(
      kickLength < 30 ~ "<30",
      kickLength >= 30 & kickLength <= 39 ~ "30-39",
      kickLength >= 40 & kickLength <= 49 ~ "40-49",
      kickLength >= 50 ~ "50+"
    ),
    distance_group = factor(
      distance_group,
      levels = c("<30", "30-39", "40-49", "50+")
    )
  )

# 2. Overall kicker statistics
kicker_profile <- field_goals %>%
  group_by(kickerId) %>%
  summarise(
    Attempts = n(),
    Makes = sum(made),
    Misses = Attempts - Makes,
    FG_Percentage = Makes / Attempts * 100,
    Average_Distance = mean(kickLength),
    Longest_Kick = max(kickLength),
    .groups = "drop"
  )

# 3. Field goal success by distance
distance_success <- field_goals %>%
  group_by(distance_group) %>%
  summarise(
    Attempts = n(),
    Makes = sum(made),
    Misses = Attempts - Makes,
    FG_Percentage = Makes / Attempts * 100,
    .groups = "drop"
  )

# 4. Logistic regression model
# Predict the probability of making a kick based on distance
logit_model <- glm(
  made ~ kickLength,
  data = field_goals,
  family = binomial
)

# Expected probability for every field goal attempt
field_goals$expected_prob <- predict(
  logit_model,
  newdata = field_goals,
  type = "response"
)

# 5. Actual makes, expected makes, and field goals above expected
kicker_expected <- field_goals %>%
  group_by(kickerId) %>%
  summarise(
    Attempts = n(),
    Makes = sum(made),
    Expected_Makes = sum(expected_prob),
    .groups = "drop"
  ) %>%
  mutate(
    FG_Above_Expected = Makes - Expected_Makes,
    FG_Percentage = Makes / Attempts * 100
  ) %>%
  left_join(
    players %>%
      select(nflId, displayName) %>%
      distinct(nflId, .keep_all = TRUE),
    by = c("kickerId" = "nflId")
  ) %>%
  select(
    kickerId,
    displayName,
    Attempts,
    Makes,
    Expected_Makes,
    FG_Above_Expected,
    FG_Percentage
  )

# 6. Select the top 5 kickers by field goals above expected
top_5_expected <- kicker_expected %>%
  filter(!is.na(displayName)) %>%
  arrange(desc(FG_Above_Expected)) %>%
  slice_head(n = 5)

# 7. Predicted probabilities for the football field visualization
field_probabilities <- data.frame(
  kickLength = 20:60
) %>%
  mutate(
    probability = predict(
      logit_model,
      newdata = .,
      type = "response"
    ),
    x_right = 127 - kickLength,
    x_left = kickLength - 7
  )

# Create both sides of the field probability shading
field_probabilities_long <- bind_rows(
  field_probabilities %>%
    select(kickLength, probability, x = x_left),
  field_probabilities %>%
    select(kickLength, probability, x = x_right)
)

# 8. Successful field goals made by the top 5 kickers
top_5_made <- field_goals %>%
  filter(
    kickerId %in% top_5_expected$kickerId,
    made == 1
  ) %>%
  left_join(
    players %>%
      select(nflId, displayName) %>%
      distinct(nflId, .keep_all = TRUE),
    by = c("kickerId" = "nflId")
  )

# 9. Top 5 kickers' performance by distance group
kicker_distance <- field_goals %>%
  group_by(kickerId, distance_group) %>%
  summarise(
    Attempts = n(),
    Makes = sum(made),
    FG_Percentage = Makes / Attempts * 100,
    .groups = "drop"
  ) %>%
  left_join(
    players %>%
      select(nflId, displayName) %>%
      distinct(nflId, .keep_all = TRUE),
    by = c("kickerId" = "nflId")
  )

top_5_distance <- kicker_distance %>%
  filter(
    kickerId %in% top_5_expected$kickerId,
    Attempts >= 10
  ) %>%
  group_by(distance_group) %>%
  arrange(desc(FG_Percentage), .by_group = TRUE) %>%
  slice_head(n = 5) %>%
  ungroup()

Bar Chart Visual

The bar chart compares the actual number of field goals each kicker made with the number my logistic regression model expected them to make based on kick distance. I used this comparison because a kicker who makes longer, more difficult field goals should not necessarily be judged the same way as a kicker who attempts shorter kicks.

Justin Tucker ranked first, making approximately 9 more field goals than the model expected. Jason Myers, Brandon McManus, Josh Lambo, and Graham Gano also made more field goals than expected. These results suggest that the five kickers performed better than the model predicted, given the distances of their attempts.

This metric helps compare kickers while accounting for kick distance, although it does not account for other factors that can affect a kick, such as weather, pressure, or the conditions of the play.

ggplot(
  top_5_expected,
  aes(
    x = reorder(displayName, FG_Above_Expected),
    y = FG_Above_Expected
  )
) +
  geom_col(
    fill = "#247BA0",
    width = 0.7
  ) +
  geom_text(
    aes(label = sprintf("+%.1f", FG_Above_Expected)),
    hjust = -0.2,
    size = 4
  ) +
  coord_flip() +
  scale_y_continuous(
    limits = c(0, 11),
    breaks = seq(0, 10, 2),
    expand = expansion(mult = c(0, 0.05))
  ) +
  labs(
    title = "The NFL's Top 5 Kickers Above Expected",
    subtitle = "Actual field goals made minus expected makes",
    x = NULL,
    y = "Field Goals Above Expected",
    caption = "Expected makes estimated using logistic regression on kick distance."
  ) +
  theme_minimal(base_size = 13) +
  theme(
    plot.title = element_text(
      size = 18,
      face = "bold"
    ),
    plot.subtitle = element_text(
      size = 11,
      color = "gray40"
    ),
    axis.text.y = element_text(
      face = "bold",
      color = "gray20"
    ),
    panel.grid.major.y = element_blank(),
    panel.grid.minor = element_blank(),
    plot.caption = element_text(
      color = "gray40",
      hjust = 0
    )
  )

Football Field Visual

The football field visualization provides a closer look at the successful field goals made by the five kickers who ranked highest in field goals above expected. Each dot represents a successful kick, and the colors distinguish between kickers. The position of each dot represents the kick’s location along the length of the field, based on the available distance information.

This visualization complements the bar chart by showing the distances at which these kickers successfully made field goals. It helps illustrate that the top performers were successful on field goal attempts of varying distances. However, because the dataset does not provide the exact left-to-right location of each kick, the dots are spread vertically for visibility rather than representing the actual lateral location of the kicks.

top_5_made <- field_goals %>%
  filter(
    kickerId %in% top_5_expected$kickerId,
    made == 1
  ) %>%
  left_join(
    players %>% select(nflId, displayName),
    by = c("kickerId" = "nflId")
  )

ggplot() +
  gg_field() +
  
  # Predicted probability across the field
  geom_tile(
    data = field_probabilities_long,
    aes(
      x = x,
      y = 26.665,
      fill = probability
    ),
    width = 1,
    height = 53.33,
    alpha = 0.7
  ) +
  
  # Made field goals by the top 5 kickers
  geom_jitter(
    data = top_5_made,
    aes(
      x = absoluteYardlineNumber,
      y = 26.665,
      color = displayName
    ),
    height = 8,
    width = 0.2,
    size = 2,
    alpha = 0.7
  ) +
  
  # Probability color scale
  scale_fill_gradient2(
    name = "Probability of Make",
    low = "red",
    mid = "yellow",
    high = "green",
    midpoint = 0.75,
    limits = c(0.4, 1),
    labels = scales::percent
  ) +
  
  labs(
    title = "Top 5 Kickers: Successful Field Goals by Distance",
    subtitle = "Background shows predicted probability; dots show made field goals",
    color = "Kicker"
  )

Conclusion

Overall, the analysis shows that kick distance is an important factor in field goal success. By comparing actual makes with expected makes, I was able to identify kickers who performed better than predicted based on distance alone. This demonstrates why evaluating kickers using multiple metrics provides a more complete picture of their performance than field goal percentage alone.

Logo