Introduction

For this project I wanted to figure out who the best NFL kicker really is, using field goal and extra point attempts from the 2022 NFL Big Data Bowl dataset (2018 to 2020 seasons). A simple accuracy rate, makes divided by attempts, is a good starting point, but it has an obvious flaw: it treats a 20 yard kick the same as a 55 yard kick. A kicker who only ever attempts short kicks can post a better percentage than a kicker who is constantly asked to attempt 50 plus yarders, without actually being the better kicker. So instead of stopping at raw accuracy, I built a metric that adjusts for how hard each kick actually was, rewarding kickers more for making long kicks and penalizing them less for missing them, since misses become statistically more likely as distance increases.

My approach: first calculate the league wide make rate at each kick distance using every field goal attempt in the dataset. That becomes an “expected” make probability for a kick of that length. Then for each kicker, I compare what they actually made against what an average kicker would have been expected to make given the same distances they attempted. A kicker who consistently makes more kicks than expected, especially at longer distances, ranks higher than someone with a flashy overall percentage built entirely on chip shots. I also required a minimum number of attempts and more than one season played so a kicker with a handful of lucky attempts does not rank above someone with a large, proven sample. Beyond that core metric, I tested whether the ranking survives accounting for small sample size, whether kickers perform differently in high pressure situations, and whether they are consistent from season to season.

1. Loading and Merging the Data

The first step is pulling in the four non-tracking files (plays, players, games, and the PFF scouting data) and merging them into a single data frame keyed on game, play, and kicker IDs. I also ran a quick sanity check against the reference counts from the assignment prompt to confirm the merge and filtering logic were working correctly before moving forward.

setwd("C:/Users/Rprut/Downloads/Sports Analytics/NFLBDB2022/")

source("https://raw.githubusercontent.com/ptallon/SportsAnalytics_Fall2026/refs/heads/main/SharedCode.R")

load_packages(c("data.table", "dplyr", "ggplot2", "scales", "tidyr", "ggrepel"))

suppressMessages(library(data.table))

plays   <- fread("plays.csv")
players <- fread("players.csv")
games   <- fread("games.csv")
pff     <- fread("PFFScoutingData.csv")

full_df <- plays |>
  left_join(players, by = c("kickerId" = "nflId")) |>
  left_join(games, by = "gameId") |>
  left_join(pff, by = c("gameId", "playId")) |>
  select(where(~ !all(is.na(.)))) |>
  data.frame()

# check counts against assignment reference numbers
table(plays$specialTeamsPlayType)
## 
## Extra Point  Field Goal     Kickoff        Punt 
##        3488        2657        7843        5991
table(plays$specialTeamsResult)
## 
##     Blocked Kick Attempt             Blocked Punt                   Downed 
##                       61                       39                      834 
##               Fair Catch        Kick Attempt Good     Kick Attempt No Good 
##                     1645                     5470                      585 
##    Kickoff Team Recovery                   Muffed Non-Special Teams Result 
##                       16                      214                      101 
##            Out of Bounds                   Return                Touchback 
##                      651                     5207                     5156

2. Narrowing to Field Goal Attempts

Field goals (and extra points, which share the same kicking mechanics) are the cleanest play type for this question, since the outcome is binary and distance is the clearest driver of difficulty. I filtered down to field goal attempts only, keeping made kicks, missed kicks, and blocked kicks as the three possible outcomes.

# blocks kept as misses for now, a block isn't purely a kicker skill failure
fg_df <- full_df |>
  filter(specialTeamsPlayType == "Field Goal",
         specialTeamsResult %in% c("Kick Attempt Good",
                                    "Kick Attempt No Good",
                                    "Blocked Kick Attempt")) |>
  mutate(made = as.integer(specialTeamsResult == "Kick Attempt Good")) |>
  filter(!is.na(kickLength))

3. Building the League Wide Difficulty Curve

To adjust for distance, I needed a baseline: how likely is any kicker, league average, to make a kick from a given distance? I fit a logistic regression of make probability on kick distance using every attempt in the dataset. This curve becomes the “expected” probability of success for each individual attempt, which every later metric is built on top of.

league_model <- glm(made ~ kickLength, data = fg_df, family = binomial)
summary(league_model)
## 
## Call:
## glm(formula = made ~ kickLength, family = binomial, data = fg_df)
## 
## Coefficients:
##              Estimate Std. Error z value Pr(>|z|)    
## (Intercept)  6.593214   0.329670   20.00   <2e-16 ***
## kickLength  -0.115666   0.007219  -16.02   <2e-16 ***
## ---
## Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
## 
## (Dispersion parameter for binomial family taken to be 1)
## 
##     Null deviance: 2185.4  on 2603  degrees of freedom
## Residual deviance: 1846.3  on 2602  degrees of freedom
## AIC: 1850.3
## 
## Number of Fisher Scoring iterations: 5
fg_df <- fg_df |>
  mutate(expected_prob = predict(league_model, newdata = fg_df, type = "response"))

ggplot(fg_df, aes(x = kickLength, y = expected_prob)) +
  geom_line(linewidth = 1, color = "steelblue") +
  geom_jitter(aes(y = made), alpha = 0.15, height = 0.03, size = 1) +
  labs(title = "League-Wide Field Goal Make Probability by Distance",
       subtitle = "Points jittered vertically; line is the fitted model",
       x = "Kick Distance (yards)",
       y = "Predicted Probability of Success") +
  theme_minimal()

As expected, the probability of success declines steadily as distance increases, dropping from close to automatic inside 30 yards to roughly a coin flip past 60.

4. Points Above Expected: A Distance-Adjusted Ranking

With a baseline difficulty curve in hand, I summarized each kicker’s attempts by season and then by career, comparing their actual makes to the number the model expected given the distances they attempted. The difference, points above expected, is the core metric for this analysis. I required kickers to have played more than one season and attempted at least 20 kicks so the ranking is not dominated by small, noisy samples.

kicker_season <- fg_df |>
  group_by(displayName, possessionTeam, season) |>
  summarise(
    attempts       = n(),
    actual_makes   = sum(made),
    expected_makes = sum(expected_prob),
    avg_distance   = mean(kickLength, na.rm = TRUE),
    .groups = "drop"
  )

kicker_career <- kicker_season |>
  group_by(displayName) |>
  summarise(
    seasons        = n_distinct(season),
    attempts       = sum(attempts),
    actual_makes   = sum(actual_makes),
    expected_makes = sum(expected_makes),
    avg_distance   = mean(avg_distance),
    .groups = "drop"
  ) |>
  mutate(
    points_above_expected = actual_makes - expected_makes,
    pae_per_attempt        = points_above_expected / attempts,
    actual_pct              = percent(actual_makes / attempts, accuracy = 0.1)
  ) |>
  filter(seasons > 1, attempts >= 20) |>
  arrange(-pae_per_attempt) |>
  data.frame()

The two plots below visualize this ranking two different ways. The first shows the top 10 kickers ranked by points above expected per attempt, with their raw accuracy labeled on each bar for context. The second plots actual makes against expected makes directly, with bubble size showing attempt volume, so a kicker who is both accurate and heavily used stands out clearly above and to the right.

top10 <- kicker_career |> slice_head(n = 10)

ggplot(top10, aes(x = reorder(displayName, pae_per_attempt),
                   y = pae_per_attempt, fill = pae_per_attempt > 0)) +
  geom_col() +
  geom_text(aes(label = actual_pct), hjust = -0.1, size = 3) +
  coord_flip() +
  scale_fill_manual(values = c("firebrick", "forestgreen"), guide = "none") +
  labs(title = "Top 10 Kickers: Distance-Adjusted Performance (2018-2020)",
       subtitle = "Points Above Expected per attempt, vs. a league-average kicker at the same distances",
       x = NULL,
       y = "Points Above Expected per Attempt") +
  theme_minimal() +
  theme(plot.title = element_text(hjust = 0.5))

ggplot(top10, aes(x = expected_makes, y = actual_makes, size = attempts, label = displayName)) +
  geom_abline(slope = 1, intercept = 0, linetype = "dashed", color = "gray50") +
  geom_point(alpha = 0.7, color = "steelblue") +
  ggrepel::geom_text_repel(size = 3) +
  labs(title = "Actual vs. Expected Field Goals Made",
       subtitle = "Above the dashed line = outperforming expectation given kick distances attempted",
       x = "Expected Makes (league-average kicker)",
       y = "Actual Makes",
       size = "Attempts") +
  theme_minimal()

At this stage, Graham Gano leads the per-attempt rate, with Justin Tucker, Josh Lambo, and Jason Myers close behind. Gano’s lead comes from a relatively small sample (45 attempts), which raises the question of whether that top spot is a real skill edge or a small sample effect. The next section addresses this directly.

5. Shrinkage Estimator: Does the Ranking Survive Small Samples

A kicker’s accuracy is a proportion, and proportions estimated from small samples are noisier than ones estimated from large samples. To account for this, I applied a shrinkage estimator that pulls every kicker’s accuracy toward the league average, weighted by how many attempts they have. A kicker with few attempts gets pulled further toward average; a kicker with many attempts barely moves. This is a standard way to separate a real effect from a lucky streak.

league_avg_rate <- sum(fg_df$made) / nrow(fg_df)
k <- 30  # pseudo-attempts of league-average prior

shrinkage_table <- fg_df |>
  group_by(displayName) |>
  summarise(attempts = n(), actual_makes = sum(made), .groups = "drop") |>
  filter(attempts >= 20) |>
  mutate(
    actual_pct = round(actual_makes / attempts * 100, 1),
    shrunk_pct = round((actual_makes + k * league_avg_rate) / (attempts + k) * 100, 1)
  ) |>
  arrange(-shrunk_pct) |>
  mutate(shrunk_rank = row_number()) |>
  select(displayName, attempts, actual_pct, shrunk_pct, shrunk_rank)

This changes the picture. After shrinkage, Justin Tucker and Josh Lambo are essentially tied for first (92.3 percent both), with Gano dropping to third. Gano’s raw rate advantage was partly a small sample effect, exactly the kind of result this test is designed to catch. Tucker and Lambo being this close matters for the final decision: when two kickers are statistically even, the one with more attempts behind the number (Tucker, with 92 attempts versus Lambo’s 57) is the safer pick, since there is more evidence supporting it.

6. Clutch Performance

A common claim in kicker rankings is that some kickers are more reliable under pressure than others. I defined a clutch attempt as any field goal in the 4th quarter or overtime with the score within one possession (8 points) either way, and compared each kicker’s distance-adjusted performance in clutch versus non-clutch situations.

fg_df <- fg_df |>
  mutate(
    kicking_team_score = ifelse(possessionTeam == homeTeamAbbr,
                                 preSnapHomeScore, preSnapVisitorScore),
    opponent_score      = ifelse(possessionTeam == homeTeamAbbr,
                                  preSnapVisitorScore, preSnapHomeScore),
    score_diff = kicking_team_score - opponent_score,
    clutch = quarter >= 4 & abs(score_diff) <= 8
  )

table(fg_df$clutch)
## 
## FALSE  TRUE 
##  2161   443
# table used to compare clutch vs normal attempts by kicker
clutch_summary <- fg_df |>
  filter(displayName %in% kicker_career$displayName[1:15]) |>
  group_by(displayName, clutch) |>
  summarise(
    attempts       = n(),
    actual_makes   = sum(made),
    expected_makes = sum(expected_prob),
    pct_made       = percent(actual_makes / attempts, accuracy = 1),
    pae            = actual_makes - expected_makes,
    .groups = "drop"
  ) |>
  pivot_wider(
    names_from  = clutch,
    values_from = c(attempts, actual_makes, expected_makes, pct_made, pae),
    names_glue  = "{.value}_{ifelse(clutch, 'clutch', 'normal')}"
  )

Beyond the kicker by kicker comparison, I also wanted to know whether the league as a whole performs differently under pressure, after controlling for distance. A pooled logistic regression across every attempt gives a much more statistically powerful test than any single kicker’s small clutch sample.

clutch_model <- glm(made ~ expected_prob + clutch, data = fg_df, family = binomial)
summary(clutch_model)
## 
## Call:
## glm(formula = made ~ expected_prob + clutch, family = binomial, 
##     data = fg_df)
## 
## Coefficients:
##               Estimate Std. Error z value Pr(>|z|)    
## (Intercept)   -3.57265    0.32323  -11.05   <2e-16 ***
## expected_prob  6.52628    0.40413   16.15   <2e-16 ***
## clutchTRUE     0.08067    0.15810    0.51     0.61    
## ---
## Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
## 
## (Dispersion parameter for binomial family taken to be 1)
## 
##     Null deviance: 2185.4  on 2603  degrees of freedom
## Residual deviance: 1900.0  on 2601  degrees of freedom
## AIC: 1906
## 
## Number of Fisher Scoring iterations: 5

The clutch indicator comes back nowhere near statistically significant (p = 0.61). This result held up even after testing a stricter clutch definition earlier in the project, with the direction of the (non-significant) effect flipping between the two definitions, which is consistent with noise rather than a real effect. The chart below visualizes clutch versus non-clutch performance for the top 10 kickers for reference, followed by a per-kicker significance test on clutch attempts only.

clutch_plot_df <- fg_df |>
  filter(displayName %in% kicker_career$displayName[1:10]) |>
  group_by(displayName, clutch) |>
  summarise(pae_rate = (sum(made) - sum(expected_prob)) / n(),
            attempts = n(), .groups = "drop") |>
  mutate(situation = ifelse(clutch, "Clutch", "Normal"))

ggplot(clutch_plot_df, aes(x = reorder(displayName, pae_rate), y = pae_rate, fill = situation)) +
  geom_col(position = "dodge") +
  coord_flip() +
  labs(title = "Clutch vs. Non-Clutch Performance (Distance-Adjusted)",
       subtitle = "Clutch = Q4/OT and one-score game",
       x = NULL, y = "Points Above Expected per Attempt",
       fill = NULL) +
  theme_minimal()

# per kicker test on clutch attempts only, small samples so treat as suggestive
clutch_sig <- fg_df |>
  filter(clutch, displayName %in% kicker_career$displayName[1:15]) |>
  group_by(displayName) |>
  summarise(attempts = n(), made = sum(made), expected_p = mean(expected_prob), .groups = "drop") |>
  filter(attempts >= 5) |>
  rowwise() |>
  mutate(p_value = binom.test(made, attempts, p = expected_p)$p.value) |>
  ungroup() |>
  arrange(p_value)

No individual kicker clears the conventional 0.05 significance threshold on their clutch attempts either. Justin Tucker comes closest, with a perfect 17 for 17 clutch record against an expected make rate of about 85 percent (p = 0.094), which is a notable descriptive fact even though it does not reach statistical significance. Overall, clutch performance does not appear to be a reliably measurable skill in this dataset, for any kicker.

7. Season to Season Consistency

The last angle I tested was consistency. Two kickers can post the same career average while one is steady every season and the other swings between great and poor years. A reliable “best kicker” candidate should ideally be consistent as well as accurate, so I calculated the standard deviation of each kicker’s season by season accuracy.

consistency <- kicker_season |>
  mutate(season_pct = actual_makes / attempts) |>
  group_by(displayName) |>
  summarise(seasons = n(), mean_pct = mean(season_pct), sd_pct = sd(season_pct), .groups = "drop") |>
  filter(seasons > 1) |>
  arrange(sd_pct)

Consistency and quality are not the same thing. Matt Gay is the most consistent kicker in the dataset but with a below average mean accuracy, steady but steadily mediocre rather than steadily great. Justin Tucker ranks 7th out of 42 kickers for consistency while also posting one of the highest mean accuracies, a strong combination of being both good and reliable. It is worth noting these standard deviations come from only 2 to 3 seasons per kicker, so they should be treated as a directional signal rather than a precise ranking.

Conclusion

Taking all of these results together, I land on Justin Tucker as the best kicker in this dataset, though the margin is closer than a clean headline would suggest, and that nuance is worth stating directly rather than smoothing over. Across the different tests:

Tucker is the best supported answer mainly because he combines being elite with being proven over a large, reliable sample, not because he decisively wins every individual category. Josh Lambo is a legitimate alternate pick for anyone who weights rate of success over total volume, but his case rests on noticeably less data. Howeverm he could still be a valid choice for anyone who weights rate of success over total volume, and that tradeoff between rewarding rate versus total value is really the central judgment call in deciding what “best kicker” means in the first place.