Assignment Description

This assignment uses the 2022 NFL Big Data Bowl data, which covers every special teams play from the 2018–2020 NFL seasons, to answer one question: who is the best kicker in the NFL? There is no single right definition of “best,” so the first job is to choose one and defend it.

Most kicker rankings come down to one number: field goal percentage. That number hides when and where the kicks happened. A 90% kicker who plays half his games in a dome and rarely kicks with the game on the line has had an easier job than an 87% kicker who kicks outdoors in December and in close fourth quarters.

For this assignment I define the best kicker as the most reliable kicker in hard situations. Using every field goal attempt from the 2018–2020 seasons, I measure how accurate each kicker stays:

  • under pressure (the 4th quarter or overtime in a one-score game),
  • on the road,
  • outdoors, where wind, cold and rain come into play,

and how often his kicks get blocked. These combine into a single reliability score.

Description of Project

Data

load_packages <- function(pkgs) {
  missing <- pkgs[!pkgs %in% rownames(installed.packages())]
  if (length(missing) > 0) install.packages(missing, dependencies = TRUE, quiet = TRUE,
                                           repos = "https://cloud.r-project.org")
  invisible(suppressPackageStartupMessages(lapply(pkgs, library, character.only = TRUE)))
}
load_packages(c("data.table", "dplyr", "tidyr", "ggplot2", "scales", "knitr",
                "kableExtra", "plotly"))

# --- 2022 NFL Big Data Bowl files (special teams, 2018-2020) ---
plays   <- fread(file.path(data_dir, "plays.csv"))
games   <- fread(file.path(data_dir, "games.csv"))
players <- fread(file.path(data_dir, "players.csv"))

# --- Kick-level data with distance and game-time weather (one row per made/missed FG or XP) ---
kick_data <- fread(file.path(data_dir, "kick_model_data.csv"))

# --- Stadium roof types from the WeatherData repository (github.com/ThompsonJamesBliss/WeatherData) ---
wx_url   <- "https://raw.githubusercontent.com/ThompsonJamesBliss/WeatherData/master/data/"
wx_games <- fread(paste0(wx_url, "games.csv"))
stadiums <- fread(paste0(wx_url, "stadium_coordinates.csv")) |>
  distinct(StadiumName, RoofType)                       # NYG/NYJ share a stadium

roofs <- wx_games |>
  select(gameId = game_id, StadiumName) |>
  left_join(stadiums, by = "StadiumName")

I combined three sources:

  1. The 2022 Big Data Bowl files (plays, games, players): the result of every kick, the score and quarter before the snap, which team was kicking, and who was at home.
  2. A kick-level data file built from the same Big Data Bowl plays, which adds each kick’s distance and the game-time weather (wind, temperature, rain, snow) and flags whether the game was played under a dome or closed roof.
  3. The WeatherData repository by Tom Bliss of the NFL, which lists the stadium for every game and whether its roof is outdoor, indoor or retractable.

Checking the data

Before building anything, I merged the four Big Data Bowl files (games, plays, players, PFFScoutingData) into one data frame and checked the play counts against the ones given in the assignment.

pff <- fread(file.path(data_dir, "PFFScoutingData.csv"))

df <- plays |>
  left_join(games, by = "gameId") |>
  left_join(pff, by = c("gameId", "playId")) |>
  left_join(players |> select(nflId, displayName, Position), by = c("kickerId" = "nflId"))

# the joins should not add or drop any plays
stopifnot(nrow(df) == nrow(plays))

bind_rows(
  df |> count(Category = specialTeamsPlayType) |> mutate(Variable = "specialTeamsPlayType"),
  df |> count(Category = specialTeamsResult)   |> mutate(Variable = "specialTeamsResult")
) |>
  select(Variable, Category, Plays = n) |>
  kable(caption = paste0("Play counts in the merged data (", format(nrow(df), big.mark = ","),
                         " plays, ", ncol(df), " columns)"), format.args = list(big.mark = ",")) |>
  kable_styling(bootstrap_options = c("striped", "condensed"), full_width = FALSE) |>
  collapse_rows(columns = 1, valign = "top")
Play counts in the merged data (19,979 plays, 51 columns)
Variable Category Plays
specialTeamsPlayType Extra Point 3,488
Field Goal 2,657
Kickoff 7,843
Punt 5,991
specialTeamsResult Blocked Kick Attempt 61
Blocked Punt 39
Downed 834
Fair Catch 1,645
Kick Attempt Good 5,470
Kick Attempt No Good 585
Kickoff Team Recovery 16
Muffed 214
Non-Special Teams Result 101
Out of Bounds 651
Return 5,207
Touchback 5,156

Every count matches: 19,979 plays, including 2,657 field goals and 3,488 extra points. The merge keeps one row per play, so no plays were duplicated or lost along the way.

Which kicks count

fg <- plays |>
  filter(specialTeamsPlayType == "Field Goal",
         specialTeamsResult != "Non-Special Teams Result") |>          # drop fakes / botched snaps
  left_join(games, by = "gameId") |>
  left_join(players |> select(nflId, displayName), by = c("kickerId" = "nflId")) |>
  left_join(roofs, by = "gameId") |>
  left_join(kick_data |> select(gameId, playId, D, e, indoor, crosswind, temp, rain, snow),
            by = c("gameId", "playId")) |>
  mutate(
    made     = specialTeamsResult == "Kick Attempt Good",
    blocked  = specialTeamsResult == "Blocked Kick Attempt",
    home     = possessionTeam == homeTeamAbbr,
    margin   = ifelse(home, preSnapHomeScore - preSnapVisitorScore,
                            preSnapVisitorScore - preSnapHomeScore),
    pressure = quarter >= 4 & abs(margin) <= 8,                          # 4th qtr / OT, one-score game
    # game-level roof status from the kick data; fall back to the stadium's roof type
    outdoor  = coalesce(indoor == 0, RoofType == "Outdoor"),
    # blocked kicks have no recorded length: yards to the goal line + 17
    to_goal  = ifelse(yardlineSide == possessionTeam, 100 - yardlineNumber, yardlineNumber),
    distance = coalesce(as.numeric(D), as.numeric(kickLength), to_goal + 17)
  )

league <- fg |>
  summarize(attempts = n(),
            overall  = mean(made),
            pressure = mean(made[pressure]),
            road     = mean(made[!home]),
            outdoor  = mean(made[outdoor]),
            blocked  = mean(blocked))
  • Field goals only, from 2018–2020: 2,643 attempts.
  • Blocked kicks are kept and count as misses. Plays marked “Non-Special Teams Result” (fakes and botched snaps) are dropped, since the kicker never really got a kick off.
  • Kickers need at least 25 attempts across the three seasons, so one perfect afternoon can’t top the list. That leaves the kickers ranked below.

The metric

For each kicker I calculate five numbers:

Measure Definition League average
Overall accuracy makes ÷ attempts 83.9%
Pressure accuracy 4th quarter or overtime, kicking team within 8 points 83.2%
Road accuracy kicks where the kicking team was the visitor 83.1%
Outdoor accuracy kicks with no dome or closed roof 83.5%
Blocked rate blocked kicks ÷ attempts 1.4%

Reliability score = the average of the four accuracies − the blocked rate, shown on a 0–100 scale.

Some of these splits are small. A kicker may have only 3–19 pressure kicks in three seasons, and a 3-for-3 kicker should not automatically look perfect. So each split is smoothed: I add five “average” attempts at the league rate for that situation before dividing. A kicker with lots of attempts barely moves; one with only a handful is pulled toward the league average until he has shown more.

m <- 5                                               # phantom attempts used for smoothing
smooth <- function(makes, n, league_rate) (makes + m * league_rate) / (n + m)

kickers <- fg |>
  group_by(kickerId, kicker = displayName) |>
  summarize(attempts   = n(),
            makes      = sum(made),
            avg_dist   = mean(distance, na.rm = TRUE),
            pr_n  = sum(pressure),  pr_m  = sum(made[pressure]),
            rd_n  = sum(!home),     rd_m  = sum(made[!home]),
            hm_n  = sum(home),      hm_m  = sum(made[home]),
            out_n = sum(outdoor),   out_m = sum(made[outdoor]),
            blocked = sum(blocked),
            .groups = "drop") |>
  filter(attempts >= 25) |>
  mutate(overall      = makes / attempts,
         pressure_acc = smooth(pr_m,  pr_n,  league$pressure),
         road_acc     = smooth(rd_m,  rd_n,  league$road),
         outdoor_acc  = smooth(out_m, out_n, league$outdoor),
         home_raw     = hm_m / hm_n,
         road_raw     = rd_m / rd_n,
         blocked_rate = blocked / attempts,
         reliability  = 100 * ((overall + pressure_acc + road_acc + outdoor_acc) / 4 - blocked_rate),
         rank         = min_rank(desc(reliability)),
         raw_rank     = min_rank(desc(overall))) |>
  arrange(rank)

top10 <- kickers |> filter(rank <= 10)

Data Visualization


Visualization 1: The ten most reliable kickers

blue <- "#2a78d6"

ggplot(top10, aes(x = reliability, y = reorder(kicker, reliability))) +
  geom_col(fill = blue, width = 0.7) +
  geom_text(aes(label = sprintf("%.1f", reliability)), hjust = -0.2, size = 3.8) +
  scale_x_continuous(limits = c(0, 100), expand = expansion(mult = c(0, 0.02))) +
  labs(title = "Top 10 NFL Kickers by Reliability Score, 2018-2020",
       subtitle = "Average of overall, pressure, road and outdoor accuracy, minus blocked rate",
       x = "Reliability score (0-100)", y = NULL) +
  theme_minimal(base_size = 13) +
  theme(plot.title = element_text(face = "bold", hjust = 0.5),
        plot.subtitle = element_text(hjust = 0.5),
        panel.grid.major.y = element_blank())

top10 |>
  transmute(Rank = rank, Kicker = kicker, Attempts = attempts,
            `Raw accuracy rank` = raw_rank,
            Overall  = percent(overall, 0.1),
            Pressure = paste0(percent(pressure_acc, 0.1), " (", pr_n, ")"),
            Road     = paste0(percent(road_acc, 0.1), " (", rd_n, ")"),
            Outdoor  = paste0(percent(outdoor_acc, 0.1), " (", out_n, ")"),
            Blocked  = blocked,
            Score    = round(reliability, 1)) |>
  kable(caption = "Top 10 by reliability score. Split accuracies are smoothed; attempts in parentheses.",
        align = "rlrrrrrrrr")
Top 10 by reliability score. Split accuracies are smoothed; attempts in parentheses.
Rank Kicker Attempts Raw accuracy rank Overall Pressure Road Outdoor Blocked Score
1 Josh Lambo 57 1 94.7% 93.5% (8) 87.2% (25) 96.6% (49) 0 93.0
2 Justin Tucker 94 3 92.6% 96.2% (17) 93.7% (40) 91.4% (86) 2 91.3
3 Graham Gano 45 2 93.3% 88.5% (11) 89.0% (21) 93.6% (39) 0 91.1
4 Rodrigo Blankenship 36 12 86.1% 91.6% (5) 91.6% (17) 89.9% (13) 0 89.8
5 Jason Myers 80 4 91.2% 83.3% (12) 93.8% (41) 89.8% (72) 0 89.5
6 Jason Sanders 83 15 85.5% 95.3% (13) 92.3% (45) 83.6% (67) 0 89.2
7 Mason Crosby 71 9 88.7% 89.8% (13) 84.4% (39) 93.2% (51) 0 89.0
8 Nick Folk 42 7 90.5% 92.4% (6) 93.2% (22) 88.8% (38) 1 88.8
9 Harrison Butker 85 6 90.6% 84.2% (13) 92.0% (43) 90.7% (79) 1 88.2
10 Wil Lutz 86 8 89.5% 88.5% (11) 88.5% (46) 87.3% (33) 1 87.3

Josh Lambo comes out on top with a score of 93.0. He made 94.7% of his kicks and never had one blocked. Justin Tucker is a close second and has the best pressure numbers of the group (17 of 17 in close fourth quarters). The bigger story is who moves: Jason Sanders ranks only #15 on plain accuracy but #6 on reliability, because he was strong in exactly the situations this score rewards.


Visualization 2: Accuracy by situation for the top 10

situ_cols <- c("Overall" = "#2a78d6", "Pressure" = "#eb6834",
               "Road" = "#1baf7a", "Outdoor" = "#eda100")

top10 |>
  select(kicker, rank, Overall = overall, Pressure = pressure_acc,
         Road = road_acc, Outdoor = outdoor_acc) |>
  pivot_longer(Overall:Outdoor, names_to = "situation", values_to = "acc") |>
  mutate(situation = factor(situation, levels = names(situ_cols))) |>
  ggplot(aes(x = acc, y = reorder(kicker, -rank), colour = situation)) +
  geom_vline(xintercept = league$overall, linetype = "dashed", colour = "grey60") +
  geom_point(size = 3.5, position = position_dodge(width = 0.6)) +
  scale_colour_manual(values = situ_cols) +
  scale_x_continuous(labels = percent_format(1)) +
  labs(title = "How the Top 10 Hold Up in Hard Situations",
       subtitle = "Dashed line = league-wide field goal accuracy",
       x = "Field goal accuracy (pressure, road and outdoor are smoothed)", y = NULL,
       colour = "Situation") +
  theme_minimal(base_size = 13) +
  theme(plot.title = element_text(face = "bold", hjust = 0.5),
        plot.subtitle = element_text(hjust = 0.5),
        legend.position = "top")

I used dots rather than bars here because every accuracy falls between about 75% and 97%. Bars starting at zero would make the differences almost invisible. Most of the top 10 stay above the league average in every situation, which is the point of a reliability ranking: there is no situation where they fall apart.


Visualization 3: Home vs. road accuracy

ha <- kickers |>
  mutate(gap = road_raw - home_raw,
         kicker = reorder(kicker, gap))

ggplot(ha) +
  geom_segment(aes(x = home_raw, xend = road_raw, y = kicker, yend = kicker),
               colour = "grey70", linewidth = 1) +
  geom_point(aes(x = home_raw, y = kicker, colour = "Home"), size = 3) +
  geom_point(aes(x = road_raw, y = kicker, colour = "Road"), size = 3) +
  scale_colour_manual(values = c("Home" = "#2a78d6", "Road" = "#eb6834")) +
  scale_x_continuous(labels = percent_format(1)) +
  labs(title = "Home vs. Road Field Goal Accuracy",
       subtitle = "Sorted from biggest road drop-off (bottom) to biggest road improvement (top)",
       x = "Field goal accuracy (raw)", y = NULL, colour = NULL) +
  theme_minimal(base_size = 12) +
  theme(plot.title = element_text(face = "bold", hjust = 0.5),
        plot.subtitle = element_text(hjust = 0.5),
        legend.position = "top")

Across the league, kickers made 84.8% at home and 83.1% on the road, a small gap. Individually it varies a lot. Sam Ficken dropped from 86% at home to 65% on the road, while Tyler Bass was actually better away (88% vs. 67%). I expected dome kickers to struggle most on the road, but the pattern is mixed: most of the biggest drop-offs belong to kickers with outdoor home stadiums. Home-field comfort seems to be about the individual kicker more than the roof.


Visualization 4: Does playing outdoors matter?

fg |>
  filter(!is.na(outdoor)) |>
  mutate(bucket = cut(distance, c(0, 29, 39, 49, 70),
                      labels = c("Under 30", "30-39", "40-49", "50+")),
         setting = ifelse(outdoor, "Outdoor", "Dome / closed roof")) |>
  filter(!is.na(bucket)) |>
  group_by(bucket, setting) |>
  summarize(acc = mean(made), n = n(), .groups = "drop") |>
  ggplot(aes(x = bucket, y = acc, fill = setting)) +
  geom_col(position = position_dodge(width = 0.8), width = 0.75) +
  geom_text(aes(label = percent(acc, 1)), position = position_dodge(width = 0.8),
            vjust = -0.4, size = 3.6) +
  scale_fill_manual(values = c("Dome / closed roof" = "#2a78d6", "Outdoor" = "#eb6834")) +
  scale_y_continuous(labels = percent_format(1), limits = c(0, 1.05)) +
  labs(title = "League-Wide Accuracy by Distance, Indoors vs. Outdoors",
       x = "Kick distance (yards)", y = "Field goal accuracy", fill = NULL) +
  theme_minimal(base_size = 13) +
  theme(plot.title = element_text(face = "bold", hjust = 0.5), legend.position = "top")

The gap between indoor and outdoor kicking shows up mostly in the 30–39 yard range. Short kicks are nearly automatic everywhere. From 50+ yards outdoor kickers do slightly better, likely because coaches only send kickers out for long tries when the conditions allow it. Because conditions can matter, I score outdoor accuracy separately instead of letting easier indoor kicks boost a kicker’s overall number.


Visualization 5: Blocked kicks

kickers |>
  filter(blocked > 0) |>
  ggplot(aes(x = blocked_rate, y = reorder(kicker, blocked_rate))) +
  geom_col(fill = "#2a78d6", width = 0.7) +
  geom_text(aes(label = paste0(blocked, " of ", attempts)), hjust = -0.15, size = 3.4) +
  scale_x_continuous(labels = percent_format(0.1), expand = expansion(mult = c(0, 0.25))) +
  labs(title = "Blocked Field Goal Rate (kickers with at least one block)",
       x = "Blocked kicks / attempts", y = NULL) +
  theme_minimal(base_size = 12) +
  theme(plot.title = element_text(face = "bold", hjust = 0.5),
        panel.grid.major.y = element_blank())

Blocks are rare: 37 of 2,643 field goals (1.4%), and 15 of the 38 ranked kickers never had one blocked. A block is partly the protection unit’s fault, but a kicker whose kicks come out low gets blocked more often, so a small penalty is fair. Because the counts are so small, the blocked rate only nudges the score and never decides it.


Visualization 6: Consistency season by season

A three-season total can hide one great year and two average ones. This chart splits each top-10 kicker’s accuracy by season.

season_acc <- fg |>
  filter(kickerId %in% top10$kickerId) |>
  group_by(kicker = displayName, season) |>
  summarize(acc = mean(made), n = n(), .groups = "drop") |>
  left_join(top10 |> select(kicker, rank), by = "kicker") |>
  mutate(kicker = reorder(gsub(" ", "\n", kicker), rank))

season_sd <- season_acc |>
  group_by(kicker) |>
  summarize(seasons = n(), spread = max(acc) - min(acc), .groups = "drop")

ggplot(season_acc, aes(x = kicker, y = acc, fill = factor(season))) +
  geom_col(position = position_dodge(width = 0.85), width = 0.8) +
  geom_text(aes(label = n), position = position_dodge(width = 0.85),
            vjust = -0.4, size = 2.8, colour = "grey30") +
  geom_hline(yintercept = league$overall, linetype = "dashed", colour = "grey40") +
  scale_fill_manual(values = c("2018" = "#9ec5f0", "2019" = "#2a78d6", "2020" = "#0d3d75")) +
  scale_y_continuous(labels = percent_format(1), limits = c(0, 1.08),
                     breaks = seq(0, 1, 0.25)) +
  labs(title = "Field Goal Accuracy by Season, Top 10 Reliable Kickers",
       subtitle = "Ordered by reliability rank. Numbers above bars = attempts. Dashed line = league average.",
       x = NULL, y = "Field goal accuracy", fill = "Season") +
  theme_minimal(base_size = 12) +
  theme(plot.title = element_text(face = "bold", hjust = 0.5),
        plot.subtitle = element_text(hjust = 0.5),
        legend.position = "top", panel.grid.major.x = element_blank())

steady <- season_sd |> filter(seasons == 3) |> arrange(spread) |> slice(1)
swingy <- season_sd |> filter(seasons == 3) |> arrange(-spread) |> slice(1)

Even the most reliable kickers have off years: 5 of the top 10 dipped below the league average in at least one season. Harrison Butker was the steadiest kicker with all three seasons: his best and worst years were only 3 percentage points apart. Mason Crosby swung the most (17 points). Some kickers only appear in one or two seasons because they changed teams or were injured, and a few seasons rest on small numbers of attempts (see the numbers above the bars), so one bad season can look worse than it really was.


Visualization 7: Down the middle or squeaking by?

Made vs. missed treats a kick that barely slips inside the upright the same as one that splits the posts. The kick-level file includes where each field goal crossed the goal line (offset from the center of the uprights, in yards, measured from the ball-tracking data). The uprights are 18.5 feet apart, so a kick is between the posts when it crosses less than about 3.08 yards from center.

post <- 18.5 / 2 / 3                                  # half the goalpost width, in yards

offsets <- fg |>
  filter(!is.na(e)) |>
  mutate(result = ifelse(made, "Good", "No good"))

ggplot(offsets, aes(x = e, y = distance, colour = result)) +
  annotate("rect", xmin = -post, xmax = post, ymin = -Inf, ymax = Inf,
           fill = "#2a78d6", alpha = 0.06) +
  geom_vline(xintercept = c(-post, post), linetype = "dashed", linewidth = 1) +
  geom_point(alpha = 0.55, size = 1.8, position = position_jitter(height = 0.3, seed = 1)) +
  scale_colour_manual(values = c("Good" = "#2a78d6", "No good" = "#eb6834")) +
  scale_x_continuous(limits = c(-10, 10), breaks = seq(-9, 9, 3)) +
  labs(title = "Where Every Field Goal Crossed the Goal Line, 2018-2020",
       subtitle = "Dashed lines = uprights. Negative = left of center, positive = right of center.",
       x = "Offset from center of the uprights (yards)", y = "Kick distance (yards)", colour = NULL) +
  theme_minimal(base_size = 12) +
  theme(plot.title = element_text(face = "bold", hjust = 0.5),
        plot.subtitle = element_text(hjust = 0.5), legend.position = "top")

precision <- offsets |>
  group_by(kickerId) |>
  summarize(med_offset = median(abs(e)),
            middle     = mean(abs(e) < 1),            # within 1 yard of dead center
            n_tracked  = n(), .groups = "drop") |>
  inner_join(kickers |> select(kickerId, kicker, rank), by = "kickerId") |>
  filter(n_tracked >= 20) |>
  mutate(p_rank = min_rank(med_offset)) |>
  arrange(p_rank)

precision |>
  slice(1:10) |>
  transmute(`Precision rank` = p_rank, Kicker = kicker,
            `Median offset (yds)` = round(med_offset, 2),
            `Within 1 yd of center` = percent(middle, 1),
            `Tracked kicks` = n_tracked,
            `Reliability rank` = rank) |>
  kable(caption = "Most precise kickers: median distance from the center of the uprights (kickers with 20+ tracked field goals)",
        align = "rlrrrr") |>
  kable_styling(bootstrap_options = c("striped", "condensed"), full_width = FALSE)
Most precise kickers: median distance from the center of the uprights (kickers with 20+ tracked field goals)
Precision rank Kicker Median offset (yds) Within 1 yd of center Tracked kicks Reliability rank
1 Matt Prater 0.84 58% 78 24
2 Younghoe Koo 0.91 53% 57 13
3 Justin Tucker 0.94 52% 91 2
4 Josh Lambo 0.94 54% 57 1
5 Harrison Butker 0.99 50% 84 9
6 Rodrigo Blankenship 1.02 50% 36 4
7 Robbie Gould 1.03 49% 77 19
8 Jake Elliott 1.08 48% 66 22
9 Wil Lutz 1.11 44% 85 10
10 Stephen Gostkowski 1.13 46% 63 33

The makes cluster tightly around the center, and the misses sit just outside the dashed lines. Most misses are near-misses, not shanks. The few orange points between the posts are kicks that came up short or were deflected.

Matt Prater has the most precise median kick: 0.84 yards from dead center, with 58% of his kicks crossing within a yard of the middle. But he ranks only #24 on reliability, with 82.3% accuracy overall. When his kicks were good they were usually perfect, but he missed too often for that to make him the best. My top two reliability kickers come next: Josh Lambo (#4 on precision) and Justin Tucker (#3). They weren’t squeaking kicks inside the uprights; they were splitting them. Graham Gano is the opposite case: #32 on precision, because more of his makes slipped in near a post. Accuracy and precision overlap, but they are not the same thing.

Adjusting for Distance: A Binary Logit Model

The reliability score has one blind spot: it treats a 25-yard kick and a 55-yard kick the same. A kicker whose coach sends him out for long attempts will miss more often, and that can look like a weakness when it is really a sign of confidence in his leg. To check this, I modeled the chance that a field goal is good.

The outcome is binary (made or missed), so ordinary regression doesn’t fit: it can predict probabilities below 0 or above 1. A binary logit model (glm(..., family = binomial)) predicts the log-odds of a make, so every prediction stays between 0 and 1. I built it up in three steps, adding one group of variables at a time:

  1. Model 1: distance only.
  2. Model 2: distance plus conditions: outdoor stadium, crosswind (mph), cold (degrees below 50°F) and rain or snow.
  3. Model 3: Model 2 plus the game situation from my reliability score: pressure kick and home kick.
fg_model <- fg |>
  mutate(crosswind = coalesce(as.numeric(crosswind), 0),
         cold      = pmax(0, 50 - coalesce(as.numeric(temp), 70)),   # 0 indoors / warm days
         wet       = (coalesce(rain, 0L) + coalesce(snow, 0L)) > 0) |>
  filter(!is.na(distance), !is.na(outdoor))

model1 <- glm(made ~ distance, data = fg_model, family = binomial)
model2 <- glm(made ~ distance + outdoor + crosswind + cold + wet, data = fg_model, family = binomial)
model3 <- update(model2, . ~ . + pressure + home)

# odds ratios: exp(coefficient). Below 1 lowers the odds of a make, above 1 raises them
tidy_model <- function(m, name) {
  s <- summary(m)$coefficients
  data.frame(term = rownames(s), model = name,
             cell = sprintf("%.3f%s", exp(s[, 1]),
                            ifelse(s[, 4] < 0.001, "***", ifelse(s[, 4] < 0.01, "**",
                                   ifelse(s[, 4] < 0.05, "*", "")))))
}

labels <- c(distance = "Distance (per yard)", outdoorTRUE = "Outdoor stadium",
            crosswind = "Crosswind (per mph)", cold = "Cold (per degree below 50 F)",
            wetTRUE = "Rain or snow", pressureTRUE = "Pressure kick", homeTRUE = "Home kick")

bind_rows(tidy_model(model1, "Model 1"), tidy_model(model2, "Model 2"), tidy_model(model3, "Model 3")) |>
  filter(term != "(Intercept)") |>
  pivot_wider(names_from = model, values_from = cell, values_fill = "") |>
  mutate(term = labels[term]) |>
  bind_rows(data.frame(term = "AIC (lower is better)",
                       `Model 1` = sprintf("%.0f", AIC(model1)),
                       `Model 2` = sprintf("%.0f", AIC(model2)),
                       `Model 3` = sprintf("%.0f", AIC(model3)), check.names = FALSE)) |>
  rename(Variable = term) |>
  kable(caption = "Odds ratios from the binary logit models (*** p < 0.001, ** p < 0.01, * p < 0.05)",
        align = "lrrr") |>
  kable_styling(bootstrap_options = c("striped", "condensed"), full_width = FALSE)
Odds ratios from the binary logit models (*** p < 0.001, ** p < 0.01, * p < 0.05)
Variable Model 1 Model 2 Model 3
Distance (per yard) 0.894*** 0.895*** 0.895***
Outdoor stadium 0.801 0.808
Crosswind (per mph) 1.019 1.019
Cold (per degree below 50 F) 1.015 1.015
Rain or snow 1.030 1.029
Pressure kick 1.040
Home kick 1.187
AIC (lower is better) 1987 1992 1993
d_or  <- exp(coef(model1)["distance"])
p_out <- summary(model2)$coefficients["outdoorTRUE", 4]

Distance is the only variable that matters. Each extra yard multiplies the odds of a make by 0.894, which means about 10.6% lower odds per yard. The effect compounds instead of adding up in a straight line: a 50-yard kick has roughly 33% of the odds of a 40-yard kick. None of the condition or situation variables is significant once distance is in the model (outdoor p = 0.12), and Model 1 has the lowest AIC. In other words, in this data the gap between indoor and outdoor kicking from Visualization 4 is mostly explained by which kicks get attempted, not the conditions themselves.

That finding lets me ask a cleaner question: given the distance of every kick he attempted, how many more field goals did each kicker make than an average NFL kicker would have? I call this field goals over expected (FGOE). For each kick, Model 1 gives the probability that an average kicker makes it. FGOE is the sum of (made − expected) over all of a kicker’s attempts. I also show it per 100 attempts so kickers with different workloads can be compared.

fg_model$expected <- predict(model1, fg_model, type = "response")

fgoe <- fg_model |>
  group_by(kickerId) |>
  summarize(exp_pct = mean(expected),
            fgoe    = sum(made - expected),
            .groups = "drop") |>
  inner_join(kickers |> select(kickerId, kicker, attempts, overall, avg_dist, reliability, rank),
             by = "kickerId") |>
  mutate(fgoe100 = 100 * fgoe / attempts,
         x_rank  = min_rank(desc(fgoe100))) |>
  arrange(x_rank)

rank_cor <- cor(fgoe$x_rank, fgoe$rank, method = "spearman")
big_mover <- fgoe |> mutate(jump = rank - x_rank) |> arrange(-jump) |> slice(1)
big_faller <- fgoe |> mutate(jump = rank - x_rank) |> arrange(jump) |> slice(1)

Visualization 8: Field goals over expected

ggplot(fgoe, aes(x = fgoe100, y = reorder(kicker, fgoe100),
                 fill = ifelse(fgoe100 >= 0, "Above expected", "Below expected"))) +
  geom_col(width = 0.75) +
  geom_vline(xintercept = 0, colour = "grey30") +
  geom_text(aes(label = sprintf("%+.1f", fgoe),
                hjust = ifelse(fgoe100 >= 0, -0.2, 1.2)), size = 3.2) +
  scale_fill_manual(values = c("Above expected" = "#2a78d6", "Below expected" = "#eb6834")) +
  scale_x_continuous(expand = expansion(mult = 0.15)) +
  labs(title = "Field Goals Made Over Expected per 100 Attempts, 2018-2020",
       subtitle = "Expected = make probability for each kick's distance from the logit model. Labels show total FGOE.",
       x = "Field goals over expected per 100 attempts", y = NULL, fill = NULL) +
  theme_minimal(base_size = 12) +
  theme(plot.title = element_text(face = "bold", hjust = 0.5),
        plot.subtitle = element_text(hjust = 0.5),
        legend.position = "top", panel.grid.major.y = element_blank())

Graham Gano leads on a per-attempt basis (+11.2 per 100 kicks), and Justin Tucker has the most total value: 8.4 more field goals than an average kicker would have made on the same kicks, worth about 25 points. The biggest change from my reliability ranking is Brandon McManus, who rises from #25 to #5. His raw accuracy (85.7%) looks ordinary, but his average attempt was 41.4 yards, among the longest in the league. Going the other way, Austin Seibert falls from #15 to #32 because his kicks were shorter than average.


Visualization 9: Reliability vs. distance-adjusted accuracy (interactive)

Hover over a point to see the kicker’s numbers. Kickers in the top-right corner are strong on both measures.

g <- ggplot(fgoe, aes(x = fgoe100, y = reliability, size = attempts,
                      text = paste0("<b>", kicker, "</b>",
                                    "<br>Reliability: ", sprintf("%.1f", reliability), " (rank ", rank, ")",
                                    "<br>FGOE per 100: ", sprintf("%+.1f", fgoe100), " (rank ", x_rank, ")",
                                    "<br>Accuracy: ", percent(overall, 0.1),
                                    "<br>Avg distance: ", sprintf("%.1f", avg_dist), " yds",
                                    "<br>Attempts: ", attempts))) +
  geom_vline(xintercept = 0, linetype = "dashed", colour = "grey60") +
  geom_hline(yintercept = mean(fgoe$reliability), linetype = "dashed", colour = "grey60") +
  geom_point(colour = "#2a78d6", alpha = 0.7) +
  geom_text(data = fgoe |> filter(x_rank <= 5 | rank <= 5),
            aes(label = kicker), size = 3, nudge_y = 0.8, colour = "grey20") +
  scale_size_continuous(range = c(2, 7), guide = "none") +
  labs(title = "Reliability Score vs. Field Goals Over Expected",
       x = "Field goals over expected per 100 attempts", y = "Reliability score") +
  theme_minimal(base_size = 12) +
  theme(plot.title = element_text(face = "bold", hjust = 0.5))

ggplotly(g, tooltip = "text") |> config(displayModeBar = FALSE)

The two rankings mostly agree (rank correlation = 0.83). Graham Gano, Josh Lambo, Justin Tucker, and Jason Myers are in the top five of both. That gives me more confidence in the result: these kickers come out on top whether you reward performance in hard situations or performance on hard kicks.

Full Rankings

kickers |>
  transmute(Rank = rank, Kicker = kicker, Attempts = attempts,
            `Avg distance` = round(avg_dist, 1),
            Overall  = percent(overall, 0.1),
            Pressure = percent(pressure_acc, 0.1),
            Road     = percent(road_acc, 0.1),
            Outdoor  = percent(outdoor_acc, 0.1),
            Blocked  = blocked,
            Score    = round(reliability, 1)) |>
  kable(caption = "All kickers with at least 25 field goal attempts, 2018-2020",
        align = "rlrrrrrrrr")
All kickers with at least 25 field goal attempts, 2018-2020
Rank Kicker Attempts Avg distance Overall Pressure Road Outdoor Blocked Score
1 Josh Lambo 57 37.8 94.7% 93.5% 87.2% 96.6% 0 93.0
2 Justin Tucker 94 38.1 92.6% 96.2% 93.7% 91.4% 2 91.3
3 Graham Gano 45 39.9 93.3% 88.5% 89.0% 93.6% 0 91.1
4 Rodrigo Blankenship 36 35.7 86.1% 91.6% 91.6% 89.9% 0 89.8
5 Jason Myers 80 39.4 91.2% 83.3% 93.8% 89.8% 0 89.5
6 Jason Sanders 83 39.1 85.5% 95.3% 92.3% 83.6% 0 89.2
7 Mason Crosby 71 39.5 88.7% 89.8% 84.4% 93.2% 0 89.0
8 Nick Folk 42 36.9 90.5% 92.4% 93.2% 88.8% 1 88.8
9 Harrison Butker 85 36.1 90.6% 84.2% 92.0% 90.7% 1 88.2
10 Wil Lutz 86 37.6 89.5% 88.5% 88.5% 87.3% 1 87.3
11 Matt Bryant 34 41.4 82.4% 89.5% 89.1% 84.8% 0 86.4
12 Dustin Hopkins 86 39.2 84.9% 92.0% 89.7% 83.8% 1 86.4
13 Younghoe Koo 57 36.8 91.2% 77.0% 90.8% 85.9% 0 86.2
14 Aldrick Rosas 52 36.2 86.5% 84.7% 84.9% 84.4% 0 85.1
15 Austin Seibert 34 38.9 79.4% 90.7% 84.0% 82.4% 0 84.1
16 Chris Boswell 63 36.1 87.3% 85.8% 81.5% 85.8% 1 83.5
17 Randy Bullock 77 39.1 85.7% 79.8% 80.8% 85.8% 0 83.0
18 Ryan Succop 62 36.2 83.9% 90.3% 79.4% 86.5% 2 81.8
19 Robbie Gould 79 37.0 83.5% 86.5% 80.7% 84.8% 2 81.3
20 Greg Zuerlein 93 39.4 81.7% 89.2% 76.5% 81.8% 1 81.2
21 Tyler Bass 28 40.1 78.6% 83.3% 86.4% 76.5% 0 81.2
22 Jake Elliott 66 38.2 81.8% 77.4% 82.8% 80.9% 0 80.7
23 Ka’imi Fairbairn 96 38.0 84.4% 87.1% 76.8% 72.9% 0 80.3
24 Matt Prater 79 38.1 82.3% 78.7% 81.9% 81.8% 1 79.9
25 Brandon McManus 84 41.4 85.7% 73.1% 83.3% 85.6% 2 79.5
26 Cody Parkey 54 37.6 81.5% 76.3% 85.8% 81.1% 1 79.3
27 Matt Gay 44 40.2 79.5% 82.2% 82.7% 79.4% 1 78.7
28 Daniel Carlson 77 35.9 83.1% 79.8% 74.8% 76.7% 0 78.6
29 Cairo Santos 51 37.4 80.4% 71.5% 87.2% 80.4% 1 77.9
30 Dan Bailey 72 37.2 81.9% 79.7% 75.9% 76.8% 1 77.2
31 Sam Ficken 39 39.6 76.9% 93.0% 68.9% 76.0% 1 76.1
32 Michael Badgley 62 38.8 79.0% 74.5% 72.9% 83.7% 1 75.9
33 Stephen Gostkowski 65 38.7 78.5% 83.3% 76.6% 76.1% 2 75.5
34 Stephen Hauschka 55 39.9 78.2% 74.2% 74.7% 80.4% 1 75.0
35 Brett Maher 62 38.6 75.8% 86.9% 76.1% 72.1% 2 74.5
36 Adam Vinatieri 49 38.0 77.6% 69.7% 80.6% 79.0% 2 72.7
37 Zane Gonzalez 64 38.4 78.1% 65.6% 79.3% 78.7% 2 72.3
38 Joey Slye 55 40.6 76.4% 69.7% 78.8% 74.1% 2 71.1

Conclusion

  • The most reliable kicker of 2018–2020 was Josh Lambo, with Justin Tucker and Graham Gano close behind. All three were accurate overall and held up under pressure, on the road and outdoors.
  • Raw accuracy and reliability mostly agree at the very top, but the middle of the table shuffles. Kickers who pile up makes in comfortable conditions slide, and kickers who deliver in close games and bad weather climb.
  • Situation matters more for some kickers than others. The home/road chart shows double-digit drop-offs for a few kickers, which a single percentage would never reveal.
  • Across the league, distance explains whether a kick is made; conditions add little. The logit model shows about 10.6% lower odds per extra yard, while roof, wind, cold, rain/snow, pressure and home field are not significant once distance is accounted for.
  • The answer holds up after adjusting for distance. Graham Gano, Josh Lambo, Justin Tucker, and Jason Myers rank in the top five on both reliability and field goals over expected. If I had to pick one kicker, Justin Tucker stands out: near the top of both rankings, with the most total field goals over expected, the strongest pressure record of the group, and one of the most centered kicks in the league (#3 on precision).

Limitations and next steps. Pressure and outdoor samples are small, even with smoothing. The logit model covers three seasons with only about 2,643 kicks, so it can’t reliably separate small weather effects, and stadium weather is measured at the nearest weather station, not on the field. Next steps would be adding kicker fixed effects to the model, interaction terms (for example distance × crosswind, since wind should matter more on long kicks), and more features from the full ball-tracking data, such as how much each kick curved in flight and how far it drifted in the first half-second after contact.