This assignment uses the 2022 NFL Big Data Bowl data, which covers every special teams play from the 2018–2020 NFL seasons, to answer one question: who is the best kicker in the NFL? There is no single right definition of “best,” so the first job is to choose one and defend it.
Most kicker rankings come down to one number: field goal percentage. That number hides when and where the kicks happened. A 90% kicker who plays half his games in a dome and rarely kicks with the game on the line has had an easier job than an 87% kicker who kicks outdoors in December and in close fourth quarters.
For this assignment I define the best kicker as the most reliable kicker in hard situations. Using every field goal attempt from the 2018–2020 seasons, I measure how accurate each kicker stays:
and how often his kicks get blocked. These combine into a single reliability score.
load_packages <- function(pkgs) {
missing <- pkgs[!pkgs %in% rownames(installed.packages())]
if (length(missing) > 0) install.packages(missing, dependencies = TRUE, quiet = TRUE,
repos = "https://cloud.r-project.org")
invisible(suppressPackageStartupMessages(lapply(pkgs, library, character.only = TRUE)))
}
load_packages(c("data.table", "dplyr", "tidyr", "ggplot2", "scales", "knitr",
"kableExtra", "plotly"))
# --- 2022 NFL Big Data Bowl files (special teams, 2018-2020) ---
plays <- fread(file.path(data_dir, "plays.csv"))
games <- fread(file.path(data_dir, "games.csv"))
players <- fread(file.path(data_dir, "players.csv"))
# --- Kick-level data with distance and game-time weather (one row per made/missed FG or XP) ---
kick_data <- fread(file.path(data_dir, "kick_model_data.csv"))
# --- Stadium roof types from the WeatherData repository (github.com/ThompsonJamesBliss/WeatherData) ---
wx_url <- "https://raw.githubusercontent.com/ThompsonJamesBliss/WeatherData/master/data/"
wx_games <- fread(paste0(wx_url, "games.csv"))
stadiums <- fread(paste0(wx_url, "stadium_coordinates.csv")) |>
distinct(StadiumName, RoofType) # NYG/NYJ share a stadium
roofs <- wx_games |>
select(gameId = game_id, StadiumName) |>
left_join(stadiums, by = "StadiumName")
I combined three sources:
plays,
games, players): the result of every kick, the
score and quarter before the snap, which team was kicking, and who was
at home.Before building anything, I merged the four Big Data Bowl files
(games, plays, players,
PFFScoutingData) into one data frame and checked the play
counts against the ones given in the assignment.
pff <- fread(file.path(data_dir, "PFFScoutingData.csv"))
df <- plays |>
left_join(games, by = "gameId") |>
left_join(pff, by = c("gameId", "playId")) |>
left_join(players |> select(nflId, displayName, Position), by = c("kickerId" = "nflId"))
# the joins should not add or drop any plays
stopifnot(nrow(df) == nrow(plays))
bind_rows(
df |> count(Category = specialTeamsPlayType) |> mutate(Variable = "specialTeamsPlayType"),
df |> count(Category = specialTeamsResult) |> mutate(Variable = "specialTeamsResult")
) |>
select(Variable, Category, Plays = n) |>
kable(caption = paste0("Play counts in the merged data (", format(nrow(df), big.mark = ","),
" plays, ", ncol(df), " columns)"), format.args = list(big.mark = ",")) |>
kable_styling(bootstrap_options = c("striped", "condensed"), full_width = FALSE) |>
collapse_rows(columns = 1, valign = "top")
| Variable | Category | Plays |
|---|---|---|
| specialTeamsPlayType | Extra Point | 3,488 |
| Field Goal | 2,657 | |
| Kickoff | 7,843 | |
| Punt | 5,991 | |
| specialTeamsResult | Blocked Kick Attempt | 61 |
| Blocked Punt | 39 | |
| Downed | 834 | |
| Fair Catch | 1,645 | |
| Kick Attempt Good | 5,470 | |
| Kick Attempt No Good | 585 | |
| Kickoff Team Recovery | 16 | |
| Muffed | 214 | |
| Non-Special Teams Result | 101 | |
| Out of Bounds | 651 | |
| Return | 5,207 | |
| Touchback | 5,156 |
Every count matches: 19,979 plays, including 2,657 field goals and 3,488 extra points. The merge keeps one row per play, so no plays were duplicated or lost along the way.
fg <- plays |>
filter(specialTeamsPlayType == "Field Goal",
specialTeamsResult != "Non-Special Teams Result") |> # drop fakes / botched snaps
left_join(games, by = "gameId") |>
left_join(players |> select(nflId, displayName), by = c("kickerId" = "nflId")) |>
left_join(roofs, by = "gameId") |>
left_join(kick_data |> select(gameId, playId, D, e, indoor, crosswind, temp, rain, snow),
by = c("gameId", "playId")) |>
mutate(
made = specialTeamsResult == "Kick Attempt Good",
blocked = specialTeamsResult == "Blocked Kick Attempt",
home = possessionTeam == homeTeamAbbr,
margin = ifelse(home, preSnapHomeScore - preSnapVisitorScore,
preSnapVisitorScore - preSnapHomeScore),
pressure = quarter >= 4 & abs(margin) <= 8, # 4th qtr / OT, one-score game
# game-level roof status from the kick data; fall back to the stadium's roof type
outdoor = coalesce(indoor == 0, RoofType == "Outdoor"),
# blocked kicks have no recorded length: yards to the goal line + 17
to_goal = ifelse(yardlineSide == possessionTeam, 100 - yardlineNumber, yardlineNumber),
distance = coalesce(as.numeric(D), as.numeric(kickLength), to_goal + 17)
)
league <- fg |>
summarize(attempts = n(),
overall = mean(made),
pressure = mean(made[pressure]),
road = mean(made[!home]),
outdoor = mean(made[outdoor]),
blocked = mean(blocked))
For each kicker I calculate five numbers:
| Measure | Definition | League average |
|---|---|---|
| Overall accuracy | makes ÷ attempts | 83.9% |
| Pressure accuracy | 4th quarter or overtime, kicking team within 8 points | 83.2% |
| Road accuracy | kicks where the kicking team was the visitor | 83.1% |
| Outdoor accuracy | kicks with no dome or closed roof | 83.5% |
| Blocked rate | blocked kicks ÷ attempts | 1.4% |
Reliability score = the average of the four accuracies − the blocked rate, shown on a 0–100 scale.
Some of these splits are small. A kicker may have only 3–19 pressure kicks in three seasons, and a 3-for-3 kicker should not automatically look perfect. So each split is smoothed: I add five “average” attempts at the league rate for that situation before dividing. A kicker with lots of attempts barely moves; one with only a handful is pulled toward the league average until he has shown more.
m <- 5 # phantom attempts used for smoothing
smooth <- function(makes, n, league_rate) (makes + m * league_rate) / (n + m)
kickers <- fg |>
group_by(kickerId, kicker = displayName) |>
summarize(attempts = n(),
makes = sum(made),
avg_dist = mean(distance, na.rm = TRUE),
pr_n = sum(pressure), pr_m = sum(made[pressure]),
rd_n = sum(!home), rd_m = sum(made[!home]),
hm_n = sum(home), hm_m = sum(made[home]),
out_n = sum(outdoor), out_m = sum(made[outdoor]),
blocked = sum(blocked),
.groups = "drop") |>
filter(attempts >= 25) |>
mutate(overall = makes / attempts,
pressure_acc = smooth(pr_m, pr_n, league$pressure),
road_acc = smooth(rd_m, rd_n, league$road),
outdoor_acc = smooth(out_m, out_n, league$outdoor),
home_raw = hm_m / hm_n,
road_raw = rd_m / rd_n,
blocked_rate = blocked / attempts,
reliability = 100 * ((overall + pressure_acc + road_acc + outdoor_acc) / 4 - blocked_rate),
rank = min_rank(desc(reliability)),
raw_rank = min_rank(desc(overall))) |>
arrange(rank)
top10 <- kickers |> filter(rank <= 10)
blue <- "#2a78d6"
ggplot(top10, aes(x = reliability, y = reorder(kicker, reliability))) +
geom_col(fill = blue, width = 0.7) +
geom_text(aes(label = sprintf("%.1f", reliability)), hjust = -0.2, size = 3.8) +
scale_x_continuous(limits = c(0, 100), expand = expansion(mult = c(0, 0.02))) +
labs(title = "Top 10 NFL Kickers by Reliability Score, 2018-2020",
subtitle = "Average of overall, pressure, road and outdoor accuracy, minus blocked rate",
x = "Reliability score (0-100)", y = NULL) +
theme_minimal(base_size = 13) +
theme(plot.title = element_text(face = "bold", hjust = 0.5),
plot.subtitle = element_text(hjust = 0.5),
panel.grid.major.y = element_blank())
top10 |>
transmute(Rank = rank, Kicker = kicker, Attempts = attempts,
`Raw accuracy rank` = raw_rank,
Overall = percent(overall, 0.1),
Pressure = paste0(percent(pressure_acc, 0.1), " (", pr_n, ")"),
Road = paste0(percent(road_acc, 0.1), " (", rd_n, ")"),
Outdoor = paste0(percent(outdoor_acc, 0.1), " (", out_n, ")"),
Blocked = blocked,
Score = round(reliability, 1)) |>
kable(caption = "Top 10 by reliability score. Split accuracies are smoothed; attempts in parentheses.",
align = "rlrrrrrrrr")
| Rank | Kicker | Attempts | Raw accuracy rank | Overall | Pressure | Road | Outdoor | Blocked | Score |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Josh Lambo | 57 | 1 | 94.7% | 93.5% (8) | 87.2% (25) | 96.6% (49) | 0 | 93.0 |
| 2 | Justin Tucker | 94 | 3 | 92.6% | 96.2% (17) | 93.7% (40) | 91.4% (86) | 2 | 91.3 |
| 3 | Graham Gano | 45 | 2 | 93.3% | 88.5% (11) | 89.0% (21) | 93.6% (39) | 0 | 91.1 |
| 4 | Rodrigo Blankenship | 36 | 12 | 86.1% | 91.6% (5) | 91.6% (17) | 89.9% (13) | 0 | 89.8 |
| 5 | Jason Myers | 80 | 4 | 91.2% | 83.3% (12) | 93.8% (41) | 89.8% (72) | 0 | 89.5 |
| 6 | Jason Sanders | 83 | 15 | 85.5% | 95.3% (13) | 92.3% (45) | 83.6% (67) | 0 | 89.2 |
| 7 | Mason Crosby | 71 | 9 | 88.7% | 89.8% (13) | 84.4% (39) | 93.2% (51) | 0 | 89.0 |
| 8 | Nick Folk | 42 | 7 | 90.5% | 92.4% (6) | 93.2% (22) | 88.8% (38) | 1 | 88.8 |
| 9 | Harrison Butker | 85 | 6 | 90.6% | 84.2% (13) | 92.0% (43) | 90.7% (79) | 1 | 88.2 |
| 10 | Wil Lutz | 86 | 8 | 89.5% | 88.5% (11) | 88.5% (46) | 87.3% (33) | 1 | 87.3 |
Josh Lambo comes out on top with a score of 93.0. He made 94.7% of his kicks and never had one blocked. Justin Tucker is a close second and has the best pressure numbers of the group (17 of 17 in close fourth quarters). The bigger story is who moves: Jason Sanders ranks only #15 on plain accuracy but #6 on reliability, because he was strong in exactly the situations this score rewards.
situ_cols <- c("Overall" = "#2a78d6", "Pressure" = "#eb6834",
"Road" = "#1baf7a", "Outdoor" = "#eda100")
top10 |>
select(kicker, rank, Overall = overall, Pressure = pressure_acc,
Road = road_acc, Outdoor = outdoor_acc) |>
pivot_longer(Overall:Outdoor, names_to = "situation", values_to = "acc") |>
mutate(situation = factor(situation, levels = names(situ_cols))) |>
ggplot(aes(x = acc, y = reorder(kicker, -rank), colour = situation)) +
geom_vline(xintercept = league$overall, linetype = "dashed", colour = "grey60") +
geom_point(size = 3.5, position = position_dodge(width = 0.6)) +
scale_colour_manual(values = situ_cols) +
scale_x_continuous(labels = percent_format(1)) +
labs(title = "How the Top 10 Hold Up in Hard Situations",
subtitle = "Dashed line = league-wide field goal accuracy",
x = "Field goal accuracy (pressure, road and outdoor are smoothed)", y = NULL,
colour = "Situation") +
theme_minimal(base_size = 13) +
theme(plot.title = element_text(face = "bold", hjust = 0.5),
plot.subtitle = element_text(hjust = 0.5),
legend.position = "top")
I used dots rather than bars here because every accuracy falls between about 75% and 97%. Bars starting at zero would make the differences almost invisible. Most of the top 10 stay above the league average in every situation, which is the point of a reliability ranking: there is no situation where they fall apart.
ha <- kickers |>
mutate(gap = road_raw - home_raw,
kicker = reorder(kicker, gap))
ggplot(ha) +
geom_segment(aes(x = home_raw, xend = road_raw, y = kicker, yend = kicker),
colour = "grey70", linewidth = 1) +
geom_point(aes(x = home_raw, y = kicker, colour = "Home"), size = 3) +
geom_point(aes(x = road_raw, y = kicker, colour = "Road"), size = 3) +
scale_colour_manual(values = c("Home" = "#2a78d6", "Road" = "#eb6834")) +
scale_x_continuous(labels = percent_format(1)) +
labs(title = "Home vs. Road Field Goal Accuracy",
subtitle = "Sorted from biggest road drop-off (bottom) to biggest road improvement (top)",
x = "Field goal accuracy (raw)", y = NULL, colour = NULL) +
theme_minimal(base_size = 12) +
theme(plot.title = element_text(face = "bold", hjust = 0.5),
plot.subtitle = element_text(hjust = 0.5),
legend.position = "top")
Across the league, kickers made 84.8% at home and 83.1% on the road, a small gap. Individually it varies a lot. Sam Ficken dropped from 86% at home to 65% on the road, while Tyler Bass was actually better away (88% vs. 67%). I expected dome kickers to struggle most on the road, but the pattern is mixed: most of the biggest drop-offs belong to kickers with outdoor home stadiums. Home-field comfort seems to be about the individual kicker more than the roof.
fg |>
filter(!is.na(outdoor)) |>
mutate(bucket = cut(distance, c(0, 29, 39, 49, 70),
labels = c("Under 30", "30-39", "40-49", "50+")),
setting = ifelse(outdoor, "Outdoor", "Dome / closed roof")) |>
filter(!is.na(bucket)) |>
group_by(bucket, setting) |>
summarize(acc = mean(made), n = n(), .groups = "drop") |>
ggplot(aes(x = bucket, y = acc, fill = setting)) +
geom_col(position = position_dodge(width = 0.8), width = 0.75) +
geom_text(aes(label = percent(acc, 1)), position = position_dodge(width = 0.8),
vjust = -0.4, size = 3.6) +
scale_fill_manual(values = c("Dome / closed roof" = "#2a78d6", "Outdoor" = "#eb6834")) +
scale_y_continuous(labels = percent_format(1), limits = c(0, 1.05)) +
labs(title = "League-Wide Accuracy by Distance, Indoors vs. Outdoors",
x = "Kick distance (yards)", y = "Field goal accuracy", fill = NULL) +
theme_minimal(base_size = 13) +
theme(plot.title = element_text(face = "bold", hjust = 0.5), legend.position = "top")
The gap between indoor and outdoor kicking shows up mostly in the 30–39 yard range. Short kicks are nearly automatic everywhere. From 50+ yards outdoor kickers do slightly better, likely because coaches only send kickers out for long tries when the conditions allow it. Because conditions can matter, I score outdoor accuracy separately instead of letting easier indoor kicks boost a kicker’s overall number.
kickers |>
filter(blocked > 0) |>
ggplot(aes(x = blocked_rate, y = reorder(kicker, blocked_rate))) +
geom_col(fill = "#2a78d6", width = 0.7) +
geom_text(aes(label = paste0(blocked, " of ", attempts)), hjust = -0.15, size = 3.4) +
scale_x_continuous(labels = percent_format(0.1), expand = expansion(mult = c(0, 0.25))) +
labs(title = "Blocked Field Goal Rate (kickers with at least one block)",
x = "Blocked kicks / attempts", y = NULL) +
theme_minimal(base_size = 12) +
theme(plot.title = element_text(face = "bold", hjust = 0.5),
panel.grid.major.y = element_blank())
Blocks are rare: 37 of 2,643 field goals (1.4%), and 15 of the 38 ranked kickers never had one blocked. A block is partly the protection unit’s fault, but a kicker whose kicks come out low gets blocked more often, so a small penalty is fair. Because the counts are so small, the blocked rate only nudges the score and never decides it.
A three-season total can hide one great year and two average ones. This chart splits each top-10 kicker’s accuracy by season.
season_acc <- fg |>
filter(kickerId %in% top10$kickerId) |>
group_by(kicker = displayName, season) |>
summarize(acc = mean(made), n = n(), .groups = "drop") |>
left_join(top10 |> select(kicker, rank), by = "kicker") |>
mutate(kicker = reorder(gsub(" ", "\n", kicker), rank))
season_sd <- season_acc |>
group_by(kicker) |>
summarize(seasons = n(), spread = max(acc) - min(acc), .groups = "drop")
ggplot(season_acc, aes(x = kicker, y = acc, fill = factor(season))) +
geom_col(position = position_dodge(width = 0.85), width = 0.8) +
geom_text(aes(label = n), position = position_dodge(width = 0.85),
vjust = -0.4, size = 2.8, colour = "grey30") +
geom_hline(yintercept = league$overall, linetype = "dashed", colour = "grey40") +
scale_fill_manual(values = c("2018" = "#9ec5f0", "2019" = "#2a78d6", "2020" = "#0d3d75")) +
scale_y_continuous(labels = percent_format(1), limits = c(0, 1.08),
breaks = seq(0, 1, 0.25)) +
labs(title = "Field Goal Accuracy by Season, Top 10 Reliable Kickers",
subtitle = "Ordered by reliability rank. Numbers above bars = attempts. Dashed line = league average.",
x = NULL, y = "Field goal accuracy", fill = "Season") +
theme_minimal(base_size = 12) +
theme(plot.title = element_text(face = "bold", hjust = 0.5),
plot.subtitle = element_text(hjust = 0.5),
legend.position = "top", panel.grid.major.x = element_blank())
steady <- season_sd |> filter(seasons == 3) |> arrange(spread) |> slice(1)
swingy <- season_sd |> filter(seasons == 3) |> arrange(-spread) |> slice(1)
Even the most reliable kickers have off years: 5 of the top 10 dipped below the league average in at least one season. Harrison Butker was the steadiest kicker with all three seasons: his best and worst years were only 3 percentage points apart. Mason Crosby swung the most (17 points). Some kickers only appear in one or two seasons because they changed teams or were injured, and a few seasons rest on small numbers of attempts (see the numbers above the bars), so one bad season can look worse than it really was.
Made vs. missed treats a kick that barely slips inside the upright the same as one that splits the posts. The kick-level file includes where each field goal crossed the goal line (offset from the center of the uprights, in yards, measured from the ball-tracking data). The uprights are 18.5 feet apart, so a kick is between the posts when it crosses less than about 3.08 yards from center.
post <- 18.5 / 2 / 3 # half the goalpost width, in yards
offsets <- fg |>
filter(!is.na(e)) |>
mutate(result = ifelse(made, "Good", "No good"))
ggplot(offsets, aes(x = e, y = distance, colour = result)) +
annotate("rect", xmin = -post, xmax = post, ymin = -Inf, ymax = Inf,
fill = "#2a78d6", alpha = 0.06) +
geom_vline(xintercept = c(-post, post), linetype = "dashed", linewidth = 1) +
geom_point(alpha = 0.55, size = 1.8, position = position_jitter(height = 0.3, seed = 1)) +
scale_colour_manual(values = c("Good" = "#2a78d6", "No good" = "#eb6834")) +
scale_x_continuous(limits = c(-10, 10), breaks = seq(-9, 9, 3)) +
labs(title = "Where Every Field Goal Crossed the Goal Line, 2018-2020",
subtitle = "Dashed lines = uprights. Negative = left of center, positive = right of center.",
x = "Offset from center of the uprights (yards)", y = "Kick distance (yards)", colour = NULL) +
theme_minimal(base_size = 12) +
theme(plot.title = element_text(face = "bold", hjust = 0.5),
plot.subtitle = element_text(hjust = 0.5), legend.position = "top")
precision <- offsets |>
group_by(kickerId) |>
summarize(med_offset = median(abs(e)),
middle = mean(abs(e) < 1), # within 1 yard of dead center
n_tracked = n(), .groups = "drop") |>
inner_join(kickers |> select(kickerId, kicker, rank), by = "kickerId") |>
filter(n_tracked >= 20) |>
mutate(p_rank = min_rank(med_offset)) |>
arrange(p_rank)
precision |>
slice(1:10) |>
transmute(`Precision rank` = p_rank, Kicker = kicker,
`Median offset (yds)` = round(med_offset, 2),
`Within 1 yd of center` = percent(middle, 1),
`Tracked kicks` = n_tracked,
`Reliability rank` = rank) |>
kable(caption = "Most precise kickers: median distance from the center of the uprights (kickers with 20+ tracked field goals)",
align = "rlrrrr") |>
kable_styling(bootstrap_options = c("striped", "condensed"), full_width = FALSE)
| Precision rank | Kicker | Median offset (yds) | Within 1 yd of center | Tracked kicks | Reliability rank |
|---|---|---|---|---|---|
| 1 | Matt Prater | 0.84 | 58% | 78 | 24 |
| 2 | Younghoe Koo | 0.91 | 53% | 57 | 13 |
| 3 | Justin Tucker | 0.94 | 52% | 91 | 2 |
| 4 | Josh Lambo | 0.94 | 54% | 57 | 1 |
| 5 | Harrison Butker | 0.99 | 50% | 84 | 9 |
| 6 | Rodrigo Blankenship | 1.02 | 50% | 36 | 4 |
| 7 | Robbie Gould | 1.03 | 49% | 77 | 19 |
| 8 | Jake Elliott | 1.08 | 48% | 66 | 22 |
| 9 | Wil Lutz | 1.11 | 44% | 85 | 10 |
| 10 | Stephen Gostkowski | 1.13 | 46% | 63 | 33 |
The makes cluster tightly around the center, and the misses sit just outside the dashed lines. Most misses are near-misses, not shanks. The few orange points between the posts are kicks that came up short or were deflected.
Matt Prater has the most precise median kick: 0.84 yards from dead center, with 58% of his kicks crossing within a yard of the middle. But he ranks only #24 on reliability, with 82.3% accuracy overall. When his kicks were good they were usually perfect, but he missed too often for that to make him the best. My top two reliability kickers come next: Josh Lambo (#4 on precision) and Justin Tucker (#3). They weren’t squeaking kicks inside the uprights; they were splitting them. Graham Gano is the opposite case: #32 on precision, because more of his makes slipped in near a post. Accuracy and precision overlap, but they are not the same thing.
The reliability score has one blind spot: it treats a 25-yard kick and a 55-yard kick the same. A kicker whose coach sends him out for long attempts will miss more often, and that can look like a weakness when it is really a sign of confidence in his leg. To check this, I modeled the chance that a field goal is good.
The outcome is binary (made or missed), so ordinary regression
doesn’t fit: it can predict probabilities below 0 or above 1. A
binary logit model
(glm(..., family = binomial)) predicts the
log-odds of a make, so every prediction stays between 0 and 1.
I built it up in three steps, adding one group of variables at a
time:
fg_model <- fg |>
mutate(crosswind = coalesce(as.numeric(crosswind), 0),
cold = pmax(0, 50 - coalesce(as.numeric(temp), 70)), # 0 indoors / warm days
wet = (coalesce(rain, 0L) + coalesce(snow, 0L)) > 0) |>
filter(!is.na(distance), !is.na(outdoor))
model1 <- glm(made ~ distance, data = fg_model, family = binomial)
model2 <- glm(made ~ distance + outdoor + crosswind + cold + wet, data = fg_model, family = binomial)
model3 <- update(model2, . ~ . + pressure + home)
# odds ratios: exp(coefficient). Below 1 lowers the odds of a make, above 1 raises them
tidy_model <- function(m, name) {
s <- summary(m)$coefficients
data.frame(term = rownames(s), model = name,
cell = sprintf("%.3f%s", exp(s[, 1]),
ifelse(s[, 4] < 0.001, "***", ifelse(s[, 4] < 0.01, "**",
ifelse(s[, 4] < 0.05, "*", "")))))
}
labels <- c(distance = "Distance (per yard)", outdoorTRUE = "Outdoor stadium",
crosswind = "Crosswind (per mph)", cold = "Cold (per degree below 50 F)",
wetTRUE = "Rain or snow", pressureTRUE = "Pressure kick", homeTRUE = "Home kick")
bind_rows(tidy_model(model1, "Model 1"), tidy_model(model2, "Model 2"), tidy_model(model3, "Model 3")) |>
filter(term != "(Intercept)") |>
pivot_wider(names_from = model, values_from = cell, values_fill = "") |>
mutate(term = labels[term]) |>
bind_rows(data.frame(term = "AIC (lower is better)",
`Model 1` = sprintf("%.0f", AIC(model1)),
`Model 2` = sprintf("%.0f", AIC(model2)),
`Model 3` = sprintf("%.0f", AIC(model3)), check.names = FALSE)) |>
rename(Variable = term) |>
kable(caption = "Odds ratios from the binary logit models (*** p < 0.001, ** p < 0.01, * p < 0.05)",
align = "lrrr") |>
kable_styling(bootstrap_options = c("striped", "condensed"), full_width = FALSE)
| Variable | Model 1 | Model 2 | Model 3 |
|---|---|---|---|
| Distance (per yard) | 0.894*** | 0.895*** | 0.895*** |
| Outdoor stadium | 0.801 | 0.808 | |
| Crosswind (per mph) | 1.019 | 1.019 | |
| Cold (per degree below 50 F) | 1.015 | 1.015 | |
| Rain or snow | 1.030 | 1.029 | |
| Pressure kick | 1.040 | ||
| Home kick | 1.187 | ||
| AIC (lower is better) | 1987 | 1992 | 1993 |
d_or <- exp(coef(model1)["distance"])
p_out <- summary(model2)$coefficients["outdoorTRUE", 4]
Distance is the only variable that matters. Each extra yard multiplies the odds of a make by 0.894, which means about 10.6% lower odds per yard. The effect compounds instead of adding up in a straight line: a 50-yard kick has roughly 33% of the odds of a 40-yard kick. None of the condition or situation variables is significant once distance is in the model (outdoor p = 0.12), and Model 1 has the lowest AIC. In other words, in this data the gap between indoor and outdoor kicking from Visualization 4 is mostly explained by which kicks get attempted, not the conditions themselves.
That finding lets me ask a cleaner question: given the distance of every kick he attempted, how many more field goals did each kicker make than an average NFL kicker would have? I call this field goals over expected (FGOE). For each kick, Model 1 gives the probability that an average kicker makes it. FGOE is the sum of (made − expected) over all of a kicker’s attempts. I also show it per 100 attempts so kickers with different workloads can be compared.
fg_model$expected <- predict(model1, fg_model, type = "response")
fgoe <- fg_model |>
group_by(kickerId) |>
summarize(exp_pct = mean(expected),
fgoe = sum(made - expected),
.groups = "drop") |>
inner_join(kickers |> select(kickerId, kicker, attempts, overall, avg_dist, reliability, rank),
by = "kickerId") |>
mutate(fgoe100 = 100 * fgoe / attempts,
x_rank = min_rank(desc(fgoe100))) |>
arrange(x_rank)
rank_cor <- cor(fgoe$x_rank, fgoe$rank, method = "spearman")
big_mover <- fgoe |> mutate(jump = rank - x_rank) |> arrange(-jump) |> slice(1)
big_faller <- fgoe |> mutate(jump = rank - x_rank) |> arrange(jump) |> slice(1)
ggplot(fgoe, aes(x = fgoe100, y = reorder(kicker, fgoe100),
fill = ifelse(fgoe100 >= 0, "Above expected", "Below expected"))) +
geom_col(width = 0.75) +
geom_vline(xintercept = 0, colour = "grey30") +
geom_text(aes(label = sprintf("%+.1f", fgoe),
hjust = ifelse(fgoe100 >= 0, -0.2, 1.2)), size = 3.2) +
scale_fill_manual(values = c("Above expected" = "#2a78d6", "Below expected" = "#eb6834")) +
scale_x_continuous(expand = expansion(mult = 0.15)) +
labs(title = "Field Goals Made Over Expected per 100 Attempts, 2018-2020",
subtitle = "Expected = make probability for each kick's distance from the logit model. Labels show total FGOE.",
x = "Field goals over expected per 100 attempts", y = NULL, fill = NULL) +
theme_minimal(base_size = 12) +
theme(plot.title = element_text(face = "bold", hjust = 0.5),
plot.subtitle = element_text(hjust = 0.5),
legend.position = "top", panel.grid.major.y = element_blank())
Graham Gano leads on a per-attempt basis (+11.2 per 100 kicks), and Justin Tucker has the most total value: 8.4 more field goals than an average kicker would have made on the same kicks, worth about 25 points. The biggest change from my reliability ranking is Brandon McManus, who rises from #25 to #5. His raw accuracy (85.7%) looks ordinary, but his average attempt was 41.4 yards, among the longest in the league. Going the other way, Austin Seibert falls from #15 to #32 because his kicks were shorter than average.
Hover over a point to see the kicker’s numbers. Kickers in the top-right corner are strong on both measures.
g <- ggplot(fgoe, aes(x = fgoe100, y = reliability, size = attempts,
text = paste0("<b>", kicker, "</b>",
"<br>Reliability: ", sprintf("%.1f", reliability), " (rank ", rank, ")",
"<br>FGOE per 100: ", sprintf("%+.1f", fgoe100), " (rank ", x_rank, ")",
"<br>Accuracy: ", percent(overall, 0.1),
"<br>Avg distance: ", sprintf("%.1f", avg_dist), " yds",
"<br>Attempts: ", attempts))) +
geom_vline(xintercept = 0, linetype = "dashed", colour = "grey60") +
geom_hline(yintercept = mean(fgoe$reliability), linetype = "dashed", colour = "grey60") +
geom_point(colour = "#2a78d6", alpha = 0.7) +
geom_text(data = fgoe |> filter(x_rank <= 5 | rank <= 5),
aes(label = kicker), size = 3, nudge_y = 0.8, colour = "grey20") +
scale_size_continuous(range = c(2, 7), guide = "none") +
labs(title = "Reliability Score vs. Field Goals Over Expected",
x = "Field goals over expected per 100 attempts", y = "Reliability score") +
theme_minimal(base_size = 12) +
theme(plot.title = element_text(face = "bold", hjust = 0.5))
ggplotly(g, tooltip = "text") |> config(displayModeBar = FALSE)
The two rankings mostly agree (rank correlation = 0.83). Graham Gano, Josh Lambo, Justin Tucker, and Jason Myers are in the top five of both. That gives me more confidence in the result: these kickers come out on top whether you reward performance in hard situations or performance on hard kicks.
kickers |>
transmute(Rank = rank, Kicker = kicker, Attempts = attempts,
`Avg distance` = round(avg_dist, 1),
Overall = percent(overall, 0.1),
Pressure = percent(pressure_acc, 0.1),
Road = percent(road_acc, 0.1),
Outdoor = percent(outdoor_acc, 0.1),
Blocked = blocked,
Score = round(reliability, 1)) |>
kable(caption = "All kickers with at least 25 field goal attempts, 2018-2020",
align = "rlrrrrrrrr")
| Rank | Kicker | Attempts | Avg distance | Overall | Pressure | Road | Outdoor | Blocked | Score |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Josh Lambo | 57 | 37.8 | 94.7% | 93.5% | 87.2% | 96.6% | 0 | 93.0 |
| 2 | Justin Tucker | 94 | 38.1 | 92.6% | 96.2% | 93.7% | 91.4% | 2 | 91.3 |
| 3 | Graham Gano | 45 | 39.9 | 93.3% | 88.5% | 89.0% | 93.6% | 0 | 91.1 |
| 4 | Rodrigo Blankenship | 36 | 35.7 | 86.1% | 91.6% | 91.6% | 89.9% | 0 | 89.8 |
| 5 | Jason Myers | 80 | 39.4 | 91.2% | 83.3% | 93.8% | 89.8% | 0 | 89.5 |
| 6 | Jason Sanders | 83 | 39.1 | 85.5% | 95.3% | 92.3% | 83.6% | 0 | 89.2 |
| 7 | Mason Crosby | 71 | 39.5 | 88.7% | 89.8% | 84.4% | 93.2% | 0 | 89.0 |
| 8 | Nick Folk | 42 | 36.9 | 90.5% | 92.4% | 93.2% | 88.8% | 1 | 88.8 |
| 9 | Harrison Butker | 85 | 36.1 | 90.6% | 84.2% | 92.0% | 90.7% | 1 | 88.2 |
| 10 | Wil Lutz | 86 | 37.6 | 89.5% | 88.5% | 88.5% | 87.3% | 1 | 87.3 |
| 11 | Matt Bryant | 34 | 41.4 | 82.4% | 89.5% | 89.1% | 84.8% | 0 | 86.4 |
| 12 | Dustin Hopkins | 86 | 39.2 | 84.9% | 92.0% | 89.7% | 83.8% | 1 | 86.4 |
| 13 | Younghoe Koo | 57 | 36.8 | 91.2% | 77.0% | 90.8% | 85.9% | 0 | 86.2 |
| 14 | Aldrick Rosas | 52 | 36.2 | 86.5% | 84.7% | 84.9% | 84.4% | 0 | 85.1 |
| 15 | Austin Seibert | 34 | 38.9 | 79.4% | 90.7% | 84.0% | 82.4% | 0 | 84.1 |
| 16 | Chris Boswell | 63 | 36.1 | 87.3% | 85.8% | 81.5% | 85.8% | 1 | 83.5 |
| 17 | Randy Bullock | 77 | 39.1 | 85.7% | 79.8% | 80.8% | 85.8% | 0 | 83.0 |
| 18 | Ryan Succop | 62 | 36.2 | 83.9% | 90.3% | 79.4% | 86.5% | 2 | 81.8 |
| 19 | Robbie Gould | 79 | 37.0 | 83.5% | 86.5% | 80.7% | 84.8% | 2 | 81.3 |
| 20 | Greg Zuerlein | 93 | 39.4 | 81.7% | 89.2% | 76.5% | 81.8% | 1 | 81.2 |
| 21 | Tyler Bass | 28 | 40.1 | 78.6% | 83.3% | 86.4% | 76.5% | 0 | 81.2 |
| 22 | Jake Elliott | 66 | 38.2 | 81.8% | 77.4% | 82.8% | 80.9% | 0 | 80.7 |
| 23 | Ka’imi Fairbairn | 96 | 38.0 | 84.4% | 87.1% | 76.8% | 72.9% | 0 | 80.3 |
| 24 | Matt Prater | 79 | 38.1 | 82.3% | 78.7% | 81.9% | 81.8% | 1 | 79.9 |
| 25 | Brandon McManus | 84 | 41.4 | 85.7% | 73.1% | 83.3% | 85.6% | 2 | 79.5 |
| 26 | Cody Parkey | 54 | 37.6 | 81.5% | 76.3% | 85.8% | 81.1% | 1 | 79.3 |
| 27 | Matt Gay | 44 | 40.2 | 79.5% | 82.2% | 82.7% | 79.4% | 1 | 78.7 |
| 28 | Daniel Carlson | 77 | 35.9 | 83.1% | 79.8% | 74.8% | 76.7% | 0 | 78.6 |
| 29 | Cairo Santos | 51 | 37.4 | 80.4% | 71.5% | 87.2% | 80.4% | 1 | 77.9 |
| 30 | Dan Bailey | 72 | 37.2 | 81.9% | 79.7% | 75.9% | 76.8% | 1 | 77.2 |
| 31 | Sam Ficken | 39 | 39.6 | 76.9% | 93.0% | 68.9% | 76.0% | 1 | 76.1 |
| 32 | Michael Badgley | 62 | 38.8 | 79.0% | 74.5% | 72.9% | 83.7% | 1 | 75.9 |
| 33 | Stephen Gostkowski | 65 | 38.7 | 78.5% | 83.3% | 76.6% | 76.1% | 2 | 75.5 |
| 34 | Stephen Hauschka | 55 | 39.9 | 78.2% | 74.2% | 74.7% | 80.4% | 1 | 75.0 |
| 35 | Brett Maher | 62 | 38.6 | 75.8% | 86.9% | 76.1% | 72.1% | 2 | 74.5 |
| 36 | Adam Vinatieri | 49 | 38.0 | 77.6% | 69.7% | 80.6% | 79.0% | 2 | 72.7 |
| 37 | Zane Gonzalez | 64 | 38.4 | 78.1% | 65.6% | 79.3% | 78.7% | 2 | 72.3 |
| 38 | Joey Slye | 55 | 40.6 | 76.4% | 69.7% | 78.8% | 74.1% | 2 | 71.1 |
Limitations and next steps. Pressure and outdoor samples are small, even with smoothing. The logit model covers three seasons with only about 2,643 kicks, so it can’t reliably separate small weather effects, and stadium weather is measured at the nearest weather station, not on the field. Next steps would be adding kicker fixed effects to the model, interaction terms (for example distance × crosswind, since wind should matter more on long kicks), and more features from the full ball-tracking data, such as how much each kick curved in flight and how far it drifted in the first half-second after contact.