This code through explores ridgeline plots using the ggridges and ggplot2 packages. This allows you to visually stack 6 or more density graphs vertically on a shared horizontal axis. By doing so you can visually see how the data overlaps, and incidentally gives the visual effect of a mountain range.
Using previously prepared data from the Lahman package, I will compare the distribution of inflation-adjusted salary cost per win according to each team’s season-ending rank.
Ridgeline plots are perfect to use when you need to compare a moderate to large number of groups. Because the distributions overlap slightly, Ridgeline plots use less space than displaying each group in a separate graph. Ridgeline plots also work best when there is a clear ranking in groups. If there is no clear ranking in groups they tend to overlap and will only provide a messy plot with little to no insight. Something to note is that density ridges require a numeric x-axis.
Specifically, you’ll learn how to:
Create a ridgeline plot with ggridges Interpret the
resulting distributions
We will be using the following packages:
Before creating the ridgeline plot, the salary and team data must be prepared and joined. The following code is adapted from a previous analysis and creates the final.table dataset used in this tutorial.
# Load data
data(Teams)
data(Salaries)
# Salaries adjusted for inflation at a constant annual increase rate of 3%
Salaries <- Salaries %>%
mutate(salary.adj = salary * (1.03)^(max(yearID) - yearID))
# Total adjusted team budget
salarysummary <- Salaries %>%
group_by(yearID, teamID) %>%
summarise(
n = n(),
adjteambudget = sum(salary.adj, na.rm = TRUE),
.groups = "drop"
)
# Join Salaries & Teams
team.data <- salarysummary %>%
left_join(
Teams,
by = c("yearID", "teamID")
)
# Calculate Cost per Win
team.data <- team.data %>%
mutate(
cost.per.win = adjteambudget / W / 100000
)
# Select, Filter, & Arrange for Final Table
final.table <- team.data %>%
filter(n >= 25) %>%
arrange(cost.per.win) %>%
select(
yearID,
teamID,
name,
lgID,
Rank,
adjteambudget,
cost.per.win
)A ridgeline plot ac be demonstrated by the distribution of one team’s inflation-adjusted salary cost per win across multiple seasons
ridge.data <- final.table %>%
filter(
name %in% c(
"Arizona Diamondbacks",
"Atlanta Braves",
"Boston Red Sox",
"Chicago Cubs",
"Los Angeles Dodgers",
"New York Yankees",
"San Francisco Giants",
"St. Louis Cardinals"
)
)
ggplot(
ridge.data,
aes(
x = cost.per.win,
y = name,
fill = name
)
) +
geom_density_ridges(
alpha = 0.7,
scale = 1.5
) +
labs(
title = "Salary Cost per Win by MLB Team",
subtitle = "Distribution of inflation-adjusted cost per win",
x = "Cost per Win ($100,000s)",
y = "Team"
) +
theme_minimal() +
theme(
legend.position = "none"
)
Each ridge represents the distribution of one team’s inflation-adjusted
salary cost per win across multiple seasons. Teams whose ridges are
concentrated farther to the left generally earned wins at a lower salary
cost. Wider ridges indicate greater variation in cost per win across
seasons.
The height of a ridge represents the concentration of observations rather than the number of seasons. Therefore, the graph is most useful for comparing the location and shape of the teams’ distributions.
A better demonstration would compare the distribution of cost per win across different years.
# create new data frame from final.table data set to keep ears dividable by 5
ridge.data2 <- final.table %>%
filter(yearID %% 5 == 0)
# create the ridgeline plot
ggplot(
ridge.data,
aes(
x = cost.per.win,
y = factor(yearID),
fill = factor(yearID)
)
) +
geom_density_ridges(
alpha = 0.7,
scale = 1.5
) +
labs(
title = "MLB Salary Cost per Win Over Time",
subtitle = "Distribution among teams at five-year intervals",
x = "Cost per Win ($100,000s)",
y = "Year"
) +
theme_minimal() +
theme(
legend.position = "none"
)Ridgeline plots are incredibly useful when the groups have a natural order. In this example, the groups are baseball seasons arranged chronologically. Each ridge displays the distribution of inflation-adjusted salary cost per win among MLB teams during that season. Five-year intervals are used to keep the graph readable while still showing changes over time.
Learn more about [package, technique, dataset] with the following:
Resource I Data-to-Viz: Ridgeline Plot
Resource II Lahman Baseball Database
Resource III ggridges Documentation
This code through references and cites the following sources:
Friendly, M., Dalzell, C., Monkman, M., & Murphy, D. (2026). Lahman: Sean Lahman baseball database (R package version 14.0-0). https://CRAN.R-project.org/package=Lahman
Holtz, Y. (n.d.). Ridgeline plot. From Data to Viz. https://www.data-to-viz.com/graph/ridgeline.html
Wilke, C. O. (2025). ggridges: Ridgeline plots in ggplot2 (R package version 0.5.7). https://wilkelab.org/ggridges/