Introduction

This code through explores ridgeline plots using the ggridges and ggplot2 packages. This allows you to visually stack 6 or more density graphs vertically on a shared horizontal axis. By doing so you can visually see how the data overlaps, and incidentally gives the visual effect of a mountain range.


Content Overview

Using previously prepared data from the Lahman package, I will compare the distribution of inflation-adjusted salary cost per win according to each team’s season-ending rank.


Why You Should Care

Ridgeline plots are perfect to use when you need to compare a moderate to large number of groups. Because the distributions overlap slightly, Ridgeline plots use less space than displaying each group in a separate graph. Ridgeline plots also work best when there is a clear ranking in groups. If there is no clear ranking in groups they tend to overlap and will only provide a messy plot with little to no insight. Something to note is that density ridges require a numeric x-axis.


Learning Objectives

Specifically, you’ll learn how to:

Create a ridgeline plot with ggridges Interpret the resulting distributions



Packages Needed

We will be using the following packages:

library(dplyr)
library(Lahman)
library(ggplot2)
library(ggridges)


Preparing the data

Before creating the ridgeline plot, the salary and team data must be prepared and joined. The following code is adapted from a previous analysis and creates the final.table dataset used in this tutorial.

# Load data

data(Teams)
data(Salaries) 


# Salaries adjusted for inflation at a constant annual increase rate of 3%

Salaries <- Salaries %>% 
  mutate(salary.adj = salary * (1.03)^(max(yearID) - yearID))



# Total adjusted team budget

salarysummary <- Salaries %>% 
  group_by(yearID, teamID) %>%
  summarise(
    n = n(),
    adjteambudget = sum(salary.adj, na.rm = TRUE),
    .groups = "drop"
            )


# Join Salaries & Teams

team.data <- salarysummary %>%
  left_join(
    Teams, 
    by = c("yearID", "teamID")
  )


# Calculate Cost per Win

team.data <- team.data %>%
  mutate(
    cost.per.win = adjteambudget / W / 100000
  )



# Select, Filter, & Arrange for Final Table

final.table <- team.data %>%
  filter(n >= 25) %>%
  arrange(cost.per.win) %>%
  select(
    yearID,
    teamID,
    name,
    lgID,
    Rank,
    adjteambudget,
    cost.per.win
    )


Creating the Ridgeline Plot

A ridgeline plot ac be demonstrated by the distribution of one team’s inflation-adjusted salary cost per win across multiple seasons

ridge.data <- final.table %>%
  filter(
    name %in% c(
      "Arizona Diamondbacks",
      "Atlanta Braves",
      "Boston Red Sox",
      "Chicago Cubs",
      "Los Angeles Dodgers",
      "New York Yankees",
      "San Francisco Giants",
      "St. Louis Cardinals"
    )
  )

ggplot(
  ridge.data,
  aes(
    x = cost.per.win,
    y = name,
    fill = name
  )
) +
  geom_density_ridges(
    alpha = 0.7,
    scale = 1.5
  ) +
  labs(
    title = "Salary Cost per Win by MLB Team",
    subtitle = "Distribution of inflation-adjusted cost per win",
    x = "Cost per Win ($100,000s)",
    y = "Team"
  ) +
  theme_minimal() +
  theme(
    legend.position = "none"
  )

Each ridge represents the distribution of one team’s inflation-adjusted salary cost per win across multiple seasons. Teams whose ridges are concentrated farther to the left generally earned wins at a lower salary cost. Wider ridges indicate greater variation in cost per win across seasons.

The height of a ridge represents the concentration of observations rather than the number of seasons. Therefore, the graph is most useful for comparing the location and shape of the teams’ distributions.


A better demonstration would compare the distribution of cost per win across different years.

# create new data frame from final.table data set to keep ears dividable by 5

ridge.data2 <- final.table %>%
  filter(yearID %% 5 == 0) 

# create the ridgeline plot

ggplot(
  ridge.data,
  aes(
    x = cost.per.win,
    y = factor(yearID),
    fill = factor(yearID)
  )
) +
  geom_density_ridges(
    alpha = 0.7,
    scale = 1.5
  ) +
  labs(
    title = "MLB Salary Cost per Win Over Time",
    subtitle = "Distribution among teams at five-year intervals",
    x = "Cost per Win ($100,000s)",
    y = "Year"
  ) +
  theme_minimal() +
  theme(
    legend.position = "none"
  )

Ridgeline plots are incredibly useful when the groups have a natural order. In this example, the groups are baseball seasons arranged chronologically. Each ridge displays the distribution of inflation-adjusted salary cost per win among MLB teams during that season. Five-year intervals are used to keep the graph readable while still showing changes over time.




Further Resources

Learn more about [package, technique, dataset] with the following:




Works Cited

This code through references and cites the following sources:


Friendly, M., Dalzell, C., Monkman, M., & Murphy, D. (2026). Lahman: Sean Lahman baseball database (R package version 14.0-0). https://CRAN.R-project.org/package=Lahman

Holtz, Y. (n.d.). Ridgeline plot. From Data to Viz. https://www.data-to-viz.com/graph/ridgeline.html

Wilke, C. O. (2025). ggridges: Ridgeline plots in ggplot2 (R package version 0.5.7). https://wilkelab.org/ggridges/