Introduction

This code-through explores the visualization between MLB team spending and Post Season wins, using ridge plots.The package used for this is ggridges.


Content Overview

Specifically, we’ll explain and demonstrate the extraction of postseason performance and team salary with the R Studio baseball package, “Lahman”. We will use ridge plots to highlight stats and information that may suggest whether team spending correlates with the amount of post season wins. This code-through will cover the 2016 MLB season.


Why You Should Care

This topic is valuable because a ridge plot can help one better visual multiple statistics in comparison to other graphs, like a traditional line chart for example- as various sets of data can become overcrowded and difficult to read. This can also be an entertaining code-through to go through, specifically for baseball fans- to see if teams are “buying the win” or if they are just smart spenders.


Learning Objectives

This code-through can show one

  • Basic Ridgeplots

  • Customized Ridge plots with statistics

  • Post season data and player salaries using the “Lahman” package

  • Distributional plotting across the graph



2016 MLB POST SEASON

Below, we will show how to gather and code 2016 MLB Post season data with visuals.


Further Exposition

The data below is provided by the R studio baseball and basketball package Lahman, and the ggridges package, which is essentially an extension for ggplot2 commands- allowing us to create various geoms.


The base ridge plot

This first code and ridge plot is constructed to show you the filtration of 2016 post season data, including only the twelve post season teams- those being:

  • Chicago Cubs (CHN)

  • Cleveland Guardians (CLE)

  • Toronto Bluejays (TOR)

  • Los Angeles Dodgers (LAN)

  • Washington Nationals (WAS)

  • San Francisco Giants (SFN)

  • Texas Rangers (TEX)

  • New York Mets (NYN)

  • Boston Red Sox (BOS)

  • Baltimore Orioles (BAL)

It also shows each teams total postseason wins, alongside player salaries in their respective teams.

Our x-axis covers Player Salary (in millions), and our y-axis covers team rankings by the most post-season wins in the 2016 season. The left side of the ridge plot indicates the least amount (base pay) a player received versus the right side of the ridge plot representing high salaries. With “Lahman” package having records of wins and losses by each team, it is not difficult to sort out postseason teams with commands like filter, teamIDWinner, and teamIDloser.

# Some code

# The packages used for this code-through
library(pander)
library(kableExtra)
library(Lahman)
library(dplyr)
library(ggplot2)
library(ggridges)


# Draw out 2016 PostSeason Teams and Wins

postseason_2016 <- SeriesPost %>%
  filter(yearID == 2016) %>%
  tidyr::pivot_longer(cols = c(teamIDwinner, teamIDloser), 
                      names_to = "role", 
                      values_to = "teamID") %>%

# How many wins in the postseason?
mutate(p_wins = if_else(role == "teamIDwinner", wins, losses)) %>%
  group_by(teamID) %>%
  summarize(postseason_wins = sum(p_wins), .groups = 'drop')


# 2016 Player salary

salary_dist_2016 <- Salaries %>%
  filter(yearID == 2016) %>%
  inner_join(postseason_2016, by = "teamID") %>%

  mutate(team_label = paste0(teamID, " (", postseason_wins, " PS Wins)"))


# Ridgeplot time

ggplot(salary_dist_2016, aes(x = salary, y = reorder(team_label, postseason_wins))) +
  geom_density_ridges(alpha = 0.7, fill = "skyblue") +
  scale_x_continuous(
    labels = function(x) paste0("$", x / 1000000, "M"),
    breaks = c(0, 15000000, by= 30000000),
    limits = c(0,35000000)
                 )+
  theme_ridges() +
  theme(
  axis.text.x= element_text(angle = 45, hjust= 1, vjust = 1)
          ) +                  
  labs(
    title = "2016 MLB Postseason Teams: Player Salary Distribution",
    x = "Player Salary ($) in Millions ",
    y = "Teams (by Postseason Wins)"
  )


Advanced Examples

In the advanced examples, we cover gradient colors for our ridge plot, which can make reading and understanding data easier. This highlights if teams with star players on their rosters also had the most wins, or if teams with average pay held that title. We are able to create the gradience on the ridge plot with the theme function(s). Yellow indicates the highest amount spent, $40M, whereas indigo/purple represents the least amount spent, $0-$10M.

# This code is to show a concentration in Payroll $

ggplot(salary_dist_2016, aes(x = salary, y = reorder(team_label, postseason_wins), fill = ..x..)) +
  geom_density_ridges_gradient(scale = 1.5, rel_min_height = 0.01) +
  scale_fill_viridis_c(name = "Salary ($)", 
    option = "C", 
    labels = function(x) paste0("$", x / 1000000, "M")
  ) +
  scale_x_continuous(
    labels = function(x) paste0("$", x / 1000000, "M"), 
    breaks = seq(0, 30000000, by = 15000000),          
    limits = c(0, 35000000)
  ) +
  theme_ridges(grid = TRUE) +
  theme(
    axis.text.x = element_text(angle = 45, hjust = 1, vjust = 1) 
  ) +
  labs(
    title = "Concentrated Salary vs. Postseason Success",
    subtitle = "2016 MLB Postseason participants",
    x = "Individual Player Salary",
    y = "Team & Postseason Wins"
  )


What’s more, it can also be used for showing statistics within the ridge plot, like player salary medians, using functions that ggridges already has.The light line on the ridge plots indicates the salary medians for each postseason team. We are able to compute this with the quantile function.

Looking at the postseason and world-champion winning team, the Chicago Cubs, we can see that a lot of their players had a middle-class amount of pay, with most of their players being on the lower end pay side. Our big spenders, the Los Angeles Dodgers, shows a relatively spread-out graph- indicating they had a good number of expensive players. However, spending did not come in clutch in this case, as they would lose to the Chicago Cubs in the 2016 World Series. Our lowest spending team, the Cleveland Guardians, made it all the way to the American League championship, despite not having those “expensive” players.

# We can also add medians to the ridgeplot

ggplot(salary_dist_2016, aes(x = salary, y = reorder(team_label, postseason_wins))) +
  geom_density_ridges(
    quantile_lines = TRUE, 
    quantiles = 2, # Shows the exact median line
    alpha = 0.7, 
    fill = "#2c3e50", 
    color = "white"
  ) +
  scale_x_continuous(
    labels = function(x) paste0("$", x / 1000000, "M"),
    breaks = seq(0, 30000000, by = 15000000),
    limits = c(0, 35000000)
  ) +
  theme_ridges(grid = TRUE) +
  theme(
    axis.text.x = element_text(angle = 45, hjust = 1, vjust = 1)
  ) +
  labs(
    title = "Median Player Salaries Among 2016 PostSeason Teams",
    subtitle = "Vertical lines represent team median player salary for each team",
    x = "Player Salary",
    y = "Team & Postseason Wins"
  )


Most notably, its valuable for grouping teams by the amount of postseason wins- to measure whether the amount of team spending correlates with wins.

As we categorize teams by the amount of postseason wins they had for the 2016 season, we can see that teams with zero postseason wins and teams with 11 post seasons have a ridge plot that is almost identical. In seeing this, its evident that although spending may help teams progress- luck, player skill, and series match ups play a far greater role than the amount spent.

# Groups by Postseason wins 

ggplot(salary_dist_2016, aes(x = salary, y = as.factor(postseason_wins), fill = as.factor(postseason_wins))) +
  geom_density_ridges(alpha = 0.6, scale = 1.2) +
  scale_fill_brewer(palette = "Spectral", direction = -1, guide = "none") +
  scale_x_continuous(
    labels = function(x) paste0("$", x / 1000000, "M"),
    breaks = seq(0, 30000000, by = 15000000),
    limits = c(0, 35000000)
  ) +
  theme_ridges(grid = TRUE) +
  theme(
    axis.text.x = element_text(angle = 45, hjust = 1, vjust = 1)
  ) +
  labs(
    title = "Does Playoff Success Determine a Higher Salary?",
    subtitle = "Player salary distributions by 2016 postseason wins",
    x = "Player Salary",
    y = "Total Postseason Wins Achieved"
  )



Further Resources

Learn more about [package, technique, dataset] with the following:




Works Cited

This code through references and cites the following sources: