install.packages("PrettyCols", repos = "https://r-project.org")
# Load packages
library(tidyverse)
library(here)
library(scales)
library(usmap)
library(prettyunits)
library(PrettyCols)
# Set theme for plotsWho Votes?
State-Level Turnout in the 2024 Election
1 Overview
In this session of the Election Lab we’ll continue our exploration of who votes by investigating state-level voter turnout in the 2024 presidential election.
By the end of this session, you should be able to:
Explore imported datasets using the
glimpse()andview()functionsVisualize spatial data by making beautiful choropleth maps with the {usmap} package
Analyze the variation in voter turnout across the 50 states and DC
2 Setup
Let’s get setup for our analysis by loading some packages we’ll need. As before, we’ll load {tidyverse} and {here}. To enhance our data visualizations, we’ll draw on some functions from the {scales} package, some beautiful color palettes available in Nicola Rennie’s {PrettyCols} package, and, to leverage the geospatial dimensions of our data, the exceptionally helpful {usmap} package.
As a reminder, if you get an error saying the package does not exist, use the install.packages("package_name") function in the Console and then re-run the chunk below.
3 Data Import
We’ll start by importing some more data made available by Michael McDonald and his team at the University of Florida’s Election Lab.
turnout_2024 <-read_csv(here("data",
"turnout_2024_tidy.csv"))4 Data Inspection
Let’s take a look to get a sense of what’s inside the dataset.
# look at the first 10 rows
turnout_2024# A tibble: 51 × 12
state state_abv total_ballots_counted vap noncitizen_pct ineligible_prison
<chr> <chr> <dbl> <dbl> <dbl> <dbl>
1 Alab… AL 2270000 4.02e6 3.08 19073
2 Alas… AK 340981 5.61e5 3.41 1357
3 Ariz… AZ 3428011 5.95e6 8.18 31033
4 Arka… AR 1190172 2.39e6 4.39 18001
5 Cali… CA 16140044 3.06e7 15.0 95827
6 Colo… CO 3241120 4.73e6 6.26 16412
7 Conn… CT 1788981 2.91e6 8.12 6308
8 Dela… DE 518330 8.36e5 6.35 2800
9 Dist… DC 328871 5.62e5 9.87 0
10 Flor… FL 11004209 1.87e7 11.4 81978
# ℹ 41 more rows
# ℹ 6 more variables: ineligible_probation <dbl>, ineligible_parole <dbl>,
# ineligible_felons_total <dbl>, vep <dbl>, vep_turnout_rate <dbl>,
# vap_turnout_rate <dbl>
There’s 12 variables or columns in this dataset, which can’t be fully displayed inside the Console window. In scenarios like this, we have a few options.
One option is to use the glimpse() function from the {dplyr} package to transpose the dataset (i.e., flip it 90 degrees to the left). This flipped perspective provides us with a full view of the column names, which are arrayed vertically with dollar signs $ on the left side, and a preview of the first several observations or rows.
# glimpse variable names, their data class, and first several observations
glimpse(turnout_2024)Rows: 51
Columns: 12
$ state <chr> "Alabama", "Alaska", "Arizona", "Arkansas", "C…
$ state_abv <chr> "AL", "AK", "AZ", "AR", "CA", "CO", "CT", "DE"…
$ total_ballots_counted <dbl> 2270000, 340981, 3428011, 1190172, 16140044, 3…
$ vap <dbl> 4020661, 560512, 5951329, 2391789, 30601993, 4…
$ noncitizen_pct <dbl> 3.08, 3.41, 8.18, 4.39, 14.98, 6.26, 8.12, 6.3…
$ ineligible_prison <dbl> 19073, 1357, 31033, 18001, 95827, 16412, 6308,…
$ ineligible_probation <dbl> 3420, 2310, 60350, 44540, 0, 0, 0, 3250, 0, 13…
$ ineligible_parole <dbl> 6870, 1170, 6790, 22580, 0, 0, 0, 310, 0, 3810…
$ ineligible_felons_total <dbl> 29363, 4837, 98173, 85121, 95827, 16412, 6308,…
$ vep <dbl> 3867406, 536581, 5366526, 2201764, 25922868, 4…
$ vep_turnout_rate <dbl> 58.70, 63.55, 63.88, 54.06, 62.26, 73.40, 67.0…
$ vap_turnout_rate <dbl> 56.46, 60.83, 57.60, 49.76, 52.74, 68.55, 61.4…
Another option is to pop over to the Console command line and pass the dataset’s name into the view() function. This allows you to see the full dataset in spreadsheet format as you would with Excel or Google Sheets.
5 Exploratory Data Analysis
With our data imported we’re ready to start exploring it for answers, insights, and new questions about 2024 turnout across the states. Recall that the purpose of EDA is to better understand our data by examining variation within our main variables of interest and any potential relationships between them.
5.1 Use {dplyr} verbs to wrangle some information from the data
Data scientists often want to know the range of a variable; that is, what are its lowest and highest observed values? In our case that means asking: which states had the highest and lowest turnout rates during the 2024 election? Use the select() and arrange() functions to find out.
# lowest turnout
turnout_2024 |>
select(state, vep_turnout_rate) |>
arrange(vep_turnout_rate)# A tibble: 51 × 2
state vep_turnout_rate
<chr> <dbl>
1 Hawaii 50.3
2 Oklahoma 53.5
3 Arkansas 54.1
4 West Virginia 55.6
5 Texas 56.8
6 Mississippi 57.8
7 Tennessee 57.9
8 Alabama 58.7
9 Indiana 59.0
10 New Mexico 59.7
# ℹ 41 more rows
# highest turnout
turnout_2024 |>
select(state, vep_turnout_rate) |>
arrange(vep_turnout_rate)# A tibble: 51 × 2
state vep_turnout_rate
<chr> <dbl>
1 Hawaii 50.3
2 Oklahoma 53.5
3 Arkansas 54.1
4 West Virginia 55.6
5 Texas 56.8
6 Mississippi 57.8
7 Tennessee 57.9
8 Alabama 58.7
9 Indiana 59.0
10 New Mexico 59.7
# ℹ 41 more rows
slice_* your data
Use the Console to look up the documentation for ?slice. Does the family of slice_*() functions offer another way to conduct the same operations as the previous code chunk? If so, create a new chunk below and show how to find the states with the lowest and highest turnout rates.
5.2 Use {ggplot2} to visualize state variation
To get a fuller picture let’s use {ggplot2} to visualize 2024 voter turnout across the states using a dot plot. To visualize the changing value of turnout — a continuous variable — with color, we can employ some of the customized palettes in the {PrettyCols} package. Run the view_all_palettes(type = "seq") in the Console to see the available options.
turnoutplot_2024<-turnout_2024 |>
ggplot(aes(
x = vep_turnout_rate,
y = reorder(state, vep_turnout_rate)
)) +
geom_point(aes(color = vep_turnout_rate)) +
labs(y = "",
x = "Turnout Rate",
title = "2024 Voter Turnout by State") +
scale_color_pretty_c(palette = "Pinks") +
theme(legend.position = "none")5.3 Mapping turnout with {usmap}
To further improve our understanding of our data, let’s leverage its spatial dimensions to visualize state-level turnout rates with a choropleth map. We’ll use the main function of the {usmap} package: plot_usmap().
We’ll also scale our color variable with another set of palettes called the viridis color maps. These palettes are common in R visualizations, are more accessible to those with colorblindness, and work well with geographic maps.
plotUSmap<-plot_usmap(data = turnout_2024,
values = "vep_turnout_rate",
color = "grey80") +
scale_fill_viridis_c() +
theme(legend.position = "right")+
labs(fill = "Turnout \n Rate")Compare the visual results to your inspection of the data with {dplyr} verbs above. Can you use the map to confirm which states had the highest and lowest turnout rates?
5.4 Beyond turnout
We can also feed other variables to the plot_usmap() function, such as noncitizen_pct or the ineligible_felons_total columns in our dataset. Copy and paste the code chunk above and make the necessary modifications in the code chunks below. Try some other viridis color scales too.
5.4.1 Noncitizens
5.4.2 Ineligible felons
The above map reveals that most states literally pale in comparison to the numbers of ineligible felons in places like Texas, Florida, Georgia, and California. We can confirm just how different these states (and others) are by returning to our dot plot-type visualization we did above.
6 Summary
Look back at the initial choropleth map we did above, displaying the variation in voter turnout across the states. Observing variation like this prompts a natural but critical question for political scientists: why does voter turnout vary across the states?
What do you think may explain variation in voter turnout across the states? Work in teams to brainstorm a list of potential influences which may affect whether state-level turnout goes up or down, and add them below. Include a brief rationale about why or how you think the items on your list positively or negatively affect the turnout rate.