Exploratory data analysis helps us learn what is in a data set before we model it.
What does each row represent?
Which variables need cleaning?
Are values missing, unusual, or implausible?
Which comparisons deserve a closer look?
Cleaning and renaming columns
clean_names() first converts the original labels to lowercase snake case. rename() then gives the measures short, readable names we can use throughout the lesson.
# A tibble: 8 × 2
state fatal_collisions
<chr> <dbl>
1 North Dakota 23.9
2 South Carolina 23.9
3 West Virginia 23.8
4 Arkansas 22.4
5 Kentucky 21.4
6 Montana 21.4
7 Louisiana 20.5
8 Oklahoma 19.9
Collision-rate distribution
The histogram counts states in rate ranges. The density plot shows the same distribution as a smooth curve. patchwork places the two plots side by side.
plot_state_comparison <- drivers |>ggplot(aes(x =reorder(state, fatal_collisions), y = fatal_collisions)) +geom_col(fill ="navy") +coord_flip() +labs(title ="Fatal-collision rates by state",x =NULL,y ="Drivers involved in fatal collisions per billion miles" ) +theme_minimal(base_size =12)
Comparing all states: rendered plot
A relationship to explore
Does the share of drivers who were speeding rise with the fatal-collision rate?
Code
plot_speeding_collision <- drivers |>ggplot(aes(x = speeding, y = fatal_collisions)) +geom_point(size =2.5, alpha =0.75, color ="navy") +geom_smooth(method ="lm", se =FALSE, color ="red") +labs(title ="Speeding and fatal-collision rates",x ="Drivers involved in fatal collisions who were speeding (%)",y ="Fatal collisions per billion miles" ) +theme_minimal(base_size =16)
Speeding and fatal-collision rates: rendered plot
Label unusual observations
Labels can make a plot easier to discuss. Here, geom_label_repel() labels only the three states with the largest collision rates and prevents overlapping labels.
plot_alcohol_collision <- drivers |>ggplot(aes(x = alcohol_impaired, y = fatal_collisions)) +geom_point(size =2.5, alpha =0.75, color ="navy") +geom_smooth(method ="lm", se =FALSE, color ="red") +labs(title ="Alcohol impairment and fatal-collision rates",x ="Drivers involved in fatal collisions who were alcohol-impaired (%)",y ="Fatal collisions per billion miles" ) +theme_minimal(base_size =16)
Alcohol impairment and collision rates: rendered plot
Insurance premiums and losses
Code
plot_premium_losses <- drivers |>ggplot(aes(x = premium, y = losses)) +geom_point(size =2.5, alpha =0.75, color ="navy") +geom_smooth(method ="lm", se =FALSE, color ="red") +scale_x_continuous(labels = scales::dollar) +scale_y_continuous(labels = scales::dollar) +labs(title ="Insurance premiums and insurer losses",x ="Average car-insurance premium",y ="Insurer losses per insured driver" ) +theme_minimal(base_size =16)
Insurance premiums and losses: rendered plot
Stitching relationship plots with GGally
GGally::ggpairs() stitches scatterplots, distributions, and correlations into one grid. Use it to screen several relationships, then make one focused chart for your audience.
Code
plot_relationship_grid <- drivers |>select(fatal_collisions, speeding, alcohol_impaired, premium, losses) |>ggpairs(lower =list(continuous =wrap("points", color ="navy")),diag =list(continuous =wrap("densityDiag", fill ="navy", color ="navy")) )
GGally relationship grid: rendered plot
Ridge plots with ggridges
Ridge plots compare the distribution of one numeric measure across groups. Here, each ridge shows fatal-collision rates for a broad U.S. region.
Code
region_lookup <-tibble(state =c(state.name, "District of Columbia"),region =c(as.character(state.region), "South"))plot_ridges <- drivers |>left_join(region_lookup, by ="state") |>ggplot(aes(x = fatal_collisions, y = region)) +geom_density_ridges(fill ="navy", color ="navy", alpha =0.7) +labs(title ="Fatal-collision rates by broad U.S. region",x ="Drivers involved in fatal collisions per billion miles",y =NULL ) +theme_minimal(base_size =16)
A correlation describes how two numeric variables move together in this data set. It does not establish why they move together.
Correlation heat map
A heat map uses color to show the direction and strength of correlations. GGally::ggcorr() draws the plot directly from a correlation matrix. We keep only one triangular half, since every correlation otherwise appears twice.