Week 6: Visualization in Practice with ggplot2
Student Summary and Lab Activity
1 Overview
This week connects two complementary skills. First, you will use the grammar of graphics implemented in ggplot2 to construct plots layer by layer. Second, you will apply data-visualization principles to decide whether a graph communicates quantities, distributions, and relationships accurately.
The goal is not simply to produce a plot that runs. A successful visualization should match the analytical question, represent the data faithfully, make comparisons easy, and remain interpretable to its intended audience.
1.2 Learning outcomes
By the end of this lab, you should be able to:
- construct statistical graphics by combining data, aesthetic mappings, geometries, scales, annotations, and themes;
- distinguish an aesthetic mapping from a fixed graphical setting;
- select displays for quantities, distributions, group comparisons, and relationships;
- use transformations when they improve the visibility and interpretation of skewed data;
- identify graphical choices that exaggerate, conceal, or distort differences;
- improve comparisons through meaningful ordering, common axes, adjacency, and direct display of the data; and
- produce readable graphics using clear labels and color-blind-friendly encodings.
2 Reading summary
2.1 Chapter 8: Building plots with ggplot2
ggplot2 uses a grammar of graphics: a plot is assembled from components rather than selected as a fixed chart type. The basic structure is:
data |>
ggplot(aes(x = x_variable, y = y_variable)) +
geom_name()The principal components are:
| Component | Purpose | Examples |
|---|---|---|
| Data | Supplies the observations and variables | murders, heights |
| Aesthetic mapping | Connects variables to visible properties | aes(x = population, y = total, color = region) |
| Geometry | Determines how observations are drawn | geom_point(), geom_col(), geom_histogram() |
| Scale | Controls how mapped values appear on axes or in legends | scale_x_log10(), scale_color_manual() |
| Annotation | Adds explanatory information not necessarily mapped from variables | labs(), annotate(), geom_abline() |
| Theme | Controls non-data elements and overall appearance | theme_minimal() |
2.1.1 Mapped versus fixed properties
Place a graphical property inside aes() when its values should vary according to a variable. Place it outside aes() when every observation should receive the same setting.
# Color is mapped to a variable; ggplot2 creates a legend.
ggplot(murders, aes(population, total, color = region)) +
geom_point()
# Every point is assigned the same navy color; no variable is mapped to color.
ggplot(murders, aes(population, total)) +
geom_point(color = "#003B5C")2.1.2 Common geometries
- Use
geom_point()for relationships between two quantitative variables. - Use
geom_col()when bar heights are already stored in the data; usegeom_bar()when counts should be computed from observations. - Use
geom_histogram()orgeom_density()to examine a quantitative distribution. - Use
geom_boxplot()for compact comparisons of distributions, preferably supplemented with raw observations when feasible. - Use transparency through
alphaand jitter throughgeom_jitter()to reduce overplotting.
2.1.3 Scales and transformations
A log scale can make multiplicative relationships and right-skewed data easier to inspect. A transformation should have an analytical reason and must be clearly communicated in the axis label. Do not use a transformation simply to make a figure look more balanced.
2.2 Chapter 9: Designing truthful and interpretable graphics
Chapter 9 moves from plot construction to graphical judgment. Its central principles include the following.
- Use visual cues people can compare accurately. Position along a common axis and aligned length are generally easier to judge than angle, area, brightness, or color intensity.
- Use zero appropriately. Bars encode quantity through length, so their baseline should normally be zero. Position-based displays, such as scatterplots or dot plots, need not always include zero when a restricted range clarifies meaningful variation.
- Do not distort quantities. Area-based symbols must scale by area rather than radius, but position or length is usually a clearer encoding.
- Order categories meaningfully. Alphabetical order is rarely the most informative choice. Order categories by the displayed value or another quantity connected to the question.
- Show the data. A mean with an error bar can conceal distributional shape, overlap, outliers, sample size, and within-group variability.
- Make comparisons easy. Use common axes, align plots in the direction of the comparison, and place the groups being compared next to each other.
- Use transformations deliberately. Log scales can reveal structure in right-skewed data and display multiplicative changes symmetrically.
- Use accessible colors. Do not depend on red-green differences alone. Combine color with position, shape, labels, or facets when needed.
- Avoid pseudo-three-dimensional displays. They introduce perspective and occlusion without adding reliable information.
- Match precision and complexity to the audience. Remove unnecessary digits and ensure that titles, axes, legends, and transformations are understandable to the intended reader.
3 Lab activity
3.1 Expected time and submission
- Estimated time: 90–120 minutes
- Submit: the completed
.qmdfile and its rendered HTML file - Evidence required: executable R code, all requested figures, and concise written interpretations
- Reproducibility: restart R and render the document before submitting it
3.2 AI-use classification: AI-Permitted
Complete an initial attempt before consulting generative AI. You may use AI to explain an error message, locate relevant documentation, or critique a visualization you have already created. You may not ask AI to complete the entire lab or make the analytical decisions for you.
If you use AI, add a brief disclosure at the end that states what you asked, what suggestion you used or rejected, and how you verified the result. You remain responsible for the code, figures, and interpretations.
3.3 Setup
The lab uses ggplot2, dplyr, and the dslabs datasets. Install a missing package from the Console before rendering; do not place installation commands in the submitted document.
library(ggplot2)
library(dplyr)
library(dslabs)
colorblind_palette <- c(
"#0072B2", "#D55E00", "#009E73", "#CC79A7",
"#E69F00", "#56B4E9", "#F0E442", "#000000"
)
theme_set(theme_minimal(base_size = 12))3.4 Part A: Constructing plots with the grammar of graphics
3.4.1 Exercise 1: Data, mappings, and geometry
Adapted from Chapter 8, Exercises 1–7.
- Inspect the structure and variable names of
murders. - Create a
ggplotobject namedpassociated withmurdersbut with no geometry. Print it and explain why the result is a blank plotting area. - Add a scatterplot layer with population on the horizontal axis and total murders on the vertical axis.
- In two or three sentences, identify the data, mapping, and geometry used in your completed plot.
# 1. Structure, variable names, and a check for missing values
str(murders)'data.frame': 51 obs. of 5 variables:
$ state : chr "Alabama" "Alaska" "Arizona" "Arkansas" ...
$ abb : chr "AL" "AK" "AZ" "AR" ...
$ region : Factor w/ 4 levels "Northeast","South",..: 2 4 4 2 4 4 1 2 2 2 ...
$ population: num 4779736 710231 6392017 2915918 37253956 ...
$ total : num 135 19 232 93 1257 ...
names(murders)[1] "state" "abb" "region" "population" "total"
colSums(is.na(murders)) state abb region population total
0 0 0 0 0
# 2. A ggplot object with data but no mapping or geometry
p <- ggplot(data = murders)
p# 3. Add a scatterplot layer
p +
geom_point(aes(x = population / 10^6, y = total)) +
labs(
title = "Total gun murders vs. state population, 2010",
x = "Population (millions)",
y = "Total gun murders"
)Interpretation:
murders has 51 rows (the 50 states plus the District of Columbia), five variables (state, abb, region, population, total), and no missing values. Printing p gives a blank panel because the object knows only which data to use. It has no aesthetic mapping telling it which variables go on the axes and no geometry telling it how to draw them.
In the completed plot, the data is murders, the mapping is aes(x = population / 10^6, y = total) (population in millions), and the geometry is geom_point(), which draws one point per state. States with more people have more murders in total. Most points are crowded near the origin because a few very populous states stretch both axes.
3.4.2 Exercise 2: Mapped versus fixed color
Adapted from Chapter 8, Exercises 8–13.
Create two versions of the scatterplot from Exercise 1:
- Assign the same navy color to every point.
- Map point color to
regionusingaes()and applycolorblind_palettewithscale_color_manual(). - Label the axes and legend clearly.
- Explain why one version produces a legend and the other does not.
# Version 1: color is a fixed setting (outside aes)
ggplot(murders, aes(x = population / 10^6, y = total)) +
geom_point(color = "navy", size = 2) +
labs(
title = "Total gun murders vs. state population, 2010",
x = "Population (millions)",
y = "Total gun murders"
)# Version 2: color is mapped to a variable (inside aes)
ggplot(murders, aes(x = population / 10^6, y = total, color = region)) +
geom_point(size = 2) +
scale_color_manual(values = colorblind_palette) +
labs(
title = "Total gun murders vs. state population by region, 2010",
x = "Population (millions)",
y = "Total gun murders",
color = "Region"
)Explanation:
In Version 1, color = "navy" is placed outside aes(). It is a fixed setting applied to every point regardless of the data, so no variable is linked to color, there is nothing to decode, and ggplot2 draws no legend.
In Version 2, color = region is placed inside aes(), so it is a mapping. ggplot2 builds a color scale that assigns each of the four regions a color (the first four colors of colorblind_palette, via scale_color_manual()). The legend is the key the reader needs to translate colors back to regions.
3.4.3 Exercise 3: Labels, scales, and annotations
Adapted from Chapter 8, Exercises 9 and 14–16.
- Create a scatterplot of total murders against population.
- Use state abbreviations as text labels.
- Place both axes on log10 scales.
- Add an informative title, axis labels that identify the log scales, and a caption naming
dslabsas the data source. - Explain what becomes easier to see after the transformation and what a reader could misunderstand if the transformation were not labeled.
# National murder rate (murders per million people), used as a reference line
us_rate_per_million <- sum(murders$total) / sum(murders$population) * 10^6
ggplot(murders, aes(x = population / 10^6, y = total, label = abb)) +
geom_abline(
intercept = log10(us_rate_per_million), slope = 1,
linetype = "dashed", color = "grey50"
) +
geom_text(size = 3) +
scale_x_log10() +
scale_y_log10() +
labs(
title = "US gun murders vs. population by state, 2010",
subtitle = "Dashed line = national average rate (about 30 murders per million people)",
x = "Population in millions (log10 scale)",
y = "Total gun murders (log10 scale)",
caption = "Source: dslabs package, murders dataset (FBI data)"
)Interpretation:
On the original scale, most states were squeezed into one corner. On log10 axes, equal distances represent equal multiplicative changes (1 to 10 million is the same distance as 10 to 100 million), so the states spread out and every label is readable. The points fall close to a straight line with slope 1, which means murders grow roughly in proportion to population. The dashed line shows which states are above the national rate (e.g., DC, LA, MO, MD) and which are below it (e.g., VT, NH, HI, ID).
Without the “log10 scale” labels, a reader would assume linear axes and underestimate the differences. The gap between California (1,257 murders) and Vermont (2 murders) looks moderate here but is more than 600-fold.
3.4.4 Exercise 4: Histograms and analytical choices
Adapted from Chapter 8, Exercises 17–20.
Using the heights data:
- Create two histograms of
height, one withbinwidth = 1and another withbinwidth = 3. - Give both plots the same horizontal limits and informative labels.
- Describe one feature that appears more or less prominent when the bin width changes.
- State why bin width is an analytical choice rather than a cosmetic setting.
# Check for missing heights before plotting
sum(is.na(heights$height))[1] 0
# coord_cartesian() fixes the visible range without dropping any observations
ggplot(heights, aes(x = height)) +
geom_histogram(binwidth = 1, fill = colorblind_palette[1], color = "white") +
coord_cartesian(xlim = c(48, 85)) +
labs(
title = "Self-reported student heights (bin width = 1 inch)",
x = "Height (inches)",
y = "Number of students"
)ggplot(heights, aes(x = height)) +
geom_histogram(binwidth = 3, fill = colorblind_palette[1], color = "white") +
coord_cartesian(xlim = c(48, 85)) +
labs(
title = "Self-reported student heights (bin width = 3 inches)",
x = "Height (inches)",
y = "Number of students"
)Comparison:
With binwidth = 1, the histogram shows a spike at 72 inches (106 students, compared with 78 at 71 inches and 34 at 73 inches), likely from people rounding their height to “six feet”. It also shows a few isolated values between 50 and 55 inches, separated from the rest by an empty stretch up to 58 inches. With binwidth = 3, the spike disappears, the distribution looks smooth and roughly symmetric around 68–69 inches, and the low values blend into the left tail.
Bin width is an analytical choice because it decides which features a reader can see. Bins that are too narrow show noise and rounding as if they were real structure, and bins that are too wide hide outliers, gaps, or multiple peaks. A different bin width can lead to a different conclusion.
3.4.5 Exercise 5: Comparing distributions
Adapted from Chapter 8, Exercises 21–24.
- Create density plots of height for the groups represented by
sex. - Use a color-blind-friendly palette,
alphatransparency, clear labels, and a descriptive title. - Then create a boxplot of height by
sexwith jittered observations overlaid. - Compare what the density plot and the boxplot-plus-points reveal. Do not make causal claims.
# Density plot by sex
ggplot(heights, aes(x = height, fill = sex)) +
geom_density(alpha = 0.4, color = NA) +
scale_fill_manual(values = colorblind_palette[c(2, 1)]) +
labs(
title = "Self-reported heights of female and male students",
x = "Height (inches)",
y = "Density",
fill = "Sex"
)# Boxplot with jittered observations
# outlier.shape = NA avoids drawing outliers twice, since every point is shown
ggplot(heights, aes(x = sex, y = height)) +
geom_boxplot(outlier.shape = NA, width = 0.5) +
geom_jitter(aes(color = sex), width = 0.2, alpha = 0.3, size = 1) +
scale_color_manual(values = colorblind_palette[c(2, 1)], guide = "none") +
labs(
title = "Self-reported heights by sex, with every student shown",
x = "Sex",
y = "Height (inches)"
)Comparison:
The density plot shows the shape of each distribution. Both are roughly bell-shaped, the male distribution sits to the right of the female one (medians about 69 vs. 65 inches), and the two overlap a great deal. Each curve has an area of 1, so the plot hides how many students are in each group, and the smoothing hides individual extreme values.
The boxplot with jittered points shows medians and quartiles directly and every observation. It reveals that the groups are very different in size (812 male vs. 238 female respondents). It also shows unusual reports in both groups: very short heights (under 56 inches) from 5 female and 4 male respondents, and female reports of 77–79 inches. Some of these may be reporting errors. In this sample, male students tend to report greater heights, with substantial overlap. The data describe these self-reports and do not explain why the difference exists.
3.5 Part B: Applying visualization principles
3.5.1 Exercise 6: Meaningful category order
Adapted from Chapter 9, Exercises 3–5.
The following code constructs state-level measles rates for 1967.
measles_1967 <- us_contagious_diseases |>
filter(
year == 1967,
disease == "Measles",
!is.na(population),
!is.na(count),
weeks_reporting > 0
) |>
mutate(
rate = count / population * 10000 * 52 / weeks_reporting
)- Create a horizontal bar chart with states in the default order.
- Create a revised chart that orders states by
rate. - Use a zero baseline, readable labels, and a title that includes the year and measurement unit.
- Explain which version better supports identification of the highest- and lowest-rate states.
# Number of states kept after filtering (the 1967 data has no missing values)
nrow(measles_1967)[1] 51
# Default (alphabetical) order
ggplot(measles_1967, aes(x = rate, y = state)) +
geom_col(fill = colorblind_palette[1]) +
scale_x_continuous(expand = expansion(mult = c(0, 0.05))) +
labs(
title = "Measles rate by state, 1967 (alphabetical order)",
x = "Measles cases per 10,000 people per year",
y = NULL
) +
theme(axis.text.y = element_text(size = 8))# Ordered by rate
ggplot(measles_1967, aes(x = rate, y = reorder(state, rate))) +
geom_col(fill = colorblind_palette[1]) +
scale_x_continuous(expand = expansion(mult = c(0, 0.05))) +
labs(
title = "Measles rate by state, 1967 (ordered by rate)",
x = "Measles cases per 10,000 people per year",
y = NULL
) +
theme(axis.text.y = element_text(size = 8))Evaluation:
The ordered version is better. In alphabetical order, the reader has to scan all 51 bars and compare lengths that may be far apart to find the extremes. Ordered by rate, the ranking is shown by position: the highest rates (Washington about 17.7, North Dakota about 14.5, Texas about 12.5 cases per 10,000) are at the top, the lowest (Georgia about 0.1, DC and Connecticut about 0.4) are at the bottom, and neighboring bars are directly comparable. Alphabetical order only helps look up a single state. Both charts start the bars at zero (geom_col() uses a zero baseline, and expand removes the gap), so bar length is proportional to the rate.
3.5.2 Exercise 7: Show the data, not only an average
Adapted from Chapter 9, Exercises 6–7.
- Calculate state murder rates per 100,000 residents.
- Reproduce a bar chart containing only the mean state rate for each region.
- Replace it with a boxplot that shows the regional distributions, orders regions by median rate, and overlays jittered state observations.
- Explain why choosing a region based only on the regional mean would be poorly supported. Refer specifically to within-region variability and overlap.
# 1. Murder rate per 100,000 residents
murders_rate <- murders |>
mutate(rate = total / population * 100000)
region_summary <- murders_rate |>
group_by(region) |>
summarize(
mean_rate = mean(rate),
median_rate = median(rate),
q1 = quantile(rate, 0.25),
q3 = quantile(rate, 0.75),
min_rate = min(rate),
max_rate = max(rate),
n_states = n()
)
region_summary# A tibble: 4 × 8
region mean_rate median_rate q1 q3 min_rate max_rate n_states
<fct> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl> <int>
1 Northeast 1.85 1.80 0.828 2.71 0.320 3.60 9
2 South 4.42 3.40 3.00 4.23 1.46 16.5 17
3 North Central 2.18 1.97 0.995 2.72 0.595 5.36 12
4 West 1.83 1.29 0.887 3.11 0.515 3.63 13
# 2. Bar chart of the mean state rate only
ggplot(region_summary, aes(x = region, y = mean_rate)) +
geom_col(fill = colorblind_palette[1], width = 0.6) +
labs(
title = "Mean state gun murder rate by region, 2010",
x = "Region",
y = "Mean murder rate (per 100,000 residents)"
)# 3. Boxplot ordered by median, with every state shown
ggplot(murders_rate, aes(x = reorder(region, rate, FUN = median), y = rate)) +
geom_boxplot(outlier.shape = NA, width = 0.5) +
geom_jitter(aes(color = region), width = 0.15, size = 2, alpha = 0.8) +
geom_text(
data = filter(murders_rate, abb == "DC"),
aes(label = "DC"), nudge_x = 0.25, size = 3.5
) +
scale_color_manual(values = colorblind_palette, guide = "none") +
labs(
title = "State gun murder rates by region, 2010",
subtitle = "Regions ordered by median state rate; each point is one state",
x = "Region",
y = "Murder rate (per 100,000 residents)",
caption = "Source: dslabs package, murders dataset"
)Evaluation:
The bar chart reduces each region to one number: South 4.4, North Central 2.2, Northeast 1.85, West 1.83. The boxplot shows why ranking regions on these means is poorly supported.
- Within-region variability is large. State rates range from about 1.5 to 16.5 in the South, 0.6 to 5.4 in North Central, 0.5 to 3.6 in the West, and 0.3 to 3.6 in the Northeast. One mean says nothing about this spread.
- Overlap. The Northeast, North Central, and West distributions overlap almost completely (their middle 50% all span roughly 0.8–3.1), so their mean differences are small compared with the overlap and variation among states. The South’s middle 50% (3.0–4.2) sits higher. But the lowest Southern states (West Virginia 1.5, Kentucky 2.7, Alabama 2.8) fall inside the other regions’ ranges, and several Western and Northeastern states (e.g., Arizona 3.6, Pennsylvania 3.6) have higher rates than some Southern states.
- The mean is pulled by outliers. DC (16.5) raises the Southern mean (4.4) well above its median (3.4). The West and Northeast have almost equal means but different medians (1.29 vs. 1.80), so ordering by mean and ordering by median do not agree.
The choice of state within a region can matter as much as the choice of region, and a single mean hides that.
3.5.3 Exercise 8: Visualization audit and redesign
Synthesizes Chapter 9 principles.
Return to one figure you created in Exercises 1–7. Audit it using the questions below, then revise the figure.
- Does the geometry match the analytical question?
- Are quantities encoded through position or length when feasible?
- Is zero included when length is the visual cue?
- Are categories ordered to support the comparison?
- Are relevant observations or distributions visible?
- Are scales, units, transformations, and data sources labeled?
- Can the figure be interpreted without distinguishing red from green?
- Are the title and precision appropriate for the intended audience?
Figure audited: the ordered measles bar chart from Exercise 6. Analytical question: Which states had the highest and lowest measles rates in 1967, and is there a geographic pattern?
| Audit question | Ordered measles chart (Exercise 6) |
|---|---|
| 1. Geometry matches question? | Yes. Bars suit one quantity per state. |
| 2. Position/length encoding? | Yes, rate is encoded by bar length. |
| 3. Zero with length cues? | Yes, bars start at zero. |
| 4. Categories ordered? | Ordered by rate, but the 51 states are one long list, so a regional pattern cannot be seen. |
| 5. Observations visible? | Every state is shown, but nothing tells the reader that 20 states’ rates are scaled up from fewer than 52 weeks of reports (Kansas reported only 34 weeks). |
| 6. Scales, units, source labeled? | Unit is labeled; no data source and no national benchmark. |
| 7. Readable without red–green? | Yes, a single color is used. |
| 8. Title and precision appropriate? | The title describes the chart but not the finding. |
# Add Census regions from `murders`; DC is spelled differently in the two datasets
measles_region <- measles_1967 |>
mutate(
state = as.character(state),
state = if_else(state == "District Of Columbia", "District of Columbia", state)
) |>
left_join(select(murders, state, region), by = "state")
# Order the region panels by median rate, highest first
region_medians <- tapply(measles_region$rate, measles_region$region, median)
measles_region <- mutate(
measles_region,
region = factor(region, levels = names(sort(region_medians, decreasing = TRUE)))
)
# Every state received a region
sum(is.na(measles_region$region))[1] 0
# Population-weighted US rate, used as a reference line
us_measles_rate <- sum(measles_region$rate * measles_region$population) /
sum(measles_region$population)
ggplot(measles_region, aes(x = rate, y = reorder(state, rate))) +
geom_col(fill = colorblind_palette[1]) +
geom_vline(xintercept = us_measles_rate, linetype = "dashed", color = "grey30") +
geom_text(
data = filter(measles_region, weeks_reporting < 40),
aes(label = paste0("rate estimated from ", weeks_reporting, " of 52 weeks")),
hjust = -0.05, size = 3
) +
facet_grid(region ~ ., scales = "free_y", space = "free_y") +
scale_x_continuous(expand = expansion(mult = c(0, 0.05))) +
labs(
title = "Measles rates in 1967 tended to be highest in the West\nand lowest in the Northeast",
subtitle = sprintf(
"Cases per 10,000 people per year. Dashed line = US rate (%.1f).\nRegions ordered by median rate; states ordered by rate within each region.",
us_measles_rate
),
x = "Measles cases per 10,000 people per year",
y = NULL,
caption = paste(
"Rates for states reporting fewer than 52 weeks are scaled to a full year.",
"Source: dslabs::us_contagious_diseases; regions from dslabs::murders",
sep = "\n"
)
) +
theme(
plot.title = element_text(size = 11, face = "bold"),
plot.subtitle = element_text(size = 9),
axis.text.y = element_text(size = 7),
strip.text.y = element_text(angle = 0, face = "bold")
)Audit summary:
- Grouped by region with facets (Q4). A single list of 51 states hid any geographic pattern. With states grouped by region, panels ordered by median rate, and states still ordered by rate within each panel, the reader can see that Western states tended to have high rates (median 6.2) and Northeastern states low rates (median 0.7), while still finding individual extremes.
- US reference line (Q6, Q8). The dashed line at the population-weighted national rate (3.0) gives the reader a benchmark, so each bar reads as above or below the national level rather than only as a number.
- Flagging estimated rates (Q5, Q6). The rate formula scales partial-year counts to a full year. The caption states this for all 20 affected states, and Kansas, whose rate is based on only 34 of 52 weeks, is labeled directly. This keeps a less reliable value from looking as solid as the others.
- Informative title and source caption (Q6, Q8). The title states the main pattern, and the caption documents the data sources, including where the regions came from.
The zero baseline, length encoding, ordering by rate, and single color-blind-safe color from the original were kept.
3.6 Graduate extension
Complete this section if assigned for the graduate version of the course.
3.6.1 Exercise 9: Sensitivity of distribution displays
Using one quantitative variable from heights or murders, construct a small sensitivity analysis:
- Compare at least three defensible histogram bin widths or density bandwidth adjustments.
- Keep all other mappings and scales constant.
- Identify which features persist across the displays and which depend on the smoothing or binning choice.
- Defend one display for exploratory analysis and one for communication to a general audience. They may be the same, but the justification must address the audience and the risk of overinterpretation.
# Same variable, mapping, x-range, and y-scale for every panel; only bin width changes.
# The y-axis is density (not count) because counts per bin grow with wider bins.
# center = 0 centers every bin on a whole inch, matching the bins in Exercise 4.
bin_widths <- c(0.5, 1, 2, 4)
for (bw in bin_widths) {
print(
ggplot(heights, aes(x = height)) +
geom_histogram(
aes(y = after_stat(density)),
binwidth = bw, center = 0,
fill = colorblind_palette[1], color = "white"
) +
coord_cartesian(xlim = c(48, 85), ylim = c(0, 0.25)) +
labs(
title = paste0(
"Self-reported student heights (bin width = ", bw,
if (bw == 1) " inch)" else " inches)"
),
x = "Height (inches)",
y = "Density"
)
)
}# Share of heights reported as whole inches
mean(heights$height == round(heights$height))[1] 0.7838095
# Kernel density with three bandwidth adjustments on the same axes
ggplot(heights, aes(x = height)) +
geom_density(aes(color = "adjust = 0.25"), adjust = 0.25, linewidth = 0.7) +
geom_density(aes(color = "adjust = 1 (default)"), adjust = 1, linewidth = 0.7) +
geom_density(aes(color = "adjust = 3"), adjust = 3, linewidth = 0.7) +
scale_color_manual(values = colorblind_palette[c(2, 1, 3)]) +
coord_cartesian(xlim = c(48, 85)) +
labs(
title = "Density estimates of self-reported student heights, three bandwidths",
x = "Height (inches)",
y = "Density",
color = "Bandwidth"
)Sensitivity analysis:
Features that persist in every display: one main peak near 68–69 inches, most students between about 60 and 76 inches, and a long lower tail reaching about 50 inches. These are properties of the data, not of the display settings.
Features that depend on the choice:
- Rounding artifacts. About 78% of heights are reported as whole inches. At bin width 0.5 the histogram becomes a “picket fence” of tall bars at whole inches and near-empty half-inch bins, and with
adjust = 0.25the density curve has a spike at every inch. At bin width 1, one heaping effect remains: 72 inches is clearly higher than 71 or 73. These spikes reflect how people report height, not real features of height, and they disappear at bin width 2 or 4 and at the default or wider bandwidths. - Low outliers and the gap. Narrow bins clearly show the isolated values at 50–55 inches and the empty stretch between 55 and 58 inches. Wide bins and smooth densities absorb them into the left tail.
- Shape at wide settings. At bin width 4, the bins are coarse and the peak is blunt. With
adjust = 3, the peak is lower and the curve wider, so the distribution looks more spread out than the data show.
For exploratory analysis: bin width 1. An analyst needs to see heaping at whole inches (a data-collection feature) and implausibly short values that may need checking. An analyst also knows that small bumps may be noise, so the risk of overreading them is low.
For a general audience: bin width 2. A general reader tends to treat every bump as meaningful. Bin width 2 keeps the true center, spread, and tail, but removes rounding spikes that could be misread as “many people are exactly six feet tall”. It is not wide enough to distort the spread the way bin width 4 begins to.
3.6.2 Exercise 10: Redesign under competing goals
Create two visualizations from the same data and analytical question:
- an exploratory version intended for a data analyst; and
- a communication version intended for a general audience.
The two figures may differ in annotation, labeling, transformation, displayed detail, or use of direct labels. Explain what information you preserved, simplified, or emphasized in each version. Neither version may distort the underlying quantities.
Data and question: murders, How do state gun murder rates differ across US regions?
# Reuse the rates from Exercise 7 and the national rate from Exercise 3
us_rate <- us_rate_per_million / 10 # murders per 100,000 residents
# Fix one region order (ascending median) for every layer of both plots
region_order <- region_summary$region[order(region_summary$median_rate)]
murders_ordered <- mutate(murders_rate, region = factor(region, levels = region_order))
summary_ordered <- mutate(region_summary, region = factor(region, levels = region_order))
# ---- Exploratory version (for a data analyst) ----
ggplot(murders_ordered, aes(x = region, y = rate)) +
geom_boxplot(outlier.shape = NA, width = 0.5, color = "grey40") +
geom_text(
aes(label = abb, color = region),
position = position_jitter(width = 0.2, seed = 1),
size = 2.8
) +
geom_point(
data = summary_ordered,
aes(x = region, y = mean_rate),
shape = 4, size = 4, stroke = 1.2, inherit.aes = FALSE
) +
geom_text(
data = summary_ordered,
aes(x = region, y = 0.25, label = paste0("n = ", n_states)),
size = 3, inherit.aes = FALSE
) +
scale_y_log10() +
scale_color_manual(values = colorblind_palette, guide = "none") +
labs(
title = "State gun murder rate by region, 2010 (exploratory)",
subtitle = "Box = median and IQR; X = mean; labels = states; regions ordered by median",
x = "Region",
y = "Murders per 100,000 residents (log10 scale)",
caption = "Source: dslabs::murders"
)# ---- Communication version (for a general audience) ----
ggplot(murders_ordered, aes(x = rate, y = region)) +
geom_vline(xintercept = us_rate, linetype = "dashed", color = "grey40") +
# Benchmark label sits above the top row so it does not cover any states
annotate(
"text", x = us_rate + 0.2, y = Inf, vjust = 1.2, hjust = 0,
label = sprintf("US average: %.1f", us_rate),
size = 3.5, color = "grey30"
) +
geom_point(
position = position_jitter(height = 0.12, seed = 1),
color = colorblind_palette[1], alpha = 0.6, size = 2.5
) +
geom_point(
data = summary_ordered,
aes(x = median_rate, y = region),
shape = "|", size = 9, color = colorblind_palette[2], inherit.aes = FALSE
) +
annotate("text", x = 16.45, y = "South", vjust = -1.3, hjust = 1,
label = "Washington, DC", size = 3.5) +
annotate("text", x = 7.74, y = "South", vjust = -1.3, hjust = 0.5,
label = "Louisiana", size = 3.5) +
annotate("text", x = 0.32, y = "Northeast", vjust = 3.6, hjust = 0,
label = "Vermont", size = 3.5) +
scale_x_continuous(limits = c(0, 17), breaks = seq(0, 16, 4)) +
# annotations use plain text for y, so set the category order explicitly
scale_y_discrete(limits = region_order, expand = expansion(add = c(0.6, 1))) +
labs(
title = "Southern states tend to have higher gun murder rates",
subtitle = "Each dot is a state. Orange bar = the typical (median) state in each region.",
x = "Gun murders per 100,000 residents (2010)",
y = NULL,
caption = "Source: FBI data via the dslabs R package"
) +
theme(
plot.title = element_text(face = "bold"),
panel.grid.major.y = element_blank()
)Design justification:
Exploratory version. An analyst needs detail for checking assumptions. I preserved every state, labeled by abbreviation so outliers can be identified. I also showed the median and IQR, the mean as an X (so the gap between mean and median is visible), and the number of states per region. A log10 scale, labeled on the axis, spreads out the right-skewed rates so DC does not compress the other states.
Communication version. A general reader needs one clear message and enough context not to misread it. I simplified the chart by removing the boxplot, the mean, the sample sizes, and the log scale. The axis is linear and starts at zero, so distances correspond directly to rates. I emphasized the takeaway in the title, marked the typical state with a median bar described in plain words, added the US average as a benchmark, and labeled three reference states (DC, Louisiana, Vermont). Every state remains a dot, so the reader can still see the overlap between regions, which guards against the misreading that every Southern state has a high rate.
In both versions regions are ordered by median, all values are the same state-level rates, and the only transformation (log10) is labeled. Neither version distorts the quantities.
4 Final verification checklist
Before submitting, confirm that:
5 AI-use disclosure
- Tool and purpose: I used Claude (Anthropic) after completing my own initial attempt. I used it to check my reasoning and interpretations, review grammar and clarity, review my rendered HTML against the lab instructions, and suggest small formatting fixes.
- What I accepted, modified, or rejected: I accepted some wording improvements to my interpretations after checking them against my own output. I accepted two formatting fixes: showing population in millions on the Exercise 1 x-axis (to replace scientific notation), and moving the “US average” label in the Exercise 10 communication chart so it no longer covers the data points. I also restored the missing Sources heading. I did not accept suggestions that did not accurately reflect my analysis. The analytical choices, code, and conclusions are my own.
- How I verified the result: I restarted R, rendered the document with Quarto without errors, and inspected every figure and numerical result in the HTML output to confirm that it was displayed correctly and supported my interpretations.
6 Sources
This student summary and lab were developed from:
- Irizarry, R. A. Introduction to Data Science, Chapter 8: ggplot2.
- Irizarry, R. A. Introduction to Data Science, Chapter 9: Data visualization principles.
The lab questions select, combine, and adapt concepts and exercises from both chapters for instructional use.