Week 2: Principles of data visualisationWeek 3: Grammar of graphics; aesthetics and attributes- Week 4: Major visualisation tools
- Week 5: Customising visualisations (scales, themes, and labels)
By the end of the session, you should be able to:
Download the Week 4 files from NOW and move them into your RStudio Project folder.
Start with the question and the variables involved.
| Question | Typical variable types | Useful starting point |
|---|---|---|
| What values occur? | one continuous variable | histogram, density plot |
| Are two measures related? | two continuous variables | scatterplot |
| Do groups differ? | categorical group + continuous outcome | jitter/boxplot/violin |
| How many observations are in each group? | one categorical variable | bar plot |
Plot type follows the question and variable types, not the other way round.
Before choosing a plot, identify whether each variable is continuous or categorical.
| Variable combination | What it supports | Examples |
|---|---|---|
| one continuous variable | distribution of values | histogram, density plot |
| one categorical variable | counts or proportions | bar chart |
| continuous + continuous | relationship | scatterplot, density contours |
| categorical + continuous | group comparison | jitter, boxplot, violin, raincloud |
| categorical + categorical | counts across groups | grouped bar chart, heatmap |
The geom should match the structure of the variables you want to compare.
For each visualisation today, ask three things:
A plot type is a design decision: it makes some comparisons easier and others harder.
Last week, we used principles as a checklist for judging plots (Cleveland & McGill, 1984; Hartwig & Dearing, 1979; Tufte, 1983). This week we use the same checklist to choose plot types.
| Week 2 principle | Week 4 plot-choice question |
|---|---|
| Show the data and variation | Do we need raw points, distributions, or uncertainty? |
| Make comparisons easy | Which plot makes the target comparison quickest to see? |
| Avoid misleading summaries | What does this plot hide, average away, or imply? |
| Label enough for interpretation | Are variables, groups, and units clear? |
Choosing a plot type is how we put the Week 2 principles into practice.
Some visual encodings make comparisons easier than others (Cleveland & McGill, 1984).
What types of data visualisations have you seen in journal articles, news, reports, dashboards, or social media?
You do not need to know the correct name. Describe or sketch the idea: what does it look like, and what does it seem to show?
There are many named chart types. We will not read or memorise this list. Use it as a quick reminder that visualisation tools vary with the question, the variables, and the data structure.
Amounts and proportions
Distributions
Relationships
Change over time
Places, networks, and structures
The name matters less than the comparison: What are we trying to see?
Think of it like this:
ggplot2 Has Many Geomsgeom_) control visual encoding of aesthetics layergeom_ are part of ggplot2 (below)geoms in other packages such as ggdist, ggbeeswarm, and ggridges[1] abline area bar bin_2d [5] bin2d blank boxplot col [9] column contour contour_filled count [13] crossbar curve density density_2d [17] density_2d_filled density2d density2d_filled dotplot [21] errorbar errorbarh freqpoly function [25] hex histogram hline jitter [29] label line linerange map [33] path point pointrange polygon [37] qq qq_line quantile raster [41] rect ribbon rug segment [45] sf sf_label sf_text smooth [49] spoke step text tile [53] violin vline
Use geometries as tools for comparison, not as a menu of effects.
| Geometry | Function | What it shows |
|---|---|---|
geom_point() |
Scatter plot | Relationship between two variables |
geom_line() |
Line plot | Trends over time or ordered categories |
geom_bar() |
Bar chart | Counts or values for categories |
geom_histogram() |
Histogram | Distribution of a single variable |
geom_boxplot() |
Boxplot | Summary of distribution (median, quartiles, outliers) |
geom_text() |
Text labels | Add labels to points or bars |
A histogram answers: what values occur, and how often? It reveals distribution shape, but the answer depends on binning.
ggplot(data = d_spellname,
mapping = aes(x = rt)) +
geom_histogram(bins = 30)Open 01_choosing_plot_types.Rmd.
Complete now:
Stop before Question 3.
A scatterplot answers: do two continuous variables vary together? It reveals individual observations, but dense regions can overplot.
ggplot(data = d_spellname,
mapping = aes(x = dur,
y = rt)) +
geom_point() +
stat_smooth(method = "lm")A time plot answers: how does a value change across an ordered sequence? It reveals trends, cycles, and departures from the pattern.
ggplot(data = passenger_data,
mapping = aes(x = year,
y = passengers,
colour = month)) +
geom_point() +
stat_smooth()A trend line answers: what is the overall pattern? It reveals direction, but it can hide local variation and individual cases.
ggplot(data = d_spellname,
mapping = aes(x = dur,
y = rt)) +
geom_point() +
stat_smooth(method = "lm")Marginal distributions answer two questions at once: how are the variables related, and how is each variable distributed?
p <- ggplot(data = d_spellname,
mapping = aes(x = dur,
y = rt)) +
geom_point() +
stat_smooth(method = "lm")
ggExtra::ggMarginal(p)Density lines answer: where are observations concentrated? They reveal clusters, but individual points disappear.
ggplot(data = d_spellname,
mapping = aes(x = dur,
y = rt)) +
geom_density_2d()Filled density regions answer the same question with colour. They reveal broad structure, but exact values become less visible.
ggplot(data = d_spellname,
mapping = aes(x = dur,
y = rt)) +
geom_density_2d_filled()A heatmap answers: which combinations occur most often? It reveals counts across two dimensions, but not individual observations.
ggplot(data = data_char,
mapping = aes(x = nchar,
y = nphon,
fill = n)) +
geom_tile(colour = "grey90")We will return to interactivity in the Shiny part of the module.
library(plotly)
plot <- ggplot(data = d_spellname,
mapping = aes(x = dur,
y = rt)) +
geom_point()
ggplotly(plot)Each plot changes the question we can answer. This is the Week 2 point again: plots reveal structure that summaries can leave out (Anscombe, 1973; Matejka & Fitzmaurice, 2017).
| Plot choice | Strongest question | Main trade-off |
|---|---|---|
| Scatterplot | Do two variables vary together? | Can overplot dense data |
| Trend line | What is the overall pattern? | Hides individual variation |
| Marginal distributions | What is the relationship and each variable’s spread? | More visually complex |
| Density or heatmap | Where are observations concentrated? | Individual observations disappear |
Choose the plot that makes the important comparison easiest to see.
Return to 01_choosing_plot_types.Rmd.
Complete now:
If you finish early, complete the Principle check.
Look at the plot first, before thinking about code. Bar plots with error bars can make summaries feel more complete than they are (Correll & Gleicher, 2014).
Visualisation suggests \(\dots\)
Now watch what happens to the y-axis when we use dots.
Boxplots were introduced as part of exploratory data analysis (Tukey, 1977). They compress a distribution into a small set of robust summaries.
See Tukey (1977)
Read the parts as:
A boxplot is compact, but it hides individual observations unless we add raw points.
library(ggdist)
ggplot(d_spellname,
aes(x = modality, y = rt)) +
# half violin (the "cloud")
stat_halfeye(
adjust = 0.45, # smoothness
justification = -0.2, # horiz. shift
.width = 0, # no interv bars
point_colour = NA) +
# boxplot
geom_boxplot(
width = 0.15,
outlier.shape = NA,
alpha = 0.5) +
# jittered points (the "rain")
geom_jitter(
width = 0.1,
alpha = 0.5,
size = 1) +
coord_flip() # horizontal orientation
Open 02_group_comparisons_and_raw_data.Rmd.
Complete the file and use the ranking task near the end to decide which plot is most honest for the question.
This is homework/reference material.
You’ve already seen these:
# Count number of identical observations count(data, group) # Calculate descriptive statistics summarise(data, mean = mean(rt)) # Remove rows with missing data (i.e. NA) drop_na()
These wrangling tasks can be managed in ggplot and tidyverse, more specifically dplyr.
dplyr and tidyr have useful functions for preparing data for plots.# Transforms dataframes into a long format pivot_longer(data, cols) # Transforms dataframes into a wide format pivot_wider(data, names_from, values_from) # Selects and removes variables select(data, var1, var2) # Retains and removes observations filter(data, condition) # Creates new variables mutate(data, new_var = old_var)
Data wrangling is homework/reference: 03_data_wrangling_for_visualisation_homework.Rmd. Focus on how reshaping data makes it match what the geom needs.
For support with this week’s material, use:
Andrews, M. (2021). Doing data science in R: An introduction for Social Scientists. SAGE Publications Ltd.
Anscombe, F. J. (1973). Graphs in statistical analysis. The American Statistician, 27, 17–21.
Cleveland, W. S., & McGill, R. (1984). Graphical perception: Theory, experimentation, and application to the development of graphical methods. Journal of the American Statistical Association, 79(387), 531–554.
Correll, M., & Gleicher, M. (2014). Error bars considered harmful: Exploring alternate encodings for mean and error. IEEE Transactions on Visualization and Computer Graphics, 20(12), 2142–2151. https://doi.org/10.1109/TVCG.2014.2346298
Hartwig, F., & Dearing, B. E. (1979). Exploratory data analysis. Sage.
Matejka, J., & Fitzmaurice, G. (2017). Same stats, different graphs: Generating datasets with varied appearance and identical statistics through simulated annealing. Proceedings of the 2017 CHI Conference on Human Factors in Computing Systems, 1290–1294.
Tufte, E. R. (1983). The visual display of information. Graphics Press.
Tukey, J. W. (1977). Exploratory data analysis (Vol. 2).
Wickham, H., & Grolemund, G. (2016). R for data science: Import, tidy, transform, visualize, and model data. O’Reilly Media, Inc.