From Grammar To Plot Choice

  • Week 2: Principles of data visualisation
  • Week 3: Grammar of graphics; aesthetics and attributes
  • Week 4: Major visualisation tools
  • Week 5: Customising visualisations (scales, themes, and labels)

What You Should Take Away

By the end of the session, you should be able to:

  • choose a geometry based on the question and variable types
  • compare options for distributions, relationships, and group comparisons
  • recognise when a summary plot hides important variation
  • use simple data wrangling when it makes a clearer plot possible

Download the Week 4 files from NOW and move them into your RStudio Project folder.

Before thinking about code, look at the plot first.

  • What is this plot trying to compare?
  • Which group seems faster or slower?
  • What makes the comparison easy?
  • What makes the comparison difficult?
  • What would you try instead?

Start With Question And Variable Type

Start with the question and the variables involved.

Question Typical variable types Useful starting point
What values occur? one continuous variable histogram, density plot
Are two measures related? two continuous variables scatterplot
Do groups differ? categorical group + continuous outcome jitter/boxplot/violin
How many observations are in each group? one categorical variable bar plot

Plot type follows the question and variable types, not the other way round.

Variable Type Is A Plot Choice Shortcut

Before choosing a plot, identify whether each variable is continuous or categorical.

Variable combination What it supports Examples
one continuous variable distribution of values histogram, density plot
one categorical variable counts or proportions bar chart
continuous + continuous relationship scatterplot, density contours
categorical + continuous group comparison jitter, boxplot, violin, raincloud
categorical + categorical counts across groups grouped bar chart, heatmap

The geom should match the structure of the variables you want to compare.

The Main Question For Every Plot

For each visualisation today, ask three things:

  • What question does this plot answer?
  • What does this plot reveal clearly?
  • What does this plot hide or make difficult to compare?

A plot type is a design decision: it makes some comparisons easier and others harder.

Week 2 Principles, Week 4 Decisions

Last week, we used principles as a checklist for judging plots (Cleveland & McGill, 1984; Hartwig & Dearing, 1979; Tufte, 1983). This week we use the same checklist to choose plot types.

Week 2 principle Week 4 plot-choice question
Show the data and variation Do we need raw points, distributions, or uncertainty?
Make comparisons easy Which plot makes the target comparison quickest to see?
Avoid misleading summaries What does this plot hide, average away, or imply?
Label enough for interpretation Are variables, groups, and units clear?

Choosing a plot type is how we put the Week 2 principles into practice.

Choose Encodings That Support Comparison

Some visual encodings make comparisons easier than others (Cleveland & McGill, 1984).

  • Continuous quantities: position on a common scale is usually easiest to compare.
  • Groups/categories: use distinct positions, hues, shapes, linetypes, or facets.
  • Use size carefully: it can imply magnitude or importance.
  • Avoid unnecessary mappings: extra aesthetics should make the comparison clearer.

What Plot Types Do You Know?

What types of data visualisations have you seen in journal articles, news, reports, dashboards, or social media?

You do not need to know the correct name. Describe or sketch the idea: what does it look like, and what does it seem to show?

A Snapshot Of The Visualisation Landscape

There are many named chart types. We will not read or memorise this list. Use it as a quick reminder that visualisation tools vary with the question, the variables, and the data structure.

Amounts and proportions

  • bar chart
  • column chart
  • stacked bar chart
  • pie chart
  • donut chart
  • waffle chart

Distributions

  • histogram
  • density plot
  • frequency polygon
  • box plot
  • violin plot
  • raincloud plot

Relationships

  • scatterplot
  • bubble chart
  • line chart
  • contour plot
  • heatmap
  • pair plot

Change over time

  • time-series line chart
  • area chart
  • streamgraph
  • timeline
  • calendar heatmap
  • Gantt chart

Places, networks, and structures

  • choropleth map
  • dot map
  • flow map
  • treemap
  • sunburst chart
  • dendrogram
  • network graph
  • Sankey diagram
  • parallel coordinates
  • radar chart
  • word cloud
  • ternary plot

The name matters less than the comparison: What are we trying to see?

A Decision Tree For Plot Choice data-to-viz.com/

Choosing Visualisation Tools

Geoms Are Drawing Tools

  • choice depends on visualisation goals, number and type of variables (and your subject domain)
  • what kind of visual element you want to draw?
  • a geometry (or geom) is the type of plot layer that tells R how to represent the data visually.

Think of it like this:

  • you have data (e.g., heights of children, test scores, etc.).
  • you want to show it visually.
  • the geom decides how it appears: as points, lines, bars, boxes, etc.
  • geoms are your drawing toolkit

ggplot2 Has Many Geoms

  • Geometries (geom_) control visual encoding of aesthetics layer
  • ~50 geometries: geom_ are part of ggplot2 (below)
  • more geoms in other packages such as ggdist, ggbeeswarm, and ggridges
  • many can be combined
 [1] abline            area              bar               bin_2d           
 [5] bin2d             blank             boxplot           col              
 [9] column            contour           contour_filled    count            
[13] crossbar          curve             density           density_2d       
[17] density_2d_filled density2d         density2d_filled  dotplot          
[21] errorbar          errorbarh         freqpoly          function         
[25] hex               histogram         hline             jitter           
[29] label             line              linerange         map              
[33] path              point             pointrange        polygon          
[37] qq                qq_line           quantile          raster           
[41] rect              ribbon            rug               segment          
[45] sf                sf_label          sf_text           smooth           
[49] spoke             step              text              tile             
[53] violin            vline            

Common Geoms And What They Show

Use geometries as tools for comparison, not as a menu of effects.

Geometry Function What it shows
geom_point() Scatter plot Relationship between two variables
geom_line() Line plot Trends over time or ordered categories
geom_bar() Bar chart Counts or values for categories
geom_histogram() Histogram Distribution of a single variable
geom_boxplot() Boxplot Summary of distribution (median, quartiles, outliers)
geom_text() Text labels Add labels to points or bars

Same Data, Different Views

Three Plotting Problems We Will Practise

  • Univariate distributions
  • Bivariate distributions
  • Group comparisons

Show One Variable: Histograms

A histogram answers: what values occur, and how often? It reveals distribution shape, but the answer depends on binning.

ggplot(data = d_spellname,
       mapping = aes(x = rt)) +
  geom_histogram(bins = 30)

Exercise 1: Choosing Plot Types

Open 01_choosing_plot_types.Rmd.

Complete now:

  • Load packages and data
  • Question 1: distribution of age
  • Choose between histogram and scatterplot for the age question
  • Question 2: distribution of reaction times

Stop before Question 3.

Show Two Variables: Scatterplots

A scatterplot answers: do two continuous variables vary together? It reveals individual observations, but dense regions can overplot.

ggplot(data = d_spellname,
       mapping = aes(x = dur,
                     y = rt)) +
  geom_point() +
  stat_smooth(method = "lm")

Ordered Data: Time Trends

Scatterplots With A Trend Line

A trend line answers: what is the overall pattern? It reveals direction, but it can hide local variation and individual cases.

ggplot(data = d_spellname,
       mapping = aes(x = dur,
                     y = rt)) +
  geom_point() +
  stat_smooth(method = "lm")

Add Marginal Distributions

Marginal distributions answer two questions at once: how are the variables related, and how is each variable distributed?

p <- ggplot(data = d_spellname,
       mapping = aes(x = dur,
                     y = rt)) +
  geom_point() +
  stat_smooth(method = "lm")

ggExtra::ggMarginal(p)

Two-Dimensional Density Lines

Density lines answer: where are observations concentrated? They reveal clusters, but individual points disappear.

ggplot(data = d_spellname,
       mapping = aes(x = dur,
                     y = rt)) +
  geom_density_2d()

Filled Density Regions

Filled density regions answer the same question with colour. They reveal broad structure, but exact values become less visible.

ggplot(data = d_spellname,
       mapping = aes(x = dur,
                     y = rt)) +
  geom_density_2d_filled()

Heatmaps Show Counts Across Two Variables

A heatmap answers: which combinations occur most often? It reveals counts across two dimensions, but not individual observations.

ggplot(data = data_char,
       mapping = aes(x = nchar,
                     y = nphon,
                     fill = n)) +
  geom_tile(colour = "grey90")

Optional Preview: Interactive Scatterplots

We will return to interactivity in the Shiny part of the module.

library(plotly)

plot <- ggplot(data = d_spellname,
       mapping = aes(x = dur,
                     y = rt)) +
  geom_point()

ggplotly(plot)

Same Data, Different Questions

Each plot changes the question we can answer. This is the Week 2 point again: plots reveal structure that summaries can leave out (Anscombe, 1973; Matejka & Fitzmaurice, 2017).

Plot choice Strongest question Main trade-off
Scatterplot Do two variables vary together? Can overplot dense data
Trend line What is the overall pattern? Hides individual variation
Marginal distributions What is the relationship and each variable’s spread? More visually complex
Density or heatmap Where are observations concentrated? Individual observations disappear

Choose the plot that makes the important comparison easiest to see.

Continue Exercise 1: Relationships

Return to 01_choosing_plot_types.Rmd.

Complete now:

  • Question 3: relationship between two reaction-time measures
  • Choose between scatterplot and separate histograms for the relationship question
  • Question 4: whether the relationship differs by group

If you finish early, complete the Principle check.

Compare Groups Without Hiding The Data

  • Function: distribution of values for two or more groups (often closely tied to statistical descriptions)
  • Variable type: continuous
  • Examples: (jitter) dots, box plot, violin plot, beeswarm plots, barplot (pie chart), dynamite plots

What Does This Plot Hide?

Look at the plot first, before thinking about code. Bar plots with error bars can make summaries feel more complete than they are (Correll & Gleicher, 2014).

  • What does this visualisation suggest?
  • What information is missing?
  • Can we see spread, outliers, or sample size?
  • What would you add or change?

Why Dynamite Plots Can Mislead

Visualisation suggests \(\dots\)

  • normal distribution
  • same number of observations in each group
  • presence of data where there are none?
  • absence of data above the errorbars

Now watch what happens to the y-axis when we use dots.

Summary Bars Hide Raw Variation

Raw Points Reveal Variation

Jitter Reduces Overplotting

Combine Raw Data With Summary Estimates

Boxplots Summarise Distributions

Boxplots Work Better With Raw Points

Reading A Boxplot

Boxplots were introduced as part of exploratory data analysis (Tukey, 1977). They compress a distribution into a small set of robust summaries.

See @tukey1977exploratory

See Tukey (1977)

Read the parts as:

  • middle line: median
  • box: middle 50% of the data
  • lower edge: first quartile, Q1
  • upper edge: third quartile, Q3
  • box height: interquartile range, IQR
  • whiskers: values within 1.5 IQR from the box
  • points beyond whiskers: possible outliers

A boxplot is compact, but it hides individual observations unless we add raw points.

Raincloud Plots: Distribution And Raw Data

library(ggdist)

ggplot(d_spellname, 
       aes(x = modality, y = rt)) +
  # half violin (the "cloud")
  stat_halfeye(
    adjust = 0.45,        # smoothness
    justification = -0.2, # horiz. shift 
    .width = 0,           # no interv bars
    point_colour = NA) +
  # boxplot
  geom_boxplot(
    width = 0.15,
    outlier.shape = NA,
    alpha = 0.5) +
  # jittered points (the "rain")
  geom_jitter(
    width = 0.1,
    alpha = 0.5,
    size = 1) +
  coord_flip() # horizontal orientation 

Exercise 2: Group Comparisons

Open 02_group_comparisons_and_raw_data.Rmd.

Complete the file and use the ranking task near the end to decide which plot is most honest for the question.

Homework/Reference: Data Wrangling For Visualisation

This is homework/reference material.

  • Data come in various formats.
  • For some visualisations, the format of the data needs to be changed.

You’ve already seen these:

# Count number of identical observations 
count(data, group)

# Calculate descriptive statistics
summarise(data, mean = mean(rt))

# Remove rows with missing data (i.e. NA) 
drop_na()

These wrangling tasks can be managed in ggplot and tidyverse, more specifically dplyr.

Wrangling Tools To Recognise

  • dplyr and tidyr have useful functions for preparing data for plots.
  • You do not need to master data wrangling in this module, but these tools are useful to recognise.
# Transforms dataframes into a long format
pivot_longer(data, cols)

# Transforms dataframes into a wide format
pivot_wider(data, names_from, values_from)

# Selects and removes variables
select(data, var1, var2)

# Retains and removes observations
filter(data, condition)

# Creates new variables
mutate(data, new_var = old_var)

Data wrangling is homework/reference: 03_data_wrangling_for_visualisation_homework.Rmd. Focus on how reshaping data makes it match what the geom needs.

Homework

  • Data wrangling is homework/reference: 03_data_wrangling_for_visualisation_homework.Rmd. Focus on how reshaping data makes it match what the geom needs.
  • Bring in data and anything you’ve already done for the formative assessment.
  • Complete recommended reading

Recommended Reading

References

Andrews, M. (2021). Doing data science in R: An introduction for Social Scientists. SAGE Publications Ltd.

Anscombe, F. J. (1973). Graphs in statistical analysis. The American Statistician, 27, 17–21.

Cleveland, W. S., & McGill, R. (1984). Graphical perception: Theory, experimentation, and application to the development of graphical methods. Journal of the American Statistical Association, 79(387), 531–554.

Correll, M., & Gleicher, M. (2014). Error bars considered harmful: Exploring alternate encodings for mean and error. IEEE Transactions on Visualization and Computer Graphics, 20(12), 2142–2151. https://doi.org/10.1109/TVCG.2014.2346298

Hartwig, F., & Dearing, B. E. (1979). Exploratory data analysis. Sage.

Matejka, J., & Fitzmaurice, G. (2017). Same stats, different graphs: Generating datasets with varied appearance and identical statistics through simulated annealing. Proceedings of the 2017 CHI Conference on Human Factors in Computing Systems, 1290–1294.

Tufte, E. R. (1983). The visual display of information. Graphics Press.

Tukey, J. W. (1977). Exploratory data analysis (Vol. 2).

Wickham, H., & Grolemund, G. (2016). R for data science: Import, tidy, transform, visualize, and model data. O’Reilly Media, Inc.