From last week to this week

Last week:

  • module overview and assessment
  • RMarkdown, code chunks, citations, and rendering to HTML
  • organising files in an RStudio Project

This week:

  • why visualisation matters for behavioural data
  • first ggplot2 plots
  • principles for clear and honest visualisation

Get the Week 2 files

  1. Download psyc40940-week-02-foundations.zip from the Week 2 section in the NOW Learning Room.
  2. Unzip the file on your computer.
  3. Move the contents into your module RStudio Project folder.
  4. Open 1_scatterplots.Rmd from inside RStudio.

Keep the .Rmd files and the data folder together. The code expects them to stay in the same project folder.

Today’s session

Today we will use simple plots to connect practical ggplot2 work with the principles of clear and responsible visualisation.

We will:

  • start with the Week 2 files and a short ggplot2 walkthrough
  • complete an exploratory scatterplot exercise
  • use Anscombe’s quartet and the Datasaurus Dozen to show why plots matter
  • discuss media examples where design choices shape interpretation
  • finish with a checklist of visualisation principles and common pitfalls

By the end of the session, take away that plots help us reason about evidence. They can reveal patterns, question summaries, and communicate findings, but design choices also shape what people believe.

What is data visualisation?

  • Using graphics to understand data
  • Using graphics to communicate evidence
  • Exploratory: what patterns are in the data?
  • Explanatory: what should the audience take away?

Common plot types:

  • scatterplots, line plots, bar plots
  • histograms, boxplots, density plots
  • pie charts, maps, heatmaps

The data

  • Picture naming task
  • Written and spoken responses
  • Outcome: response onset time
  • Data/materials: Roeser et al. (2025)

Exploring data

# A tibble: 72 × 4
   ppt_id ppt_vocab modality    rt
    <dbl>     <dbl> <chr>    <dbl>
 1     40     0.95  speech   1358.
 2     41     1     speech   1274.
 3     42     0.925 speech   1393.
 4     43     0.9   speech    735.
 5     44     0.975 speech   1188.
 6     45     1     speech   1345.
 7     46     0.925 speech   1520.
 8     47     0.962 speech   1513.
 9     48     0.975 speech   1479.
10     50     0.825 speech   1212.
# ℹ 62 more rows

Discuss in pairs (2 mins)

What is the first thing you notice in each plot? What did you learn from it?

Building up a plot: data and aesthetics

ggplot(data = d_vocab,
       aes(x = ppt_vocab, y = rt))

Building up a plot: add points

ggplot(data = d_vocab,
       aes(x = ppt_vocab, y = rt)) +
  geom_point()

Building up a plot: add a trend

ggplot(data = d_vocab,
       aes(x = ppt_vocab, y = rt)) +
  geom_point() +
  stat_smooth(method = "lm")

Building up a plot: add a group

ggplot(data = d_vocab,
       aes(x = ppt_vocab, y = rt,
           colour = modality)) +
  geom_point() +
  stat_smooth(method = "lm")

A more polished version

We can add labels, colour choices, and a cleaner theme later.

What to recognise today

  • data = ...: which dataset should R use?
  • aes(...): which variables go on the plot?
  • geom_point(): draw points
  • stat_smooth(): add a trend line
  • labs(): write clearer labels
ggplot(data = d_vocab,
       aes(x = ppt_vocab, y = rt)) +
  geom_point() +
  stat_smooth(method = "lm") +
  labs(x = "Vocabulary score",
       y = "Average reaction time (ms)")

Next week we will slow down and unpack the layer logic properly.

Creating an exploratory plot

15 minutes: work on 1_scatterplots.Rmd

Aim: complete the main blanks, render the document, and write short interpretations.

Start in class and finish at home if needed.

Why data visualisation?

“[data visualisation] forces us to notice what we never expected to see.” (Tukey, 1977)

Behavioural data are usually noisy and full of individual differences. A good visualisation helps us ask better questions before decide for a summary or model.

  • See what summaries hide: distributions, clusters, outliers, and non-linear patterns
  • Check the story: does the plot support the claim made by the mean, correlation, or model?
  • Choose better analyses: visual patterns can suggest transformations, grouping variables, or model problems
  • Communicate responsibly: design choices can make effects look larger, smaller, clearer, or more uncertain

Plots are not decoration. They are part of how we reason about evidence.

Anscombe’s quartet (Anscombe, 1973)

x
y
y ~ x
Data set Mean SD Mean SD Correlation Intercept Slope
1 9 3.32 7.5 2.03 0.82 3 0.5
2 9 3.32 7.5 2.03 0.82 3 0.5
3 9 3.32 7.5 2.03 0.82 3 0.5
4 9 3.32 7.5 2.03 0.82 3 0.5

Anscombe’s quartet

Anscombe’s quartet

The datasaurus dozen

Matejka & Fitzmaurice (2017): see link

Visualisation is not neutral

Anscombe’s quartet and the Datasaurus Dozen make the same point:

  • the same summary statistics can describe very different data
  • a mean, standard deviation, correlation, or model line can hide the pattern
  • plotting the data can reveal outliers, clusters, curves, and unusual cases
  • the visualisation changes what we notice and what we think needs explaining

Plots are not decoration after the analysis. They are part of how we understand the data.

How plots shape beliefs

Next, we will look at examples from media and public communication.

For each example, ask:

  • what does the plot make easy to believe?
  • what design choice creates that impression?
  • what information is missing, hidden, or emphasised?
  • how could the plot be changed to support a fairer comparison?

Design choices can persuade or mislead, even when the numbers are real (Pandey et al., 2014).

Pair activity: media examples

Pair activity: main points

  • Truncated y-axis: small numerical changes can look dramatic when the baseline is hidden.
  • Wrong chart type: a pie chart implies parts of one whole, so it misleads when categories overlap or add to more than 100%.
  • Election maps: large areas can look more important than densely populated areas.
  • Inverted axis: reversing the axis can make increases look like decreases.

Visual design choices, data choices, and wording can change the story people see.

Same numbers, different beliefs

A bar can make the mean look like a container for the data.

Viewers can treat values inside the bar as more likely than equally distant values outside it (Correll & Gleicher, 2014).

Better alternative: show the distribution, not only the summary.

Encoding choices matter

Cleveland and McGill showed that some visual comparisons are easier than others (Cleveland & McGill, 1984).

Easier:

  • position on a common scale
  • aligned lengths

Harder:

  • angles
  • areas
  • colour intensity

Principles for this module

A useful plot should help the viewer reason about the evidence, not just decorate the analysis.

Drawing on exploratory data analysis and visualisation principles (Cleveland & McGill, 1984; Hartwig & Dearing, 1979; Tufte, 1983), use these checks:

  • Show the data and variation: make distributions, outliers, clusters, and uncertainty visible where they matter.
  • Make comparisons easy: arrange the plot around the question the viewer needs to answer.
  • Use scales honestly: avoid truncated, inverted, unequal, or area-based encodings that distort effects.
  • Label enough for interpretation: make variables, units, groups, and the intended comparison clear.
  • Remove competing design: decoration, heavy gridlines, and excessive colour should not compete with the evidence.
  • Stay skeptical and open: every plot is a choice; use visualisation for discovery, not only confirmation.

Same data, different design

What changes in your interpretation?

Make comparisons easy

Good plots make the intended comparison visually easy.

Use labels to reduce effort

Good labels answer:

  • What is measured?
  • What are the units?
  • What do colours or groups mean?
  • What should the viewer compare?

Common visualisation pitfalls

Use the principles as a checklist when judging plots.

Pitfall Why it matters Principle
Truncated or inverted axes changes apparent size or direction of effects avoid distortion
Overplotting hides density and individual observations show the data
Poor chart type makes the wrong comparison easy make comparisons easy
Too much decoration competes with the evidence remove competing design
Missing uncertainty makes estimates look more precise show the data and variation
Weak labels or legends increases interpretation effort label enough for interpretation

What’s wrong with these? (1)

If time: 2 minutes in pairs which pitfall is most important in each example.

Sources: A: Ke (2024); B: Rubiah et al. (2024).

What’s wrong with these? (2)

If time: 2 minutes in pairs which pitfall is most important in each example.

Sources: A: CBSN; B: Gong & Liu (2022).

What’s wrong with these? (3)

If time: 2 minutes in pairs which pitfall is most important in each example.

What’s wrong with these? (4)

If time: 1 minutes in pairs which pitfall is most important in this example.

Recommended Reading

Homework

Before next week:

  • complete the Week 2 visualisation exercises if you did not finish them in class
  • continue looking for a behavioural dataset for the formative assessment
  • bring dataset questions to the tutorial checkpoint on 21 October

Suggestion for formative assessment: rework (some) figures reported in van Lieburg et al. (2023): article, code and data.

References

Andrews, M. (2021). Doing data science in R: An introduction for Social Scientists. SAGE Publications Ltd.

Anscombe, F. J. (1973). Graphs in statistical analysis. The American Statistician, 27, 17–21.

Cleveland, W. S., & McGill, R. (1984). Graphical perception: Theory, experimentation, and application to the development of graphical methods. Journal of the American Statistical Association, 79(387), 531–554.

Correll, M., & Gleicher, M. (2014). Error bars considered harmful: Exploring alternate encodings for mean and error. IEEE Transactions on Visualization and Computer Graphics, 20(12), 2142–2151. https://doi.org/10.1109/TVCG.2014.2346298

Gong, R., & Liu, B. (2022). [Retracted] monitoring of sports health indicators based on wearable nanobiosensors. Advances in Materials Science and Engineering, 2022(1), 3802603. https://doi.org/10.1155/2022/3802603

Hartwig, F., & Dearing, B. E. (1979). Exploratory data analysis. Sage.

Ke, Y. (2024). Examining simultaneous pausing on the cognitive writing process: A micro-formative writing assessment. Current Psychology, 43(1), 39–50. https://doi.org/10.1007/s12144-023-04429-z

Matejka, J., & Fitzmaurice, G. (2017). Same stats, different graphs: Generating datasets with varied appearance and identical statistics through simulated annealing. Proceedings of the 2017 CHI Conference on Human Factors in Computing Systems, 1290–1294.

Pandey, A. V., Manivannan, A., Nov, O., Satterthwaite, M., & Bertini, E. (2014). The persuasive power of data visualization. IEEE Transactions on Visualization and Computer Graphics, 20(12), 2211–2220. https://doi.org/10.1109/TVCG.2014.2346419

Roeser, J., Aros Muñoz, P., & Torrance, M. (2025). Written picture naming norms to assess spelling difficulty. OSF. https://doi.org/10.17605/OSF.IO/JVHRZ

Rubiah, R., Degeng, I. N. S., Setyosari, P., & Kuswandi, D. (2024). The effect of problem-based learning assisted with concept mapping founded on cognitive style on the creativity of writing exposition text. Creativity Studies, 17(2), 419–434. https://doi.org/10.3846/cs.2024.16302

Tufte, E. R. (1983). The visual display of quantitative information. Graphics Press.

Tufte, E. R. (2001). The visual display of quantitative information (2nd ed.). Graphics Press.

Tukey, J. W. (1977). Exploratory data analysis (Vol. 2).

van Lieburg, R., Sijyeniyo, E., Hartsuiker, R. J., & Bernolet, S. (2023). The development of abstract syntactic representations in beginning L2 learners of Dutch. Journal of Cultural Cognitive Science, 7, 289–309. https://doi.org/10.1007/s41809-023-00131-5

Wickham, H., & Grolemund, G. (2016). R for data science: Import, tidy, transform, visualize, and model data. O’Reilly Media, Inc.