Today

  • Calculate a mean and SD by hand for a small example.
  • Set up tidyverse, the exercise script, and the workshop data.
  • Inspect a data frame using View(), glimpse(), and summary().
  • Calculate descriptive statistics for reaction time.
  • Practise the same workflow on a second data set.

Check the online lecture before the workshop. Weekly problem sets begin in Week 3 and are completed after the relevant workshop.

Link To The Lecture

Workshop Frame

For each analysis, use the same basic workflow:

  1. Check what one row represents.
  2. Identify numeric variables and grouping variables.
  3. Remove missing rows for the variables used in the summary.
  4. Calculate and interpret the summary.

Running an analysis is useful; understanding what the output means is the point. If you don’t understand the output, please ask!

Start In RStudio

Work from inside the RStudio Project folder you created last week.

  1. Open the .Rproj file for this module.
  2. Download the Week 2 exercise script from NOW: week_02_workshop.R.
  3. Move week_02_workshop.R into your RStudio Project folder.
  4. Open week_02_workshop.R from inside RStudio.

Do this before Part A. For learning R, run one line at a time and inspect the console after each line.

Packages And Tidyverse

Some R functions are available as soon as R starts. Other functions come from packages.

Think of packages like apps on a phone: they add extra tools to base R.

For this workshop, we use tidyverse because it gives us the functions we need to load, inspect, clean, and summarise data.

You should have installed tidyverse last week. Only run this line if you have not installed it yet:

install.packages("tidyverse")

Load tidyverse at the start of each script:

library(tidyverse)

Installation only needs to happen once. library(tidyverse) is needed each time you start a new R session.

Mean And SD In Words

Before we write formulas or code, we need the idea.

  • The mean is a typical value: where the scores are centred.
  • The SD describes spread: how far scores usually are from the mean.
  • A small SD means scores are close to the mean.
  • A large SD means scores are more spread out.

The maths and R code are just precise ways to express this idea.

Later workshops reuse this same logic: build a calculation from named quantities, then compare it with what R reports.

Manual Mean

Functions such as mean() do the arithmetic for us, but it is useful to see the pieces once.

Formula

For scores 6, 8, and 10:

\[ \bar{x} = \frac{6 + 8 + 10}{3} = 8 \]

Total divided by number of scores.

R code

scores <- c(6, 8, 10)

sum(scores)
# 24

length(scores)
# 3

manual_mean <- sum(scores) / length(scores)
manual_mean
# 8

mean(scores)
# 8

The formula and the R code do the same calculation.

Manual SD: Distances From The Mean

SD starts by measuring how far each score is from the mean.

Formula

For scores 6, 8, and 10:

\[ \bar{x} = 8 \] \[ x_i - \bar{x} = -2, 0, 2 \] \[ (x_i - \bar{x})^2 = 4, 0, 4 \]

We square the distances so negative and positive distances do not cancel out.

R code

scores <- c(6, 8, 10)

mean_score <- mean(scores)
mean_score
# 8

deviations <- scores - mean_score
deviations
# -2  0  2

squared_deviations <- deviations^2
squared_deviations
# 4  0  4

Manual SD: Final Calculation

The sample SD divides by n - 1, then takes the square root.

Formula

\[ s = \sqrt{\frac{\sum(x_i - \bar{x})^2}{n - 1}} \] \[ s = \sqrt{\frac{4 + 0 + 4}{3 - 1}} \] \[ s = \sqrt{4} = 2 \]

R code

sum_of_squares <- sum(squared_deviations)
sum_of_squares
# 8

variance <- sum_of_squares / (length(scores) - 1)
variance
# 4

manual_sd <- sqrt(variance)
manual_sd
# 2

sd(scores)
# 2

Now open week_02_workshop.R and complete Part A. Run one line at a time and check how the numbers change in the console.

Before Part B: Data Folder

Part B uses the Blomkvist data. Before starting Part B in week_02_workshop.R:

  1. Download workshop_data.zip from the Data sets folder on NOW.
  2. Move workshop_data.zip into your RStudio Project folder.
  3. Unzip it there. Your RStudio Project folder should now contain a folder called workshop_data.

The script uses paths such as workshop_data/analysis_blomkvist.csv, so the unzipped workshop_data folder needs to be directly inside the RStudio Project folder.

Load The Data

blomkvist <- read_csv("workshop_data/analysis_blomkvist.csv")

If read_csv() gives an error, first check that:

  • you opened the .Rproj file from last week.
  • you unzipped workshop_data.zip.
  • the folder is called workshop_data.
  • the workshop_data folder is directly inside the RStudio Project folder.
  • the file name inside that folder is exactly analysis_blomkvist.csv.

Other file formats exist, but today we practise CSV files. RStudio’s Import Dataset menu can help discover import code for unfamiliar files.

Follow along in Part B of week_02_workshop.R while I show the same steps on the slides.

Data Frames And Demonstration Data

A data frame is the usual shape for data analysis in R: rows are observations, columns are variables, and a cell is one value for one observation on one variable.

blomkvist <- read_csv("workshop_data/analysis_blomkvist.csv")

Each row is one participant. Key columns are:

  • id: participant ID
  • age: age in years
  • medicine: medication score
  • smoker: smoking category
  • meds_group: medication group
  • age_group: age category
  • rt: hand reaction time
  • rt_foot: foot reaction time

For today, the main question is: does average rt differ across age_group?

Variable Types

glimpse() shows variable types next to each column name.

  • <dbl>: numeric values, such as age or rt.
  • <chr>: text values, such as labels or IDs.
  • <lgl>: logical values, TRUE or FALSE.
  • categorical variables often arrive as <chr> and define groups.

For summaries today, rt needs to be numeric and age_group needs to define groups.

Inspect The Data

Before summarising, check what the data frame contains.

blomkvist
View(blomkvist)
glimpse(blomkvist)
summary(blomkvist$rt)
  • running the data frame name prints a preview in the console.
  • View() opens a spreadsheet-like view in RStudio.
  • glimpse() shows variable names, types, and preview values.
  • summary() gives a quick numeric summary of one variable.

What does one row represent, which variables are numeric, and which variables could define groups?

Counting Categories

count(blomkvist, age_group)
count(blomkvist, age_group, meds_group)

count() is descriptive: it checks how many observations are in each category before we summarise or test them.

Remove Missing Values

drop_na() keeps rows that have data for the variables named.

blomkvist_rt <- drop_na(blomkvist, rt, age_group, meds_group)

Overall Summary

This gives one row for the whole data frame.

summarise(blomkvist_rt,
  mean_rt = mean(rt),
  sd_rt = sd(rt),
  n = n()
)

More Descriptive Statistics

Mean and SD are the main focus today. Minimum and maximum help us see the range of values.

summarise(blomkvist_rt,
  mean_rt = mean(rt),
  sd_rt = sd(rt),
  min_rt = min(rt),
  max_rt = max(rt),
  n = n()
)

What do the minimum and maximum reaction times tell us about the range of values?

Summary By One Factor

Add .by = age_group to calculate the same summary separately for each age group.

summarise(blomkvist_rt,
  mean_rt = mean(rt),
  sd_rt = sd(rt),
  min_rt = min(rt),
  max_rt = max(rt),
  n = n(),
  .by = age_group
)

Which age group has the slower average reaction time?

Summary By Two Factors

Use .by = c(age_group, meds_group) when the summary should be split by two grouping variables.

summarise(blomkvist_rt,
  mean_rt = mean(rt),
  sd_rt = sd(rt),
  n = n(),
  .by = c(age_group, meds_group)
)

Do the same age-group pattern and medication-group pattern appear together?

Extract Values From A Table

Descriptive tables are still R objects. To extract values, first save a small table.

summary_age <- summarise(blomkvist_rt,
  mean_rt = mean(rt),
  sd_rt = sd(rt),
  n = n(),
  .by = age_group
)

summary_age[1, "mean_rt"]
summary_age$mean_rt
summary_age$mean_rt[1]
  • [] extracts by row and column: [row, column].
  • $ extracts one named column as a vector.
  • Combining $ and [1] extracts the first value from that column.

We use the same extraction idea later for model objects, where the output is harder to read.

Exercise Script

Complete the remaining exercises in week_02_workshop.R.

Part C is independent practice with blanks. It uses analysis_chinese_ldt.csv to answer concrete questions about reaction time, list_group, and accuracy_group.

Optional: Part D is independent write-your-own-code practice. Use the same logic, but decide and write the full code block yourself.

Optional Stretch Script

Only open week_02_stretch.R if the core workshop script felt comfortable. There is no expectation that you complete the stretch script, and it is not required preparation for the exam.

Reading / Follow-Up

Core R Functions

Function What it does
read_csv() reads a CSV file into R
View(), glimpse() shows variables, types, and preview values
count() counts rows in categories
drop_na() removes rows with missing values in named variables
summarise() calculates descriptive statistics
n() counts rows inside summarise()
mean() calculates an average
sd() calculates the sample standard deviation
[] extracts values by position, row, or column
$ extracts one named column from a data frame

References

Andrews, M. (2021). Doing data science in R: An Introduction for Social Scientists. SAGE Publications Ltd.