Today

  • Revisit the independent-samples t-test from Week 3: two different groups of people, one numeric outcome.
  • Calculate the standard error from group means, SDs, and sample sizes.
  • Reconstruct the t-value, p-value, and confidence interval.
  • Compare the manual calculations with t.test() output.

Check the online lecture before the workshop. Complete the Week 4 problem set after this workshop on Friday.

Link To The Lecture

Workshop Frame

The goal is not to memorise formulas. The goal is to understand what R is reporting.

  1. Run the t-test and save the output.
  2. Calculate descriptive statistics by group.
  3. Use the descriptives to calculate the standard error.
  4. Use the standard error to calculate t, p, and CI.
  5. Compare with the values stored in the t-test object.

Core Message

For the equal-variance independent-samples t-test, the main output can be reconstructed from the descriptive statistics:

Start with Then calculate
group means mean difference
group SDs pooled variance / standard error
group sample sizes df and standard error

Once we have mean, SD, and n for each group, we can reconstruct the standard error, t-value, p-value, and confidence interval.

Start In RStudio

Work from inside your RStudio Project folder.

  1. Download week_04_workshop.R from NOW.
  2. Move week_04_workshop.R into your RStudio Project folder.
  3. Check that the workshop_data folder is also inside your RStudio Project folder.
  4. Open your .Rproj file for this module.
  5. Open week_04_workshop.R from inside RStudio.

Run one line at a time and check the console after each line.

Demonstration Data

Part A of week_04_workshop.R follows the slides. We will complete it together while working through the Blomkvist example.

blomkvist <- read_csv("workshop_data/analysis_blomkvist.csv")
blomkvist <- drop_na(blomkvist, rt, meds_group)

Is mean hand reaction time different for medicated and not medicated participants?

The Blomkvist data are adapted from a reaction-time and aging study using the Nintendo Wii Balance Board (Blomkvist et al., 2017).

Group Descriptives

  • We can calculate the hypothesis test from the descriptives.
  • For this, we need the difference in the means and its standard error.
  • For the standard error, we need the group SDs and sample sizes.
meds_summary <- summarise(blomkvist,
  mean_rt = mean(rt),
  sd_rt = sd(rt),
  n = n(),
  .by = meds_group
)

Group Descriptives: Output

meds_summary
# A tibble: 2 × 4
  meds_group    mean_rt sd_rt     n
  <chr>           <dbl> <dbl> <int>
1 medicated        739.  219.    70
2 not medicated    581.  155.   102

Exercise Checkpoint 1

Open week_04_workshop.R and complete Part A1 with me.

Stop after you have created and printed meds_summary.

At this point, we have the group means, SDs, and sample sizes.

Extract Values From The Summary

First extract the whole mean_rt column:

meds_summary$mean_rt
# 738.8143 581.3268

Then add [1] to take the first value only:

meds_summary$mean_rt[1]
# 738.8143

Now save that value as mean_x:

mean_x <- meds_summary$mean_rt[1]

Extract Values From The Summary

For the second mean, use [2]:

mean_y <- meds_summary$mean_rt[2]

Use the same pattern for the two SDs and the two sample sizes.

Exercise Checkpoint 2

Continue Part A2 in week_04_workshop.R with me.

Stop after you have extracted mean_x, mean_y, sd_x, sd_y, n_x, and n_y from meds_summary.

Now we have the ingredients needed for the manual reconstruction.

Standard Error: Idea

The standard error is the uncertainty in the mean difference.

For this t-test, it combines two things:

  1. how much the scores vary inside the groups: SDs
  2. how much data we have in each group: n

The point is not to memorise this formula. The point is to see how to break a formula into smaller R steps and check each result.

\[ SE = \sqrt{\frac{(n_x - 1)s_x^2 + (n_y - 1)s_y^2}{n_x + n_y - 2}}\sqrt{\frac{1}{n_x} + \frac{1}{n_y}} \]

Pooled Numerator

Math:

\[ (n_x - 1)s_x^2 + (n_y - 1)s_y^2 \]

Break this into the x-group part and y-group part first.

R:

pooled_x <- (n_x - 1) * sd_x^2
pooled_x
# 3312858

pooled_y <- (n_y - 1) * sd_y^2
pooled_y
# 2422699

pooled_numerator <- pooled_x + pooled_y
pooled_numerator
# 5735557

Pooled Denominator And Variance

Math:

\[ \frac{\text{pooled numerator}}{n_x + n_y - 2} \]

R:

pooled_denominator <- n_x + n_y - 2
pooled_denominator
# 170

pooled_variance <- pooled_numerator / pooled_denominator
pooled_variance
# 33738.57

Pooled SD

Math:

\[ \sqrt{\text{pooled variance}} \]

R:

pooled_sd <- sqrt(pooled_variance)
pooled_sd
# 183.6806

Sample-Size Part

Math:

\[ \sqrt{\frac{1}{n_x} + \frac{1}{n_y}} \]

R:

sample_size_part <- sqrt(1 / n_x + 1 / n_y)
sample_size_part
# 0.1552084

Standard Error

Math:

\[ SE = \text{pooled SD} \times \text{sample-size part} \]

R:

se <- pooled_sd * sample_size_part
se
# 28.50877

The standard error is the bridge from descriptives to the rest of the t-test output.

t-Value

A t-value is the observed mean difference divided by its standard error.

mean_difference <- mean_x - mean_y
mean_difference
# 157.4875

t_value <- mean_difference / se
t_value
# 5.524177

The t-value says how large the observed difference is relative to the uncertainty in that difference.

p-Value

Once we have the t-value and df, the p-value comes from the t-distribution.

df <- n_x + n_y - 2
df
# 170

p_value <- 2 * pt(abs(t_value), df = df, lower.tail = FALSE)
p_value
# 0.0000001222335

This is the same pt() logic as Week 3, now using the t-value reconstructed from the group means, SDs, and sample sizes.

Confidence Interval: Idea

A confidence interval starts with the observed mean difference and adds uncertainty around it.

For a 95% confidence interval, we use the middle 95% of the relevant t-distribution to decide how far the interval should extend.

The wider the uncertainty, the wider the confidence interval.

Confidence Interval: Calculation

Use the mean difference, standard error, and a t-distribution cutoff called tau.

tau <- qt(.975, df = df)
tau
# 1.974017

ci_lower <- mean_difference - tau * se
ci_lower
# 101.2107

ci_upper <- mean_difference + tau * se
ci_upper
# 213.7643

Why .975 For A 95% CI?

qt() works with cumulative probability: the area to the left of a t-value.

A 95% confidence interval keeps the middle 95% of the t-distribution. That leaves 5% outside the interval:

  • 2.5% in the lower tail
  • 2.5% in the upper tail

For the upper edge of the interval, we want the point with 97.5% to the left and 2.5% to the right.

qt(.975, df = df)

The .975 is not the confidence level. It is the upper cutoff for the middle 95%.

Exercise Checkpoint 3

Continue Part A3 in week_04_workshop.R with me.

Complete the standard error, t-value, p-value, and 95% CI from scratch.

This is the part where we build the t-test output from means, SDs, and sample sizes.

Run The Actual t-Test

Now run t.test() and compare the output to the values we calculated manually.

m_meds <- t.test(rt ~ meds_group, data = blomkvist, var.equal = TRUE)
Code What it extracts
m_meds$estimate group means
m_meds$stderr standard error
m_meds$statistic t-value
m_meds$parameter degrees of freedom
m_meds$p.value p-value
m_meds$conf.int confidence interval

Exercise Checkpoint 4

Complete Part A4 in week_04_workshop.R with me.

Run the actual t-test, extract the stored values, and compare them with the manual calculations.

Independent Practice

Now work on Part B independently: calculate descriptives, run the Chinese lexical-decision accuracy-group t-test, and extract the key values from the output.

Part C asks you to write a few short lines yourself, using the same pt(), qt(), and $ extraction patterns.

Optional Stretch Script

Only open week_04_stretch.R if the core workshop script felt comfortable. There is no expectation that you complete the stretch script, and it is not required preparation for the exam.

The stretch script includes the extra 99% CI practice and an optional mu = ... challenge that connects the non-zero null hypothesis from Week 3 with the Week 4 reconstruction logic.

Reading / Problem Set

Recommended reading: Andrews (2021), Chapter 8, Statistical Models and Statistical Inference.

Problem-set reminder: complete the Week 4 problem set after this workshop on Friday. Prioritise the logic that the main t.test() outputs can be reconstructed from group means, SDs, and sample sizes.

Core R Functions

Function What it does
summarise() calculates group descriptives
$ and [ ] extract values from objects and tables
sqrt() calculates a square root
^2 squares a value
pt() gives areas under a t-distribution
qt() gives t-distribution cutoff values
t.test() runs an independent-samples t-test

Leave With

Before the problem set, you should be able to:

  • explain that m$stderr is the standard error of the mean difference.
  • calculate t = mean difference / standard error.
  • use pt() to reconstruct a two-sided p-value.
  • use qt() to reconstruct a confidence interval.
  • compare manual calculations with values stored in a t.test() object.

References

Andrews, M. (2021). Doing data science in R: An Introduction for Social Scientists. SAGE Publications Ltd.

Blomkvist, A. W., Eika, F., Rahbek, M. T., Eikhof, K. D., Hansen, M. D., Søndergaard, M., Ryg, J., Andersen, S., & Jørgensen, M. G. (2017). Reference data on reaction time and aging using the Nintendo Wii Balance Board: A cross-sectional study of 354 subjects from 20 to 99 years of age. PLoS One, 12(12), e0189598.