Today

  • Connect the lecture idea of comparing two population means to R code.
  • Calculate grouped descriptives with summarise().
  • Run an independent-samples t-test with t.test().
  • Extract t, df, p, and the confidence interval from the t-test object.
  • Use pt() to understand how a t-value and df become evidence about statistical significance.
  • Use mu = ... when the null hypothesis is not a zero difference.

Check the online lecture before the workshop. Complete the Week 3 problem set after this workshop.

Link To The Lecture

Workshop Frame

For each independent two-group analysis, use the same workflow:

  1. Identify the numeric outcome and the grouping variable.
  2. Calculate mean, SD, and n for each group.
  3. Run the t-test.
  4. Extract and interpret the result.

Running the code is only the first step. The important part is knowing what the output says about the research question.

Start In RStudio

Work from inside the RStudio Project folder you created in Week 1.

  1. Download the Week 3 exercise script from NOW: week_03_workshop.R.
  2. Move week_03_workshop.R into your RStudio Project folder.
  3. Check that the workshop_data folder is also inside your RStudio Project folder.
  4. Open your .Rproj file for this module.
  5. Open week_03_workshop.R from inside RStudio.

For learning R, run one line at a time and inspect the console after each line.

Demonstration Data

Part A uses the Blomkvist data from the workshop_data folder.

blomkvist <- read_csv("workshop_data/analysis_blomkvist.csv")
blomkvist <- drop_na(blomkvist, rt, age_group)

Each row is one participant. Today we ask:

Is mean hand reaction time different for participants aged 60 and over versus participants under 60?

The Blomkvist data are adapted from a reaction-time and aging study using the Nintendo Wii Balance Board (Blomkvist et al., 2017).

Before The Test

Calculate the descriptives that make the t-test easier to interpret. The n = n() line also checks how many participants are in each group:

summarise(blomkvist,
  mean_rt = mean(rt),
  sd_rt = sd(rt),
  n = n(),
  .by = age_group
)

Which age group has the slower mean reaction time?

Formula Syntax

For an independent-groups t-test in R, the formula has this structure:

outcome ~ group

For the Blomkvist example:

rt ~ age_group
  • rt is the numeric outcome.
  • age_group defines the two independent groups.
  • The outcome goes on the left of ~; the group goes on the right.

The t-Test

m_age <- t.test(rt ~ age_group, data = blomkvist, var.equal = TRUE)
m_age

var.equal = TRUE matches the equal-variance t-test introduced in the lecture.

The printed output is useful, but saving the test as an object lets us extract the exact values we need.

Extract Results

In R, the dollar sign operator $ extracts a named value from an object.

Read m_age$statistic as: from the object m_age, extract the named value statistic.

Code Value
m_age$statistic t-value
m_age$parameter degrees of freedom
m_age$p.value p-value
m_age$conf.int confidence interval for the mean difference

Exercise Script

Open week_03_workshop.R.

Part A repeats the Blomkvist t-test from the slides. Complete it together with me, one line at a time.

Part B asks the same kind of question for the Chinese lexical-decision data: is mean reaction time different for high accuracy and lower accuracy participants?

Standard t-Distribution

The t-distribution is the reference curve for the t-value when the null hypothesis is true.

  • The total area under the curve is 1, or 100%.
  • The distribution is centred at 0 and symmetric.
  • Degrees of freedom control the shape: smaller df means heavier tails.
  • As df gets larger, the t-distribution becomes more like a standard Normal distribution.

A t-value tells us where we are on this curve.

Why This Curve Matters

To decide whether a result is statistically significant, we ask:

If the null hypothesis were true, how unlikely would it be to get a t-value this far from 0?

For the usual independent-samples t-test, the null hypothesis is that the population mean difference is 0: mu = 0.

If that null hypothesis is true, the expected t-value is also 0. Large positive or negative t-values are more surprising under the null hypothesis.

Degrees of freedom depend on sample size. Smaller samples give smaller df, and smaller df makes the tails of the t-distribution heavier.

The t-distribution tells us how much probability is in the tails for the df of our test.

Introducing pt()

pt() turns a t-value and degrees of freedom into an area under the t-distribution.

The key inputs are:

pt(t_value, df = degrees_of_freedom)

In the next examples, we use df = 10 so the shape of the curve stays the same.

What pt() Gives Us

By default, pt() gives the area to the left of a t-value.

pt(1.5, df = 10)
# 0.918

The answer is the area below 1.5 in a t-distribution with 10 degrees of freedom.

Right-Tail Areas

By default, pt() gives the lower tail: the area to the left of the t-value.

Sometimes we need the area to the right of a t-value.

Right-Tail Areas In R

To get the area to the right, set lower.tail = FALSE.

pt(1.5, df = 10, lower.tail = FALSE)
# 0.082

This gives the same right-tail area as subtracting the left-tail area from 1:

1 - pt(1.5, df = 10)
# 0.082

Use lower.tail = FALSE in code because it asks R for the right-tail probability directly and is more numerically precise for very small probabilities.

Areas Between Two t-Values

To get the area between two values, subtract two left-tail areas.

pt(1.5, df = 10) - pt(-1.5, df = 10)
# 0.835

Think: area below 1.5 minus area below -1.5.

Two-Sided p-Values With pt()

For a two-sided test, “as or more extreme” means both tails.

pt(-1.5, df = 10) + pt(1.5, df = 10, lower.tail = FALSE)
# 0.165

2 * pt(1.5, df = 10, lower.tail = FALSE)
# 0.165

Because the t-distribution is symmetric, the shortcut doubles one tail.

p-Values From A t-Test

The same area-under-the-curve idea explains where the p-value in a t-test comes from.

t_value <- abs(m_age$statistic)
df_value <- m_age$parameter

2 * pt(t_value, df = df_value, lower.tail = FALSE)

m_age$p.value

This recreates the two-sided p-value stored in the t-test object.

Practice With pt()

Open Part C of week_03_workshop.R.

Use pt() to work out areas below a t-value, above a t-value, between two t-values, and in both tails. These are the same area-under-the-curve ideas used for p-values.

The three moves are: below with pt(), above with lower.tail = FALSE, and between two values with subtraction.

Non-Null Hypotheses

The usual t-test tests whether the population mean difference is 0.

Use mu = ... to test a different hypothesised difference:

t.test(rt ~ age_group, data = blomkvist, mu = 200, var.equal = TRUE)

This is the same model, but the comparison value is no longer zero.

More Practice

Complete Parts C and D of week_03_workshop.R. Part C connects t-values and df to statistical significance; Part D shows how to test a null hypothesis that is not a zero difference.

Only open week_03_stretch.R if the core workshop script felt comfortable. There is no expectation that you complete the stretch script, and it is not required preparation for the exam.

Reading / Problem Set

Recommended reading: Andrews (2021), Chapter 8, Statistical Models and Statistical Inference.

Problem-set reminder: complete the Week 3 problem set after this workshop on Friday. Focus on explaining what each result means, not only on running the code.

Core R Functions

Function What it does
read_csv() reads a prepared data file
drop_na() removes incomplete rows for named variables
summarise() calculates grouped descriptives
mean() / sd() / n() calculate mean, SD, and sample size
t.test() compares two group means
$ dollar sign operator: extracts named values from an R object
abs() removes the sign from a number
pt() gets probabilities from a t-distribution

Leave With

Before the problem set, you should be able to:

  • create grouped descriptives with summarise(..., .by = group).
  • run t.test(y ~ group, data = d, var.equal = TRUE).
  • extract t, df, p, and CI using the dollar sign operator $.
  • explain how pt() links a t-value and df to a p-value.
  • use mu = ... when the null hypothesis is about a non-zero difference.

References

Andrews, M. (2021). Doing data science in R: An Introduction for Social Scientists. SAGE Publications Ltd.

Blomkvist, A. W., Eika, F., Rahbek, M. T., Eikhof, K. D., Hansen, M. D., Søndergaard, M., Ryg, J., Andersen, S., & Jørgensen, M. G. (2017). Reference data on reaction time and aging using the Nintendo Wii Balance Board: A cross-sectional study of 354 subjects from 20 to 99 years of age. PLoS One, 12(12), e0189598.