1 How to use this report

This file explains the full research process and places the relevant R code directly below each step. Group members can use it in two ways:

  1. Click Knit in RStudio to render the complete report as HTML.
  2. Run the code chunks one at a time, from top to bottom, to see how every result is produced.

The chunks must be run in order because later steps use objects created earlier. Put this .Rmd file and the course CSV file in the same folder before starting.

The analysis follows the same structure as the current group presentation:

  • study context and design;
  • data limitations and coding decisions;
  • RQ1: oral-contraceptive use and dysmenorrhea at Wave 8;
  • RQ2: types of activities missed across Waves 6–8;
  • RQ3: number of activity types missed across Waves 6–8;
  • conclusions and limits of interpretation.

2 Step 1: Prepare R

2.1 What this step does

The first chunk installs any missing packages, loads them and sets a consistent graph style. The code uses only common packages that are suitable for this course.

2.2 Why it is needed

  • dplyr organises and summarises the data.
  • tidyr changes data between wide and long formats.
  • ggplot2 produces the three graphs.
  • scales formats proportions as percentages.
  • knitr displays result tables in the rendered report.

2.3 Run this code

required_packages <- c("dplyr", "tidyr", "ggplot2", "scales", "knitr")

missing_packages <- required_packages[
  !required_packages %in% rownames(installed.packages())
]

if (length(missing_packages) > 0) {
  install.packages(
    missing_packages,
    repos = "https://cloud.r-project.org"
  )
}

invisible(lapply(required_packages, library, character.only = TRUE))

knitr::opts_chunk$set(
  echo = TRUE,
  message = FALSE,
  warning = FALSE,
  fig.align = "center",
  out.width = "90%"
)

theme_set(
  theme_minimal(base_size = 12) +
    theme(
      plot.title = element_text(face = "bold"),
      legend.position = "bottom",
      panel.grid.minor = element_blank()
    )
)

2.4 What you should see

The chunk normally produces no printed output. If a package is missing, R installs it first. A red installation message does not necessarily mean an error. Check whether the final line in the Console begins with Error.

3 Step 2: Find and import the CSV file

3.1 What this step does

The code searches the current folder for a CSV filename containing both “menstrual” and “pain”. It then imports the file as raw_data.

3.2 Why this version is useful

The code does not depend on an absolute Windows path such as C:/Users/.... This makes the file easier to share with group members. It also works with filenames such as:

  • MenstrualPain.csv;
  • MenstrualPain (2).csv;
  • 01-MenstrualPain-2-.csv.

3.3 Run this code

csv_files <- list.files(
  path = ".",
  pattern = "(menstrual.*pain|pain.*menstrual).*\\.csv$",
  full.names = TRUE,
  ignore.case = TRUE
)

if (length(csv_files) == 0) {
  stop(
    paste(
      "No menstrual-pain CSV file was found.",
      "Put the CSV and this Rmd file in the same folder,",
      "then run the chunk again."
    )
  )
}

data_file <- csv_files[1]

raw_data <- read.csv(
  data_file,
  stringsAsFactors = FALSE,
  check.names = FALSE,
  fileEncoding = "UTF-8-BOM"
)

cat("Data file used:", basename(data_file), "\n")
## Data file used: MenstrualPain.csv
cat("Rows:", nrow(raw_data), "\n")
## Rows: 24
cat("Columns:", ncol(raw_data), "\n")
## Columns: 10

3.4 What you should see

The supplied file should have 24 rows and 10 columns. If R says that no file was found, check the Files pane in RStudio and confirm that the .Rmd and .csv files are in the same folder.

4 Step 3: Understand the supplied file

4.1 What this step does

The next chunk displays the first rows and the original column names.

4.2 Why this matters

The supplied CSV is unusual. It contains three separate tables placed side by side:

  1. pain severity by wave and oral-contraceptive use;
  2. activity type missed by wave;
  3. number of activity types missed by wave.

The repeated headings Wave and Count do not refer to one single rectangular dataset. We therefore separate the file by column position before analysing it.

4.3 Run this code

names(raw_data)
##  [1] "Wave"                            "Oral contraception"             
##  [3] "Pain severity"                   "Count"                          
##  [5] "Wave"                            "Activity type missed"           
##  [7] "Count"                           "Wave"                           
##  [9] "Number of activity types missed" "Count"
head(raw_data, 8)
##   Wave Oral contraception      Pain severity Count Wave Activity type missed
## 1    6                yes       very painful     1    6    school/university
## 2    6                yes      quite painful     1    6                 work
## 3    6                yes   a little painful     1    6    social activities
## 4    6                yes not at all painful     1    6   sports or exercise
## 5    6                 no       very painful    78    6    total respondents
## 6    6                 no      quite painful   147    7    school/university
## 7    6                 no   a little painful   220    7                 work
## 8    6                 no not at all painful   195    7    social activities
##   Count Wave Number of activity types missed Count
## 1   115    6                            none   424
## 2     0    6                             one   122
## 3   123    6                             two    61
## 4   148    6                           three    40
## 5   647    6                            four     0
## 6   293    7                            none   887
## 7    67    7                             one   293
## 8   168    7                             two   107

4.4 What you should notice

  • Columns 1–4 form the pain table.
  • Columns 5–7 form the activity table.
  • Columns 8–10 form the number-of-activities table.
  • The shorter tables contain empty cells near the bottom of the CSV.

5 Step 4: Separate and clean the three tables

5.1 What this step does

This code separates the side-by-side tables, gives them clear variable names, removes empty rows and converts counts and wave numbers to numeric values.

5.2 Why it is needed

Each research question uses a different part of the CSV. Keeping three clean datasets reduces the risk of using the wrong denominator or treating blank cells as observations.

5.3 Run this code

# Table 1: pain severity and oral-contraceptive use
pain_data <- raw_data[, 1:4]
names(pain_data) <- c(
  "Wave", "Oral_contraception", "Pain_severity", "Count"
)

pain_data <- pain_data |>
  filter(!is.na(Wave), !is.na(Pain_severity)) |>
  mutate(
    Wave = as.integer(Wave),
    Count = as.numeric(Count),
    Oral_contraception = tolower(trimws(Oral_contraception)),
    Pain_severity = tolower(trimws(Pain_severity))
  )

# Table 2: activity types missed
activity_data <- raw_data[, 5:7]
names(activity_data) <- c("Wave", "Activity_type", "Count")

activity_data <- activity_data |>
  filter(!is.na(Wave), !is.na(Activity_type)) |>
  mutate(
    Wave = as.integer(Wave),
    Count = as.numeric(Count),
    Activity_type = tolower(trimws(Activity_type))
  )

# Table 3: number of activity types missed
number_data <- raw_data[, 8:10]
names(number_data) <- c("Wave", "Number_missed", "Count")

number_data <- number_data |>
  filter(!is.na(Wave), !is.na(Number_missed)) |>
  mutate(
    Wave = as.integer(Wave),
    Count = as.numeric(Count),
    Number_missed = tolower(trimws(Number_missed))
  )

cat("Pain-table rows:", nrow(pain_data), "\n")
## Pain-table rows: 24
cat("Activity-table rows:", nrow(activity_data), "\n")
## Activity-table rows: 15
cat("Number-table rows:", nrow(number_data), "\n")
## Number-table rows: 15

5.4 What you should see

  • pain_data: 24 rows;
  • activity_data: 15 rows;
  • number_data: 15 rows.

6 Step 5: Check that the cleaning worked

6.1 What this step does

The code checks the expected number of rows and confirms that all three waves are present. It also displays each cleaned table.

6.2 Why it is needed

Data checks should happen before the statistical analysis. If the file structure changes or the wrong CSV is selected, it is safer for the code to stop than to produce misleading results.

6.3 Run this code

stopifnot(
  nrow(pain_data) == 24,
  nrow(activity_data) == 15,
  nrow(number_data) == 15,
  all(pain_data$Wave %in% 6:8),
  all(activity_data$Wave %in% 6:8),
  all(number_data$Wave %in% 6:8)
)

knitr::kable(pain_data, caption = "Cleaned pain table")
Cleaned pain table
Wave Oral_contraception Pain_severity Count
6 yes very painful 1
6 yes quite painful 1
6 yes a little painful 1
6 yes not at all painful 1
6 no very painful 78
6 no quite painful 147
6 no a little painful 220
6 no not at all painful 195
7 yes very painful 58
7 yes quite painful 47
7 yes a little painful 65
7 yes not at all painful 21
7 no very painful 249
7 no quite painful 322
7 no a little painful 403
7 no not at all painful 177
8 yes very painful 63
8 yes quite painful 107
8 yes a little painful 149
8 yes not at all painful 55
8 no very painful 140
8 no quite painful 208
8 no a little painful 290
8 no not at all painful 103
knitr::kable(activity_data, caption = "Cleaned activity table")
Cleaned activity table
Wave Activity_type Count
6 school/university 115
6 work 0
6 social activities 123
6 sports or exercise 148
6 total respondents 647
7 school/university 293
7 work 67
7 social activities 168
7 sports or exercise 151
7 total respondents 1341
8 school/university 154
8 work 101
8 social activities 173
8 sports or exercise 170
8 total respondents 1111
knitr::kable(number_data, caption = "Cleaned number-of-activities table")
Cleaned number-of-activities table
Wave Number_missed Count
6 none 424
6 one 122
6 two 61
6 three 40
6 four 0
7 none 887
7 one 293
7 two 107
7 three 42
7 four 11
8 none 767
8 one 163
8 two 125
8 three 40
8 four 16

6.4 What you should notice

Wave 6 corresponds to approximate mean age 14, Wave 7 to age 16 and Wave 8 to age 18. In the activity table, the Wave 6 work count is recorded as zero even though work was not measured at age 14. We correct this later by treating it as unavailable.

7 Step 6: Explain the study and the research questions

7.1 Original study

Cameron et al. (2024) analysed data from the Longitudinal Study of Australian Children. The original design was a prospective, population-based cohort study. Participants reported menstrual pain and activities missed because of their periods at approximate mean ages 14, 16 and 18.

The longitudinal design allows population-level patterns to be compared across adolescence. The national cohort is also more informative than a small convenience sample. However, self-report, attrition, non-response and possible exposure misclassification limit the interpretation. The study is observational, so an association cannot establish causation.

7.2 Limitation of our supplied file

Our CSV contains aggregated category counts rather than one row for each participant. We can calculate group percentages and fit a simple unadjusted model using the counts. We cannot:

  • identify the same participant across waves;
  • adjust for socio-economic status, previous pain or other individual confounders;
  • reproduce the article’s weighted analysis;
  • interpret wave-level patterns as individual change.

7.3 Research questions

RQ1. At Wave 8, was oral-contraceptive use associated with the probability of dysmenorrhea?

RQ2. How did the reported types of activities missed because of periods differ across Waves 6, 7 and 8?

RQ3. How did the distribution of the number of activity types missed differ across Waves 6, 7 and 8?

8 Step 7: Recode pain for RQ1

8.1 What this step does

The original pain variable has four categories. We create a binary outcome for the Wave 8 comparison:

  • Dysmenorrhea: very painful or quite painful;
  • No dysmenorrhea: a little painful or not at all painful.

8.2 Why this coding is used

The binary outcome matches the definition used for this group analysis and allows a binary logistic regression. The disadvantage is that it removes information about variation within the two combined groups.

8.3 Run this code

pain_binary <- pain_data |>
  mutate(
    Dysmenorrhea = if_else(
      Pain_severity %in% c("very painful", "quite painful"),
      "Dysmenorrhea",
      "No dysmenorrhea"
    )
  )

knitr::kable(
  pain_binary,
  caption = "Pain categories after binary dysmenorrhea coding"
)
Pain categories after binary dysmenorrhea coding
Wave Oral_contraception Pain_severity Count Dysmenorrhea
6 yes very painful 1 Dysmenorrhea
6 yes quite painful 1 Dysmenorrhea
6 yes a little painful 1 No dysmenorrhea
6 yes not at all painful 1 No dysmenorrhea
6 no very painful 78 Dysmenorrhea
6 no quite painful 147 Dysmenorrhea
6 no a little painful 220 No dysmenorrhea
6 no not at all painful 195 No dysmenorrhea
7 yes very painful 58 Dysmenorrhea
7 yes quite painful 47 Dysmenorrhea
7 yes a little painful 65 No dysmenorrhea
7 yes not at all painful 21 No dysmenorrhea
7 no very painful 249 Dysmenorrhea
7 no quite painful 322 Dysmenorrhea
7 no a little painful 403 No dysmenorrhea
7 no not at all painful 177 No dysmenorrhea
8 yes very painful 63 Dysmenorrhea
8 yes quite painful 107 Dysmenorrhea
8 yes a little painful 149 No dysmenorrhea
8 yes not at all painful 55 No dysmenorrhea
8 no very painful 140 Dysmenorrhea
8 no quite painful 208 Dysmenorrhea
8 no a little painful 290 No dysmenorrhea
8 no not at all painful 103 No dysmenorrhea

9 Step 8: Calculate the Wave 8 group counts

9.1 What this step does

We keep Wave 8, add the counts for the two dysmenorrhea categories and add the counts for the two no-dysmenorrhea categories. The result is a 2 × 2 table comparing pill users and non-users.

9.2 Why Wave 8 is used

The group research question concerns pill use and dysmenorrhea at age 18. Restricting the analysis to one wave also avoids incorrectly treating repeated observations across waves as independent.

9.3 Run this code

rq1_counts <- pain_binary |>
  filter(Wave == 8) |>
  group_by(Oral_contraception, Dysmenorrhea) |>
  summarise(Count = sum(Count), .groups = "drop") |>
  tidyr::pivot_wider(
    names_from = Dysmenorrhea,
    values_from = Count,
    values_fill = 0
  ) |>
  rename(
    dys = Dysmenorrhea,
    no_dys = `No dysmenorrhea`
  ) |>
  mutate(
    total = dys + no_dys,
    Pill = factor(
      Oral_contraception,
      levels = c("no", "yes")
    )
  ) |>
  arrange(Pill)

knitr::kable(
  rq1_counts |>
    select(Pill, dys, no_dys, total),
  col.names = c(
    "Pill use", "Dysmenorrhea", "No dysmenorrhea", "Total"
  ),
  caption = "Wave 8 counts for RQ1"
)
Wave 8 counts for RQ1
Pill use Dysmenorrhea No dysmenorrhea Total
no 348 393 741
yes 170 204 374

9.4 Expected counts

  • Non-users: 348 with dysmenorrhea and 393 without dysmenorrhea, total 741.
  • Pill users: 170 with dysmenorrhea and 204 without dysmenorrhea, total 374.

10 Step 9: Fit the logistic regression for RQ1

10.1 What this step does

The model compares the odds of dysmenorrhea among pill users with the odds among non-users. Non-users are the reference group because no is the first factor level.

The hypotheses are:

  • Null hypothesis: pill use is not associated with dysmenorrhea at Wave 8, so the odds ratio equals 1.
  • Alternative hypothesis: the odds ratio differs from 1.

10.2 Why logistic regression is appropriate here

The response is binary and the explanatory variable is binary. The four cell counts are sufficiently large, so sparse cells are not an obvious problem. The model is still unadjusted and observational.

10.3 Run this code

rq1_model <- glm(
  cbind(dys, no_dys) ~ Pill,
  family = binomial(link = "logit"),
  data = rq1_counts
)

summary(rq1_model)
## 
## Call:
## glm(formula = cbind(dys, no_dys) ~ Pill, family = binomial(link = "logit"), 
##     data = rq1_counts)
## 
## Coefficients:
##             Estimate Std. Error z value Pr(>|z|)  
## (Intercept) -0.12161    0.07361  -1.652   0.0985 .
## Pillyes     -0.06071    0.12729  -0.477   0.6334  
## ---
## Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
## 
## (Dispersion parameter for binomial family taken to be 1)
## 
##     Null deviance: 0.22765  on 1  degrees of freedom
## Residual deviance: 0.00000  on 0  degrees of freedom
## AIC: 17.425
## 
## Number of Fisher Scoring iterations: 2

10.4 How to read the output

The row called Pillyes is the pill-user comparison. Its estimate is on the log-odds scale, so it is converted to an odds ratio in the next step. The p-value tests the null hypothesis that this coefficient equals zero, which is equivalent to an odds ratio of 1.

11 Step 10: Calculate RQ1 probabilities, confidence intervals and odds ratio

11.1 What this step does

The first part obtains each group’s estimated probability and 95% confidence interval. The second part converts the pill coefficient to an odds ratio and calculates its 95% confidence interval.

11.2 Run this code

# Estimated group probabilities
new_groups <- data.frame(
  Pill = factor(c("no", "yes"), levels = c("no", "yes"))
)

prediction <- predict(
  rq1_model,
  newdata = new_groups,
  type = "link",
  se.fit = TRUE
)

z_value <- qnorm(0.975)

rq1_results <- rq1_counts |>
  mutate(
    probability = plogis(prediction$fit),
    lower = plogis(prediction$fit - z_value * prediction$se.fit),
    upper = plogis(prediction$fit + z_value * prediction$se.fit),
    Pill_label = factor(
      if_else(as.character(Pill) == "yes", "Yes", "No"),
      levels = c("No", "Yes")
    )
  )

# Odds ratio, confidence interval and p-value
coefficient_table <- summary(rq1_model)$coefficients

log_odds_ratio <- coefficient_table["Pillyes", "Estimate"]
standard_error <- coefficient_table["Pillyes", "Std. Error"]
p_value <- coefficient_table["Pillyes", "Pr(>|z|)"]

rq1_model_result <- data.frame(
  Odds_ratio = exp(log_odds_ratio),
  CI_lower = exp(log_odds_ratio - z_value * standard_error),
  CI_upper = exp(log_odds_ratio + z_value * standard_error),
  p_value = p_value
)

rq1_probability_table <- rq1_results |>
  transmute(
    `Pill use` = Pill_label,
    Dysmenorrhea = dys,
    Total = total,
    Probability = scales::percent(probability, accuracy = 0.1),
    `95% CI` = paste0(
      scales::percent(lower, accuracy = 0.1),
      " to ",
      scales::percent(upper, accuracy = 0.1)
    )
  )

rq1_test_table <- rq1_model_result |>
  transmute(
    `Odds ratio` = round(Odds_ratio, 2),
    `95% CI` = paste0(round(CI_lower, 2), " to ", round(CI_upper, 2)),
    `p-value` = format.pval(p_value, digits = 2, eps = 0.00001)
  )

knitr::kable(
  rq1_probability_table,
  caption = "Estimated probability of dysmenorrhea by pill use"
)
Estimated probability of dysmenorrhea by pill use
Pill use Dysmenorrhea Total Probability 95% CI
No 348 741 47.0% 43.4% to 50.6%
Yes 170 374 45.5% 40.5% to 50.5%
knitr::kable(
  rq1_test_table,
  caption = "Unadjusted logistic-regression result"
)
Unadjusted logistic-regression result
Odds ratio 95% CI p-value
0.94 0.73 to 1.21 0.63

11.3 Expected result and meaning

  • Non-users: 47.0%, with 95% CI 43.4%–50.6%.
  • Pill users: 45.5%, with 95% CI 40.5%–50.5%.
  • Odds ratio: 0.94, with 95% CI 0.73–1.21.
  • p-value: .63.

The confidence interval includes 1 and the p-value is large. The data provide little evidence of an association between reported pill use and dysmenorrhea at Wave 8.

This is not proof that the two groups are identical. It also does not show that the pill has no effect. Reasons for using the pill and other confounders have not been adjusted for.

12 Step 11: Produce the RQ1 graph

12.1 What this graph shows

Each point is the estimated probability of dysmenorrhea in one pill-use group. The vertical bars are 95% confidence intervals. Showing uncertainty prevents the small difference between the points from being overstated.

12.2 Run this code

p1 <- ggplot(
  rq1_results,
  aes(x = Pill_label, y = probability)
) +
  geom_errorbar(
    aes(ymin = lower, ymax = upper),
    width = 0.08,
    linewidth = 0.8,
    colour = "#0072B2"
  ) +
  geom_point(size = 3.7, colour = "#0072B2") +
  geom_text(
    aes(
      y = upper + 0.035,
      label = paste0(
        scales::percent(probability, accuracy = 0.1),
        "\n(n = ", total, ")"
      )
    ),
    size = 4
  ) +
  scale_y_continuous(
    labels = scales::percent_format(accuracy = 1),
    limits = c(0, 0.62),
    breaks = seq(0, 0.6, by = 0.1)
  ) +
  labs(
    title = "Dysmenorrhea at Wave 8 by Oral Contraceptive Use",
    x = "Oral contraceptive use",
    y = "Estimated probability of dysmenorrhea",
    caption = paste(
      "Points show estimated probabilities;",
      "bars show 95% confidence intervals."
    )
  )

p1

13 Step 12: Prepare RQ2 activity percentages

13.1 What this step does

The activity table includes a total respondents row for each wave. We separate those totals and join them back to the activity counts so that each percentage uses the correct denominator.

The Wave 6 work value is changed to NA because work was not measured at age 14. Leaving it as zero would incorrectly suggest that work was measured and nobody missed it.

13.2 Run this code

# Denominator for each wave
activity_totals <- activity_data |>
  filter(Activity_type == "total respondents") |>
  transmute(Wave, Total = Count)

# Percentages for each activity type
activity_results <- activity_data |>
  filter(Activity_type != "total respondents") |>
  left_join(activity_totals, by = "Wave") |>
  mutate(
    Age = recode(Wave, `6` = 14L, `7` = 16L, `8` = 18L),
    Activity_label = recode(
      Activity_type,
      "school/university" = "School or university",
      "work" = "Work",
      "social activities" = "Social activities",
      "sports or exercise" = "Sport or exercise"
    ),
    Count = if_else(
      Wave == 6 & Activity_type == "work",
      NA_real_,
      Count
    ),
    proportion = Count / Total,
    percentage = 100 * proportion,
    Activity_label = factor(
      Activity_label,
      levels = c(
        "School or university", "Work",
        "Social activities", "Sport or exercise"
      )
    )
  )

activity_display <- activity_results |>
  select(Activity_label, Age, percentage) |>
  tidyr::pivot_wider(names_from = Age, values_from = percentage) |>
  mutate(across(where(is.numeric), ~ round(.x, 1)))

knitr::kable(
  activity_display,
  col.names = c("Activity type", "Age 14", "Age 16", "Age 18"),
  caption = "Percentage missing each activity type"
)
Percentage missing each activity type
Activity type Age 14 Age 16 Age 18
School or university 17.8 21.8 13.9
Work NA 5.0 9.1
Social activities 19.0 12.5 15.6
Sport or exercise 22.9 11.3 15.3

13.3 How to read the table

The percentages for different activities should not be added. A participant could report missing more than one activity, so the categories overlap.

14 Step 13: Calculate the percentage missing at least one activity

14.1 What this step does

The third supplied table includes the number who missed no activities. We subtract this count from the total number of respondents to calculate how many missed at least one activity.

14.2 Run this code

none_counts <- number_data |>
  filter(Number_missed == "none") |>
  transmute(Wave, None = Count)

at_least_one <- activity_totals |>
  left_join(none_counts, by = "Wave") |>
  mutate(
    Age = recode(Wave, `6` = 14L, `7` = 16L, `8` = 18L),
    At_least_one = Total - None,
    proportion = At_least_one / Total,
    percentage = 100 * proportion
  )

knitr::kable(
  at_least_one |>
    transmute(
      Age,
      Count = paste0(At_least_one, "/", Total),
      Percentage = round(percentage, 1)
    ),
  caption = "Participants missing at least one activity"
)
Participants missing at least one activity
Age Count Percentage
14 223/647 34.5
16 454/1341 33.9
18 344/1111 31.0

14.3 Expected result and meaning

  • Age 14: 223/647, or 34.5%.
  • Age 16: 454/1,341, or 33.9%.
  • Age 18: 344/1,111, or 31.0%.

About one-third reported missing at least one activity at each age. These are descriptive wave-level percentages, not individual trajectories.

15 Step 14: Produce the RQ2 graph

15.1 What this graph shows

The graph compares the percentage reporting each type of missed activity. Separate lines make it possible to see that the activity pattern differs across waves.

15.2 Run this code

activity_colours <- c(
  "School or university" = "#1B9E77",
  "Work" = "#D95F02",
  "Social activities" = "#7570B3",
  "Sport or exercise" = "#E7298A"
)

p2 <- ggplot(
  activity_results,
  aes(
    x = Age,
    y = proportion,
    colour = Activity_label,
    group = Activity_label
  )
) +
  geom_line(linewidth = 0.9, na.rm = TRUE) +
  geom_point(size = 3, na.rm = TRUE) +
  scale_colour_manual(values = activity_colours, drop = FALSE) +
  scale_x_continuous(breaks = c(14, 16, 18)) +
  scale_y_continuous(
    labels = scales::percent_format(accuracy = 1),
    limits = c(0, 0.25),
    breaks = seq(0, 0.25, by = 0.05)
  ) +
  labs(
    title = "Activities Missed Because of Periods",
    x = "Approximate mean age (years)",
    y = "Percentage of respondents",
    colour = "Activity type",
    caption = paste(
      "Participants could report more than one activity.",
      "Work was not measured at age 14."
    )
  )

p2

15.3 Interpretation

School or university absence was highest at age 16. Sport or exercise was the most commonly reported activity at age 14. Work absence increased from 5.0% at age 16 to 9.1% at age 18.

The graph describes the sample but does not test whether the wave differences are statistically significant. Different participants may have contributed at different waves.

16 Step 15: Prepare the RQ3 distribution

16.1 What this step does

For each wave, the code divides each category count by the number of participants with a classified number-of-activities response.

16.2 Why this denominator differs at Wave 7

The five Wave 7 categories sum to 1,340, while the activity table reports 1,341 total respondents. One Wave 7 record therefore does not have a complete number-of-types classification. Using 1,340 ensures that the displayed distribution sums to 100%.

16.3 Run this code

number_results <- number_data |>
  group_by(Wave) |>
  mutate(Available_total = sum(Count)) |>
  ungroup() |>
  mutate(
    Age = recode(Wave, `6` = 14L, `7` = 16L, `8` = 18L),
    Number = recode(
      Number_missed,
      "none" = "None",
      "one" = "One",
      "two" = "Two",
      "three" = "Three",
      "four" = "Four"
    ),
    Number = factor(
      Number,
      levels = c("None", "One", "Two", "Three", "Four")
    ),
    proportion = Count / Available_total,
    percentage = 100 * proportion,
    stack_order = match(
      as.character(Number),
      c("Four", "Three", "Two", "One", "None")
    )
  ) |>
  arrange(Age, stack_order) |>
  group_by(Age) |>
  mutate(label_y = cumsum(proportion) - proportion / 2) |>
  ungroup()

number_display <- number_results |>
  select(Number, Age, percentage) |>
  arrange(Number, Age) |>
  tidyr::pivot_wider(names_from = Age, values_from = percentage) |>
  mutate(across(where(is.numeric), ~ round(.x, 1)))

knitr::kable(
  number_display,
  col.names = c("Number missed", "Age 14", "Age 16", "Age 18"),
  caption = "Distribution of the number of activity types missed"
)
Distribution of the number of activity types missed
Number missed Age 14 Age 16 Age 18
None 65.5 66.2 69.0
One 18.9 21.9 14.7
Two 9.4 8.0 11.3
Three 6.2 3.1 3.6
Four 0.0 0.8 1.4

16.4 Expected result and meaning

The percentage reporting no missed activity types was 65.5% at age 14, 66.2% at age 16 and 69.0% at age 18. Among those affected, missing one type was the most common category.

Four activity types could not be reported at age 14 because work was not measured. The age-14 zero in the Four category is therefore structural rather than an observed absence of this outcome.

17 Step 16: Produce the RQ3 graph

17.1 What this graph shows

Each stacked bar shows the full distribution of the number of activity types missed at one wave. The bars sum to 100% using the available classified responses.

17.2 Run this code

number_colours <- c(
  "None" = "#D9D9D9",
  "One" = "#9ECAE1",
  "Two" = "#4292C6",
  "Three" = "#756BB1",
  "Four" = "#CB181D"
)

p3 <- ggplot(
  number_results,
  aes(x = factor(Age), y = proportion, fill = Number)
) +
  geom_col(width = 0.7) +
  geom_text(
    data = number_results |> filter(proportion >= 0.025),
    aes(
      y = label_y,
      label = scales::percent(proportion, accuracy = 0.1)
    ),
    size = 3.2
  ) +
  scale_fill_manual(values = number_colours, drop = FALSE) +
  scale_y_continuous(
    labels = scales::percent_format(accuracy = 1),
    limits = c(0, 1),
    breaks = seq(0, 1, by = 0.2),
    expand = expansion(mult = c(0, 0))
  ) +
  labs(
    title = "Number of Activity Types Missed Because of Periods",
    x = "Approximate mean age (years)",
    y = "Percentage of respondents",
    fill = "Number missed",
    caption = paste(
      "Wave 7 had one respondent without a classified",
      "number-of-types response."
    )
  )

p3

18 Step 17: Save the figures and result tables

18.1 What this step does

The report already displays the figures and tables. This final code also saves them as separate files that can be placed into the PowerPoint presentation.

18.2 Run this code

dir.create("figures", showWarnings = FALSE)
dir.create("results", showWarnings = FALSE)

ggsave(
  "figures/RQ1_oral_contraception_dysmenorrhea.png",
  plot = p1,
  width = 8,
  height = 5.6,
  dpi = 300
)

ggsave(
  "figures/RQ2_activity_types.png",
  plot = p2,
  width = 8.5,
  height = 5.3,
  dpi = 300
)

ggsave(
  "figures/RQ3_number_of_activities.png",
  plot = p3,
  width = 8.5,
  height = 5.3,
  dpi = 300
)

write.csv(
  rq1_results,
  "results/RQ1_group_results.csv",
  row.names = FALSE
)

write.csv(
  rq1_model_result,
  "results/RQ1_logistic_regression.csv",
  row.names = FALSE
)

write.csv(
  at_least_one,
  "results/RQ2_at_least_one_activity.csv",
  row.names = FALSE
)

write.csv(
  activity_results,
  "results/RQ2_activity_percentages.csv",
  row.names = FALSE
)

write.csv(
  number_results,
  "results/RQ3_number_distribution.csv",
  row.names = FALSE
)

cat("Figures saved in:", normalizePath("figures"), "\n")
## Figures saved in: C:\MenstrualPain\MenstrualPain\figures
cat("Result tables saved in:", normalizePath("results"), "\n")
## Result tables saved in: C:\MenstrualPain\MenstrualPain\results

19 Step 18: Bring the results together

19.1 RQ1 conclusion

At Wave 8, dysmenorrhea prevalence was 47.0% among non-users and 45.5% among pill users. The unadjusted odds ratio was 0.94 (95% CI 0.73–1.21, p = .63). The data provide little evidence of an association between reported pill use and dysmenorrhea.

19.2 RQ2 conclusion

About one-third reported missing at least one activity at each age. The types of affected activities differed across waves, but these comparisons are descriptive and do not establish individual change.

19.3 RQ3 conclusion

Most participants reported no missed activity types. Among those affected, missing one type was the most common result. The upper end of the age-14 distribution is not directly comparable because work was not measured.

19.4 Claims that should be avoided

The project should not claim that:

  • oral contraception caused, prevented or had no effect on dysmenorrhea;
  • individual participants improved as they became older;
  • missed activities were caused specifically by pain;
  • the simple model provides reliable individual predictions.

The conclusions should mention the observational design, possible confounding, self-report, attrition, question routing and the aggregated unweighted data.

20 Step 19: Suggested five-minute presentation order

Time Content Main point
0:00–0:40 Context and purpose Menstrual symptoms may affect education and daily participation.
0:40–1:20 Original design Explain the LSAC cohort, strengths and limitations.
1:20–1:50 Research questions and methods Explain why RQ1 uses logistic regression and RQ2–RQ3 are descriptive.
1:50–3:05 RQ1 Report probabilities, OR, 95% CI and p-value.
3:05–4:15 RQ2 and RQ3 Describe the two activity patterns without claiming individual change.
4:15–5:00 Conclusion State what the evidence supports and what it cannot establish.

21 Reference

Cameron, L., Mikocka-Walus, A., Sciberras, E., Druitt, M., Stanley, K., & Evans, S. (2024). Menstrual pain in Australian adolescent girls and its impact on regular activities: A population-based cohort analysis based on Longitudinal Study of Australian Children survey data. Medical Journal of Australia, 220(9), 466–471. https://doi.org/10.5694/mja2.52288

22 Reproducibility information

This final chunk records the R version and package versions used to render the report. It can help identify differences if a group member receives a different result.

sessionInfo()
## R version 4.4.1 (2024-06-14 ucrt)
## Platform: x86_64-w64-mingw32/x64
## Running under: Windows 11 x64 (build 26200)
## 
## Matrix products: default
## 
## 
## locale:
## [1] LC_COLLATE=Chinese (Simplified)_China.utf8 
## [2] LC_CTYPE=Chinese (Simplified)_China.utf8   
## [3] LC_MONETARY=Chinese (Simplified)_China.utf8
## [4] LC_NUMERIC=C                               
## [5] LC_TIME=Chinese (Simplified)_China.utf8    
## 
## time zone: Australia/Sydney
## tzcode source: internal
## 
## attached base packages:
## [1] stats     graphics  grDevices utils     datasets  methods   base     
## 
## other attached packages:
## [1] knitr_1.51    scales_1.4.0  ggplot2_4.0.3 tidyr_1.3.2   dplyr_1.2.1  
## 
## loaded via a namespace (and not attached):
##  [1] gtable_0.3.6       jsonlite_2.0.0     compiler_4.4.1     tidyselect_1.2.1  
##  [5] jquerylib_0.1.4    textshaping_1.0.5  systemfonts_1.3.2  yaml_2.3.12       
##  [9] fastmap_1.2.0      R6_2.6.1           generics_0.1.4     tibble_3.3.1      
## [13] bslib_0.12.0       pillar_1.11.1      RColorBrewer_1.1-3 rlang_1.3.0       
## [17] cachem_1.1.0       xfun_0.60          sass_0.4.10        S7_0.2.2          
## [21] otel_0.2.0         cli_3.6.6          withr_3.0.3        magrittr_2.0.5    
## [25] digest_0.6.39      grid_4.4.1         rstudioapi_0.19.0  lifecycle_1.0.5   
## [29] vctrs_0.7.3        evaluate_1.0.5     glue_1.8.1         farver_2.1.2      
## [33] ragg_1.5.2         rmarkdown_2.31     purrr_1.2.2        tools_4.4.1       
## [37] pkgconfig_2.0.3    htmltools_0.5.9