DATA 101 — Homework 3: Data Wrangling

Author

Enda Hughes

Instructions

You will be working on the olympic_gymnasts dataset from TidyTuesday. Please DO NOT change the setup code below:

olympics <- readr::read_csv('https://raw.githubusercontent.com/rfordatascience/tidytuesday/master/data/2021/2021-07-27/olympics.csv')

olympic_gymnasts <- olympics |> 
  filter(!is.na(age)) |>             # only keep athletes with known age
  filter(sport == "Gymnastics") |>   # keep only gymnasts
  mutate(
    medalist = case_when(             # add column for success in medaling
      is.na(medal) ~ FALSE,           # NA values go to FALSE
      !is.na(medal) ~ TRUE            # non-NA values (Gold, Silver, Bronze) go to TRUE
    )
  )

More information about the dataset can be found at:
https://github.com/rfordatascience/tidytuesday/blob/master/data/2021/2021-07-27/readme.md


Question 1: Create a subset dataset with the following columns only: name, sex, age, team, year, and medalist. Call this new dataset df.

df <- olympic_gymnasts |>
  select(name, sex, age, team, year, medalist)

df
# A tibble: 25,528 × 6
   name                    sex     age team     year medalist
   <chr>                   <chr> <dbl> <chr>   <dbl> <lgl>   
 1 Paavo Johannes Aaltonen M        28 Finland  1948 TRUE    
 2 Paavo Johannes Aaltonen M        28 Finland  1948 TRUE    
 3 Paavo Johannes Aaltonen M        28 Finland  1948 FALSE   
 4 Paavo Johannes Aaltonen M        28 Finland  1948 TRUE    
 5 Paavo Johannes Aaltonen M        28 Finland  1948 FALSE   
 6 Paavo Johannes Aaltonen M        28 Finland  1948 FALSE   
 7 Paavo Johannes Aaltonen M        28 Finland  1948 FALSE   
 8 Paavo Johannes Aaltonen M        28 Finland  1948 TRUE    
 9 Paavo Johannes Aaltonen M        32 Finland  1952 FALSE   
10 Paavo Johannes Aaltonen M        32 Finland  1952 TRUE    
# ℹ 25,518 more rows

Question 2: From df, create df2 that only contains gymnasts from the years 2008, 2012, and 2016.

df2 <- df |>
  filter(year %in% c(2008, 2012, 2016))

df2
# A tibble: 2,703 × 6
   name              sex     age team     year medalist
   <chr>             <chr> <dbl> <chr>   <dbl> <lgl>   
 1 Nstor Abad Sanjun M        23 Spain    2016 FALSE   
 2 Nstor Abad Sanjun M        23 Spain    2016 FALSE   
 3 Nstor Abad Sanjun M        23 Spain    2016 FALSE   
 4 Nstor Abad Sanjun M        23 Spain    2016 FALSE   
 5 Nstor Abad Sanjun M        23 Spain    2016 FALSE   
 6 Nstor Abad Sanjun M        23 Spain    2016 FALSE   
 7 Katja Abel        F        25 Germany  2008 FALSE   
 8 Katja Abel        F        25 Germany  2008 FALSE   
 9 Katja Abel        F        25 Germany  2008 FALSE   
10 Katja Abel        F        25 Germany  2008 FALSE   
# ℹ 2,693 more rows

Question 3: Group df2 by year (2008, 2012, and 2016) and summarize the mean of the age in each group.

df2 |>
  group_by(year) |>
  summarize(mean_age = mean(age), .groups = "drop")
# A tibble: 3 × 2
   year mean_age
  <dbl>    <dbl>
1  2008     21.6
2  2012     21.9
3  2016     22.2

Question 4: Using the full olympic_gymnasts dataset, group by year and find the mean of the age for each year. Call this dataset oly_year.
(Optional: After creating the dataset, find the minimum average age and the year it occurred).

oly_year <- olympic_gymnasts |>
  group_by(year) |>
  summarize(mean_age = mean(age), .groups = "drop")

oly_year
# A tibble: 29 × 2
    year mean_age
   <dbl>    <dbl>
 1  1896     24.3
 2  1900     22.2
 3  1904     25.1
 4  1906     24.7
 5  1908     23.2
 6  1912     24.2
 7  1920     26.7
 8  1924     27.6
 9  1928     25.6
10  1932     23.9
# ℹ 19 more rows
# Optional: display the year with the minimum average age
oly_year |>
  slice_min(mean_age, n = 1, with_ties = FALSE)
# A tibble: 1 × 2
   year mean_age
  <dbl>    <dbl>
1  1988     19.9

Question 5 (Open-ended): Create a question that requires you to use at least two dplyr verbs (e.g., filter, select, mutate, group_by, summarize, arrange). Write the code that answers your question, and below the chunk, write a brief reflection on your question choice and findings.

Question: Among gymnast records from 2008, 2012, and 2016, how did average age differ between medalists and non-medalists in each year?

age_by_medalist <- olympic_gymnasts |>
  filter(year %in% c(2008, 2012, 2016)) |>
  group_by(year, medalist) |>
  summarize(
    mean_age = mean(age),
    observations = n(),
    .groups = "drop"
  ) |>
  arrange(year, desc(medalist))

age_by_medalist
# A tibble: 6 × 4
   year medalist mean_age observations
  <dbl> <lgl>       <dbl>        <int>
1  2008 TRUE         21.1           72
2  2008 FALSE        21.7          922
3  2012 TRUE         20.8           66
4  2012 FALSE        22.0          782
5  2016 TRUE         21.8           66
6  2016 FALSE        22.2          795

Reflection & Discussion:
I chose this question because it compares age with a clear measure of Olympic success while also allowing the comparison to be repeated across three recent Olympic years. In all three years, medalist records had a lower average age than non-medalist records. The gap was about 0.59 years in 2008, 1.12 years in 2012, and 0.42 years in 2016. Both groups were oldest on average in 2016. These results describe rows in the dataset rather than unique athletes, so gymnasts who competed in multiple events may appear more than once. Therefore, the findings show an association in the recorded events and do not establish that being younger causes a gymnast to win a medal.