DATA 101 — Homework 3: Data Wrangling

Author

Dawit Merdassa

Instructions

You will be working on the olympic_gymnasts dataset from TidyTuesday. Please DO NOT change the setup code below:

olympics <- readr::read_csv('https://raw.githubusercontent.com/rfordatascience/tidytuesday/master/data/2021/2021-07-27/olympics.csv')

olympic_gymnasts <- olympics |> 
  filter(!is.na(age)) |>             # only keep athletes with known age
  filter(sport == "Gymnastics") |>   # keep only gymnasts
  mutate(
    medalist = case_when(             # add column for success in medaling
      is.na(medal) ~ FALSE,           # NA values go to FALSE
      !is.na(medal) ~ TRUE            # non-NA values (Gold, Silver, Bronze) go to TRUE
    )
  )

More information about the dataset can be found at:
https://github.com/rfordatascience/tidytuesday/blob/master/data/2021/2021-07-27/readme.md


Question 1: Create a subset dataset with the following columns only: name, sex, age, team, year, and medalist. Call this new dataset df.

# Your code here
df <- olympic_gymnasts |> 
  select(name, sex, age, team, year, medalist)

Question 2: From df, create df2 that only contains gymnasts from the years 2008, 2012, and 2016.

# Your code here
df2 <- df |> 
  filter(year %in% c(2008, 2012, 2016))

Question 3: Group df2 by year (2008, 2012, and 2016) and summarize the mean of the age in each group.

# Your code here
df2 |> 
  group_by(year) |> 
  summarize(mean_age = mean(age))
# A tibble: 3 × 2
   year mean_age
  <dbl>    <dbl>
1  2008     21.6
2  2012     21.9
3  2016     22.2

Question 4: Using the full olympic_gymnasts dataset, group by year and find the mean of the age for each year. Call this dataset oly_year.
(Optional: After creating the dataset, find the minimum average age and the year it occurred).

# Your code here
oly_year <- olympic_gymnasts |> 
  group_by(year) |> 
  summarize(mean_age = mean(age))

min_age_year <- oly_year |>
  arrange(mean_age) |>
  slice(1)

Question 5 (Open-ended): Create a question that requires you to use at least two dplyr verbs (e.g., filter, select, mutate, group_by, summarize, arrange). Write the code that answers your question, and below the chunk, write a brief reflection on your question choice and findings.

Example question: Do gymnasts aged 20 or younger win more medals compared to older gymnasts?

# Your R code here
medal_rates <- olympic_gymnasts |> 
  filter(year >= 2000, year <= 2016) |> 
  group_by(year) |> 
  summarize(
    total_gymnasts = n(),
    medalists = sum(medalist),
    medal_rate = medalists / total_gymnasts
  ) |> 
  arrange(desc(medal_rate))

medal_rates
# A tibble: 5 × 4
   year total_gymnasts medalists medal_rate
  <dbl>          <int>     <int>      <dbl>
1  2012            848        66     0.0778
2  2016            861        66     0.0767
3  2008            994        72     0.0724
4  2000           1144        72     0.0629
5  2004           1151        72     0.0626

Reflection & Discussion:
(I chose this question because I was curious about which year had the most successful gymnasts. I compared the medal rates for year and found it interesting to see the differences. This helped me practice using dplyr to work the data.)