olympics <- readr::read_csv('https://raw.githubusercontent.com/rfordatascience/tidytuesday/master/data/2021/2021-07-27/olympics.csv')
olympic_gymnasts <- olympics |>
filter(!is.na(age)) |> # only keep athletes with known age
filter(sport == "Gymnastics") |> # keep only gymnasts
mutate(
medalist = case_when( # add column for success in medaling
is.na(medal) ~ FALSE, # NA values go to FALSE
!is.na(medal) ~ TRUE # non-NA values (Gold, Silver, Bronze) go to TRUE
)
)DATA 101 — Homework 3: Data Wrangling
Instructions
You will be working on the olympic_gymnasts dataset from TidyTuesday. Please DO NOT change the setup code below:
More information about the dataset can be found at:
https://github.com/rfordatascience/tidytuesday/blob/master/data/2021/2021-07-27/readme.md
Question 1: Create a subset dataset with the following columns only: name, sex, age, team, year, and medalist. Call this new dataset df.
# Your code here
df <- olympic_gymnasts |>
select(name, sex, age, team, year, medalist)Question 2: From df, create df2 that only contains gymnasts from the years 2008, 2012, and 2016.
# Your code here
df2 <- df |>
filter(year %in% c(2008, 2012, 2016))Question 3: Group df2 by year (2008, 2012, and 2016) and summarize the mean of the age in each group.
# Your code here
df2 |>
group_by(year) |>
summarize(mean_age = mean(age))# A tibble: 3 × 2
year mean_age
<dbl> <dbl>
1 2008 21.6
2 2012 21.9
3 2016 22.2
Question 4: Using the full olympic_gymnasts dataset, group by year and find the mean of the age for each year. Call this dataset oly_year.
(Optional: After creating the dataset, find the minimum average age and the year it occurred).
# Your code here
oly_year <- olympic_gymnasts |>
group_by(year) |>
summarize(mean_age = mean(age))
min_age_year <- oly_year |>
arrange(mean_age) |>
slice(1)Question 5 (Open-ended): Create a question that requires you to use at least two dplyr verbs (e.g., filter, select, mutate, group_by, summarize, arrange). Write the code that answers your question, and below the chunk, write a brief reflection on your question choice and findings.
Example question: Do gymnasts aged 20 or younger win more medals compared to older gymnasts?
# Your R code here
medal_rates <- olympic_gymnasts |>
filter(year >= 2000, year <= 2016) |>
group_by(year) |>
summarize(
total_gymnasts = n(),
medalists = sum(medalist),
medal_rate = medalists / total_gymnasts
) |>
arrange(desc(medal_rate))
medal_rates# A tibble: 5 × 4
year total_gymnasts medalists medal_rate
<dbl> <int> <int> <dbl>
1 2012 848 66 0.0778
2 2016 861 66 0.0767
3 2008 994 72 0.0724
4 2000 1144 72 0.0629
5 2004 1151 72 0.0626
Reflection & Discussion:
(I chose this question because I was curious about which year had the most successful gymnasts. I compared the medal rates for year and found it interesting to see the differences. This helped me practice using dplyr to work the data.)