olympics <- readr::read_csv('https://raw.githubusercontent.com/rfordatascience/tidytuesday/master/data/2021/2021-07-27/olympics.csv')
olympic_gymnasts <- olympics |>
filter(!is.na(age)) |> # only keep athletes with known age
filter(sport == "Gymnastics") |> # keep only gymnasts
mutate(
medalist = case_when( # add column for success in medaling
is.na(medal) ~ FALSE, # NA values go to FALSE
!is.na(medal) ~ TRUE # non-NA values (Gold, Silver, Bronze) go to TRUE
)
)DATA 101 — Homework 3: Data Wrangling
Instructions
You will be working on the olympic_gymnasts dataset from TidyTuesday. Please DO NOT change the setup code below:
More information about the dataset can be found at:
https://github.com/rfordatascience/tidytuesday/blob/master/data/2021/2021-07-27/readme.md
Question 1: Create a subset dataset with the following columns only: name, sex, age, team, year, and medalist. Call this new dataset df.
df <- olympic_gymnasts |>
select(name, sex, age, team, year, medalist)
df# A tibble: 25,528 x 6
name sex age team year medalist
<chr> <chr> <dbl> <chr> <dbl> <lgl>
1 Paavo Johannes Aaltonen M 28 Finland 1948 TRUE
2 Paavo Johannes Aaltonen M 28 Finland 1948 TRUE
3 Paavo Johannes Aaltonen M 28 Finland 1948 FALSE
4 Paavo Johannes Aaltonen M 28 Finland 1948 TRUE
5 Paavo Johannes Aaltonen M 28 Finland 1948 FALSE
6 Paavo Johannes Aaltonen M 28 Finland 1948 FALSE
7 Paavo Johannes Aaltonen M 28 Finland 1948 FALSE
8 Paavo Johannes Aaltonen M 28 Finland 1948 TRUE
9 Paavo Johannes Aaltonen M 32 Finland 1952 FALSE
10 Paavo Johannes Aaltonen M 32 Finland 1952 TRUE
# i 25,518 more rows
Question 2: From df, create df2 that only contains gymnasts from the years 2008, 2012, and 2016.
df2 <- df |>
filter(year %in% c(2008, 2012, 2016))
df2# A tibble: 2,703 x 6
name sex age team year medalist
<chr> <chr> <dbl> <chr> <dbl> <lgl>
1 Nstor Abad Sanjun M 23 Spain 2016 FALSE
2 Nstor Abad Sanjun M 23 Spain 2016 FALSE
3 Nstor Abad Sanjun M 23 Spain 2016 FALSE
4 Nstor Abad Sanjun M 23 Spain 2016 FALSE
5 Nstor Abad Sanjun M 23 Spain 2016 FALSE
6 Nstor Abad Sanjun M 23 Spain 2016 FALSE
7 Katja Abel F 25 Germany 2008 FALSE
8 Katja Abel F 25 Germany 2008 FALSE
9 Katja Abel F 25 Germany 2008 FALSE
10 Katja Abel F 25 Germany 2008 FALSE
# i 2,693 more rows
Question 3: Group df2 by year (2008, 2012, and 2016) and summarize the mean of the age in each group.
df2 |>
group_by(year) |>
summarize(mean_age = mean(age))# A tibble: 3 x 2
year mean_age
<dbl> <dbl>
1 2008 21.6
2 2012 21.9
3 2016 22.2
Question 4: Using the full olympic_gymnasts dataset, group by year and find the mean of the age for each year. Call this dataset oly_year.
(Optional: After creating the dataset, find the minimum average age and the year it occurred).
oly_year <- olympic_gymnasts |>
group_by(year) |>
summarize(mean_age = mean(age))
oly_year# A tibble: 29 x 2
year mean_age
<dbl> <dbl>
1 1896 24.3
2 1900 22.2
3 1904 25.1
4 1906 24.7
5 1908 23.2
6 1912 24.2
7 1920 26.7
8 1924 27.6
9 1928 25.6
10 1932 23.9
# i 19 more rows
# Optional: show the year with the lowest average age
oly_year |>
slice_min(mean_age, n = 1, with_ties = FALSE)# A tibble: 1 x 2
year mean_age
<dbl> <dbl>
1 1988 19.9
Question 5 (Open-ended): Create a question that requires you to use at least two dplyr verbs (e.g., filter, select, mutate, group_by, summarize, arrange). Write the code that answers your question, and below the chunk, write a brief reflection on your question choice and findings.
My question: How does the medalist rate among gymnasts age 20 or younger compare with the rate among gymnasts older than 20?
age_medal_summary <- olympic_gymnasts |>
mutate(age_group = if_else(age <= 20, "20 or younger", "Older than 20")) |>
group_by(age_group) |>
summarize(
observations = n(),
medalist_rate = mean(medalist)
) |>
arrange(desc(medalist_rate))
age_medal_summary# A tibble: 2 x 3
age_group observations medalist_rate
<chr> <int> <dbl>
1 Older than 20 16676 0.0918
2 20 or younger 8852 0.0741
Reflection & Discussion:
I chose this question because age may relate to both experience and physical demands in gymnastics. I used mutate() to create two age groups, group_by() and summarize() to calculate the share of observations associated with a medal, and arrange() to make the comparison easy to read. The results show that gymnasts older than 20 had a medalist rate of 9.2%, compared with 7.4% for gymnasts age 20 or younger. This suggests that the older group had the higher medalist rate in these data. Because each row represents an Olympic event participation rather than one unique athlete, this comparison describes event-level observations and should not be interpreted as a causal effect of age.