Welcome to the PSYC3361 coding W1 self test. The test assesses your ability to use the coding skills covered in the Week 1 online coding modules.
In particular, it assesses your ability to…
It is IMPORTANT to document the code that you write so that someone who is looking at your code can understand what it is doing. Above each chunk, write a few sentences outlining which packages/functions you have chosen to use and what the function is doing to your data. Where relevant, also write a sentence that interprets the output of your code.
Your notes should also document the troubleshooting process you went through to arrive at the code that worked.
For each of the challenges below, the documentation is JUST AS IMPORTANT as the code.
Good luck!!
Jenny
library(tidyverse)
birthweight_data = read_csv("./data/birthweight_data.csv")
glimpse(birthweight_data)
## Rows: 788
## Columns: 5
## $ true_ID <dbl> 3100, 3101, 3102, 3103, 3104, 3105, 3106, 3107, 3108, …
## $ birthweight <dbl> 3030, 3710, 3770, 3660, 3800, 3540, 3400, 3650, 3460, …
## $ gestation_age_w <chr> "39", "40", "42", "38", "39", "41", "37", "39", "39", …
## $ child_ethn <chr> "Middle-Eastern", "Caucasian", "African/African-Americ…
## $ plurality <chr> "singleton", "singleton", "singleton", "singleton", "s…
birthweight_by_plurality = birthweight_data %>%
group_by(plurality) %>%
summarise(sum(birthweight))
print(birthweight_by_plurality)
## # A tibble: 2 × 2
## plurality `sum(birthweight)`
## <chr> <dbl>
## 1 singleton 2416589
## 2 twin 101670
birthweight_data %>%
group_by(child_ethn) %>%
summarise(min(gestation_age_w))
## # A tibble: 10 × 2
## child_ethn `min(gestation_age_w)`
## <chr> <chr>
## 1 Aboriginal/Torres Strait Islander 33
## 2 African/African-American 26
## 3 Caucasian 26
## 4 East Asian 33
## 5 Hispanic/Latino 37
## 6 Middle-Eastern 28
## 7 Missing 36
## 8 Polynesian/Melanesian 28
## 9 South Asian 28
## 10 South-East Asian 29
https://www.rdocumentation.org/packages/dplyr/versions/1.0.10/topics/group_by
https://www.rdocumentation.org/packages/dplyr/versions/0.7.8/topics/summarise
Group by takes an ungrouped dataframe and returns a grouped dataframe. In data analysis, mostly one wants to perform operations on groups, not on the whole table. For example, we often care about the differences between groups in an experiment, or in a between subjects design, the differences between two trial types.
This is the purpose of the group_by function. This is very powerful, because it allows us to store all our data in one dataframe, but still be able to make comparisons between subsets of this data. Using the earlier analogy of an RCT, this would mean that we don’t have to keep the control and intervention group’s data separate, we can keep them in the same file with a column that indicates group membership. Storing all data in a 2d table is best practice, and the group_by function is both the reason for this and critical to the functioning of this system.
Another fundamental operation in data analysis is summarising data. For example, in an RCT it is very common to compare the mean scores of two different groups on some outcome measure. This is a form of summarisation.
Summarise takes a column or columns as input and returns a single value for these data. It also take a function to apply to these data, which is how it obtains the single value representing the input data.
write_csv(birthweight_by_plurality, "birthweight_by_plurality.csv")