Welcome to the PSYC3361 coding W1 self test. The test assesses your ability to use the coding skills covered in the Week 1 online coding modules.

In particular, it assesses your ability to…

  • choose packages/functions
  • read in data
  • group_by and summarise
  • make notes using RMarkdown
  • insert pictures in an Rmd document
  • write data to csv

It is IMPORTANT to document the code that you write so that someone who is looking at your code can understand what it is doing. Above each chunk, write a few sentences outlining which packages/functions you have chosen to use and what the function is doing to your data. Where relevant, also write a sentence that interprets the output of your code.

Your notes should also document the troubleshooting process you went through to arrive at the code that worked.

For each of the challenges below, the documentation is JUST AS IMPORTANT as the code.

Good luck!!

Jenny

1. customise your Rmd document by adding your name as the author, a table of contents and choosing a theme that you like.

In the YAML section above, the document has been customised with a title, name, date and HTML output with the “cerulean” theme and a floating table of contents.

2. load the packages you will need

This chunk of code loads the tidyverse package.

library(tidyverse)

3. read the birthweight data

This chunk of code reads the file “birthweight_data.csv” and gives it to the variable birthweight.

birthweight <- read_csv(file = "birthweight_data.csv")

I could not call “birthweight_data.csv” at first. I had to move the csv file from within the data file, in project, to just a stand alone file within project.

4. calculate the mean birthweight separately for twins and singletons

This pipe groups the data into twins and singletons (via the plurality variable) and then summarises the birthweights of each into calculated means. Then, using “print(mean_birthweight)” this pipe is displayed.

mean_birthweight_summary <- birthweight %>% 
  group_by(plurality) %>% 
  summarise(
    mean_birth = mean(birthweight)
  ) %>% 
ungroup()

print(mean_birthweight_summary)
## # A tibble: 2 × 2
##   plurality mean_birth
##   <chr>          <dbl>
## 1 singleton      3248.
## 2 twin           2311.

Thus, the mean birthweight is 3248g for singletons, and is 2311g for twins (These values are assumed to be in grams).

5. identify the earliest (i.e. the minimum value) gestational age for each ethicity group

This pipe groups the data into ethnicities (via the child_ethn variable) and then summarises the gestation age of each into the minimum value. Then, using “print(min_gest_age_ethn)” this pipe is displayed.

min_gest_age_ethn <- birthweight %>% 
  group_by(child_ethn) %>% 
  summarise(min_gest = min(gestation_age_w)) %>% 
  ungroup()

print(min_gest_age_ethn)
## # A tibble: 10 × 2
##    child_ethn                        min_gest
##    <chr>                             <chr>   
##  1 Aboriginal/Torres Strait Islander 33      
##  2 African/African-American          26      
##  3 Caucasian                         26      
##  4 East Asian                        33      
##  5 Hispanic/Latino                   37      
##  6 Middle-Eastern                    28      
##  7 Missing                           36      
##  8 Polynesian/Melanesian             28      
##  9 South Asian                       28      
## 10 South-East Asian                  29

I had difficulty here figuring out how to either sort or filter through the data with maybe an ‘If()’ case sitution. After further thought I realised that there was a likely chance that R had a function that did this for me - I did some research into different R functions and saw that my theory was correct!

6. write some notes about how group_by and summarise work with the pipe below, including a link to documentation or a blog post that you think is useful

When using a pipe, the group_by function is used to group/categorise data with according to the selected grouping variable/s, and the summarise function is then used to return one row of specified summary statistics for each of the group variables.

This can be visualised using pipe in question 3:

“mean_birthweight_summary <- birthweight %>% group_by(plurality) %>% summarise( mean_birth = mean(birthweight) ) %>% ungroup()

print(mean_birthweight_summary)”

Where the dataset with grouped in singletons and twins via the specified grouping variable, plurality, and then the summarise function calculate the mean birthweight for each of these groups.

A useful post, that goes further into this can be accessed here.

7. download a picture of a baby from the internet and insert it into your document below

I have sourced an image of a baby from the website “People.com”, referenced below.

Andaloro, A. (2023, May 17). Luna to oliver: See the most popular baby names in 2022. Peoplemag. https://people.com/parents/most-popular-baby-names-2022-revealed/

8. write the summary of mean birthweight by twins/singletons that you made in step 4 above to a new csv file

Here I write the pipe summary of mean twin/singleton birthweight, “mean_birthweight_summary”, that I created in question 4 to a new csv document. I have named this document “mean_birth_plurality.csv”.

write_csv(mean_birthweight_summary, file = "mean_birthweight_plurality.csv")

9. Knit your document and publish the output to RPubs