This is my Week 1 Self Test Document

Excited to try this out!

Loading the needed packages

I’ve chosen to load up tidyverse as its very helpful for piping and wrangling data.

# install.packages("tidyverse") if you don't already have package on your device/cloudspace 
library(tidyverse) #loads tidyverse and its helpful shortcuts 
## ── Attaching core tidyverse packages ──────────────────────── tidyverse 2.0.0 ──
## ✔ dplyr     1.1.2     ✔ readr     2.1.4
## ✔ forcats   1.0.0     ✔ stringr   1.5.0
## ✔ ggplot2   3.4.2     ✔ tibble    3.2.1
## ✔ lubridate 1.9.2     ✔ tidyr     1.3.0
## ✔ purrr     1.0.1     
## ── Conflicts ────────────────────────────────────────── tidyverse_conflicts() ──
## ✖ dplyr::filter() masks stats::filter()
## ✖ dplyr::lag()    masks stats::lag()
## ℹ Use the conflicted package (<http://conflicted.r-lib.org/>) to force all conflicts to become errors
library(here) # loads the here package for locating files
## here() starts at /cloud/project

Reading Data

For the self test this week we are looking at the birth weight data that is available in the self test cloud project files. To look at the data briefly i’m going to use the glimpse fucntion from tidyverse. I think its from tidyverse im not 100% sure actually.

frames <- read_csv(here::here("data", "birthweight_data.csv")) #to start your work space looking at the data (now called frames )
## Rows: 788 Columns: 5
## ── Column specification ────────────────────────────────────────────────────────
## Delimiter: ","
## chr (3): gestation_age_w, child_ethn, plurality
## dbl (2): true_ID, birthweight
## 
## ℹ Use `spec()` to retrieve the full column specification for this data.
## ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
glimpse(frames) #running this will give you an overview of the data, not 
## Rows: 788
## Columns: 5
## $ true_ID         <dbl> 3100, 3101, 3102, 3103, 3104, 3105, 3106, 3107, 3108, …
## $ birthweight     <dbl> 3030, 3710, 3770, 3660, 3800, 3540, 3400, 3650, 3460, …
## $ gestation_age_w <chr> "39", "40", "42", "38", "39", "41", "37", "39", "39", …
## $ child_ethn      <chr> "Middle-Eastern", "Caucasian", "African/African-Americ…
## $ plurality       <chr> "singleton", "singleton", "singleton", "singleton", "s…

Glimpse is kind of hell to look at but the data is there

Further inspection

Now we gotta calculate the mean birthweight separately for twins and singletons.We are going to use the pipe function.

frames %>% 
  group_by (plurality) %>% #this separates the singletons and twins
  summarise (birthweight = mean (birthweight)) # then this gives the average birthweight 
## # A tibble: 2 × 2
##   plurality birthweight
##   <chr>           <dbl>
## 1 singleton       3248.
## 2 twin            2311.

More exploring

Now we gotta look at the earliest (i.e. the minimum value) gestational age for each ethicity group. I’ve never done this before.

frames %>% 
  group_by (child_ethn) %>% #This groups by ethnicity 
  summarise (gestational_age_w = min(gestation_age_w)) #this finds min gestational age in weeks #initially i tried using minimum_value to find the eariest age, but when that didn't work I googled how to find the min value in R. Makes sense it would be faster to type out 
## # A tibble: 10 × 2
##    child_ethn                        gestational_age_w
##    <chr>                             <chr>            
##  1 Aboriginal/Torres Strait Islander 33               
##  2 African/African-American          26               
##  3 Caucasian                         26               
##  4 East Asian                        33               
##  5 Hispanic/Latino                   37               
##  6 Middle-Eastern                    28               
##  7 Missing                           36               
##  8 Polynesian/Melanesian             28               
##  9 South Asian                       28               
## 10 South-East Asian                  29

Group By and Summarise are linked together to work as one command to run through use if the pipe function %>% . It acts as a “and then” for the code instructions. So in this case it could be read as group by ethnicity and then find the minimum gestational age.

Baby Picture

I gotta find a picture to now load into this document. I’ll find one, download it and then add it to the cloud project.

I tried to do this inside a coding chunk for way too long… eventually figured out it needed to just be in the text of the Rmd document.

Save summary of data to a new csv file

I done this before but ive forgotten how so i’m going to make some educated guesses then look it up…

# frames %>% 
 # group_by (plurality) %>% #this separates the singletons and twins
  #summarise (birthweight = mean (birthweight)) %>% # gives the average birthweight 
  #save( file = "twins_single_birthweight_sum.csv") #this was me attempting to save the info to a csv file without looking it up. It did create a file but it doesn't contain the info on it 

Lets try the actual way now!

my_twin_singles_birthweight_summary <- frames %>% #this turns the output into the 'object of "my_twin_singles_birthweight_summary" that can be printed and saved 
  group_by (plurality) %>% #this separates the singletons and twins
  summarise (
    birthweight = mean (birthweight) #gives the average birthweight #now also making this easier to read 
    ) %>% 
  ungroup () # forgot to ungroup for later investigations earlier thats important 
write_csv(my_twin_singles_birthweight_summary, file = "my_twin_singles_birthweight_summary.csv") #now this will save as a csv file with the info on it this time #R now likes 'file' rather than 'path' that was in the original video 

Publishing to Rpubs

Now for publishing. First you must check that knitting has been working - it has. Now in the top right corner I’ll use the publish button. https://rpubs.com/sally_obryan/1193802 Its now saved at this URL