Excited to try this out!
I’ve chosen to load up tidyverse as its very helpful for piping and wrangling data.
# install.packages("tidyverse") if you don't already have package on your device/cloudspace
library(tidyverse) #loads tidyverse and its helpful shortcuts
## ── Attaching core tidyverse packages ──────────────────────── tidyverse 2.0.0 ──
## ✔ dplyr 1.1.2 ✔ readr 2.1.4
## ✔ forcats 1.0.0 ✔ stringr 1.5.0
## ✔ ggplot2 3.4.2 ✔ tibble 3.2.1
## ✔ lubridate 1.9.2 ✔ tidyr 1.3.0
## ✔ purrr 1.0.1
## ── Conflicts ────────────────────────────────────────── tidyverse_conflicts() ──
## ✖ dplyr::filter() masks stats::filter()
## ✖ dplyr::lag() masks stats::lag()
## ℹ Use the conflicted package (<http://conflicted.r-lib.org/>) to force all conflicts to become errors
library(here) # loads the here package for locating files
## here() starts at /cloud/project
For the self test this week we are looking at the birth weight data that is available in the self test cloud project files. To look at the data briefly i’m going to use the glimpse fucntion from tidyverse. I think its from tidyverse im not 100% sure actually.
frames <- read_csv(here::here("data", "birthweight_data.csv")) #to start your work space looking at the data (now called frames )
## Rows: 788 Columns: 5
## ── Column specification ────────────────────────────────────────────────────────
## Delimiter: ","
## chr (3): gestation_age_w, child_ethn, plurality
## dbl (2): true_ID, birthweight
##
## ℹ Use `spec()` to retrieve the full column specification for this data.
## ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
glimpse(frames) #running this will give you an overview of the data, not
## Rows: 788
## Columns: 5
## $ true_ID <dbl> 3100, 3101, 3102, 3103, 3104, 3105, 3106, 3107, 3108, …
## $ birthweight <dbl> 3030, 3710, 3770, 3660, 3800, 3540, 3400, 3650, 3460, …
## $ gestation_age_w <chr> "39", "40", "42", "38", "39", "41", "37", "39", "39", …
## $ child_ethn <chr> "Middle-Eastern", "Caucasian", "African/African-Americ…
## $ plurality <chr> "singleton", "singleton", "singleton", "singleton", "s…
Glimpse is kind of hell to look at but the data is there
Now we gotta calculate the mean birthweight separately for twins and singletons.We are going to use the pipe function.
frames %>%
group_by (plurality) %>% #this separates the singletons and twins
summarise (birthweight = mean (birthweight)) # then this gives the average birthweight
## # A tibble: 2 × 2
## plurality birthweight
## <chr> <dbl>
## 1 singleton 3248.
## 2 twin 2311.
Now we gotta look at the earliest (i.e. the minimum value) gestational age for each ethicity group. I’ve never done this before.
frames %>%
group_by (child_ethn) %>% #This groups by ethnicity
summarise (gestational_age_w = min(gestation_age_w)) #this finds min gestational age in weeks #initially i tried using minimum_value to find the eariest age, but when that didn't work I googled how to find the min value in R. Makes sense it would be faster to type out
## # A tibble: 10 × 2
## child_ethn gestational_age_w
## <chr> <chr>
## 1 Aboriginal/Torres Strait Islander 33
## 2 African/African-American 26
## 3 Caucasian 26
## 4 East Asian 33
## 5 Hispanic/Latino 37
## 6 Middle-Eastern 28
## 7 Missing 36
## 8 Polynesian/Melanesian 28
## 9 South Asian 28
## 10 South-East Asian 29
Group By and Summarise are linked together to work as one command to run through use if the pipe function %>% . It acts as a “and then” for the code instructions. So in this case it could be read as group by ethnicity and then find the minimum gestational age.
I gotta find a picture to now load into this document. I’ll find one, download it and then add it to the cloud project.
I tried to do this inside a coding chunk for way too long… eventually figured out it needed to just be in the text of the Rmd document.
I done this before but ive forgotten how so i’m going to make some educated guesses then look it up…
# frames %>%
# group_by (plurality) %>% #this separates the singletons and twins
#summarise (birthweight = mean (birthweight)) %>% # gives the average birthweight
#save( file = "twins_single_birthweight_sum.csv") #this was me attempting to save the info to a csv file without looking it up. It did create a file but it doesn't contain the info on it
Lets try the actual way now!
my_twin_singles_birthweight_summary <- frames %>% #this turns the output into the 'object of "my_twin_singles_birthweight_summary" that can be printed and saved
group_by (plurality) %>% #this separates the singletons and twins
summarise (
birthweight = mean (birthweight) #gives the average birthweight #now also making this easier to read
) %>%
ungroup () # forgot to ungroup for later investigations earlier thats important
write_csv(my_twin_singles_birthweight_summary, file = "my_twin_singles_birthweight_summary.csv") #now this will save as a csv file with the info on it this time #R now likes 'file' rather than 'path' that was in the original video
Now for publishing. First you must check that knitting has been working - it has. Now in the top right corner I’ll use the publish button. https://rpubs.com/sally_obryan/1193802 Its now saved at this URL