library(tidyverse)
## ── Attaching core tidyverse packages ──────────────────────── tidyverse 2.0.0 ──
## ✔ dplyr 1.2.1 ✔ readr 2.2.0
## ✔ forcats 1.0.1 ✔ stringr 1.6.0
## ✔ ggplot2 4.0.3 ✔ tibble 3.3.1
## ✔ lubridate 1.9.5 ✔ tidyr 1.3.2
## ✔ purrr 1.2.2
## ── Conflicts ────────────────────────────────────────── tidyverse_conflicts() ──
## ✖ dplyr::filter() masks stats::filter()
## ✖ dplyr::lag() masks stats::lag()
## ℹ Use the conflicted package (<http://conflicted.r-lib.org/>) to force all conflicts to become errors
library(readxl)
library(pastecs)
##
## Attaching package: 'pastecs'
##
## The following objects are masked from 'package:dplyr':
##
## first, last
##
## The following object is masked from 'package:tidyr':
##
## extract
- From the data you have chosen, select a variable that you are
interested in
County_Data <- read_excel("County Data.xlsx")
## New names:
## • `Unreliable` -> `Unreliable...4`
## • `95% CI - Low` -> `95% CI - Low...7`
## • `95% CI - High` -> `95% CI - High...8`
## • `National Z-Score` -> `National Z-Score...9`
## • `95% CI - Low` -> `95% CI - Low...39`
## • `95% CI - High` -> `95% CI - High...40`
## • `National Z-Score` -> `National Z-Score...41`
## • `Unreliable` -> `Unreliable...42`
## • `95% CI - Low` -> `95% CI - Low...44`
## • `95% CI - High` -> `95% CI - High...45`
## • `National Z-Score` -> `National Z-Score...46`
## • `95% CI - Low` -> `95% CI - Low...69`
## • `95% CI - High` -> `95% CI - High...70`
## • `National Z-Score` -> `National Z-Score...71`
## • `95% CI - Low` -> `95% CI - Low...73`
## • `95% CI - High` -> `95% CI - High...74`
## • `National Z-Score` -> `National Z-Score...75`
## • `National Z-Score` -> `National Z-Score...77`
## • `National Z-Score` -> `National Z-Score...84`
## • `National Z-Score` -> `National Z-Score...86`
## • `National Z-Score` -> `National Z-Score...90`
## • `National Z-Score` -> `National Z-Score...94`
## • `National Z-Score` -> `National Z-Score...98`
## • `National Z-Score` -> `National Z-Score...100`
## • `National Z-Score` -> `National Z-Score...107`
## • `95% CI - Low` -> `95% CI - Low...115`
## • `95% CI - High` -> `95% CI - High...116`
## • `National Z-Score` -> `National Z-Score...117`
## • `95% CI - Low` -> `95% CI - Low...119`
## • `95% CI - High` -> `95% CI - High...120`
## • `National Z-Score` -> `National Z-Score...130`
## • `95% CI - Low` -> `95% CI - Low...132`
## • `95% CI - High` -> `95% CI - High...133`
## • `National Z-Score` -> `National Z-Score...134`
## • `95% CI - Low` -> `95% CI - Low...152`
## • `95% CI - High` -> `95% CI - High...153`
## • `National Z-Score` -> `National Z-Score...154`
## • `National Z-Score` -> `National Z-Score...156`
## • `National Z-Score` -> `National Z-Score...158`
## • `95% CI - Low` -> `95% CI - Low...161`
## • `95% CI - High` -> `95% CI - High...162`
## • `National Z-Score` -> `National Z-Score...163`
## • `National Z-Score` -> `National Z-Score...165`
## • `Population` -> `Population...167`
## • `95% CI - Low` -> `95% CI - Low...169`
## • `95% CI - High` -> `95% CI - High...170`
## • `National Z-Score` -> `National Z-Score...171`
## • `Population` -> `Population...173`
## • `95% CI - Low` -> `95% CI - Low...175`
## • `95% CI - High` -> `95% CI - High...176`
## • `National Z-Score` -> `National Z-Score...177`
## • `National Z-Score` -> `National Z-Score...181`
## • `National Z-Score` -> `National Z-Score...185`
## • `95% CI - Low` -> `95% CI - Low...187`
## • `95% CI - High` -> `95% CI - High...188`
## • `National Z-Score` -> `National Z-Score...189`
## • `95% CI - Low` -> `95% CI - Low...197`
## • `95% CI - High` -> `95% CI - High...198`
## • `National Z-Score` -> `National Z-Score...199`
## • `National Z-Score` -> `National Z-Score...223`
## • `National Z-Score` -> `National Z-Score...225`
## • `National Z-Score` -> `National Z-Score...226`
## • `Health Group` -> `Health Group...227`
## • `Health Group Range` -> `Health Group Range...228`
## • `National Z-Score` -> `National Z-Score...229`
## • `Health Group` -> `Health Group...230`
## • `Health Group Range` -> `Health Group Range...231`
## • `95% CI - Low` -> `95% CI - Low...233`
## • `95% CI - High` -> `95% CI - High...234`
## • `# Deaths` -> `# Deaths...256`
## • `95% CI - Low` -> `95% CI - Low...258`
## • `95% CI - High` -> `95% CI - High...259`
## • `# Deaths` -> `# Deaths...281`
## • `95% CI - Low` -> `95% CI - Low...283`
## • `95% CI - High` -> `95% CI - High...284`
## • `# Deaths` -> `# Deaths...306`
## • `95% CI - Low` -> `95% CI - Low...308`
## • `95% CI - High` -> `95% CI - High...309`
## • `95% CI - Low` -> `95% CI - Low...332`
## • `95% CI - High` -> `95% CI - High...333`
## • `95% CI - Low` -> `95% CI - Low...335`
## • `95% CI - High` -> `95% CI - High...336`
## • `95% CI - Low` -> `95% CI - Low...340`
## • `95% CI - High` -> `95% CI - High...341`
## • `95% CI - Low` -> `95% CI - Low...343`
## • `95% CI - High` -> `95% CI - High...344`
## • `# Deaths` -> `# Deaths...345`
## • `95% CI - Low` -> `95% CI - Low...347`
## • `95% CI - High` -> `95% CI - High...348`
## • `95% CI - Low` -> `95% CI - Low...372`
## • `95% CI - High` -> `95% CI - High...373`
## • `95% CI - Low` -> `95% CI - Low...379`
## • `95% CI - High` -> `95% CI - High...380`
## • `95% CI - Low` -> `95% CI - Low...382`
## • `95% CI - High` -> `95% CI - High...383`
## • `95% CI - Low` -> `95% CI - Low...408`
## • `95% CI - High` -> `95% CI - High...409`
## • `95% CI - Low` -> `95% CI - Low...413`
## • `95% CI - High` -> `95% CI - High...414`
## • `95% CI - Low` -> `95% CI - Low...417`
## • `95% CI - High` -> `95% CI - High...418`
## • `95% CI - Low` -> `95% CI - Low...441`
## • `95% CI - High` -> `95% CI - High...442`
## • `95% CI - Low` -> `95% CI - Low...444`
## • `95% CI - High` -> `95% CI - High...445`
## • `95% CI - Low` -> `95% CI - Low...448`
## • `95% CI - High` -> `95% CI - High...449`
## • `95% CI - Low` -> `95% CI - Low...452`
## • `95% CI - High` -> `95% CI - High...453`
## • `95% CI - Low` -> `95% CI - Low...459`
## • `95% CI - High` -> `95% CI - High...460`
## • `95% CI - Low` -> `95% CI - Low...463`
## • `95% CI - High` -> `95% CI - High...464`
## • `Average Grade Performance` -> `Average Grade Performance...474`
## • `Average Grade Performance (AIAN)` -> `Average Grade Performance
## (AIAN)...475`
## • `Average Grade Performance (Asian)` -> `Average Grade Performance
## (Asian)...476`
## • `Average Grade Performance (Black)` -> `Average Grade Performance
## (Black)...477`
## • `Average Grade Performance (Hispanic)` -> `Average Grade Performance
## (Hispanic)...478`
## • `Average Grade Performance (White)` -> `Average Grade Performance
## (White)...479`
## • `Average Grade Performance` -> `Average Grade Performance...480`
## • `Average Grade Performance (AIAN)` -> `Average Grade Performance
## (AIAN)...481`
## • `Average Grade Performance (Asian)` -> `Average Grade Performance
## (Asian)...482`
## • `Average Grade Performance (Black)` -> `Average Grade Performance
## (Black)...483`
## • `Average Grade Performance (Hispanic)` -> `Average Grade Performance
## (Hispanic)...484`
## • `Average Grade Performance (White)` -> `Average Grade Performance
## (White)...485`
## • `Segregation Index` -> `Segregation Index...486`
## • `95% CI - Low` -> `95% CI - Low...493`
## • `95% CI - High` -> `95% CI - High...494`
## • `95% CI - Low` -> `95% CI - Low...496`
## • `95% CI - High` -> `95% CI - High...497`
## • `Segregation Index` -> `Segregation Index...515`
## • `95% CI - Low` -> `95% CI - Low...517`
## • `95% CI - High` -> `95% CI - High...518`
## • `95% CI - Low` -> `95% CI - Low...542`
## • `95% CI - High` -> `95% CI - High...543`
## • `95% CI - Low` -> `95% CI - Low...567`
## • `95% CI - High` -> `95% CI - High...568`
## • `95% CI - Low` -> `95% CI - Low...591`
## • `95% CI - High` -> `95% CI - High...592`
## • `95% CI - Low` -> `95% CI - Low...594`
## • `95% CI - High` -> `95% CI - High...595`
## • `95% CI - Low` -> `95% CI - Low...614`
## • `95% CI - High` -> `95% CI - High...615`
## • `95% CI - Low` -> `95% CI - Low...619`
## • `95% CI - High` -> `95% CI - High...620`
## • `Population` -> `Population...623`
- Use pastecs::stat.desc to describe the variable. Include a few
sentences about what the variable is and what it’s measuring.
pastecs::stat.desc(County_Data$`Primary Care Physicians Rate`)
## nbr.val nbr.null nbr.na min max range
## 232.0000000 16.0000000 23.0000000 0.0000000 230.0834100 230.0834100
## sum median mean SE.mean CI.mean.0.95 var
## 9840.4798200 37.6597450 42.4158613 1.9934944 3.9277554 921.9726041
## std.dev coef.var
## 30.3640018 0.7158643
- Remove NA’s if needed using dplyr:filter (or anything similar)
County_Data_clean<-County_Data |> drop_na(`Primary Care Physicians Rate`) %>% filter(`Primary Care Physicians Rate`>0)
- Provide a histogram of the variable (as shown in this lesson)
hist(County_Data_clean$`Primary Care Physicians Rate`)

- transform the variable using the log transformation or square root
transformation (whatever is more appropriate) using dplyr::mutate or
something similar
County_Data_clean<-County_Data_clean %>% mutate(`Primary Care Physicians Rate SQRT`=sqrt(`Primary Care Physicians Rate`))
- provide a histogram of the transformed variable
hist(County_Data_clean$`Primary Care Physicians Rate SQRT`)
