Module 3 Quarto Project Myrie

Author

Joseph Myrie

Published

June 10, 2024

Introduction

This cohort data set has 5 variables including: smoking status, gender, age, cardiac status, and then cost. One could imagine this to be a cohort of patients that was assess for cardiac morbidities and then followed to see what their yearly healthcare expenditures were. We could make the assumption that costs will be higher for those with a positive smoking and cardiac status, also for those who are male, and for those who are older.

Overall this dataset is pretty simple with only 5 variables, but there does not seem to be alot of missing data which is great. There is also are large patient sample which is also great.

library(readr)
cohort_1_ <- read_csv("C:/Users/myriej01/Downloads/cohort (1).csv")
Rows: 5000 Columns: 5
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
dbl (5): smoke, female, age, cardiac, cost

ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
summary(cohort_1_)
     smoke            female           age           cardiac     
 Min.   :0.0000   Min.   :0.000   Min.   :18.00   Min.   :0.000  
 1st Qu.:0.0000   1st Qu.:0.000   1st Qu.:30.00   1st Qu.:0.000  
 Median :0.0000   Median :0.000   Median :41.00   Median :0.000  
 Mean   :0.1016   Mean   :0.487   Mean   :41.47   Mean   :0.038  
 3rd Qu.:0.0000   3rd Qu.:1.000   3rd Qu.:53.00   3rd Qu.:0.000  
 Max.   :1.0000   Max.   :1.000   Max.   :65.00   Max.   :1.000  
      cost      
 Min.   : 8478  
 1st Qu.: 9389  
 Median : 9664  
 Mean   : 9672  
 3rd Qu.: 9925  
 Max.   :11326  

Methods

First, we tried to see what relationships there were in the dataset by producing some scatter plots of the variables. First, we assessed the relationship of Age vs Cost by Cardiac status.

library(ggplot2)
ggplot(cohort_1_, aes(x = cohort_1_$age, y = cohort_1_$cost, color = cohort_1_$cardiac)) +
  geom_point() +
  labs(title = "Scatterplot of Age vs. Cost by Cardiac Status", x = "Age", y = "Cost")

<Age vs Cost by cardiac status>

We next made a scatterplot of Age vs Cost by Smoking status.

ggplot(cohort_1_, aes(x = cohort_1_$age, y = cohort_1_$cost, color = cohort_1_$smoke)) +
  geom_point() +
  labs(title = "Scatterplot of Age vs. Cost by Smoking Status", x = "Age", y = "Cost")

<Age vs Cost by Smoking Status>

We next looked at the relationship between smoking status and cost by cardiac status.

ggplot(cohort_1_, aes(x = cohort_1_$smoke, y = cohort_1_$cost, color = cohort_1_$cardiac)) +
  geom_point() +
  labs(title = "Scatterplot of Smoking Status vs. Cost by Cardiac Status", x = "Smoking Status", y = "Cost")

Results

First we have a table of the summary statics below:

summary(cohort_1_)
     smoke            female           age           cardiac     
 Min.   :0.0000   Min.   :0.000   Min.   :18.00   Min.   :0.000  
 1st Qu.:0.0000   1st Qu.:0.000   1st Qu.:30.00   1st Qu.:0.000  
 Median :0.0000   Median :0.000   Median :41.00   Median :0.000  
 Mean   :0.1016   Mean   :0.487   Mean   :41.47   Mean   :0.038  
 3rd Qu.:0.0000   3rd Qu.:1.000   3rd Qu.:53.00   3rd Qu.:0.000  
 Max.   :1.0000   Max.   :1.000   Max.   :65.00   Max.   :1.000  
      cost      
 Min.   : 8478  
 1st Qu.: 9389  
 Median : 9664  
 Mean   : 9672  
 3rd Qu.: 9925  
 Max.   :11326  

Next we ran a linear on the data looking at relationship between cardiac status, gender, and age with cost.

model <- lm(cohort_1_$cost ~ cohort_1_$cardiac + cohort_1_$female + cohort_1_$age, data = cohort_1_)
summary(model)

Call:
lm(formula = cohort_1_$cost ~ cohort_1_$cardiac + cohort_1_$female + 
    cohort_1_$age, data = cohort_1_)

Residuals:
    Min      1Q  Median      3Q     Max 
-836.42 -177.77  -31.19  133.48 1062.31 

Coefficients:
                   Estimate Std. Error t value Pr(>|t|)    
(Intercept)       9037.3877    12.6722  713.17   <2e-16 ***
cohort_1_$cardiac  478.1643    19.8784   24.05   <2e-16 ***
cohort_1_$female  -289.3979     7.6024  -38.07   <2e-16 ***
cohort_1_$age       18.2698     0.2774   65.85   <2e-16 ***
---
Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1

Residual standard error: 265.5 on 4996 degrees of freedom
Multiple R-squared:  0.5656,    Adjusted R-squared:  0.5653 
F-statistic:  2168 on 3 and 4996 DF,  p-value: < 2.2e-16
par(mfrow = c(2, 2))
# Residuals vs Fitted plot
plot(model, 1)

<linear model of cardiac status, gender, and age by cost>

When we look estimates for our variables we see there female status is correlated with lower cost. Cardiac status is strongly associated with increased costs, and age is associated with increased costs, but not as much as cardiac status. Our residuals versus fitted plot looks pretty good. Data points are spread out pretty uniformly until cost get to about 10,250.