This cohort data set has 5 variables including: smoking status, gender, age, cardiac status, and then cost. One could imagine this to be a cohort of patients that was assess for cardiac morbidities and then followed to see what their yearly healthcare expenditures were. We could make the assumption that costs will be higher for those with a positive smoking and cardiac status, also for those who are male, and for those who are older.
Overall this dataset is pretty simple with only 5 variables, but there does not seem to be alot of missing data which is great. There is also are large patient sample which is also great.
Rows: 5000 Columns: 5
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
dbl (5): smoke, female, age, cardiac, cost
ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
summary(cohort_1_)
smoke female age cardiac
Min. :0.0000 Min. :0.000 Min. :18.00 Min. :0.000
1st Qu.:0.0000 1st Qu.:0.000 1st Qu.:30.00 1st Qu.:0.000
Median :0.0000 Median :0.000 Median :41.00 Median :0.000
Mean :0.1016 Mean :0.487 Mean :41.47 Mean :0.038
3rd Qu.:0.0000 3rd Qu.:1.000 3rd Qu.:53.00 3rd Qu.:0.000
Max. :1.0000 Max. :1.000 Max. :65.00 Max. :1.000
cost
Min. : 8478
1st Qu.: 9389
Median : 9664
Mean : 9672
3rd Qu.: 9925
Max. :11326
Methods
First, we tried to see what relationships there were in the dataset by producing some scatter plots of the variables. First, we assessed the relationship of Age vs Cost by Cardiac status.
library(ggplot2)ggplot(cohort_1_, aes(x = cohort_1_$age, y = cohort_1_$cost, color = cohort_1_$cardiac)) +geom_point() +labs(title ="Scatterplot of Age vs. Cost by Cardiac Status", x ="Age", y ="Cost")
<Age vs Cost by cardiac status>
We next made a scatterplot of Age vs Cost by Smoking status.
ggplot(cohort_1_, aes(x = cohort_1_$age, y = cohort_1_$cost, color = cohort_1_$smoke)) +geom_point() +labs(title ="Scatterplot of Age vs. Cost by Smoking Status", x ="Age", y ="Cost")
<Age vs Cost by Smoking Status>
We next looked at the relationship between smoking status and cost by cardiac status.
ggplot(cohort_1_, aes(x = cohort_1_$smoke, y = cohort_1_$cost, color = cohort_1_$cardiac)) +geom_point() +labs(title ="Scatterplot of Smoking Status vs. Cost by Cardiac Status", x ="Smoking Status", y ="Cost")
Results
First we have a table of the summary statics below:
summary(cohort_1_)
smoke female age cardiac
Min. :0.0000 Min. :0.000 Min. :18.00 Min. :0.000
1st Qu.:0.0000 1st Qu.:0.000 1st Qu.:30.00 1st Qu.:0.000
Median :0.0000 Median :0.000 Median :41.00 Median :0.000
Mean :0.1016 Mean :0.487 Mean :41.47 Mean :0.038
3rd Qu.:0.0000 3rd Qu.:1.000 3rd Qu.:53.00 3rd Qu.:0.000
Max. :1.0000 Max. :1.000 Max. :65.00 Max. :1.000
cost
Min. : 8478
1st Qu.: 9389
Median : 9664
Mean : 9672
3rd Qu.: 9925
Max. :11326
Next we ran a linear on the data looking at relationship between cardiac status, gender, and age with cost.
model <-lm(cohort_1_$cost ~ cohort_1_$cardiac + cohort_1_$female + cohort_1_$age, data = cohort_1_)summary(model)
Call:
lm(formula = cohort_1_$cost ~ cohort_1_$cardiac + cohort_1_$female +
cohort_1_$age, data = cohort_1_)
Residuals:
Min 1Q Median 3Q Max
-836.42 -177.77 -31.19 133.48 1062.31
Coefficients:
Estimate Std. Error t value Pr(>|t|)
(Intercept) 9037.3877 12.6722 713.17 <2e-16 ***
cohort_1_$cardiac 478.1643 19.8784 24.05 <2e-16 ***
cohort_1_$female -289.3979 7.6024 -38.07 <2e-16 ***
cohort_1_$age 18.2698 0.2774 65.85 <2e-16 ***
---
Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
Residual standard error: 265.5 on 4996 degrees of freedom
Multiple R-squared: 0.5656, Adjusted R-squared: 0.5653
F-statistic: 2168 on 3 and 4996 DF, p-value: < 2.2e-16
par(mfrow =c(2, 2))# Residuals vs Fitted plotplot(model, 1)
<linear model of cardiac status, gender, and age by cost>
When we look estimates for our variables we see there female status is correlated with lower cost. Cardiac status is strongly associated with increased costs, and age is associated with increased costs, but not as much as cardiac status. Our residuals versus fitted plot looks pretty good. Data points are spread out pretty uniformly until cost get to about 10,250.