NOTE: there is code in the hypothesis testing guide document to help create code to run these types of tests, but use this link for more detail and for other methods of graphing: https://www.datanovia.com/en/lessons/ancova-in-r/
The “iris.csv” dataset is commonly used for examples in R. The dataset utilizes measurements of flower parts across different species of irises. Please download and bring in the “iris.csv” file.
For this dataset, I would like you to do the following: 1. Code: run the ANCOVA to determine how ‘variety’ and ‘sepal.width’ influence ‘sepal.length’ NOTE: don’t worry about testing assumptions for this (you would need to do this normally)
library(ggplot2)
## Warning: package 'ggplot2' was built under R version 4.4.3
iris <- read.csv("~/Desktop/BIN510-assignments/iris.csv")
model <- aov(sepal.length ~ variety + sepal.width, data = iris)
summary(model)
## Df Sum Sq Mean Sq F value Pr(>F)
## variety 2 63.21 31.606 164.8 < 2e-16 ***
## sepal.width 1 10.95 10.953 57.1 4.19e-12 ***
## Residuals 146 28.00 0.192
## ---
## Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
Download and bring in the “biomassTill.csv” file. This file contains data on how different soil tilling methods impact the growth (biomass) of the plants. But the researchers also manipulated the amount of water received by each plant, which would obviously influence any measurement of growth.
For this dataset, I would like you to do the following:
biomass <- read.csv("~/Desktop/BIN510-assignments/biomassTill.csv")
biomass$DVS <- as.factor(biomass$DVS)
model <- aov(Biomass ~ Tillage * DVS, data = biomass)
summary(model)
## Df Sum Sq Mean Sq F value Pr(>F)
## Tillage 2 36447 18224 2.686 0.0796 .
## DVS 4 711964 177991 26.238 4.71e-11 ***
## Tillage:DVS 8 34574 4322 0.637 0.7422
## Residuals 43 291703 6784
## ---
## Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
interaction.plot(biomass$DVS, biomass$Tillage, biomass$Biomass,
xlab = "Watering Regime (DVS)",
ylab = "Mean Biomass",
trace.label = "Tillage")
Download and bring in the “sleep.csv” file. This file contains data on participants in a sleep study and documenting the occurance of non-restorative sleep. Here, data in the “NRS” column are either a 0 or a 1, 0 meaning NRS not observed, 1 meaning NRS observed.
For this dataset, I would like you to do the following:
sleep <- read.csv("~/Desktop/BIN510-assignments/sleep.csv")
model <- glm(NRS ~ Age_2013, data = sleep, family = binomial)
summary(model)
##
## Call:
## glm(formula = NRS ~ Age_2013, family = binomial, data = sleep)
##
## Coefficients:
## Estimate Std. Error z value Pr(>|z|)
## (Intercept) 0.868077 0.068995 12.58 <2e-16 ***
## Age_2013 -0.031965 0.001079 -29.63 <2e-16 ***
## ---
## Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
##
## (Dispersion parameter for binomial family taken to be 1)
##
## Null deviance: 98557 on 90188 degrees of freedom
## Residual deviance: 97701 on 90187 degrees of freedom
## AIC: 97705
##
## Number of Fisher Scoring iterations: 4
ggplot(sleep, aes(x = Age_2013, y = NRS)) +
geom_jitter(height = 0.05, width = 0) +
stat_smooth(method = "glm", method.args = list(family = "binomial"),
se = TRUE) +
labs(x = "Age", y = "Probability of Non-Restorative Sleep") +
theme_classic()
## `geom_smooth()` using formula = 'y ~ x'